Deep convolutional neural networks predicting variant pathogenicity using three-dimensional (3D) protein structures
Deep convolutional neural networks analyzing voxelized 3D protein structures enhance pathogenicity prediction by integrating structural and evolutionary data, addressing limitations of prior knowledge-dependent methods and improving prediction accuracy.
Patent Information
- Application Number
- JP2025116672
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-24
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-15
- Estimated Expiration
- 2042-04-14
AI Technical Summary
Existing methods for predicting the pathogenicity of genetic variants rely heavily on prior knowledge and are limited by the availability of clinical data, leading to potential overfitting and suboptimal performance, especially when analyzing protein structures.
Utilizing deep convolutional neural networks to analyze multi-channel voxelized representations of three-dimensional protein structures, incorporating structural and evolutionary information to predict variant pathogenicity.
Improves the accuracy of pathogenicity prediction by leveraging structural data and evolutionary patterns, reducing reliance on prior knowledge and enhancing the model's ability to generalize across diverse protein sequences.
Smart Images

Figure 0007755105000001 
Figure 0007755105000002 
Figure 0007755105000003
Abstract
Description
[Technical Field]
[0001] (Priority application) This application claims priority to U.S. Non-Provisional Patent Application No. 17 / 468,411 (Attorney Docket No. ILLM 1037-3 / IP-2051A-US), entitled "Deep Convolutional Neural Networks to Predict Variant Pathogenicity using Three-Dimensional (3D) Protein Structures," filed on September 7, 2021, which is a continuation of U.S. Non-Provisional Patent Application No. 17 / 232,056 (Attorney Docket No. ILLM 1037-2 / IP-2051-US), entitled "Deep Convolutional Neural Networks to Predict Variant Pathogenicity using Three-Dimensional (3D) Protein Structures," filed on April 15, 2021.
[0002] This application also claims priority to U.S. Non-Provisional Patent Application No. 17 / 703,935, entitled "Multi-channel Protein Voxelization To Predict Variant Pathogenicity Using Deep Convolutional Neural Networks," filed March 24, 2022 (Attorney Docket No. ILLM 1047-2 / IP-2142-US), which in turn claims priority to or the benefit of U.S. Provisional Patent Application No. 63 / 175,495, entitled "Multi-channel Protein Voxelization To Predict Variant Pathogenicity Using Deep Convolutional Neural Networks," filed April 15, 2021 (Attorney Docket No. ILLM 1047-1 / IP-2142-PRV).
[0003] This application also claims priority to U.S. Non-Provisional Patent Application No. 17 / 703,958, entitled "Efficient Voxelization For Deep Learning," filed March 24, 2022 (Attorney Docket No. ILLM 1048-2 / IP-2143-US), which in turn claims priority to or the benefit of U.S. Provisional Patent Application No. 63 / 175,767, entitled "Efficient Voxelization For Deep Learning," filed April 16, 2021 (Attorney Docket No. ILLM 1048-1 / IP-2143-PRV).
[0004] The priority application is incorporated herein by reference for all purposes.
[0005] (Related Applications) This application is related to a concurrently filed PCT patent application entitled "Artificial Intelligence-based Analysis of Protein Three-Dimensional (3D) Structures" (Attorney Docket No. ILLM 1037-4 / IP-2051-PCT), which is incorporated herein by reference for all purposes.
[0006] FIELD OF THE INVENTION The disclosed technology relates to artificial intelligence-based computers and digital data processing systems and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using deep convolutional neural networks to analyze multi-channel voxelized data.
[0007] (built-in) The following are incorporated by reference for all purposes as if fully set forth herein:
[0008] Sundaram, L et al., Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018);
[0009] Jaganathan, K. et al., Predicting splicing from primary sequence with deep learning.Cell 176, 535-548 (2019);
[0010] U.S. Patent Application No. 62 / 573 (144) entitled "TRAINING A DEEP PATHOGENICITY CLASSIFIER USING LARGE-SCALE BENIGN TRAINING DATA," filed October 16, 2017 (Attorney Docket No. ILLM 1000-1 / IP-1611-PRV);
[0011] U.S. Patent Application No. 62 / 573,149, entitled "PATHOGENICITY CLASSIFIER BASED ON DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs)," filed October 16, 2017 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV);
[0012] U.S. Patent Application Serial No. 62 / 573 (153) entitled "DEEP SEMI-SUPERVISED LEARNING THAT GENERATES LARGE-SCALE PATHOGENIC TRAINING DATA," filed October 16, 2017 (Attorney Docket No. ILLM 1000-3 / IP-1613-PRV).
[0013] U.S. Patent Application No. 62 / 582,898, entitled "PATHOGENICITY CLASSIFICATION OF GEOMIC DATA USING DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs)," filed November 7, 2017 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV);
[0014] U.S. Patent Application No. 16 / 160,903, entitled "DEEP LEARNING-BASED TECHNIQUES FOR TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS," filed October 15, 2018 (Attorney Docket No. ILLM 1000-5 / IP-1611-US);
[0015] U.S. Patent Application No. 16 / 160,986, entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS FOR VARIANT CLASSIFICATION," filed October 15, 2018 (Attorney Docket No. ILLM 1000-6 / IP-1612-US);
[0016] U.S. Patent Application No. 16 / 160,968, entitled "SEMI-SUPERVISED LEARNING FOR TRAINING AN ENSEMBLE OF DEEP CONVOLUTIONAL NEURAL NETWORKS," filed October 15, 2018 (Attorney Docket No. ILLM 1000-7 / IP-1613-US); and
[0017] U.S. Patent Application No. 16 / 407,149, filed May 8, 2019, entitled "DEEP LEARNING-BASED TECHNIQUES FOR PRE-TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS" (Attorney Docket No. ILLM 1010-1 / IP-1734-US). [Background technology]
[0018] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or problems associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which themselves may also correspond to embodiments of the claimed technology.
[0019] Genomics, broadly defined, also known as functional genomics, aims to characterize the function of all genomic elements of an organism by using genome-scale assays such as genome sequencing, transcriptome profiling, and proteomics. Genomics has emerged as a data-driven science that operates not by testing preconceived models and hypotheses, but by discovering novel properties from the examination of genome-scale data. Applications of genomics include finding associations between genotypes and phenotypes, discovering biomarkers for patient stratification, predicting gene function, and mapping biochemically active genomic regions such as transcriptional enhancers.
[0020] Genomics data is too large and complex to be mined solely by visual inspection of pairwise correlations. Instead, analytical tools are needed to support the discovery of unexpected relationships, derive novel hypotheses and models, and make predictions. Unlike some algorithms in which assumptions and domain expertise are hard-coded, machine learning algorithms are designed to automatically discover patterns in data. Therefore, machine learning algorithms are well-suited for data-driven science, particularly genomics. However, the performance of machine learning algorithms can strongly depend on how the data is represented—i.e., how each variable (also called a feature) is calculated. For example, to classify tumors as malignant or benign from fluorescence microscopy images, a preprocessing algorithm can detect cells, identify cell types, and generate a list of cell counts for each cell type.
[0021] A machine learning model can take estimated cell counts, an example of hand-designed features, as input features for classifying tumors. The central problem is that classification performance depends heavily on the quality and relevance of these features. For example, relevant visual features such as cell morphology, intercellular distance, or intraorgan localization are not captured in cell counts, and this incomplete representation of the data can reduce classification accuracy.
[0022] Deep learning, a subdivision of machine learning, addresses this problem by embedding feature computation within the machine learning model itself, generating an end-to-end model. This achievement was made possible by the development of deep neural networks, machine learning models that involve successive elementary operations that compute increasingly complex features by taking the results of previous operations as input. Deep neural networks can improve prediction accuracy by discovering relevant features of increasing complexity, such as cell morphology and spatial organization of cells in the above example. The construction and training of deep neural networks has been made possible by the explosion of data, advances in algorithms, and substantial increases in computing power, particularly through the use of graphics processing units (GPUs).
[0023] The goal of supervised learning is to obtain a model that takes features as input and returns a prediction of a so-called target variable. An example of a supervised learning problem is predicting whether an intron will be spliced (target) or not, given RNA features such as the presence or absence of canonical splice site sequences, the location of splicing branch points, or the length of the intron. Learning a machine learning model refers to learning its parameters, which generally involves minimizing a loss function over training data with the goal of making accurate predictions on unknown data.
[0024] For many supervised learning problems in computational biology, input data can be represented as a table with multiple columns, or features, each of which contains numerical or categorical data potentially useful for making predictions. While some input data are naturally represented as tabular features (e.g., temperature or time), other input data must first be transformed using a process called feature extraction to fit the tabular representation (e.g., converting deoxyribonucleic acid (DNA) sequences to k-mer counts). For intron-splicing prediction problems, the presence or absence of canonical splice site sequences, the location of splicing branch points, and intron length can be preprocessed features collected in tabular form. Tabular data is the standard for a wide range of supervised machine learning models, ranging from simple linear models such as logistic regression to more flexible nonlinear models such as neural networks and many others.
[0025] Logistic regression is a binary classifier, i.e., a supervised learning model that predicts a binary target variable. Specifically, logistic regression predicts the probability of a positive class by calculating a weighted sum of input features mapped to the [0, 1] interval using a sigmoid activation function. The parameters of logistic regression, or other linear classifiers that use different activation functions, are the weights in the weighted sum. Linear classifiers fail when the weighted sum of input features cannot sufficiently distinguish between classes, such as whether an intron is spliced or not. To improve prediction performance, new input features can be manually added by transforming or combining existing features in new ways, for example, by taking exponentiation or pairwise products.
[0026] Neural networks automatically learn these nonlinear feature transformations using hidden layers. Each hidden layer can be thought of as multiple linear models with their output transformed by a nonlinear activation function, such as a sigmoid function or the more general rectified linear unit (ReLU). Together, these layers organize the input features into related complex patterns, easing the task of distinguishing between two classes.
[0027] Deep neural networks use many hidden layers, and when each neuron receives input from all neurons in the previous layer, the layer is said to be fully connected. Neural networks are typically trained using stochastic gradient descent, an algorithm suitable for training models on very large datasets. Implementation of neural networks using modern deep learning frameworks allows rapid prototyping with different architectures and datasets. Fully connected neural networks can be used in several genomics applications, including predicting the proportion of spliced exons for a given sequence from sequence features such as the presence of splice factor binding motifs or sequence conservation, prioritizing potential disease-causing genetic variants, and predicting cis-regulatory elements in a given genomic region using features such as chromatin marks, gene expression, and evolutionary conservation.
[0028] For effective prediction, local dependencies in spatial and longitudinal data must be considered. For example, shuffling DNA sequences or image pixels severely disrupts information patterns. These local dependencies set spatial or longitudinal data apart from tabular data, where feature ordering is arbitrary. Consider the problem of classifying genomic regions as bound versus unbound by a specific transcription factor, where binding regions are defined as high-confidence binding events in chromatin immunoprecipitation following sequencing (ChIP-seq) data. Transcription factors bind to DNA by recognizing sequence motifs. Fully connected layers based on sequence-derived features, such as the number of k-mer instances in a sequence or position weight matrix (PWM) matches, can be used for this task. Because k-mer or PWM instance frequencies are robust to shifting motifs within a sequence, such models can generalize well to sequences with the same motif located at different positions. However, they cannot recognize patterns where transcription factor binding depends on the combination of multiple motifs with well-defined intervals. Furthermore, the number of possible k-mers increases exponentially with k-mer length, which poses both conservation and overfitting challenges.
[0029] A convolutional layer is a special form of a fully connected layer in which the same fully connected layer is applied locally to every sequence position, for example, within a 6-bp window. This approach can also be viewed as scanning a sequence using multiple PWMs, for example, for the transcription factors GATA1 and TAL1. By using the same model parameters across positions, the total number of parameters is dramatically reduced, allowing the network to detect motifs at positions not seen during training. Each convolutional layer scans the sequence with several filters by generating a scalar value at every position that quantizes the match between the filter and the sequence. As with fully connected neural networks, a nonlinear activation function (typically ReLU) is applied in each layer. Next, a pooling operation is applied, which aggregates activations within successive bins across the position axis, typically taking the maximum or average activation for each channel. Pooling reduces the effective sequence length and coarsens the signal. Subsequent convolutional layers can then construct the output of the previous layer and detect whether the GATA1 and TAL1 motifs were present within a certain distance range. Finally, the output of the convolutional layer can be used as input to a fully connected neural network to perform the final prediction task. Thus, different types of neural network layers (e.g., fully connected and convolutional layers) can be combined within a single neural network.
[0030] Convolutional neural networks (CNNs) can predict various molecular phenotypes based solely on DNA sequence. Applications include classification of transcription factor binding sites and prediction of molecular phenotypes such as chromatin features, DNA contact maps, DNA methylation, gene expression, translation efficiency, RBP binding, and microRNA (miRNA) targets. In addition to predicting molecular phenotypes from sequence, convolutional neural networks can be applied to more technical tasks traditionally addressed by hand-designed bioinformatics pipelines. For example, convolutional neural networks can predict guide RNA specificity, denoise ChIP-seq, improve Hi-C data resolution, predict laboratory origin from DNA sequence, and call genetic variants. Convolutional neural networks have also been used to model long-range dependencies in genomes. Although interacting regulatory elements may be located far apart on unfolded linear DNA sequences, these elements are often proximal in actual 3D chromatin structures. Thus, modeling molecular phenotypes from linear DNA sequences can be improved by allowing long-range dependencies, albeit with a rough approximation of chromatin, and allowing the model to implicitly learn aspects of 3D organization such as promoter-enhancer loops. This is achieved by using extended convolutions with receptive fields up to 32 kb. Extended convolutions also allow splice sites to be predicted from sequences using receptive fields as small as 10 kb, thereby enabling integration of gene sequences over distances as long as a typical human intron (see Jaganathan, K. et al., Predicting splicing from primary sequence with deep learning. Cell 176, 535-548 (2019)).
[0031] Different types of neural networks can be characterized by their parameter-sharing schemes. For example, fully connected layers have no parameter sharing, while convolutional layers impose translational invariance by applying the same filter at every position in their input. Recurrent neural networks (RNNs) are an alternative to convolutional neural networks for processing sequential data, such as DNA sequences or time series, that implement a different parameter-sharing scheme. Recurrent neural networks apply the same operation to each sequence element. This operation takes as input the memory of the previous sequence element and the new input. It updates the memory and optionally emits an output that is passed to subsequent layers or used directly as a model prediction. By applying the same model at each sequence element, recurrent neural networks are invariant to positional indices in the processed sequence. For example, recurrent neural networks can detect open reading frames in DNA sequences, regardless of their position in the sequence. This task requires the recognition of a specific sequence of inputs, such as a start codon followed by an in-frame stop codon.
[0032] The main advantage of recurrent neural networks over convolutional neural networks is that, theoretically, they can carry information through infinitely long sequences via memory. Furthermore, recurrent neural networks can naturally process sequences of widely varying lengths, such as mRNA sequences. However, convolutional neural networks combined with various tricks (such as dilated convolutions) can achieve performance comparable to or even better than recurrent neural networks for sequence modeling tasks such as audio synthesis and machine translation. Recurrent neural networks can aggregate the outputs of convolutional neural networks to predict single-cell DNA methylation status, RBP binding, transcription factor binding, and DNA accessibility. Furthermore, because recurrent neural networks apply sequential operations, they cannot be easily parallelized and are therefore much slower to compute than convolutional neural networks.
[0033] Although most of the human genetic code is common to all humans, each human has a unique genetic code. In some cases, the human genetic code may contain outliers, called genetic variants, that may be common among a relatively small group of individuals in the human population. For example, a particular human protein may contain a particular sequence of amino acids, but variants of that protein may differ by one amino acid in an otherwise identical particular sequence.
[0034] Genetic variants can be pathogenic and can result in disease. Although most such genetic variants have been depleted from the genome by natural selection, the ability to identify which genetic variants are likely pathogenic can help researchers focus on these genetic variants to gain an understanding of the corresponding diseases and their diagnosis, treatment, or cure. The clinical interpretation of millions of human genetic variants remains unclear. Some of the most frequent pathogenic variants are single-nucleotide missense mutations that change the amino acid of a protein. However, not all missense mutations are pathogenic.
[0035] Models that can directly predict molecular phenotypes from biological sequences can be used as in silico perturbation tools to investigate the association between genetic and phenotypic variation and have emerged as new methods for quantitative trait locus identification and variant prioritization. These approaches are crucial given that the majority of variants identified by genome-wide association studies of complex phenotypes are non-coding, making it difficult to estimate their effects and contribution to the phenotype. Furthermore, linkage disequilibrium results in blocks of co-inherited variants, making it difficult to accurately identify individual causal variants. Therefore, sequence-based deep learning models that can be used as matching tools to assess the impact of such variants offer a promising approach for discovering potential drivers of complex phenotypes. One example is predicting the effects of non-coding single-nucleotide variants and short insertions or deletions (indels) indirectly from the differences between two variants on transcription factor binding, chromatin accessibility, or gene expression prediction. Another example is predicting novel splice site creation from the sequence or quantitative effect of genetic variants on splicing.
[0036] To predict the pathogenicity of missense variants from protein sequence and sequence conservation data, an end-to-end deep learning approach for variant effect prediction is applied (see Sundaram, L. et al., Predicting the clinical impact of human mutations with deep neural networks. Nat. Genet. 50, 1161-1170 (2018) and referred to herein as "PrimateAI"). PrimateAI uses a deep neural network trained on known pathogenic variants with data augmentation using cross-species information. In particular, PrimateAI uses wild-type and mutant protein sequences to compare differences and determine the pathogenicity of the mutation using a trained deep neural network. Such an approach, utilizing protein sequences for pathogenicity prediction, is promising because it can avoid the circularity problem and overfitting to prior knowledge. However, the amount of clinical data available in ClinVar is relatively small compared to the amount of data sufficient to effectively train a deep neural network. To overcome this data shortage, PrimateAI uses common human variants and primate-derived variants as benign data, and simulated variants based on trinucleotide context as unlabeled data.
[0037] PrimateAI outperforms conventional methods when trained directly on sequence alignments. PrimateAI learns important protein domains, conserved amino acid positions, and sequence dependencies directly from training data consisting of approximately 120,000 human samples. PrimateAI substantially outperforms other variant pathogenicity prediction tools in distinguishing benign and pathogenic de novo mutations in candidate developmental disorder genes and in replicating prior knowledge in ClinVar. These results suggest that PrimateAI represents an important step forward for variant classification tools that can reduce reliance on prior knowledge of clinical reports.
[0038] Central to protein biology is understanding how structural elements give rise to observed functions. The plethora of protein structural data enables the development of computational methods to systematically derive the rules governing structure-function relationships. However, the performance of these methods critically depends on the choice of protein structural representation.
[0039] Protein sites are microenvironments within a protein structure that are distinguished by their structural or functional role. Sites can be defined by their three-dimensional (3D) location and the local neighborhood around this location where the structure or function resides. Central to rational protein engineering is understanding how the structural arrangement of amino acids creates functional characteristics within a protein site. Determining the structural and functional roles of individual amino acids within a protein provides information to aid in the manipulation and modification of protein function. Identifying functionally or structurally important amino acids allows for focused engineering efforts, such as site-directed mutagenesis, to alter the functional properties of a target protein. Alternatively, this knowledge can help avoid engineering designs that abolish desired functions.
[0040] Because it is well established that structure is much more conserved than sequence, the increase in protein structural data provides an opportunity to systematically study the underlying patterns governing structure-function relationships using data-driven approaches. A fundamental aspect of any computational protein analysis is how protein structural information is represented. The performance of machine learning methods often depends more on the choice of data representation than on the machine learning algorithm used. A good representation efficiently captures the most important information, while a poor representation produces a noisy distribution lacking the underlying pattern.
[0041] The plethora of protein structures and recent success of deep learning algorithms provide an opportunity to develop tools for automatically extracting task-specific representations of protein structure. Thus, an opportunity arises to predict variant pathogenicity using multi-channel voxelized representations of 3D protein structures as input to deep neural networks. [Brief explanation of the drawings]
[0042] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings: [Figure 1] 1 is a flow chart illustrating the process of a system for determining pathogenicity of a variant, according to various embodiments of the disclosed technology. [Figure 2] FIG. 1 is a schematic diagram illustrating an exemplary reference amino acid sequence of a protein and alternative amino acid sequences of the protein, according to one embodiment of the disclosed technology. [Figure 3] FIG. 3 shows an amino acid unit classification of atoms of amino acids in the reference amino acid sequence of FIG. 2 according to one embodiment of the disclosed technology. [Figure 4] FIG. 4 shows the assignment of amino acid units of the 3D atomic coordinates of the alpha carbon atoms classified in FIG. 3 on an amino acid basis, according to one embodiment of the disclosed technology. [Figure 5] FIG. 10 is a diagram that schematically illustrates a process for determining per-voxel distance values, according to one implementation of the disclosed technology. [Figure 6] FIG. 1 shows an example of a distance channel of 21 amino acid units, according to one embodiment of the disclosed technology. [Figure 7] FIG. 1 is a schematic diagram of a distance channel tensor, in accordance with one implementation of the disclosed technology; [Figure 8] FIG. 3 illustrates one-hot encoding of the reference amino acid and alternative amino acid from FIG. 2 according to one embodiment of the disclosed technology. [Figure 9]FIG. 1 is a schematic diagram of voxelized one-hot encoded reference amino acids and voxelized one-hot encoded variant / alternative amino acids, according to one embodiment of the disclosed technology. [Figure 10] FIG. 8 is a diagram that schematically illustrates the concatenation process for voxel-wise concatenation of the distance channel tensor and reference allele tensor of FIG. 7, in accordance with one implementation of the disclosed technology. [Figure 11] FIG. 11 is a diagram illustrating a schematic illustration of a concatenation process for concatenating the distance channel tensor of FIG. 7, the reference allele tensor of FIG. 10, and the alternative allele tensor on a voxel-by-voxel basis, in accordance with one implementation of the disclosed technology. [Figure 12] 1 is a flow diagram illustrating the process of a system for determining and assigning nearest-neighbor atoms general amino acid conservation frequencies to voxels (voxelization), according to one implementation of the disclosed technology. [Figure 13] FIG. 10 illustrates a diagram of a voxel to nearest amino acid according to one embodiment of the disclosed technology. [Figure 14] FIG. 1 shows an exemplary multiple sequence alignment of reference amino acid sequences across 99 species, according to one embodiment of the disclosed technology. [Figure 15] FIG. 10 shows an example of determining a general amino acid conservation frequency sequence for a particular voxel according to one embodiment of the disclosed technology. [Figure 16] FIG. 16 illustrates the respective global amino acid conservation frequencies determined for each voxel using the positional frequency logic described in FIG. 15, according to one embodiment of the disclosed technology. [Figure 17] FIG. 1 illustrates an evolutionary profile for each voxelized voxel, according to one implementation of the disclosed technology. [Figure 18] FIG. 10 illustrates an example of an evolutionary profile tensor, according to one implementation of the disclosed technology. [Figure 19] 1 is a flow diagram illustrating a process of a system for determining amino acid-by-amino acid conservation frequencies of nearest neighbor atoms and assigning them to voxels (voxelization), according to one implementation of the disclosed technology. [Figure 20]10A-10C illustrate various examples of voxelized annotation channels concatenated with distance channel tensors, according to one implementation of the disclosed technology. [Figure 21] FIG. 1 illustrates different combinations and permutations of input channels that can be provided as input to a pathogenicity classifier for determining the pathogenicity of a target variant, according to one embodiment of the disclosed technology. [Figure 22] 1A-1C illustrate different methods of calculating the disclosed distance channel, according to various implementations of the disclosed technology. [Figure 23] 1A-1C illustrate different examples of evolutionary channels in accordance with various embodiments of the disclosed technology. [Figure 24] 1A-1C illustrate different examples of annotation channels in accordance with various implementations of the disclosed technology. [Figure 25] 10A-10C illustrate different examples of structural confidence channels in accordance with various implementations of the disclosed technology. [Figure 26] FIG. 1 illustrates an exemplary processing architecture of a pathogenicity classifier, according to one implementation of the disclosed technology. [Figure 27] FIG. 1 illustrates an exemplary processing architecture of a pathogenicity classifier, according to one implementation of the disclosed technology. [Figure 28] Figure 10 demonstrates the classification superiority of the disclosed PrimateAI 3D over PrimateAI, using PrimateAI as a benchmark model. [Figure 29] Figure 10 demonstrates the classification superiority of the disclosed PrimateAI 3D over PrimateAI, using PrimateAI as a benchmark model. [Figure 30] Figure 10 demonstrates the classification superiority of the disclosed PrimateAI 3D over PrimateAI, using PrimateAI as a benchmark model. [Figure 31] Figure 10 demonstrates the classification superiority of the disclosed PrimateAI 3D over PrimateAI, using PrimateAI as a benchmark model. [Figure 32A] 1 illustrates the disclosed efficient voxelization process, in accordance with various implementations of the disclosed technology. [Figure 32B] 1 illustrates the disclosed efficient voxelization process, in accordance with various implementations of the disclosed technology. [Figure 33] FIG. 10 illustrates how atoms are associated with voxels that contain the atoms, according to one implementation of the disclosed technology. [Figure 34] FIG. 10 illustrates generating a voxel-to-atom mapping from an atom-to-voxel mapping to identify nearest neighbor atoms for each voxel, according to one implementation of the disclosed technology. [Figure 35A] FIG. 10 illustrates how the disclosed efficient voxelization has a runtime complexity of O(number of atoms) versus O(number of atoms * number of voxels) without using the disclosed efficient voxelization. [Figure 35B] FIG. 10 illustrates how the disclosed efficient voxelization has a runtime complexity of O(number of atoms) versus O(number of atoms * number of voxels) without using the disclosed efficient voxelization. [Figure 36] FIG. 1 illustrates an exemplary computer system that can be used to implement the disclosed techniques. DETAILED DESCRIPTION OF THE INVENTION
[0043] The following discussion is presented to enable any person skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0044] The detailed description of various embodiments may be better understood when read in conjunction with the accompanying drawings. To the extent that the figures illustrate diagrams of functional blocks of various embodiments, the functional blocks are not necessarily indicative of a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., modules, processors, or memories) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or block of random access memory, hard disk, etc.) or in multiple pieces of hardware. Similarly, a program may be a stand-alone program, may be incorporated as a subroutine within an operating system, may be a function within an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentality shown in the figures.
[0045] The processing engines and databases in the figures designated as modules can be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some modules may be implemented on different processors, computers, or servers, or may be spread across multiple different processors, computers, or servers. In addition, it will be understood that some of the modules may operate in parallel or in a different order than shown in the figures without affecting the functionality achieved. The modules in the figures may also be considered flowchart steps in a method. Also, a module need not necessarily have all its code located contiguously in memory. Some portions of code may be separated from other portions of code, with code from other modules or other functions located between them.
[0046] Protein structure-based pathogenicity determination Figure 1 is a flow diagram illustrating a process 100 of a system for determining pathogenicity of a variant. In step 102, a sequence accessor 104 of the system accesses a reference amino acid sequence and an alternative amino acid sequence. In 112, a 3D structure generator 114 of the system generates a 3D protein structure of the reference amino acid sequence. In some embodiments, the 3D protein structure is a homology model of a human protein. In one embodiment, the so-called SwissModel homology modeling pipeline provides a public repository of predicted human protein structures. In another embodiment, the so-called HHpred homology modeling uses a tool called Modeller to predict the structure of a target protein from a template structure.
[0047] Proteins are represented by a collection of atoms and their coordinates in 3D space. Amino acids can have various atoms such as carbon atoms, oxygen (O) atoms, nitrogen (N) atoms, and hydrogen (H) atoms. The atoms can be further classified as side chain atoms and backbone atoms. Backbone carbon atoms are the alpha carbon (C α ) atoms and beta carbon (C β ) atoms.
[0048] In step 122, the coordinate classifier 124 of the system classifies the 3D atomic coordinates of the 3D protein structure on an amino acid basis. In one embodiment, classifying the amino acid units includes assigning the 3D atomic coordinates to 21 amino acid categories (including stop or gap amino acid categories). In one example, classifying the amino acid units of alpha carbon atoms can list each alpha carbon atom under each of the 21 amino acid categories. In another example, classifying the amino acid units of beta carbon atoms can list each beta carbon atom under each of the 21 amino acid categories.
[0049] In yet another example, the classification of amino acid units for oxygen atoms can list each oxygen atom under each of the 21 amino acid categories. In yet another example, the classification of amino acid units for nitrogen atoms can list each nitrogen atom under each of the 21 amino acid categories. In yet another example, the classification of amino acid units for hydrogen atoms can list each hydrogen atom under each of the 21 amino acid categories.
[0050] Those skilled in the art will appreciate that in various embodiments, the classification of amino acid units may include a subset of the 21 amino acid categories and a subset of different atomic elements.
[0051] In step 132, the system's voxel grid generator 134 instantiates a voxel grid. The voxel grid can have any resolution, e.g., 3x3x3, 5x5x5, 7x7x7, etc. The voxels within the voxel grid can be any size, e.g., 1 Angstrom (Å) on each side, 2 Å on each side, 3 Å on each side, etc. Those skilled in the art will understand that since voxels are cubes, these exemplary dimensions refer to cubic dimensions. Those skilled in the art will also understand that these exemplary dimensions are non-limiting and that voxels can have any cubic dimensions.
[0052] In step 142, the system's voxel grid centerer 144 centers the voxel grid at the amino acid level on the reference amino acid undergoing the target mutation. In one embodiment, the voxel grid is centered on the atomic coordinates of a particular atom of the reference amino acid undergoing the target mutation, for example, the 3D atomic coordinates of the alpha carbon atom of the reference amino acid undergoing the target mutation.
[0053] Distance Channel A voxel in the voxel grid can have multiple channels (or features). In one embodiment, a voxel in the voxel grid has multiple distance channels (e.g., 21 distance channels each for one of the 21 amino acid categories (including the stop or gap amino acid category)). In step 152, a distance channel generator 154 of the system generates distance channels per amino acid for the voxels in the voxel grid. A distance channel is generated independently for each of the 21 amino acid categories.
[0054] For example, consider the alanine (A) amino acid category. Further, consider, for example, a voxel grid that is 3x3x3 in size and has 27 voxels. Then, in one embodiment, the alanine distance channel contains 27 distance values for each of the 27 voxels in the voxel grid. The 27 distance values in the alanine distance channel are: is measured from the center of each of the 27 voxels in the voxel grid to their nearest neighbor atom in the alanine amino acid category.
[0055] In one example, the alanine amino acid category includes only alpha carbon atoms, so the nearest neighbor atoms are the alanine alpha carbon atoms that are closest to each of the 27 voxels in the voxel grid. In another example, the alanine amino acid category includes only beta carbon atoms, so the nearest neighbor atoms are the alanine beta carbon atoms that are closest to each of the 27 voxels in the voxel grid.
[0056] In yet another example, the alanine amino acid category contains only oxygen atoms, and therefore the nearest neighbor atoms are the alanine oxygen atoms that are closest to each of the 27 voxels in the voxel grid. In yet another example, the alanine amino acid category contains only nitrogen atoms, and therefore the nearest neighbor atoms are the alanine nitrogen atoms that are closest to each of the 27 voxels in the voxel grid. In yet another example, the alanine amino acid category contains only hydrogen atoms, and therefore the nearest neighbor atoms are the alanine hydrogen atoms that are closest to each of the 27 voxels in the voxel grid.
[0057] Similar to the alanine distance channel, distance channel generator 154 generates distance channels (i.e., sets of distance values in voxels) for each of the remaining amino acid categories. In other embodiments, distance channel generator 154 generates distance channels for only a subset of the 21 amino acid categories.
[0058] In other embodiments, the selection of nearest neighbor atoms is not limited to a particular atom type, i.e., within the target amino acid category, the nearest neighbor atom to a particular voxel is selected regardless of the atomic component of the nearest neighbor atom, and the distance value of the particular voxel is calculated for inclusion in the distance channel of the target amino acid category.
[0059] In yet other embodiments, distance channels are generated on an atomic element basis. Instead of, or in addition to, having distance channels for amino acid categories, distance values can be generated for atomic element categories regardless of the amino acid to which the atoms belong. For example, consider that the atoms of amino acids in a reference amino acid sequence span seven atomic elements: carbon, oxygen, nitrogen, hydrogen, calcium, iodine, and sulfur. Then, voxels in a voxel grid are configured to have seven distance channels, such that each of the seven distance channels has 27 distance values per voxel that specify the distance to nearest atoms only within the corresponding atomic element category. In other embodiments, distance channels can be generated for only a subset of the seven atomic elements. In still other embodiments, atomic element category and distance channel generation are performed for the same atomic element, e.g., alpha carbon (C α ) atoms and beta carbon (C β ) can be further layered into atomic variations.
[0060] In yet other embodiments, distance channels can be generated based on atom type, for example, a distance channel for only side chain atoms and a distance channel for only backbone atoms.
[0061] The nearest neighbor atoms can be searched within a predetermined maximum scanning radius (e.g., 6 Angstroms (Å)) from the voxel center, and multiple atoms may be nearest to the same voxel within the voxel grid.
[0062] Distances are calculated between the 3D coordinates of the voxel center and the 3D atomic coordinates of the atom, and distance channels are generated using a voxel grid centered at the same location (e.g., centered on the 3D atomic coordinates of the alpha carbon atom of the reference amino acid undergoing the target mutation).
[0063] The distance can be a Euclidean distance. The distance can also be parameterized by the atom size (or influence) (e.g., by using the Lennard-Jones potential and / or van der Waals atomic radius of the atom in question). The distance value can also be normalized by the maximum scan radius or by the maximum observed distance value of the farthest nearest neighbor atom within the target amino acid category, target atomic element category, or target atom type category. In some embodiments, the distance between a voxel and an atom is calculated based on the polar coordinates of the voxel and the atom. The polar coordinates are parameterized by the angle between the voxel and the atom. In one embodiment, this angle information is used to generate an angle channel for the voxel (i.e., independently of the distance channel). In some embodiments, the angle between the nearest neighbor atom and a neighboring atom (e.g., a backbone atom) can be used as a feature encoded with the voxel.
[0064] Reference allele and alternative allele channels Voxels in the voxel grid can also have reference allele and alternative allele channels. In step 162, a one-hot encoder 164 of the system generates a reference one-hot encoding of a reference amino acid in the reference amino acid sequence and an alternative one-hot encoding of an alternative amino acid in the alternative amino acid sequence. The reference amino acid undergoes a target mutation. The alternative amino acid is the target mutation. The reference amino acid and the alternative amino acid are located at the same positions in the reference amino acid sequence and the alternative amino acid sequence, respectively. The reference amino acid sequence and the alternative amino acid sequence have the same position-by-position amino acid composition, with one exception. The exception is the position with the reference amino acid in the reference amino acid sequence and the alternative amino acid in the alternative amino acid sequence.
[0065] In step 172, a concatenator 174 of the system concatenates the amino acid-based distance channel with the reference and alternative one-hot encodings. In another embodiment, the concatenator 174 concatenates the atomic element-based distance channel with the reference and alternative one-hot encodings. In yet another embodiment, the concatenator 174 concatenates the atom type-based distance channel with the reference and alternative one-hot encodings.
[0066] In step 182, the system's runtime logic 184 processes the concatenated amino acid-by-amino acid / atom element-by-atom type-by-atom distance channels and the reference and alternative one-hot encodings through a pathogenicity classifier (pathogenicity determination engine) to determine the pathogenicity of the target variant, which is then inferred as the pathogenicity determination of the underlying nucleotide variant that generated the target variant at the amino acid level. The pathogenicity classifier is trained using a labeled dataset of benign and pathogenic variants, for example, using a backpropagation algorithm. Further details regarding labeled datasets of benign and pathogenic variants and exemplary architectures and training of pathogenicity classifiers can be found in commonly owned U.S. Patent Application Nos. 16 / 160,903, 16 / 160986, 16 / 160968, and 16 / 407149.
[0067] FIG. 2 schematically illustrates a reference amino acid sequence 202 of a protein 200 and an alternative amino acid sequence 212 of the protein 200. The protein 200 includes N amino acids. The amino acid positions in the protein 200 are labeled 1, 2, 3, ... N. In the illustrated example, position 16 is a position that experiences an amino acid variant 214 (mutation) caused by an underlying nucleotide variant. For example, for the reference amino acid sequence 202, position 1 has the reference amino acid phenylalanine (F), position 16 has the reference amino acid glycine (G) 204, and position N (e.g., the last amino acid of sequence 202) has the reference amino acid leucine (L). Although not illustrated for clarity, the remaining positions in the reference amino acid sequence 202 contain various amino acids in an order specific to the protein 200. The alternative amino acid sequence 212 is identical to the reference amino acid sequence 202 except for the variant 214 at position 16, which contains the alternative amino acid alanine (A) 214 instead of the reference amino acid glycine (G) 204.
[0068] Figure 3 shows the amino acid unit classification of the atoms of the amino acids in the reference amino acid sequence 202, also referred to herein as "atom classification 300." Certain types of amino acids among the 20 natural amino acids listed in column 302 may be repeated in a protein. That is, certain types of amino acids may occur more than once in a protein. A protein may also have some undetermined amino acids, which are classified by the 21st stop or gap amino acid category. The right column of Figure 3 shows the alpha carbons (C α ) atom counts.
[0069] Specifically, FIG. 3 shows the alpha carbon (C α) atom classification of the amino acid unit. Column 308 in Figure 3 lists the total number of alpha carbon atoms observed for reference amino acid sequence 202 in each of the 21 amino acid categories. For example, column 308 lists 11 alpha carbon atoms observed for the alanine (A) amino acid category. Since each amino acid has only one alpha carbon atom, this means that alanine occurs 11 times in reference amino acid sequence 202. In another example, arginine (R) appears 35 times in reference amino acid sequence 202. The total number of alpha carbon atoms across the 21 amino acid categories is 828.
[0070] Figure 4 shows the assignment of the 3D atomic coordinates of the alpha carbon atoms of the reference amino acid sequence 202 to amino acid units based on the atom classification 300 of Figure 3. This is referred to herein as "atomic coordinate bucketing 400." In Figure 4, lists 404-440 tabulate the 3D atomic coordinates of the alpha carbon atoms bucketed into each of the 21 amino acid categories.
[0071] In the illustrated embodiment, bucketing 400 in Figure 4 follows classification 300 in Figure 3. For example, in Figure 3, the alanine amino acid category has 11 alpha carbon atoms, and therefore, in Figure 4, the alanine amino acid category has 11 3D atomic coordinates of the corresponding 11 alpha carbon atoms from Figure 3. This classification-to-bucketing logic flows from Figure 3 to Figure 4 for other amino acid categories as well. However, this classification-to-bucketing logic is for representational purposes only, and in other embodiments, the disclosed techniques need not perform classification 300 and bucketing 400 to locate nearest neighbor atoms voxel-wise, but can perform fewer, additional, or different steps. For example, in some embodiments, the disclosed technology can locate nearest neighbor atoms voxel-wise by using a sorting and searching algorithm that returns nearest neighbor atoms voxel-wise from one or more databases in response to a search query configured to accept query parameters such as sorting criteria (e.g., amino acid-wise, atomic element-wise, atom type-wise), a predetermined maximum scanning radius, and distance type (e.g., Euclidean, Mahalanobis, normalized, non-normalized). In various embodiments of the disclosed technology, multiple sorting and searching algorithms from the current or future art can be similarly used by those skilled in the art to locate nearest neighbor atoms voxel-wise.
[0072] In Figure 4, the 3D atomic coordinates are represented by Cartesian coordinates x, y, z, although any type of coordinate system, such as spherical or cylindrical coordinates, may be used, and claimed subject matter is not limited in this respect. In some embodiments, one or more databases may contain information regarding the 3D atomic coordinates of the alpha carbon atoms of amino acids and other atoms in proteins. Such databases may be searchable by particular protein.
[0073] As noted above, voxels and voxel grids are 3D entities. However, for clarity, the figures show, and the description will discuss, voxels and voxel grids in two-dimensional (2D) format. For example, a 3x3x3 voxel grid of 27 voxels is shown and described herein as a 3x3 2D pixel grid with 9 2D pixels. Those skilled in the art will understand that the 2D format is used for representational purposes only and is intended to cover its 3D counterpart (i.e., a 2D pixel represents a 3D voxel, and a 2D pixel grid represents a 3D voxel grid). Also, the figures are not to scale. For example, a voxel of size 2 angstroms (Å) is depicted using a single pixel.
[0074] Voxel-wise distance calculation FIG. 5 illustrates a schematic process for determining voxel-wise distance values, also referred to herein as "voxel-wise distance calculation 500." In the illustrated example, voxel-wise distance values are calculated only for the alanine (A) distance channel. However, as described above with respect to FIG. 1, the same distance calculation logic can be performed for each of the 21 amino acid categories to generate 21 amino acid-wise distance channels, and can be further extended to other atom types, such as beta carbon atoms, and other atomic elements, such as oxygen, nitrogen, and hydrogen. In some embodiments, atoms are randomly rotated before distance calculation to make the pathogenicity classifier training invariant to atom orientation.
[0075] 5, voxel grid 522 has nine voxels 514 identified by indices (1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), and (3,3). Voxel grid 522 is centered, for example, on the 3D atomic coordinate 532 of the alpha carbon atom of the glycine (G) amino acid at position 16 of reference amino acid sequence 202 because, as described above with respect to FIG. 2, position 16 in alternative amino acid sequence 212 experiences a mutation that mutates the glycine (G) amino acid to an alanine (A) amino acid. Additionally, the center of voxel grid 522 coincides with the center of voxel (2,2).
[0076] The central voxel grid 522 is used for voxel-by-voxel distance calculations for each of the 21 amino acid-by-amino acid distance channels. For example, starting with the alanine (A) distance channel, the distances between the 3D coordinates of the centers of each of the nine voxels 514 and the 3D atomic coordinates 402 of the 11 alanine alpha carbon atoms are measured to locate the nearest alanine alpha carbon atom for each of the nine voxels 514. The nine distance values for the nine distances between the nine voxels 514 and their respective nearest alanine alpha carbon atoms are then used to construct the alanine distance channel. The resulting alanine distance channel arranges the nine alanine distance values in the same order as the nine voxels 514 in the voxel grid 522.
[0077] The above process is performed for each of the 21 amino acid categories. For example, the central voxel grid 522 is similarly used to calculate the arginine (R) distance channel, measuring the distance between the 3D coordinates of the center of each of the nine voxels 514 and the 3D atomic coordinates 404 of the 35 arginine alpha carbon atoms to identify the location of the nearest arginine alpha carbon atom for each of the nine voxels 514. The nine distance values for the nine distances between the nine voxels 514 and their respective nearest arginine alpha carbon atoms are then used to construct the arginine distance channel. The resulting arginine distance channel arranges the nine arginine distance values in the same order as the nine voxels 514 in the voxel grid 522. The 21 amino acid-based distance channels are encoded voxel-by-voxel to form a distance channel tensor.
[0078] Specifically, in the illustrated example, distance 512 is the distance between the center of voxel (1,1) in voxel grid 522 and Cα A5 The nearest alpha carbon atom (C α ) atom. Therefore, the value assigned to voxel (1,1) is distance 512. A4 The atom is located at the center of voxel (1,2) with the nearest C α Therefore, the value assigned to voxel (1,2) is the value between the center of voxel (1,2) and Cα A4 In another example, the distance between the Cα A6 The atom is located at the center of the voxel (2,1) with the nearest C α Therefore, the value assigned to voxel (2,1) is the value between the center of voxel (2,1) and Cα A6 In another example, the distance between the Cα A6 The atoms are also located at the centers of voxels (3,2) and (3,3) and are closest to C α Therefore, the value assigned to voxel (3,2) is the center of voxel (3,2) and Cα A6 The value assigned to voxel (3,3) is the distance between the center of voxel (3,3) and Cα A6The distance between the atoms. In some embodiments, the distance value assigned to voxel 514 may be a normalized distance. For example, the distance value assigned to voxel (1,1) may be distance 512 divided by maximum distance 502 (a predetermined maximum scan radius). In some embodiments, the nearest neighbor atom distance may be a Euclidean distance, and the nearest neighbor atom distance may be normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance (e.g., maximum distance 502).
[0079] As noted above, for amino acids having an alpha carbon atom, the distance can be the nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding amino acid. Additionally, for amino acids having a beta carbon atom, the distance can be the nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding amino acid. Similarly, for amino acids having backbone atoms, the distance can be the nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding amino acid. Similarly, for amino acids having side chain atoms, the distance can be the nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding amino acid. In some embodiments, the distance can additionally / alternatively include the distance to the second, third, fourth nearest atom, etc.
[0080] Amino acid distance channel FIG. 6 shows an example of 21 amino acid-based distance channels 600. Each column in FIG. 6 corresponds to one of the 21 amino acid-based distance channels 602-642. Each amino acid-based distance channel includes a distance value for each of the voxels 514 in the voxel grid 522. For example, the amino acid-based distance channel 602 for alanine (A) includes a distance value for each of the voxels 514 in the voxel grid 522. As described above, the voxel grid 522 is a 3D grid with a volume of 3×3×3 and includes 27 voxels. Similarly, while FIG. 6 shows the voxels 514 in two dimensions (e.g., 9 voxels in a 3×3 grid), each amino acid-based distance channel may include 27 voxel-based distance values for the 3×3×3 voxel grid.
[0081] directional encoding In some embodiments, the disclosed technology uses the directionality parameter to identify the directionality of a reference amino acid within reference amino acid sequence 202. In some embodiments, the disclosed technology uses the directionality parameter to identify the directionality of an alternative amino acid within alternative amino acid sequence 212. In some embodiments, the disclosed technology uses the directionality parameter to identify positions within protein 200 that will experience a target mutation at the amino acid level.
[0082] As described above, all distance values in the 21 amino acid unit distance channels 602-642 are measured from their respective nearest neighbor atoms to voxels 514 in the voxel grid 522. These nearest neighbor atoms originate from one of the reference amino acids in the reference amino acid sequence 202. These original reference amino acids containing nearest neighbor atoms can be classified into two categories: (1) the original reference amino acids preceding the variant-experiencing reference amino acid 204 in the reference amino acid sequence 202, and (2) the original reference amino acids following the variant-experiencing reference amino acid 204 in the reference amino acid sequence 202. The original reference amino acids in the first category can be referred to as preceding reference amino acids. The original reference amino acids in the second category can be referred to as following reference amino acids.
[0083] The directionality parameter is applied to distance values in distance channels 602-642 of 21 amino acid units measured from the nearest neighbor atom from the preceding reference amino acid. In one embodiment, the directionality parameter is multiplied by such distance values. The directionality parameter can be any number, such as -1.
[0084] As a result of the application of the directionality parameter, the 21 amino acid distance channel 600 contains several distance values that indicate to the pathogenicity classifier which end of the protein 200 is the start end and which end is the end end, which also allows the pathogenicity classifier to reconstruct the protein sequence from the 3D protein structure information provided by the distance channel and the reference and allele channels.
[0085] Distance Channel Tensor FIG. 7 is a schematic diagram of a distance channel tensor 700. Distance channel tensor 700 is a voxelized representation of the amino acid-based distance channel 600 from FIG. 6. In distance channel tensor 700, the 21 amino acid-based distance channels 602-642 are concatenated voxel-by-voxel, like the RGB channels of a color image. The voxelized dimensionality of distance channel tensor 700 is 21×3×3×3 (where 21 indicates the 21 amino acid categories and 3×3×3 indicates a 3D voxel grid with 27 voxels). However, FIG. 7 is a 2D depiction of the dimensionality 21×3×3.
[0086] One-hot coding Figure 8 shows one-hot encoding 800 of reference amino acids 204 and alternative amino acids 214. In Figure 8, the left column is one-hot encoding 802 of reference amino acid glycine (G) 204, with 1 for the glycine amino acid category and 0 for all other amino acid categories. In Figure 8, the right column is one-hot encoding 804 of variant / alternative amino acid alanine (A) 214, with 1 for the alanine amino acid category and 0 for all other amino acid categories.
[0087] 9 is a schematic diagram of a voxelized one-hot encoded reference amino acid 902 and a voxelized one-hot encoded variant / alternative amino acid 912. The voxelized one-hot encoded reference amino acid 902 is a voxelized representation of the one-hot encoding 802 of the reference amino acid glycine (G) 204 from FIG. 8. The voxelized one-hot encoded alternative amino acid 912 is a voxelized representation of the one-hot encoding 804 of the variant / alternative amino acid alanine (A) 214 from FIG. 8. The voxelized one-hot encoded reference amino acid 902 has a voxelized dimensionality of 21×1×1×1 (where 21 represents 21 amino acid categories). However, FIG. 9 is a 2D depiction with a dimensionality of 21×1×1. Similarly, the voxelized one-hot encoded alternative amino acid 912 has a voxelized dimensionality of 21×1×1×1 (where 21 represents 21 amino acid categories). However, Figure 9 is a 2D depiction with dimensionality 21 × 1 × 1.
[0088] Reference allele tensor Figure 10 schematically illustrates a concatenation process 1000 for voxel-wise concatenation of the distance channel tensor 700 of Figure 7 with a reference allele tensor 1004. The reference allele tensor 1004 is a voxel-wise collection (repeated / cloned / duplicated) of the voxelized one-hot encoded reference amino acids 902 from Figure 9. That is, the multiple copies of the voxelized one-hot encoded reference amino acids 902 are voxel-wise concatenated to each other according to the spatial arrangement of the voxels 514 within the voxel grid 522, such that the reference allele tensor 1004 has a corresponding copy of the voxelized one-hot encoded reference amino acid 910 for each of the voxels 514 within the voxel grid 522.
[0089] The concatenation process 1000 produces a concatenation tensor 1010. The voxelized dimensionality of the reference allele tensor 1004 is 21×3×3×3 (where 21 indicates 21 amino acid categories and 3×3×3 indicates a 3D voxel grid with 27 voxels). However, FIG. 10 is a 2D representation of the reference allele tensor 1004, which has dimensionality 21×3×3. The voxelized dimensionality of the concatenation tensor 1010 is 42×3×3×3. However, FIG. 10 is a 2D representation of the concatenation tensor 1010, which has dimensionality 42×3×3.
[0090] Alternative allele tensor Figure 11 shows a schematic diagram of a concatenation process 1100 for voxel-wise concatenation of the distance channel tensor 700 of Figure 7, the reference allele tensor 1004 of Figure 10, and an alternative allele tensor 1104. The alternative allele tensor 1104 is a voxel-wise collection (repeated / cloned / duplicated) of the voxelized one-hot encoded alternative amino acids 912 from Figure 9. That is, the multiple copies of the voxelized one-hot encoded alternative amino acids 912 are concatenated voxel-wise to each other according to the spatial arrangement of the voxels 514 in the voxel grid 522, such that the alternative allele tensor 1104 has a corresponding copy of the voxelized one-hot encoded alternative amino acid 910 for each of the voxels 514 in the voxel grid 522.
[0091] The concatenation process 1100 produces a concatenation tensor 1110. The voxelized dimensionality of the alternative allele tensor 1104 is 21×3×3×3 (where 21 indicates 21 amino acid categories and 3×3×3 indicates a 3D voxel grid with 27 voxels). However, FIG. 11 is a 2D depiction of the alternative allele tensor 1104, which has dimensionality 21×3×3. The voxelized dimensionality of the concatenation tensor 1110 is 63×3×3×3. However, FIG. 11 is a 2D depiction of the concatenation tensor 1110, which has dimensionality 63×3×3.
[0092] In some embodiments, the runtime logic 184 processes the concatenated tensor 1110 through a pathogenicity classifier to determine the pathogenicity of the variant / alternative amino acid alanine (A) 214, which is then inferred as the pathogenicity determination of the underlying nucleotide variant that generates the variant / alternative amino acid alanine (A) 214.
[0093] Evolutionarily conserved channels Predicting the functional consequences of variants relies, at least in part, on the assumption that amino acids important to a protein family have been conserved through evolution by negative selection (i.e., amino acid changes at these sites have been deleterious in the past), and that mutations at these sites are likely to be pathogenic (disease-causing) in humans. Generally, homologous sequences of a target protein are collected and aligned, and a conservation metric is calculated based on the weighted frequency of different amino acids observed at the target position in the alignment.
[0094] Thus, the disclosed technology connects the distance channel tensor 700, the reference allele tensor 1004, and the alternative allele tensor 1104 with evolutionary channels. One example of an evolutionary channel is the general amino acid conservation frequency. Another example of an evolutionary channel is the per-amino acid conservation frequency.
[0095] In some embodiments, evolutionary channels are constructed using position-specific weight matrices (PWMs). In other embodiments, evolutionary channels are constructed using position-specific frequency matrices (PSFMs). In still other embodiments, evolutionary channels are constructed using computational tools such as SIFT, PolyPhen, and PANTHER-PSEC. In still other embodiments, evolutionary channels are conserved channels based on evolutionary conservation. Conservation is related to conservation because it also reflects the effects of negative selection that has acted to prevent evolutionary change at a given site in a protein.
[0096] Pan-amino acid evolutionary profile Figure 12 is a flow diagram illustrating a process 1200 of a system for determining and assigning nearest-neighbor amino acid conservation frequencies to voxels (voxelization), according to one embodiment of the disclosed technology. Figures 12, 13, 14, 15, 16, 17, and 18 are described in parallel.
[0097] In step 1202, the system's similar sequence finder 1204 searches for amino acid sequences that are similar (homologous) to the reference amino acid sequence 202. Similar amino acid sequences can be selected from multiple species, such as primates, mammals, and vertebrates.
[0098] In step 1212, the system's aligner 1214 positionally aligns the reference amino acid sequence 202 with similar amino acid sequences, i.e., the aligner 1214 performs a multiple sequence alignment. Figure 14 shows an exemplary multiple sequence alignment 1400 of the reference amino acid sequence 202 across 99 species. In some embodiments, the multiple sequence alignment 1400 can be partitioned to generate, for example, a first position frequency matrix 1402 for primates, a second position frequency matrix 1412 for mammals, and a third position frequency matrix 1422 for primates. In other embodiments, a single position frequency matrix is generated across the 99 species.
[0099] In step 1222, the system's global amino acid conservation frequency calculator 1224 uses the multiple sequence alignment to determine the global amino acid conservation frequency of the reference amino acids in the reference amino acid sequence 202.
[0100] In step 1232, the system's nearest atom finder 1234 finds the nearest atom to voxel 514 within voxel grid 522. In some implementations, the voxel-wise search for nearest neighbor atoms may not be limited to any particular amino acid category or atom type. That is, voxel-wise nearest neighbor atoms can be selected across amino acid categories and types, as long as they are the atoms closest to the respective voxel centers. In other implementations, the voxel-wise search for nearest neighbor atoms may be limited to only particular atom categories, such as only specific atomic elements, such as oxygen, nitrogen, and hydrogen, or only alpha carbon atoms, or only beta carbon atoms, or only side chain atoms, or only backbone atoms.
[0101] In step 1242, the amino acid selector 1244 of the system selects a reference amino acid in the reference amino acid sequence 202 that contains the nearest atom identified in step 1232. Such a reference amino acid can be referred to as a nearest reference amino acid. Figure 13 shows an example of locating a nearest atom 1302 to a voxel 514 in the voxel grid 522 and mapping a nearest reference amino acid 1312 that contains the nearest atom 1302 to the voxel 514 in the voxel grid 522, respectively. This is identified in Figure 13 as "voxel to nearest amino acid mapping 1300."
[0102] In step 1252, the system's voxelizer 1254 voxelizes the global amino acid conservation frequencies of the nearest reference amino acids. Figure 15 shows an example of determining the global amino acid conservation frequency sequence for the first voxel (1,1) in the voxel grid 522, also referred to herein as "voxel-by-voxel evolutionary profile determination 1500."
[0103] 13, the nearest reference amino acid mapped to the first voxel (1,1) is the aspartic acid (D) amino acid at position 15 in reference amino acid sequence 202. A multiple sequence alignment of reference amino acid sequence 202 with, for example, 99 homologous amino acid sequences from 99 species is then analyzed at position 15. Such position-specific and cross-species analysis reveals how many instances of amino acids from each of the 21 amino acid categories are found at position 15 across the 100 aligned amino acid sequences (i.e., reference amino acid sequence 202 + 99 homologous amino acid sequences).
[0104] In the example shown in FIG. 15 , the aspartic acid (D) amino acid is found at position 15 in 96 of the 100 aligned amino acid sequences. Therefore, the aspartic acid amino acid category 1504 is assigned a pan-amino acid conservation frequency of 0.96. Similarly, in the example shown, the valine (V) acid amino acid is found at position 15 in 4 of the 100 aligned amino acid sequences. Therefore, the valine acid amino acid category 1514 is assigned a pan-amino acid conservation frequency of 0.04. Because no examples of amino acids from other amino acid categories are found at position 15, the remaining amino acid categories are assigned a pan-amino acid conservation frequency of 0. In this manner, each of the 21 amino acid categories is assigned a respective pan-amino acid conservation frequency, which can be encoded in the pan-amino acid conservation frequency array 1502 for the first voxel (1,1).
[0105] FIG. 16 shows the respective pan-amino acid conservation frequencies 1612-1692 determined for each of the voxels 514 within the voxel grid 522 using the positional frequency logic described in FIG. 15, also referred to herein as the "voxel to evolutionary profile mapping 1600."
[0106] The per-voxel evolutionary profile 1602 is then used by the voxelizer 1254 to generate the voxelized per-voxel evolutionary profile 1700 shown in Figure 17. In many cases, each of the voxels 514 in the voxel grid 522 will have a different global amino acid conservation frequency sequence and therefore a different per-voxel voxelized evolutionary profile, as the voxels regularly map to different nearest-neighbor atoms and thus different nearest-neighbor reference amino acids. Of course, if two or more voxels have the same nearest-neighbor atoms, and thereby the same nearest-neighbor reference amino acid, then the same global amino acid conservation frequency sequence and the same per-voxel voxelized evolutionary profile will be assigned to each of the two or more voxels.
[0107] 18 shows an example of an evolutionary profile tensor 1800 in which the voxelized per-voxel evolutionary profiles 1700 are connected to each other voxel-by-voxel according to the spatial arrangement of the voxels 514 within the voxel grid 522. The voxelized dimensionality of the evolutionary profile tensor 1800 is 21×3×3×3 (where 21 indicates 21 amino acid categories and 3×3×3 indicates a 3D voxel grid with 27 voxels). However, FIG. 18 is a 2D depiction of the evolutionary profile tensor 1800 with dimensionality 21×3×3.
[0108] In step 1262, the concatenator 174 concatenates the evolutionary profile tensor 1800 voxel-wise with the distance channel tensor 700. In some implementations, the evolutionary profile tensor 1800 is concatenated voxel-wise with the concatenation tensor 1110 to generate a further concatenation tensor (not shown) of dimensionality 84x3x3x3.
[0109] In step 1272, the runtime logic 184 processes a further concatenated tensor of dimensionality 84x3x3x3 through a pathogenicity classifier to determine the pathogenicity of the target variant, which is then inferred as the pathogenicity determination of the underlying nucleotide variant that generates the target variant at the amino acid level.
[0110] Amino acid-by-amino acid evolutionary profile 19 is a flow diagram illustrating a system process 1900 for determining the amino acid-by-amino acid conservation frequencies of nearest neighbor atoms and assigning them to voxels (voxelization). In FIG. 19, steps 1202 and 1212 are the same as in FIG. 12.
[0111] In step 1922, the system's per-amino acid conservation frequency calculator 1924 uses the multiple sequence alignment to determine the per-amino acid conservation frequency of the reference amino acids in the reference amino acid sequence 202.
[0112] In step 1932, the system's nearest atom finder 1934 finds, for each voxel 514 in the voxel grid 522, 21 nearest atoms across each of the 21 amino acid categories. Each of the 21 nearest neighbor atoms is different from the others because they are selected from different amino acid categories. This leads to the selection of 21 unique nearest reference amino acids for the particular voxel, which in turn leads to the generation of 21 unique position frequency matrices for the particular voxel, which in turn leads to the determination of the conservation frequency for each of the 21 unique amino acids for the particular voxel.
[0113] In step 1942, the system's amino acid selector 1944 selects, for each voxel 514 in the voxel grid 522, 21 reference amino acids in the reference amino acid sequence 202 that contain the 21 nearest neighbor atoms identified in step 1932. Such reference amino acids may be referred to as nearest neighbor reference amino acids.
[0114] In step 1952, the system's voxelizer 1954 voxelizes the pen-amino acid conservation frequencies of the 21 nearest reference amino acids identified for a particular voxel in step 1942. Because the 21 nearest reference amino acids correspond to different underlying nearest neighbor atoms, they are necessarily located at 21 different positions in the reference amino acid sequence 202. Thus, for a particular voxel, 21 position frequency matrices can be generated for the 21 nearest reference amino acids. The 21 position frequency matrices can be generated across multiple species whose homologous amino acid sequences are positionally aligned with the reference amino acid sequence 202, as described above with respect to Figures 12-15.
[0115] Using the 21 position frequency matrices, 21 position-specific conservation scores can then be calculated for the 21 nearest reference amino acids identified for a particular voxel. These 21 position-specific conservation scores form the pen-amino acid conservation frequency for a particular voxel, similar to the pan-amino acid conservation frequency array 1502 of Figure 12. However, because the 21 nearest reference amino acids across the 21 amino acid categories will necessarily have different positions that result in different position frequency matrices and thereby different per-amino acid conservation frequencies, array 1502 has many zero entries, but each element (feature) in the per-amino acid conservation frequency array has a value (e.g., a floating-point number).
[0116] The above process is performed for each voxel 514 in the voxel grid 522, and the resulting voxel-wise amino acid conservation frequencies are voxelized, tensorized, connected, and processed for pathogenicity determination in the same manner as the general amino acid conservation frequencies described with respect to Figures 12-18.
[0117] Annotation Channel 20 shows various examples of voxelized annotation channels 2000 concatenated with distance channel tensors 700. In some embodiments, the voxelized annotation channels are one-hot indicators of different protein annotations, such as whether an amino acid (residue) is part of a transmembrane region, a signal peptide, an active site, or any other binding site, or whether the residue undergoes a post-translational modification PathRatio (see Pei P, Zhang A: A Topological Measurement for Weighted Protein Interaction Network. CSB 2005, 268-278). Additional examples of annotation channels can be found in the specific embodiments section below and in the claims.
[0118] The voxelized annotation channels are arranged voxel-by-voxel such that voxels can have the same annotation sequence, such as voxelized reference allele and alternative allele sequences (e.g., annotation channels 2002, 2004, 2006), or voxels can have respective annotation sequences, such as voxelized per-voxel evolutionary profile 1700 (e.g., annotation channels 2012, 2014, 2016 (as indicated by different colors)).
[0119] The annotation channels are voxelized, tensorized, concatenated, and processed for pathogenicity determination, similar to the pan-amino acid conservation frequency described with respect to Figures 12-18.
[0120] Structural Reliability Channel The disclosed technology can also concatenate various voxelized structure confidence channels with the distance channel tensor 700. Some examples of structure confidence channels include the GMQE score (provided by SwissModel); B-factor; the homology model temperature factor string (which indicates how well a residue satisfies (physical) constraints in a protein structure); the normalized number of aligned template proteins to the residue closest to the center of the voxel (for an alignment provided by HHpred, e.g., a voxel is closest to the residue to which three of the six template structures align, meaning the feature has a value of 3 / 6=0.5); the minimum, maximum, and average TM score; and the predicted TM score of the template protein structure aligned to the residue closest to the voxel (continuing the example above, suppose three template structures have TM scores of 0.5, 0.5, and 1.5, the minimum is 0.5, the average is 2 / 3, and the maximum is 1.5). A TM score can be provided for each protein template by HHpred. Additional examples of structural confidence channels can be found in the specific implementations section and claims below.
[0121] The voxelized structural confidence channels are arranged voxel by voxel such that voxels can have the same structural confidence sequences, such as voxelized reference allele and alternative allele sequences, or voxels can have respective structural confidence sequences, such as voxelized per-voxel evolutionary profile 1700.
[0122] The structural confidence channel is voxelized, tensorized, concatenated, and processed for pathogenicity determination, similar to the general amino acid conservation frequency described with respect to Figures 12-18.
[0123] Pathogenicity classifier FIG. 21 illustrates different combinations and permutations of input channels that may be provided as inputs 2102 to a pathogenicity classifier 2108 for determining 2106 the pathogenicity of a target variant. One of the inputs 2102 may be a distance channel 2104 generated by a distance channel generator 2272. FIG. 22 illustrates various methods for calculating the distance channel 2104. In one embodiment, the distance channel 2104 is generated based on distances 2202 between voxel centers and atoms across multiple atomic elements, regardless of amino acid. In some embodiments, the distances 2202 are normalized by the maximum scan radius to generate normalized distances 2202a. In another embodiment, the distance channel 2104 is generated based on distances 2212 between voxel centers and alpha carbon atoms on an amino acid basis. In some embodiments, the distances 2212 are normalized by the maximum scan radius to generate normalized distances 2212a. In yet another embodiment, the distance channel 2104 is generated based on distances 2222 between voxel centers and beta carbon atoms on an amino acid basis. In some embodiments, distance 2222 is normalized by the maximum scan radius to generate normalized distance 2222a. In yet another embodiment, distance channel 2104 is generated based on distance 2232 between the voxel center and side chain atoms on an amino acid basis. In some embodiments, distance 2232 is normalized by the maximum scan radius to generate normalized distance 2232a. In yet another embodiment, distance channel 2104 is generated based on distance 2242 between the voxel center and backbone atoms on an amino acid basis. In some embodiments, distance 2242 is normalized by the maximum scan radius to generate normalized distance 2242a. In yet another embodiment, distance channel 2104 is generated based on distance 2252 (one feature) between the voxel center and each nearest neighbor atom, regardless of atom type and amino acid type. In yet another embodiment, distance channel 2104 is generated based on distance 2262 (one feature) between the voxel center and atoms from non-standard amino acids. In some embodiments, the distance between a voxel and an atom is calculated based on the polar coordinates of the voxel and the atom.The polar coordinates are parameterized by the angle between the voxel and the atom. In one implementation, this angle information is used to generate the angle channel of the voxel (i.e., independently of the distance channel). In some implementations, the angle between the nearest neighbor atom and a neighboring atom (e.g., a backbone atom) can be used as a feature encoded with the voxel.
[0124] Another of the inputs 2102 can be a feature 2114 that indicates missing atoms within a specified radius.
[0125] Another of the inputs 2102 can be one-hot encoding of a reference amino acid 2124. Another of the inputs 2102 can be one-hot encoding of a variant / alternative amino acid 2134.
[0126] Another of the inputs 2102 can be evolutionary channels 2144 generated by the evolutionary profile generator 2372 shown in Figure 23. In one embodiment, the evolutionary channels 2144 can be generated based on the global amino acid conservation frequencies 2302. In another embodiment, the evolutionary channels 2144 can be generated based on the global amino acid conservation frequencies 2312.
[0127] Another of the inputs 2102 can be features 2154 that indicate deleted residues or a deletion evolutionary profile.
[0128] Another of the inputs 2102 may be annotation channels 2164 generated by annotation generator 2472 shown in Figure 24. In one embodiment, annotation channel 2164 may be generated based on molecular processing annotations 2402. In another embodiment, annotation channel 2164 may be generated based on region annotations 2412. In yet another embodiment, annotation channel 2164 may be generated based on site annotations 2422. In yet another embodiment, annotation channel 2164 may be generated based on amino acid modification annotations 2432. In yet another embodiment, annotation channel 2164 may be generated based on secondary structure annotations 2442.
[0129] In yet another embodiment, the annotation channel 2164 can be generated based on the experiment information annotations 2452 .
[0130] Another of the inputs 2102 may be a structure confidence channel 2174 generated by a structure confidence generator 2572 shown in FIG. 25. In one embodiment, the structure confidence 2174 may be generated based on a global model quality estimate (GMQE) 2502. In another embodiment, the structure confidence 2174 may be generated based on a qualitative model energy analysis (QMEAN) score 2512. In yet another embodiment, the structure confidence 2174 may be generated based on a temperature factor 2522. In yet another embodiment, the structure confidence 2174 may be generated based on a template modeling score 2542. Examples of the template modeling score 2542 include a minimum template modeling score 2542a, an average template modeling score 2542b, and a maximum template modeling score 2542c.
[0131] Those skilled in the art will understand that any permutation and combination of input channels may be concatenated as inputs for processing through the pathogenicity classifier 2108 for pathogenicity determination 2106 of the target variant. In some embodiments, only a subset of the input channels may be concatenated. The input channels may be concatenated in any order. In one embodiment, the input channels may be concatenated into a single tensor by a tensor generator (input encoder) 2110. This single tensor may then be provided as an input to the pathogenicity classifier 2108 for pathogenicity determination 2106 of the target variant.
[0132] In one embodiment, the pathogenicity classifier 2108 uses a convolutional neural network (CNN) with multiple convolutional layers. In another embodiment, the pathogenicity classifier 2108 uses a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), and a gated recurrent unit (GRU). In yet another embodiment, the pathogenicity classifier 2108 uses both a CNN and an RNN. In yet another embodiment, the pathogenicity classifier 2108 uses a graph convolutional neural network that models dependencies in graph-structured data. In yet another embodiment, the pathogenicity classifier 2108 uses a variational autoencoder (VAE). In yet another embodiment, the pathogenicity classifier 2108 uses a generative adversarial network (GAN). In yet another embodiment, the pathogenicity classifier 2108 can also be a self-attention based language model, such as those implemented by Transformation and BERT.
[0133] In yet other embodiments, the pathogenicity classifier 2108 may use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth-separable convolution, pointwise convolution, 1x1 convolution, group convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. It may use one or more loss functions such as logistic regression / logarithmic loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. It can use any parallel, efficient, and compression scheme, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). It can include upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (e.g., LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., rectified linear units (ReLU), leaky ReLU, exponential linear units (ELU), sigmoid, and hyperbolic tangent functions (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, attention mechanisms, and Gaussian error linear units.
[0134] The pathogenicity classifier 2108 trains using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that the pathogenicity classifier 2108 may use to train include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that the pathogenicity classifier 2108 may use to train include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad. In other embodiments, the pathogenicity classifier 2108 may train by unsupervised learning, semi-supervised learning, self-learning, reinforcement learning, multitask learning, multimodal learning, transfer learning, knowledge distillation, etc.
[0135] Figure 26 illustrates an exemplary processing architecture 2600 of the pathogenicity classifier 2108 according to one implementation of the disclosed technology. The processing architecture 2600 includes a cascade of processing modules 2606, 2610, 2614, 2618, 2622, 2626, 2630, 2634, 2638, and 2642, each of which may include 1D convolution (1x1x1 CONV), 3D convolution (3x3x3 CONV), ReLU nonlinearity, and batch normalization (BN). Other examples of processing modules include a fully connected (FC) layer, a dropout layer, a smoothing layer, and a final softmax layer that generates exponentially normalized scores for target variants belonging to the benign and pathogenic classes. In Figure 26, "64" indicates the number of convolution filters applied by a particular processing module. In Figure 26, the size of the input voxel 2602 is 15x15x15x8. FIG. 26 also shows the volumetric dimensions of each of the intermediate inputs 2604 , 2608 , 2612 , 2616 , 2620 , 2624 , 2628 , 2632 , 2636 , and 2640 generated by the processing architecture 2600 .
[0136] Figure 27 illustrates an exemplary processing architecture 2700 of the pathogenicity classifier 2108, according to one implementation of the disclosed technology. The processing architecture 2700 includes a cascade of processing modules 2708, 2714, 2720, 2726, 2732, 2738, 2744, 2750, 2756, 2762, 2768, 2774, and 2780, such as 1D convolution (CONV 1D), 3D convolution (CONV 3D), ReLU nonlinearity, and batch normalization (BN). Other examples of processing modules include a fully connected (dense) layer, a dropout layer, a smoothing layer, and a final softmax layer that generates exponentially normalized scores for target variants belonging to the benign and pathogenic classes. In Figure 27, "64" and "32" indicate the number of convolution filters applied by a particular processing module. In Figure 27, the size of the input voxels 2704 provided by the input layer 2702 is 7x7x7x10. Figure 27 also shows the volumetric dimensions of each of the intermediate inputs 2710, 2716, 2722, 2728, 2734, 2740, 2746, 2752, 2758, 2764, 2770, 2776, and 2782 and the resulting intermediate outputs 2706, 2712, 2718, 2724, 2730, 2736, 2742, 2748, 2754, 2760, 2766, 2772, 2778, and 2784 generated by the processing architecture 2700.
[0137] Those skilled in the art will understand that other current and future artificial intelligence, machine learning, and deep learning models, datasets, and techniques can be incorporated into the disclosed variant pathogenicity classifier without departing from the spirit of the disclosed technology.
[0138] Performance Results as Objective Indicators of Inventiveness and Nonobviousness
[0139] The variant pathogenicity classifier disclosed herein makes pathogenicity predictions based on 3D protein structure and is referred to as "PrimateAI 3D." "Primate AI" is a previously disclosed, co-owned variant pathogenicity classifier that makes pathogenicity predictions based on protein sequence. Further details about PrimateAI can be found in co-owned U.S. Patent Application Nos. 16 / 160,903, 16 / 160,986, 16 / 160,968, and 16 / 407,149, and in Sundaram, L. et al., "Predicting the clinical impact of human mutations with deep neural networks." Nat. Genet. 50, 1161-1170 (2018).
[0140] Figures 28, 29, 30, and 31 use PrimateAI as a benchmark model to demonstrate the classification superiority of PrimateAI 3D over PrimateAI. The performance results in Figures 28, 29, 30, and 31 are generated based on the classification task of accurately distinguishing benign variants from pathogenic variants across multiple validation sets. PrimateAI 3D is trained on a training set that is different from the multiple validation sets. PrimateAI 3D is trained on common human variants and primate-derived variants, which are used as benign datasets, while simulated variants based on trinucleotide context are used as unlabeled or pseudopathogenic datasets.
[0141] De novo developmental delay disorder (De novo DDD) is an example of a validation set used to compare the classification accuracy of Primate AI 3D against Primate AI. The De novo DDD validation set labels variants from individuals with DDD as pathogenic and the same variants from healthy relatives of individuals with DDD as benign. A similar labeling scheme is used with the Autism Spectrum Disorder (ASD) validation set, shown in Figure 31.
[0142] BRCA1 is another example of a validation set used to compare the classification accuracy of Primate AI 3D to Primate AI. The BRCA1 validation set labels synthetically generated reference amino acid sequences that simulate the BRCA1 gene protein as benign variants and synthetically modified allele amino acid sequences that simulate the BRCA1 gene protein as pathogenic variants. A similar labeling scheme is used with different validation sets for the TP53 gene, the TP53S3 gene and its variants, and other genes and their variants, as shown in Figure 31.
[0143] Figure 28 identifies the performance of the benchmark PrimateAI model with the horizontal bar (labeled "PAI") and the disclosed PrimateAI 3D model with the horizontal bar (labeled "ens10_7x7x7x2_hhpred_evo+alt"). The horizontal bar labeled "ens10_7x7x7x2_hhpred_evo+alt_paisum" shows the virulence prediction derived by combining the respective virulence predictions of the disclosed PrimateAI 3D model and the benchmark PrimateAI model. In the legend, "ens10" refers to an ensemble of 10 PrimateAI 3D models, each trained with a different seed training dataset and randomly initialized with different weights and biases. Additionally, "7x7x7x2" refers to the size of the voxel grid used to encode the input channels during training of the ensemble of 10 PrimateAI 3D models. For a given variant, an ensemble of 10 PrimateAI 3D models will each generate 10 pathogenicity predictions, which are then combined (e.g., by averaging) to generate a final pathogenicity prediction for the given variant. This logic applies equally to ensembles of different group sizes.
[0144] Also in Figure 28, the y-axis has different validation sets, and the x-axis has p-values. A larger p-value, i.e., a longer horizontal bar, indicates a higher accuracy in distinguishing benign variants from pathogenic variants. As demonstrated by the p-values in Figure 28, PrimateAI 3D outperforms PrimateAI across the majority of validation sets (the only exception is the tp53s3_A549 validation set). That is, the horizontal bar of PrimateAI 3D (labeled "ens10_7x7x7x2_hhpred_evo+alt") is consistently longer than the horizontal bar of PrimateAI (labeled "PAI").
[0145] Also in Figure 28, the "average" category along the y-axis calculates the average of the p-values determined for each of the validation sets. Even in the average category, PrimateAI 3D outperforms PrimateAI.
[0146] In Figure 29, PrimateAI is represented by the horizontal bar (labeled "PAI"), an ensemble of 20 PrimateAI 3D models trained on a voxel grid of size 3x3x3 is represented by the horizontal bar labeled "ns20_3x3x3x2_evo+alt", an ensemble of 10 PrimateAI 3D models trained on a voxel grid of size 7x7x7 is represented by the horizontal bar labeled "ens10_7x7x7x2_evo+alt", an ensemble of 20 PrimateAI 3D models trained on a voxel grid of size 7x7x7 is represented by the horizontal bar labeled "ens20_7x7x7x2_evo+alt", and an ensemble of 20 PrimateAI 3D models trained on a voxel grid of size 17x17x17. The ensemble of 3D models is represented by a horizontal bar labeled "ens20_17x17x17x2_evo+alt".
[0147] Also in Figure 29, the y-axis has different validation sets, and the x-axis has p-values. As before, a larger p-value, i.e., a longer horizontal bar, indicates a higher accuracy in distinguishing benign from pathogenic variants. As demonstrated by the p-values in Figure 20, different configurations of PrimateAI 3D outperform PrimateAI across the majority of the validation set. That is, the horizontal bars representing ensembles of multiple PrimateAI 3D models (labeled "ns20_3x3x3x2_evo+alt," "ens10_7x7x7x2_evo+alt," "ens20_7x7x7x2_evo+alt," and "ens20_17x17x17x2_evo+alt") are mostly longer than the horizontal bars of PrimateAI (labeled "PAI").
[0148] Also in Figure 29, the "average" category along the y-axis calculates the average of the p-values determined for each of the validation sets. Even in the average category, different configurations of PrimateAI 3D outperform PrimateAI.
[0149] In Figure 30, the darkly shaded vertical bars represent PrimateAI ("PrimateAI(v1)") and the lightly shaded vertical bars represent PrimateAI 3D ("PrimateAI 3D"). In Figure 30, the y-axis has the p-value and the x-axis has the different validation sets. In Figure 30, without exception, PrimateAI 3D consistently outperforms PrimateAI across all of the validation sets; that is, the PrimateAI 3D vertical bars are always longer than the PrimateAI vertical bars.
[0150] Figure 31 identifies the performance of the benchmark PrimateAI model with the vertical bar labeled "PAI-v1_plain" and the disclosed PrimateAI 3D model with the vertical bar labeled "PAI-3D-orig_plain." The vertical bar labeled "PAI-3D-orig_paisum" shows the virulence prediction derived by combining the respective virulence predictions of the disclosed PrimateAI 3D model and the benchmark PrimateAI model. In Figure 31, the y-axis has p-values and the x-axis has different validation sets.
[0151] As demonstrated by the p-values in Figure 31, PrimateAI 3D outperforms PrimateAI across the majority of the validation set (the only exception being the tp53s3_A549_p53NULL_Nutlin-3 validation set), i.e., PrimateAI 3D columns are consistently longer than PrimateAI columns.
[0152] Also in Figure 31, a separate "Mean" chart calculates the average of the p-values determined for each of the validation sets. In the mean chart, PrimateAI 3D also outperforms PrimateAI.
[0153] The average statistics may be biased by outliers. To address this, a separate "Method Rank" chart is also shown in Figure 31. Higher ranks indicate lower classification accuracy. Similarly in the Method Rank chart, PrimateAI 3D outperforms PrimateAI by having more counts of the lower ranks 1 and 2, compared to PrimateAI, which has all 3s.
[0154] 28-31, it is also clear that superior classification accuracy can be achieved by combining PrimateAI 3D with PrimateAI: a protein can be fed into PrimateAI as an amino acid sequence to generate a first output, and the same protein can be fed into PrimateAI 3D as a 3D voxelized protein structure to generate a second output, and the first and second outputs can be combined or analyzed as a whole to generate a final pathogenicity prediction for the variants experienced by the protein.
[0155] Efficient Voxelization FIG. 32 is a flow chart illustrating an efficient voxelization process 3200 that efficiently identifies nearest neighbor atoms for each voxel.
[0156] Here, distance channels are discussed again. As described above, the reference amino acid sequence 202 can contain different types of atoms, such as alpha carbon atoms, beta carbon atoms, oxygen atoms, nitrogen atoms, and hydrogen atoms. Therefore, as described above, distance channels can be arranged by nearest alpha carbon atoms, nearest beta carbon atoms, nearest oxygen atoms, nearest nitrogen atoms, nearest hydrogen atoms, and so on. For example, in FIG. 6, each of the nine voxels 514 has a 21-amino acid distance channel for the nearest alpha carbon atom. FIG. 6 can be further expanded so that each of the nine voxels 514 also has a 21-amino acid distance channel for the nearest beta carbon atom, regardless of atom type and amino acid type, and also has a nearest general atom distance channel for the nearest atom. In this way, each of the nine voxels 514 can have 43 distance channels.
[0157] We now describe the number of distance calculations required to identify the nearest neighbor atom per voxel for inclusion in a distance channel. Consider the example in Figure 3, which shows a total of 828 alpha carbon atoms distributed across 21 amino acid categories. To calculate the amino acid-wise distance channels 602-642 in Figure 6, i.e., to determine 189 distance values, the distance from each of the nine voxels 514 to each of the 828 alpha carbon atoms is measured. * We get 828 = 7,452 distance calculations. For a 3D image with 27 voxels, this means 27 * This gives a distance calculation of 828 = 22,356. If the 828 beta carbon atoms are also included, this number becomes 27 * This increases the distance calculation to 1656 = 44,712.
[0158] This is illustrated by Figure 35A, where the runtime complexity of identifying nearest neighbor atoms per voxel for a single protein voxelization is O(number of atoms) * Furthermore, when distance channels are computed across various attributes (e.g., different features or channels per voxel, such as annotation channels and structural confidence channels), the runtime complexity of single protein voxelization is O(number of atoms). * Number of voxels * The number of attributes increases.
[0159] As a result, distance calculations can be the most computationally intensive part of the voxelization process, taking valuable computational resources away from important runtime tasks like model training and model inference. For example, consider the case of model training using a training dataset of 7,000 proteins. Generating distance channels for multiple voxels across multiple amino acids, atoms, and attributes can involve over 100 voxelizations per protein, resulting in approximately 800,000 voxelizations in a single training iteration (epoch). A training run of 20-40 epochs with rotation of atomic coordinates at each epoch can result in as many as 32 million voxelizations.
[0160] In addition to the high computational cost, the size of the data for 32 million voxelizations is too large to fit into main memory (e.g., over 20 TB for a 15 × 15 × 15 voxel grid). Considering the iterative training runs for parameter optimization and ensemble training, the memory footprint of the voxelization process is too large to store on disk, making the voxelization process part of the model training and not a pre-computation step.
[0161] The disclosed technology is based on the O (atomic number * The present invention provides an efficient voxelization process that achieves up to approximately a 100x speedup over a runtime complexity of O(number of voxels). The disclosed efficient voxelization process reduces the runtime complexity for single protein voxelization to O(number of atoms). For features or channels that differ per voxel, the disclosed efficient voxelization process reduces the runtime complexity for single protein voxelization to O(number of atoms). * As a result, the voxelization process becomes as fast as model training, shifting the computational bottleneck from voxelization back to computing neural network weights on processors such as GPUs, ASICs, TPUs, FPGAs, and CGRAs.
[0162] In some embodiments of the disclosed efficient voxelization process involving large voxel grids, the runtime complexity for single protein voxelization is O(number of atoms + voxel) and O(number of atoms + voxel) for different features or channels per voxel. * The "+voxel" complexity is observed when the number of atoms is very small compared to the number of voxels, e.g., one atom in a 100x100x100 voxel grid (i.e., 1 million voxels per atom). In such scenarios, the runtime is dominated by the overhead of the huge number of voxels, e.g., allocating memory for 1 million voxels, initializing 1 million voxels to 0, etc.
[0163] We will now describe in detail the disclosed efficient voxelization process. Figures 32A, 32B, 33, 34, and 35B will be described in parallel.
[0164] Starting with Figure 32A, in step 3202, each atom (e.g., each of the 828 alpha carbon atoms and each of the 828 beta carbon atoms) is associated with a voxel (e.g., one of the nine voxels 514) that contains the atom. The term "contains" refers to the 3D atomic coordinates of the atom located within the voxel. A voxel that contains an atom is also referred to herein as an "atom-containing voxel."
[0165] Figures 32B and 33 illustrate how voxels containing particular atoms are selected. Figure 33 uses 2D atomic coordinates as a representation of 3D atomic coordinates. Note that the voxel grid 522 is equally spaced, with each of the voxels 514 being equally spaced, with the same step size (e.g., 1 Angstrom (Å) or 2 Å).
[0166] Also in Figure 33, voxel grid 522 has indices [0,1,2] along a first dimension (e.g., the x-axis) and indices [0,1,2] along a second dimension (e.g., the y-axis). Also in Figure 33, each voxel 514 in voxel grid 522 is identified by a voxel index [voxel 0, voxel 1, ..., voxel 8] and by a voxel center index [(1,1), (1,2), ..., (3,3)].
[0167] Also, in Figure 33, the central coordinates of the voxel centers along the first dimension, i.e., first dimension voxel coordinates, are identified. Also, in Figure 33, the central coordinates of the voxel centers along the second dimension, i.e., second dimension voxel coordinates, are identified.
[0168] First, in step 3202a (step 1 in Figure 33), the 3D atomic coordinates of a particular atom (1.7456, 2.14323) are quantized to generate quantized 3D atomic coordinates (1.7, 2.1). Quantization can be achieved by rounding or truncating bits.
[0169] Then, in step 3202b (step 2 in FIG. 33), the voxel coordinates (or voxel centers or voxel center coordinates) of voxel 514 are assigned to dimension-wise quantized 3D atomic coordinates. For the first dimension, quantized atomic coordinate 1.7 is assigned to voxel 1 because it covers first-dimensional voxel coordinates ranging from 1 to 2 and is centered at 1.5 in the first dimension. Note that voxel 1 has index 1 along the first dimension, as opposed to having index 0 along the second dimension.
[0170] For the second dimension, starting with voxel 1, voxel grid 522 is traversed along the second dimension. This assigns quantized atomic coordinate 2.5 to voxel 7 because voxel 7 covers the range of second-dimensional voxel coordinates from 2 to 3 and is centered at 2.5 in the second dimension. Note that voxel 7 has index 2 along the second dimension, as opposed to having index 1 along the first dimension.
[0171] Next, in step 3202c (step 3 in FIG. 33), a dimension index corresponding to the assigned voxel coordinate is selected. That is, for voxel 1, index 1 is selected along the first dimension, and for voxel 7, index 2 is selected along the second dimension. Those skilled in the art will understand that the above steps can be similarly performed for the third dimension to select a dimension index along the third dimension.
[0172] Next, in step 3202d (step 4 in FIG. 33), a running sum is generated based on positionally weighting the selected dimension indices by powers of the radix. The general concept behind a positional numbering system is that numbers are represented by powers of a radix (or base); for example, binary numbers are base 2, ternary numbers are base 3, octal numbers are base 8, and hexadecimal numbers are base 16. This is often called a weighted numbering system because each position is weighted by a power of the radix. The set of valid digits for a positional numbering system is equal in size to the radix of that system. For example, the decimal system has 10 digits, 0 through 9, while the ternary system has 3 digits, 0, 1, and 2. The largest valid number in a radix system is 1 less than the radix (thus, 8 is not a valid number in any radix system less than 9). Any decimal integer can be represented exactly in any other integer radix system, and vice versa.
[0173] Returning to the example of Figure 33, the selected dimension indices 1 and 2 are converted to a single integer by position-wise multiplying them by their respective powers of base 3 and summing the results of the position-wise multiplications. Base 3 is chosen here because 3D atomic coordinates have three dimensions (although Figure 33 only shows 2D atomic coordinates along two dimensions for simplicity).
[0174] Index 2 is located at the rightmost bit (i.e., the least significant bit), so it is multiplied by 3 to the power of 0, which gives 2. Index 1 is located at the second bit from the right (i.e., the second least significant bit), so it is multiplied by 3 to the power of 1, which gives 3. This results in a running sum of 5.
[0175] Next, in step 3202e (step 5 in FIG. 33), voxel indices of voxels containing particular atoms are selected based on the cumulative sum, i.e., the cumulative sum is interpreted as voxel indices of voxels containing particular atoms.
[0176] In step 3212, after each atom is associated with the atom-containing voxel, each atom is further associated with one or more voxels in the vicinity of the atom-containing voxel, also referred to herein as “neighboring voxels.” Neighboring voxels may be selected based on being within a predetermined radius (e.g., 5 angstroms (Å)) of the atom-containing voxel. In other implementations, neighboring voxels may be selected based on being contiguous neighbors of the atom-containing voxel (e.g., neighboring voxels above, below, right, and left). The resulting associations associating each atom with the atom-containing voxel and the neighboring voxels are encoded in atom-to-voxel mapping 3402, also referred to herein as element-to-cell mapping. In one example, a first alpha carbon atom is associated with a first subset of voxels 3404 that includes the atom-containing voxel and the neighboring voxels of the first alpha carbon atom. In another example, the second alpha carbon atom is associated with a second subset of voxels 3406 that includes the atom-containing voxel and neighboring voxels of the second alpha carbon atom.
[0177] Note that no distance calculations are performed to determine atom-containing voxels and neighboring voxels. Atom-containing voxels are selected by a spatial arrangement of voxels that allows quantized 3D atom coordinates to be assigned to corresponding regularly spaced voxel centers in the voxel grid (without using distance calculations). And neighboring voxels are selected by being spatially adjacent to the atom-containing voxel in the voxel grid (again without using distance calculations).
[0178] In step 3222, each voxel is mapped to the atoms associated with it in steps 3202 and 3212. In one implementation, this mapping is encoded in voxel-to-atom mapping 3412, which is generated based on atom-to-voxel mapping 3402 (e.g., by applying a voxel-based sort key to atom-to-voxel mapping 3402). Voxel-to-atom mapping 3412 is also referred to herein as "cell-to-element mapping." In one example, in steps 3202 and 3212, a first voxel is mapped to a first subset of alpha carbon atoms 3414 that includes the alpha carbon atom associated with the first voxel. In another example, in steps 3202 and 3212, a second voxel is mapped to a second subset of alpha carbon atoms 3416 that includes the alpha carbon atom associated with the second voxel.
[0179] In step 3232, for each voxel, the distance between the voxel and the atoms mapped to the voxel in step 3222 is calculated. Step 3232 has a runtime complexity of O(number of atoms) because the distance to a particular atom is measured only once from each voxel to which a particular atom is uniquely mapped in voxel-to-atom mapping 3412. This is true if neighboring voxels are not considered. Without neighbors, the constant factor implied in big-O notation is 1. For neighbors, because the number of neighbors is constant for each voxel, big-O notation equals the number of neighbors + 1, and therefore the runtime complexity remains O(number of atoms). In contrast, in Figure 35A, the distance to a particular atom is measured redundantly for the number of voxels (e.g., 27 distances for a particular atom due to 27 voxels).
[0180] In Figure 35B, based on the voxel-to-atom mapping 3412, each voxel is mapped to a respective subset of 828 atoms (not including distance calculations to neighboring voxels), as indicated by the respective ellipses for each voxel. The respective subsets are largely non-overlapping, with a few exceptions. A small amount of overlap exists due to the few instances where multiple atoms map to the same voxel, as indicated by the prime symbol "'" and the yellow overlap between the ellipses in Figure 35B. This minimal overlap has an additive, not multiplicative, effect on the runtime complexity of O (number of atoms). This overlap is the result of considering neighboring voxels after determining the voxel containing the atom. Without neighboring voxels, there would be no overlap because an atom is associated with only one voxel. However, when neighbors are considered, each neighbor can potentially be associated with the same atom (unless there are other atoms of the same amino acid that are closer).
[0181] In step 3242, for each voxel, the nearest atom to the voxel is identified based on the distance calculated in step 3232. In one implementation, this identification is encoded in a voxel-to-nearest atom mapping 3422, also referred to herein as a "cell-to-nearest atom mapping." In one example, a first voxel is mapped to the second alpha carbon atom as its nearest alpha carbon atom 3424. In another example, a second voxel is mapped to the 31st alpha carbon atom as its nearest alpha carbon atom 3426.
[0182] Furthermore, as the voxel-wise distances are calculated using the techniques described above, the atom type and amino acid type classifications of the atoms and the corresponding distance values are stored to generate a classified distance channel.
[0183] Once the distances to the nearest neighbor atoms have been identified using the techniques described above, these distances can be encoded in a distance channel for voxelization and subsequent processing by the pathogenicity classifier 2108.
[0184] Computer Systems 36 illustrates an exemplary computer system 3600 that may be used to implement the disclosed techniques. The computer system 3600 includes at least one central processing unit (CPU) 3672 that communicates with a number of peripheral devices via a bus subsystem 3655. These peripheral devices may include, for example, a storage subsystem 3610 including memory devices and a file storage subsystem 3636, a user interface input device 3638, a user interface output device 3676, and a network interface subsystem 3674. The input and output devices enable user interaction with the computer system 3600. The network interface subsystem 3674 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0185] In one embodiment, the pathogenicity classifier 2108 is communicatively linked to the memory subsystem 3610 and the user interface input device 3638 .
[0186] The user interface input devices 3638 can include pointing devices such as a keyboard, a mouse, a trackball, a touchpad, or a graphics tablet, a scanner, a touchscreen integrated into a display, audio input devices such as a voice recognition system and a microphone, and other types of input devices. In general, use of the term "input device" is intended to encompass all possible types of devices and ways of inputting information into the computer system 3600.
[0187] The user interface output devices 3676 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a cathode ray tube (CRT), a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display such as an audio output device. In general, use of the term "output device" is intended to encompass all possible types of devices and ways for outputting information from the computer system 3600 to a user or to another machine or computer system.
[0188] The storage subsystem 3610 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 3678.
[0189] The processor 3678 may be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 3678 may be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 3678 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX36 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.
[0190] The memory subsystem 3622 used in memory subsystem 3610 may include multiple memories, including a main random access memory (RAM) 3632 for storing instructions and data during program execution, and a read only memory (ROM) 3634 in which fixed instructions are stored. The file storage subsystem 3636 may provide persistent storage for program and data files and may include a hard disk drive, associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular embodiment may be stored by the file storage subsystem 3636 in the storage subsystem 3610 or in another machine accessible by the processor.
[0191] Bus subsystem 3655 provides a mechanism for allowing the various components and subsystems of computer system 3600 to communicate with each other as intended. Although bus subsystem 3655 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0192] The computer system 3600 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 3600 shown in Figure 36 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 3600 can have more or fewer components than the computer system shown in Figure 36.
[0193] Specific Embodiment 1 The following embodiments can be implemented as a system, method, or product. One or more features of an embodiment can be combined with a base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of a repeat listing of these options should not be construed as limiting the combinations taught in the preceding sections. These descriptions are incorporated herein by reference into each of the following embodiments.
[0194] The disclosed techniques use 3D data as input, although in other embodiments, 1D data, 2D data (e.g., pixels and 2D atomic coordinates), 4D data, 5D data, etc. may be used as well.
[0195] In some embodiments, the system includes a memory that stores amino acid-by-amino acid distance channels for a plurality of amino acids in the protein. Each of the amino acid-by-amino acid distance channels has a voxel-by-voxel distance value for a voxel within the plurality of voxels. The voxel-by-voxel distance value specifies a distance from a corresponding voxel within the plurality of voxels to an atom of a corresponding amino acid within the plurality of amino acids. The system further includes a pathogenicity determination engine configured to process a tensor including the amino acid-by-amino acid distance channels and alternative alleles of the protein expressed by the variant. The pathogenicity determination engine can also be configured to determine the pathogenicity of the variant based at least in part on the tensor.
[0196] In some embodiments, the system further comprises a distance channel generator that centers a voxel grid of voxels on the alpha carbon atom of each residue of an amino acid. The distance channel generator can center the voxel grid on the alpha carbon atom of a residue of a particular amino acid that is located at a variant amino acid in the protein.
[0197] The system may be configured to encode the directionality of an amino acid and the position of a particular amino acid in a tensor by multiplying the voxel-wise distance value of the amino acid preceding the particular amino acid by a directionality parameter. The distance may be the nearest atom distance from the corresponding voxel center in the voxel grid to the nearest atom of the corresponding amino acid. In some embodiments, the nearest atom distance may be a Euclidean distance. The nearest atom distance may be normalized by dividing the Euclidean distance by the maximum nearest atom distance. The amino acid may have an alpha carbon atom, and in some embodiments, the distance may be the nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding amino acid. The amino acid may have a beta carbon atom, and in some embodiments, the distance may be the nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding amino acid. The amino acid may have backbone atoms, and in some embodiments, the distance may be the nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding amino acid. The amino acids have side chain atoms, and in some embodiments, the distance can be the nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding amino acid.
[0198] The system may further be configured to encode in the tensor a nearest-neighbor atom channel specifying the distance from each voxel to the nearest atom. The nearest-neighbor atom may be selected regardless of the amino acid and atomic component of the amino acid. In some embodiments, the distance is a Euclidean distance. The distance may be normalized by dividing the Euclidean distance by the maximum distance. The amino acid may include a non-standard amino acid. The tensor may further include an absent atom channel specifying atoms not found within a predetermined radius of the voxel center, and the absent atom channel may be one-hot encoded. In some embodiments, the tensor may further include one-hot encoding of alternative alleles encoded on a voxel-by-voxel basis in each of the amino acid-based distance channels. The tensor may further include a reference allele for the protein. In some embodiments, the tensor may further include one-hot encoding of reference alleles encoded on a voxel-by-voxel basis in each of the amino acid-based distance channels. The tensor may further include an evolutionary profile specifying the conservation level of the amino acid across multiple species.
[0199] The system may further include an evolutionary profile generator that, for each voxel, selects nearest atoms across amino acids and atom categories, selects a global amino acid conservation frequency array for amino acid residues containing the nearest atoms, and makes the global amino acid conservation frequency array available as one of the evolutionary profiles. The global amino acid conservation frequency array may be constructed for specific positions of residues observed in multiple species. The global amino acid conservation frequency array may identify whether there is a missing conservation frequency for a specific amino acid. In some embodiments, the evolutionary profile generator may select, for each voxel, each nearest atom in each of the amino acids, select a per-amino acid conservation frequency for each residue of the amino acids containing the nearest atoms, and make the per-amino acid conservation frequency available as one of the evolutionary profiles. The per-amino acid conservation frequency may be constructed for specific positions of residues observed in multiple species. The per-amino acid conservation frequency may identify whether there is a missing conservation frequency for a specific amino acid.
[0200] In some embodiments of the system, the tensor can further include an amino acid annotation channel. The annotation channel can be one-hot encoded in the tensor. The annotation channel can be a molecular process annotation including initiator methionine, signal, transit peptide, propeptide, chain, and peptide. The annotation channel can be a region annotation including topological domain, transmembrane, intramembrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled-coil, motif, and compositional bias. The annotation channel can be a site annotation including active site, metal binding, binding site, and site. The annotation channel can be an amino acid modification annotation including non-canonical residues, modified residues, lipidation, glycosylation, disulfide bonds, and crosslinks. The annotation channel can be a secondary structure annotation including helices, turns, and beta strands. The annotation channel can be an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residues, and non-terminal residues.
[0201] In some embodiments of the system, the tensor further includes an amino acid structure confidence channel that specifies a quality of the structure of each of the amino acids. The structure confidence channel can be a global model quality estimate (GMQE). The structure confidence channel can include a qualitative model energy analysis (QMEAN) score.
[0202] The structure confidence channel can be a temperature factor that specifies the degree to which a residue satisfies the physical constraints of the respective protein structure. The structure confidence channel can be a template structure alignment that specifies the degree to which a residue of a nearest atom to a voxel has an aligned template structure. The structure confidence channel can be a template modeling score of the aligned template structures. The structure confidence channel can be the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0203] In some embodiments, the system may further include a tensor generator that concatenates the amino acid unit distance channel for the alpha carbon atom with the one-hot encoding of alternative alleles on a voxel-by-voxel basis to generate a tensor. The tensor generator may concatenate the amino acid unit distance channel for the beta carbon atom with the one-hot encoding of alternative alleles on a voxel-by-voxel basis to generate a tensor. The tensor generator may concatenate the amino acid unit distance channel for the alpha carbon atom, the amino acid unit distance channel for the beta carbon atom, and the one-hot encoding of alternative alleles on a voxel-by-voxel basis to generate a tensor. The tensor generator may concatenate the amino acid unit distance channel for the alpha carbon atom, the amino acid unit distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, and the global amino acid conservation frequency on a voxel-by-voxel basis to generate a tensor. The tensor generator may concatenate the amino acid unit distance channel for the alpha carbon atom, the amino acid unit distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the global amino acid conservation frequency, and the annotation channel on a voxel-by-voxel basis to generate a tensor. The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-per-distance channel for the alpha carbon atom, the amino acid-per-distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the general amino acid conservation frequency, the annotation channel, and the structural confidence channel. The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-per-distance channel for the alpha carbon atom, the amino acid-per-distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, and the per-amino acid conservation frequency for each amino acid. The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-per-distance channel for the alpha carbon atom, the amino acid-per-distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the per-amino acid conservation frequency for each of the amino acids, and the annotation channel.The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-per-distance channel for the alpha carbon atom, the amino acid-per-distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the per-amino acid conservation frequency for each of the amino acids, the annotation channel, and the structural confidence channel. The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-per-distance channel for the alpha carbon atom, the amino acid-per-distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, and the one-hot encoding of the reference allele. The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-per-distance channel for the alpha carbon atom, the amino acid-per-distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of the reference allele, and the general amino acid conservation frequency. The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-per-distance channel for the alpha carbon atom, the amino acid-per-distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of the reference allele, the general amino acid conservation frequency, and the annotation channel. The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of the reference allele, the general amino acid conservation frequency, the annotation channel, and the structural confidence channel.The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of the reference allele, and the per-amino acid conservation frequency for each of the amino acids.The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of the reference allele, the per-amino acid conservation frequency for each of the amino acids, and the annotation channel.The tensor generator can generate a tensor by voxel-wise concatenation of the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of the reference allele, the per-amino acid conservation frequency for each of the amino acids, the annotation channel, and the structure confidence channel.
[0204] In some embodiments, the system may further include an atomic rotation engine that rotates amino acid atoms before generating the distance channels for each amino acid. The pathogenicity determination engine may be a neural network. In certain embodiments, the pathogenicity determination engine may be a convolutional neural network. The convolutional neural network may use 1x1x1 convolution, 3x3x3 convolution, a regularized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer. The 1x1x1 convolution and the 3x3x3 convolution may be three-dimensional convolutions.
[0205] In some implementations, a 1x1x1 convolutional layer can process a tensor and generate an intermediate output that is a convolutional representation of the tensor. A sequence of 3x3x3 convolutional layers can process the intermediate output and generate a flattened output. A fully connected layer can process the flattened output and generate a denormalized output. A softmax classification layer can process the denormalized output and generate an exponentially normalized output that identifies the likelihood that the variant is pathogenic or benign. A sigmoid layer can process the denormalized output and generate a normalized output that identifies the likelihood that the variant is pathogenic. Voxels, atoms, and distances can have three-dimensional coordinates. Tensors can have at least three dimensions, intermediate outputs can have at least three dimensions, and flattened outputs can have one dimension.
[0206] In some embodiments, the pathogenicity determination engine is a recurrent neural network. In other embodiments, the pathogenicity determination engine is an attention-based neural network. In yet other embodiments, the pathogenicity determination engine is a gradient boosted tree. In yet other embodiments, the pathogenicity determination engine is a state vector machine.
[0207] In other embodiments, the system can include a memory that stores per-atom-category distance channels for amino acids in a protein. An amino acid can have atoms of multiple atom categories, and an atom category within the multiple atom categories can identify an atomic element of the amino acid. The per-atom-category distance channels can have per-voxel distance values for voxels within the multiple voxels. The per-voxel distance values can identify a distance from a corresponding voxel within the multiple voxels to an atom within a corresponding atom category within the multiple atom categories. The system can further include a pathogenicity determination engine configured to process the per-atom-category distance channels and a tensor including alternative alleles of a protein expressed by the variant, and to determine a pathogenicity of the variant based at least in part on the tensor.
[0208] The system may further include a distance channel generator that centers a voxel grid of voxels on each atom of each atom category within the plurality of atom categories. The distance channel generator may center the voxel grid on an alpha carbon atom of a residue of at least one variant amino acid in the protein. The distance may be a nearest atom distance from a corresponding voxel center in the voxel grid to a nearest atom in the corresponding atom category. The nearest atom distance may be a Euclidean distance. The nearest atom distance may be normalized by dividing the Euclidean distance by a maximum nearest atom distance. The distance may be a nearest atom distance from a corresponding voxel center in the voxel grid to a nearest atom, regardless of the amino acid and the atom category of the amino acid. The nearest atom distance may be a Euclidean distance. The nearest atom distance may be normalized by dividing the Euclidean distance by a maximum nearest atom distance.
[0209] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0210] Clause 1 1. storing distance channels in amino acid units for a plurality of amino acids in a protein; each of the amino acid-based distance channels having a voxel-based distance value for a voxel within the plurality of voxels; the per-voxel distance values identify distances from corresponding voxels in the plurality of voxels to atoms of corresponding amino acids in the plurality of amino acids; processing tensors containing distance channels in amino acid units and alternative alleles of proteins expressed by the variants; and determining pathogenicity of the variant based at least in part on the tensor.
[0211] 2. The computer-implemented method of claim 1, further comprising centering the voxel grid of voxels on the alpha carbon atom of each residue of the amino acid.
[0212] 3. The computer-implemented method of claim 2, further comprising centering the voxel grid on the alpha carbon atom of a residue of a particular amino acid that corresponds to at least one variant amino acid in the protein.
[0213] 4. The computer-implemented method of claim 3, further comprising encoding in the tensor the directionality of amino acids and the position of a particular amino acid by multiplying the directionality parameter by a voxel-wise distance value for the amino acid preceding the particular amino acid.
[0214] 5. The computer-implemented method of claim 3, wherein the distance is a nearest atom distance from a corresponding voxel center in a voxel grid to a nearest atom of the corresponding amino acid.
[0215] 6. The computer-implemented method of claim 5, wherein the nearest neighbor atomic distance is a Euclidean distance.
[0216] 7. The computer-implemented method of claim 6, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0217] 8. The computer-implemented method of claim 5, wherein the amino acids have an alpha carbon atom and the distance is the nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding amino acid.
[0218] 9. The computer-implemented method of claim 5, wherein the amino acids have beta carbon atoms and the distance is the nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding amino acid.
[0219] 10. The computer-implemented method of claim 5, wherein the amino acids have backbone atoms and the distance is the nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding amino acid.
[0220] 11. The computer-implemented method of claim 5, wherein the amino acids have side chain atoms and the distance is the nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding amino acid.
[0221] 12. The computer-implemented method of claim 3, further comprising encoding in the tensor a nearest atom channel that specifies the distance from each voxel to the nearest atom, the nearest atom being selected without regard to amino acid and atomic component of the amino acid.
[0222] 13. The computer-implemented method of claim 12, wherein the distance is a Euclidean distance.
[0223] 14. The computer-implemented method of claim 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0224] 15. The computer-implemented method of claim 12, wherein the amino acids include non-standard amino acids.
[0225] 16. The computer-implemented method of claim 1, wherein the tensor further includes absent atom channels that specify atoms not found within a predetermined radius of the voxel center, and the absent atom channels are one-hot coded.
[0226] 17. The computer-implemented method of claim 1, wherein the tensor further comprises one-hot encoding of alternative alleles encoded voxel-wise into each of the amino acid-wise distance channels.
[0227] 18. The computer-implemented method of claim 1, wherein the tensor further comprises a reference allele of the protein.
[0228] 19. The computer-implemented method of claim 18, wherein the tensor further comprises one-hot encoding of the reference alleles encoded voxel-wise into each of the amino acid-wise distance channels.
[0229] 20. The computer-implemented method of claim 1, wherein the tensor further comprises an evolutionary profile specifying the level of conservation of amino acids across multiple species.
[0230] 21. For each voxel, selecting nearest neighbor atoms across amino acids and atom categories; selecting a general amino acid conservation frequency sequence for residues of amino acids containing nearest neighbor atoms; and making the pan-amino acid conservation frequency sequence available as one of the evolutionary profiles.
[0231] 22. The computer-implemented method of claim 21, wherein a global amino acid conservation frequency sequence is constructed for particular positions of residues as observed in multiple species.
[0232] 23. The computer-implemented method of claim 21, wherein the pan-amino acid conservation frequency sequence identifies whether there is a missing conservation frequency for a particular amino acid.
[0233] 24. For each voxel, selecting each nearest neighbor atom in each of the amino acids; selecting a conservation frequency for each amino acid for each residue of amino acids including nearest neighbor atoms; and making the conservation frequency for each amino acid available as one of the evolutionary profiles.
[0234] 25. The computer-implemented method of claim 24, wherein the conservation frequency for each amino acid is constructed for a particular position of the residue as observed in multiple species.
[0235] 26. The computer-implemented method of claim 24, wherein the conservation frequency for each amino acid identifies whether a missing conservation frequency exists for a particular amino acid.
[0236] 27. The computer-implemented method of claim 1, wherein the tensor further includes an annotation channel for amino acids, the annotation channel being one-hot coded in the tensor.
[0237] 28. The computer-implemented method of claim 27, wherein the annotation channel is a molecular processing annotation including initiator methionine, signal, transit peptide, propeptide, chain, and peptide.
[0238] 29. The computer-implemented method of claim 27, wherein the annotation channels are region annotations including topological domain, transmembrane, intramembrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
[0239] 30. The computer-implemented method of claim 27, wherein the annotation channels are site annotations including active sites, metal binding, binding sites, and sites.
[0240] 31. The computer-implemented method of claim 27, wherein the annotation channel is amino acid modification annotations including non-canonical residues, modified residues, lipidation, glycosylation, disulfide bonds, and cross-links.
[0241] 32. The computer-implemented method of claim 27, wherein the annotation channel is a secondary structure annotation including helices, turns, and beta strands.
[0242] 33. The computer-implemented method of claim 27, wherein the annotation channel is an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residues, and non-terminal residues.
[0243] 34. The computer-implemented method of claim 1, wherein the tensor further comprises an amino acid structural confidence channel that specifies the quality of each of the amino acid structures.
[0244] 35. The computer-implemented method of claim 34, wherein the structural confidence channel is a global model quality estimate (GMQE).
[0245] 36. The computer-implemented method of claim 34, wherein the structural confidence channel comprises a qualitative model energy analysis (QMEAN) score.
[0246] 37. The computer-implemented method of claim 34, wherein the structural confidence channel is a temperature factor that specifies the degree to which residues satisfy the physical constraints of the respective protein structure.
[0247] 38. The computer-implemented method of claim 34, wherein the structural confidence channel is a template structure alignment that identifies the degree to which residues of nearest neighbor atoms to a voxel have aligned template structures.
[0248] 39. The computer-implemented method of claim 38, wherein the structure confidence channel is a template modeling score for the aligned template structure.
[0249] 40. The computer-implemented method of claim 39, wherein the structural confidence channel is the smallest one of the template modeling scores, the average of the template modeling scores, and the largest one of the template modeling scores.
[0250] 41. The computer-implemented method of claim 1, further comprising voxel-wise concatenation of amino acid-wise distance channels for alpha carbon atoms with one-hot encoding of alternative alleles to generate tensors.
[0251] 42. The computer-implemented method of claim 41, further comprising voxel-wise concatenation of amino acid-wise distance channels for beta carbon atoms with one-hot encoding of alternative alleles to generate tensors.
[0252] 43. The computer-implemented method of claim 42, further comprising concatenating, on a voxel-by-voxel basis, the alpha carbon atom-based amino acid distance channel, the beta carbon atom-based amino acid distance channel, and the one-hot encoding of alternative alleles to generate a tensor.
[0253] 44. The computer-implemented method of claim 43, further comprising the step of concatenating the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, and the general amino acid conservation frequency sequence on a voxel-by-voxel basis to generate a tensor.
[0254] 45. The computer-implemented method of claim 44, further comprising concatenating the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the general amino acid conservation frequency sequence, and the annotation channel on a voxel-by-voxel basis to generate a tensor.
[0255] 46. The computer-implemented method of claim 45, further comprising the step of concatenating, on a voxel-by-voxel basis, the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the general amino acid conservation frequency sequence, the annotation channel, and the structural confidence channel to generate a tensor.
[0256] 47. The computer-implemented method of claim 46, further comprising concatenating the amino acid-wise distance channel for the alpha carbon atom, the amino acid-wise distance channel for the beta carbon atom, the amino acid-wise distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, and the amino acid-wise conservation frequency for each of the amino acids on a voxel-by-voxel basis to generate a tensor.
[0257] 48. The computer-implemented method of claim 47, further comprising concatenating, on a voxel-by-voxel basis, the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the amino acid-by-amino acid conservation frequency for each of the amino acids, and the annotation channel to generate a tensor.
[0258] 49. The computer-implemented method of claim 48, further comprising voxel-wise concatenation of a per-amino acid distance channel for the alpha carbon atom, a per-amino acid distance channel for the beta carbon atom, one-hot encoding of alternative alleles, per-amino acid conservation frequency for each of the amino acids, an annotation channel, and a structural confidence channel to generate a tensor.
[0259] 50. The computer-implemented method of claim 49, further comprising voxel-wise concatenation of the amino acid-wise distance channel for the alpha carbon atom, the amino acid-wise distance channel for the beta carbon atom, the one-hot encoding of the alternative allele, and the one-hot encoding of the reference allele to generate a tensor.
[0260] 51. The computer-implemented method of claim 50, further comprising concatenating the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of the alternative allele, the one-hot encoding of the reference allele, and the general amino acid conservation frequency sequence on a voxel-by-voxel basis to generate a tensor.
[0261] 52. The computer-implemented method of claim 51, further comprising concatenating the alpha carbon atom amino acid unit distance channel, the beta carbon atom amino acid unit distance channel, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the pan-amino acid conservation frequency sequence, and the annotation channel on a voxel-by-voxel basis to generate a tensor.
[0262] 53. The computer-implemented method of claim 52, further comprising the step of concatenating the alpha carbon atom amino acid unit distance channel, the beta carbon atom amino acid unit distance channel, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the pan-amino acid conservation frequency sequence, the annotation channel, and the structural confidence channel on a voxel-by-voxel basis to generate a tensor.
[0263] 54. The computer-implemented method of claim 53, further comprising concatenating, on a voxel-by-voxel basis, the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of the alternative allele, the one-hot encoding of the reference allele, and the amino acid-by-amino acid conservation frequency for each of the amino acids to generate a tensor.
[0264] 55. The computer-implemented method of claim 54, further comprising concatenating, on a voxel-by-voxel basis, the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the amino acid-by-amino acid conservation frequency for each of the amino acids, and the annotation channel to generate a tensor.
[0265] 56. The computer-implemented method of claim 55, further comprising concatenating, on a voxel-by-voxel basis, the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the amino acid-by-amino acid conservation frequency for each of the amino acids, the annotation channel, and the structural confidence channel to generate a tensor.
[0266] 57. The computer-implemented method of claim 1, further comprising rotating atoms of amino acids before distance channels per amino acid are generated.
[0267] 58. The computer-implemented method of claim 1, further comprising using a 1x1x1 convolution, a 3x3x3 convolution, a regularized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer in the convolutional neural network.
[0268] 59. The computer-implemented method of claim 58, wherein the 1x1x1 convolution and the 3x3x3 convolution are three-dimensional convolutions.
[0269] 60. The computer-implemented method of claim 58, wherein a 1x1x1 convolutional layer processes the tensors and produces intermediate outputs that are convolutional representations of the tensors, a sequence of 3x3x3 convolutional layers processes the intermediate outputs and produces flattened outputs, a fully connected layer processes the flattened outputs and produces unnormalized outputs, and a softmax classification layer processes the unnormalized outputs and produces exponentially normalized outputs that identify the likelihood that the variants are pathogenic and benign.
[0270] 61. The computer-implemented method of claim 60, wherein a sigmoid layer processes the unnormalized output and generates a normalized output that identifies the likelihood that the variant is pathogenic.
[0271] 62. The computer-implemented method of claim 60, wherein the voxels, atoms, and distances have three-dimensional coordinates, the tensors have at least three dimensions, the intermediate output has at least three dimensions, and the flattened output has one dimension.
[0272] 63. A step of storing distance channels per atom category of amino acids in a protein; Amino acids have atoms from multiple atom categories, an atom category of the plurality of atom categories identifying an atomic element of the amino acid; each of the atomic category-based distance channels has a voxel-based distance value for a voxel within the plurality of voxels; the per-voxel distance values identify distances from corresponding voxels in the plurality of voxels to atoms in corresponding atom categories in the plurality of atom categories; processing tensors containing distance channels per atomic category and alternative alleles of proteins expressed by the variants; and determining pathogenicity of the variant based at least in part on the tensor.
[0273] 64. The computer-implemented method of claim 63, further comprising centering the voxel grid of voxels on each atom of each atom category within the plurality of atom categories.
[0274] 65. The computer-implemented method of claim 64, further comprising centering the voxel grid on the alpha carbon atom of a residue of at least one variant amino acid in the protein.
[0275] 66. The computer-implemented method of claim 65, wherein the distance is a nearest atom distance from a corresponding voxel center in a voxel grid to a nearest atom in a corresponding atom category.
[0276] 67. The computer-implemented method of claim 66, wherein the nearest neighbor atomic distance is a Euclidean distance.
[0277] 68. The computer-implemented method of claim 67, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0278] 69. The computer-implemented method of claim 68, wherein the distance is the nearest atom distance from the corresponding voxel center to the nearest neighbor atom in the voxel grid, regardless of the amino acid and atom category of the amino acid.
[0279] 70. The computer-implemented method of claim 69, wherein the nearest-neighbor atom distance is a Euclidean distance.
[0280] 71. The computer-implemented method of claim 70, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0281] 1. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, storing distance channels in amino acid units for a plurality of amino acids in the protein; each of the amino acid-based distance channels having a voxel-based distance value for a voxel within the plurality of voxels; processing a tensor including per-voxel distance channels, wherein the per-voxel distance values identify distances from corresponding voxels within the plurality of voxels to atoms of corresponding amino acids within the plurality of amino acids, and alternative alleles of the protein expressed by the variants; and determining pathogenicity of the variant based at least in part on the tensor.
[0282] 2. The computer-readable medium of claim 1, wherein the operation further comprises centering the voxel grid of voxels on the alpha carbon atom of each residue of the amino acid.
[0283] 3. The computer-readable medium of claim 2, wherein the operation further comprises centering the voxel grid on the alpha carbon atom of a residue of a particular amino acid that corresponds to at least one variant amino acid in the protein.
[0284] 4. The computer-readable medium of claim 3, wherein the operations further include encoding in the tensor the directionality of amino acids and the position of a particular amino acid by multiplying the distance value in voxels for the amino acid preceding the particular amino acid with the directionality parameter.
[0285] 5. The computer-readable medium of claim 3, wherein the distance is a nearest atom distance from a corresponding voxel center in a voxel grid to a nearest atom of the corresponding amino acid.
[0286] 6. The computer-readable medium of claim 5, wherein the nearest neighbor atomic distance is a Euclidean distance.
[0287] 7. The computer-readable medium of claim 6, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0288] 8. The computer-readable medium of claim 5, wherein the amino acids have an alpha carbon atom and the distance is a nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding amino acid.
[0289] 9. The computer-readable medium of claim 5, wherein the amino acids have a beta carbon atom and the distance is a nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding amino acid.
[0290] 10. The computer-readable medium of claim 5, wherein the amino acids have backbone atoms and the distance is a nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding amino acid.
[0291] 11. The computer-readable medium of claim 5, wherein the amino acids have side chain atoms and the distance is a nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding amino acid.
[0292] 12. The computer-readable medium of claim 3, wherein the operations further include encoding, in the tensor, a nearest atom channel that specifies a distance from each voxel to a nearest neighbor atom, the nearest neighbor atom being selected without regard to amino acid and atomic component of the amino acid.
[0293] 13. The computer-readable medium of claim 12, wherein the distance is a Euclidean distance.
[0294] 14. The computer-readable medium of claim 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0295] 15. The computer-readable medium of claim 12, wherein the amino acids include non-standard amino acids.
[0296] 16. The computer-readable medium of claim 1, wherein the tensor further includes an absent atom channel that specifies atoms not found within a predetermined radius of the voxel center, the absent atom channel being one-hot coded.
[0297] 17. The computer-readable medium of claim 1, wherein the tensor further comprises one-hot encoding of alternative alleles encoded voxel-wise into each of the amino acid-wise distance channels.
[0298] 18. The computer-readable medium of claim 1, wherein the tensor further comprises a reference allele of the protein.
[0299] 19. The computer-readable medium of claim 18, wherein the tensor further comprises one-hot encoding of reference alleles encoded voxel-wise into each of the amino acid-wise distance channels.
[0300] 20. The computer-readable medium of claim 1, wherein the tensor further comprises an evolutionary profile identifying the level of conservation of amino acids across multiple species.
[0301] 21. The operation is, for each voxel: selecting nearest neighbor atoms across amino acids and atom categories; selecting a general amino acid conservation frequency sequence for residues of amino acids containing nearest neighbor atoms; and making the pan-amino acid conservation frequency sequence available as one of the evolutionary profiles.
[0302] 22. The computer-readable medium of claim 21, wherein a global amino acid conservation frequency sequence is constructed for particular positions of residues as observed in multiple species.
[0303] 23. The computer-readable medium of claim 21, wherein the pan-amino acid conservation frequency sequence identifies whether there is a missing conservation frequency for a particular amino acid.
[0304] 24. The operation is, for each voxel: selecting each nearest neighbor atom in each of the amino acids; selecting a conservation frequency for each amino acid for each residue of amino acids including nearest neighbor atoms; and making the conservation frequency for each amino acid available as one of the evolutionary profiles.
[0305] 25. The computer-readable medium of claim 24, wherein the conservation frequency for each amino acid is constructed for a particular position of the residue as observed in multiple species.
[0306] 26. The computer-readable medium of claim 24, wherein the conservation frequency for each amino acid identifies whether a missing conservation frequency exists for a particular amino acid.
[0307] 27. The computer-readable medium of claim 1, wherein the tensor further comprises an annotation channel for amino acids, the annotation channel being one-hot encoded in the tensor.
[0308] 28. The computer-readable medium of claim 27, wherein the annotation channel is a molecular processing annotation including initiator methionine, signal, transit peptide, propeptide, chain, and peptide.
[0309] 29. The computer-readable medium of claim 27, wherein the annotation channels are region annotations including topological domain, transmembrane, intramembrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
[0310] 30. The computer-readable medium of claim 27, wherein the annotation channels are site annotations including active sites, metal binding, binding sites, and sites.
[0311] 31. The computer-readable medium of claim 27, wherein the annotation channel is an amino acid modification annotation including non-canonical residues, modified residues, lipidation, glycosylation, disulfide bonds, and cross-links.
[0312] 32. The computer-readable medium of claim 27, wherein the annotation channel is a secondary structure annotation including helices, turns, and beta strands.
[0313] 33. The computer-readable medium of claim 27, wherein the annotation channel is an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflicts, non-adjacent residues, and non-terminal residues.
[0314] 34. The computer-readable medium of claim 1, wherein the tensor further comprises an amino acid structural confidence channel that specifies the quality of the structure of each of the amino acids.
[0315] 35. The computer-readable medium of claim 34, wherein the structural confidence channel is a global model quality estimation (GMQE).
[0316] 36. The computer-readable medium of claim 34, wherein the structural confidence channel comprises a qualitative model energy analysis (QMEAN) score.
[0317] 37. The computer-readable medium of claim 34, wherein the structural confidence channel is a temperature factor that specifies the degree to which residues satisfy the physical constraints of the respective protein structure.
[0318] 38. The computer-readable medium of claim 34, wherein the structural confidence channel is a template structure alignment that identifies the degree to which residues of nearest neighbor atoms to a voxel have aligned template structures.
[0319] 39. The computer-readable medium of claim 38, wherein the structure confidence channel is a template modeling score for the aligned template structure.
[0320] 40. The computer-readable medium of claim 39, wherein the structural confidence channel is the smallest one of the template modeling scores, the average of the template modeling scores, and the largest one of the template modeling scores.
[0321] 41. The computer-readable medium of claim 1, wherein the operations further comprise concatenating amino acid-wise distance channels for alpha carbon atoms with one-hot encoding of alternative alleles voxel-wise to generate a tensor.
[0322] 42. The computer-readable medium of claim 41, wherein the operations further comprise concatenating amino acid-wise distance channels for beta carbon atoms with one-hot encoding of alternative alleles voxel-wise to generate a tensor.
[0323] 43. The computer-readable medium of claim 42, wherein the operations further comprise voxel-wise concatenation of the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, and the one-hot encoding of alternative alleles to generate a tensor.
[0324] 44. The computer-readable medium of claim 43, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a distance channel of amino acids per alpha carbon atom, a distance channel of amino acids per beta carbon atom, one-hot encoding of alternative alleles, and a pan-amino acid conservation frequency sequence to generate a tensor.
[0325] 45. The computer-readable medium of claim 44, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a distance channel of amino acids per alpha carbon atom, a distance channel of amino acids per beta carbon atom, one-hot encoding of alternative alleles, a pan-amino acid conservation frequency sequence, and an annotation channel to generate a tensor.
[0326] 46. The computer-readable medium of claim 45, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a distance channel of amino acids per alpha carbon atom, a distance channel of amino acids per beta carbon atom, one-hot encoding of alternative alleles, a general amino acid conservation frequency sequence, an annotation channel, and a structural confidence channel to generate a tensor.
[0327] 47. The computer-readable medium of claim 46, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, the per-amino acid distance channel for the alpha carbon atom, the per-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, and the per-amino acid conservation frequency for each of the amino acids to generate a tensor.
[0328] 48. The computer-readable medium of claim 47, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a per-amino acid distance channel for the alpha carbon atom, a per-amino acid distance channel for the beta carbon atom, one-hot encoding of alternative alleles, per-amino acid conservation frequency for each of the amino acids, and an annotation channel to generate a tensor.
[0329] 49. The computer-readable medium of claim 48, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a per-amino acid distance channel for the alpha carbon atom, a per-amino acid distance channel for the beta carbon atom, a one-hot encoding of alternative alleles, a per-amino acid conservation frequency for each of the amino acids, an annotation channel, and a structural confidence channel to generate a tensor.
[0330] 50. The computer-readable medium of claim 49, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a distance channel per amino acid for the alpha carbon atom, a distance channel per amino acid for the beta carbon atom, a one-hot encoding of the alternative allele, and a one-hot encoding of the reference allele to generate a tensor.
[0331] 51. The computer-readable medium of claim 50, wherein the operations further comprise concatenating the alpha carbon atom amino acid unit distance channel, the beta carbon atom amino acid unit distance channel, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, and the pan-amino acid conservation frequency sequence on a voxel-by-voxel basis to generate a tensor.
[0332] 52. The computer-readable medium of claim 51, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, the amino acid-by-amino acid distance channel for the alpha carbon atom, the amino acid-by-amino acid distance channel for the beta carbon atom, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the pan-amino acid conservation frequency sequence, and the annotation channel to generate a tensor.
[0333] 53. The computer-readable medium of claim 52, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a distance channel of amino acids per alpha carbon atom, a distance channel of amino acids per beta carbon atom, one-hot encoding of alternative alleles, one-hot encoding of reference alleles, a pan-amino acid conservation frequency sequence, an annotation channel, and a structural confidence channel to generate a tensor.
[0334] 54. The computer-readable medium of claim 53, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, the per-amino acid distance channel for the alpha carbon atom, the per-amino acid distance channel for the beta carbon atom, the one-hot encoding of the alternative allele, the one-hot encoding of the reference allele, and the per-amino acid conservation frequency for each of the amino acids to generate a tensor.
[0335] 55. The computer-readable medium of claim 54, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a distance channel per amino acid for the alpha carbon atom, a distance channel per amino acid for the beta carbon atom, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a per-amino acid conservation frequency for each of the amino acids, and an annotation channel to generate a tensor.
[0336] 56. The computer-readable medium of claim 55, wherein the operations further comprise concatenating, on a voxel-by-voxel basis, a distance channel per amino acid for the alpha carbon atom, a distance channel per amino acid for the beta carbon atom, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a per-amino acid conservation frequency for each of the amino acids, an annotation channel, and a structural confidence channel to generate a tensor.
[0337] 57. The computer-readable medium of claim 1, wherein the operation further comprises rotating atoms of the amino acid before distance channels for the amino acid units are generated.
[0338] 58. The computer-readable medium of claim 1, wherein the operations further include using a 1x1x1 convolution, a 3x3x3 convolution, a regularized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer in the convolutional neural network.
[0339] 59. The computer-readable medium of claim 58, wherein the 1x1x1 convolution and the 3x3x3 convolution are three-dimensional convolutions.
[0340] 60. The computer-readable medium of claim 58, wherein a layer of 1x1x1 convolution processes the tensor and produces intermediate outputs that are convolutional representations of the tensor, a sequence of layers of 3x3x3 convolution processes the intermediate outputs and produces flattened outputs, a fully connected layer processes the flattened outputs and produces unnormalized outputs, and a softmax classification layer processes the unnormalized outputs and produces exponentially normalized outputs that identify the likelihood that the variant is pathogenic and benign.
[0341] 61. The computer-readable medium of claim 60, wherein a sigmoid layer processes the unnormalized output and generates a normalized output that identifies the likelihood that the variant is pathogenic.
[0342] 62. The computer-readable medium of claim 60, wherein the voxels, atoms, and distances have three-dimensional coordinates, the tensors have at least three dimensions, the intermediate output has at least three dimensions, and the flattened output has one dimension.
[0343] 63. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, Storing distance channels for each atom category of amino acids in a protein; Amino acids have atoms from multiple atom categories, an atom category of the plurality of atom categories identifying an atomic element of the amino acid; each of the atomic category-based distance channels has a voxel-based distance value for a voxel within the plurality of voxels; the per-voxel distance values identify distances from corresponding voxels in the plurality of voxels to atoms in corresponding atom categories in the plurality of atom categories; processing tensors containing distance channels per atomic category and alternative alleles of proteins expressed by the variants; and determining pathogenicity of the variant based at least in part on the tensor.
[0344] 64. The computer-readable medium of claim 63, wherein the operations further comprise centering the voxel grid of voxels on each atom of each atom category within the plurality of atom categories.
[0345] 65. The computer-readable medium of claim 64, wherein the operation further comprises centering the voxel grid on an alpha carbon atom of a residue of at least one variant amino acid in the protein.
[0346] 66. The computer-readable medium of claim 65, wherein the distance is a nearest atom distance from a corresponding voxel center in a voxel grid to a nearest atom in a corresponding atom category.
[0347] 67. The computer-readable medium of claim 66, wherein the nearest neighbor atomic distance is a Euclidean distance.
[0348] 68. The computer-readable medium of claim 67, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0349] 69. The computer-readable medium of claim 68, wherein the distance is the nearest atom distance from the corresponding voxel center to the nearest neighbor atom in the voxel grid, regardless of the amino acid and atom category of the amino acid.
[0350] 70. The computer-readable medium of claim 69, wherein the nearest neighbor atomic distance is a Euclidean distance.
[0351] 71. The computer-readable medium of claim 70, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0352] Specific Embodiment 2 In some embodiments, the system includes a voxelizer that accesses a three-dimensional structure of a reference amino acid sequence of the protein and fits a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid basis to generate amino acid-based distance channels. Each amino acid-based distance channel has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels. The three-dimensional distance value specifies the distance from a corresponding voxel in the three-dimensional grid of voxels to an atom of a corresponding reference amino acid in the reference amino acid sequence. The system further includes an alternative allele encoder that encodes an alternative allelic amino acid for each voxel in the three-dimensional grid of voxels. The alternative allelic amino acid is a one-hot encoded three-dimensional representation of a variant amino acid expressed by a variant nucleotide. The system further includes an evolutionary conservation encoder that encodes an evolutionarily conserved sequence for each voxel in the three-dimensional grid of voxels. The evolutionarily conserved sequence can be a three-dimensional representation of amino acid-specific conservation frequencies across multiple species. The amino acid-specific conservation frequencies can be selected according to amino acid proximity to the corresponding voxel. The system further includes a convolutional neural network configured to apply a three-dimensional convolution to a tensor including distance channels of amino acid units encoded by alternative allelic amino acids and their respective evolutionarily conserved sequences. The convolutional neural network can also be configured to determine pathogenicity of the variant nucleotide based at least in part on the tensor.
[0353] The voxelizer can center the three-dimensional grid of voxels on the alpha carbon atom of each reference amino acid residue in the reference amino acid sequence. The voxelizer can center the three-dimensional grid of voxels on the alpha carbon atom of a particular reference amino acid residue that is located at the variant amino acid residue.
[0354] In some embodiments, the system may be further configured to encode the orientation of the reference amino acid and the position of the particular reference amino acid within the reference amino acid sequence by multiplying the three-dimensional distance value for the reference amino acid preceding the particular reference amino acid by a directionality parameter in the tensor. The distance may be a nearest atom distance from a corresponding voxel center in a three-dimensional grid of voxels to the nearest atom of the corresponding reference amino acid. The nearest atom distance may be a Euclidean distance, which may be normalized by dividing the Euclidean distance by the maximum nearest atom distance.
[0355] In some embodiments, the reference amino acids can have alpha carbon atoms, and the distance can be the nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding reference amino acid. In some embodiments, the reference amino acids can have beta carbon atoms, and the distance can be the nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding reference amino acid. In some embodiments, the reference amino acids can have backbone atoms, and the distance can be the nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding reference amino acid. In some embodiments, the amino acids can have side chain atoms, and the distance can be the nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding reference amino acid.
[0356] In some embodiments, the system may be further configured to encode in the tensor a nearest-neighbor atom channel that specifies the distance from each voxel to the nearest atom. The nearest atom may be selected regardless of the amino acid and atomic component of the amino acid. The distance may be a Euclidean distance and may be normalized by dividing the Euclidean distance by the maximum distance. The amino acid may include a non-standard amino acid. The tensor may further include an absent atom channel that specifies atoms that are not found within a predetermined radius of the voxel center. The absent atom channel may be one-hot encoded.
[0357] In some embodiments, the system can further comprise a reference allele encoder that encodes a reference allele amino acid by amino acid position on a voxel-by-voxel basis for each three-dimensional distance value. The reference allele amino acid can be a one-hot encoded three-dimensional representation of the reference amino acid sequence. The amino acid-specific conservation frequency can identify the conservation level of each amino acid across multiple species.
[0358] In some embodiments, the evolutionary conservation encoder can select the nearest atom to the corresponding voxel across the reference amino acid and atom category, and can select a global amino acid conservation frequency for the residue of the reference amino acid that includes the nearest atom, and can use a three-dimensional representation of the global amino acid conservation frequency as the evolutionary conservation sequence. The global amino acid conservation frequency can be configured for a specific position of the residue as observed in multiple species. The global amino acid conservation frequency can identify whether there is a missing conservation frequency for a specific reference amino acid.
[0359] In some embodiments, the evolutionary conservation encoder can select the nearest atom to each corresponding voxel in each reference amino acid, and select the respective per-amino acid conservation frequency for each residue of the reference amino acid containing the nearest atom, and use the three-dimensional representation of the per-amino acid conservation frequency as the evolutionary conservation sequence. The per-amino acid conservation frequency can be configured for a specific position of the residue as observed in multiple species. The per-amino acid conservation frequency can identify whether there is a missing conservation frequency for a specific reference amino acid.
[0360] In some embodiments, the system may further comprise an annotation encoder that encodes one or more annotation channels into each three-dimensional distance value on a voxel-by-voxel basis. An annotation channel may be a one-hot encoded three-dimensional representation of residue annotations, and may be molecular process annotations including initiator methionine, signal, transit peptide, propeptide, chain, and peptide. In some embodiments, an annotation channel may be region annotations including topological domain, transmembrane, intramembrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled-coil, motif, and compositional bias, or site annotations including active site, metal binding, binding site, and site. In some embodiments, an annotation channel may be amino acid modification annotations including non-canonical residues, modified residues, lipidation, glycosylation, disulfide bonds, and crosslinks, or secondary structure annotations including helices, turns, and beta-strands. The annotation channel can be experimental information annotation including mutagenesis, sequence uncertainty, sequence conflicts, non-adjacent residues, and non-terminal residues.
[0361] In some embodiments, the system may further comprise a structure confidence encoder that encodes one or more structure confidence channels into each three-dimensional distance value on a voxel-by-voxel basis. The structure confidence channels may be a three-dimensional representation of a confidence score that specifies the quality of each residue structure. The structure confidence channel may be a global model quality estimate (GMQE), a qualitative model energy analysis (QMEAN) score, a temperature factor that specifies the degree to which a residue satisfies the physical constraints of each protein structure, a template structure alignment that specifies the degree to which the residue of the nearest atom to the voxel has an aligned template structure, a template modeling score of the aligned template structure, or a minimum of the template modeling score, an average of the template modeling score, and a maximum of the template modeling score.
[0362] In some embodiments, the system may further comprise an atom rotation engine that rotates atoms before distance channels of amino acid units are generated.
[0363] The convolutional neural network may use 1x1x1 convolutions, 3x3x3 convolutions, normalized linear unit activation layers, batch normalization layers, fully connected layers, dropout regularization layers, and softmax classification layers. The 1x1x1 convolutions and 3x3x3 convolutions may be 3D convolutions. In some embodiments, a 1x1x1 convolutional layer may process a tensor and generate an intermediate output that is a convolutional representation of the tensor. A sequence of 3x3x3 convolutional layers may process the intermediate output and generate a flattened output. A fully connected layer may process the flattened output and generate a denormalized output. A softmax classification layer may process the denormalized output and generate an exponentially normalized output that identifies the likelihood that a variant nucleotide is pathogenic or benign.
[0364] In some embodiments, a sigmoid layer can process the unnormalized output and generate a normalized output that identifies the likelihood that the variant nucleotide is pathogenic. The convolutional neural network can be an attention-based neural network. The tensor can include a distance channel for amino acid units further encoded by a reference allele amino acid, can include a distance channel for amino acid units further encoded by an annotation channel, or can include a distance channel for amino acid units further encoded by a structure confidence channel.
[0365] In some embodiments, the system may include a voxelizer that accesses a three-dimensional structure of a reference amino acid sequence of the protein and fits a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid basis to generate per-atom-category distance channels. The atoms span multiple atom categories that identify atomic elements of the amino acids. Each per-atom-category distance channel has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels. The three-dimensional distance value specifies a distance from a corresponding voxel in the three-dimensional grid of voxels to an atom of the corresponding atom category in the multiple atom categories. The system further includes an alternative allele encoder that encodes an alternative allele amino acid for each voxel in the three-dimensional grid of voxels. The alternative allele amino acid is a three-dimensional representation of a one-hot encoding of a variant amino acid expressed by a variant nucleotide. The system further includes an evolutionary conservation encoder that encodes an evolutionarily conserved sequence for each voxel in the three-dimensional grid of voxels. The evolutionarily conserved sequence may be a three-dimensional representation of an amino acid-specific conservation frequency across multiple species. The amino acid-specific conservation frequencies can be selected according to amino acid proximity to the corresponding voxel. The system further includes a convolutional neural network configured to apply a three-dimensional convolution to a tensor including distance channels per atomic category encoded by alternative allelic amino acids and their respective evolutionarily conserved sequences, and determine pathogenicity of the variant nucleotide based at least in part on the tensor.
[0366] In some embodiments, the system includes a voxelizer that accesses a three-dimensional structure of a reference amino acid sequence of the protein and fits a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid basis to generate amino acid-based distance channels. Each amino acid-based distance channel can have a three-dimensional distance value for each voxel in the three-dimensional grid of voxels. The three-dimensional distance value can specify a distance from a corresponding voxel in the three-dimensional grid of voxels to an atom of a corresponding reference amino acid in the reference amino acid sequence. The system further includes an alternative allele encoder that encodes an alternative allelic amino acid for each voxel in the three-dimensional grid of voxels. The alternative allelic amino acid is a one-hot encoded three-dimensional representation of a variant amino acid expressed by a variant nucleotide. The system further includes an evolutionary conservation encoder that encodes an evolutionarily conserved sequence for each voxel in the three-dimensional grid of voxels. The evolutionarily conserved sequence can be a three-dimensional representation of amino acid-specific conservation frequencies across multiple species. The amino acid-specific conservation frequencies can be selected according to amino acid proximity to the corresponding voxel. The system further includes a tensor generator configured to generate a tensor including distance channels for amino acid units encoded by alternative allelic amino acids and their respective evolutionarily conserved sequences.
[0367] In some embodiments, the system includes a voxelizer that accesses a three-dimensional structure of a reference amino acid sequence of the protein and fits a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid basis to generate per-atom-category distance channels. The atoms can span multiple atom categories that identify atomic elements of the amino acids. Each per-atom-category distance channel can have a three-dimensional distance value for each voxel in the three-dimensional grid of voxels. The three-dimensional distance value can identify a distance from a corresponding voxel in the three-dimensional grid of voxels to an atom of the corresponding atom category in the multiple atom categories. The system further includes an alternative allele encoder that encodes an alternative allele amino acid for each voxel in the three-dimensional grid of voxels. The alternative allele amino acid is a one-hot encoded three-dimensional representation of a variant amino acid expressed by a variant nucleotide. The system further includes an evolutionary conservation encoder that encodes an evolutionarily conserved sequence for each voxel in the three-dimensional grid of voxels. The evolutionarily conserved sequence can be a three-dimensional representation of an amino acid-specific conservation frequency across multiple species. The amino acid-specific conservation frequencies can be selected according to amino acid proximity to the corresponding voxel. The system further includes a tensor generator configured to generate a tensor including distance channels per atomic category encoded with alternative allelic amino acids and their respective evolutionarily conserved sequences.
[0368] Clause 2 1. Accessing a three-dimensional structure of a reference amino acid sequence of a protein and fitting a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid-by-amino acid basis to generate distance channels on an amino acid-by-amino acid basis; each of the amino acid-based distance channels has a three-dimensional distance value for each voxel in a three-dimensional grid of voxels; the three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atom of the corresponding reference amino acid in the reference amino acid sequence; encoding an alternative allele channel into each voxel in a three-dimensional grid of voxels, the alternative allele channel being a one-hot coded three-dimensional representation of a variant amino acid expressed by a variant nucleotide; encoding an evolutionary conservation channel into each sequence of three-dimensional distance values across a distance channel in units of amino acids on a voxel location basis, the evolutionary conservation channel being a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, the amino acid-specific conservation frequencies being selected according to amino acid proximity to a corresponding voxel; applying a three-dimensional convolution to a tensor comprising distance channels of amino acid units encoded with alternative allele channels and respective evolutionarily conserved channels; and determining a pathogenicity of the variant nucleotide based at least in part on the tensor.
[0369] 2. The computer-implemented method of clause 1, further comprising centering the three-dimensional grid of voxels on the alpha carbon atom of each residue of the reference amino acid in the reference amino acid sequence.
[0370] 3. The computer-implemented method of clause 2, further comprising centering the three-dimensional grid of voxels on the alpha carbon atom of the residue of the particular reference amino acid that corresponds to the variant amino acid.
[0371] 4. The computer-implemented method of clause 3, further comprising encoding the directionality of the reference amino acid and the position of the particular reference amino acid in the reference amino acid sequence by multiplying the three-dimensional distance value for the reference amino acid preceding the particular reference amino acid by a directionality parameter in the tensor.
[0372] 5. The computer-implemented method of clause 4, wherein the distance is a nearest atom distance from a corresponding voxel center in a three-dimensional grid of voxels to a nearest atom of the corresponding reference amino acid.
[0373] 6. The computer-implemented method of clause 5, wherein the nearest neighbor atomic distance is a Euclidean distance.
[0374] 7. The computer-implemented method of clause 6, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0375] 8. The computer-implemented method of clause 5, wherein the reference amino acid has an alpha carbon atom and the distance is the nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding reference amino acid.
[0376] 9. The computer-implemented method of clause 5, wherein the reference amino acid has a beta carbon atom and the distance is the nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding reference amino acid.
[0377] 10. The computer-implemented method of clause 5, wherein the reference amino acid has backbone atoms and the distance is the nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding reference amino acid.
[0378] 11. The computer-implemented method of clause 5, wherein the amino acids have side chain atoms and the distance is the nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding reference amino acid.
[0379] 12. The computer-implemented method of clause 3, further comprising encoding in the tensor a nearest atom channel that specifies the distance from each voxel to the nearest atom, the nearest atom being selected without regard to amino acid and atomic component of the amino acid.
[0380] 13. The computer-implemented method of clause 12, wherein the distance is a Euclidean distance.
[0381] 14. The computer-implemented method of clause 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0382] 15. The computer-implemented method of clause 12, wherein the amino acids include non-standard amino acids.
[0383] 16. The computer-implemented method of clause 1, wherein the tensor further includes an absent atom channel that specifies atoms that are not found within a predetermined radius of the voxel center.
[0384] 17. The computer-implemented method of clause 16, wherein the absent atomic channels are one-hot coded.
[0385] 18. The computer-implemented method of clause 1, further comprising voxel-wise encoding a reference allele channel for each voxel in the three-dimensional grid of voxels.
[0386] 19. The computer-implemented method of clause 18, wherein the reference allele amino acid is a one-hot coded three-dimensional representation of the reference amino acid undergoing the variant amino acid.
[0387] 20. The computer-implemented method of clause 1, wherein the amino acid-specific conservation frequency identifies the level of conservation of each amino acid across multiple species.
[0388] 21. Selecting nearest neighbor atoms to corresponding voxels across reference amino acids and atom categories; selecting a pan-amino acid conservation frequency for residues of the reference amino acid that include the nearest atom; 21. The computer-implemented method of clause 20, further comprising using a three-dimensional representation of the general amino acid conservation frequencies as an evolutionary conservation channel.
[0389] 22. The computer-implemented method of clause 21, wherein a global amino acid conservation frequency is constructed for a particular position of the residue as observed in multiple species.
[0390] 23. The computer-implemented method of clause 21, wherein the pan-amino acid conservation frequency identifies whether there is a missing conservation frequency for a particular reference amino acid.
[0391] 24. Selecting the nearest atom to each corresponding voxel in each of the reference amino acids; selecting a conservation frequency for each amino acid for each residue of the reference amino acid that includes the nearest atom; and using the three-dimensional representation of the conservation frequency per amino acid as an evolutionary conservation channel.
[0392] 25. The computer-implemented method of clause 24, wherein a conservation frequency for each amino acid is constructed for a particular position of the residue as observed in multiple species.
[0393] 26. The computer-implemented method of clause 24, wherein the conservation frequency for each amino acid identifies whether there is a missing conservation frequency for a particular reference amino acid.
[0394] 27. The computer-implemented method of clause 1, further comprising voxel-wise encoding one or more annotation channels into each voxel in a three-dimensional grid of voxels, wherein the annotation channels are three-dimensional representations of one-hot encoding of the residual annotations.
[0395] 28. The computer-implemented method of clause 27, wherein the annotation channel is a molecular processing annotation including initiator methionine, signal, transit peptide, propeptide, chain, and peptide.
[0396] 29. The computer-implemented method of clause 27, wherein the annotation channels are region annotations including topological domain, transmembrane, intramembrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
[0397] 30. The computer-implemented method of clause 27, wherein the annotation channel is a site annotation including an active site, a metal binding site, a binding site, and a site.
[0398] 31. The computer-implemented method of clause 27, wherein the annotation channel is an amino acid modification annotation including non-canonical residues, modified residues, lipidation, glycosylation, disulfide bonds, and cross-links.
[0399] 32. The computer-implemented method of clause 27, wherein the annotation channel is a secondary structure annotation including helices, turns, and beta strands.
[0400] 33. The computer-implemented method of clause 27, wherein the annotation channel is an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residues, and non-terminal residues.
[0401] 34. The computer-implemented method of clause 1, further comprising voxel-wise encoding one or more structural confidence channels into each voxel in a three-dimensional grid of voxels, the structural confidence channels being three-dimensional representations of confidence scores identifying the quality of each residue structure.
[0402] 35. The computer-implemented method of clause 34, wherein the structural confidence channel is a global model quality estimate (GMQE).
[0403] 36. The computer-implemented method of clause 34, wherein the structural confidence channel is a qualitative model energy analysis (QMEAN) score.
[0404] 37. The computer-implemented method of clause 34, wherein the structural confidence channel is a temperature factor that specifies the degree to which residues satisfy the physical constraints of the respective protein structure.
[0405] 38. The computer-implemented method of clause 34, wherein the structure confidence channel is a template structure alignment that identifies the degree to which residues of nearest neighbor atoms to a voxel align the template structure.
[0406] 39. The computer-implemented method of clause 38, wherein the structure confidence channel is a template modeling score for the aligned template structure.
[0407] 40. The computer-implemented method of clause 39, wherein the structural confidence channel is the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0408] 41. The computer-implemented method of clause 1, further comprising rotating atoms before the distance channels for each amino acid are generated.
[0409] 42. The computer-implemented method of clause 1, further comprising using a 1x1x1 convolution, a 3x3x3 convolution, a regularized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer in the convolutional neural network.
[0410] 43. The computer-implemented method of clause 42, wherein 1x1x1 convolution and 3x3x3 convolution are three-dimensional convolutions.
[0411] 44. The computer-implemented method of clause 42, wherein a 1x1x1 convolutional layer processes the tensor and generates intermediate outputs that are convolutional representations of the tensor; a sequence of 3x3x3 convolutional layers processes the intermediate outputs and generates flattened outputs; a fully connected layer processes the flattened outputs and generates denormalized outputs; and a softmax classification layer processes the denormalized outputs and generates exponentially normalized outputs that identify the likelihood that the variant nucleotide is pathogenic and benign.
[0412] 45. The computer-implemented method of clause 44, wherein a sigmoid layer processes the unnormalized output and generates a normalized output that identifies the likelihood that the variant nucleotide is pathogenic.
[0413] 46. The computer-implemented method of clause 1, wherein the convolutional neural network is an attention-based neural network.
[0414] 47. The computer-implemented method of clause 1, wherein the tensor includes a distance channel in amino acid units further encoded with a reference allele channel.
[0415] 48. The computer-implemented method of clause 1, wherein the tensor includes a distance channel for amino acid units further encoded with an annotation channel.
[0416] 49. The computer-implemented method of clause 1, wherein the tensor includes an amino acid-based distance channel further encoded with a structural confidence channel.
[0417] 50. Accessing a three-dimensional structure of a reference amino acid sequence of the protein and fitting a three-dimensional grid of voxels on atoms in the three-dimensional structure on an amino acid basis to generate distance channels per atom category; Atoms span multiple atom categories, an atom category of the plurality of atom categories identifying an atomic element of the amino acid; each of the atomic category-based distance channels has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels; the 3D distance values specify the distances from corresponding voxels in the 3D grid of voxels to atoms of corresponding atom categories in the multiple atom categories; encoding an alternative allele channel into each voxel in a three-dimensional grid of voxels, the alternative allele channel being a one-hot coded three-dimensional representation of a variant amino acid expressed by a variant nucleotide; encoding an evolutionary conservation channel into each sequence of three-dimensional distance values across the atomic category-based distance channel on a voxel location basis; The evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, where amino acid-specific conservation frequencies are selected according to amino acid proximity to the corresponding voxel. applying a three-dimensional convolution to a tensor comprising distance channels per atomic category encoded with alternative allele channels and respective evolutionarily conserved channels; and determining a pathogenicity of the variant nucleotide based at least in part on the tensor.
[0418] 51. Accessing a three-dimensional structure of a reference amino acid sequence of a protein and fitting a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid basis to generate distance channels on an amino acid basis; each of the amino acid-based distance channels has a three-dimensional distance value for each voxel in a three-dimensional grid of voxels; the three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atom of the corresponding reference amino acid in the reference amino acid sequence; encoding an alternative allele channel into each voxel in a three-dimensional grid of voxels, the alternative allele channel being a one-hot coded three-dimensional representation of a variant amino acid expressed by a variant nucleotide; encoding an evolutionary conservation channel into each sequence of three-dimensional distance values across the amino acid-based distance channel on a voxel-location basis; The evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, where amino acid-specific conservation frequencies are selected according to amino acid proximity to the corresponding voxel. generating a tensor including distance channels for amino acid units coded using alternative allele channels and respective evolutionarily conserved channels.
[0419] 52. Accessing a three-dimensional structure of a reference amino acid sequence of the protein and fitting a three-dimensional grid of voxels on atoms in the three-dimensional structure on an amino acid basis to generate distance channels per atom category; Atoms span multiple atom categories, an atom category of the plurality of atom categories identifying an atomic element of the amino acid; each of the atomic category-based distance channels has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels; the 3D distance values specify the distances from corresponding voxels in the 3D grid of voxels to atoms of corresponding atom categories in the multiple atom categories; encoding an alternative allele channel into each voxel in a three-dimensional grid of voxels, the alternative allele channel being a one-hot coded three-dimensional representation of a variant amino acid expressed by a variant nucleotide; encoding an evolutionary conservation channel on a voxel-location basis into each sequence of three-dimensional distance values across atomic-category distance channels, the evolutionary conservation channel being a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, the amino acid-specific conservation frequencies being selected according to amino acid proximity to the corresponding voxel; generating a tensor including distance channels per atomic category encoded using alternative allele channels and respective evolutionarily conserved channels.
[0420] 1. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, accessing a three-dimensional structure of a reference amino acid sequence of the protein and fitting a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid-by-amino acid basis to generate distance channels on an amino acid-by-amino acid basis; each of the amino acid-based distance channels has a three-dimensional distance value for each voxel in a three-dimensional grid of voxels; the three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atom of the corresponding reference amino acid in the reference amino acid sequence; encoding an alternative allele channel into each voxel in a three-dimensional grid of voxels, the alternative allele channel being a one-hot coded three-dimensional representation of a variant amino acid expressed by a variant nucleotide; encoding an evolutionary conservation channel into each sequence of three-dimensional distance values across the amino acid-based distance channel on a voxel-location basis; The evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, where amino acid-specific conservation frequencies are selected according to amino acid proximity to the corresponding voxel. applying a three-dimensional convolution to a tensor comprising distance channels of amino acid units encoded with alternative allele channels and respective evolutionarily conserved channels; and determining a pathogenicity of the variant nucleotide based at least in part on the tensor.
[0421] 2. The computer-readable medium of clause 1, wherein the operation further comprises centering the three-dimensional grid of voxels on the alpha carbon atom of each residue of the reference amino acid in the reference amino acid sequence.
[0422] 3. The computer-readable medium of clause 2, wherein the operation further comprises centering the three-dimensional grid of voxels on the alpha carbon atom of a residue of the particular reference amino acid that corresponds to the variant amino acid.
[0423] 4. The computer-readable medium of clause 3, wherein the operation further includes encoding, in the tensor, the orientation of the reference amino acid and the position of the particular reference amino acid within the reference amino acid sequence by multiplying the three-dimensional distance value of the reference amino acid preceding the particular reference amino acid by a directionality parameter.
[0424] 5. The computer-readable medium of clause 4, wherein the distance is a nearest atom distance from a corresponding voxel center in a three-dimensional grid of voxels to a nearest atom of the corresponding reference amino acid.
[0425] 6. The computer-readable medium of clause 5, wherein the nearest neighbor atomic distance is a Euclidean distance.
[0426] 7. The computer-readable medium of clause 6, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0427] 8. The computer-readable medium of clause 5, wherein the reference amino acid has an alpha carbon atom, and the distance is a nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding reference amino acid.
[0428] 9. The computer-readable medium of clause 5, wherein the reference amino acid has a beta carbon atom, and the distance is a nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding reference amino acid.
[0429] 10. The computer-readable medium of clause 5, wherein the reference amino acid has backbone atoms and the distance is a nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding reference amino acid.
[0430] 11. The computer-readable medium of clause 5, wherein the amino acid has a side chain atom and the distance is a nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding reference amino acid.
[0431] 12. The computer-readable medium of clause 3, wherein the operations further include encoding, in the tensor, a nearest atom channel that specifies a distance from each voxel to a nearest atom, the nearest atom being selected without regard to amino acid and atomic component of the amino acid.
[0432] 13. The computer-readable medium of clause 12, wherein the distance is a Euclidean distance.
[0433] 14. The computer-readable medium of clause 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0434] 15. The computer-readable medium of clause 12, wherein the amino acids include non-standard amino acids.
[0435] 16. The computer-readable medium of clause 1, wherein the tensor further includes an absent atom channel that specifies atoms not found within a predetermined radius of the voxel center.
[0436] 17. The computer-readable medium of clause 16, wherein the absent atomic channels are one-hot encoded.
[0437] 18. The computer-readable medium of clause 1, wherein the operations further include voxel-by-voxel encoding the reference allele channel into each voxel in the three-dimensional grid of voxels.
[0438] 19. The computer-readable medium of clause 18, wherein the reference allele amino acid is a one-hot encoded three-dimensional representation of the reference amino acid undergoing the variant amino acid.
[0439] 20. The computer-readable medium of clause 1, wherein the amino acid-specific conservation frequency identifies the level of conservation of each amino acid across multiple species.
[0440] 21. The operation selects nearest neighbor atoms to corresponding voxels across reference amino acids and atom categories; selecting a pan-amino acid conservation frequency for residues of the reference amino acid that include the nearest atom; and using a three-dimensional representation of the general amino acid conservation frequency as an evolutionary conservation channel.
[0441] 22. The computer-readable medium of clause 21, wherein a global amino acid conservation frequency is constructed for a particular position of the residue as observed in multiple species.
[0442] 23. The computer-readable medium of clause 21, wherein the general amino acid conservation frequency identifies whether or not there is a missing conservation frequency for a particular reference amino acid.
[0443] 24. The operation comprises selecting each nearest atom to the corresponding voxel in each of the reference amino acids; selecting a conservation frequency for each amino acid for each residue of the reference amino acid that includes the nearest atom; and using the three-dimensional representation of conservation frequencies per amino acid as an evolutionary conservation channel.
[0444] 25. The computer-readable medium of clause 24, wherein a conservation frequency for each amino acid is constructed for a particular position of the residue as observed in multiple species.
[0445] 26. The computer-readable medium of clause 24, wherein the conservation frequency for each amino acid identifies whether a missing conservation frequency exists for a particular reference amino acid.
[0446] 27. The computer-readable medium of clause 1, wherein the operation further includes voxel-wise encoding one or more annotation channels into each voxel in a three-dimensional grid of voxels, the annotation channels being a three-dimensional representation of a one-hot encoding of the residual annotation.
[0447] 28. The computer-readable medium of clause 27, wherein the annotation channel is a molecular processing annotation including initiator methionine, signal, transit peptide, propeptide, chain, and peptide.
[0448] 29. The computer-readable medium of clause 27, wherein the annotation channels are region annotations including topological domain, transmembrane, intramembrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
[0449] 30. The computer-readable medium of clause 27, wherein the annotation channel is a site annotation including an active site, a metal binding, a binding site, and a site.
[0450] 31. The computer-readable medium of clause 27, wherein the annotation channel is an amino acid modification annotation including non-canonical residues, modified residues, lipidation, glycosylation, disulfide bonds, and cross-links.
[0451] 32. The computer-readable medium of clause 27, wherein the annotation channel is a secondary structure annotation including helices, turns, and beta strands.
[0452] 33. The computer-readable medium of clause 27, wherein the annotation channel is an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflicts, non-adjacent residues, and non-terminal residues.
[0453] 34. The computer-readable medium of clause 1, wherein the operations further include voxel-wise encoding one or more structure confidence channels into each voxel in the three-dimensional grid of voxels, the structure confidence channels being three-dimensional representations of confidence scores identifying the quality of each residue structure.
[0454] 35. The computer-readable medium of clause 34, wherein the structural confidence channel is a global model quality estimation (GMQE).
[0455] 36. The computer-readable medium of clause 34, wherein the structural confidence channel is a qualitative model energy analysis (QMEAN) score.
[0456] 37. The computer-readable medium of clause 34, wherein the structural confidence channel is a temperature factor that specifies the degree to which residues satisfy the physical constraints of the respective protein structure.
[0457] 38. The computer-readable medium of clause 34, wherein the structure confidence channel is a template structure alignment that identifies the degree to which residues of nearest neighbor atoms to a voxel align the template structure.
[0458] 39. The computer-readable medium of clause 38, wherein the structure confidence channel is a template modeling score for the aligned template structure.
[0459] 40. The computer-readable medium of clause 39, wherein the structural confidence channel is the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0460] 41. The computer-readable medium of clause 1, wherein the operation further comprises rotating the atoms before the distance channels of the amino acid units are generated.
[0461] 42. The computer-readable medium of clause 1, wherein the operations further include using a 1x1x1 convolution, a 3x3x3 convolution, a regularized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer in the convolutional neural network.
[0462] 43. The computer-readable medium of clause 42, wherein 1x1x1 convolution and 3x3x3 convolution are three-dimensional convolutions.
[0463] 44. The computer-readable medium of clause 42, wherein a layer of 1x1x1 convolution processes the tensor and generates intermediate outputs that are convolutional representations of the tensor; a sequence of layers of 3x3x3 convolution processes the intermediate outputs and generates flattened outputs; a fully connected layer processes the flattened outputs and generates denormalized outputs; and a softmax classification layer processes the denormalized outputs and generates exponentially normalized outputs that identify the likelihood that the variant nucleotide is pathogenic and benign.
[0464] 45. The computer-readable medium of clause 44, wherein a sigmoid layer processes the non-normalized output and generates a normalized output that identifies the likelihood that the variant nucleotide is pathogenic.
[0465] 46. The computer-readable medium of clause 1, wherein the convolutional neural network is an attention-based neural network.
[0466] 47. The computer-readable medium of clause 1, wherein the tensor includes a distance channel in amino acid units further encoded with a reference allele channel.
[0467] 48. The computer-readable medium of clause 1, wherein the tensor includes a distance channel for amino acid units further encoded with an annotation channel.
[0468] 49. The computer-readable medium of clause 1, wherein the tensor includes an amino acid-based distance channel further encoded with a structural confidence channel.
[0469] 50. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, accessing a three-dimensional structure of a reference amino acid sequence of the protein and fitting a three-dimensional grid of voxels onto atoms in the three-dimensional structure on an amino acid basis to generate distance channels per atom category; Atoms span multiple atom categories, an atom category of the plurality of atom categories identifying an atomic element of the amino acid; each of the atomic category-based distance channels has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels; the 3D distance values specify the distances from corresponding voxels in the 3D grid of voxels to atoms of corresponding atom categories in the multiple atom categories; encoding an alternative allele channel into each voxel in a three-dimensional grid of voxels, the alternative allele channel being a one-hot coded three-dimensional representation of a variant amino acid expressed by a variant nucleotide; encoding an evolutionary conservation channel on a voxel-location basis into each sequence of three-dimensional distance values across atomic-category distance channels, the evolutionary conservation channel being a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, the amino acid-specific conservation frequencies being selected according to amino acid proximity to the corresponding voxel; applying a three-dimensional convolution to a tensor comprising distance channels per atomic category encoded with alternative allele channels and respective evolutionarily conserved channels; and determining a pathogenicity of the variant nucleotide based at least in part on the tensor.
[0470] 51. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, accessing a three-dimensional structure of a reference amino acid sequence of the protein and fitting a three-dimensional grid of voxels to atoms in the three-dimensional structure on an amino acid-by-amino acid basis to generate distance channels on an amino acid-by-amino acid basis; each of the amino acid-based distance channels has a three-dimensional distance value for each voxel in a three-dimensional grid of voxels; the three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atom of the corresponding reference amino acid in the reference amino acid sequence; encoding alternative allele channels by amino acid position into each three-dimensional distance value of each amino acid unit distance channel; the alternative allele channel is a three-dimensional representation of one-hot coding of variant amino acids expressed by variant nucleotides; encoding an evolutionary conservation channel into each sequence of three-dimensional distance values across the amino acid-based distance channel on a voxel-location basis; The evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, where amino acid-specific conservation frequencies are selected according to amino acid proximity to the corresponding voxel. and generating a tensor including distance channels for amino acid units coded using alternative allele channels and their respective evolutionarily conserved channels.
[0471] 52. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, accessing a three-dimensional structure of a reference amino acid sequence of the protein and fitting a three-dimensional grid of voxels onto atoms in the three-dimensional structure on an amino acid basis to generate distance channels per atom category; Atoms span multiple atom categories, an atom category of the plurality of atom categories identifying an atomic element of the amino acid; each of the atomic category-based distance channels has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels; the 3D distance values specify the distances from corresponding voxels in the 3D grid of voxels to atoms of corresponding atom categories in the multiple atom categories; encoding an alternative allele channel into each voxel in a three-dimensional grid of voxels, the alternative allele channel being a one-hot coded three-dimensional representation of a variant amino acid expressed by a variant nucleotide; encoding an evolutionary conservation channel on a voxel-location basis into each sequence of three-dimensional distance values across atomic-category distance channels, the evolutionary conservation channel being a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, the amino acid-specific conservation frequencies being selected according to amino acid proximity to the corresponding voxel; generating a tensor including distance channels per atomic category encoded with alternative allele channels and respective evolutionarily conserved channels.
[0472] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0473] Specific Embodiment 3 Clause 3 1. A computer-implemented method for efficiently determining which elements of a sequence are nearest to uniformly spaced cells in a grid, wherein the elements have element coordinates and the cells have dimensional cell indices and cell coordinates, the method comprising the steps of generating an element-to-cell mapping that maps a subset of cells to each of the elements, wherein the subset of cells mapped to a particular element in the sequence includes the nearest cell in the grid and one or more neighboring cells in the grid; The nearest neighbor cells are selected based on matching the element coordinates of a particular element to the cell coordinates; the neighboring cells are selected based on being consecutively adjacent to a nearest neighbor cell and within a distance proximity range from a particular element, and generating a cell-to-element mapping that maps a subset of elements to each of the cells, wherein the subset of elements mapped to a particular cell in the grid includes elements in a sequence that are mapped to the particular cell by the element-to-cell mapping; determining, for each of the cells, the nearest element in the sequence using a cell-to-element mapping; and determining the nearest neighbor element to the particular cell based on the distance between the particular cell and an element in the subset of elements.
[0474] 2. The computer-implemented method of clause 1, wherein matching element coordinates of a particular element to cell coordinates further comprises truncating the fractional portion of the element coordinates to generate truncated element coordinates.
[0475] 3. Matching element coordinates of a particular element to cell coordinates is For a first dimension, matching a first truncated element coordinate in the truncated element coordinates to a first cell coordinate of a first cell in the grid and selecting a first dimension index for the first cell; For a second dimension, matching a second truncated element coordinate in the truncated element coordinates to a second cell coordinate of a second cell in the grid and selecting a second dimension index for the second cell; For a third dimension, matching a third truncated element coordinate in the truncated element coordinates to a third cell coordinate of a third cell in the grid and selecting a third dimension index of the third cell; using the selected first, second, and third dimension indices to generate a cumulative sum based on position-wise weighting of the selected first, second, and third dimension indices by a power of the radix; and using the cumulative sum as a cell index for selecting a nearest neighbor cell.
[0476] 4. The computer-implemented method of clause 1, wherein the distance is calculated between a cell coordinate of a particular cell and an element coordinate of an element in the subset of elements.
[0477] 5. The computer-implemented method of clause 1, wherein the sequence is a protein sequence of amino acids.
[0478] 6. The computer-implemented method of clause 5, wherein the elements are atoms of an amino acid.
[0479] 7. Generating a mapping from elements to cells, and generating a mapping from cells to elements, and using the mapping from cells to elements, determining that for each cell, the nearest neighbor element is O(a * determining that the algorithm has a runtime complexity of (f + v), a is the number of atoms, f is the number of amino acids, v is the number of cells, * 7. The computer-implemented method of claim 6, wherein:
[0480] 8. The computer-implemented method of clause 7, wherein the atom comprises an alpha carbon atom.
[0481] 9. The computer-implemented method of clause 7, wherein the atom comprises a beta carbon atom.
[0482] 10. The computer-implemented method of clause 7, wherein the atoms include non-carbon atoms.
[0483] 11. The computer-implemented method of clause 1, wherein the cells are three-dimensional voxels.
[0484] 12. The computer-implemented method of clause 11, wherein the cell coordinates are three-dimensional coordinates.
[0485] 13. The computer-implemented method of clause 12, wherein the element coordinates are three-dimensional coordinates.
[0486] 14. The computer-implemented method of clause 1, wherein the neighboring cells are selected based on being within an index neighborhood range from a nearest neighbor cell.
[0487] 15. The computer-implemented method of clause 1, wherein the neighboring cells are selected based on being within a neighborhood of cells in a grid that includes a nearest neighbor cell.
[0488] 16. The computer-implemented method of clause 1, wherein the sequence includes M elements and the subset of elements includes N elements, where M>>N.
[0489] 17. A computer-implemented method for efficiently determining which atoms in a protein are nearest neighbors to voxels in a grid, the atoms having three-dimensional (3D) atomic coordinates and the voxels having 3D voxel coordinates, the method comprising: generating an atom-to-voxel mapping that maps a selected containing voxel to each of the atoms based on matching the 3D atomic coordinates of a particular atom of the protein to the 3D voxel coordinates in the grid; 1. A computer-implemented method comprising: generating a voxel-to-atom mapping that maps a subset of atoms to each of the voxels, wherein the subset of atoms mapped to a particular voxel in the grid includes atoms in a protein that are mapped to the particular voxel by the atom-to-voxel mapping; and determining, for each of the voxels, a nearest atom in the protein using the voxel-to-atom mapping.
[0490] 18. The computer-implemented method of clause 17, wherein the steps of clause 17 have a runtime complexity of O (number of atoms).
[0491] Other implementations of the methods described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the above-described methods. Yet another implementation of the methods described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the above-described methods.
[0492] While the present invention has been disclosed with reference to the preferred embodiments and examples set forth above, it should be understood that these examples are intended in an illustrative and not a limiting sense. Modifications and combinations will readily occur to those skilled in the art, and such modifications and combinations are deemed to be within the spirit of the invention and the scope of the following claims.
Claims
1. accessing a three-dimensional structure of a reference amino acid sequence of a protein; mapping atoms of the three-dimensional structure to voxels of the three-dimensional voxel grid by associating each atom with an individual voxel of the three-dimensional voxel grid; defining, for each voxel in the three-dimensional voxel grid on an amino acid basis, a distance channel per amino acid that includes a three-dimensional distance value that specifies the distance from each respective voxel to a corresponding atom mapped to each respective voxel; encoding an alternative allele channel into each voxel in a three-dimensional grid voxel grid, the alternative allele channel being a three-dimensional representation of a one-hot encoding of variant amino acids expressed by variant nucleotides; encoding an evolutionary conservation channel into each sequence of three-dimensional distance values across a distance channel on a voxel-location basis in units of amino acids, the evolutionary conservation channel being a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, the amino acid-specific conservation frequencies being selected according to amino acid proximity to a corresponding voxel; applying a three-dimensional convolution to a tensor comprising distance channels of amino acid units encoded with alternative allele channels and respective evolutionarily conserved channels; and determining a pathogenicity of the variant nucleotide based at least in part on the tensor.
2. 10. The computer-implemented method of claim 1, further comprising centering the three-dimensional voxel grid on an alpha carbon atom of each residue of a reference amino acid in the reference amino acid sequence.
3. 3. The computer-implemented method of claim 2, further comprising centering the three-dimensional voxel grid on the alpha carbon atom of a residue of the particular reference amino acid that corresponds to the variant amino acid.
4. 4. The computer-implemented method of claim 3, further comprising encoding the directionality of the reference amino acid and the position of the particular reference amino acid in the reference amino acid sequence by multiplying the three-dimensional distance value for the reference amino acid preceding the particular reference amino acid by a directionality parameter in the tensor.
5. 5. The computer-implemented method of claim 4, wherein the distances identified by the three-dimensional distance values are nearest atom distances from corresponding voxel centers in a three-dimensional voxel grid to nearest atoms of the corresponding reference amino acids.
6. The computer-implemented method of claim 5 , wherein the nearest neighbor atom distance is a Euclidean distance.
7. The computer-implemented method of claim 6 , wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
8. 6. The computer-implemented method of claim 5, wherein the reference amino acid has an alpha carbon atom, and the distance specified by the three-dimensional distance value is the nearest alpha carbon atom distance from the corresponding voxel center to the nearest alpha carbon atom of the corresponding reference amino acid.
9. 6. The computer-implemented method of claim 5, wherein the reference amino acid has a beta carbon atom, and the distance specified by the three-dimensional distance value is a nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding reference amino acid.
10. 6. The computer-implemented method of claim 5, wherein the reference amino acid has backbone atoms, and the distance specified by the three-dimensional distance value is a nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding reference amino acid.
11. 6. The computer-implemented method of claim 5, wherein the reference amino acid has a side chain atom, and the distance specified by the three-dimensional distance value is a nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding reference amino acid.
12. 4. The computer-implemented method of claim 3, further comprising encoding, in the tensor, a nearest atom channel that specifies a distance from each voxel to a nearest atom, wherein the nearest atom is selected without regard to the reference amino acid and the atomic component of the reference amino acid.
13. The computer-implemented method of claim 12 , wherein the distance is a Euclidean distance.
14. The computer-implemented method of claim 13 , wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
15. 13. The computer-implemented method of claim 12, wherein the reference amino acids include non-standard amino acids.
16. The computer-implemented method of claim 1 , wherein the tensor further includes an absent atom channel that specifies atoms that are not found within a predetermined radius of the voxel center.
17. The computer-implemented method of claim 16 , wherein the absent atomic channels are one-hot coded.
18. The computer-implemented method of claim 1 , further comprising voxel-by-voxel encoding a reference allele channel for each voxel in a three-dimensional voxel grid.
19. 20. The computer-implemented method of claim 18, wherein the reference allele amino acid for the reference allele channel is a one-hot coded three-dimensional representation of the reference amino acid undergoing the variant amino acid.
20. 10. The computer-implemented method of claim 1, wherein the amino acid-specific conservation frequencies identify the level of conservation of each amino acid across multiple species.
21. selecting nearest neighbor atoms to corresponding voxels across the reference amino acids from the reference amino acid sequence and atom categories; selecting a pan-amino acid conservation frequency for residues of the reference amino acid that include the nearest atom; and using a three-dimensional representation of the general amino acid conservation frequencies as an evolutionary conservation channel.
22. 22. The computer-implemented method of claim 21, wherein a global amino acid conservation frequency is constructed for a particular position of a residue observed in multiple species.
23. 22. The computer-implemented method of claim 21, wherein the general amino acid conservation frequency identifies whether there is a missing conservation frequency for a particular reference amino acid.
24. selecting the nearest atom to the corresponding voxel in each of the reference amino acids; selecting a conservation frequency for each amino acid for each residue of the reference amino acid that includes the nearest atom; and using a three-dimensional representation of the conservation frequency for each amino acid as an evolutionary conservation channel.
25. 25. The computer-implemented method of claim 24, wherein a conservation frequency for each amino acid is constructed for a particular position of the residue observed in multiple species.
26. 25. The computer-implemented method of claim 24, wherein the conservation frequency for each amino acid identifies whether there is a missing conservation frequency for a particular reference amino acid.
27. 2. The computer-implemented method of claim 1, further comprising voxel-wise encoding one or more annotation channels into each voxel in a three-dimensional voxel grid, wherein the one or more annotation channels are a three-dimensional representation of a one-hot encoding of the residual annotation.
28. 28. The computer-implemented method of claim 27, wherein one or more annotation channels are molecular process annotations including initiator methionine, signal, transit peptide, propeptide, chain, and peptide.
29. 28. The computer-implemented method of claim 27, wherein one or more annotation channels are region annotations including topological domain, transmembrane, intramembrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
30. 28. The computer-implemented method of claim 27, wherein the one or more annotation channels are site annotations including active sites, metal binding, binding sites, and sites.
Citation Information
Patent Citations
Deep Convolutional Neural Networks for Variant Classification
JP2020530918A
Artificial Intelligence-Based Quality Scoring
JP2022524562A
Multi-channel protein voxelization to predict variant pathogenicity using deep convolutional neural networks
JP2024513995A
JPP6834029B
Systems and methods for applying a convolutional network to spatial data
US20160196480A1