Deep convolutional neural network for predicting mutant pathogenicity using three-dimensional (3D) protein structures
A deep learning approach using voxelized 3D protein structures and evolutionary conservation improves the prediction of pathogenicity in genetic variants, addressing data limitations and enhancing accuracy in clinical interpretation.
Patent Information
- Application Number
- JP2023563032
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-24
- Filing Date
- 2022-04-14
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2042-04-14
AI Technical Summary
Existing methods for predicting the pathogenicity of genetic variants rely heavily on prior knowledge and are limited by the availability of clinical data, leading to potential overfitting and reduced accuracy, especially when dealing with non-coding variants and complex phenotypes.
A deep learning approach using multi-channel voxelized representations of 3D protein structures is employed, integrating evolutionary conservation and structural information to predict the pathogenicity of genetic variants, leveraging a deep convolutional neural network trained on both human and primate data to improve classification accuracy.
The method significantly enhances the prediction of pathogenicity by reducing reliance on prior knowledge and improving performance over conventional methods, particularly in distinguishing benign and pathogenic mutations, and provides a more robust tool for clinical interpretation.
Smart Images

Figure 0007712387000001 
Figure 0007712387000002 
Figure 0007712387000003
Abstract
Description
Technical Field
[0001] (Priority Application) This application claims priority to U.S. Non-Provisional Application No. 17 / 468,411, filed on September 7, 2021, entitled "Deep Convolutional Neural Networks to Predict Variant Pathogenicity using Three-Dimensional (3D) Protein Structures" (Attorney Docket No. ILLM 1037-3 / IP-2051A-US), which is a continuation of U.S. Non-Provisional Application No. 17 / 232,056, filed on April 15, 2021, entitled "Deep Convolutional Neural Networks to Predict Variant Pathogenicity using Three-Dimensional (3D) Protein Structures" (Attorney Docket No. ILLM 1037-2 / IP-2051-US).
[0002] This application also claims priority to U.S. Non-Provisional Application No. 17 / 703,935, filed on March 24, 2022, entitled "Multi-channel Protein Voxelization To Predict Variant Pathogenicity Using Deep Convolutional Neural Networks" (Attorney Docket No. ILLM 1047-2 / IP-2142-US), which in turn claims priority or the benefit thereof to U.S. Provisional Application No. 63 / 175,495, filed on April 15, 2021, entitled "Multi-channel Protein Voxelization To Predict Variant Pathogenicity Using Deep Convolutional Neural Networks" (Attorney Docket No. ILLM 1047-1 / IP-2142-PRV).
[0003] This application also claims the benefit of priority of U.S. Non-Provisional Patent Application No. 17 / 703,958, entitled "Efficient Voxelization For Deep Learning", filed on Mar. 24, 2022 (Attorney Docket No. ILLM 1048-2 / IP-2143-US), which in turn claims the priority or benefit of U.S. Provisional Patent Application No. 63 / 175,767, entitled "Efficient Voxelization For Deep Learning", filed on Apr. 16, 2021 (Attorney Docket No. ILLM 1048-1 / IP-2143-PRV).
[0004] The priority applications are hereby incorporated by reference herein for all purposes.
[0005] (Related Applications) This application is related to the concurrently filed PCT Patent Application, entitled "Artificial Intelligence-based Analysis of Protein Three-Dimensional(3D)Structures" (Attorney Docket No. ILLM 1037-4 / IP-2051-PCT). The related application is hereby incorporated by reference herein for all purposes.
[0006] (Field of the Invention) The disclosed technology relates to artificial intelligence computers and digital data processing systems, and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for inference with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using deep convolutional neural networks to analyze multi-channel voxelized data.
[0007] (Incorporation) The following is hereby incorporated by reference herein for all purposes as if fully set forth herein.
[0008] Sundaram, L et al., Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018);
[0009] Jaganathan, K. et al., Predicting splicing from primary sequence with deep learning. Cell 176, 535-548 (2019);
[0010] U.S. Patent Application No. 62 / 573,144 (144 years) titled "TRAINING A DEEP PATHOGENICITY CLASSIFIER USING LARGE-SCALE BENIGN TRAINING DATA" filed on October 16, 2017 (Attorney Docket No. ILLM 1000-1 / IP-1611-PRV);
[0011] U.S. Patent Application No. 62 / 573,149 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV) titled "PATHOGENICITY CLASSIFIER BASED ON DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs)" filed on October 16, 2017;
[0012] U.S. Patent Application No. 62 / 573,153 (Attorney Docket No. ILLM 1000-3 / IP-1613-PRV) titled "DEEP SEMI-SUPERVISED LEARNING THAT GENERATES LARGE-SCALE PATHOGENIC TRAINING DATA" filed on October 16, 2017.
[0013] U.S. Patent Application No. 62 / 582,898 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV) entitled "PATHOGENICITY CLASSIFICATION OF GENOMIC DATA USING DEEP CONVOLUTIONAL NEURAL NETWORKS(CNNs)", filed on November 7, 2017;
[0014] U.S. Patent Application No. 16 / 160,903 (Attorney Docket No. ILLM 1000-5 / IP- 1611-US) entitled "DEEP LEARNING-BASED TECHNIQUES FOR TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS", filed on October 15, 2018;
[0015] U.S. Patent Application No. 16 / 160,986 (Attorney Docket No. ILLM 1000-6 / IP- 1612-US) entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS FOR VARIANT CLASSIFICATION", filed on October 15, 2018;
[0016] U.S. Patent Application No. 16 / 160,968 (Attorney Docket No. ILLM 1000-7 / IP-1613-US) entitled "SEMI- SUPERVISED LEARNING FOR TRAINING AN ENSEMBLE OF DEEP CONVOLUTIONAL NEURAL NETWORKS", filed on October 15, 2018; and
[0017] U.S. Patent Application No. 16 / 407,149 (Attorney Docket No. ILLM 1010-1 / IP-1734-US) entitled "DEEP LEARNING-BASED TECHNIQUES FOR PRE-TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS", filed on May 8, 2019.
BACKGROUND ART
[0018] The subject matter discussed in this section should not be assumed to be prior art simply as a result of mention in this section. Similarly, the problems mentioned in this section, or problems associated with the subject matter provided as background, should not be assumed to have been previously recognized in the prior art. The subject matter of this section simply represents different approaches, which in themselves may also correspond to embodiments of the claimed technology.
[0019] Genomics in a broad sense, also called functional genomics, aims to characterize the functions of all genomic elements of an organism by using genome-scale assays such as genome sequencing, transcriptome profiling, and proteomics. Genomics has emerged as a data-driven science, which operates by discovering new characteristics from the investigation of genome-scale data rather than by testing pre-conceived models and hypotheses. Applications of genomics include finding associations between genotypes and phenotypes, discovering biomarkers for patient stratification, predicting gene functions, and mapping biochemically active genomic regions such as transcriptional enhancers.
[0020] Genomic data is too large and complex to be mined by visual inspection of pairwise correlations alone. Instead, analytical tools are needed to support the discovery of unexpected relationships, derive new hypotheses and models, and make predictions. Unlike some algorithms where assumptions and domain expertise are hard-coded, machine learning algorithms are designed to automatically detect patterns in data. Therefore, machine learning algorithms are suitable for data-driven sciences, especially genomics. However, the performance of machine learning algorithms can strongly depend on how the data is represented, i.e., how each variable (also called a feature) is computed. For example, to classify tumors as malignant or benign from fluorescence microscope images, a preprocessing algorithm can detect cells, identify cell types, and generate a list of cell counts for each cell type.
[0021] A machine learning model can take the estimated cell count, which is an example of a manually designed feature, as an input feature for classifying tumors. The central problem is that the classification performance highly depends on the quality and relevance of these features. For example, related visual features such as cell morphology, distance between cells, or localization within an organ are not captured in cell counting, and this incomplete representation of data may reduce the classification accuracy.
[0022] Deep learning, a subfield of machine learning, addresses this problem by embedding the calculation of features into the machine learning model itself to generate an end-to-end model. This achievement was realized by the development of deep neural networks, which are machine learning models that include successive basic operations that calculate increasingly complex features by taking the results of previous operations as inputs. Deep neural networks can improve prediction accuracy by discovering highly complex related features such as cell morphology and the spatial configuration of cells in the above example. The construction and learning of deep neural networks have been made possible by the explosion of data, algorithmic advancements, and a substantial increase in computing power, especially through the use of graphics processing units (GPUs).
[0023] The goal of supervised learning is to obtain a model that takes features as inputs and returns predictions of so-called target variables. An example of a supervised learning problem is to predict whether an intron is spliced (the target) considering features on RNA such as the presence or absence of a standard splice site sequence, the position of the splicing branch point, or the length of the intron. What a machine learning model learns refers to learning its parameters, which generally includes minimizing a loss function with respect to the training data for the purpose of making accurate predictions for unknown data.
[0024] For many supervised learning problems in computational biology, the input data can be represented as a table with multiple columns or features, each containing numerical or categorical data that is potentially useful for making predictions. Some input data is naturally represented as tabular features (e.g., temperature or time), while other input data needs to be first transformed using a process called feature extraction in order to fit into a tabular representation (e.g., converting a deoxyribonucleic acid (DNA) sequence to k-mer counts). For the intron-splicing prediction problem, the presence or absence of standard splice site sequences, the location of splicing branch points, and intron length can be preprocessed features collected in tabular form. Tabular data is the standard for a wide range of supervised machine learning models, ranging from simple linear models such as logistic regression to more flexible non-linear models such as neural networks and many others.
[0025] Logistic regression is a binary classifier, i.e., a supervised learning model that predicts a binary target variable. Specifically, logistic regression predicts the probability of the positive class by calculating the weighted sum of input features mapped to the [0,1] interval using the sigmoid function, which is a type of activation function. The parameters of logistic regression, or other linear classifiers that use different activation functions, are the weights in the weighted sum. Linear classifiers fail when, for example, the class of whether an intron is spliced or not cannot be adequately discriminated by the weighted sum of input features. To improve prediction performance, new input features can be added manually by transforming or combining existing features in new ways, e.g., by taking powers or pairwise products.
[0026] Neural networks use hidden layers to automatically learn these non-linear feature transformations. Each hidden layer can be thought of as a collection of linear models with outputs transformed by a non-linear activation function such as the sigmoid function or the more general rectified linear unit (ReLU). At the same time, these layers organize the input features into complex patterns relevant to the task of distinguishing between two classes, facilitating the task of discrimination.
[0027] Deep neural networks use many hidden layers and are said to be fully connected when each neuron receives input from all neurons in the previous layer. Neural networks generally learn using stochastic gradient descent, an algorithm well-suited for learning models on very large datasets. Implementations of neural networks using modern deep learning frameworks enable rapid prototyping with different architectures and datasets. Fully connected neural networks can be used for several genomics applications, including predicting the proportion of spliced exons for a given sequence from sequence features such as the presence or sequence conservation of splice factor binding motifs, prioritizing gene variants that cause potential diseases, and predicting cis-regulatory elements using features such as chromatin marks, gene expression, and evolutionary conservation in a given genomic region.
[0028] For effective prediction, local dependencies in spatial and longitudinal data must be considered. For example, shuffling of DNA sequences or pixels of an image severely disrupts the information pattern. These local dependencies define spatial or longitudinal data, distinct from tabular data where the ordering of features is arbitrary. Consider the problem of classifying genomic regions as bound or unbound by a specific transcription factor, where the binding regions are defined as high-confidence binding events in chromatin immunoprecipitation following sequencing (ChIP-seq) data. Transcription factors bind to DNA by recognizing sequence motifs. A fully connected layer based on sequence-derived features such as the number or position of k-mer instances in the sequence or position weight matrix (PWM) matches can be used for this task. Such models can generalize well to sequences with the same motif located at different positions because k-mer or PWM instance frequencies are robust to shifting of motifs within the sequence. However, they are unable to recognize patterns that depend on combinations of multiple motifs with clearly defined intervals for transcription factor binding. Furthermore, the number of possible k-mers increases exponentially with k-mer length, which poses problems of both conservation and overfitting.
[0029] The convolutional layer is a special form of the fully connected layer where the same fully connected layer is applied locally, e.g., within a 6bp window, to all sequence positions. This approach can also be considered as scanning the sequence using multiple PWMs, e.g., for transcription factors GATA1 and TAL1. By using the same model parameters across positions, the total number of parameters is dramatically reduced and the network can detect motifs at positions not seen during learning. Each convolutional layer scans the sequence using several filters by generating scalar values quantifying the match between the filter and the sequence at all positions. As in the case of fully connected neural networks, a non-linear activation function (generally ReLU) is applied at each layer. Next, a pooling operation is applied which aggregates the activations within consecutive bins across the position axis and generally takes the maximum or average activation for each channel. Pooling reduces the effective sequence length and coarsens the signal. Subsequent convolutional layers are constructed from the output of the previous layer and can detect whether GATA1 and TAL1 motifs are present within a certain distance range. Finally, the output of the convolutional layer can be used as input to a fully connected neural network to perform the final prediction task. Thus, different types of neural network layers (e.g., fully connected and convolutional layers) can be combined within a single neural network.
[0030] Convolutional neural networks (CNNs) can predict various molecular phenotypes based only on DNA sequences. Applications include the classification of transcription factor binding sites and the prediction of molecular phenotypes such as chromatin features, DNA contact maps, DNA methylation, gene expression, translation efficiency, RBP binding, and microRNA (miRNA) targets. In addition to predicting molecular phenotypes from sequences, convolutional neural networks can be applied to more technical tasks that have traditionally been addressed by manually designed bioinformatics pipelines. For example, convolutional neural networks can predict the specificity of guide RNAs, denoise ChIP-seq, improve Hi-C data resolution, predict laboratory origin from DNA sequences, and call gene variants. Convolutional neural networks have also been used to model long-range dependencies in the genome. Interacting regulatory elements can be located far apart on the un-folded linear DNA sequence, but these elements are often proximal in the actual 3D chromatin conformation. Thus, modeling molecular phenotypes from linear DNA sequences, while a rough approximation of chromatin, can be improved by enabling long-range dependencies and allowing the model to implicitly learn aspects of 3D configurations such as promoter-enhancer loops. This is achieved by using dilated convolutions with receptive fields of up to 32 kb. Dilated convolutions also enable splice sites to be predicted from sequences using a receptive field of 10 kb, thereby allowing the integration of gene sequences over distances the same length as typical human introns (see Jaganathan, K. et al., Predicting splicing from primary sequence with deep learning. Cell 176, 535-548 (2019)).
[0031] Different types of neural networks can be characterized by their parameter sharing schemes. For example, fully connected layers do not have parameter sharing, while convolutional layers impose translational invariance by applying the same filter at all positions of their inputs. Recurrent neural networks (RNNs) are an alternative to convolutional neural networks for processing sequential data such as DNA sequences or time series, implementing different parameter sharing schemes. Recurrent neural networks apply the same operation to each sequence element. This operation takes as input the memory of the previous sequence element and the new input. This updates the memory and optionally emits an output, which is passed to subsequent layers or used directly as a model prediction. By applying the same model at each sequence element, recurrent neural networks are invariant to the position index in the processed sequence. For example, a recurrent neural network can detect open reading frames in a DNA sequence regardless of their position in the sequence. This task requires the recognition of a specific set of inputs such as a start codon and subsequent in-frame stop codons.
[0032] The main advantage of recurrent neural networks over convolutional neural networks is that, theoretically, they can inherit information through memory over infinitely long sequences. Furthermore, recurrent neural networks can naturally process sequences of widely varying lengths, such as mRNA sequences. However, convolutional neural networks combined with various tricks (such as dilated convolutions) can achieve performance comparable to or even better than recurrent neural networks for sequence modeling tasks such as audio synthesis and machine translation. Recurrent neural networks can aggregate the outputs of convolutional neural networks to predict single-cell DNA methylation status, RBP binding, transcription factor binding, and DNA accessibility. Furthermore, since recurrent neural networks apply sequential operations, they cannot be easily parallelized and are thus much slower computationally than convolutional neural networks.
[0033] Most of the human genetic code is common to all humans, but each human has a unique genetic code. In some cases, the human genetic code can contain outliers called genetic variants that can be common among relatively small groups of individuals in the human population. For example, a particular human protein may contain a particular sequence of amino acids, but a variant of that protein may differ by one amino acid in the same particular sequence in other respects.
[0034] Genetic variants can be pathogenic and can cause disease. Most of these genetic variants have been depleted from the genome by natural selection, but the ability to identify which genetic variants are likely to be pathogenic can help researchers focus on these genetic variants and gain an understanding of the corresponding diseases and their diagnosis, treatment, or cure. The clinical interpretation of millions of human genetic variants remains unclear. Some of the most frequent pathogenic variants are missense mutations of single nucleotides that change the amino acids of a protein. However, not all missense mutations are pathogenic.
[0035] Models that can directly predict molecular phenotypes from biological sequences can be used as in silico perturbation tools to investigate the relationship between genetic and phenotypic variations, and have emerged as new methods for quantitative trait locus identification and mutant prioritization. Considering that the majority of mutants identified by genome-wide association studies of complex phenotypes are non-coding, which makes it difficult to infer their effects and contributions to phenotypes, these approaches are very important. Furthermore, linkage disequilibrium results in blocks of co-inherited mutants, which makes it difficult to accurately identify individual causal mutants. Therefore, sequence-based deep learning models that can be used as collation tools to evaluate the effects of such mutants provide a promising approach for finding potential drivers of complex phenotypes. As an example, with regard to transcription factor binding, chromatin accessibility, or gene expression prediction, it is possible to indirectly predict the effects of non-coding single nucleotide mutants and short insertions or deletions (indels) from the differences between two mutants. Another example is the prediction of novel splice site creation from the sequence or quantitative effects of gene mutants on splicing.
[0036] To predict the pathogenicity of missense variants from protein sequences and sequence conservation data, an end-to-end deep learning approach for variant effect prediction is applied (see Sundaram, L. et al., Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018), referred to herein as "PrimateAI"). PrimateAI uses a deep neural network trained on known pathogenic variants using data augmentation that uses inter-species information. In particular, PrimateAI uses the sequences of wild-type and mutant proteins to compare differences and uses the trained deep neural network to determine the pathogenicity of the mutation. Such an approach that utilizes protein sequences for pathogenicity prediction is promising because it can avoid the circularity problem and overfitting to previous knowledge. However, the number of clinical data available in ClinVar is relatively small compared to the number of data sufficient to effectively train a deep neural network. To overcome this data deficiency, PrimateAI used common human variants and variants from primates as benign data and variants simulated based on the trinucleotide context as unlabeled data.
[0037] PrimateAI performs better than conventional methods when learning directly on sequence alignments. PrimateAI directly learns important protein domains, conserved amino acid positions, and sequence dependencies from training data consisting of approximately 120,000 human samples. PrimateAI substantially outperforms the performance of other variant pathogenicity prediction tools in distinguishing benign and pathogenic de novo mutations in candidate developmental disorder genes and in reproducing prior knowledge in ClinVar. These results suggest that PrimateAI represents an important advance for variant classification tools that can reduce dependence on prior knowledge in clinical reporting.
[0038] At the heart of protein biology is the understanding of how structural elements give rise to the functions they observe. Excessive protein structure data enables the development of computational methods to systematically derive the rules governing structure-function relationships. However, the performance of these methods depends critically on the choice of protein structure representation.
[0039] Protein sites are microenvironments within a protein structure, distinguished by their structural or functional roles. Sites can be defined by their three-dimensional (3D) position and the local neighborhood around this position where the structure or function exists. At the heart of rational protein engineering is the understanding of how the structural arrangement of amino acids creates functional features within protein sites. Determining the structural and functional roles of individual amino acids within a protein provides information to assist in the manipulation and modification of protein function. Identifying functionally or structurally important amino acids enables focused engineering efforts such as site-directed mutagenesis to alter the functional properties of the target protein. Alternatively, this knowledge can help avoid engineering designs that disable the desired function.
[0040] Since it has been established that structure is much more conserved than sequence, the increasing amount of protein structure data provides an opportunity to systematically study the underlying patterns governing structure-function relationships using a data-driven approach. A fundamental aspect of any computational protein analysis is how protein structure information is represented. The performance of machine learning methods often depends more on the choice of data representation than on the machine learning algorithm used. A good representation efficiently captures the most important information, while an inadequate representation generates a noisy distribution without underlying patterns.
[0041] Excessive protein structures and the recent success of deep learning algorithms provide an opportunity to develop tools for automatically extracting task-specific representations of protein structures. Therefore, there is an opportunity to predict mutant pathogenicity using a multi-channel voxelized representation of the 3D protein structure as input to a deep neural network.
Brief Description of the Drawings
[0042] In the drawings, like reference characters generally refer to like parts throughout different views. Also, the drawings are not necessarily to scale; instead, emphasis is placed on illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
Figure 19
Figure 20
Figure 21
Figure 22
Figure 23
Figure 24
Figure 25
Figure 26
Figure 27
Figure 28
Figure 29
Figure 30
Figure 31
Figure 32A
Figure 32B
Figure 33
Figure 34
Figure 35A
Figure 35B
Figure 36
DETAILED DESCRIPTION
[0043] The following discussion is presented to enable one skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Accordingly, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0044] The detailed description of various embodiments can be better understood when read in conjunction with the accompanying drawings. To the extent that the figures show diagrams of functional blocks of various embodiments, the functional blocks do not necessarily represent a division between hardware circuits. Thus, for example, one or more of the functional blocks (e.g., modules, processors, or memories) may be implemented in a single piece of hardware (e.g., a block of a general-purpose signal processor or random access memory, a hard disk, etc.) or multiple pieces of hardware. Similarly, the program may be a stand-alone program, may be incorporated as a subroutine within an operating system, may be a function within an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and means shown in the drawings.
[0045] The processing engine and database of the figure designated as modules can be implemented in hardware or software and need not be divided exactly into the same blocks as shown in the figure. Some modules may be implemented on different processors, computers, or servers, or may span multiple different processors, computers, or servers. Additionally, it will be understood that some portions of the modules may be operated in parallel with or in a different order than those shown in the figure without affecting the functions to be achieved. The modules of the figure can also be considered as flowchart steps in a method. Also, the modules do not necessarily have to have all the code arranged adjacent to each other in memory. Some portions of the code may be separated from other portions of the code with code from other modules or other functions arranged in between.
[0046] Pathogenicity determination based on protein structure Figure 1 is a flow diagram showing process 100 of a system for determining the pathogenicity of a variant. At step 102, sequence accessor 104 of the system accesses a reference amino acid sequence and an alternative amino acid sequence. At 112, 3D structure generator 114 of the system generates a 3D protein structure of the reference amino acid sequence. In some embodiments, the 3D protein structure is a homology model of a human protein. In one embodiment, a so-called SwissModel homology modeling pipeline provides a public repository of predicted human protein structures. In another embodiment, so-called HHpred homology modeling uses a tool called Modeller to predict the structure of a target protein from a template structure.
[0047] A protein is represented by a set of atoms and their coordinates in 3D space. An amino acid can have various atoms such as carbon atoms, oxygen (O) atoms, nitrogen (N) atoms, and hydrogen (H) atoms. Atoms can be further classified as side chain atoms and backbone atoms. Backbone carbon atoms can include alpha carbon (C α ) atoms and beta carbon (C β ) atoms.
[0048] At step 122, coordinate classifier 124 of the system classifies the 3D atomic coordinates of the 3D protein structure on an amino acid basis. In one embodiment, the classification of amino acid units includes assigning 3D atomic coordinates to 21 amino acid categories (including a stop or gap amino acid category). In one example, the classification of amino acid units of alpha carbon atoms can list alpha carbon atoms respectively under each of the 21 amino acid categories. In another example, the classification of amino acid units of beta carbon atoms can list beta carbon atoms respectively under each of the 21 amino acid categories.
[0049] In yet another example, the classification of amino acid units with oxygen atoms can list the oxygen atoms respectively under each of the 21 amino acid categories. In yet another example, the classification of amino acid units with nitrogen atoms can list the nitrogen atoms respectively under each of the 21 amino acid categories. In yet another example, the classification of amino acid units with hydrogen atoms can list the hydrogen atoms respectively under each of the 21 amino acid categories.
[0050] One of ordinary skill in the art will understand that in various embodiments, the classification of amino acid units can include a subset of the 21 amino acid categories and a subset of different atomic elements.
[0051] In step 132, the voxel grid generator 134 of the system instantiates a voxel grid. The voxel grid can have any resolution, for example, 3×3×3, 5×5×5, 7×7×7, etc. The voxels in the voxel grid can be of any size, for example, 1 angstrom (Å) per side, 2 Å per side, 3 Å per side, etc. One of ordinary skill in the art will understand that since the voxels are cubes, these exemplary dimensions refer to cube dimensions. Also, one of ordinary skill in the art will understand that these exemplary dimensions are non-limiting and that the voxels can have any cube dimensions.
[0052] In step 142, the voxel grid centering unit 144 of the system centers the voxel grid on a reference amino acid experiencing a target variant at the amino acid level. In one embodiment, the voxel grid is centered on the atomic coordinates of a specific atom of the reference amino acid experiencing the target variant, for example, the 3D atomic coordinates of the alpha carbon atom of the reference amino acid experiencing the target variant.
[0053] Distance channel Voxels within the voxel grid can have multiple channels (or features). In one embodiment, voxels within the voxel grid have multiple distance channels (e.g., 21 distance channels for each of 21 amino acid categories (including stop or gap amino acid categories)). In step 152, the distance channel generator 154 of the system generates distance channels in amino acid units for voxels within the voxel grid. The distance channels are generated independently for each of the 21 amino acid categories.
[0054] For example, consider the alanine (A) amino acid category. Further, for example, consider that the voxel grid is 3×3×3 in size and has 27 voxels. Then, in one embodiment, the alanine distance channel includes 27 distance values for each of the 27 voxels within the voxel grid. The 27 distance values in the alanine distance channel are measured from the center of each of the 27 voxels in the voxel grid to the nearest atom in the alanine amino acid category.
[0055] In one example, the alanine amino acid category includes only alpha carbon atoms, and thus the nearest atoms are the alanine alpha carbon atoms closest to each of the 27 voxels in the voxel grid. In another example, the alanine amino acid category includes only beta carbon atoms, and thus the nearest atoms are the alanine beta carbon atoms closest to each of the 27 voxels in the voxel grid.
[0056] In yet another example, the alanine amino acid category contains only oxygen atoms, and thus, the nearest neighbor atoms are the alanine oxygen atoms that are closest to each of the 27 voxels within the voxel grid. In yet another example, the alanine amino acid category contains only nitrogen atoms, and thus, the nearest neighbor atoms are the alanine nitrogen atoms that are closest to each of the 27 voxels within the voxel grid. In yet another example, the alanine amino acid category contains only hydrogen atoms, and thus, the nearest neighbor atoms are the alanine hydrogen atoms that are closest to each of the 27 voxels within the voxel grid.
[0057] Similar to the alanine distance channel, the distance channel generator 154 generates a distance channel (i.e., a set of distance values per voxel) for each of the remaining amino acid categories. In other embodiments, the distance channel generator 154 generates distance channels only for a subset of the 21 amino acid categories.
[0058] In other embodiments, the selection of the nearest neighbor atoms is not limited to a particular atom type. That is, within the target amino acid category, the nearest neighbor atoms to a particular voxel are selected regardless of the atomic element of the nearest neighbor atoms, and the distance value of the particular voxel is calculated for inclusion in the distance channel of the target amino acid category.
[0059] In yet other embodiments, the distance channels are generated on an atomic element basis. Instead of or in addition to having distance channels for amino acid categories, distance values can be generated for atomic element categories regardless of the amino acid to which the atoms belong. For example, consider that the atoms of an amino acid in a reference amino acid sequence span seven atomic elements, carbon, oxygen, nitrogen, hydrogen, calcium, iodine, and sulfur. Then, the voxels within the voxel grid are configured to have seven distance channels, such that each of the seven distance channels has a distance value in 27 voxel units that specifies the distance to the nearest neighbor atoms only within the corresponding atomic element category. In other embodiments, distance channels can be generated for only a subset of the seven atomic elements. In yet other embodiments, the atomic element category and distance channel generation can be further hierarchically structured into variant forms of the same atomic element, e.g., alpha carbon (C α ) atoms and beta carbon (C β ) atoms.
[0060] In yet other embodiments, the distance channels can be generated based on atom type, e.g., a distance channel for only side chain atoms and a distance channel for only backbone atoms.
[0061] The nearest neighbor atoms can be searched for within a predetermined maximum scanning radius (e.g., 6 angstroms (Å)) from the voxel center. Also, multiple atoms can be closest to the same voxel within the voxel grid.
[0062] The distance is calculated between the 3D coordinates of the voxel center and the 3D atomic coordinates of the atom. Also, the distance channels are generated using a voxel grid centered at the same location (e.g., centering on the 3D atomic coordinates of the alpha carbon atom of the reference amino acid experiencing the target variant).
[0063] The distance can be the Euclidean distance. Also, the distance can be parameterized by atomic size (or the influence of the atom) (e.g., by using the Lennard-Jones potential and / or the van der Waals atomic radius of the atom in question). Also, the distance value can be normalized by the maximum scanning radius or by the maximum observed distance value of the farthest nearest neighbor atom within the target amino acid category or the target atomic element category or the target atomic type category. In some embodiments, the distance between a voxel and an atom is calculated based on the polar coordinates of the voxel and the atom. The polar coordinates are parameterized by the angle between the voxel and the atom. In one embodiment, this angle information is used to generate an angle channel of the voxel (i.e., independent of the distance channel). In some embodiments, the angle between the nearest neighbor atom and an adjacent atom (e.g., a backbone atom) can be used as a feature encoded using the voxel.
[0064] Reference allele and alternative allele channels The voxels within the voxel grid can also have reference allele and alternative allele channels. In step 162, the one-hot encoder 164 of the system generates a reference one-hot encoding of a reference amino acid in a reference amino acid sequence and an alternative one-hot encoding of an alternative amino acid in an alternative amino acid sequence. The reference amino acid experiences the target variant. The alternative amino acid is the target variant. The reference amino acid and the alternative amino acid are located at the same position in the reference amino acid sequence and the alternative amino acid sequence, respectively. The reference amino acid sequence and the alternative amino acid sequence have the same amino acid composition per position unit, except for one exception. The exception is the position having the reference amino acid in the reference amino acid sequence and the alternative amino acid in the alternative amino acid sequence.
[0065] In step 172, the system's linker 174 links the amino acid unit distance channel with the reference and alternative one-hot encodings. In another embodiment, the linker 174 links the atomic element unit distance channel with the reference one-hot encoding and the alternative one-hot encoding. In yet another embodiment, the linker 174 links the atomic type unit distance channel with the reference one-hot encoding and the alternative one-hot encoding.
[0066] In step 182, the system's runtime logic 184 processes the linked amino acid unit / atomic element unit / atomic type unit distance channel along with the reference and alternative one-hot encodings through a pathogenicity classifier (pathogenicity determination engine) to determine the pathogenicity of the target variant, which is then presumed as the pathogenicity determination of the nucleotide variant that forms the basis for generating the target variant at the amino acid level. The pathogenicity classifier is trained using a labeled dataset of benign and pathogenic variants, for example, using an error backpropagation algorithm. Further details regarding the labeled dataset of benign and pathogenic variants, as well as an exemplary architecture and training of the pathogenicity classifier, can be found in U.S. Patent Applications Nos. 16 / 160,903, 16 / 160986, 16 / 160968, and 16 / 407149 related to the present application.
[0067] Figure 2 schematically shows the reference amino acid sequence 202 of protein 200 and the alternative amino acid sequence 212 of protein 200. Protein 200 contains N amino acids. The positions of the amino acids in protein 200 are labeled 1, 2, 3... N. In the illustrated example, position 16 is the position that experiences amino acid variant 214 (mutation) caused by the underlying nucleotide variant. For example, for the reference amino acid sequence 202, position 1 has the reference amino acid phenylalanine (F), position 16 has the reference amino acid glycine (G) 204, and position N (for example, the last amino acid of sequence 202) has the reference amino acid leucine (L). Although not illustrated for clarity, the remaining positions in the reference amino acid sequence 202 contain various amino acids in an order specific to protein 200. The alternative amino acid sequence 212 is the same as the reference amino acid sequence 202 except for the variant 214 at position 16 that contains the alternative amino acid alanine (A) 214 instead of the reference amino acid glycine (G) 204.
[0068] Figure 3 shows the classification of the amino acid units of the atoms of the amino acids in the reference amino acid sequence 202, also referred to herein as "atomic classification 300". Certain types of amino acids among the 20 natural amino acids listed in column 302 can repeat in a protein. That is, a particular type of amino acid can be present more than once in a protein. A protein can also have several undetermined amino acids classified by a 21st stop or gap amino acid category. The right column in Figure 3 contains the count of alpha carbon (C α ) atoms from different amino acids.
[0069] Specifically, Figure 3 shows the alpha carbon (C α)Shows the classification of amino acid units of atoms. Column 308 in FIG. 3 enumerates the total number of alpha carbon atoms observed for the reference amino acid sequence 202 in each of the 21 amino acid categories. For example, column 308 enumerates the 11 alpha carbon atoms observed for the alanine (A) amino acid category. Since each amino acid has only one alpha carbon atom, this means that alanine occurs 11 times in the reference amino acid sequence 202. In another example, arginine (R) appears 35 times in the reference amino acid sequence 202. The total number of alpha carbon atoms across the 21 amino acid categories is 828.
[0070] FIG. 4 shows the assignment of amino acid units of the 3D atomic coordinates of the alpha carbon atoms of the reference amino acid sequence 202 based on the atomic classification 300 of FIG. 3. This is referred to herein as "atomic coordinate bucketing 400". In FIG. 4, lists 404 to 440 represent the 3D atomic coordinates of the alpha carbon atoms bucketed into each of the 21 amino acid categories in tabular form.
[0071] In the illustrated embodiment, the binning 400 of FIG. 4 follows the classification 300 of FIG. 3. For example, in FIG. 3, the alanine amino acid category has 11 alpha carbon atoms, and thus, in FIG. 4, the alanine amino acid category has the 11 3D atomic coordinates of the corresponding 11 alpha carbon atoms from FIG. 3. The logic from this classification to binning flows from FIG. 3 to FIG. 4 for other amino acid categories as well. However, this logic from classification to binning is for illustrative purposes only, and in other embodiments, the disclosed technology need not perform the classification 300 and binning 400 to locate the nearest neighbor atoms in voxel units and can perform fewer steps, additional steps, or different steps. For example, in some embodiments, the disclosed technology can locate the nearest neighbor atoms in voxel units by using sorting and search algorithms that respond to a search query configured to accept query parameters such as sorting criteria (e.g., amino acid unit, atomic element unit, atomic type unit), a predetermined maximum search radius, and type of distance (e.g., Euclidean, Mahalanobis, normalized, non-normalized) to return the nearest neighbor atoms in voxel units from one or more databases. In various embodiments of the disclosed technology, a plurality of sorting and search algorithms from current or future technological fields can similarly be used by those skilled in the art to locate the nearest neighbor atoms in voxel units.
[0072] In FIG. 4, the 3D atomic coordinates are represented by Cartesian coordinates x, y, z, but any type of coordinate system such as spherical or cylindrical coordinates may be used and the claimed subject matter is not limited in this regard. In some embodiments, one or more databases may contain information regarding the 3D atomic coordinates of alpha carbon atoms and other atoms of amino acids in a protein. Such databases may be searchable by a particular protein.
[0073] As described above, voxels and voxel grids are 3D entities. However, for clarity, the drawings show voxels and voxel grids in a two-dimensional (2D) format, and the description considers them as such. For example, a 3×3×3 voxel grid of 27 voxels is shown and described herein as a 3×3 2D pixel grid having 9 2D pixels. One of ordinary skill in the art will understand that the 2D format is used only for illustrative purposes and is intended to cover the 3D counterparts (i.e., 2D pixels represent 3D voxels and 2D pixel grids represent 3D voxel grids). Also, the drawings are not to scale. For example, a voxel of size 2 angstroms (Å) is depicted using a single pixel.
[0074] Distance Calculation in Voxel Units FIG. 5 schematically shows a process for determining distance values in voxel units, also referred to herein as "distance calculation in voxel units 500". In the illustrated example, the distance values in voxel units are calculated only for the alanine (A) distance channel. However, as described above with respect to FIG. 1, the same distance calculation logic can be executed for each of the 21 amino acid categories to generate 21 amino acid unit distance channels and further extended to other atom types such as beta carbon atoms and other atomic elements such as oxygen, nitrogen, and hydrogen. In some embodiments, the atoms are randomly rotated prior to distance calculation to make the learning of the pathogenicity classifier invariant to atomic orientation.
[0075] In FIG. 5, the voxel grid 522 has nine voxels 514 identified by the indices (1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), and (3,3). The voxel grid 522 is centered, for example, at the 3D atomic coordinates 532 of the alpha carbon atom of the glycine (G) amino acid at position 16 of the reference amino acid sequence 202, but as described above with respect to FIG. 2, in the alternative amino acid sequence 212, position 16 experiences a mutant that mutates the glycine (G) amino acid to an alanine (A) amino acid. Also, the center of the voxel grid 522 coincides with the center of the voxel (2,2).
[0076] The central voxel grid 522 is used for voxel - unit distance calculations for each of the 21 - amino - acid - unit distance channels. For example, starting from the alanine (A) distance channel, the distance between the 3D coordinates of the center of each of the nine voxels 514 and the 3D atomic coordinates 402 of the 11 alanine alpha - carbon atoms is measured to locate the nearest - neighbor alanine alpha - carbon atom for each of the nine voxels 514. Then, nine distance values for the nine distances between the nine voxels 514 and their respective nearest - neighbor alanine alpha - carbon atoms are used to construct the alanine distance channel. The resulting alanine distance channel arranges nine alanine distance values in the same order as the nine voxels 514 within the voxel grid 522.
[0077] The above process is performed for each of the 21 amino acid categories. For example, the central 3D coordinates of each of the nine voxels 514 and the 3D atomic coordinates 404 of the 35 arginine alpha carbon atoms are used to calculate the arginine (R) distance channel in a similar manner using the central voxel grid 522 to identify the position of the nearest arginine alpha carbon atom for each of the nine voxels 514 by measuring the distance between them. Then, the nine distance values for the nine distances between the nine voxels 514 and their respective nearest arginine alpha carbon atoms are used to construct the arginine distance channel. The resulting arginine distance channel arranges the nine arginine distance values in the same order as the nine voxels 514 within the voxel grid 522. The distance channels for the 21 amino acid units are encoded in voxel units to form a distance channel tensor.
[0078] Specifically, in the illustrated example, the distance 512 is between the center of the voxel (1,1) of the voxel grid 522 and the nearest alpha carbon (C A5 ) atom, which is a Cα α atom in the list 402. Thus, the value assigned to the voxel (1,1) is the distance 512. In another example, the Cα A4 atom is the nearest C α atom to the center of the voxel (1,2). Thus, the value assigned to the voxel (1,2) is the distance between the center of the voxel (1,2) and the Cα A4 atom. In yet another example, the Cα A6 atom is the nearest C α atom to the center of the voxel (2,1). Thus, the value assigned to the voxel (2,1) is the distance between the center of the voxel (2,1) and the Cα A6 atom. In yet another example, the Cα A6 atom is also the nearest C α atom to the centers of the voxels (3,2) and (3,3). Thus, the value assigned to the voxel (3,2) is the distance between the center of the voxel (3,2) and the Cα A6 atom, and the value assigned to the voxel (3,3) is the distance between the center of the voxel (3,3) and the Cα A6It is the distance between atoms. In some embodiments, the distance value assigned to voxel 514 can be a normalized distance. For example, the distance value assigned to voxel (1,1) may be the distance 512 divided by the maximum distance 502 (a predetermined maximum scan radius). In some embodiments, the nearest neighbor atom distance may be the Euclidean distance, and the nearest neighbor atom distance may be normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance (e.g., the maximum distance 502, etc.).
[0079] As described above, for an amino acid having an alpha carbon atom, the distance can be the nearest neighbor alpha carbon atom distance from the corresponding voxel center to the nearest neighbor alpha carbon atom of the corresponding amino acid. Further, for an amino acid having a beta carbon atom, the distance can be the nearest neighbor beta carbon atom distance from the corresponding voxel center to the nearest neighbor beta carbon atom of the corresponding amino acid. Similarly, for an amino acid having a backbone atom, the distance can be the nearest neighbor backbone atom distance from the corresponding voxel center to the nearest neighbor backbone atom of the corresponding amino acid. Similarly, for an amino acid having a side chain atom, the distance can be the nearest neighbor side chain atom distance from the corresponding voxel center to the nearest neighbor side chain atom of the corresponding amino acid. In some embodiments, the distance can additionally / alternatively include distances to the second, third, fourth nearest atoms, etc.
[0080] Distance channel of amino acid unit FIG. 6 shows an example of a distance channel 600 of 21 amino acid units. Each column in FIG. 6 corresponds to one of the distance channels 602-642 of 21 amino acid units. The distance channel for each amino acid unit includes distance values for each of the voxels 514 of the voxel grid 522. For example, the distance channel 602 for alanine (A) includes distance values for each of the voxels 514 of the voxel grid 522. As described above, the voxel grid 522 is a 3D grid of volume 3×3×3 and includes 27 voxels. Similarly, FIG. 6 shows the voxels 514 in 2D (e.g., 9 voxels of a 3×3 grid), but the distance channel for each amino acid unit can include distance values for 27 voxel units for the 3×3×3 voxel grid.
[0081] Directional encoding In some embodiments, the disclosed techniques use directional parameters to identify the directionality of reference amino acids within a reference amino acid sequence 202. In some embodiments, the disclosed techniques use directional parameters to identify the directionality of alternative amino acids within an alternative amino acid sequence 212. In some embodiments, the disclosed techniques use directional parameters to identify positions within a protein 200 that experience a target variant at the amino acid level.
[0082] As described above, all distance values within the distance channels 602-642 of 21 amino acid units are measured from their respective nearest neighbor atoms to the voxels 514 within the voxel grid 522. These nearest neighbor atoms are derived from one of the reference amino acids in the reference amino acid sequence 202. These original reference amino acids containing the nearest neighbor atoms can be classified into two categories: (1) the original reference amino acids preceding the reference amino acid 204 that experiences a variant in the reference amino acid sequence 202 and (2) the original reference amino acids following the variant-experiencing reference amino acid 204 in the reference amino acid sequence 202. The original reference amino acids in the first category can be called preceding reference amino acids. The original reference amino acids in the second category can be called subsequent reference amino acids.
[0083] The directionality parameter is applied to the distance values in the 21 amino acid unit distance channels 602 - 642 measured from the nearest neighbor atoms derived from the reference amino acid. In one embodiment, the directionality parameter is multiplied by such distance values. The directionality parameter can be any number, such as -1.
[0084] As a result of the application of the directionality parameter, the 21 amino acid unit distance channel 600 contains several distance values that indicate to the pathogenicity classifier which end of the protein 200 is the start end and which end is the end end. This also enables the pathogenicity classifier to reconstruct the protein sequence from the 3D protein structure information provided by the distance channel as well as the reference channel and the allele channel.
[0085] Distance channel tensor FIG. 7 is a schematic diagram of the distance channel tensor 700. The distance channel tensor 700 is a voxelized representation of the amino acid unit distance channel 600 from FIG. 6. In the distance channel tensor 700, the 21 amino acid unit distance channels 602 - 642 are concatenated in voxel units like the RGB channels of a color image. The voxelized dimensionality of the distance channel tensor 700 is 21×3×3×3 (where 21 indicates 21 amino acid categories and 3×3×3 indicates a 3D voxel grid with 27 voxels). However, FIG. 7 is a 2D depiction with dimensions 21×3×3.
[0086] One - hot encoding FIG. 8 shows the one - hot encoding 800 of the reference amino acid 204 and the alternative amino acid 214. In FIG. 8, the left column is the one - hot encoding 802 of the reference amino acid glycine (G) 204, where 1 is for the glycine amino acid category and 0 is for all other amino acid categories. In FIG. 8, the right column is the one - hot encoding 804 of the mutant / alternative amino acid alanine (A) 214, where 1 is for the alanine amino acid category and 0 is for all other amino acid categories.
[0087] Figure 9 is a schematic diagram of the voxelized one-hot encoded reference amino acid 902 and the voxelized one-hot encoded mutant / alternative amino acid 912. The voxelized one-hot encoded reference amino acid 902 is a voxelized representation of the one-hot encoding 802 of the reference amino acid glycine (G) 204 from Figure 8. The voxelized one-hot encoded alternative amino acid 912 is a voxelized representation of the one-hot encoding 804 of the mutant / alternative amino acid alanine (A) 214 from Figure 8. The voxelized dimensionality of the voxelized one-hot encoded reference amino acid 902 is 21×1×1×1 (where 21 indicates 21 amino acid categories). However, Figure 9 is a 2D depiction with dimensionality 21×1×1. Similarly, the voxelized dimensionality of the voxelized one-hot encoded alternative amino acid 912 is 21×1×1×1 (where 21 indicates 21 amino acid categories). However, Figure 9 is a 2D depiction with dimensionality 21×1×1.
[0088] Reference allele tensor Figure 10 schematically shows a concatenation process 1000 that concatenates the distance channel tensor 700 of Figure 7 and the reference allele tensor 1004 on a voxel-by-voxel basis. The reference allele tensor 1004 is a set (repetition / cloning / copying) of voxel units of the voxelized one-hot encoded reference amino acid 902 from Figure 9. That is, multiple copies of the voxelized one-hot encoded reference amino acid 902 are concatenated with each other on a voxel-by-voxel basis according to the spatial arrangement of the voxels 514 within the voxel grid 522 such that the reference allele tensor 1004 has corresponding copies of the voxelized one-hot encoded reference amino acid 910 for each of the voxels 514 within the voxel grid 522.
[0089] The linking process 1000 generates a linking tensor 1010. The voxelized dimensionality of the reference allele tensor 1004 is 21×3×3×3 (where 21 represents 21 amino acid categories and 3×3×3 represents a 3D voxel grid with 27 voxels). However, FIG. 10 is a 2D depiction of the reference allele tensor 1004 having a dimensionality of 21×3×3. The voxelized dimensionality of the linking tensor 1010 is 42×3×3×3. However, FIG. 10 is a 2D depiction of the linking tensor 1010 having a dimensionality of 42×3×3.
[0090] Alternative allele tensor FIG. 11 schematically shows a linking process 1100 that links the distance channel tensor 700 of FIG. 7, the reference allele tensor 1004 of FIG. 10, and the alternative allele tensor 1104 on a voxel-by-voxel basis. The alternative allele tensor 1104 is a set (repetition / cloning / copying) of voxelized one-hot encoded alternative amino acids 912 from FIG. 9. That is, multiple copies of the voxelized one-hot encoded alternative amino acids 912 are linked to each other on a voxel-by-voxel basis according to the spatial arrangement of the voxels 514 within the voxel grid 522 such that the alternative allele tensor 1104 has corresponding copies of the voxelized one-hot encoded alternative amino acids 910 for each of the voxels 514 within the voxel grid 522.
[0091] The linking process 1100 generates a linking tensor 1110. The voxelized dimensionality of the alternative allele tensor 1104 is 21×3×3×3 (where 21 represents 21 amino acid categories and 3×3×3 represents a 3D voxel grid with 27 voxels). However, FIG. 11 is a 2D depiction of the alternative allele tensor 1104 having a dimensionality of 21×3×3. The voxelized dimensionality of the linking tensor 1110 is 63×3×3×3. However, FIG. 11 is a 2D depiction of the linking tensor 1110 having a dimensionality of 63×3×3.
[0092] In some embodiments, the runtime logic 184 processes the concatenated tensor 1110 via a pathogenicity classifier to determine the pathogenicity of the variant / alternative amino acid alanine (A) 214, which is then inferred as the pathogenicity determination of the nucleotide variant underlying the generation of the variant / alternative amino acid alanine (A) 214.
[0093] Evolutionarily conserved channel Predicting the functional consequences of variants depends, at least in part, on the assumption that amino acids important to the protein family are conserved through evolution by negative selection (i.e., amino acid changes at these sites have been detrimental in the past), and the assumption that mutations at these sites are likely to be pathogenic (disease-causing) in humans. Generally, homologous sequences of the target protein are collected and aligned, and the conservation metric is calculated based on the weighted frequencies of the different amino acids observed at the target positions in the alignment.
[0094] Accordingly, the disclosed techniques connect the distance channel tensor 700, the reference allele tensor 1004, and the alternative allele tensor 1104 to an evolutionary channel. An example of an evolutionary channel is the pan-amino acid conservation frequency. Another example of an evolutionary channel is the conservation frequency for each amino acid.
[0095] In some embodiments, the evolutionary channel is constructed using a position-specific weight matrix (PWM). In other embodiments, the evolutionary channel is constructed using a position-specific frequency matrix (PSFM). In still other embodiments, the evolutionary channel is constructed using computational tools such as SIFT, PolyPhen, and PANTHER-PSEC. In still other embodiments, the evolutionary channel is a conservation channel based on evolutionary conservation. Conservation is relevant because it also reflects the effect of negative selection that has acted to prevent evolutionary changes at a given site in a protein.
[0096] Pan-amino acid evolutionary profile FIG. 12 is a flowchart showing a process 1200 of a system for determining the pan - amino - acid conservation frequency of neighboring atoms and assigning (voxelizing) them to voxels, according to one embodiment of the disclosed technology. FIGS. 12, 13, 14, 15, 16, 17, and 18 are described in parallel.
[0097] In step 1202, a similar - sequence finder 1204 of the system searches for amino - acid sequences that are similar (homologous) to a reference amino - acid sequence 202. Similar amino - acid sequences can be selected from multiple species such as primates, mammals, and vertebrates.
[0098] In step 1212, an aligner 1214 of the system aligns the reference amino - acid sequence 202 positionally with the similar amino - acid sequences, i.e., the aligner 1214 performs a multiple - sequence alignment. FIG. 14 shows an exemplary multiple - sequence alignment 1400 of the reference amino - acid sequence 202 across 99 species. In some embodiments, the multiple - sequence alignment 1400 can be divided to generate, for example, a first position - frequency matrix 1402 for primates, a second position - frequency matrix 1412 for mammals, and a third position - frequency matrix 1422 for primates. In other embodiments, a single position - frequency matrix is generated across 99 species.
[0099] In step 1222, a pan - amino - acid conservation - frequency calculator 1224 of the system uses the multiple - sequence alignment to determine the pan - amino - acid conservation frequency of the reference amino - acids in the reference amino - acid sequence 202.
[0100] In step 1232, the system's nearest neighbor atom finder 1234 finds the nearest neighbor atoms to the voxels 514 within the voxel grid 522. In some embodiments, the search for nearest neighbor atoms per voxel may not be limited to any particular amino acid category or atom type. That is, the nearest neighbor atoms per voxel can be selected across amino acid categories and amino acid types as long as they are the atoms closest to their respective voxel centers. In other embodiments, the search for nearest neighbor atoms per voxel can be limited to only specific atomic elements such as oxygen, nitrogen, and hydrogen, or only alpha carbon atoms, or only beta carbon atoms, or only side chain atoms, or only backbone atoms, etc., i.e., only to specific atomic categories.
[0101] In step 1242, the system's amino acid selector 1244 selects the reference amino acids in the reference amino acid sequence 202 that contain the nearest neighbor atoms identified in step 1232. Such reference amino acids can be referred to as nearest neighbor reference amino acids. FIG. 13 shows an example of locating the nearest neighbor atoms 1302 to the voxels 514 within the voxel grid 522 and mapping the nearest neighbor reference amino acids 1312 that contain the nearest neighbor atoms 1302 to the voxels 514 within the voxel grid 522 respectively. This is identified as "mapping 1300 from voxel to nearest neighbor amino acid" in FIG. 13.
[0102] In step 1252, the system's voxelizer 1254 voxelizes the pan - amino acid conservation frequency of the nearest neighbor reference amino acids. FIG. 15 shows an example of determining the pan - amino acid conservation frequency array for the first voxel (1,1) in the voxel grid 522, which is also referred to herein as "voxel - by - voxel evolutionary profile determination 1500".
[0103] Referring to FIG. 13, the nearest neighbor reference amino acid mapped to the first voxel (1, 1) is the aspartic acid (D) amino acid at position 15 in the reference amino acid sequence 202. Next, a multiple sequence alignment of the reference amino acid sequence 202 with, for example, 99 homologous amino acid sequences of 99 is analyzed at position 15. Such position-specific and inter-species analysis reveals how many examples of amino acids from each of the 21 amino acid categories are found at position 15 across 100 aligned amino acid sequences (i.e., the reference amino acid sequence 202 + 99 homologous amino acid sequences).
[0104] In the example shown in FIG. 15, the aspartic acid (D) amino acid is found at position 15 in 96 out of 100 aligned amino acid sequences. Thus, a pan-amino acid conservation frequency of 0.96 is assigned to the aspartic acid amino acid category 1504. Similarly, in the example shown, the valine (V) amino acid is found at position 15 in 4 out of 100 aligned amino acid sequences. Thus, a pan-amino acid conservation frequency of 0.04 is assigned to the valine amino acid category 1514. Since examples of amino acids from other amino acid categories are not detected at position 15, a pan-amino acid conservation frequency of 0 is assigned to the remaining amino acid categories. In this way, each of the 21 amino acid categories is assigned a respective pan-amino acid conservation frequency, which can be encoded in the pan-amino acid conservation frequency array 1502 for the first voxel (1, 1).
[0105] FIG. 16 shows the respective pan-amino acid conservation frequencies 1612 - 1692 determined for each of the voxels 514 within the voxel grid 522 using the position frequency logic described in FIG. 15, which is also referred to herein as "mapping from voxels to evolutionary profiles 1600".
[0106] Next, the per-voxel evolutionary profile 1602 is used by the voxelizer 1254 to generate the voxelized per-voxel evolutionary profile 1700 shown in FIG. 17. In many cases, each of the voxels 514 within the voxel grid 522 has a different pan-amino acid conservation frequency array, and thus, since the voxels are regularly mapped to different nearest neighbor atoms and thus different nearest neighbor reference amino acids, each voxel has a different voxelized evolutionary profile. Of course, if two or more voxels have the same nearest neighbor atom and thus the same nearest neighbor reference amino acid, the same pan-amino acid conservation frequency array and the same voxelized evolutionary profile per voxel are assigned to each of the two or more voxels.
[0107] FIG. 18 shows an example of an evolutionary profile tensor 1800 in which the voxelized per-voxel evolutionary profiles 1700 are connected to each other in voxel units according to the spatial arrangement of the voxels 514 within the voxel grid 522. The voxelized dimensionality of the evolutionary profile tensor 1800 is 21×3×3×3 (where 21 represents 21 amino acid categories and 3×3×3 represents a 3D voxel grid having 27 voxels). However, FIG. 18 is a 2D depiction of the evolutionary profile tensor 1800 having dimensions 21×3×3.
[0108] In step 1262, the concatenator 174 concatenates the evolutionary profile tensor 1800 with the distance channel tensor 700 in voxel units. In some embodiments, the evolutionary profile tensor 1800 is concatenated with the concatenation tensor 1110 in voxel units to generate a further concatenation tensor (not shown) having dimensions 84×3×3×3.
[0109] In step 1272, the runtime logic 184 processes a further concatenation tensor having dimensions 84×3×3×3 via a pathogenicity classifier to determine the pathogenicity of the target variant, which is then presumed to be the pathogenicity determination of the nucleotide variant that underlies the generation of the target variant at the amino acid level.
[0110] Evolutionary profile for each amino acid FIG. 19 is a flowchart showing a process 1900 of a system for determining the conservation frequency for each amino acid of the nearest neighbor atoms and assigning (voxelizing) them to voxels. In FIG. 19, steps 1202 and 1212 are the same as those in FIG. 12.
[0111] In step 1922, the per - amino - acid conservation frequency calculator 1924 of the system uses the multiple sequence alignment to determine the per - amino - acid conservation frequency of the reference amino acids in the reference amino acid sequence 202.
[0112] In step 1932, the nearest neighbor atom finder 1934 of the system finds 21 nearest neighbor atoms for each of the 21 amino acid categories for each voxel 514 within the voxel grid 522. Each of the 21 nearest neighbor atoms is different from each other since they are selected from different amino acid categories. This leads to the selection of 21 unique nearest neighbor reference amino acids for a particular voxel, which in turn leads to the generation of 21 unique position frequency matrices for a particular voxel, which in turn leads to the determination of 21 unique per - amino - acid conservation frequencies for a particular voxel.
[0113] In step 1942, the amino acid selector 1944 of the system selects 21 reference amino acids within the reference amino acid sequence 202 that contain the 21 nearest neighbor atoms identified in step 1932 for each of the voxels 514 within the voxel grid 522. Such reference amino acids can be referred to as nearest neighbor reference amino acids.
[0114] In step 1952, the voxelizer 1954 of the system voxelizes the pen - amino acid conservation frequencies of the 21 nearest - neighbor reference amino acids identified for a particular voxel in step 1942. Since the 21 nearest - neighbor reference amino acids correspond to different underlying nearest - neighbor atoms, they are necessarily located at 21 different positions in the reference amino acid sequence 202. Thus, for a particular voxel, 21 position - frequency matrices can be generated for the 21 nearest - neighbor reference amino acids. The 21 position - frequency matrices can be generated across multiple species whose homologous amino acid sequences are positionally aligned with the reference amino acid sequence 202 as described above with respect to FIGS. 12 - 15.
[0115] Next, using the 21 position - frequency matrices, 21 position - specific conservation scores can be calculated for the 21 nearest - neighbor reference amino acids identified for a particular voxel. These 21 position - specific conservation scores form the pen - amino acid conservation frequencies for the particular voxel, similar to the pan - amino acid conservation frequency array 1502 of FIG. 12. However, since the 21 nearest - neighbor reference amino acids across 21 amino - acid categories result in different position - frequency matrices, and thus necessarily have different positions for different amino - acid - specific conservation frequencies, while array 1502 has many 0 entries, each element (feature) in the amino - acid - specific conservation frequency array has a value (e.g., a floating - point number).
[0116] The above process is performed for each of the voxels 514 within the voxel grid 522, and the resulting amino - acid - specific conservation frequencies per voxel are voxelized, tensored, concatenated, and processed for pathogenicity determination, similar to the pan - amino acid conservation frequencies described with respect to FIGS. 12 - 18.
[0117] Annotation channel Figure 20 shows various examples of the voxelized annotation channel 2000 that is concatenated with the distance channel tensor 700. In some embodiments, the voxelized annotation channel is a one-hot indicator of different protein annotations, e.g., whether an amino acid (residue) is part of a transmembrane region, signal peptide, active site, or any other binding site, or whether the residue undergoes a post-translational modification PathRatio (see Pei P, Zhang A: A Topological Measurement for Weighted Protein Interaction Network. CSB 2005,268-278). Additional examples of annotation channels can be found in the section on specific embodiments and claims below.
[0118] The voxelized annotation channel is arranged in voxel units such that voxels can have the same annotation sequence, such as voxelized reference and alternative allele sequences (e.g., annotation channels 2002, 2004, 2006), or such that voxels can have respective annotation sequences, such as the voxelized per-voxel evolutionary profile 1700 (e.g., annotation channels 2012, 2014, 2016 (as indicated by different colors)).
[0119] The annotation channel is voxelized, tensored, concatenated, and processed for pathogenicity determination, similar to the pan-amino acid conservation frequency described with respect to FIGS. 12 - 18.
[0120] Structure reliability channel The disclosed technology can also connect various voxelized structural reliability channels to the distance channel tensor 700. Some examples of structural reliability channels include the GMQE score (provided by SwissModel); the B factor; the temperature factor column of the homology model (indicating how well the residues satisfy the (physical) constraints in the protein structure); the normalized number of alignment template proteins for the residues closest to the center of the voxel (alignment provided by HHpred, for example, the voxel is closest to the residues where 3 out of 6 template structures align, meaning the feature has a value of 3 / 6 = 0.5); the minimum, maximum, and average TM scores; and the predicted TM score of the template protein structure that aligns with the residues closest to the voxel (continuing the above example, assuming the 3 template structures have TM scores of 0.5, 0.5, and 1.5, the minimum is 0.5, the average is 2 / 3, and the maximum is 1.5). The TM score can be provided by HHpred for each protein template. Additional examples of structural reliability channels can be found in the following specific embodiments section and claims.
[0121] The voxelized structural reliability channels are arranged on a voxel-by-voxel basis such that the voxels can have the same structural reliability sequence, such as voxelized reference alleles and alternative allele sequences, or such that the voxels can have their own structural reliability sequences, such as the voxelized per-voxel evolutionary profile 1700.
[0122] The structural reliability channels are voxelized, tensorized, concatenated, and processed for pathogenicity determination, similar to the general amino acid conservation frequency described with respect to FIGS. 12 - 18.
[0123] Pathogenicity classifier FIG. 21 shows different combinations and permutations of input channels that can be provided as input 2102 to a pathogenicity classifier 2108 for pathogenicity determination 2106 of a target variant. One of the inputs 2102 can be a distance channel 2104 generated by a distance channel generator 2272. FIG. 22 shows various methods for calculating the distance channel 2104. In one embodiment, the distance channel 2104 is generated based on the distances 2202 between voxel centers and atoms across a plurality of atomic elements, regardless of the amino acids. In some embodiments, the distances 2202 are normalized by a maximum scanning radius to generate normalized distances 2202a. In another embodiment, the distance channel 2104 is generated based on the distances 2212 between voxel centers and alpha carbon atoms on an amino acid basis. In some embodiments, the distances 2212 are normalized by a maximum scanning radius to generate normalized distances 2212a. In yet another embodiment, the distance channel 2104 is generated based on the distances 2222 between voxel centers and beta carbon atoms on an amino acid basis. In some embodiments, the distances 2222 are normalized by a maximum scanning radius to generate normalized distances 2222a. In yet another embodiment, the distance channel 2104 is generated based on the distances 2232 between voxel centers and side chain atoms on an amino acid basis. In some embodiments, the distances 2232 are normalized by a maximum scanning radius to generate normalized distances 2232a. In yet another embodiment, the distance channel 2104 is generated based on the distances 2242 between voxel centers and backbone atoms on an amino acid basis. In some embodiments, the distances 2242 are normalized by a maximum scanning radius to generate normalized distances 2242a. In yet another embodiment, the distance channel 2104 is generated based on the distances 2252 (one feature) between voxel centers and their respective nearest neighbor atoms, regardless of atomic type and amino acid type. In yet another embodiment, the distance channel 2104 is generated based on the distances 2262 (one feature) between voxel centers and atoms from non-standard amino acids. In some embodiments, the distance between a voxel and an atom is calculated based on the polar coordinates of the voxel and the atom.The polar coordinates are parameterized by the angle between the voxel and the atom. In one embodiment, this angle information is used to generate an angle channel of the voxel (i.e., independent of the distance channel). In some embodiments, the angle between the nearest neighbor atom and the adjacent atom (e.g., the backbone atom) can be used as a feature encoded using the voxel.
[0124] Another of the inputs 2102 can be a feature 2114 indicating a missing atom within a specified radius.
[0125] Another of the inputs 2102 can be a one-hot encoding 2124 of the reference amino acid. Another of the inputs 2102 can be a one-hot encoding 2134 of the mutant / alternative amino acid.
[0126] Another of the inputs 2102 can be an evolutionary channel 2144 generated by the evolutionary profile generator 2372 shown in FIG. 23. In one embodiment, the evolutionary channel 2144 can be generated based on the pan-amino acid conservation frequency 2302. In another embodiment, the evolutionary channel 2144 can be generated based on the pan-amino acid conservation frequency 2312.
[0127] Another of the inputs 2102 can be a feature 2154 indicating a deletion residue or a deletion evolutionary profile.
[0128] Another one of the inputs 2102 can be an annotation channel 2164 generated by an annotation generator 2472 shown in FIG. 24. In one embodiment, the annotation channel 2164 can be generated based on the molecular processing annotation 2402. In another embodiment, the annotation channel 2164 can be generated based on the region annotation 2412. In yet another embodiment, the annotation channel 2164 can be generated based on the site annotation 2422. In yet another embodiment, the annotation channel 2164 can be generated based on the amino acid modification annotation 2432. In yet another embodiment, the annotation channel 2164 can be generated based on the secondary structure annotation 2442.
[0129] In yet another embodiment, the annotation channel 2164 can be generated based on the experimental information annotation 2452.
[0130] Another one of the inputs 2102 can be a structure reliability channel 2174 generated by a structure reliability generator 2572 shown in FIG. 25. In one embodiment, the structure reliability 2174 can be generated based on the global model quality estimation (GMQE) 2502. In another embodiment, the structure reliability 2174 can be generated based on the qualitative model energy analysis (QMEAN) score 2512. In yet another embodiment, the structure reliability 2174 can be generated based on the temperature factor 2522. In yet another embodiment, the structure reliability 2174 can be generated based on the template modeling score 2542. Examples of the template modeling score 2542 include a minimum template modeling score 2542a, an average template modeling score 2542b, and a maximum template modeling score 2542c.
[0131] One of ordinary skill in the art will understand that any permutation and combination of input channels can be concatenated to an input for processing through a pathogenicity classifier 2108 for pathogenicity determination 2106 of a target variant. In some embodiments, only a subset of the input channels can be concatenated. The input channels can be concatenated in any order. In one embodiment, the input channels can be concatenated into a single tensor by a tensor generator (input encoder) 2110. This single tensor can then be provided as an input to the pathogenicity classifier 2108 for pathogenicity determination 2106 of the target variant.
[0132] In one embodiment, the pathogenicity classifier 2108 uses a convolutional neural network (CNN) having a plurality of convolutional layers. In another embodiment, the pathogenicity classifier 2108 uses a recurrent neural network (RNN) such as a long short-term memory network (LSTM), a bidirectional LSTM (Bi-LSTM), and a gated recurrent unit (GRU). In yet another embodiment, the pathogenicity classifier 2108 uses both a CNN and an RNN. In yet another embodiment, the pathogenicity classifier 2108 uses a graph convolutional neural network that models dependencies in graph-structured data. In yet another embodiment, the pathogenicity classifier 2108 uses a variational autoencoder (VAE). In yet another embodiment, the pathogenicity classifier 2108 uses an adversarial generative network (GAN). In yet another embodiment, the pathogenicity classifier 2108 can be a self-attention based language model such as those implemented by, for example, transformers and BERT.
[0133] In still other embodiments, the pathogenic classifier 2108 can use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or atrous convolution, transposed convolution, depthwise separable convolution, pointwise convolution, 1x1 convolution, grouped convolution, flattened convolution, spatial and cross-channel convolution, shuffled grouped convolution, spatially separable convolution, and inverse convolution. It can use one or more loss functions such as logistic regression / log loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smooth L1 loss, and Huber loss. It can use any parallel, efficient, and compression methods such as TFRecord, compressed encoding (e.g., PNG), sharpening, parallel calls for map transformation, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). It can include upsampling layers, downsampling layers, regression connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, dilated connections, activation functions (e.g., non-linear transformation functions such as rectified linear unit (ReLU), Leaky ReLU, exponential linear unit (ELU), sigmoid, and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, attention mechanisms, and Gaussian error linear units.
[0134] The pathogenicity classifier 2108 is trained using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that can be used for the pathogenicity classifier 2108 to learn include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used for the pathogenicity classifier 2108 to learn include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad. In other embodiments, the pathogenicity classifier 2108 can be trained by unsupervised learning, semi-supervised learning, self-learning, reinforcement learning, multi-task learning, multi-modal learning, transfer learning, knowledge distillation, and the like.
[0135] FIG. 26 shows an exemplary processing architecture 2600 of the pathogenicity classifier 2108 according to one embodiment of the disclosed technology. The processing architecture 2600 includes a cascade of processing modules 2606, 2610, 2614, 2618, 2622, 2626, 2630, 2634, 2638, and 2642, each of which can include 1D convolution (1x1x1 CONV), 3D convolution (3x3x3 CONV), ReLU non-linearity, and batch normalization (BN). Other examples of processing modules include fully connected (FC) layers, dropout layers, flattening layers, and a final softmax layer that generates exponentially normalized scores for target variants belonging to the benign class and the pathogenic class. In FIG. 26, "64" indicates the number of convolution filters applied by a particular processing module. In FIG. 26, the size of the input voxel 2602 is 15×15×15×8. FIG. 26 also shows the respective volume dimensions of the intermediate inputs 2604, 2608, 2612, 2616, 2620, 2624, 2628, 2632, 2636, and 2640 generated by the processing architecture 2600.
[0136] FIG. 27 shows an exemplary processing architecture 2700 of a pathogenicity classifier 2108 according to one embodiment of the disclosed technology. The processing architecture 2700 includes a cascade of processing modules 2708, 2714, 2720, 2726, 2732, 2738, 2744, 2750, 2756, 2762, 2768, 2774, and 2780 such as 1D convolution (CONV 1D), 3D convolution (CONV 3D), ReLU non-linearity, and batch normalization (BN). Other examples of processing modules include fully connected (dense) layers, dropout layers, flattening layers, and a final softmax layer that generates exponentially normalized scores for target variants belonging to the benign and pathogenic classes. In FIG. 27, "64" and "32" indicate the number of convolution filters applied by specific processing modules. In FIG. 27, the size of the input voxels 2704 supplied by the input layer 2702 is 7×7×7×108. FIG. 27 also shows the respective volume dimensions of the intermediate inputs 2710, 2716, 2722, 2728, 2734, 2740, 2746, 2752, 2758, 2764, 2770, 2776, and 2782 generated by the processing architecture 2700, and the resulting intermediate outputs 2706, 2712, 2718, 2724, 2730, 2736, 2742, 2748, 2754, 2760, 2766, 2772, 2778, and 2784.
[0137] Those skilled in the art will understand that other current and future artificial intelligence, machine learning, and deep learning models, datasets, and techniques can be incorporated into the disclosed variant pathogenicity classifier without departing from the spirit of the disclosed technology.
[0138] Performance results as objective indicators of inventiveness and non-obviousness
[0139] The variant pathogenicity classifier disclosed herein performs pathogenicity prediction based on 3D protein structure and is referred to as "PrimateAI 3D". "Primate AI" is a previously disclosed variant pathogenicity classifier related to sharing that performs pathogenicity prediction based on protein sequences. Further details about PrimateAI can be found in co-owned U.S. Patent Applications Nos. 16 / 160,903, 16 / 160986, 16 / 160,968, and 16 / 407,149, and Sundaram, L. et al., Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161 - 1170 (2018).
[0140] Figures 28, 29, 30, and 31 use PrimateAI as a benchmark model to demonstrate the superiority of the classification of PrimateAI 3D over PrimateAI. The performance results in Figures 28, 29, 30, and 31 are generated based on a classification task that accurately distinguishes pathogenic variants from benign variants across multiple validation sets. PrimateAI 3D is trained on a training set different from the multiple validation sets. PrimateAI 3D is trained on common human variants and primate-derived variants used as a benign dataset, while variants simulated based on trinucleotide context are used as an unlabeled or pseudo-pathogenic dataset.
[0141] The new developmental delay disorder (new DDD) is an example of a validation set used to compare the classification accuracy of Primate AI 3D over Primate AI. The new DDD validation set labels variants from individuals with DDD as pathogenic and the same variants from healthy relatives of individuals with DDD as benign. A similar labeling scheme is used with the autism spectrum disorder (ASD) validation set shown in Figure 31.
[0142] BRCA1 is another example of a validation set used to compare the classification accuracy of PrimateAI 3D against PrimateAI. The BRCA1 validation set labels synthetically generated reference amino acid sequences that simulate the protein of the BRCA1 gene as benign variants, and synthetically modified allelic amino acid sequences that simulate the protein of the BRCA1 gene as pathogenic variants. Similar labeling schemes are used with different validation sets of the TP53 gene, the TP53S3 gene and its variants, and other genes and their variants shown in FIG. 31.
[0143] FIG. 28 identifies the performance of the benchmark PrimateAI model with a horizontal bar (labeled "PAI") and the performance of the disclosed PrimateAI 3D model with a horizontal bar (labeled "ens10_7x7x7x2_hhpred_evo+alt"). The horizontal bar labeled "ens10_7x7x7x2_hhpred_evo+alt_paisum" shows the pathogenicity prediction derived by combining the respective pathogenicity predictions of the disclosed PrimateAI 3D model and the benchmark PrimateAI model. In the legend, "ens10" refers to an ensemble of 10 PrimateAI 3D models, each trained on a different seed learning dataset and randomly initialized with different weights and biases. Also, "7×7×7×2" indicates the size of the voxel grid used to encode the input channels during the training of the ensemble of 10 PrimateAI 3D models. For a given variant, the ensemble of 10 PrimateAI 3D models each generate 10 pathogenicity predictions, which are then combined (e.g., by averaging) to generate the final pathogenicity prediction for the given variant. This logic also applies to ensembles of different group sizes.
[0144] Also, in Figure 28, the y-axis has different validation sets and the x-axis has p-values. The larger the p-value, that is, the longer the horizontal bar, the higher the accuracy in distinguishing benign variants from pathogenic variants. As demonstrated by the p-values in Figure 28, PrimateAI 3D performs better than PrimateAI across most of the validation sets (the only exception being the tp53s3_A549 validation set). That is, the horizontal bar of PrimateAI 3D (labeled "ens10_7x7x7x2_hhpred_evo+alt") is consistently longer than the horizontal bar of PrimateAI (labeled "PAI").
[0145] Also, in Figure 28, the "mean" category along the y-axis calculates the mean of the p-values determined for each of the validation sets. Even in the mean category, PrimateAI 3D performs better than PrimateAI.
[0146] In Figure 29, PrimateAI is represented by the horizontal bar (labeled "PAI"), an ensemble of 20 PrimateAI 3D models trained on a 3x3x3 voxel grid is represented by the horizontal bar labeled "ns20_3x3x3x2_evo+alt", an ensemble of 10 PrimateAI 3D models trained on a 7x7x7 voxel grid is represented by the horizontal bar labeled "ens10_7x7x7x2_evo+alt", an ensemble of 20 PrimateAI 3D models trained on a 7x7x7 voxel grid is represented by the horizontal bar labeled "ens20_7x7x7x2_evo+alt", and an ensemble of 20 PrimateAI 3D models trained on a 17x17x17 voxel grid is represented by the horizontal bar labeled "ens20_17x17x17x2_evo+alt".
[0147] Also, in Figure 29, the y-axis has different validation sets and the x-axis has p-values. Similar to before, the larger the p-value, i.e., the longer the horizontal bar, the higher the accuracy in distinguishing benign variants from pathogenic variants. As demonstrated by the p-values in Figure 20, different configurations of PrimateAI 3D perform better than PrimateAI across most of the validation sets. That is, the horizontal bars representing the ensemble of multiple PrimateAI 3D models (labeled "ns20_3x3x3x2_evo+alt", "ens10_7x7x7x2_evo+alt", "ens20_7x7x7x2_evo+alt", and "ens20_17x17x17x2_evo+alt") are mostly longer than the horizontal bar of PrimateAI (labeled "PAI").
[0148] Also, in Figure 29, the "Mean" category along the y-axis calculates the mean of the p-values determined for each of the validation sets. Also in the mean category, different configurations of PrimateAI 3D perform better than PrimateAI.
[0149] In Figure 30, the thick shaded vertical bars represent PrimateAI ("PrimateAI(v1)") and the thin shaded vertical bars represent PrimateAI 3D ("PrimateAI 3D"). In Figure 30, the y-axis has p-values and the x-axis has different validation sets. In Figure 30, without exception, PrimateAI 3D consistently performs better than PrimateAI across all of the validation sets. That is, the vertical bars of PrimateAI 3D are always longer than the vertical bars of PrimateAI.
[0150] Figure 31 discriminates between the performance of the benchmark PrimateAI model with vertical bars labeled "PAI-v1_plain" and the performance of the disclosed PrimateAI 3D model with vertical bars labeled "PAI-3D-orig_plain". The vertical bar labeled "PAI-3D-orig_paisum" shows the pathogenicity prediction derived by combining the respective pathogenicity predictions of the disclosed PrimateAI 3D model and the benchmark PrimateAI model. In Figure 31, the y-axis has p-values and the x-axis has different validation sets.
[0151] As demonstrated by the p-values in Figure 31, PrimateAI 3D is more performant than PrimateAI over most of the validation sets (the only exception being the tp53s3_A549_p53NULL_Nutlin-3 validation set). That is, the vertical bars for PrimateAI 3D are consistently longer than those for PrimateAI.
[0152] Also in Figure 31, a separate "mean" chart calculates the mean of the p-values determined for each of the validation sets. Also in the mean chart, PrimateAI 3D is more performant than PrimateAI.
[0153] Mean statistics can be biased by outliers. To address this, a separate "method rank" chart is also shown in Figure 31. A higher rank indicates lower classification accuracy. Similarly in the method rank chart, PrimateAI 3D is more performant than PrimateAI by having more counts of the lower ranks 1 and 2 for all of Primate AI having 3.
[0154] In FIGS. 28 to 31, it is also clear that by combining PrimateAI 3D with PrimateAI, excellent classification accuracy can be obtained. That is, a protein can be supplied to PrimateAI as an amino acid sequence to generate a first output, and the same protein can be supplied to PrimateAI 3D as a 3D voxelized protein structure to generate a second output. By combining or analyzing the first and second outputs as a whole, a final pathogenicity prediction of the variants experienced by the protein can be generated.
[0155] Efficient voxelization FIG. 32 is a flowchart showing an efficient voxelization process 3200 that efficiently identifies nearest neighbor atoms for each voxel.
[0156] Here, the distance channel will be described again. As described above, the reference amino acid sequence 202 can contain different types of atoms such as alpha carbon atoms, beta carbon atoms, oxygen atoms, nitrogen atoms, and hydrogen atoms. Therefore, as described above, the distance channel can be arranged by the nearest neighbor alpha carbon atom, the nearest neighbor beta carbon atom, the nearest neighbor oxygen atom, the nearest neighbor nitrogen atom, the nearest neighbor hydrogen atom, etc. For example, in FIG. 6, each of the nine voxels 514 has a distance channel of 21 amino acid units for the nearest neighbor alpha carbon atom. FIG. 6 can be further extended so that each of the nine voxels 514 also has a distance channel of 21 amino acid units for the nearest neighbor beta carbon atom, and each of the nine voxels 514 also has the closest general atomic distance channel for the nearest neighbor atom, regardless of the type of atom and the type of amino acid. In this way, each of the nine voxels 514 can have 43 distance channels.
[0157] Next, the number of distance calculations required to identify the nearest neighbor atoms for each voxel for inclusion in the distance channel will be described. Consider the example of FIG. 3 showing a total of 828 alpha carbon atoms distributed across 21 amino acid categories. To calculate the distance channels 602-642 of the amino acid units in FIG. 6, i.e., to determine 189 distance values, the distance from each of the 9 voxels 514 to each of the 828 alpha carbon atoms is measured, and 9 * 828 = 7,452 distance calculations are obtained. In the case of 3D with 27 voxels, this results in 27 * 828 = 22,356 distance calculations. If 828 beta carbon atoms are also included, this number increases to 27 * 1656 = 44,712 distance calculations.
[0158] This means that the runtime complexity for identifying the nearest neighbor atoms for each voxel for single protein voxelization is O(number of atoms * number of voxels), as illustrated by FIG. 35A. Further, when the distance channels are calculated across various attributes (e.g., different features or channels per voxel such as annotation channels and structure confidence channels), the runtime complexity of single protein voxelization increases to O(number of atoms * number of voxels * number of attributes).
[0159] As a result, distance calculations can become the most computationally consuming part of the voxelization process, removing valuable computational resources from important runtime tasks such as model learning and model inference. For example, consider the case of model learning using a training dataset of 7,000 proteins. Generating distance channels for multiple voxels across multiple amino acids, atoms, and attributes can involve over 100 voxelizations per protein, resulting in approximately 800,000 voxelizations in a single learning iteration (epoch). Learning runs of 20-40 epochs with rotation of atomic coordinates in each epoch can result in 32 million voxelizations.
[0160] In addition to the high computational cost, the size of the data for 32 million voxels is too large to fit into main memory (e.g., exceeding 20 TB for a 15×15×15 voxel grid). Considering the execution of iterative learning for parameter optimization and ensemble learning, the memory footprint of the voxelization process is too large to store on disk, and the voxelization process is not made part of the model learning and a pre-computation step.
[0161] The disclosed technology provides an efficient voxelization process that achieves a speedup of up to about 100 times for a runtime complexity of O(number of atoms * number of voxels). The disclosed efficient voxelization process reduces the runtime complexity for single-protein voxelization to O(number of atoms). For the case of different features or channels per voxel, the disclosed efficient voxelization process reduces the runtime complexity for single-protein voxelization to O(number of atoms * number of attributes). As a result, the voxelization process becomes as fast as the model learning, and the computational bottleneck is shifted from voxelization to calculating neural network weights on processors such as GPUs, ASICs, TPUs, FPGAs, CGRAs.
[0162] In some embodiments of the disclosed efficient voxelization process with a large voxel grid, the runtime complexity for single-protein voxelization is O(number of atoms + number of voxels) and O(number of atoms * number of attributes + number of voxels) for the case of different features or channels per voxel. The "+ number of voxels" complexity is observed when the number of atoms is very small compared to the number of voxels, e.g., when there is 1 atom in a 100×100×100 voxel grid (i.e., 1 million voxels per atom). In such a scenario, the runtime is dominated by the overhead of a huge number of voxels, e.g., for allocating memory for 1 million voxels and initializing 1 million voxels to 0.
[0163] Here, the details of the disclosed efficient voxelization process will be described. FIGS. 32A, 32B, 33, 34, and 35B are described in parallel.
[0164] Starting from FIG. 32A, in step 3202, each atom (e.g., each of the 828 alpha carbon atoms and each of the 828 beta carbon atoms) is associated with a voxel containing the atom (e.g., one of the nine voxels 514). The term "containing" refers to the 3D atomic coordinates of the atoms located within the voxel. The voxel containing the atom is also referred to herein as an "atom-containing voxel".
[0165] FIGS. 32B and 33 illustrate how a voxel containing a particular atom is selected. FIG. 33 uses 2D atomic coordinates as representing 3D atomic coordinates. Note that for the voxel grid 522, each of the voxels 514 is equally spaced with the same step size (e.g., 1 angstrom (Å) or 2 Å).
[0166] Also, in FIG. 33, the voxel grid 522 has indices [0, 1, 2] along the first dimension (e.g., the x-axis) and indices [0, 1, 2] along the second dimension (e.g., the y-axis). Also, in FIG. 33, each voxel 514 within the voxel grid 522 is identified by a voxel index [voxel 0, voxel 1,..., voxel 8] and by a voxel center index [(1, 1), (1, 2),..., (3, 3)].
[0167] Also, in FIG. 33, the center coordinates of the voxel centers along the first dimension, i.e., the voxel coordinates of the first dimension, are identified. Also, in FIG. 33, the center coordinates of the voxel centers along the second dimension, i.e., the voxel coordinates of the second dimension, are identified.
[0168] First, in step 3202a (step 1 in FIG. 33), the 3D atomic coordinates (1.7456, 2.14323) of a specific atom are quantized to generate quantized 3D atomic coordinates (1.7, 2.1). Quantization can be achieved by rounding or truncating bits.
[0169] Next, in step 3202b (step 2 in FIG. 33), the voxel coordinates (or voxel center or voxel center coordinates) of voxel 514 are assigned to the dimension-based quantized 3D atomic coordinates. For the first dimension, the quantized atomic coordinate 1.7 covers the voxel coordinates in the first dimension in the range from 1 to 2 and is centered at 1.5 in the first dimension, so it is assigned to voxel 1. Note that voxel 1 has index 1 along the first dimension, as opposed to having index 0 along the second dimension.
[0170] For the second dimension, starting from voxel 1, the voxel grid 522 is traversed along the second dimension. As a result, the quantized atomic coordinate 2.5 is assigned to voxel 7, because voxel 7 covers the voxel coordinates in the second dimension in the range from 2 to 3 and is centered at 2.5 in the second dimension. Note that voxel 7 has index 2 along the second dimension, as opposed to having index 1 along the first dimension.
[0171] Next, in step 3202c (step 3 in FIG. 33), the dimension index corresponding to the assigned voxel coordinates is selected. That is, for voxel 1, index 1 is selected along the first dimension, and for voxel 7, index 2 is selected along the second dimension. Those skilled in the art will understand that the above steps can be similarly performed for the third dimension to select the dimension index along the third dimension.
[0172] Next, in step 3202d (step 4 in FIG. 33), a cumulative sum is generated based on weighting the selected dimension index by a power of the radix in position units. The general concept behind the positional numbering system is that a numerical value is represented by a power of a radix (or base), for example, a binary number has a radix of 2, a ternary number has a radix of 3, an octal number has a radix of 8, and a hexadecimal number has a radix of 16. Since each position is weighted by a power of the radix, it is often called a weighted numbering system. The set of valid digits for a positional numbering system is of a size equal to the radix of that system. For example, in decimal there are 10 digits from 0 to 9, and in ternary there are 3 digits 0, 1, 2. The largest valid number in a radix system is one less than the radix (thus, 8 is not a valid number in any radix system smaller than 9). Any decimal integer can be exactly represented in any other integer radix system, and vice versa.
[0173] Returning to the example of FIG. 33, the selected dimension indexes 1 and 2 are converted to a single integer by multiplying them by the respective powers of 3 in position units and summing the results of the position unit multiplications. Since the 3D atomic coordinates have three dimensions, radix 3 is chosen here (however, FIG. 33 shows only 2D atomic coordinates along two dimensions for simplicity).
[0174] Since index 2 is located in the rightmost bit (i.e., the least significant bit), it is multiplied by 3 to the power of 0 to obtain 2. Since index 1 is located in the second bit from the right (i.e., the second bit from the least significant bit), it is multiplied by 3 to the power of 1 to obtain 3. As a result, the cumulative sum is 5.
[0175] Next, in step 3202e (step 5 in FIG. 33), based on the cumulative sum, the voxel index of the voxel containing the specific atom is selected. That is, the cumulative sum is interpreted as the voxel index of the voxel containing the specific atom.
[0176] In step 3212, after each atom is associated with an atom-containing voxel, each atom is further associated with one or more voxels in the vicinity of the atom-containing voxel, which is also referred to herein as a "neighboring voxel". The neighboring voxels can be selected based on being within a predetermined radius (e.g., 5 angstroms (Å)) of the atom-containing voxel. In other embodiments, the neighboring voxels can be selected based on being continuously adjacent to the atom-containing voxel (e.g., the upper, lower, right, and left adjacent voxels). The resulting association that associates each atom with the atom-containing voxel and the neighboring voxels is encoded in a mapping 3402 from atoms to voxels, which is also referred to herein as a mapping from elements to cells. In one example, a first alpha-carbon atom is associated with a first subset of voxels 3404 that includes the atom-containing voxel and the neighboring voxels of the first alpha-carbon atom. In another example, a second alpha-carbon atom is associated with a second subset of voxels 3406 that contains the atom-containing voxel and the neighboring voxels of the second alpha-carbon atom.
[0177] Note that no distance calculations are performed to determine the atom-containing voxel and the neighboring voxels. The atom-containing voxel is selected by the spatial arrangement of the voxels that allows for the assignment of the quantized 3D atomic coordinates to the corresponding regularly spaced voxel centers within the voxel grid (without using distance calculations). Also, the neighboring voxels are selected by being spatially adjacent to the atom-containing voxel within the voxel grid (again without using distance calculations).
[0178] In step 3222, each voxel is mapped to the atoms associated in steps 3202 and 3212. In one embodiment, this mapping is encoded by a voxel-to-atom mapping 3412, which is generated based on a mapping 3402 from atoms to voxels (e.g., by applying a voxel-based sort key to the mapping 3402 from atoms to voxels). The voxel-to-atom mapping 3412 is also referred to herein as a "cell-to-element mapping". In one example, in steps 3202 and 3212, the first voxel is mapped to a first subset 3414 of alpha carbon atoms including the alpha carbon atom associated with the first voxel. In another example, in steps 3202 and 3212, the second voxel is mapped to a second subset 3416 of alpha carbon atoms including the alpha carbon atom associated with the second voxel.
[0179] In step 3232, for each voxel, the distance between the voxel and the atoms mapped to the voxel in step 3222 is calculated. Step 3232 has a runtime complexity of O(number of atoms) because the distance to a particular atom is measured only once from each voxel to which the particular atom is uniquely mapped in the voxel-to-atom mapping 3412. This is true when neighboring voxels are not considered. Without a neighborhood, the constant factor implied by big-O notation is 1. In a neighborhood, since the number of neighbors is constant for each voxel, the big-O notation is equal to the number of neighbors + 1, and thus the runtime complexity remains O(number of atoms). In contrast, in FIG. 35A, the distance to a particular atom is measured repeatedly the number of voxels (e.g., 27 distances for a particular atom due to 27 voxels).
[0180] In FIG. 35B, based on the mapping 3412 from voxels to atoms, each voxel is mapped to a respective subset of 828 atoms (not including distance calculations to neighboring voxels), as indicated by respective ellipses for each voxel. Each subset, with some exceptions, mostly does not overlap. As shown by the prime symbol “’” and the yellow overlap between ellipses in FIG. 35B, there is some overlap due to several cases where multiple atoms are mapped to the same voxel. This minimal overlap has an additive effect on the O(number of atoms) runtime complexity and does not have a multiplicative effect. This overlap is the result of considering neighboring voxels after determining the voxels containing atoms. Without neighboring voxels, an atom is associated with only one voxel, so no overlap can exist. However, when considering neighbors, each neighbor can potentially be associated with the same atom (unless there are other atoms of the same amino acid closer).
[0181] In step 3242, for each voxel, based on the distances calculated in step 3232, the nearest neighbor atom to the voxel is identified. In one embodiment, this identification is encoded in the mapping 3422 from voxels to nearest neighbor atoms, also referred to herein as “mapping from cell to nearest neighbor element”. In one example, a first voxel is mapped to a second alpha carbon atom as its nearest neighbor alpha carbon atom 3424. In another example, a second voxel is mapped to a 31st alpha carbon atom as its nearest neighbor alpha carbon atom 3426.
[0182] Furthermore, since the distance per voxel is calculated using the techniques described above, the classification of atom types and amino acid types of atoms and the corresponding distance values are stored to generate a classified distance channel.
[0183] Once the distances to the nearest neighbor atoms are identified using the techniques described above, these distances can be encoded in a distance channel for voxelization and subsequent processing by the pathogenicity classifier 2108.
[0184] Computer system FIG. 36 shows an exemplary computer system 3600 that may be used to implement the disclosed technology. The computer system 3600 includes at least one central processing unit (CPU) 3672 that communicates with a number of peripheral devices via a bus subsystem 3655. These peripheral devices can include, for example, a memory subsystem 3610 that includes a memory device and a file storage subsystem 3636, a user interface input device 3638, a user interface output device 3676, and a network interface subsystem 3674. The input and output devices enable user interaction with the computer system 3600. The network interface subsystem 3674 provides an interface to an external network that includes an interface to a corresponding interface device within other computer systems.
[0185] In one embodiment, the pathogenic classifier 2108 is communicatively linked to the memory subsystem 3610 and the user interface input device 3638.
[0186] The user interface input device 3638 can include a keyboard, a mouse, a trackball, a touchpad, or a pointing device such as a graphics tablet, a scanner, a touch screen incorporated into a display, an audio input device such as a voice recognition system and a microphone, and other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and means for inputting information into the computer system 3600.
[0187] The user interface output device 3676 can include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem can include an LED display, a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visible image. The display subsystem can also provide a non-visual display such as an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and means for outputting information from the computer system 3600 to the user or another machine or computer system.
[0188] The memory subsystem 3610 stores programming and data constructs that provide some or all of the functionality of some of the modules and methods described herein. These software modules are generally executed by the processor 3678.
[0189] Processor 3678 can be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). Processor 3678 can be hosted by deep learning cloud platforms such as Google Cloud Platform (trademark), Xilinx (trademark), and Cirrascale (trademark). Examples of Processor 3678 include Google's Tensor Processing Unit (TPU) (trademark), rackmount solutions such as the GX4 Rackmount Series (trademark), GX36 Rackmount Series (trademark), NVIDIA DGX-1 (trademark), Microsoft’ Stratix V FPGA (trademark), Graphcore's Intelligent Processor Unit (IPU) (trademark), Qualcomm's Zeroth Platform (trademark) with Snapdragon processors (trademark), NVIDIA's Volta (trademark), NVIDIA's DRIVE PX (trademark), NVIDIA's JETSON TX1 / TX2 MODULE (trademark), Intel's Nirvana (trademark), Movidius VPU (trademark), Fujitsu DPI (trademark), ARM's DynamicIQ (trademark), IBM TrueNorth (trademark), Lambda GPU Server with Testa V100s (trademark), and others.
[0190] The memory subsystem 3622 used in the memory subsystem 3610 can include a number of memories, such as a main random access memory (RAM) 3632 for storing instructions and data during program execution, and a read only memory (ROM) 3634 in which fixed instructions are stored. The file storage subsystem 3636 can provide a persistent storage device for program and data files, and can include a hard disk drive, a related removable medium, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functions of a particular embodiment can be stored by the file storage subsystem 3636 within the storage subsystem 3610, or in other machines accessible by the processor.
[0191] The bus subsystem 3655 provides a mechanism for enabling the various components and subsystems of the computer system 3600 to communicate with each other as intended. Although the bus subsystem 3655 is shown schematically as a single bus, alternative embodiments of the bus subsystem can use multiple buses.
[0192] The computer system 3600 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Since computers and networks are constantly changing in nature, the description of the computer system 3600 shown in FIG. 36 is intended only as a specific example for the purpose of illustrating a preferred embodiment of the present invention. Many other configurations of the computer system 3600 can have more or fewer components than the computer system shown in FIG. 36.
[0193] Specific Embodiment 1 The following embodiments can be implemented as a system, method, or product. One or more features of an embodiment can be combined with a base embodiment. Embodiments that are not mutually exclusive are taught to be combinable. One or more features of an embodiment can be combined with other embodiments. The present disclosure periodically notifies users of these options. Omissions from some embodiments of the repeated listing of these options should not be construed as limiting the combinations taught in the foregoing sections. These descriptions are incorporated herein by reference to each of the following implementations.
[0194] The disclosed technology uses 3D data as input, but in other embodiments, 1D data, 2D data (e.g., pixels and 2D atomic coordinates), 4D data, 5D data, etc. can be used as well.
[0195] In some embodiments, the system includes a memory that stores distance channels for amino acid units for a plurality of amino acids in a protein. Each of the distance channels for amino acid units has distance values in voxel units for voxels within a plurality of voxels. The distance values in voxel units specify the distances from corresponding voxels within the plurality of voxels to atoms of corresponding amino acids within the plurality of amino acids. The system further includes a pathogenicity determination engine configured to process a tensor that includes the distance channels for amino acid units and alternative alleles of a protein expressed by a variant. The pathogenicity determination engine can also be configured to determine the pathogenicity of the variant based at least in part on the tensor.
[0196] In some embodiments, the system further comprises a distance channel generator that centers a voxel grid of voxels on the alpha carbon atom of each residue of an amino acid. The distance channel generator can center the voxel grid on the alpha carbon atom of a residue of a specific amino acid located at a variant amino acid in a protein.
[0197] The system can be configured to encode the directionality of amino acids and the position of a particular amino acid in a tensor by multiplying the distance value in voxel units of the amino acid preceding a particular amino acid by a directionality parameter. The distance can be the nearest neighbor atomic distance from the corresponding voxel center in the voxel grid to the nearest neighbor atom of the corresponding amino acid. In some embodiments, the nearest neighbor atomic distance can be the Euclidean distance. The nearest neighbor atomic distance can be normalized by dividing the Euclidean distance by the maximum nearest neighbor atomic distance. The amino acid can have an alpha carbon atom, and in some embodiments, the distance can be the nearest neighbor alpha carbon atom distance from the corresponding voxel center to the nearest neighbor alpha carbon atom of the corresponding amino acid. The amino acid can have a beta carbon atom, and in some embodiments, the distance can be the nearest neighbor beta carbon atom distance from the corresponding voxel center to the nearest neighbor beta carbon atom of the corresponding amino acid. The amino acid can have backbone atoms, and in some embodiments, the distance can be the nearest neighbor backbone atom distance from the corresponding voxel center to the nearest neighbor backbone atom of the corresponding amino acid. The amino acid can have side chain atoms, and in some embodiments, the distance can be the nearest neighbor side chain atom distance from the corresponding voxel center to the nearest neighbor side chain atom of the corresponding amino acid.
[0198] The system can be further configured to encode, in the tensor, a nearest neighbor atom channel that specifies the distance from each voxel to the nearest neighbor atom. The nearest neighbor atoms can be selected regardless of the amino acid and the atomic elements of the amino acid. In some embodiments, the distance is the Euclidean distance. The distance can be normalized by dividing the Euclidean distance by a maximum distance. The amino acid can include non-standard amino acids. The tensor can further include an absent atom channel that specifies atoms not found within a predetermined radius of the voxel center, and the absent atom channel can be one-hot encoded. In some embodiments, the tensor can further include a one-hot encoding of alternative alleles that are encoded in voxel units for each of the distance channels of the amino acid units. The tensor can further include a reference allele of the protein. In some embodiments, the tensor can further include a one-hot encoding of reference alleles that are encoded in voxel units for each of the distance channels of the amino acid units. The tensor can further include an evolutionary profile that specifies the conservation level of amino acids across multiple species.
[0199] The system can further include an evolutionary profile generator that, for each voxel, selects the nearest neighbor atoms across amino acids and atom categories, selects a pan - amino acid conservation frequency array for the amino acid residues containing the nearest neighbor atoms, and makes the pan - amino acid conservation frequency array available as one of the evolutionary profiles. The pan - amino acid conservation frequency array can be constructed for specific positions of residues as observed in multiple species. The pan - amino acid conservation frequency array can identify whether there is a missing conservation frequency for a particular amino acid. In some embodiments, the evolutionary profile generator can, for each voxel, select the respective nearest neighbor atoms in each of the amino acids, select the conservation frequency for each amino acid for each of the respective residues of the amino acids containing the nearest neighbor atoms, and make the conservation frequency for each amino acid available as one of the evolutionary profiles. The conservation frequency for each amino acid can be constructed for specific positions of residues as observed in multiple species. The conservation frequency for each amino acid can identify whether there is a missing conservation frequency for a particular amino acid.
[0200] In some embodiments of the system, the tensor can further include an annotation channel for amino acids. The annotation channel can be one-hot encoded in the tensor. The annotation channel can be a molecular processing annotation including initiator methionine, signal, transport peptide, propeptide, chain, and peptide. The annotation channel can be a region annotation including topological domain, transmembrane, in-membrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias. The annotation channel can be a site annotation including active site, metal binding, binding site, and site. The annotation channel can be an amino acid modification annotation including non-standard residue, modified residue, lipidation, glycosylation, disulfide bond, and crosslinking. The annotation channel can be a secondary structure annotation including helix, turn, and beta strand. The annotation channel can be an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residue, and non-terminal residue.
[0201] In some embodiments of the system, the tensor can further include an amino acid structure reliability channel that specifies the quality of each structure of the amino acids. The structure reliability channel can be a global model quality estimation (GMQE). The structure reliability channel can include a qualitative model energy analysis (QMEAN) score.
[0202] The structure reliability channel can be a temperature factor that specifies the extent to which a residue satisfies the physical constraints of the respective protein structure. The structure reliability channel can be a template structure alignment that specifies the extent to which the residues of the nearest neighbor atoms to the voxel have an aligned template structure. The structure reliability channel can be a template modeling score of the aligned template structure. The structure reliability channel can be the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0203] In some embodiments, the system can further include a tensor generator that concatenates the distance channels of amino acid units of alpha carbon atoms with the one-hot encoding of alternative alleles in voxel units to generate a tensor. The tensor generator can concatenate the distance channels of amino acid units for beta carbon atoms with the one-hot encoding of alternative alleles in voxel units to generate a tensor. The tensor generator can concatenate the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, and the one-hot encoding of alternative alleles in voxel units to generate a tensor. The tensor generator can concatenate the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, and the general amino acid conservation frequency in voxel units to generate a tensor. The tensor generator can concatenate the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the general amino acid conservation frequency, and the annotation channel in voxel units to generate a tensor. The tensor generator can concatenate the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the general amino acid conservation frequency, the annotation channel, and the structural confidence channel in voxel units to generate a tensor. The tensor generator can concatenate the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, and the conservation frequency for each amino acid for each amino acid in voxel units to generate a tensor. The tensor generator can concatenate the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the conservation frequency for each amino acid for each amino acid, and the annotation channel in voxel units to generate a tensor.The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the per-amino-acid conservation frequency for each of the amino acids, the annotation channel, and the structural confidence channel. The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, and the one-hot encoding of reference alleles. The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, and the pan-amino-acid conservation frequency. The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the pan-amino-acid conservation frequency, and the annotation channel. The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the pan-amino-acid conservation frequency, the annotation channel, and the structural confidence channel. The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, and the per-amino-acid conservation frequency for each of the amino acids.The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the conservation frequency for each amino acid, and the annotation channel. The tensor generator can generate a tensor by concatenating, in voxel units, the distance channel of amino acid units for alpha carbon atoms, the distance channel of amino acid units for beta carbon atoms, the one-hot encoding of alternative alleles, the one-hot encoding of reference alleles, the conservation frequency for each amino acid, the annotation channel, and the structural confidence channel.
[0204] In some embodiments, the system can further comprise an atomic rotation engine that rotates the atoms of the amino acid before the distance channel of the amino acid unit is generated. The pathogenicity determination engine can be a neural network. In certain embodiments, the pathogenicity determination engine can be a convolutional neural network. The convolutional neural network can use 1x1x1 convolutions, 3×3×3 convolutions, a normalized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer. The 1x1x1 convolution and the 3×3×3 convolution can be 3D convolutions.
[0205] In some embodiments, a 1×1×1 convolutional layer can process a tensor and generate an intermediate output that is a convolutional representation of the tensor. A sequence of 3×3×3 convolutional layers can process the intermediate output and generate a flattened output. A fully connected layer can process the flattened output and generate an unnormalized output. A softmax classification layer can process the unnormalized output and generate an exponentially normalized output that discriminates the likelihood that the variant is pathogenic and benign. A sigmoid layer can process the unnormalized output and generate a normalized output that discriminates the likelihood that the variant is pathogenic. A voxel, an atom, and a distance can have three-dimensional coordinates. A tensor can have at least three dimensions, an intermediate output can have at least three dimensions, and a flattened output can have one dimension.
[0206] In some embodiments, the pathogenicity determination engine is a recurrent neural network. In other embodiments, the pathogenicity determination engine is an attention-based neural network. In still other embodiments, the pathogenicity determination engine is a gradient boosted tree. In still other embodiments, the pathogenicity determination engine is a state vector machine.
[0207] In other embodiments, the system can include a memory that stores a distance channel of atomic category units for amino acids in a protein. An amino acid can have atoms of a plurality of atomic categories, and the atomic categories within the plurality of atomic categories can identify the atomic elements of the amino acid. The distance channel of atomic category units can have distance values in voxel units for voxels within a plurality of voxels. The distance value in voxel units can identify the distance from a corresponding voxel within the plurality of voxels to an atom within a corresponding atomic category within the plurality of atomic categories. The system can further include a pathogenicity determination engine configured to process a tensor that includes the distance channel of atomic category units and an alternative allele of a protein expressed by a variant and determine the pathogenicity of the variant based at least in part on the tensor.
[0208] The system can further comprise a distance channel generator that centers a voxel grid of voxels on each atom of each atom category within a plurality of atom categories. The distance channel generator can place the center of the voxel grid on the alpha carbon atom of the residue of at least one mutant amino acid in the protein. The distance can be the nearest neighbor atom distance from the corresponding voxel center within the voxel grid to the nearest neighbor atom within the corresponding atom category. The nearest neighbor atom distance can be the Euclidean distance. The nearest neighbor atom distance can be normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance. The distance can be the nearest neighbor atom distance from the corresponding voxel center within the voxel grid to the nearest neighbor atom, regardless of the amino acid and the atom category of the amino acid. The nearest neighbor atom distance can be the Euclidean distance. The nearest neighbor atom distance can be normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0209] Other embodiments of the methods described in this section can include a non-transitory computer-readable storage medium that stores instructions executable by a processor to perform any of the methods described above. Still other embodiments of the methods described in this section can include a system that includes a memory and one or more processors operable to execute instructions stored in the memory to perform any of the methods described above.
[0210] Clause 1 1. Storing a distance channel of amino acid units for a plurality of amino acids in a protein, and each of the distance channels of amino acid units has distance values of voxel units for voxels within a plurality of voxels, the distance values of voxel units specify the distances from the corresponding voxels within the plurality of voxels to the atoms of the corresponding amino acids within the plurality of amino acids, processing a tensor that includes the distance channels of amino acid units and alternative alleles of a protein expressed by a mutant, and A computer-implemented method comprising: determining the pathogenicity of a variant based at least in part on a tensor.
[0211] 2. The computer-implemented method of claim 1, further comprising centering a voxel grid of voxels on the alpha carbon atom of each residue of the amino acids.
[0212] 3. The computer-implemented method of claim 2, further comprising centering a voxel grid on the alpha carbon atom of the residue of a specific amino acid corresponding to at least one mutant amino acid in the protein.
[0213] 4. The computer-implemented method of claim 3, further comprising encoding, in the tensor, the directionality of the amino acid and the position of the specific amino acid by multiplying the distance value in voxel units for the amino acid preceding the specific amino acid using a directionality parameter.
[0214] 5. The computer-implemented method of claim 3, wherein the distance is the nearest neighbor atom distance from the corresponding voxel center in the voxel grid to the nearest neighbor atom of the corresponding amino acid.
[0215] 6. The computer-implemented method of claim 5, wherein the nearest neighbor atom distance is the Euclidean distance.
[0216] 7. The computer-implemented method of claim 6, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0217] 8. The computer-implemented method of claim 5, wherein the amino acid has an alpha carbon atom and the distance is the nearest neighbor alpha carbon atom distance from the corresponding voxel center to the nearest neighbor alpha carbon atom of the corresponding amino acid.
[0218] 9. The computer-implemented method according to claim 5, wherein the amino acid has a beta carbon atom and the distance is the nearest beta carbon atom distance from the corresponding voxel center to the nearest beta carbon atom of the corresponding amino acid.
[0219] 10. The computer-implemented method according to claim 5, wherein the amino acid has a backbone atom and the distance is the nearest backbone atom distance from the corresponding voxel center to the nearest backbone atom of the corresponding amino acid.
[0220] 11. The computer-implemented method according to claim 5, wherein the amino acid has a side chain atom and the distance is the nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding amino acid.
[0221] 12. The computer-implemented method according to claim 3, further comprising the step of encoding a nearest neighbor atom channel that specifies the distance from each voxel to the nearest neighbor atom, wherein the nearest neighbor atom is selected regardless of the amino acid and the atomic elements of the amino acid.
[0222] 13. The computer-implemented method according to claim 12, wherein the distance is the Euclidean distance.
[0223] 14. The computer-implemented method according to claim 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0224] 15. The computer-implemented method according to claim 12, wherein the amino acid includes non-standard amino acids.
[0225] 16. The computer-implemented method according to claim 1, wherein the tensor further includes an absent atom channel that specifies atoms not found within a predetermined radius of the voxel center, and the absent atom channel is one-hot encoded.
[0226] 17. The computer-implemented method according to claim 1, wherein the tensor further includes a one-hot encoding of alternative alleles encoded in voxel units for each of the distance channels of the amino acid units.
[0227] 18. The computer-implemented method according to claim 1, wherein the tensor further comprises a reference allele of the protein.
[0228] 19. The computer-implemented method according to claim 18, wherein the tensor further comprises a one-hot encoding of the reference allele encoded in voxel units for each of the distance channels of the amino acid units.
[0229] 20. The computer-implemented method according to claim 1, wherein the tensor further comprises an evolutionary profile that identifies the conservation level of amino acids across multiple species.
[0230] 21. For each voxel, selecting the nearest neighbor atoms across amino acids and atom categories; selecting a pan-amino acid conservation frequency array for the residues of the amino acids containing the nearest neighbor atoms; and making the pan-amino acid conservation frequency array available as one of the evolutionary profiles. The computer-implemented method according to claim 20 further comprises the steps of:
[0231] 22. The computer-implemented method according to claim 21, wherein the pan-amino acid conservation frequency array is constructed for specific positions of residues as observed in multiple species.
[0232] 23. The computer-implemented method according to claim 21, wherein the pan-amino acid conservation frequency array identifies whether there is a missing conservation frequency for a particular amino acid.
[0233] 24. For each voxel, selecting each of the nearest neighbor atoms in each of the amino acids; selecting, for each residue of the amino acids containing the nearest neighbor atoms, the conservation frequency for each amino acid; and making the conservation frequency for each amino acid available as one of the evolutionary profiles. The computer-implemented method according to claim 21 further comprises the steps of:
[0234] 25. The computer-implemented method according to claim 24, wherein the conservation frequency for each amino acid is configured for a specific position of a residue as observed in a plurality of species.
[0235] 26. The computer-implemented method according to claim 24, wherein the conservation frequency for each amino acid determines whether there is a missing conservation frequency for a specific amino acid.
[0236] 27. The computer-implemented method according to claim 1, wherein the tensor further includes an annotation channel for amino acids, and the annotation channel is one-hot encoded in the tensor.
[0237] 28. The computer-implemented method according to claim 27, wherein the annotation channel is a molecular processing annotation including initiator methionine, signal, transport peptide, propeptide, chain, and peptide.
[0238] 29. The computer-implemented method according to claim 27, wherein the annotation channel is a region annotation including topological domain, transmembrane, intra-membrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
[0239] 30. The computer-implemented method according to claim 27, wherein the annotation channel is a site annotation including active site, metal binding, binding site, and site.
[0240] 31. The computer-implemented method according to claim 27, wherein the annotation channel is an amino acid modification annotation including non-standard residue, modified residue, lipidation, glycosylation, disulfide bond, and crosslinking.
[0241] 32. The computer-implemented method according to claim 27, wherein the annotation channel is a secondary structure annotation including helix, turn, and beta strand.
[0242] 33. The computer-implemented method according to claim 27, wherein the annotation channel is experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residues, and non-terminal residues.
[0243] 34. The computer-implemented method according to claim 1, wherein the tensor further includes an amino acid structure reliability channel that specifies the quality of each structure of the amino acid.
[0244] 35. The computer-implemented method according to claim 34, wherein the structure reliability channel is global model quality estimation (GMQE).
[0245] 36. The computer-implemented method according to claim 34, wherein the structure reliability channel includes a qualitative model energy analysis (QMEAN) score.
[0246] 37. The computer-implemented method according to claim 34, wherein the structure reliability channel is a temperature factor that specifies the degree to which a residue satisfies the physical constraints of each protein structure.
[0247] 38. The computer-implemented method according to claim 34, wherein the structure reliability channel is a template structure alignment that specifies the degree to which the residues of the nearest neighbor atoms to the voxel have an aligned template structure.
[0248] 39. The computer-implemented method according to claim 38, wherein the structure reliability channel is a template modeling score of the aligned template structure.
[0249] 40. The computer-implemented method according to claim 39, wherein the structure reliability channel is one of the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0250] 41. The computer-implemented method according to claim 1, further comprising concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom with the one-hot encoding of the alternative alleles to generate a tensor.
[0251] 42. The computer-implemented method according to claim 41, further comprising concatenating, in voxel units, the distance channel of the amino acid unit for the beta carbon atom with the one-hot encoding of the alternative alleles to generate a tensor.
[0252] 43. The computer-implemented method according to claim 42, further comprising a step of concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom unit, the distance channel of the amino acid unit for the beta carbon atom unit, and the one-hot encoding of the alternative alleles to generate a tensor.
[0253] 44. The computer-implemented method according to claim 43, further comprising a step of concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative alleles, and the general amino acid conservation frequency array to generate a tensor.
[0254] 45. The computer-implemented method according to claim 44, further comprising concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative alleles, the general amino acid conservation frequency array, and the annotation channel to generate a tensor.
[0255] 46. The computer-implemented method according to claim 45, further comprising a step of concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative alleles, the general amino acid conservation frequency array, the annotation channel, and the structural reliability channel to generate a tensor.
[0256] 47. The computer-implemented method according to claim 46, further comprising concatenating, in voxel units, the distance channel of the amino acid unit of the alpha carbon atom, the distance channel of the amino acid unit of the beta carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative allele, and the conservation frequency for each amino acid for each amino acid to generate a tensor.
[0257] 48. The computer-implemented method according to claim 47, further comprising the step of concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative allele, the conservation frequency for each amino acid for each amino acid, and the annotation channel to generate a tensor.
[0258] 49. The computer-implemented method according to claim 48, further comprising concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative allele, the conservation frequency for each amino acid for each amino acid, the annotation channel, and the structural reliability channel to generate a tensor.
[0259] 50. The computer-implemented method according to claim 49, further comprising concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative allele, and the one-hot encoding of the reference allele.
[0260] 51. The computer-implemented method according to claim 50, further comprising concatenating, in voxel units, the distance channel of the amino acid unit for the alpha carbon atom, the distance channel of the amino acid unit for the beta carbon atom, the one-hot encoding of the alternative allele, the one-hot encoding of the reference allele, and the pan-amino acid conservation frequency array to generate a tensor.
[0261] 52. The computer-implemented method according to claim 51, further comprising generating a tensor by concatenating, in voxel units, a distance channel of amino acid units of alpha carbon atoms, a distance channel of amino acid units of beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a pan-amino acid conservation frequency array, and an annotation channel.
[0262] 53. The computer-implemented method according to claim 52, further comprising the step of generating a tensor by concatenating, in voxel units, a distance channel of amino acid units of alpha carbon atoms, a distance channel of amino acid units of beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a pan-amino acid conservation frequency array, an annotation channel, and a structural reliability channel.
[0263] 54. The computer-implemented method according to claim 53, further comprising the step of generating a tensor by concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, and the per-amino acid conservation frequency for each amino acid.
[0264] 55. The computer-implemented method according to claim 54, further comprising the step of generating a tensor by concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, the per-amino acid conservation frequency for each amino acid, and an annotation channel.
[0265] 56. The computer-implemented method according to claim 55, further comprising generating a tensor by concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a per-amino-acid conservation frequency for each of the amino acids, an annotation channel, and a structural confidence channel.
[0266] 57. The computer-implemented method according to claim 1, further comprising the step of rotating the atoms of the amino acid before the distance channel of the amino acid units is generated.
[0267] 58. The computer-implemented method according to claim 1, further comprising the step of using, in a convolutional neural network, a 1x1x1 convolution, a 3×3×3 convolution, a rectified linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer.
[0268] 59. The 1x1x1 convolution and the 3×3×3 convolution are 3D convolutions, the computer-implemented method according to claim 58.
[0269] 60. The layer of 1x1x1 convolutions processes the tensor and generates an intermediate output that is a convolutional representation of the tensor, the sequence of layers of 3×3×3 convolutions processes the intermediate output and generates a flattened output, the fully connected layer processes the flattened output and generates an unnormalized output, and the softmax classification layer processes the unnormalized output and generates an exponentially normalized output that discriminates the likelihood that the variant is pathogenic and benign, the computer-implemented method according to claim 58.
[0270] 61. The sigmoid layer processes the unnormalized output and generates a normalized output that discriminates the likelihood that the variant is pathogenic, the computer-implemented method according to claim 60.
[0271] 62. The computer-implemented method according to claim 60, wherein the voxels, atoms, and distances have three-dimensional coordinates, the tensor has at least three dimensions, the intermediate output has at least three dimensions, and the flattened output has one dimension.
[0272] 63. A step of storing a distance channel of atomic category units of amino acids in a protein, wherein the amino acid has atoms of a plurality of atomic categories, among the plurality of atomic categories, the atomic category identifies an atomic element of the amino acid, each of the distance channels of atomic category units has distance values per voxel with respect to voxels in a plurality of voxels, the distance value per voxel identifies the distance from the corresponding voxel in the plurality of voxels to the atom in the corresponding atomic category among the plurality of atomic categories, a step of processing a tensor including the distance channel of atomic category units and an alternative allele of a protein expressed by a mutant, a step of determining the pathogenicity of the mutant, at least partially based on the tensor. A computer-implemented method comprising:
[0273] 64. The computer-implemented method according to claim 63, further comprising a step of centering a voxel grid of voxels on each atom of each atomic category in a plurality of atomic categories.
[0274] 65. The computer-implemented method according to claim 64, further comprising centering a voxel grid on the alpha-carbon atom of the residue of at least one mutant amino acid in the protein.
[0275] 66. The computer-implemented method according to claim 65, wherein the distance is the nearest neighbor atom distance from the corresponding voxel center in the voxel grid to the nearest neighbor atom in the corresponding atomic category.
[0276] 67. The computer-implemented method according to claim 66, wherein the nearest neighbor atom distance is the Euclidean distance.
[0277] 68. The computer-implemented method according to claim 67, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0278] 69. The computer-implemented method according to claim 68, wherein the distance is the nearest neighbor atom distance from the corresponding voxel center in the voxel grid to the nearest neighbor atom, regardless of the amino acid and the atomic category of the amino acid.
[0279] 70. The computer-implemented method according to claim 69, wherein the nearest neighbor atom distance is the Euclidean distance.
[0280] 71. The computer-implemented method according to claim 70, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0281] 1. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, store a distance channel of amino acid units for a plurality of amino acids in a protein; each of the distance channels of amino acid units has distance values of voxel units for voxels in a plurality of voxels, the distance values of voxel units identify the distances from the corresponding voxels in a plurality of voxels to the atoms of the corresponding amino acids in a plurality of amino acids, and process a tensor including the distance channel of amino acid units and alternative alleles of the protein expressed by the variant; and determine the pathogenicity of the variant based at least in part on the tensor. A computer-readable medium that configures a computer to perform the operations.
[0282] 2. The computer-readable medium according to claim 1, wherein the operation further includes centering a voxel grid of voxels on the alpha carbon atom of each residue of the amino acid.
[0283] 3. The computer-readable medium according to claim 2, wherein the operation further comprises centering a voxel grid on the alpha carbon atom of the residue of a specific amino acid corresponding to at least one variant amino acid in the protein.
[0284] 4. The computer-readable medium according to claim 3, wherein the operation further comprises encoding, in the tensor, the directionality of the amino acid and the position of the specific amino acid by multiplying the distance value in voxel units for the amino acid preceding the specific amino acid by a directionality parameter.
[0285] 5. The computer-readable medium according to claim 3, wherein the distance is the nearest neighbor atomic distance from the corresponding voxel center in the voxel grid to the nearest neighbor atom of the corresponding amino acid.
[0286] 6. The computer-readable medium according to claim 5, wherein the nearest neighbor atomic distance is the Euclidean distance.
[0287] 7. The computer-readable medium according to claim 6, wherein the nearest neighbor atomic distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atomic distance.
[0288] 8. The computer-readable medium according to claim 5, wherein the amino acid has an alpha carbon atom and the distance is the nearest neighbor alpha carbon atom distance from the corresponding voxel center to the nearest neighbor alpha carbon atom of the corresponding amino acid.
[0289] 9. The computer-readable medium according to claim 5, wherein the amino acid has a beta carbon atom and the distance is the nearest neighbor beta carbon atom distance from the corresponding voxel center to the nearest neighbor beta carbon atom of the corresponding amino acid.
[0290] 10. The computer-readable medium according to claim 5, wherein the amino acid has a backbone atom and the distance is the nearest neighbor backbone atom distance from the corresponding voxel center to the nearest neighbor backbone atom of the corresponding amino acid.
[0291] 11. The computer-readable medium according to claim 5, wherein the amino acid has side chain atoms, and the distance is the nearest side chain atom distance from the corresponding voxel center to the nearest side chain atom of the corresponding amino acid.
[0292] 12. The computer-readable medium according to claim 3, wherein the operation further includes encoding, in the tensor, a nearest neighbor atom channel that specifies the distance from each voxel to the nearest neighbor atom, and the nearest neighbor atom is selected regardless of the amino acid and the atomic elements of the amino acid.
[0293] 13. The computer-readable medium according to claim 12, wherein the distance is a Euclidean distance.
[0294] 14. The computer-readable medium according to claim 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0295] 15. The computer-readable medium according to claim 12, wherein the amino acid includes non-standard amino acids.
[0296] 16. The computer-readable medium according to claim 1, wherein the tensor further includes an absent atom channel that specifies atoms not found within a predetermined radius of the voxel center, and the absent atom channel is one-hot encoded.
[0297] 17. The computer-readable medium according to claim 1, wherein the tensor further includes a one-hot encoding of alternative alleles encoded in voxel units for each of the distance channels of the amino acid units.
[0298] 18. The computer-readable medium according to claim 1, wherein the tensor further includes a reference allele of the protein.
[0299] 19. The computer-readable medium according to claim 18, wherein the tensor further includes a one-hot encoding of reference alleles encoded in voxel units for each of the distance channels of the amino acid units.
[0300] 20. The computer-readable medium according to claim 1, wherein the tensor further includes an evolutionary profile that identifies the conservation levels of amino acids across multiple species.
[0301] 21. The operation further includes, for each voxel, selecting a nearest neighbor atom across amino acids and atom categories, and selecting a general amino acid conservation frequency array for the residue of the amino acid that includes the nearest neighbor atom, and making the general amino acid conservation frequency array available as one of the evolutionary profiles. The computer-readable medium according to claim 20.
[0302] 22. The computer-readable medium according to claim 21, wherein the general amino acid conservation frequency array is configured for specific positions of residues as observed in multiple species.
[0303] 23. The computer-readable medium according to claim 21, wherein the general amino acid conservation frequency array identifies whether there is a missing conservation frequency for a specific amino acid.
[0304] 24. The operation further includes, for each voxel, selecting each nearest neighbor atom in each of the amino acids, and selecting, for each residue of the amino acid that includes the nearest neighbor atom, a conservation frequency for each amino acid, and making the conservation frequency for each amino acid available as one of the evolutionary profiles. The computer-readable medium according to claim 21.
[0305] 25. The computer-readable medium according to claim 24, wherein the conservation frequency for each amino acid is configured for specific positions of residues as observed in multiple species.
[0306] 26. The computer-readable medium according to claim 24, wherein the conservation frequency for each amino acid identifies whether there is a missing conservation frequency for a specific amino acid.
[0307] 27. The tensor further includes an annotation channel for amino acids, and the annotation channel is one-hot encoded in the tensor, the computer-readable medium according to claim 1.
[0308] 28. The annotation channel is a molecular processing annotation including initiator methionine, signal, transport peptide, propeptide, chain, and peptide, the computer-readable medium according to claim 27.
[0309] 29. The annotation channel is a region annotation including topological domain, transmembrane, in-membrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias, the computer-readable medium according to claim 27.
[0310] 30. The annotation channel is a site annotation including active site, metal binding, binding site, and site, the computer-readable medium according to claim 27.
[0311] 31. The annotation channel is an amino acid modification annotation including non-standard residue, modified residue, lipidation, glycosylation, disulfide bond, and crosslinking, the computer-readable medium according to claim 27.
[0312] 32. The annotation channel is a secondary structure annotation including helix, turn, and beta strand, the computer-readable medium according to claim 27.
[0313] 33. The annotation channel is an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residue, and non-terminal residue, the computer-readable medium according to claim 27.
[0314] 34. The computer-readable medium according to claim 1, wherein the tensor further includes an amino acid structure reliability channel that identifies the quality of each structure of the amino acids.
[0315] 35. The computer-readable medium according to claim 34, wherein the structure reliability channel is a global model quality estimation (GMQE).
[0316] 36. The computer-readable medium according to claim 34, wherein the structure reliability channel includes a qualitative model energy analysis (QMEAN) score.
[0317] 37. The computer-readable medium according to claim 34, wherein the structure reliability channel is a temperature factor that identifies the degree to which residues satisfy the physical constraints of each protein structure.
[0318] 38. The computer-readable medium according to claim 34, wherein the structure reliability channel is a template structure alignment that identifies the degree to which residues of neighboring atoms to a voxel have an aligned template structure.
[0319] 39. The computer-readable medium according to claim 38, wherein the structure reliability channel is a template modeling score of the aligned template structure.
[0320] 40. The computer-readable medium according to claim 39, wherein the structure reliability channel is one of the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0321] 41. The computer-readable medium according to claim 1, wherein the operation further includes concatenating a distance channel of amino acid units for alpha carbon atoms with a one-hot encoding of alternative alleles in voxel units to generate a tensor.
[0322] 42. The computer-readable medium according to claim 41, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for beta carbon atoms with a one-hot encoding of alternative alleles to generate a tensor.
[0323] 43. The computer-readable medium according to claim 42, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, and a one-hot encoding of alternative alleles to generate a tensor.
[0324] 44. The computer-readable medium according to claim 43, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, and a general amino acid conservation frequency array to generate a tensor.
[0325] 45. The computer-readable medium according to claim 44, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a general amino acid conservation frequency array, and an annotation channel to generate a tensor.
[0326] 46. The computer-readable medium according to claim 45, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a general amino acid conservation frequency array, an annotation channel, and a structural reliability channel to generate a tensor.
[0327] 47. The computer-readable medium according to claim 46, wherein the operation further includes concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, and a per-amino-acid conservation frequency for each of the amino acids to generate a tensor.
[0328] 48. The computer-readable medium according to claim 47, wherein the operation further includes concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a per-amino-acid conservation frequency for each of the amino acids, and an annotation channel to generate a tensor.
[0329] 49. The computer-readable medium according to claim 48, wherein the operation further includes concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a per-amino-acid conservation frequency for each of the amino acids, an annotation channel, and a structural confidence channel to generate a tensor.
[0330] 50. The computer-readable medium according to claim 49, wherein the operation further includes concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, and a one-hot encoding of reference alleles to generate a tensor.
[0331] 51. The computer-readable medium according to claim 50, wherein the operation further includes concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, and a pan-amino-acid conservation frequency array to generate a tensor.
[0332] 52. The computer-readable medium according to claim 51, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a pan-amino acid conservation frequency array, and an annotation channel to generate a tensor.
[0333] 53. The computer-readable medium according to claim 52, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a pan-amino acid conservation frequency array, an annotation channel, and a structural confidence channel to generate a tensor.
[0334] 54. The computer-readable medium according to claim 53, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, and the per-amino acid conservation frequency for each amino acid to generate a tensor.
[0335] 55. The computer-readable medium according to claim 54, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, the per-amino acid conservation frequency for each amino acid, and an annotation channel to generate a tensor.
[0336] 56. The computer-readable medium of claim 55, wherein the operation further comprises concatenating, in voxel units, a distance channel of amino acid units for alpha carbon atoms, a distance channel of amino acid units for beta carbon atoms, a one-hot encoding of alternative alleles, a one-hot encoding of reference alleles, a per-amino acid conservation frequency for each amino acid, an annotation channel, and a structural confidence channel to generate a tensor.
[0337] 57. The computer-readable medium of claim 1, wherein the operation further comprises rotating the atoms of the amino acid before the distance channel of amino acid units is generated.
[0338] 58. The computer-readable medium of claim 1, wherein the operation further comprises using, in a convolutional neural network, a 1x1x1 convolution, a 3×3×3 convolution, a rectified linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer.
[0339] 59. The computer-readable medium of claim 58, wherein the 1x1x1 convolution and the 3×3×3 convolution are 3D convolutions.
[0340] 60. The layer of 1x1x1 convolutions processes a tensor and generates an intermediate output that is a convolutional representation of the tensor. The sequence of layers of 3×3×3 convolutions processes the intermediate output and generates a flattened output. The fully connected layer processes the flattened output and generates an unnormalized output. The softmax classification layer processes the unnormalized output and generates an exponentially normalized output that identifies the likelihood that the variant is pathogenic and benign. The computer-readable medium of claim 58.
[0341] 61. The computer-readable medium of claim 60, wherein a sigmoid layer processes the unnormalized output and generates a normalized output that identifies the likelihood that the variant is pathogenic.
[0342] 62. The computer-readable medium according to claim 60, wherein the voxels, atoms, and distances have three-dimensional coordinates, the tensor has at least three dimensions, the intermediate output has at least three dimensions, and the flattened output has one dimension.
[0343] 63. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, store a distance channel of atomic category units of amino acids in a protein; wherein the amino acids have atoms of a plurality of atomic categories, wherein the atomic categories among the plurality of atomic categories identify atomic elements of the amino acids, each of the distance channels of atomic category units has distance values in voxel units for voxels in a plurality of voxels, wherein the distance values in voxel units identify the distances from corresponding voxels in the plurality of voxels to atoms in corresponding atomic categories among the plurality of atomic categories; process a tensor including the distance channels of atomic category units and alternative alleles of a protein expressed by a mutant; and determine the pathogenicity of the mutant based at least in part on the tensor.
[0344] 64. The computer-readable medium according to claim 63, wherein the operation further includes centering a voxel grid of voxels on each atom of each atomic category in the plurality of atomic categories.
[0345] 65. The computer-readable medium according to claim 64, wherein the operation further includes centering a voxel grid on an alpha carbon atom of a residue of at least one mutant amino acid in the protein.
[0346] 66. The computer-readable medium according to claim 65, wherein the distance is the nearest neighbor atom distance from the corresponding voxel center in the voxel grid to the nearest neighbor atom in the corresponding atom category.
[0347] 67. The computer-readable medium according to claim 66, wherein the nearest neighbor atom distance is the Euclidean distance.
[0348] 68. The computer-readable medium according to claim 67, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0349] 69. The computer-readable medium according to claim 68, wherein the distance is the nearest neighbor atom distance from the corresponding voxel center in the voxel grid to the nearest neighbor atom, regardless of the amino acid and the atom category of the amino acid.
[0350] 70. The computer-readable medium according to claim 69, wherein the nearest neighbor atom distance is the Euclidean distance.
[0351] 71. The computer-readable medium according to claim 70, wherein the nearest neighbor atom distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atom distance.
[0352] Specific Embodiment 2 In some embodiments, the system includes a voxelizer that accesses the three-dimensional structure of a reference amino acid sequence of a protein and fits a three-dimensional grid of voxels to atoms within the three-dimensional structure on an amino acid basis to generate a distance channel for amino acid units. Each of the distance channels for amino acid units has three-dimensional distance values for each voxel within the three-dimensional grid of voxels. The three-dimensional distance values specify the distance from the corresponding voxel within the three-dimensional grid of voxels to the atoms of the corresponding reference amino acid within the reference amino acid sequence. The system further includes an alternative allele encoder that encodes alternative allele amino acids for each voxel within the three-dimensional grid of voxels. The alternative allele amino acids are a three-dimensional representation of the one-hot encoding of mutant amino acids expressed by mutant nucleotides. The system further includes an evolutionary conservation encoder that encodes an evolutionary conservation sequence for each voxel within the three-dimensional grid of voxels. The evolutionary conservation sequence can be a three-dimensional representation of amino acid-specific conservation frequencies across multiple species. The amino acid-specific conservation frequencies can be selected according to the amino acid proximity to the corresponding voxel. The system further includes a convolutional neural network configured to apply three-dimensional convolution to a tensor that includes the alternative allele amino acids and the distance channels for amino acid units encoded with their respective evolutionary conservation sequences. The convolutional neural network can also be configured to determine the pathogenicity of mutant nucleotides based at least in part on the tensor.
[0353] The voxelizer can place the center of the three-dimensional grid of voxels on the alpha carbon atom of each residue of the reference amino acids within the reference amino acid sequence. The voxelizer can place the center of the three-dimensional grid of voxels on the alpha carbon atom of a specific reference amino acid residue located at a mutant amino acid residue.
[0354] In some embodiments, the system may be further configured to encode the directionality of reference amino acids and the position of a particular reference amino acid within a reference amino acid sequence by multiplying, in a tensor, a directionality parameter by a three-dimensional distance value for a reference amino acid preceding the particular reference amino acid. The distance may be the nearest neighbor atomic distance from the corresponding voxel center in a three-dimensional grid of voxels to the nearest neighbor atom of the corresponding reference amino acid. The nearest neighbor atomic distance may be a Euclidean distance and can be normalized by dividing the Euclidean distance by a maximum nearest neighbor atomic distance.
[0355] In some embodiments, a reference amino acid can have an alpha carbon atom, and the distance can be the nearest neighbor alpha carbon atom distance from the corresponding voxel center to the nearest neighbor alpha carbon atom of the corresponding reference amino acid. In some embodiments, a reference amino acid can have a beta carbon atom, and the distance can be the nearest neighbor beta carbon atom distance from the corresponding voxel center to the nearest neighbor beta carbon atom of the corresponding reference amino acid. In some embodiments, a reference amino acid can have a backbone atom, and the distance can be the nearest neighbor backbone atom distance from the corresponding voxel center to the nearest neighbor backbone atom of the corresponding reference amino acid. In some embodiments, an amino acid can have a side chain atom, and the distance can be the nearest neighbor side chain atom distance from the corresponding voxel center to the nearest neighbor side chain atom of the corresponding reference amino acid.
[0356] In some embodiments, the system may be further configured to encode, in a tensor, a nearest neighbor atom channel that specifies the distance from each voxel to a nearest neighbor atom. The nearest neighbor atom can be selected regardless of the amino acid and the atomic elements of the amino acid. The distance may be a Euclidean distance and can be normalized by dividing the Euclidean distance by a maximum distance. The amino acids can include non-standard amino acids. The tensor can further include an absent atom channel that specifies atoms not found within a predetermined radius of a voxel center. The absent atom channel can be one-hot encoded.
[0357] In some embodiments, the system can further comprise a reference allele encoder that encodes reference allele amino acids for each three-dimensional distance value, in voxel units, by amino acid position. The reference allele amino acids can be a three-dimensional representation of the one-hot encoding of a reference amino acid sequence. The amino acid-specific conservation frequency can identify the conservation level of each amino acid across multiple species.
[0358] In some embodiments, the evolutionary conservation encoder can select the nearest neighbor atoms to the corresponding voxels across reference amino acids and atom categories, select the pan-amino acid conservation frequency for the residues of the reference amino acids containing the nearest neighbor atoms, and use the three-dimensional representation of the pan-amino acid conservation frequency as the evolutionary conservation sequence. The pan-amino acid conservation frequency can be configured for specific positions of residues as observed across multiple species. The pan-amino acid conservation frequency can identify whether there is a missing conservation frequency for a particular reference amino acid.
[0359] In some embodiments, the evolutionary conservation encoder can select the nearest neighbor atoms to the corresponding voxels for each of the reference amino acids, select the conservation frequency for each amino acid for each of the residues of the reference amino acids containing the nearest neighbor atoms, and use the three-dimensional representation of the conservation frequency for each amino acid as the evolutionary conservation sequence. The conservation frequency for each amino acid can be configured for specific positions of residues as observed across multiple species. The conservation frequency for each amino acid can identify whether there is a missing conservation frequency for a particular reference amino acid.
[0360] In some embodiments, the system can further comprise an annotation encoder that encodes one or more annotation channels in voxel units for each three-dimensional distance value. The annotation channel can be a three-dimensional representation of the one-hot encoding of residue annotations and can be a molecular processing annotation including initiator methionine, signal, transport peptide, propeptide, chain, and peptide. In some embodiments, the annotation channel can be a region annotation including topological domain, transmembrane, in-membrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias, or can be a site annotation including active site, metal binding, binding site, and site. In some embodiments, the annotation channel can be an amino acid modification annotation including non-standard residue, modified residue, lipidation, glycosylation, disulfide bond, and cross-link, or can be a secondary structure annotation including helix, turn, and beta strand. The annotation channel can be an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residue, and non-terminal residue.
[0361] In some embodiments, the system can further comprise a structure reliability encoder that encodes one or more structure reliability channels in voxel units for each three-dimensional distance value. The structure reliability channel can be a three-dimensional representation of a confidence score that identifies the quality of each residue structure. The structure reliability channel can be a global model quality estimation (GMQE), can be a qualitative model energy analysis (QMEAN) score, can be a temperature factor that identifies the extent to which a residue satisfies the physical constraints of its respective protein structure, can be a template structure alignment that identifies the extent to which the residues of the nearest neighbor atoms to a voxel have an aligned template structure, can be a template modeling score of the aligned template structure, or can be the minimum one of the template modeling scores, the average of the template modeling scores, and the maximum one of the template modeling scores.
[0362] In some embodiments, the system can further comprise an atomic rotation engine that rotates atoms before the distance channels of amino acid units are generated.
[0363] The convolutional neural network can use 1x1x1 convolutions, 3×3×3 convolutions, a normalized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer. The 1x1x1 convolution and the 3×3×3 convolution can be three-dimensional convolutions. In some embodiments, the 1×1×1 convolution layer can process a tensor and generate an intermediate output that is a convolutional representation of the tensor. A sequence of 3×3×3 convolution layers can process the intermediate output and generate a flattened output. The fully connected layer can process the flattened output and generate an unnormalized output. The softmax classification layer can process the unnormalized output and generate an exponentially normalized output that discriminates the likelihood that the mutant nucleotide is pathogenic and benign.
[0364] In some embodiments, the sigmoid layer can process the unnormalized output and generate a normalized output that discriminates the likelihood that the mutant nucleotide is pathogenic. The convolutional neural network can be an attention-based neural network. The tensor can include a distance channel of amino acid units further encoded by reference allele amino acids, can include a distance channel of amino acid units further encoded by an annotation channel, or can include a distance channel of amino acid units further encoded by a structural confidence channel.
[0365] In some embodiments, the system can include a voxelizer that accesses the three-dimensional structure of a reference amino acid sequence of a protein and fits a three-dimensional grid of voxels over the atoms in the three-dimensional structure on an amino acid basis to generate distance channels per atom category. The atoms span a plurality of atom categories that identify atomic elements of the amino acids. Each distance channel per atom category has three-dimensional distance values for each voxel in the three-dimensional grid of voxels. The three-dimensional distance values specify the distance from the corresponding voxel in the three-dimensional grid of voxels to the atoms of the corresponding atom category within the plurality of atom categories. The system further includes an alternative allele encoder that encodes alternative allele amino acids for each voxel in the three-dimensional grid of voxels. The alternative allele amino acids are a three-dimensional representation of a one-hot encoding of mutant amino acids expressed by mutant nucleotides. The system further includes an evolutionary conservation encoder that encodes an evolutionary conservation sequence for each voxel in the three-dimensional grid of voxels. The evolutionary conservation sequence can be a three-dimensional representation of amino acid-specific conservation frequencies across a plurality of species. The amino acid-specific conservation frequencies can be selected according to the amino acid proximity to the corresponding voxel. The system further includes a convolutional neural network configured to apply three-dimensional convolution to a tensor including the alternative allele amino acids and the distance channels per atom category encoded with their respective evolutionary conservation sequences and to determine the pathogenicity of the mutant nucleotides based at least in part on the tensor.
[0366] In some embodiments, the system includes a voxelizer that accesses the three-dimensional structure of the reference amino acid sequence of the protein and fits a three-dimensional grid of voxels to the atoms within the three-dimensional structure on an amino acid basis to generate a distance channel for amino acid units. Each of the distance channels for amino acid units can have three-dimensional distance values for each voxel within the three-dimensional grid of voxels. The three-dimensional distance value can specify the distance from the corresponding voxel within the three-dimensional grid of voxels to the atoms of the corresponding reference amino acid within the reference amino acid sequence. The system further includes an alternative allele encoder that encodes alternative allele amino acids for each voxel within the three-dimensional grid of voxels. The alternative allele amino acid is a three-dimensional representation of the one-hot encoding of the mutant amino acid expressed by the mutant nucleotide. The system further comprises an evolutionary conservation encoder that encodes an evolutionary conservation sequence for each voxel within the three-dimensional grid of voxels. The evolutionary conservation sequence can be a three-dimensional representation of the amino acid-specific conservation frequency across multiple species. The amino acid-specific conservation frequency can be selected according to the amino acid proximity to the corresponding voxel. The system further includes a tensor generator configured to generate a tensor including the alternative allele amino acids and the distance channels for amino acid units encoded with their respective evolutionary conservation sequences.
[0367] In some embodiments, the system includes a voxelizer that accesses the three-dimensional structure of a reference amino acid sequence of a protein and fits a three-dimensional grid of voxels to the atoms in the three-dimensional structure on an amino acid basis to generate a distance channel in atomic category units. Atoms can span a plurality of atomic categories that identify atomic elements of amino acids. Each of the distance channels in atomic category units can have a three-dimensional distance value for each voxel in the three-dimensional grid of voxels. The three-dimensional distance value can specify the distance from the corresponding voxel in the three-dimensional grid of voxels to the atoms of the corresponding atomic category in the plurality of atomic categories. The system further includes an alternative allele encoder that encodes alternative allele amino acids for each voxel in the three-dimensional grid of voxels. The alternative allele amino acids are a three-dimensional representation of the one-hot encoding of mutant amino acids expressed by mutant nucleotides. The system further comprises an evolutionary conservation encoder that encodes an evolutionary conservation sequence for each voxel in the three-dimensional grid of voxels. The evolutionary conservation sequence can be a three-dimensional representation of amino acid-specific conservation frequencies across a plurality of species. The amino acid-specific conservation frequencies can be selected according to the amino acid proximity to the corresponding voxel. The system further includes a tensor generator configured to generate a tensor including the alternative allele amino acids and the distance channels in atomic category units encoded by their respective evolutionary conservation sequences.
[0368] Clause 2 1. A step of accessing the three-dimensional structure of a reference amino acid sequence of a protein and fitting a three-dimensional grid of voxels to the atoms in the three-dimensional structure on an amino acid basis to generate a distance channel in amino acid units; Each of the distance channels in amino acid units has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels, The three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atoms of the corresponding reference amino acid in the reference amino acid sequence, Encoding a substitution allele channel into each voxel within a three-dimensional grid of voxels, wherein the substitution allele channel is a three-dimensional representation of a one-hot encoding of a mutant amino acid expressed by a mutant nucleotide; Encoding an evolutionary conservation channel into each sequence of three-dimensional distance values across an amino acid unit distance channel based on voxel position, wherein the evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, and the amino acid-specific conservation frequencies are selected according to the proximity of the amino acid to the corresponding voxel; Applying a three-dimensional convolution to a tensor comprising the substitution allele channel and the amino acid unit distance channels encoded with respective evolutionary conservation channels; Determining the pathogenicity of a mutant nucleotide based at least in part on the tensor. A computer-implemented method comprising:
[0369] 2. The computer-implemented method of clause 1, further comprising centering a three-dimensional grid of voxels on the alpha-carbon atom of each residue of a reference amino acid in a reference amino acid sequence.
[0370] 3. The computer-implemented method of clause 2, further comprising centering a three-dimensional grid of voxels on the alpha-carbon atom of the residue of a specific reference amino acid corresponding to a mutant amino acid.
[0371] 4. The computer-implemented method of clause 3, further comprising encoding the directionality of a reference amino acid in a reference amino acid sequence and the position of a specific reference amino acid by multiplying a directional parameter by three-dimensional distance values for reference amino acids preceding the specific reference amino acid in the tensor.
[0372] 5. The computer-implemented method of clause 4, wherein the distance is the nearest neighbor atom distance from the corresponding voxel center within the three-dimensional grid of voxels to the nearest neighbor atom of the corresponding reference amino acid.
[0373] 6. The computer-implemented method according to clause 5, wherein the nearest neighbor atomic distance is the Euclidean distance.
[0374] 7. The computer-implemented method according to clause 6, wherein the nearest neighbor atomic distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atomic distance.
[0375] 8. The computer-implemented method according to clause 5, wherein the reference amino acid has an alpha carbon atom, and the distance is the nearest neighbor alpha carbon atom distance from the corresponding voxel center to the nearest neighbor alpha carbon atom of the corresponding reference amino acid.
[0376] 9. The computer-implemented method according to clause 5, wherein the reference amino acid has a beta carbon atom, and the distance is the nearest neighbor beta carbon atom distance from the corresponding voxel center to the nearest neighbor beta carbon atom of the corresponding reference amino acid.
[0377] 10. The computer-implemented method according to clause 5, wherein the reference amino acid has a backbone atom, and the distance is the nearest neighbor backbone atom distance from the corresponding voxel center to the nearest neighbor backbone atom of the corresponding reference amino acid.
[0378] 11. The computer-implemented method according to clause 5, wherein the amino acid has a side chain atom, and the distance is the nearest neighbor side chain atom distance from the corresponding voxel center to the nearest neighbor side chain atom of the corresponding reference amino acid.
[0379] 12. The computer-implemented method according to clause 3, further comprising encoding, in the tensor, a nearest neighbor atom channel that specifies the distance from each voxel to the nearest neighbor atom, and the nearest neighbor atom is selected regardless of the amino acid and the atomic elements of the amino acid.
[0380] 13. The computer-implemented method according to clause 12, wherein the distance is the Euclidean distance.
[0381] 14. The computer-implemented method according to clause 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0382] 15. The computer-implemented method according to clause 12, wherein the amino acid includes non-standard amino acids.
[0383] 16. The computer-implemented method according to clause 1, wherein the tensor further includes a non-present atom channel that designates atoms not found within a predetermined radius of the voxel center.
[0384] 17. The computer-implemented method according to clause 16, wherein the non-present atom channel is one-hot encoded.
[0385] 18. The computer-implemented method according to clause 1, further comprising encoding a reference allele channel in voxel units for each voxel within a three-dimensional grid of voxels.
[0386] 19. The computer-implemented method according to clause 18, wherein the reference allele amino acid is a three-dimensional representation of a one-hot encoding of a reference amino acid experiencing a mutant amino acid.
[0387] 20. The computer-implemented method according to clause 1, wherein the amino acid-specific conservation frequency identifies the conservation level of each amino acid across multiple species.
[0388] 21. Selecting the nearest neighbor atom to the corresponding voxel across reference amino acids and atom categories, and selecting a pan-amino acid conservation frequency for a residue of the reference amino acid including the nearest neighbor atom, and using a three-dimensional representation of the pan-amino acid conservation frequency as an evolutionary conservation channel. The computer-implemented method according to clause 20 further includes these steps.
[0389] 22. The computer-implemented method according to clause 21, wherein the pan-amino acid conservation frequency is configured for specific positions of residues as observed in multiple species.
[0390] 23. The computer-implemented method according to clause 21, which determines whether there is a missing conservation frequency for a specific reference amino acid in the pan-amino acid conservation frequency.
[0391] 24. The step of selecting the nearest neighbor atom for each corresponding voxel in each of the reference amino acids, and the step of selecting the conservation frequency for each amino acid for each residue of the reference amino acid containing the nearest neighbor atom, and the step of using the three-dimensional representation of the conservation frequency for each amino acid as an evolutionary conservation channel, further included in the computer-implemented method according to clause 21.
[0392] 25. The computer-implemented method according to clause 24, wherein the conservation frequency for each amino acid is configured for a specific position of a residue as observed in multiple species.
[0393] 26. The computer-implemented method according to clause 24, which determines whether there is a missing conservation frequency for a specific reference amino acid in the conservation frequency for each amino acid.
[0394] 27. Further including encoding one or more annotation channels in each voxel within the three-dimensional grid of voxels on a voxel-by-voxel basis, wherein the annotation channel is a three-dimensional representation of the one-hot encoding of the residual annotation, in the computer-implemented method according to clause 1.
[0395] 28. The computer-implemented method according to clause 27, wherein the annotation channel is a molecular processing annotation including initiator methionine, signal, transport peptide, propeptide, chain, and peptide.
[0396] 29. The computer-implemented method according to clause 27, wherein the annotation channel is a region annotation including topological domain, transmembrane, intra-membrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
[0397] 30. The annotation channel is the computer-implemented method according to clause 27, which is site annotation including active sites, metal binding, binding sites, and sites.
[0398] 31. The annotation channel is the computer-implemented method according to clause 27, which is amino acid modification annotation including non-standard residues, modified residues, lipidation, glycosylation, disulfide bonds, and cross-linking.
[0399] 32. The annotation channel is the computer-implemented method according to clause 27, which is secondary structure annotation including helices, turns, and beta strands.
[0400] 33. The annotation channel is the computer-implemented method according to clause 27, which is experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residues, and non-terminal residues.
[0401] 34. The computer-implemented method according to clause 1 further includes encoding one or more structure reliability channels in each voxel of a three-dimensional grid of voxels on a voxel-by-voxel basis, and the structure reliability channel is a three-dimensional representation of a confidence score that specifies the quality of each residue structure.
[0402] 35. The structure reliability channel is the global model quality estimation (GMQE), and it is the computer-implemented method according to clause 34.
[0403] 36. The structure reliability channel is the qualitative model energy analysis (QMEAN) score, and it is the computer-implemented method according to clause 34.
[0404] 37. The structure reliability channel is the temperature factor that specifies the degree to which residues satisfy the physical constraints of each protein structure, and it is the computer-implemented method according to clause 34.
[0405] 38. The structural reliability channel is the computer-implemented method according to clause 34, which is a template structure alignment that identifies the degree to which the residues of the nearest neighbor atoms to the voxel align the template structure.
[0406] 39. The structural reliability channel is the computer-implemented method according to clause 38, which is the template modeling score of the aligned template structure.
[0407] 40. The structural reliability channel is the computer-implemented method according to clause 39, which is the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0408] 41. The computer-implemented method according to clause 1, further comprising rotating the atoms before generating the distance channel of the amino acid units.
[0409] 42. The computer-implemented method according to clause 1, further comprising using a 1x1x1 convolution, a 3×3×3 convolution, a normalized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer in a convolutional neural network.
[0410] 43. The 1x1x1 convolution and the 3×3×3 convolution are 3D convolutions, which is the computer-implemented method according to clause 42.
[0411] 44. The 1x1x1 convolution layer processes the tensor, generates an intermediate output that is the convolutional representation of the tensor, the sequence of 3×3×3 convolution layers processes the intermediate output, generates a flattened output, the fully connected layer processes the flattened output, generates an unnormalized output, and the softmax classification layer processes the unnormalized output, generates an exponentially normalized output that identifies the likelihood that the mutant nucleotide is pathogenic and benign, which is the computer-implemented method according to clause 42.
[0412] 45. The computer-implemented method according to clause 44, wherein the sigmoid layer processes the unnormalized output and generates a normalized output that identifies the likelihood that the mutant nucleotide is pathogenic.
[0413] 46. The computer-implemented method according to clause 1, wherein the convolutional neural network is an attention-based neural network.
[0414] 47. The computer-implemented method according to clause 1, wherein the tensor includes a distance channel of amino acid units further encoded in a reference allele channel.
[0415] 48. The computer-implemented method according to clause 1, wherein the tensor includes a distance channel of amino acid units further encoded in an annotation channel.
[0416] 49. The computer-implemented method according to clause 1, wherein the tensor includes a distance channel of amino acid units further encoded in a structural confidence channel.
[0417] 50. Accessing the three-dimensional structure of the reference amino acid sequence of the protein and fitting a three-dimensional grid of voxels on the atoms in the three-dimensional structure on an amino acid basis to generate a distance channel in atomic category units; The atoms span multiple atomic categories, Among the multiple atomic categories, the atomic category identifies the atomic elements of the amino acid, Each of the distance channels in atomic category units has three-dimensional distance values for each voxel in the three-dimensional grid of voxels, The three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atoms of the corresponding atomic category in the multiple atomic categories, Encoding an alternative allele channel into each voxel in the three-dimensional grid of voxels, wherein the alternative allele channel is a three-dimensional representation of the one-hot encoding of the mutant amino acid expressed by the mutant nucleotide. Encoding an evolutionary conservation channel in each sequence of three-dimensional distance values across distance channels for atomic categories, based on voxel positions; The evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, and the amino acid-specific conservation frequencies are selected according to the proximity of amino acids to the corresponding voxels; Applying three-dimensional convolution to a tensor including a distance channel for atomic categories encoded in an alternative allele channel and each evolutionary conservation channel; Determining the pathogenicity of mutant nucleotides, at least in part based on the tensor. A computer-implemented method comprising:
[0418] Accessing the three-dimensional structure of a reference amino acid sequence of a protein, fitting a three-dimensional grid of voxels to atoms in the three-dimensional structure in amino acid units, and generating a distance channel for amino acid units; Each of the distance channels for amino acid units has three-dimensional distance values for each voxel in the three-dimensional grid of voxels; The three-dimensional distance values specify the distance from the corresponding voxel in the three-dimensional grid of voxels to the atoms of the corresponding reference amino acid in the reference amino acid sequence; Encoding an alternative allele channel in each voxel in the three-dimensional grid of voxels, wherein the alternative allele channel is a three-dimensional representation of the one-hot encoding of the mutant amino acid expressed by the mutant nucleotide; Encoding an evolutionary conservation channel in each sequence of three-dimensional distance values across the distance channel for amino acid units, based on voxel positions; The evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, and the amino acid-specific conservation frequencies are selected according to the proximity of amino acids to the corresponding voxels; Generating a tensor including a distance channel for amino acid units encoded using the alternative allele channel and each evolutionary conservation channel. A computer-implemented method comprising:
[0419] 52. Accessing the three-dimensional structure of the reference amino acid sequence of the protein, fitting a three-dimensional grid of voxels on the atoms within the three-dimensional structure on an amino acid basis, and generating a distance channel in atomic category units; Atoms span multiple atomic categories, Among the multiple atomic categories, the atomic categories identify the atomic elements of the amino acids, Each of the distance channels in atomic category units has three-dimensional distance values for each voxel within the three-dimensional grid of voxels, The three-dimensional distance value specifies the distance from the corresponding voxel within the three-dimensional grid of voxels to the atoms of the corresponding atomic category within the multiple atomic categories, Encoding an alternative allele channel in each voxel within the three-dimensional grid of voxels, wherein the alternative allele channel is a three-dimensional representation of the one-hot encoding of the mutant amino acid expressed by the mutant nucleotide; Encoding an evolutionary conservation channel based on voxel positions for each sequence of three-dimensional distance values across the distance channels in atomic category units, wherein the evolutionary conservation channel is a three-dimensional representation of the amino acid-specific conservation frequencies across multiple species, and the amino acid-specific conservation frequencies are selected according to the amino acid proximity to the corresponding voxel; Generating a tensor including the distance channels in atomic category units encoded using the alternative allele channel and each evolutionary conservation channel. A computer-implemented method comprising.
[0420] 1. One or more computer-readable media storing computer-executable instructions, which when executed on one or more processors, Accessing the three-dimensional structure of the reference amino acid sequence of the protein, fitting a three-dimensional grid of voxels to the atoms within the three-dimensional structure in amino acid units, and generating a distance channel in amino acid units; Each of the distance channels of amino acid units has three-dimensional distance values for each voxel within a three-dimensional grid of voxels, The three-dimensional distance value specifies the distance from the corresponding voxel within the three-dimensional grid of voxels to the atom of the corresponding reference amino acid in the reference amino acid sequence, Encoding an alternative allele channel for each voxel within a three-dimensional grid of voxels, wherein the alternative allele channel is a three-dimensional representation of a one-hot encoding of a mutant amino acid expressed by a mutant nucleotide, the encoding step; Encoding an evolutionary conservation channel for each sequence of three-dimensional distance values across the distance channels of amino acid units on a voxel-by-voxel basis; The evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, and the amino acid-specific conservation frequencies are selected according to the amino acid proximity to the corresponding voxel, Applying three-dimensional convolution to a tensor including the alternative allele channel and the distance channels of amino acid units encoded by each evolutionary conservation channel; Determining the pathogenicity of the mutant nucleotide based at least in part on the tensor, a computer-readable medium for configuring a computer to perform an operation including.
[0421] 2. The computer-readable medium according to clause 1, wherein the operation further includes centering the three-dimensional grid of voxels on the alpha carbon atom of each residue of the reference amino acid in the reference amino acid sequence.
[0422] 3. The computer-readable medium according to clause 2, wherein the operation further includes centering the three-dimensional grid of voxels on the alpha carbon atom of the residue of the specific reference amino acid corresponding to the mutant amino acid.
[0423] 4. The operation further includes encoding the directionality of the reference amino acids in the reference amino acid sequence and the position of a specific reference amino acid by multiplying a directionality parameter by the three-dimensional distance value of the reference amino acids preceding the specific reference amino acid in the tensor, the computer-readable medium according to clause 3.
[0424] 5. The computer-readable medium according to clause 4, wherein the distance is the nearest neighbor atomic distance from the corresponding voxel center to the nearest neighbor atom of the corresponding reference amino acid within a three-dimensional grid of voxels.
[0425] 6. The computer-readable medium according to clause 5, wherein the nearest neighbor atomic distance is the Euclidean distance.
[0426] 7. The computer-readable medium according to clause 6, wherein the nearest neighbor atomic distance is normalized by dividing the Euclidean distance by the maximum nearest neighbor atomic distance.
[0427] 8. The computer-readable medium according to clause 5, wherein the reference amino acid has an alpha carbon atom and the distance is the nearest neighbor alpha carbon atom distance from the corresponding voxel center to the nearest neighbor alpha carbon atom of the corresponding reference amino acid.
[0428] 9. The computer-readable medium according to clause 5, wherein the reference amino acid has a beta carbon atom and the distance is the nearest neighbor beta carbon atom distance from the corresponding voxel center to the nearest neighbor beta carbon atom of the corresponding reference amino acid.
[0429] 10. The computer-readable medium according to clause 5, wherein the reference amino acid has a backbone atom and the distance is the nearest neighbor backbone atom distance from the corresponding voxel center to the nearest neighbor backbone atom of the corresponding reference amino acid.
[0430] 11. The computer-readable medium according to clause 5, wherein the amino acid has a side chain atom and the distance is the nearest neighbor side chain atom distance from the corresponding voxel center to the nearest neighbor side chain atom of the corresponding reference amino acid.
[0431] 12. The operation further includes encoding, in the tensor, a nearest neighbor atom channel that specifies the distance from each voxel to the nearest neighbor atom, where the nearest neighbor atom is selected regardless of the amino acid and the atomic elements of the amino acid, the computer-readable medium according to clause 3.
[0432] 13. The computer-readable medium according to clause 12, wherein the distance is the Euclidean distance.
[0433] 14. The computer-readable medium according to clause 13, wherein the distance is normalized by dividing the Euclidean distance by the maximum distance.
[0434] 15. The computer-readable medium according to clause 12, wherein the amino acid includes non-standard amino acids.
[0435] 16. The computer-readable medium according to clause 1, wherein the tensor further includes an absent atom channel that specifies atoms not found within a predetermined radius of the voxel center.
[0436] 17. The computer-readable medium according to clause 16, wherein the absent atom channel is one-hot encoded.
[0437] 18. The computer-readable medium according to clause 1, wherein the operation further includes encoding, on a voxel-by-voxel basis, a reference allele channel for each voxel within a three-dimensional grid of voxels.
[0438] 19. The computer-readable medium according to clause 18, wherein the reference allele amino acid is a three-dimensional representation of the one-hot encoding of the reference amino acid experiencing the mutant amino acid.
[0439] 20. The computer-readable medium according to clause 1, wherein the amino acid-specific conservation frequency specifies the conservation level of each amino acid across multiple species.
[0440] 21. The operation includes the step of selecting the nearest neighbor atoms to the corresponding voxels across the reference amino acids and atomic categories, and selecting a pan - amino acid conservation frequency for residues of a reference amino acid including the nearest neighbor atoms; using a three - dimensional representation of the pan - amino acid conservation frequency as an evolutionary conservation channel, further comprising the computer - readable medium according to clause 20.
[0441] 22. The computer - readable medium according to clause 21, wherein the pan - amino acid conservation frequency is configured for specific positions of residues as observed in multiple species.
[0442] 23. The computer - readable medium according to clause 21, wherein the pan - amino acid conservation frequency identifies whether there is a missing conservation frequency for a specific reference amino acid.
[0443] 24. The operation further includes selecting each nearest neighbor atom to a corresponding voxel for each of the reference amino acids, selecting a conservation frequency for each amino acid for each residue of the reference amino acid including the nearest neighbor atoms, using a three - dimensional representation of the conservation frequency for each amino acid as an evolutionary conservation channel, further comprising the computer - readable medium according to clause 21.
[0444] 25. The computer - readable medium according to clause 24, wherein the conservation frequency for each amino acid is configured for specific positions of residues as observed in multiple species.
[0445] 26. The computer - readable medium according to clause 24, wherein the conservation frequency for each amino acid identifies whether there is a missing conservation frequency for a specific reference amino acid.
[0446] 27. The operation further includes encoding, on a per - voxel basis, one or more annotation channels in each voxel within a three - dimensional grid of voxels, the annotation channel being a three - dimensional representation of a one - hot encoding of a residual annotation, the computer - readable medium according to clause 1.
[0447] 28. The computer-readable medium according to clause 27, wherein the annotation channel is a molecular processing annotation including initiator methionine, signal, transport peptide, propeptide, chain, and peptide.
[0448] 29. The computer-readable medium according to clause 27, wherein the annotation channel is a region annotation including topological domain, transmembrane, intra-membrane, domain, repeat, calcium binding, zinc finger, deoxyribonucleic acid (DNA) binding, nucleotide binding, region, coiled coil, motif, and compositional bias.
[0449] 30. The computer-readable medium according to clause 27, wherein the annotation channel is a site annotation including active site, metal binding, binding site, and site.
[0450] 31. The computer-readable medium according to clause 27, wherein the annotation channel is an amino acid modification annotation including non-standard residue, modified residue, lipidation, glycosylation, disulfide bond, and cross-linking.
[0451] 32. The computer-readable medium according to clause 27, wherein the annotation channel is a secondary structure annotation including helix, turn, and beta strand.
[0452] 33. The computer-readable medium according to clause 27, wherein the annotation channel is an experimental information annotation including mutagenesis, sequence uncertainty, sequence conflict, non-adjacent residue, and non-terminal residue.
[0453] 34. The computer-readable medium according to clause 1, wherein the operation further includes encoding one or more structure reliability channels in each voxel within a three-dimensional grid of voxels on a voxel-by-voxel basis, and the structure reliability channel is a three-dimensional representation of a confidence score specifying the quality of each residue structure.
[0454] 35. The structural reliability channel is the computer-readable medium according to clause 34, which is the global model quality estimation (GMQE).
[0455] 36. The structural reliability channel is the computer-readable medium according to clause 34, which is the qualitative model energy analysis (QMEAN) score.
[0456] 37. The structural reliability channel is the computer-readable medium according to clause 34, which is the temperature factor that specifies the degree to which residues satisfy the physical constraints of their respective protein structures.
[0457] 38. The structural reliability channel is the computer-readable medium according to clause 34, which is the template structure alignment that specifies the degree to which the residues of the nearest neighbor atoms to the voxel align the template structure.
[0458] 39. The structural reliability channel is the computer-readable medium according to clause 38, which is the template modeling score of the aligned template structure.
[0459] 40. The structural reliability channel is the computer-readable medium according to clause 39, which is the minimum of the template modeling scores, the average of the template modeling scores, and the maximum of the template modeling scores.
[0460] 41. The operation further includes rotating atoms before the distance channel of amino acid units is generated, in the computer-readable medium according to clause 1.
[0461] 42. The operation further includes using a 1x1x1 convolution, a 3×3×3 convolution, a normalized linear unit activation layer, a batch normalization layer, a fully connected layer, a dropout regularization layer, and a softmax classification layer in a convolutional neural network, in the computer-readable medium according to clause 1.
[0462] 43. The 1x1x1 convolution and the 3×3×3 convolution are 3D convolutions, in the computer-readable medium according to clause 42.
[0463] A layer of 44.1x1x1 convolution processes a tensor and generates an intermediate output that is a convolutional representation of the tensor. A sequence of 3×3×3 convolution layers processes the intermediate output and generates a flattened output. A fully connected layer processes the flattened output and generates an unnormalized output. A softmax classification layer processes the unnormalized output and generates an exponentially normalized output that discriminates the likelihood that a mutant nucleotide is pathogenic and benign. The computer-readable medium according to clause 42.
[0464] 45. The computer-readable medium according to clause 44, wherein a sigmoid layer processes the unnormalized output and generates a normalized output that identifies the likelihood that a mutant nucleotide is pathogenic.
[0465] 46. The computer-readable medium according to clause 1, wherein the convolutional neural network is an attention-based neural network.
[0466] 47. The computer-readable medium according to clause 1, wherein the tensor includes a distance channel of amino acid units further encoded in a reference allele channel.
[0467] 48. The computer-readable medium according to clause 1, wherein the tensor includes a distance channel of amino acid units further encoded in an annotation channel.
[0468] 49. The computer-readable medium according to clause 1, wherein the tensor includes a distance channel of amino acid units further encoded in a structural confidence channel.
[0469] 50. One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, access the three-dimensional structure of a reference amino acid sequence of a protein, fit a three-dimensional grid of voxels on atoms in the three-dimensional structure on an amino acid basis to generate a distance channel in atomic category units, and The atoms span multiple atom categories, wherein the atom categories among the multiple atom categories identify atomic elements of amino acids, each of the distance channels for atom category units has three-dimensional distance values for each voxel within a three-dimensional grid of voxels, the three-dimensional distance values specify the distances from corresponding voxels within the three-dimensional grid of voxels to the atoms of the corresponding atom categories within the multiple atom categories, encoding an alternative allele channel for each voxel within a three-dimensional grid of voxels, wherein the alternative allele channel is a three-dimensional representation of a one-hot encoding of a mutant amino acid expressed by a mutant nucleotide; encoding an evolutionary conservation channel on a voxel-position basis for each sequence of three-dimensional distance values across the distance channels for atom category units, wherein the evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, and the amino acid-specific conservation frequencies are selected according to amino acid proximity to the corresponding voxels; applying three-dimensional convolution to a tensor including the alternative allele channel and the distance channels for atom category units encoded with respective evolutionary conservation channels; determining the pathogenicity of a mutant nucleotide based at least in part on the tensor; a computer-readable medium configuring a computer to perform an operation including these steps.
[0470] One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, access the three-dimensional structure of a reference amino acid sequence of a protein, fit a three-dimensional grid of voxels to the atoms within the three-dimensional structure in amino acid units, and generate a distance channel for amino acid units; each of the distance channels for amino acid units has three-dimensional distance values for each voxel within a three-dimensional grid of voxels, The three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atom of the corresponding reference amino acid in the reference amino acid sequence, encoding an alternative allele channel for each three-dimensional distance value of the distance channel for each amino acid unit according to the amino acid position, and the alternative allele channel is a three-dimensional representation of the one-hot encoding of the mutant amino acid expressed by the mutant nucleotide, encoding an evolutionary conservation channel for each sequence of three-dimensional distance values across the distance channel for each amino acid unit based on the voxel position, and the evolutionary conservation channel is a three-dimensional representation of the amino acid-specific conservation frequency across multiple species, and the amino acid-specific conservation frequency is selected according to the amino acid proximity to the corresponding voxel, generating a tensor including the distance channel for each amino acid unit encoded using the alternative allele channel and each respective evolutionary conservation channel. A computer-readable medium configuring a computer to perform operations including these steps.
[0471] One or more computer-readable media storing computer-executable instructions that, when executed on one or more processors, access the three-dimensional structure of the reference amino acid sequence of the protein, fit a three-dimensional grid of voxels on the atoms in the three-dimensional structure on an amino acid basis to generate a distance channel for each atom category, and the atoms span multiple atom categories, the atom categories among the multiple atom categories identify the atomic elements of the amino acid, each of the distance channels for each atom category has a three-dimensional distance value for each voxel in the three-dimensional grid of voxels, the three-dimensional distance value specifies the distance from the corresponding voxel in the three-dimensional grid of voxels to the atom of the corresponding atom category in the multiple atom categories, Encoding a substitution allele channel into each voxel within a three-dimensional grid of voxels, wherein the substitution allele channel is a three-dimensional representation of a one-hot encoding of a mutant amino acid expressed by a mutant nucleotide; Encoding an evolutionary conservation channel voxel-position based into each sequence of three-dimensional distance values across a distance channel in atomic category units, wherein the evolutionary conservation channel is a three-dimensional representation of amino acid-specific conservation frequencies across multiple species, and the amino acid-specific conservation frequencies are selected according to amino acid proximity to the corresponding voxel; Generating a tensor including a distance channel in atomic category units encoded using the substitution allele channel and respective evolutionary conservation channels; and configuring a computer for performing operations including the foregoing. A computer-readable medium
[0472] Other embodiments of the methods described in this section can include a non-transitory computer-readable storage medium storing instructions executable by a processor for performing any of the methods described above. Still other embodiments of the methods described in this section can include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the methods described above.
[0473] Specific Embodiment 3 Clause 3 1. A computer-implemented method for efficiently determining which elements of a sequence are nearest neighbors to evenly-spaced cells within a grid, where the elements have element coordinates and the cells have cell indices and cell coordinates in dimensional units, the method including generating an element-to-cell mapping that maps a subset of cells to each of the elements, where the subset of cells mapped to a particular element in the sequence includes the nearest neighbor cell and one or more neighboring cells within the grid; The nearest neighbor cell is selected based on matching the element coordinates of the particular element to the cell coordinates; The neighboring cells are continuously adjacent to the nearest neighbor cell and are selected based on being within a distance proximity range from a specific element, and for each of the cells, generating a cell-to-element mapping that maps a subset of the elements, wherein the subset of elements mapped to a specific cell within the grid includes the elements within the sequence that are mapped to the specific cell by the element-to-cell mapping, step, Using the cell-to-element mapping to determine, for each of the cells, the nearest neighbor element within the sequence, step, The nearest neighbor element to a specific cell is determined based on the distance between the specific cell and the elements within the subset of elements, and a computer-implemented method including the step.
[0474] 2. The computer-implemented method according to clause 1, wherein matching the element coordinates of a specific element to cell coordinates further includes generating truncated element coordinates by truncating the fractional part of the element coordinates.
[0475] 3. Matching the element coordinates of a specific element to cell coordinates, For the first dimension, matching the first truncated element coordinate within the truncated element coordinates to the first cell coordinate of the first cell within the grid and selecting the first dimension index of the first cell, step, For the second dimension, matching the second truncated element coordinate within the truncated element coordinates to the second cell coordinate of the second cell within the grid and selecting the second dimension index of the second cell, step, For the third dimension, matching the third truncated element coordinate within the truncated element coordinates to the third cell coordinate of the third cell within the grid and selecting the third dimension index of the third cell, step, Using the selected first, second, and third dimension indices to generate a cumulative sum based on weighting the selected first, second, and third dimension indices by powers of the base in position units, step, Using the cumulative sum as a cell index for the selection of the nearest neighbor cell, and the computer-implemented method according to clause 2 further including the step.
[0476] 4. The distance is calculated between the cell coordinates of a specific cell and the element coordinates of elements within a subset of elements, the computer-implemented method according to claim 1.
[0477] 5. The computer-implemented method according to claim 1, wherein the array is a protein sequence of amino acids.
[0478] 6. The computer-implemented method according to claim 5, wherein the elements are atoms of amino acids.
[0479] 7. A step of generating a mapping from elements to cells, a step of generating a mapping from cells to elements, and using the mapping from cells to elements, for each of the cells, to determine that the nearest neighbor elements have a runtime complexity of O(a * f + v), the method including: where a is the number of atoms, where f is the number of amino acids, where v is the number of cells, * which is a multiplication operation, the computer-implemented method according to claim 6.
[0480] 8. The computer-implemented method according to claim 7, wherein the atoms include alpha carbon atoms.
[0481] 9. The computer-implemented method according to claim 7, wherein the atoms include beta carbon atoms.
[0482] 10. The computer-implemented method according to claim 7, wherein the atoms include non-carbon atoms.
[0483] 11. The computer-implemented method according to claim 1, wherein the cells are 3D voxels.
[0484] 12. The computer-implemented method according to claim 11, wherein the cell coordinates are 3D coordinates.
[0485] 13. The computer-implemented method according to claim 12, wherein the element coordinates are 3D coordinates.
[0486] 14. The computer-implemented method according to claim 1, wherein the neighboring cells are selected based on being within an index neighborhood range from the nearest neighbor cells.
[0487] 15. The computer-implemented method according to claim 1, wherein the neighborhood cells are selected based on being within a cell neighborhood within a grid including the nearest neighbor cells.
[0488] 16. The computer-implemented method according to claim 1, wherein the sequence includes M elements, a subset of the elements includes N elements, and M >> N.
[0489] 17. A computer-implemented method for efficiently determining which atoms in a protein are nearest neighbors to voxels in a grid, wherein the atoms have three-dimensional (3D) atomic coordinates, the voxels have 3D voxel coordinates, and the method comprises: generating a mapping from atoms to voxels that maps inclusion voxels selected based on matching the 3D atomic coordinates of a particular atom of the protein to 3D voxel coordinates in the grid for each of the atoms; for each of the voxels, generating a mapping from voxels to atoms that maps a subset of the atoms, wherein the subset of the atoms mapped to a particular voxel in the grid includes the atoms in the protein mapped to the particular voxel by the mapping from atoms to voxels; and determining the nearest neighbor atoms in the protein for each of the voxels using the mapping from voxels to atoms.
[0490] 18. The computer-implemented method according to claim 17, wherein the steps of claim 17 have a runtime complexity of O(number of atoms).
[0491] Other embodiments of the methods described in this section can include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the methods described above. Yet another embodiment of the methods described in this section can include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the methods described above.
[0492] The present invention has been disclosed with reference to the preferred embodiments and examples described above, and it should be understood that these examples are intended in an illustrative rather than a limiting sense. Those skilled in the art will readily occur to changes and combinations, and such changes and combinations are considered to be within the spirit of the present invention and the scope of the following claims.
Claims
1. A system comprising: a memory storing distance channels for amino acid units of a plurality of amino acids within an amino acid sequence of a protein, each of the distance channels for the amino acid units having distance values for voxels within a plurality of voxels, the distance values for the voxels specifying distances from corresponding voxels within the plurality of voxels to atoms of corresponding amino acids within the plurality of amino acids; a neural network-based variant pathogenicity classifier executed on at least one processor coupled to the memory, the variant pathogenicity classifier processing, as input, a tensor including the distance channels for the amino acid units and alternative allele amino acids of a protein expressed by a variant, and being trained to classify the variant as benign or pathogenic based at least in part on the tensor.
2. The system of claim 1, further comprising a distance channel generator executed on at least one processor coupled to the memory, the distance channel generator calculating the distance values for the voxels by centering a voxel grid of the voxels on alpha carbon atoms of respective residues of the corresponding amino acids and specifying distances between centers of the voxels within the voxel grid and the atoms of the corresponding amino acids.
3. The system of claim 2, wherein the distance channel generator centers the voxel grid on an alpha carbon atom of a residue of a specific amino acid corresponding to at least one variant amino acid in the protein.
4. The system of claim 3, further configured to encode a directionality of the corresponding amino acid and a position of the specific amino acid by multiplying, in the tensor, distance values for voxels of a preceding amino acid preceding the specific amino acid by using a directionality parameter.
5. The system of claim 3, wherein the distance is a nearest neighbor atom distance from a corresponding voxel center within the voxel grid to a nearest neighbor atom of the corresponding amino acid.
6. The system according to claim 5, wherein the corresponding amino acid has an alpha carbon atom, and the distance is the nearest alpha carbon atom distance from the center of the corresponding voxel to the nearest alpha carbon atom of the corresponding amino acid.
7. The system according to claim 5, wherein the corresponding amino acid has a beta carbon atom, and the distance is the nearest beta carbon atom distance from the center of the corresponding voxel to the nearest beta carbon atom of the corresponding amino acid.
8. The system according to claim 5, wherein the corresponding amino acid has a backbone atom, and the distance is the nearest backbone atom distance from the center of the corresponding voxel to the nearest backbone atom of the corresponding amino acid.
9. The system according to claim 3, further configured to encode a nearest neighbor atom channel that specifies the distance from each voxel to the nearest neighbor atom, wherein the nearest neighbor atom is selected regardless of the amino acid and the atomic elements of the amino acid.
10. The system according to claim 1, wherein the tensor further includes an absent atom channel that specifies atoms not found within a predetermined radius of the voxel center, and the absent atom channel is one-hot encoded.
11. The system according to claim 1, wherein the tensor further includes a one-hot encoding of the alternative allele amino acids encoded in voxel units for each of the distance channels of the amino acid units.
12. The system according to claim 1, wherein the tensor further includes a one-hot encoding of the reference allele amino acids encoded in voxel units for each of the distance channels of the amino acid units.
13. The system according to claim 1, wherein the tensor further includes an evolutionary profile that specifies the conservation level of the corresponding amino acids across a plurality of species having sequences homologous to the amino acid sequence of the protein.
14. An evolutionary profile generator executed on at least one processor coupled to the memory, for each of the voxels, using multiple sequence alignment to determine the pan-amino acid conservation frequency, selecting nearest neighbor atoms across the plurality of amino acids and atom categories, selecting a pan-amino acid conservation frequency sequence for residues of the amino acids containing the nearest neighbor atoms, voxelizing the pan-amino acid conservation frequency for the residues of the amino acids, and The system according to claim 13, further comprising an evolutionary profile generator that makes the pan - amino acid conservation frequency array available as one of the evolutionary profiles.
15. The system according to claim 14, wherein the pan - amino acid conservation frequency array is constructed for specific positions of residues as observed in the plurality of species.
16. For each of the voxels, the evolutionary profile generator uses the multiple sequence alignment to determine the conservation frequency for each amino acid, selects the respective nearest - neighbor atoms in each of the plurality of amino acids, for each residue of the plurality of amino acids containing the nearest - neighbor atoms, selects the conservation frequency for each amino acid, voxelizes the conservation frequency for each amino acid for each of the respective residues of the plurality of amino acids and makes the conservation frequency for each amino acid available as one of the evolutionary profiles. The system according to claim 14.
17. The system according to claim 16, wherein the conservation frequency for each amino acid is constructed for specific positions of residues as observed in the plurality of species.
18. The tensor is as follows: an annotation channel for the corresponding amino acid that annotates the features of the corresponding amino acid and is one - hot encoded in the tensor, or a structural reliability channel for the corresponding amino acid that specifies the quality of the structure of each of the corresponding amino acids The system according to claim 1, further comprising one or more of these.
19. A system comprising: a memory that stores an atom - category unit distance channel for atoms in amino acids of a protein, wherein the amino acids have atoms of a plurality of atom categories, the atom categories among the plurality of atom categories identify the atomic elements of the amino acids, each of the atom - category unit distance channels has voxel - unit distance values for voxels within a plurality of voxels, the voxel - unit distance values specify the distance from the corresponding voxels within the plurality of voxels to the atoms within the corresponding atom categories among the plurality of atom categories, and a memory; a neural - network - based mutant pathogenicity classifier executed on at least one processor coupled to the memory, wherein the mutant pathogenicity classifier processing, as input, a tensor comprising the distance channel of the atomic category unit and the alternative allele amino acids of the protein expressed by the variant; a neural network-based variant pathogenicity classifier trained to classify the variant as benign or pathogenic based at least in part on the tensor. **Claim 20** storing distance channels of amino acid units for a plurality of amino acids within an amino acid sequence of a protein, each of the distance channels of the amino acid units having distance values per voxel for voxels within a plurality of voxels, the distance values per voxel specifying distances from corresponding voxels within the plurality of voxels to atoms of corresponding amino acids within the plurality of amino acids; processing, as input, a tensor comprising the distance channel of the amino acid unit and an alternative allele of a protein expressed by a variant; classifying the variant as benign or pathogenic based at least in part on the tensor. A computer-implemented method comprising: A system comprising:
Citation Information
Patent Citations
Deep Convolutional Neural Networks for Variant Classification
JP2020530918A
Artificial Intelligence-Based Quality Scoring
JP2022524562A
Multi-channel protein voxelization to predict variant pathogenicity using deep convolutional neural networks
JP2024513995A
JPP6834029B
Systems and methods for applying a convolutional network to spatial data
US20160196480A1