Image-based variant pathogenicity determination

JP2025504942A5Pending Publication Date: 2026-02-06ILLUMINA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024544789
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-01-28
Filing Date
2023-01-27
Publication Date
2026-02-06

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Techniques for classifying protein structures, such as techniques for classifying the pathogenicity of a protein structure associated with a nucleotide variant, are described herein. Such classification is based on a two-dimensional image obtained from a three-dimensional image of the protein structure. For some embodiments, a multi-view convolutional neural network (CNN) for classifying a protein structure based on an input of a two-dimensional image obtained from a three-dimensional image of the protein structure is described herein. In some embodiments, a computer-implemented method for determining pathogenicity of a variant includes accessing a structural rendition of amino acids, capturing an image of a portion of the structural rendition that includes a target amino acid from the amino acids, and determining the pathogenicity of a nucleotide variant that mutates the target amino acid to an alternative amino acid based on the image.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims the benefit of and priority to U.S. Provisional Patent Application No. 63 / 304,544, entitled "IMAGE-BASED VARIANT PATHOGENICITY DETERMINATION," filed January 28, 2022, which is incorporated herein by reference in its entirety.

[0002] The disclosed technology relates to artificial intelligence-based computers and digital data processing systems and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using techniques for transforming the context of an artificial neural network (ANN) or another type of computing system that is trainable through machine learning.

[0003] Additionally, the disclosed technology relates to pre-processing of inputs for artificial intelligence type computers and digital data processing systems, corresponding data processing methods, products for emulation of intelligence, as well as the actual pre-processed inputs themselves.

[0004] Built-in The following are incorporated by reference for all purposes as if fully set forth herein: U.S. Provisional Patent Application No. 63 / 253,122, entitled “PROTEIN STRUCTURE-BASED PROTEIN LANGUAGE MODELS,” filed on October 6, 2021 (Attorney Docket No. ILLM1050-1 / IP-2164-PRV); U.S. Provisional Patent Application No. 63 / 281,579, entitled “PREDICTING VARIANT PATHOGENICITY FROM EVOLUTIONARY CONSERVATION USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURE VOXELS,” filed on November 19, 2021 (Attorney Docket No. ILLM1060-1 / IP-2270-PRV); U.S. Provisional Patent Application No. 63 / 281,592, entitled “COMBINED AND TRANSFER LEARNING OF A VARIANT PATHOGENICITY PREDICTOR USING GAPED AND NON-GAPED PROTEIN SAMPLES,” filed on November 19, 2021 (Attorney Docket No. ILLM1061-1 / IP-2271-PRV); U.S. patent application Ser. No. 62 / 573,144, entitled “TRAINING A DEEP PATHOGENICITY CLASSIFIER USING LARGE-SCALE BENIGN TRAINING DATA,” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-1 / IP-1611-PRV); U.S. Patent Application No. 62 / 573,149, entitled “PATHOGENICITY CLASSIFIER BASED ON DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs),” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV); U.S. patent application Ser. No. 62 / 573,153, entitled “DEEP SEMI-SUPERVISED LEARNING THAT GENERATES LARGE-SCALE PATHOGENIC TRAINING DATA,” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-3 / IP-1613-PRV); U.S. Patent Application No. 62 / 582,898, entitled “PATHOGENICITY CLASSIFICATION OF GENOMIC DATA USING DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs),” filed on November 7, 2017 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV); U.S. Patent Application No. 16 / 160,903, entitled “DEEP LEARNING-BASED TECHNIQUES FOR TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on October 15, 2018 (Attorney Docket No. ILLM1000-5 / IP-1611-US); U.S. Patent Application Serial No. 16 / 160,986, entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS FOR VARIANT CLASSIFICATION," filed on October 15, 2018 (Attorney Docket No. ILLM 1000-6 / IP-1612-US); U.S. Patent Application Serial No. 16 / 160,968, entitled “SEMI-SUPERVISED LEARNING FOR TRAINING AN ENSEMBLE OF DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on October 15, 2018 (Attorney Docket No. ILLM 1000-7 / IP-1613-US); U.S. Patent Application No. 16 / 407,149, entitled “DEEP LEARNING-BASED TECHNIQUES FOR PRE-TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed May 8, 2019 (Attorney Docket No. ILLM 1010-1 / IP-1734-US); U.S. Patent Application No. 17 / 232,056, entitled “DEEP CONVOLUTIONAL NEURAL NETWORKS TO PREDICT VARIANT PATHOGENICITY USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURES,” filed on April 15, 2021 (Attorney Docket No. ILLM 1037-2 / IP-2051-US); U.S. Patent Application No. 63 / 175,495, entitled “MULTI-CHANNEL PROTEIN VOXELIZATION TO PREDICT VARIANT PATHOGENICITY USING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on April 15, 2021 (Attorney Docket No. ILLM 1047-1 / IP-2142-PRV); Sundaram,L.et al.Predicting the clinical impact of human mutation with deep neural networks.Nat.Genet.50,1161-1170(2018), Jaganathan,K.et al.Predicting splicing from primary sequence with deep learning.Cell 176,535-548(2019), U.S. Patent Application No. 63 / 175,767, entitled “EFFICIENT VOXELIZATION FOR DEEP LEARNING,” filed on April 16, 2021 (Attorney Docket No. ILLM 1048-1 / IP-2143-PRV); and U.S. Patent Application No. 17 / 468,411, entitled “ARTIFICIAL INTELLIGENCE-BASED ANALYSIS OF PROTEIN THREE-DIMENSIONAL (3D) STRUCTURES,” filed on September 7, 2021 (Attorney Docket No. ILLM 1037-3 / IP-2051A-US). [Background technology]

[0005] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to embodiments of the claimed technology.

[0006] Genomics in the broad sense, also called functional genomics, aims to characterize the function of all genomic elements of an organism by using genome-scale assays such as genome sequencing, transcriptome profiling, and proteomics. Genomics has emerged as a data-driven science and operates not by testing preconceived models and hypotheses, but by discovering novel properties from the exploration of genome-scale data. Applications of genomics include finding associations between genotypes and phenotypes, discovering biomarkers for patient stratification, predicting gene functions, and mapping biochemically active genomic regions such as transcriptional enhancers.

[0007] Genomics data is too large and complex to be mined solely by visual inspection of pairwise correlations. Instead, analytical tools are needed to support the discovery of unexpected relationships, derive novel hypotheses and models, and make predictions. Unlike some algorithms in which assumptions and domain expertise are hard-coded, machine learning algorithms are designed to automatically detect patterns in data. Thus, machine learning algorithms are well suited for data-driven science, especially genomics. However, the performance of machine learning algorithms can be highly dependent on how the data is represented, i.e., how each variable (also called a feature) is calculated. For example, to classify tumors as malignant or benign from a fluorescent microscopy image, a pre-processing algorithm can detect cells, identify cell types, and generate a list of cell counts for each cell type.

[0008] The machine learning model can take estimated cell counts, which are an example of hand-designed features, as input features to classify tumors. The central problem is that classification performance is highly dependent on the quality and relevance of these features. For example, relevant visual features such as cell morphology, distance between cells or localization within an organ are not captured in the cell counts, and this incomplete representation of the data can reduce classification accuracy.

[0009] Deep learning, a sub-discipline of machine learning, addresses this problem by embedding feature computation into the machine learning model itself, generating an end-to-end model. This result has been achieved through the development of deep neural networks, which are machine learning models that involve successive primitive operations that compute increasingly complex features by taking the results of previous operations as input. Deep neural networks can improve prediction accuracy by discovering relevant features of high complexity, such as cell morphology and spatial organization of cells in the above example. The construction and training of deep neural networks has been made possible by the explosion of data, advances in algorithms, and substantial increases in computing power, especially through the use of graphical processing units (GPUs).

[0010] The goal of supervised learning is to obtain a model that takes features as input and returns a prediction of a so-called target variable. An example of a supervised learning problem is the problem of predicting whether an intron will be spliced ​​out (target) or not, given features on the RNA such as the presence or absence of canonical splice site sequences, the location of splicing branch points, or the intron length. Training a machine learning model refers to learning its parameters, which generally involves minimizing a loss function on the training data with the goal of making accurate predictions on unknown data.

[0011] For many supervised learning problems in computational biology, the input data can be represented as a table with multiple columns or features, each of which contains numerical or categorical data that is potentially useful for making predictions. Some input data are naturally represented as tabular features (e.g., temperature or time), while other input data must first be transformed using a process called feature extraction (e.g., converting deoxyribonucleic acid (DNA) sequences to k-mer counts) to fit into a tabular representation. For intron-splicing prediction problems, the presence or absence of canonical splice site sequences, the location of splicing branch points, and intron lengths can be preprocessed features collected in tabular form. Tabular data is the norm for a wide range of supervised machine learning models, ranging from simple linear models such as logistic regression to more flexible nonlinear models such as neural networks and many others.

[0012] Logistic regression is a binary classifier, i.e., a supervised learning model that predicts a binary target variable. Specifically, logistic regression predicts the probability of a positive class by calculating a weighted sum of input features that are mapped to the [0,1] interval using a sigmoid function, a type of activation function. The parameters of logistic regression, or other linear classifiers that use different activation functions, are the weights in the weighted sum. Linear classifiers fail when the weighted sum of input features cannot sufficiently distinguish the classes, for example, whether an intron is spliced ​​out or not. To improve the prediction performance, new input features can be added manually by transforming or combining existing features in new ways, for example, by taking exponentiations or pairwise products.

[0013] Neural networks automatically learn these nonlinear feature transformations using hidden layers, each of which can be thought of as multiple linear models with outputs transformed by a nonlinear activation function such as a sigmoid function or the more common rectified-linear unit (ReLU). Together, these layers organize the input features into related complex patterns, facilitating the task of distinguishing between two classes.

[0014] Deep neural networks use many hidden layers, and when each neuron receives input from all neurons in the previous layer, the layer is said to be fully connected. Neural networks are generally trained using stochastic gradient descent, an algorithm suitable for training models on very large data sets. Implementation of neural networks using modern deep learning frameworks allows rapid prototyping with different architectures and data sets. Fully connected neural networks can be used for several genomics applications, including predicting the proportion of exons spliced ​​in for a given sequence from sequence features such as the presence of splice factor binding motifs or sequence conservation, prioritizing potentially disease-causing genetic variants, and predicting cis-regulatory elements in a given genomic region using features such as chromatin marks, gene expression, and evolutionary conservation.

[0015] For effective prediction, local dependencies in spatial and longitudinal data must be considered. For example, shuffling of DNA sequences or image pixels severely disrupts information patterns. These local dependencies set spatial or longitudinal data apart from tabular data, where feature ordering is arbitrary. Consider the problem of classifying genomic regions as bound vs. unbound by a particular transcription factor, where binding regions are defined as high-confidence binding events in chromatin immunoprecipitation following sequencing (ChIP-seq) data. Transcription factors bind to DNA by recognizing sequence motifs. Fully connected layers based on sequence-derived features such as the number of k-mer instances in a sequence or position weight matrix (PWM) matches can be used for this task. Such models can generalize well to sequences with the same motif located at different positions, because k-mer or PWM instance frequencies are robust to shifting motifs in the sequence. However, they cannot recognize patterns where transcription factor binding depends on the combination of multiple motifs with distinct intervals. Furthermore, the number of possible k-mers grows exponentially with the k-mer length, which poses both conservation and overfitting challenges.

[0016] A convolutional layer is a special form of a fully connected layer in which the same fully connected layer is applied locally, for example within a 6 bp window, to all sequence positions. This approach can also be viewed as scanning the sequence using multiple PWMs, for example for the transcription factors GATA1 and TAL1. By using the same model parameters across positions, the total number of parameters is dramatically reduced and the network can detect motifs at positions not seen during training. Each convolutional layer scans the sequence with several filters by generating a scalar value at every position that quantizes the match between the filter and the sequence. As in a fully connected neural network, a nonlinear activation function (typically ReLU) is applied at each layer. A pooling operation is then applied, which aggregates the activations in successive bins across the position axis, typically taking the maximum or average activation for each channel. Pooling reduces the effective sequence length and coarsens the signal. Subsequent convolutional layers can construct the output of the previous layer and detect whether the GATA1 and TAL1 motifs were present within a certain distance range. Finally, the output of the convolutional layers can be used as input to a fully connected neural network to perform the final prediction task. Thus, different types of neural network layers (e.g., fully connected and convolutional layers) can be combined within a single neural network.

[0017] Convolutional neural networks (CNNs) can predict various molecular phenotypes based on DNA sequence alone. Applications include classification of transcription factor binding sites, as well as prediction of molecular phenotypes such as chromatin features, DNA contact maps, DNA methylation, gene expression, translation efficiency, RBP binding, and microRNA (miRNA) targets. In addition to predicting molecular phenotypes from sequences, convolutional neural networks can be applied to more technical tasks traditionally addressed by hand-designed bioinformatics pipelines. For example, convolutional neural networks can predict guide RNA specificity, denoise ChIP-seq, improve Hi-C data resolution, predict laboratory origin from DNA sequences, and call genetic variants. Convolutional neural networks have also been used to model long-range dependencies in genomes. Although interacting regulatory elements may be located far apart on unfolded linear DNA sequences, these elements are often proximal in real 3D chromatin conformations. Thus, modeling of molecular phenotypes from linear DNA sequences can be improved by allowing long-range dependencies, albeit with a crude approximation of chromatin, and allowing the model to implicitly learn aspects of 3D organization such as promoter-enhancer loops. This is achieved by using dilated convolutions with receptive fields of up to 32 kb. Dilated convolutions also allow splice sites to be predicted from sequences using receptive fields of 10 kb, thereby allowing integration of gene sequences over distances as long as a typical human intron (see Jaganathan, K. et al. Predicting splicing from primary sequence with deep learning. Cell 176, 535-548 (2019)).

[0018] Different types of neural networks can be characterized by their parameter sharing schemes. For example, fully connected layers have no parameter sharing, while convolutional layers impose translational invariance by applying the same filter at every position of their input. Recurrent neural networks (RNNs) are an alternative to convolutional neural networks for processing sequential data such as DNA sequences or time series that implement a different parameter sharing scheme. Recurrent neural networks apply the same operation to each sequence element. This operation takes as input the memory of the previous sequence element and the new input. It updates the memory and optionally emits an output that is either passed to subsequent layers or used directly as a model prediction. By applying the same model at each sequence element, recurrent neural networks are invariant to position index in the processed sequence. For example, recurrent neural networks can detect open reading frames in DNA sequences regardless of their position in the sequence. This task requires the recognition of a specific sequence of inputs, such as a start codon followed by an in-frame stop codon.

[0019] The main advantage of recurrent neural networks over convolutional neural networks is that, in theory, they can carry over information through infinitely long sequences via memory. Furthermore, recurrent neural networks can naturally handle sequences of widely varying length, such as mRNA sequences. However, convolutional neural networks combined with various tricks (such as dilated convolutions) can reach performance comparable to or even better than recurrent neural networks for sequence modeling tasks such as audio synthesis and machine translation. Recurrent neural networks can aggregate the output of convolutional neural networks to predict single-cell DNA methylation status, RBP binding, transcription factor binding, and DNA accessibility. Furthermore, recurrent neural networks apply sequential operations, so they cannot be easily parallelized and are therefore much slower computationally than convolutional neural networks.

[0020] Although most of the human genetic code is common to all humans, each human has a unique genetic code. In some cases, the human genetic code may contain outliers, called genetic variants, that may be common among a relatively small group of individuals in the human population. For example, a particular human protein may contain a particular sequence of amino acids, but variants of that protein may differ by one amino acid in an otherwise identical particular sequence.

[0021] Gene variants can be pathogenic and can lead to disease. Most of such gene variants have been depleted from the genome by natural selection, but the ability to identify which gene variants are likely to be pathogenic can help researchers focus on these gene variants to gain understanding of the corresponding diseases and their diagnosis, treatment, or cure. The clinical interpretation of millions of human gene variants remains unclear. Some of the most frequent pathogenic variants are single nucleotide missense mutations that change the amino acid of a protein. However, not all missense mutations are pathogenic.

[0022] Models that can predict molecular phenotypes directly from biological sequences can be used as in silico perturbation tools to investigate the association between genetic and phenotypic variations and have emerged as new methods for quantitative trait locus identification and variant prioritization. These approaches are of great importance considering that the majority of variants identified by genome-wide association studies of complex phenotypes are non-coding, which makes it difficult to estimate their effect and contribution to the phenotype. Furthermore, linkage disequilibrium results in blocks of co-inherited variants, which makes it difficult to pinpoint individual causal variants. Thus, sequence-based deep learning models that can be used as matching tools to evaluate the impact of such variants provide a promising approach to find potential drivers of complex phenotypes. One example is predicting the effect of non-coding single nucleotide variants and short insertions or deletions (indels) indirectly from the differences between two variants on transcription factor binding, chromatin accessibility or gene expression prediction. Another example is predicting novel splice site generation from the quantitative effect of gene variants on sequence or splicing.

[0023] To predict the pathogenicity of missense variants from protein sequence and sequence conservation data, an end-to-end deep learning approach for variant effect prediction is applied (see Sundaram, L. et al. Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018), referred to herein as "PrimateAI"). PrimateAI uses a deep neural network trained on variants of known pathogenicity with data augmentation using cross-species information. In particular, PrimateAI uses wild-type and mutant protein sequences to compare differences and determine the pathogenicity of the variant using a trained deep neural network. Such an approach utilizing protein sequences for pathogenicity prediction is promising because it can avoid the circularity problem and overfitting to prior knowledge. However, the number of clinical data available in ClinVar is relatively small compared to the number of data sufficient to effectively train a deep neural network. To overcome this data scarcity, PrimateAI uses common human variants and primate-derived variants as benign data, and simulated variants based on trinucleotide context as unlabeled data.

[0024] PrimateAI outperforms conventional methods when trained directly on sequence alignments. PrimateAI learns important protein domains, conserved amino acid positions, and sequence dependencies directly from training data consisting of approximately 120,000 human samples. PrimateAI substantially outperforms other variant pathogenicity prediction tools in distinguishing benign and pathogenic de novo mutations in candidate developmental disorder genes and in reproducing prior knowledge in ClinVar. These results suggest that PrimateAI is an important step forward for variant classification tools that can reduce reliance on prior knowledge of clinical reports.

[0025] Central to protein biology is the understanding of how structural elements give rise to observed functions. The more than ample protein structural data allows the development of computational methods to systematically derive the rules governing structure-function relationships. However, the performance of these methods critically depends on the choice of protein structural representation.

[0026] Protein sites are microenvironments within a protein structure that are differentiated by their structural or functional role. Sites can be defined by a three-dimensional (3D) location and the local neighborhood around this location where the structure or function resides. Central to rational protein engineering is the understanding of how the structural arrangement of amino acids creates functional features within a protein site. Determination of the structural and functional roles of individual amino acids in a protein provides information to aid in the manipulation and modification of protein function. Identification of functionally or structurally important amino acids allows for focused engineering efforts such as site-directed mutagenesis to modify the functional properties of a target protein. Alternatively, this knowledge can help avoid engineering designs that disable desired functions.

[0027] Since it is established that structure is much more conserved than sequence, the increase in protein structural data provides an opportunity to systematically study the underlying patterns governing structure-function relationships using data-driven approaches. A fundamental aspect of any computational protein analysis is how protein structural information is represented. The performance of machine learning methods often depends more on the choice of data representation than on the machine learning algorithm used. A good representation efficiently captures the most important information, whereas a poor representation produces a noisy distribution lacking the underlying pattern.

[0028] The ample supply of protein structures and recent successes of deep learning algorithms provide an opportunity to develop tools to automatically extract task-specific representations of protein structures, thus predicting variant pathogenicity using multi-channel voxelized representations of 3D protein structures as input to deep neural networks. Summary of the Invention [Means for solving the problem]

[0029] In some embodiments, a multi-view convolutional neural network (CNN) for classifying protein structures is described herein. In some embodiments, a processed input for such a multi-view CNN is described herein. In providing such and other embodiments, the systems and methods described herein overcome some technical problems in classifying protein structures through artificial intelligence type computers and digital data processing systems, and corresponding data processing methods and products for emulation of intelligence. In addition, the techniques described herein provide specific technical solutions to at least overcome the technical problems mentioned herein, as well as other technical problems not described herein but recognized by those skilled in the art.

[0030] With respect to some embodiments, described herein are computerized methods for classifying protein structures via a trainable computing system, and computerized methods for preprocessing input for a trainable computing system, as well as non-transitory computer-readable storage media for performing the technical operations of the computerized methods. The non-transitory computer-readable storage medium tangibly stores or has computer-readable instructions that, when executed by one or more devices (e.g., one or more personal computers or servers), cause at least one processor to perform: (1) a method for classifying protein structures via a trainable computing system, (2) a method for preprocessing input for a trainable computing system, or (3) a method for performing a combination of the method for classifying protein structures and the method for preprocessing input.

[0031] In some embodiments, a computer-implemented method for determining pathogenicity of a variant includes accessing a structural rendition of amino acids of a protein, capturing a plurality of images of a portion of the structural rendition that includes a target amino acid from the amino acids, and determining pathogenicity of a nucleotide variant that mutates the target amino acid to an alternative amino acid based at least in part on the plurality of images. [Brief description of the drawings]

[0032] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings, in which: [Figure 1] 1 illustrates one embodiment of a multi-view convolutional neural network for 3D shape recognition. [Diagram 2]1 illustrates a system including a two-dimensional (2D) image classification network that generates classifications of protein structures, according to some embodiments of the present disclosure. [Figure 3A] 1 illustrates a 2D image of a protein structure generated from a three-dimensional (3D) image of the protein structure, according to some embodiments of the present disclosure. [Figure 3B] 1 illustrates a 2D image of a protein structure generated from a three-dimensional (3D) image of the protein structure, according to some embodiments of the present disclosure. [Figure 3C] 1 illustrates a 2D image of a protein structure generated from a three-dimensional (3D) image of the protein structure, according to some embodiments of the present disclosure. [Figure 3D] 1 illustrates a 2D image of a protein structure generated from a three-dimensional (3D) image of the protein structure, according to some embodiments of the present disclosure. [Figure 4] 1 illustrates a method of using a 2D image classification network to generate classifications of protein structures, according to some embodiments of the present disclosure. [Diagram 5] 1 illustrates a method of using a 2D image classification network to generate classifications of protein structures, according to some embodiments of the present disclosure. [Figure 6] 1 illustrates a method of using a 2D image classification network to generate classifications of protein structures, according to some embodiments of the present disclosure. [Figure 7] 1 illustrates a system, including a convolutional neural network, for generating classifications of protein structures, according to some embodiments of the present disclosure. [Figure 8] 1 illustrates a method for generating classifications of protein structures using a convolutional neural network (CNN), according to some embodiments of the present disclosure. [Figure 9] 1 illustrates a method for generating classifications of protein structures using a convolutional neural network (CNN), according to some embodiments of the present disclosure. [Figure 10] 1 shows a method for determining pathogenicity of a variant according to some embodiments of the present disclosure. [Figure 11] 1 shows a method for determining pathogenicity of a variant according to some embodiments of the present disclosure. [Figure 12] 1 shows a method for determining pathogenicity of a variant according to some embodiments of the present disclosure. [Figure 13] 1 illustrates a block diagram of an exemplary aspect of a computing system according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0033] The following discussion is presented to enable those skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.

[0034] Multi-view convolutional neural networks for 3D shape recognition FIG. 1 illustrates one implementation of a multi-view convolutional neural network (also known as a multi-view CNN) for three-dimensional (3D) shape recognition. FIG. 1 can be found in a paper entitled "Multi-view Convolutional Neural Networks for 3D Shape Recognition" by Hang Su et al., published May 5, 2015 in connection with the 2015 IEEE International Conference on Computer Vision (ICCV). FIG. 1 illustrates a 3D shape rendered from multiple different views and the views passing through multiple corresponding first convolutional neural networks (CNNs) to extract view-based features. The features are then pooled across the views and passed through a second convolutional neural network (CNN) to obtain a compact shape descriptor.

[0035] Genomics and Deep Learning Genomics in the broad sense, also called functional genomics, aims to characterize the function of all genomic elements of an organism by using genome-scale assays such as genome sequencing, transcriptome profiling, and proteomics. Genomics has emerged as a data-driven science and operates not by testing preconceived models and hypotheses, but by discovering novel properties from the exploration of genome-scale data. Applications of genomics include finding associations between genotypes and phenotypes, discovering biomarkers for patient stratification, predicting gene functions, and mapping biochemically active genomic regions such as transcriptional enhancers.

[0036] Genomics data is too large and complex to be mined solely by visual inspection of pairwise correlations. Instead, analytical tools are needed to support the discovery of unexpected relationships, derive novel hypotheses and models, and make predictions. Unlike some models in which assumptions and domain expertise are hard-coded, machine learning models are designed to automatically detect patterns in the data. Thus, machine learning models are well suited for data-driven science, especially genomics. However, the performance of machine learning models can be highly dependent on how the data is represented, i.e., how each variable (also called a feature) is calculated. For example, to classify tumors as malignant or benign from a fluorescent microscopy image, a pre-processing model can detect cells, identify cell types, and generate a list of cell counts for each cell type.

[0037] The machine learning model can take estimated cell counts, which are an example of hand-designed features, as input features to classify tumors. The central problem is that classification performance is highly dependent on the quality and relevance of these features. For example, relevant visual features such as cell morphology, distance between cells or localization within an organ are not captured in the cell counts, and this incomplete representation of the data can reduce classification accuracy.

[0038] Deep learning, a sub-discipline of machine learning, addresses this problem by embedding feature computation into the machine learning model itself, generating an end-to-end model. This result has been achieved through the development of deep neural networks, which are machine learning models that involve successive primitive operations that compute increasingly complex features by taking the results of previous operations as input. Deep neural networks can improve prediction accuracy by discovering relevant features of high complexity, such as cell morphology and spatial organization of cells in the above example. The construction and training of deep neural networks has been made possible by the explosion of data, advances in models, and substantial increases in computing power, especially through the use of graphical processing units (GPUs).

[0039] Techniques for classifying protein structures, such as techniques for classifying pathogenicity of protein structures associated with nucleotide variants, are described herein. Such classification is based on two-dimensional images obtained from a three-dimensional image of the protein structure. For some embodiments, a multi-view convolutional neural network (CNN) for classifying protein structures based on input of two-dimensional images obtained from a three-dimensional image of the protein structure is described herein. In some embodiments, a computer-implemented method for determining pathogenicity of a variant includes accessing a structural rendition of amino acids of a protein, capturing a plurality of images of a portion of the structural rendition that includes a target amino acid from the amino acids, and determining pathogenicity of a nucleotide variant that mutates the target amino acid to an alternative amino acid based at least in part on the plurality of images. The portion is captured by zooming in on the portion. In other embodiments, the portion is captured by filtering out other portions.

[0040] The actions in the figures disclosed herein may be performed at least in part with and / or by one or more processors configured to receive or retrieve information, process information, store results, and transmit results. Other embodiments may perform the actions in a different order than illustrated in the figures and / or with different, fewer, or additional actions. In some embodiments, multiple actions may be combined. For convenience, the figures are described with reference to a system that performs the method. A system is not necessarily part of a method. The actions in the figures disclosed herein may be performed in parallel or sequentially.

[0041] In one embodiment, the phenotype decision logic (e.g., a pathogenicity classifier) ​​is a multilayer perceptron (MLP). In another embodiment, the phenotype decision logic is a feedforward neural network. In yet another embodiment, the phenotype decision logic is a fully-connected neural network. In a further embodiment, the phenotype decision logic is a fully convolution neural network. In yet a further embodiment, the phenotype decision logic is a semantic segmentation neural network. In yet another further embodiment, the phenotype decision logic is a generative adversarial network (GAN) (e.g., CycleGAN, StyleGAN, pixelRNN, text-2-image, DiscoGAN, IsGAN).

[0042] In one embodiment, the phenotype determination logic is a convolutional neural network (CNN) having multiple convolution layers. In another embodiment, the phenotype determination logic is a recurrent neural network (RNN), such as a long short-term memory network (LSTM), a bi-directional LSTM (Bi-LSTM), or a gated recurrent unit (GRU). In yet another embodiment, the phenotype determination logic includes both a CNN and an RNN.

[0043] In yet other embodiments, the phenotyping logic may use 1D convolution, 2D convolution, 3D convolution, 4D convolution, 5D convolution, dilated or expanded convolution, transposed convolution, depth separable convolution, point-wise convolution, 1×1 convolution, group convolution, flattened convolution, spatial and cross-channel phenotyping logic convolution, shuffled grouped convolution, spatially separable convolution, and deconvolution. The phenotyping logic may use one or more loss functions, such as logistic regression / logarithmic loss, multiclass cross-entropy / softmax loss, binary cross-entropy loss, mean squared error loss, L1 loss, L2 loss, smoothed L1 loss, and Huber loss. The representation-determination logic can use any parallel, efficient, and compact scheme, such as TFRecords, compression encoding (e.g., PNG), sharpening, parallel calls to map transforms, batching, prefetching, model parallel, data parallel, and synchronous / asynchronous stochastic gradient descent (SGD). The representation-determination logic can include upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (such as LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., nonlinear transformation functions such as rectified linear unit (ReLU), leaky ReLU, exponential linear unit (ELU), sigmoid, and hyperbolic tangent function (tanh)), batch normalization layers, normalization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, and attention mechanisms (e.g., self-attention).

[0044] The phenotype decision logic may be a rule-based model, a linear regression model, a logistic regression model, an Elastic Net model, a support vector machine (SVM), a random forest (RF), a decision tree, and a boosted decision tree (e.g., XGBoost), or some other tree-based logic (e.g., metric tree, kd tree, R-tree, universal B-tree, X-tree, ball tree, locality sensitive hashing, and inverted index). The phenotype decision logic may be an ensemble of multiple models in some embodiments.

[0045] The phenotype decision logic is trained using a backpropagation-based gradient update technique. Exemplary gradient descent techniques that can be used to train the phenotype decision logic include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the phenotype decision logic include Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.

[0046] Effectively designed specifications include Vision Transformer(ViT), Bidirectional Transformer(BERT), Detection Transformer(DETR), Deformable DETR, UP-DETR, DeiT, Swin, GPT, iGPT, GPT-2, GPT-3, BERT, SpanBERT, RoBERTa, XLNet, ELECTRA, UniLM, BART, T5, ER NIE(THU), KnowBERT, DeiT-Ti, DeiT-S, DeiT-B, T2T-ViT-14, T2T-ViT-19, T2T-ViT-24, PVT-Small, PVT-Medium, PV T-Large, TNT-S, TNT-B, CPVT-S, CPVT-S-GAP, CPVT-B, Swin-T, Swin-S, Swin-B, Twins-SVT-S, Twins-SVT-B, Twin- SVT-L, Shuffle-T, Shuffle-S, Shuffle-B, XCiT-S12 / 16, CMT-S, CMT-B, VOLO-D1, VOLO-D2, VOLO-D3, VOLO-D4, MoCo v3、ACT、TSP、Max-DeepLab、VisTR、SETR、Hand-Transformer、HOT-Net、METRO、Image Transformer、Taming transformer、TransGAN、IPT、TTSR、STTN、Masked Transformer, CLIP, DALL-E, Cogview, UniT, ASH, TinyBert, FullyQT, ConvBert, FCOS, Faster R-CNN+FPN、DETR-DC5、TSP-FCOS、TSP-RCNN、ACT+MKDD(L=32)、ACT+MKDD(L=16)、SMCA、Efficient DETR, UP-DETR, UP-DETR, ViTB / 16-FRCNN, ViT-B / 16-FRCNN, PVT-Small+RetinaNet, Swin-T+Retin aNet, Swin-T+ATSS, PVT-Small+DETR, TNT-S+DETR, YOLOS-Ti, YOLOS-S, and YOLOS-B are all slightly different.

[0047] FIG. 2 illustrates a system 200 including a 2D image classification network 202 (e.g., a network implemented by a convolutional neural network) that generates a classification of a protein structure (e.g., pathogenicity of the protein structure). The system 200 includes a plurality of virtual cameras 204 configured to generate a plurality of 2D images of a protein structure 206 from a 3D image of the protein structure 208. See, for example, the 2D images shown in FIGS. 3A, 3B, 3C, and 3D. Each camera of the plurality of virtual cameras 204 is configured to capture a respective image of the plurality of 2D images of the protein structure 206. The 2D image classification network 202 is configured to process the plurality of 2D images of the protein structure 206 to generate a classification of the protein structure 210.

[0048] For purposes of this disclosure, a virtual camera should be understood to be a feature of animation software that operates in a manner similar to a camera. The software that includes a virtual camera includes calculations that determine the rendering of objects based on the location and angle of the virtual camera according to the software. Like a camera, a virtual camera may use features such as panning, zooming, focusing, and changing focus.

[0049] Also, for purposes of this disclosure, a protein viewer may include any type of software, with or without a virtual camera, for viewing a protein or protein structure, or for viewing a derivative of a protein or protein structure. Also, for purposes of this disclosure, a protein structure should be understood to be a structural portion of one or more proteins. For example, a protein structure may include a portion of an amino acid, one or more amino acids, or one or more residues, where a residue is an amino acid when linked into a protein chain.

[0050] In some implementations, such as the implementation shown in FIG. 2, the multiple virtual cameras 204 are part of a protein viewer 212. The protein viewer 212 is configured to select a 3D image of the protein structure 208 and to magnify (e.g., zoom in, increase resolution) the 3D image of the protein structure. The protein viewer 212 is also configured to color the 3D image of the protein structure 208 according to the protein viewer's coloring parameters. The protein viewer 212 is also configured to capture multiple 2D images of the protein structure 206 from different perspectives via the multiple virtual cameras 204 (e.g., after the 3D image of the protein structure 208 has been magnified (e.g., zoomed in) and colored, etc.).

[0051] 2, the system 200 includes a computing system 214 configured to train the 2D image classification network 202. To train the 2D image classification network 202, the computing system 214 is configured to iterate, for a 3D image of a protein structure (see, e.g., the 3D image of the protein structure 208), the steps of (1) generating a plurality of 2D images of the protein structure from the 3D image of the protein structure, and (2) processing the plurality of 2D images by the 2D image classification network to generate a classification of the protein structure. To train the 2D image classification network 202, the computing system 214 is also configured to adjust parameters of the 2D image classification network 202 according to the generated classification of the protein structure after each iteration of steps (1) and (2).

[0052] In some embodiments, the protein structure comprises a target amino acid.

[0053] In some embodiments, the protein structure classification comprises a pathogenicity score.

[0054] In some embodiments, the 2D image classification network 202 is configured to process a plurality of 2D images of the protein structure 206 to determine the pathogenicity of a nucleotide variant that results in the substitution of a target amino acid with an alternative amino acid in the protein structure. In such embodiments, the classification comprises the pathogenicity of the nucleotide variant.

[0055] In some embodiments, the pathogenicity of a nucleotide variant is represented by a pathogenicity score. For the purposes of this disclosure, it should be understood that pathogenicity is a property that causes disease. Thus, a pathogenicity score provides the degree to which something causes disease. In some embodiments, a pathogenicity score relates to the degree to which a nucleotide variant causes disease compared to other nucleotide variants in the genome of a population (such as a human population).

[0056] In some embodiments, the plurality of virtual cameras 204 are configured to generate a plurality of graphical representations of amino acids in the protein structure from the 3D graphical representations of the amino acids. In such embodiments, the 3D image of the protein structure 208 includes a 3D graphical representation of the amino acids, and each image of the plurality of 2D images of the protein structure 206 includes a respective graphical representation of an amino acid of the plurality of graphical representations of amino acids.

[0057] In some embodiments, the amino acid comprises a target amino acid. In some embodiments, the 2D image classification network 202 is configured to process a plurality of 2D images of the protein structure 206 to determine the pathogenicity of a nucleotide variant that results in the replacement of the target amino acid with an alternative amino acid. In some such embodiments, the classification of the protein structure comprises a pathogenicity score.

[0058] Also, as mentioned herein, in some embodiments, the multiple virtual cameras 204 are part of the protein viewer 212. Also, the multiple virtual cameras 204 are configured to generate multiple graphical representations of amino acids in the protein structure. In some embodiments, the protein viewer 212 is configured to: (1) select a 3D image of the protein structure 208; (2) zoom in on the 3D image of the protein structure 208; (3) color each amino acid in the multiple graphical representations of amino acids in the 3D image of the protein structure 208 by amino acid type; and (4) capture multiple 2D images of the protein structure 206 from different perspectives via the multiple virtual cameras 204 (e.g., after the 3D image is zoomed in and colored).

[0059] In some such embodiments where the multiple virtual cameras 204 are configured to generate multiple graphical representations of amino acids in a protein structure, the coloring parameters are amino acid types and the protein viewer 212 is configured to color the 3D image of the protein structure 208 by amino acid type. In other embodiments, the coloring parameters are used to adjust protein structural elements such as atom distributions.

[0060] Also, in some embodiments, the coloring parameters are used to represent the conservation of the protein structure, and the protein viewer 212 is configured to color the 3D image of the protein structure 208 according to the conservation of the protein structure, and in some such embodiments, the coloring parameters are structural qualities, and the protein viewer 212 is configured to color the 3D image of the protein structure 208 according to the structural qualities.

[0061] Some examples of structure quality / confidence information include GMQE score (provided by SwissModel), B-factor; homology model temperature factor column (indicating how well a residue satisfies a (physical) constraint in a protein structure); normalized number of aligned template proteins to the residue closest to the center of the voxel (for an alignment provided by HHpred, e.g., a voxel is closest to the residue that 3 of the 6 template structures align to, meaning the feature has a value 3 / 6=0.5); minimum, maximum, and average TM scores; and predicted TM score of the template protein structure that aligns to the residue closest to the voxel (continuing the example above, suppose the 3 template structures have TM scores 0.5, 0.5, and 1.5, the minimum is 0.5, the average is 2 / 3, and the maximum is 1.5). TM scores can be provided for each protein template by HHpred.

[0062] In some embodiments, the 3D image of the protein structure 208 includes a graphical representation of atoms. In some embodiments, the 3D image of the protein structure 208 includes a graphical representation of residues. Also, in some embodiments, the 3D image of the protein structure 208 includes a graphical representation of atoms and residues.

[0063] Suitable embodiments of system 200 perform any one of the methods described herein. For example, system 200 may provide a computer-implemented method for determining pathogenicity of a variant, including accessing a structural rendition of amino acids, capturing an image of a portion of the structural rendition that includes a target amino acid from the amino acids, and determining, based on the image, the pathogenicity of a nucleotide variant that mutates the target amino acid to an alternative amino acid.

[0064] 3A, 3B, 3C, and 3D illustrate 2D images 302, 304, 306, and 308 of a protein structure generated from a 3D image of the protein structure. As shown with 2D images 302, 304, 306, and 308, in an embodiment using a protein viewer (see, e.g., protein viewer 212), the viewer is configured to color the 3D image of the protein structure with coloring parameters of the protein viewer and then capture via multiple virtual cameras (see, e.g., multiple virtual cameras 204), multiple 2D images from different perspectives, and the like. The protein structure colored by the viewer includes a target amino acid, as shown as 2D images 302, 304, 306, and 308. Processing the multiple 2D images according to some of the embodiments disclosed herein also includes processing the multiple 2D images by a 2D image classification network to determine the pathogenicity of a nucleotide variant that replaces a target amino acid in the protein structure with an alternative amino acid.

[0065] 4, 5, and 6 show methods 400, 500, and 600, respectively. Each of methods 400, 500, and 600 uses a 2D image classification network (such as 2D image classification network 202 shown in FIG. 2) to generate a classification of protein structures. Any one of the 2D image classification networks described herein may be used by methods 400, 500, and 600.

[0066] As shown by FIG. 4, the method 400 begins in step 402 with generating a plurality of 2D images of the protein structure from a 3D image of the protein structure (see, for example, the plurality of 2D images of the protein structure 206 and the 3D image of the protein structure 208 shown in FIG. 2, and the 2D images 302, 304, 306, and 308 shown in FIGS. 3A, 3B, 3C, and 3D, respectively). The method 400 then continues in step 404 with processing the plurality of 2D images by a 2D image classification network to generate a classification of the protein structure. In some embodiments of the method 400, the protein structure includes a target amino acid. In some embodiments of the method 400, the classification of the protein structure includes generating a pathogenicity score. More specifically, in some embodiments of the method 400, the processing of the plurality of 2D images in step 404 includes processing the plurality of 2D images by a 2D image classification network to determine the pathogenicity of a nucleotide variant that replaces a target amino acid in the protein structure with an alternative amino acid. In such embodiments, the protein structural classification may also include a pathogenicity score.

[0067] In some embodiments of method 400, generating the plurality of 2D images includes generating a plurality of graphical representations of amino acids in the protein structure from a 3D graphical representation of the amino acids. The 3D image of the protein structure includes a 3D graphical representation of amino acids, and each image of the plurality of 2D images includes a respective graphical representation of an amino acid of the plurality of graphical representations of amino acids. In such embodiments, the amino acid includes a target amino acid. Also, in some examples, processing the plurality of 2D images includes processing the plurality of 2D images by a 2D image classification network to determine a pathogenicity of a nucleotide variant that results in the substitution of the target amino acid with an alternative amino acid. In such examples, the classification of the protein structure may include a pathogenicity score.

[0068] Also, in those embodiments in which a plurality of graphical representations of amino acids of a protein structure are generated, as illustrated by the steps of method 500 shown in FIG. 5, and in other embodiments, generating the plurality of 2D images may include: (1) selecting, by a protein viewer (e.g., see protein viewer 212), a 3D image of the protein structure (e.g., see step 502); (2) zooming, by the protein viewer, in the 3D image of the protein structure (e.g., see step 504); and (3) coloring, by the protein viewer, the 3D image of the protein structure with coloring parameters of the protein viewer (e.g., see step 506). In some examples, the coloring of the 3D image of the protein structure is by amino acid type. In such cases, the coloring includes coloring each amino acid of the plurality of graphical representations of amino acids by amino acid type. Also, in such embodiments where multiple graphical representations of amino acids of the protein structure are generated, as well as in other embodiments, generating the multiple 2D images may include (4) capturing, by the protein viewer, the multiple 2D images from different perspectives after the 3D image is magnified and colored (see, e.g., step 508). As discussed above, the coloring parameter may be amino acid type, and in such examples, the coloring of the 3D image of the protein structure is by amino acid type. The coloring parameter may also be protein structure conservation, and in such examples, the coloring of the 3D image of the protein structure is by protein structure conservation. The coloring parameter may also be structural quality, and in such examples, the coloring of the 3D image of the protein structure is by structural quality.

[0069] 4 and 5, in some embodiments, step 402 of method 400 may include steps 502-508 of method 500. Also, with reference to FIGS. 4 and 5, in some embodiments of method 400, the 3D image of the protein structure includes a graphical representation of atoms. In some embodiments of methods 400 and 500, the 3D image of the protein structure includes a graphical representation of residues. Also, the 3D image of the protein structure may include a graphical representation of atoms and residues.

[0070] 6 illustrates a method 600 that includes training a 2D image classification network. Training the 2D image classification network includes, for different 3D images of different protein structures, repeating generating a plurality of different 2D images of the different protein structures from the 3D images of the different protein structures (step 602). Training the 2D image classification network also includes, for different 3D images of different protein structures, repeating processing the plurality of different 2D images by the 2D image classification network to generate classifications of the different protein structures (step 604). In other words, the method 600 includes, for different 3D images of different protein structures, repeating the steps of (1) generating a plurality of different 2D images of the different protein structures from the 3D images of the different protein structures (step 602) and (2) processing the plurality of different 2D images by the 2D image classification network to generate classifications of the different protein structures (step 604). Training the 2D image classification network also includes adjusting parameters of the 2D image classification network according to the generated classifications of the different protein structures after each iteration of steps 602 and 604 (step 606). In the training shown in method 600, processing the plurality of 2D images may include processing the plurality of 2D images by the 2D image classification network to determine pathogenicity of nucleotide variants that result in the replacement of a target amino acid with an alternative amino acid in the protein structure. The classification of the protein structure may also include a pathogenicity score.

[0071] FIG. 7 illustrates a system 700 including a convolutional neural network 702. The combination of the convolutional neural networks 702 is configured to generate a classification of protein structures. For some embodiments, the convolutional neural networks described herein are multi-view convolutional neural networks for classifying protein structures. In some embodiments, the multi-view CNN is somewhat similar to the multi-view CNN described in Hang Su et al. "Multi-view Convolutional Neural Networks for 3D Shape Recognition." 2015 IEEE International Conference on Computer Vision (ICCV) (Su, Hang et al. "Multi-view Convolutional Neural Networks for 3D Shape Recognition." 2015 IEEE International Conference on Computer Vision (ICCV) (2015):945-953.), published May 5, 2015. In general, for purposes of this disclosure, a convolutional neural network or CNN is a class of artificial neural networks that employ an operation called convolution. Convolutional networks are a special type of neural network that use convolution instead of common matrix multiplication in at least one of their layers. In some implementations, CNNs are regularized versions of multilayer perceptrons. Multilayer perceptrons are fully connected networks in that each neuron in one layer is connected to every neuron in the next layer.

[0072] System 700 also includes a plurality of virtual cameras 204 (see, e.g., FIG. 2) configured to generate a plurality of 2D images of protein structure 206 from a 3D image of protein structure 208 (see, e.g., FIG. 2 and images 3A, 3B, 3C, and 3D). Each camera of the plurality of virtual cameras 204 is configured to capture a respective image of the plurality of 2D images of protein structure 206.

[0073] 7, the convolutional neural network 702 includes a first convolutional layer 704 including a plurality of convolutional neural networks (see, e.g., CNNs 706a, 706b, and 706c) configured to generate a plurality of first-layer outputs (see, e.g., outputs 708a, 708b, and 708c) based on a plurality of 2D images of the protein structure 206. Each output of the plurality of first-layer outputs is for a respective image of the plurality of 2D images of the protein structure 206. Each convolutional neural network of the plurality of convolutional neural networks is for a respective image of the plurality of 2D images.

[0074] The convolutional neural network 702 also includes a pooling layer 710 configured to combine multiple first layer outputs into a second layer input 712. The convolutional neural network 702 also includes a second convolutional layer 714 including a second layer convolutional neural network 716 configured to generate a classification of the protein structure 210 based on the second layer input 712.

[0075] 7, in some implementations of system 700, the multiple virtual cameras 204 are part of the protein viewer 212. Similar to system 200, in some implementations of system 700, the protein viewer 212 may be configured to (1) select a 3D image of a protein structure 208, (2) zoom in on the 3D image of the protein structure, (3) color each amino acid of multiple graphical representations of amino acids in the 3D image of the protein structure by amino acid type, and (4) capture multiple 2D images of the protein structure 206 from different perspectives via the multiple virtual cameras 204 (e.g., after the 3D image has been zoomed in and colored). In a more generalized embodiment, the multiple virtual cameras are part of a protein viewer, and the protein viewer is configured to: (1) select a 3D image of a protein structure; (2) magnify the 3D image of the protein structure; (3) color the 3D image of the protein structure according to coloring parameters of the protein viewer; and (4) capture multiple 2D images from different perspectives via the multiple virtual cameras (e.g., after the 3D image has been magnified and colored).

[0076] As described above, in some implementations of the system 700, the coloring parameter may be an amino acid type and the coloring of the 3D image of the protein structure may be by the amino acid type. Also, the coloring parameter may be a conservation of the protein structure and the coloring of the 3D image of the protein structure may be by the conservation of the protein structure. Furthermore, in some examples, the coloring parameter may be a structural quality and the coloring of the 3D image of the protein structure may be by the structural quality.

[0077] Also as shown in FIG. 7, in some implementations of the system 700, the system includes a 2D image classification network 720 including a first convolutional layer 704, a pooling layer 710, and a second convolutional layer 714. The system 700 also includes a computing system 214 configured to train the 2D image classification network 720. In such implementations and other implementations, the computing system 214 is configured to tune parameters of the 2D image classification network 720. For example, the computing system 214 may be configured to tune parameters of the multiple convolutional neural networks (see, e.g., CNNs 706a, 706b, and 706c) in the first convolutional layer 704. Also, for example, the computing system 214 may be configured to tune parameters of the second layer convolutional neural network 716 in the second convolutional layer 714. Also, for example, the computing system 214 may be configured to tune parameters of the pooling layer 710. Such adjustments may be made in the training of the 2D image classification network 720 performed by the computing system 214, or in some other process performed by the computing system.

[0078] Specifically, in some embodiments, the computing system 214 is configured to train the 2D image classification network 720 by repeating, for different 3D images of different protein structures, the steps of: (1) generating different 2D images of the different protein structures from the 3D images of the different protein structures via the multiple virtual cameras 204; (2) generating different first-layer outputs based on the different 2D images via a plurality of convolutional neural networks (see, e.g., CNNs 706a, 706b, and 706c); (3) combining the different first-layer outputs into different second-layer inputs via a pooling layer 710; and (4) generating classifications of the different protein structures based on the different second-layer inputs via a second-layer convolutional neural network 716. The computing system 214 is also configured to train the 2D image classification network 720 by being configured to adjust parameters of the 2D image classification network 720 according to the classifications of the different protein structures generated after each iteration of steps (1), (2), (3), and (4).

[0079] Similar to system 200, in system 700, the protein structure may include a target amino acid and the classification of the protein structure may include a pathogenicity score. In such an embodiment, for example, the 2D image classification network 202 may be configured to process a plurality of 2D images to determine the pathogenicity of a nucleotide variant that substitutes a target amino acid with an alternative amino acid.

[0080] Also similar to system 200, according to system 700, the plurality of virtual cameras 204 may be configured to generate a plurality of graphical representations of amino acids in the protein structure from the 3D graphical representations of the amino acids. Also, the 3D image of the protein structure includes a 3D graphical representation of the amino acids, and each image of the plurality of 2D images includes a respective graphical representation of an amino acid of the plurality of graphical representations of amino acids. Moreover, in such an embodiment, the amino acid includes a target amino acid. Additionally, in such an embodiment, system 700 may include a 2D image classification network 720 (instead of network 202), which may be configured to process the plurality of 2D images to determine the pathogenicity of a nucleotide variant that results in the substitution of a target amino acid with an alternative amino acid. Also, in such an embodiment, the classification of the protein structure may include a pathogenicity score.

[0081] Also, similar to system 200, according to system 700, the 3D image of the protein structure may include a graphical representation of the atoms. Also, the 3D image of the protein structure may include a graphical representation of the residues.

[0082] Suitable embodiments of system 700 perform any one of the methods described herein using a CNN, such as through a computer-implemented method. For example, system 700 may provide a computer-implemented method for determining pathogenicity of a variant that includes accessing a structural rendition of amino acids, capturing an image of a portion of the structural rendition that includes a target amino acid from the amino acids, and determining, based on the image, via a CNN, the pathogenicity of a nucleotide variant that mutates the target amino acid to an alternative amino acid.

[0083] 8, 9, and 10 show methods 800, 900, and 1000, respectively. Each of methods 800, 900, and 1000 uses a convolutional neural network (such as convolutional neural network 702 shown in FIG. 7, which is part of 2D image classification network 720) to generate a classification of protein structures. Any one of the 2D image classification networks described herein, including convolutional neural networks, may be used by methods 800, 900, and 1000.

[0084] As shown by FIG. 8, the method 800 begins in step 402 (also in method 400) with generating a plurality of 2D images of the protein structure from a 3D image of the protein structure. See, for example, the plurality of 2D images of the protein structure 206 and the 3D image of the protein structure 208 shown in FIG. 2, and the 2D images 302, 304, 306, and 308 shown in FIGS. 3A, 3B, 3C, and 3D, respectively. Next, in step 802, the method 800 continues with processing the plurality of 2D images using a plurality of convolutional neural networks (see, for example, CNNs 706a, 706b, and 706c) to generate a plurality of first-layer outputs (see, for example, outputs 708a, 708b, and 708c). Each output of the plurality of first-layer outputs is for a respective image of the plurality of 2D images, and each convolutional neural network of the plurality of convolutional neural networks is for a respective image of the plurality of 2D images.

[0085] Following generation of the plurality of first-layer outputs in step 802, the method 800 continues with combining the plurality of first-layer outputs with second-layer inputs in step 804 (see, e.g., second-layer inputs 712). Also, in step 806, the method 800 continues with processing the second-layer inputs using a second-layer convolutional neural network to generate a classification of protein structures. See, e.g., second-layer convolutional neural network 716.

[0086] In some embodiments of method 800, the protein structure comprises a target amino acid. In some embodiments of method 800, the classification of the protein structure comprises a pathogenicity score.

[0087] In some embodiments of method 800, processing the plurality of 2D images includes processing the plurality of 2D images by a 2D image classification network to determine pathogenicity of nucleotide variants that result in the substitution of a target amino acid with an alternative amino acid in the protein structure. Also in such embodiments, the 2D image classification network includes a plurality of convolutional neural networks and a second layer convolutional neural network. Also in such embodiments, the classification of the protein structure includes a pathogenicity score.

[0088] In some embodiments of method 800, similar to some embodiments of method 400, generating the plurality of 2D images in step 402 includes generating a plurality of graphical representations of amino acids in the protein structure from a 3D graphical representation of the amino acids. In such embodiments, the 3D image of the protein structure includes a 3D graphical representation of amino acids, and each image of the plurality of 2D images includes a respective graphical representation of an amino acid of the plurality of graphical representations of amino acids. Also, in such embodiments, the amino acid may include a target amino acid. Also, in such embodiments, processing the plurality of 2D images may include processing the plurality of 2D images by a 2D image classification network to determine a pathogenicity of a nucleotide variant that results in the substitution of the target amino acid with an alternative amino acid. In such and other examples, the 2D image classification network includes a plurality of convolutional neural networks and a second layer convolutional neural network. Also, the classification of the protein structure includes a pathogenicity score.

[0089] FIG. 9 illustrates a method 900 that includes training a 2D image classification network having a plurality of convolutional neural networks and a second layer convolutional neural network (see, e.g., 2D image classification network 720 shown in FIG. 7). Training such a 2D image classification network includes iterating, for different 3D images of different protein structures, generating a plurality of different 2D images of the different protein structures from the 3D images of the different protein structures (step 902). Training the 2D image classification network also includes iterating, for different 3D images of different protein structures, processing the plurality of different 2D images using a plurality of convolutional neural networks (see, e.g., CNNs 706a, 706b, and 706c) to generate a plurality of different first layer outputs (step 904). Training the 2D image classification network also includes iterating, for different 3D images of different protein structures, combining the plurality of different first layer outputs with different second layer inputs (step 906). Training the 2D image classification network also includes iteratively processing different second-layer inputs with the second-layer convolutional neural network for different 3D images of different protein structures to generate classifications of the different protein structures (step 908). See, e.g., second-layer convolutional neural network 716.

[0090] In other words, method 900 includes repeating, for different 3D images of different protein structures, the steps of: (1) generating a plurality of different 2D images of the different protein structures from the 3D images of the different protein structures (step 902); (2) processing the plurality of different 2D images using a plurality of convolutional neural networks to generate a plurality of different first-layer outputs (step 902); (3) combining the plurality of different first-layer outputs with different second-layer inputs (step 906); and (4) processing the different second-layer inputs using a second-layer convolutional neural network to generate classifications of the different protein structures (step 908).

[0091] In method 900, training the 2D image classification network also includes adjusting parameters of the 2D image classification network according to the generated classifications of the different protein structures after each iteration of steps 902, 904, 906, and 908 (step 910). Also, in the training shown in method 900, processing the plurality of 2D images may include processing the plurality of 2D images by the 2D image classification network to determine pathogenicity of nucleotide variants that result in the substitution of a target amino acid with an alternative amino acid in the protein structure. Also, the classification of the protein structure may include a pathogenicity score.

[0092] In some implementations of method 900, adjusting parameters of the 2D image classification network in step 910 includes adjusting parameters of a plurality of convolutional neural networks. In some implementations of method 900, adjusting parameters of the 2D image classification network in step 910 includes adjusting parameters of a second layer convolutional neural network. In some implementations of method 900, adjusting parameters of the 2D image classification network in step 910 includes adjusting parameters of a pooling layer. Also, in some such implementations, the pooling layer implements a combination of the different first layer outputs to the second layer inputs in step 906.

[0093] 10, 11, and 12 show methods 1000, 1100, and 1200, respectively, for determining pathogenicity of a variant, such as a nucleotide variant.

[0094] For the purposes of this disclosure, it should be understood that a variant is not necessarily limited to a nucleotide variant. Also, a nucleotide variant may include one or more nucleotide variants. Unless otherwise specified herein, the term "variant" refers to a nucleic acid sequence that differs from a nucleic acid reference. Exemplary nucleic acid sequence variants include, but are not limited to, single nucleotide polymorphisms (SNPs), short deletion and insertion polymorphisms (indels), copy number variations (CNVs), microsatellite markers, or short tandem repeats and structural variations. In this application, amino acid substitutions include single and / or multiple amino acid substitutions.

[0095] The method 1000 begins by accessing a structural rendition of an amino acid of a protein (step 1002). The method 1000 continues by capturing a plurality of images of a portion of the structural rendition that includes the amino acid to a target amino acid (step 1004). Finally, the method 1000 concludes by determining, based at least in part on the plurality of images, a pathogenicity of a nucleotide variant that mutates the target amino acid to an alternative amino acid (step 1006).

[0096] In some implementations of the method 1000, the images in the plurality of images are captured from multiple perspectives, e.g., the images are captured from multiple perspectives, multiple orientations, multiple positions, multiple zoom levels, etc., or some combination thereof.

[0097] In some embodiments of method 1000, the portion of the structural rendition includes the target amino acid and a number of additional amino acids adjacent to the target amino acid. Also, in some embodiments of method 1000, the structural rendition is a 3D structural rendition.

[0098] Method 1100 includes steps 1102 and 1104 that may be combined with some steps of method 1000. Step 1102 includes activating a feature configuration of a structural rendition prior to capturing a plurality of images in step 1004 of method 1000. Also, in step 1104, a plurality of images of the portion of the structural rendition having the feature configuration activated in step 1004 are captured.

[0099] In some embodiments of the method 1100, the feature is to display residue details of the amino acid. In some embodiments of the method 1100, the feature is to display atomic details of the amino acid. In some embodiments of the method 1100, the feature is to display polymer details of the amino acid. In some embodiments of the method 1100, the feature is to display ligand details of the amino acid. In some embodiments of the method 1100, the feature is to display chain details of the amino acid. In some embodiments of the method 1100, the feature is to display surface effects of the amino acid. In such examples, the surface effects include transparency and color coding for electrostatic and hydrophobic values. In some embodiments of the method 1100, the feature is to display a density map of the amino acid. In some embodiments of the method 1100, the feature is to display a supramolecular assembly of the amino acid. In some embodiments of the method 1100, the feature is to display a sequence alignment of the amino acid. In some embodiments of the method 1100, the feature is to display docking results of the amino acid. In some embodiments of the method 1100, the feature is displaying a trajectory of an amino acid. In some embodiments of the method 1100, the feature is displaying a conformational ensemble of amino acids. In some embodiments of the method 1100, the feature is displaying a secondary structure of an amino acid. In some embodiments of the method 1100, the feature is displaying a tertiary structure of an amino acid. In some embodiments of the method 1100, the feature is displaying a quaternary structure of an amino acid. In some embodiments of the method 1100, the feature is setting a zoom factor for displaying the structural rendition. In some embodiments of the method 1100, the feature is setting a lighting condition for displaying the structural rendition. In some embodiments of the method 1100, the feature is setting a visibility range for displaying the structural rendition. In some embodiments of the method 1100, the feature is color coding the structural rendition by amino acid.In some embodiments of method 1100, the feature configuration is color-coding the structural rendition by evolutionary conservation of amino acids. In some embodiments of method 1100, the feature configuration is color-coding the structural rendition by structural quality of amino acids. In some embodiments of method 1100, the feature configuration is color-coding the structural rendition to identify gaps / missing amino acids.

[0100] In some embodiments of method 1100, the pathogenicity determiner determines the pathogenicity of the nucleotide variant by processing a plurality of images as input and generating a pathogenicity score for the alternative amino acid as output. In some such examples, the pathogenicity determiner is a neural network. In some such examples, the neural network is a convolutional neural network. In some such examples using a CNN, the pathogenicity determiner processes each image in the plurality of images through a respective CNN to generate a respective feature map for each image. Also, in some examples, each convolutional neural network is a respective instance of the same convolutional neural network. Also, in some examples, the pathogenicity determiner combines each feature map into a pooled feature map and processes the pooled feature map through a final convolutional neural network to generate a pathogenicity score for the alternative amino acid.

[0101] Also, in some examples, each image in the plurality of images is combined into a combined representation for processing by the pathogenicity determiner, and in some such examples, each image is arranged as a respective color channel in the combined representation, and intensity values ​​in each image may be summed pixel by pixel into the combined representation.

[0102] Method 1200 includes the steps of method 1000 (and in some embodiments, also the steps of method 1100), as well as steps 1202, 1204, 1206, and 1208. One of the objectives of method 1200 is to train a pathogenicity determiner from method 1000. Also, such steps may be used to train any of the 2D image classification networks described herein. Step 1202 includes capturing a respective plurality of images for each target amino acid mutated to a respective alternative amino acid. Step 1204 includes assigning a benign truth label to a first subset of the plurality of images based on the benignity of the corresponding alternative amino acid. In other words, step 1204 includes assigning a benign truth label to the first subset of the plurality of images if the corresponding alternative amino acid is benign. Step 1206 includes assigning a pathogenic truth label to a second subset of the plurality of images based on the pathogenicity of the corresponding alternative amino acid. Step 1208 includes training a pathogenicity determiner using the plurality of images with assigned benign and pathogenic truth labels.

[0103] 13 shows a block diagram of an exemplary embodiment of a computing system 1300 that may include, be, or be part of any one of the electronic or computing systems described herein. FIG 13 illustrates a portion of the computing system 1300 on which a set of instructions for causing a machine to perform any one or more of the methodologies discussed herein is executed.

[0104] In some embodiments, the computing system 1300 corresponds to a host system that includes, is coupled to, or utilizes memory, or is used to perform operations performed by any one of the computing devices, data processors, and user interface devices described herein. In alternative embodiments, the machine is connected (e.g., networked) to other machines in a LAN, an intranet, an extranet, or the Internet. In some embodiments, the machine operates in the capacity of a server or a client machine in a client-server network environment, as a peer machine in a peer-to-peer (or distributed) network environment, or as a server or a client machine in a cloud computing infrastructure or environment. In some embodiments, the machine is a personal computer (PC), a tablet PC, a cellular phone, a web appliance, a server, or any machine capable of executing (in sequence or otherwise) a set of instructions that specify actions to be taken by the machine. Furthermore, although a single machine is illustrated, the term "machine" is also intended to include any collection of machines that individually or jointly execute a set (or sets) of instructions to perform any one or more of the methodologies discussed herein.

[0105] The computing system 1300 includes a processing device 1302, a main memory 1304 (e.g., read-only memory (ROM), flash memory, dynamic random-access memory (DRAM), etc.), a static memory 1306 (e.g., flash memory, static random-access memory (SRAM), etc.), and a data storage system 1310, which communicate with each other via a bus 1330. The processing device 1302 represents one or more general-purpose processing devices, such as a microprocessor, a central processing unit, etc. More specifically, the processing device is a microprocessor or a processor that implements other instruction sets, or a processor that implements a combination of instruction sets. Alternatively, the processing device 1302 is one or more special-purpose processing devices, such as an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a digital signal processor (DSP), a network processor, etc. The processing device 1302 is configured to execute instructions 1314 for performing the operations or steps discussed herein. In some implementations, the computing system 1300 includes a network interface device 1308 for communicating over a communications network 1340 shown in FIG.

[0106] The data storage system 1310 includes a machine-readable storage medium 1312 (also known as a computer-readable medium) on which one or more sets of instructions 1314 or software embodying any one or more of the methodologies or functions described herein are stored. The instructions 1314 also reside, completely or at least partially, within the main memory 1304 or within the processing device 1302 during execution thereof by the computing system 1300, with the main memory 1304 and the processing device 1302 also constituting machine-readable storage media.

[0107] In some implementations, the instructions 1314 include instructions for performing functions corresponding to any one of a computing device, a data processor, a user interface device, and an I / O device described herein. Although the machine-readable storage medium 1312 is shown in the exemplary implementation as being a single medium, the term "machine-readable storage medium" should be interpreted to include a single medium or multiple media that store one or more sets of instructions. The term "machine-readable storage medium" should also be interpreted to include any medium capable of storing or encoding a set of instructions for execution by a machine, causing the machine to perform any one or more of the methodologies of the present disclosure. Thus, the term "machine-readable storage medium" should be interpreted to include solid-state memory, optical media, magnetic media, and the like.

[0108] Also as shown, computing system 1300 includes a user interface 1320, which in some implementations includes a display, for example, performing functionality corresponding to any one of the user interface devices disclosed herein. A user interface, such as user interface 1320, or a user interface device as described herein, includes any space or equipment where interaction between a human and a machine occurs. A user interface as described herein allows operation and control of a machine from a human user, while the machine provides feedback information to the user. Examples of user interfaces (UIs) or user interface devices include computer operating systems (such as graphic user interfaces or GUIs), machine operator controls, and interactive aspects of process control. A UI as described herein includes one or more layers, including a human-machine interface (HMI) that interfaces the machine with physical input and output hardware.

[0109] It should also be understood that the methodologies discussed herein are computer-implemented methods and, in some embodiments, may be implemented by the computing system 1300. For example, the computer-implemented method includes processing a plurality of missense variants by an artificial neural network (ANN) to generate a plurality of missense scores associated with a respective pathogenicity for each missense variant of the plurality of missense variants. The computer-implemented method also includes processing a plurality of indels by the ANN to generate a plurality of indel scores associated with a respective pathogenicity for each indel of the plurality of indels, and further processing the plurality of indel scores and the plurality of missense scores to be applied to one or more curve-forming functions. The computer-implemented method also includes applying the processed scores to the curve-forming functions to generate respective indel curves and respective missense curves, and determining the difference between the respective indel curves and the respective missense curves. The computer-implemented method also includes determining one or more scaling functions to reduce the difference between the curves, and modifying the ANN or the output of the ANN according to the scaling functions. Modifying the ANN or the output of the ANN according to a scaling function includes enhancing the scores of the plurality of indels according to a scaling function to provide an increased accuracy of pathogenicity for each indel of the plurality of indels.

[0110] Some portions of the above detailed description are presented in terms of models and symbolic representations of operations on data bits within a computer memory. These models and representations are the manner used by those skilled in the data processing arts to most effectively convey the substance of their work to others skilled in the art. The models are herein, and generally, considered to be self-consistent sequences of operations leading to a given result. The operations require physical manipulations of physical quantities. Usually, although not necessarily, these quantities take the form of electrical or magnetic signals capable of being stored, combined, compared, and otherwise manipulated. It has proven convenient at times, primarily for reasons of common usage, to refer to these signals as bits, values, elements, symbols, characters, terms, numbers, or the like.

[0111] It should be noted, however, that these and similar terms are to be associated with the appropriate physical quantities and are merely convenient labels applied to these quantities. This disclosure may refer to actions and processes of a computing system, or similar electronic computing device, that manipulate and convert data represented as physical (electronic) quantities in the computing system's registers and memory into other data that is similarly represented as physical quantities in the computing system's memory or registers or other such information storage systems.

[0112] The present disclosure also relates to an apparatus for performing the operations herein. The apparatus may be specially constructed for the intended purposes or may include a general-purpose computer selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a computer-readable storage medium coupled to a computer system bus, such as any type of disk, including optical disks, CD-ROMs, and magneto-optical disks, read-only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical cards, or any type of medium suitable for storing electronic instructions.

[0113] The models and representations presented herein are not inherently related to any particular computer or other apparatus. Various general-purpose systems may be used with programs in accordance with the teachings herein, or it may prove convenient to construct a more specialized apparatus to implement the methods. The structure of a variety of these systems will become apparent as described below. In addition, the present disclosure is not described with reference to any particular programming language. It will be understood that a variety of programming languages ​​can be used to implement the teachings of the present disclosure as described herein.

[0114] The present disclosure may be provided as a computer program product or software that may include a machine-readable medium having stored thereon instructions that may be used to program a computing system (or other electronic device) to perform a process according to the present disclosure. A machine-readable medium includes any mechanism for storing information in a form readable by a machine (e.g., a computer). In some embodiments, the machine-readable (e.g., computer-readable) medium includes a machine (e.g., computer) readable storage medium, such as a read-only memory ("ROM"), a random access memory ("RAM"), a magnetic disk storage medium, an optical storage medium, a flash memory component, or the like.

[0115] While the present invention has been described in conjunction with specific embodiments illustrated herein, it is apparent that many alternatives, combinations, modifications, and variations will be apparent to those skilled in the art. Accordingly, the exemplary embodiments of the invention described herein are intended to be illustrative only and not in a limiting sense. Various changes may be made without departing from the spirit and scope of the invention.

[0116] As used herein, "logic" (e.g., a curve forming function) may be realized in the form of a computer product including a non-transitory computer-readable storage medium with computer usable program code to perform the method steps described herein. "Logic" may be realized in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the exemplary method steps. "Logic" may be realized in the form of a means for performing one or more of the method steps described herein. The means may include (i) hardware modules, (ii) software modules running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing certain techniques described herein, the software modules being stored in a computer-readable storage medium (or multiple such media). In one embodiment, logic performs a data processing function. The logic may be a general-purpose, single-core or multi-core processor with a computer program specifying the function, a digital signal processor with a computer program, configurable logic such as an FPGA with a configuration file, special purpose circuitry such as a state machine, or any combination thereof. Also, the computer program product may embody computer program and configuration file portions of logic.

[0117] Terms The disclosed technology can be implemented as a system, a method, or a product. One or more features of the embodiments can be combined with the base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform users of these options. The omission from some embodiments of the enumeration of repeating these options should not be interpreted as limiting the combinations taught in the preceding section. These descriptions are incorporated herein by reference in each of the following implementations.

[0118] One or more embodiments and provisions of the disclosed technology, or elements thereof, can be realized in the form of a computer product including a non-transitory computer-readable storage medium with computer usable program code for performing the illustrated method steps. Furthermore, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be realized in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be realized in the form of a means for performing one or more of the method steps described herein, which means can include (i) a hardware module, (ii) a software module running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing the particular technology described herein, and the software module is stored in a computer-readable storage medium (or multiple such media).

[0119] The clauses described in this section can be combined as features. For the sake of brevity, combinations of features are not listed separately and are not repeated for each base set of features. The reader will understand how features specified in the clauses described in this section can be easily combined with sets of basic features specified as embodiments in other sections of this application. These clauses are not meant to be mutually exclusive, exhaustive, or restrictive, and the disclosed technology is not limited to these clauses, but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.

[0120] Other implementations of the provisions described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the provisions described in this section. Yet another implementation of the provisions described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the provisions described in this section.

[0121] The inventors disclose the following provisions: 1. A computer-implemented method comprising: generating a plurality of two-dimensional images of the protein structure from the three-dimensional image of the protein structure; and processing a plurality of two-dimensional images with a two-dimensional image classification network to generate a classification of protein structures. 2. The computer-implemented method of clause 1, wherein the protein structure comprises a target amino acid. 3. The computer-implemented method of claim 1, wherein the classification of the protein structure includes a pathogenicity score. 4. The computer-implemented method of clause 1, wherein processing the plurality of two-dimensional images includes processing the plurality of two-dimensional images by a two-dimensional image classification network to determine pathogenicity of a nucleotide variant that results in the replacement of a target amino acid with an alternative amino acid in a protein structure. 5. The computer-implemented method of claim 4, wherein the classification of the protein structure includes a pathogenicity score. 6. generating the plurality of two-dimensional images includes generating a plurality of graphical representations of amino acids in the protein structure from the three-dimensional graphical representation of the amino acids; a three-dimensional image of the protein structure comprising a three-dimensional graphical representation of amino acids; 2. The computer-implemented method of claim 1, wherein each image of the plurality of two-dimensional images comprises a respective graphical representation of an amino acid of a plurality of graphical representations of amino acids. 7. The computer-implemented method of clause 6, wherein the amino acids include target amino acids. 8. The computer-implemented method of clause 7, wherein processing the plurality of two-dimensional images includes processing the plurality of two-dimensional images by a two-dimensional image classification network to determine pathogenicity of nucleotide variants that result in the replacement of a target amino acid with an alternative amino acid. 9. The computer-implemented method of claim 8, wherein the classification of the protein structure includes a pathogenicity score. 10. The generation of multiple two-dimensional images selecting a three-dimensional image of a protein structure by a protein viewer; Magnifying a three-dimensional image of a protein structure with a protein viewer; coloring, with a protein viewer, the three-dimensional image of the protein structure by amino acid type, including coloring each amino acid of the plurality of graphical representations of amino acids by amino acid type; and capturing a plurality of two-dimensional images from different viewpoints after the three-dimensional image is magnified and colored by the protein viewer. 11. The generation of multiple two-dimensional images selecting a three-dimensional image of a protein structure by a protein viewer; Magnifying a three-dimensional image of a protein structure with a protein viewer; coloring, by the protein viewer, the three-dimensional image of the protein structure according to coloring parameters of the protein viewer; and capturing a plurality of two-dimensional images from different viewpoints after the three-dimensional image is magnified and colored by a protein viewer. 12. The computer-implemented method of claim 11, wherein the coloring parameter is an amino acid type and the coloring of the three-dimensional image of the protein structure is by amino acid type. 13. The computer-implemented method of claim 11, wherein the coloring parameter is a conservation of protein structure and the coloring of the three-dimensional image of the protein structure is in accordance with the conservation of the protein structure. 14. The computer-implemented method of claim 11, wherein the coloring parameter is a structural quality and the coloring of the three-dimensional image of the protein structure is according to the structural quality. 15. The computer-implemented method of claim 11, wherein the three-dimensional image of the protein structure includes a graphical representation of atoms. 16. The computer-implemented method of claim 11, wherein the three-dimensional image of the protein structure includes a graphical representation of the residues. 17. The computer-implemented method of claim 11, wherein the three-dimensional image of the protein structure includes a graphical representation of atoms and residues. 18. 2D image classification network, For different 3D images of different protein structures, (1) generating a plurality of different two-dimensional images of different protein structures from a three-dimensional image of the different protein structures; and (2) repeating the steps of processing different plurality of two-dimensional images with the two-dimensional image classification network to generate classifications of different protein structures; 2. The computer-implemented method of claim 1, further comprising, after each iteration of steps (1) and (2), adjusting parameters of the two-dimensional image classification network according to the generated classifications of different protein structures, thereby training. 19. The computer-implemented method of clause 18, wherein processing the plurality of two-dimensional images includes processing the plurality of two-dimensional images by a two-dimensional image classification network to determine pathogenicity of a nucleotide variant that results in the replacement of a target amino acid with an alternative amino acid in a protein structure. 20. The computer-implemented method of clause 19, wherein the classification of the protein structure includes a pathogenicity score. 21. A computer-implemented method comprising: generating a plurality of two-dimensional images of the protein structure from the three-dimensional image of the protein structure; Processing a plurality of two-dimensional images using a plurality of convolutional neural networks to generate a plurality of first layer outputs; each output of the plurality of first layer outputs is for a respective image of the plurality of two-dimensional images; generating a plurality of convolutional neural networks, each convolutional neural network for a respective image of the plurality of two-dimensional images; combining a plurality of first layer outputs into a second layer input; and processing the second layer inputs using a second layer convolutional neural network to generate a classification of protein structures. 22. The computer-implemented method of clause 21, wherein the protein structure comprises a target amino acid. 23. The computer-implemented method of claim 21, wherein the classification of the protein structure includes a pathogenicity score. 24. The computer-implemented method of clause 21, wherein processing the plurality of two-dimensional images includes processing the plurality of two-dimensional images by a two-dimensional image classification network to determine pathogenicity of a nucleotide variant that results in the replacement of a target amino acid with an alternative amino acid in a protein structure, the two-dimensional image classification network including a plurality of convolutional neural networks and a second layer convolutional neural network. 25. The computer-implemented method of claim 24, wherein the classification of the protein structure includes a pathogenicity score. 26. generating the plurality of two-dimensional images includes generating a plurality of graphical representations of amino acids in the protein structure from the three-dimensional graphical representation of the amino acids; a three-dimensional image of the protein structure comprising a three-dimensional graphical representation of amino acids; 22. The computer-implemented method of claim 21, wherein each image of the plurality of two-dimensional images comprises a respective graphical representation of an amino acid of the plurality of graphical representations of amino acids. 27. The computer-implemented method of clause 26, wherein the amino acids include target amino acids. 28. The computer-implemented method of clause 27, wherein processing the plurality of two-dimensional images includes processing the plurality of two-dimensional images by a two-dimensional image classification network to determine pathogenicity of a nucleotide variant that replaces a target amino acid with an alternative amino acid, the two-dimensional image classification network including a plurality of convolutional neural networks and a second layer convolutional neural network. 29. The computer-implemented method of clause 28, wherein the classification of the protein structure includes a pathogenicity score. 30. The generation of multiple two-dimensional images is selecting a three-dimensional image of a protein structure by a protein viewer; Magnifying a three-dimensional image of a protein structure with a protein viewer; coloring, with a protein viewer, the three-dimensional image of the protein structure by amino acid type, including coloring each amino acid of the plurality of graphical representations of amino acids by amino acid type; and capturing a plurality of two-dimensional images from different viewpoints after the three-dimensional image is magnified and colored by the protein viewer. 31. The generation of multiple two-dimensional images is selecting a three-dimensional image of a protein structure by a protein viewer; Magnifying a three-dimensional image of a protein structure with a protein viewer; coloring, by the protein viewer, the three-dimensional image of the protein structure according to coloring parameters of the protein viewer; and capturing a plurality of two-dimensional images from different viewpoints after the three-dimensional image is magnified and colored by the protein viewer. 32. The computer-implemented method of claim 31, wherein the coloring parameter is an amino acid type and the coloring of the three-dimensional image of the protein structure is by amino acid type. 33. The computer-implemented method of claim 31, wherein the coloring parameter is a conservation of protein structure and the coloring of the three-dimensional image of the protein structure is in accordance with the conservation of the protein structure. 34. The computer-implemented method of claim 31, wherein the coloring parameter is a structural quality and the coloring of the three-dimensional image of the protein structure is according to the structural quality. 35. The computer-implemented method of clause 31, wherein the three-dimensional image of the protein structure includes a graphical representation of atoms. 36. The computer-implemented method of clause 35, wherein the three-dimensional image of the protein structure includes a graphical representation of the residues. 37. the plurality of convolutional neural networks and the second layer convolutional neural network are part of a two-dimensional image classification network; The method comprises: For different 3D images of different protein structures, (1) generating a plurality of different two-dimensional images of different protein structures from a three-dimensional image of the different protein structures; (2) processing different two-dimensional images using a plurality of convolutional neural networks to generate different first layer outputs; (3) combining different first layer outputs with different second layer inputs; (4) repeating the step of processing different second-layer inputs using a second-layer convolutional neural network to generate classifications of different protein structures; 22. The computer-implemented method of claim 21, further comprising, after each iteration of steps (1), (2), (3), and (4), training the two-dimensional image classification network by adjusting parameters of the network according to the generated classifications of the different protein structures. 38. The computer-implemented method of clause 37, wherein tuning parameters of the two-dimensional image classification network includes tuning parameters of a plurality of convolutional neural networks. 39. The computer-implemented method of clause 37, wherein tuning parameters of the two-dimensional image classification network includes tuning parameters of a second layer convolutional neural network. 40. The computer-implemented method of clause 37, wherein adjusting parameters of the two-dimensional image classification network includes adjusting parameters of a pooling layer, the pooling layer performing a combination of outputs of different first layers to inputs of a second layer. 41. A system comprising: a plurality of virtual cameras configured to generate a plurality of two-dimensional images of the protein structure from the three-dimensional image of the protein structure, a plurality of virtual cameras, each camera of the plurality of virtual cameras configured to capture a respective image of the plurality of two-dimensional images; a two-dimensional image classification network configured to process a plurality of two-dimensional images to generate a classification of protein structures. 42. The system according to clause 41, wherein the protein structure comprises a target amino acid. 43. The system of claim 41, wherein the classification of the protein structure includes a pathogenicity score. 44. The system of claim 41, wherein the two-dimensional image classification network is configured to process the plurality of two-dimensional images to determine the pathogenicity of a nucleotide variant that results in the replacement of a target amino acid with an alternative amino acid in a protein structure, and the classification includes the pathogenicity of the nucleotide variant. 45. The system according to clause 44, wherein the pathogenicity of the nucleotide variant is represented by a pathogenicity score. 46. a plurality of virtual cameras configured to generate a plurality of graphical representations of amino acids in the protein structure from the three-dimensional graphical representation of the amino acids; a three-dimensional image of the protein structure comprising a three-dimensional graphical representation of amino acids; 42. The system of claim 41, wherein each image of the plurality of two-dimensional images comprises a respective graphical representation of an amino acid of the plurality of graphical representations of amino acids. 47. The system according to clause 46, wherein the amino acid comprises a target amino acid. 48. The system described in clause 47, wherein the two-dimensional image classification network is configured to process a plurality of two-dimensional images to determine the pathogenicity of a nucleotide variant that results in the replacement of a target amino acid with an alternative amino acid. 49. The system according to clause 48, wherein the classification of the protein structure includes a pathogenicity score. 50. A plurality of virtual cameras are part of a protein viewer, the protein viewer comprising: selecting a three-dimensional image of a protein structure; Magnifying a three-dimensional image of a protein structure; coloring each amino acid of the plurality of graphical representations of amino acids by amino acid type in the three-dimensional image of the protein structure; and capturing a plurality of two-dimensional images from different viewpoints via a plurality of virtual cameras. 51. A plurality of virtual cameras are part of a protein viewer, the protein viewer comprising: selecting a three-dimensional image of a protein structure; Magnifying a three-dimensional image of a protein structure; coloring the three-dimensional image of the protein structure according to coloring parameters of the protein viewer; and capturing a plurality of two-dimensional images from different viewpoints via a plurality of virtual cameras. 52. The system described in clause 51, wherein the colouring parameter is an amino acid type and the protein viewer is configured to colour the three-dimensional image of the protein structure by amino acid type. 53. The system described in clause 51, wherein the colouring parameter is a conservation of protein structure, and the protein viewer is configured to colour the three-dimensional image of the protein structure according to the conservation of the protein structure. 54. The system described in clause 51, wherein the colouring parameter is a structural quality and the protein viewer is configured to colour the three-dimensional image of the protein structure according to the structural quality. 55. The system of claim 51, wherein the three-dimensional image of the protein structure includes a graphical representation of atoms. 56. The system according to clause 51, wherein the three-dimensional image of the protein structure includes a graphical representation of the residues. 57. The system of claim 51, wherein the three-dimensional image of the protein structure includes a graphical representation of atoms and residues. 58.To train a 2D image classification network, For different 3D images of different protein structures, (1) generating a plurality of different two-dimensional images of different protein structures from a three-dimensional image of the different protein structures; and (2) repeating the steps of processing different plurality of two-dimensional images with the two-dimensional image classification network to generate classifications of different protein structures; Clause 43. The system of clause 42, further comprising a computing system configured to: after each iteration of steps (1) and (2), adjust parameters of the two-dimensional image classification network according to the generated classifications of the different protein structures. 59. The system of clause 58, wherein the two-dimensional image classification network is configured to process a plurality of two-dimensional images to determine the pathogenicity of a nucleotide variant that results in the replacement of a target amino acid with an alternative amino acid. 60. The system according to clause 59, wherein the classification of the protein structure includes a pathogenicity score. 61. A system comprising: a plurality of virtual cameras configured to generate a plurality of two-dimensional images of the protein structure from the three-dimensional image of the protein structure, a plurality of virtual cameras, each camera of the plurality of virtual cameras configured to capture a respective image of the plurality of two-dimensional images; a first convolutional layer including a plurality of convolutional neural networks configured to generate a plurality of first layer outputs based on a plurality of two-dimensional images, each output of the plurality of first layer outputs is for a respective image of the plurality of two-dimensional images; a first convolutional layer, each of the plurality of convolutional neural networks being for a respective image of the plurality of two-dimensional images; a pooling layer configured to combine a plurality of first layer outputs into a second layer input; a second convolutional layer including a second layer convolutional neural network configured to generate a classification of protein structures based on the second layer input. 62. The system according to clause 61, wherein the protein structure comprises a target amino acid. 63. The system according to clause 61, wherein the classification of the protein structure includes a pathogenicity score. 64. Equipped with a two-dimensional image classification network; a two-dimensional image classification network configured to process the plurality of two-dimensional images to determine pathogenicity of a nucleotide variant that substitutes a target amino acid with an alternative amino acid; 63. The system of claim 62, wherein the two-dimensional image classification network includes a first convolutional layer, a pooling layer, and a second convolutional layer. 65. The system according to clause 64, wherein the classification of the protein structure includes a pathogenicity score. 66. a plurality of virtual cameras configured to generate a plurality of graphical representations of amino acids in the protein structure from the three-dimensional graphical representation of the amino acids; a three-dimensional image of the protein structure comprising a three-dimensional graphical representation of amino acids; 62. The system of claim 61, wherein each image of the plurality of two-dimensional images includes a respective graphical representation of an amino acid among the plurality of graphical representations of amino acids. 67. The system according to clause 66, wherein the amino acid comprises a target amino acid. 68. Equipped with a two-dimensional image classification network; a two-dimensional image classification network configured to process the plurality of two-dimensional images to determine pathogenicity of a nucleotide variant that substitutes a target amino acid with an alternative amino acid; 68. The system of claim 67, wherein the two-dimensional image classification network includes a first convolutional layer, a pooling layer, and a second convolutional layer. 69. The system according to clause 68, wherein the classification of the protein structure includes a pathogenicity score. 70. A plurality of virtual cameras are part of a protein viewer, the protein viewer comprising: selecting a three-dimensional image of a protein structure; Magnifying a three-dimensional image of a protein structure; coloring each amino acid of the plurality of graphical representations of amino acids by amino acid type in the three-dimensional image of the protein structure; 69. The system of claim 68, configured to capture a plurality of two-dimensional images from different viewpoints via a plurality of virtual cameras. 71. A plurality of virtual cameras are part of a protein viewer, the protein viewer comprising: selecting a three-dimensional image of a protein structure; Magnifying a three-dimensional image of a protein structure; coloring the three-dimensional image of the protein structure according to coloring parameters of the protein viewer; 62. The system of claim 61, configured to capture a plurality of two-dimensional images from different viewpoints via a plurality of virtual cameras. 72. The system according to clause 71, wherein the colouring parameter is an amino acid type and the colouring of the three-dimensional image of the protein structure is according to the amino acid type. 73. The system according to claim 71, wherein the colouring parameter is the conservation of protein structure and the colouring of the three-dimensional image of the protein structure is according to the conservation of the protein structure. 74. The system according to clause 71, wherein the colouring parameter is a structural quality and the colouring of the three-dimensional image of the protein structure is according to the structural quality. 75. The system of claim 61, wherein the three-dimensional image of the protein structure includes a graphical representation of atoms. 76. The system according to clause 75, wherein the three-dimensional image of the protein structure includes a graphical representation of the residues. 77. a two-dimensional image classification network including a first convolutional layer, a pooling layer, and a second convolutional layer; To train a 2D image classification network, For different 3D images of different protein structures, (1) generating a plurality of different two-dimensional images of different protein structures from a three-dimensional image of the different protein structures via a plurality of virtual cameras; (2) generating different first layer outputs based on the different two-dimensional images via a plurality of convolutional neural networks; (3) combining different first layer outputs with different second layer inputs via a pooling layer; and (4) repeating the step of generating, via a second layer convolutional neural network, different protein structure classifications based on the different second layer inputs; and and after each iteration of steps (1), (2), (3), and (4), adjusting parameters of the two-dimensional image classification network according to the generated classifications of the different protein structures. 78. The system of clause 77, wherein the computing system is configured to tune parameters of the multiple convolutional neural networks. 79. The system of clause 77, further comprising tuning parameters of a second layer convolutional neural network when the computing system is configured to tune parameters of the two-dimensional image classification network. 80. The system of clause 77, wherein the computing system is configured to adjust parameters of the pooling layer. 81. A computer-implemented method for determining pathogenicity of a variant, comprising: accessing a structural rendition of the amino acids of a protein; capturing a plurality of images of a portion of the structural rendition that includes the target amino acid from the amino acids; and determining a pathogenicity of a nucleotide variant that mutates a target amino acid to an alternative amino acid based at least in part on the plurality of images. 82. The computer-implemented method of clause 81, wherein images in the plurality of images are captured from multiple perspectives. 83. The computer-implemented method of clause 82, wherein the images are captured from multiple viewpoints. 84. The computer-implemented method of clause 82, wherein the images are captured from multiple orientations. 85. The computer-implemented method of clause 82, wherein the images are captured from multiple positions. 86. The computer-implemented method of clause 82, wherein images are captured from multiple zoom levels. 87. The computer-implemented method of clause 81, further comprising activating a feature configuration of the structural rendition prior to capturing the plurality of images, and capturing the plurality of images of the portion of the structural rendition using the activated feature configuration. 88. The computer-implemented method of clause 87, wherein the feature is displaying residue details of amino acids. 89. The computer-implemented method of claim 87, wherein the feature is displaying atomic details of amino acids. 90. The computer-implemented method of claim 87, wherein the feature is to display polymer details of amino acids. 91. The computer-implemented method of claim 87, wherein the characteristic configuration is to display ligand details of amino acids. 92. The computer-implemented method of claim 87, wherein the feature is to display amino acid chain details. 93. The computer-implemented method of claim 87, wherein the feature represents a surface effect of an amino acid. 94. The computer-implemented method of claim 93, wherein the surface effects include transparency and color coding for electrostatic and hydrophobic values. 95. The computer-implemented method of claim 87, wherein the feature configuration is to display a density map of amino acids. 96. The computer-implemented method of clause 87, wherein the feature represents a supramolecular assembly of amino acids. 97. The computer-implemented method of claim 87, wherein the feature is displaying an alignment of amino acid sequences. 98. The computer-implemented method of claim 87, wherein the feature configuration is to display the amino acid docking results. 99. The computer-implemented method of claim 87, wherein the feature is to display a trajectory of amino acids. 100. The computer-implemented method of clause 87, wherein the feature configuration represents a conformational ensemble of amino acids. 101. The computer-implemented method of clause 87, wherein the feature represents a secondary structure of amino acids. 102. The computer-implemented method of clause 87, wherein the feature represents a tertiary structure of an amino acid. 103. The computer-implemented method of clause 87, wherein the feature represents a quaternary structure of an amino acid. 104. The computer-implemented method of clause 87, wherein the feature configuration is setting a zoom factor for displaying the structural rendition. 105. The computer-implemented method of clause 87, wherein the feature is setting lighting conditions for displaying the structural rendition. 106. The computer-implemented method of clause 87, wherein the feature is to set a visible range for displaying the structural rendition. 107. The computer-implemented method of clause 87, wherein the feature configuration is color-coding the structural rendition by amino acid. 108. The computer-implemented method of clause 87, wherein the feature configuration is color-coding the structural renditions by evolutionary conservation of amino acids. 109. The computer-implemented method of clause 87, wherein the feature configuration is color-coding the structural rendition by structural quality of the amino acids. 110. The computer-implemented method of clause 81, wherein the pathogenicity determiner determines the pathogenicity of the nucleotide variant by processing a plurality of images as input and generating pathogenicity scores for alternative amino acids as output. 111. The computer-implemented method of clause 110, wherein the pathogenicity determiner is a neural network. 112. The computer-implemented method of clause 111, wherein the neural network is a convolutional neural network. 113. The computer-implemented method of clause 112, wherein the pathogenicity determiner processes each image in the plurality of images through a respective convolutional neural network to generate a respective feature map for each image. 114. The computer-implemented method of clause 113, wherein each convolutional neural network is a respective instance of the same convolutional neural network. 115. The computer-implemented method of clause 114, wherein the pathogenicity determiner combines each feature map into a pooled feature map and processes the pooled feature map through a final convolutional neural network to generate pathogenicity scores for the alternative amino acids. 116. The computer-implemented method of claim 113, wherein each image in the plurality of images is combined into a combined representation for processing by the pathogenicity determiner. 117. The computer-implemented method of clause 116, wherein each image is arranged as a respective color channel in the combined representation. 118. The computer-implemented method of clause 116, wherein intensity values ​​in each image are summed pixel by pixel into a pixel of the combined representation. 119. The computer-implemented method of clause 81, further comprising capturing a respective plurality of images for each target amino acid to be mutated to a respective alternative amino acid. 120. The computer-implemented method of clause 119, further comprising assigning a benign truth label to a first subset of each of the plurality of images based on the benignity of the corresponding alternative amino acids. 121. The computer-implemented method of clause 120, further comprising assigning a pathogenicity truth label to a second subset of each of the plurality of images based on the pathogenicity of the corresponding alternative amino acid. 122. The computer-implemented method of clause 121, further comprising training a pathogenicity determiner using the plurality of images each having an assigned benign and pathogenic truth label. 123. The computer-implemented method of clause 81, wherein the portion of the structural rendition comprises a target amino acid and a number of additional amino acids adjacent to the target amino acid. 124. The computer-implemented method of clause 81, wherein the structural rendition is a three-dimensional (3D) structural rendition. 125. A computer-implemented method comprising: Accessing the image data; and Modifying the image data to encode a plurality of properties / characteristics of the protein; and using the modified image data for further processing. 126. The computer-implemented method of clause 125, wherein modifying the image data includes modifying red, blue, and green (R, G, B) values ​​of pixels of the image data. [Explanation of symbols]

[0122] 200 Systems 202 2D Image Classification Network 204 Virtual Camera 206 2D images of protein structures 208 3D images of protein structures 210 Classification of Protein Structures 212 Protein Viewer 214 Computing Systems 700 System 702 Convolutional Neural Network (CNN) 704 First Convolutional Layer 706a, 706b, 706c CNN 708a, 708b, 708c Output of the first layer 710 Pooling Layer 712 Second layer input 714 Second convolution layer 716 CNN 720 2D Image Classification Network 1300 Computing System 1302 Processing Device 1304 Memory 1306 Static Memory 1308 Network Interface Device 1310 Data Storage System 1312 Machine-readable medium 1314 Instructions 1320 User Interface 1330 Bus 1340 Network

Claims

1. 1. A computer-implemented method for determining pathogenicity of a variant, comprising: accessing a three-dimensional structural rendition of the amino acids of a protein; generating a plurality of two-dimensional images of a portion of said three-dimensional structural rendition comprising a target amino acid from said amino acids; and determining a pathogenicity score of a nucleotide variant that mutates the target amino acid to an alternative amino acid using a convolutional neural network that processes at least the plurality of two-dimensional images.

2. generating a plurality of two-dimensional images of a portion of said three-dimensional structural rendition comprising said amino acid sequence and a target amino acid sequence; The computer-implemented method of claim 1 , comprising using a plurality of virtual cameras configured to capture respective images of the plurality of two-dimensional images.

3. The computer-implemented method of claim 1 , wherein the plurality of two-dimensional images are captured from multiple viewpoints, multiple orientations, multiple positions, or multiple zoom levels.

4. activating a feature of the three-dimensional structural rendition prior to capturing the plurality of two-dimensional images; The computer-implemented method of claim 1 , further comprising capturing the plurality of two-dimensional images of the portion of the three-dimensional structural rendition with the activated feature configuration.

5. 5. The computer-implemented method of claim 4, wherein the feature configuration displays one or more of: a residue detail of the amino acid, an atomic detail of the amino acid, a polymer detail of the amino acid, a ligand detail of the amino acid, a chain detail of the amino acid, a surface effect of the amino acid, a density map of the amino acid, a supramolecular assembly of the amino acid, a sequence alignment of the amino acid, a docking result of the amino acid, a trajectory of the amino acid, a conformational ensemble of the amino acid, a secondary structure of the amino acid, a tertiary structure of the amino acid, or a quaternary structure of the amino acid.

6. The computer-implemented method of claim 5 , wherein the surface effects include transparency and color coding for electrostatic and hydrophobic values.

7. The characteristic configuration is setting one or more of a zoom factor for displaying the three-dimensional structural rendition, a lighting condition for displaying the three-dimensional structural rendition, or a visibility range for displaying the three-dimensional structural rendition; or 5. The computer-implemented method of claim 4, comprising color-coding the three-dimensional structural rendition by one or more of the amino acid, the evolutionary conservation of the amino acid, or the structural quality of the amino acid.

8. 1. A system comprising: at least one processor; a non-transitory computer-readable storage medium having instructions stored thereon; The instructions, when executed by the at least one processor, provide the system with: accessing a three-dimensional structural rendition of the amino acids of a protein; generating a plurality of two-dimensional images of a portion of said three-dimensional structural rendition comprising a target amino acid from said amino acids; and determining a pathogenicity score of a nucleotide variant that mutates the target amino acid to an alternative amino acid using a convolutional neural network that processes at least the plurality of two-dimensional images.

9. 9. The system of Claim 8, wherein the convolutional neural network determines the pathogenicity scores for the nucleotide variants by processing the plurality of two-dimensional images as inputs and generating the pathogenicity scores for the alternative amino acids as outputs.

10. 10. The system of claim 8, wherein the convolutional neural network includes one or more convolutional layers and a pooling layer.

11. 11. The system of claim 10, wherein a first convolutional layer of the one or more convolutional layers comprises a respective convolutional neural network.

12. 12. The system of claim 11, wherein the convolutional neural network processes each image in the plurality of two-dimensional images through the respective convolutional neural network to generate a respective feature map for the each image.

13. 13. The system of claim 12, wherein the respective convolutional neural networks are respective instances of the same convolutional neural network.

14. 14. The system of claim 13, wherein the convolutional neural network combines the respective feature maps into a pooled feature map and processes the pooled feature map through a final convolutional neural network to generate the pathogenicity score for the alternative amino acid.

15. 13. The system of claim 12, wherein the respective images in the plurality of two-dimensional images are combined into a combined representation for processing by the convolutional neural network.

16. The system of claim 15 , wherein the respective images are arranged as respective color channels in the combined representation.

17. The system of claim 15 , wherein intensity values ​​in each of the images are summed pixel by pixel into the pixel of the combined representation.

18. 9. The system of claim 8, wherein the portion of the three-dimensional structural rendition comprises the target amino acid and several additional amino acids adjacent to the target amino acid.

19. A non-transitory computer-readable storage medium having instructions stored thereon, The instructions, when executed by at least one processor, cause a computing device to: accessing a three-dimensional structural rendition of the amino acids of a protein; generating a plurality of two-dimensional images of a portion of said three-dimensional structural rendition comprising a target amino acid from said amino acids; and determining a pathogenicity score of a nucleotide variant that mutates the target amino acid to an alternative amino acid using a convolutional neural network that processes at least the plurality of two-dimensional images.

20. When executed by the at least one processor, the computing device activating a feature of the three-dimensional structural rendition prior to capturing the plurality of two-dimensional images; 20. The non-transitory computer-readable storage medium of claim 19, further storing instructions to: capture the plurality of two-dimensional images of the portion of the three-dimensional structural rendition using the activated feature configuration.