Mask patterns for protein language models to predict pathogenicity
Patent Information
- Application Number
- JP2024539695
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-10-27
- Filing Date
- 2022-12-23
- Publication Date
- 2026-01-06
AI Technical Summary
Existing technologies are difficult to effectively use protein sequence data to predict pathways of lesion variation, especially in the case of insufficient data, and traditional methods lack accuracy and efficiency.
Using a protein language model based on Transformer architecture, through masked language modeling, combining multi-sequence alignment (MSA) and deep neural networks, we predict the variation pathways in protein sequences, and using large-scale text data to train the model to reduce the computational amount and improve the prediction accuracy.
It realizes efficient and accurate prediction of protein variation under the condition of less clinical data, significantly improving the accuracy and efficiency of pathway prediction and reducing calculation costs.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] (Priority Application) This application claims the benefit and priority of the following: U.S. Patent Application No. 17 / 975,536, entitled "MASK PATTERN FOR PROTEIN LANGUAGE MODELS," filed on October 27, 2022 (Attorney Docket No. ILLM1063-2 / IP-2296-US1); U.S. Patent Application No. 17 / 975,547, entitled “PATHOGENICITY LANGUAGE MODEL,” filed on October 27, 2022 (Attorney Docket No. ILLM1063-3 / IP-2296-US2); U.S. Provisional Patent Application No. 63 / 294,813, entitled “PERIODIC MASK PATTERN FOR REVELATION LANGUAGE MODELS,” filed on December 29, 2021 (Attorney Docket No. ILLM1063-1 / IP-2296-PRV); U.S. Provisional Patent Application No. 63 / 294,816, entitled “CLASSIFYING MILLIONS OF VARIANTS OF UNCERTAIN SIGNIFICANCE USING PRIMATE SEQUENCING AND DEEP LEARNING,” filed on December 29, 2021 (Attorney Docket No. ILLM1064-1 / IP-2297-PRV); U.S. Provisional Patent Application No. 63 / 294,820, entitled “IDENTIFYING GENES WITH DIFFERENTIAL SELECTIVE CONSTRAINT BETWEEN HUMANS AND NONHUMAN PRIMATES,” filed on December 29, 2021 (Attorney Docket No. ILLM1065-1 / IP-2298-PRV); U.S. Provisional Patent Application No. 63 / 294,827, entitled “DEEP LEARNING NETWORK FOR EVOLUTIONARY CONSERVATION,” filed on December 29, 2021 (Attorney Docket No. ILLM1066-1 / IP-2299-PRV); U.S. Provisional Patent Application No. 63 / 294,828, entitled “INTER-MODEL PREDICTION SCORE RECALIBRATION,” filed on December 29, 2021 (Attorney Docket No. ILLM1067-1 / IP-2301-PRV), and U.S. Provisional Patent Application No. 63 / 294,830, entitled “SPECIES-DIFFERENTIABLE EVOLUTIONARY PROFILES,” filed on December 29, 2021 (Attorney Docket No. ILLM1068-1 / IP-2302-PRV).
[0002] The priority application is hereby incorporated by reference in its entirety as if fully set forth herein.
[0003] (Technical field) The disclosed technology relates to artificial intelligence based computers and digital data processing systems and corresponding data processing methods and products for mimicking intelligence (i.e., knowledge-based systems, inference systems, and knowledge acquisition systems), including systems for reasoning with uncertainty (e.g., fuzzy logic systems), adaptive systems, machine learning systems, and artificial neural networks. In particular, the disclosed technology relates to using neural networks to analyze ordered data.
[0004] (Built-in) The following are incorporated by reference for all purposes as if fully set forth herein: Sundaram,L.et al.Predicting the clinical impact of human mutation with deep neural networks.Nat.Genet.50,1161-1170(2018), Jaganathan,K.et al.Predicting splicing from primary sequence with deep learning.Cell 176,535-548(2019), Concurrently filed U.S. patent application entitled "PATHOGENICITY LANGUAGE MODEL" (Attorney Docket No. ILLM1063-3 / IP-2296-US2); U.S. Patent Application No. 62 / 573,144, entitled “TRAINING A DEEP PATHOGENICITY CLASSIFIER USING LARGE-SCALE BENIGN TRAINING DATA,” filed on October 16, 2017 (Attorney Docket No. ILLM1000-1 / IP-1611-PRV); U.S. Patent Application No. 62 / 573,149, entitled “PATHOGENICITY CLASSIFIER BASED ON DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs),” filed on October 16, 2017 (Attorney Docket No. ILLM 1000-2 / IP-1612-PRV); U.S. Patent Application No. 62 / 573,153, entitled “DEEP SEMI-SUPERVISED LEARNING THAT GENERATES LARGE-SCALE PATHOGENIC TRAINING DATA,” filed on October 16, 2017 (Attorney Docket No. ILLM1000-3 / IP-1613-PRV); U.S. Patent Application No. 62 / 582,898, entitled “PATHOGENICITY CLASSIFICATION OF GENOMIC DATA USING DEEP CONVOLUTIONAL NEURAL NETWORKS (CNNs),” filed on November 7, 2017 (Attorney Docket No. ILLM 1000-4 / IP-1618-PRV); U.S. Patent Application No. 16 / 160,903, entitled “DEEP LEARNING-BASED TECHNIQUES FOR TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on October 15, 2018 (Attorney Docket No. ILLM1000-5 / IP-1611-US); U.S. Patent Application Serial No. 16 / 160,986, entitled "DEEP CONVOLUTIONAL NEURAL NETWORKS FOR VARIANT CLASSIFICATION," filed on October 15, 2018 (Attorney Docket No. ILLM1000-6 / IP-1612-US); U.S. Patent Application Serial No. 16 / 160,968, entitled “SEMI-SUPERVISED LEARNING FOR TRAINING AN ENSEMBLE OF DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on October 15, 2018 (Attorney Docket No. ILLM1000-7 / IP-1613-US); U.S. patent application Ser. No. 16 / 160,978, entitled “DEEP LEARNING-BASED SPLICE SITE CLASSIFICATION,” filed on October 15, 2018 (Attorney Docket No. ILLM1001-4 / IP-1680-US); U.S. Patent Application No. 16 / 407,149, entitled “DEEP LEARNING-BASED TECHNIQUES FOR PRE-TRAINING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on May 8, 2019 (Attorney Docket No. ILLM1010-1 / IP-1734-US); U.S. Patent Application No. 17 / 232,056, entitled “DEEP CONVOLUTIONAL NEURAL NETWORKS TO PREDICT VARIANT PATHOGENICITY USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURES,” filed on April 15, 2021 (Attorney Docket No. ILLM 1037-2 / IP-2051-US); U.S. Patent Application No. 63 / 175,495, entitled “MULTI-CHANNEL PROTEIN VOXELIZATION TO PREDICT VARIANT PATHOGENICITY USING DEEP CONVOLUTIONAL NEURAL NETWORKS,” filed on April 15, 2021 (Attorney Docket No. ILLM 1047-1 / IP-2142-PRV); U.S. Patent Application No. 63 / 175,767, entitled “EFFICIENT VOXELIZATION FOR DEEP LEARNING,” filed on April 16, 2021 (Attorney Docket No. ILLM 1048-1 / IP-2143-PRV); U.S. Patent Application No. 17 / 468,411, entitled “ARTIFICIAL INTELLIGENCE-BASED ANALYSIS OF PROTEIN THREE-DIMENSIONAL (3D) STRUCTURES,” filed on September 7, 2021 (Attorney Docket No. ILLM 1037-3 / IP-2051A-US); U.S. Provisional Patent Application No. 63 / 253,122, entitled “PROTEIN STRUCTURE-BASED PROTEIN LANGUAGE MODELS,” filed on October 6, 2021 (Attorney Docket No. ILLM1050-1 / IP-2164-PRV); U.S. Provisional Patent Application No. 63 / 281,579, entitled “PREDICTING VARIANT PATHOGENICITY FROM EVOLUTIONARY CONSERVATION USING THREE-DIMENSIONAL (3D) PROTEIN STRUCTURE VOXELS,” filed on November 19, 2021 (Attorney Docket No. ILLM1060-1 / IP-2270-PRV), and U.S. Provisional Patent Application No. 63 / 281,592, entitled “COMBINED AND TRANSFER LEARNING OF A VARIANT PATHOGENICITY PREDICTOR USING GAPED AND NON-GAPED PROTEIN SAMPLES,” filed on November 19, 2021 (Attorney Docket No. ILLM1061-1 / IP-2271-PRV). [Background technology]
[0005] The subject matter discussed in this section should not be assumed to be prior art merely as a result of its mention in this section. Similarly, it should not be assumed that the problems mentioned in this section, or associated with the subject matter provided as background, have been previously recognized in the prior art. The subject matter in this section merely represents different approaches, which as such may also correspond to embodiments of the claimed technology.
[0006] The proliferation of available biological sequence data has led to multiple computational approaches to infer the three-dimensional structure, biological function, fitness, and evolutionary history of proteins from sequence data. So-called protein language models, such as models based on the Transformer architecture, have been trained on large ensembles of protein sequences by using masked language modeling objects that fill in masked amino acids in a sequence by considering the surrounding amino acids.
[0007] Protein language models capture long-range dependencies, learn rich representations of protein sequences, and can be employed for multiple tasks, for example, they can predict structural contacts from a single sequence in an unsupervised manner.
[0008] Protein sequences are derived from ancestral proteins and can be grouped into families of homologous proteins that share similar structures and functions. Analysis of multiple sequence alignments (MSA) of homologous proteins provides important information about functional and structural constraints. Statistics of MSA columns representing amino acid sites identify functional residues that are conserved during evolution. Correlation of amino acid usage between MSA columns contains important information about functional sectors and structural contacts.
[0009] Language models were initially developed for natural language processing and operate on a simple but powerful principle: language models gain language understanding by learning to fill in missing words in sentences, similar to sentence completion tasks in standardized tests. Language models develop powerful inference capabilities by applying this principle across large text corpora. The Bidirectional Encoder Representations from Transformers (BERT) model instantiated this principle using Transformers, a class of neural networks in which attention is a key component of the learning system. In Transformers, each token in the input sentence can "join" all other tokens by exchanging activation patterns that correspond to the intermediate outputs of neurons in the neural network.
[0010] Protein language models such as the MSA Transformer have been trained to perform inference from the MSAs of evolutionarily related sequences. The MSA Transformer interleaves attention per sequence ("row") with attention per site ("column") to incorporate epistasis. Epistasis leads to the co-evolution of specific protein positions. The effect of a mutation at one site depends on the presence or absence of mutations at other sites that affect the mutation. The combination of row attention heads in the MSA Transformer has led to state-of-the-art unsupervised structural contact prediction.
[0011] To predict the pathogenicity of missense variants from protein sequence and sequence conservation data, an end-to-end deep learning approach for variant effect prediction is applied (see Sundaram, L. et al. Predicting the clinical impact of human mutation with deep neural networks. Nat. Genet. 50, 1161-1170 (2018), referred to herein as "PrimateAI"). PrimateAI uses a deep neural network trained on variants of known pathogenicity with data augmentation using cross-species information. In particular, PrimateAI uses wild-type and mutant protein sequences to compare differences and determine the pathogenicity of the variant using a trained deep neural network. Such an approach utilizing protein sequences for pathogenicity prediction is promising because it can avoid the circularity problem and overfitting to prior knowledge. Compared to the sufficient number of data to effectively train a deep neural network, the number of clinical data available in ClinVar is relatively small. To overcome this data scarcity, PrimateAI uses common human variants and primate-derived variants as benign data and mutation-rate-matched samples of unlabeled data based on trinucleotide context as unknown data.
[0012] Opportunities arise to use protein language models and MSA for variant pathogenicity prediction. More accurate variant pathogenicity predictions may be obtained. [Brief description of the drawings]
[0013] In the drawings, like reference characters generally refer to like parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the disclosed technology. In the following description, various embodiments of the disclosed technology are described with reference to the following drawings, in which: [Figure 1]FIG. 1 is a high-level diagram showing various aspects of the disclosed technology, illustrating in particular generating a masked MSA and processing the masked MSA through the disclosed PrimateAI language model to generate phenotype predictions. [Diagram 2] 1 illustrates one embodiment of applying the disclosed periodically spaced mask grid to an MSA to generate the disclosed partially masked MSA. [Diagram 3] 1 shows one embodiment of one-hot tokens defined for a 20 residue one-hot vector, a gap residue one-hot vector, and a mask one-hot vector. [Figure 4] 1 illustrates one embodiment of channel embedding defined for a 20 residue channel embedding set, a gap channel embedding set, and a mask channel embedding set. [Diagram 5] 1 illustrates trimming, padding, and masking of an MSA according to various embodiments of the disclosed technology. [Figure 6] 1 illustrates one embodiment of generating the disclosed MSA representation. [Figure 7] 1 illustrates an example architecture of the disclosed PrimateAI language model. [Figure 8] 1 shows details of the disclosed mask representation. [Figure 9] Shows the various components of the PrimateAI language model. [Figure 10] 1 illustrates one embodiment of the disclosed expressive output head for use with the disclosed PrimateAI language model. [Figure 11] 1 is a computer-implemented method of logic flow for the PrimateAI language model according to one embodiment of the disclosed technology. [Figure 12] 1 is a system configured to implement a PrimateAI language model according to one embodiment of the disclosed technology. [Figure 13] 13 shows a performance evaluation of the language modeling portion of the disclosed PrimateAI language model with other language models. [Figure 14]Illustrates the first-place training accuracy of the disclosed PrimateAI language model. [Figure 15] A computer system that can be used for compilation and runtime execution of the disclosed PrimateAI language models. [Figure 16] Illustrates a comparison between UniRef50 HHblits MSA and human HHblits MSA. [Figure 17] We demonstrate training of a PrimateAI language model using the LAMB optimizer with gradient pre-standardization. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] The following discussion is presented to enable those skilled in the art to make and use the disclosed technology and is provided in the context of a particular application and its requirements. Various modifications to the disclosed embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other embodiments and applications without departing from the spirit and scope of the disclosed technology. Thus, the disclosed technology is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features disclosed herein.
[0015] The detailed description of the various embodiments can be better understood when read in conjunction with the accompanying drawings. To the extent that the figures show diagrams of functional blocks of the various embodiments, the functional blocks do not necessarily show a division between hardware circuitry. Thus, for example, one or more of the functional blocks (e.g., a module, a processor, or a memory) may be implemented in a single piece of hardware (e.g., a general-purpose signal processor or a block of random access memory, a hard disk, etc.) or in multiple pieces of hardware. Similarly, a program may be a stand-alone program, may be incorporated as a subroutine in an operating system, may be a function in an installed software package, etc. It should be understood that the various embodiments are not limited to the arrangements and instrumentalities shown in the figures.
[0016] The processing engines and databases in the figures designated as modules can be implemented in hardware or software and need not be divided into exactly the same blocks as shown in the figures. Some modules may be implemented on different processors, computers or servers, or even spread across multiple different processors, computers or servers.
[0017] In addition, it will be understood that some of the modules may be operated in parallel or in a different order than that shown in the figures without affecting the functionality achieved. The modules in the figures may also be considered as flowchart steps in a method. Also, modules do not necessarily have to have all code located contiguously in memory. Some portions of code may be separate from other portions of code, with code from other modules or other functions located between them.
[0018] Introduction The disclosed PrimateAI language model uses a masked language modeling objective for training on sequences: during training, residues at different positions in the sequence are replaced with mask tokens, and the PrimateAI language model is trained to predict the original residues at those positions.
[0019] Masked language modeling allows training on large amounts of unlabeled data. A fill-in-the-blank multiple sequence alignment (MSA) Transformer simultaneously classifies multiple masked locations in the MSA during training. A larger number of masked locations can add more masked language modelling (MLM) gradients that inform the optimization, thereby allowing for higher learning rates and faster training.
[0020] However, fill-in-the-blank pathogenicity prediction is fundamentally different from traditional MLM because the classification at a masked location depends on the predicted values of residues at other masked locations: the classification score can often be the average of the conditional predictions across all possible combinations of residues at other masked locations.
[0021] The PrimateAI language model avoids this averaging by exposing masked tokens at other mask locations before making predictions. The PrimateAI language model achieves state-of-the-art clinical performance and denoising accuracy while requiring 50 times less computation for training than a conventional MSA Transformer. Various aspects of the disclosed technology, discussed below, contribute to the 50 times reduction in training computation. Examples of such aspects include the periodically spaced mask grid, mask exposing, and the architecture of the PrimateAI language model.
[0022] The PrimateAI language model can be thought of as an MSA Transformer for fill-in-the-blank residue classification. In one embodiment, the PrimateAI language model is trained end-to-end on the MSA of UniRef50 proteins to minimize unsupervised MLM objects. The PrimateAI language model outputs classification scores for alternative and reference residues, which serve as input to the PrimateAI three-dimensional (3D) rank loss.
[0023] Phenotype prediction FIG. 1 is a high-level diagram 100 illustrating various aspects of the disclosed technology, in particular generating a masked MSA 140 and processing the masked MSA 140 through a disclosed PrimateAI language model (e.g., a phenotype predictor 150 or a pathogenic language model) to generate a phenotype prediction 160.
[0024] In one embodiment, the MSA dataset 110 includes a multiple sequence alignment (MSA) 120 for each sequence in the UniRef50 database, which is retrieved by searching the UniClust30 database. The MSA 120 is an alignment of multiple homologous protein sequences to a target protein. From the MSA 120, the degree of homology can be inferred and evolutionary relationships between sequences can be studied. Since real protein sequences are likely to have insertions, deletions, and substitutions, the sequences are aligned by minimizing a Levenshtein distance-like metric across all sequences. In some embodiments, a heuristic alignment scheme is used. For example, tools such as JackHMMER and HHblits can increase the number and diversity of sequences returned by performing the search and alignment steps iteratively.
[0025] Due to the difference in mutations in organisms with recent ancestors that are significantly affected by the electromechanical sensitivity of proteins to mutations, it is difficult to incorporate nearby evolution.To avoid this, the MSA used by the disclosed technology contains a variety of proteins that align with the query sequence.Using a variety of sequences from many species reduces the impact of electromechanical sensitivity on prediction, since differences are more highly determined by natural selection.
[0026] In some embodiments, the MSA dataset 110 may include 26,000,000 MSAs generated by using protein homology detection software HHblits. In other embodiments, an additional set of MSAs may be generated for 19,071 human proteins using HHblits. One of skill in the art will appreciate that the disclosed techniques can search, generate, and otherwise utilize any number of MSAs.
[0027] In some embodiments, UniRef50 MSAs in which the query sequence possesses a rare amino acid may be excluded from the MSA dataset 110, thereby retaining only MSAs in the MSA dataset 110 that contain the 20 most abundant residues. In other embodiments, only non-query sequences that contain the 20 most common residues and gaps may be included in the MSA, which in turn represents deletions to the query sequence.
[0028] In some implementations, the MSA provided as input to the PrimateAI language model may have a fixed size of 1024 sequences. Of the 1024 sequences, up to 1023 non-query sequences may be randomly sampled from the filtered sequences if the MSA depth is greater than 1024. If the MSA depth is less than 1024, the MSA may be padded with zeros to fill the input. The MSA depth refers to the number of protein sequences in the MSA. For example, an MSA transformer may be trained with a fixed input MSA depth of 1024 sequences. This makes the model easier to process because the tensors input to the model have a fixed shape. If the full MSA depth is less than 1024, padding may be added to increase its size to 1024. If the full MSA depth is greater than 1024, 1023 sequences may be randomly sampled from the full MSA depth. One query sequence can be kept such that the remaining MSA has a depth of 1024 (1023 randomly sampled sequences and one query sequence).
[0029] The masking logic 130 may apply one or more masks to the MSA 120 to generate a masked MSA 140. The masks may be arranged in a periodic, aperiodic, regular, or irregular manner. The masks are not limited to periodically spaced masks or regular grids or arrays of masks. The masks may be irregularly shaped, may be linear or curved, and may be arranged in irregular, non-uniformly spaced patterns. A mask is regular in shape if the distance between adjacent masks is fixed or the same. A mask is irregular in shape if the distance between adjacent masks varies.
[0030] Phenotype predictor 150 (e.g., PrimateAI language model) may process masked MSA 140 and generate phenotype prediction 160. In one embodiment, phenotype prediction 160 outputs the identities of masked residues in masked MSA 140. In other embodiments, phenotype prediction 160 may be used for variant pathogenicity prediction, protein contact map generation, protein functionality prediction, etc.
[0031] Please note that parts of this application refer to proteins interchangeably as "sequence," "residue sequence," "amino acid sequence," and "chain of amino acids." Also, please note that parts of this application use "amino acid" and "residue" interchangeably. Furthermore, please note that parts of this application use "periodically spaced mask set," "periodically spaced mask," "mask grid," "periodically spaced mask grid," "periodic mask pattern," and "fixed mask pattern" interchangeably.
[0032] The sequences shown in the figures are protein sequences comprising amino acid residues, in other embodiments the sequences may instead comprise DNA, RNA, carbohydrates, lipids, or any other linear or branched biopolymers.
[0033] Having described the disclosed technique at a high level using FIG. 1, a specific implementation of the disclosed periodically spaced mask grid, masking logic 130, will now be considered.
[0034] Periodically spaced mask grid FIG. 2 illustrates one embodiment of applying a disclosed periodically spaced mask grid 210 to an MSA 220 to produce a disclosed partially masked MSA 230.
[0035] The columns of the periodically spaced mask grid 210 correspond to residue positions, also referred to herein as ordinal positions. For example, in Figure 2, the periodically spaced mask grid 210 has nine columns corresponding to nine residue positions (i.e., r=9).
[0036] The periodically spaced mask grid 210 has elements (or units or tokens) that are masks. In Figure 2, such mask elements are illustrated by boxes with a filled "?" symbol. The periodically spaced mask grid 210 also has elements (or units or tokens) that are not masks. In Figure 2, such non-mask elements are illustrated by boxes filled with a diagonal line pattern.
[0037] The rows of the periodically spaced mask grid 210 include elements that are masks and elements that are not masks. The rows of the periodically spaced mask grid 210 are referred to herein as mask distributions. For example, in Figure 2, there are five mask distributions 1-5 (i.e., m mask distributions, m=5).
[0038] Each mask distribution has k periodically spaced masks. For example, in FIG. 2, mask distributions 1-4 each have three masks (i.e., k=3), and mask distribution 5 has two masks (i.e., k=2).
[0039] The k periodic interval masks in the mask distributions are at k ordinal positions starting at varying offsets from the first residue position in the periodic interval mask grid 210. For example, in FIG. 2, the k periodic interval masks of the first mask distribution are located at the third, sixth, and ninth ordinal positions and start at an offset of two from the first residue position in the periodic interval mask grid 210. The k periodic interval masks of the second mask distribution are located at the first, fourth, and seventh ordinal positions and start at an offset of zero from the first residue position in the periodic interval mask grid 210. The k periodic interval masks of the third mask distribution are located at the second, fifth, and eighth ordinal positions and start at an offset of one from the first residue position in the periodic interval mask grid 210. The k periodic interval masks of the fourth mask distribution are located at the third, sixth, and ninth ordinal positions and start at an offset of two from the first residue position in the periodic interval mask grid 210. The k periodically spaced masks of the fifth mask distribution are located at the fourth and seventh ordinal positions and begin at an offset of three from the first residue position in the periodically spaced mask grid 210 .
[0040] The masks in periodically spaced mask grid 210 are periodic because the masks have regular intervals between them and repeat at regular intervals, i.e., the masks are repetitive at regular intervals. The masks in periodically spaced mask grid 210 are also periodic because the masks have ordered patterns.
[0041] The masks in the periodically spaced mask grid 210 may have a lattice pattern, a diagonal pattern, a hexagonal pattern, a diamond pattern, a rectangular pattern, a square pattern, a triangular pattern, a convex pattern, a concave pattern, and / or a polygonal pattern.
[0042] In one embodiment, each of the k periodically spaced masks of the mask distribution in the periodically spaced mask grid 210 has the same stride (e.g., stride=3 in FIG. 2). In another embodiment, the k periodically spaced masks across the mask distribution in the periodically spaced mask grid 210 have a diagonal pattern. In other embodiments, the stride can be any number, such as 16, or any number in the range of 8 to 64, or within that range or subrange of that range. As used herein, the term "stride" refers to the distance between adjacent masks.
[0043] In other embodiments, the masks in the periodically spaced mask grid 210 are quasi-periodic, so that the masks have an ordered pattern, but the masks are not repeated at precisely regular intervals.
[0044] Details of how the mask is encoded for processing by the PrimateAI language model are now considered with reference to Figures 3 and 4. After describing Figures 3 and 4, discussion will return to Figure 2 to consider how the disclosed partially masked MSAs are generated.
[0045] mask A mask token defines a mask. A mask token is configured to mask or replace original residues in the MSA to which it is applied. A mask token is a special or auxiliary token in the sense that it is different from the 20 residue tokens used to define the 20 naturally occurring residues. A mask token is also different from the gap residue tokens used to define gap residues. A gap residue is a residue whose identity is unresolved (or unknown) and therefore cannot be reliably classified as any of the 21 known residues. A gap residue is encoded by a gap residue token.
[0046] The mask token may be defined by the same encoding logic that defines the 20 residue tokens and the gap residue token such that the mask token is encoded as the 22nd residue.
[0047] 3 shows one embodiment of a one-hot token 300 defined for 20 residue one-hot vectors 301, 302, 303, 304, 305, 306, 307, 308, 309, 310, 311, 312, 313, 314, 315, 316, 317, 318, 319, and 320, a gap residue one-hot vector 321, and a mask one-hot vector 322. The one-hot token 300 is encoded in a 22-bit binary vector in which one of the bits is hot (i.e., 1) while the others are 0. In some embodiments, a one-hot encoder (not shown) generates the one-hot token 300.
[0048] 4 illustrates one embodiment of channel embeddings 400 (or learned embeddings) defined for 20 residue channel embedding sets 401, 402, 403, 404, 405, 406, 407, 408, 409, 410, 411, 412, 413, 414, 415, 416, 417, 418, 419, and 420, a gap channel embedding set 421, and a mask channel embedding set 422. The channel embeddings 400 span the 21 known residues. The channel embedding set 421 spans the gap residues. The mask channel embedding set 422 spans the mask residues. The channel embeddings 400 are tensors with height, width, and depth dimensions, and each channel embedding set may contain N channel embeddings, where N is an integer, such as 94.
[0049] In some implementations, an embedding generator (not shown (eg, a multi-layer perceptron)) generates the channel embedding 400 .
[0050] In some implementations, the embedding generator may be trained with the PrimateAI language model to learn and generate the channel embeddings 400. During inference, a lookup table may store the mapping between the one-hot tokens 300 and the channel embeddings 400. The lookup table may be accessed during inference to replace residue tokens, gap tokens, and mask tokens with the corresponding channel embeddings.
[0051] In other embodiments, the encoding of the mask tokens (e.g., one-hot or channel embedding) may vary depending on a variety of factors, including the location of the mask (i.e., residue position), the residue type to which the mask is applied, the sequence type to which the mask is applied, the sequence number to which the mask is applied, and the species type of the sequence to which the mask is applied.
[0052] In other implementations, the mask tokens may be encoded using other schemes. Examples include quantitative or numeric data types, qualitative data types, discrete data types, continuous data types (with lower and upper bounds), integer data types (with lower and upper bounds), nominal data types, ordered or ranked data types, categorical data types, interval data types, and ratio data types. For example, the encoding may be based on continuous values such as multiple bits, real values between 0 and 1, floating point numbers, Red, Green, Blue (RGB) values between 0 and 256, hexadecimal values of CSS colors (e.g., #F0F8FF), categorical color values of CSS colors, respective values of other CSS property groups and properties, sizes of specific dimensions (e.g., height and width), sets of different values and data types, and the like, or any combination thereof.
[0053] Discussion now turns back to FIG. 2 and considers how the disclosed partially masked MSAs are generated.
[0054] Partially Masked MSA MSA 220 has p rows and r columns. The p columns correspond to p protein sequences. The r columns correspond to r residue positions (e.g., r=16 in FIG. 2). Periodically spaced mask grid 210 may have a different number of rows and columns (i.e., different shapes) than MSA 220. In some embodiments, periodically spaced mask grid 210 may have the same number of rows and columns (i.e., the same shapes) as MSA 220.
[0055] The periodically spaced mask grid 210 may be applied (or overlaid) 212 anywhere on the MSA 220. For example, the periodically spaced mask grid 210 may be applied such that the periodically spaced mask grid 210 is centered on a particular column of the MSA 220 that contains a residue of interest 214 (red) at a position of interest 216 (red). In another example, the periodically spaced mask grid 210 may be applied such that the periodically spaced mask grid 210 is positioned on a particular row of the MSA 220 (e.g., a query sequence such as sequence 1 in FIG. 2) that contains a residue of interest 214 at a position of interest 216.
[0056] In one embodiment, the periodically spaced mask grid 210 is applied to a subset of arrays in the MSA 220 that spans a window of arrays 222 (e.g., five arrays in FIG. 2). In some embodiments, the periodically spaced mask grid 210 may be applied onto the MSA 220 in a left-adjacent or right-adjacent manner. In other embodiments, the periodically spaced mask grid 210 may be applied onto the MSA 220 in a portion-by-portion manner, across portions (e.g., quadrants) of the MSA 220 simultaneously or sequentially.
[0057] Residues of MSA 220 that are overlaid with non-masked elements of periodically spaced mask grid 210 remain unchanged and are referred to herein as unmasked residues. Conversely, residues of MSA 220 that are overlaid with masked elements of periodically spaced mask grid 210 are changed to mask tokens and are referred to herein as masked residues.
[0058] The combination or aggregation of unmasked and masked residues forms a partially masked MSA 230. A partially masked MSA 230 may be defined as an MSA that includes some residues that are unmasked and some residues that are masked. A partially masked MSA 230 may also be defined as an MSA that includes some sequences that include masked residues and some sequences that do not include any masked residues.
[0059] A portion (or patch) of the partially masked MSA 230 may be trimmed (or selected or extracted) to generate trimmed portion 232 (blue dashed outline in FIG. 2). In some embodiments, trimmed portion 232 may include (i) the masked residues within a window of sequence 222, (ii) some unmasked residues contiguously adjacent to the masked residues within a neighborhood that coincides with (or defines) the boundary of trimmed portion 232, and (iii) some additional sequence portion that extends beyond the window of sequence 222 and does not include any masked residues.
[0060] MSA Trimming, Padding, and Masking Figure 5 illustrates trimming, padding, and masking of MSA 500 according to various embodiments of the disclosed technology. In Figure 5, the residue of interest at the position of interest in the query sequence is indicated by an X, the mask location is indicated by black fill, the padding is indicated by gray fill, and the trimming region is indicated by black dashed line. In these examples, the mask stride is 3 and the trimming window width is 6 residues.
[0061] In panel A, the position of interest is to the right of the center of the crop region, away from the MSA edge. In panel B, the crop region is shifted to the right of the position of interest to avoid going beyond the MSA edge. In panel C, the MSA for the short protein is padded to fill the crop region. In panel D, the crop region is shifted to the right of the position of interest to minimize padding, and the MSA is padded to fill the crop region.
[0062] In some embodiments, the positions of interest are randomly sampled from positions in the query sequence during training or selected by a user during inference. To maximize information about the positions of interest, in some embodiments, a trimming window with a size of 256 residues is selected so that the position of interest is centered. However, the trimming window can be shifted to avoid padding zeros and increase information about the position of interest if the position of interest is near the edge of the MSA. If the query sequence is shorter than the trimming window, zeros can be padded to fill the window size.
[0063] In some embodiments, if the protein length L is shorter than the query sequence, then there is a smaller probability ρ sample is assigned to the MSAs sampled during training, e.g.
[0064]
number
[0065] UniRef50 proteins used for training often have short sequences, whereas the majority of human proteins have long sequences. Figure 16 illustrates a comparison between UniRef50 HHblits MSA and human HHblits MSA. Many of the proteins in UniRef50 HHblits MSA have short sequences, whereas only a few human proteins in the MSA are short. Therefore, sampling of longer UniRef50 proteins during training can be increased so that the sampled distribution of short and long proteins is closer to the distribution of human proteins. Increasing the sampling of long sequence UniRef50 proteins also increases computational efficiency. If only short sequence UniRef50 proteins are used as inputs, the inputs will be padded to a fixed input shape, which means that computation during the training process is wasted on padding instead of adding gradients to model optimization.
[0066] The probability of sampling a non-query sequence to be included in the first sequence of the MSA can also be adjusted (e.g., f=32). In one embodiment, a periodically spaced mask grid 210 is applied to penalize the occurrence of gaps in the first sequence. The probability that a non-query sequence is masked ρ mask decreases with increasing number of gap tokens, e.g.,
[0067]
number
[0068] MSA representation FIG. 6 illustrates one embodiment of generating 600 a disclosed MSA representation. Panel A shows MSA 220. Panel B shows partially masked MSA 230. In this example, a periodically spaced mask grid 210 is applied to the first four sequences of MSA 220 and has a stride of 3. The partially masked MSA 230 is generated as a result of applying the periodically spaced mask grid 210 to MSA 220. In panel C, unmasked and masked residues in the partially masked MSA 230 are replaced with their corresponding ones from the channel embedding 400. In one embodiment, the corresponding ones from the channel embedding 400 are summed with a positional embedding for the residue string. The positional embeddings can be learned and generated during training of the PrimateAI language model. The sum of the corresponding ones from the channel embedding 400 and the positional embedding is split into chunks 640. In panel D, chunks 640 are concatenated in the channel dimension into a stack 660 and then linearly projected (670) to form an MSA representation 680. In some implementations, the linear projection 670 uses multiple one-dimensional (1D) convolution filters.
[0069] The channel embedding 400 is also referred to herein as a learned embedding. In one embodiment, the masked and unmasked residues in the partially masked MSA 230 are converted to a learned embedding by using a lookup table that stores learned embeddings corresponding to the masked and unmasked residues.
[0070] The position embedding is also referred to herein as the residue position embedding. The sum of the position embedding and the corresponding one of the channel embeddings 400 is also referred to herein as the embedded representation of the partially masked MSA 230. The learned embedding is concatenated with the residue position embedding to generate the embedded representation.
[0071] The embedded representation is chunked into a series of chunks 640. The chunks in the series are concatenated into a stack 660.
[0072] The MSA representation 680 is also referred to herein as the projected (or compressed) representation of the embedded representation. The projected representation has m rows and r columns. The stack 660 is converted to the projected representation by using a convolution operation according to one embodiment. Note that the projected representation is not compressed at this stage in the sense of making the data smaller. If the rows were not stacked, the projected representation would be "compressed" or "smaller" compared to the embedded representation, which is why row stacking reduces the computational requirements. However, the projected representation is not smaller than the model input in terms of feature dimension.
[0073] In one embodiment, a fixed mask pattern is applied to the first 32 sequences of the MSA. The MSA tokens are encoded by a learned 96-channel embedding, which is summed with a 96-channel position embedding learned for the residue columns before layer normalization. To reduce computational requirements, the embedding for the 1024 sequences in the MSA is divided into 32 chunks, each containing 32 sequences, at periodic intervals along the sequence axis. These chunks are then concatenated in the channel dimension and blended by linear projection. In the context of this application, chunks may be referred to as different non-overlapping rows of the MSA. In other embodiments, the MSA may be "chunked" in other manners, such as by column or some other irregular pattern.
[0074] PrimateAI language model 7 illustrates an example architecture 700 of the PrimateAI language model. The PrimateAI language model includes a cascade of axial attention blocks 710 (e.g., 12 axial attention blocks). The cascade of axial attention blocks 710 receives the MSA representation 680 as input and produces an updated MSA representation 720 as output. Each axial attention block includes residues that add a coupled row-wise gated self-attention layer 712, a coupled column-wise gated self-attention layer 714, and a transition layer 716.
[0075] In one embodiment, there are 12 heads in the combined row-wise gated self-attention layer 712. In one embodiment, there are 12 heads in the combined column-wise gated self-attention layer 714. Each head generates 64 channels and sums the channels across the 12 heads (768). In one embodiment, the transition layer 716 projects up to 3072 channels for GELU activation.
[0076] The present technology discloses a modified axially gated self-attention to include combined attention instead of triangular attention, which has a high computational cost. Combined attention is the sum of dot product similarities between keys and values across non-padded rows divided by the square root of the number of non-padded rows, which substantially reduces the computational load.
[0077] Next, consider the disclosed mask expressions.
[0078] Mask expression The mask expression exposes unknown values at other mask locations after a cascade of axial attention blocks 710. The mask expression collects features aligned with the mask sites. For each masked residue in a row, the mask expression exposes the embedded target tokens in other masked locations in that row.
[0079] The mask rendering combines the updated 768-channel MSA representation as updated MSA representation 720 with the 96-channel target embedded representation (token embedding) 690 at locations indicated by a Boolean mask 770 that labels the locations of the mask tokens. The Boolean mask 770, which is a fixed mask pattern with stride 16, is applied row-wise to collect features from the MSA representation and the target token embedding at the mask token locations.
[0080] Feature collection reduces the row length from 256 to 16, which dramatically reduces the computational cost of the attention block following mask rendering. For each location in each row of the collected MSA representation, the row is concatenated with the corresponding row from the collected target token embedding, and the location is also masked in the target token embedding. The MSA representation and the partially rendered target embedding are concatenated in the channel dimension and blended by linear projection.
[0081] After mask rendering 730, the now-informed MSA representation 740 is propagated through the remaining row-wise gated self-attention layers (e.g., row-wise gated self-attention layer 750 and row-wise gated self-attention layer 756) and transition layer 754. Attention is only applied to features at mask locations because the residues are known for other positions from the MSA representation 680 provided as input to the PrimateAI language model. Thus, attention only needs to be applied to mask locations where there is new information from the mask rendering. In some cases, the transition layer 754 and row-wise gated self-attention layer 756 may be repeated four times, as shown by the repeat loop 752 in FIG. 7.
[0082] After interpretation of the masked representation by self-attention, a masked collection operation 760 collects features from the resulting MSA representation at positions where the target token embedding remains masked. The collected MSA representation 772 is converted by the output head 780 into predictions 790 for 21 candidates in the amino acid and gap token vocabulary. The output head 780 includes a transition layer and a perceptron.
[0083] 8 shows details 800 of the disclosed mask representation. Mask representation allows for more information during subsequent training, improving the accuracy of predicting each residue of interest.
[0084] The first step is to collect (804, 830, 862) all tokens in the mask locations 802, 860 marked by dots. The term collect is used interchangeably with the term aggregate herein. This is done for the updated MSA representation 720, the periodically spaced mask grid 210, and the tokens in the embedded representation (embedded tokens) 690.
[0085] In Fig. 8, dashed lines and colors indicate how the MSA tiles 806 and embedding tiles 844 are selected. Feature collection reduces the row length from 256 to 16 (from 6 to 2 in Fig. 8), which dramatically reduces the computational cost of the attention block following mask representation. Each collected representation is tiled or replicated / cloned (808, 830, 866) by the number of masks in the row. In the example shown in Fig. 8, there are two masks per row. Thus, as a result of cloning 808 and 866, respectively, there are two tiles that are concatenated as clones in the cloned MSA tiles 810 and embedding tiles 870.
[0086] Mask reveal 830 is the removal of all masks in a tile except for the mask at a single location. The top tile of the collected masks is masked at a first location of interest 834 and unmasked at all other locations of interest 836. The second tile is masked at a second location of interest 838 and unmasked at all other locations of interest 832. Mask reveal reveals the other tokens in the row for each masked location in the row. In some implementations, the locations are masked in the same manner in both training and inference. This results in higher performance than just changing the location of interest to mask during inference. The location of the location of interest in the input is selected to maximize the input information, for example, when the location of interest is in the center of the mask, more adjacent columns of the MSA are included in the input processed by the PrimateAI language model.
[0087] The remaining mask after mask expression 830 is then applied 868 to the embedding tile 844 to generate a cloned masked embedding tile 870. The cloned masked embedding tile 870 is concatenated 872 with the cloned MSA tile 810 to generate a concatenated tile 873. The concatenated tile 873 is linearly projected 874 to generate the notified MSA representation 740.
[0088] PrimateAI language model components and training FIG. 9 illustrates various components 900 of the PrimateAI language model, according to one implementation. The components may include combined row-wise gated self-attention, row-wise gated self-attention, and column-wise gated self-attention. The PrimateAI language model may also use combined attention. Axial attention creates separate attention maps for each row and column of the input. Arrays in an MSA typically have similar three-dimensional structures. Direct binding analysis takes advantage of this fact to learn structural contact information. To exploit this shared structure, it is beneficial to combine row attention maps between arrays in an MSA. As an added benefit, combined attention reduces the memory footprint of row attention.
[0089] In an implementation with recomputation, combined attention reduces the memory footprint of row attention to O(ML 2 ) to O(L 2 ) where M is the number of rows, d is the hidden dimension, and Q m , K m Let be the matrix of queries and keys for the mth row of the input. The combined row attention, before softmax is applied, is defined as
[0090]
number
[0091] The final model uses square-root normalization. In other implementations, the model may use mean normalization. In such implementations, the denominator 1(M,d) is the normalization constant in standard scaled dot-product attention.
[0092]
number
[0093]
number
[0094]
number
[0095] In FIG. 9, the dimensions are sequence, s=32, residue, r=256, attention head, h=12, and channels, c=64 and c MSA = 768.
[0096] In one implementation, the PrimateAI language model may be trained on four A100 graphical processing units (GPUs). The optimizer step is for a batch size of 80 MSA and is split into four gradient aggregations to fit the batch into 40 GB of A100 memory. The PrimateAI language model is trained with the LAMB optimizer using the following parameters: β_1=0.9, β_2=0.999, ∈=10-6, and a weight decay of 0.01. The gradients are pre-normalized by division by their global L2 norm before applying the LAMB optimizer. Training is normalized by dropout with probability 0.1, which is applied after activation and before residue connections.
[0097] FIG. 17 illustrates training of the PrimateAI language model using the LAMB optimizer with gradient pre-standardization. Residue blocks are initiated as identity operations that accelerate convergence and enable the PrimateAI language model. "AdamW" refers to the ADAM optimizer with weight decay, "ReZeRO" refers to the zero redundancy optimizer, and "LR" refers to the LAMB optimizer with gradient pre-standardization. See Large Batch Optimization for Deep Learning Training BERT in 76 minutes, Yang You, Jing Li, Sashank Reddi, et al., International Conference on Learning Representations (ICLR) 2020. As illustrated, the LAMB optimizer with gradient pre-standardization shows better performance (e.g., higher accuracy rates over fewer training iterations) and is more effective for a range of learning rates compared to using the ADAMW optimizer and the zero redundancy optimizer.
[0098] Axial dropout may be applied to the self-attention block before the residual connections. Softmax post-spatial gating in column-wise attention is followed by column-wise dropout, while softmax post-spatial gating in row-wise attention is followed by row-wise dropout. Softmax post-spatial gating allows for modulation on the exponentially normalized scores or probabilities produced by softmax.
[0099] In one embodiment, the PrimateAI language model may be trained for 100,000 parameter updates. The learning rate is η = 5 × 10 over the first 5,000 steps. -6 So η = 5 × 10 -4 Then, it increases linearly to a peak value of η=10 -4The 32-bit precision linearly decays to 16-bit precision. Automatic mixed precision (AMP) can be applied to cast preferred arithmetic from 32-bit to 16-bit precision during training and inference. This increases throughput and reduces memory consumption without impacting performance. In addition, the zero-redundancy optimizer reduced memory usage by sharding the optimizer state across multiple GPUs.
[0100] Expression output head FIG. 10 shows one embodiment of an expression output head 780 that may be used by the disclosed PrimateAI language model. The collected MSA representations 772 may be converted by the output head 780 into predictions 790 of 21 candidates in the amino acid vocabulary, including gap tokens. In one embodiment, the amino acid vocabulary may be enumerated, and the amino acid enumerations used to index into a learned dictionary of embeddings. In other embodiments, one-hot embeddings of amino acids may be used and combined with linear projections. In some embodiments, the expression output head 780 may comprise a transition layer 1002, a gate 1004, a layer normalization block 1006, a linear block 1008, a GELU block 1010, and another linear block 1012. The dimension is the channel c MSA = 768, and vocabulary size, v = 21.
[0101] method FIG. 11 is a computer-implemented method 1100 of the logic flow of the PrimateAI language model according to one embodiment of the disclosed technology.
[0102] At action 1102, a multiple sequence alignment (MSA) 220 may be accessed. The MSA may have p rows and r columns. The p rows may correspond to p protein sequences. The r columns may correspond to r residue positions.
[0103] At action 1104, a periodically spaced mask grid 210 may be accessed. The periodically spaced mask grid 210 may have m mask distributions. Each of the m mask distributions may have k periodically spaced masks at k ordinal positions starting at various offsets from a first residue position in the mask grid.
[0104] In action 1106, m mask distributions may be applied to m protein sequences in the p protein sequences to generate partially masked MSA 230 that includes masked and unmasked residues, where p>m. In various embodiments, p>=m.
[0105] In action 1108, the masked and unmasked residues may be converted into a channel embedding 400 (or learned embedding), which may be concatenated with the residue position embedding to generate an embedded representation (embedded token) 690 of the partially masked MSA 230.
[0106] At action 1110, the embedded representation (embedding tokens) 690 may be chunked (or split) into a series of chunks 640, the chunks in the series of chunks 640 may be concatenated into a stack 650, and the stack 650 may be converted into a compressed representation of the embedded representation (embedding tokens) 690 as an MSA representation 680. The compressed representation as an MSA representation 680 may have m rows and r columns.
[0107] At action 1112, axial attention (e.g., by axial attention block 710) may be applied iteratively (or sequentially) across m rows and r columns of the compressed representation, and the applied axial attention may be interleaved (using a transition layer) to generate an updated MSA representation 720 of (or from) the compressed representation as MSA representation 680. The updated MSA representation 720 may have m rows and r columns.
[0108] At action 1114, k updated representation tiles (e.g., cloned MSA tile 810) may be aggregated from updated MSA representation 720. Each of the k updated representation tiles (e.g., cloned MSA tile 810) may include updated representation features of updated MSA representation 720 corresponding to masked residues. Each of the k updated representation tiles may have m rows and k columns. A given column in the k columns of a given updated representation tile of MSA tile 806 may include a respective subset of the updated representation features. Each subset may be located at a given ordinal position in the k ordinal positions. A given ordinal position may be represented by a given column.
[0109] At action 1116, k embedding tiles 870 corresponding to the k updated representation tiles (e.g., cloned MSA tiles 810) may be aggregated from the embedded representation (embedding tokens) 690. Each of the k embedding tiles 844 may include embedded features in a first chunk of a set of chunks that is a transform of masked residues. Each of the k embedding tiles may have m rows and k columns. A given column in the k columns of a given embedding tile may include a respective subset of embedded features. Each subset may be located at a given ordinal position in the k ordinal positions. A given ordinal position may be represented by a given column.
[0110] At action 1118, k Boolean tiles (e.g., at the first point of interest 834 and the second point of interest 838) may be applied to the k embedding tiles to generate k Booleanized (partially revealed) embedding tiles. Each of the k Boolean tiles may have m rows and k columns. Each of the k Boolean tiles may cause obscuration of a corresponding one of the k columns in a corresponding one of the k embedding tiles and may cause revealing of others of the k columns in a corresponding one of the k embedding tiles. Each of the k Booleanized embedding tiles may have m rows and k columns.
[0111] At action 1120, the k Boolean (partially rendered) embedded tiles 870 may be concatenated with k updated representation tiles (e.g., cloned MSA tiles 810) to generate k concatenated tiles 873, and the k concatenated tiles 873 may be converted to k compressed tile representations (notified MSA representation 740) of the k concatenated tiles 873. Each of the k compressed tile representations may have m rows and k columns.
[0112] At action 1122, self-attention (e.g., row-wise gated self-attention layer 750, transition layer 754, and row-wise gated self-attention layer 756) may be iteratively applied to the k compressed tile representations 740 to generate interpretations of compressed tile features in the k compressed tile representations that correspond to embedding features in the k embedding tiles represented by the k Boolean tiles.
[0113] At action 1124, interpreted features corresponding to embedding features in the k embedding tiles that are occluded by the k Boolean tiles may be aggregated from the interpretation to generate an aggregated representation of the interpretation (collected MSA representation 772). The aggregated representation may have m rows and k columns.
[0114] In action 1126, the collected MSA representations 772 may be converted into masked residue identities (eg, predictions 790).
[0115] system FIG. 12 is a system 1200 configured to implement the PrimateAI language model according to one embodiment of the disclosed technology.
[0116] The memory 1202 may store a multiple sequence alignment (MSA) having a plurality of masked residues.
[0117] The chunking logic 1204 may be configured to chunk the MSA into a series of chunks.
[0118] The first attention logic 1206 may be configured to attend to the representation of the sequence of chunks and generate a first attention output.
[0119] The first aggregation logic 1208 may be configured to generate a first aggregated output that includes features in the first attention output that correspond to masked residues in the plurality of masked residues. The features, in one embodiment, include elements of the MSA, such as one-hot encodings of amino acids in the MSA.
[0120] The mask expression logic 1210 may be configured to generate a signaled output based on the first aggregated output and, for each subset, a Boolean mask that alternates between concealing a given subset of masked residues and expressing the remaining subset of masked residues.
[0121] The second attention logic 1212 may be configured to pay attention to the signaled output and generate a second attention output based on the masked residues represented by the Boolean mask.
[0122] The second aggregation logic 1214 may be configured to generate a second aggregated output that includes features in the second attention output that correspond to the masked residues that were concealed by the Boolean mask.
[0123] The output logic 1216 can be configured to generate an identification of the masked residues based on the second aggregated output.
[0124] In summary, in some embodiments, the system comprises chunking logic to chunk (or split) a multiple sequence alignment (MSA) into chunks; first attention logic to attend to a representation of the chunks and generate a first attention output; first aggregation logic to generate a first aggregated output including features in the first attention output corresponding to masked residues in the plurality of masked residues; mask expression logic to generate an informed output based on the first aggregated output and the Boolean mask; second attention logic to attend to the informed output and generate a second attention output based on the masked residues expressed by the Boolean mask; second aggregation logic to generate a second aggregated output including features in the second attention output corresponding to the masked residues hidden by the Boolean mask; and output logic to generate an identification of the masked residue based on the second aggregated output.
[0125] Objective Indicators of Inventiveness and Nonobviousness Figure 13 shows EVE (J. Frazer et al., Disease variant prediction with deep generative models of evolutionary data. Nature 599, 91-95 (2021) (Evolutionary model of Variant Effect) “EVE * The replicated VAE parts of the model (labeled with "PrimateAI LM+EVE") and their combined scores (labeled with "PrimateAI LM+EVE") *Figure 13 shows a performance evaluation 1300 of the language modeling part of the PrimateAI language model (LM) compared to a selection of competitive unsupervised methods (labeled "-only"). Performance is further compared to a selection of competitive unsupervised methods (ESMlv, SIFT, LIST-S2). Starting from the top left, in a clockwise direction, the individual panels correspond to evaluations for DDD vs. UKBB, Assays, ClinVar, ASD, CHD, DDD, and UKBB. For Assays and UKBB, summary statistics are given in terms of the absolute value of the correlation (|corr|) between the score and an empirical measure of pathogenicity, i.e., the mean phenotype (UKBB) or the assay score (Assay). For DDD, we calculate the Wilcoxon rank sum P-value for control and case distributions across all datasets. For ClinVar, we measure the AUC averaged over all genes.
[0126] Evaluation Dataset Saturation mutagenesis assay The performance of the PrimateAI language model is compared using deep mutation scanning assays for the following nine genes: amyloid-beta, YAP1, MSH2, SYUA, VKOR1, PTEN, BRCA1, TP53, and ADRB2. Several assays for genes for which prediction scores for some classifiers are not available are excluded from the evaluation analysis, including TPMT, RASH, CALM1, UBE2I, SUM01, TPK1, and MAPK1. Assays for KRAS (due to different transcript sequences), SLCO1B1 (only 137 variants), and amyloid-beta are also excluded. The performance of the PrimateAI language model is evaluated by calculating the absolute Spearman rank correlation between the model prediction score and the assay score for each assay individually, and then taking the average across all assays.
[0127] UK Biobank The UK Biobank (UKBB) dataset contains 61 phenotypes across 100 genes. Evaluating for common variants of all methods reduces the number to 41 phenotypes across 42 genes. For each gene / phenotype pair, calculate the absolute Spearman rank correlation between the predicted pathogenicity score and the quantitative phenotype score. Only gene / phenotype pairs with at least 10 variants were included in the evaluation (14 phenotypes across 16 genes). This confirmed that the evaluation was robust to this choice of threshold.
[0128] ClinVar We benchmark the performance of PrimateAI language models in classifying clinical labels of ClinVar missense variants as benign or pathogenic. Variants labeled as "benign" and "likely benign" are both considered benign, and the same for variants labeled as "pathogenic" and "likely pathogenic" (both considered pathogenic). To ensure high-quality labels, we only include ClinVar variants with a review status of 1 star or higher (including "criteria provided, single submitter", "criteria provided, multiple submitters, no discrepancies", "reviewed by expert panel", and "clinical practice guideline"). This reduced the number of variants from 36,705 to 22,165 for pathogenic and from 41,986 to 39,560 for the benign class. For each gene, we calculate the area under the receiver operating characteristic curve and then report the average AUC across all genes.
[0129] DDD / ASD / CHD de novo missense variants To evaluate the performance of the deep learning network in a clinical setting, we obtain de novo mutations from published studies on intellectual disabilities, including autism spectrum disorder (ASD) and developmental disorders (DDD). ASD included 2,127 patients with at least one de novo missense (DNM) mutation. Collectively, there are a total of 3,135 DNM mutations. This reduced after requiring that all methods have predictions for their variants to 517 patients with at least one DNM variant and a total of 558 DNM variants. For DDD, 17,952 patients had at least one de novo missense variant (26,880 variants in total), which reduced to 5,872 patients (6,398 variants) after requiring the availability of predictions for all methods. A DNM variant set from patients with congenital heart disorder (CHD) is obtained consisting of 1,839 de novo missense variants from 1,342 patients (reduced to 314 variants from 299 patients after requesting the availability of predictions for all methods). For all three datasets of de novo variants from affected patients, a shared set of DNM variants from healthy controls is used, which includes 1,823 DNM variants from 1,215 healthy controls with at least one DNM variant, collected from multiple studies. It is reduced to 250 variants (235 patients) after requesting the availability of variant prediction scores for all methods. For each disease set of DNMs, a Mann-Whitney U test is applied to evaluate how well each classifier can distinguish the DNM set of patients from the DNM set of controls.
[0130] Methods for comparison Predictions from other methods are evaluated using rank scores downloaded from the database of functional predictions dbNSFP4.2a.To avoid drastic reduction of the number of common variants, methods with incomplete score sets (methods with less than 67 of 71,000,000 possible missense variants in hg38) are removed, except for Polyphen2, due to its widespread adoption.We include the following methods (method abbreviations) for comparison: BayesDel_noAF (BayesDel), CADD_raw (CADD), DANN, DEOGEN2, LIST-S2, M-CAP, MutationTaster_converted (MutationTaster), PROVEAN_converted (PROVEAN), Polyphen2_HVAR (Polyphen2, due to better performance than Polyphen2 HDIV), PrimateAI, Revel (REVEL), SIFT_converted (SIFT), VEST4, fathmm-MKL_coding (fathmm-MKL, the best performing among fathmm models for the given benchmark).
[0131] Application of EVE to more proteins In the original publication, EVE is only applied to a small set of disease-associated genes in ClinVar. To generate the disclosed language model-based training dataset, it is essential to extend the predictions of EVE to as many proteins as possible. Due to the unavailability of the EVE source code, a similar method DeepSequence is applied, converting DeepSequence scores to EVE scores by fitting a Gaussian mixture model. The latest version of UniRef100 is used, but otherwise follows the alignment depth and sequence coverage filtering steps described in EVE. At least one prediction in 18,920 proteins and a total of 50.2M predicted variants out of 71.2M possible missense variants are achieved. To validate the disclosed replication, the replicated EVE model is evaluated using the published variants from EVE. Scores from the replicated EVE models yield comparable performance to the published EVE software for all benchmarking datasets, e.g., both methods achieve a 0.41 mean absolute correlation for Assay and a 0.22 mean absolute correlation for UKBB.
[0132] Benchmarking the PrimateAI language model against other sequence-only models for pathogenicity prediction PrimateAI language models fall into a class of methods that are only trained to model protein sequences, but perform surprisingly well as pathogenicity predictors. Although they do not achieve the best overall performance by themselves, they become important features or components in classifiers that incorporate more diverse data. Figure 13 summarizes the evaluation performance of the PrimateAI language model against other such sequence-only methods: ESMlv, EVE, LIST-S2, and SIFT for pathogenicity prediction. Our language model outperforms another language model, ESMlv, on all test datasets, except for the assay, which uses only 1 / 50 of the training time. This is particularly impressive because PrimateAI LM does not rely on any fine-tuning of the assay.
[0133] Combining PrimateAI language model with EVE The language model is trained to model the entire region of the protein. EVE trains a separate model for each human protein and all similar sequences. This, and the differences in model architecture and training algorithms, suggest that the models extract separate features from their inputs. Therefore, we expected that the scores from EVE and our language model are complementary and combining the scores could lead to improved performance. We found that simply taking the average of the pathogenicity scores already performed better than either of the two methods alone. Using more sophisticated combinations, such as ridge regression, did not lead to any further improvement. The resulting performance is shown in FIG. 13, where the combined score leads to a 6.6% (or 6.8%) performance improvement in mean correlation across assays compared to PrimateAI LM (or compared to replicated EVE), a 1.4% (or 1.7%) improvement in mean AUC in ClinVar, and an increase in P-values of 11% (29%) for DDD, 3% (26%) for ASD, and 17% (23%) for CHD.
[0134] #1 Training Accuracy FIG. 14 illustrates the first-rank training accuracy 1400 of the PrimateAI language model. An ensemble of six PrimateAI language model networks was trained with different random seeds for training data sampling and model parameter initialization. Their first-rank accuracy during training is shown in FIG. 14 for the masked locations of the query sequence and all sequences in the UniRef50 MSA. The first-rank accuracy for the query sequence is much lower than that for all sequences because the query sequence does not contain gap tokens, which are easier to predict than residues because gap tokens often form long continuous segments in the MSA. The accuracy of the PrimateAI language model for the query sequence continues to improve with training. In some implementations, convergence can be accelerated by adding auxiliary losses to each layer of the PrimateAI language model.
[0135] Entropy and Pathogenicity Score The scores of the PrimateAI language model can be tabulated for future reference, rather than rerunning the model each time the score is needed. For example, the fill-in-the-blank predictions of the PrimateAI language model can be provided for locations of interest at all sites in 19,071 human proteins, summing predictions for 2,057,437,040 variants at 108,286,160 positions. Those skilled in the art will understand that these numbers will change if, for example, a small number of human proteins not included here are included. In some embodiments, the PrimateAI language models can be ensemble to generate an average score that performs better than the individual model scores. For example, each prediction can be made by an ensemble of six models, each model contributing at least four inferences with different random seeds for sampling and ordering sequences in human MSA. The inference logits can be averaged by taking the average of the predictions grouped by random seeds, and then taking the average of the averages.
[0136] The pathogenicity prediction of a variant can be evaluated using the relative values of the logits of the reference and alternative amino acids, or by subtracting the logit value of the reference amino acid from the logit value of the alternative amino acid. The probabilities are normalized across all possible residues, ignoring gap tokens, resulting in Σ r p r = 1, and the probability of the rth residue p r is obtained from the ensemble logits. The log difference captures how unlikely the variant amino acid is compared to the reference amino acid. However, the score does not take into account the predictions of the other 18 possible amino acids, which contains information about the language model internal estimate of protein-site conservation, and the convergence of the language model. The probability p r Amino acid prediction S=-Σ r p r log(p r ) to capture variant-independent, site-dependent contributions to the pathogenicity score. Specifically, the score s for alternative residues at a given site was calculated using alt is given by the ordinary logarithmic difference of the alt and reference logits at that site minus the entropy over the amino acids at a given site, i.e., s alt =log(p alt )-log(p ref )-S.
[0137] The entropy term is small whenever the probability over all amino acids is dominated by a single term, and large whenever the model is uncertain about a residue and assigns high values to multiple residues. Physically, in this case the site is less conserved and more prone to mutation. This should lead to less pathogenic signals. The score adjustment by entropy incorporates the model's internal estimate of amino acid conservation. A given log difference between a residue and a reference will be considered more pathogenic whenever it is associated with a highly conserved site. The score adjustment additionally incorporates the lack of convergence associated with very poorly trained models.
[0138] As used herein, "logic" (e.g., masking logic) may be implemented in the form of a computer product including a non-transitory computer-readable storage medium with computer usable program code to perform the method steps described herein. "Logic" may be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the exemplary method steps. "Logic" may be implemented in the form of a means for performing one or more of the method steps described herein. The means may include (i) hardware modules, (ii) software modules running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing certain techniques described herein, and the software modules being stored in a computer-readable storage medium (or multiple such media). In one embodiment, the logic implements a data processing function. The logic may be a general-purpose, single-core or multi-core processor with a computer program that specifies the function, a digital signal processor with a computer program, configurable logic such as an FPGA with a configuration file, special-purpose circuitry such as a state machine, or any combination thereof. Also, the computer program product may embody computer program and configuration file portions of logic.
[0139] Computer Systems 15 is a computer system 1500 that may be used for compilation and runtime execution of the PrimateAI language model. Computer system 1500 includes at least one central processing unit (CPU) 1572 that communicates with a number of peripheral devices via a bus subsystem 1555. These peripheral devices may include, for example, a storage subsystem 1515 including memory devices and a file storage subsystem 1536, user interface input devices 1538, user interface output devices 1576, and a network interface subsystem 1574. The input and output devices enable user interaction with computer system 1500. Network interface subsystem 1574 provides an interface to external networks, including interfaces to corresponding interface devices in other computer systems.
[0140] In one embodiment, the phenotype predictor 150 (e.g., the PrimateAI language model) is communicatively linked to the memory subsystem 1515 and the user interface input device 1538.
[0141] User interface input devices 1538 can include pointing devices such as a keyboard, a mouse, a trackball, a touch pad, or a graphics tablet, a scanner, a touch screen integrated into a display, audio input devices such as a voice recognition system and a microphone, as well as other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and manners for inputting information into computer system 1500.
[0142] The user interface output devices 1576 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as an LED display, a Cathode Ray Tube (CRT), a Liquid Crystal Display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and manners for outputting information from the computer system 1500 to a user or to another machine or computer system.
[0143] The storage subsystem 1515 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by the processor 1578.
[0144] The processor 1578 can be a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), and / or a coarse-grained reconfigurable architecture (CGRA). The processor 1578 can be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processor 1578 include Google's Tensor Processing Unit (TPU)™, rackmount solutions such as the GX4 Rackmount Series™, GX15 Rackmount Series™, NVIDIA DGX-1™, Microsoft's Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm's Zeroth Platform™ with Snapdragon processors™, NVIDIA's Volta™, NVIDIA's DRIVE PX™, NVIDIA's JETSON TX1 / TX2 MODULE™, Intel's Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM's DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.
[0145] The memory subsystem 1522 used in the storage subsystem 1515 may include several memories including a main random access memory (RAM) 1532 for storing instructions and data during program execution, and a read only memory (ROM) 1534 in which fixed instructions are stored. The file storage subsystem 1536 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular embodiment may be stored by the file storage subsystem 1536 in the storage subsystem 1515 or in another machine accessible by the processor.
[0146] Bus subsystem 1555 provides a mechanism for allowing the various components and subsystems of computer system 1500 to communicate with each other as intended. Although bus subsystem 1555 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0147] The computer system 1500 itself can be of various types, including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a loosely distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 1500 shown in Figure 15 is intended only as a specific example for purposes of illustrating a preferred embodiment of the present invention. Many other configurations of computer system 1500 can have more or fewer components than the computer system illustrated in Figure 15.
[0148] Terms The disclosed technology can be implemented as a system, a method, or a product. One or more features of the embodiments can be combined with the base embodiment. Non-mutually exclusive embodiments are taught as combinable. One or more features of the embodiments can be combined with other embodiments. The present disclosure will periodically inform the user of these options. The omission from some embodiments of the recurring list of these options should not be interpreted as limiting the combinations taught in the preceding section. These descriptions are incorporated herein by reference in each of the following implementations.
[0149] One or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of a computer product including a non-transitory computer-readable storage medium with computer usable program code for performing the illustrated method steps. Furthermore, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of an apparatus including a memory and at least one processor coupled to the memory and operative to perform the illustrated method steps. Furthermore, in another aspect, one or more embodiments and provisions of the disclosed technology, or elements thereof, can be implemented in the form of a means for performing one or more of the method steps described herein, which means can include (i) a hardware module, (ii) a software module running on one or more hardware processors, or (iii) a combination of hardware and software modules, any of (i)-(iii) implementing the particular technology described herein, and the software module is stored in a computer-readable storage medium (or multiple such media).
[0150] The clauses described in this section can be combined as features. For the sake of brevity, combinations of features are not individually listed and are not repeated for each base set of features. The reader will understand how the features specified in the clauses described in this section can be easily combined with the sets of basic features specified as embodiments in other sections of this application. These and other features, aspects, and advantages of the disclosed technology will become apparent from the following detailed description of exemplary embodiments thereof, which should be read in conjunction with the accompanying drawings. These clauses are not meant to be mutually exclusive, exhaustive, or limiting, and the disclosed technology is not limited to these clauses, but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.
[0151] Other implementations of the provisions described in this section may include a non-transitory computer-readable storage medium storing instructions executable by a processor to perform any of the provisions described in this section. Yet another implementation of the provisions described in this section may include a system including a memory and one or more processors operable to execute instructions stored in the memory to perform any of the provisions described in this section.
[0152] The present inventors disclose the following items:
[0153] Clause Set 1 Clause 1. A computer-implemented method of variant pathogenicity prediction, comprising: accessing a multiple sequence alignment that aligns the query residue sequence to a plurality of non-query residue sequences; applying a set of periodically spaced masks to a first set of residues at a first set of positions in the multiple sequence alignment, the first set of residues including residues of interest at positions of interest in the query residue sequence; Trimming a portion of the multiple sequence alignment, the portion of the multiple sequence alignment being: (i) a set of periodically spaced masks at a first set of positions; (ii) a second set of residues at a second set of positions in the multiple sequence alignment to which the set of periodic spacing masks is not applied; and generating a pathogenicity prediction for the variant at the position of interest based on a portion of the multiple sequence alignment.
[0154] Clause 2. The computer-implemented method of clause 1, wherein the multiple sequence alignment aligns a query residue sequence to a plurality of non-query residue sequences along a position-by-position dimension and along a sequence-by-sequence dimension.
[0155] Clause 3. The computer-implemented method of clause 2, wherein a set of periodically spaced masks is applied along a dimension for each sequence within a window of sequences in the multiple sequence alignment.
[0156] Clause 4. The computer-implemented method of clause 3, wherein the set of periodically spaced masks is applied along a dimension per position within a window of positions spanning the window of sequences in the multiple sequence alignment.
[0157] Clause 5. The computer-implemented method of clause 4, wherein the portion spans a window of positions across the multiple sequence alignment.
[0158] Clause 6. The computer-implemented method of clause 4, wherein the portion spans a window of positions across a subset of sequences in the multiple sequence alignment.
[0159] Clause 7. The computer-implemented method of clause 1, wherein the portion has a predetermined width and a predetermined height.
[0160] Clause 8. The computer-implemented method of clause 7, wherein the portion is padded to compensate for the multiple sequence alignment having a width less than a predetermined width of the portion.
[0161] Clause 9. The computer-implemented method of clause 7, wherein the portion is padded to compensate for the multiple sequence alignment having a height less than a predetermined height of the portion.
[0162] Clause 10. The computer-implemented method of clause 2, wherein the periodically spaced mask set is distributed along each dimension of the array into periodically spaced mask subsets.
[0163] Clause 11. The computer-implemented method of clause 10, wherein the masked subset of periodic intervals corresponds to sequences within a window of sequences.
[0164] Clause 12. The computer-implemented method of clause 11, wherein consecutive masks in the periodically spaced mask subsets corresponding to a given sequence within a window of sequences are separated by unmasked residues in the given sequence.
[0165] Clause 13. The computer-implemented method of clause 12, wherein successive masks are spaced such that the number of unmasked residues is the same across the sequence within the window of sequence.
[0166] Clause 14. The computer-implemented method of clause 12, wherein the number of unmasked residues between which successive masks are spaced varies across the sequence within a window of sequence.
[0167] Clause 15. The computer-implemented method of clause 12, wherein the starting position in a given sequence at which a masked subset of a corresponding periodic interval begins varies between sequences within a window of sequences.
[0168] Clause 16. The computer-implemented method of clause 12, wherein the starting positions follow a diagonal pattern across the array within a window of the array.
[0169] Clause 17. The computer-implemented method of clause 14, wherein the starting positions follow a diagonal pattern that begins to repeat at least once across the array within a window of the array.
[0170] Clause 18. The computer-implemented method of clause 17, wherein the starting positions follow a diagonal pattern that repeats at least once across the array within a window of the array.
[0171] Clause 19. The computer-implemented method of clause 1, wherein the periodically spaced mask set has a pattern.
[0172] Clause 20. The computer-implemented method of clause 19, wherein the pattern is a diagonal pattern.
[0173] Clause 21. The computer-implemented method of clause 19, wherein the pattern is a hexagonal pattern.
[0174] Clause 22. The computer-implemented method of clause 19, wherein the pattern is a diamond pattern.
[0175] Clause 23. The computer-implemented method of clause 19, wherein the pattern is a rectangular pattern.
[0176] Clause 24. The computer-implemented method of clause 19, wherein the pattern is a square pattern. Clause 25. The computer-implemented method of clause 19, wherein the pattern is a triangle pattern.
[0177] Clause 26. The computer-implemented method of clause 19, wherein the pattern is a convex pattern.
[0178] Clause 27. The computer-implemented method of clause 19, wherein the pattern is a concave pattern.
[0179] Clause 28. The computer-implemented method of clause 19, wherein the pattern is a polygonal pattern.
[0180] Clause 29. The computer-implemented method of clause 19, further comprising right-shifting a cropping window used for cropping to minimize padding of the portion.
[0181] Clause 30. The computer-implemented method of clause 29, further comprising left-shifting the cropping window to minimize padding of the portion.
[0182] Clause 31. The computer-implemented method of clause 1, further comprising configuring the cropping window to position the location of interest in a central column of the portion.
[0183] Clause 32. The computer-implemented method of clause 31, further comprising configuring the cropping window to position the location of interest adjacent the center column.
[0184] Clause 33. The computer-implemented method of clause 1, further comprising, in part, replacing a set of periodically spaced masks at a first set of positions with the learned mask embeddings, and in part, replacing a second set of residues at a second set of positions with the learned residue embeddings.
[0185] Clause 34. The computer-implemented method of clause 33, wherein the one-hot encoding generator generates a learned mask embedding and a learned residue embedding.
[0186] Clause 35. The computer-implemented method of clause 34, wherein the learned mask embedding and the learned residue embedding are selected from a lookup table.
[0187] Clause 36. The computer-implemented method of clause 1, further comprising, in part, replacing a set of periodically spaced masks at the first set of positions and a second set of residues at the second set of positions with the learned position embeddings.
[0188] Clause 37. The computer-implemented method of clause 36, further comprising chunking the portion into multiple chunks using the learned mask embedding, the learned residue embedding, and the learned position embedding.
[0189] Clause 38. The computer-implemented method of clause 37, further comprising processing the multiple chunks as an aggregate and generating an alternative representation of the portion.
[0190] Clause 39. The computer-implemented method of clause 38, wherein the linear projection layer processes the chunks as an aggregate using a filter bank of 1×1 convolutions to generate alternative representations of the portions.
[0191] Clause 40. The computer-implemented method of clause 39, further comprising processing the alternative representations of the portion through a cascade of attention blocks to generate updated alternative representations of the portion.
[0192] Clause 41. The computer-implemented method of clause 40, wherein an attention block in a cascade of attention blocks uses self-attention.
[0193] Clause 42. The computer-implemented method of clause 41, wherein each of the attention blocks comprises combined row-wise gated self-attention followed by column-wise gated self-attention followed by transition logic.
[0194] Clause 43. The computer-implemented method of clause 40, wherein the attention block uses cross-attention.
[0195] Clause 44. The computer-implemented method of clause 40, wherein the mask rendering block processes the updated alternative representation of the portion to generate a notified alternative representation of the portion.
[0196] Clause 45. The computer-implemented method of clause 44, wherein the mask expression block collects features aligned with the masked locations in the row and, for each mask in the row, expresses target tokens embedded in other masked locations in the row.
[0197] Clause 46. The computer-implemented method of clause 44, wherein the mask collection block processes the notified alternative representations of the portion to generate a collected alternative representation of the portion.
[0198] Clause 47. The computer-implemented method of clause 46, wherein the mask collection block processes the informed alternative representations through a cascade of transition logic and row-wise gated self-attention blocks that collect features for which the target embedding remains masked.
[0199] Clause 48. The computer-implemented method of clause 47, wherein the output block processes the collected alternative representations of the portion and predicts the identity of residues masked by a set of periodically spaced masks.
[0200] Clause 49. The computer-implemented method of clause 48, wherein the output block includes transition logic and perceptron logic.
[0201] Clause 50. The computer-implemented method of clause 48, wherein the probability of applying the periodically spaced mask subset to a non-sequence within a window of a sequence is proportional to (1-number of gap tokens in the non-sequence)^2.
[0202] Clause 51. The computer-implemented method of clause 1, further comprising generating a pathogenicity prediction for the variant based on the difference between the log probability of the variant and the log probability of the corresponding reference amino acid minus the entropy evaluated over the amino acid unit predictions.
[0203] Clause Set 2 Clause 1. A computer-implemented method comprising: accessing a multiple sequence alignment (MSA), the MSA having p rows and r columns, the p rows corresponding to the p protein sequences and the r columns corresponding to the r residue positions; accessing a mask grid having m mask distributions, each of the m mask distributions having k periodically spaced masks at k ordinal positions beginning at varying offsets from a first residue position in the mask grid; applying the m mask distributions to m protein sequences in the p protein sequences to generate a partially masked MSA including masked and unmasked residues, where p>m; converting the masked and unmasked residues into learned embeddings and concatenating the learned embeddings with residue position embeddings to generate an embedded representation of the partially masked MSA; chunking the embedded representation into a series of chunks, concatenating the chunks in the series of chunks into a stack, and converting the stack into a compressed representation of the embedded representation, the compressed representation having m rows and r columns; iteratively applying axial attention across the m rows and r columns of the compressed representation and interleaving the applied axial attention to generate an updated representation of the compressed representation, the updated representation having m rows and r columns; and aggregating k updated representation tiles from the updated representation, each of the k updated representation tiles including updated representation features of the updated representation corresponding to the masked residues, each of the k updated representation tiles having m rows and k columns; aggregating, where a given column within the k columns of a given updated representation tile includes a respective subset of the updated representation features, each subset located at a given ordinal position within the k ordinal positions, the given ordinal position being represented by a given column; and aggregating, from the embedded representation, k embedding tiles corresponding to the k updated representation tiles, where each of the k embedding tiles includes embedding features within a first chunk of a set of chunks that is a transform of masked residues, each of the k embedding tiles having m rows and k columns, where a given column within the k columns of a given embedding tile includes a respective subset of the embedding features, each subset located at a given ordinal position within the k ordinal positions, the given ordinal position being represented by a given column; applying k Boolean tiles to the k embedding tiles to generate k Booleanized embedding tiles, each of the k Boolean tiles having m rows and k columns, each of the k Boolean tiles causing obscuration of a corresponding one of the k columns in a corresponding one of the k embedding tiles and causing exposure of other of the k columns in the corresponding one of the k embedding tiles, each of the k Booleanized embedding tiles having m rows and k columns; concatenating the k boolean embedded tiles with the k updated representation tiles to generate k concatenated tiles and converting the k concatenated tiles into k compressed tile representations of the k concatenated tiles, each of the k compressed tile representations having m rows and k columns; iteratively applying self-attention to the k compressed tile representations to generate interpretations of compressed tile features in the k compressed tile representations that correspond to embedding features in the k embedding tiles that are represented by the k Boolean tiles; aggregating interpreted features from the interpretation that correspond to embedding features in the k embedding tiles that are occluded by the k Boolean tiles to generate an aggregated representation of the interpretation, the aggregated representation having m rows and k columns; and converting the aggregated representation into masked residue identities.
[0204] Clause 2. The computer-implemented method of clause 1, further comprising converting the 20 naturally occurring residues, the gap residues, and the mask into respective one-hot encoded vectors using a one-hot encoding scheme.
[0205] Clause 3. The computer-implemented method of clause 2, further comprising training a neural network to generate a respective learned embedding for each one-hot encoded vector.
[0206] Clause 4. The computer-implemented method of clause 3, wherein the masked and unmasked residues are converted to learned embeddings based on a lookup table that maps each one-hot encoded vector to a respective learned embedding.
[0207] Clause 5. The computer-implemented method of clause 4, wherein the residue position embedding specifies the order in which residues are arranged within the p protein sequence.
[0208] Clause 6. The computer-implemented method of clause 1, wherein the chunks are concatenated into stacks along a channel dimension.
[0209] Clause 7. The computer-implemented method of clause 1, wherein the stack is converted to a compressed representation by processing the stack through a linear projection.
[0210] Clause 8. The computer-implemented method of clause 7, wherein the linear projection uses a plurality of one-dimensional (1D) convolution filters.
[0211] Clause 9. The computer-implemented method of clause 8, wherein the k connected tiles are converted into k compressed tiled representations by processing the k connected tiles through a linear projection.
[0212] Clause 10. The computer-implemented method of clause 1, wherein the aggregated representation is converted to masked residue identities by processing the aggregated representation through an expression output head.
[0213] Clause 11. The computer-implemented method of clause 1, wherein p=m.
[0214] Clause 12. The computer-implemented method of clause 1, wherein each of the k Boolean tiles causes concealment of a corresponding one of the k columns in a corresponding one of the k embedding tiles and causes exposure of at least some of the others of the k columns in the corresponding one of the k embedding tiles.
[0215] Clause 13. The computer-implemented method of clause 1, wherein each of the k Boolean tiles causes concealment of a corresponding subset of the k columns in a corresponding one of the k embedding tiles and causes exposure of at least some of others of the k columns in the corresponding one of the k embedding tiles.
[0216] Clause 14. The computer-implemented method of clause 1, wherein masks for at least some of the k periodic intervals of the m mask distributions start at the same offset from the first residue position.
[0217] Clause 15. A system comprising: a memory for storing a multiple sequence alignment (MSA) having a plurality of masked residues; and chunking logic configured to chunk the MSA into a series of chunks; a first attention logic configured to attend to a representation of the sequence of chunks and generate a first attention output; first aggregation logic configured to generate a first aggregated output including features in the first attention output corresponding to masked residues in the plurality of masked residues; a mask expression logic configured to generate an informed output based on the first aggregated output and a Boolean mask that alternates, for each subset, between concealing a given subset of the masked residues and expressing a remaining subset of the masked residues; and a second attention logic configured to attend to the informed output and generate a second attention output based on the masked residues expressed by the Boolean mask. second aggregation logic configured to generate a second aggregated output including features in the second attention output that correspond to masked residues that are obscured by the Boolean mask; and and output logic configured to generate an identification of the masked residues based on the second aggregated output.
[0218] Clause 16. The system of clause 15, wherein the first attention logic uses axial attention.
[0219] Clause 17. The system of clause 15, wherein the second attention logic uses self-attention.
[0220] Clause 18. A computer-implemented method comprising: accessing a multiple sequence alignment (MSA), the MSA having p rows and r columns, the p rows corresponding to the p protein sequences and the r columns corresponding to the r residue positions; accessing a mask grid, the mask grid having m mask distributions, each of the m mask distributions having masks at k periodic intervals at k ordinal positions; applying the m mask distributions to m protein sequences in the p protein sequences to generate partially masked MSAs including masked and unmasked residues, where p>m; converting the masked and unmasked residues into learned embeddings and concatenating the learned embeddings with the residue position embeddings to generate an embedded representation of the partially masked MSA; chunking the embedded representation into a series of chunks, concatenating the chunks in the series of chunks into a stack, and converting the stack into a compressed representation of the embedded representation; iteratively applying axial attention across the m rows and r columns of the compressed representation, interleaving the applied axial attention to generate an updated representation of the compressed representation; aggregating k updated representation tiles from the updated representation, each of the k updated representation tiles including updated representation features of the updated representation corresponding to the masked residues; aggregating k embedding tiles from the embedded representation corresponding to the k updated representation tiles, each of the k embedding tiles including an embedding feature in a first chunk of the set of chunks that is a transformation of the masked residues; applying k Boolean tiles to the k embedding tiles to generate k Booleanized embedding tiles, each of the k Boolean tiles causing concealment of a corresponding one of the k columns in the corresponding one of the k embedding tiles and revealing of other of the k columns in the corresponding one of the k embedding tiles; concatenating the k boolean embedded tiles with the k updated representation tiles to generate k concatenated tiles and converting the k concatenated tiles into k compressed tile representations of the k concatenated tiles; iteratively applying self-attention to the k compressed tile representations to generate interpretations of compressed tile features in the k compressed tile representations that correspond to embedding features in the k embedding tiles that are represented by the k Boolean tiles; A computer-implemented method comprising: aggregating interpreted features from the interpretations corresponding to embedding features in the k embedding tiles that are masked by the k Boolean tiles to generate an aggregated representation of the interpretation; and converting the aggregated representation into masked residue identities.
[0221] Clause 19. The computer-implemented method of clause 18, wherein masks for at least some of the k periodic intervals of the m mask distributions start at varying offsets from a first residue position in the mask grid.
[0222] Clause 20. The computer-implemented method of clause 19, wherein masks for at least some of the k periodic intervals of the m mask distributions start at the same offset from the first residue position.
[0223] Clause 21. The computer-implemented method of clause 18, wherein the compressed representation has m rows and r columns.
[0224] Clause 22. The computer-implemented method of clause 18, wherein the updated representation has m rows and r columns.
[0225] Clause 23. A computer-implemented method as described in Clause 18, wherein each of the k updated representation tiles has m rows and k columns, and a given column within the k columns of a given updated representation tile includes a respective subset of the updated representation features, each subset being located at a given ordinal position within the k ordinal positions, and the given ordinal position being represented by the given column.
[0226] Clause 24. The computer-implemented method of clause 18, wherein each of the k embedding tiles has m rows and k columns, and a given column within the k columns of a given embedding tile includes a respective subset of embedding features, each subset being located at a given ordinal position within the k ordinal positions, the given ordinal position being represented by the given column.
[0227] Clause 25. The computer-implemented method of clause 18, wherein each of the k Boolean tiles has m rows and k columns.
[0228] Clause 26. The computer-implemented method of clause 18, wherein each of the k Boolean embedding tiles has m rows and k columns.
[0229] Clause 27. The computer-implemented method of clause 18, wherein each of the k compressed tiled representations has m rows and k columns.
[0230] Clause 28. The computer-implemented method of clause 18, wherein the aggregated representation has m rows and k columns.
[0231] Clause Set 3 Clause 1. A computer-implemented method of variant pathogenicity prediction, comprising: accessing a multiple sequence alignment that aligns the query residue sequence to a plurality of non-query residue sequences; applying a set of periodically spaced masks to a first set of residues at a first set of positions in the multiple sequence alignment, the first set of residues including residues of interest at positions of interest in the query residue sequence; Trimming a portion of the multiple sequence alignment, the portion of the multiple sequence alignment being: (i) a set of periodically spaced masks at a first set of positions; (ii) a second set of residues at a second set of positions in the multiple sequence alignment to which the set of periodic spacing masks is not applied; and generating a pathogenicity prediction for the variant at the position of interest based on a portion of the multiple sequence alignment.
[0232] Clause 2. The computer-implemented method of clause 1, wherein the multiple sequence alignment aligns a query residue sequence to multiple non-query residue sequences along a position-by-position dimension and along a sequence-by-sequence dimension.
[0233] Clause 3. The computer-implemented method of clause 2, wherein a set of periodically spaced masks is applied along a dimension for each sequence within a window of sequences in the multiple sequence alignment.
[0234] Clause 4. The computer-implemented method of clause 3, wherein the set of periodically spaced masks is applied along a dimension per position within a window of positions spanning the window of sequences in the multiple sequence alignment.
[0235] Clause 5. The computer-implemented method of clause 4, wherein the portion spans a window of positions across the multiple sequence alignment.
[0236] Clause 6. The computer-implemented method of clause 4, wherein the portion spans a window of positions across a subset of sequences in the multiple sequence alignment.
[0237] Clause 7. The computer-implemented method of clause 1, wherein the portion has a predetermined width and a predetermined height.
[0238] Clause 8. The computer-implemented method of clause 7, wherein the portion is padded to compensate for the multiple sequence alignment having a width less than a predetermined width of the portion.
[0239] Clause 9. The computer-implemented method of clause 7, wherein the portion is padded to compensate for the multiple sequence alignment having a height less than a predetermined height of the portion.
[0240] Clause 10. The computer-implemented method of clause 2, wherein the periodically spaced mask set is distributed along each dimension of the array into periodically spaced mask subsets.
[0241] Clause 11. The computer-implemented method of clause 10, wherein the masked subset of periodic intervals corresponds to sequences within a window of sequences.
[0242] Clause 12. The computer-implemented method of clause 11, wherein consecutive masks in the periodically spaced mask subsets corresponding to a given sequence within a window of sequences are separated by unmasked residues in the given sequence.
[0243] Clause 13. The computer-implemented method of clause 12, wherein successive masks are spaced such that the number of unmasked residues is the same across the sequence within the window of sequence.
[0244] Clause 14. The computer-implemented method of clause 12, wherein the number of unmasked residues between which successive masks are spaced varies across the sequence within a window of sequence.
[0245] Clause 15. The computer-implemented method of clause 12, wherein the starting position in a given sequence at which a masked subset of a corresponding periodic interval begins varies between sequences within a window of sequences.
[0246] Clause 16. The computer-implemented method of clause 12, wherein the starting positions follow a diagonal pattern across the array within a window of the array.
[0247] Clause 17. The computer-implemented method of clause 14, wherein the starting positions follow a diagonal pattern that begins to repeat at least once across the array within a window of the array.
[0248] Clause 18. The computer-implemented method of clause 17, wherein the starting positions follow a diagonal pattern that repeats at least once across the array within a window of the array.
[0249] Clause 19. The computer-implemented method of clause 1, wherein the periodically spaced mask set has a pattern.
[0250] Clause 20. The computer-implemented method of clause 19, wherein the pattern is a diagonal pattern.
[0251] Clause 21. The computer-implemented method of clause 19, wherein the pattern is a hexagonal pattern.
[0252] Clause 22. The computer-implemented method of clause 19, wherein the pattern is a diamond pattern.
[0253] Clause 23. The computer-implemented method of clause 19, wherein the pattern is a rectangular pattern.
[0254] Clause 24. The computer-implemented method of clause 19, wherein the pattern is a square pattern. Clause 25. The computer-implemented method of clause 19, wherein the pattern is a triangle pattern.
[0255] Clause 26. The computer-implemented method of clause 19, wherein the pattern is a convex pattern.
[0256] Clause 27. The computer-implemented method of clause 19, wherein the pattern is a concave pattern.
[0257] Clause 28. The computer-implemented method of clause 19, wherein the pattern is a polygonal pattern.
[0258] Clause 29. The computer-implemented method of clause 19, further comprising right-shifting a cropping window used for cropping to minimize padding of the portion.
[0259] Clause 30. The computer-implemented method of clause 29, further comprising left-shifting the cropping window to minimize padding of the portion.
[0260] Clause 31. The computer-implemented method of clause 1, further comprising configuring the cropping window to position the location of interest in a central column of the portion.
[0261] Clause 32. The computer-implemented method of clause 31, further comprising configuring the cropping window to position the location of interest adjacent the center column.
[0262] Clause 33. The computer-implemented method of clause 1, further comprising, in part, replacing a set of periodically spaced masks at a first set of positions with the learned mask embeddings, and in part, replacing a second set of residues at a second set of positions with the learned residue embeddings.
[0263] Clause 34. The computer-implemented method of clause 33, wherein the one-hot encoding generator generates a learned mask embedding and a learned residue embedding.
[0264] Clause 35. The computer-implemented method of clause 34, wherein the learned mask embedding and the learned residue embedding are selected from a lookup table.
[0265] Clause 36. The computer-implemented method of clause 1, further comprising, in part, replacing a set of periodically spaced masks at the first set of positions and a second set of residues at the second set of positions with the learned position embeddings.
[0266] Clause 37. The computer-implemented method of clause 36, further comprising chunking the portion into multiple chunks using the learned mask embedding, the learned residue embedding, and the learned position embedding.
[0267] Clause 38. The computer-implemented method of clause 37, further comprising processing the multiple chunks as an aggregate and generating an alternative representation of the portion.
[0268] Clause 39. The computer-implemented method of clause 38, wherein the linear projection layer processes the chunks as an aggregate using a filter bank of 1×1 convolutions to generate alternative representations of the portions.
[0269] Clause 40. The computer-implemented method of clause 39, further comprising processing the alternative representations of the portion through a cascade of attention blocks to generate updated alternative representations of the portion.
[0270] Clause 41. The computer-implemented method of clause 40, wherein an attention block in a cascade of attention blocks uses self-attention.
[0271] Clause 42. The computer-implemented method of clause 41, wherein each of the attention blocks comprises combined row-wise gated self-attention followed by column-wise gated self-attention followed by transition logic.
[0272] Clause 43. The computer-implemented method of clause 40, wherein the attention block uses cross-attention.
[0273] Clause 44. The computer-implemented method of clause 40, wherein the mask rendering block processes the updated alternative representation of the portion to generate a notified alternative representation of the portion.
[0274] Clause 45. The computer-implemented method of clause 44, wherein the mask expression block collects features aligned with the masked locations in the row and, for each mask in the row, expresses target tokens embedded in other masked locations in the row.
[0275] Clause 46. The computer-implemented method of clause 44, wherein the mask collection block processes the notified alternative representations of the portion to generate a collected alternative representation of the portion.
[0276] Clause 47. The computer-implemented method of clause 46, wherein the mask collection block processes the informed alternative representations through a cascade of transition logic and row-wise gated self-attention blocks that collect features for which the target embedding remains masked.
[0277] Clause 48. The computer-implemented method of clause 47, wherein the output block processes the collected alternative representations of the portion and predicts the identity of residues masked by a set of periodically spaced masks.
[0278] Clause 49. The computer-implemented method of clause 48, wherein the output block includes transition logic and perceptron logic.
[0279] Clause 50. The computer-implemented method of clause 48, wherein the probability of applying the periodically spaced mask subset to a non-sequence within a window of a sequence is proportional to (1-number of gap tokens in the non-sequence)^2.
[0280] Clause 51. The computer-implemented method of clause 1, further comprising generating a pathogenicity prediction for the variant based on the difference between the log probability of the variant and the log probability of the corresponding reference amino acid minus the entropy evaluated over the amino acid unit predictions.
[0281] Clause Set 4 Clause 1. A computer-implemented method comprising: accessing a multiple sequence alignment (MSA), the MSA having p rows and r columns, the p rows corresponding to the p protein sequences and the r columns corresponding to the r residue positions; accessing a mask grid having m mask distributions, each of the m mask distributions having k periodically spaced masks at k ordinal positions beginning at varying offsets from a first residue position in the mask grid; applying the m mask distributions to m protein sequences in the p protein sequences to generate a partially masked MSA including masked and unmasked residues, where p>m; converting the masked and unmasked residues into learned embeddings and concatenating the learned embeddings with residue position embeddings to generate an embedded representation of the partially masked MSA; chunking the embedded representation into a series of chunks, concatenating the chunks in the series of chunks into a stack, and converting the stack into a compressed representation of the embedded representation, the compressed representation having m rows and r columns; iteratively applying axial attention across the m rows and r columns of the compressed representation and interleaving the applied axial attention to generate an updated representation of the compressed representation, the updated representation having m rows and r columns; and aggregating k updated representation tiles from the updated representation, each of the k updated representation tiles including updated representation features of the updated representation corresponding to the masked residues, each of the k updated representation tiles having m rows and k columns; aggregating, where a given column within the k columns of a given updated representation tile includes a respective subset of the updated representation features, each subset located at a given ordinal position within the k ordinal positions, the given ordinal position being represented by a given column; and aggregating, from the embedded representation, k embedding tiles corresponding to the k updated representation tiles, where each of the k embedding tiles includes embedding features within a first chunk of a set of chunks that is a transform of masked residues, each of the k embedding tiles having m rows and k columns, where a given column within the k columns of a given embedding tile includes a respective subset of the embedding features, each subset located at a given ordinal position within the k ordinal positions, the given ordinal position being represented by a given column; applying k Boolean tiles to the k embedding tiles to generate k Booleanized embedding tiles, each of the k Boolean tiles having m rows and k columns, each of the k Boolean tiles causing obscuration of a corresponding one of the k columns in a corresponding one of the k embedding tiles and causing exposure of other of the k columns in the corresponding one of the k embedding tiles, each of the k Booleanized embedding tiles having m rows and k columns; concatenating the k boolean embedded tiles with the k updated representation tiles to generate k concatenated tiles and converting the k concatenated tiles into k compressed tile representations of the k concatenated tiles, each of the k compressed tile representations having m rows and k columns; iteratively applying self-attention to the k compressed tile representations to generate interpretations of compressed tile features in the k compressed tile representations that correspond to embedding features in the k embedding tiles that are represented by the k Boolean tiles; aggregating interpreted features from the interpretation that correspond to embedding features in the k embedding tiles that are occluded by the k Boolean tiles to generate an aggregated representation of the interpretation, the aggregated representation having m rows and k columns; and converting the aggregated representation into masked residue identities.
[0282] Clause 2. The computer-implemented method of clause 1, further comprising converting the 20 naturally occurring residues, the gap residues, and the mask into respective one-hot encoded vectors using a one-hot encoding scheme.
[0283] Clause 3. The computer-implemented method of clause 2, further comprising training a neural network to generate a respective learned embedding for each one-hot encoded vector.
[0284] Clause 4. The computer-implemented method of clause 3, wherein the masked and unmasked residues are converted to learned embeddings based on a lookup table that maps each one-hot encoded vector to a respective learned embedding.
[0285] Clause 5. The computer-implemented method of clause 4, wherein the residue position embedding specifies the order in which residues are arranged within the p protein sequence.
[0286] Clause 6. The computer-implemented method of clause 1, wherein the chunks are concatenated into stacks along a channel dimension.
[0287] Clause 7. The computer-implemented method of clause 1, wherein the stack is converted to a compressed representation by processing the stack through a linear projection.
[0288] Clause 8. The computer-implemented method of clause 7, wherein the linear projection uses a plurality of one-dimensional (1D) convolution filters.
[0289] Clause 9. The computer-implemented method of clause 8, wherein the k connected tiles are converted into k compressed tiled representations by processing the k connected tiles through a linear projection.
[0290] Clause 10. The computer-implemented method of clause 1, wherein the aggregated representation is converted to masked residue identities by processing the aggregated representation through an expression output head.
[0291] Clause 11. The computer-implemented method of clause 1, wherein p=m.
[0292] Clause 12. The computer-implemented method of clause 1, wherein each of the k Boolean tiles causes concealment of a corresponding one of the k columns in a corresponding one of the k embedding tiles and causes exposure of at least some of the others of the k columns in the corresponding one of the k embedding tiles.
[0293] Clause 13. The computer-implemented method of clause 1, wherein each of the k Boolean tiles causes concealment of a corresponding subset of the k columns in a corresponding one of the k embedding tiles and causes exposure of at least some of others of the k columns in the corresponding one of the k embedding tiles.
[0294] Clause 14. The computer-implemented method of clause 1, wherein masks for at least some of the k periodic intervals of the m mask distributions start at the same offset from the first residue position.
[0295] Clause 15. A system comprising: a memory for storing a multiple sequence alignment (MSA) having a plurality of masked residues; and chunking logic configured to chunk the MSA into a series of chunks; a first attention logic configured to attend to a representation of the sequence of chunks and generate a first attention output; first aggregation logic configured to generate a first aggregated output including features in the first attention output corresponding to masked residues in the plurality of masked residues; a mask expression logic configured to generate an informed output based on the first aggregated output and a Boolean mask that alternates, for each subset, between concealing a given subset of the masked residues and expressing a remaining subset of the masked residues; and a second attention logic configured to attend to the informed output and generate a second attention output based on the masked residues expressed by the Boolean mask. second aggregation logic configured to generate a second aggregated output including features in the second attention output that correspond to masked residues that are obscured by the Boolean mask; and and output logic configured to generate an identification of the masked residues based on the second aggregated output.
[0296] Clause 16. The system of clause 15, wherein the first attention logic uses axial attention.
[0297] Clause 17. The system of clause 15, wherein the second attention logic uses self-attention.
[0298] Clause 18. A computer-implemented method comprising: accessing a multiple sequence alignment (MSA), the MSA having p rows and r columns, the p rows corresponding to the p protein sequences and the r columns corresponding to the r residue positions; accessing a mask grid, the mask grid having m mask distributions, each of the m mask distributions having masks at k periodic intervals at k ordinal positions; applying the m mask distributions to m protein sequences in the p protein sequences to generate partially masked MSAs including masked and unmasked residues, where p>m; converting the masked and unmasked residues into learned embeddings and concatenating the learned embeddings with the residue position embeddings to generate an embedded representation of the partially masked MSA; chunking the embedded representation into a series of chunks, concatenating the chunks in the series of chunks into a stack, and converting the stack into a compressed representation of the embedded representation; iteratively applying axial attention across the m rows and r columns of the compressed representation, interleaving the applied axial attention to generate an updated representation of the compressed representation; aggregating k updated representation tiles from the updated representation, each of the k updated representation tiles including updated representation features of the updated representation corresponding to the masked residues; aggregating k embedding tiles from the embedded representation corresponding to the k updated representation tiles, each of the k embedding tiles including an embedding feature in a first chunk of the set of chunks that is a transformation of the masked residues; applying k Boolean tiles to the k embedding tiles to generate k Booleanized embedding tiles, each of the k Boolean tiles causing concealment of a corresponding one of the k columns in the corresponding one of the k embedding tiles and revealing of other of the k columns in the corresponding one of the k embedding tiles; concatenating the k boolean embedded tiles with the k updated representation tiles to generate k concatenated tiles and converting the k concatenated tiles into k compressed tile representations of the k concatenated tiles; iteratively applying self-attention to the k compressed tile representations to generate interpretations of compressed tile features in the k compressed tile representations that correspond to embedding features in the k embedding tiles that are represented by the k Boolean tiles; A computer-implemented method comprising: aggregating interpreted features from the interpretations corresponding to embedding features in the k embedding tiles that are masked by the k Boolean tiles to generate an aggregated representation of the interpretation; and converting the aggregated representation into masked residue identities.
[0299] Article 19. 19. The computer-implemented method of claim 18, wherein masks for at least some of the k periodic intervals of the m mask distributions start at varying offsets from a first residue position in the mask grid.
[0300] Clause 20. The computer-implemented method of clause 19, wherein masks for at least some of the k periodic intervals of the m mask distributions start at the same offset from the first residue position.
[0301] Clause 21. The computer-implemented method of clause 18, wherein the compressed representation has m rows and r columns.
[0302] Clause 22. The computer-implemented method of clause 18, wherein the updated representation has m rows and r columns.
[0303] Clause 23. A computer-implemented method as described in Clause 18, wherein each of the k updated representation tiles has m rows and k columns, and a given column within the k columns of a given updated representation tile includes a respective subset of the updated representation features, each subset being located at a given ordinal position within the k ordinal positions, and the given ordinal position being represented by the given column.
[0304] Clause 24. The computer-implemented method of clause 18, wherein each of the k embedding tiles has m rows and k columns, and a given column within the k columns of a given embedding tile includes a respective subset of embedding features, each subset being located at a given ordinal position within the k ordinal positions, the given ordinal position being represented by the given column.
[0305] Clause 25. The computer-implemented method of clause 18, wherein each of the k Boolean tiles has m rows and k columns.
[0306] Clause 26. The computer-implemented method of clause 18, wherein each of the k Boolean embedding tiles has m rows and k columns.
[0307] Clause 27. The computer-implemented method of clause 18, wherein each of the k compressed tiled representations has m rows and k columns.
[0308] Clause 28. The computer-implemented method of clause 18, wherein the aggregated representation has m rows and k columns.
Claims
1. 1. A computer-implemented method for variant pathogenicity prediction, comprising: accessing a multiple sequence alignment that aligns the query residue sequence to multiple non-query residue sequences; applying a set of periodically spaced masks to a first set of residues at a first set of positions in the multiple sequence alignment, the first set of residues including residues of interest at positions of interest in the query residue sequence; Trimming a portion of said multiple sequence alignment, said portion of said multiple sequence alignment comprising: (i) a set of masks at the periodic intervals at the first set of positions; (ii) a second set of residues at a second set of positions in the multiple sequence alignment to which the set of periodic spacing masks is not applied; generating a pathogenicity prediction for the variant at the position of interest based on the portion of the multiple sequence alignment.
2. 2. The computer-implemented method of claim 1, wherein the multiple sequence alignment aligns the query residue sequence to the plurality of non-query residue sequences along a position-by-position dimension and along a sequence-by-sequence dimension.
3. The computer-implemented method of claim 2 , wherein the set of periodically spaced masks is applied along a dimension for each sequence within a window of sequences in the multiple sequence alignment.
4. 4. The computer-implemented method of claim 3, wherein the set of periodically spaced masks is applied along a dimension per position within a window of positions spanning a window of the sequences in the multiple sequence alignment.
5. The computer-implemented method of any one of claims 1 to 4, wherein the portion has a predetermined width and a predetermined height.
6. 6. The computer-implemented method of claim 5, wherein the portion is padded to compensate for multiple sequence alignments having a width less than the predetermined width of the portion.
7. The computer-implemented method of claim 2 , wherein the periodically spaced mask set is distributed along each dimension of the array into periodically spaced mask subsets.
8. The computer-implemented method of claim 7 , wherein the masked subset of periodic intervals corresponds to sequences within a window of sequences.
9. The computer-implemented method of claim 1 , wherein the periodically spaced mask set has a pattern.
10. The computer-implemented method of claim 9 , further comprising right-shifting a cropping window used for said cropping to minimize padding of said portion.
11. The computer-implemented method of claim 10 , further comprising left-shifting the cropping window to minimize the padding of the portion.
12. The computer-implemented method of claim 1 , further comprising configuring a cropping window to position the location of interest in a center column of the portion.
13. The computer-implemented method of claim 12 , further comprising configuring the cropping window to position the location of interest adjacent to the center column.
14. 2. The computer-implemented method of claim 1, further comprising: replacing, in the portion, the periodically spaced mask set at the first set of positions with a learned mask embedding; and replacing, in the portion, the second residue set at the second set of positions with a learned residue embedding.
15. 15. The computer-implemented method of claim 14, further comprising replacing, in the portion, the set of periodically spaced masks at the first set of positions and the second set of residues at the second set of positions with learned positional embeddings.
16. 16. The computer-implemented method of claim 15, further comprising chunking the portion into multiple chunks using a learned mask embedding, the learned residue embedding, and the learned position embedding.
17. The computer-implemented method of claim 16 , further comprising processing the chunks as an aggregate and generating an alternative representation of the portion.
18. 2. The computer-implemented method of claim 1, further comprising generating the pathogenicity prediction for the variant based on the difference between the log-probability of the variant and the log-probability of the corresponding reference amino acid minus entropy evaluated over amino acid-by-amino acid predictions.
19. 1. A system comprising one or more processors coupled to a memory, the memory loaded with computer instructions for predicting variant pathogenicity, the computer instructions, when executed on the one or more processors, accessing a multiple sequence alignment that aligns the query residue sequence to multiple non-query residue sequences; applying a set of periodically spaced masks to a first set of residues at a first set of positions in the multiple sequence alignment, the first set of residues including residues of interest at positions of interest in the query residue sequence; Trimming a portion of said multiple sequence alignment, said portion of said multiple sequence alignment comprising: (i) a set of masks at the periodic intervals at the first set of positions; (ii) a second set of residues at a second set of positions in the multiple sequence alignment to which the set of periodic spacing masks is not applied; generating a pathogenicity prediction for the variant at the position of interest based on the portion of the multiple sequence alignment.
20. 1. A non-transitory computer-readable storage medium having recorded thereon computer program instructions for predicting variant pathogenicity, the computer program instructions, when executed on a processor, comprising: accessing a multiple sequence alignment that aligns the query residue sequence to multiple non-query residue sequences; applying a set of periodically spaced masks to a first set of residues at a first set of positions in the multiple sequence alignment, the first set of residues including residues of interest at positions of interest in the query residue sequence; Trimming a portion of said multiple sequence alignment, said portion of said multiple sequence alignment comprising: (i) a set of masks at the periodic intervals at the first set of positions; (ii) a second set of residues at a second set of positions in the multiple sequence alignment to which the set of periodic spacing masks is not applied; generating a pathogenicity prediction for the variant at the position of interest based on the portion of the multiple sequence alignment.