Methods and systems for protein modeling and prediction
A computational method using amino acid residue abundance-based fingerprints and machine learning classifiers enhances protein folding and interaction predictions, addressing accuracy and scalability issues, facilitating drug and antibody development.
Patent Information
- Application Number
- PCT/US2025/032224
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-06
- Filing Date
- 2025-06-04
- Publication Date
- 2025-12-11
AI Technical Summary
Existing methodologies for predicting protein folding and protein-protein interactions are limited in accuracy, speed, and scalability, necessitating improved methods and systems for protein sequence-based predictions.
A computational approach involving the generation of protein fingerprints based on amino acid residue abundance, using machine learning classifiers to predict protein folding and interactions, without requiring chemical information about the sequence or residues, and employing synthetic templates to enhance structural diversity in protein structure prediction.
The method achieves high accuracy in protein folding and interaction predictions, with AUC values of at least 0.7, enabling applications in drug development, antibody engineering, and protein engineering for therapeutic and diagnostic purposes.
Smart Images

Figure US2025032224_11122025_PF_FP_ABST
Abstract
Description
METHODS AND SYSTEMS FOR PROTEIN MODELING AND PREDICTIONCROSS-REFERENCE
[0001] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 656,779 filed on June 6, 2024, the entire contents of which are incorporated herein by reference.BACKGROUND
[0002] The structure and interactions of proteins are fundamental to the understanding of various biological functions. For example, protein folding and protein foldability can provide important insights into therapeutic design for biologies, understanding of cancer biology, etc. Identification of active protein folding prediction and protein-protein interaction prediction have, therefore, remained crucial areas of study in computational biology. Existing methodologies are, however, often limited in accuracy, speed, and scalability. Therefore, there remains a need for improved methods and systems to predict protein folding and interactions based on a given protein sequence.SUMMARY
[0003] In some aspects, described herein are computational methods for proteins. In some embodiments, the methods comprise obtaining a sequence of a first protein. In some embodiments, the methods comprise computing an abundance of a plurality of amino acid residues within the sequence to generate a first fingerprint for the first protein, wherein the first fingerprint comprises a plurality of vectors that are ranked based at least in part on the abundance of the plurality of amino acid residues. In some embodiments, the methods comprise providing the first fingerprint for use in modeling, classifying or predicting at least one of protein folding ability or protein-protein interactions. In some embodiments, the first fingerprint comprises a fingerprint matrix, a fingerprint vector, or a fingerprint tensor. In some embodiments, the first fingerprint is a fingerprint matrix.
[0004] In some embodiments, the plurality of vectors are ranked solely on the abundance of the plurality of amino acid residues, without using or requiring chemical information about either the sequence or the amino acid residues within the sequence. In some embodiments, providing the first fingerprint comprises computing an interaction fingerprint using the first fingerprint and a second fingerprint for a second protein.
[0005] In some embodiments, the second fingerprint comprises a plurality of vectors that are ranked based at least in part on an abundance of a plurality of amino acid residues within a sequence of the second protein. In some embodiments, the method comprises: providing theinteraction fingerprint to a trained machine learning classifier and using the trained machine learning classifier to generate a classification indicative of an interaction or lack thereof between the first protein and the second protein.
[0006] In some embodiments, an accuracy of the classification has an area under the curve (AUC) value of at least 0.7 (e.g. at least 0.9, at least 0.96, or at least 0.97). In some embodiments, the plurality of vectors in the first fingerprint are associated with a plurality of window segments defined over a length of the sequence of the first protein. In some embodiments, the abundance is computed within each window segment of the plurality of window segments.
[0007] In some embodiments, the abundance is computed across the plurality of window segments. In some embodiments, a length of the plurality of window segments is about 1% to about 100% (e.g. 1% to 99%, 20% to 80%, 5% to 50%, or 1% to 20%) of the length of the sequence of the first protein. In some embodiments, a length of each of the plurality of window segments is about 3 to about 100 amino acid residues in length. In some embodiments, a length of each of the plurality of window segments is about 100 to about 1000 amino acid residues in length. In some embodiments, a length of each of the plurality of window segments is at least 100 amino acid residues (e.g., at least 150 amino acid residues, at least 200 amino acid residues, at least 300 amino acid residues, or at least 500 amino acid residues) in length. In some embodiments, a length of the sequence analyzed is about 20 amino acid residues to about 2000 amino acid residues in length. In some embodiments, a length of the sequence analyzed is about 50 amino acid residues to about 1000 amino acid residues in length. In some embodiments, a length of the sequence analyzed is about 100 amino acid residues to about 2000 amino acid residues in length.
[0008] In some embodiments, each of the plurality of window segments overlaps an alternate one of the plurality of window segments by at least one amino acid residue. In some embodiments, each of the plurality of window segments do not overlap. In some embodiments, the plurality of window segments comprises both overlapping and non-overlapping segments. In some embodiments, one or more of the plurality of windows is compared to the complete sequence.
[0009] In some embodiments, plurality of vectors in the first fingerprint are based at least in part on the abundance of each of the twenty naturally occurring amino acid residues. In some embodiments, plurality of vectors in the first fingerprint are based at least in part on an abundance of one or more of a plurality of codons encoding for a protein sequence. In some embodiments, the first fingerprint comprises abundance of each naturally occurring amino acid residue in the first protein sequence and an abundance of each naturally occurring amino acid ineach of the plurality of window segments. In some embodiments, the plurality of vectors are ranked by abundance in ascending order. In some embodiments, the plurality of vectors are ranked by abundance in descending order. In some embodiments, the plurality of vectors are ranked by abundance according to a standard (ordinal) ranking procedure.
[0010] In some embodiments, the plurality of vectors are ranked by abundance according to a fractional ranking procedure. In some embodiments, the plurality of vectors are ranked by abundance according to a dense ranking procedure. In some embodiments, the plurality of vectors are ranked by abundance according to a competition ranking procedure. In some embodiments, the plurality of vectors are ranked by abundance according to a modified competition ranking procedure. In some embodiments, the plurality of vectors are ranked by abundance according to a random ranking procedure. In some embodiments, the plurality of vectors are ranked by abundance according to a ranking procedure comprising one or more secondary ranking criteria. In some embodiments, one or more of the plurality of vectors is subtracted from a total protein sequence.
[0011] In some embodiments, the predicting is used to predict protein folding ability. In some embodiments, the predicting is used to predict protein-protein interaction. In some embodiments, the method comprises predicting, using a machine learning model, a novel protein sequence capable of interacting with the first protein sequence. In some embodiments, the novel protein sequence interacts with a defined target segment of the first protein sequence.
[0012] In some embodiments, the defined target segment is input by a user. In some embodiments, the defined target segment is identified using the machine learning model. In some embodiments, interaction with the defined target segment is expected to agonize and / or antagonize a protein having the first protein sequence upon formation of a protein-protein complex. In some embodiments, the novel protein sequence is a sequence of an engineered antibody.
[0013] In some embodiments, the antibody has a therapeutic effect when administered to a subject suffering from a condition. In some embodiments, the novel protein sequence comprises an improved variant of an initial sequence provided to the system by a user. In some embodiments, the method comprises iteratively generating by the machine learning model, a plurality of improved sequences until a threshold interaction value is achieved.
[0014] In some embodiments, the threshold interaction value is a threshold agonist effect on a protein having the first protein sequence. In some embodiments, the threshold interaction value is a threshold antagonist effect on a protein having the first protein sequence. In some embodiments, the method is used to predict or classify protein-protein interaction of a proteinprotein complex comprising at least three proteins. In some embodiments, the first proteinsequence is an engineered protein sequence, and prediction or classification of foldability is used to screen a plurality of engineered protein sequence candidates.
[0015] In some embodiments, the engineered protein sequence is a sequence of an enzyme. In some embodiments, a plurality of enzyme sequence candidates are screened for enzymatic activity and / or enzyme stability. In some embodiments, the method comprises automatically identifying, by a machine learning model, an enzyme sequence having an optimized enzymatic activity and / or an optimized enzyme stability.
[0016] In some embodiments, the optimized enzyme sequence is optimized for stability in conditions reflective of a large-scale biomanufacturing reactor. In some embodiments, the optimized enzyme sequence is optimized for enzymatic activity in conditions reflective of a large-scale biomanufacturing reactor. In some embodiments, the optimized enzyme sequence is optimized for both enzymatic activity and stability in a biomanufacturing reactor.
[0017] In some embodiments, the engineered protein sequence is an antibody sequence. In some embodiments, the antibody sequence is optimized for binding of one or more target sited. In some embodiments, the optimized enzyme antibody sequence is optimized for both binding of a first target and non-binding of a second target.
[0018] In some cases, the classification is used in conjunction with or as part of a pharmaceutical and / or medical application. In some embodiments, the classification is used as part of a method comprising drug development (e.g., for designing proteins that serve as drugs or therapeutic agents to target specific diseases). In some embodiments, the classification is used as part of a method comprising antibody development (e.g., by engineering monoclonal antibodies for targeted therapy in diseases like cancer). In some embodiments, the classification is used as part of a method comprising antibody selection (e.g., by selecting antibodies based on affinity and / or neutralization). In some embodiments, the classification is used as part of a method comprising vaccine development (e.g., developing protein-based vaccines, for example, subunit vaccines). In some embodiments, the classification is used as part of a method comprising development of an enzyme replacement therapy (e.g, by creating enzymes to replace deficient or dysfunctional ones in patients.) In some embodiments, the classification is used as part of a diagnostic method (e.g., by designing proteins for use in diagnostic assays, biosensors, and imaging technologies). In some embodiments, the classification is used as part of a method comprising development of a gene therapy method (e.g., by engineering proteins to deliver therapeutic genes or modify gene expression). In some embodiments, the classification is used as part of a method comprising cell therapy development (e.g., by designing proteins to enhance the efficacy and specificity of cell-based therapies). In some embodiments, the classification is used as part of a method comprising development of a biosimilar (e.g., by developing protein-based drugs that are highly similar to existing biologic therapies for competitive markets). In some embodiments, the classification is used as part of a method comprising development of a hormone replacement (e.g., by engineering hormones for therapeutic use, such as insulin analogs). In some embodiments, the classification is used as part of a method comprising development of a personalized medical treatment (e.g., by prediction of allosteric regions specific to proteins from a particular subject or treatment candidate).
[0019] In some embodiments, the classification is used as part of a research and / or development method. In some embodiments, the classification is used as part of a method comprising development of structural biology (e.g., by engineering proteins to understand their structurefunction relationships). In some embodiments, the classification is used as part of a method comprising synthetic biology development (e.g., by creating synthetic proteins to explore new biological functions and pathways). In some embodiments, the classification is used as part of a method comprising development of a protein library (e.g., by developing libraries of engineered proteins for high-throughput screening in research). In some embodiments, the classification is used as part of a method comprising designing proteins to study and manipulate protein-protein interactions. In some embodiments, the classification is used as part of a method comprising optogenetic development (e.g, by engineering light-responsive proteins for controlling cellular processes with light).
[0020] In some embodiments, an accuracy of the classification has an area under the curve (AUC) value of at least 0.7, 0.8, or 0.9, determined using a receiver operating characteristic curve. In some embodiments, an identity of one or more proteins is used solely as an input to any of the methods described herein, wherein classifications or predictions are made without using or requiring chemical information about either the sequence or the amino acid residues within the sequence.
[0021] In another aspect, described herein are non-transitory computer readable media storing instructions that, when executed by a processor of a computer system, cause the computer system to perform any of the methods described herein. In another aspect, described herein are computer systems comprising: a processor, and a non-transitory computer readable storage medium, storing instructions that, when executed by a processor of a computer system, cause the computer system to perform any of the methods described herein.
[0022] Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
[0023] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure.Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0024] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The novel features of the present disclosure are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present disclosure will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the present disclosure are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0026] FIG. 1 illustrates an example workflow for classifying or predicting protein foldability and / or protein-protein interactions according to embodiments described herein.
[0027] FIG. 2 illustrates a further example workflow for classifying or predicting protein foldability and / or protein-protein interactions according to embodiments described herein.
[0028] FIG. 3 illustrates another example workflow for classifying or predicting protein foldability and / or protein-protein interactions according to embodiments described herein.
[0029] FIG. 4 illustrates an example workflow for classifying whether two arbitrary protein sequences will interact according to embodiments described herein.
[0030] FIG. 5 illustrates an example of a receiver operating curve calculated for analysis of protein-protein interaction prediction using a method implemented according to embodiments described herein.
[0031] FIG. 6 shows a computer system that is programmed or otherwise configured to implement methods provided herein.DETAILED DESCRIPTION
[0032] While various embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the present disclosure. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed.
[0033] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0034] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0035] Certain inventive embodiments herein contemplate numerical ranges. When ranges are present, the ranges include the range endpoints. Additionally, every sub range and value within the range is present as if explicitly written out. The term “about” or “approximately” may mean within an acceptable error range for the particular value, which will depend in part on how the value is measured or determined, e.g., the limitations of the measurement system. For example, “about” may mean within 1 or more than 1 standard deviation, per the practice in the art. Alternatively, “about” may mean a range of up to 20%, up to 10%, up to 5%, or up to 1% of a given value. Where particular values are described in the application and claims, unless otherwise stated the term “about” meaning within an acceptable error range for the particular value may be assumed.
[0036] As used herein, “fingerprint” is used to describe a unique identifier or series of identifiers (for example, a vector containing various information about a sequence or about interaction between a first sequence and a second sequence) can be computed for a particular property of interest.
[0037] As used herein, the term “3D structure” refers to the three-dimensional shape of a protein molecule as defined by the geometric arrangement its atoms, which is also commonly referred as the molecule's “tertiary structure.”
[0038] As used herein, the term “root mean squared deviation” or “RMSD” refers to the average distance measured in Angstroms (A) between atoms in superimposed protein structures, where a lower value indicates a better structural match.
[0039] As used herein, the term “high resolution model” refers to a predicted 3D structure indistinguishable from a known experimental structure of that protein, which is defined as an RMSD less than 2 A.
[0040] As used herein, the term “conformational space” refers to the set of all chemically- plausible structural arrangements for a protein.
[0041] As used herein, the term “synthetic template” refers to an unnatural structure template that represents an alternate chemically-plausible conformation for a protein that has not been observed in known experimental structures.
[0042] As used herein, the term “accelerated conformational sampling” refers to the process of using synthetic templates to overcome energetic and kinetic barriers to discover native, low energy conformations that were otherwise undiscoverable by conventional structure prediction techniques.
[0043] Provided herein are methods and systems implementing methods, which allow for efficient and accurate classification or prediction of: protein-protein interaction and / or a lack thereof between a plurality of protein sequences, and / or of protein foldability, based at least in part on their sequence identities. An example workflow of one such method is illustrated in FIG. 1. In the example workflow, a protein sequence about which a user seeks to make a prediction or classification is obtained 101. A computation is then performed to generate a “fingerprint” of the protein based at least in part on obtained sequence 102. The computed fingerprint can then be used to classify or predict protein folding and / or to predict interaction with a second protein 103, for example, using a trained machine learning classifier. In some cases, the computed fingerprint is a matrix. In some cases the computed fingerprint is a tensor. In some embodiments, the fingerprint is a vector. In some cases, the fingerprint comprises a combination of: vector(s), tensor(s), a matrix, and / or a plurality of matrices.
[0044] A guiding principle of structural biology can be that a protein's amino acid sequence defines its tertiary (3D) structure, which can be the shape of a protein molecule as defined by the geometric arrangement of its atoms. Over the past half century, a revolution in biology and computer science has revealed the intimate relationship between a protein's biological function and its 3D structure and the feasibility of predicting a structure from sequence alone. Accordingly, there can be a pressing need for improved basic research and diagnostic tools and processes to improve protein engineering and drug development. These activities require high- resolution protein structures. Most protein structures may be determined by X-raycrystallography: a slow, expensive, and unpredictable process. Much progress has been made in high-throughput structure determination, but a library of high-resolution structures for representative model organisms, including the human proteome, can be beyond the reach of current experimental methods.
[0045] New protein sequences may be discovered at a rate far exceeding that of protein structures. Considering that it can take months or years to solve a protein structure and the minimum cost per structure (for a structural genomics facility) can be high, protein structure prediction can be an important tool to close the sequence-structure gap. Current structure prediction techniques may be classified as either ab initio or template-based.
[0046] Ab initio prediction methods generally rely on an energetic force field (a mathematical model that estimates how biophysical forces affect thermodynamic stability) to fold protein sequences from basic principles. Due to efficiency concerns, ab initio prediction can be only applicable to short protein sequences (less than 150 residues) and often require weeks of central processing unit (CPU) time. Most limiting can be that current ab initio predictions often regarded as low-resolution models, which may be insufficient for use in drug screening, drug design, ligand docking, or protein engineering by directed mutagenesis.
[0047] Template-based prediction methods, which use known experimental protein structures to guide the prediction process, may be capable of producing “high resolution models” (indistinguishable from an experimental structure) and may be used for industrial and pharmaceutical bio-simulation needs. Template-based methods can be further divided into homology modeling processes, which select templates based on sequence similarity to the query sequence, and protein threading processes, which select templates based on sequence-implied structural similarities with the query. The fundamental limitation of these methods can be that they require a known experimental structure that can be sufficiently similar to the target query.
[0048] The quality of a template-based prediction can be reliant on how similar the conformations of the selected templates may be to the native fold of the target protein. Templates may be collected from protein structure databases and consist of entire or partial protein structures. The protein modeling community uses a variety of private and public structure databases combined with individualized sequence alignment, fold recognition, and domain search techniques to search for the best templates. The query sequence can be then mapped onto the selected templates to derive the final target protein model. Current approaches treat templates as static, rigid conformations; however, proteins may be naturally flexible and plastic. As such, statistically significant templates may not adequately approximate the overall range of the target protein conformation, which leads template-based processes to converge toward suboptimal, non-native protein folds.
[0049] In some embodiments methods described herein can be directed to discovering the native 3D structure for a given protein by using synthetic templates to represent the range of putative structural conformations not captured in known experimental structures. A synthetic template can be derived from a template identified from conventional template-based prediction methods by perturbing it into alternate chemically-plausible conformations. The use of synthetic templates increases the sampled structural diversity of the prediction process well beyond that achieved by conventional ab initio and template-based techniques, with the effect of increasing the discovery rate of high resolution models with low energy conformations.
[0050] In one embodiment, a coarse grain mathematical model based on a mass-spring representation of a protein structure can be used to calculate the resonant frequencies and associated modes of motion for a template, which may be then used to extrapolate numerous candidate synthetic templates. This population can be reduced to a single representative conformation using an energetic scoring function to produce one synthetic template. The process can be applied to templates whose conformations may be not already energetically optimal.
[0051] In yet another embodiment, the synthetic template process can be applied to proteinprotein complexes, where one or more proteins may be bound to each other, such as antibodies comprised of heavy and light protein chains. Proteins in a complex exhibit a different mobility when compared to the individual proteins in isolation. As such, creating synthetic templates for a protein complex template ensures the process focuses on sampling chemically-plausible conformations that can be achieved by the complex as a whole.
[0052] In yet another embodiment, the synthetic template process can be applied to predict conformational changes induced by site-directed mutations or genetic variations. Predictions may be not limited to the immediate residues in contact with the mutation site. Affected residues can be several Angstroms away from the mutation or variant site. This analysis can accelerate protein engineering studies by predicting whether putative mutations may cause disruptive rearrangements to the protein's structure. In some cases, processes described herein can lead to improvements human health by identifying structural mechanisms for how somatic mutations destabilize protein folds leading to destructive changes to cellular regulation such as those found in many cancers. These new mechanisms and high-resolution structure predictions may be invaluable to developing the next generation of structure-based drugs and biotherapeutics.
[0053] The present disclosure comprises an accelerated conformational sampling process to construct high-resolution models for a target protein sequence, where synthetic templates may be used to construct new, low energy conformations that were otherwise undiscoverable by conventional structure prediction techniques. The accelerated conformational sampling process comprises creating synthetic templates using normal mode analysis (NMA), a structural mass-spring model that describes protein motion at resonant frequencies, to build biologically relevant conformations from a given structure template. The process can generate a synthetic template within minutes; as such, it can be applied to dozens of templates associated with a structure prediction. A structure prediction can be said to undergo accelerated conformational sampling when synthetic templates may be used in the prediction's downstream conformational search and refinement steps.
[0054] An example structure prediction process begins by performing two sequence-based searches for each query protein sequence. The first search, a sequence similarity search against a non-redundant database of known protein sequences, identifies similar sequences and constructs a sequence profile matrix of amino acid frequencies observed along the sequence. The sequence profile matrix can be used by two separate machine learning models that predict the secondary structure classifications and the relative solvent accessibility for each residue. The second search can be a local pairwise sequence alignment against a non-redundant structural feature database, where the identified homologous proteins may be used to predict internal contacts for each residue. The information from both searches, as well as any additional structural classification or solvent accessibility information, can be consolidated into a collected feature matrix. Next, one or more different protein threading models align the query's collective features against the non- redundant structure database to select statistically significant original templates. The number of different protein threading models necessary to generate a significant number of original templates may be between one and forty, preferably between six and twelve and most preferably eight. Nearly 100 original templates may be typically collected.
[0055] Normal mode analysis (NMA), which can be based on a harmonic approximation to describe the fluctuation of a potential energy function around a structure template can be used to calculate the normal modes of motion for each template selected from the threading process. First a potential energy can be established where contacting residues may be identified using a cutoff distance. The six lowest frequency modes representing trivial translational and rotational degrees of freedom may be ignored. The remaining normal modes describe collective motions throughout the protein structure, where each residue can be assigned an independent vector describing its velocity through Cartesian space.
[0056] Next, new conformations may be constructed by perturbing the template along linear combinations of the ten lowest frequency normal modes. A discrete set of weights (ranging from -80 to 80, incremented by values of 20) may be selected to limit the number of new conformations to approximately 3000 per template-amounting to over 300,000 perturbed conformations for each structure prediction. Each conformation can be evaluated with an energetic scoring function, whose most important factors may be short-range residue contacts,hydrogen bonding, and secondary structure propensities, and then compared against the score for the original template. The synthetic template can be identified as the conformation with the most negative difference in energy score compared to the original template. Synthetic templates that improve the accuracy of the prediction may be associated with large, negative differences in score compared to their respective original templates. To identify these useful synthetic templates, the entire set can be sorted based on their score difference and those in the 65th percentile may be selected to replace their original template counterparts.
[0057] Processes that use a reduced protein representation, such as that in use by NMA in the accelerated conformational sampling technique, do not sacrifice overall model accuracy since an all-atom model refinement remains part of the prediction process. In fact, NMA can be more adept than all-atom approaches like molecular dynamics (MD) at modeling biological motions occurring at the micro- to millisecond time scale. NMA can generate these conformations in minutes of process time while MD requires weeks or months of process time, which greatly improves the practicality of the applying the accelerated conformational sampling process on a large scale.
[0058] Selected templates may be then used to start fourteen Monte Carlo Markov Chain (MCMC) simulations in order to search the for low energy conformations of the query protein. Physical pairwise distance restraints and internal residue contacts may be calculated from the templates and then mapped onto the query protein using the threading alignment. Each simulation has a “decoy pool,” a set of chemically plausible conformations that can be updated with lower energy conformations over the course of the simulation. The global pool can be a square matrix where a list of decoys can be maintained over a range of simulation temperatures. The decoy list length can be dependent on protein chain length: 40 decoys for proteins less than 165 residues, 50 decoys for proteins between 165 and 240 residues, 60 decoys between 240 and 300, 70 decoys between 300 and 400, and 80 decoys for proteins greater than 400 residues. The simulations proceed for a predetermined number of simulation cycles (ab initio: 250, templatebased: 500) or 50 CPU hours. For each temperature, a local decoy pool can be constructed and each individual decoy can be randomly perturbed for as many residues that may be allowed to move in that decoy. Local data structures may be used such that this step can be performed in parallel.
[0059] Perturbations may be limited to five basic movements: 1) moving two or three consecutive bonds, 2) moving two consecutive instances of the first type, 3) translating six to twelve consecutive bonds as a rigid group, 4) inducing a sequence-shift by permuting a three bonds fragment with a two bond fragment, 5) and constructing a random conformation between two distant residues. The old and new decoys may be evaluated by a knowledge-based potentialenergy function that considers correlations between contacting residue types, local structural stiffness caused by hydrogen bonding, template-derived distance and contact restraints, long range pairwise interactions, electrostatic interactions, and contact order (average sequence separation between contacting residues). Random moves may be accepted and rejected based on the Metropolis criterion as follows. The potential energy of the change AU can be calculated for the move and if AU<0 the movement can be automatically accepted. If instead AUX), then calculate a probability cutoff to accept the move as Pacc=e-AU / T where T can be a temperature value at the stage in the simulation. A random number r can be generated between 0 and 1. If Pacc>r, then the move can be accepted. If accepted, the new decoy replaces the old. If rejected and the simulation can be template-based, a single recovery step can be attempted where the fragment can be repositioned by rotation and translation. Afterwards, the global pool can be synchronized with the updated local pools and the decoys may be saved in a trajectory file. In total, tens of thousands of chemically-plausible conformations may be recorded across all simulations.
[0060] Next, all of the decoy trajectories may be collected and clustered by the k-means method into ten clusters. The center-most representative of each cluster can be selected for final refinement, in which backbone and side chain atoms may be added to the selected scaffold and a MCMC energy minimization repacks the side chains and relaxes the backbone atoms into the final low energy conformation, resulting in the final predicted structures.
[0061] Given the stochastic nature of the MCMC process, one cannot estimate how much additional computation effort can be required — it may be weeks, months, or longer. Simulation methods may be also prone to blockage by thermodynamic and kinetic barriers in the folding energy landscape and may never discover a distant native conformation.
[0062] The construction of high-resolution models can be hindered by selecting templates whose conformations may be structurally distant to the biologically relevant conformation of the target, regardless if the template represents the correct topological fold. Given the limited number of known structures in the PDB, this situation can be expected to occur often. In addition, selecting a single, useful synthetic template from a population of thousands of unnatural perturbations can be challenging. One could use Molecular dynamics, Monte Carlo (MC) perturbation, and all forms of NMA to perturb a template; however, the blind inclusion of suboptimal synthetic templates may decrease the quality of the resulting prediction. One could also use an alternative scoring function to select a putative synthetic template.
[0063] For template-based structure prediction methods that rely on a single template, the target protein sequence can be simply mapped onto the synthetic template scaffold rather than theoriginal template if the synthetic template has a lower energetic score than the original. This can be a direct one-to-one replacement.
[0064] Different techniques may be required for template-based structure prediction methods that rely on multiple templates. For n templates, there may be up to n individual synthetic template substitutions possible. Replacing all top ranked threading templates with synthetic templates can result in a decreased prediction accuracy (the structural similarity as measured by RMSD between the predicted structure and the known experimental structures may be worse than that for the unmodified approach). All putative synthetic templates may be evaluated by an energetic scoring function that approximates the biophysical forces that mediate protein folding, which can be calculated by performing one round of MCMC simulation. The synthetic templates may be re-ranked using the energetic score difference between each synthetic template and its original template.
[0065] A method for creating NMA-based perturbations for synthetic template construction can be to perturb a protein structure along the principal vector defining a normal mode or along discrete linear combinations of modes. The method can be straightforward to apply; however, its use may not scale well beyond pairwise combinations of normal modes. Also, since normal modes may be conventionally linear vectors, implausible bond stretching and bending may occur as the deformation magnitude increases. Energetic minimization processes, like those present in hybrid ab initio structure prediction, may be used to correct chemically-implausible deformations.
[0066] Biomacromolecular structures can be represented by either Cartesian coordinates or by an internal coordinate system describing bonds and bond angles within the atomic structure. A method for creating NMA-based perturbations can be to map normal modes from linear Cartesian vectors to changes in torsional angles, which describe the rotations of the protein backbone around the bonds containing each residue's carbon-alpha atom, and perturb the structure strictly along torsion angles. This “rigid geometry” approach conserves bond lengths and angles and always generates a chemically-plausible conformation. High-strain torsion angles may be still possible, but they remain more energetically favorable than impossibly long bonds or contorted angles.
[0067] An embodiment of an internal coordinate system involves dividing a structure into rigid elements whose motion can be restricted to rotations about the bonds of the protein chain. A rigid element can be a minimal unit of a protein over which its associated atoms remain fixed during perturbation. Each rigid element may contain one or more atoms and its local frame of reference can be centered on a key atom. The 3D conformation of a protein structure can be defined by a series of rigid elements, each with an independent rotation and translation relativeto its preceding element. Amino acid side chains may be represented as a rotamer, a conformational isomer whose atomic positions may be defined by torsion angles, that can be defined using the original structure or selected from a library of rotamers. Carbonyl and nitrogen groups may be represented by separate rigid elements even though the connecting peptide bond constraints the atoms into a fixed planar group. In addition, the restricted geometry of proline residues and the steric interactions between neighboring atoms may further constrain rotatable bonds.
[0068] The position and orientation of a local frame can be described relative to its prior element along the protein backbone by a 4 / 4 rotational and translational matrix. Similarly, the subsequent local frame can be described relative to the current frame by a 4x4 rotational and translation matrix. Multiplication of these matrices can be used to derive the orientation of any given reference frame in the chain. If atom 1 can be placed at a reference frame origin and the subsequent reference frame origin for atom 2 can be referenced by vector V=[tx, ty, tz] defined along the bond between atoms 1 and 2, then the rotation around V (such as a torsion angle) can be described by a rotation matrix R and a translation matrix T. The R and T matrices may be combined as a single rotation-translation matrix (RT).
[0069] RT describes a right-handed coordinate system whose axes satisfy the right-hand rule. Obtaining global Cartesian coordinates for an atom with position V=[xi, yi, zi, 1] in a local frame then involves the chained matrix multiplications of all the prior RT matrices:
[0070] NMA-based perturbations distort a protein structure along the vector defining a principal normal mode or a discrete linear combination of modes. To satisfy the rigid geometry approximation, normal modes can be converted from displacements in Cartesian space to changes in internal dihedral angles. The Cartesian displacement coordinates may be related to the internal rotations as defined by the Wilson s-vector method: s=Bd, where s can be a vector of internal coordinates, d can be a vector of Cartesian displacement coordinates, and B can be a matrix of constants determined by the geometry of the molecule. The normal modes may be mapped to a quaternion [tx, ty, tz, 9] that describes a vector representing the bond between two atoms with torsion angle 9. The rotation matrix determined from this quaternion can be used to update the original rotation matrix R and to construct a new RT matrix that includes the normal mode distortion.
[0071] An embodiment for exploring the chemically-plausible conformational space along an arbitrary number of normal modes can be to use a rapidly exploring random tree (RRT) data structure combined with a collision detection process. An RRT encourages incremental perturbations towards unexplored portions of an original template's conformational space and efficiently supports the search along a combination of dozens of normal modes, rather thansimple pairwise combinations. In addition to using a rigid geometry approximation, a chemical restraint process and a collision detection process screen for movements that cannot physically occur, which stops further searching along that particular combination of normal modes. The approach replaces the use of a time-consuming energetic calculation with a fast validation of atomic geometry.
[0072] The RRT structure defines a conformation as a vector of weights of size equivalent to the number of normal modes used for the exploration (effectively an arbitrary number, but typically 20 modes for modeling large, concerted, motions and 50 modes for modeling local fluctuations). The sampling of conformational space increases as the tree structure grows, which involves creating a random conformation by generating a new set of weights, incrementally perturbing the nearest saved conformation toward the new conformation and adding new plausible conformations to the tree. Each incremental change can be evaluated for preserving chemical linkages (hydrogen bonds between strands in a beta sheet and disulfide bonds) and preventing atomic clashes (described later). The linkage geometry can be determined from the unperturbed structure. To determine whether a linkage restraint can be conserved, the Cartesian coordinates for atoms involved in linkages may be calculated and the distance between pairwise atoms involved in the linkages may be evaluated. For collisions involving side chain atoms, alternate rotamers may be screened for poses that resolve the clash. If the collision cannot be resolved, the last valid conformation can be added as a leaf node to the previous saved conformation and the cycle repeats itself. RRT expansion ends when the tree reaches a maximum size, a series of sequential attempts to expand the tree fail (typically 100 attempts), or a time limit can be exceeded (typically ten minutes).
[0073] Collision detection can be implemented as a one-dimensional Sweep and Prune (SAP) process. SAP can be a sorting-based technique that examines the overlap between bounding boxes around rigid elements. The process can be optimal for collections where all elements move incrementally, such as the concerted motions described by normal modes. For a given conformation, the extrema of the bounding boxes may be projected onto a vector defined by the original template's N-terminal nitrogen and C-terminal carbon, and then a linked list of the elements can be sorted by their projected minima. Potential overlapping pairs of elements may be identified based on whether their minima overlap, then all atom pairs between the two elements may be tested for collisions by identifying overlapping hard spheres with van der Waals radii (typically 80% of ideal values to account for allowable soft penetration). Subsequent collision detection for incremental movements requires less time because the elements may be now partially sorted. This embodiment can be the first SAP process to consider an all-atomrepresentation of a protein with residue-specific side chains, which can be a critical requirement for accurately detecting atomic clashes within a protein structure.
[0074] The RRT process can be used to sample the accessible conformational space for an original template along a set of many normal modes, which creates a set of putative synthetic templates with greater structural diversity than those achievable using simple pairwise combinations of normal modes. Finally, the established process for selecting a synthetic template can be applied to this resulting set. This process can be repeated for every template selected by protein threading and the associated synthetic templates may be integrated into the prediction process as previously described.
[0075] In some embodiments, a conformational sampling method for predicting the structure of an amino acid sequence can be applied to an amino acid sequence. Any method described herein may independently and in any combination comprise: creating a sequence profile matrix of the amino acid sequence; determining an alignment of each respective residue of the amino acid sequence against the sequence profile matrix; identifying internal contacts for one or more residues in the plurality of residues using the alignment; collecting the features of the sequence profile matrix and internal contacts for each residue in the plurality of residues into a collected feature matrix; aligning the collected feature matrix using one or more threading models with a structural feature database of original templates; selecting a plurality of optimally aligned original templates from the aligning; calculating normal modes of motion for each original template in the plurality of optimally aligned original templates; perturbing each respective original template in the plurality of optimally aligned original templates, for each pair of calculated normal modes, thereby collectively creating a plurality of synthetic templates; scoring the energy difference between each original template and the corresponding synthetic template; selecting a subset of synthetic templates from the plurality of synthetic templates based on satisfaction of a predetermined cut-off criterion; replacing or supplementing the original templates in the plurality of optimally aligned original templates with the corresponding selected subset of synthetic templates to generate a plurality of modeling templates; calculating distance and contact restraints within modeling templates of the plurality of modeling templates; performing, with modeling templates in the plurality of modeling templates, Markov Chain Monte Carlo simulations, thereby obtaining simulation results; clustering the simulation results, wherein the clustering comprises a plurality of clusters, each cluster in the plurality of clusters representing models from the performing; selecting representative models of each cluster in the plurality of clusters; refining the representative models of the selecting by energy minimization; and / or selecting the lowest energy refined representative model as the predicted structure of the amino acid sequence.
[0076] According to some embodiments, the methods described herein can comprise designing and selecting an amino-acid sequence having a desired affinity to a molecular surface of interest of a molecular entity. Such methods may be carried out by providing atomic coordinates of atoms of a definable molecular surface of interest that forms a part of the molecular entity. In some cases, the alternate surface can be the surface of a second protein.
[0077] Another example workflow of a method according to this disclosure is illustrated in FIG. 2. In this example workflow, a protein sequence about which a user seeks to make a prediction or classification is obtained 201. A computation of the abundance of one or more residues in the sequence (for example, of each of the 20 naturally occurring amino acids present in the sequence) is performed 202. The resulting abundances are then ranked 203, and a protein fingerprint matrix is generated based at least in part on the computed abundances and the protein sequence 204. The computed fingerprint matrix can then be used to classify or predict protein folding and / or to predict interaction with a second protein 205, for example, using a trained machine learning classifier.
[0078] In some cases, matrices used in any step of any of the methods described herein may be n-dimensional. In some embodiments, n is about 1 to about 20. In some embodiments, n is about 1 to about 2, about 1 to about 3, about 1 to about 4, about 1 to about 5, about 1 to about 10, about 1 to about 15, about 1 to about 20, about 2 to about 3, about 2 to about 4, about 2 to about 5, about 2 to about 10, about 2 to about 15, about 2 to about 20, about 3 to about 4, about 3 to about 5, about 3 to about 10, about 3 to about 15, about 3 to about 20, about 4 to about 5, about 4 to about 10, about 4 to about 15, about 4 to about 20, about 5 to about 10, about 5 to about 15, about 5 to about 20, about 10 to about 15, about 10 to about 20, or about 15 to about 20. In some embodiments, n is about 1, about 2, about 3, about 4, about 5, about 10, about 15, or about 20. In some embodiments, n is at least about 1, about 2, about 3, about 4, about 5, about 10, or about 15. In some embodiments, n is at most about 2, about 3, about 4, about 5, about 10, about 15, or about 20.
[0079] Methods disclosed herein which are particularly useful for classification of one or more protein-protein interaction(s) can comprise one or more, or all of the steps of the workflow illustrated in FIG. 3. For example, a workflow for predicting a protein-protein interaction and / or a lack thereof may comprise obtaining a first protein sequence about which a user seeks to make a prediction or classification 301, segmenting the sequence into a plurality of windows 302, computing an abundance of one or more amino acid residues (for example, by computing abundance of each of the 20 naturally occurring amino acids within the sequence, and / or within a subset of, and / or each of the plurality of windows) 303, ranking the residues based at least in part on the computed abundance(s) 304, generating a protein fingerprint based at least in part onthe ranked abundance(s) 305, computing an interaction fingerprint (e.g. an interaction matrix) between the protein sequence and a second protein sequence 306, and / or classifying or predicting protein-protein interaction and / or a lack thereof, for example, using a trained machine learning classifier.
[0080] Described herein are computational methods for proteins. The methods can comprise obtaining a sequence of a first protein. The methods can comprise computing an abundance of a plurality of amino acid residues within the sequence to generate a first fingerprint matrix for the first protein, wherein the first fingerprint matrix can comprise a plurality of vectors that can be ranked based at least in part on the abundance of the plurality of amino acid residues. The methods comprise providing the first fingerprint matrix for use in modeling, classifying or predicting at least one of protein folding ability or protein-protein interactions. The plurality of vectors can be ranked solely on the abundance of the plurality of amino acid residues, without using or requiring chemical information about either the sequence or the amino acid residues within the sequence. An interaction fingerprint matrix can be computed using the first fingerprint matrix for the first protein and a second fingerprint matrix for a second protein.
[0081] The second fingerprint matrix can comprise a plurality of vectors that can be ranked based at least in part on an abundance of a plurality of amino acid residues within a sequence of the second protein. The method may comprise providing the interaction fingerprint matrix to a trained machine learning classifier and using the trained machine learning classifier to generate a classification indicative of an interaction or lack thereof between the first protein and the second protein.
[0082] An accuracy of the classification can have an area under the curve (AUC) value of at least 0.75, 0.79, and / or 0.969. The plurality of vectors in the first fingerprint matrix can be associated with a plurality of window segments defined over a length of the sequence of the first protein.
[0083] The abundance may be computed within each window segment of the plurality of window segments. The abundance may be computed across the plurality of window segments. A length of the plurality of window segments may be about 1% to about 99% of the length of the sequence of the first protein.
[0084] A length of each of the plurality of window segment may be about 3 to about 100 amino acid residues in length. In some embodiments, a length of each of the plurality of window segments is about 100 to about 1000 amino acid residues in length. In some embodiments, a length of each of the plurality of window segments is at least 100 amino acid residues (e.g., at least 150 amino acid residues, at least 200 amino acid residues, at least 300 amino acid residues, or at least 500 amino acid residues) in length. In some cases, each of the plurality of windowsegments overlaps an alternate one of the plurality of window segments by at least one amino acid residue. In some cases, each of the plurality of window segments do not overlap. The plurality of window segments can comprise both overlapping and non-overlapping segments.
[0085] The plurality of vectors in the first fingerprint matrix can be based at least in part on the abundance of each of twenty naturally occurring amino acid residues. The first fingerprint can comprise an abundance of each naturally occurring amino acid residue in the first protein sequence and an abundance of each naturally occurring amino acid in each of the plurality of window segments. The plurality of vectors can be ranked by abundance in ascending order. Alternatively, the plurality of vectors can be ranked by abundance in descending order. The plurality of vectors can be ranked by abundance according to a standard (ordinal) ranking procedure. The plurality of vectors can be ranked by abundance according to a fractional ranking procedure. The plurality of vectors can be ranked by abundance according to a dense ranking procedure.
[0086] The plurality of vectors can be ranked by abundance according to a competition ranking procedure. The plurality of vectors can be ranked by abundance according to a modified competition ranking procedure. The plurality of vectors can be ranked by abundance according to a random ranking procedure. The plurality of vectors can be ranked by abundance according to a ranking procedure comprising one or more secondary ranking criteria.
[0087] The fingerprint matrices may be used to predict protein folding ability. The fingerprint matrices may be used to predict protein-protein interaction. The method can comprise predicting, using a machine learning model, a novel protein sequence capable of interacting with the first protein sequence. The novel protein sequence can interact with a defined target segment of the first protein sequence.
[0088] The defined target segment may be input by a user. The defined target segment may be identified using the machine learning model. Interaction with the defined target segment may be expected to agonize and / or antagonize a protein having the first protein sequence upon formation of a protein-protein complex.
[0089] The novel protein sequence may be a sequence of an engineered antibody. The antibody can have a therapeutic effect when administered to a subject suffering from a condition. The novel protein sequence can comprise an improved variant of an initial sequence provided to the system by a user. The method can comprise iteratively generating by the machine learning model, a plurality of improved sequences until a threshold interaction value may be achieved.
[0090] The threshold interaction value may be a threshold agonist effect on a protein having the first protein sequence. The threshold interaction value may be a threshold antagonist effect on a protein having the first protein sequence. The method may be used to predict or classify protein-protein interaction of a protein-protein complex comprising at least three proteins. The first protein sequence may be an engineered protein sequence and prediction or classification of foldability may be used to screen a plurality of engineered protein sequence candidates.
[0091] The engineered protein sequence may be a sequence of an enzyme. A plurality of enzyme sequence candidates can be screened for enzymatic activity and / or enzyme stability. The method can comprise automatically identifying, by a machine learning model, an enzyme sequence having an optimized enzymatic activity and / or an optimized enzyme stability.
[0092] The optimized enzyme sequence may be optimized for stability in conditions reflective of a large-scale biomanufacturing reactor. The optimized enzyme sequence may be optimized for enzymatic activity in conditions reflective of a large-scale biomanufacturing reactor. The optimized enzyme sequence may be optimized for both enzymatic activity and stability in a biomanufacturing reactor.
[0093] An accuracy of the classification can have an area under the curve (AUC) value of at least 0.7, 0.8, or 0.9, determined using a receiver operating characteristic curve. An identity of one or more proteins may be used solely as an input to any of the methods described herein, wherein classifications or predictions can be made without using or requiring chemical information about either the sequence or the amino acid residues within the sequence.
[0094] A visual depiction of a workflow according to this disclosure, which is similar to the embodiment illustrated in FIG. 3 is shown in FIG. 4. The example in FIG. 4, illustrates a binary classification (interaction or non-interaction; foldable protein sequence or non-foldable sequence; etc.). When used for binary classification, methods described herein provide for reliable and accurate classification which can be described using a receiver operating characteristic (ROC) curve, for example, as illustrated in FIG. 5.
[0095] In some embodiments, an area under a ROC curve for one or more binary classifications provided by any of the methods described herein is at least about 0.8 to about 0.995. In some embodiments, an area under a ROC curve for one or more binary classifications provided by any of the methods described herein is at least about 0.7 to about 0.85, about 0.7 to about 0.9, about 0.7 to about 0.95, about 0.7 to about 0.97, about 0.7 to about 0.99, about 0.7 to about 0.995, about 0.85 to about 0.9, about 0.85 to about 0.95, about 0.85 to about 0.97, about 0.85 to about 0.99, about 0.85 to about 0.995, about 0.9 to about 0.95, about 0.9 to about 0.97, about 0.9 to about 0.99, about 0.9 to about 0.995, about 0.95 to about 0.97, about 0.95 to about 0.99, about 0.95 to about 0.995, about 0.97 to about 0.99, about 0.97 to about 0.995, or about 0.99 to about 0.995. In some embodiments, an area under a ROC curve for one or more binary classifications provided by any of the methods described herein is at least about 0.7, about 0.8, about 0.85, about 0.9, about 0.95, about 0.97, about 0.99, or about 0.995. In some embodiments, an areaunder a ROC curve for one or more binary classifications provided by any of the methods described herein is at least at least about 0.8, about 0.85, about 0.9, about 0.95, about 0.97, or about 0.99.
[0096] Methods described herein can also provide for continuous classification (e.g. by providing an interaction strength estimate) or may provide additional qualitative information about foldability and / or interaction (e.g. identification of residues of each sequence involved in interaction, three-dimensional structural information about a protein-protein complex, and / or a predicted sequence of a protein which interacts with the input protein). Such applications may be accomplished using one or more machine learning classifiers, for example, by training a classifier to provide any of the quantities and / or qualitative parameters described herein. In some cases, the same machine learning classifier can be used for both binary and non-binary classification. Alternatively, different classifiers may be configured to perform each function separately. Multiple classifiers may also be combined in series or parallel as needed in order to accomplish a particular application.Machine Learning Algorithms
[0097] Disclosed herein are platforms, systems, and methods that provide protein classification using machine learning algorithm(s) based in part on a sequence of the protein. In particular, in some aspects, the machine learning algorithms include deep learning neural networks configured for evaluating protein-protein interaction and / or protein foldability, or for evaluating, designing, and / or refining a novel protein sequence which interacts with a target protein. The algorithms can include or be used to assist in any of the operations of systems or methods described herein.
[0098] The development of each machine learning algorithm spans three phases: (1) dataset creation and curation, (2) algorithm training, and (3) adapting design elements for product performance and useability. The dataset used for training the algorithm can be generated by obtaining interaction and / or folding information of a variety of known samples (e.g. from various databases) that are then curated and / or labeled with particular quantities or metrics of interest. Each algorithm then undergoes training using the training dataset, which can include one or more different metrics desired in an output, such as foldability, interactability, interaction regions, identification of critical amino acids, affinity and / or interaction strength.
[0099] A machine learning model can comprise a supervised, semi-supervised, unsupervised, or self-supervised machine learning model. In some cases, the one or more ML approaches perform classification or clustering of the sequence classification data. In some examples, the machine learning approach comprises a classical machine learning method, such as, but not limited to, support vector machine (SVM) (e.g., one-class SVM, linear or radial kernels, etc.), K-nearest neighbor (KNN), isolation forest, random forest, logistic regression, AdaBoost classifier, extratrees classifier, extreme gradient boosting, gaussian process classifier, gradient boosting classifier, light gradient boosting, linear discriminant analysis, naive Bayes, quadratic discriminant analysis, ridge classifier, or any combination thereof. In some examples, the machine learning approach comprises a deep leaning method (e.g., deep neural network (DNN)), such as, but not limited to a fully-connected network, convolutional neural network (CNN) (e.g., one-class CNN), recurrent neural network (RNN), transformer, graph neural network (GNN), convolutional graph neural network (CGNN), multi-level perceptron (MLP), or any combination thereof.
[0100] In some aspects, a classical ML method comprises one or more algorithms that learns from existing observations (e.g., known features) to predict outputs. In some aspects, the one or more algorithms perform clustering of data. In some examples, the classical ML algorithms for clustering comprise K-means clustering, mean-shift clustering, density-based spatial clustering of applications with noise (DBSCAN), expectation-maximization (EM) clustering (e.g., using Gaussian mixture models (GMM)), agglomerative hierarchical clustering, or any combination thereof. In some aspects, the one or more algorithms perform classification of data. In some examples, the classical ML algorithms for classification comprise logistic regression, naive Bayes, KNN, random forest, isolation forest, decision trees, gradient boosting, support vector machine (SVM), or any combination thereof. In some examples, the SVM comprises a one-class SMV or a multi-class SVM.
[0101] In some aspects, the deep learning method comprises one or more algorithms that learns by extracting new features to predict outputs. In some aspects, the deep learning method comprises one or more layers. In some aspects, the deep learning method comprises a neural network (e.g., DNN comprising more than one layer). Neural networks generally comprise connected nodes in a network, which can perform functions, such as transforming or translating input data. In some aspects, the output from a given node is passed on as input to another node. The nodes in the network generally comprise input units in an input layer, hidden units in one or more hidden layers, output units in an output layer, or a combination thereof. In some aspects, an input node is connected to one or more hidden units. In some aspects, one or more hidden units is connected to an output unit. The nodes can generally take in input through the input units and generate an output from the output units using an activation function. In some aspects, the input or output comprises a tensor, a matrix, a vector, an array, or a scalar. In some aspects, the activation function is a Rectified Linear Unit (ReLU) activation function, a sigmoid activation function, a hyperbolic tangent activation function, or a Softmax activation function.
[0102] The connections between nodes further comprise weights for adjusting input data to a given node (e.g., to activate input data or deactivate input data). In some aspects, the weights areleamed by the neural network. In some aspects, the neural network is trained to learn weights using gradient-based optimizations. In some aspects, the gradient-based optimization comprises one or more loss functions. In some aspects, the gradient-based optimization is gradient descent, conjugate gradient descent, stochastic gradient descent, or any variation thereof (e.g., adaptive moment estimation (Adam)). In some further aspects, the gradient in the gradient-based optimization is computed using backpropagation. In some aspects, the nodes are organized into graphs to generate a network (e.g., graph neural networks). In some aspects, the nodes are organized into one or more layers to generate a network (e.g., feed forward neural networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), etc.). In some aspects, the CNN comprises a one-class CNN or a multi-class CNN.
[0103] In some aspects, the neural network comprises one or more recurrent layers. In some aspects, the one or more recurrent layers are one or more long short-term memory (LSTM) layers or gated recurrent units (GRUs). In some aspects, the one or more recurrent layers perform sequential data classification and clustering in which the data ordering is considered (e.g., time series data). In such aspects, future predictions are made by the one or more recurrent layers according to the sequence of past events. In some aspects, the recurrent layer retains or “remembers” important information, while selectively “forgets” what is not essential to the classification.
[0104] In some aspects, the neural network comprise one or more convolutional layers. In some aspects, the input and the output are a tensor representing variables or attributes in a data set (e.g., features), which may be referred to as a feature map (or activation map). In such aspects, the one or more convolutional layers are referred to as a feature extraction phase. In some aspects, the convolutions are one-dimensional (ID) convolutions, two dimensional (2D) convolutions, three dimensional (3D) convolutions, or any combination thereof. In further aspects, the convolutions are ID transpose convolutions, 2D transpose convolutions, 3D transpose convolutions, or any combination thereof.
[0105] The layers in a neural network can further comprise one or more pooling layers before or after a convolutional layer. In some aspects, the one or more pooling layers reduces the dimensionality of a feature map using filters that summarize regions of a matrix. In some aspects, this down samples the number of outputs, and thus reduces the parameters and computational resources needed for the neural network. In some aspects, the one or more pooling layers comprises max pooling, min pooling, average pooling, global pooling, norm pooling, or a combination thereof. In some aspects, max pooling reduces the dimensionality of the data by taking the maximums values in the region of the matrix. In some aspects, this helps capture the most significant one or more features. In some aspects, the one or more poolinglayers is one dimensional (ID), two dimensional (2D), three dimensional (3D), or any combination thereof.
[0106] The neural network can further comprise of one or more flattening layers, which can flatten the input to be passed on to the next layer. In some aspects, an input (e.g., feature map) is flattened by reducing the input to a one-dimensional array. In some aspects, the flattened inputs can be used to output a classification of an object. In some aspects, the classification comprises a binary classification or multi-class classification of visual data (e.g., images, videos, etc.) or non-visual data (e.g., measurements, audio, text, etc.). In some aspects, the classification comprises binary classification of an image. In some aspects, the classification comprises multiclass classification of a plurality of metrics. In some aspects, the classification comprises binary classification of a measurement. In some examples, the binary classification of a measurement comprises a classification of a system’s performance using the physical measurements described herein (e.g., normal or abnormal, normal or anormal).
[0107] The neural networks can further comprise of one or more dropout layers. In some aspects, the dropout layers are used during training of the neural network (e.g., to perform binary or multi-class classifications). In some aspects, the one or more dropout layers randomly set some weights as 0 (e.g., about 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% of weights). In some aspects, the setting some weights as 0 also sets the corresponding elements in the feature map as 0. In some aspects, the one or more dropout layers can be used to avoid the neural network from overfitting.
[0108] The neural network can further comprise one or more dense layers, which comprises a fully connected network. In some aspects, information is passed through a fully connected network to generate a predicted classification of an object. In some aspects, the error associated with the predicted classification of the object is also calculated. In some aspects, the error is backpropagated to improve the prediction. In some aspects, the one or more dense layers comprises a Softmax activation function. In some aspects, the Softmax activation function converts a vector of numbers to a vector of probabilities. In some aspects, these probabilities are subsequently used in classifications, such as classifications of one or a plurality of image features comprised in one or a plurality of images.Computer systems
[0109] The present disclosure provides computer systems that are programmed to implement methods of the disclosure. FIG. 6 shows a computer system 601 that is programmed or otherwise configured to classify proteins based on sequence information. The computer system 601 can regulate various aspects of determining protein-protein interaction and / or proteinfoldability of the present disclosure, such as, for example, by implementing any of the methods described herein. The computer system 601 can be an electronic device of a user or a computer system that is remotely located with respect to the electronic device. The electronic device can be a mobile electronic device.
[0110] The computer system 601 includes a central processing unit (CPU, also “processor” and “computer processor” herein) 605, which can be a single core or multi core processor, or a plurality of processors for parallel processing. The computer system 601 also includes memory or memory location 610 (e.g., random-access memory, read-only memory, flash memory), electronic storage unit 615 (e.g., hard disk), communication interface 620 (e.g., network adapter) for communicating with one or more other systems, and peripheral devices 625, such as cache, other memory, data storage and / or electronic display adapters. The memory 610, storage unit 615, interface 620 and peripheral devices 625 are in communication with the CPU 605 through a communication bus (solid lines), such as a motherboard. The storage unit 615 can be a data storage unit (or data repository) for storing data. The computer system 601 can be operatively coupled to a computer network (“network”) 630 with the aid of the communication interface 620. The network 630 can be the Internet, an internet and / or extranet, or an intranet and / or extranet that is in communication with the Internet. The network 630 in some cases is a telecommunication and / or data network. The network 630 can include one or more computer servers, which can enable distributed computing, such as cloud computing. The network 630, in some cases with the aid of the computer system 601, can implement a peer-to-peer network, which may enable devices coupled to the computer system 601 to behave as a client or a server. [OHl] The CPU 605 can execute a sequence of machine-readable instructions, which can be embodied in a program or software. The instructions may be stored in a memory location, such as the memory 610. The instructions can be directed to the CPU 605, which can subsequently program or otherwise configure the CPU 605 to implement methods of the present disclosure. Examples of operations performed by the CPU 605 can include fetch, decode, execute, and writeback.
[0112] The CPU 605 can be part of a circuit, such as an integrated circuit. One or more other components of the system 601 can be included in the circuit. In some cases, the circuit is an application specific integrated circuit (ASIC).
[0113] The storage unit 615 can store files, such as drivers, libraries and saved programs. The storage unit 615 can store user data, e.g., user preferences and user programs. The computer system 601 in some cases can include one or more additional data storage units that are external to the computer system 601, such as located on a remote server that is in communication with the computer system 601 through an intranet or the Internet.
[0114] The computer system 601 can communicate with one or more remote computer systems through the network 630. For instance, the computer system 601 can communicate with a remote computer system of a user (e.g., a database server comprising a trained classification algorithm). Examples of remote computer systems include personal computers (e.g., portable PC), slate or tablet PC’s (e.g., Apple® iPad, Samsung® Galaxy Tab), telephones, Smart phones (e.g., Apple® iPhone, Android-enabled device, Blackberry®), or personal digital assistants. The user can access the computer system 601 via the network 630.
[0115] Methods as described herein can be implemented by way of machine (e.g., computer processor) executable code stored on an electronic storage location of the computer system 601, such as, for example, on the memory 610 or electronic storage unit 615. The machine executable or machine readable code can be provided in the form of software. During use, the code can be executed by the processor 605. In some cases, the code can be retrieved from the storage unit 615 and stored on the memory 610 for ready access by the processor 605. In some situations, the electronic storage unit 615 can be precluded, and machine-executable instructions are stored on memory 610.
[0116] The code can be pre-compiled and configured for use with a machine having a processer adapted to execute the code, or can be compiled during runtime. The code can be supplied in a programming language that can be selected to enable the code to execute in a pre-compiled or as-compiled fashion.
[0117] Aspects of the systems and methods provided herein, such as the computer system 601, can be embodied in programming. Various aspects of the technology may be thought of as “products” or “articles of manufacture” typically in the form of machine (or processor) executable code and / or associated data that is carried on or embodied in a type of machine readable medium. Machine-executable code can be stored on an electronic storage unit, such as memory (e.g., read-only memory, random-access memory, flash memory) or a hard disk.“Storage” type media can include any or all of the tangible memory of the computers, processors or the like, or associated modules thereof, such as various semiconductor memories, tape drives, disk drives and the like, which may provide non-transitory storage at any time for the software programming. All or portions of the software may at times be communicated through the Internet or various other telecommunication networks. Such communications, for example, may enable loading of the software from one computer or processor into another, for example, from a management server or host computer into the computer platform of an application server. Thus, another type of media that may bear the software elements includes optical, electrical and electromagnetic waves, such as used across physical interfaces between local devices, through wired and optical landline networks and over various air-links. The physical elements that carrysuch waves, such as wired or wireless links, optical links or the like, also may be considered as media bearing the software. As used herein, unless restricted to non-transitory, tangible “storage” media, terms such as computer or machine “readable medium” refer to any medium that participates in providing instructions to a processor for execution.
[0118] Hence, a machine readable medium, such as computer-executable code, may take many forms, including but not limited to, a tangible storage medium, a carrier wave medium or physical transmission medium. Non-volatile storage media include, for example, optical or magnetic disks, such as any of the storage devices in any computer(s) or the like, such as may be used to implement the databases, etc. shown in the drawings. Volatile storage media include dynamic memory, such as main memory of such a computer platform. Tangible transmission media include coaxial cables; copper wire and fiber optics, including the wires that comprise a bus within a computer system. Carrier-wave transmission media may take the form of electric or electromagnetic signals, or acoustic or light waves such as those generated during radio frequency (RF) and infrared (IR) data communications. Common forms of computer-readable media therefore include for example: a floppy disk, a flexible disk, hard disk, magnetic tape, any other magnetic medium, a CD-ROM, DVD or DVD-ROM, any other optical medium, punch cards paper tape, any other physical storage medium with patterns of holes, a RAM, a ROM, a PROM and EPROM, a FLASH-EPROM, any other memory chip or cartridge, a carrier wave transporting data or instructions, cables or links transporting such a carrier wave, or any other medium from which a computer may read programming code and / or data. Many of these forms of computer readable media may be involved in carrying one or more sequences of one or more instructions to a processor for execution.
[0119] The computer system 601 can include or be in communication with an electronic display 635 that comprises a user interface (UI) 640 for providing, for example, sequence information inputs, sequence manipulation tools, and / or prediction or classification outputs. Examples of UFs include, without limitation, a graphical user interface (GUI) and web-based user interface.
[0120] Methods and systems of the present disclosure can be implemented by way of one or more algorithms. An algorithm can be implemented by way of software upon execution by the central processing unit 605. The algorithm can, for example, be configured to implement any of the methods described herein.ExamplesExample 1. Classification of Protein Foldability using methods described herein:
[0121] A sequence of a protein is provided by a user to a system implementing methods of prediction and / or classification as described herein, for example, using the workflows of FIG. 1 or FIG. 2. A protein fingerprint is calculated based on sequence identity alone (e.g. without useof chemical or physical information such as Van der Waals forces or computation of the stability of any particular conformational landscape). A machine learning classifier as described herein to predict protein foldability. Despite minimal input data, the method is able to determine protein folding ability quickly and accurately based on the sequence alone.Example 2. Classification of Protein-Protein Interaction using methods described herein:
[0122] A method implementing the workflow illustrated in FIG. 4 was implemented using a machine learning classifier which was trained using a curated protein-protein interaction databases as a training set. Test set protein fingerprints were calculated based on sequence identity alone (e.g. without use of chemical or physical information such as Van der Waals forces or computation of the stability of any particular conformational landscape) and interaction fingerprints for a plurality of test proteins were calculated. A machine learning classifier was then used to predict whether or not first and second proteins for each pair tested would interact. Despite minimal input data, the method was able to determine interaction or a lack thereof quickly and accurately based on sequence identities alone. Accuracy was validated using receiver operating characteristic curves, such as the one shown in FIG. 5, for known test sequences. Accuracy varied from training database to training database, however, the area under the ROC curves for each database tested ranged from 0.75-0.969. For example, one database tested yielded: an area under the curve of 0.969, an accuracy of 0.91, an Fl-score of 0.92, and a recall of 0.91.Example 3. Prediction of a Protein Sequence that interacts with a target protein using methods described herein:
[0123] A researcher seeking to design a new biotherapeutic antibody inputs a target protein sequence into a method described herein (e.g. using the workflows of FIG. 3 and / or FIG. 4). A trained machine learning model iteratively screens a plurality of candidate sequences to identify one or more sequences expected to agonize or antagonize the target protein to at least a threshold target level. The one or more sequences are then output to the researcher, who subsequently synthesizes one or more antibodies having the one or more outputted sequences for further in-vitro and / or in-vivo testing.
[0124] While preferred embodiments of the present disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the present disclosure be limited by the specific examples provided within the specification. While the present disclosure has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from thepresent disclosure. Furthermore, it shall be understood that all aspects of the present disclosure are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the present disclosure described herein may be employed in practicing the present disclosure. It is therefore contemplated that the present disclosure shall also cover any such alternatives, modifications, variations, or equivalents. It is intended that the following claims define the scope of the present disclosure and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A computational method for proteins, comprising:(a) obtaining a sequence of a first protein;(b) computing an abundance of a plurality of amino acid residues within the sequence to generate a first fingerprint for the first protein, wherein the first fingerprint comprises a plurality of vectors that are ranked based at least in part on the abundance of the plurality of amino acid residues; and(c) providing the first fingerprint for use in modeling, classifying or predicting at least one of protein folding ability or protein-protein interactions.
2. The method of claim 1, wherein the first fingerprint comprises a fingerprint matrix, a fingerprint vector, or a fingerprint tensor.
3. The method of claim 2, wherein the first fingerprint is a fingerprint matrix.
4. The method of claim 1, wherein the plurality of vectors are ranked solely on the abundance of the plurality of amino acid residues, without using or requiring chemical information about either the sequence or the amino acid residues within the sequence.
5. The method of claim 1, wherein (c) comprises computing an interaction fingerprint using the first fingerprint and a second fingerprint for a second protein.
6. The method of claim 5, wherein the second fingerprint comprises a plurality of vectors that are ranked based at least in part on an abundance of a plurality of amino acid residues within a sequence of the second protein.
7. The method of claim 5, further comprising: providing the interaction fingerprint to a trained machine learning classifier and using the trained machine learning classifier to generate a classification indicative of an interaction or lack thereof between the first protein and the second protein.
8. The method of claim 7, wherein an accuracy of the classification has an area under the curve (AUC) value of at least 0.7.
9. The method of claim 1, wherein the plurality of vectors in the first fingerprint are associated with a plurality of window segments defined over a length of the sequence of the first protein.
10. The method of claim 9, wherein the abundance is computed within each window segment of the plurality of window segments.
11. The method of claim 9, wherein the abundance is computed across the plurality of window segments.
12. The method of any one of claims 9-11, wherein a length of the plurality of window segments is about 1% to about 99% of the length of the sequence of the first protein.
13. The method of any one of claims 9-12, wherein a length of each of the plurality of window segment is about 3 to about 100 amino acid residues in length.
14. The method of any one of claims 9-13, wherein each of the plurality of window segments overlaps an alternate one of the plurality of window segments by at least one amino acid residue.
15. The method of any one of claims 9-13, each of the plurality of window segments do not overlap.
16. The method of any one of claims 9-13, wherein the plurality of window segments comprises both overlapping and non-overlapping segments.
17. The method of claim 1, wherein plurality of vectors in the first fingerprint are based at least in part on the abundance each of twenty naturally occurring amino acid residues.
18. The method of claim 9-13, wherein the first fingerprint comprises abundance of each naturally occurring amino acid residue in the first protein sequence and an abundance of each naturally occurring amino acid in each of the plurality of window segments.
19. The method of any one of the preceding claims, wherein the plurality of vectors are ranked by abundance in ascending order.
20. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance in descending order.
21. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance according to a standard (ordinal) ranking procedure.
22. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance according to a fractional ranking procedure.
23. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance according to a dense ranking procedure.
24. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance according to a competition ranking procedure.
25. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance according to a modified competition ranking procedure.
26. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance according to a random ranking procedure.
27. The method of any of claims 1-19, wherein the plurality of vectors are ranked by abundance according to a ranking procedure comprising one or more secondary ranking criteria.
28. The method of any of claims 1-19, wherein the predicting of (c) is used to predict protein folding ability, protein folding region, and / or affinity.
29. The method of any of claims 1-19, wherein the predicting of (c) is used to predict protein-protein interaction.
30. The method of claim 29, further comprising predicting, using a machine learning model (e.g. a generative model), a novel protein sequence capable of interacting with the firstprotein sequence, and / or which has a target foldability, folding region, and / or affinity value.
31. The method of claim 30, wherein the novel protein sequence interacts with a defined target segment of the first protein sequence.
32. The method of claim 31, wherein the defined target segment is input by a user.
33. The method of claim 31, wherein the defined target segment is identified using the machine learning model.
34. The method of claim 32 or 33, wherein interaction with the defined target segment is expected to agonize and / or antagonize a protein having the first protein sequence upon formation of a protein-protein complex.
35. The method of claim 34, wherein the novel protein sequence is a sequence of an engineered antibody.
36. The method of claim 35, wherein the antibody has a therapeutic effect when administered to a subject suffering from a condition.
37. The method of claim 30, wherein the novel protein sequence comprises an improved variant of an initial sequence provided to the system by a user.
38. The method of claim 37, comprising iteratively generating by the machine learning model, a plurality of improved sequences until a threshold interaction value is achieved.
39. The method of claim 38, wherein the threshold interaction value is a threshold agonist effect on a protein having the first protein sequence.
40. The method of claim 39, wherein the threshold interaction value is a threshold antagonist effect on a protein having the first protein sequence.
41. The method of any one of claims 1-19, wherein the method is used to predict or classify protein-protein interaction of a protein-protein complex comprising at least three proteins.
42. The method of any one of claims 1-41, wherein the first protein sequence is an engineered protein sequence and prediction or classification of foldability is used to screen a plurality of engineered protein sequence candidates.
43. The method of claim 42, wherein the engineered protein sequence is a sequence of an enzyme.
44. The method of claim 43, wherein a plurality of enzyme sequence candidates are screened for enzymatic activity and / or enzyme stability.
45. The method of claim 44, wherein the method comprises automatically identifying, by a machine learning model, an enzyme sequence having an optimized enzymatic activity and / or an optimized enzyme stability.
46. The method of claim 45, wherein the optimized enzyme sequence is optimized for stability in conditions reflective of a large-scale biomanufacturing reactor.
47. The method of claim 46, wherein the optimized enzyme sequence is optimized for enzymatic activity in conditions reflective of a large-scale biomanufacturing reactor.
48. The method of claim 47, wherein the optimized enzyme sequence is optimized for both enzymatic activity and stability in a biomanufacturing reactor.
49. The method of claim 7, wherein an accuracy of the classification has an area under the curve (AUC) value of at least 0.7 (e.g., at least 0.9), determined using a receiver operating characteristic curve.
50. The method of any one of claim 1-49, wherein sequence identity of one or more proteins is provided or received as a sole input, wherein classifications or predictions are madewithout using or requiring chemical information about either the sequence or the amino acid residues within the sequence.
51. The method of any one of claims 9-12 or 14-50, wherein a length of each of the plurality of window segment is at least 100 amino acid residues in length (e.g., at least 150, 200, 300, or 500 residues in length).
52. The method of any one of the preceding claims, wherein the modeling, classifying or predicting is used in a drug development application, an antibody engineering application, a vaccine development application, as part of a diagnostic tool, for development of personalized medicine, for generation of a synthetic protein library, and / or to design an light-responsive protein sequence.
53. A non-transitory computer readable medium storing instructions that, when executed by a processor of a computer system, cause the computer system to: perform the method of any one of claims 1-52.
54. A computer system comprising: a processor, and a non-transitory computer readable storage medium, storing instructions that, when executed by a processor of a computer system, cause the computer system to: perform the method of any one of claims 1-52.
Citation Information
Patent Citations
Biological sequence fingerprints
US20190130064A1
Method, apparatus, device and storage medium for predicting protein binding site
US20190156915A1