Spatiotemporal determination of polypeptide structure
Patent Information
- Application Number
- JP2023571791
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-21
- Filing Date
- 2022-05-19
- Publication Date
- 2025-05-27
AI Technical Summary
Existing methods for predicting polypeptide structure, such as X-ray crystallography, are limited in accurately determining structures of poorly expressed or poorly folded proteins, hindering drug discovery for related diseases.
A method combining molecular dynamics simulations, machine learning algorithms, and data encoding to generate polypeptide structures, utilizing residue-specific and pairwise features, and incorporating known structural data from techniques like X-ray crystallography and NMR to predict structures.
Enables accurate prediction of polypeptide structures, including epitope mapping, allowing for the development of therapeutic agents that bind to dynamic surfaces of interest, overcoming limitations of static structural techniques.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] cross reference This application claims the benefit of European Application No. 21382464.2, filed May 21, 2021, which is incorporated by reference in its entirety.
[0002] Sequence Listing This application contains a Sequence Listing that has been submitted electronically in ASCII format and is incorporated by reference in its entirety. The ASCII copy, created on May 13, 2022, is named 199589-701601_ST25.txt and is 7,000 bytes in size. [Background technology]
[0003] background Accurate prediction of polypeptide structure has the potential to unlock the druggability of proteins involved in various diseases. Although techniques such as X-ray crystallography can be used to elucidate polypeptide structures, such techniques are hindered in poorly expressed or poorly folded proteins. Therefore, predicting structures using in silico techniques holds promise in unlocking the therapeutic potential of such proteins. Summary of the Invention [Means for solving the problem]
[0004] overview Provided herein is a method of in silico polypeptide structure generation, comprising: (a) performing a molecular dynamics (MD) simulation of a polypeptide to generate output data as a function of time, the output data including tertiary structure conformational information of the polypeptide; (b) encoding the output data into a function to generate a vector map, the vector map including (i) at least one residue-specific feature obtained from the MD simulation for amino acids in the polypeptide, and (ii) at least one pairwise feature obtained from the MD simulation for at least two amino acids in the polypeptide; and (c) applying a machine learning algorithm to the vector map to generate a predicted polypeptide structure based on the at least one residue-specific feature and the at least one pairwise feature. In some embodiments, the vector map includes a D-dimensional array, where D is the number of (i) residue-specific features and (ii) pairwise features. In some embodiments, the machine learning algorithm is an unsupervised algorithm. In some embodiments, the machine learning algorithm is a supervised algorithm. In some embodiments, known structural data can be used as input prior to performing the MD simulation. Such structural data may be obtained, for example, from static structures generated through X-ray crystallography, or from dynamic structures generated, for example, from NMR. In some embodiments, the residue-specific properties include Coulombic energy, van der Waals energy, residue labels, GRAVY scores, or any combination thereof. In some embodiments, the pairwise properties include Coulombic energy between at least two amino acids, van der Waals energy between at least two amino acids, distance between at least two amino acids, or any combination thereof. In some embodiments, the function is a continuous-time dynamic graph function. In some embodiments, the function is a discrete-time dynamic graph function. In some embodiments, the MD simulation includes replica exchange molecular dynamics.In some embodiments, the MD simulation comprises Monte Carlo dynamics. In some embodiments, the encoding comprises dynamic residue embedding. In some embodiments, the method further comprises generating a second function derived from the function, the second function comprising a static protein embedding based on the dynamic residue embedding. In some embodiments, the method further comprises encoding data from a crystal structure into the function. In some embodiments, the method further comprises attributing the predicted structure to a database. In some embodiments, the method further comprises linking the predicted structure to a disease state in the database. In some embodiments, the method further comprises selecting an intervention based on the predicted structure and the disease state. A system comprising a computer readable memory is also disclosed. In some embodiments, the computer readable memory comprises instructions for performing the method of in silico polypeptide structure generation described herein.
[0005] Provided herein is a method for generating epitope structures, comprising: (a) preparing a polypeptide sequence; (b) calculating an index score for a plurality of epitope structures in the polypeptide sequence, wherein the index score is calculated based on at least two of the structural salience parameters, disorder parameters, or conservation parameters of the epitopes, (i) the conservation parameters are calculated based on the conservation of at least two amino acid residues in a multiple sequence alignment comprising the target polypeptide, (ii) the disorder parameters and the structural salience parameters are obtained from a molecular dynamics (MD) simulation of a homology model comprising an aggregated structure of a homolog of the target polypeptide, and (iii) the index score is proportional to the structural salience parameters and the conservation parameters, and inversely proportional to the disorder parameters; (c) ranking the index scores and selecting the epitope structure with the highest index score from the plurality of epitope structures. In some embodiments, the method further comprises generating a paratope structure that is predicted to specifically bind to the epitope structure. In some embodiments, the method further comprises generating a therapeutic comprising said paratope structure. In some embodiments, the therapeutic is a small molecule. In some embodiments, the therapeutic is a polypeptide. In some embodiments, the polypeptide is an antibody. In some embodiments, the polypeptide is a nanobody. In some embodiments, the molecular dynamics simulation is a replicate exchange molecular dynamics simulation. In some embodiments, the structural prominence parameter is determined by the solvent accessible surface area of exposed amino acids in the target polypeptide. In some embodiments, the structural prominence parameter is determined by an atomic volume map of the target polypeptide. In some embodiments, the disorder parameter is determined by the root mean square fluctuation of alpha carbons in the backbone of the target polypeptide.In some embodiments, the disorder parameter is determined by the N-H bond order in the backbone of the target polypeptide. In some embodiments, the method further comprises generating a free energy surface representation of the target polypeptide based on a homology model, thereby determining a represented conformation of the target polypeptide at a free energy minimum. In some embodiments, the method further comprises bundling the represented conformations based on the magnitude of the representation at a given free energy minimum. In some embodiments, the method further comprises, prior to calculating the index score, generating a graph network including graph nodes and graph edges, the graph nodes including alpha carbons of the polypeptide, and the graph edges including interactions between at least two alpha carbon atoms in the backbone of the polypeptide. In some embodiments, the method further comprises applying a clustering algorithm to the graph network. In some embodiments, the clustering algorithm is selected from the group consisting of K-means clustering, t-distribution stochastic neighbor embedding, and any combination thereof. In some embodiments, the method further comprises applying empirical data to the index score. In some embodiments, the empirical data is an IC of binding of an antibody to an epitope of the target polynucleotide. 50 In some embodiments, the method further comprises a homology model that is a solvation model of the target polypeptide. Also disclosed is a system comprising a computer readable memory. In some embodiments, the computer readable memory comprises instructions for performing the method of generating epitope structures described herein.
[0006] Also disclosed herein are polypeptides comprising a paratopic structure, said paratopic structure being obtainable by the methods of generating an epitopic structure as described herein. [Brief description of the drawings]
[0007] The novel features of the illustrative embodiments are set forth with particularity in the appended claims. A better understanding of the features and advantages will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the disclosed systems and methods are utilized, and the accompanying drawings.
[0008] [Figure 1] FIG. 1 depicts an exemplary workflow for generating predicted polypeptide structures and protein therapeutics in silico consistent with embodiments described herein.
[0009] [Diagram 2] FIG. 2 depicts a schematic diagram of using the methods described herein to evaluate the druggability of binding sites and generate therapeutic polypeptides.
[0010] [Diagram 3] FIG. 3 shows an illustration of the spatio-temporal graphing of the individual graph functions as a function of time.
[0011] [Figure 4] FIG. 4 depicts interactions between pairs of polypeptides.
[0012] [Diagram 5] Figure 5 shows a cartoon representation of the conformation of a single polypeptide, where the underlying CA atoms are represented as spheres and secondary structures such as alpha helices are present.
[0013] [Figure 6] 6 shows an exemplary graphical representation of a single polypeptide conformation, where nodes represent CA atoms and edges represent interactions between adjacent CA atoms.
[0014] [Figure 7A]Figures 7A, 7B and 7C show exemplary graph functions as a function of time. Figure 7A is a snapshot of the evolution of a pairwise CA graph taken at time t = 20 ns (time frame = 100), where each node is colored according to the type of residue (i.e., node label), the size of each node is proportional to the degree of association, and the width of each edge is proportional to the associated weight. Figure 7B is a snapshot of the evolution of a pairwise CA graph taken at time t = 50 ns (time frame = 500), where each node is colored according to the type of residue (i.e., node label), the size of each node is proportional to the degree of association, and the width of each edge is proportional to the associated weight. Figure 7C summarizes the transition from t = 0 to t = 20 ns in a schematic manner. [Figure 7B] Figures 7A, 7B and 7C show exemplary graph functions as a function of time. Figure 7A is a snapshot of the evolution of a pairwise CA graph taken at time t = 20 ns (time frame = 100), where each node is colored according to the type of residue (i.e., node label), the size of each node is proportional to the degree of association, and the width of each edge is proportional to the associated weight. Figure 7B is a snapshot of the evolution of a pairwise CA graph taken at time t = 50 ns (time frame = 500), where each node is colored according to the type of residue (i.e., node label), the size of each node is proportional to the degree of association, and the width of each edge is proportional to the associated weight. Figure 7C summarizes the transition from t = 0 to t = 20 ns in a schematic manner. [Figure 7C]Figures 7A, 7B and 7C show exemplary graph functions as a function of time. Figure 7A is a snapshot of the evolution of a pairwise CA graph taken at time t = 20 ns (time frame = 100), where each node is colored according to the type of residue (i.e., node label), the size of each node is proportional to the degree of association, and the width of each edge is proportional to the associated weight. Figure 7B is a snapshot of the evolution of a pairwise CA graph taken at time t = 50 ns (time frame = 500), where each node is colored according to the type of residue (i.e., node label), the size of each node is proportional to the degree of association, and the width of each edge is proportional to the associated weight. Figure 7C summarizes the transition from t = 0 to t = 20 ns in a schematic manner.
[0015] [Figure 8] 8 depicts a surface representation of an exemplary polypeptide onto which information generated from the Druggability Index has been grafted. The shaded surface indicates potential epitopes generated using the Druggability Index.
[0016] [Figure 9] Figure 9 represents an exemplary output from the druggability index calculation procedure for several natural and non-natural variants of a target protein. Grey shading represents potentially druggable sites. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0017] Detailed Description For example, disclosed herein is a method for generating polypeptide structures in silico using big data generated from molecular dynamics simulations as a function of time. An exemplary workflow for generating predicted polypeptide structures using the methods described herein is illustrated in FIG. 1. Such data can be processed using the machine learning algorithms described herein to generate more comprehensive predicted polypeptide structures compared to using existing methods of structure prediction. By varying the polypeptide structure as a function of time and sampling rare conformations along the potential energy well, the predicted structure can more closely match the dynamics present in the polypeptide when present in its natural environment. Using this method, binding surfaces and epitopes present through dynamic movements of residues separated by significant sequence space can be accurately mapped, which may enable the generation of robust therapeutics that can interact with these epitopes. FIG. 2 illustrates a schematic for evaluating the druggability of predicted epitope sites in a polypeptide and producing protein therapeutics according to the methods described herein.
[0018] definition The section headings used herein may be for organizational purposes and should not be construed as limiting the subject matter described. In some instances, the section headings may not be construed as limiting the subject matter described.
[0019] As used in this specification and claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. For example, the term "a polypeptide" includes a plurality of polypeptides, including a mixture of polypeptides.
[0020] As used herein, the term "about" or "approximately" when referring to a measurable value such as an amount or concentration is meant to encompass a variation of + / -20%, including + / - 10%, 5%, 1%, 0.5% or even 0.1% of the specified amount, unless otherwise specified.
[0021] As used herein, the term "comprising" is intended to mean that the compositions and methods include the recited elements but do not exclude other elements. When used to define compositions and methods, "consisting essentially of" is intended to mean excluding other elements that have some essential importance to the combination for the intended use. Thus, a composition consisting essentially of elements as defined herein does not exclude trace contaminants from the isolation and purification methods and pharma- ceutically acceptable carriers such as phosphate buffered saline, preservatives, etc. "Consisting of" is intended to mean excluding more than trace amounts of other ingredients and substantial method steps for administering the compositions of the present disclosure. The embodiments defined by each of these transition terms are within the scope of the present disclosure.
[0022] The terms "subject," "host," "individual," and "patient" interchangeably refer to an animal, typically a mammal. Any suitable mammal can be treated with the compositions described herein. Non-limiting examples of mammals include humans, non-human primates (e.g., apes, gibbons, chimpanzees, orangutans, monkeys, macaques, etc.), domestic animals (e.g., dogs and cats), farm animals (e.g., horses, cows, goats, sheep, pigs), and laboratory animals (e.g., mice, rats, rabbits, guinea pigs). In some embodiments, the mammal can be a human. The mammal can be of any age or at any stage of development (e.g., adult, teen, child, infant, or intrauterine mammal). The mammal can be male or female. The mammal can be a pregnant female. In some embodiments, the subject can be a human. In some cases, the human may be greater than about 1 day to about 10 months old, about 9 months to about 24 months old, about 1 year to about 8 years old, about 5 years to about 25 years old, about 20 years to about 50 years old, about 1 year to about 130 years old, or about 30 years to about 100 years old. The human may be greater than about: 1, 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, or 120 years old. The human may be less than about: 1, 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, 100, 110, 120, or 130 years old.
[0023] The terms "treat", "treatment", and the like, may be used herein to mean obtaining a desired pharmacological effect, physiological effect, or any combination thereof. In some instances, treatment may reverse adverse effects resulting from a disease or disorder. In some instances, treatment may stabilize a disease or disorder. In some instances, treatment may slow the progression of a disease or disorder. In some instances, treatment may cause regression of a disease or disorder. In some instances, treatment may prevent the development of a disease or disorder. In some embodiments, the effect of treatment may be measured. In some instances, measurements may be compared before and after administration of the composition. For example, a subject may have pre-treatment medical images that are compared to post-treatment images to show regression of cancer. In some instances, a subject may have improved blood test results after treatment that are compared to blood tests before treatment. In some instances, measurements may be compared to a standard.
[0024] The terms "protein", "peptide" and "polypeptide" refer interchangeably and in their broadest sense to a compound of two or more subunit amino acids, amino acid analogs or peptidomimetics. The subunits may be linked by peptide bonds. In another embodiment, the subunits may be linked by other bonds, such as esters, ethers, etc. A protein or peptide may contain at least two amino acids, and the maximum number of amino acids that may make up the sequence of a protein or peptide may be unlimited. As used herein, the term "amino acid" may refer to natural, unnatural or synthetic amino acids. Natural, unnatural or synthetic amino acids may include glycine and both D and L optical isomers, amino acid analogs and peptidomimetics. As used herein, the term "fusion protein" may refer to a protein composed of domains from more than one naturally occurring or recombinantly produced protein, generally each domain performing a different function. In this regard, a linker may refer to a protein fragment that can be used to link these domains together, to preserve the conformation of the fused protein domains as needed, and / or to prevent unfavorable interactions between the fused protein domains that may impair the respective functions of the fused protein domains.
[0025] "Homology" refers to the sequence similarity between two peptides or two nucleic acid molecules. Homology is determined by comparing the positions in each sequence that are aligned for comparison. When a position in the compared sequences is the same base or amino acid, the molecules are identical at that position. Sequence homology refers to the % identity of a sequence to a reference sequence. In practical terms, a homologous sequence has at least 50%, 60%, 70%, 80%, 85%, 90%, 92%, 95%, 96%, 97%, 98% or 99% identity to a reference sequence when aligned using a known computer program such as the Bestfit program. When using Bestfit or any other sequence alignment program to determine whether a particular sequence is, for example, 95% identical to a reference sequence, parameters can be set so that the percentage of identity can be calculated over the entire length of the reference sequence and gaps in sequence homology up to 5% of the entire reference sequence can be allowed. An "unrelated" sequence shares less than 40% identity, or less than 25% identity, with one of the sequences of the present disclosure.
[0026] The term "epitope" refers to a moiety or structure on a polypeptide to which a moiety (eg, a polypeptide immunoglobulin, an antibody, etc.) specifically binds.
[0027] The term "paratope" refers to the structure of a moiety (eg, a polypeptide immunoglobulin, an antibody, etc.) that specifically binds to an epitope.
[0028] The term "supervised learning" refers to deep learning training methods in which a machine is provided with data from a human source. The term "unsupervised learning" refers to deep learning training methods in which a machine is not provided with data from a human source.
[0029] The term "semi-supervised learning" refers to a deep learning training method in which a machine is provided with small amounts of data from a human source that is then compared to larger amounts of data from other sources available to the machine.
[0030] I. Generation of data from molecular dynamics simulations Disclosed herein is a method for predicting polypeptide structure using data input generated from molecular dynamics simulation. Molecular dynamics simulation can be performed in silico to model conformational and biophysical features of polypeptide structure. Molecular dynamics simulation can enable structural dynamics such that the secondary and tertiary structure of the polypeptide can change within the timeline of the simulation along allowed conformations. Generally, allowed conformations are conformations that correspond to local minima along various free energy wells. Thus, molecular dynamics simulation can be used to visualize and sample biologically relevant conformations that static structural techniques (e.g., X-ray crystallography) may not sample. Exemplary molecular dynamics simulations for inclusion in the methods described herein include, but are not limited to, classical dynamics, replica exchange molecular dynamics, metadynamics, Langevin dynamics, and Monte Carlo dynamics.
[0031] Provided herein is a method for modeling and predicting polypeptide structures using data generated from molecular dynamics simulations. As described herein, data generated from molecular dynamics simulations is used as input for machine learning to iterate among allowed and rare structure conformations to generate more robust and comprehensive predicted polypeptide structures. Such data can include residue-specific biophysical properties associated with a single residue in a molecular dynamics simulation and pairwise properties associated with a set of biophysical properties associated with interactions between at least two residues in a molecular dynamics simulation. Examples of residue-specific biophysical properties generated using molecular dynamics simulations include the total average hydropathy (GRAVY) score, residue identity or label, Coulombic energy, van der Waals energy, solvent accessible surface area (SASA), side chain order parameter (S2), and the like. Examples of pairwise biophysical properties generated using molecular dynamics simulations include the distance between given residues, Coulombic energy, van der Waals energy, fraction of native contacts (Q), and the like.
[0032] Such properties generated from molecular dynamics can be generated from a given conformation as a function of time. Thus, a data set of biophysical properties as a function of time from a set of polypeptide structures can be generated from molecular dynamics simulations and used as input to machine learning algorithms. This data is arranged in a graph format prior to embedding. Length
number
number
[0033] The function can be the conditional log-probability of a certain set of temporal random walks. These are random walks that preserve the time order or the temporal edges, i.e. along the path of such a walk, the timestamps of successive edges do not decrease. Moreover, such a function can be modeled using a pre-trained skip-gram model.
number
number
[0034] After generating the graphical representation described herein, the data is embedded for input into the machine learning algorithms described herein. In some examples, manifold learning techniques, such as t-SNE (t-SNE) can be used. Embeddings described herein can include dynamic residue embeddings and static protein embeddings.
[0035]
number
[0036]
number
[0037] II. Machine Learning Tensor representations generated from dynamic and static embeddings, respectively
number
number
[0038] In some embodiments, the polypeptide structure may be generated using unstructured computation, artificial intelligence, or deep learning. In some cases, unstructured computation may be used so that the computation can be performed iteratively. Furthermore, the computation of the polypeptide structure may rely on artificial intelligence or deep learning. For example, the methods described herein, such as random forest, may use deep learning to generate a Gini impurity score that can be used to define probes with improved predictive value.
[0039] In some embodiments, the methods of structure prediction described herein can use a combination of machine learning and computational intelligence techniques, such as deep neural networks, as well as supervised, semi-supervised, and unsupervised learning techniques. In some embodiments, the methods of structure prediction described herein use supervised algorithms (such as, by way of non-limiting example, linear regions, random forest classification, decision tree learning, ensemble learning, bootstrap aggregating, and the like). In some embodiments, the methods of structure prediction described herein use unsupervised algorithms (such as, by way of non-limiting example, clustering or correlation).
[0040] In some embodiments, the methods of structure prediction described herein may be configured to utilize one or more exemplary AI / machine learning techniques selected from, but not limited to, decision trees, boosting, support vector machines, neural networks, nearest neighbor algorithms, naive Bayes, bagging, random forests, etc. In some embodiments, and optionally in combination with any of the embodiments described above or below, the exemplary neural network technique may be one of, but not limited to, a feed-forward neural network, a radial basis function network, a recurrent neural network, a convolutional network (e.g., U-net), or other suitable networks. In some embodiments, and optionally in combination with any of the embodiments described above or below, an exemplary implementation of a neural network may be performed as follows: a. Define the neural network architecture / model; b. forwarding the input data to an exemplary neural network model; c. Incrementally training the example model; d. Determine the accuracy for a particular number of time steps; e. applying the example trained model to process newly received input data; f. Continue to train the example trained model on a predetermined periodic basis, as needed and in parallel.
[0041] In some embodiments, and optionally in combination with any of the embodiments described above or below, the exemplary trained neural network model may specify the neural network by at least a neural network topology, a set of activation functions, and connection weights. For example, the topology of the neural network may include the arrangement of nodes of the neural network and the connections between such nodes. In some embodiments, and optionally in combination with any of the embodiments described above or below, the exemplary trained neural network model may also be specified to include other parameters, including but not limited to bias values / functions and / or aggregation functions. For example, the activation functions of the nodes may be step functions, sine functions, continuous or piecewise linear functions, sigmoid functions, hyperbolic tangent functions, or other types of mathematical functions that represent thresholds at which the nodes are activated. In some embodiments, and optionally in combination with any of the embodiments described above or below, the exemplary aggregation functions may be mathematical functions that integrate (e.g., sum, product, etc.) the input signals to the nodes. In some embodiments, and optionally in combination with any of the embodiments described above or below, the output of the exemplary aggregation functions may be used as input to the exemplary activation functions. In some embodiments, and optionally in combination with any of the embodiments above or below, the bias may be a constant value or function that may be used by the aggregation function and / or activation function to make a node more likely or less likely to be activated.
[0042] In some embodiments, a machine learning model for structure prediction processes the biophysical properties encoded in the embedding by applying parameters of the machine learning model to generate a model output, which in some embodiments can be decoded to generate one or more numerical output values and / or vectors indicative of the polypeptide structure.
[0043] In some embodiments, the parameters of the machine learning model may be trained based on the known polypeptide structure. For example, biophysical properties may be paired with the target structure and / or measurements to form training pairs, such as past biophysical properties and observed structures that represent data points in the relationship between past biophysical properties and structures. In some embodiments, the biophysical properties may be provided to the machine learning model, e.g., encoded in an embedding, to generate data representative of the polypeptide structure. In some embodiments, the optimization problem associated with the machine learning model may then compare the polypeptide structure to known outputs of training pairs that include past biophysical properties to determine the error of the polypeptide structure. In some embodiments, the optimization problem may use a loss function, such as, for example, hinge loss, multi-class SVM loss, cross-entropy loss, negative log-likelihood, or other suitable classification loss function to determine the error of the polypeptide structure based on the known structure.
[0044] In some embodiments, the known output may be obtained after the machine learning model generates a prediction, such as in an online learning scenario. In such a scenario, the machine learning model may receive the biophysical properties and generate a model output vector to generate data representative of the polypeptide structure. The user may then provide feedback by modifying, adjusting, removing, and / or validating the predicted structure via a suitable feedback mechanism, such as, for example, a user interface device (e.g., a keyboard, a mouse, a touch screen, a user interface or other interface mechanism of a user device, or any suitable combination thereof). The feedback may be paired with the biophysical properties to form a training pair, and the optimization problem may use the feedback to determine the error of the polypeptide structure.
[0045] In some embodiments, based on the error, the optimization problem may update the parameters of the machine learning model using a suitable training algorithm, such as, for example, backpropagation for the predictive machine learning model. In some embodiments, the backpropagation may include any suitable minimization algorithm, such as a gradient method of a loss function on the weights of the predictive machine learning model. Examples of suitable gradient methods include, for example, stochastic gradient descent, batch gradient descent, mini-batch gradient descent, or other suitable gradient descent techniques. As a result, the optimization problem may update the parameters of the machine learning model based on the error of the predicted structure to train the machine learning model and model the correlation between the biophysical properties and the polypeptide structure to generate more accurate predictions of the structure based on the biophysical properties.
[0046] III. Generation of Polypeptide Therapeutics As described herein, data generated from molecular dynamics simulations using the machine learning algorithms described herein can be used to predict robust and comprehensive polypeptide structures. Knowledge of such structures can be used to effectively and accurately map the dynamic surface of a polypeptide of interest involved in a disease or condition. By accurately modeling the surface of a polypeptide as a function of time, novel therapeutics can be generated that can bind to and interact with epitopes of a polypeptide of interest. Thus, such therapeutics can be generated with paratope structures configured to bind to epitopes and are useful for treating a disease or condition. Figure 4 represents an illustration of predicted epitope and paratope structures using the methods described herein. Furthermore, by capturing the dynamic structure of a polypeptide using the methods described herein, rare biologically relevant conformations can be predicted that may not be present in static structures such as those generated by X-ray crystallography. Furthermore, iterations using machine learning using the methods described herein allow for robust simulations beyond the capabilities of molecular dynamics simulations alone, which allows for sampling of rare and short-lived (but biologically relevant) conformations that generate epitopes.
[0047] Any combination of data can be utilized as described above to generate predicted polypeptide structures using any of the above machine learning algorithms. Additionally, additional inputs can be used to provide additional information useful in elucidating biologically relevant epitope conformations. For example, evolutionary covariance between related or homologous polypeptides can be used to determine conservation between residues separated by significant distances in primary and secondary structures. Without wishing to be bound by theory, the methods described herein utilize evolutionary coupling between pairs of residues as input to determine whether they share biological functions (e.g., are present in the same binding epitope). With such inputs, dynamic modeling can be performed to determine whether such residues are present in dynamic structures with minimal entropy penalty. Thus, evolutionary coupling and dynamics / disorder parameters are balanced to harvest rare but biologically relevant conformations that give rise to such epitopes.
[0048] When evolutionary linkage is used, the method described herein includes generating a multiple sequence alignment to determine the homology between amino acid sequences. The identity between a reference sequence (query sequence, i.e., the sequence of the present disclosure) and a subject sequence, also called global sequence alignment, can be determined using the FASTDB computer program based on the algorithm of Brutlag et al. (Comp. App. Biosci. 6:237-245 (1990)). In some embodiments, the parameters using FASTDB amino acid alignment can include: scoring scheme=PAM (percentage of mutations allowed) 0, k-tuple=2, mismatch penalty=1, binding penalty=20, randomization group length=0, cutoff score=1, window size=sequence length, gap penalty=5, gap size penalty=0.05, window size=500 or the length of the subject sequence, whichever is shorter. According to this embodiment, if the subject sequence is shorter than the query sequence due to N- or C-terminal deletions, but not due to internal deletions, the FASTDB program can perform manual corrections on the results to account for the fact that the N- and C-terminal truncations of the subject sequence are not considered when calculating the overall percent identity. For subject sequences that are N- and C-terminally truncated relative to the query sequence, the percent identity is corrected by calculating the number of residues of the query sequence that are present on the N- and C-terminal sides of the subject sequence that are not matched / aligned with the corresponding subject residues as a percentage of the total bases of the query sequence. The determination of whether a residue is matched / aligned can be determined by the results of FASTDB sequence alignment. This percentage is then subtracted from the percent identity calculated by the FASTDB program using the specified parameters to arrive at a final percent identity score. This final percent identity score can be used in this embodiment. In some cases, for the purpose of manually adjusting the percent identity score, only residues to the N- and C-termini of the subject sequence that are not matched / aligned with the query sequence can be considered.That is, only the query residue positions outside the farthest N- and C-terminal residues of the subject sequence can be considered for this manual correction. For example, to determine percent identity, a 90-residue subject sequence can be aligned with a 100-residue query sequence. The deletion occurs at the N-terminus of the subject sequence, and therefore the FASTDB alignment does not show the first 10 residues matched / aligned at the N-terminus. The 10 unpaired residues represent 10% of the sequence (number of unmatched N- and C-terminal residues / total number of residues in the query sequence), so 10% can be subtracted from the percent identity score calculated by the FASTDB program. If the remaining 90 residues are perfectly matched, the final percent identity can be 90%. In another example, a 90-residue subject sequence can be compared with a 100-residue query sequence. This time, the deletion can be an internal deletion, so there can be no residues at the N- or C-terminus of the subject sequence that cannot be matched / aligned with the query. In this case, the percent identity calculated by FASTDB cannot be manually corrected, and again, only residue positions outside the N- and C-termini of the subject sequence that cannot be matched / aligned with the query sequence as displayed in the FASTDB alignment can be manually corrected.
[0049] In some instances, known structures can be used in conjunction with sequences as input for the methods described herein. For example, structures stored in protein structure databases can be accessed and used as input for determining novel epitopes. In some instances, empirical structural data can be used as input. For example, the static structure of the target polypeptide obtained by X-ray crystallography can be used as input. Also, dynamic structures obtained using techniques such as circular dichroism and NMR (e.g., 2D NMR, 3D NMR, solid-state NMR, etc.) can be used as input.
[0050] An exemplary workflow for predicting epitope structure is shown below. · The protein sequence (or list) is fed into the algorithm. · Multiple sequence alignments (MSA) are performed to assess the evolutionary coupling (EC) between pairs of amino acid residues in the analyzed sequences. Evolutionary coupling reports on the probability that any pair of amino acid residues in a given sequence has evolved in a coupled fashion and is therefore evolutionarily significant and likely has a biological role. · A protein homology 3D model similar to the X-ray crystallography or NMR structure (or a model from a protein sequence list) is calculated. A solvated 3D model of the protein (using SPC or TIP3 water models) is generated and the remaining unneutralized charges are neutralized by the addition of monovalent positive (Na+) and negative (CL-) ions such that the net charge (sum of all charges) of the simulated system is equal to 0. The solvated system is subjected to replica exchange molecular dynamics (REMD) simulations. In the REMD simulations, Any number of simulation replicas (>2) are started. The number itself depends on the size of the system and scales with the number of atoms, e.g. a 25000 atom system might require 25 replicas running for 500 nanoseconds each. b.All replicas receive a copy of the original force field assigned to the simulation, for which the torsion angle potential, dihedral angle potential and selected non-bonded terms are linearly scaled by a factor proportional to the number of replicas. The first replica in the set receives all forces, while the last replica is exposed to a modified force field scaled by a significant factor equal to 0.5. When performing REMD, the free energy surface (FES) of the conformational protein space is reconstructed such that the most representative structures belonging to different free energy wells can be identified and bundled together as a 3D protein ensemble. The newly constructed 3D protein ensemble is then subjected to a subdomain identification procedure that evaluates the geometric and spatiotemporal fit of the target protein fragments using the following metrics: a. Structural disorder of individual protein fragments resulting from: i. Protein backbone HN bond order parameter (S2). ii. The root mean square fluctuation (RMSD) of the CA atoms in the protein backbone. b. Structural prominence resulting from: i. Solvent accessible surface exposed amino acids (SASA). ii. Atomic Volume Map (AVM). A graph network is constructed in which every CA atom in the original 3D protein molecule is represented by a node, whereas its interactions with neighboring CA backbone atoms are represented by graph edges. In this representation, A graph node is assigned. i.RMSF and S2. ii.The sum of the intra-residue interaction energies calculated from the REMD protocol. iii. Combined SASA and AVM. b. Graph edges are assigned. i. Intra-residue interaction energies estimated from the REMD protocol. ii. The EC probabilities obtained from step 2 of the algorithm. A graph node clustering algorithm is applied to the graph from step 8 so that clusters of amino acid residues that share similar spatiotemporal (dynamics) and structural prominences can be identified and flagged as subdomains. The clustering algorithm: aK-means clustering. bt-Distributed Stochastic Neighbor Embedding (t-SNE) c. and equivalents. For every clustered class, a composite druggability index (DI) is devised and calculated. The score is the sum of structural prominence from SASA and AVM, the sum of evolutionary conservation from EC divided by the sum of RMSF and the inverse of S2. High scores indicate domains that are protruding and solvent exposed but undergo small structural transitions throughout the molecular dynamics. Furthermore, the addition of the EC component allows prioritization of sites with strongly conserved evolutionary features. Low scores represent domains with poor structural prominence, high dynamics and, importantly, low evolutionary conservation. DI score is IC 50 It can be further strengthened by the addition of manually curated data on antibody-epitope interactions, such as binding values, which can be obtained from personally performed experiments or through automated literature searches using natural language processing (NLP) methods.
[0051] Following prediction of epitope surfaces using the methods described herein, methods are provided herein in which protein therapeutics are designed in silico to contain paratope structures configured to bind and interact with the predicted epitope structures. Protein therapeutics can be synthesized using standard FMOC protein synthesis or other standard peptide synthesis techniques used in the art. Alternatively, some protein therapeutics can be expressed in microorganisms such as Escherichia coli from DNA vectors. In such embodiments, a polynucleotide sequence encoding a polypeptide of interest is subcloned into an expression vector for overexpression in the microorganism. Successful subcloning of the polynucleotide sequences may be confirmed by sequencing using commercially readily available methods including, but not limited to, capillary sequencing, bisulfite-free sequencing, bisulfite sequencing, TET-assisted bisulfite (TAB) sequencing, ACE sequencing, high-throughput sequencing, Maxam-Gilbert sequencing, massively parallel signature sequencing, Polony sequencing, 454 Pyrosequencing, Sanger sequencing, Illumina sequencing, SOLiD sequencing, Ion Torrent semiconductor sequencing, DNA nanoball sequencing, Heliscope single molecule sequencing, single molecule real-time (SMRT) sequencing, nanopore sequencing, shotgun sequencing, RNA sequencing, Enigma sequencing or any combination of these.
[0052] Such protein therapeutics contain high potency of binding to predicted epitopes based on the robust structural sampling methods provided herein, and thus, such therapeutic polypeptides exhibit biologically relevant activity against the protein of interest when administered to a subject.
[0053] Moreover, such protein therapeutics are expected to have high specificity and selectivity for the protein of interest.In some cases, the protein of interest can have at least about 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% specificity for the target of interest, for example, as determined by in vitro competitive assay. In some cases, a protein of interest can have at least about 60%, 61%, 62%, 63%, 64%, 65%, 66%, 67%, 68%, 69%, 70%, 71%, 72%, 73%, 74%, 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or 100% selectivity for a target of interest among other proteins as determined, for example, in an in vitro competition assay.
[0054] The therapeutic peptides described herein can be delivered at any dose required to achieve a biological outcome (e.g., treatment of a disease or condition in a subject). Given the high degree of specificity and potency of therapeutics made using the methods described herein, the dose of therapeutic required to achieve a biological outcome is lower than therapeutics made against the same target using comparable methods (i.e., molecular dynamics, mutagenesis, or using static structures alone).
[0055] system
[0056] Also disclosed herein is a system for carrying out the methods described herein. The system can include a computer readable memory that stores instructions for carrying out the methods described herein. For example, the computer readable memory can include instructions for in silico determination of polypeptide structure as described herein. In some embodiments, the computer readable memory can include instructions for epitope determination as described herein.
[0057] The system may further comprise a computer system utilizing the computer readable memory. The computer system may include a processor operatively coupled to the computer readable memory and may be configured to execute instructions to perform the methods described herein. The computer system may further include user input and output means such as a keyboard, monitor and mouse.
[0058] The systems described herein can be configured to access a database. For example, the systems can be configured to access a local or online (e.g., cloud) database, such as a protein structure database, a protein sequence database, a homology database, a nucleic acid sequence database, etc.
[0059] Upon performing the methods described herein, the system may further include data obtained by performing the methods described herein. For example, the system upon performing the methods described herein may include a druggability index score for determining novel epitopes. Example 5 herein provides an exemplary output of such data that may be stored on the system after performing the methods described herein. The system may include structural information obtained from MD simulations described herein. Additionally, the system may include empirical structural data, such as protein structures obtained from NMR or X-ray crystallography. The system may include optimized polypeptide structures obtained using the in silico methods described herein.
[0060] Such systems may include storage means for storing or transferring data obtained by the methods described herein, in some examples, the systems may include means for transmitting data obtained by the methods described herein to an external database (e.g., a local database or an online database). EXAMPLES
[0061] In order to obtain a better understanding of the present disclosure and its many advantages, the following examples are given by way of illustration without limiting the scope of the disclosure.
[0062] Example 1: Generation of Polypeptide Structures Using Molecular Dynamics Simulations
[0063] To model the conformational dynamics of a polypeptide sequence, an example polypeptide sequence is input into a molecular dynamics simulation. The model is solvated using the TIP3 water model and monovalent Na + and Cl -Ions are used to neutralize the charge. FIG. 5 represents a ribbon representation of a polypeptide structure generated using molecular dynamics at a single time point. The conformations of each structure at a given time point are arranged into a graph function. FIG. 6 represents an exemplary graph function generated from molecular dynamics for a single conformation at a single time point. In this representation, the nodes represent individual CA atoms, while the edges represent pairwise interactions between residues. FIG. 7A and FIG. 7B show exemplary continuous-time dynamic graphs. FIG. 7A is a snapshot of the evolution of a pairwise CA graph taken at time t=20 ns (time window=100), where each node is colored according to the type of residue (i.e., node label), the size of each node is proportional to the degree of association (i.e., the number of edges or neighbors to which the node is connected), and the width of each edge is proportional to the associated weight (i.e., the magnitude of the pairwise property). Figure 7B is a snapshot of the evolution of the pairwise CA graph taken at time t = 50 ns (time window = 500), where each node is colored according to the residue type (i.e., node label), the size of each node is proportional to the degree of association, and the width of each edge is proportional to the associated weight. Figure 7C summarizes the transition from t = 0 to t = 20 ns in a schematic manner.
[0064] Example 2: Encoding graph functions for machine learning implementations
[0065] The data generated from the molecular dynamics simulation performed in Example 1 and converted into a graph format is encoded into a vector table to implement a machine learning algorithm.
number
[0066] Example 3: Optimizing dynamic graph display using machine learning
[0067] Manifold learning techniques, including t-distribution stochastic neighbor embedding, are applied to the encoded data generated in Example 2 to generate an optimized dynamic graph representation based on the encoded data. An unsupervised learning algorithm is used to iteratively generate the dynamic graph representation to generate a predicted polypeptide structure.
[0068] Example 4: Prediction of epitope binding surfaces
[0069] Evolutionary covariance is determined in silico by performing multiple sequence alignment of protein homologues. Focusing on pairwise conservation of residues, an evolutionary coupling report is calculated for any pair of respective amino acids based on the probability that the two amino acids have evolved in a coupled manner. Replica exchange molecular dynamics simulation is performed to generate data from MD simulation as described in Example 1. Structural disorder parameters are calculated based on the RMSD fluctuation of CA atoms in MD simulation, while structural prominence parameters are calculated based on the solvent accessible surface area of exposed amino acids and atomic volume mapping of polypeptide.
[0070] A graph function is generated as described above for Example 2 and embedded into a vector containing the structural disorder parameters, the structural prominence parameters and data from the evolutionary binding reports. A clustering algorithm is run using machine learning to generate an optimized polypeptide structure as described above for Example 3.
[0071] Clustered residues sharing similar structural disorder parameters, structural prominence parameters and evolutionary coupling are grouped together and a composite druggability index score is calculated for the clustered residues. The druggability index score is proportional to the structural prominence parameters and evolutionary coupling and inversely proportional to the structural disorder parameters. The druggability index can be mapped onto the predicted structure to identify putative epitopes. FIG. 8 represents a surface view of an exemplary polypeptide onto which information generated from the druggability index has been grafted. The shaded surface indicates potential epitopes generated using the druggability index. FIG. 9 represents an exemplary output from the druggability index calculation procedure for several natural and non-natural variants of a target protein. The grey shading represents potentially druggable sites.
[0072] Example 5 - Druggability Index Calculation for Exemplary α-Synuclein Epitopes
[0073] To illustrate the use of disorder parameters for elucidation of novel epitopes, α-synuclein variants were collected for epitope determination. In this study, the effect of mutations at H50 on the druggability of novel epitopes was investigated. H50 is a residue that, when mutated, can lead to aggregation of α-synuclein. Therefore, being able to design therapeutics that target novel epitopes in α-synuclein variants with mutations at H50 provides significant therapeutic implications.
[0074] As described above for Example 4, structural disorder and structural prominence parameters were calculated based on the RMSD fluctuations of CA atoms in the MD simulations of each α-synuclein variant. Table 2 represents the various parameters calculated from the MD simulations. In the binary disorder prediction, a value of 1 indicates a disordered residue, and a value of 0 indicates an ordered residue. In the disorder propensity, a normalized magnitude of disorder is provided, with higher values representing a higher likelihood that a given disorder is disordered. Then, a binary druggability index prediction is calculated, with a value of 1 indicating a disordered epitope-binding residue, a value of 0 indicating a disordered residue other than the epitope-binding residue, and X indicating a residue that is irrelevant to epitope binding. Finally, a normalized epitope binding propensity is determined based on the disorder propensity, with higher values representing a higher likelihood that a given residue will provide a druggable epitope, while X indicates a residue that is irrelevant to epitope binding.
[0075] As shown in Table 2 below, disorder parameters can be used to determine epitope binding propensity per residue for each variant. Notably, residues along the C-terminus with high relative disorder propensity are predicted to have high epitope binding propensity across all variants tested (values above 0.9 are underlined in Table 2), and thus would be attractive targets for therapeutic targeting. Furthermore, although mutations in H50 itself (residues shown in bold in Table 2) are believed to play a role in the aggregation propensity of α-synuclein, the H50Y variant provided herein (SEQ ID NO: 9) showed a dramatic reduction in disorder along the H50 aggregation surface compared to variants with wild-type H50 residues. Thus, the H50 aggregation surface does not appear to be a druggable epitope for variants of α-synuclein with H50Y mutations.
[0076] Although exemplary embodiments have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes and substitutions will occur to those skilled in the art. It should be understood that various alternatives to the embodiments described herein may be used. The following claims define the scope of the disclosure, and it is intended that methods and structures within the scope of these claims and their equivalents be covered by the scope of the invention.
[0077] [Table 2-1] [Table 2-2] [Table 2-3] [Table 2-4] [Table 2-5] [Table 2-6] [Table 2-7] [Table 2-8] [Table 2-9] [Table 2-10]
Claims
1. A method for generating an epitope structure, comprising: a) preparing a polypeptide sequence; and b) calculating an exponential score for a plurality of epitope structures in the polypeptide sequence, wherein the exponential score is calculated based on at least two of a structural protrusion parameter, a disorder parameter, or a conservation parameter of the epitope, wherein (i) the conservation parameter is calculated based on the conservation of at least two amino acid residues in a multiple sequence alignment including the polypeptide sequence, wherein (ii) the disorder parameter and the structural protrusion parameter are obtained from a molecular dynamics (MD) simulation of a homology model including the collective structure of homologs of the polypeptide sequence, and wherein (iii) the exponential score is proportional to the structural protrusion parameter and the conservation parameter and inversely proportional to the disorder parameter; and c) ranking the exponential scores and selecting an epitope structure having the highest exponential score among the plurality of epitope structures. A method comprising the above steps.
2. The method according to claim 1, further comprising generating a paratope structure predicted to specifically bind to the epitope structure.
3. The method according to claim 2, further comprising producing a therapeutic agent comprising the paratope structure, wherein the therapeutic agent is produced by chemical synthesis or expression in a host cell.
4. The method according to claim 3, wherein the therapeutic agent is a small molecule or a polypeptide.
5. The method according to claim 4, wherein the polypeptide is an antibody and optionally a nanobody.
6. The method according to claim 1, wherein the molecular dynamics simulation is a replica exchange molecular dynamics simulation.
7. The method according to claim 1, wherein the structural protrusion parameter is determined by the solvent accessible surface area of exposed amino acids in the polypeptide sequence or an atomic volume map of the polypeptide sequence.
8. The method according to claim 1, wherein the disorder parameter is determined by the root mean square fluctuation of α-carbons in the backbone of the polypeptide sequence.
9. The method according to claim 1, wherein the disorder parameter is determined by the N-H bond order in the backbone of the polypeptide sequence. The method according to claim 1, further comprising the step of generating a free energy surface representation of the polypeptide sequence based on the aggregation structure of the homologs, thereby determining the displayed three-dimensional structure of the polypeptide sequence at the free energy minimum value.
11. The method according to claim 10, further comprising the step of bundling the displayed three-dimensional structure based on the size of the display at a given free energy minimum value.
12. The method according to claim 1, further comprising the step of generating a graph network including graph nodes and graph edges before calculating the exponential score, wherein the graph nodes include the alpha carbons of the polypeptide, and the graph edges include the interactions between at least two alpha carbon atoms in the backbone of the polypeptide.
13. The method according to claim 12, further comprising the step of applying a clustering algorithm to the graph network.
14. The method according to claim 13, wherein the clustering algorithm is selected from the group consisting of K-means clustering, t-distributed stochastic neighbor embedding, and any combination thereof.
15. The method according to claim 1, further comprising the step of applying empirical data to the exponential score.
16. The method according to claim 15, wherein the empirical data comprises the IC 50 of the binding of an antibody to the epitope of the polypeptide sequence.
17. The method according to claim 1, wherein the aggregation structure of the homologs of the polypeptide is generated from a solvation model of the polypeptide.
18. The method according to claim 1, further comprising providing the structure of the polypeptide.
19. The method according to claim 18, wherein the structure is an NMR structure.
20. A polypeptide comprising a paratope structure, wherein the paratope structure is obtained by the method according to claim 2.