MHC-peptide binding free energy calculation method and MHC-peptide screening method
By calculating the MHC-peptide binding free energy, utilizing anchor amino acid template selection and molecular dynamics simulation, the problems of data dependence and bias in existing technologies are resolved, achieving high-precision peptide affinity prediction and screening, and supporting the development of immunotherapy.
Patent Information
- Application Number
- CN202510371832.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-09-26
AI Technical Summary
Existing MHC-peptide affinity calculation methods rely on large amounts of training data, suffer from data bias and model deviation, and are difficult to make accurate quantitative predictions, especially in the prediction of low-affinity or special peptide segments.
The MHC-peptide binding free energy calculation method is used to obtain known MHC-peptide data, utilize template selection of anchor amino acids and homology modeling, combined with molecular dynamics simulation, calculate the binding free energy, and screen out high-affinity peptides.
It improves the accuracy and reliability of peptide affinity prediction, avoids data offset and model bias, and can accurately screen out strong binding peptides to support the design and development of immunotherapy.
Smart Images

Figure CN120708731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biomedicine technology, and in particular to a method for calculating MHC-peptide binding free energy and a method for screening MHC-peptides. Background Art
[0002] Currently, the primary method for calculating the affinity between antigenic peptides and MHC (major histocompatibility complex) is protein sequence-based machine learning techniques. These methods employ data modeling based on large-scale training datasets, using MHC protein and peptide sequences as model inputs and IC50 or KD values measured using competitive binding assays or wet-label assays such as SPR and BLI as model outputs. Alternatively, affinity values are classified into strong, moderate, weak, and non-binding categories using thresholds such as 50 nmol or 500 nmol. Commonly used machine learning prediction methods include software such as NetMHC, NetMHCpan, and MHCflurry. However, these methods share several drawbacks: 1. These machine learning or artificial intelligence methods rely on extensive training data. When training data is insufficient or biased, their predictions suffer from poor generalization performance. 2. These methods are susceptible to the poor accuracy of low-affinity data within the experimental data, resulting in poor affinity prediction performance. Therefore, better performance is often achieved by converting regression problems into classification problems during the training or output stages. For example, the NetMHC series software ranks the affinity results predicted by the model in a reference data set on one side, and recommends that users use the ranking values (Rank 0.5% and 2%) as screening indicators for strong and weak binding antigen peptides in order to avoid the problem of inaccurate relative affinity predictions.
[0003] However, in actual use, such as the development of tumor therapeutic vaccines, the affinity of the selected antigen peptide will directly affect the vaccine antigen capacity, immune response rate, and the incidence of autoimmune reactions. Therefore, a calculation method that can accurately calculate and predict the affinity of antigen peptides is of great significance for the development of immunotherapy.
[0004] Existing technical solutions include sequence-based machine learning methods and structure-based computational methods. Wet laboratory methods such as competitive binding assays, BLI, SPR and other detection methods usually have low detection throughput, long cycle times and are expensive.
[0005] Sequence-based machine learning methods mainly involve two steps: 1. Inputting MHC protein sequences and polypeptide sequences into feature construction. The construction methods include feature encoding of amino acids or atoms and selecting important sequence features. 2. Using a series of machine learning algorithms such as logistic regression, SVM support vector machine, random forest and deep learning algorithms such as CNN, GNN, LSTM, Transformer, etc., the experimentally measured affinity values are used as regression tasks, or the affinity values are converted into classification tasks according to thresholds for modeling.
[0006] These sequence-based machine learning methods are often affected by data size and bias. For example, HLA-A*02, the most common type in the human population, has the most historical data, so machine learning algorithms typically perform better than those for rarer HLA types. Specifically, for several common HLA types, including HLA-A*02, such as HLA-A*06 and HLA-A*11, most prediction models typically achieve an accuracy rate above 0.9 and an AUC above 0.8. However, for rare types such as HLA-A*26, the accuracy and AUC of the prediction models are generally below 0.5. Given the polymorphism and structural similarity of MHC proteins, this suggests that current machine learning methods are only overfitting on a subset of datasets, rather than truly learning the general principles of interactions between MHC and peptides. Studies have reported that methods that rely on solid-phase peptide synthesis, such as mass spectrometry detection, cannot capture the binding of certain peptides and MHC in vitro. For example, when the peptide terminus is V, the intracellular peptide binding detection method Episcan can capture this interaction, but it cannot be detected in in vitro experiments. When there are too many hydrophobic amino acids in the peptide, the hydrophobicity of the peptide increases and the solubility decreases, which makes peptide synthesis, dissolution and detection experiments difficult. The peptide binding pocket of MHC prefers hydrophobic amino acid anchors. These situations will cause bias in historical data, which will cause bias and overfitting in machine learning methods that rely on large-scale data sets.
[0007] Model development that incorporates structural information into machine learning methods typically involves feature selection to reduce the dimensionality of the input features, thereby reducing the complexity of the learning task. Alternatively, structural information can be used to refine amino acid-level features to the atomic level or embed features to improve predictive model performance. However, these approaches have not yet achieved significant progress.
[0008] First-principles molecular dynamics simulations can fully exploit the structural information of MHC-peptide interactions to calculate physicochemical interaction energies at the atomic level, thereby estimating the binding free energy of protein-protein interactions. Common molecular dynamics simulation methods decompose the binding free energy into van der Waals forces, electrostatic interactions, solvation free energy, and interaction entropy. Force fields are used to simulate the MHC-peptide interaction process at the picosecond or nanosecond scale. Commonly used software and energy calculation methods include Amber, NAMD, Gromacs, and MMGBSA / PBSA. Traditional molecular mechanics computational methods are limited in simulating protein-protein / protein-peptide systems, primarily due to the lack of accurate force fields to describe the complex interactions of MHC-peptide complexes and the lack of reliable free energy prediction methods. These factors are crucial for accurate affinity prediction, which is key to accelerating immunotherapy drug design and treatment. Currently, physics-based calculations of MHC-peptide binding free energies rely on empirical molecular force fields such as AMBER, CHARMM, and OPLS. However, these force fields ignore key quantum effects such as electrostatic polarization and charge transfer, limiting the accuracy of the calculations. Studies have shown that using quantum mechanics (QM) methods to treat electrostatic interactions can significantly improve the accuracy of protein-ligand docking. In addition, the molecular force field of the point charge model has defects in dealing with interactions with close distances between atoms, which makes it difficult for calculation methods based on current force fields to improve the calculation accuracy of MHC-peptide complexes. Currently, the most commonly used methods for calculating protein-ligand binding free energy are the MM / GBSA and MM / PBSA methods. These two methods use approximate normal mode vibration methods when calculating the contribution of entropy change in binding free energy. The two main disadvantages of this method are: (1) the amount of calculation is large; (2) the reliability of the calculation results is very poor. Therefore, in most calculations of protein-ligand binding free energy, the contribution of entropy change is often ignored. Higher-precision calculation methods such as FEP, TI and quantum chemical calculation methods are not practical at present due to the high computational cost.
[0009] In summary, although existing machine learning-based MHC-peptide affinity calculation methods can make predictions to a certain extent, they have several significant shortcomings:
[0010] 1. Dependence on Extensive Training Data: Existing machine learning models require extensive training data to ensure accurate predictions. Insufficient or unbalanced training data significantly reduces the model's generalization ability, leading to inaccurate or unstable affinity predictions. Furthermore, insufficient experimental data particularly impacts predictions for low-affinity peptides, making it difficult to cover all possible scenarios.
[0011] 2. Data bias and model bias: Existing models are often trained on experimental datasets that may contain measurement errors or selection bias. Consequently, models are prone to overfitting or systematic errors when data are unevenly distributed, particularly in the prediction of low-affinity or specific peptides.
[0012] 3. Limited quantitative prediction of affinity: Existing technologies can only provide qualitative predictions of affinity, but lack precision in quantitative calculations. In particular, the accuracy and applicability of existing methods remain limited in terms of affinity ranking and specific numerical predictions of affinity values. Summary of the Invention
[0013] In order to address the shortcomings of existing machine learning-based MHC-peptide affinity calculation methods, such as reliance on large amounts of training data, data offset and model bias, and limited quantitative prediction of affinity, the present invention proposes a method for calculating MHC-peptide binding free energy.
[0014] The technical solution adopted by the present invention is a method for calculating the MHC-peptide binding free energy, comprising:
[0015] S100, acquiring data, acquiring known MHC-peptide data and known structural information of the target MHC-peptide;
[0016] S210: Determine a preferred polypeptide, perform sequence similarity calculation on the polypeptides in the MHC-peptide data and the polypeptides in the target MHC-peptide, and perform match score calculation on the anchor amino acids of the two, obtain a degree of similarity based on the results of the sequence similarity calculation and the match score calculation, and select a preferred polypeptide with the best match based on the degree of similarity;
[0017] S220, determining a preferred MHC, selecting the MHC most similar to the MHC in the target MHC-peptide from the MHC-peptide data as the preferred MHC;
[0018] S230, forming an MHC-peptide template, forming an MHC-peptide template according to the preferred polypeptide and the preferred MHC;
[0019] S300, perform homology modeling on the MHC-peptide template to obtain candidate MHC-peptides;
[0020] S400, performing molecular dynamics simulation on the candidate MHC-peptide to obtain simulation results;
[0021] S500: Calculate the binding free energy of the MHC and the polypeptide in the candidate MHC-peptide according to the simulation results to obtain the binding free energy.
[0022] Preferably, step S210 includes:
[0023] The edit distance value between the peptides in the MHC-peptide data and the peptides in the target MHC-peptide is calculated by the edit distance;
[0024] Compare and match the anchor amino acids of the peptides in the MHC-peptide data with the anchor amino acids of the peptides in the target MHC-peptide, and calculate the similarity contribution score by using the mismatch penalty and / or the match score;
[0025] Subtract the similarity contribution score from the edit distance value to obtain the similarity degree;
[0026] The polypeptide with the smallest similarity is selected as the preferred polypeptide.
[0027] Preferably, in step S300, the specific steps of obtaining the candidate MHC-peptide include:
[0028] Modification of non-anchor amino acids in the MHC-peptide template;
[0029] Multiple predicted structures are generated, and the predicted structure with the best score is selected as the candidate MHC-peptide.
[0030] Preferably, the step of modifying the non-anchor amino acids in the MHC-peptide template includes: in the modification process involving multiple non-anchor amino acids, all discontinuous non-anchor amino acids are modified at once, and continuous non-anchor amino acids are modified one by one.
[0031] Preferably, step S500 includes:
[0032] S510, selecting amino acids in the candidate MHC-peptide to form an amino acid set;
[0033] S520, performing an alanine scan for each amino acid in the amino acid set to calculate its van der Waals interaction, electrostatic interaction, solvation free energy, and non-solvation free energy;
[0034] S530, calculating the entropy contribution of each amino acid using an interaction entropy calculation method;
[0035] S540. Perform binding free energy calculations based on the amino acid set, including: S541. Summing the energy terms in step S520 and step S530 to obtain the original binding free energy; S542. Extracting the van der Waals interaction energy and summing the entropy contribution to obtain the primary improved binding free energy; S543. Calculating the primary improved binding free energy of the amino acids corresponding to the MHC and polypeptide, respectively, and taking the arithmetic mean thereof as the final improved binding free energy.
[0036] Preferably, step S510 includes:
[0037] Amino acids whose distance from any atom on the MHC in the candidate MHC-peptide is within a cutoff distance of the polypeptide in the candidate MHC-peptide are selected, and these amino acids and all amino acids on the polypeptide in the candidate MHC-peptide are used as the amino acid set for alanine scanning. The cutoff distance ranges from 5 angstroms to 12 angstroms.
[0038] Preferably, after step S100 and before step S210, the following steps are included:
[0039] Supplementation of missing heavy atoms for MHC-peptide data;
[0040] For the MHC-peptide data, other irrelevant components that aided crystallization, except for MHC and peptide, were removed;
[0041] Filter MHC-peptide data with a resolution below 3 angstroms.
[0042] Preferably, step S400 includes:
[0043] Before performing molecular dynamics simulations, hydrogen atoms are added to check for atomic overlap. If any atomic overlap occurs, it is repaired manually and / or automatically corrected using the optimization command.
[0044] After ensuring that hydrogen atoms are added and the structure is correct, water molecules are added, and ions are added to the simulation system to make the charge of the simulation system zero;
[0045] Molecular dynamics simulations were performed to obtain the simulation results.
[0046] In order to solve the defect that polypeptides that strongly bind to MHC cannot be screened out, the present invention proposes an MHC-peptide screening method for screening polypeptides obtained by the above-mentioned MHC-peptide binding free energy calculation method.
[0047] Preferably, polypeptides are screened based on improved binding free energy and the free energy contribution of the anchor amino acid.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] The present application discloses a method for calculating the free energy of MHC-peptide binding, in which a template construction technique is used to obtain the initial structure of any unknown candidate MHC-peptide by adjusting the amino acids in the known MHC-peptide data. Constructing an accurate MHC-peptide template is a prerequisite for accurately calculating affinity. The present application adopts a template selection method based on anchor amino acids, aiming to select templates by combining highly conserved structural features of the groove (such as hydrophobic pockets and α-helix-β-sheet combinations). This method is different from the traditional template matching method based on sequence similarity. It particularly emphasizes the selective binding of the first and last amino acids (anchor amino acids) of the polypeptide to the two ends of the MHC binding groove, which helps to improve the accuracy and reliability of modeling.
[0050] In the MHC-peptide complex structure, the peptide binds to the binding groove at the base of the beta sheet, formed by two alpha helices (helices 65-85 and 90-120) of the MHC. This groove can accommodate peptides of 8-12 amino acids and is relatively stable. The peptide's binding position is fixed, and the RMSD (root mean square deviation) of the peptide's spatial position relative to the MHC varies by less than 1 Å across different MHC-peptides. The smaller the edit distance of the peptide sequence, the smaller the RMSD. Therefore, the initial structure of any unknown candidate MHC-peptide can be obtained by modifying a portion of the peptide's amino acids using known MHC-peptide data. The experimental dataset, which includes 327 peptides with reliable IC50 experimental data, shows a correlation of 0.55 between theoretically calculated and experimental affinity values. The correlation is particularly high, reaching 0.76, for peptides with an edit distance to the template less than 2.
[0051] This result shows that by rationally selecting the template and performing appropriate data correction, affinity calculation results that are consistent with experimental observations can be obtained. These findings make this application not only an important contribution to the progress in the field of immunology, but also provide strong support for the design and development of peptide-based vaccines and immunotherapy. At the same time, it avoids the defect of relying on a large amount of training data in the MHC-peptide affinity calculation method based on machine learning, and no data offset and model deviation will occur during the calculation process, and the quantitative prediction ability of affinity is better. Compared with the prior art, the MHC-peptide binding free energy calculation method disclosed in this application demonstrates its powerful performance in predicting polypeptide affinity, laying the foundation for further understanding the immune recognition mechanism and strengthening the research on MHC-peptide interactions.
[0052] The present application also discloses an MHC-peptide screening method, which is used to screen polypeptides obtained by the above-mentioned MHC-peptide binding free energy calculation method, thereby screening out polypeptides with higher affinity.
[0053] Compared with the prior art, the MHC-peptide screening method disclosed in the present application can improve the accuracy of screening out strongly binding peptides. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The present invention is described in detail below with reference to the embodiments and accompanying drawings, in which:
[0055] Figure 1 A flowchart of a method for calculating MHC-peptide binding free energy according to an embodiment of the present invention is shown;
[0056] Figure 2 A schematic flow chart of a method for calculating MHC-peptide binding free energy according to an embodiment of the present invention is shown;
[0057] Figure 1a A comparison diagram of various energy terms of the template structure, constructed 4K7F, and negative control pMHC complexes in one embodiment is shown;
[0058] Figure 1b A diagram showing score distribution of peptide template selection and detailed information record of template alignment in one embodiment is shown;
[0059] Figure 2a shows a polypeptide enthalpy energy distribution diagram in one embodiment;
[0060] Figure 2b A graph comparing correlation coefficients between templates and negative controls in one embodiment is shown;
[0061] Figure 2c A diagram showing the RMSD (root mean square deviation) fluctuation range of the peptide portion in one embodiment is shown;
[0062] Figure 2d A comparison diagram of RMSD (root mean square deviation) in one embodiment is shown;
[0063] Figure 2e It shows that two negative control peptide structures constructed in one embodiment exhibit different unstable kinetic behaviors;
[0064] Figure 3a A correlation distribution diagram showing the calculated results of the van der Waals force and interaction entropy of a polypeptide and the experimental results in one embodiment is shown;
[0065] Figure 3b A correlation distribution diagram of the sum of the van der Waals force and interaction entropy of peptide amino acids and the experimental values in one embodiment is shown;
[0066] Figure 4a shows a polypeptide energy distribution diagram in one embodiment;
[0067] Figure 4bAn energy distribution diagram of MHC hotspot amino acids in one embodiment is shown;
[0068] Figure 5A Shown are the peptide chain sequences, theoretical calculations, and ELISA test results in one embodiment;
[0069] Figure 5B A graph showing a positive correlation between theoretical energy and affinity measured by BLI experiments in one embodiment is shown. DETAILED DESCRIPTION
[0070] To make the objectives, technical solutions, and advantages of the present invention more apparent, embodiments of the present invention will be described in further detail below with reference to the accompanying drawings. Examples of embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar components or components having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0071] The present invention discloses a method for calculating MHC-peptide binding free energy, comprising:
[0072] S100, acquiring data, acquiring known MHC-peptide data and known structural information of the target MHC-peptide;
[0073] S210: Determine a preferred polypeptide, perform sequence similarity calculation on the polypeptides in the MHC-peptide data and the polypeptides in the target MHC-peptide, and perform match score calculation on the anchor amino acids of the two, obtain a degree of similarity based on the results of the sequence similarity calculation and the match score calculation, and select a preferred polypeptide with the best match based on the degree of similarity;
[0074] S220, determining a preferred MHC, selecting the MHC most similar to the MHC in the target MHC-peptide from the MHC-peptide data as the preferred MHC;
[0075] S230, forming an MHC-peptide template, forming an MHC-peptide template according to the preferred polypeptide and the preferred MHC;
[0076] S300, perform homology modeling on the MHC-peptide template to obtain candidate MHC-peptides;
[0077] S400, performing molecular dynamics simulation on the candidate MHC-peptide to obtain simulation results;
[0078] S500: Calculate the binding free energy of the MHC and the polypeptide in the candidate MHC-peptide according to the simulation results to obtain the binding free energy.
[0079] Step S100 acquires data. MHC-peptide data can be obtained from public databases such as PDB, IMGT, TCR3d, VDJdb, etc. MHC-peptide data are usually obtained by analyzing techniques such as cryo-electron microscopy and crystal diffraction, and stored as pdb files or other file formats. The pdb structure files of the collected MHC-peptide data will be used in the next step. Preferably, after acquiring the MHC-peptide data, these template files need to be preprocessed, such as supplementing missing heavy atoms, removing other irrelevant crystallization-aiding components other than MHC and polypeptides from the MHC-peptide data, removing low-quality templates based on crystal resolution, etc. The commonly used resolution that meets the requirements for removing low-quality templates based on crystal resolution should be below 2-3 angstroms.
[0080] The step S200 of obtaining the MHC-peptide template includes step S210 of determining a preferred polypeptide, step S220 of determining a preferred MHC, and step S230 of forming the MHC-peptide template.
[0081] Determining the preferred peptide in step S210 is based on an optimal template matching method using anchor amino acids. The target MHC-peptide to be calculated includes the target peptide and its corresponding MHC. Therefore, the more similar the selected template is to the target MHC-peptide, the fewer amino acids need to be modified in step S300, resulting in more accurate results. Current homology modeling methods mostly use sequence alignment, selecting templates based on sequence similarity algorithms such as edit distance, PAM, and BLOSUM scoring matrices. In the present invention, due to the highly conserved nature of MHC-peptide binding, peptides bind to a structurally stable MHC binding groove. This groove, consisting of a pair of alpha helices and a group of beta sheets, forms a binding pocket that prefers hydrophobic amino acids and is conserved across different MHC types. Several amino acids at the beginning and end of the peptide are referred to as anchor amino acids due to their physicochemical properties, enabling them to selectively bind to the ends of the binding groove. It is generally believed that amino acids in the middle of the peptide interact less with the MHC. Therefore, the present application optimizes the alignment and template selection method based on sequence similarity, and adds a mismatch penalty or matching score for anchor amino acids on the basis of the sequence similarity algorithm. For example, when the target sequence and the polypeptide in the candidate template are identical at the 1st, 2nd, 3rd or the penultimate and 2nd amino acids, a matching score of 1 point is given, especially when the second amino acid is identical, a score of 2 points is given.
[0082] The specific method for scoring peptide similarity is as follows: the similarity between two peptide sequences can be evaluated using the Levenshtein edit distance. When selecting the most suitable MHC-peptide template for each peptide, templates with the same anchor amino acids as the target peptide are preferred. This is because using these templates can avoid changing the anchor amino acids when constructing candidate MHC-peptides, thereby minimizing the interference that amino acid modifications may cause in energy calculations. To optimize this method, the Levenshtein edit distance is first calculated between each pair of peptides, and amino acids that require substitution, insertion, and deletion are considered separately, and only those templates that require only substitution are retained. Next, the similarity of the anchor amino acids between the target and template peptides is evaluated, and the anchor amino acids are defined as the first three and last two amino acids of the peptide. If there are n amino acid matches, the similarity contributes n points.
[0083] The final template match score, called the MatchScore, is calculated by subtracting the anchor amino acid similarity score from the Levenshtein edit distance. Each peptide is then matched to the template that achieves the lowest MatchScore, indicating the shortest edit distance and the highest anchor amino acid similarity. This optimized match score ensures that the selected MHC-peptide template is optimal for constructing MHC-peptides with minimal deviation from the original sequence, thereby improving the accuracy of subsequent affinity calculations and predictions.
[0084] Modification design of template matching method based on anchor amino acid:
[0085] The peptide sequence matching method can be changed to other sequence matching scoring algorithms based on PAM, BLOSUM matrix, etc., and the implementation software can be replaced with other commonly used sequence matching algorithms and software such as blastp and clustalw;
[0086] The matching score or penalty adjustment of the anchor amino acid can be used to adjust the details of the scheme, such as selectively incorporating the first and last 1-3 amino acids of the peptide into the calculation, and changing the score or penalty according to the position.
[0087] Step S230 determines the preferred MHC. The preferred MHC is selected by screening the MHC with the most similar gene family according to the HLA nomenclature, or screening the most similar MHC from the MHC-peptide data according to the similarity calculation method, and the MHC that is exactly the same as the target MHC is preferentially selected.
[0088] In step S230, an MHC-peptide template is formed by combining the preferred MHC and the preferred polypeptide. When considering amino acid substitutions, insertions, and deletions, the rule is to modify only a few amino acids within the polypeptide in the subsequent step S300, while minimizing the need to alter the anchor amino acids. Amino acids within the MHC binding groove should not be modified. Based on these considerations, an MHC-peptide is selected as the MHC-peptide template.
[0089] The present invention uses an anchor amino acid-based template selection method, aiming to select templates by incorporating highly conserved structural features of the groove, such as the hydrophobic pocket and the α-helix-β-sheet combination. This method differs from traditional template matching methods based on sequence similarity and has the following innovations:
[0090] Conservation of the binding groove: Utilize the conserved structure of the peptide and MHC binding groove for template selection to reduce unnecessary amino acid modifications and improve model accuracy.
[0091] Anchor amino acid selection: Special emphasis is placed on the selective binding of the first and last amino acids of the peptide (anchor amino acids) to both ends of the MHC binding groove, which helps to improve the accuracy and reliability of modeling.
[0092] By incorporating the conserved nature of anchor amino acids, this paper proposes a precise method for MHC-peptide modeling. This method leverages the limited interaction between amino acids in the peptide's middle region and the MHC molecule to minimize the impact of amino acid modifications, focusing on preserving the binding sites at the peptide's head and tail ends. Furthermore, after modifying individual amino acids, local side chain energy optimization is performed, thereby improving the rationality of MHC-peptide binding throughout the modeling process.
[0093] Step S300 obtains a candidate MHC-peptide. Based on homology modeling, software such as Modeller and Rossetta can be used to make necessary amino acid modifications based on the MHC-peptide template selected in the previous step to construct a candidate MHC-peptide. When making amino acid modifications, if multiple amino acid modifications are involved, they can be modified all at once or one by one in steps. Generally, the two methods are adopted under the following conditions: when the amino acids to be modified are discontinuous or spatially far apart, they can be modified all at once; when the amino acids to be modified are close in distance, they can be modified one by one. After each amino acid modification, it is necessary to perform energy minimization side chain optimization for this amino acid, while keeping the coordinates of other unmodified amino acids unchanged. Different software will output multiple optional predicted structures as candidate MHC-peptides, and the predicted structure with the best score is usually selected based on the corresponding evaluation score.
[0094] Specific steps for constructing the initial structure of candidate MHC-peptide:
[0095] Using MODELLER version 10.2, each target peptide and the selected MHC-peptide template PDB structure were input into the software. MODELLER performed a sequence alignment between the target peptide and the MHC-peptide and constructed an initial protein structure. This structure was then optimized through side-chain alignment and energy minimization to generate five predicted final structures.
[0096] To evaluate these structures, DOPE (Discrete Optimized Protein Energy) and GA341 scores were used, which are used to assess the structural and energetic feasibility of the modeled complexes. The structures with the best scores were ultimately selected as the final candidate MHC-peptide for the target peptide, thus ensuring the optimality of the structural prediction and the accuracy of the potential biological function.
[0097] Changes to the target pMHC structure modeling method based on homology modeling can be designed using homology modeling software such as Modeller and Rosetta. Software such as Amber, CHARMM-GUI, and Gromacs can also be used to perform atomic modifications to achieve the purpose of homology modeling. Alternatively, programming languages such as Python and Fortran can be used to create scripts to perform atomic modifications and adjustments in the PDB file, followed by molecular dynamics methods for local structure optimization.
[0098] Step S400 obtains the simulation results. Specifically, the system can be prepared before performing molecular dynamics simulation. First, use the tLEap tool in the Amber software to prepare the system, start tLEap and load the required molecular structure file, such as loading the molecular structure file in PDB format through the loadPdb command. tLEap will initialize the system based on these files and prepare for subsequent simulations. Next, hydrogen atoms are added. In most PDB files, hydrogen atoms are usually not included, so it is necessary to use the addHydrogens command in tLEap to add hydrogen atoms to the system. Hydrogen atoms are crucial to the accuracy of the simulation because they affect the structure and dynamic behavior of the molecule. After completing the addition of hydrogen atoms, it is necessary to check whether there is any atomic overlap, which can be verified using the check command. If overlap is found, tLEap will give feedback, and the user can manually repair it according to the prompts or automatically correct it through optimization commands (such as minimize).
[0099] After ensuring that the hydrogen atoms have been added and the structure has been modified correctly, the next step is to add a solvent box. Use the solvateBox command in tLEap to add water molecules to the system. Common water models include TIP3P and SPC / E. When adding water molecules, ensure that the water box is large enough to avoid boundary effects between molecules. It is generally recommended that the water box be larger than the size of the molecule itself. In addition, to simulate physiological conditions and ensure charge balance of the system, ions need to be added to the system. Use the addIons command in tLEap to add sodium ions (Na+) and chloride ions (Cl-). Ensure charge balance in the simulation by setting appropriate ion concentrations and bringing the system charge to zero. If the total charge of the system is not zero, tLEap will automatically correct it.
[0100] After these steps are complete, the system is ready for molecular dynamics simulations. Next, the system can be pre-equilibrated and heated, typically by gradually heating the system from a low temperature state to the target temperature to avoid unphysical behavior caused by sudden temperature changes. After equilibration and heating, the system is simulated for a short time to minimize the energy. The actual simulation is then performed, typically with a simulation time of 1-50 nanoseconds and 3-5 trajectory repetitions to ensure the reliability of the simulation results.
[0101] The specific procedures and parameter settings for molecular dynamics simulation are as follows:
[0102] The candidate MHC-peptide complexes were parameterized using the AMBERff14SB force field. The complexes were dissolved in a periodic box filled with TIP3P water molecules, ensuring a minimum buffer distance between any atom in the complex and the edge of the box of To neutralize the charge of the system, an appropriate number of sodium or chloride ions were added. Long-range electrostatic interactions were handled using the Particle Mesh Ewald (PME) algorithm, and the cutoff distance for non-bonded interactions was set to
[0103] Each system was energy minimized using the steepest descent method and the conjugate gradient method. Initially, the complex was subjected to Force constant restraints were applied to allow the solvent molecules to adjust around the complex; these restraints were subsequently removed to optimize the entire system, particularly to address potential atomic collisions caused by the sudden change. The system was then slowly heated to the target temperature of 300 K in the NVT ensemble (over 300 ps), followed by a 1 ns simulation in the NPT ensemble with continuous restraints applied to all atoms of the complex. The temperature was controlled using Langevin dynamics (collision frequency of 1.0 ps-1), and the pressure was maintained at 1.0 atm using a Berendsen constant pressure cell.
[0104] Finally, a 50 ns unconstrained production simulation was performed, with trajectory data saved every 1 ps. These 50 ns trajectories were used for binding free energy calculations, and all 50,000 saved frames were used for entropy calculations, while 1,000 frames were uniformly selected from these for enthalpy calculations.
[0105] Changes to the design of the molecular dynamics simulation process:
[0106] The software used to implement molecular dynamics simulation can be replaced, such as Amber, Gromacs, NAMD, CHARMM-GUI or their derivative software;
[0107] Specific settings of the simulation process, such as the specific force field, water box type, shape and size selection, each energy optimization, and specific parameters of the simulation process such as simulation time, step size, etc.
[0108] Step S500 obtains the binding free energy. Based on the ASGBIE method and the improved binding free energy calculation, alanine scanning is a method used in wet experiments to study the contribution of a specific amino acid to the interaction. In the calculation, the amino acids that contribute to the interaction in the complex trajectory obtained by the previous step simulation are modified one by one to alanine. Since the alanine side chain is very small, it can be considered that it has no interaction with the surrounding amino acids. By calculating the free energy difference before and after the modification, the specific value of the contribution of the amino acid at that position to the binding free energy can be obtained.
[0109] Specifically, first, amino acids on the MHC protein on the candidate MHC-peptide whose atoms are within 5 angstroms from the polypeptide are selected, and these amino acids and all amino acids on the polypeptide on the candidate MHC-peptide are used as the amino acid set requiring alanine scanning.
[0110] The van der Waals interaction, electrostatic interaction, solvation free energy, and desolvation free energy were calculated for each amino acid using the MMGBSA method.
[0111] The entropy contribution of each amino acid was calculated using the developed interaction entropy calculation method;
[0112] Calculation of improved binding free energy: Specifically, the individual energy terms in the above steps are calculated using the amino acid sets on the MHC and peptide, respectively. The sum of all energy terms is then calculated as the original binding free energy, a standard approach in the field. The "improved binding free energy" is calculated by summing the van der Waals interaction and entropy. This method eliminates the bias caused by electrostatic interactions and incorporates entropy, making it an improved method consistent with the MHC-peptide system. The average of the "improved free energies" for the amino acid sets on the MHC and peptide is used as the "final improved binding free energy." Simultaneously, the sum of the "improved binding free energies" of the peptide anchor amino acids is calculated as a reference for subsequent calculations.
[0113] The specific calculation formula of the ASGBIE method is as follows:
[0114] In each MHC-peptide system, the binding energy was evaluated using the ASGBIE (generalized Born model and alanine scanning of interaction entropy) method. The contribution of a specific mutant residue (x) to the binding energy was determined using the alanine scanning technique, where the difference in binding free energy between the alanine mutant and the wild-type complex reflects the effect of the mutation, as shown in Equation 1. Equation 1 calculates the binding free energy of the alanine mutant and wild-type complex:
[0115]
[0116] Furthermore, Equation 1 can be broken down into its various components, as shown in Equation 2. Although the MHC-peptide interface is usually composed of a large number of residues, there are usually only a few key residues that play an important role in binding. Assuming that the distance from the interface is more than The overall binding energy is thus obtained by summing the distances within the MHC or peptide The contribution of each residue in the range is calculated as shown in Equation 3:
[0117]
[0118] By performing alanine scanning of MHC (MHC-AS) or peptide (Peptide-AS) residues, hotspot residues critical for binding can be identified and their individual contributions to the overall binding free energy calculated. This approach provides a detailed analysis of the energy contributions in MHC-peptide complexes, contributing to a deeper understanding of the interaction mechanisms between molecules.
[0119] Changes in the design of the binding free energy calculation method:
[0120] The ASGBIE method used in this application can be replaced by other binding free energy calculation methods such as MMGBSA / MMPBSA;
[0121] The improved binding free energy method used in this application can be changed to other energy term addition methods, and different degrees of weights are assigned to each energy term as a specific adjustment strategy to achieve better consistency with the experimental value.
[0122] The present invention further utilizes and improves a novel binding free energy calculation method (ASGBIE method) based on the structural modeling of the target MHC-peptide. This method accurately calculates the binding free energy between the peptide and the MHC and takes into account the influence of interaction entropy. Simultaneously, calculations are performed from the amino acids on both the MHC and the peptide. The improved calculation method takes into account the interaction energy contributions of both and eliminates the inaccuracy of traditional electrostatic interaction calculations. This method helps predict and evaluate the stability of different MHC-peptide bindings and optimizes the design of molecules related to immune responses.
[0123] Example 1: Verification of the multi-template calculation method for 4K7F.
[0124] In order to verify the template-based pMHC (wherein "pMHC" in this application means "MHC-peptide") structure construction method, the pMHC structure in the PDB database was used to predict the new structure. First, 4K7F was selected as the prediction target, and five other known pMHC structures (1I7U, 3H7B, 5HHQ, 5FA3, 5E00) were selected from the PDB as construction templates. All selected known structures and target structures belong to the HLA-A02:01 type MHC, but there are differences of 6 to 8 amino acids in their peptide chain sequences (see Table 1). This selection facilitates the study of the effects of different editing distances on the prediction of the target pMHC structure. At the same time, the 3OXS structure was selected for molecular dynamics simulation to analyze and compare the binding modes of peptides consisting of 10 and 9 amino acids.
[0125] A 50-nanosecond simulation was performed on the known 4K7F structure selected from the PDB and the predicted pMHC structure constructed using the Modeller software. During the entire simulation, the RMSD (root mean square deviation) of the peptide portion of the selected known 4K7F structure and the 4K7F pMHC structure constructed using 1I7U, 3H7B, and 5FA3 as templates fluctuated within a range of approximately 1 angstrom ( Figure 2c ), showing good structural stability and highly consistent with the known 4K7F structure ( Figure 2d The peptides in the 4K7F pMHC structure constructed using 5E00 and 5HHQ as templates showed a slightly higher average RMSD of approximately 1.5 angstroms. Stability analysis confirmed that template-based molecular dynamics simulations are effective in predicting and validating the structure of the pMHC complex and its dynamic interactions.
[0126] like Figure 1aAs shown, the ASGBIE method was used to calculate the global energy contributions of hotspot amino acids in MHC molecules and peptides. Comparative analysis of the predicted target structure with the reference structure 4K7F revealed high similarity in all evaluated energy terms, with van der Waals interactions for MHC hotspot amino acids approximately 45 kcal / mol, enthalpy values of approximately 63 kcal / mol, and a total energy of 23 kcal / mol. These energy terms significantly outperform those of two negative control peptides, supporting the hypothesis that the predicted structure possesses interaction stability comparable to known structures. In contrast, the constructed negative control pMHC complex exhibited insufficient interaction stability.
[0127] In the analysis of hot spot amino acids of the peptide, the calculated van der Waals interactions, enthalpy values, and overall energy contributions also showed good consistency between the reference structure and the target structure. Figure 2a and Figure 4a The results show that within each template structure, the peptides exhibit binding modes in which the second, third, and terminal amino acids contribute most to the interaction energy, while the central amino acids of the peptides exhibit different interaction modes between different templates. Notably, the first amino acids of the 9-amino acid peptide of 3H7B and the 10-amino acid peptide of 3OXS exhibit strong interactions. The energy profiles of peptides derived from various template mutations are highly consistent with those of the classic 4K7F structure, indicating that selecting any template with 5 to 8 amino acid mutations will generate an energy profile similar to the original 4K7F structure without changing the peptide's binding mode.
[0128] Based on the unique energy distribution of the reference template used for modeling, this result confirms that the energy distribution of the peptide derived in molecular dynamics simulations is less affected by the choice of template. Correlation analysis based on the peptide interaction pattern further revealed that the energy distribution pattern of the 4K7F construct is more similar to that of the known target structure 4K7F than when the template is used alone. The Pearson correlation coefficient of the energy distribution pattern between the 4K7F constructs generated by different templates and the reference structure exceeds 0.9, showing strong similarity, while the correlation coefficient with their respective original templates is less than 0.8, and the correlation coefficient between templates and the negative control is less than 0.7 ( Figure 2b This pattern is also reflected in the energy distribution of MHC hotspot amino acids ( Figure 2b and Figure 4b ).
[0129] Example 2: Computational verification of negative peptides.
[0130] To evaluate the binding of peptides with lower MHC binding ability, a negative control peptide (SRYWAIRTR) was selected, which was experimentally shown to have an affinity for HLA-A02:01 below 50,000 nM, and a new pMHC construct was constructed with it.
[0131] In molecular dynamics simulations, the two negative control peptide structures constructed showed different unstable dynamic behaviors ( Figure 2e In one case, the first half of one of the negative peptides moved out of the MHC binding groove during the simulation, indicating that there was insufficient interaction between the anchoring amino acids and the MHC, resulting in unstable binding. This phenomenon is consistent with the larger RMSD of the peptide ( Figure 2d In another case, although the position of the peptide remained relatively unchanged and its RMSD did not exceed 2 angstroms, the two α helices of the MHC binding groove expanded outward, making the groove looser, reflecting that the structural stability of the MHC binding groove itself was affected ( Figure 2d This change in the pMHC structure likely results from repulsive forces induced by the negative control peptide, further impacting the structural stability of the entire complex. This structural change correlated with peptide stability highlights the importance of considering peptide-MHC interactions when assessing pMHC affinity and stability.
[0132] In the analysis of hotspot amino acids of the peptides, the van der Waals interactions, enthalpy values, and overall energy contributions of the negative control peptide did not show significant differences from those of the other peptides in Example 1. This may be due to the small number of amino acids in the peptides and the changes in the amino acid composition of the peptides, which are more likely to affect the accuracy of the energy calculation.
[0133] Both negative control peptides exhibited significant differences in MHC and peptide energy profiles compared to any of the structures in Example 1, as well as to each other. This demonstrates that, despite the imperfect accuracy of molecular dynamics calculations of peptide total energy, analysis of peptide binding patterns at the amino acid level reveals the mechanism of MHC-peptide interaction. An accurate description of energy profiles is not only fundamental to MHC-peptide affinity but also demonstrates the potential for distinguishing positive from negative peptides through hotspot amino acid energy analysis.
[0134] Example 3: Large-scale simulation verification of 426 polypeptides.
[0135] A publicly available benchmark dataset was used, which contains a collection of 426 peptides with affinity data, and is specifically designed to evaluate peptide affinity prediction. The corresponding pMHC known structure templates were obtained from the Protein Data Bank (PDB), where 187 different pMHC structures were selected from 397 HLA-A type pMHCs, all classified as HLA-A*0201 type MHC. Of these structures, 12 / 426 peptides have known structures. In order to determine the best template for each peptide prediction, the peptide similarity scoring method based on edit distance in the signed patent scheme was used. This method combines the edit distance of the peptide and the weight of the anchor amino acid, and is able to select structural templates with the same anchor site and requiring the least amino acid mutations. The amino acid edit distance of the test peptides ranged from 1 to 7, and their match scores (Match Score) with the corresponding templates ranged from -5 to 6, with most peptides having a match score of 3. The score distribution for each peptide template selection and detailed information on template alignment are recorded in Figure 1b middle.
[0136] The initial structure of each peptide was constructed according to the method described previously, and a 60-nanosecond molecular dynamics (MD) simulation was performed on it. The ASGBIE method was then used to calculate the energies of the hotspot amino acids of the MHC molecule and the peptide, respectively. The theoretical energy was positively correlated with the experimentally measured affinity, as shown in Table 2. When analyzing the consistency of various energy parameters including van der Waals forces, charge interactions, and interaction entropy with the experimental values, it was found that van der Waals interactions were the main factor mediating the MHC-peptide interaction, and the correlations between MHC and peptide hotspot amino acids were 0.32 and 0.45, respectively. The charge interaction of MHC hotspot amino acids is the main source of calculation error, with a correlation of only 0.08, while the correlation of peptide hotspot amino acids was 0.38. The calculation method combining van der Waals forces and peptide amino acid interaction entropy (excluding the less accurate charge interaction part) had the highest correlation with the experimental results, reaching 0.48 ( Figure 3a ), exceeding the correlation of free energy change (ΔG) of 0.39, which indicates the important role of interaction entropy.
[0137] Furthermore, pMHC energies calculated based on peptide amino acids agree better with experimental data than those calculated based on MHC hotspot amino acids. The most accurate calculation of pMHC binding energy (ΔG) was achieved by averaging the van der Waals forces and entropy values between MHC and peptide hotspot amino acids, with a correlation coefficient of 0.39. Energy values calculated using peptide hotspot amino acids indicate that the total energy of negative peptides is less than 8 kcal / mol.
[0138] Example 4: Bias of experimental data and its impact on the accuracy of computational models.
[0139] When the IC50 benchmarks for these 426 peptides were reviewed and compared with experimental evidence from the Immune Epitope Database (IEDB), it was discovered that the experimental affinity data for some peptides were incorrectly annotated as IC50 rather than IC50. Furthermore, the IC50 / EC50 affinity values for some peptides were inconsistent with the qualitative annotations in the IEDB. In the IEDB, when multiple experimental entries for a peptide are available, the IC50 result is prioritized, and the qualitative description with the highest frequency among the multiple entries is considered representative of the peptide's final affinity status. In this dataset, 22 peptides were classified as "negative" in the IEDB, but their IC50 values were below 500 nM (inconsistent category 1); 55 peptides were classified as "positive," but their IC50 values were above 500 nM (inconsistent category 2); and an additional 22 peptides were annotated as EC50 (inconsistent category 3). A total of 99 peptides had ambiguous classifications.
[0140] After removing these ambiguous entries from the dataset, the correlation between energy calculations and experimental affinity values was improved, with the sum of the van der Waals force and interaction entropy of peptide amino acids (excluding electrostatic interactions with lower accuracy) showing the highest correlation with the experimental value, reaching 0.55 ( Figure 3b This finding highlights the significant impact of uncertainty in benchmark datasets on evaluation accuracy, particularly for machine learning models that rely on data authenticity. Correlation analysis between theoretically calculated values and experimental affinity data for these three categories of ambiguous peptides revealed a correlation of 0.3 for inconsistency category 1, indicating that these peptides are in a critical state of high-affinity MHC interaction. For peptides with experimental IC50 values below 500 nM, the binding free energy was mostly concentrated around 6.5 kcal / mol. This suggests that challenges exist in the consistency and judgment criteria of experimental results, leading to data uncertainty.
[0141] In contrast, the correlations for inconsistent categories 2 and 3 were -0.07 and -0.28, respectively, indicating that for peptides with contradictory results, there were large discrepancies between the experimental results and theoretical calculations. These discrepancies may come from errors in the experiment itself or in the data recording process. Peptides labeled as EC50 may originate from experimental designs that are different from IC50 experiments. The presence of these uncertain data suggests that other peptides may also have similar labeling errors. Combined with potential experimental differences (such as the differences found by Ding et al. in different experimental tests of a set of HBV peptides), this poses a substantial risk that machine learning-based methods may introduce significant bias.
[0142] Example 5: Practical screening and verification of HBV polypeptides based on competitive binding ELISA experiments.
[0143] Using the patented method, we calculated the affinity of peptides derived from the HBV virus. Two of these peptides have publicly reported experimental validation results and will serve as positive and negative controls for this experiment. However, no corresponding competitive binding assays or other affinity assays have been performed for the two peptides tested.
[0144] The principle of this experiment is to utilize the instability of MHC and its subunits in the absence of peptides to competitively bind the test peptide with the reference peptide. When the test peptide has a strong affinity, more MHC-peptide complexes will be retained during multiple rounds of UV irradiation and elution and captured by the ELISA technology, thereby showing a higher detection value. The relevant detection reagent products are purchased from Legend Biotech, the product model is: Flex-T TM HLA-A*02:01Monomer UVX and LEGEND MAX TM Flex-T TM Human Class IPeptide Exchange ELISA Kit.
[0145] The experimental results are shown in the table. Among the three different binding free energy calculation methods, the binding free energy of the positive control peptide was higher than that of the negative control peptide. The improved MHC-peptide free energy calculation method achieved the largest difference between theoretically calculated values, demonstrating that the improved calculation method has better ability to distinguish between positive and negative peptides. Of the two test peptides, test peptide 1 had a higher binding free energy than peptide 2. Their ELISA experimental values were 1.127 and 0.118, respectively, consistent with the theoretically calculated values. This indicates that the patented method accurately distinguishes unknown test peptides. Furthermore, a comparison of the ELISA competitive binding results of the four peptides with the reference OD values revealed that the positive control peptide had the largest experimental and theoretically calculated values, indicating high affinity for MHC. The negative peptide had the smallest theoretically calculated and experimental values, indicating medium to low affinity. Of the two test peptides, test peptide 1 had a medium-to-high affinity, while test peptide 2 had a medium-to-low affinity.
[0146] Example 6: Screening and verification of HBV polypeptides based on BLI experiments.
[0147] Using the patented method, we calculated the affinity of peptides derived from the HBV virus. Two of these peptides had publicly reported experimental validation results and served as the positive and negative controls for this experiment. However, we did not find any experimental records for a competitive binding assay or other affinity assay for one of the peptides being tested.
[0148] The experimental principle is to use biolayer interferometry (BLI) to determine the affinity of the test peptide for the MHC molecule. The test peptide and reference peptide are each bound to a BLI sensor probe preloaded with MHC. BLI is used to monitor the binding and dissociation processes in real time, recording the association rate (ka) and dissociation rate (kd), and calculating the binding constant (KD). When the test peptide has a higher affinity, it will exhibit a lower dissociation rate and a higher binding constant, reflecting a more stable MHC-peptide complex. During the experiment, the quantitative nature of the peptide affinity is further verified by binding the test peptide solution to the MHC probe at a gradient concentration.
[0149] The experimental results are shown in the table. Among the various free energy calculation methods based on molecular dynamics, the binding free energy of the positive control peptide is higher than that of the negative control peptide, and the improved MHC-peptide free energy calculation method shows the largest theoretical calculation difference, which reflects the superiority of the improved method in the ability to distinguish between positive and negative control peptides.
[0150] For the test peptide, its binding free energy is 38.907, the sum of the van der Waals interaction and entropy is 39.578, and the binding free energy obtained by the improved free energy calculation method is 33.194. In the BLI experiment, the binding constant (Ka) of the test peptide is 8.37×10 2 1 / Ms, and the dissociation rate (Kdis) was 1.79×10 -4 1 / s, and the equilibrium dissociation constant (KD) is 2.14×10 -7 M, indicating that it has a medium to high binding ability. The binding free energy of the negative control peptide is 26.527, the sum of the van der Waals interaction and the entropy value is 23.491, and the binding free energy calculated by the improved free energy calculation method is 16.817. Since the BLI experiment failed to detect its Ka, Kdis, and KD values, it shows that the binding ability of this peptide to MHC is low. The binding free energy of the positive control peptide is 32.811, the sum of the van der Waals interaction and the entropy value is 34.314, and the binding free energy calculated by the improved free energy calculation method is 27.223. In the BLI experiment, the positive control peptide showed the highest Ka value (1.13×10 3 1 / Ms), the lowest Kdis value (4.82×10-71 / s), and the lowest KD value (4.27×10 -10 / M), indicating that it has the highest affinity with MHC.
[0151] Combining theoretical calculations and BLI experimental results, the binding ability of the tested peptide was between the positive control and the negative control. The improved free energy calculation method was highly consistent with the BLI experimental data, demonstrating the ability to accurately distinguish the binding ability of peptides to MHC.
[0152] In some embodiments, step S210 includes:
[0153] The edit distance value between the peptides in the MHC-peptide data and the peptides in the target MHC-peptide is calculated by the edit distance;
[0154] Compare and match the anchor amino acids of the peptides in the MHC-peptide data with the anchor amino acids of the peptides in the target MHC-peptide, and calculate the similarity contribution score by using the mismatch penalty and / or the match score;
[0155] Subtract the similarity contribution score from the edit distance value to obtain the similarity degree;
[0156] The polypeptide with the smallest similarity is selected as the preferred polypeptide.
[0157] In some embodiments, in step S300, the specific steps of obtaining the candidate MHC-peptide include:
[0158] Modification of non-anchor amino acids in the MHC-peptide template;
[0159] Multiple predicted structures are generated, and the predicted structure with the best score is selected as the candidate MHC-peptide.
[0160] In some specific embodiments, the step of modifying the non-anchor amino acids in the MHC-peptide template includes: in the modification process involving multiple non-anchor amino acids, all discontinuous non-anchor amino acids are modified at once, and continuous non-anchor amino acids are modified one by one.
[0161] In some embodiments, step S500 includes:
[0162] S510, selecting amino acids in the candidate MHC-peptide to form an amino acid set;
[0163] S520, performing an alanine scan for each amino acid in the amino acid set to calculate its van der Waals interaction, electrostatic interaction, solvation free energy, and non-solvation free energy;
[0164] S530, calculating the entropy contribution of each amino acid using an interaction entropy calculation method;
[0165] S540. Perform binding free energy calculations based on the amino acid set, including: S541. Summing the energy terms in step S520 and step S530 to obtain the original binding free energy; S542. Extracting the van der Waals interaction energy and summing the entropy contribution to obtain the primary improved binding free energy; S543. Calculating the primary improved binding free energy of the amino acids corresponding to the MHC and polypeptide, respectively, and taking the arithmetic mean thereof as the final improved binding free energy.
[0166] In some specific embodiments, step S510 includes:
[0167] Amino acids whose distance from any atom on the MHC in the candidate MHC-peptide is within a cutoff distance of the polypeptide in the candidate MHC-peptide are selected, and these amino acids and all amino acids on the polypeptide in the candidate MHC-peptide are used as the amino acid set for alanine scanning. The cutoff distance ranges from 5 angstroms to 12 angstroms.
[0168] From the perspective of physics, the effective range of van der Waals interaction (including Lennard-Jones potential) and Coulomb force usually does not exceed Therefore, the distance The atoms above are considered to interact very weakly or not at all;
[0169] From experience, mainstream molecular dynamics simulation software has long-range correction, and the cutoff distance is generally set to Does not affect the calculation results;
[0170] Here you can choose An arbitrary threshold within 500 nm is used as the cutoff distance. A larger value will include more amino acids in the calculation, but at the same time, atoms farther away will have little effect on the results due to their weak interactions. However, the computational effort will increase exponentially, so 5 or 8 are usually used as acceptable thresholds. A value less than 5 is generally considered to have a significant impact on the accuracy of the results.
[0171] In some embodiments, after step S100 and before step S210, the following steps are included:
[0172] Supplementation of missing heavy atoms for MHC-peptide data;
[0173] For the MHC-peptide data, other irrelevant components that aided crystallization, except for MHC and peptide, were removed;
[0174] Filter MHC-peptide data with a resolution below 3 angstroms.
[0175] It is generally believed that: high-resolution structure: The side chain conformation, water molecules and ion coordination details can be clearly analyzed. Medium resolution structure: The main chain conformation is clear, but some side chains (such as long chain residues) may be blurred. Low-resolution structure: Only the main chain direction can be identified, and the side chain position needs to rely on modeling optimization. Since the subsequent steps involve modification and optimization simulation of amino acids, a larger screening threshold can be accepted. Of course, if the threshold is reduced, the quality of the template structure will be better and the calculation results will be more accurate.
[0176] In some embodiments, step S400 includes:
[0177] Before performing molecular dynamics simulations, hydrogen atoms are added to check for atomic overlap. If any atomic overlap occurs, it is repaired manually and / or automatically corrected using the optimization command.
[0178] After ensuring that hydrogen atoms are added and the structure is correct, water molecules are added, and ions are added to the simulation system to make the charge of the simulation system zero;
[0179] Molecular dynamics simulations were performed to obtain the simulation results.
[0180] The present application also discloses an MHC-peptide screening method for screening polypeptides obtained by the above-mentioned MHC-peptide binding free energy calculation method.
[0181] Candidate peptides are ranked and screened based on the improved binding free energy: sorting by improved binding free energy from large to small. Top 5, 10, 15, 30, etc. can be selected as the basis for screening target peptides as needed. Empirically, the improved binding free energy should not be less than 15 kcal / mol, and the free energy contribution of the anchor amino acid should not be less than 8-10 kcal / mol. The RMSD of the peptide should not exceed 2.5 angstroms for a long time during the simulation. Otherwise, it is necessary to observe whether the peptide has escaped from the MHC binding pocket during the simulation. This phenomenon usually indicates that the peptide has low affinity, and in this case the calculated binding free energy will be very low or impossible to calculate. The above screening criteria and methods will improve the accuracy of screening strong binding peptides.
[0182] In some embodiments, polypeptides are screened for improved binding free energy and free energy contribution of the anchor amino acid.
[0183] In this specification, the use of terms such as "Embodiment 1," "this embodiment," and "in one embodiment" indicates that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in the invention or at least one embodiment or example of the invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example; furthermore, the specific features, structures, materials, or characteristics described may be appropriately combined in any one or more embodiments or examples.
[0184] In the description of this specification, the terms "connect," "install," "fix," "dispose," and "have" are to be understood in a broad sense. For example, "connect" can mean a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, or it can be internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on specific circumstances.
[0185] In the description of this specification, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0186] The above description of the embodiments is to facilitate ordinary technicians in this technical field to understand and apply the technology of this case. People familiar with the technology in this field can obviously make various modifications to these examples easily and apply the general principles described here to other embodiments without having to go through creative work. Therefore, this case is not limited to the above embodiments. Modifications to the following situations should all be within the scope of protection of this case: ① A new technical solution implemented based on the technical solution of the present invention and combined with existing common knowledge, the technical effect produced by the new technical solution does not exceed the technical effect of the present invention; ② The equivalent replacement of some features of the technical solution of the present invention with the known technology, the technical effect produced is the same as the technical effect of the present invention; ③ The technical solution of the present invention is expandable, and the substantive content of the expanded technical solution does not exceed the technical solution of the present invention; ④ The equivalent transformation made by the content of the description and drawings of the present invention is directly or indirectly applied to other related technical fields.
Claims
1. A method for calculating MHC-peptide binding free energy, characterized in that: include: S100, acquiring data, acquiring known MHC-peptide data and known structural information of the target MHC-peptide; S210: Determine a preferred polypeptide, perform sequence similarity calculation on the polypeptides in the MHC-peptide data and the polypeptides in the target MHC-peptide, and perform match score calculation on the anchor amino acids of the two, obtain a degree of similarity based on the results of the sequence similarity calculation and the match score calculation, and select the preferred polypeptide with the best match based on the degree of similarity; S220, determining a preferred MHC, selecting the MHC most similar to the MHC in the target MHC-peptide from the MHC-peptide data as the preferred MHC; S230, forming an MHC-peptide template, forming the MHC-peptide template according to the preferred polypeptide and the preferred MHC; S300, performing homology modeling on the MHC-peptide template to obtain a candidate MHC-peptide; S400, performing molecular dynamics simulation on the candidate MHC-peptide to obtain simulation results; S500 , performing binding free energy calculation on the MHC and the polypeptide in the candidate MHC-peptide according to the simulation result to obtain binding free energy.
2. The method for calculating MHC-peptide binding free energy according to claim 1, wherein: In the step S210, it includes: Calculating the edit distance value between the polypeptide in the MHC-peptide data and the polypeptide in the target MHC-peptide by the edit distance; Comparing and matching the anchor amino acid of the polypeptide in the MHC-peptide data with the anchor amino acid of the polypeptide in the target MHC-peptide, and calculating the similarity contribution score by mismatch penalty and / or match score; Subtract the similarity contribution score from the edit distance value to obtain the similarity degree; The polypeptide with the smallest degree of similarity is selected as the preferred polypeptide.
3. The method for calculating MHC-peptide binding free energy according to claim 1, wherein: In step S300, the specific steps of obtaining the candidate MHC-peptide include: modifying non-anchor amino acids in the MHC-peptide template; A plurality of predicted structures are generated, and the predicted structure with the best score is selected as the candidate MHC-peptide.
4. The method for calculating MHC-peptide binding free energy according to claim 3, wherein: The step of modifying the non-anchor amino acids in the MHC-peptide template includes: in the modification process involving multiple non-anchor amino acids, all discontinuous non-anchor amino acids are modified at once, and continuous non-anchor amino acids are modified one by one.
5. A method for calculating MHC-peptide binding free energy according to any one of claims 1 to 4, characterized in that: In the step S500, it includes: S510, selecting amino acids in the candidate MHC-peptide to form an amino acid set; S520, performing an alanine scan for each amino acid in the amino acid set to calculate its van der Waals interaction, electrostatic interaction, solvation free energy, and non-solvation free energy; S530, calculating the entropy contribution of each amino acid using an interaction entropy calculation method; S540. Perform binding free energy calculations based on the amino acid set, including: S541. Summing the energy terms in step S520 and step S530 to obtain the original binding free energy; S542. Extracting the van der Waals interaction energy and summing the entropy contribution to obtain the primary improved binding free energy; S543. Calculating the primary improved binding free energies of the amino acids corresponding to the MHC and polypeptide, respectively, and taking the arithmetic mean thereof as the final improved binding free energy.
6. The method for calculating MHC-peptide binding free energy according to claim 5, wherein: In the step S510, it includes: Amino acids whose distance from any atom on the MHC in the candidate MHC-peptide is within a cutoff distance of the polypeptide in the candidate MHC-peptide are selected, and these amino acids and all amino acids on the polypeptide in the candidate MHC-peptide are used as the amino acid set for alanine scanning. The cutoff distance ranges from 5 angstroms to 12 angstroms.
7. A method for calculating MHC-peptide binding free energy according to any one of claims 1 to 4, characterized in that: After step S100 and before step S210, the following steps are included: Supplement missing heavy atoms to the MHC-peptide data; The MHC-peptide data were removed from other irrelevant components that aided crystallization except for MHC and peptide; The MHC-peptide data were screened for resolution below 3 angstroms.
8. A method for calculating MHC-peptide binding free energy according to any one of claims 1 to 4, characterized in that: In the step S400, it includes: Before performing molecular dynamics simulations, hydrogen atoms are added to check for atomic overlap. If any atomic overlap occurs, it is repaired manually and / or automatically corrected using the optimization command. After ensuring that hydrogen atoms are added and the structure is correct, water molecules are added, and ions are added to the simulation system to make the charge of the simulation system zero; Molecular dynamics simulations were performed to obtain the simulation results.
9. A method for screening MHC-peptides, characterized in that: Used for screening polypeptides obtained by the MHC-peptide binding free energy calculation method according to any one of claims 1 to 8.
10. The MHC-peptide screening method according to claim 9, characterized in that: The polypeptides are screened based on improved binding free energy and free energy contribution of the anchor amino acid.