MHC-peptide binding free energy calculation method and MHC-peptide screening method
Patent Information
- Application Number
- PCT/CN2025/085268
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2026-10-01
Smart Images

Figure CN2025085268_01102026_PF_FP_ABST
Abstract
Description
A method for calculating the binding free energy of MHC-peptides and a method for screening MHC-peptides. Technical Field
[0001] This invention relates to the field of biomedical technology, and in particular to a method for calculating the binding free energy of MHC-peptides and a method for screening MHC-peptides. Background Technology
[0002] Currently, the main method for calculating the affinity between antigen peptides and MHC (major histocompatibility complex) is based on protein sequence-based machine learning techniques. This involves data modeling on large-scale training datasets, using MHC protein and peptide sequences as model inputs, and outputting the model's predicted values as IC50 or KD values determined using competitive binding experiments or wet assays such as SPR and BLI. Alternatively, affinity experimental values can be categorized into strong, intermediate, weak, and non-binding labels using different thresholds such as 50 nmol or 500 nmol. Commonly used software for this type of machine learning prediction includes NetMHC, NetMHCpan, and MHCflurry. However, they share the following drawbacks: 1. Training these machine learning or artificial intelligence methods relies heavily on large amounts of training data. When training data is insufficient or biased, the generalization performance of these methods will decrease; 2. These methods are affected by the poor detection accuracy of low-affinity detection data in experimental data, resulting in poor prediction performance of affinity values. Therefore, it is common practice to transform the regression problem into a classification problem during the training or output phase to achieve better performance. For example, the NetMHC software series sorts the affinity results predicted by the model in a reference dataset on one side and suggests that users use the ranking value (Rank 0.5% and 2%) as the screening index for strong and weak binding antigen peptides in order to avoid the problem that the relative affinity prediction is not accurate.
[0003] However, in practical applications, such as the development of therapeutic cancer vaccines, the affinity of the selected antigenic peptide directly affects the vaccine antigen capacity, immune response rate, and incidence of autoimmune reactions. Therefore, accurate calculation methods for predicting antigenic peptide affinity are of great significance for immunotherapy development.
[0004] Existing technical solutions include sequence-based machine learning methods and structure-based computational methods. Wet experimental methods, such as competitive binding experiments, BLI, and SPR detection methods, typically have low throughput, long cycles, and high costs.
[0005] Sequence-based machine learning methods mainly involve two steps: 1. Inputting MHC protein and peptide sequences to construct features, including encoding amino acids or atoms as features and selecting important sequence features; 2. Using a series of machine learning algorithms such as logistic regression, support vector machine (SVM), random forest, and deep learning algorithms such as CNN, GNN, LSTM, and Transformer, the experimentally determined affinity values are used as regression tasks, or the affinity values are converted into classification tasks based on thresholds for modeling.
[0006] These sequence-based machine learning methods are often affected by data scale and bias. For example, the most common HLA-A*02 in the population has the most historical data, so machine learning models usually perform better than those for rarer HLA types. Specifically, for common HLA types, including HLA-A*02, such as HLA-A*06 and HLA-A*11, most prediction models can achieve an accuracy of over 0.9 and an AUC of over 0.8. However, for rare types such as HLA-A*26, the accuracy and AUC of prediction models are usually below 0.5. Considering the polymorphism and structural similarity of MHC proteins, this situation indicates that current machine learning methods are only overfitting on partial datasets, rather than truly learning the general rules of interaction between MHC and peptides. Studies have shown that solid-phase peptide synthesis methods, such as mass spectrometry, cannot capture the binding of certain peptides to MHC in vitro. For example, when a peptide has a V-terminus, the intracellular peptide binding detection method Episcan can capture this interaction, but it cannot be detected in in vitro experiments. When there are too many hydrophobic amino acids in a peptide, the peptide's hydrophobicity increases and its solubility decreases, which makes peptide synthesis, dissolution, and detection experiments difficult. Furthermore, the peptide binding pocket of MHC tends to be more inclined towards hydrophobic amino acid anchor sites. These factors can all cause bias in historical data, leading to bias and overfitting in machine learning methods that rely on large-scale datasets.
[0007] Model development that incorporates structural information into machine learning methods typically involves feature selection to reduce the dimensionality of input features, thereby decreasing the complexity of the learning task. Alternatively, structural information can be used to refine or embed features at the amino acid level at the atomic level to improve predictive model performance, but these methods have not yet yielded significant progress.
[0008] First-principles molecular dynamics simulations can fully utilize the structural information of MHC-peptide interactions to calculate physicochemical interaction energies at the atomic level, thereby calculating the binding free energy of protein-protein interactions. Commonly used molecular dynamics simulation methods decompose the binding free energy into van der Waals forces, electrostatic interactions, solvation free energy, and interaction entropy, simulating the MHC-peptide interaction process using force fields on picosecond or nanosecond scales. Commonly used software and energy calculation methods include Amber, NAMD, Gromacs, and MMGBSA / PBSA. Traditional molecular dynamics calculation methods are limited in simulating protein-protein / protein-peptide systems, mainly due to the lack of accurate force fields to describe the complex interactions of MHC-peptide complexes and the lack of reliable free energy prediction methods. These factors are crucial for accurately predicting affinity, and improving affinity prediction is key to accelerating immunotherapy drug design and treatment. Currently, MHC-peptide binding free energy calculations based on physics principles rely on empirical molecular force fields such as AMBER, CHARMM, and OPLS, but these force fields neglect key quantum effects, such as electrostatic polarization and charge transfer, limiting computational accuracy. Studies have shown that using quantum mechanical (QM) methods to handle electrostatic interactions can significantly improve the accuracy of protein-ligand docking. Furthermore, the molecular force field of the point charge model has limitations in handling interactions between closely spaced atoms, making it difficult for current force-field-based calculation methods to improve the computational accuracy of MHC-peptide complexes. Currently, the most commonly used methods for calculating protein-ligand binding free energy are the MM / GBSA and MM / PBSA methods. These two methods use an approximate normal vibrational method to calculate the contribution of entropy change to the binding free energy. The two main drawbacks of this method are: (1) very high computational cost; and (2) poor reliability of the calculation results. Therefore, in most calculations of protein-ligand binding free energy, the contribution of entropy change is often ignored. Higher-precision calculation methods such as FEP and TI, as well as quantum chemical calculation methods, are currently impractical due to their excessively high computational costs.
[0009] In summary, while existing machine learning-based methods for calculating MHC-peptide affinity can make predictions to some extent, they have several significant drawbacks:
[0010] 1. Reliance on large amounts of training data: Existing machine learning models require a large amount of training data to ensure prediction accuracy. If the training data is insufficient or imbalanced, the model's generalization ability will decrease significantly, leading to inaccurate or unstable predicted affinity results. Furthermore, insufficient experimental data particularly affects the prediction of low-affinity peptides, making it difficult to cover all possible scenarios.
[0011] 2. Data Bias and Model Bias: Since the training data for existing models are usually derived from experimental datasets, which may contain measurement errors or selection biases, the models are prone to overfitting or generating systematic errors when the data distribution is uneven, especially in the prediction of low affinity or certain specific peptides.
[0012] 3. Limited quantitative prediction of affinity: Most existing technologies can only provide qualitative predictions of affinity, while lacking precision in quantitative calculations. In particular, the accuracy and application scope of existing methods remain limited in affinity ranking and specific numerical predictions of affinity values. Summary of the Invention
[0013] To address the shortcomings of existing machine learning-based methods for calculating MHC-peptide affinity, such as reliance on large amounts of training data, data bias and model deviation, and limited quantitative prediction of affinity, this invention proposes a method for calculating MHC-peptide binding free energy.
[0014] The technical solution adopted in this invention is a method for calculating the MHC-peptide binding free energy, comprising:
[0015] S100. Acquire data, including known MHC-peptide data and known structural information of the target MHC-peptide.
[0016] S210. Determine the preferred peptide: perform sequence similarity calculation on the peptides in the MHC-peptide data and the peptides in the target MHC-peptide, and perform matching and scoring calculation on the anchor amino acids of the two. Obtain the similarity degree based on the calculation results of sequence similarity calculation and matching and scoring calculation, and select the best matching preferred peptide based on the similarity degree.
[0017] S220. Determine the preferred MHC: Select the MHC that is most similar to the MHC in the target MHC peptide from the MHC-peptide data as the preferred MHC.
[0018] S230, Forming an MHC-peptide template: Forming an MHC-peptide template based on a preferred polypeptide and a preferred MHC;
[0019] S300. Homology modeling is performed on MHC-peptide templates to obtain candidate MHC-peptides;
[0020] S400. Perform molecular dynamics simulations on the candidate MHC-peptide and obtain the simulation results;
[0021] S500. Based on the simulation results, calculate the binding free energy of MHC and polypeptide in the candidate MHC-peptide to obtain the binding free energy.
[0022] Preferably, step S210 includes:
[0023] Edit distance is calculated between peptides in MHC-peptide data and peptides in target MHC-peptides.
[0024] The anchor amino acids of peptides in MHC-peptide data are compared and matched with the anchor amino acids of peptides in target MHC-peptides, and the similarity contribution score is calculated by mismatch penalty and / or matching score.
[0025] The similarity score is obtained by subtracting the similarity contribution score from the edit distance value;
[0026] The peptide with the lowest similarity is selected as the preferred peptide.
[0027] Preferably, in step S300, the specific steps for obtaining the candidate MHC-peptide include:
[0028] Modify the non-anchor amino acids in the MHC-peptide template;
[0029] Multiple predicted structures are generated, and the predicted structure with the best score is selected as the candidate MHC-peptide.
[0030] Preferably, the step of modifying the non-anchor amino acids in the MHC-peptide template includes: in the process of modifying multiple non-anchor amino acids, modifying all discontinuous non-anchor amino acids at once, and modifying continuous non-anchor amino acids one by one.
[0031] Preferably, step S500 includes:
[0032] S510. Select amino acids from candidate MHC-peptides to form an amino acid set;
[0033] S520. Perform a scan of alanine in the amino acid set to calculate its van der Waals interactions, electrostatic interactions, solvation free energy and nonsolvation free energy.
[0034] S530. Calculate the entropy contribution of each amino acid using the interaction entropy calculation method;
[0035] S540. Calculate the binding free energy based on the amino acid set, including: S541. Summing up the energy terms in steps S520 and S530 to obtain the original binding free energy; S542. Extracting the van der Waals interaction energy and the entropy contribution and summing them to obtain the primary improved binding free energy; S543. Calculating the primary improved binding free energy of the corresponding amino acids for MHC and peptides respectively, and taking their arithmetic mean as the final improved binding free energy.
[0036] Preferably, step S510 includes:
[0037] Select any amino acid on the MHC atom of the candidate MHC peptide that is within the cutoff distance of the peptide in the candidate MHC peptide. Collect these amino acids together with all amino acids on the peptide in the candidate MHC peptide as the amino acid set for alanine scanning. The cutoff distance ranges from 5 Å to 12 Å.
[0038] Preferably, after step S100 and before step S210, the procedure includes:
[0039] To supplement missing heavy atoms in MHC-peptide data;
[0040] Remove irrelevant components that aid crystallization other than MHC and peptides from MHC-peptide data;
[0041] Screen MHC-peptide data with a resolution below 3 angstroms.
[0042] Preferably, step S400 includes:
[0043] Before performing molecular dynamics simulations, hydrogen atoms are added to check for atomic overlap. If atomic overlap is found, it is manually repaired and / or automatically corrected through optimization commands.
[0044] After ensuring that the addition of hydrogen atoms and structural corrections were correct, water molecules were added, and ions were added to the simulation system to make the charge of the simulation system zero.
[0045] Perform molecular dynamics simulations to obtain simulation results.
[0046] To address the limitation of failing to screen for peptides that strongly bind to MHC, this invention proposes an MHC-peptide screening method for screening peptides obtained by the aforementioned MHC-peptide binding free energy calculation method.
[0047] Preferably, peptides are screened based on improved binding free energy and the free energy contribution of anchor amino acids.
[0048] Compared with the prior art, the present invention has the following beneficial effects:
[0049] This application discloses a method for calculating the binding free energy of MHC-peptides, which utilizes template construction technology to obtain the initial structure of any unknown candidate MHC-peptide by adjusting the amino acids in known MHC-peptide data. Constructing an accurate MHC-peptide template is a prerequisite for accurate affinity calculation. This application employs a template selection method based on anchor amino acids, aiming to select templates through highly conserved structural features of the binding groove (such as hydrophobic pockets and α-helix-β-sheet combinations). This method differs from traditional template matching methods based on sequence similarity, particularly emphasizing the selective binding of the first and last amino acids (anchor amino acids) of the peptide to both ends of the MHC binding groove, thus contributing to improved modeling accuracy and reliability.
[0050] In the complex structure of MHC-peptides, the peptide binds to two alpha helices formed by MHC at positions 65–85 and 90–120, and to a binding groove at the bottom of the beta-sheet. This groove can accommodate 8–12 amino acids of peptide and is relatively stable. Benefiting from the fixed binding position of MHC-bound peptides, the RMSD (root mean square deviation) of the peptide's spatial position relative to MHC is less than 1 Å in different MHC-peptides. Furthermore, the smaller the edit distance of the peptide sequence, the smaller the RMSD. Therefore, by using the structure of known MHC-peptide data, the initial structure of any unknown candidate MHC-peptide can be obtained by modifying some amino acids on the peptide. In the experimental dataset, which includes 327 peptides with reliable IC50 experimental data, the correlation between theoretical affinity calculations and experimental values is 0.55, especially for peptides with an edit distance of less than 2 Å, where the correlation is even higher, reaching 0.76.
[0051] This result demonstrates that by appropriately selecting templates and performing suitable data correction, affinity calculation results consistent with experimental observations can be obtained. These findings not only make a significant contribution to the advancement of immunology but also provide strong support for the design and development of peptide-based vaccines and immunotherapies. Furthermore, it avoids the drawbacks of machine learning-based MHC-peptide affinity calculation methods that rely on large amounts of training data, preventing data bias and model deviation during calculation and offering superior quantitative affinity prediction capabilities. Compared to existing technologies, the MHC-peptide binding free energy calculation method disclosed in this application demonstrates its powerful performance in peptide affinity prediction, laying the foundation for further understanding of immune recognition mechanisms and enhancing research on MHC-peptide interactions.
[0052] This application also discloses an MHC-peptide screening method, which screens peptides obtained by the above-mentioned MHC-peptide binding free energy calculation method, thereby enabling the screening of peptides with higher affinity.
[0053] Compared with existing technologies, the MHC-peptide screening method disclosed in this application can improve the accuracy of screening for strongly binding peptides. Attached Figure Description
[0054] The present invention will now be described in detail with reference to the embodiments and accompanying drawings, wherein:
[0055] Figure 1 shows a flowchart of a method for calculating the MHC-peptide binding free energy according to an embodiment of the present invention;
[0056] Figure 2 shows a flowchart of a method for calculating the MHC-peptide binding free energy according to an embodiment of the present invention;
[0057] Figure 1a shows a comparison of the energy terms of the template structure, the constructed 4K7F, and the negative control pMHC complex in one embodiment.
[0058] Figure 1b shows a record of the score distribution for peptide template selection and detailed information on template alignment in one embodiment;
[0059] Figure 2a shows the enthalpy energy distribution of the peptide in one embodiment;
[0060] Figure 2b shows a comparison of correlation coefficients between templates and with the negative control in one embodiment;
[0061] Figure 2c shows the RMSD (root mean square deviation) fluctuation range of the peptide portion in one embodiment;
[0062] Figure 2d shows a comparison of RMSD (root mean square deviation) in one embodiment;
[0063] Figure 2e shows the different unstable kinetic behaviors of two negative control peptide structures constructed in one embodiment;
[0064] Figure 3a shows the correlation distribution between the calculated results of the van der Waals forces and interaction entropy of the peptides in one embodiment and the experimental results.
[0065] Figure 3b shows the correlation distribution of the sum of van der Waals forces and interaction entropies of peptide amino acids in one embodiment with experimental values.
[0066] Figure 4a shows a peptide energy distribution diagram in one embodiment;
[0067] Figure 4b shows the energy distribution of MHC hotspot amino acids in one embodiment;
[0068] Figure A shows a table of peptide chain sequences, theoretical calculations, and ELISA results in one embodiment;
[0069] Figure B shows a graph illustrating a positive correlation between theoretical energy and affinity measured by the BLI experiment in one embodiment. Detailed Implementation
[0070] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Examples of embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar components or components having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.
[0071] This invention discloses a method for calculating the MHC-peptide binding free energy, comprising:
[0072] S100. Acquire data, including known MHC-peptide data and known structural information of the target MHC-peptide.
[0073] S210. Determine the preferred peptide: perform sequence similarity calculation on the peptides in the MHC-peptide data and the peptides in the target MHC-peptide, and perform matching and scoring calculation on the anchor amino acids of the two. Obtain the similarity degree based on the calculation results of sequence similarity calculation and matching and scoring calculation, and select the best matching preferred peptide based on the similarity degree.
[0074] S220. Determine the preferred MHC: Select the MHC that is most similar to the MHC in the target MHC peptide from the MHC-peptide data as the preferred MHC.
[0075] S230, Forming an MHC-peptide template: Forming an MHC-peptide template based on a preferred polypeptide and a preferred MHC;
[0076] S300. Homology modeling is performed on MHC-peptide templates to obtain candidate MHC-peptides;
[0077] S400. Perform molecular dynamics simulations on the candidate MHC-peptide and obtain the simulation results;
[0078] S500. Based on the simulation results, calculate the binding free energy of MHC and polypeptide in the candidate MHC-peptide to obtain the binding free energy.
[0079] Step S100 involves acquiring data. MHC-peptide data can be obtained from public databases such as PDB, IMGT, TCR3d, and VDJdb. MHC-peptide data is typically obtained through techniques such as cryo-electron microscopy and crystal diffraction, and is stored as a PDB file or other file formats. The collected MHC-peptide data's PDB structure file will be used in the next step. Preferably, after acquiring the MHC-peptide data, these template files need to undergo some preprocessing, such as supplementing missing heavy atoms, removing irrelevant crystallization aids other than MHC and peptides from the MHC-peptide data, and removing low-quality templates according to crystal resolution. A commonly used resolution for removing low-quality templates according to crystal resolution is below 2-3 Å.
[0080] The step S200 to obtain the MHC-peptide template includes step S210 to determine the preferred polypeptide, step S220 to determine the preferred MHC, and step S230 to form the MHC-peptide template.
[0081] Step S210 determines the preferred peptide based on the optimal template matching method using anchor amino acids. The target MHC-peptide to be calculated includes both the target peptide and its corresponding MHC. Therefore, the more similar the selected template is to the target MHC-peptide, the fewer amino acids need to be modified in step S300, and the more accurate the result. Most current homology modeling methods use sequence alignment, selecting templates based on sequence similarity algorithms such as edit distance, PAM, and BLOSUM scoring matrix. In this invention, because the binding mode of MHC to peptides is highly conserved, the peptide binds to a structurally stable MHC binding groove. This binding groove consists of a bottom composed of a pair of alpha helices and a set of beta sheets, forming a binding pocket that favors hydrophobic amino acids. This conserved structure exists across different types of MHC. Several amino acids at the beginning and end of the peptide can selectively bind to the two ends of the binding groove due to their physicochemical properties; these are called anchor amino acids. It is generally believed that the amino acids in the middle of the peptide have less interaction with the MHC. Therefore, this application optimizes the alignment and template selection method based on sequence similarity. It adds a mismatch penalty or matching score for anchor amino acids to the sequence similarity algorithm. For example, when the target sequence and the peptide in the candidate template are the same in the 1st, 2nd, 3rd or the 1st and 2nd from the end, a matching score of 1 is given. In particular, when the 2nd amino acid is the same, a score of 2 is given.
[0082] The specific method for peptide similarity scoring is as follows: the similarity between two peptide sequences can be evaluated using the Levenshtein edit distance. When selecting the most suitable MHC-peptide template for each peptide, templates with the same anchor amino acids as the target peptide are preferred. This is because using these templates avoids changing the anchor amino acids during the construction of candidate MHC-peptides, thereby minimizing the interference that amino acid modifications may cause to energy calculations. To optimize this method, the Levenshtein edit distance between each pair of peptides is first calculated, and the amino acids that need to be replaced, inserted, and deleted are considered separately, retaining only those templates that only need to be replaced. Next, the similarity of anchor amino acids between the target peptide and the template peptide is evaluated. Anchor amino acids are defined as the first three and last two amino acids of the peptide. If there are n matching amino acids, the similarity contributes n points.
[0083] The final template matching score, called the MatchScore, is calculated by subtracting the anchor amino acid similarity score from the Levenshtein edit distance. Each peptide is then matched to the template that results in the lowest MatchScore, indicating the closest edit distance and the highest anchor amino acid similarity. This optimized matching score ensures that the selected MHC-peptide template is best suited for constructing an MHC-peptide with minimal deviation from the original sequence, thereby improving the accuracy of subsequent affinity calculations and predictions.
[0084] Modified design of template matching method based on anchor amino acids:
[0085] The peptide sequence matching method can be changed to other sequence matching scoring algorithms such as PAM and BLOSUM matrix, and the implementation software can be replaced with other commonly used sequence matching algorithms and software such as blastp and clustalw;
[0086] The matching score or penalty score of anchor amino acids can be adjusted in detail, such as selectively including the first 1-3 amino acids of the peptide in the calculation, and different scoring or penalty score changes can be made according to the position.
[0087] Step S230 determines the preferred MHC. The selection of preferred MHC is based on the HLA nomenclature rule to screen the MHC with the most similar gene family, or based on the similarity calculation method to screen the most similar MHC from MHC-peptide data, with priority given to MHC that is exactly the same as the target MHC.
[0088] Step S230 forms an MHC-peptide template by combining a preferred MHC and a preferred polypeptide. When considering which amino acids need to be replaced, inserted, or deleted, the rule is to modify only a few amino acids in the middle of the polypeptide in subsequent step S300, without altering the anchor amino acids, and without modifying the amino acids in the binding groove of the MHC. Based on these rules, an MHC-peptide is selected as the MHC-peptide template.
[0089] This invention employs a template selection method based on anchor amino acids, aiming to select templates by leveraging highly conserved structural features of binding grooves (such as hydrophobic pockets and α-helix-β-sheet combinations). This method differs from traditional template matching methods based on sequence similarity and has the following innovative aspects:
[0090] Conservation of binding groove: Template selection is performed by utilizing the conserved structure of the peptide-MHC binding groove, which reduces unnecessary amino acid modifications and improves the accuracy of the model.
[0091] Anchor amino acid selection: Special emphasis is placed on the selective binding of the first and last amino acids of the peptide (anchor amino acids) to both ends of the MHC binding groove, which helps to improve the accuracy and reliability of modeling.
[0092] By leveraging the conserved properties of anchor amino acids, this invention proposes a precise method for modeling MHC-peptides. This method utilizes the limited interaction between intermediate amino acids and MHC molecules to reduce the impact of amino acid modifications, focusing on preserving the binding regions at the head and tail of the peptide. Furthermore, it optimizes the side chain energy locally after modifying a single amino acid, thereby improving the rationality of MHC-peptide binding throughout the modeling process.
[0093] Step S300 obtains candidate MHC peptides. Based on homology modeling, software such as Modeller and Rossetta can be used. Necessary amino acid modifications are made to the selected MHC peptide template from the previous step to construct the candidate MHC peptide. When modifying amino acids, if multiple amino acids are involved, they can be modified all at once or stepwise. Generally, the choice between these two methods is based on the following conditions: when the amino acids to be modified are discontinuous or spatially distant, all can be modified at once; when the amino acids to be modified are close together, they can be modified one by one. After each amino acid modification, energy-minimizing side-chain optimization is required for that amino acid, while keeping the coordinates of other unmodified amino acids unchanged. Different software will output multiple predictive structures as candidate MHC peptides; usually, the predictive structure with the best score is selected based on the corresponding evaluation score.
[0094] The specific steps for constructing the initial structure of candidate MHC peptides are as follows:
[0095] Using Modeller version 10.2, each target peptide and the selected MHC-peptide template PDB structure were input into the software. Modeller performed sequence alignment between the target peptide and the MHC-peptide and constructed an initial protein structure. Subsequently, this structure was optimized through side chain adjustments and energy minimization, generating five predicted final structures.
[0096] To evaluate these structures, the DOPE (Discrete Optimized Protein Energy) score and the GA341 score were used to assess the structural and energy feasibility of the modeled complexes. The structure with the highest score was ultimately selected as the final candidate MHC-peptide for the target peptide, thus ensuring optimal structural prediction and accuracy of potential biological function.
[0097] Modification design of target pMHC structure modeling method based on homology modeling: It can be done using homology modeling software including Modeller and Rosetta, or by modifying atoms using software including Amber, CHARMM-GUI, and Gromacs to achieve homology modeling, or by using programming languages such as Python and Fortran to write custom program scripts to modify and adjust atoms in PDB files, and then use molecular dynamics methods for local structure optimization.
[0098] Step S400 involves obtaining simulation results. Specifically, system preparation is performed before conducting molecular dynamics simulations. First, the tLEap tool in Amber software is used for system preparation. Launch tLEap and load the required molecular structure files, such as PDB format molecular structure files, using the loadPdb command. tLEap will initialize the system based on these files and prepare for subsequent simulations. Next, hydrogen atoms are added. In most PDB files, hydrogen atoms are usually not included, so the addHydrogens command in tLEap needs to be used to add them to the system. Hydrogen atoms are crucial for simulation accuracy because they affect the molecular structure and dynamic behavior. After adding hydrogen atoms, it is necessary to check for atomic overlap, which can be verified using the check command. If overlap is found, tLEap will provide feedback, and the user can manually repair it based on the prompts or automatically correct it using optimization commands (such as minimize).
[0099] After ensuring the addition of hydrogen atoms and structural corrections are correct, the next step is to add a solvent box. Use the `solvateBox` command in tLEap to add water molecules to the system. Commonly used water models include TIP3P and SPC / E. When adding water molecules, ensure the water box is large enough to avoid intermolecular boundary effects; it is generally recommended that the water box size be larger than the size of the molecules themselves. In addition, to simulate physiological conditions and ensure charge balance in the system, ions need to be added. Use the `addIons` command in tLEap to add sodium ions (Na+) and chloride ions (Cl-). By setting appropriate ion concentrations and bringing the system charge to zero, ensure the simulated charge balance. If the total charge of the system is not zero, tLEap will automatically correct it.
[0100] After these steps are completed, the system is ready for molecular dynamics simulations. Next, the system can be pre-equilibrated and heated, typically by gradually increasing the temperature from a low temperature to the target temperature to avoid non-physical behaviors caused by sudden temperature changes. After equilibrium and heating, a short simulation is performed to bring the system to a state of energy minimization, followed by the formal simulation process. This process typically lasts 1-50 nanoseconds and involves 3-5 trajectory repetitions to ensure the reliability of the simulation results.
[0101] The specific procedures and parameter settings for molecular dynamics simulations are as follows:
[0102] The candidate MHC-peptide complex was parameterized using an AMBERff14SB force field. The complex was dissolved in a periodic box filled with TIP3P water molecules, ensuring a minimum buffer distance between any atom in the complex and the edge of the box. To neutralize the system's charge, appropriate amounts of sodium or chloride ions were added. Long-range electrostatic interactions were handled using the Particle Mesh Ewald (PME) algorithm, with the cutoff distance for non-bonded interactions set to...
[0103] Energy minimization was performed on each system using the steepest descent method and the conjugate gradient method. Initially, the complex was subjected to... Force constants were initially imposed to allow solvent molecules to tune around the complex; these constraints were then removed to optimize the entire system, particularly addressing potential atomic collisions caused by abrupt changes. The system was then slowly heated to the target temperature of 300 K (over 300 ps) in the NVT ensemble, followed by a 1 ns simulation in the NPT ensemble, with continuous constraints applied to all atoms of the complex. Temperature was determined using Langevin dynamics (collision frequency 1.0 ps). -1 The pressure is regulated by the / ) and maintained at 1.0 atm through the Berendsen constant pressure tank.
[0104] Finally, a 50ns unconstrained production simulation was performed, with trajectory data saved every 1ps. This 50ns trajectory was used for combined free energy calculation, all 50,000 saved frames were used for entropy calculation, and 1,000 frames were evenly selected from them for enthalpy calculation.
[0105] Modification of the molecular dynamics simulation process:
[0106] The software used to perform molecular dynamics simulations can be changed, such as Amber, Gromacs, NAMD, CHARMM-GUI, or their derivatives;
[0107] The specific settings for the simulation process, such as the choice of specific force field, water box type, shape and size, each energy optimization, and specific parameters of the simulation process, such as simulation duration and step size.
[0108] Step S500 obtains the binding free energy. Based on the ASGBIE method and an improved binding free energy calculation, the alanine scan is a method used in wet experiments to study the contribution of a specific amino acid to the interaction. In the calculation, the amino acids that contribute to the interaction in the complex trajectory obtained in the previous step are modified one by one to alanine. Since the alanine side chain is very small, it can be considered that it does not interact with the surrounding amino acids. By calculating the difference in free energy before and after the modification, the specific value of the contribution of the amino acid at that position to the binding free energy can be obtained.
[0109] Specifically, firstly, select any amino acid on the MHC protein of the candidate MHC-peptide that is within 5 angstroms of the polypeptide. Then, combine these amino acids with all the amino acids on the polypeptide of the candidate MHC-peptide to form the set of amino acids that need to be scanned for alanine.
[0110] For each amino acid, the van der Waals interaction, electrostatic interaction, solvation free energy, and non-solvation free energy were calculated using the MMGBSA method for alanine scans.
[0111] The entropy contribution of each amino acid was calculated using the developed interaction entropy calculation method.
[0112] The improved binding free energy is calculated as follows: Specifically, this involves calculating each energy term from the above steps separately using the amino acid sets on both MHC and the peptide; then, the sum of all energy terms is calculated as the original binding free energy, a method conforming to standard practice in the field; the sum of van der Waals interactions and entropy values is used as the "improved binding free energy," a method that eliminates bias from electrostatic interactions and incorporates entropy values, consistent with improved methods for MHC-peptide systems; the average of the "improved free energies" of the amino acid sets on both MHC and the peptide is used as the "final improved binding free energy." Simultaneously, the sum of the "improved binding free energies" of the peptide anchor amino acids is calculated as a subsequent reference indicator.
[0113] The specific calculation formula for the ASGBIE method is as follows:
[0114] In each MHC-peptide system, the binding energy was assessed using the ASGBIE (Generalized Born Model and Alanine Scan of Interaction Entropy) method. The contribution of a specific mutant residue (x) to the binding energy was determined using the alanine scan technique, where the difference in binding free energy between the alanine mutant and the wild-type complex reflects the effect of the mutation, as shown in Equation 1. Equation 1 calculates the binding free energy of the alanine mutant and the wild-type complex:
[0115] Furthermore, Equation 1 can be broken down into different components, as shown in Equation 2. Although the MHC-peptide interface is typically composed of a large number of residues, usually only a few key residues play an important role in binding. Assuming the distance from the interface exceeds... The effect of the residues on the interaction is negligible. Therefore, the overall binding energy is determined by summing the MHC or intrapeptide distance. The contribution of each residue within the range is calculated as shown in Equation 3:
[0116] By scanning MHC (MHC-AS) or peptide (Peptide-AS) residues for alanine residues, hotspot residues crucial for binding can be identified, and their individual contributions to the overall binding free energy can be calculated. This method provides a detailed analysis of the energy contributions in MHC-peptide complexes, contributing to a deeper understanding of intermolecular interaction mechanisms.
[0117] Design changes incorporating free energy calculation methods:
[0118] The ASGBIE method used in this application can be replaced by other combined free energy calculation methods such as MMGBSA / MMPBSA.
[0119] The improved binding free energy method used in this application can be replaced with other energy term addition methods, assigning different weights to each energy term as a specific adjustment strategy to achieve better consistency with experimental values.
[0120] This invention further improves upon the structural modeling of target MHC-peptides by employing a novel and improved method for calculating binding free energy (ASGBIE method). This method can accurately calculate the binding free energy between the peptide and MHC, and can account for the influence of interaction entropy. It calculates the binding free energy separately from the amino acids on both MHC and the peptide, and the improved calculation method takes into account the energy contribution of their interaction, eliminating the inaccuracies of traditional electrostatic interaction calculations. This helps predict and evaluate the stability of different MHC-peptide bindings and optimize the molecular design related to immune responses.
[0121] Example 1: Verification of the multi-template calculation method for 4K7F.
[0122] To validate the template-based pMHC (wherein this application, "pMHC" refers to "MHC-peptide") structure construction method, novel structures were predicted using pMHC structures from the PDB database. First, 4K7F was selected as the prediction target, and five other known pMHC structures (1I7U, 3H7B, 5HHQ, 5FA3, and 5E00) were chosen from the PDB as templates. All selected known and target structures belong to the HLA-A02:01 type MHC, but their peptide sequences differ by 6 to 8 amino acids (see Table 1). This selection facilitates the study of the impact of different edit distances on the prediction of the target pMHC structure. Simultaneously, the 3OXS structure was selected for molecular dynamics simulations to analyze and compare the binding modes of peptides composed of 10 and 9 amino acids.
[0123] Simulations were performed for 50 nanoseconds on known 4K7F structures selected from the PDB and predicted pMHC structures constructed using Modeller software. Throughout the simulations, the RMSD (root mean square deviation) of the peptide moieties of both the selected known 4K7F structures and the 4K7F pMHC structures constructed using 1I7U, 3H7B, and 5FA3 as templates fluctuated within approximately 1 Å (Fig. 2c), demonstrating good structural stability and high consistency with known 4K7F structures (Fig. 2d). The average RMSD of the peptides in the 4K7F pMHC structures constructed using 5E00 and 5HHQ as templates was slightly higher, approximately 1.5 Å. Stability analysis confirmed that template-based molecular dynamics simulations can effectively predict and validate the structure and dynamic interactions of pMHC complexes.
[0124] As shown in Figure 1a, the ASGBIE method was used to calculate the overall energy contribution of hotspot amino acids in MHC molecules and peptides. Comparative analysis of the predicted target structure and the reference structure 4K7F showed that all assessed energy terms exhibited high similarity, with van der Waals interactions of MHC hotspot amino acids at approximately 45 kcal / mol, an enthalpy of approximately 63 kcal / mol, and a total energy of 23 kcal / mol. These energy terms were significantly better than those of the two negative control peptides, supporting the hypothesis that the predicted structure possesses interaction stability comparable to known structures. In contrast, the constructed negative control pMHC complex exhibited insufficient interaction stability.
[0125] In the hotspot amino acid analysis of the peptides, the calculated van der Waals interactions, enthalpy values, and overall energy contributions also showed good consistency between the reference and target structures. Figures 2a and 4a show that the peptides exhibited binding patterns in each template structure, with the second, third, and terminal amino acids contributing the most to the interaction energy, while the central amino acids of the peptides exhibited different interaction patterns between different templates. Notably, the first amino acid of the 9-peptide of 3H7B and the 10-peptide of 3OXS showed strong interactions. The energy distribution of peptides derived from various template mutations was highly consistent with the energy distribution of the classic 4K7F structure, indicating that the energy distribution generated by selecting any template with 5 to 8 amino acid mutations is similar to the original 4K7F structure without changing the peptide binding pattern.
[0126] Based on the unique energy distribution of the reference template used in the modeling, this result confirms that the peptide energy distribution derived in molecular dynamics simulations is less affected by template selection. Correlation analysis based on peptide interaction patterns further reveals that the energy distribution patterns of the 4K7F constructs are more similar to those of the known target 4K7F structure compared to using a template alone. The Pearson correlation coefficients between the energy distribution patterns of 4K7F constructs generated from different templates and the reference structure exceed 0.9, showing strong similarity, while the correlation coefficients with their respective original templates are below 0.8, and the correlation coefficients between templates and with the negative control are all below 0.7 (Figure 2b). This pattern is also reflected in the energy distribution of MHC hotspot amino acids (Figures 2b and 4b).
[0127] Example 2: Calculation and verification of negative peptides.
[0128] To evaluate the binding effect of peptides with low MHC binding capacity, a negative control peptide (SRYWAIRTR) was selected. This peptide was experimentally shown to have an affinity for HLA-A02:01 of less than 50,000 nM, and a new pMHC structure was constructed using it.
[0129] In molecular dynamics simulations, the two constructed negative control peptide structures exhibited different unstable kinetic behaviors (Fig. 2e). In one case, the first half of one of the negative peptides migrated out of the MHC binding groove during the simulation, indicating a lack of sufficient interaction between the anchored amino acid and MHC, leading to binding instability. This phenomenon is consistent with the peptide's larger RMSD (Fig. 2d). In another case, although the peptide's position remained relatively unchanged and its RMSD did not exceed 2 Å, the two α-helices of the MHC binding groove expanded outward, making the groove looser, reflecting an impact on the structural stability of the MHC binding groove itself (Fig. 2d). This change may be due to some repulsive forces induced by the negative control peptide, further affecting the structural stability of the entire complex. This structural change related to peptide stability highlights the importance of considering peptide-MHC interactions when assessing the affinity and stability of pMHC.
[0130] In the hotspot amino acid analysis of the peptide, the van der Waals interactions, enthalpy, and total energy contribution of the negative control peptide did not show significant differences from the other peptides in Example 1. This may be due to the smaller number of amino acids in the peptide and the variation in the amino acid composition of the peptide, which are more likely to affect the accuracy of energy calculations.
[0131] Both negative control peptides exhibited significant differences in MHC and peptide energy distribution compared to any structure in Example 1, as well as between themselves. This indicates that, although the accuracy of total peptide energy calculations via molecular dynamics is not perfect, analysis of peptide binding patterns at the amino acid level reflects the MHC-peptide interaction mechanism. Accurate energy distribution descriptions not only form the basis of MHC-peptide affinity but also demonstrate the potential to distinguish between positive and negative peptides through hotspot amino acid energy analysis.
[0132] Example 3: Validation of large-scale simulation calculations of the 426 peptide.
[0133] A publicly available benchmark dataset containing affinity data for 426 peptides was used, specifically designed for evaluating peptide affinity predictions. The corresponding known pMHC structure templates were obtained from the Protein Data Bank (PDB), from which 187 distinct pMHC structures were selected from 397 HLA-A type pMHCs, all classified as HLA-A*0201 type MHCs. Of these structures, 12 out of 426 peptides had known structures. To determine the optimal template for each peptide prediction, a peptide similarity scoring method based on edit distance, as described in a signed patent scheme, was used. This method combines the edit distance of the peptide with the weights of the anchor amino acids, enabling the selection of structural templates with the same anchor points and requiring the fewest amino acid mutations. The amino acid edit distances of the test peptides ranged from 1 to 7, and their match scores with the corresponding templates ranged from -5 to 6, with most peptides having a match score of 3. The score distribution for each peptide template selection and detailed template alignment information are recorded in Figure 1b.
[0134] The initial structures of each peptide were constructed according to the previously described method, and molecular dynamics (MD) simulations were performed on them for 60 nanoseconds. The energies of hotspot amino acids in MHC molecules and peptides were then calculated using the ASGBIE method. The theoretical energies were positively correlated with the experimentally measured affinities, as shown in Table 2. Analysis of the consistency between various energy parameters, including van der Waals forces, charge interactions, and interaction entropy, and experimental values revealed that van der Waals interactions are the main mediator of MHC-peptide interactions, with correlations of 0.32 and 0.45 for MHC and peptide hotspot amino acids, respectively. The charge interactions of MHC hotspot amino acids were the main source of calculation error, with a correlation of only 0.08, while the correlation for peptide hotspot amino acids was 0.38. The method combining van der Waals forces and peptide amino acid interaction entropy (excluding the less accurate charge interaction component) showed the highest correlation with experimental results, reaching 0.48 (Figure 3a), exceeding the correlation of 0.39 for free energy change (ΔG), indicating the important role of interaction entropy.
[0135] Furthermore, the pMHC energy calculated based on peptide amino acids showed better agreement with experimental data than that calculated based on MHC hotspot amino acids. The most accurate method for calculating the pMHC binding energy (ΔG) was to average the van der Waals forces and entropy values of MHC and peptide hotspot amino acids, achieving a correlation coefficient of 0.39. The energy values calculated using peptide hotspot amino acids indicated that the total energy of the anion peptide was below 8 kcal / mol.
[0136] Example 4: Bias in experimental data and its impact on the accuracy of the computational model.
[0137] In reviewing the IC50 benchmarks of these 426 peptides and comparing them with experimental evidence from the Immunoeptope Database (IEDB), it was found that the experimental affinity data of some peptides were incorrectly labeled as IC50 instead of IC50. Furthermore, there were inconsistencies between the IC50 / EC50 affinity values of some peptides and the qualitative annotations in IEDB. In IEDB, when a peptide has multiple experimental entries, the IC50 result is prioritized, and the qualitative description with the highest frequency among the multiple entries is considered representative of the peptide's final affinity state. In this dataset, 22 peptides were classified as "negative" in IEDB but had IC50 values below 500 nM (classified as inconsistency category 1); 55 peptides were classified as "positive" but had IC50 values above 500 nM (inconsistency category 2); and another 22 peptides were labeled as EC50 (inconsistency category 3). In total, the classification of 99 peptides was ambiguous.
[0138] After removing these ambiguous entries from the dataset, the correlation between energy calculations and experimental affinity values improved. The sum of van der Waals forces and interaction entropies of peptide amino acids (excluding less accurate electrostatic interactions) showed the highest correlation with experimental values, reaching 0.55 (Figure 3b). This finding highlights the significant impact of uncertain data in benchmark datasets on evaluation accuracy, particularly for machine learning models that rely on data accuracy. Correlation analysis between theoretical calculations and experimental affinity data for these three categories of ambiguous peptides showed a correlation of 0.3 for inconsistency category 1, indicating that these peptides are in a critical state of high-affinity MHC interactions. For peptides with experimental IC50 below 500 nM, the binding free energy is mostly concentrated around 6.5 kcal / mol. This indicates challenges in the consistency of experimental results and judgment criteria, leading to data uncertainty.
[0139] In contrast, the correlations for inconsistency categories 2 and 3 were -0.07 and -0.28, respectively, indicating a significant discrepancy between experimental results and theoretical calculations for peptides with conflicting records. This discrepancy may stem from errors in the experiment itself or in the data recording process. Peptides labeled as EC50 may originate from experimental designs different from those used for IC50 experiments. The existence of these uncertainties suggests that similar labeling errors may exist for other peptides as well. Combined with potential experimental discrepancies (e.g., the differences found by Ding et al. in different experimental detections of a group of HBV peptides), this poses a substantial risk of introducing significant bias into machine learning-based methods.
[0140] Example 5: Actual screening and validation of HBV peptides based on competitive binding ELISA assay.
[0141] The peptides derived from HBV virus were calculated using the method described in the patent. Two of these peptides have publicly reported experimental validation results and will serve as the positive and negative control peptides in this experiment. No corresponding competitive binding assays or other affinity detection methods were found for the two test peptides.
[0142] The principle of this experiment is based on the instability of MHC and its subunits in the absence of peptides. The test peptide competitively binds to a reference peptide. When the test peptide has a strong affinity, more MHC-peptide complexes remain and are captured by ELISA technology after multiple rounds of UV irradiation and elution, resulting in a higher detection value. The relevant detection reagents were purchased from Legend Biotech; the product model is Flex-T. TM HLA-A*02:01 Monomer UVX and LEGEND MAX TM Flex-T TM Human Class I Peptide Exchange ELISA Kit.
[0143] The experimental results are shown in the table. Among the three different methods for calculating binding free energy, the binding free energy of the positive control peptide was higher than that of the negative control peptide. Furthermore, the improved MHC-Peptide free energy calculation method showed the largest theoretical calculation difference, demonstrating its better ability to distinguish between positive and negative control peptides. Of the two test peptides, test peptide 1 had a higher binding free energy than peptide 2, with ELISA experimental values of 1.127 and 0.118, respectively, consistent with the theoretical calculations. This indicates that the method of this patent has accurate distinguishing ability for unknown test peptides. In addition, a comparison of the ELISA competitive binding experiment results of the four peptides with the reference OD values showed that the positive control peptide had the largest experimental and theoretical calculation values, indicating high affinity for MHC. The negative peptide had the smallest theoretical and experimental values, indicating medium or low affinity. Of the two test peptides, test peptide 1 had moderately high affinity, and test peptide 2 had moderately low affinity.
[0144] Example 6: Screening and verification of HBV peptides based on BLI experiment.
[0145] The peptides derived from HBV virus were calculated using the method described in the patent. Two of these peptides have publicly reported experimental validation results and will serve as the positive and negative control peptides in this experiment. No corresponding competitive binding assays or other affinity detection methods were found for the remaining peptide.
[0146] The experimental principle utilizes biolayer interferometry (BLI) to determine the affinity of the test peptide for MHC molecules. By binding the test peptide and a reference peptide to a BLI sensor probe pre-loaded with MHC, the binding and dissociation processes are monitored in real time using BLI, recording the binding rate (ka) and dissociation rate (kd) to calculate the binding constant (KD). Higher affinity of the test peptide will result in a lower dissociation rate and a higher binding constant, reflecting a more stable MHC-peptide complex. During the experiment, gradient concentrations of test peptide solutions are used to bind to the MHC probe, further validating the quantitative characteristics of peptide affinity.
[0147] The experimental results are shown in the table. Among the various free energy calculation methods based on molecular dynamics, the binding free energy of the positive control peptide is higher than that of the negative control peptide. Furthermore, the improved MHC-peptide free energy calculation method shows the largest theoretical calculation difference, demonstrating the superior ability of the improved method to distinguish between positive and negative control peptides.
[0148] For the test peptide, its binding free energy was 38.907, the sum of van der Waals interaction and entropy was 39.578, and the binding free energy calculated by the improved free energy calculation method was 33.194. In the BLI experiment, the binding constant (Ka) of the test peptide was 8.37 × 10²¹ / Ms, and the dissociation rate (Kdis) was 1.79 × 10²¹ / Ms. -4 1 / s, equilibrium dissociation constant (KD) is 2.14 × 10 -7 M indicates that it has a moderately high binding capacity. The binding free energy of the negative control peptide was 26.527, the sum of van der Waals interaction and entropy was 23.491, and the binding free energy calculated by the improved free energy method was 16.817. Since the Ka, Kdis, and KD values were not detected in the BLI experiment, it indicates that the peptide has a low binding capacity to MHC. The binding free energy of the positive control peptide was 32.811, the sum of van der Waals interaction and entropy was 34.314, and the binding free energy calculated by the improved free energy method was 27.223. In the BLI experiment, the positive control peptide exhibited the highest Ka value (1.13 × 10⁻⁶). 3 1 / Ms), the lowest Kdis value (4.82×10 -7 1 / s), and the lowest KD value. This indicates that it has the highest affinity for MHC.
[0149] Based on the combined theoretical calculations and BLI experimental results, the binding ability of the tested peptides falls between that of the positive and negative controls. The improved free energy calculation method shows high consistency with the BLI experimental data, demonstrating its ability to accurately distinguish the binding ability of peptides to MHC.
[0150] In some embodiments, step S210 includes:
[0151] Edit distance is calculated between peptides in MHC-peptide data and peptides in target MHC-peptides.
[0152] The anchor amino acids of peptides in MHC-peptide data are compared and matched with the anchor amino acids of peptides in target MHC-peptides, and the similarity contribution score is calculated by mismatch penalty and / or matching score.
[0153] The similarity score is obtained by subtracting the similarity contribution score from the edit distance value;
[0154] The peptide with the lowest similarity is selected as the preferred peptide.
[0155] In some embodiments, the specific steps for obtaining candidate MHC-peptides in step S300 include:
[0156] Modify the non-anchor amino acids in the MHC-peptide template;
[0157] Multiple predicted structures are generated, and the predicted structure with the best score is selected as the candidate MHC-peptide.
[0158] In some specific embodiments, the step of modifying non-anchor amino acids in the MHC-peptide template includes: in the process of modifying multiple non-anchor amino acids, modifying all discontinuous non-anchor amino acids at once, and modifying continuous non-anchor amino acids one by one.
[0159] In some embodiments, step S500 includes:
[0160] S510. Select amino acids from candidate MHC-peptides to form an amino acid set;
[0161] S520. Perform a scan of alanine in the amino acid set to calculate its van der Waals interactions, electrostatic interactions, solvation free energy and nonsolvation free energy.
[0162] S530. Calculate the entropy contribution of each amino acid using the interaction entropy calculation method;
[0163] S540. Calculate the binding free energy based on the amino acid set, including: S541. Summing up the energy terms in steps S520 and S530 to obtain the original binding free energy; S542. Extracting the van der Waals interaction energy and the entropy contribution and summing them to obtain the primary improved binding free energy; S543. Calculating the primary improved binding free energy of the corresponding amino acids for MHC and peptides respectively, and taking their arithmetic mean as the final improved binding free energy.
[0164] In some specific embodiments, step S510 includes:
[0165] Select any amino acid on the MHC atom of the candidate MHC peptide that is within the cutoff distance of the peptide in the candidate MHC peptide. Collect these amino acids together with all amino acids on the peptide in the candidate MHC peptide as the amino acid set for alanine scanning. The cutoff distance ranges from 5 Å to 12 Å.
[0166] From a physics perspective, the effective range of van der Waals interactions (including Lennard-Jones potential energy) and Coulomb forces typically does not exceed [a certain value]. Therefore, distance is usually considered The atoms mentioned above are considered to have very weak or no interaction;
[0167] Based on experience, most mainstream molecular dynamics simulation software has long-range corrections, and the cutoff distance is generally set to... It does not affect the calculation results;
[0168] Here you can choose An arbitrary threshold is used as the cutoff distance. The larger this value, the more amino acids will be included in the calculation. However, the interaction between atoms that are far apart is weak and has little impact on the result. But the computational workload will increase exponentially. Therefore, 5 or 8 is taken as an acceptable threshold. A value less than 5 is generally considered to have a significant impact on the accuracy of the result.
[0169] In some embodiments, after step S100 and before step S210, the following is included:
[0170] To supplement missing heavy atoms in MHC-peptide data;
[0171] Remove irrelevant components that aid crystallization other than MHC and peptides from MHC-peptide data;
[0172] Screen MHC-peptide data with a resolution below 3 angstroms.
[0173] It is generally believed that: High-resolution structures: It can clearly resolve details such as side chain conformation, water molecule and ion coordination. Medium resolution structure: The main chain conformation is clear, but some side chains (such as long chain residues) may be blurred. Low-resolution structure: It can only identify the main chain direction; the side chain positions require modeling and optimization. Because subsequent steps involve modification and optimization simulations of amino acids, a relatively large screening threshold is acceptable. Of course, reducing the threshold will improve the quality of the template structure and lead to more accurate calculation results.
[0174] In some embodiments, step S400 includes:
[0175] Before performing molecular dynamics simulations, hydrogen atoms are added to check for atomic overlap. If atomic overlap is found, it is manually repaired and / or automatically corrected through optimization commands.
[0176] After ensuring that the addition of hydrogen atoms and structural corrections were correct, water molecules were added, and ions were added to the simulation system to make the charge of the simulation system zero.
[0177] Perform molecular dynamics simulations to obtain simulation results.
[0178] This application also discloses an MHC-peptide screening method for screening peptides obtained by the above-mentioned MHC-peptide binding free energy calculation method.
[0179] Candidate peptides are ranked and screened based on improved binding free energy: They are sorted from highest to lowest improved binding free energy, with the Top 5, 10, 15, and 30 peptides selected as the primary criteria for selection. The empirically assessed improved binding free energy should not be lower than 15 kcal / mol, and the free energy contribution of anchor amino acids should not be lower than 8-10 kcal / mol. The RMSD of the peptide should not exceed 2.5 Å for an extended period during simulation; otherwise, it is necessary to observe whether the peptide has broken free from the MHC binding pocket during the simulation. This phenomenon usually indicates low affinity of the peptide, and in this case, the calculated binding free energy will be very low or impossible to calculate. Following these screening criteria and methods will improve the accuracy of identifying strongly binding peptides.
[0180] In some embodiments, peptides are screened based on improved binding free energy and the free energy contribution of anchor amino acids.
[0181] In the description of this specification, the use of terms such as "Embodiment 1," "this embodiment," or "in one embodiment" indicates that the specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example; moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in one or more embodiments or examples.
[0182] In the description of this specification, the terms "connection," "installation," "fixing," "setting," and "having" are interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be a connection within two components. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0183] In the description of this specification, relational terms such as “first” and “second” are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0184] The above description of the embodiments is intended to enable those skilled in the art to understand and apply the technology of this invention. Those skilled in the art can easily make various modifications to these examples and apply the general principles described herein to other embodiments without creative effort. Therefore, this invention is not limited to the above embodiments. Modifications in the following situations should be within the scope of protection of this invention: ① New technical solutions implemented based on the technical solution of this invention and combined with existing common knowledge, where the technical effects of the new technical solution do not exceed the technical effects of this invention; ② Equivalent substitutions of some features of the technical solution of this invention using known technology, resulting in the same technical effects as those of this invention; ③ Extendable technical solutions based on the technical solution of this invention, where the substantive content of the extended technical solution does not exceed the technical solution of this invention; ④ Equivalent transformations made using the content of this specification and drawings, directly or indirectly applied to other related technical fields.
Claims
1. A method for calculating the MHC-peptide binding free energy, characterized in that, include: S100. Acquire data, including known MHC-peptide data and known structural information of the target MHC-peptide. S210. Determine the preferred polypeptide by performing sequence similarity calculation on the polypeptide in the MHC-peptide data and the polypeptide in the target MHC-peptide, and performing matching and scoring calculation on the anchor amino acids of the two, obtaining the similarity degree based on the calculation results of the sequence similarity calculation and the matching and scoring calculation, and selecting the best matching preferred polypeptide based on the similarity degree. S220. Determine the preferred MHC: Select the MHC that is most similar to the MHC in the target MHC-peptide from the MHC-peptide data as the preferred MHC. S230, Forming an MHC-peptide template: The MHC-peptide template is formed according to the preferred polypeptide and the preferred MHC. S300. Homology modeling is performed on the MHC-peptide template to obtain candidate MHC-peptides; S400. Perform molecular dynamics simulations on the candidate MHC-peptide and obtain the simulation results; S500. Based on the simulation results, calculate the binding free energy of MHC and polypeptide in the candidate MHC-peptide to obtain the binding free energy.
2. The method for calculating the MHC-peptide binding free energy according to claim 1, characterized in that, Step S210 includes: The edit distance value between the peptide in the MHC-peptide data and the peptide in the target MHC-peptide is calculated by editing distance; The anchor amino acids of the peptides in the MHC-peptide data and the anchor amino acids of the peptides in the target MHC-peptide are compared and matched, and the similarity contribution score is calculated by mismatch penalty and / or matching score. The similarity score is obtained by subtracting the similarity contribution score from the edit distance value; The polypeptide with the lowest degree of similarity is selected as the preferred polypeptide.
3. The method for calculating the MHC-peptide binding free energy according to claim 1, characterized in that, In step S300, the specific steps for obtaining candidate MHC-peptides include: Modify the non-anchor amino acids in the MHC-peptide template; Multiple predicted structures are generated, and the predicted structure with the best score is selected as the candidate MHC-peptide.
4. The method for calculating the MHC-peptide binding free energy according to claim 3, characterized in that, The step of modifying the non-anchor amino acids in the MHC-peptide template includes: in the process of modifying multiple non-anchor amino acids, modifying all discontinuous non-anchor amino acids at once, and modifying continuous non-anchor amino acids one by one.
5. A method for calculating the MHC-peptide binding free energy according to any one of claims 1 to 4, characterized in that, Step S500 includes: S510. Select amino acids from the candidate MHC-peptides to form an amino acid set; S520. Perform a scan of alanine in the amino acid set to calculate its van der Waals interaction, electrostatic interaction, solvation free energy and non-solvation free energy; S530. Calculate the entropy contribution of each amino acid using the interaction entropy calculation method; S540. Calculate the binding free energy based on the amino acid set, including: S541. Summing up the energy terms in steps S520 and S530 to obtain the original binding free energy; S542. Extracting the van der Waals interaction energy and the entropy contribution and summing them to obtain the primary improved binding free energy; S543. Calculating the primary improved binding free energy of the corresponding amino acids for MHC and polypeptide, and taking their arithmetic mean as the final improved binding free energy.
6. The method for calculating the MHC-peptide binding free energy according to claim 5, characterized in that, Step S510 includes: Select any amino acid within the cutoff distance of any MHC atom in the candidate MHC peptide from the polypeptide cutoff distance of the candidate MHC peptide. Collect these amino acids together with all amino acids in the polypeptide of the candidate MHC peptide as the amino acid set for alanine scanning. The cutoff distance ranges from 5 Å to 12 Å.
7. A method for calculating the MHC-peptide binding free energy according to any one of claims 1 to 4, characterized in that, After step S100 and before step S210, the following is included: The missing heavy atoms were supplemented in the MHC-peptide data; Remove irrelevant components that aid crystallization other than MHC and polypeptides from the MHC-peptide data; Screening for MHC-peptide data with a resolution below 3 angstroms.
8. A method for calculating the MHC-peptide binding free energy according to any one of claims 1 to 4, characterized in that, Step S400 includes: Before performing molecular dynamics simulations, hydrogen atoms are added to check for atomic overlap. If atomic overlap is found, it is manually repaired and / or automatically corrected through optimization commands. After ensuring that the addition of hydrogen atoms and structural corrections were correct, water molecules were added, and ions were added to the simulation system to make the charge of the simulation system zero. Perform molecular dynamics simulations to obtain simulation results.
9. A method for screening MHC-peptides, characterized in that, Used for screening peptides obtained by the MHC-peptide binding free energy calculation method according to any one of claims 1 to 8.
10. The MHC-peptide screening method according to claim 9, characterized in that, The peptides were screened based on the improved binding free energy and the free energy contribution of the anchor amino acids.