Method and device for predicting bonding of protein fragment-major histocompatibility complex
By integrating protein structure prediction, structure binding simulation, affinity analysis and machine learning models to predict peptide-MHC binding, the time-consuming and costly problems of traditional experiments are solved, and efficient and accurate treatment plans are achieved.
Patent Information
- Application Number
- CN202410313911.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-19
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional peptide-MHC binding experiments require a lot of time and money, and it is difficult to efficiently confirm whether protein fragments in cancer cells bind to MHC, which affects the efficiency and accuracy of treatment.
By combining protein structure prediction models, structure binding simulation tools, affinity analysis models and machine learning models, peptide-MHC binding is predicted, integrated scores are obtained, and experimental requirements are reduced.
Rapidly and accurately predict peptide-MHC binding without the need for experiments, saving time and costs, and improving treatment efficiency and accuracy.
Smart Images

Figure CN120673833A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a prediction method and apparatus, and more particularly to a protein fragment-major histocompatibility complex bonding prediction method and apparatus. Background Art
[0002] Cancer treatment has been a widely explored research topic, with a variety of innovative treatments emerging in recent years. Among these, immunotherapy, which uses specially formulated antigen vaccines to stimulate the body's immune system to actively mark and eliminate cancer cells, has garnered significant attention. The core challenge of this treatment approach lies in confirming whether protein fragments (peptides) in cancer cells can effectively bind to the major histocompatibility complex (MHC). Each cancer patient possesses a different type of MHC, which in turn leads to a great diversity in different protein fragments. Once peptide-MHC successfully binds and is recognized by T cells, it can trigger an immune response, prompting the immune system to recognize and kill cancer cells.
[0003] Traditional peptide-MHC binding experiments often require significant time and money, as cell culture is required to confirm whether the mutated protein fragment binds to MHC. Faced with this challenge, developing methods that can predict peptide-MHC binding is crucial, saving time and costs while also improving the efficiency and accuracy of treatment. Summary of the Invention
[0004] The present invention relates to a protein fragment-major histocompatibility complex (MHC) binding prediction method and device. Multiple sets of structural binding indicators obtained from different protein structure prediction models and affinity indices obtained from different affinity analysis models are simultaneously input into a machine learning model. After inference by the machine learning model, an integrated score is generated to successfully predict whether the peptide sequence and MHC sequence can stably bind. This not only saves time and costs, but also improves the efficiency and accuracy of treatment.
[0005] According to one aspect of the present invention, a method for predicting protein fragment (peptide)-major histocompatibility complex (MHC) binding is proposed. The peptide-MHC binding prediction method includes the following steps: Obtain a peptide sequence and an MHC sequence. Based on the peptide sequence and the MHC sequence, obtain at least one protein structure file (Protein Data Bank, PDB) using at least one protein structure prediction model. Based on the protein structure file, obtain at least one structure binding pointer using at least one structure binding simulation tool. Obtain a peptide sequence and a human leukocyte antigen (HLA) type corresponding to the MHC sequence. Based on the peptide sequence and the HLA type, obtain at least one affinity index using at least one affinity analysis model. Input the structure binding pointer and the affinity pointer into a machine learning model to obtain an integration score.
[0006] According to another aspect of the present invention, a device for predicting protein fragment (peptide)-major histocompatibility complex (MHC) binding is provided. The peptide-MHC binding prediction device includes an input unit, at least one protein structure prediction model, at least one structure binding simulation tool, a classification unit, at least one affinity analysis model, and a machine learning model. The input unit is configured to obtain a peptide sequence and an MHC sequence. The protein structure prediction model is configured to obtain at least one protein structure file (Protein Data Bank, PDB) based on the peptide sequence and the MHC sequence. The structure binding simulation tool is configured to obtain at least one structure binding index based on the protein structure file. The classification unit is configured to obtain a human leukocyte antigen (HLA) type corresponding to the MHC sequence. The affinity analysis model is configured to obtain at least one affinity index based on the peptide sequence and the HLA type. The machine learning model is configured to obtain an integration score based on the structure binding index and the affinity index.
[0007] In order to better understand the above and other aspects of the present invention, embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 FIG. 1 is a block diagram of a protein fragment (peptide)-major histocompatibility complex (MHC) binding prediction device according to an embodiment.
[0009] Figure 2 FIG1 is a flowchart of a protein structure analysis method for predicting peptide-MHC bonds according to an embodiment.
[0010] Figure 3 FIG1 is a flowchart of a protein sequence analysis method for predicting peptide-MHC binding according to an embodiment.
[0011] Figure 4 This is a machine learning flowchart of a peptide-MHC binding prediction method according to an embodiment.
[0012] Reference numerals:
[0013] 110: Input unit
[0014] 120m: Protein structure prediction model
[0015] 130n:Structural bonding simulation tool
[0016] 210: Classification unit
[0017] 230k: Affinity analysis model
[0018] 300: Machine Learning Model
[0019] 1000: Bond prediction device
[0020] AXik: Affinity Index
[0021] BXijmnt: structure combined with pointer
[0022] HLAp: Human Leukocyte Antigen Type
[0023] MHCj:MHC sequence
[0024] PDBijm: protein structure file
[0025] PPTi:peptide sequence
[0026] S110, S120, S130, S210, S220, S230, S300: Steps
[0027] SCij: integrated score
[0028] RK:Indicator Importance Ranking DETAILED DESCRIPTION
[0029] The technical terms used in this specification are based on customary terms in the technical field. If some terms are explained or defined in this specification, the interpretation of these terms shall be based on the explanations or definitions in this specification. Each embodiment of the present invention has one or more technical features. Under the premise of possible implementation, those with ordinary knowledge in this technical field may selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.
[0030] Please refer to Figure 1 , which is a block diagram of a protein fragment (peptide)-major histocompatibility complex (MHC) binding prediction device 1000 according to one embodiment. The peptide-MHC binding prediction device 1000 includes an input unit 110, at least one protein structure prediction model 120m, at least one structure binding simulation tool 130n, a classification unit 210, at least one affinity analysis model 230k, and a machine learning model 300. The input unit 110 is used to input various data, such as a wireless communication module, a wired communication module, a transmission line, or a storage device.
[0031] The protein structure prediction model 120m is used to perform protein structure prediction, such as "AlphaFold2", "Colabfold", "Pandora" or other models.
[0032] AlphaFold2's protein structure prediction model 120m is a deep learning model developed by DeepMind for protein structure prediction. It achieved outstanding results in the competition, providing an advanced solution for protein structure prediction.
[0033] Pandora's protein structure prediction model 120m is a fast, modular and highly flexible prediction model that uses MODELLER's modeling script and is specifically designed to generate peptide-MHC structures.
[0034] The structural binding simulation tool 130n is used to perform protein binding simulation and analysis, such as "AutoDockVina", "PISA (Proteins, Interfaces, Structures and Assemblies)", "Local Distance Difference Test (LDDT)" or other tools.
[0035] The PISA structural binding simulation tool 130n is a tool for complex and interface analysis in the field of protein structure, promoted by the Protein Data Bank in Europe (PDBe). Its primary goal is to provide researchers with a deeper understanding of protein structure, interfaces, and assembly, thereby promoting the understanding of biological and biochemical processes.
[0036] AutoDock Vina's structural binding simulation tool 130n is a computational molecular docking software. It's an improved version of AutoDock and was developed by Scripps Research Institute. It's widely used in drug design and biomolecular interaction research to predict how small molecules bind to proteins.
[0037] Other tools include "CONTACT", which calculates the number of atomic bonds based on van der Waals forces. By using experimental data, the radius of each interatomic bond is obtained, and the number of bonds generated is then calculated.
[0038] The affinity analysis model 230k is used for affinity prediction, such as “HLAthena”, “MHCflurry”, “NetMHCpan”, “NetMHCIIpan” or other models.
[0039] The affinity analysis model 230k is used to predict the binding between MHC molecules and antigenic peptides. These tools combine extensive experimental data with deep learning models and features derived from diverse biological perspectives to provide predictions for different MHC species and peptides. These tools, based on diverse biological hypotheses and data techniques, are constantly updated with technological advancements, contributing to a deeper understanding of peptide-MHC interactions.
[0040] For example, the input unit 110 obtains a peptide sequence PPTi and an MHC sequence MHCj. The peptide sequence PPTi and the MHC sequence MHCj are input into the protein structure prediction model 120m.
[0041] The protein structure prediction model 120m obtains at least one protein structure file (Protein Data Bank, PDBijm) based on the peptide sequence PPTi and the MHC sequence MHCj. The protein structure file PDBijm is input into the structure binding simulation tool 130n.
[0042] The structure binding simulation tool 130n obtains at least one structure binding pointer BXijmnt according to the protein structure file PDBijm.
[0043] On the other hand, after the input unit 110 inputs the MHC sequence MHCj into the classification unit 210, the classification unit 210 classifies the MHC sequence MHCj to obtain the human leukocyte antigen (HLA) type HLAp. The human leukocyte antigen type HLAp is divided into "type 1" and "type 2".
[0044] The peptide sequence PPTi and the human leukocyte antigen type HLAp are input into the affinity analysis model 230k. The affinity analysis model 230k obtains at least one affinity index AXik based on the peptide sequence PPTi and the human leukocyte antigen type HLAp.
[0045] After obtaining the above-mentioned structure binding pointer BXijmnt and affinity pointer AXik, the structure binding pointer BXijmnt and affinity index AXik are input into the machine learning model 300. The machine learning model 300 obtains an integration score SCij according to the structure binding pointer BXijmnt and affinity index AXik.
[0046] In other words, the peptide-MHC binding prediction device 1000 of the present invention can rapidly obtain an integrated score for peptide-MHC binding prediction without the need for experimentation, taking into account both protein structure prediction and affinity prediction, for researchers' reference. This not only saves time and costs, but also improves the efficiency and accuracy of treatment.
[0047] Three flow charts are provided below to explain in detail the operation of the peptide-MHC bonding prediction device 1000 .
[0048] Please refer to Figure 2 , which is a protein structure analysis flow chart of a peptide-MHC bonding prediction method according to an embodiment of the present invention. The protein structure analysis flow chart includes steps S110, S120, and S130.
[0049] In step S110 , the peptide sequence PPTi and the MHC sequence MHCj to be analyzed are obtained.
[0050] Next, in step S120, at least one protein structure file (Protein Data Bank, PDB) PDBijm is obtained using at least one protein structure prediction model 120m according to the peptide sequence PPTi and the MHC sequence MHCj. Figure 2As shown, the protein structure prediction model 120m includes, for example, "AlphaFold2," "Colabfold," "Pandora," or other models. Each protein structure prediction model 120m obtains a protein structure file PDBijm based on the peptide sequence PPTi and the MHC sequence MHCj to be predicted. The "AlphaFold2" protein structure prediction model 120m obtains the protein structure file PDBijm of "PDB1." The "Colabfold" protein structure prediction model 120m obtains the protein structure file PDBijm of "PDB2." The "Pandora" protein structure prediction model 120m obtains the protein structure file PDBijm of "PDB3."
[0051] Then, in step S130, the structure binding pointer BXijmnt is obtained by using the structure binding simulation tool 130n according to the protein structure file PDBijm. Figure 2 As shown, the structural integration simulation tool 130n includes, for example, “AutoDock Vina”, “PISA (Proteins, Interfaces, Structures and Assemblies)” or “Local Distance Difference Test (LDDT)”.
[0052] If the structure binding simulation tool 130n is "AutoDock Vina", the structure binding pointer BXijmnt of "AutoDock Vina" includes "binding score" and / or "number of hydrogen bonds (H-Bond)". The structure binding simulation tool 130n of "AutoDock Vina" will obtain a set of structure binding pointers BXijmnt of "binding score" and "number of hydrogen bonds" based on the protein structure file PDBijm of "PDB1". The structure binding simulation tool 130n of "AutoDock Vina" will also obtain another set of structure binding pointers BXijmnt of "binding score" and "number of hydrogen bonds" based on the protein structure file PDBijm of "PDB2". The structure binding simulation tool 130n of "AutoDock Vina" will also obtain another set of structure binding pointers BXijmnt of "binding score" and "number of hydrogen bonds" based on the protein structure file PDBijm of "PDB3".
[0053] If the structure binding simulation tool 130n is "PISA", the structure binding pointer BXijmnt of "PISA" includes "p-value", "number of hydrogen bonds (H-Bond)" and / or "folding energy (Energy for folding)". The structure binding simulation tool 130n of "PISA" will obtain a set of structure binding pointers BXijmnt of "p-value", "number of hydrogen bonds" and "folding energy" based on the protein structure file PDBijm of "PDB1". The structure binding simulation tool 130n of "PISA" will also obtain another set of structure binding pointers BXijmnt of "p-value", "number of hydrogen bonds" and "folding energy" based on the protein structure file PDBijm of "PDB2". The structure binding simulation tool 130n of "PISA" will also obtain another set of structure binding pointers BXijmnt of "p-value", "number of hydrogen bonds" and "folding energy" based on the protein structure file PDBijm of "PDB3".
[0054] If the structure binding simulation tool 130n is a "local distance difference test (LDDT)", the structure binding pointer BXijmnt of the "local distance difference test (LDDT)" includes the "protein structure file distance difference (Iddt with other PDB)". The structure binding simulation tool 130n of the "local distance difference test (LDDT)" will obtain a set of structure binding pointers BXijmnt of "protein structure file distance difference" based on the protein structure file PDBijm of "PDB1". The structure binding simulation tool 130n of the "local distance difference test (LDDT)" will also obtain another set of structure binding pointers BXijmnt of "protein structure file distance difference" based on the protein structure file PDBijm of "PDB2". The structure binding simulation tool 130n of the "local distance difference test (LDDT)" will also obtain another set of structure binding pointers BXijmnt of "protein structure file distance difference" based on the protein structure file PDBijm of "PDB3".
[0055] According to the above protein structure analysis process, these different protein structure prediction models 120m and different structure binding simulation tools 130n can be interleaved and combined to obtain multiple sets of structure binding indicators BXijmnt. These structure binding indicators BXijmnt take into account the different protein structure prediction models 120m and different structure binding simulation tools 130n, making the protein structure analysis more complete.
[0056] Please refer to Figure 3This is a protein sequence analysis process for a peptide-MHC binding prediction method according to one embodiment. The protein sequence analysis process includes steps S210, S220, and S230. In step S210, the peptide sequence PPTi to be analyzed and the human leukocyte antigen (HLA) type HLAp corresponding to the MHC sequence MHCj are obtained. In this step, the classification unit 210 classifies the peptide based on the MHC sequence MHCj to identify a "type 1" HLAp and a "type 2" HLAp.
[0057] Next, in step S220 , it is determined whether the human leukocyte antigen type HLAp belongs to “type 1” or “type 2”.
[0058] Then, in step S230, at least one affinity index AXik is obtained based on the peptide sequence PPTi and the human leukocyte antigen type HLAp using the affinity analysis model 230k. Figure 3 As shown, if the human leukocyte antigen type HLAp is "type 1", the affinity analysis model 230k includes "HLAthena", "MHCflurry" or "NetMHCpan".
[0059] If the affinity analysis model 230k is "HLAthena," the affinity index AXik includes "Probability." The affinity analysis model 230k of "HLAthena" obtains a set of "Probability" affinity indexes AXik based on the peptide sequence PPTi and the "type 1" human leukocyte antigen type HLAp.
[0060] If affinity analysis model 230k is "MHCflurry," affinity index AXik includes "Probability" and / or "Bonding Affinity." The affinity analysis model 230k for "MHCflurry" generates a set of affinity indexes AXik for "Probability" and "Bonding Affinity" based on the peptide sequence PPTi and the "type 1" human leukocyte antigen (HLAp).
[0061] If affinity analysis model 230k is "NetMHCpan," affinity index AXik includes "Probability" and / or "Bonding Affinity." The "NetMHCpan" affinity analysis model 230k generates a set of affinity indexes AXik for "Probability" and "Bonding Affinity" based on the peptide sequence PPTi and the "type 1" human leukocyte antigen (HLAp).
[0062] like Figure 3 As shown, if the human leukocyte antigen type HLAp is "type II", the affinity analysis model 230k includes "NetMHCIIpan".
[0063] If affinity analysis model 230k is "NetMHCIIpan," affinity index AXik includes "Probability" and / or "Bonding Affinity." The "NetMHCIIpan" affinity analysis model 230k generates a set of affinity indexes AXik for "Probability" and "Bonding Affinity" based on the peptide sequence PPTi and the "Type II" human leukocyte antigen (HLAp).
[0064] According to the above protein sequence analysis process, these different affinity analysis models 230k can correspondingly obtain multiple sets of affinity indices AXik. These affinity indices AXik take into account different affinity analysis models 230k, making protein sequence analysis more complete.
[0065] Please refer to Figure 4 , which is a machine learning process of a peptide-MHC binding prediction method according to an embodiment. The machine learning process of the peptide-MHC binding prediction method includes step S300.
[0066] In step S300, the structural binding index BXijmnt and affinity index AXik are input to the machine learning model 300 to obtain an integration score SCij and a pointer importance ranking RK. The aforementioned multiple sets of structural binding indexes BXijmnt and affinity indexes AXik are simultaneously input to the machine learning model 300. After inference by the machine learning model 300, an integration score SCij and a pointer importance ranking RK are obtained. The integration score SCij represents whether the peptide sequence PPTi and the MHC sequence MHCj can stably bind. The index importance ranking RK indicates the importance ranking of the aforementioned multiple sets of structural binding indexes BXijmnt and affinity indexes AXik for researchers' reference.
[0067] In order to more clearly illustrate the operation of the peptide-MHC bonding prediction method of the present invention, two embodiments are described in detail below.
[0068] Implementation Example 1:
[0069] The human leukocyte antigen type HLAp of "type 1" is HLA-A*02:01, and the MHC sequence MHCj and peptide sequence PPTi are as follows:
[0070] Sequence A:
[0071] GSHSMRYFFTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRMEPRAPWIEQEGPEYWDGETRKVKAHSQTHRVDLGTLRGYYNQSEAGSHTVQRMYGCDVGSDWRFLRGYHQYAYDGKDYIALKEDLRSWTAAD MAAQTTKHKWEAAHVAEQLRAYLEGTCVEWLRRYLENGKETLQRTDAPKTHMTHHAVSDHEATLRCWALSFYPAEITLTWQRDGEDQTQDTELVETRPAGDGTFQKWAAVVVPSGQEQRYTCHVQHEGLPKPLTLRWE
[0072] Sequence B:
[0073] MIQRTPKIQVYSRHPAENGKSNFLNCYVSGFHPSDIEVDLLKNGERIEKVEH SDLFSKDWSFYLLYYTEFTPTEKDEYACRVNHVTLSQPKIVKWDRDM
[0074] Peptide:
[0075] ALGIGILTV
[0076] The FASTA format is a text file with the extension .fasta. Each sequence will have a line of description starting with >, as follows:
[0077] >4EUP_1|Chains A,D|HLAclass I histocompatibility antigen,A-2alphachain|Homo sapiens(9606)
[0078] GSHSMRYFFTSVSRPGRGEPRFIAVGYVDDTQFVRFDSDAASQRMEPRAPWIEQEGPEYWDGETRKVKAHSQTHRVDLGTLRGYYNQSEAGSHTVQRMYGCDVGSDWRFLRGYHQYAYDGKDYIALKEDLRSWTAAD MAAQTTKHKWEAAHVAEQLRAYLEGTCVEWLRRYLENGKETLQRTDAPKTHMTHHAVSDHEATLRCWALSFYPAEITLTWQRDGEDQTQDTELVETRPAGDGTFQKWAAVVVPSGQEQRYTCHVQHEGLPKPLTLRWE
[0079] >4EUP_2|Chains B,E|Beta-2-microglobulin|Homo sapiens(9606)
[0080] MIQRTPKIQVYSRHPAENGKSNFLNCYVSGFHPSDIEVDLLKNGERIEKVEHSDLSFSKDWSFYLLYYTEFTPTEKDEYACRVNHVTLSQPKIVKWDRDM
[0081] >4EUP_3|Chains C,F|Melanoma antigen recognized by T-cells 1|Homosapiens(9606)
[0082] ALGIGILTV
[0083] Implementation Example 2:
[0084] The human leukocyte antigen type HLAp of "type 2" is HLA-DQ0602, and the MHC sequence MHCj and peptide sequence PPTi are as follows:
[0085] Sequence A:
[0086] HVASCGVNLYQFYGPSGQYTHEFDGDEQFYVDLERKETAWRWPEFSKFGGFDPQGALRNMAVAKHNLNIMIKRYN
[0087] Sequence B:
[0088] FVFQFKGMCYFTNGTERVRLVTRYIYNREEYARFDSDVGVYRAVTPQGRPDAEYWNSQKEVLEGTRAELDTVCRHNYEVAFRGIL
[0089] Peptide:
[0090] NLPSTKVSWAA
[0091] The FASTA format is a text file with the extension .fasta, as follows:
[0092] >1UVQ_1|Chain A|HLA CLASS II HISTOCOMPATIBILITY ANTIGEN|HOMO SAPIENS(9606)
[0093] HVASCGVNLYQFYGPSGQYTHEFDGDEQFYVDLERKETAWRWPEFSKFGGFDPQGALRNMAVAKHNLNIMIKRYN
[0094] >1UVQ_2|Chain B|HLA CLASS II HISTOCOMPATIBILITY ANTIGEN|HOMO SAPIENS(9606)
[0095] FVFQFKGMCYFTNGTERVRLVTRYIYNREEYARFDSDVGVYRAVTPQGRPD AEYWNSQKEVLEGTRAELDTVCRHNYEVAFRGIL
[0096] >1UVQ_3|Chain C|OREXIN|HOMO SAPIENS (9606)
[0097] NLPSTKVSWAA
[0098] The protein structure prediction model 120m is as follows:
[0099] The protein structure prediction model 120m of "AlphaFold2" is a deep learning model developed by DeepMind. The input is a FASTA file. The protein structure prediction model 120m of "AlphaFold2" will refer to multiple different protein structure databases, generate features, and make predictions with 5 different models. The protein structure file PDBijm with the best effect is taken as the output.
[0100] Taking the first embodiment as an example, the input is in FASTA file format, and the protein structure file PDBijm is obtained as follows (partial):
[0101] ATOM 1N GLY A 1 -15.669 -8.879 9.204 1.00 89.38N
[0102] ATOM 2H GLY A 1 -15.484 -8.182 8.497 1.00 89.38H
[0103] ATOM 3H2 GLY A 1 -16.623 -8.790 9.522 1.00 89.38H
[0104] ATOM 4H3 GLY A 1 -15.508 -9.777 8.769 1.00 89.38H
[0105] ATOM 5CA GLY A 1 -14.718 -8.710 10.313 1.00 89.38C
[0106] ATOM 6HA2 GLY A 1 -13.714 -8.885 9.927 1.00 89.38H
[0107] ATOM 7HA3 GLY A 1 -14.934 -9.428 11.105 1.00 89.38H
[0108] ATOM 8C GLY A1 -14.799 -7.301 10.862 1.00 89.38C
[0109] ATOM 9O GLY A 1 -15.891 -6.745 10.932 1.00 89.38O
[0110] The protein structure prediction model "ColabFold" is based on the AlphaFold model. It simplifies the steps and resources required to search protein databases. It can be executed on Google Colab. The input method is to connect the sequences with colons. Taking the implementation example 1 as an example, the input format is as follows:
[0111] GSH...(omitted)...WE:MIQ...(omitted)...RDM:ALGIGILTV
[0112] The predicted protein structure file PDBijm is as follows (partial):
[0113] ATOM 1N GLY A1 6.787 8.354-17.701 1.00 89.06N
[0114] ATOM 2H GLY A1 6.703 8.154-16.713 1.00 89.06H
[0115] ATOM 3H2 GLY A1 7.566 8.980-17.848 1.00 89.06H
[0116] ATOM 4H3 GLY A1 6.929 7.477-18.182 1.00 89.06H
[0117] ATOM 5CAGLY A1 5.528 8.974-18.151 1.00 89.06C
[0118] ATOM 6HA2 GLY A1 4.739 8.226-18.090 1.00 89.06H
[0119] ATOM 7HA3 GLY A1 5.626 9.321-19.180 1.00 89.06H
[0120] ATOM 8C GLY A1 5.173 10.143 -17.255 1.00 89.06C
[0121] ATOM 9O GLY A 1 6.031 10.969 -16.968 1.00 89.06O
[0122] The protein structure prediction model 120m of "Pandora" is based on the Modeller model. Taking Implementation Example 2 as an example, Sequence A, Sequence B, and the peptide sequence PPTi are input, and the predicted protein structure file PDBijm is obtained as follows (partial):
[0123] ATOM 1N HIS M 1 -7.980 23.611 17.433 1.00 31.32N
[0124] ATOM 2CA HIS M 1 -8.706 22.443 17.969 1.00 31.32C
[0125] ATOM 3ND1 HIS M 1 -9.861 23.806 20.764 1.00 31.32N
[0126] ATOM 4CG HIS M 1 -10.143 23.959 19.424 1.00 31.32C
[0127] ATOM 5CB HIS M 1 -10.121 22.847 18.417 1.00 31.32C
[0128] ATOM 6NE2 HIS M 1 -10.346 25.955 20.458 1.00 31.32N
[0129] ATOM 7CD2 HIS M 1 -10.437 25.278 19.256 1.00 31.32C
[0130] ATOM 8CE1 HIS M 1-9.997 25.030 21.334 1.00 31.32C
[0131] ATOM 9C HIS M 1-8.846 21.405 16.906 1.00 31.32O
[0132] Other models
[0133] The present invention is not limited to the use of any model. As long as it can input peptide-MHC and output the predicted protein structure file PDBijm, it can be included in the process for use, such as RosettaFold.
[0134] Continuing with the above implementation example 1, we can obtain three protein structure files PDBijm: "Structure A", "Structure B", and "Structure C" predicted by "AlphaFold2", "ColabFold", and "Pandora" respectively. In implementation example 2, we can also predict three protein structure files PDBijm: "Structure A", "Structure B", and "Structure C" by "AlphaFold2", "ColabFold", and "Pandora" respectively.
[0135] The affinity analysis model 230k is as follows:
[0136] Affinity Analysis Model 230k is an algorithm that takes the peptide sequence PPTi and the human leukocyte antigen type HLAp as input. The input data is the peptide sequence PPTi. It uses the following tools to make predictions based on the input human leukocyte antigen type HLAp. Each prediction tool generates a probability of binding between the peptide sequence PPTi and the MHC sequence MHCj, with higher probability values being preferred. All tools except "HLAthena" predict a continuous value (BondingAffinity) that represents the peptide-MHC binding strength, with lower values being preferred.
[0137] Affinity analysis model 230k for "HLAthena":
[0138] Continuing with the above embodiment 1, the peptide sequence PPTi is placed into the affinity analysis model 230k of "HLAthena", the human leukocyte antigen type HLAp is HLA-A*02:01, and the "probability" value is 0.98.
[0139] Affinity analysis model 230k for "MHCFlurry":
[0140] Continuing with the above Example 1, the peptide sequence PPTi was placed into the affinity analysis model 230k of "MHCFlurry", and the human leukocyte antigen type HLAp was HLA-A*02:01, resulting in a "probability" value of 0.52 and a "binding affinity" value of 25.98.
[0141] Affinity analysis model 230k of "NetMHCpan":
[0142] Continuing with the above Example 1, the peptide sequence PPTi was placed into the affinity analysis model 230k of NetMHCpan, and the human leukocyte antigen HLAp type was HLA-A*02:01, resulting in a "probability" value of 0.78 and a "binding affinity" value of 44.70.
[0143] Affinity analysis model 230k of "NetMHCIIpan":
[0144] Continuing with Example 2 above, the peptide sequence PPTi was placed into the affinity analysis model 230k of "NetMHCIIpan" and the human leukocyte antigen type HLAp was HLA-DQA1-0102-HLA-DQB1-0602. The resulting "Probability" value was 0.0023 and the "Adhesion Affinity" value was 3246.80.
[0145] Other models:
[0146] The present invention is not limited to the use of any tool, such as a user's own prediction model, which takes the peptide sequence PPTi as input and knows which human leukocyte antigen type HLAp it will bind to, and obtains the "adhesion affinity" and "probability" of binding corresponding to the peptide sequence PPTi.
[0147] As shown in Table 1 below, based on the above embodiment 1, the "adhesion affinity" and "probability" of the affinity analysis model 230k of "NetMHCpan" and "MHCflurry" as well as the "probability" of the affinity analysis model 230k of "HLAthena" are obtained.
[0148]
[0149]
[0150] Table 1
[0151] As shown in Table 2 below, according to the above Example 2, the "adhesion affinity" and "probability" are obtained from the affinity analysis model 230k of "NetMHCIIpan".
[0152] NetMHCIIpan Binding Affinity 3246.80 NetMHCIIpan Probability 0.0023
[0153] Table 2
[0154] The structural combination simulation tool 130n is as follows:
[0155] "PISA (Proteins, Interfaces, Structures and Assemblies)" structural integration simulation tool 130n:
[0156] The data input into the structure binding simulation tool 130n of "PISA" is the protein structure file PDBijm, which contains chain A, chain B and peptide sequence PPTi.
[0157] According to Example 1, the "PISA" structural binding simulation tool 130n analyzes the different results obtained by different protein structure prediction models 120m. For the binding analysis of the "type 1" human leukocyte antigen type HLAp and the peptide sequence PPTi, only the relationship between Chain A and the peptide sequence PPTi needs to be analyzed, as only Chain A and the peptide sequence PPTi are actually relevant.
[0158] After the protein structure file PDBijm of "Structure A" (obtained by the "AlphaFold2" protein structure prediction model 120m) is analyzed by the "PISA" structure binding simulation tool 130n, the structure binding pointer BXijmnt of "number of hydrogen bonds" is obtained to be 9, the structure binding pointer BXijmnt of "p-value" is obtained to be 0.39, and the structure binding pointer BXijmnt of "folding energy" is obtained to be -11.8.
[0159] After the protein structure file PDBijm of "Structure B" (obtained by the "ColabFold" protein structure prediction model 120m) was analyzed by the "PISA" structure binding simulation tool 130n, the structure binding pointer BXijmnt of "number of hydrogen bonds" was obtained to be 10, the structure binding pointer BXijmnt of "p-value" was 0.527, and the structure binding pointer BXijmnt of "folding energy" was -11.1.
[0160] After the protein structure file PDBijm of "Structure C" (obtained by the "Pandora" protein structure prediction model 120m) was analyzed by the "PISA" structure binding simulation tool 130n, the structure binding pointer BXijmnt of "number of hydrogen bonds" was obtained to be 16, the structure binding pointer BXijmnt of "p-value" was obtained to be 0.445, and the structure binding pointer BXijmnt of "folding energy" was obtained to be -12.2.
[0161] According to the second embodiment, different protein structure files (PDBijm) generated by different protein structure prediction models 120m are analyzed using the "PISA" structure binding simulation tool 130n. Because the binding site of the "type II" human leukocyte antigen type HLAp and the peptide sequence PPTi consists of two distinct structures, chain A and chain B, separate calculations are required, followed by integration of the two results.
[0162] After the protein structure file PDBijm of "Structure D" (obtained by the "AlphaFold2" protein structure prediction model 120m) was analyzed by the "PISA" structure binding simulation tool 130n, the structure binding pointer BXijmnt of "folding energy" was obtained as -6.25, the structure binding pointer BXijmnt of "number of hydrogen bonds" was obtained as 17, and the structure binding pointer BXijmnt of "p-value" was obtained as 0.440.
[0163] The calculation of "folding energy" is: (-5.3 + -7.2) / 2 = -6.25
[0164] The calculation of “p-value” is: (min of AC, BC) = 0.440
[0165] The calculation of “number of hydrogen bonds” is: 13+4=17
[0166] After the protein structure file PDBijm of "Structure E" (obtained by the "Colabfold" protein structure prediction model 120m) was analyzed by the "PISA" structure binding simulation tool 130n, the structure binding index BXijmnt of "folding energy" was obtained as -6.4, the structure binding index BXijmnt of "number of hydrogen bonds" was 13, and the structure binding index BXijmnt of "p-value" was 0.444.
[0167] After the protein structure file PDBijm of "Structure F" (obtained by the "Pandora" protein structure prediction model 120m) was analyzed by the "PISA" structure binding simulation tool 130n, the structure binding pointer BXijmnt of "folding energy" was obtained as -4.3, the structure binding pointer BXijmnt of "number of hydrogen bonds" was obtained as 21, and the structure binding pointer BXijmnt of "p-value" was obtained as 0.530.
[0168] "AutoDock Vina" structural integration simulation tool 130n:
[0169] According to the first embodiment, the different results obtained by different protein structure prediction models 120m are analyzed by the structure binding simulation tool 130n of "AutoDock Vina".
[0170] When the protein structure file "Structure A" (PDBijm) (obtained from the "AlphaFold2" protein structure prediction model 120m) was analyzed using the "AutoDock Vina" structure binding simulation tool 130n, the molecular simulation tables were integrated and the scores were averaged, resulting in a value of -5.73. Furthermore, the structure binding indicator "number of hydrogen bonds," BXijmnt, was calculated and summed to 23.
[0171] When the protein structure file PDBijm of "Structure B" (obtained by the "Colabfold" protein structure prediction model 120m) is analyzed by the structure binding simulation tool 130n of "AutoDock Vina", the structure binding pointer BXijmnt of "binding score" is -6.256 and the structure binding pointer BXijmnt of "hydrogen bond number" is 40.
[0172] When the protein structure file PDBijm of "Structure C" (obtained by the "Pandora" protein structure prediction model 120m) is analyzed by the structure binding simulation tool 130n of "AutoDock Vina", the structure binding pointer BXijmnt of "binding score" is -5.367 and the structure binding pointer BXijmnt of "hydrogen bond number" is 32.
[0173] According to the second embodiment, the different results obtained by different protein structure prediction models 120m are analyzed by the structure binding simulation tool 130n of "AutoDock Vina".
[0174] Analyzing the protein structure file PDBijm of "Structure D" (obtained by the "AlphaFold2" protein structure prediction model 120m), the structure binding pointer BXijmnt of "binding score" is -3.067 and the structure binding pointer BXijmnt of "hydrogen bond number" is 25.
[0175] When the protein structure file "Structure E" is analyzed using PDBijm, the structure binding index BXijmnt for "binding score" is -4.389 and the structure binding index BXijmnt for "number of hydrogen bonds" is 38.
[0176] When the protein structure file PDBijm of "Structure F" is analyzed by the "pandora" structure binding simulation tool 130n, the structure binding index BXijmnt of "binding score" is -7.067 and the structure binding index BXijmnt of "hydrogen bond number" is 43.
[0177] "LDDT" structural combination simulation tool 130n:
[0178] The "LDDT" structure-binding simulation tool 130n can be used to compare the similarity between two protein structures predicted from an input sequence. Continuing with Example 1 and Example 2, we calculated the distance differences (Iddt with other PDB files) between each pair of protein structure files, as shown in Tables 3 and 4 below:
[0179] Implementation Example 1:
[0180] Structure A,Structure B 96.85% Structure A,Structure C 91.20% Structure B,Structure C 90.33%
[0181] Table 3
[0182] Implementation Example 2:
[0183] Structure D,Structure E 97.06% Structure E,Structure F 77.02% Structure D,Structure F 79.36%
[0184] Table 4
[0185] From the above process, it can be seen that if there are m protein structure prediction models 120m, m x 5 + C (m, 2) structure binding indexes BXijmnt can be generated. After analysis by the structure binding simulation tools "PISA", "AutoDock Vina" and "LDDT" 130n, the relevant structure binding indexes BXijmnt are obtained as shown in Tables 5 to 8 below:
[0186] Implementation Example 1:
[0187]
[0188] Table 5
[0189]
[0190] Table 6
[0191] Implementation Example 2:
[0192]
[0193]
[0194] Table 7
[0195]
[0196] Table 8
[0197] Operation of Machine Learning Model 300:
[0198] Continuing from the above implementation example 1, a total of 18 protein structure-related features of the structure-binding index BXijmnt and 5 affinity indices AXik corresponding to the "first type" human leukocyte antigen type HLAp are integrated; continuing from the above implementation example 2, a total of 18 protein structure-related features of the structure-binding index BXijmnt and 2 affinity indices AXik corresponding to the "first type" human leukocyte antigen type HLAp are integrated. Figure 4After collecting the above-mentioned relevant information, the prediction of the machine learning model 300 can be started. This machine learning model 300 needs to be trained with relevant data in advance. The peptide-MHC binding data set can be collected from the Protein Data Bank database, including the peptide sequence PPTi, the human leukocyte antigen type HLAp as input, and whether it will bind as a training mark (from the laboratory results). The machine learning model 300 is, for example, Random Forest, Logistic Regression, XGBoost, etc. Due to the different protein-related characteristics, the data of the "first type" human leukocyte antigen type HLAp and the "second type" human leukocyte antigen type HLAp need to be trained separately. The trained model can obtain the pointer importance ranking RK, which can be used as a basis for judging which models / tools are more important.
[0199] The above content provides different features for implementing some embodiments or examples of the present invention. The specific examples of the above description components and configurations (e.g., the numerical values or names mentioned) are to simplify / illustrate some embodiments of the present invention. Of course, these components and configurations are only examples and do not constitute limitations. In addition, some embodiments of the present invention may repeat reference symbols and / or letters in various examples. This repetition is for the purpose of simplicity and clarity and does not in itself indicate the relationship between the various embodiments and / or configurations discussed.
[0200] In summary, although the present invention has been disclosed above with reference to the embodiments, these are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations can be made without departing from the spirit and scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. A method for predicting protein fragment (peptide)-major histocompatibility complex (MHC) bonding, comprising: obtaining a peptide sequence and an MHC sequence; Obtaining at least one protein structure file (Protein Data Bank, PDB) based on the peptide sequence and the MHC sequence using at least one protein structure prediction model; Obtaining at least one structure binding pointer using at least one structure binding simulation tool based on the at least one protein structure file; Obtaining the peptide sequence and a human leukocyte antigen (HLA) type corresponding to the MHC sequence; Obtaining at least one affinity index using at least one affinity analysis model based on the peptide sequence and the human leukocyte antigen type; as well as The at least one structural binding indicator and the at least one affinity indicator are input into a machine learning model to obtain an integrated score.
2. The peptide-MHC bonding prediction method according to claim 1, wherein The at least one protein structure prediction model includes AlphaFold2, Colabfold or Pandora.
3. The peptide-MHC bonding prediction method according to claim 1, wherein: The at least one structural combination simulation tool includes AutoDock Vina, PISA (Proteins, Interfaces, Structures and Assemblies) or Local Distance Difference Test (LDDT).
4. The peptide-MHC bonding prediction method according to claim 3, wherein: If the at least one structural binding simulation tool is AutoDock Vina, the at least one structural binding indicator includes a binding score or a number of hydrogen bonds (H-Bond).
5. The peptide-MHC bonding prediction method according to claim 1, wherein If the at least one structural binding simulation tool is PISA, the at least one structural binding indicator includes p-value, hydrogen bond number (H-Bond) or folding energy.
6. The peptide-MHC bonding prediction method according to claim 1, wherein If the at least one structure binding simulation tool is a local distance difference test (LDDT), the at least one structure binding indicator includes a protein structure file distance difference (Iddt with other PDB).
7. The peptide-MHC bonding prediction method according to claim 1, wherein If the human leukocyte antigen type is type 1, the at least one affinity analysis model includes HLAthena, MHCflurry or NetMHCpan.
8. The peptide-MHC bonding prediction method according to claim 7, wherein: If the at least one affinity analysis model is HLAthena, the at least one affinity indicator includes probability.
9. The peptide-MHC bonding prediction method according to claim 7, wherein: If the at least one affinity analysis model is MHCflurry, the at least one affinity index includes probability or bonding affinity.
10. The peptide-MHC bonding prediction method according to claim 7, wherein: If the at least one affinity analysis model is NetMHCpan, the at least one affinity indicator includes probability or bonding affinity.
11. The peptide-MHC bonding prediction method according to claim 1, wherein If the human leukocyte antigen type is type II, the at least one affinity analysis model includes NetMHCIIpan.
12. The peptide-MHC bonding prediction method according to claim 11, wherein: If the at least one affinity analysis model is NetMHCIIpan, the at least one affinity index includes probability or bonding affinity.
13. A protein fragment (peptide)-major histocompatibility complex (MHC) bonding prediction device, comprising: An input unit for obtaining a peptide sequence and an MHC sequence; At least one protein structure prediction model, used to obtain at least one protein structure file (Protein Data Bank, PDB) based on the peptide sequence and the MHC sequence; at least one structure binding simulation tool for obtaining at least one structure binding pointer based on the at least one protein structure file; a classification unit for obtaining a human leukocyte antigen (HLA) type corresponding to the MHC sequence; at least one affinity analysis model for obtaining at least one affinity index based on the peptide sequence and the human leukocyte antigen type; and A machine learning model is used to obtain an integration score based on the at least one structural binding indicator and the at least one affinity index.