Structure-based model for evaluating equilibrium dissociation constants of peptide ligands for target proteins
By constructing a structure-based evaluation model for the equilibrium dissociation constant of peptide ligand molecules and target proteins, and using machine learning algorithms to screen for highly efficient peptide ligands, the problem of low efficiency in peptide drug screening in existing technologies has been solved. This has enabled a rapid and economical peptide screening method that can be applied to antiviral and antibacterial research.
Patent Information
- Application Number
- CN202210993384.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2042-08-18
AI Technical Summary
There are no systematic reports on existing virtual screening methods for peptide drugs targeting specific functional regions of target proteins, resulting in low drug screening efficiency and high costs.
A structure-based model for evaluating the equilibrium dissociation constant of peptide ligand molecules and target proteins was constructed. Using machine learning algorithms and experimental data, a peptide screening system was built based on the interaction data between peptide ligands and target proteins to screen out highly efficient peptide ligands.
This method improves the efficiency of peptide ligand screening, provides a rapid and reliable virtual screening method for peptides, and offers a reference for research in the fields of antiviral and antibacterial applications.
Smart Images

Figure CN115331729B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to peptide ligand screening, antiviral peptides for livestock and poultry, and animal immunology in the fields of basic veterinary medicine, specifically to a model for evaluating the equilibrium dissociation constant of peptide ligand molecules and target proteins based on structure. Background Technology
[0002] Pathogens typically initiate host invasion through the interaction of their own proteins with host proteins. In-depth research into the peptide regions of viral interaction proteins can provide insights into viral pathogenesis. These interaction regions are usually peptides of about 5 to 20 amino acid residues (aa), playing roles in related protein recognition, regulation, and signal transduction. Interfering with the interaction between viral-related interaction regions and host proteins can reduce viral load and alleviate symptoms; therefore, studying these interaction region amino acid peptides has become an important strategy for the screening, design, and development of antiviral peptide ligand drugs. Immunologically, biodisplay technology is mainly used to screen affinity peptide ligands, but this method is costly and time-consuming. Therefore, using virtual screening methods to study the interaction regions between viral peptides and target proteins will further improve the efficiency of drug screening and reduce the corresponding costs. Currently, however, there are no systematic reports on methods for virtual screening of peptide drugs targeting specific functional regions of target proteins. Summary of the Invention
[0003] To address the shortcomings of existing technologies, the purpose of this invention is to provide a structure-based model for evaluating the equilibrium dissociation constant of peptide ligand molecules and target proteins. Using actual experimental data and docking data between peptide ligands and target protein molecules, a peptide screening system based on machine learning algorithms is constructed, providing a new method for targeted virtual screening of peptide ligands and offering a reference for the establishment of other related drug screening systems.
[0004] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0005] The model for evaluating the equilibrium dissociation constant of a structure-targeted peptide ligand molecule and its target protein includes the following steps:
[0006] The reactivity data of amino acid peptide ligands and target proteins in the interaction region were collected through experiments. At the same time, the structural data of the amino acid peptide were analyzed. The overall data were classified into two categories, Active (A) and Unactive (UA), based on the actual equilibrium dissociation constant of the amino acid peptide and its receptor. A 1940×13 data matrix containing 1940 samples and 13 feature data was constructed by combining the classification information and feature information.
[0007] A random subset of the collected and classified data was constructed, including feature libraries targeting IgG series peptides and αβ42 series peptides. The machine learning algorithm was trained using the IgG series peptide feature library, while the system's predictive performance was initially evaluated using the αβ42 series peptide feature library. Important features were selected based on the Mean Decrease Gini coefficient provided by the machine learning algorithm. The machine learning algorithm was further trained based on the selected important features, and relevant parameters were adjusted to optimize the system accordingly.
[0008] An independent dataset containing important feature data for the PEDV S protein was constructed. An optimized machine learning classifier was used to further predict the collected independent data. The prediction results were compared with the actual peptide equilibrium dissociation constants to evaluate the performance of the classifier in practical applications.
[0009] In this scheme, the data used for system construction is obtained using the rDock program, including amino acid peptide classification data and their corresponding structural feature data, including a matrix of 1940 samples with 23 features such as INTER, INTER.POLAR, INTER.REPUL, INTER.ROT, INTER.VDW, INTER.NORM, INTRA, INTRA.DIHEDRAL, INTRA.DIHEDRAL0, INTRA.POLAR, INTRA.POLAR0, INTRA.REPUL, INTRA.REPUL0, INTRA.VDW, INTRA.VDW0, INTRA.NORM, RESTRSR, RESTR.NORM, SYSTEM, SYSTEM.DIHEDRAL, SYSTEM.NORM, HEAVY, and NORM. Features with a majority score of zero were removed, including INTER.POLAR, INTER.REPUL, INTRA.POLAR0, INTRA.REPUL, INTRA.REPUL0, RESTRSR, RESTR.NORM, SYSTEM, SYSTEM.DIHEDRAL, and SYSTEM.NORM, leaving 13 features in the entire dataset.
[0010] In this scheme, the equilibrium dissociation constant data between the peptide ligand and the receptor were obtained using SPR, and KD was set to 1 × 10⁻⁶ based on the actual ELISA experimental results. -5 Using a threshold, all samples are divided into A(KD≤1×10). -5 ) and UA(KD>1×10 -5 Two groups.
[0011] In this approach, training and testing datasets are constructed based on all samples in the dataset. The targeted IgG series peptide feature library is imported into a machine learning algorithm for training to obtain important feature information of the system and corresponding training parameters. Specifically:
[0012] The system is trained using a constructed training dataset and default parameters of a machine learning algorithm. Representative and important features are selected based on the ranking of average node impurity reduction and significance. The number of features used on nodes in the machine learning algorithm is optimized based on the obtained important features, and the optimized machine learning classifier system is rebuilt.
[0013] In this approach, system performance is evaluated by calculating the system's sensitivity, specificity, accuracy, Kappa value, and Matthews' correlation coefficient (MCC). The specific calculation formulas are as follows:
[0014]
[0015]
[0016]
[0017]
[0018] In addition, the receiver operating characteristic (ROC) is used to evaluate the relationship between sensitivity and specificity, and its area under the curve (AUC) is also calculated to evaluate system performance.
[0019] In this approach, the independent dataset consists of a batch of peptides redesigned based on pathogen structural proteins. These peptides, validated by ELISA and SPR experiments, are classified according to the aforementioned classification criteria. Combined with their corresponding equilibrium dissociation constant data, a new independent dataset is created. This dataset is used to further validate the prediction accuracy of the optimized machine learning classifier system, thereby evaluating its performance in practical applications.
[0020] This invention uses proteins of different sizes, such as IgG, αβ42, and PEDVS, as research subjects. Two peptide ligand libraries were designed and constructed, and molecular docking operations were performed between the peptide ligands and target proteins. The equilibrium dissociation constants of their interactions were determined using surface plasmon resonance (SPR) technology. Rapid screening and verification of the peptide ligand equilibrium dissociation constants were performed using enzyme-linked immunosorbent assay (ELISA). A relevant prediction system was constructed using machine learning algorithms, and the system was validated using independent data. This invention constructs a method for predicting actual equilibrium dissociation constants based on key information about peptide ligand-target protein interactions; it provides a rapid method for virtual peptide screening and a new and reliable approach for obtaining peptides with high equilibrium dissociation constants.
[0021] The beneficial effects of this invention are:
[0022] To address the lack of methods for predicting the equilibrium dissociation constant of the interaction between peptide ligands and target proteins, this invention analyzes the amino acid peptide information of the interaction region between the protein and its peptide ligands, synthesizes the corresponding viral protein peptides, collects peptide equilibrium dissociation constant data and their structural feature score data to construct a dataset, and builds an evaluation model for the equilibrium dissociation constant of a specific region of the target protein with its peptide ligands.
[0023] The prediction method of this invention can effectively and rapidly predict the peptide regions on unknown viral proteins that bind to their ligands based on the structural information characteristics of the amino acid peptides that bind to the ligands.
[0024] This invention can effectively improve the efficiency of peptide ligand screening for target proteins and is beneficial for application research in antiviral, antibacterial and other related fields. It also provides a reference for the construction of related prediction systems. Attached Figure Description
[0025] Figure 1 A schematic diagram of the process for predicting the equilibrium dissociation constant of protein-peptide interactions according to the present invention.
[0026] Figure 2 The distribution of scores for the 13 features after filtering.
[0027] Figure 3 Results of peptide affinity test in ELISA detection.
[0028] Figure 4 Distribution of peptide ligand KD values in SPR assay.
[0029] Figure 5 Linear relationship between peptide ligand ELISA OD value and SPR KD value.
[0030] Figure 6ROC curves and AUC values of machine learning classifiers built based on different numbers of important features. Detailed Implementation
[0031] The specific embodiments of the present invention will be further described in detail below with reference to examples.
[0032] Example 1. Obtaining and processing characteristic information of peptide ligand-target protein interaction.
[0033] (1) Construction of virtual peptide library: Based on the protein structure database (Protein Data Bank, PDB https: / / www.rcsb.org / ) provided with the IgG and αβ42 protein crystal structures, the peptide ligands that interact with the active regions of IgG and αβ42 were predicted using a molecular docking program. The peptide ligands with higher scores were selected to construct an information library containing 1940 peptide ligands with a length of 6 amino acids.
[0034] (2) Analysis of interaction forces between peptide ligands and target proteins: Using the obtained peptide ligand information database and the protein crystal structures of IgG, αβ42, and PEDV S proteins, the intermolecular and intramolecular interactions between peptide ligands and target proteins, as well as the energy and non-physical constraints of the flexible regions of active sites, were analyzed using rDock software. The following interaction force scores were obtained, including INTER, INTER.POLAR, INTER.REPUL, INTER.ROT, INTER.VDW, INTER.NORM, INTRA, INTRA.DIHEDRAL, INTRA.DIHEDRAL0, INTRA.POLAR, INTRA.POLAR0, INTRA.REPUL, INTRA.REPUL0, INTRA.VDW, INTRA.V The dataset contains 23 feature metrics, including DW0, INTRA.NORM, RESTRSR, RESTR.NORM, SYSTEM, SYSTEM.DIHEDRAL, SYSTEM.NORM, HEAVY, and NORM. Features with a majority of zero scores, such as INTER.POLAR, INTER.REPUL, INTRA.POLAR0, INTRA.REPUL, INTRA.REPUL0, RESTRSR, RESTR.NORM, SYSTEM, SYSTEM.DIHEDRAL, and SYSTEM.NORM, were removed. This leaves 13 features in the dataset. The density distribution of these 13 features is shown in the image. Figure 2 .
[0035] (3) Peptide synthesis: The obtained peptides (information library of 1940 peptide ligands) were given to a biotechnology company for synthesis. To facilitate ELSA detection, the N-terminus of the peptides was labeled with biotin.
[0036] Example 2. Detection of the equilibrium dissociation constant between viral protein peptides and target proteins
[0037] The equilibrium dissociation constant of the interaction between the peptide ligand and the target protein was determined using ELISA and SPR techniques.
[0038] (1) ELISA detection of synthetic peptides and target proteins
[0039] The affinity of the synthesized peptides for IgG, αβ42, and PEDV S proteins was verified by indirect ELISA, with the specific steps as follows:
[0040] 1) Coat ELISA plates with IgG, αβ42, and PEDVS proteins at a final concentration of 20 μg / mL using coating buffer. After incubation at 37°C for 2 h, block with 5% skim milk.
[0041] 2) Dilute the synthesized biotin-labeled polypeptide (obtained in Example 1) to 1 μg / mL with PBS and add it to the ELISA plate prepared in 1), 50 μL / well, and incubate at 37°C for 1 h. Use the corresponding mouse monoclonal antibody or polyclonal antibody of the target protein as a positive control and PBS buffer as a negative control.
[0042] 3) Add HRP-labeled avidin antibody diluted 1:4000, and use HRP-labeled anti-mouse antibody as secondary antibody. Add positive control and incubate at 37℃ for 30 min.
[0043] 4) Add 100 μL of 3,3',5,5'-tetramethylbenzidine (TMB) substrate solution to each well and allow to develop color at room temperature for 10 min.
[0044] 5) Add 50 μL of stop solution to each well to terminate the reaction, and measure the OD value at 450 nm using a microplate reader. If the ratio of the OD of the peptide reaction well to the OD of the negative well is greater than 2.1, it is considered positive, meaning that the peptide and the target protein can undergo an affinity reaction.
[0045] The results show ( Figure 3 a) Of the 1400 peptides, 1400 showed good affinity for the target protein, while 540 showed no affinity; the overall OD values ranged from 0.01 to 2.49. Figure 3 b).
[0046] (2) SPR detection of synthetic peptides and target proteins
[0047] The equilibrium dissociation constant KD value between the synthesized peptide and the target proteins (IgG, αβ42, and PEDVS proteins) was detected using a Biacore X100 instrument. The specific steps are as follows:
[0048] 1) The target protein is coupled to the chip required by the instrument using the EDC / NHS method;
[0049] 2) The synthesized peptide (Example 1) was diluted to six different concentrations using purchased HBS-EP buffer and loaded into the instrument at a rate of 30 μL / min. The changes in the resonance signal between the peptide and the target protein on the chip were detected. After the peptide and protein had fully reacted, the unbound peptide was washed away by passing HBS-EP buffer through the chip. Then, all peptides on the chip were completely eluted with 0.25% SDS solution before detecting the second synthesized peptide. This cycle was repeated until all peptides were detected.
[0050] The results show ( Figure 4 The KD values of all peptides are between 2.72 × 10⁻⁶. -11 ~4.67×10 -3 Within the range.
[0051] Example 3. Dataset Construction and Processing
[0052] Combining ELISA OD values and SPR KD values, it was found that ( Figure 5 These two values are negatively correlated, with a correlation coefficient of -0.90. Data shows that KD = 1 × 10⁻⁶. -5 Its affinity threshold is when KD ≤ 1 × 10 -5 When the peptide ligand and target protein have good affinity, a good response can be achieved, and this is classified as the Active(A) group; when the KD value is >1×10 -5 When there is no affinity between the peptide ligand and the target protein, it is classified as Unactive (UA).
[0053] Combining the information on the regional interaction between peptide ligands and target proteins from Example 1 with the above classification information, a dataset for machine learning is constructed; with KD = 1 × 10 -5 Using a threshold, all samples are divided into Active(A, KD≤1×10) classes. -5 ) and Unactive(UA, KD>1×10 -5 Two groups of peptide ligand equilibrium dissociation constants were used to verify the experimental results and collect characteristic data on the equilibrium dissociation constants. A dataset of 1940 (number of samples) × 13 (number of features) was constructed, and a peptide feature library targeting IgG series and a peptide feature library targeting αβ42 series were constructed based on this dataset. The dataset was then imported into a machine learning algorithm construction system.
[0054] Simultaneously, an independent dataset was constructed for the PEDV S protein, containing information on the regional interaction forces between the peptide ligand and the target protein, as well as classification information, to be used for subsequent evaluation of the accuracy of the system's actual predictions.
[0055] Example 4. Construction, parameter optimization, and validation of a machine learning system for peptide ligand protein equilibrium dissociation constants.
[0056] (1) Establishment of a system classifier targeting IgG series peptide feature library and a system classifier targeting αβ42 series peptide feature library
[0057] To build the classifier for the machine learning system, the experimental dataset was randomly divided into a feature library targeting IgG series peptides and a feature library targeting αβ42 series peptides in a 7:3 ratio. Specifically, when training with 4 features, the feature library targeting IgG series peptides contained 986 alpha values and 372 analytes, while the feature library targeting αβ42 series peptides contained 414 alpha values and 168 analytes; when training with 3 features, the feature library targeting IgG series peptides contained 986 alpha values and 372 analytes, while the feature library targeting αβ42 series peptides contained 414 alpha values and 168 analytes; and when training with 3 features, the feature library targeting IgG series peptides contained 977 alpha values and 381 analytes, while the feature library targeting αβ42 series peptides contained 423 alpha values and 159 analytes.
[0058] (2) Preliminary establishment of machine learning classification system and selection of important features
[0059] The experimental dataset containing all 13 equilibrium dissociation constants, constant features, and classification information was imported into a machine learning algorithm for learning. The system was built using the default parameter settings. The importance of each feature was ranked according to the average Gini coefficient reduction provided by the algorithm. Important features were determined based on the significance of feature importance. The results (Table 1) show that INTRA.VDW0, INTRA.DIHEDRAL0, HEAVY, and INTER.ROT are the four features that significantly affect the system's prediction results.
[0060] Table 1 Importance of various physicochemical characteristics of peptide ligands
[0061] feature Average Gini coefficient reduction INTRA.VDW0 214.00* INTRA.DIHEDRAL0 193.43* HEAVY 161.88* INTER.ROT 152.18* NORM 87.86 INTRA.DIHEDRAL 82.12 INTRA.VDW 58.70 INTER 56.92 INTER.VDW 54.53 INTER.NORM 51.90 INTRA.POLAR 49.84 INTRA 47.15 NTRA.norm 46.82
[0062] * P < 0.05 indicates feature importance.
[0063] (3) Formal establishment and optimization of machine learning classification system
[0064] Four key features from the target IgG series peptide feature library were selected, and 4, 3, and 2 features were chosen according to their importance to train the machine learning system. Each system had 500 branches. To prevent overfitting, each classification system underwent 10-fold cross-validation 10 times, and the number of features at each branch node in the machine learning algorithm was optimized.
[0065] The system was evaluated by calculating its accuracy, Kappa value, and Matthews's correlation coefficient (MCC).
[0066] In addition, the receiver operating characteristic (ROC) is used to assess the relationship between sensitivity and specificity, and its area under the curve (AUC) is also calculated to evaluate system performance.
[0067] The specific calculation formula is as follows:
[0068]
[0069]
[0070]
[0071]
[0072] Nine machine learning algorithm systems were constructed with different numbers of features and features per node. The results are shown in Table 2. Machine learning systems constructed using 4, 3, and 2 important features, respectively, achieved high accuracy and Kappa values regardless of the number of features used on each branch node. Specifically, when the system was constructed with 4 features and 4 features were used on each branch node, the system was optimal (accuracy 99.03%, Kappa value 0.9755); when the system was constructed with 3 features and 3 features were used on each node, the system was optimal (accuracy 98.95%, Kappa value 0.9736); and when the system was constructed with 2 features and 2 features were used on each node, the system was optimal (accuracy 99.15%, Kappa value 0.9786). The highest accuracy and Kappa value were achieved when using 2 features and 2 features per node.
[0073] Table 2 shows the performance of classification systems using machine learning classifiers built with different numbers of importance features.
[0074]
[0075] Further research revealed that when the number of node features was optimized, the three optimized machine learning systems were validated using data from a targeted αβ42 series peptide feature library. Validation metrics included test accuracy, Matthews's correlation coefficient (MCC), and AUC value.
[0076] As shown in Table 3, the prediction accuracy using the optimized 4-feature and 3-feature systems was 98.79%, while the prediction accuracy using the optimized 2-feature system was 98.96%. The MCC value of the systems built with 4 and 3 features was 0.971, while the MCC value of the system built with 2 features was 0.974. Therefore, the above analysis shows that although the classifier systems exhibit similarities when using different feature models and different numbers of node features, the classifier achieves optimal performance when trained with two features and applied to both nodes.
[0077] ROC curves were plotted and AUC values were calculated for the optimal classifier systems constructed with 4, 3, and 2 features respectively (i.e., the systems with the highest accuracy and Kappa value). The results show ( Figure 6 The AUC values of the machine learning classification systems built with 4 and 3 features are both 0.9996. The machine learning classifier system built with 2 features has better performance and its AUC area is also larger at 0.9998.
[0078] Table 3. Predictive performance of different machine learning classification systems on the feature library of αβ42 series peptides.
[0079]
[0080] Example 5. Independent Data Validation of a Machine Learning System for Peptide-Ligand Protein Equilibrium Dissociation Constant
[0081] For the PEDV S protein sequence (GenBank Accession No. KF664124) protein crystal structure, an independent database containing four important features and classification information was constructed according to the methods in Examples 1 to 3. Peptide ligand equilibrium dissociation constant category information and equilibrium dissociation constant constant feature information were collected to construct an independent dataset of 1120 (number of samples) × 4 (number of important features) for system validation testing.
[0082] The optimized classifier was used to predict independent datasets, and the system accuracy was predicted. The results (Table 4) show that, compared with the accuracy of other systems, the system built using 4 features has a higher machine learning prediction accuracy (71.07%), while the machine learning prediction accuracy of systems built using 3 and 2 features is only 64.29% and 60.71%, respectively. Furthermore, the three constructed systems performed well in predicting type A peptide ligands. When the machine learning classification systems were built using 4, 3, and 2 features, the accuracy rates of correctly predicting positive results as a percentage of true positives were 80.43%, 70.21%, and 65.96%, respectively.
[0083] Therefore, although the system built with two features has good performance, it may cause overfitting due to the small number of prediction features considered, resulting in a decrease in prediction accuracy. On the other hand, using all four important features to build a machine learning algorithm classifier is more suitable for actual prediction.
[0084] Table 4. Accuracy of different machine learning systems in predicting independent datasets
[0085]
Claims
1. A structure-based model for evaluating the equilibrium dissociation constant of peptide ligands and target proteins, characterized in that, The construction of this evaluation model includes the following steps: (1) Acquisition of characteristic data on the interaction between peptide ligands and their receptor proteins: Using specific proteins as research objects, a series of peptide ligand molecules are designed, molecular docking operations are performed between peptide ligands and target proteins, and characteristic information on the binding of peptide ligands to proteins is obtained. (2) Acquisition of peptide ligand equilibrium dissociation constant data: The equilibrium dissociation constant of the interaction between peptide ligand and target protein was determined by ELISA and SPR technology; (3) Construction of the algorithm system dataset and the independent validation dataset; the algorithm system dataset includes a feature library of targeted IgG series peptides and a feature library of targeted αβ42 series peptides; Based on the experimental verifiability of peptide ligand equilibrium dissociation constant and the characteristic data of equilibrium dissociation constant, a dataset is constructed to build a feature library of targeted IgG series peptides and a feature library of targeted αβ42 series peptides. The dataset is then imported into a machine learning algorithm construction system. (4) A machine learning algorithm classifier system was constructed using the targeted IgG series peptide feature library, important feature data were screened, relevant parameters were optimized, and the system performance was evaluated by testing the targeted αβ42 series peptide feature library. Four important features were selected from the target IgG series peptide feature library. Based on their importance, 4, 3, and 2 features were selected respectively to train the machine learning system. Each system had 500 branches. To prevent overfitting, each classification system underwent 10x10 cross-validation, and the number of features at each branch node in the machine learning algorithm was optimized. (5) Evaluate the actual prediction performance of the system using the constructed independent validation dataset; Using the Spike protein of porcine epidemic diarrhea virus as the research object, we collected information on the category of peptide ligand equilibrium dissociation constants and characteristic data of equilibrium dissociation constants, constructed an independent dataset, and used it to optimize the system's validation test.
2. The evaluation model as described in claim 1, characterized in that, The four corresponding key features are INTRA.VDW0, INTRA.DIHEDRAL0, HEAVY, and INTER.ROT.
3. The evaluation model as described in claim 1, characterized in that, Nine machine learning algorithm systems were constructed based on different numbers of features and features at different nodes. The systems were evaluated by calculating their sensitivity, specificity, accuracy, Kappa value, and Matthew correlation coefficient (MCC). The specific calculation formulas are as follows: 。 4. The application of the evaluation model as described in any one of claims 1-3 in the screening of peptide ligands for target proteins.
5. The application of the evaluation model as described in any one of claims 1-3 in peptide drug screening.
Citation Information
Patent Citations
Systems and methods for predicting potential inhibitors of target protein
US20220084627A1
Computer screening method for chemical small molecule medication targeting RNA
WO2019232748A1