A post-processing method for predicting protein-ATP binding residues

By combining protein sequence and three-dimensional structural information, using I-LBR and COACH programs to generate a probability matrix and construct a support vector machine prediction model, the problem of high cost and insufficient accuracy of protein ATP binding residue prediction calculation in the prior art is solved, and efficient and accurate prediction effects are achieved.

CN114121143BActive Publication Date: 2025-05-09ZHEJIANG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111254444.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-27
Publication Date
2025-05-09
Estimated Expiration
2041-10-27

AI Technical Summary

Technical Problem

The existing protein ATP-bound residue prediction methods have shortcomings in terms of calculation cost and prediction accuracy, and it is difficult to meet the requirements of practical applications.

Method used

A method for predicting protein and ATP-bound residues based on post-processing is proposed. By combining protein sequence and three-dimensional structural information, a probability matrix is ​​generated using I-LBR and COACH programs, the feature matrix is ​​merged and extracted through a sliding window, a support vector machine prediction model is constructed, and post-processed to improve prediction accuracy.

Benefits of technology

ATP-bound residue prediction of protein ATP-bound residues with low calculation cost and high prediction accuracy is achieved. By taking into account both sequence and structural information, the relationship between residues is further captured and the prediction accuracy is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114121143B_ABST
    Figure CN114121143B_ABST
Patent Text Reader

Abstract

A method for predicting protein residues bound to ATP based on a post - processing approach generates a probability matrix of size \(L\times1\) for the sequence information and three - dimensional structure information of a protein, respectively, for protein - ligand binding residues; combines the two probability matrices into a matrix of size \(L\times2\), and uses a sliding window on it to obtain a feature matrix \(M\) of size \(17\times2\) for each residue in the protein; secondly, collects protein information with existing ATP - binding residue labels from the Protein Data Bank (PDB), constructs a training sample set after processing, and trains a prediction model using the support vector machine algorithm; thirdly, inputs the feature matrix \(M\) of the residues of the protein to be tested into the trained model and outputs the prediction result; when the C atoms of four residues among the predicted ATP - binding residues are not in the same plane, post - processes the residues predicted to be non - ATP - binding. The present invention has a low computational cost and high prediction accuracy. α When the C atoms of four residues among the predicted ATP - binding residues are not in the same plane, post - processes the residues predicted to be non - ATP - binding. The present invention has a low computational cost and high prediction accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics, deep learning and computer applications, and in particular to a method for predicting protein and ATP binding residues based on a post-processing approach. Background Art

[0002] As one of the small ligand molecules of proteins, ATP interacts with proteins through binding residues, changes protein structure, and further causes changes in protein activity. Studies have shown that protein ligand binding residues tend to cluster in space to form protein ligand binding sites, which are important targets for antibacterial and anticancer chemotherapy. Therefore, accurately locating protein ATP binding residues can not only help us understand the interaction between proteins and ATP, but also has important significance for protein function analysis and drug discovery.

[0003] At present, the methods for predicting protein ATP binding sites through deep learning include: ATPsite (Chen K, Mizianty MJ, Kurgan L. ATPsite: sequence-based prediction of ATP-binding residues. Proteome Science. 2011. Chen K et al. Sequence-based prediction of ATP binding residues. Proteomics Science, 2011), TargetATPsite (Yu DJ, Hu J, Huang Y, Shen HB, Qi Y, Tang ZM, Yang JY. TargetATPsite: a template-free method for ATP-binding sites prediction with residue evolution image sparse representation and classifier ensemble [J]. Journal of Computational Chemistry, 2013. Yu DJ et al. A template-free ATP binding site prediction method based on residue evolution image sparse representation and classifier ensemble [J]. Journal of Computational Chemistry, 2013) and ATPbind (Hu J, Li Y, Zhang Y, Yu DJ. ATPbind: Accurate Protein-ATP Binding Site Prediction by Combining Sequence-Profiling and Structure-Based Comparisons[J].Journal of Chemical Information&Modeling,2018. Hu J et al. Combining sequence analysis and structure-based comparison to accurately predict protein-ATP binding sites[J].Journal of Chemical Information & Modeling,2018). Compared with traditional machine learning methods, deep learning-based methods can automatically extract amino acid features and hidden patterns in protein sequences, and have achieved good results; however, these methods do not utilize the potential feature information contained in protein structure information. It is believed that full utilization of protein structure information can help improve the prediction accuracy of protein ATP binding residues.

[0004] In summary, the existing protein ATP binding residue prediction methods are still far from the requirements of practical applications in terms of computational cost and prediction accuracy, and are in urgent need of improvement. Summary of the invention

[0005] In order to overcome the deficiencies of existing protein ATP binding residue prediction methods in terms of computational cost and prediction accuracy, the present invention proposes a protein and ATP binding residue prediction method based on a post-processing method with low computational cost and high prediction accuracy.

[0006] The technical solution adopted by the present invention to solve its technical problem is:

[0007] A method for predicting protein-ATP binding residues based on post-processing, the method comprising the following steps:

[0008] 1) Input the sequence information and three-dimensional structure information of a protein with L amino acid residues for which ATP binding residues are to be predicted, denoted as P and S respectively;

[0009] 2) For the protein sequence information P, use the I-LBR program to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F1;

[0010] 3) For the three-dimensional structure information S of the protein, the COACH program is used to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F2;

[0011] 4) Combine F1 and F2 into a matrix of size L×2, denoted as F3, and concatenate a zero matrix of size 8×2 before the first row and after the last row, respectively. The concatenated matrix is ​​denoted as M fea ;

[0012] 5) Use a window of size 17×2 and a step size of 1 in the matrix M fea Slide up and down, and each time you slide, take the residue corresponding to the 8th row in the window as the prediction target, and extract a feature matrix of size 17×2, denoted as M;

[0013] 6) Collect protein information with existing ATP binding residue labels from the protein structure database PDB; for each residue in each protein, generate its feature matrix M through steps 1) to 5), and then combine it with the label information indicating whether it binds ATP to form corresponding sample data; use the sample data corresponding to all residues of all proteins to construct a training sample set;

[0014] 7) Based on the training sample set constructed in step 6), a prediction model of protein and ATP binding residues is trained using a support vector machine algorithm;

[0015] 8) For any protein to be tested, obtain the feature matrix M of each residue through step 1) to step 5), and input it into the prediction model trained by step 7), and output the probability P of the residue binding ATP; if the P value is greater than the threshold T, the residue is predicted to be an ATP-binding residue, otherwise it is predicted to be a non-ATP-binding residue;

[0016] 9) When there are four C residues in the predicted ATP binding residues in step 8) α When the atoms are not in the same plane, all residues predicted to be non-ATP binding in step 8) are post-processed as follows:

[0017] 9.1) For C α Any four predicted ATP-binding residues whose atoms are not in the same plane are identified according to their C α The atomic coordinates are calculated by the following equations:

[0018]

[0019] Among them, (x, y, z) are the spatial coordinates of the center O of the sphere, (x i ,y i , z i ), i = 1, 2, 3, 4, is the C of the i-th residue among the four predicted ATP binding residues α Atomic coordinates;

[0020] 9.2) The maximum value of all sphere radii R calculated in step 9.1) is recorded as R max and the center of the corresponding sphere is denoted as O max ;

[0021] 9.3) For each non-ATP binding residue predicted in step 8), calculate its C α The atom and the sphere center O obtained in step 9.2) max The Euclidean distance of o If d o Less than R max , the prediction result of the residue is changed to ATP binding residue.

[0022] The technical concept of the present invention is as follows: for a protein to be predicted for ATP binding residues with a given number of amino acid residues L, firstly, for the sequence information and three-dimensional structure information of the protein, the I-LBR program and the COACH program are used to generate a probability matrix of protein ligand binding residues with a size of L×1 respectively; then, the two probability matrices are merged into a matrix with a size of L×2, and a sliding window is used thereon to obtain a feature matrix M with a size of 17×2 for each residue in the protein; secondly, protein information with existing ATP binding residue labels is collected from the protein structure database PDB, and a training sample set is constructed after processing, and a prediction model is trained using a support vector machine algorithm; thirdly, the feature matrix M of the protein residues to be tested is input into the trained model, and a prediction result is output; when there are four residues with C in the predicted ATP binding residues, the prediction result is output; α When the atoms are not in the same plane, post-processing is performed on the residues predicted to be non-ATP binding: first calculate any four C α Atoms not in the same plane as the C of the predicted ATP binding residue α atoms on the sphere, and then calculate the C α The Euclidean distance from the atom to the center of the sphere is calculated, and finally, whether the predicted non-ATP binding residue is in the sphere is determined. If it is in the sphere, the final prediction result is an ATP binding residue. If it is not in the sphere, the prediction is kept as a non-ATP binding residue. The present invention proposes a protein and ATP binding residue prediction method based on a post-processing method with low computational cost and high prediction accuracy.

[0023] The beneficial effects of the present invention are as follows: the method adopts a post-processing method and utilizes feature information based on sequence and structure consistency. Structural information is more conservative than sequence information but is also more limited in quantity, while the difference between the predicted structure and the actual structure still exists. Taking both sequence and structural information into account can further capture the relationship between residues and better ensure the accuracy of protein ATP binding residue prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 Schematic diagram of a post-processing based method for predicting protein and ATP binding residues.

[0025] Figure 2 This is a graph showing the result of predicting ATP binding residues of a protein using a post-processing based protein and ATP binding residue prediction method. DETAILED DESCRIPTION

[0026] The present invention will be further described below in conjunction with the accompanying drawings.

[0027] Reference Figure 1 and Figure 2 , a protein and ATP binding residue prediction method based on post-processing, comprising the following steps:

[0028] 1) Input the sequence information and three-dimensional structure information of a protein with L amino acid residues for which ATP binding residues are to be predicted, denoted as P and S respectively;

[0029] 2) For the protein sequence information P, use the I-LBR program (https: / / jun-csbio.github.io / I-LBR) to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F1;

[0030] 3) For the three-dimensional structure information S of the protein, the COACH program (https: / / zhanggroup.org / COACH / ) is used to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F2;

[0031] 4) Combine F1 and F2 into a matrix of size L×2, denoted as F3, and concatenate a zero matrix of size 8×2 before the first row and after the last row, respectively. The concatenated matrix is ​​denoted as M fea ;

[0032] 5) Use a window of size 17×2 and a step size of 1 in the matrix M fea Slide up and down, and each time you slide, take the residue corresponding to the 8th row in the window as the prediction target, and extract a feature matrix of size 17×2, denoted as M;

[0033] 6) Collect protein information with existing ATP binding residue labels from the protein structure database PDB (https: / / www.rcsb.org / ); for each residue in each protein, generate its feature matrix M through steps 1) to 5), and then combine it with the label information indicating whether it binds ATP to form the corresponding sample data; use the sample data corresponding to all residues of all proteins to construct a training sample set;

[0034] 7) Based on the training sample set constructed in step 6), a prediction model of protein and ATP binding residues is trained using a support vector machine algorithm;

[0035] 8) For any protein to be tested, obtain the feature matrix M of each residue through step 1) to step 5), and input it into the prediction model trained by step 7), and output the probability P of the residue binding ATP; if the P value is greater than the threshold T, the residue is predicted to be an ATP-binding residue, otherwise it is predicted to be a non-ATP-binding residue;

[0036] 9) When there are four C residues in the predicted ATP binding residues in step 8) α When the atoms are not in the same plane, all residues predicted to be non-ATP binding in step 8) are post-processed:

[0037] 9.1) For C α Any four predicted ATP-binding residues whose atoms are not in the same plane are identified according to their C α The atomic coordinates are calculated by the following equations:

[0038]

[0039] Among them, (x, y, z) are the spatial coordinates of the center O of the sphere, (x i ,y i , z i ), i = 1, 2, 3, 4, is the C of the i-th residue among the four predicted ATP binding residues α Atomic coordinates;

[0040] 9.2) The maximum value of all sphere radii R calculated in step 9.1) is recorded as R max and the center of the corresponding sphere is denoted as O max ;

[0041] 9.3) For each non-ATP binding residue predicted in step 8), calculate its C α The atom and the sphere center O obtained in step 9.2) max The Euclidean distance of o If d o Less than R max , the prediction result of the residue is changed to ATP binding residue.

[0042] This example uses the prediction of protein ATP binding residues of protein 1AOIA as an example, and a method for predicting protein and ATP binding residues based on post-processing includes the following steps:

[0043] 1) Input the sequence information and three-dimensional structure information of a protein with L amino acid residues for which ATP binding residues are to be predicted, denoted as P and S respectively;

[0044] 2) For the protein sequence information P, use the I-LBR program (https: / / jun-csbio.github.io / I-LBR) to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F1;

[0045] 3) For the three-dimensional structure information S of the protein, the COACH program (https: / / zhanggroup.org / COACH / ) is used to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F2;

[0046] 4) Combine F1 and F2 into a matrix of size L×2, denoted as F3, and concatenate a zero matrix of size 8×2 before the first row and after the last row, respectively. The concatenated matrix is ​​denoted as M fea ;

[0047] 5) Use a window of size 17×2 and a step size of 1 in the matrix M fea Slide up and down, and each time you slide, take the residue corresponding to the 8th row in the window as the prediction target, and extract a feature matrix of size 17×2, denoted as M;

[0048] 6) Collect information on 388 proteins with ATP-binding residue labels from the protein structure database PDB (https: / / www.rcsb.org / ); for each residue in each protein, generate its feature matrix M through steps 1) to 5), and then combine it with the label information indicating whether it binds ATP to form the corresponding sample data; use the sample data corresponding to all residues of all proteins to construct a training sample set;

[0049] 7) Based on the training sample set constructed in step 6), a prediction model of protein and ATP binding residues is trained using a support vector machine algorithm;

[0050] 8) For the protein 1AOIA to be tested with 332 amino acid residues, obtain the feature matrix M of each residue through steps 1) to 5), and input it into the prediction model trained by step 7), and output the probability P of the residue binding ATP; if the P value is greater than the threshold value 0.5, the residue is predicted to be an ATP-binding residue, otherwise it is predicted to be a non-ATP-binding residue; the number of ATP-binding residues in protein 1AOIA is predicted to be 10;

[0051] 9) Four of the 10 ATP binding residues predicted in step 8) have C α Atoms are not in the same plane. Post-process all residues predicted as non-ATP binding in step 8):

[0052] 9.1) For C α Any four predicted ATP-binding residues whose atoms are not in the same plane are identified according to their C α The atomic coordinates are calculated by the following equations:

[0053]

[0054] Among them, (x, y, z) are the spatial coordinates of the center O of the sphere, (x i ,y i , z i ), i = 1, 2, 3, 4, is the C of the i-th residue among the four predicted ATP binding residues α Atomic coordinates;

[0055] 9.2) The maximum value of all sphere radii R calculated in step 9.1) is Denoted as R max , the spatial coordinates of the center of the sphere are (10.281, -19.941, 57.591), denoted as O max ;

[0056] 9.3) For the remaining 322 predicted non-ATP binding residues in step 8), calculate their C α The atom and the sphere center O obtained in step 9.2) max The Euclidean distance of o If d o Less than When , the predicted result of the residue is changed to ATP binding residue; the number of ATP binding residues predicted in protein 1AOIA is 13;

[0057] Taking the prediction of ATP binding residues of protein 1AOIA as an example, the prediction of protein 1AOIA is obtained by using the above method. Figure 2 shown.

[0058] The above description is the prediction result obtained by the present invention using the ATP binding residues of protein 1AOIA as an example, and does not limit the scope of implementation of the present invention. Various modifications and improvements made thereto without departing from the scope involved in the basic content of the present invention should not be excluded from the scope of protection of the present invention.

Claims

1. A method for predicting protein and ATP binding residues based on post-processing, characterized in that: The prediction method comprises the following steps: 1) Input the sequence information and three-dimensional structure information of a protein with L amino acid residues for which ATP binding residues are to be predicted, denoted as P and S respectively; 2) For the protein sequence information P, use the I-LBR program to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F1; 3) For the three-dimensional structure information S of the protein, the COACH program is used to generate a probability matrix of protein ligand binding residues of size L×1, denoted as F2; 4) Combine F1 and F2 into a matrix of size L×2, denoted as F3, and concatenate a zero matrix of size 8×2 before the first row and after the last row, respectively. The concatenated matrix is ​​denoted as M fea ; 5) Use a window of size 17×2 and a step size of 1 in the matrix M fea Slide up and down, and each time you slide, take the residue corresponding to the 8th row in the window as the prediction target, and extract a feature matrix of size 17×2, denoted as M; 6) Collect protein information with existing ATP binding residue labels from the protein structure database PDB; for each residue in each protein, generate its feature matrix M through steps 1) to 5), and then combine it with the label information indicating whether it binds ATP to form corresponding sample data; use the sample data corresponding to all residues of all proteins to construct a training sample set; 7) Based on the training sample set constructed in step 6), a prediction model of protein and ATP binding residues is trained using a support vector machine algorithm; 8) For any protein to be tested, obtain the feature matrix M of each residue through step 1) to step 5), and input it into the prediction model trained by step 7), and output the probability Q of the residue binding ATP; if the Q value is greater than the threshold T, the residue is predicted to be an ATP-binding residue, otherwise it is predicted to be a non-ATP-binding residue; 9) When there are four C residues in the predicted ATP binding residues in step 8) α When the atoms are not in the same plane, all residues predicted to be non-ATP binding in step 8) are post-processed as follows: 9.1) For C α Any four predicted ATP-binding residues whose atoms are not in the same plane are identified according to their C α The atomic coordinates are calculated by the following equations: Among them, (x, y, z) are the spatial coordinates of the center O of the sphere, (x i ,y i ,z i ), i = 1, 2, 3, 4, is the C of the i-th residue among the four predicted ATP binding residues. α Atomic coordinates; 9.2) The maximum value of all sphere radii R calculated in step 9.1) is recorded as R max and the center of the corresponding sphere is denoted as O max ; 9.3) For each non-ATP binding residue predicted in step 8), calculate its C α The atom and the sphere center O obtained in step 9.2) max The Euclidean distance of o If d o Less than R max , the prediction result of the residue is changed to ATP binding residue.

Citation Information

Patent Citations

  • Method for predicting epitope through cost-sensitive integrating and clustering on basis of sequence

    CN105868583A

  • DNA binding residue prediction method based on convolutional neural network

    CN112149881A