Enzyme turnover number prediction method based on deep learning model and protein structure information
By constructing a data set of enzyme sequence-enzyme structure-substrate combination and enzyme turnover number mapping, and combining deep learning models, using protein structure information to predict enzyme turnover number, the problems of low prediction accuracy and limited sequence similarity in the prior art are solved, and higher prediction accuracy and generalization ability are achieved.
Patent Information
- Application Number
- CN202311512721.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-11-14
AI Technical Summary
The existing deep learning models have low accuracy in predicting enzyme turnover numbers and fail to effectively utilize the protein structure information of the enzyme, resulting in the prediction effect being limited by sequence similarity.
By constructing a data set of enzyme sequence-enzyme structure-substrate combination and enzyme turnover number mapping, combining deep learning models, including the Transformer model and a graph convolutional neural network of multi-head attention mechanism, we use protein structure information to enhance model generalization capabilities.
It significantly improves the accuracy of the prediction of enzyme turnover number, reduces the negative impact of different sequence similarities on prediction, and can predict the impact of point mutations on enzyme turnover number to a certain extent.
Smart Images

Figure CN120015107A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technology in the field of bioengineering, specifically a method for predicting enzyme turnover number based on deep learning models and protein structure information. Background Art
[0002] Accurate prediction of enzyme turnover number is crucial for rational protein modification and cell metabolism modeling. Although a large number of enzyme turnover values are available in the database, the number of enzymes with experimentally measured enzyme turnover numbers is far less than the number of sequenced proteins. The use of artificial intelligence technology can significantly reduce the experimental cost of large-scale enzyme turnover number acquisition. However, the existing deep learning model has low technical accuracy for prediction and basically does not explore the deep relationship related to enzyme turnover number in the protein structure of the enzyme. Summary of the invention
[0003] In view of the shortcomings of the prior art, the present invention proposes an enzyme turnover number prediction method based on deep learning model and protein structure information. By adding protein structure information to enzyme turnover number prediction training and using deep learning method to explore the potential relationship between the two, the generalization effect of the model can be effectively enhanced, and the problems of low accuracy of previous prediction models and being limited by sequence similarity can be solved.
[0004] The present invention is achieved through the following technical solutions:
[0005] The present invention relates to an enzyme turnover number prediction method based on a deep learning model and protein structure information. A data set with unique enzyme sequence-enzyme structure-substrate combination and enzyme turnover number mapping is constructed in an offline stage to train an enzyme turnover number prediction model. In an online stage, the enzyme turnover number of a protein structure to be processed is predicted in real time by using the trained enzyme turnover number prediction model.
[0006] The dataset of enzyme sequence-enzyme structure-substrate combination and enzyme turnover number mapping is constructed in the following way:
[0007] Step 1, obtain enzyme catalytic reaction data from the Brenda database and SABIO-RK database, including: enzyme ECnumber, species origin, sequence, substrate molecule name, substrate molecule SMILES and enzyme catalytic reaction turnover value;
[0008] Step 2: In order to solve the problem that the existing data set has high similarity of the same enzyme sequence, which will lead to falsely high prediction results, the data set was screened and reconstructed. The enzymes with the same substrate molecules and sequence similarity higher than 90% in the data set were deduplicated, and only the enzyme-substrate pairs with the longest enzyme sequence were retained to construct a data set of enzyme sequence-substrate combinations and enzyme turnover number mapping.
[0009] Step 3: Use the fast protein structure prediction software: ColabFold to obtain the protein structure predicted based on the enzyme sequence data in step 2, and construct a data set of enzyme sequence-enzyme structure-substrate combination and enzyme turnover number mapping without data feature conversion;
[0010] Step 4: Convert protein sequence data, protein structure data and substrate molecule SMILES data into initial features as model input, including:
[0011] Step 4.1, using the RDKit software package to convert the substrate molecule SMILES string into a molecular fingerprint and a contact matrix, where the molecular fingerprint represents the small molecule atoms as nodes of the contact matrix and the intermolecular chemical bonds as edges of the contact matrix; the initial characteristics of the substrate molecule are obtained;
[0012] Step 4.2: Encode the enzyme sequence, with each group of 4 adjacent amino acids being encoded starting from 0. The same amino acid combination is encoded in the same way to obtain the initial features characterizing the enzyme sequence. There are 137,026 different amino acid combinations in the constructed data set;
[0013] Step 4.3, based on the three-dimensional structure of the enzyme protein predicted in step 3, obtain the three-dimensional coordinate information of the protein amino acids;
[0014] Step 4.4: Using amino acid nodes as graph structure nodes, calculate the Euclidean distance between two amino acids based on the three-dimensional coordinates of the amino acids. When , it is considered that there is an edge between the two amino acids, where (x1, y1, z1) and (x2, y2, z2) are the corresponding three-dimensional coordinates of the amino acids, and the initial structural characteristics of the protein are obtained;
[0015] The enzyme turnover number prediction model includes: a Transformer model and a graph convolutional neural network with a multi-head attention mechanism, wherein: the initial features of the substrate molecule are first processed by the graph convolutional neural network with a multi-head attention mechanism, specifically: the initial features are: and attention mechanism For further feature extraction, where: σ is the activation function, is the adjacency matrix, is the degree matrix, H (l) is the node feature, W (l) are the graph convolutional neural network model parameters; the initial features of the enzyme sequence are then processed through the Transformer model, specifically: The attention weight feature matrix used to obtain the enzyme sequence, where softmax is the activation function, Q, K, V are the task representation vectors, and d kis the number of dimensions; finally, the graph convolutional neural network is used to process the initial features of the enzyme structure, which are: For further feature extraction, where: σ is the activation function, is the adjacency matrix, is the degree matrix, H (l) is the node feature, W (l) are the graph convolutional neural network model parameters.
[0016] The encoding module in the Transformer model obtains the initial features of the enzyme sequence, and the decoding module uses a fully connected layer to comprehensively decode the enzyme catalytic reaction matrix.
[0017] A multi-head attention module composed of multiple self-attention combinations is arranged between each layer of the graph convolutional neural network to explore the different effects of different amino acid positions of the enzyme on the enzyme turnover number, thereby enhancing the accuracy and generalization ability of the model.
[0018] The training described above is based on R 2 and RMSE were used as the accuracy criteria of the enzyme turnover number prediction model, and the model parameters were updated and optimized based on the preset learning rate.
[0019] The real-time prediction refers to: combining the substrate feature matrix, enzyme sequence matrix and enzyme structure matrix of each group of enzyme-substrate pairs output by the enzyme turnover number prediction model to obtain an enzyme catalytic reaction matrix, using a fully connected layer to operate the enzyme catalytic reaction matrix, and finally obtaining the final turnover value by: y=Wx.
[0020] The present invention relates to a system for implementing the above method, comprising: a data collection unit, a data preprocessing unit, a model building unit, a model training unit and a result prediction unit, wherein: the data collection unit obtains enzyme catalytic reaction data from the Brenda database and the SABIO-RK database; the data preprocessing unit predicts the protein structure using ColabFold, and obtains the initial features of the protein structure according to the structural information, obtains the initial features of the substrate molecule using RDKit, and continues to encode the amino acids to obtain the initial features of the protein sequence; the model building unit combines graph convolutional neural networks with different numbers of layers and Transformer to establish a deep learning model framework; the model training unit performs model training according to the initial features obtained from the data preprocessing part to obtain a trained deep learning model; the result prediction unit predicts the test set according to the enzyme and substrate molecule information of the test set and obtains an evaluation result. Technical Effects
[0021] The present invention constructs a new data set that can reduce the influence of sequence similarity on the one-to-one mapping of enzyme sequence-enzyme structure-substrate molecule combination and enzyme turnover number; adds the three-dimensional structure data of the enzyme, and adds a multi-head attention module on the basis of the graph convolutional neural network. Compared with the prior art, the present invention significantly improves the prediction effect of enzyme turnover number while significantly reducing the negative impact of different sequence similarities on enzyme turnover number prediction; to a certain extent, it can predict the influence of point mutations on enzyme turnover number. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic diagram of the process of the present invention;
[0023] Figure 2 It is a schematic diagram of the overall structure of the system;
[0024] Figure 3 It is the prediction effect diagram of the embodiment;
[0025] Figure 4 This is a diagram showing the prediction effect of the embodiment under different enzyme sequence similarities;
[0026] Figure 5 This is a diagram showing the predicted effects of the embodiment under different point mutations. DETAILED DESCRIPTION
[0027] After specific practical experiments, this embodiment uses Python to build a model in a Windows environment and uses GPU for training. The model training starts with random initialization parameters. The model prediction effect of the test set is as follows Figure 3 As shown in Table 1, the prediction accuracy is 0.6. The specific prediction results are shown in Table 1, which gives examples of the enzyme EC number, substrate name, and predicted enzyme turnover number of the five enzyme-catalyzed reactions predicted. The model prediction effect under different enzyme sequence similarities in the test set is shown in Table 1. Figure 4 as shown.
[0028] like Figure 1 As shown, the specific steps of this embodiment include:
[0029] Step 1. Collect enzyme catalytic reaction data, pre-process the collected enzyme catalytic reaction data, remove duplicate enzymes with the same substrate molecules and sequence similarity higher than 90% in the data set, retain only the enzyme-substrate pairs with the longest enzyme sequence length, and construct an enzyme catalytic reaction data set that can reduce the influence of enzyme sequence similarity. Use the ColabFold protein structure prediction method to predict the structure of the enzyme protein sequence, and construct a data set with a one-to-one mapping of enzyme sequence-enzyme structure-substrate molecule combination and enzyme turnover number based on each enzyme catalytic reaction.
[0030] Step 2: Based on the sequence data, structure data and substrate molecule data of the enzyme, the RDKit software package is used to convert the substrate molecule SMILES string into a molecular fingerprint and a contact matrix to obtain the initial features characterizing the substrate molecule; the enzyme sequence is encoded, and each 4 adjacent amino acids are grouped as a group, and the encoding starts from 0. The same amino acids are combined into the same code to obtain the initial features characterizing the enzyme sequence; the amino acid nodes are used as graph structure nodes, and the Euclidean distance between the two amino acids is calculated based on the three-dimensional coordinates of the amino acids. When , it is considered that there is an edge between the two amino acids, where (x1, y1, z1) and (x2, y2, z2) are the corresponding three-dimensional coordinates of the amino acids, and the initial structural characteristics of the protein are obtained; the initial characteristics of the enzyme sequence, enzyme structure and substrate molecule are used as input, and the input data is randomly divided into 80% training set, 10% validation set and 10% test set.
[0031] Step 3: Use the deep learning model: add the graph neural network with multi-head attention mechanism, the encoding module of Transformer and the fully connected layer as the decoding module to build an enzyme turnover number prediction model based on protein structure information. The specific prediction model structure is shown in Figure 2 .
[0032] Step 4: Based on the data set constructed in step 1 and the initial features extracted in step 2, the prediction model in step 3 is trained to obtain the model prediction accuracy of the prediction task, and the training is repeated multiple times to find the model parameters with the highest prediction accuracy.
[0033] Step 5: Use the model training weights with the best prediction accuracy to predict the enzyme turnover number, and obtain the final enzyme turnover number predicted based on protein structure information.
[0034] Table 1 Experimental data
[0035] Compared with the existing technology, this method significantly improves the prediction accuracy and significantly reduces the negative impact of different sequence similarities on the prediction of enzyme turnover number; to a certain extent, it can predict the impact of point mutations on enzyme turnover number and identify key amino acid residue sites.
[0036] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.
Claims
1. A method for predicting enzyme turnover number based on deep learning model and protein structure information, characterized in that: In the offline stage, a dataset of unique enzyme sequence-enzyme structure-substrate combinations and enzyme turnover number mappings is constructed to train an enzyme turnover number prediction model. In the online stage, the trained enzyme turnover number prediction model is used to predict the enzyme turnover number of the protein structure to be processed in real time.
2. The method for predicting enzyme turnover number based on deep learning model and protein structure information according to claim 1, characterized in that: The dataset of enzyme sequence-enzyme structure-substrate combination and enzyme turnover number mapping is constructed in the following way: Step 1, obtain enzyme catalytic reaction data from the Brenda database and SABIO-RK database, including: enzyme ECnumber, species origin, sequence, substrate molecule name, substrate molecule SMILES and enzyme catalytic reaction turnover value; Step 2: In order to solve the problem that the existing data set has a high similarity of the same enzyme sequence, which will lead to an inflated prediction effect, the data set was screened and reconstructed; the enzymes with the same substrate molecules and a sequence similarity higher than 90% in the data set were deduplicated, and only the enzyme-substrate pair with the longest enzyme sequence was retained, and a data set of enzyme sequence-substrate combination and enzyme turnover number mapping was constructed; Step 3: Use the fast protein structure prediction software ColabFold to obtain the protein structure predicted based on the enzyme sequence data in step 2, and construct a dataset of enzyme sequence-enzyme structure-substrate combination and enzyme turnover number mapping without data feature conversion; Step 4: Convert protein sequence data, protein structure data and substrate molecule SMILES data into initial features as model input, including: Step 4.1, using the RDKit software package to convert the substrate molecule SMILES string into a molecular fingerprint and a contact matrix, where the molecular fingerprint represents the small molecule atoms as nodes of the contact matrix and the intermolecular chemical bonds as edges of the contact matrix; the initial characteristics of the substrate molecule are obtained; Step 4.2, encode the enzyme sequence, with each group of 4 adjacent amino acids being encoded starting from 0, and the same amino acid combination being the same code, to obtain the initial features characterizing the enzyme sequence; there are a total of 137,026 different amino acid combinations in the constructed data set; Step 4.3, based on the three-dimensional structure of the enzyme protein predicted in step 3, obtain the three-dimensional coordinate information of the protein amino acids; Step 4.4: Using amino acid nodes as graph structure nodes, calculate the Euclidean distance between two amino acids based on the three-dimensional coordinates of the amino acids. When , it is considered that there is an edge between the two amino acids, where (x1, y1, z1) and (x2, y2, z2) are the corresponding three-dimensional coordinates of the amino acids, and the initial structural characteristics of the protein are obtained.
3. The method for predicting enzyme turnover number based on deep learning model and protein structure information according to claim 1, characterized in that: The enzyme turnover number prediction model includes: a Transformer model and a graph convolutional neural network with a multi-head attention mechanism, wherein: the initial features of the substrate molecule are first processed by the graph convolutional neural network with a multi-head attention mechanism, specifically: the initial features are: and attention mechanism For further feature extraction, where: σ is the activation function, is the adjacency matrix, is the degree matrix, H (l) is the node feature, W (l) are the graph convolutional neural network model parameters; the initial features of the enzyme sequence are then processed through the Transformer model, specifically: The attention weight feature matrix used to obtain the enzyme sequence, where softmax is the activation function, Q, K, V are the task representation vectors, and d k is the number of dimensions; finally, the graph convolutional neural network is used to process the initial features of the enzyme structure, which are: For further feature extraction, where: σ is the activation function, is the adjacency matrix, is the degree matrix, H (l) is the node feature, W (l) are the graph convolutional neural network model parameters.
4. The method for predicting enzyme turnover number based on deep learning model and protein structure information according to claim 3, characterized in that: The encoding module in the Transformer model obtains the initial features of the enzyme sequence, and the decoding module uses a fully connected layer to comprehensively decode the enzyme catalytic reaction matrix.
5. The method for predicting enzyme turnover number based on deep learning model and protein structure information according to claim 3, characterized in that: A multi-head attention module composed of multiple self-attention combinations is arranged between each layer of the graph convolutional neural network to explore the different effects of different amino acid positions of the enzyme on the enzyme turnover number, thereby enhancing the accuracy and generalization ability of the model.
6. The method for predicting enzyme turnover number based on deep learning model and protein structure information according to claim 1, characterized in that: The training described above is based on R 2 and RMSE were used as the accuracy criteria of the enzyme turnover number prediction model, and the model parameters were updated and optimized based on the preset learning rate.
7. The method for predicting enzyme turnover number based on deep learning model and protein structure information according to claim 1, characterized in that: The real-time prediction refers to: combining the substrate feature matrix, enzyme sequence matrix and enzyme structure matrix of each group of enzyme-substrate pairs output by the enzyme turnover number prediction model to obtain an enzyme catalytic reaction matrix, using a fully connected layer to operate the enzyme catalytic reaction matrix, and finally obtaining the final turnover value by: y=Wx.
8. An enzyme turnover number prediction system based on protein structure information of deep learning for implementing the method described in any one of claims 1 to 7, characterized in that: include: Data collection unit, data preprocessing unit, model building unit, model training unit and result prediction unit, among which: the data collection unit obtains enzyme catalytic reaction data from the Brenda database and the SABIO-RK database; the data preprocessing unit uses ColabFold to predict protein structure, and obtains the initial features of protein structure based on the structural information, uses RDKit to obtain the initial features of substrate molecules, and continues to encode amino acids to obtain the initial features of protein sequences; The model building unit combines graph convolutional neural networks with different numbers of layers and Transformer to establish a deep learning model framework; the model training unit performs model training based on the initial features obtained from the data preprocessing part to obtain a trained deep learning model; the result prediction unit predicts the test set based on the enzyme and substrate molecule information of the test set and obtains the evaluation results.
Citation Information
Patent Citations
Deep learning-based rational design method for enzyme modification
CN115798581A
Target protein drug binding prediction method based on meta learning and sub-graph matching
CN116504303A
Enzyme kinetic parameter prediction model training and prediction method and related equipment
CN116580776A
Methods for enzyme engineering
WO2022254192A1