A rational design method for enzyme modification based on deep learning
By using a deep learning-based enzyme modification method and optimizing enzyme functional sites using protein crystal structure datasets and amino acid residue prediction models, the problem of difficulty in improving physicochemical properties in existing enzyme modification methods has been solved. This has enabled efficient enzyme modification and improved stability, promoting its application in industrial and medical fields.
Patent Information
- Application Number
- CN202211200712.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-29
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-09-29
AI Technical Summary
Existing methods for rational design and modification of enzymes cannot directly improve their physicochemical properties and require extensive manual screening.
This deep learning-based enzyme modification method utilizes protein crystal structure datasets and amino acid residue prediction models to optimize enzyme functional sites through deep learning models, and then verifies the feedback model through biochemical experiments.
Effectively modifying natural enzymes increases the types and stability of catalytic substrates, enhances the physicochemical properties of enzymes, and promotes their application in industrial and medical fields.
Smart Images

Figure CN115798581B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of biological enzyme engineering, and particularly relates to a rational design method for enzyme modification based on deep learning. BACKGROUND
[0002] The statements in this section merely provide background information related to the application and do not necessarily constitute prior art.
[0003] As a biological catalyst, enzymes are far superior to chemical catalysts in catalytic activity and environmental protection. However, natural enzymes have the shortcomings of limited catalytic substrate types and low stability. Designing and modifying natural enzymes has important value for improving their industrial and medical applications.
[0004] Existing enzyme rational design and modification methods are either based on sequencing sequences or based on protein simulation calculation for design and modification. The enzyme rational design and modification method based on sequencing sequences cannot directly improve the function of enzymes in terms of physical and chemical properties. The enzyme rational design and modification method based on protein simulation calculation requires a large amount of manual screening work in the early stage. SUMMARY
[0005] In order to solve the above problems, the application provides a rational design method for enzyme modification based on deep learning. The application uses protein crystal structure data as a basic data set, fully utilizes the physicochemical properties of proteins, and performs rational design and modification of enzymes based on artificial intelligence, thereby avoiding the inconvenience caused by manual screening.
[0006] According to some embodiments, the application provides a rational design method for enzyme modification based on deep learning, which adopts the following technical scheme:
[0007] A rational design method for enzyme modification based on deep learning, comprising:
[0008] Constructing a protein crystal structure data set;
[0009] Adding an atomic microenvironment centered on an amino acid residue based on the protein crystal structure data set to reconstruct the protein crystal structure data set;
[0010] According to the reconstructed protein crystal structure data set, an amino acid residue prediction model trained in advance is used to predict the amino acid residues of the reconstructed protein crystal structure data set;
[0011] According to the difference between the predicted amino acid residue probability and the actual amino acid probability of the natural enzyme, the potential site for optimizing the function of the enzyme is determined.
[0012] Further, the constructing a protein crystal structure data set comprises:
[0013] Clustering proteins in the existing protein crystal structure dataset according to sequence similarity;
[0014] Unifying the crystal structure representation based on the clustered proteins;
[0015] Screening the unified protein crystal structure, and removing low-resolution protein crystal structures;
[0016] Resampling the protein abundance and amino acid abundance in the screened protein crystal structure dataset to construct a protein structure dataset conforming to the natural distribution.
[0017] Further, the protein crystal structure dataset is reconstructed by adding an atomic microenvironment centered on an amino acid residue, including:
[0018] Constructing a three-dimensional grid space for each protein structure in the protein crystal structure dataset;
[0019] Randomly sampling protein structure sites in the three-dimensional grid space, and extracting the local space of an amino acid atom as atomic structure data information of the protein structure site;
[0020] Based on the randomly sampled protein structure sites, generating charge data and solvent accessibility data of the force field of each atom in the local structure;
[0021] Combining the physicochemical property data of the protein structure site with the atomic structure data information to form a data structure tensor, and obtaining the final protein crystal structure dataset.
[0022] Further, the protein structure site is centered on the nearest amino acid to collect structure in a three-dimensional local space with a size of 20x20x20.
[0023] Further, the amino acid residue prediction model includes a first convolutional layer, a second convolutional layer, a max-pooling layer, a third convolutional layer, a max-pooling layer, a fully connected network layer, and a Softmax layer connected in turn.
[0024] Further, the first convolutional layer, the second convolutional layer, and the third convolutional layer all adopt a 3D convolutional neural network.
[0025] Further, the max-pooling layer adopts a 2x2x2 spatial output.
[0026] Further, the fully connected network layer includes two layers of fully connected neural networks, both of which adopt a ReLU function as an activation function.
[0027] Further, the training of the amino acid residue prediction model uses the difference between the predicted amino acid residue probability value and the actual amino acid residue probability value as a loss function.
[0028] Further, the method further comprises:
[0029] Based on the determination of the potential site of enzyme function optimization, the potential modification site of the rational design of enzyme modification is modified.
[0030] Compared with the prior art, the beneficial effects of the present application are:
[0031] The present application takes protein crystal structure data as a basic data set, fully utilizes the physicochemical properties of the protein, and rationally designs and modifies the enzyme based on artificial intelligence, avoiding the inconvenience brought by manual screening.
[0032] The present application proposes an enzyme rational design and modification method and system based on deep learning, which effectively modifies natural enzymes, increases the catalytic substrate type, improves the activity and stability, and promotes the application of enzymes in the industrial and medical fields.
[0033] The enzyme rational design and modification method and system based on deep learning proposed by the present application further feed the biochemical experiment verification results into the training process of the prediction model, realizing the continuous optimization and upgrading of the model. BRIEF DESCRIPTION OF DRAWINGS
[0034] The drawings accompanying the specification of the present application form part of the present application and serve to provide further understanding of the present application, and the illustrative embodiments thereof and their descriptions serve to explain the present application, and do not constitute improper limitations on the present application.
[0035] Figure 1 The present application proposes an enzyme rational design and modification method and system based on deep learning, which effectively modifies natural enzymes, increases the catalytic substrate type, improves the activity and stability, and promotes the application of enzymes in the industrial and medical fields. DETAILED DESCRIPTION
[0036] The present application will be further described below in conjunction with the drawings and embodiments.
[0037] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as generally understood by those skilled in the art to which the present application belongs.
[0038] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0039] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0040] Example 1
[0041] This embodiment provides a rational design method for enzyme modification based on deep learning, which includes the following steps:
[0042] Construct a protein crystal structure dataset;
[0043] The protein crystal structure dataset was reconstructed by adding atomic microenvironments centered on amino acid residues to the protein crystal structure dataset.
[0044] Based on the reconstructed protein crystal structure dataset, amino acid residue prediction is performed on the reconstructed protein crystal structure dataset using a pre-trained amino acid residue prediction model.
[0045] Based on the degree of difference between the predicted amino acid residue probabilities and the actual amino acid probabilities of the natural enzyme, potential sites for enzyme function optimization are identified.
[0046] like Figure 1 As shown, the method described in this embodiment specifically includes:
[0047] Protein dataset reconstruction
[0048] To avoid oversampling of certain protein classes, proteins in existing protein crystal structure datasets are clustered based on sequence similarity using the CD-HIT software. For example, the PDB database.
[0049] The clustered proteins were represented using a unified crystal structure representation method from the PDB-redo database. Low-resolution crystal structure data were removed from the protein dataset. The protein abundance and amino acid abundance distributions of the remaining dataset were resampled to construct a protein structure dataset that conforms to the natural distribution.
[0050] Input data design and construction
[0051] A three-dimensional grid space is constructed for each protein structure of the above protein data set, and protein structure sites are randomly sampled from the space. For the randomly sampled points, a 20x20x20 three-dimensional local space is collected around the nearest amino acid. For the collected local space, the distribution of oxygen, carbon, nitrogen, sulfur, hydrogen, phosphorus, iron, zinc, copper, and magnesium atoms in the local space is extracted as the atomic structure data information of the local site, and the data structure is a tensor of (10, 20, 20, 20).
[0052] For each sampled local site, the CHARMM biological information software is used to calculate and generate the charge of the local site of each atom in the local structure to obtain data with a structure of (1, 20, 20, 20), and the FreeSASA software is used to calculate and generate the solvent accessibility of each atom in the local structure to obtain data with a structure of (1, 20, 20, 20), and the physical and chemical property data of the local site is combined with the above atomic structure to form a tensor with a data structure of (12, 20, 20, 20) as input data, which is used as input data for neural network model training.
[0053] Neural network model training
[0054] A prediction model is constructed based on a 3D convolutional neural network model. The network model first includes two layers of 3D convolutional neural networks. The first layer of 3D convolutional neural network uses a 4x4x4 filter with 100 channels and no padding, and uses a ReLU function as the activation function. The second layer of convolutional neural network uses a 4x4x4 filter with 200 channels and no padding, and uses a ReLU function as the activation function. The calculation of the output value of each filter is as follows:
[0055]
[0056] ReLU(x) = max(0, x) (2)
[0057] where L represents the Lth filter, F represents the size of the filter (here F = 3), C represents the input channel (C = 12), W is the weight, b is the bias, X is the input data, (i, j, k) is the position of the output data, (m, n, d) is the position of the input data, c represents the input channel number C, which takes a value from 0 to C; x is the independent variable of the ReLU function.
[0058] Next, the above output is connected to a pooling layer, which uses maximum pooling to output the maximum data in each 2x2x2 space of the input data, and the calculation formula is as follows:
[0059] MP c,l,m,n = max({X c,i,j,kX c,i+1,j,k X c,i,j+1,k X c,i,j,k+1 X c,i+1,j+1,k X c,i,j+1,k+1 X c,i+1,j,k+1 X c,i+1,j+1,k+1}) (3)
[0060] where i = l*2, j = m*2, k = n*2.
[0061] Next is a layer of 3D convolutional neural network and a layer of max pooling layer, the 3D convolutional neural network uses 2x2x2 filter and 400 channels and no padding, using ReLU function as activation function, the calculation method of max pooling layer is the same as the pooling layer of the last layer.
[0062] The above output results are used as the input of the full connection network layer, the full connection network layer includes two layers of full connection neural network, the first layer of full connection neural network has 10800 input neurons and 1000 output neurons, the activation function uses ReLU function, and the dropout is 0.8, the second layer of full connection neural network has 1000 input neurons and 100 output neurons, the activation function is ReLU function, and the dropout is 0.5. The calculation of the output value of the full connection network is as follows:
[0063] h n = ReLU(∑ m=0 M-1 W m,n X m +b n ) (4)
[0064] where h n represents the nth output value, M represents the number of input neurons, and N represents the number of output neurons.
[0065] Next is the Softmax layer, the Softmax layer is a full connection network, using Softmax activation function dropout is 0.2. The neural network outputs the predicted probability of each amino acid through the last Softmax activation function. The difference between the predicted probability value and the actual amino acid is used as the calculation of the loss function to perform supervised learning, and the calculation of the loss function is as follows:
[0066] loss = (o - a) 2 (5)
[0067] Wherein, o is the probability distribution vector of 20 kinds of output amino acids, the data structure is (1, 20), a is the actual amino acid vector (the value of the actual amino acid position is 1, and the others are 0), the training uses the Adam optimization algorithm, and finally the amino acid residue prediction model for generating the best matching microenvironment is trained and constructed.
[0068] Screening of potential modification sites
[0069] Taking the crystal structure of the natural enzyme as the input, the probability of the predicted amino acid residue with the highest probability is compared with the probability of the actual amino acid residue at the site, and the site with a larger ratio is screened as a potential modification site for optimizing the function of the enzyme.
[0070] Biochemical experiment verification
[0071] According to the site given above, the enzyme structure is modified, and the activity, stability, function and other physicochemical properties of the enzyme are biologically verified, and the results are fed back to the data set, and the neural network prediction model is further optimized and trained.
[0072] Although the specific embodiments of the present application are described above in combination with the drawings, it is not a limitation on the protection scope of the present application, and those skilled in the art should understand that various modifications or changes made by those skilled in the art on the basis of the technical solutions of the present application without creative labor are still within the protection scope of the present application.
Claims
1. A rational design method for enzyme modification based on deep learning, characterized in that, include: Construct a protein crystal structure dataset; The protein crystal structure dataset was reconstructed by adding atomic microenvironments centered on amino acid residues to the protein crystal structure dataset. The process of reconstructing the protein crystal structure dataset by adding atomic microenvironments centered on amino acid residues includes: A three-dimensional mesh space is constructed for each protein structure in the protein crystal structure dataset; In the three-dimensional grid space, protein structural sites are randomly sampled, and the local space of amino acid atoms is extracted as the atomic structure data information of the protein structural site. Based on randomly sampled protein structural sites, charge data and solvent accessibility data of the force field of each atom in the local structure are generated. The physicochemical properties of protein structural sites are combined with the atomic structure data to form a data structure tensor, resulting in the final protein crystal structure dataset. Based on the reconstructed protein crystal structure dataset, amino acid residue prediction is performed on the reconstructed protein crystal structure dataset using a pre-trained amino acid residue prediction model. The amino acid residue prediction model includes a first convolutional layer, a second convolutional layer, a maximum pooling layer, a third convolutional layer, a maximum pooling layer, a fully connected network layer, and a Softmax layer connected in sequence. Based on the degree of difference between the predicted amino acid residue probabilities and the actual amino acid probabilities of the natural enzyme, potential sites for enzyme function optimization are identified.
2. The rational design method for enzyme modification based on deep learning as described in claim 1, characterized in that, The construction of the protein crystal structure dataset includes: The proteins in the existing protein crystal structure dataset are clustered based on sequence similarity. Based on the clustered proteins, a unified crystal structure representation is achieved; Screening is performed on uniform protein crystal structures to remove those with low resolution. The protein abundance and amino acid abundance distributions in the screened protein crystal structure dataset were resampled to construct a protein structure dataset that conforms to the natural distribution.
3. The rational design method for enzyme modification based on deep learning as described in claim 1, characterized in that, The protein structural sites are acquired by sampling the structure in a three-dimensional local space of 20×20×20, centered on the nearest amino acid.
4. The rational design method for enzyme modification based on deep learning as described in claim 1, characterized in that, The first convolutional layer, the second convolutional layer, and the third convolutional layer all employ a 3D convolutional neural network.
5. The rational design method for enzyme modification based on deep learning as described in claim 1, characterized in that, The maximum pooling layer uses a 2x2x2 spatial output.
6. The rational design method for enzyme modification based on deep learning as described in claim 1, characterized in that, The fully connected network layer consists of two fully connected neural network layers, both of which use the ReLU function as the activation function.
7. The rational design method for enzyme modification based on deep learning as described in claim 1, characterized in that, The training of the amino acid residue prediction model uses the difference between the predicted amino acid residue probability value and the actual amino acid residue probability value as the loss function.
8. The rational design method for enzyme modification based on deep learning as described in claim 1, characterized in that, Also includes: Potential modification sites are rationally designed based on identifying potential sites for enzyme function optimization.
Citation Information
Patent Citations
Method for calculating and identifying protein kinase phosphorylation specific sites
CN101710365A
Enzymes for the treatment of lignocellulosics, nucleic acids encoding them and methods for making and using them
CN104212822A