Transaminase screening method and device, electronic equipment and storage medium
By using protein-substrate binding activity prediction models and molecular dynamics simulations, the problems of long screening time and high cost of transaminases have been solved, achieving efficient transaminase screening and improving screening efficiency and accuracy.
Patent Information
- Application Number
- CN202511084317.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-18
AI Technical Summary
Existing transaminase screening methods are time-consuming, costly, and inefficient. In particular, molecular dynamics simulations cannot support high-throughput screening, making it difficult to evaluate a large number of transaminase-substrate pairs in a short period of time.
A protein-substrate binding activity prediction model is used for rapid screening, a protein structure prediction model is used to construct three-dimensional structures, and molecular dynamics simulation is used for final screening. Deep learning and physical modeling are used to improve screening efficiency.
By rapidly predicting the binding activity of transaminases and substrates, constructing the three-dimensional structure of the complex, and performing molecular dynamics simulations, the efficiency and accuracy of transaminase screening have been improved.
Smart Images

Figure CN120977440A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, and in particular to a transaminase screening method and device, an electronic device and a storage medium. BACKGROUND
[0002] With the rapid development of biotechnology, transaminase as an important biological catalyst has a wide application prospect in the fields of medicine, chemical industry, food and the like. Transaminase can catalyze the transfer of amino group from a donor molecule to an acceptor molecule to realize the amination reaction, which is particularly important in the synthesis of chiral amine compounds. However, natural transaminase often has problems such as narrow substrate specificity and low catalytic efficiency, which is difficult to meet the needs of industrial applications, so screening of transaminase with high activity and high specificity has become a research hotspot.
[0003] At present, transaminase screening mainly relies on experimental methods such as directed evolution and rational design, which are usually time-consuming, costly and inefficient. With the development of computer technology, computer-aided screening methods have gradually attracted attention. For example, CN118197417A discloses an enzyme virtual screening method, which constructs a candidate enzyme and a target small molecule structure library, uses a preset molecular docking tool for processing, and performs molecular dynamics simulation on the molecular docking results to screen positive results.
[0004] However, the existing transaminase screening method still has some technical problems: although molecular dynamics (MD) simulation can provide dynamic information of protein-substrate interaction, it has large calculation amount and long single simulation time, which cannot support high-throughput screening, and it is difficult to complete the evaluation of a large number of transaminase-substrate pairs in a short time, which is not conducive to improving the efficiency of transaminase screening. SUMMARY
[0005] In view of the above problems, the embodiments of the present application provide a transaminase screening method, device, electronic device and storage medium to solve the technical problem that the efficiency of transaminase screening is not improved.
[0006] In a first aspect, the embodiments of the present application provide a transaminase screening method, comprising: obtaining first data to be predicted, the first data comprising an amino acid sequence of a transaminase and a molecular structure of a substrate; inputting the first data into a pre-trained protein-substrate binding activity prediction model to output a first prediction result of the transaminase and the substrate in the first data; wherein the protein-substrate binding activity prediction model is obtained by training a preset enzyme-substrate pair data; performing first screening on the first data according to the first prediction result to obtain first screened first data; inputting the first data after the first screening into a pre-trained protein structure prediction model, and outputting a three-dimensional structure of a complex formed by the transaminase and the substrate; obtaining binding data of the transaminase and the substrate according to the three-dimensional structure of the complex, and performing second screening on the first data according to the binding data to obtain first data after the second screening; performing molecular dynamics simulation on the transaminase-substrate complex corresponding to the first data after the second screening to perform third screening on the first data.
[0007] Optionally, the molecular dynamics simulation on the transaminase-substrate complex corresponding to the first data after the second screening to perform third screening on the first data comprises: performing molecular dynamics simulation on the transaminase-substrate complex corresponding to the first data after the second screening to obtain state change data of the transaminase-substrate complex during the molecular dynamics simulation; performing third screening on the first data according to the state change data to obtain first data after the third screening.
[0008] Optionally, the training step of the protein-substrate binding activity prediction model comprises: obtaining training samples, wherein the training samples comprise preset enzyme-substrate pair data and corresponding real label categories; extracting features of amino acid sequences of enzymes using a protein semantic model to obtain first extracted features; extracting features of substrates using a graph neural network to obtain second extracted features; inputting the first extracted features and the second extracted features into a gradient boosting decision tree model to output predicted label categories; adjusting parameters of the gradient boosting decision tree model according to the predicted label categories and the corresponding real label categories to train the gradient boosting decision tree model; obtaining a trained protein-substrate binding activity prediction model according to the protein semantic model, the graph neural network, and the trained gradient boosting decision tree model.
[0009] Optionally, the obtaining of the training samples comprises: obtaining enzyme-substrate data pairs verified by experiments from a protein database as first positive sample data; obtaining enzyme-substrate data pairs corresponding to positive screening results obtained by molecular dynamics simulation screening as second positive sample data; obtaining enzyme-substrate data pairs removed by molecular dynamics simulation screening as negative sample data. construct the training sample according to the first positive sample data, the second positive sample data and the negative sample data.
[0010] Optionally, the obtaining the binding data of the transaminase and the substrate according to the three-dimensional structure of the complex, and screening the first data according to the binding data, comprises: calculating the binding energy of the transaminase and the substrate according to the three-dimensional structure of the complex; calculating the reaction site distance between the catalytic residue of the transaminase and the reaction site of the substrate according to the three-dimensional structure of the complex; obtaining the first data with the binding energy less than or equal to a preset binding energy threshold and the reaction site distance less than or equal to a preset distance threshold.
[0011] Optionally, the molecular dynamics simulation is performed on the transaminase-substrate complex corresponding to the first data after the second screening to obtain state change data of the transaminase-substrate complex in the molecular dynamics simulation process, comprising: performing three-dimensional structure establishment and energy minimization on the transaminase-substrate complex to obtain an initial structure of the transaminase-substrate complex; performing equilibrium simulation on the initial structure of the transaminase-substrate complex to obtain the transaminase-substrate complex in an equilibrium state; performing production stage simulation on the transaminase-substrate complex in the equilibrium state to perform structure stability and energy change analysis, and taking the results of the structure stability and energy change analysis as the state change data in the molecular dynamics simulation process.
[0012] Optionally, the third screening is performed on the first data according to the state change data to obtain the first data after the third screening, comprising: obtaining at least one target index from the state change data, the target index being binding energy, number of hydrogen bonds, root mean square deviation or root mean square fluctuation; performing the third screening on the first data according to the at least one target index to obtain the first data after the third screening.
[0013] In a second aspect, an embodiment of the present application provides a transaminase screening device, comprising: a sequence construction module configured to obtain first data to be predicted, the first data comprising an amino acid sequence of a transaminase and a molecular structure of a substrate; The first prediction module is configured to input the first data into a pre-trained protein-substrate binding activity prediction model, and output a first prediction result corresponding to the transaminase and the substrate in the first data; wherein the protein-substrate binding activity prediction model is obtained by training a preset enzyme-substrate pair data; The first screening module is configured to perform a first screening on the first data according to the first prediction result, and obtain first screened first data. The second prediction module is configured to input the first screened first data into a pre-trained protein structure prediction model, and output a three-dimensional structure of a complex formed by the transaminase and the substrate; The second screening module is configured to obtain binding data of the transaminase and the substrate according to the three-dimensional structure of the complex, and perform a screening on the first data according to the binding data, and obtain second screened first data. The third prediction screening module is configured to perform a third screening on the first data by performing a molecular dynamics simulation on a transaminase-substrate complex corresponding to the second screened first data.
[0014] In a third aspect, an electronic device is provided, including a processor and a memory coupled to the processor, the memory storing program instructions executable by the processor; and the processor executes the program instructions stored in the memory to implement the transaminase screening method described above.
[0015] In a fourth aspect, a computer readable storage medium is provided, the computer readable storage medium storing program instructions, the program instructions being executable by a processor to implement the transaminase screening method described above.
[0016] The transaminase screening method, device, electronic device and storage medium provided by the embodiments of the present application first learn the interaction between the enzyme and the substrate based on the protein-substrate binding activity prediction model, quickly predict the binding activity of the transaminase and the substrate, and realize the first screening; then, the three-dimensional structure of the complex formed by the combination of the first screened transaminase and the substrate is constructed based on the protein structure prediction model, and the second screening is realized; finally, the molecular dynamics simulation is performed on the second screened complex; in this way, the accuracy of the transaminase-substrate interaction prediction is improved by combining deep learning with physical modeling, which is conducive to improving the efficiency of transaminase screening.
[0017] These and other aspects of the present application will become more apparent in the following description of the embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0018] Figure 1A flowchart of the transaminase screening method provided by the embodiments of the present application is shown.
[0019] Figure 2 A performance evaluation diagram of the protein-substrate binding activity prediction model in the transaminase screening method provided by the embodiments of the present application is shown.
[0020] Figure 3 A complex screening result diagram of the protein structure prediction model establishment in the transaminase screening method provided by the embodiments of the present application is shown.
[0021] Figure 4 A molecular dynamics simulation result diagram in the transaminase screening method provided by the embodiments of the present application is shown.
[0022] Figure 5 A structural diagram of the transaminase screening device provided by the embodiments of the present application is shown.
[0023] Figure 6 A structural diagram of the electronic device provided by the embodiments of the present application is shown.
[0024] Figure 7 A structural diagram of the computer storage medium provided by the embodiments of the present application is shown. DETAILED DESCRIPTION
[0025] The embodiments of the present application will be described in detail below, examples of the embodiments are shown in the drawings, wherein the same or similar reference signs represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below are exemplary and are only used to explain the present application, and cannot be understood as a limitation of the present application.
[0026] In order to enable those skilled in the art to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0027] In the embodiments of the present application, it should be noted that in this document, relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations.
[0028] Also, the term "comprise", "comprising", or any other variant thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0029] In the description of the embodiments of the present application, the words "example" or "for example" are used to mean "an example of" or "for example". Any embodiment or design solution described as "example" or "for example" in the embodiments of the present application is not necessarily to be understood as preferable or having more advantages than another embodiment or design solution.
[0030] In addition, "multiple" in the embodiments of the present application refers to two or more, and therefore "multiple" in the embodiments of the present application can also be understood as "at least two". "At least one" can be understood as one or more, for example, one, two or more. For example, including at least one means including one, two or more, and does not limit which ones are included, for example, including at least one of A, B and C means that A, B, C, A and B, A and C, B and C, or A and B and C can be included.
[0031] It should be noted that in the embodiments of the present application, the "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B, which means that there are three cases of A alone, A and B together, and B alone. In addition, the character " / ", if not specially stated, generally represents a "or" relationship between the associated objects before and after.
[0032] It should be noted that in the embodiments of the present application, "connection" can be understood as electrical connection, and the connection between two electrical elements can be direct or indirect connection between the two electrical elements. For example, A is connected to B, which can be direct connection between A and B, or indirect connection between A and B through one or more other electrical elements.
[0033] An embodiment of the present application provides a transaminase screening method, please refer to Figure 1 The transaminase screening method includes the following steps S11-S16: Step S11: obtaining first data to be predicted, wherein the first data includes the amino acid sequence of the transaminase and the molecular structure of the substrate.
[0034] Wherein, for each first data, the interaction between transaminase and substrate is predicted. The amino acid sequence of transaminase and the molecular structure of substrate can be derived from public databases or self-built data sets in the laboratory. Exemplarily, the amino acid sequence of the enzyme can be provided in FASTA format; the molecular structure of the substrate can be provided in the form of a SMILES string or a molecular graph.
[0035] Step S12: input the first data into the pre-trained protein-substrate binding activity prediction model, and output the first prediction result of the corresponding transaminase and substrate in the first data; wherein the protein-substrate binding activity prediction model is obtained by training the preset enzyme-substrate pair data.
[0036] Wherein, the protein-substrate binding activity prediction model is a deep learning model for predicting transaminase-substrate binding activity, which can better capture the specific interaction between enzyme and substrate, and learn the key catalytic characteristics of enzyme and substrate interaction. Due to the relatively small amount of experimentally verified transaminase-substrate data pairs, the preset enzyme-substrate pair data is used in the training process of the protein-substrate binding activity prediction model, which can include transaminase-substrate data pairs and other enzyme-substrate pair data.
[0037] Wherein, the first prediction result can include the binding probability of transaminase and substrate and the unbinding probability of transaminase and substrate; or the first prediction result can be the binding affinity score between transaminase and substrate.
[0038] Step S13: first screening of the first data according to the first prediction result, to obtain the first screened first data.
[0039] Wherein, the first screening is usually based on the threshold of the predicted binding probability or the threshold of the binding affinity; the first data with a binding probability greater than a preset binding probability threshold is retained, or the first data with a binding affinity score greater than a preset binding affinity threshold is retained as the first screened first data.
[0040] Please refer to Figure 2 As shown, after training, the protein-substrate binding activity prediction model has better performance, wherein, Figure 2 As shown on the left is the ROC curve, Figure 2 As shown on the right is the confusion matrix.
[0041] Step S14: input the first screened first data into the pre-trained protein structure prediction model, and output the three-dimensional structure of the complex formed by the combination of transaminase and substrate.
[0042] The protein structure prediction model is used to predict and generate a complex three-dimensional structure of the transaminase-substrate. Figure 3 As shown, the modeling screening obtains a complex formed by the combination of the transaminase and the substrate.
[0043] Step S15: Obtain the binding data of the transaminase and the substrate according to the three-dimensional structure of the complex, and perform a second screening on the first data according to the binding data to obtain the first data after the second screening.
[0044] The binding data may, for example, include at least one of a binding energy or a reaction site distance, wherein the lower the binding energy, the more stable the interaction between the transaminase and the substrate; and a reasonable reaction site distance indicates that the substrate can effectively bind to the active site of the transaminase.
[0045] The calculation of the binding data can be performed according to the three-dimensional structure output by the protein structure prediction model. For example, the three-dimensional structure of the transaminase and the three-dimensional structure of the substrate can be extracted from the three-dimensional structure of the complex, and a molecular docking tool such as AutoDock, Glide, or other molecular docking software can be used to simulate the binding process of the three-dimensional structure of the transaminase and the three-dimensional structure of the substrate, and the binding energy can be calculated according to the binding process. For example, based on the three-dimensional structure of the complex, the distance between the active site of the transaminase and the binding site of the substrate in the three-dimensional structure of the complex can be directly measured.
[0046] Step S16: Perform molecular dynamics simulation on the transaminase-substrate complex corresponding to the first data after the second screening to perform a third screening on the first data.
[0047] After the first screening and the second screening, the first data remaining after the screening is subjected to molecular dynamics simulation, thereby reducing the computational load of the molecular dynamics simulation. After obtaining the simulation results of the transaminase-substrate pair in the first data after the second screening, a third screening is performed according to the simulation results.
[0048] In this embodiment, first, the interaction between the enzyme and the substrate is learned based on the protein-substrate binding activity prediction model to quickly predict the binding activity of the transaminase and the substrate, thereby achieving the first screening; then, the three-dimensional structure of the complex formed by the combination of the transaminase and the substrate after the first screening is constructed based on the protein structure prediction model, thereby achieving the second screening; finally, the complex after the second screening is subjected to molecular dynamics simulation; in this way, the combination of deep learning and physical modeling improves the accuracy of the transaminase-substrate interaction prediction and is conducive to improving the efficiency of the transaminase screening.
[0049] As an implementation manner, the protein-substrate binding activity prediction model comprises a protein semantic model, a graph neural network, and a Gradient Boosting Decision Tree (GBDT) model, and step S12 specifically comprises the following steps: Step S21: extracting features of the amino acid sequence of the transaminase in the first data by using the protein semantic model to obtain first extracted features; The protein semantic model can be a deep learning model based on a Transformer architecture, which can capture long-range dependencies and evolutionary information in the amino acid sequence and generate a high-dimensional protein feature vector. For example, the protein semantic model can be an ESM (Evolutionary Scale Modeling) model, such as an ESM-1b Transformer model.
[0050] Step S22: extracting features of the substrate in the first data by using the graph neural network to obtain second extracted features; The graph neural network (GNN) extracts 2D molecular graph features of the substrate. The GNN can process the molecular graph structure of small molecules, capture the interaction between atoms and bonds, and generate a high-dimensional feature vector of the small molecule.
[0051] Step S23: inputting the first extracted features and the second extracted features into the Gradient Boosting Decision Tree model to output a first prediction result; The Gradient Boosting Decision Tree model can first concatenate the first extracted features and the second extracted features to obtain fused features, and then predict the binding activity between the transaminase and the substrate according to the fused features by learning the features of the transaminase and the substrate.
[0052] For example, the protein-substrate binding activity prediction model can be an ESP (Enzyme Substrate Prediction) deep learning model. In some embodiments, the Gradient Boosting Decision Tree model can comprise a bidirectional conditional feature mechanism network, a concatenation network, and a Gradient Boosting Decision Tree prediction network, wherein the bidirectional conditional feature mechanism network can be used to exchange information between the first extracted features and the second extracted features, so that the features of the transaminase and the substrate influence each other, thereby better capturing the specific interaction between the transaminase and the substrate.
[0053] Specifically, the specific steps of feature fusion performed by the bidirectional conditional feature fusion network are as follows: S231: calculating a first conditional probability of the second extracted features under the condition of the first extracted features; S232: obtaining a first fusion feature according to the first weight matrix, the first conditional probability, and the first extracted feature; S233: calculating a second conditional probability of the first extracted feature under the condition of the second extracted feature; S234: obtaining a second fusion feature according to the second weight matrix, the second conditional probability, and the second extracted feature.
[0054] In the training process of the gradient boosting decision tree model, the first weight matrix and the second weight matrix can be trained.
[0055] Specifically, wherein, is the first fusion feature, is the first extracted feature, is the second extracted feature, is the first weight matrix, is the first conditional probability.
[0056] Specifically, wherein, is the second fusion feature, is the first extracted feature, is the second extracted feature, is the second weight matrix, is the second conditional probability.
[0057] Subsequently, the concatenation network concatenates the first fusion feature and the second fusion feature to obtain a fusion feature, and the gradient boosting decision tree prediction network outputs a first prediction result according to the fusion feature.
[0058] In some embodiments, the training step of the protein-substrate binding activity prediction model comprises: Step S31: obtaining a training sample, wherein the training sample comprises preset enzyme-substrate pair data and corresponding real label categories; Step S32: extracting features of the amino acid sequence of the enzyme using the protein semantic model to obtain first extracted features; Step S33: extracting features of the substrate using the graph neural network to obtain second extracted features; Step S34: inputting the fusion features of the first extracted features and the second extracted features into the gradient boosting decision tree model to output predicted label categories; Step S35: adjusting parameters of the gradient boosting decision tree model according to the predicted label categories and the corresponding real label categories to train the gradient boosting decision tree model; Step S36: obtaining a trained protein-substrate binding activity prediction model according to the protein semantic model, the graph neural network, and the trained gradient boosting decision tree model.
[0059] In some embodiments, when the gradient boosting decision tree model comprises the bidirectional conditional feature mechanism network, the concatenation network and the gradient boosting decision tree prediction network described above, during the training process, the first weight matrix and the second weight matrix in the bidirectional conditional feature mechanism network and the parameters in the gradient boosting decision tree prediction network can be trained.
[0060] In some embodiments, the step S31 specifically comprises the following steps: Step S41: obtaining the enzyme-substrate data pairs verified by experiments from a protein database as first positive sample data; Wherein, since the amount of transaminase-substrate data verified by experiments is relatively small, other enzyme-substrate data pairs verified by experiments are also used as training samples. The first positive sample data further comprises a real label category. For example, for the samples verified by experiments, the real label category can be a binding probability of 1.
[0061] For example, the protein database can include the BRENDA database (https: / / www.brenda-enzymes.org); in addition, the protein database can further include an enzyme-substrate pair database constructed according to literature data. The enzyme-substrate database can be constructed in the following manner: obtaining enzyme-substrate data pairs verified by experiments in literature data, and constructing an enzyme-substrate database according to the amino acid sequence of the enzyme and the molecular structure of the substrate in the enzyme-substrate data pairs. For example, the literature data can include paper literature and patent literature.
[0062] Step S42: obtaining the enzyme-substrate data pairs corresponding to the positive screening results screened by the molecular dynamics simulation as second positive sample data; Wherein, the enzyme-substrate data pairs with high binding activity determined by the molecular dynamics simulation can also be used as positive sample data.
[0063] Step S43: obtaining the enzyme-substrate data pairs removed by the molecular dynamics simulation screening as negative sample data; Wherein, the enzyme-substrate data pairs with low binding activity determined by the molecular dynamics simulation can be used as negative sample data.
[0064] Step S44: constructing training samples according to the first positive sample data, the second positive sample data and the negative sample data.
[0065] In the embodiment, the positive sample data is data verified by experiments or data with high binding activity determined by molecular dynamics simulation; the negative sample data is data with low binding activity determined by molecular dynamics simulation; the training data has high quality, which is beneficial to improve the prediction accuracy of the protein-substrate binding activity prediction model.
[0066] As an implementation, step S15 specifically includes the following steps: Step S51: calculating the binding energy of the transaminase and the substrate according to the three-dimensional structure of the complex; Wherein, the three-dimensional structure of the transaminase and the three-dimensional structure of the substrate can be extracted from the three-dimensional structure of the complex, the binding process of the three-dimensional structure of the transaminase and the three-dimensional structure of the substrate is simulated by using a molecular docking tool, and the binding energy is calculated according to the binding process. The molecular docking tool can be, for example, AutoDock, Glide, or other molecular docking software.
[0067] Step S52: calculating the reaction site distance between the catalytic residue of the transaminase and the reaction site of the substrate according to the three-dimensional structure of the complex; Wherein, the distance between the active site of the transaminase and the binding site of the substrate in the three-dimensional structure of the complex can be directly measured based on the three-dimensional structure of the complex. For example, the distance between the active site of the transaminase and the binding site of the substrate in the three-dimensional structure of the complex can be measured by using a molecular visualization software, such as PyMOL, VMD, or other software.
[0068] Step S53: obtaining first data with binding energy less than or equal to a preset binding energy threshold and reaction site distance less than or equal to a preset distance threshold.
[0069] Wherein, the second screening is performed according to the preset binding energy threshold and the preset distance threshold.
[0070] As an implementation, step S16 specifically includes the following steps: Step S61: performing molecular dynamics simulation on the transaminase-substrate complex corresponding to the first data after the second screening to obtain state change data of the transaminase-substrate complex during the molecular dynamics simulation; Wherein, the tool for molecular dynamics simulation is, for example, GROMACS (GROningen MAchine for Chemical Simulations, GROMACS). GROMACS supports various simulation algorithms, such as molecular dynamics, energy minimization, conformational search, and free energy calculation, and also provides functions such as molecular construction, force field parameterization, simulation setting, post-processing, and analysis.
[0071] Step S62: performing third screening on the first data according to the state change data, to obtain third screened first data.
[0072] The state change data may include, for example, changes in atomic coordinates over time, structural changes of the complex (secondary structure changes, conformation changes, etc.), kinetic parameters (atomic velocity, kinetic energy, potential energy, temperature, etc.), viscosity, and diffusion coefficient, etc.
[0073] In some embodiments, step S61 specifically includes the following steps: Step S71: performing three-dimensional structure establishment and energy minimization on the transaminase-substrate complex, to obtain an initial structure of the transaminase-substrate complex; The initial structure of the transaminase-substrate complex is obtained based on the binding site of the catalytic residue of the transaminase and the substrate, and energy minimization is performed on the initial structure of the transaminase-substrate complex to eliminate unreasonable conformations in the initial structure.
[0074] Step S72: performing equilibrium simulation on the initial structure of the transaminase-substrate complex, to obtain the transaminase-substrate complex in the equilibrium state; The molecular dynamics simulation parameters, including temperature, pressure, solvent environment, etc., are set to ensure that the experimental conditions match the physiological environment, and the initial structure of the transaminase-substrate complex is simulated to gradually adjust the parameters to achieve a thermodynamic equilibrium state.
[0075] After energy minimization and equilibrium simulation, the system reaches a stable state, and the energy deviation of the initial structure of the transaminase-substrate complex is reduced.
[0076] Step S73: performing production stage simulation on the transaminase-substrate complex in the equilibrium state to analyze the structural stability and energy change, and taking the results of the structural stability and energy change analysis as the state change data in the molecular dynamics simulation process.
[0077] In the production simulation stage, the transaminase-substrate interaction trajectory and conformation change are recorded, the structural stability and energy change are analyzed according to the above simulation trajectory and conformation change, and the state change data of the production stage simulation is obtained. The state change data may include a plurality of target indicators, which may include but are not limited to binding energy, number of hydrogen bonds, root mean square deviation, or root mean square fluctuation. For example, the production stage simulation time is less than or equal to 100 nanoseconds, for example, the production stage simulation process can last for 10 nanoseconds to 100 nanoseconds.
[0078] In some embodiments, step S62 specifically includes the following steps: Step S81: obtaining at least one target index from the state change data, the target index being binding energy, hydrogen bond number, root mean square deviation, or root mean square fluctuation; Step S82: performing third screening on the first data according to the at least one target index, to obtain third-screened first data.
[0079] Referring to Figure 4 the molecular dynamics simulation results shown in the figure, wherein, Figure 4 the upper left is a graph of RMSD (root mean square deviation) changing with time, Figure 4 the upper right is a graph of RMSF (root mean square fluctuation) changing with time, Figure 4 the lower left is a graph of hydrogen bond number distribution changing with time, Figure 4 and the lower right is a graph of reaction site distance changing with time.
[0080] The charging control circuit 100 provided in the application can be applied to Figure 1 a notebook computer 200 as shown, which comprises two hardware interfaces, and the charging control circuit 100 is connected to the two hardware interfaces respectively, and is used for controlling the charging state of the two hardware interfaces.
[0081] An embodiment of the application provides a transaminase screening device, referring to Figure 5 The transaminase screening device 200 comprises a sequence construction module 21, a first prediction module 22, a first screening module 23, a second prediction module 24, a second screening module 25, and a third prediction screening module 26. The sequence construction module 21 is used for obtaining first data to be predicted, and the first data comprises an amino acid sequence of a transaminase and a molecular structure of a substrate. The first prediction module 22 is used for inputting the first data into a pre-trained protein-substrate binding activity prediction model, and outputting a first prediction result corresponding to the transaminase and the substrate in the first data. The protein-substrate binding activity prediction model is obtained by training preset enzyme-substrate pair data. The first screening module 23 is used for performing first screening on the first data according to the first prediction result, to obtain first-screened first data. The second prediction module 24 is used for inputting the first-screened first data into a pre-trained protein structure prediction model, and outputting a three-dimensional structure of a complex formed by the transaminase and the substrate. The second screening module 25 is used for obtaining binding data of the transaminase and the substrate according to the three-dimensional structure of the complex, and performing screening on the first data according to the binding data, to obtain second-screened first data. The third prediction screening module 26 is used for performing molecular dynamics simulation on a transaminase-substrate complex corresponding to the second-screened first data, to perform third screening on the first data.
[0082] In some embodiments, the third prediction screening module 26 is further configured to perform a molecular dynamics simulation on the transaminase-substrate complex corresponding to the second screened first data, to obtain state change data of the transaminase-substrate complex during the molecular dynamics simulation; and perform a third screening on the first data according to the state change data, to obtain third screened first data.
[0083] In some embodiments, the first prediction module 22 is further configured to obtain training samples, wherein the training samples include preset enzyme-substrate pair data and corresponding real label categories; extract features of an amino acid sequence of an enzyme using a protein semantic model, to obtain first extracted features; extract features of a substrate using a graph neural network, to obtain second extracted features; input the first extracted features and the second extracted features into a gradient boosting decision tree model, to output predicted label categories; adjust parameters of the gradient boosting decision tree model according to the predicted label categories and the corresponding real label categories, to train the gradient boosting decision tree model; and obtain a trained protein-substrate binding activity prediction model according to the protein semantic model, the graph neural network, and the trained gradient boosting decision tree model.
[0084] In some embodiments, the first prediction module 22 is further configured to obtain enzyme-substrate data pairs that have been verified by experiments from a protein database, as first positive sample data; obtain enzyme-substrate data pairs corresponding to positive screening results obtained by using the molecular dynamics simulation, as second positive sample data; obtain enzyme-substrate data pairs removed by using the molecular dynamics simulation, as negative sample data; and construct training samples according to the first positive sample data, the second positive sample data, and the negative sample data.
[0085] In some embodiments, the second screening module 25 is further configured to calculate binding energy of the transaminase and the substrate according to the three-dimensional structure of the complex; calculate a reaction site distance between a catalytic residue of the transaminase and a reaction site of the substrate according to the three-dimensional structure of the complex; and obtain first data with binding energy less than or equal to a preset binding energy threshold and reaction site distance less than or equal to a preset distance threshold.
[0086] In some embodiments, the third prediction screening module 26 is further configured to perform three-dimensional structure establishment and energy minimization on the transaminase-substrate complex, to obtain an initial structure of the transaminase-substrate complex; perform equilibrium simulation on the initial structure of the transaminase-substrate complex, to obtain the transaminase-substrate complex in an equilibrium state; perform production phase simulation on the transaminase-substrate complex in the equilibrium state, to perform structure stability and energy change analysis, and use results of the structure stability and energy change analysis as state change data during the molecular dynamics simulation.
[0087] In some embodiments, the third prediction screening module 26 is further configured to obtain at least one target index from the state change data, the target index being binding energy, number of hydrogen bonds, root mean square deviation, or root mean square fluctuation; and perform third screening on the first data according to the at least one target index to obtain third screened first data.
[0088] Figure 6 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present application. As shown in the figure, the electronic device 30 includes a processor 31 and a memory 32 coupled to the processor 31. Figure 6
[0089] The memory 32 stores program instructions for implementing the transaminase screening method of any of the above embodiments.
[0090] The processor 31 is configured to execute the program instructions stored in the memory 32 to perform transaminase screening.
[0091] The processor 31 can also be referred to as a CPU (Central Processing Unit). The processor 31 can be an integrated circuit chip having signal processing capability. The processor 31 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application-Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0092] Referring to FIG. 2, Figure 7 Figure 7 FIG. 3 is a structural schematic diagram of a computer-readable storage medium according to an embodiment of the present application. The storage medium according to the embodiment of the present application stores program instructions 41 capable of implementing all the above methods. The storage medium can be non-volatile or volatile. The program instructions 41 can be stored in the above storage medium in the form of a software product, including a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the various embodiments of the present application. The above storage medium includes a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk or an optical disk, etc. various media capable of storing program codes, or a computer, a server, a mobile phone, a tablet, etc. terminal device.
[0093] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the division of the units is only a logical function division, and there can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0094] In addition, each function unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist alone physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software function unit. The above is only an embodiment of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
[0095] The above is only an embodiment of the present application, and it should be noted that for those skilled in the art, without departing from the creative concept of the present application, improvements can be made, but these are all within the protection scope of the present application.
Claims
1. A method for screening transaminases, characterized in that, The method comprises the following steps: obtaining first data to be predicted, wherein the first data comprises an amino acid sequence of a transaminase and a molecular structure of a substrate; inputting the first data into a pre-trained protein-substrate binding activity prediction model to output a first prediction result corresponding to the transaminase and the substrate in the first data; wherein the protein-substrate binding activity prediction model is obtained by training preset enzyme-substrate pair data; performing first screening on the first data according to the first prediction result to obtain first screened first data; inputting the first screened first data into a pre-trained protein structure prediction model to output a three-dimensional structure of a complex formed by the transaminase and the substrate; obtaining binding data of the transaminase and the substrate according to the three-dimensional structure of the complex, and performing second screening on the first data according to the binding data to obtain second screened first data; performing molecular dynamics simulation on a transaminase-substrate complex corresponding to the second screened first data to perform third screening on the first data.
2. The transaminase screening method according to claim 1, characterized by, The method comprises the following steps: performing molecular dynamics simulation on a transaminase-substrate complex corresponding to the second screened first data to obtain state change data of the transaminase-substrate complex during the molecular dynamics simulation; performing third screening on the first data according to the state change data to obtain third screened first data.
3. The transaminase screening method according to claim 1, wherein, The training step of the protein-substrate binding activity prediction model comprises the following steps: obtaining training samples, wherein the training samples comprise preset enzyme-substrate pair data and corresponding real label categories; extracting features of the amino acid sequence of the enzyme by using a protein semantic model to obtain first extracted features; extracting features of the substrate by using a graph neural network to obtain second extracted features; inputting the first extracted features and the second extracted features into a gradient boosting decision tree model to output predicted label categories; adjusting parameters of the gradient boosting decision tree model according to the predicted label categories and the corresponding real label categories to train the gradient boosting decision tree model; obtaining a trained protein-substrate binding activity prediction model according to the protein semantic model, the graph neural network, and the trained gradient boosting decision tree model.
4. The transaminase screening method according to claim 3, wherein, The method comprises the following steps: obtaining enzyme-substrate data pairs verified by experiments from a protein database as first positive sample data; obtaining enzyme-substrate data pairs corresponding to positive screening results obtained by molecular dynamics simulation screening as second positive sample data; obtaining enzyme-substrate data pairs removed by molecular dynamics simulation screening as negative sample data; constructing the training samples according to the first positive sample data, the second positive sample data, and the negative sample data.
5. The transaminase screening method according to claim 1, wherein, The method comprises the following steps: calculate a binding energy of the transaminase and the substrate according to a three-dimensional structure of the complex; calculate a reaction site distance between a catalytic residue of the transaminase and a reaction site of the substrate according to the three-dimensional structure of the complex; obtain first data of the binding energy being less than or equal to a preset binding energy threshold value and the reaction site distance being less than or equal to a preset distance threshold value.
6. The transaminase screening method according to claim 2, wherein, perform molecular dynamics simulation on the transaminase-substrate complex corresponding to the first data after the second screening to obtain state change data of the transaminase-substrate complex in the molecular dynamics simulation process, including: perform three-dimensional structure establishment and energy minimization on the transaminase-substrate complex to obtain an initial structure of the transaminase-substrate complex; perform equilibrium simulation on the initial structure of the transaminase-substrate complex to obtain the transaminase-substrate complex in an equilibrium state; perform production phase simulation on the transaminase-substrate complex in the equilibrium state to perform structure stability and energy change analysis, and take results of the structure stability and energy change analysis as the state change data in the molecular dynamics simulation process.
7. The transaminase screening method according to claim 6, wherein, perform third screening on the first data according to the state change data to obtain first data after the third screening, including: obtain at least one target index from the state change data, the target index being the binding energy, the number of hydrogen bonds, the root mean square deviation, or the root mean square fluctuation; perform third screening on the first data according to the at least one target index to obtain first data after the third screening.
8. A transaminase screening device, characterized by, including: a sequence construction module configured to obtain first data to be predicted, the first data including an amino acid sequence of a transaminase and a molecular structure of a substrate; a first prediction module configured to input the first data into a pre-trained protein-substrate binding activity prediction model to output first prediction results of the transaminase and the substrate corresponding to the first data, wherein the protein-substrate binding activity prediction model is obtained by training preset enzyme-substrate pair data; a first screening module configured to perform first screening on the first data according to the first prediction results to obtain first data after the first screening; a second prediction module configured to input the first data after the first screening into a pre-trained protein structure prediction model to output a three-dimensional structure of a complex formed by the transaminase and the substrate; a second screening module configured to obtain binding data of the transaminase and the substrate according to the three-dimensional structure of the complex, and perform screening on the first data according to the binding data to obtain first data after the second screening; a third prediction and screening module configured to perform molecular dynamics simulation on a transaminase-substrate complex corresponding to the first data after the second screening to perform third screening on the first data.
9. An electronic device, comprising: a processor and a memory coupled to the processor, the memory storing program instructions executable by the processor; the processor implements the transaminase screening method in any one of claims 1-7 when executing the program instructions stored in the memory.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores program instructions, and the program instructions are executed by the processor to realize the transaminase screening method in any one of claims 1-7.