Method, medium and device for constructing a general prediction model for classifying mutant enzymes and substrates
Through the two-way data augmentation strategy and the MEP protein characterization model, a general prediction model for classified mutant enzyme-substrate was constructed, which solved the limitations of the existing technology in predicting enzyme-substrate reactions, and achieved efficient prediction of the diversity and complexity of enzyme activities and screening of enzyme variants.
Patent Information
- Application Number
- CN202510397619.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The prior art has limitations in predicting enzyme-substrate reactions, especially in understanding the three-dimensional structural variations of different enzyme families and their relationship with enzyme activity. It is difficult to effectively generalize to new substrates or variants, and rely on a large amount of experimental data, increasing R&D costs and time.
A two-way data augmentation strategy and MEP protein characterization model were used to build a general prediction model for classified mutant enzyme-substrate by mining the correlation information of single-point and multi-point parallel mutations. This model uses the Transformer+CNN architecture to characterize mutant enzymes on a multi-scale basis, combines ECFP and GNN to characterize small substrate molecules, and finally builds a predictive model through integrated gradient enhancement trees and other models.
It significantly improves the ability of machine learning models to simulate and predict the diversity and complexity of enzyme activities, can effectively screen out potential efficient enzyme variants, reduces the cost and time of experimental verification, and provides a powerful tool for enzyme design and optimization.
Smart Images

Figure CN119905145B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of bioinformatics, protein engineering, and machine learning, and particularly to a method, medium, and device for constructing a general prediction model for classifying mutant enzymes and substrates. Background Art
[0002] In the field of protein engineering, enzymes, as biological catalysts, are widely used in scientific research and industrial production. The high efficiency and specificity of enzymes make them of great value in fields such as drug development, biofuel production, and environmental remediation. However, traditional enzyme engineering methods mainly rely on in-depth research on a single enzyme family, and optimize the catalytic efficiency of enzymes or change substrate specificity through site-directed mutagenesis or gene editing. This method has limitations in the face of complex biological reactions, especially in understanding the three-dimensional structure variations of different enzyme families and their relationships with enzyme activity.
[0003] Currently, there are various methods for predicting enzyme-substrate reactions. For example, SCANEER relies on sequence co-evolution analysis of the effects of single-point mutations on enzyme activity, but its algorithm can only evaluate about 47% of the variants, with limited application scope; ECNet and eUniRep are deep learning-based prediction methods, although they can provide relatively accurate predictions, they require experimental data of the target protein as a training set, limiting their application in the absence of activity assay data; MutCompute learns amino acid evolution information in protein structures through 3D convolutional neural networks (CNNs), but it relies on known structural data; methods such as UniKP, DLTKcat, and CmpdEnzymPred use deep learning methods for enzyme activity prediction, but are limited to small molecule substrate structures as inputs, restricting their effectiveness in broader applications.
[0004] Although these methods have made certain progress in specific situations, they generally face the challenge of being unable to be effectively extended to new substrates or variants. This limitation is manifested in practical applications as the inability to efficiently predict the activity of new substrates, difficulty in dealing with complex biological reactions, and dependence on a large amount of experimental data, increasing the R & D cost and time. Summary of the Invention
[0005] The present invention proposes a method for predicting mutant enzyme-substrate pairs based on a bidirectional data augmentation strategy and the MEP protein characterization model. This method significantly improves the ability of machine learning models to simulate and predict the diversity and complexity of enzyme activities by mining the correlation information of single-point and multi-point parallel mutations, and explores the interpretability of the effects of multi-dimensional mutations on enzyme activity.
[0006] The present invention is implemented through the following technical solutions:
[0007] A general prediction model construction method for classifying mutant enzymes and substrates
[0008] Step 1: Construct a dataset through data augmentation, including positive data screening and negative data generation;
[0009] The positive data screening: Screen enzyme-substrate pair data and phylogenetically inferred positive influence data from existing databases;
[0010] The negative data generation: Based on the existing mutant enzyme database, generate a large number of mutant enzyme sequences through single-point and multi-point mutations, and label their activity changes (enhanced, weakened, unchanged, disappeared). According to the positive-negative sample 1:1 pairing strategy, select negative samples that match the positive data;
[0011] Step 2: Build a multi-task coupling framework
[0012] In protein sequence characterization, use the MEP (Multi-scale Enzyme Protein) model to characterize mutant enzyme sequences. The MEP model adopts the architecture of Transformer+CNN: Use ESM-2 Transformer to obtain the global dependency representation of amino acid sequences, then focus on local mutation site features through a convolutional neural network (CNN), and finally fuse global and local information through residual connections to obtain the multi-scale feature vector of mutant enzymes;
[0013] For small molecule substrates, use two complementary characterization methods, and simultaneously support Extended Connectivity Fingerprints (ECFP) and GNN;
[0014] After completing the feature representation of enzymes and substrates, splice the mutant enzyme characterization vector (dimension dp) generated by MEP and the small molecule substrate characterization vector (dimension dm) generated by ECFP / GNN to form a joint feature vector v∈R dp+dm ; By integrating models including Gradient Boosting Decision Tree (GBDT), Support Vector Machine (SVM), Extreme Gradient Boosting (XGBoost), Random Forest (RF), and Logistic Regression, and adopting a dynamic weight fusion strategy, construct a final prediction model for classifying the interaction between mutant enzymes and substrates, and use the aforementioned spliced feature vector as the input to train the prediction model respectively;
[0015] Step 3: Adjust hyperparameters (such as learning rate, tree depth, regularization parameter, etc.) through grid search, and use 5-fold cross-validation to evaluate the performance of the prediction model to avoid overfitting.
[0016] Furthermore, the specific steps of generating negative data in step 1 are as follows:
[0017] (1) Screen out mutant enzyme data entries with reduced or absent activity through in silico simulation;
[0018] (2) Based on the co-evolutionary information of the mutation sites, the mutation sites in the negative samples are systematically replaced to ensure that the negative samples are comparable to the positive data in terms of mutation sites, that is, to avoid introducing model training bias due to significant differences in the selection of mutation sites between negative and positive samples (such as evolutionary conservation, functional importance, or structural relevance). Through systematic adjustments, this model can focus more on learning the impact of the mutation itself on the enzyme-substrate interaction, rather than being disturbed by irrelevant site differences, thereby improving the accuracy and generalization of predictions;
[0019] (3) Finally, negative samples that match the number of positive data are generated, and a 1:1 pairing strategy of positive and negative samples is adopted to ensure the balance of the data set and avoid bias problems in model training.
[0020] Further, the step 2 specifically includes protein characterization and small molecule characterization;
[0021] The protein characterization described above: the mutant enzyme sequence is characterized by using the MEP (Multi-scale Enzyme Protein) model; the MEP includes an evolutionary language model (ESM-2 Transformer) and a convolutional network, and the ESM-2 Transformer layer is fine-tuned on the basis of pre-training to extract the evolutionary conservative features and function-related clues of the mutant enzyme sequence; the multi-scale mutant enzyme feature extraction is achieved through the following technologies:
[0022] (1) Global dependency modeling: The mutant enzyme sequence is encoded using the Transformer layer of ESM-2; the Transformer self-attention mechanism calculates the global dependency of amino acids in the sequence using the following formula:
[0023] ;
[0024] Among them, Q, K, and V are query, key, and value matrices respectively, and dk is the dimension of the key vector; this mechanism captures the remote evolutionary conservation and functional relevance of mutant enzyme sequences.
[0025] (2) Local feature focusing: A two-layer 3-core convolutional neural network (CNN) is used to extract features from the local region of the mutant enzyme sequence; the convolution operation is defined as:
[0026] ;
[0027] Among them, i represents the output position of the current convolution operation (i.e., a certain node or position on the feature map); k is the index of the convolution kernel, used to distinguish different convolution kernels, so as to achieve multi-core parallel feature extraction; j represents the neighbor position in the input sequence segment covered by the convolution kernel (i.e., the position within the local sliding window relative to i); W is a 3-core convolution weight matrix, used to perform weighted summation on the local region of the input segment X; X is a local segment of the input sequence (such as a certain window of the amino acid sequence); b is the bias term. Through the sliding window strategy, CNN strengthens the representation ability of active sites and adjacent mutation regions;
[0028] (3) Feature fusion and robustness enhancement: Introduce residual connection to achieve layer-by-layer superposition of global and local features:
[0029] H(x)=F(x)+x;
[0030] Among them, H(x) is the final output of the residual block, F(x) is the transformation result of the current layer (such as a convolutional layer or a Transformer layer) on the input x, and x is the original input (the input of the skip connection); this design alleviates the problem of gradient disappearance and improves the adaptability of the model to local structural changes.
[0031] The aforementioned small molecule representation: Two complementary representation methods are used for substrate small molecules - supporting fingerprints and GNN simultaneously, comprehensively capturing the chemical environment and topological structure information of small molecules. Subsequently, a two-layer CNN with a window size of 3 is used to perform convolution on the sequence, focusing on the local patterns of mutation sites and their neighborhoods, and fusing the two features through residual connection to obtain a fixed-dimensional mutant enzyme representation; the specific method is as follows:
[0032] (1) Extended-Connectivity Circular Fingerprint (ECFP): Generate a fixed-length molecular fingerprint, encode the chemical environment of atoms and their neighborhoods in the molecule through a hash function, and capture functional group and multi-level topological structure information;
[0033] (2) Pre-trained Graph Neural Network (GNN): Encode the molecular graph structure based on the Graph Attention Network (GAT); the node feature update formula is:
[0034] ;
[0035] Among them, h' is the updated feature vector of node i, h j is the feature vector of neighbor node j, α ijis the attention weight of node i to neighbor j, representing the importance of j to i. W is a learnable parameter matrix used for linear transformation of node features, and σ is an activation function. Finally, a 100-dimensional task-specific vector is extracted to characterize the functional groups and spatial conformations of small molecules.
[0036] Furthermore, the performance of the prediction model is optimized in step 3 through the following steps:
[0037] Hyperparameter grid search: Combine and search for parameters including learning rate (0.01 - 0.2), tree depth (3 - 10), and subsample ratio (0.6 - 1.0) to screen the optimal parameter set;
[0038] Five-fold cross-validation: Divide the training set into 5 subsets, and iteratively evaluate the indicators of the prediction model on the validation set, including accuracy, ROC-AUC, and F1 score, to ensure generalization ability;
[0039] Test set validation: Finally, the prediction model is validated through an independent test set, showing high accuracy in the high-confidence region (prediction scores close to 0 or 1), and screening for efficient enzyme variants.
[0040] The present invention also provides a computer-readable storage medium for general prediction of classified mutant enzyme-substrate pairs. The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded and executed by a processor to perform the general prediction method for classified mutant enzyme-substrate pairs.
[0041] The present invention also provides a device for general prediction of classified mutant enzyme-substrate pairs. The device is equipped with a medium for running the general prediction model of classified mutant enzyme-substrate pairs.
[0042] Advantages of the present invention compared with the prior art:
[0043] 1. The data augmentation strategy in this step effectively solves the following problems by introducing negative samples and phylogenetic inference data: (1) Insufficient samples: The data scale is expanded through silicon chip simulation and phylogenetic inference, providing richer training samples; (2) Class imbalance: The 1:1 pairing strategy of positive and negative samples is adopted to ensure the balance of the dataset and avoid bias problems in model training; (3) Data diversity: Combining experimental data and inference data enhances the diversity and relevance of the dataset, providing more comprehensive information for the model.
[0044] 2. The multitask coupling framework of the present invention aims to capture information on both long-range dependencies (such as structurally / evolutionarily conserved sites) and local variations (the neighborhood of mutation sites) simultaneously. The MEP model is based on this multiscale design, demonstrating consideration for both long-range and short-range sequence features: The Transformer provides comprehensive sequence context understanding, while the CNN focuses on the patterns adjacent to mutation sites. Combining the two enables learning of long- and short-range dependencies and has shown competitiveness in fields such as mutation effect prediction. The present invention applies this hybrid architecture in the context of enzyme-substrate prediction, with particular attention to mutation sites, thus potentially better distinguishing functional changes brought about by key mutations compared to models using only the Transformer.
[0045] 3. The model constructed in the present invention has high accuracy in high-confidence regions (regions where the prediction scores are close to 0 and 1), can effectively screen out potential highly efficient enzyme variants, reducing the cost and time of experimental verification, and provides a powerful tool for enzyme design and optimization. The performance and interpretability are improved by using multi-model comparison and selection and dynamic weight fusion. This integrated optimization idea is somewhat different from a single deep learning framework, enabling the model to maintain both high accuracy and the ability to interpret the results.
[0046] 4. The present invention integrates the global modeling of the Transformer and the local feature extraction of the CNN to enhance the ability to capture mutation sites and functional changes, and combines experimental verification and phylogenetic inference data to significantly expand the scale and diversity of the training set. In addition, by jointly optimizing protein and small molecule representations, the accuracy of interaction prediction is improved. This model helps to screen out mutation sites that can significantly improve catalytic efficiency by predicting the interaction between mutant enzymes and substrates, optimize the activity of enzymes, demonstrating that the distal cooperative sites of mutant enzymes play an important role in enzymatic reactions in industrial production; it can also predict the catalytic ability of mutant enzymes towards different substrates, assisting in the design of enzymes with specific substrate selectivity. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a prediction result diagram. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The technical solution of the present invention will be further explained below through embodiments, but the protection scope of the present invention is not limited in any form by the embodiments.
[0049] A general prediction method for classifying mutant enzyme-substrate pairs
[0050] Step 1. Construct a data set through data augmentation, including positive data screening and negative data generation;
[0051] 1.1 Positive data screening: Association screening is carried out with the integrated enzyme-substrate pair dataset, including 18,351 experimentally verified data and 274,030 phylogenetically inferred data.
[0052] Screen enzyme-substrate pair data and positive influence data inferred phylogenetically from existing databases; As a specific implementation, accurately screen entries matching the Alex dataset from the D3 database, obtaining a total of 18,351 experimentally verified data and 274,030 phylogenetically inferred data. Carry out association screening with the integrated enzyme-substrate pair dataset to obtain 103 experimentally verified enzyme-substrate pair data and 250 phylogenetically inferred positive influence data. Screen out instances with reduced or disappeared activity through computational simulation as positive samples. These data have undergone strict experimental verification and phylogenetic analysis to ensure the reliability and representativeness of the data.
[0053] 1.2 Negative data generation: Based on the D3 mutant enzyme database, obtain more than 70,000 data. Generate a large number of mutant enzyme sequences through single-point and multi-point mutations, and label their activity changes (enhanced, weakened, unchanged, disappeared). Divide the 70,000 data into four categories: enhanced activity, weakened activity, unchanged activity, and disappeared activity. According to the 1:1 pairing strategy of positive and negative samples, select negative samples matching the positive data; The specific steps are as follows:
[0054] 1.2.1 Screen out mutant enzyme data entries with reduced or disappeared activity through silicon chip simulation;
[0055] 1.2.2 Systematically update the mutant sites in the negative samples according to the co-evolution information of the mutant sites to ensure the comparability of the mutant sites between the negative samples and the positive data;
[0056] Finally, generate negative samples matching the number of positive data. Adopting the 1:1 pairing strategy of positive and negative samples ensures the balance of the dataset and avoids bias problems in model training.
[0057] 1.3 Dataset construction Through the above steps, a dataset containing 206 experimentally verified enzyme-substrate pairs and 500 phylogenetically inferred enzyme-substrate pairs is finally constructed, achieving two-way expansion of the data. This two-way expansion refers to increasing both positive and negative examples simultaneously: using phylogenetic inference to expand positive examples to cover a wider range of enzyme-substrate pairs, and also introducing simulated negative examples to expand non-reciprocal pairs. Enrich the data from two directions, and enhance the diversity and relevance of the dataset by combining experimental data and inferred data, providing more comprehensive information for the model. The specific dataset structure is as follows:
[0058] Class 1 dataset: Contains 206 experimentally verified enzyme-substrate pair data for the training and validation of high-precision models;
[0059] Class 2 dataset: Contains 500 phylogenetically inferred enzyme-substrate pair data for enhancing the generalization ability of the model.
[0060] Step 2: Build a multi-task coupling framework
[0061] 2.1 In protein sequence characterization, adopt the architecture of Transformer + CNN: Use ESM-2 Transformer to obtain the global dependency representation of the amino acid sequence, then focus on the local mutation site features through the convolutional neural network (CNN), and finally fuse the global and local information through residual connections to obtain the multi-scale feature vector of the mutant enzyme. This architecture aims to capture the information of both long-range dependencies (such as structurally / evolutionarily conserved sites) and local changes (the neighborhood of mutation sites) simultaneously. The MEP model is based on this multi-scale design, reflecting the consideration of both long-range and short-range sequence features: Transformer provides a comprehensive understanding of the sequence context, while CNN focuses on the patterns adjacent to the mutation sites. Combining the two can learn long- and short-range dependencies and has shown competitiveness in fields such as mutant effect prediction. This hybrid architecture is applied to enzyme-substrate prediction, paying special attention to mutation sites, so it may better distinguish the functional changes brought about by key mutations compared to models that only use Transformer. Specifically, it includes the following steps:
[0062] Use the MEP (Multi-scale Enzyme Protein) model to characterize the mutant enzyme sequence. MEP essentially integrates the evolutionary-level language model (ESM-2 Transformer) and the convolutional network, and outputs a protein embedding vector that covers both global and local information. Specifically, the ESM-2 Transformer layer is fine-tuned on the basis of pre-training to extract the evolutionary conservation features and function-related clues of the sequence. The innovation lies in the embedding of the multi-scale protein, that is, compared with the simple Transformer embedding, the additional CNN layer makes the mutation sites be prominently characterized, which is different from most existing enzyme-substrate models. This method realizes multi-scale feature extraction through the following techniques:
[0063] 2.1.1 Global dependency modeling: Use the Transformer layer of ESM-2 to encode the mutant enzyme sequence. The self-attention mechanism of Transformer calculates the global dependencies of amino acids in the sequence through the following formula:
[0064] ;
[0065] Among them, Q, K, and V are the Query, Key, and Value matrices respectively, and dk is the dimension of the key vector. This mechanism captures the long-range evolutionary conservation and functional correlation of the sequence.
[0066] Local feature focusing: A two-layer 3-kernel convolutional neural network (CNN) is used to extract features from the local region of the sequence. The convolution operation is defined as:
[0067] ;
[0068] Among them, i represents the output position of the current convolution operation (i.e., a certain node or position on the feature map); k is the index of the convolution kernel, used to distinguish different convolution kernels, so as to achieve multi-kernel parallel feature extraction; j represents the neighbor position in the input sequence segment covered by the convolution kernel (i.e., the position within the local sliding window relative to i); W is the 3-kernel convolution weight matrix, used to perform weighted summation on the local region of the input segment X; X is the local segment of the input sequence (such as a certain window of the amino acid sequence); b is the bias term. Through the sliding window strategy, CNN strengthens the representation ability for active sites and adjacent mutation regions;
[0069] 2.1.2 Feature fusion and robustness enhancement: Introduce residual connection to achieve layer-by-layer superposition of global and local features:
[0070] H(x)=F(x)+x;
[0071] Among them, H(x) is the final output of the residual block, F(x) is the transformation result of the current layer (such as a convolutional layer or a Transformer layer) on the input x, and x is the original input (the input of the skip connection); this design alleviates the problem of gradient disappearance and improves the adaptability of the model to local structural changes.
[0072] 2.2 Two complementary representation methods are adopted for the substrate small molecule - simultaneously supporting fingerprints and GNN. This model can comprehensively capture the chemical environment and topological structure information of the small molecule. Subsequently, a two-layer CNN with a window size of 3 is used to perform convolution on the sequence, focusing on the local patterns of the mutation site and its neighborhood, and fusing the two features through residual connection to obtain a fixed-dimensional representation of the mutant enzyme. This method ensures that the model can "see" both the global sequence relationship and not miss the key local effects of mutations. Specifically, it is achieved through the following methods:
[0073] 2.2.1 Extended-Connectivity Circular Fingerprint (ECFP): Generate a fixed-length molecular fingerprint, encoding the chemical environment of atoms and their neighborhoods in the molecule through a hash function, and capturing functional group and multi-level topological structure information.
[0074] 2.2.2 Pre-trained Graph Neural Network (GNN): Encode the molecular graph structure based on the Graph Attention Network (GAT). The formula for updating node features is as follows:
[0075] ;
[0076] where h' is the updated feature vector of node i, h j is the feature vector of neighbor node j, α ij is the attention weight of node i to neighbor j, representing the importance of j to i, W is a learnable parameter matrix for linearly transforming node features, and σ is the activation function. Finally, a 100-dimensional task-specific vector is extracted to characterize the functional groups and spatial conformations of small molecules.
[0077] 2.3 After completing the feature representations of the enzyme and the substrate, concatenate the mutant enzyme characterization vector (dimension dp) generated by MEP with the small molecule characterization vector (dimension dm) generated by ECFP / GNN to form a joint feature vector v ∈ R dp+dm . By integrating models including Gradient Boosting Decision Tree (GBDT), Support Vector Machine (SVM), Extreme Gradient Boosting (XGBoost), Random Forest (RF), and Logistic Regression, and adopting a dynamic weight fusion strategy, construct the final prediction model for classifying the interaction between the mutant enzyme and the substrate; use the concatenated feature vector as the input to train the prediction model respectively, and optimize the performance mainly through the following steps:
[0078] 2.3.1 Hyperparameter grid search: Conduct a combined search for parameters such as learning rate (0.01–0.2), tree depth (3–10), and subsample ratio (0.6–1.0) to screen the optimal parameter set;
[0079] 2.3.2 Five-fold cross-validation: Divide the training set into 5 subsets, and iteratively evaluate metrics such as Accuracy, ROC-AUC, and F1-score of the prediction model on the validation set to ensure generalization ability;
[0080] 2.3.3 Test set validation: The final prediction model is validated through an independent test set, showing high accuracy in the high-confidence region (prediction scores close to 0 or 1) to screen for efficient enzyme variants.
[0081] Step 3: Adjust hyperparameters (such as learning rate, tree depth, regularization parameter, etc.) through grid search, and use 5-fold cross-validation to evaluate the performance of the prediction model to avoid overfitting. The model has high accuracy in high-confidence regions (regions where the prediction scores are close to 0 and 1), can effectively screen out potential high-efficiency enzyme variants, reduce the cost and time of experimental verification, and provide a powerful tool for enzyme design and optimization. Use multi-model comparison and selection and dynamic weight fusion to improve performance and interpretability. This integrated optimization idea can be somewhat different from a single deep learning framework, enabling the model to not only maintain high accuracy but also have the ability to explain the results.
[0082] According to the experimental test results (as shown in Table 1-2), this technical solution adjusts hyperparameters through grid search and combines 5-fold cross-validation to optimize the model performance, significantly improving the prediction accuracy and generalization ability: In cross-validation, GBDT demonstrated excellent parameter tuning efficiency with an accuracy of 0.9512 and an F1 score of 0.9394. SVM verified its comprehensive advantages in non-linear classification tasks with a ROC-AUC of 0.9845. And XGBoost achieved a 23% increase in the F1 score (0.7872 → 0.8611) through optimization, highlighting the efficiency of the gradient boosting framework. On the test set, LogisticRegression showed extremely strong reliability in the high-confidence region (prediction scores close to 0 or 1) with an accuracy of 0.88 and a ROC-AUC of 0.9541. The proportion of true predictions exceeded 90% (as Figure 1 shown), and the error rate was less than 5%, significantly reducing the false positive / false negative risk of experimental verification and increasing the enzyme variant screening efficiency by more than 60%. Through multi-model dynamic weight fusion (integrating the high F1 of GBDT, the high AUC of SVM, etc.), while maintaining interpretability, the comprehensive performance of this model is better than that of a single deep learning framework (such as the ROC-AUC of all models > 0.93), fully demonstrating the technical advantages of high accuracy, strong robustness, and practicality of this solution in enzyme design and optimization.
[0083] Table 1 is the table of cross-validation model evaluation metrics
[0084] ;
[0085] Table 2 is the table of test set model evaluation results
[0086] .
Claims
1. A method for constructing a universal prediction model for classifying mutant enzymes-substrates, characterized in that: The steps of the method are as follows: Step 1: Build a dataset through data expansion, including positive data screening and negative data generation The positive data screening: screening the enzyme-substrate pair data and phylogenetic inference positive impact data from the existing database; The negative data generation: Based on the existing mutant enzyme database, a large number of mutant enzyme sequences are generated through single-point and multi-point mutations, and their activity changes are marked. According to the 1:1 pairing strategy of positive and negative samples, negative samples matching the positive data are selected; Step 2: Build a multi-task coupling framework In terms of protein sequence characterization, the MEP model is used to characterize the mutant enzyme sequence. The MEP model adopts the Transformer+CNN architecture: the ESM-2Transformer is used to obtain the global dependency representation of the amino acid sequence, and then the convolutional neural network is used to focus on the local mutation site characteristics. Finally, the residual connection is used to fuse the global and local information to obtain the multi-scale feature vector of the mutant enzyme. Two complementary characterization methods are used for substrate small molecules, supporting both extended connection cycle fingerprints and convolutional neural networks; After completing the feature representation of the enzyme and substrate, the mutant enzyme representation vector generated by MEP is concatenated with the substrate small molecule representation vector generated by the extended connection cycle fingerprint ECFP and the convolutional neural network GNN to form a joint feature vector v∈R dp +dm ; By integrating models including gradient boosting trees, support vector machines, extreme gradient boosting, random forests and logistic regression, and adopting a dynamic weight fusion strategy, the final prediction type is constructed to classify the interaction between mutant enzymes and substrates, and the aforementioned spliced feature vectors are used as input to train the prediction models separately; Step 3: Adjust hyperparameters through grid search and evaluate model prediction performance using 5-fold cross validation to avoid overfitting.
2. The method for constructing a universal prediction model for classifying mutant enzymes and substrates according to claim 1, characterized in that: The specific steps of generating negative data in step 1 are as follows: (1) Screen out mutant enzyme data entries with reduced or absent activity through in silico simulation; (2) Based on the co-evolutionary information of the mutation sites, the mutation sites in the negative samples are systematically replaced to ensure that the negative samples and positive data are comparable in terms of mutation sites; (3) Finally, negative samples that match the number of positive data are generated, using a 1:1 pairing strategy for positive and negative samples.
3. The method for constructing a universal prediction model for classifying mutant enzymes and substrates according to claim 1, characterized in that: The step 2 specifically includes protein characterization and small molecule characterization; The protein characterization described above: the mutant enzyme sequence is characterized by using the MEP model; the MEP includes an evolutionary language model and a convolutional network, and the evolutionary language model is the ESM-2 Transformer, which realizes multi-scale mutant enzyme feature extraction by the following method: (1) Global dependency modeling: The mutant enzyme sequence is encoded using the Transformer layer of ESM-2. The Transformer self-attention mechanism calculates the global dependency of amino acids in the sequence using the following formula: Among them, Q, K, V are the matrices of query, key, and value respectively, and dk is the dimension of the key vector; (2) Local feature focusing: A two-layer three-core convolutional neural network is used to extract features from the local region of the mutant enzyme sequence; the convolution operation is defined as: Where i represents the output position of the current convolution operation; k is the index of the convolution kernel, j represents the neighbor position in the input sequence segment covered by the convolution kernel; W is the 3-kernel convolution weight matrix, which is used to perform weighted summation on the local area of the input segment X; X is the local segment of the input sequence; b is the bias term; (3) Feature fusion and robustness enhancement: Introducing residual connections to achieve layer-by-layer superposition of global and local features: H(x)=F(x)+x Among them, H(x) is the final output of the residual block, F(x) is the output of the current layer, and x is the input. This design alleviates the gradient vanishing problem and improves the model's adaptability to local structural changes. The small molecule characterization described above: two complementary characterization methods are used for substrate small molecules - fingerprint and GNN are supported at the same time to fully capture the chemical environment and topological structure information of small molecules, and then two layers of CNN with a window size of 3 are used to convolve the sequence, focusing on the local pattern of the mutation site and the neighborhood, and the two features are fused through residual connection to obtain a fixed-dimensional mutant enzyme representation; the specific method is as follows: (1) Extended connection loop fingerprint: Generate a fixed-length molecular fingerprint, encode the chemical environment of atoms and their neighborhoods in the molecule through a hash function, and capture functional groups and multi-level topological structure information; (2) Pre-trained graph neural network: Encode the molecular graph structure based on the graph attention mechanism; the node feature update formula is: Among them, h' is the updated feature vector of node i, h j is the feature vector of neighbor node j, α ij is the attention weight of node i to neighbor j, indicating the importance of j to i, W is a learnable parameter matrix used to perform linear transformation on node features, and σ is the activation function.
4. The method for constructing a universal prediction model for classifying mutant enzymes and substrates according to claim 1, characterized in that: In step 3, the performance of the prediction model is optimized by the following steps: Hyperparameter grid search: Combined search of parameters including learning rate, tree depth, and subsample ratio to select the optimal parameter set; Five-fold cross validation: Divide the training set into five subsets, and iteratively evaluate the prediction model on the validation set, including accuracy, ROC-AUC, and F1 score, to ensure generalization ability; Test set validation: The final prediction model was validated by an independent test set to screen for efficient enzyme variants.
5. A computer-readable storage medium for general prediction of classified mutant enzyme-substrate pairs, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the method for constructing a universal prediction model for classifying mutant enzymes and substrates according to any one of claims 1 to 4.
6. A device for general prediction of classified mutant enzyme-substrate pairs, characterized in that: The device is equipped with the computer-readable storage medium according to claim 5.
Citation Information
Patent Citations
Marine nutritional ingredient biosynthetic pathway mining method, device, equipment and medium
CN116072227A
Novel recombinant protease for n-terminal lysine and preparation method therefor
WO2024179188A1