Amino acid and RMSD feature improved glycosidase mutation prediction method and application thereof

By constructing the QSAR dictionary and MutShift prediction model, combining local and global structural characteristics, the problem of insufficient feature construction and differential analysis of existing enzyme mutation prediction methods is solved, and accurate prediction of glycosidase mutations and accuracy of enzyme activity prediction is achieved.

CN119943141AActive Publication Date: 2025-05-06青岛奔月生物技术有限公司 +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202411875795.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-19
Publication Date
2025-05-06
Estimated Expiration
2044-12-19

AI Technical Summary

Technical Problem

Existing machine learning-based enzyme mutation prediction methods have challenges in incomplete feature construction and insufficient analysis of mutant-wild type differences, resulting in poor prediction results.

Method used

By constructing a QSAR dictionary, the amino acid contribution scores were calculated, the local and global structural impacts were evaluated, and the MutShift prediction model was constructed, combining local mutants with wild-type offset metrics and global RMSD to achieve accurate prediction of glycosidase mutations.

Benefits of technology

It improves the accuracy of the enzyme activity prediction model, can learn and predict the increase in enzyme activity more effectively, and provides support for the research and application of glycosidase mutations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943141A_ABST
    Figure CN119943141A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of enzyme activity prediction, in particular to an amino acid and RMSD feature improved glycosidase mutation prediction method and application thereof. The method comprises the following steps: selecting amino acid features and constructing a QSAR dictionary, determining feature values and weights and calculating amino acid contribution scores, calculating mutant and wild type offset metrics and reflecting local influences, introducing protein RMSD as global features, setting an enzyme activity dichotomy target and constructing a MutShift prediction model, and determining the amino acid activity dichotomy target. The problem of data imbalance is solved, model performance is optimized, a QSAR dictionary and protein RMSD global features are introduced, and a MutShift model is trained. According to the method, the QSAR dictionary is constructed, the amino acid residue feature set is optimized, and local and global structure influence evaluation of the mutant and the wild type is combined, so that accurate prediction of glycosidase mutation is realized, and the accuracy of an enzyme activity prediction model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of enzyme activity prediction, and in particular to a method for predicting glycosidase mutations by improving amino acid and RMSD features and an application thereof. Background Art

[0002] Glycoside transferases are responsible for catalyzing the transfer of sugars from donors (such as UDP-sugar) to acceptor molecules to generate a variety of glycosylation products. In recent years, glycoside transferases catalyze the transfer of sugars to acceptor molecules in organisms to generate a variety of glycosylation products. In recent years, optimizing their functions through mutations has become a research hotspot in enzyme engineering. However, the activity of glycoside transferases is affected by many factors, and the prediction of mutation effects is complex. Although experimental screening is effective, it is costly and time-consuming, and is not suitable for large-scale application. To solve this problem, researchers have proposed a QSAR local model construction method based on model robustness, as shown in Chinese patent CN115527619A. This method uses support vector machine regression and particle swarm optimization algorithms to establish a robust and reliable local model, which is then combined into a consistency model to improve prediction performance. However, existing enzyme mutation prediction methods based on machine learning still face challenges: feature construction is not comprehensive and mainly focuses on a single dimension; the difference analysis between mutants and wild types is insufficient, and the specific effects of mutations on enzyme activity are not fully revealed. Summary of the invention

[0003] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a method for predicting glycosidase mutations by improving amino acid and RMSD characteristics and its application.

[0004] The technical solution adopted by the present invention is as follows: A method for predicting glycosidase mutations using amino acid and RMSD features, comprising the following steps: S1. Select amino acid features and construct QSAR dictionary: S11. Collect and organize data on the physical and chemical properties of amino acids, including polarity, hydrophobicity, acidity and alkalinity, molecular weight, volume, aromaticity, donor and acceptor characteristics; S12, assigning characteristic values ​​to each amino acid based on the data and constructing a QSAR dictionary; S2. Determine the eigenvalue and weight and calculate the amino acid contribution score: S21, determining the number of amino acids in the enzyme and the number of features of each amino acid; S22, obtaining characteristic values ​​from the QSAR dictionary; S23, determining feature weights through experimental data or model training; S24, calculating amino acid contribution scores; S3. Calculate the deviation metric between mutant and wild type and reflect the local impact: S31, calculate the amino acid contribution scores of the mutant and wild-type enzymes respectively; S32, calculating the deviation metric on each amino acid residue; S4. Introducing protein RMSD as a global feature: S41, determine the α-carbon atom coordinates of the wild-type and mutant enzymes; S42, compare the differences between the mutant and wild type structures, and calculate RMSD; S5. Set the enzyme activity binary classification target and build the MutShift prediction model: S51, collecting the enzyme activity data of glycosyltransferase, including the enzyme activity values ​​of mutants and wild type; S52, setting a binary classification target according to the change direction of the enzyme activity value; S53, using learning algorithms and building MutShift prediction models; S6. Solve the data imbalance problem and optimize model performance: S61, analyzing enzyme activity data to determine the number of positive and negative samples; S62. Use oversampling technology to increase the number of positive samples and balance the two types of samples; S63, training the model on the balanced data and evaluating the model performance; S7. Introduce QSAR dictionary and protein RMSD global features and train MutShift model: S71, using QSAR dictionary features as input features and introducing the MutShift prediction model; S72. Train the model using the training data set, and evaluate the model performance using the validation data set; S73, adjusting model parameters according to the evaluation results to optimize model performance; S74. Use the optimized model to predict enzyme activity and verify the accuracy of the prediction results.

[0005] This technical solution achieves accurate prediction of glycosidase mutations by constructing a QSAR dictionary, calculating amino acid contribution scores, evaluating local and global structural effects, and building and optimizing prediction models. Specifically, amino acid features are selected and a QSAR dictionary is constructed to collect and organize the physicochemical property data of amino acids; these feature values ​​and their weights are determined, and the contribution scores of amino acids in enzymes are calculated accordingly; the impact of mutations on the local structure of enzymes is reflected by calculating the offset measure of the amino acid contribution scores between mutants and wild types; the RMSD is calculated to quantify the impact of mutations on the global structure of proteins, so as to more comprehensively evaluate the effects of mutations; a binary classification target for enzyme activity is set, and a MutShift prediction model is constructed using a learning algorithm to predict changes in the enzymatic activity of glycoside transferases based on the characteristics and mutations of amino acid residues; to solve the problem of data imbalance, the number of positive samples is increased through oversampling technology, so that the model can better learn and predict the increase in enzyme activity; QSAR dictionary features are introduced as input features into the MutShift prediction model, and training is performed using a training data set, while the model performance is evaluated using a validation data set; model parameters are adjusted according to the evaluation results to optimize model performance, and finally the optimized model is used to predict enzyme activity, and the accuracy of the prediction results is verified.

[0006] In addition, the above-mentioned amino acid and RMSD feature-based glycosidase mutation prediction method according to the present invention also has the following additional technical features: According to one embodiment of the present invention, in the step S11 of collecting and arranging the physicochemical property data of amino acids, the interaction between the active site of the enzyme and the substrate is jointly affected by polarity, hydrophobicity and acidity and alkalinity, which is used to predict the change in substrate binding ability after mutation; Molecular weight and volume affect the position and force balance of amino acids in the three-dimensional structure and are used to predict the effect of mutations on enzyme structural stability; Hydrogen bond donors and acceptors directly affect the chemical reactivity of the catalytic center and are used to determine whether the enzyme retains catalytic ability after mutation; Aromaticity and hydrophobicity reflect the effect of amino acid mutations on the global conformation and are used to predict whether the enzyme is inactivated due to mutation.

[0007] In the present technical scheme, polarity describes the uniformity of molecular charge distribution; polar amino acids form hydrogen bonds and interact with water molecules or polar groups, including Ser, Thr, Asn, Gln, Tyr, His, Lys, Arg, Asp, Glu; non-polar amino acids include Ala, Val, Ile, Leu, Phe, Met, Trp, Cys, Pro, Gly; polar amino acids often appear in enzyme active sites, affecting binding and catalytic efficiency; non-polar amino acids maintain enzyme stability; hydrophobicity describes the behavior of molecules in water; hydrophobic amino acids aggregate and stay away from polar solvents, including Ala, Val, Ile, Leu, Phe, Met, Trp, Pro; hydrophilic amino acids are others; hydrophobicity affects the characteristics of substrate binding sites and is important for substrate selectivity; in protein folding, hydrophobic amino acids form a hydrophobic core to maintain enzyme structural stability; acidity and alkalinity describe the properties of amino acid side chains; acidic amino acids include Asp, Glu; basic amino acids include Lys, Arg, His; neutral amino acids are others; acidic or basic amino acids participate in acid-base catalytic reactions at the active site and affect charge interactions; molecular weight reflects the volume and structural complexity of amino acids, standardized to Trp (204); large amino acids cause steric hindrance and affect binding; small amino acids allow structural changes; volume affects the spatial conformation of the enzyme and the accessibility of the active site, standardized to Trp (163); large amino acids restrict substrate entry; small amino acids enhance enzyme flexibility; aromaticity describes whether the amino acid has an aromatic side chain, including Phe, Tyr, Trp; non-aromatic amino acids are others; aromatic groups bind to the substrate through π-π stacking or van der Waals interactions and participate in substrate recognition; whether the amino acid side chain contains hydrogen bond donors or acceptors, Ser, Thr, Asn, etc. are typical donors and acceptors, Phe, Ala, etc. do not have this property; hydrogen bond donors and acceptors form a network at the active site, which is an important factor in the specific binding of enzymes to substrates, and the number affects the binding energy and catalytic efficiency.

[0008] According to one embodiment of the present invention, in the calculation of the amino acid contribution score in step S24, the amino acid contribution score The calculation formula is as follows:

[0009] in: is the number of amino acids in the enzyme; The number of characteristics for each amino acid, including polarity and hydrophobicity; For the The j-th feature value of an amino acid is taken from the amino acid residue dictionary; For the The weight of a feature indicates the degree of influence of the feature on the change of enzyme activity; the weight is obtained through model training or experimental data; is a bias term used to adjust the predicted baseline value. This technical solution quantifies the influence of the characteristics of each amino acid in the enzyme (including polarity, hydrophobicity, etc.) on the change of enzyme activity by calculating the amino acid contribution score, wherein the score calculation formula involves the number of amino acids, the number of features, the feature value, the feature weight and the bias term.

[0010] According to one embodiment of the present invention, in the step S32 of calculating the deviation metric on each amino acid residue, the deviation metric between the mutant and the wild type is The calculation formula is as follows:

[0011] in: The mutant and wild type The difference in amino acid contribution scores on each amino acid residue directly reflects the local impact of the mutation on the enzyme molecule; For the mutant The amino acid contribution score of each amino acid residue; The wild type The amino acid contribution score of each amino acid residue.

[0012] This technical solution intuitively reflects the local impact of mutations on enzyme molecules by calculating the difference in amino acid contribution scores between mutants and wild types at each amino acid residue, thereby deriving a deviation metric.

[0013] According to one embodiment of the present invention, in the comparison of the difference between the mutant and the wild type in step S42, the root mean square deviation between the mutant and the wild type is The calculation formula is as follows:

[0014] in: is the α-carbon number, calculated by comparing the wild type with the mutant , which can quantify the effect of mutations on the global structure of the protein; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate.

[0015] This technical solution quantifies the effect of mutations on the global structure of proteins by calculating the root mean square deviation of α-carbon atoms in the mutant and wild-type protein structures.

[0016] According to one embodiment of the present invention, in the change direction of the enzyme activity value in step S52, the enzyme activity value is derived from the data of the public pdb website and the laboratory measurement, and the binary classification target is: The target value of 1 is the increase in the activity of the mutant enzyme, which is used to characterize the enhanced activity; The target value of 0 represents a decrease in the activity of the mutant enzyme and is used to characterize the decrease in activity.

[0017] This technical solution determines the enzyme activity value through the public PDB website and laboratory measurement data, and determines the direction of change of the mutant enzyme activity, that is, whether the activity is enhanced or decreased, through a binary target (target value 1 indicates an increase in enzyme activity, and target value 0 indicates a decrease in enzyme activity).

[0018] According to an embodiment of the present invention, in step S62, in which the oversampling technique is used to increase the number of positive samples, the data of enzyme activity is imbalanced, that is, the number of positive samples is much less than the number of negative samples, wherein: The number of positive samples is ; The number of negative samples is ; The goal of oversampling is to increase the number of positive samples so that the number of samples in the two categories is balanced: Number of new samples = - = (Number of positive samples, number of newly added samples).

[0019] This technical solution increases the number of positive samples through oversampling technology to balance the difference between the number of positive samples and the number of negative samples in enzyme activity data so that the number of samples in the two categories is equal.

[0020] According to one embodiment of the present invention, the QSAR dictionary features of step S71 are used as input features, and an enzyme activity prediction model is introduced through local mutant and wild-type offset metrics and global RMSD to quantify the impact of mutants on enzyme structure and establish a connection with enzyme activity changes for predicting glycoside transferase activity.

[0021] This technical solution introduces QSAR dictionary features as input, combines local mutant and wild-type offset metrics with global RMSD, and establishes an enzyme activity prediction model to quantify the effect of mutants on enzyme structure and establish a connection with enzyme activity changes, thereby predicting glycoside transferase activity.

[0022] To achieve the above objectives, the present invention also provides an application of amino acid and RMSD features to enhance glycosidase mutation prediction.

[0023] An amino acid and RMSD feature-based glycosidase mutation prediction application is disclosed, which uses the amino acid and RMSD feature-based glycosidase mutation prediction method.

[0024] According to one embodiment of the present invention, the physicochemical property data of a single amino acid is selected in the initial stage, and then the number of features is gradually increased to form a more complex feature set; the feature set is input into the MutShift prediction model for predicting enzyme activity.

[0025] This technical solution improves the glycosidase mutation prediction method by adopting amino acid and RMSD features. In the initial stage, the physicochemical property data of a single amino acid is selected, and the number of features is gradually increased to form a complex feature set. These feature sets are then input into the MutShift prediction model to predict the enzymatic activity of glycosidases.

[0026] Compared with the prior art, the present invention has the following beneficial effects: (1) By constructing a QSAR dictionary and optimizing the amino acid residue feature set, combined with the local and global structural impact assessment of mutants and wild-type, accurate prediction of glycosidase mutations can be achieved, thereby improving the accuracy of the enzyme activity prediction model; (2) By introducing oversampling technology to solve the data imbalance problem and optimizing the parameters of the MutShift prediction model, this technical solution can more effectively learn and predict the situation of increased enzyme activity, providing support for the mutation research and application of glycosidases. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a flowchart of the method of the present invention.

[0028] Figure 2 is a graph showing the effects of local and global characteristics of the mutant on enzyme activity.

[0029] Figure 3 This is a diagram of the construction and verification process of the Mutshift prediction model.

[0030] Figure 4 It is a training and screening diagram of the QSAR amino acid residue dictionary.

[0031] Figure 5This is the evaluation and screening diagram of the Mutshift prediction model. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] like Figure 1 As shown, this embodiment provides a method for predicting glycosidase mutations using amino acid and RMSD features, comprising the following steps: S1. Select amino acid features and construct QSAR dictionary: S11. Collect and organize data on the physical and chemical properties of amino acids, including polarity, hydrophobicity, acidity and alkalinity, molecular weight, volume, aromaticity, donor and acceptor characteristics; S12, assigning characteristic values ​​to each amino acid based on the data and constructing a QSAR dictionary; S2. Determine the eigenvalue and weight and calculate the amino acid contribution score: S21, determining the number of amino acids in the enzyme and the number of features of each amino acid; S22, obtaining characteristic values ​​from the QSAR dictionary; S23, determining feature weights through experimental data or model training; S24, calculating amino acid contribution scores; S3. Calculate the deviation metric between mutant and wild type and reflect the local impact: S31, calculate the amino acid contribution scores of the mutant and wild-type enzymes respectively; S32, calculating the deviation metric on each amino acid residue; S4. Introducing protein RMSD as a global feature: S41, determine the α-carbon atom coordinates of the wild-type and mutant enzymes; S42, compare the differences between the mutant and wild type structures, and calculate RMSD; S5. Set the enzyme activity binary classification target and build the MutShift prediction model: S51, collecting the enzyme activity data of glycosyltransferase, including the enzyme activity values ​​of mutants and wild type; S52, setting a binary classification target according to the change direction of the enzyme activity value; S53, using learning algorithms and building MutShift prediction models; S6. Solve the data imbalance problem and optimize model performance: S61, analyzing enzyme activity data to determine the number of positive and negative samples; S62. Use oversampling technology to increase the number of positive samples and balance the two types of samples; S63, training the model on the balanced data and evaluating the model performance; S7. Introduce QSAR dictionary and protein RMSD global features and train MutShift model: S71, using QSAR dictionary features as input features and introducing the MutShift prediction model; S72. Train the model using the training data set, and evaluate the model performance using the validation data set; S73, adjusting model parameters according to the evaluation results to optimize model performance; S74. Use the optimized model to predict enzyme activity and verify the accuracy of the prediction results.

[0034] This technical solution achieves accurate prediction of glycosidase mutations by constructing a QSAR dictionary, calculating amino acid contribution scores, evaluating local and global structural effects, and building and optimizing prediction models. Specifically, amino acid features are selected and a QSAR dictionary is constructed to collect and organize the physicochemical property data of amino acids; these feature values ​​and their weights are determined, and the contribution scores of amino acids in enzymes are calculated accordingly; the impact of mutations on the local structure of enzymes is reflected by calculating the offset measure of the amino acid contribution scores between mutants and wild types; the RMSD is calculated to quantify the impact of mutations on the global structure of proteins, so as to more comprehensively evaluate the effects of mutations; a binary classification target for enzyme activity is set, and a MutShift prediction model is constructed using a learning algorithm to predict changes in the enzymatic activity of glycoside transferases based on the characteristics and mutations of amino acid residues; to solve the problem of data imbalance, the number of positive samples is increased through oversampling technology, so that the model can better learn and predict the increase in enzyme activity; QSAR dictionary features are introduced as input features into the MutShift prediction model, and training is performed using a training data set, while the model performance is evaluated using a validation data set; model parameters are adjusted according to the evaluation results to optimize model performance, and finally the optimized model is used to predict enzyme activity, and the accuracy of the prediction results is verified.

[0035] In addition, the above-mentioned amino acid and RMSD feature-based glycosidase mutation prediction method according to the present invention also has the following additional technical features: According to one embodiment of the present invention, in the step S11 of collecting and arranging the physicochemical property data of amino acids, the interaction between the active site of the enzyme and the substrate is jointly affected by polarity, hydrophobicity and acidity and alkalinity, which is used to predict the change in substrate binding ability after mutation; Molecular weight and volume affect the position and force balance of amino acids in the three-dimensional structure and are used to predict the effect of mutations on enzyme structural stability; Hydrogen bond donors and acceptors directly affect the chemical reactivity of the catalytic center and are used to determine whether the enzyme retains catalytic ability after mutation; Aromaticity and hydrophobicity reflect the effect of amino acid mutations on the global conformation and are used to predict whether the enzyme is inactivated due to mutation.

[0036] In the present technical scheme, polarity describes the uniformity of molecular charge distribution; polar amino acids form hydrogen bonds and interact with water molecules or polar groups, including Ser, Thr, Asn, Gln, Tyr, His, Lys, Arg, Asp, Glu; non-polar amino acids include Ala, Val, Ile, Leu, Phe, Met, Trp, Cys, Pro, Gly; polar amino acids often appear in enzyme active sites, affecting binding and catalytic efficiency; non-polar amino acids maintain enzyme stability; hydrophobicity describes the behavior of molecules in water; hydrophobic amino acids aggregate and stay away from polar solvents, including Ala, Val, Ile, Leu, Phe, Met, Trp, Pro; hydrophilic amino acids are others; hydrophobicity affects the characteristics of substrate binding sites and is important for substrate selectivity; in protein folding, hydrophobic amino acids form a hydrophobic core to maintain enzyme structural stability; acidity and alkalinity describe the properties of amino acid side chains; acidic amino acids include Asp, Glu; basic amino acids include Lys, Arg, His; neutral amino acids are others; acidic or basic amino acids participate in acid-base catalytic reactions at the active site and affect charge interactions; molecular weight reflects the volume and structural complexity of amino acids, standardized to Trp (204); large amino acids cause steric hindrance and affect binding; small amino acids allow structural changes; volume affects the spatial conformation of the enzyme and the accessibility of the active site, standardized to Trp (163); large amino acids restrict substrate entry; small amino acids enhance enzyme flexibility; aromaticity describes whether the amino acid has an aromatic side chain, including Phe, Tyr, Trp; non-aromatic amino acids are others; aromatic groups bind to the substrate through π-π stacking or van der Waals interactions and participate in substrate recognition; whether the amino acid side chain contains hydrogen bond donors or acceptors, Ser, Thr, Asn, etc. are typical donors and acceptors, Phe, Ala, etc. do not have this property; hydrogen bond donors and acceptors form a network at the active site, which is an important factor in the specific binding of enzymes to substrates, and the number affects the binding energy and catalytic efficiency.

[0037] According to one embodiment of the present invention, in the calculation of the amino acid contribution score in step S24, the amino acid contribution score The calculation formula is as follows:

[0038] in: is the number of amino acids in the enzyme; The number of characteristics for each amino acid, including polarity and hydrophobicity; For the The j-th feature value of an amino acid is taken from the amino acid residue dictionary; For the The weight of a feature indicates the degree of influence of the feature on the change of enzyme activity; the weight is obtained through model training or experimental data; is a bias term used to adjust the predicted baseline value.

[0039] This technical solution quantifies the influence of the characteristics of each amino acid in the enzyme (including polarity, hydrophobicity, etc.) on the change of enzyme activity by calculating the amino acid contribution score, wherein the score calculation formula involves the number of amino acids, the number of features, the feature value, the feature weight and the bias term.

[0040] According to one embodiment of the present invention, in the step S32 of calculating the deviation metric on each amino acid residue, the deviation metric between the mutant and the wild type is The calculation formula is as follows:

[0041] in: The mutant and wild type The difference in amino acid contribution scores on each amino acid residue directly reflects the local impact of the mutation on the enzyme molecule; For the mutant The amino acid contribution score of each amino acid residue; The wild type The amino acid contribution score of each amino acid residue.

[0042] This technical solution intuitively reflects the local impact of mutations on enzyme molecules by calculating the difference in amino acid contribution scores between mutants and wild types at each amino acid residue, thereby deriving a deviation metric.

[0043] According to one embodiment of the present invention, in the comparison of the difference between the mutant and the wild type in step S42, the root mean square deviation between the mutant and the wild type is The calculation formula is as follows:

[0044] in: is the α-carbon number, calculated by comparing the wild type with the mutant , which can quantify the effect of mutations on the global structure of the protein; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate.

[0045] This technical solution quantifies the effect of mutations on the global structure of proteins by calculating the root mean square deviation of α-carbon atoms in the mutant and wild-type protein structures.

[0046] According to one embodiment of the present invention, in the change direction of the enzyme activity value in step S52, the enzyme activity value is derived from the data of the public pdb website and the laboratory measurement, and the binary classification target is: The target value of 1 is the increase in the activity of the mutant enzyme, which is used to characterize the enhanced activity; The target value of 0 represents a decrease in the activity of the mutant enzyme and is used to characterize the decrease in activity.

[0047] This technical solution determines the enzyme activity value through the public PDB website and laboratory measurement data, and determines the direction of change of the mutant enzyme activity, that is, whether the activity is enhanced or decreased, through a binary target (target value 1 indicates an increase in enzyme activity, and target value 0 indicates a decrease in enzyme activity).

[0048] According to an embodiment of the present invention, in step S62, in which the oversampling technique is used to increase the number of positive samples, the data of enzyme activity is imbalanced, that is, the number of positive samples is much less than the number of negative samples, wherein: The number of positive samples is ; The number of negative samples is ; The goal of oversampling is to increase the number of positive samples so that the number of samples in the two categories is balanced: Number of new samples = - = (Number of positive samples, number of newly added samples).

[0049] This technical solution increases the number of positive samples through oversampling technology to balance the difference between the number of positive samples and the number of negative samples in enzyme activity data so that the number of samples in the two categories is equal.

[0050] According to one embodiment of the present invention, the QSAR dictionary features of step S71 are used as input features, and an enzyme activity prediction model is introduced through local mutant and wild-type offset metrics and global RMSD to quantify the impact of mutants on enzyme structure and establish a connection with enzyme activity changes for predicting glycoside transferase activity.

[0051] This technical solution introduces QSAR dictionary features as input, combines local mutant and wild-type offset metrics with global RMSD, and establishes an enzyme activity prediction model to quantify the effect of mutants on enzyme structure and establish a connection with enzyme activity changes, thereby predicting glycoside transferase activity.

[0052] Example 1 Assume that the enzyme contains 5 amino acids, and the number of features for each amino acid is 3 (polarity, hydrophobicity, molecular weight). After entering the residue features and weights, calculate the mutant : Feature matrix T:

[0053]

[0054]

[0055]

[0056] S Rating:

[0057] This embodiment provides an application for improving glycosidase mutation prediction by using amino acid and RMSD features. In the initial stage, the physicochemical property data of a single amino acid is selected, and then the number of features is gradually increased to form a more complex feature set; the feature set is input into the MutShift prediction model for predicting enzyme activity.

[0058] like Figure 2 As shown in the figure, the local characteristics of the mutant such as polarity, hydrophobicity, acidity and alkalinity and the global characteristic RMSD jointly act on the change of enzyme activity to form a complementary mechanism; the upper part reflects that the local characteristics affect the binding of the enzyme to the substrate and the catalytic efficiency; the lower part quantifies the RMSD, reflecting the potential impact of global structural changes on enzyme activity; local characteristics describe the direct impact of the mutation site, and global characteristics provide the overall change background. The combination of the two enables the prediction model to comprehensively capture the complex effects of mutations on enzyme activity.

[0059] Figure 3The prediction and verification process of mutant enzyme activity is demonstrated; 145 sets of PDB data and 45 sets of laboratory data are collected, covering samples with increased, decreased and balanced enzyme activity; local features are extracted to quantify the properties of amino acid residues and mutation offsets, and global features are used to calculate RMSD to reflect overall conformational changes; features are used to train the Mutshift model, and its effectiveness is verified by calculating accuracy and wet experiments, and 5 mutants are tested using sucrose synthase as an example; the Mutshift model is verified to provide theoretical and experimental support for analyzing the impact of mutations.

[0060] Figure 4 Demonstrate the structured process of QSAR analysis of the activity of chemical substances; the core is to characterize the electron distribution, including polarity, acidity and basicity, and aromaticity; analyze the effect of steric hindrance on molecular weight and volume, consider the H-bond forming ability and emphasize the acceptor and donor functions; associate 255 mutants with biological deviations and integrate them into an integrated system to improve the ability to predict the biological activity of chemical substances.

[0061] The source and generation process of the feature set are based on the rational design of the protein sequence and structure of glycosidic transferase, and are generated by combining the physical and chemical properties of the molecule and data-driven analysis. The feature combination generation follows a systematic screening method, combining multiple feature sets according to rules to cover different physical and chemical properties; selecting a single feature in the initial stage, and gradually increasing it to form a complex set; using recursive feature elimination, correlation analysis, etc. to sort the importance of features and create new combinations. Model evaluation and screening Input the feature set into the graph convolutional network to predict enzyme activity, select the optimal feature combination based on the prediction accuracy and record the accuracy to evaluate its effectiveness, such as Figure 5 shown.

[0062] The best feature combination: feature set 240_6 (accuracy: 0.72). Key findings: Features such as hydrophobicity, molecular weight, and aromaticity contribute significantly to prediction accuracy, while too many features lead to performance degradation, as shown in Table 1.

[0063] Table 1 Feature set and its corresponding accuracy table

[0064] Data source: Multiple glycosidic transferase datasets, including activity change data measured by different experimental methods.

[0065] Practical application: To verify the accuracy of the model in predicting glycosidic transferase, a different mutant of SUSY sucrose synthase was selected for verification. The letters in the table below represent the amino acid and sequence of the mutation site. For example, V159T represents the mutation of valine 159 to threonine. The results show that Mutshift accurately predicts the activity of the five mutants of SUSY sucrose synthase, proving the effectiveness of the model. See Table 2 for details.

[0066] Table 2 Mutation site amino acids and sequence table

[0067] The present invention has important innovations in the field of enzyme activity prediction. Through node feature dictionary design and differential feature modeling, it not only improves the prediction accuracy, but also significantly enhances the practical application value of the model, providing an efficient mutant screening tool for glycoside transferases.

[0068] Although the present invention is described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person of ordinary skill in the art may easily think of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A method for predicting glycosidase mutations using amino acid and RMSD features, characterized in that: The steps include: S1. Select amino acid features and construct QSAR dictionary: S11. Collect and organize data on the physical and chemical properties of amino acids, including polarity, hydrophobicity, acidity and alkalinity, molecular weight, volume, aromaticity, donor and acceptor characteristics; S12, assigning characteristic values ​​to each amino acid based on the data and constructing a QSAR dictionary; S2. Determine the eigenvalue and weight and calculate the amino acid contribution score: S21, determining the number of amino acids in the enzyme and the number of features of each amino acid; S22, obtaining characteristic values ​​from the QSAR dictionary; S23, determining feature weights through experimental data or model training; S24, calculating amino acid contribution scores; S3. Calculate the deviation metric between mutant and wild type and reflect the local impact: S31, calculate the amino acid contribution scores of the mutant and wild-type enzymes respectively; S32, calculating the deviation metric on each amino acid residue; S4. Introducing protein RMSD as a global feature: S41, determine the α-carbon atom coordinates of the wild-type and mutant enzymes; S42, compare the differences between the mutant and wild type structures, and calculate RMSD; S5. Set the enzyme activity binary classification target and build the MutShift prediction model: S51, collecting the enzyme activity data of glycosyltransferase, including the enzyme activity values ​​of mutants and wild type; S52, setting a binary classification target according to the change direction of the enzyme activity value; S53, using learning algorithms and building MutShift prediction models; S6. Solve the data imbalance problem and optimize model performance: S61, analyzing enzyme activity data to determine the number of positive and negative samples; S62. Use oversampling technology to increase the number of positive samples and balance the two types of samples; S63, training the model on the balanced data and evaluating the model performance; S7. Introduce QSAR dictionary and protein RMSD global features and train MutShift model: S71, using QSAR dictionary features as input features and introducing the MutShift prediction model; S72. Train the model using the training data set, and evaluate the model performance using the validation data set; S73, adjusting model parameters according to the evaluation results to optimize model performance; S74. Use the optimized model to predict enzyme activity and verify the accuracy of the prediction results.

2. The method for predicting glycosidase mutations by using amino acid and RMSD features as claimed in claim 1, wherein: In the step S11, the physicochemical property data of amino acids are collected and sorted. The interaction between the active site of the enzyme and the substrate is affected by polarity, hydrophobicity and acidity and alkalinity, which are used to predict the change of substrate binding ability after mutation. Molecular weight and volume affect the position and force balance of amino acids in the three-dimensional structure and are used to predict the effect of mutations on enzyme structural stability; Hydrogen bond donors and acceptors directly affect the chemical reactivity of the catalytic center and are used to determine whether the enzyme retains catalytic ability after mutation; Aromaticity and hydrophobicity reflect the effect of amino acid mutations on the global conformation and are used to predict whether the enzyme is inactivated due to mutation.

3. The method for predicting glycosidase mutations by using amino acid and RMSD features as claimed in claim 1, wherein: In the step S24 of calculating the amino acid contribution score, the amino acid contribution score The calculation formula is as follows: in: is the number of amino acids in the enzyme; The number of characteristics for each amino acid, including polarity and hydrophobicity; For the The j-th feature value of an amino acid is taken from the amino acid residue dictionary; For the The weight of a feature indicates the degree of influence of the feature on the change of enzyme activity; the weight is obtained through model training or experimental data; is a bias term used to adjust the predicted baseline value.

4. The method for predicting glycosidase mutations by using amino acid and RMSD features as claimed in claim 1, wherein: In the step S32 of calculating the deviation metric for each amino acid residue, the deviation metric between the mutant and the wild type is The calculation formula is as follows: in: The mutant and wild type The difference in amino acid contribution scores on each amino acid residue directly reflects the local impact of the mutation on the enzyme molecule; For the mutant The amino acid contribution score of each amino acid residue; The wild type The amino acid contribution score of each amino acid residue.

5. The method for predicting glycosidase mutations by using amino acid and RMSD features as claimed in claim 1, wherein: In the step S42, in comparing the difference between the mutant and the wild type, the root mean square deviation between the mutant and the wild type The calculation formula is as follows: in: is the α-carbon number, calculated by comparing the wild type with the mutant , which can quantify the effect of mutations on the global structure of the protein; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate; and , respectively, in the wild-type and mutant protein structures. α-carbon atoms coordinate.

6. The method for predicting glycosidase mutations by using amino acid and RMSD features as claimed in claim 1, wherein: In the change direction of the enzyme activity value in step S52, the enzyme activity value is derived from the data of the public pdb website and the laboratory measurement, and the binary classification target is: The target value of 1 is the increase in the activity of the mutant enzyme, which is used to characterize the enhanced activity; The target value of 0 represents a decrease in the activity of the mutant enzyme and is used to characterize the decrease in activity.

7. The method for predicting glycosidase mutations by using amino acid and RMSD features as claimed in claim 1, wherein: In step S62, when the oversampling technique is used to increase the number of positive samples, the enzyme activity data is imbalanced, that is, the number of positive samples is far less than the number of negative samples, wherein: The number of positive samples is ; The number of negative samples is ; The goal of oversampling is to increase the number of positive samples so that the number of samples in the two categories is balanced: Number of new samples = - = (Number of positive samples, number of newly added samples).

8. The method for predicting glycosidase mutations by using amino acid and RMSD features as claimed in claim 1, wherein: The QSAR dictionary features of step S71 are used as input features, and the enzyme activity prediction model is introduced through the local mutant and wild-type offset measurement and the global RMSD to quantify the effect of the mutant on the enzyme structure and establish a connection with the enzyme activity change, which is used to predict the glycoside transferase activity.

9. An application of amino acid and RMSD features to improve glycosidase mutation prediction, using the amino acid and RMSD features to improve glycosidase mutation prediction method according to any one of claims 1 to 8.

10. The application of amino acid and RMSD features to enhance glycosidase mutation prediction according to claim 9, characterized in that: In the initial stage, the physicochemical property data of a single amino acid are selected, and then the number of features is gradually increased to form a more complex feature set; the feature set is input into the MutShift prediction model for predicting enzyme activity.

Citation Information

Patent Citations

  • Quantitative structure-activity relation model construction method based on model robustness

    CN115527619A

  • Chemical substance P-glycoprotein interaction QSAR (Quantitative Synthetic Aperture Radar) screening method

    CN114171111A

  • Method for high-throughput screening of food-borne adenosine deaminase inhibitor

    CN116825213A

  • Method for realizing directed evolution of activity and / or toxicity of antibacterial peptide through optimizer-predictor feedback loop

    CN118609660A

  • Method for predicting activity of enzyme mutants

    CN119028441A