Method and device for evaluating enzymatic properties of Daqu and discriminating key enzyme based on enzyme activity
Through the Daqu Enzyme Performance Evaluation method based on enzyme activity, the neural network model and SHAP algorithm are used to solve the accuracy of enzyme performance evaluation, and the accurate positioning of key enzymes is achieved, and the reliability and production optimization of enzyme performance prediction are improved.
Patent Information
- Application Number
- CN202510365716.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, Daqu enzymatic performance evaluation depends on experience and routine experiments, which is subjective and has low accuracy in identification results, making it difficult to determine key enzyme categories.
基于酶活的大曲酶学性能评估方法,通过创建酶活-酶学性能数据库,使用神经网络模型进行训练,嵌入酶活贡献模块,利用SHAP算法确定关键酶。
It improves the accuracy of enzymatic performance evaluation and can determine enzymes that play a key role in enzymatic performance, providing a more reliable enzymatic performance prediction and subsequent production optimization reference.
Smart Images

Figure CN120280048A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brewing, and specifically to a method and device for evaluating the enzymatic properties of Daqu based on enzyme activity and discriminating key enzymes. Background Art
[0002] As a key fermenting agent in the brewing fermentation process, the enzymatic properties of Daqu directly determine the fermentation efficiency and product quality. The enzymatic properties are jointly formed by the activities of various enzymes in Daqu, including multiple key enzymes such as α-amylase, glucoamylase, and cellulase. However, existing technologies mostly rely on experience and conventional experimental methods to evaluate the enzymatic properties. In this way, the subjectivity is relatively large, relying on artificial experience, the accuracy of the recognition results is not high, and the recognition results often fluctuate greatly due to human factors; moreover, it is difficult to know the categories of enzymes that play a key role in the enzymatic properties. Summary of the Invention
[0003] In order to improve the accuracy of enzymatic property evaluation and obtain the enzymes that play a key role in the enzymatic properties, the present application provides a method and device for evaluating the enzymatic properties of Daqu based on enzyme activity and discriminating key enzymes.
[0004] The technical solution adopted by the present invention to solve the above problems is as follows:
[0005] A method for evaluating the enzymatic properties of Daqu based on enzyme activity and discriminating key enzymes includes:
[0006] Step 1, create a Daqu enzyme activity-enzymatic property database, including Daqu enzyme activity data and the corresponding enzymatic properties;
[0007] Step 2, based on the neural network model, use the Daqu enzyme activity data as the input and the enzymatic properties as the output to create an enzymatic property prediction model, and train the enzymatic property prediction model based on the Daqu enzyme activity-enzymatic property database;
[0008] Step 3, embed an enzyme activity contribution module for analyzing the contribution degree of each enzyme activity data into the trained enzymatic property prediction model;
[0009] Step 4, obtain the enzyme activity data of the Daqu to be predicted, predict the enzymatic properties based on the enzymatic property prediction model, and determine the key enzymes according to the contribution degree of each enzyme activity data.
[0010] Further, the enzymatic properties include: saccharifying power, fermenting power, liquefying power, and esterifying power.
[0011] Further, the enzyme activity data includes: enzyme activity data of α-amylase, β-amylase, glucoamylase, starch debranching enzyme, cellulase, plant sucrase, trehalase, α-galactosidase, β-galactosidase, α-glucosidase, and β-glucosidase.
[0012] Further, the specific steps for training the enzymatic property prediction model based on the Daqu enzyme activity-enzymatic property database are as follows:
[0013] Divide the data in the Daqu enzyme activity-enzymatic property database into a training set and a test set;
[0014] Use the enzymatic properties in the Daqu enzyme activity-enzymatic property database as the true output labels;
[0015] Train a neural network model with the training set. During the training process, compare the predicted values of the neural network model with the corresponding true output labels, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network;
[0016] Validate the neural network model with the test set. When the corresponding loss function is less than the loss function threshold and the average error between the predicted value and the corresponding true output label is less than the error threshold, use the corresponding neural network model as the enzymatic property prediction model.
[0017] Further, when initially training the model, determine the initial weights according to the statistical characteristics of the enzyme activity data; the formula for calculating the initial weights is:
[0018]
[0019] In the formula, is the initialization weight of the i-th enzyme activity feature, N is the total number of enzyme activity features; φ i is the variance value of the i-th enzyme activity feature;
[0020] During the model training process, adjust the weight allocation of each feature in real time through the attention mechanism. The calculation formula is:
[0021]
[0022] In the formula, w i is the weight of the i-th enzyme activity feature; x i is the i-th enzyme activity feature; f(x i ) is the importance score of the i-th enzyme activity feature; N is the total number of enzyme activity features.
[0023] Further, it also includes regularly updating the Daqu enzyme activity-enzymatic property database and adjusting the enzymatic property prediction model according to the updated data.
[0024] Further, the processing steps of the enzyme activity contribution module are as follows:
[0025] Calculate the SHAP value of each enzyme activity feature on the prediction result through the SHAP algorithm;
[0026] Evaluate the contribution degree of each enzyme activity feature according to the calculated SHAP value;
[0027] Rank the contribution degrees of all enzyme activity characteristics, and select the enzymes corresponding to the top A enzyme activity data with the highest contribution degree as the key enzymes.
[0028] Furthermore, when creating the Daqu enzyme activity - enzymological performance database, it also includes: preprocessing the data.
[0029] The device for evaluating the enzymological performance of Daqu and discriminating key enzymes based on enzyme activity includes:
[0030] Enzyme activity data collection module: Obtain the enzyme activity characteristics and enzymological performance characteristics of the Daqu sample data to form a Daqu enzyme activity - enzymological performance data set;
[0031] Model construction module: Based on the neural network model, use the Daqu enzyme activity data as the input and the enzymological performance data as the output to create an enzymological performance prediction model, and train the enzymological performance prediction model based on the Daqu enzyme activity - enzymological performance data set;
[0032] Feature contribution calculation module: Analyze the contribution degrees of each enzyme activity data based on the enzymological performance prediction model;
[0033] Feature acquisition module: Obtain the enzyme activity characteristics of the Daqu to be predicted;
[0034] Prediction and identification module: Obtain the enzymological performance prediction result of the Daqu to be predicted based on the enzymological performance prediction model, and determine the key enzymes based on the contribution degrees of each enzyme activity data.
[0035] Furthermore, it also includes:
[0036] Abnormal alarm module: Used to regularly obtain the actual enzymological performance of Daqu. When the average error between the enzymological performance predicted by the enzymological performance prediction model and the actual enzymological performance is higher than the error threshold, trigger an alarm, and update the Daqu enzyme activity - enzymological performance database and the enzymological performance prediction model.
[0037] The beneficial effects of the present invention compared with the prior art are as follows: An enzymological performance prediction model is created based on the Daqu enzyme activity characteristics and the corresponding enzymological performance, and the enzymological performance determination is completed based on the enzymological performance prediction model, with higher accuracy; Based on the enzymological performance prediction model, the contribution degree of enzyme activity data is introduced, and the enzyme that plays a key role in the enzymological performance can be determined by calculating the contribution degree of enzyme activity data, which is more convenient to use; The reliability of the model is regularly inspected through the abnormal alarm module, and the Daqu enzyme activity - enzymological performance database is regularly updated, and the enzymological performance prediction model is adjusted according to the updated data to maintain the real - time effectiveness of the enzymological performance prediction model, which is more reliable to use; By determining the enzymes that play a key role in each enzymological performance, it can provide a reference for the subsequent optimization of the Daqu production process and quality improvement. Brief Description of the Drawings
[0038] Figure 1Flow chart of the evaluation method for the enzymatic properties of Daqu based on enzyme activity and the discriminant method for key enzymes;
[0039] Figure 2 Schematic diagram of the prediction results of saccharifying power;
[0040] Figure 3 Schematic diagram of the average SHAP values of different enzymes corresponding to saccharifying power;
[0041] Figure 4 Schematic diagram of the structure of the device for the evaluation of the enzymatic properties of Daqu based on enzyme activity and the discriminant of key enzymes. Detailed implementation manners
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the following further describes the present invention in detail with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] As Figure 1 shown, the evaluation method for the enzymatic properties of Daqu based on enzyme activity and the discriminant method for key enzymes include:
[0044] Step 1. Create a Daqu enzyme activity-enzymatic property database, including Daqu enzyme activity data and corresponding enzymatic properties. Among them, the enzyme activity data includes: enzyme activity data of α-amylase, β-amylase, glucoamylase, starch debranching enzyme, cellulase, plant sucrase, trehalase, α-galactosidase, β-galactosidase, α-glucosidase, β-glucosidase, etc., and the enzymatic properties include: saccharifying power, fermenting power, liquefying power, esterifying power, etc. The enzyme activity data and enzymatic properties may also include other data considered necessary by experts in the field. In this embodiment, the enzyme activity data of 19 enzymes including α-amylase and β-amylase are selected and are respectively shown as enzyme 1 to enzyme 19.
[0045] In order to improve the data accuracy, when creating the Daqu enzyme activity-enzymatic property database, this embodiment also preprocesses the data, such as removing missing values or abnormal values in the Daqu enzyme activity-enzymatic property dataset.
[0046] Step 2. Based on the neural network model, use the Daqu enzyme activity data as the input and the enzymatic properties as the output to create an enzymatic property prediction model, and train the enzymatic property prediction model based on the Daqu enzyme activity-enzymatic property database.
[0047] Neural network models such as MLP, LSTM, RNN, CNN, Transformer, etc. can also be combined models of the above models. In this embodiment, the enzymatic property prediction model is constructed based on MLP, and the attention mechanism is embedded at the same time. The number of hidden layer units is respectively: [64, 32, 32, 32].
[0048] To eliminate the influence of different enzyme activity levels, the enzyme activity data can also be normalized before being input into the model. The sklearn.preprocessing.StandardScaler function can be used to standardize the data into a normal distribution with a mean of 0 and a standard deviation of 1.
[0049] The specific steps for training the enzyme performance prediction model based on the Daqu enzyme activity - enzyme performance database are as follows: Divide the data in the Daqu enzyme activity - enzyme performance database into a training set and a test set; Use the Daqu enzyme performance in the Daqu enzyme activity - enzyme performance database as the true output label; Train the neural network model with the training set. During the training process, compare the predicted values of the neural network model with the corresponding true output labels, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network; Verify the neural network model with the test set. When the corresponding loss function is less than the loss function threshold and the average error between the predicted value and the corresponding true output label is less than the error threshold, use the corresponding neural network model as the enzyme performance prediction model.
[0050] In this embodiment, the loss function uses the logarithmic loss function, and its expression is: In the formula, and y are the predicted value and the true value of the enzyme performance respectively; n is the number of samples; The calculation formula of the cosh function is: Other loss functions can also be used, such as mean squared error, mean absolute error, etc., which are not limited here. In this embodiment, the percentage error between the predicted value and the true value is used to evaluate the model performance. Considering the volatility of the Daqu enzyme activity data, the error threshold in this embodiment is 10%, that is, it is considered that the error between the predicted value and the true value of the test set is not higher than 10%, and it can be considered that the model training is completed.
[0051] In addition, the enzyme performance prediction model of this embodiment optimizes the response to different enzyme activity characteristics through an adaptive weight allocation mechanism, specifically including:
[0052] When the model is initially trained, determine the initial weight according to the statistical characteristics of the enzyme activity data; When the model is periodically fine - tuned and updated based on the transfer learning method, determine the initial weight according to the contribution degree of each enzyme activity data; The initial weight calculation formula is:
[0053]
[0054] In the formula, is the initialization weight of the i - th enzyme activity characteristic, N is the total number of enzyme activity characteristics; When determining the initialization weight according to the statistical characteristics of the enzyme activity data, φ i is the variance value of the i - th enzyme activity characteristic; When determining the initialization weight according to the contribution degree of each enzyme activity data, φ iis the mean SHAP value of the i-th enzymatic activity feature;
[0055] During the model training process, the weight allocation of each feature is adjusted in real time through the attention mechanism, and the calculation formula is:
[0056]
[0057] In the formula, w i is the weight of the i-th enzymatic activity feature, x i is the i-th enzymatic activity feature, f(x i ) is the importance score of the i-th enzymatic activity feature, usually calculated by the attention layer neural network, and N is the total number of enzymatic activity features.
[0058] The model is regularly fine-tuned and updated based on the transfer learning method, which means that the enzymatic property prediction model is updated with the regular update of the Daqu enzymatic activity-enzymatic property database. Specifically, it includes: after accumulating new data, retraining specific layers of the model to maintain the stability of historical data; using transfer learning technology to perform incremental updates on the existing model to reduce computational costs; and ensuring the model performance through the validation set of the Daqu enzymatic activity-enzymatic property database after the update is completed. Figure 2 Shows a schematic diagram of the saccharifying power prediction result.
[0059] Step 3: To facilitate knowing the enzymes that play a key role in the determination of enzymatic properties, in this embodiment, an enzymatic activity contribution module for analyzing the contribution degree of each enzymatic activity data is embedded in the trained enzymatic property prediction model. The processing steps of the enzymatic activity contribution module are as follows:
[0060] Step 31: Calculate the SHAP value of each enzymatic activity feature to the prediction result through the SHAP algorithm;
[0061] Step 32: Evaluate the contribution degree of each enzymatic activity feature according to the calculated SHAP value;
[0062] Step 33: Sort the contribution degrees of all enzymatic activity features, and select the enzymes corresponding to the A enzymatic activity data with the highest contribution degrees as the key enzymes, where A can be determined according to actual needs.
[0063] Among them, the SHAP algorithm can be expressed as:
[0064]
[0065] In the formula, φ j is the SHAP value of the enzymatic activity feature j; N is the set of all enzymatic activities; S is the subset of enzymatic activity features excluding the enzymatic activity feature j; v(S) is the output value of the model when only including the subset S; v(S∪{j}) is the output value of the model when including the subset S and the enzymatic activity feature j; |S| is the size of the subset S; |N| is the total number of enzymatic activity features.
[0066] Step 4: Obtain the enzyme activity data of the Daqu to be predicted, classify it based on the enzymatic performance prediction model, and determine the key enzymes according to the contribution degree of each enzyme activity data.
[0067] In this embodiment, the schematic diagram of the average SHAP value of different enzymes corresponding to saccharifying power is as Figure 3 shown, where Enzyme 6, Enzyme 4, Enzyme 5, Enzyme 18, and Enzyme 15 are the 5 enzymes with the highest SHAP values and can be considered as key enzymes.
[0068] This embodiment also provides a device for evaluating the enzymatic performance of Daqu and discriminating key enzymes based on enzyme activity, as Figure 4 shown, including:
[0069] Enzyme activity data collection module: Obtain the enzyme activity characteristics and enzymatic performance characteristics of the Daqu sample data to form a Daqu enzyme activity-enzymatic performance data set;
[0070] Model construction module: Based on the neural network model, use the Daqu enzyme activity data as the input and the enzymatic performance data as the output to create an enzymatic performance prediction model, and train the enzymatic performance prediction model based on the Daqu enzyme activity-enzymatic performance data set;
[0071] Feature contribution calculation module: Analyze the contribution degree of each enzyme activity data based on the enzymatic performance prediction model;
[0072] Feature acquisition module: Obtain the enzyme activity characteristics of the Daqu to be predicted;
[0073] Prediction and recognition module: Obtain the enzymatic performance prediction result of the Daqu to be predicted based on the enzymatic performance prediction model, and determine the key enzymes according to the contribution degree of each enzyme activity data.
[0074] Furthermore, an abnormal alarm module is also set: used to regularly obtain the actual enzymatic performance of Daqu. When the average error between the enzymatic performance predicted by the enzymatic performance prediction model and the actual enzymatic performance is higher than the error threshold, trigger an alarm, and update the Daqu enzyme activity-enzymatic performance database and the enzymatic performance prediction model. The abnormal alarm module facilitates monitoring the prediction accuracy of the Daqu enzymatic performance prediction model so as to adjust the Daqu enzymatic performance prediction model in a timely manner.
Claims
1. An evaluation method for the enzymatic properties of Daqu based on enzyme activity and a key enzyme discrimination method, characterized in that, Including: Step 1: Create a Daqu enzyme activity - enzymatic property database, including Daqu enzyme activity data and corresponding enzymatic properties; Step 2: Based on the neural network model, using the Daqu enzyme activity data as input and the enzymatic properties as output, create an enzymatic property prediction model, and train the enzymatic property prediction model based on the Daqu enzyme activity - enzymatic property database; Step 3: Embed an enzyme activity contribution module for analyzing the contribution degree of each enzyme activity data into the trained enzymatic property prediction model; Step 4: Obtain the enzyme activity data of the Daqu to be predicted, predict the enzymatic properties based on the enzymatic property prediction model, and determine the key enzymes according to the contribution degree of each enzyme activity data.
2. The method for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to claim 1, wherein The enzymatic properties include: saccharifying power, fermenting power, liquefying power, and esterifying power.
3. The method for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to claim 1, wherein The enzyme activity data includes: enzyme activity data of α - amylase, β - amylase, glucoamylase, starch debranching enzyme, cellulase, plant sucrase, trehalase, α - galactosidase, β - galactosidase, α - glucosidase, and β - glucosidase.
4. The method for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to claim 1, wherein The specific steps for training the enzymatic property prediction model based on the Daqu enzyme activity - enzymatic property database are: Divide the data in the Daqu enzyme activity - enzymatic property database into a training set and a test set; Use the enzymatic properties in the Daqu enzyme activity - enzymatic property database as the true output labels; Train the neural network model with the training set. During the training process, compare the predicted values of the neural network model with the corresponding true output labels, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network; Verify the neural network model with the test set. When the corresponding loss function is less than the loss function threshold and the average error between the predicted value and the corresponding true output label is less than the error threshold, use the corresponding neural network model as the enzymatic property prediction model.
5. The method for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to claim 4, characterized in that, When initially training the model, determine the initial weights according to the statistical characteristics of the enzyme activity data; the formula for calculating the initial weights is: In the formula, is the initial weight of the i-th enzyme activity feature, and N is the total number of enzyme activity features; φ i is the variance value of the i-th enzyme activity feature; During the model training process, adjust the distribution of each feature weight in real - time through the attention mechanism, and the calculation formula is: where w i is the weight of the i-th enzymatic activity characteristic; x i is the i-th enzymatic activity characteristic; f(x i ) is the importance score of the i-th enzyme activity feature; N is the total number of enzyme activity features.
6. The method for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to claim 1, characterized in that It also includes regularly updating the Daqu enzyme activity - enzymatic property database and adjusting the enzymatic property prediction model according to the updated data.
7. The method for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to claim 1, characterized in that, The processing steps of the enzyme activity contribution module are: Calculate the SHAP value of each enzyme activity feature for the prediction result through the SHAP algorithm; Evaluate the contribution degree of each enzyme activity feature according to the calculated SHAP value; Sort the contribution degrees of all enzyme activity features, and select the enzymes corresponding to the top A enzyme activity data with the highest contribution degree as the key enzymes.
8. The method for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to any one of claims 1-7, characterized in that When creating the Daqu enzyme activity - enzymatic property database, it also includes: pre - processing the data.
9. An apparatus for evaluating the enzymatic properties of Daqu based on enzyme activity and discriminating key enzymes, characterized in that, Including: Enzyme activity data collection module: Obtain the enzyme activity features and enzymatic property features of the Daqu sample data to form a Daqu enzyme activity - enzymatic property data set; Model construction module: Based on the neural network model, using the Daqu enzyme activity data as input and the enzymatic property data as output, create an enzymatic property prediction model, and train the enzymatic property prediction model based on the Daqu enzyme activity - enzymatic property data set; Feature contribution calculation module: Analyze the contribution degree of each enzyme activity data based on the enzymatic property prediction model; Feature acquisition module: Obtain the enzyme activity features of the Daqu to be predicted; Prediction and Identification Module: Obtain the predicted results of the enzymatic properties of the Daqu to be predicted based on the enzymatic property prediction model, and determine the key enzymes based on the contribution degrees of each enzyme activity data.
10. The apparatus for evaluating the enzymatic properties of Daqu and discriminating key enzymes based on enzyme activity according to claim 9, wherein It also includes: Abnormal Alarm Module: Used to regularly obtain the actual enzymatic properties of the Daqu. When the average error between the enzymatic properties predicted by the enzymatic property prediction model and the actual enzymatic properties is higher than the error threshold, an alarm is triggered, and the Daqu enzyme activity-enzymatic property database and the enzymatic property prediction model are updated.
Citation Information
Cited By
Protease data prediction method for catalytic reaction and application
CN122369578A
Protease data prediction method for catalytic reaction and application
CN122369578B