Method and device for discriminating yeast category and key enzyme based on enzyme activity
Through the Daqu category and key enzyme discrimination method based on enzyme activity, the neural network model and SHAP algorithm are used to solve the accuracy of Daqu category recognition, and the determination of key enzymes and the optimization of Daqu production process are achieved.
Patent Information
- Application Number
- CN202510365513.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
The prior art relies on artificial experience when determining the Daqu category, resulting in low accuracy and high volatility of identification results, making it difficult to determine the categories of key enzymes.
Based on the Daqu category and key enzyme discrimination method of enzyme activity, we create an enzyme activity-category database, use a neural network model for training, embed the enzyme activity contribution module, combine the SHAP algorithm to determine the key enzyme, and set up an abnormal alarm module to ensure the accuracy and reliability of the model.
It improves the accuracy of Daqu category judgment, can determine key enzymes, provides a more reliable discrimination method, and supports subsequent Daqu production process optimization and quality improvement.
Smart Images

Figure CN120280044A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brewing, and specifically to a method and device for discriminating the categories of Daqu and key enzymes based on enzyme activity. Background Art
[0002] As a key fermenting agent in the brewing fermentation process, differences in the categories of Daqu directly affect the fermentation process and the quality of the liquor body. According to the cultivation temperature, the categories of Daqu can be divided into high-temperature Daqu (60 - 70°C), medium-high temperature Daqu (50 - 60°C), and low-temperature Daqu (40 - 50°C). In the prior art, when determining the category of Daqu, it is mainly based on sensory evaluation such as the appearance and fragrance of Daqu. Using this method, the subjectivity is relatively large, relying on artificial experience, the accuracy of the recognition result is not high, and the recognition result often fluctuates greatly due to human factors, and it is difficult to know the categories of enzymes that play a key role in classification. Summary of the Invention
[0003] In order to improve the accuracy of Daqu category discrimination and obtain the enzymes that play a key role in category discrimination, the present application provides a method and device for discriminating the categories of Daqu and key enzymes based on enzyme activity.
[0004] The technical solution adopted by the present invention to solve the above problems is as follows:
[0005] A method for discriminating the categories of Daqu and key enzymes based on enzyme activity includes:
[0006] Step 1: Create a Daqu enzyme activity-category database, including enzyme activity data and the corresponding Daqu categories;
[0007] Step 2: Based on a neural network model, using the Daqu enzyme activity data as the input and the Daqu category as the output, create a Daqu category prediction model, and train the Daqu category prediction model based on the Daqu enzyme activity-category database;
[0008] Step 3: Embed an enzyme activity contribution module for analyzing the contribution degree of each enzyme activity data into the trained Daqu category prediction model;
[0009] Step 4: Obtain the enzyme activity data of the Daqu to be classified, classify it based on the Daqu category prediction model, and determine the key enzymes according to the contribution degree of each enzyme activity data.
[0010] Further, the Daqu categories include: low-temperature Daqu, high-temperature Daqu, and medium-high temperature Daqu.
[0011] Further, the enzyme activity data includes: enzyme activity data of α-amylase, β-amylase, glucoamylase, starch debranching enzyme, cellulase, plant sucrase, trehalase, α-galactosidase, β-galactosidase, α-glucosidase, and β-glucosidase.
[0012] Further, the specific steps for training the Daqu category prediction model based on the Daqu enzyme activity-category database are as follows:
[0013] Divide the data in the Daqu enzyme activity-category database into a training set and a test set;
[0014] Use the Daqu category in the Daqu enzyme activity-category database as the true output label;
[0015] Train a neural network model with the training set. During the training process, compare the predicted values of the neural network model with the corresponding true output labels, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network;
[0016] Validate the neural network model with the test set. When the corresponding loss function is less than the loss function threshold, use the corresponding neural network model as the Daqu category prediction model.
[0017] Further, when initially training the model, determine the initial weights according to the statistical characteristics of the enzyme activity data; the formula for the initial weights is:
[0018]
[0019] In the formula, is the initialization weight of the i-th enzyme activity feature, N is the total number of enzyme activity features; φ i is the variance value of the i-th enzyme activity feature;
[0020] During the model training process, adjust the weight distribution of each feature in real time through the attention mechanism. The calculation formula is:
[0021]
[0022] In the formula, w i is the weight of the i-th enzyme activity feature, f(x i ) is the feature importance score, and N is the total number of enzyme activity features.
[0023] Further, it also includes regularly updating the Daqu enzyme activity-category database and adjusting the Daqu category prediction model according to the updated data.
[0024] Further, the processing steps of the enzyme activity contribution module are as follows:
[0025] Calculate the SHAP value of each enzyme activity feature for the prediction result through the SHAP algorithm;
[0026] Evaluate the contribution degree of each enzyme activity feature according to the calculated SHAP value;
[0027] Sort the contribution degrees of all enzyme activity features, and select the enzymes corresponding to the A enzyme activity data with the highest contribution degrees as the key enzymes.
[0028] Furthermore, when creating the Daqu enzyme activity - category database, it also includes: pre - processing the data.
[0029] The discriminant device for Daqu category and key enzymes based on enzyme activity includes:
[0030] Enzyme activity data collection module: Collect the enzyme activity data of different categories of Daqu to form a Daqu enzyme activity - category data set;
[0031] Model construction module: Based on the neural network model, using the Daqu enzyme activity data as input and the Daqu category data as output, create a Daqu category prediction model, and train the Daqu category prediction model based on the Daqu enzyme activity - category data set;
[0032] Feature contribution calculation module: Analyze the contribution degree of each enzyme activity data based on the Daqu category prediction model;
[0033] Feature acquisition module: Acquire the enzyme activity data of the Daqu to be classified;
[0034] Classification and recognition module: Obtain the category prediction result of the Daqu to be classified based on the Daqu category prediction model, and determine the key enzymes based on the contribution degree of each enzyme activity data.
[0035] Furthermore, it also includes:
[0036] Abnormal alarm module: Used to regularly obtain the actual category of Daqu. When the proportion of the category predicted by the Daqu category prediction model that is inconsistent with the actual category is higher than the preset threshold, trigger an alarm, and update the Daqu enzyme activity - category database and the Daqu category prediction model.
[0037] The beneficial effects of the present invention compared with the prior art are as follows: Create a Daqu category prediction model based on the Daqu enzyme activity characteristics and the corresponding categories, complete category determination based on the Daqu category prediction model, with higher accuracy; Introduce the contribution degree of enzyme activity data based on the Daqu category prediction model, and the key enzymes for classification can be determined by calculating the contribution degree of enzyme activity data, which is more convenient to use; Regularly check the reliability of the model through the abnormal alarm module, and regularly update the Daqu enzyme activity - category database, and adjust the Daqu category prediction model according to the updated data to maintain the real - time effectiveness of the Daqu category prediction model, which is more reliable to use; By determining the key enzymes for various categories of Daqu, it can provide a reference for the optimization of subsequent Daqu production processes and quality improvement. Description of the Drawings
[0038] Figure 1 It is a flow chart of the discriminant method for Daqu category and key enzymes based on enzyme activity;
[0039] Figure 2 It is a schematic diagram of the average SHAP value corresponding to each enzyme in the embodiment;
[0040] Figure 3 It is a schematic diagram of the confusion matrix of prediction results;
[0041] Figure 4 It is a schematic diagram of the structure of the discriminant device for Daqu categories and key enzymes based on enzyme activity. Specific implementation manners
[0042] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0043] Daqu of different categories has different microbial community structures and enzyme activity characteristics. Based on this, as Figure 1 shown, the present application provides a method for discriminating Daqu categories and key enzymes based on enzyme activity, including:
[0044] Step 1: Create a Daqu enzyme activity-category database, including enzyme activity data and corresponding Daqu categories. Among them, the enzyme activity data includes: enzyme activity data of α-amylase, β-amylase, glucoamylase, starch debranching enzyme, cellulase, plant invertase, trehalase, α-galactosidase, β-galactosidase, α-glucosidase, β-glucosidase, etc., and the Daqu categories include: low-temperature Daqu, high-temperature Daqu, and medium-high temperature Daqu.
[0045] In order to improve the data accuracy, when creating the Daqu enzyme activity-category database, this embodiment also preprocesses the data, such as removing missing values or abnormal values in the Daqu enzyme activity-category dataset, etc.
[0046] Step 2: Based on the neural network model, use the Daqu enzyme activity data as the input and the Daqu category data as the output to create a Daqu category prediction model, and train the Daqu category prediction model based on the Daqu enzyme activity-category database.
[0047] Neural network models such as MLP, LSTM, RNN, CNN, Transformer, etc. can also use a combined model of the above models. In this embodiment, the Daqu category prediction model is constructed based on MLP, and an attention mechanism is embedded at the same time, and the number of hidden layer units is: [32, 16, 8].
[0048] In order to eliminate the influence of different enzyme activity magnitudes, the enzyme activity data can also be normalized before being input into the model, and the sklearn.preprocessing.StandardScaler function can be used to standardize the data into a normal distribution with a mean of 0 and a standard deviation of 1.
[0049] The specific steps for training the Daqu category prediction model based on the Daqu enzyme activity-category database are as follows: Divide the data in the Daqu enzyme activity-category database into a training set and a test set; use the Daqu category in the Daqu enzyme activity-category database as the true output label; train a neural network model with the training set. During the training process, compare the predicted value of the neural network model with the corresponding true output label, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network; verify the neural network model with the test set. When the corresponding loss function is less than the loss function threshold, use the corresponding neural network model as the Daqu category prediction model.
[0050] In this embodiment, the loss function adopts the cross-entropy loss function, and its expression is:
[0051]
[0052] In the formula, M is the number of samples; C is the number of stages; p ij is the true stage label of the i-th sample corresponding to the j-th stage; is the probability that the i-th sample is predicted as stage j. Other loss functions can also be used, which are not limited here.
[0053] In addition, the Daqu category prediction model of this embodiment optimizes the response to different enzyme activity characteristics through an adaptive weight allocation mechanism, specifically including:
[0054] When the model is initially trained, determine the initial weights according to the statistical characteristics of the enzyme activity data; when the model is periodically fine-tuned and updated based on the transfer learning method, determine the initial weights according to the contribution degree of each enzyme activity data. The initial weight calculation formula is:
[0055]
[0056] In the formula, is the initial weight of the i-th enzyme activity characteristic, N is the total number of enzyme activity characteristics; when determining the initial weights according to the statistical characteristics of the enzyme activity data, φ i is the variance value of the enzyme activity characteristic, and when determining the initial weights according to the contribution degree of each enzyme activity data, φ i is the mean SHAP value of the enzyme activity characteristic;
[0057] During the model training process, adjust the weight allocation of each feature in real time through the attention mechanism, and the calculation formula is:
[0058]
[0059] In the formula, w i is the weight of the i-th enzyme activity characteristic, f(x i) is the feature importance score, usually calculated by the attention layer neural network, and N is the total number of enzyme activity features.
[0060] The model is regularly fine-tuned and updated based on the transfer learning method, which means that the Daqu category prediction model is updated as the Daqu enzyme activity-category database is regularly updated. Specifically, it includes: after accumulating new data, retraining specific layers of the model to retain the stability of historical data; using transfer learning technology to perform incremental updates on the existing model to reduce computational costs; and ensuring the model performance through the validation set of the Daqu enzyme activity-category database after the update is completed.
[0061] Step 3: To facilitate knowing the enzymes that play a key role in category determination, in this embodiment, an enzyme activity contribution module for analyzing the contribution degree of each enzyme activity data is embedded in the trained Daqu category prediction model. The processing steps of the enzyme activity contribution module are as follows:
[0062] Step 31: Calculate the SHAP value of each enzyme activity feature for the prediction result through the SHAP algorithm;
[0063] Step 32: Evaluate the contribution degree of each enzyme activity feature according to the calculated SHAP value;
[0064] Step 33: Sort the contribution degrees of all enzyme activity features, and select the enzymes corresponding to the A enzyme activity data with the highest contribution degrees as the key enzymes, where A can be determined according to actual needs.
[0065] Among them, the SHAP algorithm can be expressed as:
[0066]
[0067] In the formula, φ j is the SHAP value of enzyme activity data j; N is the set of all enzyme activities; S is the enzyme activity subset; v(S) is the output value of the model when only the subset S is included; |S| is the size of the subset S; and |N| is the total number of enzyme activities.
[0068] In this embodiment, the mean SHAP value of each enzyme is as Figure 2 shown, where enzyme 5, enzyme 2, and enzyme 6 are the top three enzymes in terms of ranking and can be considered as key enzymes.
[0069] Step 4: Obtain the enzyme activity data of the Daqu to be classified, perform classification based on the Daqu category prediction model, and determine the key enzymes according to the contribution degree of each enzyme activity data.
[0070] In this embodiment, a 100% classification accuracy rate can be achieved in the test set, and the confusion matrix of the test set is as Figure 3 shown.
[0071] This embodiment also provides a device for discriminating the Daqu category and key enzymes based on enzyme activity, asFigure 4 As shown in the figure, it includes:
[0072] Enzyme activity data collection module: Collect the enzyme activity data of Daqu of different categories to form a Daqu enzyme activity-category data set;
[0073] Model construction module: Based on the neural network model, using the Daqu enzyme activity data as input and the Daqu category data as output, create a Daqu category prediction model, and train the Daqu category prediction model based on the Daqu enzyme activity-category data set;
[0074] Feature contribution calculation module: Analyze the contribution degree of each enzyme activity data based on the Daqu category prediction model, and identify key enzymes according to the contribution degree;
[0075] Feature acquisition module: Obtain the enzyme activity data of the Daqu to be classified;
[0076] Classification and identification module: Obtain the category prediction result and key enzyme identification result of the Daqu to be classified based on the Daqu category prediction model.
[0077] Furthermore, an abnormal alarm module is also set up: used to regularly obtain the actual category of Daqu. When the proportion of the category predicted by the Daqu category prediction model that is inconsistent with the actual category is higher than the preset threshold, trigger an alarm, and update the Daqu enzyme activity-category database and the Daqu category prediction model. Through the abnormal alarm module, it is convenient to monitor the prediction accuracy of the Daqu category prediction model, so as to adjust the Daqu category prediction model in a timely manner.
Claims
1. A method for discriminating the category of Daqu and key enzymes based on enzyme activity, characterized in that, Including: Step 1: Create a database of Daqu enzyme activity - category, including enzyme activity data and corresponding Daqu categories; Step 2: Based on the neural network model, using the Daqu enzyme activity data as input and the Daqu category as output, create a Daqu category prediction model, and train the Daqu category prediction model based on the Daqu enzyme activity - category database; Step 3: Embed an enzyme activity contribution module for analyzing the contribution degree of each enzyme activity data into the trained Daqu category prediction model; Step 4: Obtain the enzyme activity data of the Daqu to be classified, classify it based on the Daqu category prediction model, and determine the key enzymes according to the contribution degree of each enzyme activity data.
2. The method for discriminating the category of Daqu and key enzymes based on enzyme activity according to claim 1, wherein Daqu categories include: low - temperature Daqu, high - temperature Daqu, and medium - high - temperature Daqu.
3. The method for discriminating the category of Daqu and key enzymes based on enzyme activity according to claim 1, characterized in that, Enzyme activity data includes: enzyme activity data of α - amylase, β - amylase, glucoamylase, starch debranching enzyme, cellulase, plant sucrase, trehalase, α - galactosidase, β - galactosidase, α - glucosidase, and β - glucosidase.
4. The method for discriminating the category of Daqu and key enzymes based on enzyme activity according to claim 1, wherein The specific steps for training the Daqu category prediction model based on the Daqu enzyme activity - category database are: Divide the data in the Daqu enzyme activity - category database into a training set and a test set; Take the Daqu category in the Daqu enzyme activity - category database as the true output label; Train the neural network model with the training set. During the training process, compare the predicted value of the neural network model with the corresponding true output label, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network; Verify the neural network model with the test set. When the corresponding loss function is less than the loss function threshold, take the corresponding neural network model as the Daqu category prediction model.
5. The method for discriminating the category of Daqu and key enzymes based on enzyme activity according to claim 4, wherein When initially training the model, determine the initial weights according to the statistical characteristics of the enzyme activity data; The formula for calculating the initial weights is: Wherein, is the initial weight of the i-th enzymatic activity feature, and N is the total number of enzymatic activity features; φ i is the variance value of the i-th enzymatic activity feature; During the model training process, adjust the distribution of each feature weight in real - time through the attention mechanism. The calculation formula is: where w i is the weight of the i-th enzyme activity feature, f(x i ) is the feature importance score, and N is the total number of enzyme activity features.
6. The method for discriminating the category of Daqu and key enzymes based on enzyme activity according to claim 1, wherein It also includes regularly updating the Daqu enzyme activity - category database and adjusting the Daqu category prediction model according to the updated data.
7. The method for discriminating the category of Daqu and key enzymes based on enzyme activity according to claim 1, characterized in that, The processing steps of the enzyme activity contribution module are: Calculate the SHAP value of each enzyme activity feature for the prediction result through the SHAP algorithm; Evaluate the contribution degree of each enzyme activity feature according to the calculated SHAP value; Sort the contribution degrees of all enzyme activity features, and select the enzymes corresponding to the top A enzyme activity data with the highest contribution degree as the key enzymes.
8. The discriminant method for Daqu category and key enzymes based on enzyme activity according to any one of claims 1-7, characterized in that, When creating the Daqu enzyme activity - category database, it also includes: pre - processing the data.
9. An apparatus for discriminating the types of Daqu and key enzymes based on enzyme activity, characterized in that, Including: Enzyme activity data collection module: Collect the enzyme activity data of different categories of Daqu to form a Daqu enzyme activity - category data set; Model construction module: Based on the neural network model, using the Daqu enzyme activity data as input and the Daqu category data as output, create a Daqu category prediction model, and train the Daqu category prediction model based on the Daqu enzyme activity - category data set; Feature contribution calculation module: Analyze the contribution degree of each enzyme activity data based on the Daqu category prediction model; Feature acquisition module: Obtain the enzyme activity data of the Daqu to be classified; Classification and recognition module: Obtain the category prediction result of the Daqu to be classified based on the Daqu category prediction model, and determine the key enzymes based on the contribution degree of each enzyme activity data.
10. The discriminator for Daqu category and key enzymes based on enzyme activity according to claim 9, wherein It also includes: Abnormal alarm module: It is used to regularly obtain the actual category of Daqu. When the proportion of the inconsistency between the category predicted by the Daqu category prediction model and the actual category is higher than the preset threshold, an alarm is triggered, and the Daqu enzyme activity-category database and the Daqu category prediction model are updated.