Method and device for evaluating enzymatic properties of Daqu and identifying key compounds

By creating a compound-enzyme performance database and neural network model, combining SHAP algorithm and abnormal alarm module, the systematic and intelligent identification problems of Daqu enzyme performance evaluation are solved, and the accurate identification of key compounds and production process optimization are achieved.

CN120260724APending Publication Date: 2025-07-04WULIANGYE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510498569.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the prior art, Daqu enzymatic performance evaluation lacks systematicity and intelligence, and it is difficult to identify and regulate key compounds, which affects the efficiency of the fermentation process and product quality.

Method used

Create a compound-enzyme performance database, build an enzyme performance prediction model based on neural network models, identify key compounds through compound characteristics contribution, combine SHAP algorithm and abnormal alarm module, and regularly update the model to maintain accuracy.

Benefits of technology

It realizes high-accurate enzymatic performance recognition, facilitates identification of key compounds, and supports optimization of Daqu production process and quality improvement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260724A_ABST
    Figure CN120260724A_ABST
Patent Text Reader

Abstract

The invention relates to the field of wine brewing, provides a method and a device for evaluating enzymatic properties of Daqu and identifying key compounds in order to conveniently identify the enzymatic properties and obtain the compounds playing a key role on the enzymatic properties, and establishes an enzymatic property prediction model based on compound characteristics and corresponding enzymatic property characteristics. The enzymatic performance recognition accuracy is higher; the characteristic contribution degree of the compound is introduced based on the enzymological performance prediction model, and the compound playing a key role in the enzymological performance can be determined by calculating the characteristic contribution degree of the compound, so that the use is more convenient; the reliability of the model is regularly checked through the abnormity alarm module, the compound-enzymology performance database is regularly updated, the enzymology performance prediction model is adjusted according to the updated data, the real-time effectiveness of the enzymology performance prediction model is kept, and the use is more reliable; reference can be provided for optimization and quality improvement of a subsequent Daqu production process by determining the compound playing a key role in enzymatic performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of brewing, and in particular to a method and device for evaluating the enzymatic properties of Daqu and identifying key compounds. Background Art

[0002] As a key fermenting agent in the brewing fermentation process, the enzymatic properties of Daqu directly affect the efficiency of the fermentation process and the quality of the products. Currently, the evaluation of the enzymatic properties of Daqu mostly relies on experience and experimental determination, lacking systematic and intelligent evaluation methods. In addition, there are many types of compounds in Daqu, but the specific effects of each compound on the enzymatic properties are not yet clear, resulting in great difficulties in identifying and regulating key compounds during the production process. Summary of the Invention

[0003] In order to facilitate the identification of enzymatic properties and obtain compounds that play a key role in enzymatic properties, the present application provides a method and device for evaluating the enzymatic properties of Daqu and identifying key compounds.

[0004] The technical solution adopted by the present invention to solve the above problems is as follows:

[0005] A method for evaluating the enzymatic properties of Daqu and identifying key compounds, comprising:

[0006] Step 1, create a compound-enzymatic property database, including compound characteristics and corresponding enzymatic property characteristics;

[0007] Step 2, based on a neural network model, with compound characteristics as input and enzymatic property characteristics as output, create an enzymatic property prediction model, and train the enzymatic property prediction model based on the compound-enzymatic property database;

[0008] Step 3, embed a compound feature contribution module for analyzing the contribution degree of each compound feature into the trained enzymatic property prediction model;

[0009] Step 4, obtain the compound characteristics of the Daqu to be predicted, and obtain the enzymatic property prediction result and key compound identification result of the Daqu based on the enzymatic property prediction model.

[0010] Further, the enzymatic property characteristics include: saccharifying power, fermenting power, liquefying power, and esterifying power.

[0011] Further, the compound characteristics include: the contents of esters, alcohols, aldehydes, acids, ketones, pyrazines, furans, and aromatic compounds.

[0012] Further, the specific steps of training the enzymatic property prediction model based on the compound-enzymatic property database are:

[0013] Divide the data in the compound-enzymatic property database into a training set and a test set;

[0014] Use the enzymatic performance characteristics in the compound-enzymatic performance database as the true output labels;

[0015] Train the neural network model with the training set. During the training process, compare the predicted values of the neural network model with the corresponding true output labels, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network;

[0016] Validate the neural network model with the test set. When the corresponding loss function is less than the loss function threshold and the determination coefficient value of the predicted value and the corresponding true output label is greater than the determination coefficient threshold, use the corresponding neural network model as the enzymatic performance prediction model.

[0017] Furthermore, when the model is initially trained, determine the initial weights according to the statistical characteristics of the compound features; the formula for the initial weights is:

[0018]

[0019] In the formula, is the initialization weight of the i-th compound feature, N is the total number of compound features; φ i is the variance value of the i-th compound feature;

[0020] During the model training process, adjust the weight allocation of each feature in real time through the attention mechanism, and the calculation formula is:

[0021]

[0022] In the formula, w i is the weight of the i-th compound feature, x i is the i-th compound feature; f(x i ) is the importance score of the i-th compound feature; N is the total number of compound features.

[0023] Furthermore, it also includes regularly updating the compound-enzymatic performance database and adjusting the enzymatic performance prediction model according to the updated data.

[0024] Furthermore, the processing steps of the compound feature contribution module are:

[0025] Calculate the SHAP value of each compound feature on the prediction result through the SHAP algorithm;

[0026] Evaluate the contribution degree of each compound feature according to the calculated SHAP value;

[0027] Sort the contribution degrees of all compound features, and select the compounds corresponding to the top A compound features with the highest contribution degrees as the key compounds.

[0028] Furthermore, when creating the compound-enzymatic property database, it also includes: preprocessing the data.

[0029] The device for evaluating the enzymatic properties of Daqu and identifying key compounds includes:

[0030] Data collection module: Obtain the compound characteristics and corresponding enzymatic property characteristics of Daqu samples to form a compound-enzymatic property dataset;

[0031] Model construction module: Based on the neural network model, use the compound characteristics as input and the enzymatic property characteristics as output to create an enzymatic property prediction model, and train the enzymatic property prediction model based on the compound-enzymatic property dataset;

[0032] Feature contribution calculation module: Analyze the contribution degree of each compound characteristic based on the enzymatic property prediction model;

[0033] Feature acquisition module: Obtain the compound characteristics of the Daqu to be predicted;

[0034] Prediction and identification module: Obtain the predicted enzymatic property results of the Daqu to be predicted based on the enzymatic property prediction model, and determine the key compounds based on the contribution degree of each compound characteristic.

[0035] Furthermore, it also includes:

[0036] Abnormal alarm module: Used to regularly obtain the actual enzymatic properties of Daqu. When the determination coefficient value between the results predicted by the enzymatic property prediction model and the actual enzymatic properties is lower than the preset threshold, trigger an alarm, and update the compound-enzymatic property database and the enzymatic property prediction model.

[0037] The beneficial effects of the present invention compared with the prior art are as follows: An enzymatic property prediction model is created based on the compound characteristics and corresponding enzymatic property characteristics, and the enzymatic property identification is completed based on the enzymatic property prediction model, with higher accuracy; The contribution degree of compound characteristics is introduced based on the enzymatic property prediction model, and the compounds that play a key role in the enzymatic properties can be determined by calculating the contribution degree of compound characteristics, which is more convenient to use; The reliability of the model is regularly inspected through the abnormal alarm module, and the compound-enzymatic property database is regularly updated, and the enzymatic property prediction model is adjusted according to the updated data to maintain the real-time effectiveness of the enzymatic property prediction model, which is more reliable to use; By determining the compounds that play a key role in the enzymatic properties, it can provide a reference for the optimization of subsequent Daqu production processes and quality improvement. Description of the Drawings

[0038] Figure 1 It is a flow chart of the method for evaluating the enzymatic properties of Daqu and identifying key compounds;

[0039] Figure 2Schematic diagram of the change of the loss function during the training process;

[0040] Figure 3 Schematic diagram of the prediction result;

[0041] Figure 4 Schematic diagram of the SHAP value corresponding to the compound;

[0042] Figure 5 Schematic diagram of the structure of the device for evaluating the enzymatic properties of Daqu and identifying key compounds. Detailed implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0044] As Figure 1 shown, the present application provides a method for evaluating the enzymatic properties of Daqu and identifying key compounds, including:

[0045] Step 1, create a compound-enzymatic property database, including compound characteristics and corresponding enzymatic property characteristics. Among them, the compound characteristics include: phenethyl alcohol, 4-octanol, R-2,3-butanediol, S-2,3-butanediol, ethyl hexadecanoate, ethyl caproate, trimethylpyrazine, ethyl oleate, ethyl nonanoate, pentanol, or other esters, alcohols, aldehydes, acids, ketones, pyrazines, furans, aromatic compounds, etc. commonly found in the Daqu fermentation process, and collect data in the form of compound content; the enzymatic property characteristics include: saccharifying power, fermenting power, liquefying power, esterifying power, and may also include other enzymatic properties considered by experts in the field that need to be considered.

[0046] In order to improve the data accuracy, when creating the compound-enzymatic property database, this embodiment also preprocesses the data, such as removing missing values or outliers in the compound-enzymatic property dataset.

[0047] Step 2, based on the neural network model, use the compound characteristics as the input and the enzymatic property characteristics as the output to create an enzymatic property prediction model, and train the enzymatic property prediction model based on the compound-enzymatic property database.

[0048] Neural network models such as MLP, LSTM, RNN, CNN, Transformer, etc. can also use a combined model of the above models. In this embodiment, the enzymatic property prediction model is constructed based on MLP.

[0049] To eliminate the influence of different compound magnitudes, the compound data can also be normalized before being input into the model. The sklearn.preprocessing.StandardScaler function can be used to standardize the data into a normal distribution with a mean of 0 and a standard deviation of 1.

[0050] The specific steps for training the enzymatic performance prediction model based on the compound-enzymatic performance database are as follows: Divide the data in the compound-enzymatic performance database into a training set and a test set; Use the enzymatic performance in the compound-enzymatic performance database as the true output label; Train the neural network model with the training set. During the training process, compare the predicted values of the neural network model with the corresponding true output labels, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network; Validate the neural network model with the test set. When the corresponding loss function is less than the loss function threshold, use the corresponding neural network model as the enzymatic performance prediction model.

[0051] In this embodiment, the loss function uses the logarithmic loss function, and its expression is: In the formula, and y are the predicted value and the true value of the enzymatic performance respectively; n is the number of samples; The calculation formula of the cosh function is: Other error calculation functions recognized by experts in the field, such as the mean squared error and the mean absolute error, can also be used for calculation, which is not limited here. In this embodiment, the changes in the loss function and the accuracy rate during the training process are as Figure 2 shown. Use the coefficient of determination value of the predicted value and the true value of the test set as an index to evaluate the model effect. When the mean of the coefficient of determination is greater than 0.9, it is considered that the model training is good and the training can be considered completed. The prediction results of the four enzymatic performances of saccharifying power, fermenting power, liquefying power, and esterifying power in the test set are as Figure 3 shown. In this embodiment, in the test set, for the four enzymatic performances of saccharifying power, fermenting power, liquefying power, and esterifying power, the coefficient of determination values of the predicted value and the true value are 0.94, 0.89, 0.96, and 0.98 respectively, and their mean is greater than 0.9. Therefore, the model training is completed.

[0052] In addition, the enzymatic performance prediction model of this embodiment optimizes the response to different compound characteristics through an adaptive weight allocation mechanism, specifically including:

[0053] When the model is initially trained, determine the initial weight according to the statistical characteristics of the compound characteristics; When the model is regularly fine-tuned and updated based on the transfer learning method, determine the initial weight according to the contribution degree of each compound characteristic; The initial weight calculation formula is:

[0054]

[0055] In the formula, is the initial weight of the i-th compound feature, and N is the total number of compound features; when determining the initial weight according to the statistical characteristics of the compound data, φ i is the variance value of the i-th compound feature; when determining the initial weight according to the contribution degree of each compound feature, φ i is the mean value of the SHAP values of the i-th compound feature.

[0056] During the model training process, the weight allocation of each feature is adjusted in real time through the attention mechanism, and the calculation formula is:

[0057]

[0058] In the formula, w i is the weight of the i-th compound feature, x i is the i-th compound feature, and f(x i ) is the importance score of the i-th compound feature, which is usually calculated by the neural network of the attention layer.

[0059] The model is periodically fine-tuned and updated based on the transfer learning method, which means that the enzymatic performance prediction model is updated as the compound-enzymatic performance database is periodically updated. Specifically, it includes: after accumulating new data, retraining specific layers of the model to maintain the stability of historical data; using transfer learning technology to perform incremental updates on the existing model to reduce computational costs; and ensuring the model performance through the validation set of the compound-enzymatic performance database after the update is completed.

[0060] Step 3. To facilitate knowing the compounds that play a key role in the enzymatic performance, a compound contribution module for analyzing the contribution degree of each compound data is embedded in the trained enzymatic performance prediction model in this embodiment. The processing steps of the compound contribution module are as follows:

[0061] Step 31. Calculate the SHAP value of each compound feature for the prediction result through the SHAP algorithm;

[0062] Step 32. Evaluate the contribution degree of each compound feature according to the calculated SHAP value;

[0063] Step 33. Sort the contribution degrees of all compound features, and select the compounds corresponding to the top A compound features with the highest contribution degree as the key compounds, where A can be determined according to actual needs.

[0064] Among them, the SHAP algorithm can be expressed as:

[0065]

[0066] In the formula, φ jThe SHAP value for compound feature j; N is the set of all compound features; S is a subset of compounds excluding compound feature j; v(S) is the output value of the model when only subset S is included; v(S∪{j}) is the output value of the model when subset S and compound feature j are included; |S| is the size of subset S; |N| is the total number of compounds.

[0067] Step 4: Obtain the compound features of the Daqu to be predicted, and based on the enzymatic property prediction model, obtain the enzymatic property prediction result and the key compound identification result of the Daqu.

[0068] In this embodiment, the mean SHAP value corresponding to each metabolite compound is as Figure 4 shown, where compound 7, compound 9, and compound 2 are the top three metabolite compounds with the highest mean SHAP values and can be considered key compounds.

[0069] This embodiment also provides a device for evaluating the enzymatic properties of Daqu and identifying key compounds, as Figure 5 shown, including:

[0070] Data collection module: Obtain the compound features and corresponding enzymatic property features of the Daqu sample data to form a compound-enzymatic property data set;

[0071] Model construction module: Based on the neural network model, use the compound features as input and the enzymatic property features as output to create an enzymatic property prediction model, and train the enzymatic property prediction model based on the compound-enzymatic property data set;

[0072] Feature contribution calculation module: Analyze the contribution degree of each compound feature based on the enzymatic property prediction model;

[0073] Feature acquisition module: Obtain the compound features of the Daqu to be predicted;

[0074] Prediction and identification module: Based on the enzymatic property prediction model, obtain the enzymatic property prediction result of the Daqu to be predicted, and determine the key compounds based on the contribution degree of each compound feature.

[0075] Furthermore, an abnormal alarm module is also set up: used to regularly obtain the actual enzymatic properties of the Daqu. When the determination coefficient value between the result predicted by the enzymatic property prediction model and the actual enzymatic properties is lower than the preset threshold, an alarm is triggered, and the compound-enzymatic property database and the enzymatic property prediction model are updated. The abnormal alarm module facilitates monitoring the prediction accuracy of the enzymatic property prediction model so as to adjust the enzymatic property prediction model in a timely manner.

Claims

1. Method for evaluating enzymatic properties of Daqu and identifying key compounds, characterized in that, Including: Step 1: Create a compound-enzymatic property database, including compound characteristics and corresponding enzymatic property characteristics; Step 2: Based on a neural network model, using compound characteristics as input and enzymatic property characteristics as output, create an enzymatic property prediction model, and train the enzymatic property prediction model based on the compound-enzymatic property database; Step 3: Embed a compound feature contribution module for analyzing the contribution degree of each compound feature into the trained enzymatic property prediction model; Step 4: Obtain the compound characteristics of the Daqu to be predicted, and obtain the enzymatic property prediction result and key compound identification result of the Daqu based on the enzymatic property prediction model.

2. The method for evaluating the enzymatic properties of Daqu and identifying key compounds according to claim 1, characterized in that, The enzymatic property characteristics include: saccharifying power, fermenting power, liquefying power, and esterifying power.

3. The method for evaluating the enzymatic properties of Daqu and identifying key compounds according to claim 1, wherein The compound characteristics include: the contents of esters, alcohols, aldehydes, acids, ketones, pyrazines, furans, and aromatic compounds.

4. The method for evaluating the enzymatic properties of Daqu and identifying key compounds according to claim 1, wherein The specific steps for training the enzymatic property prediction model based on the compound-enzymatic property database are: Divide the data in the compound-enzymatic property database into a training set and a test set; Use the enzymatic property characteristics in the compound-enzymatic property database as the true output labels; Train the neural network model with the training set. During the training process, compare the predicted values of the neural network model with the corresponding true output labels, calculate the corresponding loss function, and optimize the model through the backpropagation of the neural network; Verify the neural network model with the test set. When the corresponding loss function is less than the loss function threshold and the determination coefficient value between the predicted value and the corresponding true output label is greater than the determination coefficient threshold, use the corresponding neural network model as the enzymatic property prediction model.

5. The method for evaluating the enzymatic properties of Daqu and identifying key compounds according to claim 4, characterized in that, When the model is initially trained, determine the initial weights according to the statistical characteristics of the compound characteristics; the initial weight calculation formula is: In the formula, is the initial weight of the i-th compound feature, and N is the total number of compound features; φ i is the variance value of the i-th compound feature; During the model training process, adjust the weight allocation of each feature in real time through the attention mechanism, and the calculation formula is: where w i is the weight of the i-th compound feature, and x i is the i-th compound feature; f(x i ) is the importance score of the i-th compound feature; N is the total number of compound characteristics.

6. The method for evaluating the enzymatic properties of Daqu and identifying key compounds according to claim 1, characterized in that, It also includes regularly updating the compound-enzymatic property database and adjusting the enzymatic property prediction model according to the updated data.

7. The method for evaluating the enzymatic properties of Daqu and identifying key compounds according to claim 1, wherein The processing steps of the compound feature contribution module are: Calculate the SHAP value of each compound feature for the prediction result through the SHAP algorithm; Evaluate the contribution degree of each compound feature according to the calculated SHAP value; Sort the contribution degrees of all compound features, and select the compounds corresponding to the top A compound features with the highest contribution degrees as the key compounds.

8. The method for evaluating the enzymatic properties of Daqu and identifying key compounds according to any one of claims 1-7, characterized in that, When creating the compound-enzymatic property database, it also includes: preprocessing the data.

9. Device for evaluating enzymatic properties of Daqu and identifying key compounds, characterized in that, Including: Data collection module: Obtain the compound characteristics and corresponding enzymatic property characteristics of the Daqu sample data to form a compound-enzymatic property data set; Model construction module: Based on a neural network model, using compound characteristics as input and enzymatic property characteristics as output, create an enzymatic property prediction model, and train the enzymatic property prediction model based on the compound-enzymatic property data set; Feature contribution calculation module: Analyze the contribution degree of each compound feature based on the enzymatic property prediction model; Feature acquisition module: Obtain the compound characteristics of the Daqu to be predicted; Prediction and recognition module: Obtain the predicted results of the enzymatic properties of the Daqu to be predicted based on the enzymatic property prediction model, and determine the key compounds based on the contribution degrees of the characteristics of each compound.

10. The apparatus for evaluating the enzymatic properties of Daqu and identifying key compounds based on enzyme activity according to claim 9, wherein It also includes: Abnormal alarm module: Used to regularly obtain the actual enzymatic properties of the Daqu. When the determination coefficient value between the result predicted by the enzymatic property prediction model and the actual enzymatic properties is lower than the preset threshold, an alarm is triggered, and the compound-enzymatic property database and the enzymatic property prediction model are updated.