Method and device for identifying key metabolic compounds in yeast fermentation stage
The contribution of metabolic compounds in the Daqu fermentation stage was evaluated through neural network model and SHAP theory, and the problem of difficulty in identifying key metabolic compounds in the Daqu fermentation stage was solved, achieving a more accurate recognition effect.
Patent Information
- Application Number
- CN202510365712.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the identification of key metabolic compounds in the Dako fermentation stage lacks systematicity and accuracy, and it is difficult to fully reflect metabolic characteristics.
The neural network model is used to combine SHAP theory, and the data set is constructed and divided into training, verification and test sets, and the neural network model is trained. The contribution of metabolic compounds is evaluated using SHAP values and key metabolic compounds are selected.
The accurate identification of key metabolic compounds in the Dako fermentation stage has been achieved, and the systematicity and accuracy of the identification have been improved.
Smart Images

Figure CN120280046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brewing, and particularly to a method and device for identifying key metabolic compounds in the Daqu fermentation stage. Background Art
[0002] Daqu is an essential raw material in the brewing process, and its metabolites in the fermentation stage have an important impact on the final liquor flavor and quality. In traditional methods, the identification of key metabolic compounds mainly relies on experience and qualitative analysis, lacking systematicness and accuracy, and it is difficult to comprehensively reflect the metabolic characteristics of the Daqu fermentation stage. With the development of machine learning and neural network technologies, it has become possible to quantitatively analyze the metabolic compounds in the Daqu fermentation process using data-driven models. However, how to effectively identify the key compounds that make important contributions to the fermentation stage from complex metabolic compound data remains a difficult point in current research. Summary of the Invention
[0003] The technical problem to be solved by the present invention: The present invention provides a method and device for identifying key metabolic compounds in the Daqu fermentation stage, to solve the problem of difficult identification of key metabolic compounds in the existing Daqu fermentation stage.
[0004] The technical solution adopted by the present invention to solve the above technical problem: A method for identifying key metabolic compounds in the Daqu fermentation stage, comprising the following steps:
[0005] S1. Obtain the fermentation stage data of Daqu and the corresponding metabolic compound data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The metabolic compound data includes the compound name and the corresponding content;
[0006] S2. Establish a neural network model, use the metabolic compound data in the training set as input, use the fermentation stage data in the training set as output, train the neural network model, and use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model;
[0007] S3. Use the test set to test the fermentation stage prediction model, and calculate the SHAP value of the metabolic compound data in the test set based on the SHAP theory;
[0008] S4. According to the calculated SHAP value, evaluate the contribution degree of each metabolic compound data, and select the metabolic compound as the key metabolic compound of the Daqu based on the contribution degree.
[0009] Further, the fermentation stage data includes the 1st day, the 3rd day, the 8th day, the 15th day, the 21st day, and the 28th day of fermentation.
[0010] Further, the fermentation stage data includes the initial stage of fermentation, the middle stage of fermentation, and the late stage of fermentation.
[0011] Further, the metabolic compounds include ester compounds, alcohol compounds, aldehyde compounds, acid compounds, ketone compounds, pyrazine compounds, furan compounds, and aromatic compounds.
[0012] Further, the neural network model includes one or a combination of more than one of MLP, LSTM, RNN, CNN, GRU, and Transformer.
[0013] Further, the calculation formula of the SHAP value is: where, φ j is the SHAP value of the j-th metabolic compound data; N is the set of all metabolic compound data; N\{j} is all subsets excluding the j-th metabolic compound data, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model after adding the j-th metabolic compound data; |S| is the size of subset S; |N| is the total number of metabolic compound data.
[0014] Further, the loss function of the fermentation stage prediction model is the cross-entropy loss function, and the formula of the cross-entropy loss function is: where, Loss is the cross-entropy error, M is the number of samples, C is the number of fermentation stage data, p ij is the true j-th fermentation stage data corresponding to the i-th sample, is the probability that the i-th sample is predicted as the j-th fermentation stage.
[0015] Further, before training, it also includes preprocessing the fermentation stage data and the corresponding metabolic compound data, and the preprocessing includes missing value processing, outlier processing, and normalization processing.
[0016] The present invention also provides a device for identifying key metabolic compounds in the Daqu fermentation stage, which realizes the method for identifying key metabolic compounds in the Daqu fermentation stage as described above. The device includes a data acquisition module, a model training module, a model verification module, a model testing module, a metabolic compound contribution module, and a key metabolic compound identification module. The data acquisition module is used to acquire the fermentation stage data of Daqu and the corresponding metabolic compound data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The metabolic compound data includes the compound name and the corresponding content. The model training module is used to establish a neural network model, take the metabolic compound data in the training set as input, and take the fermentation stage data in the training set as output to train the neural network model. The model verification module is used to verify the trained neural network model with the validation set to obtain a fermentation stage prediction model. The model testing module is used to test the fermentation stage prediction model with the test set. The metabolic compound contribution module is used to calculate the SHAP value of the metabolic compound data in the test set. The key metabolic compound identification module evaluates the contribution degree of each metabolic compound data according to the calculated SHAP value, and selects the metabolic compound as the key metabolic compound of the Daqu based on the contribution degree.
[0017] Further, the calculation formula of the SHAP value is: where φ j is the SHAP value of the j-th metabolic compound data; N is the set of all metabolic compound data; N\{j} is the set of all subsets excluding the j-th metabolic compound data, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only the subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model brought about by adding the j-th metabolic compound data; |S| is the size of the subset S; and |N| is the total number of metabolic compound data.
[0018] Advantages of the present invention: The present invention provides a method and device for identifying key metabolic compounds in the Daqu fermentation stage. By obtaining the fermentation stage data and corresponding metabolic compound data of Daqu, constructing a data set, dividing the data set into a training set, a validation set, and a test set, establishing a neural network model, using the metabolic compound data in the training set as input and the fermentation stage data in the training set as output, training the neural network model, validating the trained neural network model using the validation set to obtain a fermentation stage prediction model, testing the fermentation stage prediction model using the test set, calculating the SHAP values of the metabolic compound data in the test set based on the SHAP theory, and evaluating the contribution degree of each metabolic compound data according to the calculated SHAP values, and selecting metabolic compounds as the key metabolic compounds of the Daqu according to the contribution degree, the problem of difficult identification of key metabolic compounds in the existing Daqu fermentation stage is solved. Description of the Drawings
[0019] Figure 1 is a schematic flowchart of a method for identifying key metabolic compounds in the Daqu fermentation stage provided by the present invention. Detailed Embodiments
[0020] In view of the problem of difficult identification of key metabolic compounds in the Daqu fermentation stage, the present invention provides a method for identifying key metabolic compounds in the Daqu fermentation stage, as Figure 1 shown, including the following steps:
[0021] S1. Obtain the fermentation stage data and corresponding metabolic compound data of Daqu, construct a data set, and divide the data set into a training set, a validation set, and a test set according to a preset ratio. The metabolic compound data includes the compound name and the corresponding content.
[0022] Specifically, the fermentation stage data includes the 1st day, the 3rd day, the 8th day, the 15th day, the 21st day, and the 28th day of fermentation, or the fermentation stage data may also include the initial stage, the middle stage, and the late stage of fermentation. The metabolic compounds include ester compounds, alcohol compounds, aldehyde compounds, acid compounds, ketone compounds, pyrazine compounds, furan compounds, and aromatic compounds. For example, the metabolic compounds include phenethyl alcohol, 4-octanol, R23-butanediol, S23-butanediol, ethyl hexadecanoate, ethyl caproate, trimethylpyrazine, ethyl oleate, ethyl nonanoate, and pentanol, etc. Usually, 70% of the data set is divided into the training set, 20% of the data set is divided into the validation set, 10% of the data set is the test set, and the data considered representative by experts is used as the test set.
[0023] S2. Establish a neural network model, use the metabolic compound data in the training set as input, and the fermentation stage data in the training set as output. Train the neural network model, and use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model.
[0024] Specifically, the neural network model includes one or a combination of MLP, LSTM, RNN, CNN, GRU, and Transformer. Before training, it also includes preprocessing the fermentation stage data and the corresponding metabolic compound data. The preprocessing includes missing value processing, outlier processing, and normalization processing. The loss function of the fermentation stage prediction model is the cross-entropy loss function, and the formula of the cross-entropy loss function is: where Loss is the cross-entropy error, M is the number of samples, C is the number of fermentation stage data, and p ij is the true j-th fermentation stage data corresponding to the i-th sample, is the probability that the i-th sample is predicted as the j-th fermentation stage.
[0025] S3. Use the test set to test the fermentation stage prediction model, and calculate the SHAP values of the metabolic compound data in the test set based on the SHAP theory;
[0026] Specifically, the formula for calculating the SHAP value is: where φ j is the SHAP value of the j-th metabolic compound data; N is the set of all metabolic compound data; N\{j} is all subsets that do not include the j-th metabolic compound data, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model brought about by adding the j-th metabolic compound data; |S| is the size of subset S; |N| is the total number of metabolic compound data.
[0027] S4. According to the calculated SHAP values, evaluate the contribution degree of each metabolic compound data, and select metabolic compounds as the key metabolic compounds of the Daqu based on the contribution degree.
[0028] Specifically, the higher the SHAP value, the greater the contribution degree. Therefore, select the top multiple metabolic compounds arranged from high to low SHAP value as the key metabolic compounds, or use the SHAP threshold as the selection condition, and the metabolic compounds exceeding the SHAP threshold as the key metabolic compounds.
[0029] The present invention also provides a device for identifying key metabolic compounds in the Daqu fermentation stage, which realizes the method for identifying key metabolic compounds in the Daqu fermentation stage as described above. The device includes a data acquisition module, a model training module, a model verification module, a model testing module, a metabolic compound contribution module, and a key metabolic compound identification module; the data acquisition module is used to acquire the fermentation stage data of Daqu and the corresponding metabolic compound data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The metabolic compound data includes the compound name and the corresponding content; the model training module is used to establish a neural network model, use the metabolic compound data in the training set as input, and use the fermentation stage data in the training set as output to train the neural network model; the model verification module is used to verify the trained neural network model using the validation set to obtain a fermentation stage prediction model; the model testing module is used to test the fermentation stage prediction model using the test set; the metabolic compound contribution module is used to calculate the SHAP values of the metabolic compound data in the test set; the key metabolic compound identification module evaluates the contribution degree of each metabolic compound data according to the calculated SHAP values, and selects metabolic compounds as the key metabolic compounds of the Daqu based on the contribution degree.
[0030] For example, taking the fermentation stage data as the initial fermentation stage, the middle fermentation stage, and the late fermentation stage, the metabolic compound data as 61 kinds of metabolic compounds and their contents, and the neural network model used as MLP as an example, obtain the 61 kinds of metabolic compounds and their contents corresponding to the initial fermentation stage, the 61 kinds of metabolic compounds and their contents corresponding to the middle fermentation stage, and the 61 kinds of metabolic compounds and their contents corresponding to the late fermentation stage, and construct a data set with this. Divide 70% of the data set into a training set, 20% of the data set into a validation set, and 10% of the data set as a test set; use the metabolic compound data in the training set as input, use the fermentation stage data in the training set as output to train the neural network model, and use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model. Use the test set to test the fermentation stage prediction model, calculate the SHAP values of the metabolic compound data in the test set, and arrange them in descending order to obtain the arrangement order of the metabolic compound contribution degrees from high to low. Select the top five metabolic compounds as the key metabolic compounds.
Claims
1. Method for identifying key metabolic compounds in the Daqu fermentation stage, characterized in that, It includes the following steps: S1. Obtain the fermentation stage data of Daqu and the corresponding metabolite data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The metabolite data includes the compound name and the corresponding content; S2. Establish a neural network model, use the metabolite data in the training set as the input, and use the fermentation stage data in the training set as the output to train the neural network model, and use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model; S3. Use the test set to test the fermentation stage prediction model, and calculate the SHAP values of the metabolite data in the test set based on the SHAP theory; S4. According to the calculated SHAP values, evaluate the contribution degree of each metabolite data, and select metabolites as the key metabolites of the Daqu based on the contribution degree.
2. The method for identifying key metabolic compounds in the Daqu fermentation stage according to claim 1, wherein The fermentation stage data includes the 1st day of fermentation, the 3rd day of fermentation, the 8th day of fermentation, the 15th day of fermentation, the 21st day of fermentation, and the 28th day of fermentation.
3. The method for identifying key metabolic compounds in the Daqu fermentation stage according to claim 1, wherein The fermentation stage data includes the initial fermentation stage, the middle fermentation stage, and the late fermentation stage.
4. The method for identifying key metabolic compounds in the Daqu fermentation stage according to claim 1, characterized in that, The metabolites include ester compounds, alcohol compounds, aldehyde compounds, acid compounds, ketone compounds, pyrazine compounds, furan compounds, and aromatic compounds.
5. The method for identifying key metabolic compounds in the Daqu fermentation stage according to claim 1, characterized in that The neural network model includes one or a combination of MLP, LSTM, RNN, CNN, GRU, and Transformer.
6. The method for identifying key metabolic compounds in the large starter fermentation stage according to claim 1, characterized in that, The calculation formula of SHAP value is as follows: Among them, φ j is the SHAP value of the data of the j-th metabolic compound; N is the set of all metabolic compound data; N\{j} is all subsets that do not include the data of the j-th metabolic compound, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only the subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model after adding the data of the j-th metabolic compound; |S| is the size of the subset S; |N| is the total number of metabolic compound data.
7. The method for identifying key metabolic compounds in the Daqu fermentation stage according to claim 1, characterized in that, The loss function of the fermentation stage prediction model is the cross-entropy loss function, and the formula of the cross-entropy loss function is: Among them, Loss is the cross-entropy error, M is the number of samples, C is the number of data in the fermentation stage, and p ij is the true j-th fermentation stage data corresponding to the i-th sample, is the probability that the i-th sample is predicted as the j-th fermentation stage.
8. The method for identifying key metabolic compounds in the Daqu fermentation stage according to claim 1, wherein Before training, it also includes preprocessing the fermentation stage data and the corresponding metabolite data. The preprocessing includes missing value processing, outlier processing, and normalization processing.
9. Key metabolic compound identification device in the Daqu fermentation stage, characterized in that, Implement the method for identifying key metabolites in the Daqu fermentation stage as described in claim 1. The device includes a data acquisition module, a model training module, a model verification module, a model testing module, a metabolite contribution module, and a key metabolite identification module; the data acquisition module is used to obtain the fermentation stage data of Daqu and the corresponding metabolite data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The metabolite data includes the compound name and the corresponding content; the model training module is used to establish a neural network model, use the metabolite data in the training set as the input, and use the fermentation stage data in the training set as the output to train the neural network model; the model verification module is used to use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model; the model testing module is used to use the test set to test the fermentation stage prediction model; the metabolite contribution module is used to calculate the SHAP values of the metabolite data in the test set; The key metabolite identification module evaluates the contribution degree of each metabolite data according to the calculated SHAP values, and selects metabolites as the key metabolites of the Daqu based on the contribution degree.
10. The key metabolite identification device in the Daqu fermentation stage according to claim 9, characterized in that, The calculation formula of SHAP value is as follows: where φ j is the SHAP value of the data of the j-th metabolic compound; N is the set of all metabolic compound data; N\{j} is all subsets that do not contain the data of the j-th metabolic compound, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model after adding the data of the j-th metabolic compound; |S| is the size of subset S; |N| is the total number of metabolic compound data.