Key microorganism identification method and device in yeast fermentation stage
Through neural network model and SHAP algorithm, the problem of difficulty in identifying key microorganisms in the Daqu fermentation stage was solved, and the precise identification and contribution evaluation of key microorganisms were achieved, supporting the optimization of the fermentation process and the improvement of the quality of liquor.
Patent Information
- Application Number
- CN202510365711.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, it is difficult to identify key microbials in the Dako fermentation stage, and traditional methods are time-consuming and labor-intensive and difficult to reflect the action pattern of microbial populations.
The neural network model is used to combine the SHAP algorithm, and the data set is constructed and divided into training, verification and test sets are divided into neural network models, and the contribution of microorganisms is evaluated using SHAP values to automatically identify key microorganisms.
It has achieved accurate identification of key microorganisms in the Dako fermentation stage, provided scientific basis to support fermentation process control and product quality optimization, and promoted the intelligent development of liquor brewing.
Smart Images

Figure CN120280045A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brewing, and particularly to a method and device for identifying key microorganisms in the Daqu fermentation stage. Background Art
[0002] Daqu is an important fermenting agent in Baijiu brewing. Its microbial community plays a key role in the Daqu fermentation stage, affecting the types and quality of fermentation products. In traditional processes, the identification of key microorganisms in the Daqu fermentation stage mainly relies on microbial culture and sequencing technologies. This method is usually time-consuming and laborious and is difficult to reflect the action rules of microbial populations. Summary of the Invention
[0003] The technical problem to be solved by the present invention: The present invention provides a method and device for identifying key microorganisms in the Daqu fermentation stage, solving the problem of difficult identification of key microorganisms in the existing Daqu fermentation stage.
[0004] The technical solution adopted by the present invention to solve the above technical problem: A method for identifying key microorganisms in the Daqu fermentation stage, comprising the following steps:
[0005] S1. Obtain the fermentation stage data and corresponding microbial abundance data of Daqu, construct a data set, and divide the data set into a training set, a validation set, and a test set. The microbial abundance data includes microbial species and quantities;
[0006] S2. Establish a neural network model, use the microbial abundance data in the training set as input, use the fermentation stage data in the training set as output, train the neural network model, and use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model;
[0007] S3. Use the test set to test the fermentation stage prediction model, and calculate the SHAP value of each microorganism in the test set based on the SHAP theory;
[0008] S4. According to the calculated SHAP values, evaluate the contribution degree of each microorganism, and select microorganisms as the key microorganisms of the Daqu according to the contribution degree.
[0009] Further, the microbial species include Bacillus, Aspergillus, Saccharomyces cerevisiae, Pediococcus acidilactici, and Staphylococcus.
[0010] Further, the fermentation stage data includes the 1st day of fermentation, the 3rd day of fermentation, the 8th day of fermentation, the 15th day of fermentation, the 21st day of fermentation, and the 28th day of fermentation.
[0011] Further, the fermentation stage data includes the initial fermentation stage, the middle fermentation stage, and the late fermentation stage.
[0012] Further, the neural network model includes one or a combination of more than one of MLP, LSTM, RNN, CNN, GRU, and Transformer.
[0013] Further, the calculation formula of the SHAP value is: where φ j is the SHAP value of the j-th microbial abundance data; N is the set of all microbial abundance data; N\{j} is all subsets excluding the j-th microbial abundance data, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only subset S is included, v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model after adding the j-th microbial abundance data; |S| is the size of subset S; |N| is the total number of microbial abundance data.
[0014] Further, the loss function of the fermentation stage prediction model is the cross-entropy loss function, and the formula of the cross-entropy loss function is: where Loss is the cross-entropy error, M is the number of samples, C is the number of fermentation stage data, p ij is the true j-th fermentation stage data corresponding to the i-th sample, is the probability that the i-th sample is predicted as the j-th fermentation stage.
[0015] Further, before training, it also includes preprocessing the fermentation stage data and the corresponding microbial abundance data, and the preprocessing includes missing value processing, outlier processing, and normalization processing.
[0016] The present invention also provides a key microorganism identification device in the Daqu fermentation stage, which realizes the key microorganism identification method in the Daqu fermentation stage as described above. The device includes a data acquisition module, a model training module, a model verification module, a model testing module, a microorganism contribution module, and a key microorganism identification module. The data acquisition module is used to acquire the fermentation stage data of Daqu and the corresponding microorganism abundance data, construct a data set, and divide the data set into a training set, a verification set, and a test set. The microorganism abundance data includes microorganism species and quantities. The model training module is used to establish a neural network model, take the microorganism abundance data in the training set as input, and take the fermentation stage data in the training set as output to train the neural network model. The model verification module is used to verify the trained neural network model with the verification set to obtain a fermentation stage prediction model. The model testing module is used to test the fermentation stage prediction model with the test set. The microorganism contribution module is used to calculate the SHAP value of each microorganism in the test set. The key microorganism identification module evaluates the contribution degree of each microorganism data according to the calculated SHAP value, and selects microorganisms as the key microorganisms of the Daqu based on the contribution degree.
[0017] Further, the calculation formula of the SHAP value is: where, φ j is the SHAP value of the j-th microorganism abundance data; N is the set of all microorganism abundance data; N\{j} is all subsets excluding the j-th microorganism abundance data, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only the subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model brought about by adding the j-th microorganism abundance data; |S| is the size of the subset S; |N| is the total number of microorganism abundance data.
[0018] The beneficial effects of the present invention: The present invention provides a key microorganism identification method and device in the Daqu fermentation stage. By acquiring the fermentation stage data of Daqu and the corresponding microorganism abundance data, constructing a data set, and dividing the data set into a training set, a verification set, and a test set, where the microorganism abundance data includes microorganism species and quantities, establishing a neural network model, taking the microorganism abundance data in the training set as input, and taking the fermentation stage data in the training set as output to train the neural network model, verifying the trained neural network model with the verification set to obtain a fermentation stage prediction model, testing the fermentation stage prediction model with the test set, calculating the SHAP value of each microorganism in the test set based on the SHAP theory, evaluating the contribution degree of each microorganism according to the calculated SHAP value, and selecting microorganisms as the key microorganisms of the Daqu based on the contribution degree, it solves the problem of difficult identification of key microorganisms in the existing Daqu fermentation stage. Description of the Drawings
[0019] Figure 1 It is a schematic flow chart of a method for identifying key microorganisms in the Daqu fermentation stage provided by the present invention. Detailed Embodiments
[0020] Aiming at the problem of difficult identification of key microorganisms in the existing Daqu fermentation stage, the present invention provides a method and device for identifying key microorganisms in the Daqu fermentation stage. By combining a neural network model and the SHAP algorithm, accurate analysis of the feature contribution degree of microbial abundance data is carried out, and key microbial populations with significant influence in the fermentation stage are automatically identified, which can provide a scientific basis for the fermentation process control and product quality optimization of Daqu, and promote the intelligent development of Baijiu brewing.
[0021] As Figure 1 shown, a method for identifying key microorganisms in the Daqu fermentation stage provided by the present invention includes the following steps:
[0022] S1. Obtain the fermentation stage data of Daqu and the corresponding microbial abundance data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The microbial abundance data includes the types and quantities of microorganisms;
[0023] Specifically, the types of microorganisms include Bacillus, Aspergillus, Saccharomyces cerevisiae, Pediococcus acidilactici, and Staphylococcus. The fermentation stage data includes the 1st day of fermentation, the 3rd day of fermentation, the 8th day of fermentation, the 15th day of fermentation, the 21st day of fermentation, and the 28th day of fermentation, or the fermentation stage data includes the initial fermentation stage, the middle fermentation stage, and the late fermentation stage. Usually, 70% of the data set is divided into the training set, 20% of the data set is divided into the validation set, 10% of the data set is the test set, and the data considered representative by experts is used as the test set.
[0024] S2. Establish a neural network model, use the microbial abundance data in the training set as the input, and use the fermentation stage data in the training set as the output. Train the neural network model, and use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model;
[0025] Specifically, the neural network model includes one or a combination of MLP, LSTM, RNN, CNN, GRU, and Transformer. Before training, it also includes preprocessing the fermentation stage data and the corresponding microbial abundance data. The preprocessing includes missing value processing, outlier processing, and normalization processing. The loss function of the fermentation stage prediction model is the cross-entropy loss function, and the formula of the cross-entropy loss function is: Among them, Loss is the cross-entropy error, M is the number of samples, C is the number of data in the fermentation stage, and p ij is the true j-th fermentation stage data corresponding to the i-th sample, and is the probability that the i-th sample is predicted as the j-th fermentation stage.
[0026] S3. Use the test set to test the fermentation stage prediction model, and calculate the SHAP value of each microorganism in the test set based on the SHAP theory;
[0027] Specifically, the calculation formula of the SHAP value is: where φ j is the SHAP value of the j-th microorganism abundance data; N is the set of all microorganism abundance data; N\{j} is all subsets that do not contain the j-th microorganism abundance data, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only the subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model after adding the j-th microorganism abundance data; |S| is the size of the subset S; |N| is the total number of microorganism abundance data.
[0028] S4. Evaluate the contribution degree of each microorganism according to the calculated SHAP value, and select microorganisms as the key microorganisms of the Daqu according to the contribution degree.
[0029] Specifically, the higher the SHAP value, the greater the contribution degree. Therefore, the top several microorganisms ranked from high to low by the SHAP value are used as key microorganisms, or the SHAP threshold is used as the selection condition, and the microorganisms exceeding the SHAP threshold are used as key microorganisms.
[0030] The present invention also provides a key microorganism recognition device in the Daqu fermentation stage, which realizes the key microorganism recognition method in the Daqu fermentation stage as described above. The device includes a data acquisition module, a model training module, a model verification module, a model testing module, a microorganism contribution module, and a key microorganism recognition module. The data acquisition module is used to acquire the fermentation stage data of Daqu and the corresponding microorganism abundance data, construct a data set, and divide the data set into a training set, a verification set, and a test set. The microorganism abundance data includes the types and quantities of microorganisms. The model training module is used to establish a neural network model, take the microorganism abundance data in the training set as input, and take the fermentation stage data in the training set as output to train the neural network model. The model verification module is used to verify the trained neural network model with the verification set to obtain a fermentation stage prediction model. The model testing module is used to test the fermentation stage prediction model with the test set. The microorganism contribution module is used to calculate the SHAP value of each microorganism in the test set. The key microorganism recognition module evaluates the contribution degree of each microorganism data according to the calculated SHAP value, and selects microorganisms as the key microorganisms of the Daqu based on the contribution degree.
[0031] For example, taking the fermentation stage data as the early fermentation stage, the middle fermentation stage, and the late fermentation stage, and the microorganism abundance data including 40 microorganisms such as Bacillus, Aspergillus, Saccharomyces cerevisiae, Pediococcus acidilactici, Staphylococcus, etc. and their corresponding quantities, and taking the neural network model as MLP as an example, the 40 microorganisms and their corresponding quantities corresponding to the early fermentation stage, the 40 microorganisms and their corresponding quantities corresponding to the middle fermentation stage, and the 40 microorganisms and their corresponding quantities corresponding to the late fermentation stage are obtained. Based on this, a data set is constructed. 70% of the data set is divided into a training set, 20% of the data set is divided into a verification set, and 10% of the data set is a test set. Taking the microorganism data in the training set as input and the fermentation stage data in the training set as output, the neural network model is trained, and the trained neural network model is verified with the verification set to obtain a fermentation stage prediction model. The fermentation stage prediction model is tested with the test set. The SHAP value of the microorganism data in the test set is calculated and arranged in descending order to obtain the arrangement order of the microorganism contribution degrees from high to low. The top five microorganisms are selected as the key microorganisms.
Claims
1. A method for identifying key microorganisms in the Daqu fermentation stage, characterized in that, It includes the following steps: S1. Obtain the fermentation stage data of Daqu and the corresponding microbial abundance data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The microbial abundance data includes microbial species and quantities; S2. Establish a neural network model, use the microbial abundance data in the training set as the input, and the fermentation stage data in the training set as the output, train the neural network model, and use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model; S3. Use the test set to test the fermentation stage prediction model, and calculate the SHAP value of each microorganism in the test set based on the SHAP theory; S4. According to the calculated SHAP values, evaluate the contribution degree of each microorganism, and select microorganisms as the key microorganisms of the Daqu based on the contribution degree.
2. The method for identifying key microorganisms in the Daqu fermentation stage according to claim 1, characterized in that, The microbial species include Bacillus, Aspergillus, Saccharomyces cerevisiae, Pediococcus acidilactici, and Staphylococcus.
3. The method for identifying key microorganisms in the Daqu fermentation stage according to claim 1, wherein The fermentation stage data includes the 1st day of fermentation, the 3rd day of fermentation, the 8th day of fermentation, the 15th day of fermentation, the 21st day of fermentation, and the 28th day of fermentation.
4. The method for identifying key microorganisms in the Daqu fermentation stage according to claim 1, characterized in that The fermentation stage data includes the initial fermentation stage, the middle fermentation stage, and the late fermentation stage.
5. The method for identifying key microorganisms in the Daqu fermentation stage according to claim 1, wherein The neural network model includes one or a combination of MLP, LSTM, RNN, CNN, GRU, and Transformer.
6. The method for identifying key microorganisms in the Daqu fermentation stage according to claim 1, characterized in that, The calculation formula of SHAP value is as follows: Among them, φ j is the SHAP value of the abundance data of the j-th microorganism; N is the set of all microorganism abundance data; N\{j} is all subsets that do not include the abundance data of the j-th microorganism, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model after adding the abundance data of the j-th microorganism; |S| is the size of subset S; |N| is the total number of microorganism abundance data.
7. The method for identifying key microorganisms in the Daqu fermentation stage according to claim 1, characterized in that The loss function of the fermentation stage prediction model is the cross-entropy loss function, and the formula of the cross-entropy loss function is: Among them, Loss is the cross-entropy error, M is the number of samples, C is the number of data in the fermentation stage, and p ij is the true j-th fermentation stage data corresponding to the i-th sample, is the probability that the i-th sample is predicted to be in the j-th fermentation stage.
8. The method for identifying key microorganisms in the Daqu fermentation stage according to claim 1, wherein Before training, it also includes preprocessing the fermentation stage data and the corresponding microbial abundance data. The preprocessing includes missing value processing, outlier processing, and normalization processing.
9. Key microorganism identification device in the large starter fermentation stage, characterized in that, Implement the method for identifying key microorganisms in the Daqu fermentation stage as described in claim 1. The device includes a data acquisition module, a model training module, a model verification module, a model testing module, a microbial contribution module, and a key microorganism identification module; The data acquisition module is used to obtain the fermentation stage data of Daqu and the corresponding microbial abundance data, construct a data set, and divide the data set into a training set, a validation set, and a test set. The microbial abundance data includes microbial species and quantities; the model training module is used to establish a neural network model, use the microbial abundance data in the training set as the input, and the fermentation stage data in the training set as the output, and train the neural network model; the model verification module is used to use the validation set to verify the trained neural network model to obtain a fermentation stage prediction model; the model testing module is used to use the test set to test the fermentation stage prediction model; the microbial contribution module is used to calculate the SHAP value of each microorganism in the test set; the key microorganism identification module, according to the calculated SHAP values, evaluates the contribution degree of each microbial data, and selects microorganisms as the key microorganisms of the Daqu based on the contribution degree.
10. The key microorganism identification device in the Daqu fermentation stage according to claim 9, characterized in that, The calculation formula of SHAP value is as follows: where φ j is the SHAP value of the abundance data of the j-th microorganism; N is the set of all microorganism abundance data; N\{j} is all subsets that do not contain the abundance data of the j-th microorganism, S is a subset in N\{j}; v(S) is the output value of the fermentation stage prediction model when only subset S is included, and v(S∪{j}) - v(S) represents the change in the output value of the fermentation stage prediction model after adding the abundance data of the j-th microorganism; |S| is the size of subset S; |N| is the total number of microorganism abundance data.