Construction method of model for predicting lactic acid content of fermented grains discharged from cellar of Maotai-flavor liquor and lactic acid content prediction method

By obtaining the physical and chemical indicators of the wine mash and the relative abundance data of the lactic acid-producing gene, and using machine learning methods to establish a prediction model, it solves the problem that it is difficult to monitor the lactic acid content in real time during the fermentation process in the cellar, and achieves rapid and accurate prediction of the lactic acid content, ensuring the stability of the fermentation environment and the quality of the sauce-flavored liquor.

CN120015127APending Publication Date: 2025-05-16KWEICHOW MOUTAI COMPANY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510009952.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to monitor and predict the lactic acid content of the wine mash during the fermentation process in the cellar in real time and simplicity, resulting in the inability to adjust production operations in time, affecting the final quality of the sauce-flavored liquor.

Method used

By obtaining the physical and chemical indicators of the wine mash and the relative abundance data of the lactic acid-producing genes when entering the cellar, a prediction model was established using machine learning methods to predict the lactic acid content of the wine mash after the cellar. The specific steps include obtaining data, performing machine learning model association, model screening and verification, and finally selecting the most accurate model for prediction.

Benefits of technology

It achieves rapid and accurate prediction of the lactic acid content of the cellar wine mash, avoids the damage to the anaerobic environment caused by sampling and detection during the fermentation process in the cellar, and provides real-time guidance without destroying the fermentation environment to ensure the balance of acidity in the cellar.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_3
    Figure SMS_3
  • Figure SMS_4
    Figure SMS_4
  • Figure SMS_5
    Figure SMS_5
Patent Text Reader

Abstract

The invention provides a method for constructing a model for predicting the lactic acid content of fermented grains discharged from a cellar of Maotai-flavor liquor. The method comprises the following steps: acquiring physicochemical indexes and lactic acid production gene relative abundance data of the fermented grains discharged from the cellar and the lactic acid content of the fermented grains discharged from the cellar when the fermented grains discharged from the cellar enter the cellar; taking the obtained relative abundance data of the lactic acid production genes or the relative abundance and physicochemical index data of the lactic acid production genes as independent variables, taking the lactic acid content as a dependent variable, and carrying out model association machine learning on the independent variables and the dependent variable; and respectively predicting the lactic acid content of the fermented grains discharged from the cellar based on the models obtained by machine learning, and screening the prediction models based on the accuracy of the obtained prediction values. Relative abundance of lactic acid-producing genes of fermented grains when entering the cellar is obtained, and the lactic acid content of the fermented grains in the cellar is subjected to correlation analysis by combining a machine learning model, so that the lactic acid content of the cellar can be quickly predicted, the anaerobic environment can be prevented from being damaged by sampling detection in the fermentation process in the cellar, and the yield of lactic acid in the cellar can be greatly improved under the condition of not damaging the fermentation environment. And predicting the lactic acid content of the fermented grains in the cellar.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of biological fermentation technology, and in particular to a method for predicting the lactic acid content of unsalted mash and a method for constructing a prediction model thereof. Background Art

[0002] Liquor is a typical representative of traditional Chinese fermented food. Its brewing process is complicated, especially for sauce-flavor liquor. During its annual production cycle, sauce-flavor liquor needs to go through multiple rounds of stacking fermentation, cellar fermentation and natural fermentation. Among them, the cellar fermentation cycle is long and is an anaerobic fermentation, which has a key impact on the final quality of sauce-flavor liquor.

[0003] During fermentation in the cellar, the metabolic activities of the dominant bacteria represented by lactic acid bacteria will cause the acid in the cellar to accumulate continuously, forming a high-acid brewing environment. The high-acid brewing environment can screen and enrich acid-resistant functional bacteria, inhibit the growth of harmful microorganisms, and ensure the normal progress of fermentation. At the same time, acid, as an important aroma and flavor substance, can give the sauce-flavored liquor a unique and rich aroma and taste. But on the other hand, excessive acid will also affect the metabolism of functional bacteria and lead to abnormal fermentation. Therefore, maintaining the balance of acidity in the cellar is of great significance to ensure normal fermentation and the quality of the final product, which requires monitoring the acidity of the mash during the fermentation process in the cellar.

[0004] The acids in the cellar fermentation are mainly organic acids, especially lactic acid. At present, the detection of lactic acid in the cellar fermentation process generally includes sampling and other steps. However, the sampling operation will destroy the anaerobic environment in the cellar due to sampling, which may lead to fermentation abnormalities. On the other hand, due to the cumbersome process, it is impossible to achieve the purpose of real-time acquisition and prediction of data, and it is impossible to provide timely guidance for production operations.

[0005] Therefore, a method is needed to timely and easily obtain the lactic acid content in the fermented grains after leaving the cellar to guide production operations. Summary of the invention

[0006] Based on this, one of the purposes of this application is to provide a method for timely and conveniently obtaining the lactic acid content in the fermented grains after leaving the cellar.

[0007] To achieve the above-mentioned purpose, the present application provides, on the one hand, a method for constructing a prediction model for lactic acid content of fermented grains of Maotai-flavor liquor, and the technical scheme adopted is as follows:

[0008] In some embodiments, the detection method of the present application comprises the following steps:

[0009] (1) obtaining the physical and chemical indexes of the fermented grains when they are put into the cellar and the relative abundance data of lactic acid-producing genes;

[0010] (2) obtaining the lactic acid content of the fermented grains removed from the cellar;

[0011] (3) Machine learning: taking the obtained relative abundance data of the lactic acid-producing gene or the relative abundance data of the lactic acid-producing gene and the physical and chemical index data as the independent variable, taking the obtained lactic acid content as the dependent variable, and performing machine learning on the association model between the independent variable and the dependent variable;

[0012] (4) Model screening: Based on the model obtained by machine learning in step (3), the lactic acid content of the uncellared mash is predicted respectively, and the prediction model is screened based on the accuracy of the predicted value of the lactic acid content.

[0013] In some embodiments, in step (1), in step (1), the physical and chemical indicators include acidity, reducing sugar content, moisture content and starch content.

[0014] In some embodiments, in step (3), the physicochemical indicators include at least one of acidity, reducing sugar content, moisture content and starch content; preferably, the physicochemical indicators include a combination of acidity, reducing sugar content, moisture content and starch content; preferably, the physicochemical indicators include reducing sugar content.

[0015] In some embodiments, in step (1), the lactic acid producing gene includes L-LDH gene and D-LDH gene.

[0016] In some embodiments, in step (3), the model in the machine learning of associating the independent variable with the dependent variable includes at least one of linear regression, decision tree, random forest, support vector machine, K nearest neighbor algorithm, gradient boosting and neural network.

[0017] In some embodiments, in step (4), screening the prediction model based on the accuracy of the obtained prediction value includes: comparing the obtained predicted value of lactic acid content with the true value of lactic acid content actually obtained in step (2), and based on the mean square error between the predicted value and the true value, selecting the model corresponding to the predicted value with the smallest mean square error as the prediction model.

[0018] On the other hand, the present application provides a method for predicting the lactic acid content of fermented grains out of the cellar, comprising the following steps:

[0019] (1) obtaining the physical and chemical indexes of the fermented grains to be predicted to be released from the cellar and the relative abundance data of lactic acid-producing genes when entering the cellar;

[0020] (2) The obtained relative abundance data of the lactic acid-producing genes or the relative abundance data of the lactic acid-producing genes and the physical and chemical index data are used as model input features, and input into the prediction model constructed by the construction method according to any one of claims 1 to 6, and the lactic acid content of the uncellared mash is predicted based on the output results.

[0021] On the other hand, a method for predicting the lactic acid content of fermented grains out of the cellar comprises the following steps:

[0022] (1) Obtaining the relative abundance data of lactic acid-producing genes when the uncellared fermented grains are put into the cellar and the lactic acid content corresponding to the uncellared fermented grains;

[0023] (2) using the obtained relative abundance data of the lactic acid-producing gene as an independent variable and the obtained lactic acid content corresponding to the uncellared mash as a dependent variable, and performing associative machine learning of a support vector machine model on the independent variable and the dependent variable;

[0024] (3) using the relative abundance data of lactic acid-producing genes when the fermented grains to be predicted to be released from the cellar enter the cellar as input features, inputting the support vector machine prediction model trained in step (2), and predicting the lactic acid content of the fermented grains to be predicted to be released from the cellar based on the output results;

[0025] In some embodiments, the lactate production genes include L-LDH gene and D-LDH gene.

[0026] On the other hand, a method for predicting the lactic acid content of fermented grains out of the cellar comprises the following steps:

[0027] (1) obtaining the acidity, reducing sugar content, water content, starch content and relative abundance data of lactic acid-producing genes of the uncellared fermented grains when entering the cellar, as well as the lactic acid content corresponding to the uncellared fermented grains;

[0028] (2) using the obtained acidity, reducing sugar content, water content, starch content and relative abundance data of lactic acid-producing genes as independent variables and the obtained lactic acid content corresponding to the uncellared fermented grains as the dependent variable, training a lactic acid content prediction model based on a neural network model;

[0029] (3) inputting the features of the acidity, reducing sugar content, water content, starch content and relative abundance of lactic acid-producing genes of the fermented grains to be predicted to be released from the cellar into the neural network prediction model trained in step (2), and predicting the lactic acid content of the fermented grains to be predicted to be released from the cellar based on the output results;

[0030] In some embodiments, the lactate production genes include L-LDH gene and D-LDH gene.

[0031] On the other hand, a method for predicting the lactic acid content of fermented grains out of the cellar comprises the following steps:

[0032] (1) Obtaining the relative abundance data of lactic acid-producing genes, reducing sugar content, and lactic acid content corresponding to the uncellared fermented grains when the uncellared fermented grains are put into the cellar;

[0033] (2) using the relative abundance data of the lactic acid-producing gene and the reducing sugar content data obtained as independent variables, and using the lactic acid content corresponding to the uncellared mash as the dependent variable, and using a K-nearest neighbor algorithm model to train a lactic acid content prediction model;

[0034] (3) Using the relative abundance data of lactic acid-producing genes and reducing sugar content data of the fermented grains to be predicted to be released from the cellar as input features, the K nearest neighbor algorithm prediction model trained in step (2) is input, and the lactic acid content of the fermented grains to be predicted to be released from the cellar is predicted based on the output results.

[0035] In some embodiments, the lactate production genes include L-LDH gene and D-LDH gene.

[0036] Beneficial effect: By obtaining the relative abundance of lactic acid-producing genes in the mash when it enters the cellar, and combining the machine learning model to make a correlation analysis of the lactic acid content of the mash in the cellar, the model with the highest accuracy among multiple machine learning models is selected to predict the lactic acid content of the mash out of the cellar. The method of the present application can quickly predict the lactic acid content out of the cellar, and can avoid the destruction of the anaerobic environment caused by sampling and testing during the fermentation process in the cellar, thereby realizing the prediction of the lactic acid content of the mash in the cellar without destroying the fermentation environment. DETAILED DESCRIPTION

[0037] The technical solution of the present invention is further described below by specific embodiments, which do not limit the protection scope of the present invention. Some non-essential modifications and adjustments made by others based on the concept of the present invention still fall within the protection scope of the present invention.

[0038] The "relative abundance of lactic acid-producing genes" in the embodiments of the present invention specifically refers to: the ratio of the number of specific gene sequences carried by microorganisms capable of producing lactic acid in the fermented grains to the total number of microbial gene sequences.

[0039] The "acidity" in the embodiments of the present invention specifically refers to the total amount of titratable acid (usually organic acid) in the fermented grains.

[0040] The "reducing sugar content" in the embodiments of the present invention specifically refers to: the amount of sugar substances in the fermented grains that can be reduced by a reducing agent.

[0041] The "water content" in the embodiment of the present invention specifically refers to: the proportion of the mass of water in the fermented grains to the total mass.

[0042] The "starch content" in the embodiments of the present invention specifically refers to the amount of undecomposed starch present in the fermented grains, calculated on a dry basis.

[0043] The "lactic acid content" in the embodiments of the present invention specifically refers to: the actual concentration of lactic acid in the fermented grains.

[0044] The "relative abundance of lactic acid bacteria" in the embodiments of the present invention specifically refers to: the ratio of the number of lactic acid bacteria in the fermented grains to the total number of the entire microbial community.

[0045] Example 1

[0046] This embodiment provides a method for predicting the lactic acid content of fermented grains out of a cellar, comprising the following steps:

[0047] 1. Obtain the physical and chemical indicators and relative abundance data of lactic acid-producing genes of the mash samples when they are put into the cellar. The physical and chemical indicators include acidity, reducing sugar content, water content and starch content.

[0048] Since the cellar undergoes a long period of anaerobic fermentation and cannot be managed by humans, the state of the mash at this critical point of entering the cellar directly determines whether the fermentation in the cellar can proceed normally, and affects the accumulation of lactic acid during the fermentation process in the cellar.

[0049] Sampling was carried out when the mash was put into the cellar and when the mash was taken out of the cellar, and physical and chemical indexes of the mash samples when entering the cellar were measured and metagenomic sequencing was performed on the mash samples when entering the cellar.

[0050] Among them, the physical and chemical indicators include acidity, reducing sugar content, moisture content, and starch content. The measurement method is as follows: weigh 200g of fermented grains sample, and use FOSS DS2500 F near-infrared spectrometer to scan and collect the spectral information of the fermented grains sample to obtain the acidity, sugar, moisture, starch and other physical and chemical index data of the fermented grains sample when entering the cellar.

[0051] The metagenomic sequencing process is as follows: DNA was extracted from the fermented grains using a DNA kit (Easy Nucleic Acid Isolation SoilDNA Kit) (OMEGA Bio-Tek in Georgia, USA), and the concentration and purity of the extracted DNA were tested using Qubit 3.0 (Thermo Fisher Scientific, Waltham, USA) and Thermo NanodropOne (Thermo Fisher Scientific, Waltham, USA), respectively. NEB Ultra TMDNA libraries were constructed using a DNA library preparation kit (New England Biolabs, MA, USA) and sequenced on an Illumina NovaSeq 6000 platform. The raw sequencing data were quality controlled using Trimmomatic (v.0.36) to obtain high-quality data for subsequent further analysis, and data were assembled on MEGAHIT (Version v1.0.6) (https: / / github.com / voutcn / megahit). Linclust software (http: / / www.nature.com / articles / s41467-018-04964-5) was used for gene clustering (default parameters were -e0.001--min-seq-id0.9-c 0.80) and redundancy removal to obtain a non-redundant gene set. The quality-controlled sequences were aligned to the gene library using bbmap software to obtain the number of sequences aligned for each gene and the abundance information of each gene in each sample. On this basis, the relative abundance of all genes identified as L-LDH and D-LDH was summed up, which was the relative abundance of the two lactate dehydrogenases in the mash samples when entering the cellar.

[0052] 2. Obtaining the lactic acid content of the fermented grains when they are put into the cellar and taken out of the cellar

[0053] The mash samples were taken when the mash was taken out of the cellar, and the lactic acid content was obtained by scanning with the aforementioned FOSS DS2500 F near-infrared spectrometer.

[0054] 3. Machine Learning

[0055] The relative abundance of acid-producing genes and four physical and chemical indicators of acidity, reducing sugar content, moisture content, and starch content of the mash entering the cellar were taken as independent variables, and the lactic acid content of the mash leaving the cellar was taken as the dependent variable. The independent variables and the dependent variables were associated with each other for machine learning, and a variety of machine learning methods were used for learning, including linear regression, decision tree, random forest, support vector machine, K nearest neighbor algorithm, gradient boosting, and neural network.

[0056] IV. Model Screening and Verification

[0057] The model obtained based on the above machine learning is used to predict the lactic acid content of the mash when it is out of the cellar, and the predicted value of the lactic acid content of the mash when it is out of the cellar is obtained. The predicted lactic acid content of the mash when it is out of the cellar is compared with the actual lactic acid content of the mash when it is out of the cellar, and the mean square error (MES) between the predicted value and the actual value of the model is used as an indicator to further improve the prediction accuracy of the model.

[0058] Among them, MSE first calculates the difference between the actual value and the predicted value of each sample, then squares these differences, and finally finds the average of all squared differences.

[0059] This embodiment selects the neural network model with the smallest mean square error as the prediction model, compares the predicted lactic acid content with the actual value, and performs absolute error analysis.

[0060] The mean square error results of each model measured in this embodiment are as follows:

[0061] Table 1 Accuracy of the model for predicting lactic acid content out of the cellar based on the relative abundance of lactic acid-producing genes and physical and chemical indicators at the time of entry

[0062]

[0063]

[0064] In the above table, the absence of mean square error means that no correlation between the independent variable and the dependent variable can be found after learning, indicating that the model is not suitable for predicting the lactic acid content of the mash out of the cellar (the same as in the following table).

[0065] The prediction results and errors of the prediction model are as follows:

[0066] Table 2 Lactate prediction accuracy based on neural model, lactate-producing genes, and physicochemical indicators

[0067]

[0068] The results showed that the neural network-based model, with the relative abundance of acidogenic genes and four physical and chemical indicators of acidity, reducing sugar content, moisture content and starch content as independent variables for prediction, had an error of 80% within 0.1 and an average absolute error of 0.064, indicating that the prediction method was highly accurate.

[0069] Embodiment 2:

[0070] Based on the invention content provided by the present invention, this embodiment provides a method for predicting the lactic acid content of fermented grains out of the cellar, comprising the following steps:

[0071] 1. Obtaining the relative abundance data of lactic acid-producing genes in the fermented grains sample when entering the cellar: Obtain the relative abundance data of acid-producing genes in the fermented grains when entering the cellar according to the method of Example 1.

[0072] 2. Obtaining the lactic acid content of the fermented grains when they are taken out of the cellar: the method is the same as that in Example 1.

[0073] 3. Machine learning: Take the relative abundance data of acid-producing genes as the independent variable, and the lactic acid content of the mash out of the cellar as the dependent variable, and perform machine learning by associating the independent variable with the dependent variable.

[0074] 4. Model screening and verification: After learning, it was found that the support vector machine model had the smallest error. The support vector machine model with the smallest mean square error was selected as the prediction model for prediction. The predicted lactic acid content was compared with the actual value to perform absolute error analysis.

[0075] The mean square error results of each machine learning model are as follows:

[0076] Table 3 Model prediction results based on relative abundance of lactic acid-producing genes and lactic acid content at cellar level

[0077]

[0078] The prediction results and errors of the prediction model are as follows:

[0079] Table 4 Accuracy of predicting lactic acid content based on support vector machine and lactic acid-producing genes

[0080]

[0081] The results showed that based on the support vector machine model, the relative abundance of lactic acid-producing genes in the mash when entering the cellar was used to predict the lactic acid content when leaving the cellar. The proportion of prediction errors within 0.1 was 70%, and the average absolute error was 0.11, which had good lactic acid prediction results.

[0082] Embodiment 3:

[0083] Based on the invention content provided by the present invention, this embodiment provides a method for predicting the lactic acid content of fermented grains out of the cellar, comprising the following steps:

[0084] 1. Obtaining the physical and chemical indicators and relative abundance data of lactic acid-producing genes of the fermented grains samples when entering the cellar: Obtain the relative abundance data of acid-producing genes, reducing sugar content, acidity, moisture and starch content of the fermented grains when entering the cellar according to the method of Example 1.

[0085] 2. Obtaining the lactic acid content of the fermented grains when they are taken out of the cellar: the method is the same as that in Example 1.

[0086] 3. Machine learning: Take "lactic acid producing genes and reducing sugar content" (A), "lactic acid producing genes, acidity and moisture" (B), and "lactic acid producing genes, acidity, moisture content and starch content" (C) as independent variables, A, B, and C as independent variables respectively, and the lactic acid content of the mash out of the cellar as the dependent variable, and perform machine learning on the association between the independent variables and the dependent variables.

[0087] The mean square error results of each machine learning model are as follows:

[0088] 4. Model screening and verification: After learning, the best model corresponding to each of the above independent variables is further used as a prediction model to predict the lactic acid content and compared with the actual value.

[0089] Table 5 Accuracy of the model for predicting lactic acid content out of the cellar based on the relative abundance of lactic acid-producing genes and the combination of physical and chemical indicators at the time of entry

[0090]

[0091] The prediction results and errors of the prediction model are as follows:

[0092] Table 6 Lactate prediction error values ​​based on the best model, lactate-producing genes, and physicochemical indicators

[0093]

[0094]

[0095] The results showed that based on the K nearest neighbor algorithm, with lactic acid genes and reducing sugar content as independent variables, 80% of the errors were within 0.1, and the mean absolute error was 0.092, indicating that when the K nearest neighbor algorithm was used for prediction based on the independent variable A, the scheme had a high accuracy rate.

[0096] Comparative Example 1:

[0097] This comparative example refers to the method of Example 1, and provides a method for predicting the lactic acid content of the uncellared fermented grains. The difference from Example 1 is that the "four physical and chemical indicators of acidity, reducing sugar content, moisture content, and starch content" of the fermented grains at the cellar entry node are used as independent variables to establish a correlation machine learning between the four physical and chemical indicators and the lactic acid content of the uncellared fermented grains, which is as follows:

[0098] 1. Obtaining the physical and chemical indicators of the fermented grains sample when entering the cellar: Obtain the physical and chemical indicators of the fermented grains when entering the cellar according to the method of Example 1, including four physical and chemical indicators: acidity, reducing sugar content, water content, and starch content.

[0099] 2. Obtaining the lactic acid content of the fermented grains when they are taken out of the cellar: the method is the same as that in Example 1.

[0100] 3. Machine learning: Take the four physical and chemical indicators of acidity, reducing sugar content, moisture content, and starch content as independent variables, and the lactic acid content of the mash out of the cellar as the dependent variable, and associate the independent variables with the dependent variables for machine learning.

[0101] 4. Model screening and verification: After learning, it was found that the random forest model had a higher accuracy. The above optimal random forest model was used as the prediction model to predict the lactic acid content and compared with the actual value.

[0102] The mean square error results of each machine learning model are as follows:

[0103] Table 7 Model prediction results based on physical and chemical indicators and lactic acid content out of the cellar

[0104]

[0105]

[0106] The prediction results of the prediction model are as follows:

[0107] Table 8 Accuracy of predicting lactic acid content based on random forest model and physical and chemical indicators

[0108]

[0109] The results showed that the proportion of lactic acid prediction errors within 0.1 was 30%, and the average absolute error was 0.20, indicating that the prediction effect of lactic acid content out of the cellar based only on physical and chemical indicators was not as good as the scheme containing the relative abundance data of acidogenic genes as independent variables.

[0110] Comparative Example 2:

[0111] This comparative example refers to the method of Example 2 to provide a method for predicting the lactic acid content of the uncellared mash. The difference from Example 2 is that the "relative abundance of lactic acid bacteria" of the uncellared mash is used as an independent variable to establish a correlation machine learning between the relative abundance of lactic acid bacteria and the lactic acid content of the uncellared mash, which is as follows:

[0112] 1. Obtaining the relative abundance of lactic acid bacteria in the mash sample when it was put into the cellar: The abundance of lactic acid bacteria was obtained by extracting DNA from the mash sample using a DNA kit (Easy Nucleic Acid Isolation Soil DNA Kit) (OMEGA Bio-Tek in Georgia, USA). The concentration and purity of the extracted DNA were tested by Qubit 3.0 (ThermoFisher Scientific, Waltham, USA) and Thermo Nanodrop One (Thermo Fisher Scientific, Waltham, USA), respectively. NEB Ultra TMDNA library preparation kit (New England Biolabs, MA, USA) was used to construct DNA libraries, and the constructed libraries were sequenced on the Illumina NovaSeq 6000 platform. Trimmomatic (v.0.36) was used to control the quality of the raw sequencing data to obtain high-quality data for subsequent further analysis, and data were assembled on MEGAHIT (Version v1.0.6) (https: / / github.com / voutcn / megahit), and further processed to construct non-redundant gene sets. BLASTP (Version 2.2.31+, http: / / blast.ncbi.nlm.nih.gov / Blast.cgi) software was further used to perform homology BLAST alignment of species with the NCBI-NR database (the threshold was set to e-value ≤ 0.0001), and Metaphlan software was used to analyze the species distribution and composition of each sample to obtain species abundance data. On this basis, the sum of the relative abundances of all species identified as lactic acid bacteria was calculated, which was the relative abundance of lactic acid bacteria.

[0113] 2. Obtaining the lactic acid content of the fermented grains when they are taken out of the cellar: the method is the same as that in Example 1.

[0114] 3. Machine learning: Take the relative abundance of lactic acid bacteria as the independent variable and the lactic acid content of the mash out of the cellar as the dependent variable, and perform machine learning by associating the independent variable with the dependent variable.

[0115] 4. Model screening and verification: After learning, it was found that the random forest model had a higher accuracy. The above optimal random forest model was used as the prediction model to predict the lactic acid content and compared with the actual value.

[0116] The mean square error results of each machine learning model are as follows:

[0117] Table 9 Model prediction results based on relative abundance of lactic acid bacteria and lactic acid content out of the cellar

[0118]

[0119] The comparison between the model prediction results and the actual values ​​is shown in the following table:

[0120] Table 10 Accuracy of predicting lactic acid content based on random forest model and relative abundance of lactic acid bacteria

[0121]

[0122] The results showed that based on the random forest model, the relative abundance of lactic acid bacteria was used as the independent variable to predict the lactic acid content. The proportion of model prediction errors within 0.1 was 50%, and the average absolute error was 0.32. The model accuracy was low, indicating that the effect of predicting the lactic acid content out of the cellar based on the relative abundance of lactic acid bacteria was not as good as the scheme using the relative abundance data of acid-producing genes as the independent variable.

[0123] Comparative Example 3:

[0124] This comparative example refers to the method of Example 1, and provides a method for predicting the lactic acid content of the uncellared fermented grains. The difference from Example 1 is that the "relative abundance of lactic acid bacteria and four physical and chemical indicators of acid, sugar, water, and starch" of the fermented grains at the cellar entry node are used as independent variables, and the correlation machine learning between the independent variables and the lactic acid content of the uncellared fermented grains is established, which is as follows:

[0125] 1. Obtain the physical and chemical index data and relative abundance data of lactic acid bacteria of the fermented grains sample when entering the cellar: Obtain the physical and chemical indexes of the fermented grains when entering the cellar according to the method of Example 1, including four physical and chemical indexes: acidity, reducing sugar content, moisture content, and starch content; obtain the relative abundance of lactic acid bacteria according to the method of Comparative Example 2.

[0126] 2. Obtaining the lactic acid content of the fermented grains when they are taken out of the cellar: the method is the same as that in Example 1.

[0127] 3. Machine learning: Take the "relative abundance of lactic acid bacteria and four physical and chemical indicators of acid, sugar, water and starch" as the independent variables, and the lactic acid content of the mash out of the cellar as the dependent variable, and associate the independent variables with the dependent variables for machine learning.

[0128] 4. Model screening and verification: After learning, it was found that the K nearest neighbor algorithm model had the highest accuracy. The above optimal K nearest neighbor algorithm model was used as the prediction model to predict the lactic acid content and compared with the actual value.

[0129] The mean square error results of each machine learning model are as follows:

[0130] Table 11 Model prediction results based on physical and chemical indicators and lactic acid content out of the cellar

[0131]

[0132] The comparison between the model prediction results and the actual values ​​is shown in the following table:

[0133] Table 12 Accuracy of predicting lactic acid content based on K nearest neighbor algorithm model and physical and chemical indicators

[0134]

[0135] The results showed that based on the K-nearest neighbor algorithm, the proportion of model prediction errors within 0.1 when the relative abundance of lactic acid bacteria and four physical and chemical indicators of acid, sugar, water and starch were used as independent variables was 50%, indicating that the prediction effect of lactic acid content out of the cellar based on the relative abundance of lactic acid bacteria and physical and chemical indicator data was poor.

[0136] By obtaining the relative abundance of lactic acid-producing genes in the mash when it enters the cellar, and combining it with a machine learning model to perform a correlation analysis on the lactic acid content of the mash in the cellar, the model with the highest accuracy among multiple machine learning models is selected to predict the lactic acid content of the mash out of the cellar. The method of the present application can quickly predict the lactic acid content of the mash out of the cellar, and can avoid the destruction of the anaerobic environment caused by sampling and testing during the fermentation process in the cellar, thereby realizing the prediction of the lactic acid content of the mash in the cellar without destroying the fermentation environment.

[0137] It is to be understood that the present invention is described by some embodiments, and it is known to those skilled in the art that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present invention. In addition, under the teachings of the present invention, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the scope of protection of the present invention.

Claims

1. A method for constructing a prediction model for lactic acid content in fermented grains of Maotai-flavor liquor, characterized in that: The steps include: (1) obtaining the physical and chemical indexes of the fermented grains when they are put into the cellar and the relative abundance data of lactic acid-producing genes; (2) obtaining the lactic acid content of the fermented grains removed from the cellar; (3) Machine learning: taking the obtained relative abundance data of the lactic acid-producing gene or the relative abundance data of the lactic acid-producing gene and the physical and chemical index data as the independent variable, taking the obtained lactic acid content as the dependent variable, and performing machine learning on the association model between the independent variable and the dependent variable; (4) Model screening: Based on the model obtained by machine learning in step (3), the lactic acid content of the uncellared mash is predicted respectively, and the prediction model is screened based on the accuracy of the predicted value of the lactic acid content.

2. The method according to claim 1, characterized in that In step (1), the physical and chemical indicators include acidity, reducing sugar content, water content and starch content.

3. The method according to claim 1, characterized in that In step (3), the physicochemical index includes at least one of acidity, reducing sugar content, moisture content and starch content; preferably, the physicochemical index includes a combination of acidity, reducing sugar content, moisture content and starch content; preferably, the physicochemical index includes reducing sugar content.

4. The method according to claim 1, characterized in that: In step (1), the lactic acid-producing gene includes L-LDH gene and D-LDH gene.

5. The method according to claim 1, characterized in that In step (3), the model for associating the independent variable with the dependent variable in the machine learning model includes at least one of linear regression, decision tree, random forest, support vector machine, K nearest neighbor algorithm, gradient boosting and neural network.

6. The method according to claim 1, characterized in that In step (4), screening the prediction model based on the accuracy of the obtained prediction value includes: comparing the obtained predicted value of lactic acid content with the true value of lactic acid content actually obtained in step (2), and based on the mean square error between the predicted value and the true value, selecting the model corresponding to the predicted value with the smallest mean square error as the prediction model.

7. A method for predicting the lactic acid content of fermented grains out of the cellar, characterized in that: The steps include: (1) obtaining the physical and chemical indexes of the fermented grains to be predicted to be released from the cellar and the relative abundance data of lactic acid-producing genes when entering the cellar; (2) The obtained relative abundance data of the lactic acid-producing genes or the relative abundance data of the lactic acid-producing genes and the physical and chemical index data are used as model input features, and input into the prediction model constructed by the construction method according to any one of claims 1 to 6, and the lactic acid content of the uncellared mash is predicted based on the output results.

8. A method for predicting the lactic acid content of fermented grains out of the cellar, characterized in that: The steps include: (1) Obtaining the relative abundance data of lactic acid-producing genes when the uncellared fermented grains are put into the cellar and the lactic acid content corresponding to the uncellared fermented grains; (2) using the obtained relative abundance data of the lactic acid-producing gene as an independent variable and the obtained lactic acid content corresponding to the uncellared mash as a dependent variable, and performing associative machine learning of a support vector machine model on the independent variable and the dependent variable; (3) using the relative abundance data of lactic acid-producing genes when the fermented grains to be predicted to be released from the cellar enter the cellar as input features, inputting the support vector machine prediction model trained in step (2), and predicting the lactic acid content of the fermented grains to be predicted to be released from the cellar based on the output results; Preferably, the lactic acid-producing genes include L-LDH gene and D-LDH gene.

9. A method for predicting the lactic acid content of fermented grains out of the cellar, characterized in that: The steps include: (1) obtaining the acidity, reducing sugar content, water content, starch content and relative abundance data of lactic acid-producing genes of the uncellared fermented grains when entering the cellar, as well as the lactic acid content corresponding to the uncellared fermented grains; (2) using the obtained acidity, reducing sugar content, water content, starch content and relative abundance data of lactic acid-producing genes as independent variables and the obtained lactic acid content corresponding to the uncellared fermented grains as the dependent variable, training a lactic acid content prediction model based on a neural network model; (3) inputting the features of the acidity, reducing sugar content, water content, starch content and relative abundance of lactic acid-producing genes of the fermented grains to be predicted to be released from the cellar into the neural network prediction model trained in step (2), and predicting the lactic acid content of the fermented grains to be predicted to be released from the cellar based on the output results; Preferably, the lactic acid-producing genes include L-LDH gene and D-LDH gene.

10. A method for predicting the lactic acid content of fermented grains out of a cellar, comprising the following steps: (1) Obtaining the relative abundance data of lactic acid-producing genes, reducing sugar content, and lactic acid content corresponding to the uncellared fermented grains when the uncellared fermented grains are put into the cellar; (2) using the relative abundance data of the lactic acid-producing gene and the reducing sugar content data obtained as independent variables, and using the lactic acid content corresponding to the uncellared mash as the dependent variable, and using a K-nearest neighbor algorithm model to train a lactic acid content prediction model; (3) using the relative abundance data of the lactic acid-producing genes and the reducing sugar content data of the fermented grains to be predicted to be released from the cellar when entering the cellar as input features, inputting the K nearest neighbor algorithm prediction model trained in step (2), and predicting the lactic acid content of the fermented grains to be predicted to be released from the cellar based on the output results; Preferably, the lactic acid-producing genes include L-LDH gene and D-LDH gene.

Citation Information

Patent Citations

  • Solid-state fermented vinegar fermentation process monitoring method based on microbial community acid production index

    CN105349690A

  • Method for using physicochemical indexes to judge structure of stacked fermented grain microorganism community

    CN108038350A

  • Method for predicting stacking fermentation process by constructing QDA model

    CN116468157A

  • Construction method of Maotai-flavor liquor yield prediction model and liquor yield prediction method

    CN119132434A