Screening method of key microorganisms in Maotai-flavor liquor fermentation process
By constructing a fermentation triplet and a Lasso linear model, key microorganisms in the fermentation process of Maotai-flavor liquor were screened, which solved the problem of insufficient understanding of microorganisms in the fermentation process of Maotai-flavor liquor in the existing technology, optimized the fermentation process, and improved the fermentation effect.
Patent Information
- Application Number
- CN202510889240.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies lack scientific understanding of the key microorganisms in the fermentation process of sauce-flavor liquor, which makes it difficult to optimize the fermentation process and unable to select the optimal fermentation microorganism combination.
A fermentation triad was constructed using metagenomic sequencing technology. Key microorganisms in the fermentation process of Maotai-flavor liquor were screened using a Lasso linear model. The fermentation process was defined as fermented mash before the addition of Daqu (fermentation starter), fermented mash after the addition of Daqu, and fermented mash after the addition of Daqu. A Lasso linear model was constructed to predict and screen potential driving microorganisms.
This study concretizes the microbial implantation process during the fermentation of Maotai-flavor liquor, accurately identifies key microorganisms, optimizes the fermentation process, and improves the fermentation effect.
Smart Images

Figure CN120796068A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of liquor brewing and microbial screening, and relates to a method for screening key microorganisms in the fermentation process of Maotai-flavor liquor. BACKGROUND
[0002] Maotai-flavor liquor is produced by a multi-stage solid fermentation process using highland millet as raw material, and is obtained after steaming, spreading, adding starter, stacking fermentation, in-pit fermentation and high-temperature distillation. Maotai is a typical representative of Maotai-flavor liquor, and its fermentation process specifically includes nine stages (down sand, sand making, 1-7 rounds of fermentation). Fermented grains (also known as fermented grains) are artificially added in the first two fermentation stages (down sand, sand making), and the starter (also known as Daqu) is added from down sand to the 6th round to initiate fermentation of the liquor.
[0003] The microorganisms in the Daqu of Maotai-flavor liquor are rich in species, can produce enzymes, metabolites to promote fermentation, and unique flavor substances. With the development of next-generation sequencing, amplicon sequencing and shotgun metagenomics can enable us to explore the composition and structure of the microorganisms in Daqu in more detail, and to understand the dynamic changes of the microbial community in the fermentation system during fermentation. However, the traditional fermentation process often selects the starter based on empirical knowledge, lacks scientific understanding of the microbial community in Daqu, and current scientific research is limited to characterizing the microbial species and abundance in the fermentation substrate and Daqu, without specifically exploring the species and functional role of key microorganisms during fermentation, which is not conducive to selecting the best combination of fermentation microorganisms and optimizing the fermentation process.
[0004] Therefore, the present application aims to provide a method for screening key microorganisms in the fermentation process of Maotai-flavor liquor based on metagenomic sequencing technology. SUMMARY
[0005] The present application aims to provide a method for screening key microorganisms in the fermentation process of Maotai-flavor liquor based on metagenomic sequencing technology.
[0006] The present application provides a method for screening key microorganisms in the fermentation process of Maotai-flavor liquor, which comprises the following steps:
[0007] (1) Constructing a fermentation triad: taking the first fermented grains sample before adding Daqu in the current fermentation stage, the Daqu sample added in the current fermentation process, and the second fermented grains sample after fermentation with the Daqu added in the current fermentation stage as a fermentation triad of one fermentation stage;
[0008] (2) taking the microbial abundance information in the first fermented grains sample, the Daqu sample and the second fermented grains sample in the fermentation triad corresponding to the corresponding fermentation stage as input features, respectively, to construct a corresponding Lasso linear model;
[0009] (3) predicting the potential driving microbial species in the fermentation triad in the corresponding fermentation stage based on the coefficients of the constructed Lasso linear model, and screening the key microorganisms in the fermentation process of Maotai-flavor liquor based on the prediction results.
[0010] Specifically, the first fermented grains sample is the fermented grains sample before the addition of Daqu in the current fermentation stage. In the specific sampling process, the first fermented grains sample corresponds to the pre-fermented grains (pre-FG) at the last fermentation time point in the last fermentation stage. The Daqu sample is the Daqu or starter added in the current fermentation stage. The second fermented grains sample is the fermented grains after the addition of Daqu in the current fermentation stage. In the specific sampling process, the second fermented grains sample corresponds to the post-fermented grains (post-FG) at any fermentation time node except the last time point in the current fermentation stage.
[0011] As described in the background section, for the fermentation process of Maotai-flavor liquor, the current research or technology is only limited to characterizing the microbial species and abundance in the fermentation substrate and Daqu. There is no related report in the prior art on how to characterize or evaluate the specific effects of the added Daqu or starter on the entire fermentation process, the variety and functional role of the key microorganisms in the specific fermentation process, or the specific effects of the added Daqu or starter on the entire fermentation process. Based on this, in order to realize whether the addition of Daqu or starter in the fermentation process of Maotai-flavor liquor has a beneficial effect on the fermentation process or whether the Daqu / starter is effectively implanted in the fermentation process, and to screen and evaluate the key microorganisms in the fermentation process of Maotai-flavor liquor, the present application first proposes the concept of fermentation triad, which defines each fermentation process of Maotai-flavor liquor in the form of fermentation triad, i.e. fermented grains before Daqu addition, Daqu and fermented grains after Daqu addition, so that the implantation process of microorganisms in the fermentation process of Maotai-flavor liquor is more specific. At the same time, based on the fermentation triad defined in the present application, the screening of key microorganisms / key driving microorganisms in the fermentation process of Maotai-flavor liquor is accurately and effectively realized by constructing a Lasso linear model.
[0012] In some embodiments of the present application, in step (2), the microorganism abundance information in the first fermented grains sample, the Daqu sample and the second fermented grains sample in the corresponding fermentation triad of the respective fermentation stage are taken as input features to construct the corresponding Lasso linear model, which includes:
[0013] The microorganism abundance information in the first fermented grains sample, the Daqu sample and the second fermented grains sample in the fermentation triad is taken as input features to construct the corresponding Lasso linear model, and the mean square error, the coefficient, and the sensitivity-specificity curve of the Lasso linear model constructed based on the microorganism abundance information in the first fermented grains sample, the Daqu sample and the second fermented grains sample in the fermentation triad as input features are obtained, respectively.
[0014] The complexity and sparsity of the corresponding Lasso linear model are adjusted by adjusting the hyperparameter λ value of the Lasso linear model, so that the mean square error of the corresponding Lasso linear model is minimized and the prediction accuracy is maximized. In this process, five-fold cross-validation is used. When the mean square error of the Lasso linear model is minimized, the corresponding hyperparameter λ value is the optimal λ value, and the Lasso linear model corresponding to the optimal λ value is the constructed Lasso linear model.
[0015] In some embodiments of the present application, in step (3), the driving microorganism species in the fermentation triad of the corresponding fermentation stage are predicted based on the coefficients of the constructed Lasso linear model, which includes:
[0016] Based on the constructed Lasso linear model, the coefficients corresponding to different microorganism species in the first fermented grains sample, the Daqu sample and the second fermented grains sample of the fermentation triad are obtained. The coefficients represent the contribution of the microorganism species in different samples to the prediction ability of the Lasso linear model. The microorganism species with absolute values of the coefficients in the range of 0-3 are selected as potential driving microorganism species in the fermentation triad.
[0017] In some embodiments of the present application, in step (3), the key microorganisms in the fermentation process of Maotai-flavor liquor are screened based on the prediction results, which includes:
[0018] The microorganism species with the abundance of the potential driving microorganism species in the fermentation triad in the range of 0.01% to 6.5% are selected as the key microorganisms in the fermentation process of Maotai-flavor liquor.
[0019] In some embodiments of the present application, in step (3), the key microorganisms in the fermentation process of Maotai-flavor liquor are screened based on the prediction results, which further includes:
[0020] The first round, the sand making, and the first round are defined as the early fermentation stage, and the second round, the third round, the fourth round, and the fifth round are defined as the late fermentation stage, based on the obtained potential driving microbial species in the fermentation triad of each fermentation stage, the average relative abundance of the potential driving microbial species in the early fermentation stage and the late fermentation stage is calculated respectively, the average relative abundance is compared, and the microbial species with obvious changes in average relative abundance in the early fermentation stage and the late fermentation stage are screened out, that is, the key microorganism in the fermentation process of Maotai-flavor liquor.
[0021] In some embodiments of the application, the microbial species with obvious changes in average relative abundance in the early fermentation stage and the late fermentation stage are the microbial species corresponding to a significance level of P<0.05 between the average relative abundance changes in the early fermentation stage and the late fermentation stage.
[0022] In some embodiments of the application, the screening method further comprises evaluating the prediction effectiveness of the Lasso linear model, comprising the following steps:
[0023] The acid concentration of the fermented grains in all fermentation triads of a certain fermentation stage is measured, and the average value of the acid concentration of the fermented grains in the corresponding fermentation stage is calculated, the acid concentration of the fermented grains in the fermentation triad is compared with the average acid concentration of the fermented grains in the corresponding fermentation stage, and the implantation of microorganisms in the Daqu sample in the corresponding fermentation stage is evaluated whether effective for the fermentation process; when the acid concentration of the fermented grains in the fermentation triad is lower than the average acid concentration of the fermented grains in the corresponding fermentation stage, the colonization process of microorganisms in the Daqu sample is defined as effective, otherwise it is ineffective.
[0024] In some embodiments of the application, the measurement method of the acid concentration is: the content of total acid is determined by titration with 0.05-0.2 mol / L NaOH, the titration end point is pH=8.2, and the acid content is calculated based on the concentration of the titrated NaOH solution, which is the acid concentration in the fermented grains.
[0025] In some embodiments of the application, the fermentation process comprises at least one fermentation stage in the sand making, the sand making, the first round, the second round, the third round, the fourth round, and the fifth round of the fermentation process of Maotai-flavor liquor.
[0026] In some embodiments of the application, the collection of the fermented grains sample comprises the following steps: simultaneously selecting the center point of the stacked fermented grains, the surface point with a depth of about 30-50 cm from the surface, the midpoint of the center point and the surface, and the upper and lower two points on the vertical line of the center point with a distance of 50-60 cm from the center point for sampling.
[0027] In some embodiments of the present application, the collecting of the fermented grains sample further comprises: sampling in 8 equal sub-regions in the same height circle corresponding to the sampling point, respectively.
[0028] In some embodiments of the present application, the screening method according to claim 1, wherein the microbial abundance information comprises: a relative abundance table of the microbial species in the first fermented grains sample, the Daqu sample and the second fermented grains sample, and the abundance table contains the classification of the microorganisms in the samples at the species level and the corresponding relative abundance.
[0029] In some embodiments of the present application, the obtaining of the microbial abundance information comprises:
[0030] extracting the total DNA in the first fermented grains sample, the Daqu sample and the second fermented grains sample; detecting the integrity and purity of the extracted DNA for detection analysis;
[0031] constructing a sequencing library, sequencing;
[0032] performing quality assessment on the original sequencing data, and filtering out low-quality reads;
[0033] annotating the obtained high-quality sequences to obtain the types of microorganisms contained in the corresponding samples, and counting the sequence numbers of each type of microorganism, and then calculating the relative abundance of each type of microorganism.
[0034] Compared with the prior art, the present application has the following beneficial effects:
[0035] Firstly, in order to realize the screening and evaluation of key microorganisms in the fermentation process of Jiang-flavor liquor, the present application first proposes the concept of fermentation triad, which defines each fermentation process of Jiang-flavor liquor in the form of fermentation triad, i.e. fermented grains before Daqu addition, Daqu and fermented grains after Daqu addition, so that the implantation process of microorganisms in the fermentation process of Jiang-flavor liquor is more specific. At the same time, based on the fermentation triad defined in the present application, the screening of key microorganisms / key driving microorganisms in the fermentation process of Jiang-flavor liquor is accurately and effectively realized by constructing a Lasso linear model. BRIEF DESCRIPTION OF DRAWINGS
[0036] Figure 1 is a sampling point schematic diagram for the fermented grains sample in the embodiments of the present application; wherein, Figure 1 (A) is a sampling point schematic diagram for each fermented grains sample, Figure 1 (B) is a schematic diagram of 8 sample points for sampling in 8 equal sub-regions in the same height circle corresponding to each sampling point in order to ensure the uniformity and representativeness of the samples;
[0037] Figure 2 The mean squared error, coefficients and sensitivity-specificity curve results of Lasso linear models corresponding to different input features in the embodiments of the present application; wherein, Figure 2 (a) represents the mean squared error of the Lasso linear model constructed by taking the microbial community results of the starter sample as the input feature, Figure 2 (b) represents the coefficients of the Lasso linear model constructed by taking the microbial community results of the starter sample as the input feature, Figure 2 (c) represents the sensitivity-specificity curve of the Lasso linear model constructed by taking the microbial community results of the starter sample as the input feature, Figure 2 (d) represents the mean squared error of the Lasso linear model constructed by taking the microbial community results of the pre-FG sample as the input feature, Figure 2 (e) represents the coefficients of the Lasso linear model constructed by taking the microbial community results of the pre-FG sample as the input feature, Figure 2 (f) represents the sensitivity-specificity curve of the Lasso linear model constructed by taking the microbial community results of the pre-FG sample as the input feature, Figure 2 (g) represents the mean squared error of the Lasso linear model constructed by taking the microbial community results of the post-FG sample as the input feature, Figure 2 (h) represents the coefficients of the Lasso linear model constructed by taking the microbial community results of the post-FG sample as the input feature, Figure 2 (i) represents the sensitivity-specificity curve of the Lasso linear model constructed by taking the microbial community results of the post-FG sample as the input feature.
[0038] Figure 3 The prediction results of the microbial community range in a single fermentation process based on the constructed Lasso linear model in the embodiments of the present application; wherein, only the most relevant predictors are included in each classification (species of starter, pre-FG and post-FG), the prediction effectiveness is measured by the effectiveness of predicting the fermentation process, the prediction performance of each result index and the variable importance and directionality of most predictors are taken as the Lasso coefficients of cross-validation, the seven key species with the highest abundance selected are marked in red, and are defined as the driving species of the fermentation process;
[0039] Figure 4 The relative abundance distribution results of the potential driving species selected in the embodiments of the present application in the entire fermentation process; wherein Figure 4(a) The stacked column chart shows the relative abundance of all driving species at each sample collection time point in each fermentation stage, the relative abundance of other species is marked as other species, and the driving species with significantly higher relative abundance is marked in red; Figure 4 (b) For the comparison of the relative abundance of the main driving species in the front-part (i.e. the first round) and the later-part (i.e. the second to fifth rounds) of fermentation, Wilcoxon test: ****, P<0.05. DETAILED DESCRIPTION
[0040] The technical solutions of the present application are further illustrated below by specific examples, which do not represent a limitation on the scope of protection of the present application. Some non-essential modifications and adjustments made by others according to the concept of the present application still fall within the scope of protection of the present application.
[0041] Example 1
[0042] Based on the summary of the present application, this example provides a specific embodiment of a method for screening key microorganisms in the fermentation process of Jiangxiang Baijiu (liquor) based on metagenomic sequencing technology, which comprises the following steps:
[0043] I. Data collection and processing
[0044] (1) Data collection: Daqu (also known as starter) samples and fermented grains (also known as FG) samples of a certain sauce-flavor liquor brewery were collected. The sample information collected in this embodiment is as follows: there are 89 Daqu samples, including 63 Daqu samples from the koji-making end and 26 Daqu samples from the wine-making end; there are 308 wine-grain samples, including 11 wine-grain samples from the Xiasha round, 49 wine-grain samples from the Zaosha round, 61 wine-grain samples from the first round (Round 1), 50 wine-grain samples from the second round (Round 2), 56 wine-grain samples from the third round (Round 3), 52 wine-grain samples from the fourth round (Round 4), and 29 wine-grain samples from the fifth round (Round 5). The Daqu samples all come from the wine production line. On the wine production line, the fermentation of Daqu and the fermentation of wine mash are not in the same workshop. After the Daqu is fermented and formed at the Stater making stage, it is put into the Fermentation stages for fermentation to play a role. The Daqu samples collected in this embodiment include both Daqu samples at the Stater making stage (called the fermentation agent at the Stater making stage) and Daqu samples at the Fermentation stage (called the fermentation agent at the Fermentation stage). Corresponding to the Daqu samples at the Fermentation stage are wine mash samples from seven stages of fermentation: Xiasha, Zaosha, and Round 1-5. The wine mash samples collected in this embodiment include wine mash samples from seven fermentation stages: Xiasha, Zaosha, and Round 1-5.
[0045] (2) Collection of mash samples: Figure 1 As shown in (A), the center point A of the piled fermented mash and the surface point B at a depth of about 30-50 cm from the surface are selected. At the same time, the midpoint C between points A and B and points D and E at 50-60 cm from point A on the same vertical line above and below point A are selected for sampling. Figure 1 As shown in (B), in order to ensure the uniformity and representativeness of the samples, 8 sampling points were collected within the same height circle. 500g of mash samples were taken from each sampling point (point A, point B, point C, point D, and point E), and the mash samples taken from the above collection points were mixed evenly to form a pile of fermented grain samples; the sampling time point was once every other day, starting from the pile fermentation, and an average of 10 samples were taken at each stage; during the sampling process, the temperature of the sampling points was measured at the same time. The samples collected in the above manner were stored in low-temperature liquid nitrogen at a temperature of -80°C.
[0046] (3) Sample processing: The collected Jiucai samples and Daqu samples were respectively used commercial kits (Guangdong Magigene Biotechnology Co., Ltd, China) to extract total DNA according to the provided manufacturer's guidelines; the integrity and purity of DNA were monitored on 1% agarose gel, and the DNA concentration and purity were detected using Qubit 3.0 (Thermo Fisher Scientific, USA) and Nanodrop One (Thermo Fisher Scientific, USA); sequencing library was prepared using (New England Biolabs, USA) and Ultra TM DNA library preparation kit; library quality evaluation was performed using Qubit 4.0 fluorometer and Qsep400 high-throughput nucleic acid protein analysis system; then, the library was sequenced on the Illumina NovaSeq 6000 platform to generate 150 bp double-end sequencing reads; after filtering out samples without corresponding metabolites, 308 Jiucai samples were used for bioinformatics analysis. Among them, the raw sequencing data was first evaluated using FastQC (v0.11.6), and then Trimmomatic (v0.38) was used to trim low-quality reads and adapters to eliminate reads with a length of less than 150 bp, adapters, leading or trailing bases with a Phred base quality (BQ) score of <20, and strings of five bases with an average BQ score of <25.
[0047] (4) Sample relative abundance information acquisition: In order to determine the taxonomic composition of the microbiome in each Jiucai sample and Daqu sample, MetaPhlAn2 (v2.6.0) was used to annotate high-quality reads using default settings to obtain the types of microorganisms contained in the corresponding samples. If the classification unit has been classified at the species level, but has not been classified at the strain level, “_unclassified” is added after the classification unit name, such as: Actinopolyspora_unclassified, and the number of sequences corresponding to each microorganism was counted, and the relative abundance of each type of microorganism was calculated.
[0048] II. Model construction
[0049] (1) Fermentation triad:
[0050] In this embodiment, the present application defines the implantation process of microorganisms in the Daqu sample in the fermentation process by defining a fermentation triad, specifically: taking the first fermented grains sample, the Daqu sample and the second fermented grains sample as a fermentation triad, and through the definition of the fermentation triad, the fermentation system (fermentation process) is better modeled to evaluate the fermentation process; wherein the first fermented grains sample refers to the pre-fermented grains (pre-FG) corresponding to the last fermentation time point of the previous fermentation stage, that is, the fermented grains sample before the addition of Daqu in the current fermentation stage, the Daqu sample refers to the starter added in the current fermentation stage, and the second fermented grains sample refers to the post-fermented grains (post-FG) corresponding to any one of the fermentation time nodes except the last time point in the current fermentation stage, that is, the fermented grains sample after the addition of Daqu in the current fermentation stage for fermentation;
[0051] (2) Construction of Lasso prediction model:
[0052] The Lasso linear model is a linear model technique for feature selection and regression analysis, which sparsifies the model coefficients by reducing some coefficients to zero, thereby realizing feature selection and model simplification. The Lasso algorithm controls the complexity and sparsity of the model by adjusting the hyperparameter λ, a larger λ value will cause more coefficients to be reduced to zero, thereby being more sparse, while a smaller λ value will cause fewer coefficients to be reduced to zero, and the model is more complex. In this embodiment, the construction steps of the Lasso prediction model are as follows:
[0053] 1) Based on the abundance information of microorganisms in the mash and daqu samples obtained in step 1, the fermentation triplet constructed above is used to respectively use the abundance information of microorganisms in the first fermentation mash sample (pre-FG), the daqu sample (starter), and the second fermentation mash sample (post-FG) in the fermentation triplet as input features to construct a corresponding Lasso linear model, and obtain the mean square error, coefficient, and sensitivity-specificity curve corresponding to the Lasso linear model constructed based on the abundance information of microorganisms in the first fermentation mash sample (pre-FG), the daqu sample (starter), and the second fermentation mash sample (post-FG) in the fermentation triplet as input features; wherein the mean square error is an indicator to measure the difference between the predicted value and the true value. For the Lasso linear model, it is the mean of the square of the difference between the predicted value and the true value of each data point; in the Lasso linear model, the coefficient refers to the weight corresponding to each feature (independent variable), which indicates the degree of influence of the feature on the target variable (dependent variable); the sensitivity-specificity curve (Receiver Operating Characteristic The ROC curve is an evaluation tool for classification problems. It evaluates model performance by plotting the relationship between sensitivity (true positive rate) and specificity (false positive rate) at different classification thresholds. The ROC curve can intuitively display the classification performance of the model at different thresholds. The closer the curve is to the upper left corner, the better the model performance, indicating that the model can achieve a higher true positive rate at a lower false positive rate, that is, it has a better ability to distinguish between positive and negative samples.
[0054] 2) Based on the obtained parameters such as the mean square error, coefficient, sensitivity-specificity curve of the Lasso linear model, the model complexity and sparsity are adjusted by adjusting the hyperparameter λ value of the Lasso linear model to minimize the mean square error of the Lasso linear model and maximize the prediction accuracy. A five-fold cross-validation is used in this process. When the error rate of the Lasso linear model is minimized, the corresponding hyperparameter λ value is the optimal λ value. Figure 2 As shown, in this embodiment, the accuracy of modeling using the Daqu sample data (using the receiver operating curve, Area Under Curve, AUC) is 50%, and the error of the model is 0.48; the accuracy of modeling using the pre_FG sample data set is 70%, and the error of the model is 0.245; the accuracy of modeling using the post_FG sample data set is 71.57%, and the error of the model is 0.222;
[0055] 3) Based on the optimal lambda value determined in step 2), the coefficients of the microbial species corresponding to the Lasso linear model constructed based on the microbial abundance information in the first fermented Jiuqu sample (pre-FG), the starter sample (starter), and the second fermented Jiuqu sample (post-FG) in the fermentation triad were obtained, respectively. The coefficients of different microbial species in the Lasso linear model represent the contribution of microbial species in different samples (pre-FG, starter, post-FG) to the prediction ability of the model. The absolute value of the coefficient in the range of 0-3 was used as the screening standard to screen out potential driving species in the fermentation triad. The results are shown in Figure 3 . Figure 3 To use the Lasso linear model with cross-validation to predict the range of microbial community in a single fermentation process; wherein only the most relevant predictors are included in each classification (species in starter, pre-FG and post-FG samples), and the prediction effectiveness is measured by the effectiveness of predicting the fermentation process. The prediction performance of each result indicator and the variable importance and direction of most predictors are used as the cross-validation Lasso coefficients. Based on the results shown in Figure 3 , a total of 53 potential driving species in the fermentation process of Jiuqu were screened out based on the Lasso linear model constructed above. Further, among the potential driving species predicted based on the Lasso linear model, the 7 species with the highest abundance (abundance range: 0.01% to 6.5%) were defined as "main driving species", which specifically include Bacillus amyloliquefaciens, Bacillus sonorensis and Staphylococcus xylosus in the starter sample (starter), and Acetobacter unclassified, Bacillus licheniformis, LactoBacillus brevis and Pediococcus lolii in the second fermented Jiuqu sample (post-FG). Figure 3 .
[0056] III. Screening of key microorganisms
[0057] The specific steps are as follows: Xiaocha, Zaocha, and one round are defined as the early fermentation stage, and two rounds, three rounds, four rounds, and five rounds are defined as the late fermentation stage. Based on the potential driving species in the fermentation triad obtained in step two, the average relative abundance of the potential driving species in the fermentation triad in the early fermentation stage and the late fermentation stage is calculated. The abundance changes of the key microorganisms screened by the Lasso prediction model constructed in the embodiment of the present application in the early fermentation stage and the late fermentation stage are compared, and the microbial species with obvious relative abundance changes (P<0.05) are screened out. The corresponding microbial species are the fermentation driving species, that is, the key microorganisms in the fermentation process. The results are shown in Figure 4 According to the results shown in Figure 4 According to the results shown in
[0058] IV. Model prediction effectiveness evaluation
[0059] At the same time, we further evaluated the results of the prediction based on the LASSO linear model, as follows:
[0060] The taste of liquor is affected by acidity, and the quality of liquor can be measured by the concentration of acidity. The colonization of microorganisms in Daqu during the fermentation process can promote the fermentation process, so the acid concentration in the fermented grains can be used to evaluate whether the microorganisms in Daqu are effective for the fermentation process, that is, whether the microorganisms in Daqu are effectively implanted in the fermentation process.
[0061] In this embodiment, the acid concentration in fermented grains is used to evaluate whether the microorganisms in the Daqu sample are effective for the fermentation process, i.e., whether the microorganisms in the Daqu sample are effective for the implantation of the fermentation process. If the acid concentration in the fermented grains is still high after the Daqu sample is added to the fermentation process, i.e., the acid concentration in the fermented grains is higher than the average acid concentration in the fermented grains at the corresponding fermentation stage, it indicates that the implantation of the microorganisms in the Daqu sample is not good. In this case, the implantation process of the microorganisms in the Daqu sample is defined as invalid. If the acid concentration in the fermented grains is lower than the average acid concentration in the fermented grains at the corresponding fermentation stage after the Daqu sample is added to the fermentation process, it indicates that the implantation of the microorganisms in the Daqu sample is good. In this case, the implantation process of the microorganisms in the Daqu sample is defined as valid.
[0062] The specific evaluation steps are as follows: Since the acid concentration is related to the fermentation stage, the acid concentration in the fermented grains of all fermentation triplets at a certain fermentation stage is measured regularly, and the average value of the acid concentration in the fermented grains at the corresponding fermentation stage is calculated. The acid concentration in the fermented grains of the fermentation triplet is compared with the average value of the acid concentration in the fermented grains at the corresponding fermentation stage to evaluate whether the microorganisms in the Daqu sample at the corresponding fermentation stage are effective for the fermentation process. When the acid concentration in the fermented grains at the corresponding fermentation stage is lower than the average acid concentration in the fermented grains at the corresponding fermentation stage, the implantation process of the microorganisms in the Daqu sample is defined as valid, otherwise it is invalid. The measurement method of the acid concentration is as follows: the total acid content is determined by titration with 0.1 mol / L NaOH, the titration end point is pH=8.2, and the acid concentration in the fermented grains is calculated based on the concentration of the titrated NaOH solution.
[0063] Based on the above method, the fitting effect of the LASSO linear model is evaluated. The results are as follows: the correlation coefficient R corresponding to the modeling using the Daqu sample data set is 0.470, the correlation coefficient R corresponding to the modeling using the pre_FG sample data set is 0.548, and the correlation coefficient R corresponding to the modeling using the post_FG sample data set is 0.587. That is, it further indicates that the LASSO linear model constructed by the present application is effective for predicting the driving microorganisms in the fermentation process of the fermented grains.
[0064] It can be understood that the present application is described by some embodiments, and those skilled in the art know that various changes or equivalent replacements can be made to these features and embodiments without departing from the spirit and scope of the present application. In addition, these features and embodiments can be modified to adapt to specific conditions and materials under the guidance of the present application without departing from the spirit and scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application are within the scope of protection of the present application.
Claims
1. A method for screening key microorganisms in the fermentation process of Maotai-flavor liquor, characterized in that: The screening method comprises the following steps: (1) Constructing a fermentation triplet: taking the first fermentation mash sample before adding Daqu in the current fermentation stage, the Daqu sample added during the current fermentation process, and the second fermentation mash sample after adding Daqu in the current fermentation stage as a fermentation triplet for one fermentation stage; (2) using the microbial abundance information of the first fermentation mash sample, the Daqu sample, and the second fermentation mash sample in the fermentation triplets corresponding to the corresponding fermentation stages as input features, respectively, to construct a corresponding Lasso linear model; (3) Based on the coefficients of the constructed Lasso linear model, the potential driving microbial species in the fermentation triad of the corresponding fermentation stage were predicted, and the key microorganisms in the fermentation process of sauce-flavor liquor were screened based on the prediction results.
2. The screening method according to claim 1, wherein In step (2), the microbial abundance information in the first fermentation mash sample, the Daqu sample, and the second fermentation mash sample in the fermentation triplet corresponding to the corresponding fermentation stage is used as input features to construct the corresponding Lasso linear model, including: The microbial abundance information of the first fermentation mash sample, Daqu sample and the second fermentation mash sample in the fermentation triplicate was used as input features to construct the corresponding Lasso linear model. The mean square error, coefficient and sensitivity-specificity curve corresponding to the Lasso linear model constructed based on the microbial abundance information of the first fermentation mash sample, Daqu sample and the second fermentation mash sample in the fermentation triplicate were obtained respectively. By adjusting the hyperparameter λ value of the Lasso linear model, the complexity and sparsity of the corresponding Lasso linear model are adjusted so that the mean square error of the corresponding Lasso linear model is minimized and the prediction accuracy is maximized. A five-fold cross-validation is used in this process. When the mean square error of the Lasso linear model is minimized, the corresponding hyperparameter λ value is the optimal λ value, and the Lasso linear model corresponding to the optimal λ value is the constructed Lasso linear model.
3. The screening method according to claim 1, wherein In step (3), the prediction of the driving microbial species in the fermentation triad of the corresponding fermentation stage based on the coefficients corresponding to the constructed Lasso linear model includes: Based on the constructed Lasso linear model, the coefficients corresponding to different microbial species in the first fermentation mash sample, the Daqu sample, and the second fermentation mash sample of the fermentation triad were obtained. The coefficients represented the contribution of the microbial species in different samples to the predictive ability of the Lasso linear model, and the microbial species with the absolute value of the coefficient in the range of 0 to 3 were selected as potential driving microbial species in the fermentation triad.
4. The screening method according to claim 1, wherein In step (3), the screening of key microorganisms in the fermentation process of Maotai-flavor liquor based on the prediction results includes: Microbial species with an abundance of potential driving microbial species in the fermentation triad within the range of 0.01% to 6.5% are selected as key microorganisms in the fermentation process of sauce-flavor liquor.
5. The screening method according to claim 1, wherein In step (3), the screening of key microorganisms in the fermentation process of Maotai-flavor liquor based on the prediction results further includes: The sand-making, sand-making and first round are defined as the early fermentation stage, and the second, third, fourth and fifth rounds are defined as the late fermentation stage. Based on the potential driving microbial species in the fermentation triplets obtained in each fermentation stage, the average relative abundance of the potential driving microbial species in the early fermentation stage and the late fermentation stage is calculated respectively, and the average relative abundance is compared to screen out the microbial species with obvious changes in the average relative abundance in the early fermentation stage and the late fermentation stage, which are the key microorganisms in the fermentation process of sauce-flavor liquor; preferably, the microbial species with obvious changes in the average relative abundance in the early fermentation stage and the late fermentation stage are: the microbial species corresponding to the significance level of P < 0.05 between the changes in the average relative abundance in the early fermentation stage and the late fermentation stage.
6. The screening method according to claim 1, wherein The screening method also includes evaluating the predictive validity of the Lasso linear model, including the following steps: The acid concentrations of the fermented mash in all fermentation triplets at a certain fermentation stage were measured, and the average acid concentration of the fermented mash in the corresponding fermentation stage was calculated. The acid concentrations of the fermented mash in the fermentation triplets were compared with the average acid concentrations of the fermented mash in the corresponding fermentation stage to evaluate whether the microorganisms in the Daqu samples at the corresponding fermentation stage were effectively colonized in the fermentation process. When the acid concentration of the fermented mash in the fermentation triplets was lower than the average acid concentration of the fermented mash in the corresponding fermentation stage, the microbial colonization process in the Daqu samples was defined as effective; otherwise, it was defined as ineffective. Preferably, the acid concentration is measured by titrating the total acid content with 0.05-0.2 mol / L NaOH, with the titration end point being pH=8.2, and calculating the acid content based on the concentration of the titrated NaOH solution, which is the acid concentration in the fermented mash.
7. The screening method according to claim 1, wherein The fermentation process includes at least one fermentation stage among the sand adding, sand making, first round, second round, third round, fourth round and fifth round of the fermentation process of the sauce-flavor liquor.
8. The screening method according to claim 1, wherein The collection of fermented mash samples comprises the following steps: simultaneously selecting the center point of the piled fermented mash, a surface point about 30-50 cm from the surface, a midpoint between the center point and the surface, and two points above and below the center point on a vertical line 50-60 cm from the center point for sampling; Preferably, the collection of the fermented mash sample further comprises: sampling in 8 equally divided areas within the same height circle of the corresponding sampling point.
9. The screening method according to claim 1, wherein The microbial abundance information includes: a relative abundance table of the first fermented mash sample, the Daqu sample, and the second fermented mash sample at the microbial species level, wherein the abundance table includes the classification of the microorganisms in the samples at the species level and their corresponding relative abundances.
10. The screening method according to claim 1, wherein The acquisition of the microbial abundance information includes: Total DNA is extracted from the first fermentation mash sample, the Daqu sample, and the second fermentation mash sample; the integrity and purity of the extracted DNA are detected and analyzed; a sequencing library is constructed and sequenced; the quality of the raw sequencing data is assessed and low-quality reads are filtered out; the obtained high-quality sequences are annotated to obtain the types of microorganisms contained in the corresponding samples, and the number of sequences corresponding to each microorganism is counted, thereby calculating the relative abundance of each type of microorganism.