Methods for predicting the age of cellar mud in strong-aroma baijiu
By detecting the bacterial microbial community in the pit mud and selecting modeling variables using the random forest regression algorithm, a pit mud age prediction model was established. This solved the problem that large instruments are required for pit mud age identification in existing technologies, and enabled rapid and accurate prediction of pit mud age.
Patent Information
- Application Number
- CN202110988371.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-26
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2041-08-26
AI Technical Summary
Existing technologies require specialized large-scale instruments and equipment for identifying the age of pit mud, and they have failed to effectively utilize microbial characteristics for rapid and accurate identification of pit mud age.
By detecting the bacterial microbial community in the pit mud, using 16S rRNA data processing and random forest regression algorithm to screen modeling variables, a predictive model is established to achieve rapid and accurate prediction of the pit mud age.
It enables rapid and accurate prediction of the age of pit mud, improves the accuracy of the model and the ease of operation, and is suitable for large-scale sample processing.
Smart Images

Figure CN113689913B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brewing technology, specifically to a method for predicting the age of cellar mud in strong-aroma baijiu (Chinese liquor) cellars. Background Technology
[0002] The unique brewing process of strong-aroma baijiu (Chinese liquor) involves fermenting the liquor in mud pits. These pits harbor a highly complex and diverse microbial community from the raw materials, the starter culture (daqu), the pit mud, and the environment. Their collaborative efforts create the rich and mellow flavor of strong-aroma baijiu. For centuries, brewing pioneers have discovered through practical experience that the microorganisms residing in the pit mud are crucial in the formation of the typical flavor profile of strong-aroma baijiu, summarizing the principles that "old pits produce exceptional aroma" and "the older the pit, the better the liquor." Further research using modern biotechnology has revealed that after years of continuous baijiu brewing, the pit mud gradually forms a healthy micro-ecosystem dominated by Clostridium and methanogens. Therefore, the continuous use of the pits is closely related to the quality of the pit mud, and establishing a method based on microbial communities is the most direct means of determining the age of the pit mud.
[0003] In the field of pit mud age identification technology, there is currently no applicable national standard. The main identification techniques proposed by researchers include: Shen Caihong et al. (ZL201110076641.1) proposed combining diffuse reflectance near-infrared spectroscopy with principal component analysis to establish characteristic projection maps and construct a database, determining the pit mud age by comparing spatial distribution. Tang Qinglan et al. (ZL201710917155.5) proposed using a pit mud age automated identification system based on the metabolic fingerprint cluster analysis of the dominant microbial community in the pit mud, capable of identifying pit mud quality and maturity. These methods provide different identification schemes for pit mud age identification, but they require specialized large-scale instruments and equipment, and do not rely on microorganisms as the most crucial and core characteristic for identification.
[0004] The microbial community composition of baijiu cellar mud is extremely complex, and key information is often lost among thousands of species. Therefore, there is significant room for improvement in the selection of important modeling variables. Developing a simple and rapid technique for identifying cellar mud age is currently a pressing need in baijiu quality assessment. Summary of the Invention
[0005] The purpose of this invention is to provide a method for predicting the age of cellar mud in strong-aroma baijiu, achieving rapid and accurate prediction of cellar mud age.
[0006] The present invention achieves the above-mentioned objective by adopting the following technical solution: a method for predicting the age of cellar mud in strong-aroma baijiu (Chinese liquor) cellars, comprising:
[0007] Step 1: Detect the bacterial microbial community of the pit mud. In the detection process, the pit mud sample data is first collected, and then the collected sample data is processed into amplicon data to obtain OTU tables.
[0008] Step 2: Select modeling variables in the OTU table using random forest regression;
[0009] Step 3: Build a prediction model based on the modeling variables;
[0010] Step 4: Predict the age of the pit mud using a prediction model.
[0011] Furthermore, in step 1, the pit mud sample data includes 16S rRNA data of the pit mud sample.
[0012] Furthermore, the specific method for collecting 16S rRNA data from the pit mud samples includes:
[0013] Step 101: Using pit mud samples of different ages as test samples, extract the genome from the pit mud to obtain mixed bacterial genome samples;
[0014] Step 102: Perform 16S rRNA sequencing on the mixed bacterial genome sample;
[0015] Step 103: Collect 16S rRNA amplicon sequencing data of bacteria from the cellar mud of strong-aroma baijiu from NCBI, DDBJ, ENA and CNGB databases, and screen them according to the comprehensiveness of information and the integrity of data.
[0016] Furthermore, in step 1, the specific method for processing the collected sample data into an amplicon data set to obtain the OTU table includes:
[0017] Step a: After splicing the unspliced two-end data, merge it with the data that does not need to be spliced, and then perform quality control;
[0018] Step b: Perform parametric clustering on the quality-controlled FASTA format data to obtain a preliminary OTU table, and then add annotations for the species names corresponding to the sequences using the SILVA database;
[0019] Step c: After statistically analyzing the differences in OTU abundance among samples, remove OTUs using different methods. Then, use the number of OTUs in the sample with the lowest abundance as the criterion for leveling. Finally, filter according to the percentage of OTUs that appear in the sample, and obtain the final OTU table.
[0020] Furthermore, in step 2, the specific methods for selecting modeling variables in the OTU table using random forest regression include:
[0021] Step 201: Use the corresponding data in the OTU table as test samples, and group the pit mud age of the test samples according to age group;
[0022] Step 201: Divide the grouped test samples into a test set and a training set according to the set ratio;
[0023] Step 202: Train the random forest discrimination model using the training set to obtain the error rate E1 of the preliminary discrimination model and the preliminary grouping;
[0024] Step 203: Test and optimize the preliminary discrimination model using the test set to obtain the error rate E2 of the final discrimination model and the final grouping;
[0025] Step 204: Compare the error rate E2 of the final grouping with the error rate E1 of the initial grouping. If E2≤E1, then obtain the contribution of all test samples to the grouping of pit mud age according to the final discrimination model. Otherwise, return to step 201 to regroup.
[0026] Step 205: Sort the samples in descending order of contribution, apply 10-fold cross-validation to the sorted samples, and then, based on the principle of parsimony, select the OTUs of the samples according to the cross-validation curve to screen out the modeling variables.
[0027] Furthermore, in step 3, the specific methods for establishing a predictive model based on the modeling variables include:
[0028] Step 301: Use the selected modeling variables as the modeling feature set, and group the pit mud age of the modeling feature set according to age group;
[0029] Step 302: Divide the grouped modeling feature set into a test set and a training set according to the set ratio;
[0030] Step 303: Train the random forest prediction model using the training set to obtain the preliminary prediction model and the error rate E3 of the preliminary grouping;
[0031] Step 304: Test and optimize the preliminary prediction model using the test set to obtain the error rate E4 of the final prediction model and the final grouping;
[0032] Step 305: Compare the error rate E4 of the final grouping with the error rate E3 of the initial grouping. If E4 ≤ E3, predict the age of the pit mud according to the final prediction model; otherwise, return to step 301 to regroup.
[0033] Furthermore, the characteristic bacterial genera for modeling include *Aminobacterium*, *Proteobacterium*, *Acetobacter*, *Escherichia*, *Acholesterolactone*, *Bacillus*, *Acetobacter*, *Symplocosomalemia*, *Porphyromonas*, *Clostridium*, *Thermophilus*, *Fastidiosipila*, and *Conostococcus*.
[0034] This invention analyzes the microbial community characteristics of pit mud of different ages, uses machine learning algorithms to mine feature variables from big data, and establishes a fast, efficient, and accurate pit age prediction model. It is easy to operate and suitable for the processing and screening of large-scale samples. By using random forest discrimination and ten-fold cross-validation to screen important variables for effective feature modeling, the feature space dimension is compressed, which effectively and reliably improves the modeling quality. In the modeling process, the error rate is also compared, which effectively improves the accuracy of the model. Attached Figure Description
[0035] Figure 1 This is a flowchart of the method for predicting the age of cellar mud in strong-aroma baijiu production according to the present invention.
[0036] Figure 2 This is the validation curve obtained using the ten-fold cross-validation method.
[0037] Figure 3 This is a diagram illustrating the accuracy after feature modeling. Detailed Implementation
[0038] The present invention provides a method for predicting the age of cellar mud in strong-aroma baijiu (Chinese liquor) cellars, comprising:
[0039] Step 1: Detect the bacterial microbial community of the pit mud. In the detection process, the pit mud sample data is first collected, and then the collected sample data is processed into amplicon data to obtain OTU tables.
[0040] Step 2: Select modeling variables in the OTU table using random forest regression;
[0041] Step 3: Build a prediction model based on the modeling variables;
[0042] Step 4: Predict the age of the pit mud using a prediction model.
[0043] In step 1, the pit mud sample data includes 16S rRNA data of the pit mud sample.
[0044] The specific methods for collecting 16S rRNA data from pit mud samples include:
[0045] Step 101: Using pit mud samples of different ages as test samples, extract the genome from the pit mud to obtain mixed bacterial genome samples;
[0046] Step 102: Perform 16S rRNA sequencing on the mixed bacterial genome sample;
[0047] Step 103: Collect 16S rRNA amplicon sequencing data of bacteria from the cellar mud of strong-aroma baijiu from NCBI, DDBJ, ENA and CNGB databases, and screen them according to the comprehensiveness of information and the integrity of data.
[0048] In step 1, the specific method for processing the collected sample data into amplicon data to obtain the OTU table includes:
[0049] Step a: After splicing the unspliced two-end data, merge it with the data that does not need to be spliced, and then perform quality control;
[0050] Step b: Perform parametric clustering on the quality-controlled FASTA format data to obtain a preliminary OTU table, and then add annotations for the species names corresponding to the sequences using the SILVA database;
[0051] Step c: After statistically analyzing the differences in OTU abundance among samples, OTUs are removed using different methods. Then, the number of OTUs in the sample with the lowest abundance is used as the criterion for leveling. Finally, the OTUs are filtered according to the fact that the frequency of OTUs (Operational Taxonomic Units) in the sample is not less than the set percentage, resulting in the final OTU table.
[0052] In step 2, the specific methods for selecting modeling variables in the OTU table using random forest regression include:
[0053] Step 201: Use the corresponding data in the OTU table as test samples, and group the pit mud age of the test samples according to age group;
[0054] Step 201: Divide the grouped test samples into a test set and a training set according to the set ratio;
[0055] Step 202: Train the random forest discrimination model using the training set to obtain the error rate E1 of the preliminary discrimination model and the preliminary grouping;
[0056] Step 203: Test and optimize the preliminary discrimination model using the test set to obtain the error rate E2 of the final discrimination model and the final grouping;
[0057] Step 204: Compare the error rate E2 of the final grouping with the error rate E1 of the initial grouping. If E2≤E1, then obtain the contribution of all test samples to the grouping of pit mud age according to the final discrimination model. Otherwise, return to step 201 to regroup.
[0058] Step 205: Sort the samples in descending order of contribution. Use the 10-fold cross-validation method on the sorted samples. Combined with the principle of parsimony, select the OTUs of the samples based on the cross-validation curve to screen out the modeling variables.
[0059] Among them, the cross-validation curve is as follows Figure 2 As shown, the horizontal axis represents the number of sample OTUs, and the vertical axis represents the cross-validation error rate.
[0060] In step 3, the specific methods for building a prediction model based on the modeling variables include:
[0061] Step 301: Use the selected modeling variables as the modeling feature set, and group the pit mud age of the modeling feature set according to age group;
[0062] Step 302: Divide the grouped modeling feature set into a test set and a training set according to the set ratio;
[0063] Step 303: Train the random forest prediction model using the training set to obtain the preliminary prediction model and the error rate E3 of the preliminary grouping;
[0064] Step 304: Test and optimize the preliminary prediction model using the test set to obtain the error rate E4 of the final prediction model and the final grouping;
[0065] Step 305: Compare the error rate E4 of the final grouping with the error rate E3 of the initial grouping. If E4 ≤ E3, predict the age of the pit mud according to the final prediction model; otherwise, return to step 301 to regroup.
[0066] The modeled bacterial genera include Aminobacterium, Proteiniphilum, Caproiciproducens, Escherichia, Acholeplasma, Bacillus, Oxobacter, Syntrophomonas, Petrimosa, Clostridium, Caloramator, Fastidiosipila, and Coprococcus. The final optimized model achieved a prediction accuracy of 90.78%.
[0067] A flowchart of an embodiment of the method for predicting the age of cellar mud in strong-aroma baijiu production according to the present invention is shown below. Figure 1 ,include:
[0068] Step S1: Detection of bacterial microbial community in pit mud. In the detection process, 16S rRNA gene sequencing data of pit mud samples are collected first, followed by high-throughput sequencing data processing.
[0069] Step S2: Select important variables for modeling from the OTU using random forest regression;
[0070] Step S3: Establish a predictive model by using the important modeling variables as the discriminant variables;
[0071] Step S4: Predict the age of the pit mud using a prediction model.
[0072] The specific implementation method is as follows:
[0073] A. Sample Collection: Cellar mud samples collected from the distillery were labeled with their year of origin and their genomes were extracted for 16S rRNA sequencing. 16S rRNA amplicon sequencing data of bacteria from the cellar mud of strong-aroma baijiu were collected from databases such as NCBI, DDBJ, ENA, and CNGB. Sample information was summarized and the dataset was screened. A total of the following sequencing data were collected: 13 samples of 1-year-old cellar mud, 8 samples of 4-year-old cellar mud, 75 samples of 6-year-old cellar mud, 8 samples of 8-year-old cellar mud, 10 samples of 10-year-old cellar mud, 1 sample of 20-year-old cellar mud, 15 samples of 30-year-old cellar mud, 20 samples of 40-year-old cellar mud, 72 samples of 50-year-old cellar mud, 38 samples of 100-year-old cellar mud, 8 samples of 300-year-old cellar mud, and 4 samples of 400-year-old cellar mud. These pit mud samples were divided into three age groups: YG (1-8 years) 104 samples; AG (10-50 years) 118 samples; AD (100-400 years) 50 samples, for a total of 272 samples.
[0074] B. Processing of 16S rRNA amplicon data: The Vsearch program was used to compare data from different sequencing intervals with data from a database containing the full 16S sequence of different species. Then, the sequences of the same species or OTUs were clustered based on the comparison results to generate an OTU table containing samples from different datasets. Finally, the species classification annotations of the 16S full-length sequences matched in the OTUs were added to the OTU table.
[0075] C. Screening of OTUs in the Sample: First, filter the OTU table obtained in step B by the sequencing volume of the sample data, selecting samples with counts greater than 10,000; then filter by OTU abundance, selecting OTUs with a relative abundance mean greater than 1 / 100,000; then flatten the OTU table according to the minimum number of sample sequences; finally, screen according to the probability of OTUs appearing in all samples greater than 80%, to obtain the final OTU table.
[0076] D. Divide the final OTU table dataset obtained in step C into a test set and a training set in a 7:3 ratio;
[0077] E. On the test set, a random forest discrimination algorithm is used to construct a preliminary grouping model to obtain the contribution of each OTU to the age grouping of the pit mud, and sort them in descending order.
[0078] F. Using the ten-fold cross-validation method, based on the cross-validation results shown in Table 2, the top 30 OTUs of importance were selected as important feature variables according to the principle of parsimony. These were used as modeling features to optimize the model construction and applied to the test set for prediction. The optimized accuracy increased from 88.48% to 90.78%, as shown in Table 3. The validation set accuracy performance is as follows: Figure 3 As shown, the horizontal axis represents the age grouping of the pit mud samples, and the vertical axis represents the corresponding accuracy rate. For example, after model prediction, 60% of an AD pit mud sample is identified as an AD pit mud sample, more than 20% is identified as an AG pit mud sample, and less than 20% is identified as a YG pit mud sample; after model prediction, more than 90% of a YG pit mud sample is identified as a YG pit mud sample, and less than 10% is identified as an AG or AD pit mud sample.
[0079] After screening using the method of this invention, the top 30 most effective modeling feature OTUs are shown in Table 1.
[0080] Table 1. Top 30 modeling features selected by the Random Forest selection method (OTUs)
[0081]
[0082]
[0083]
[0084] Table 2 shows the results of 10-fold cross-validation on the training set after the test and training sets were split in a 7:3 ratio.
[0085] serial number OTU quantity Error rate 1 1 0.414747 2 2 0.289401 3 3 0.235023 4 4 0.211982 5 6 0.173272 6 9 0.134562 7 14 0.129954 8 21 0.118894 9 31 0.117972 10 47 0.118894
[0086] Table 3 Comparison of accuracy on the test set after selecting modeling variables using tenfold cross-validation.
[0087]
[0088]
[0089] In summary, this invention achieves rapid and accurate prediction of the age of pit mud through a predictive model.
Claims
1. A method for predicting the pit age of Luzhou-flavor liquor pit mud, characterized in that, The application relates to a method for predicting the pit age of pit mud. The method comprises the following steps: Step 1, detecting the bacterial microbial community of pit mud, and obtaining an OTU table after treatment; Step 2, screening modeling variables in the OTU table through random forest regression; Step 3, establishing a prediction model according to the modeling variables; Step 4, predicting the pit age of pit mud through the prediction model. The specific method for screening modeling variables in the OTU table through random forest regression comprises the following steps: Step 201, taking the corresponding data in the OTU table as a test sample, and grouping the pit age of the test sample according to age; Step 201, dividing the grouped test sample into a test set and a training set according to a set proportion; Step 202, training a random forest discriminant model through the training set, obtaining a preliminary discriminant model and an error rate E1 of preliminary grouping; Step 203, testing and optimizing the preliminary discriminant model through the test set, obtaining a final discriminant model and an error rate E2 of final grouping; Step 204, comparing the error rate E2 of final grouping with the error rate E1 of preliminary grouping, if E2<=E1, obtaining the contribution degree of all test samples to the pit age of pit mud according to the final discriminant model, otherwise returning to step 201 for re-grouping; 2. The method for predicting the pit age of Luzhou-flavor liquor pit mud according to claim 1, characterized in that, Step 205, sorting according to the contribution degree from high to low, using ten-fold cross-validation method for the sorted sample, combining the simplicity principle, and selecting the modeling variables by taking or leaving the sample OTU according to the cross-validation curve. In step 3, the specific method for establishing a prediction model according to the modeling variables comprises the following steps: Step 301, taking the screened modeling variables as a modeling feature set, and grouping the pit age of the modeling feature set according to age; Step 302, dividing the grouped modeling feature set into a test set and a training set according to a set proportion; Step 303, training a random forest prediction model through the training set, obtaining a preliminary prediction model and an error rate E3 of preliminary grouping; Step 304, testing and optimizing the preliminary prediction model through the test set, obtaining a final prediction model and an error rate E4 of final grouping; Step 305, comparing the error rate E4 of final grouping with the error rate E3 of preliminary grouping, if E4<=E3, predicting the pit age of pit mud according to the final prediction model, otherwise returning to step 301 for re-grouping.
Citation Information
Patent Citations
Method for determining age of cellar mud of Luzhou-flavor liquor cellar
CN102226754B
Method for obtaining biological age of individual child based on oral microbial community
CN106202989A
Method for identifying pit age of pit mud
CN107513572A