Discriminating method of high-temperature yeast for making hard liquor and application thereof
Patent Information
- Application Number
- CN202510274536.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-18
AI Technical Summary
The existing technology is difficult to accurately determine the type of high-temperature daqu, which mainly depends on the experience of production personnel, and the existing methods have problems of low accuracy and strong subjectivity.
By obtaining the relative abundance value of high-temperature Daqu’s microorganisms, physical and chemical indicators, biological enzyme activity and flavor compound content as characteristic data, the stacking ensemble model is used for discrimination, and characteristic microorganisms, physical and chemical indicators and flavor compounds are screened, and a stacking ensemble model is established to accurately determine the Daqu type.
It realizes a more accurate judgment of the high-temperature large music type, improves the accuracy and reliability of the judgment, and reduces the dependence on the experience of production personnel.
Smart Images

Figure CN120340591A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of liquor detection and analysis, and relates to a method for discriminating high-temperature Daqu and its application. Background Art
[0002] In Chinese liquor, there is a saying that "qu is the bone of liquor". As the main saccharifying and fermenting agent in liquor brewing, Daqu plays roles in providing microorganisms and enzymes, providing nutrient substrates, serving as brewing raw materials, and generating flavor and precursor substances during the brewing process. The yield, quality, and flavor of liquor are closely related to Daqu. During the brewing process of Maotai-flavor liquor, high-temperature Daqu is used as the saccharifying and fermenting agent. Due to differences in the stacking space position and local environmental temperature and humidity in the fermentation warehouse, the fermentation degree of Daqu when leaving the warehouse is different, forming different characteristics. According to the obvious color differences, the Daqu leaving the fermentation warehouse can be divided into yellow Daqu, white Daqu, and black Daqu. There are differences in microorganisms, physical and chemical indexes, enzyme activities, flavor substances, etc. among the three types of Daqu. Black Daqu has a prominent Maotai flavor and a burnt aroma, but its saccharifying power is low; the indexes of yellow Daqu are relatively suitable. During the brewing process, different types of Daqu will affect the quality and taste of the produced liquor. In actual production applications, it is required that Daqu be stored for six months to ensure the stability of Daqu quality. However, after storage, yellow Daqu and black Daqu tend to be the same in shape and color and are difficult to distinguish.
[0003] At present, the identification of Daqu is mainly done by production staff through indicators such as Daqu's appearance, color and aroma. It is highly dependent on the experience of production staff, highly subjective, and not very accurate. Production staff need to undergo long-term repeated training. Machine learning is a branch of artificial intelligence that can analyze data to obtain rules and analyze unknown data to assist in making more correct and reasonable decisions. Based on this, Daqu type identification can be achieved through analysis based on indicators such as Daqu microorganisms, thermal enthalpy, absorbance, and non-volatile substances. For example, the Chinese invention patent with application number 201910120569.4 and patent name "A method for identifying the quality of Daqu based on a random forest model" provides a Daqu identification model constructed by combining high-throughput sequencing data with a random forest algorithm, but high-throughput sequencing cannot absolutely quantify the number of microorganisms, and only the relative microbial abundance obtained by high throughput cannot accurately reflect the state of Daqu, and the random forest algorithm is prone to overfitting, and may not produce a good classification effect on data with fewer features, making it difficult to accurately distinguish Daqu. At the same time, the patent mainly classifies the easily distinguishable Huangqu and Baiqu. For example, the Chinese invention patent with application number 202010518502.9 and patent name "A method for identifying the quality of Daqu" provides a method for judging the quality of Daqu based on the enthalpy value obtained by differential scanning calorimetry analysis of Daqu, but the enthalpy value is a state function that characterizes the energy of the Daqu system and is an indirect characterization value for the performance of Daqu. It is difficult to timely and quickly reflect the quality level of the entire Daqu, and it is difficult to effectively improve the accuracy of Daqu identification. For example, the Chinese invention patent with application number 202110980778.3 and patent name “A method for distinguishing the category of high-temperature Daqu” provides a method for distinguishing the category of Daqu by measuring the absorbance of the supernatant after Daqu extraction. However, this method requires preliminary sensory judgment of the category of Daqu, and is still highly dependent on production personnel. In addition, the range defined when judging by absorbance is wide, and the absorbance range between yellow qu and black qu is close, so accurate judgment cannot be made. Summary of the invention
[0004] The purpose of the present invention is to provide a method and application for distinguishing high-temperature Daqu, by obtaining the relative abundance values of microorganisms, physical and chemical indicators, biological enzyme activities and flavor compound contents of high-temperature Daqu as characteristic data, and further screening the characteristic data to obtain specific modeling characteristic data, and establishing a stacking integrated model. The stacking integrated model can more accurately distinguish the type of high-temperature Daqu.
[0005] In a first aspect, the present invention provides a method for distinguishing high-temperature Daqu, comprising the following steps:
[0006] Step 1, characteristic data acquisition: obtaining the relative abundance value of microorganisms, physical and chemical indicators, biological enzyme activity and flavor compound content of high-temperature Daqu as characteristic data;
[0007] Step 2, screening modeling characteristic data: based on the acquired characteristic data, screening characteristic microorganisms, characteristic physical and chemical indicators, characteristic enzyme activities and characteristic flavor compounds as modeling characteristic data;
[0008] Among them, the screening criteria for the characteristic microorganisms are: screening the top 5 in relative abundance value at the genus level; the screening criteria for the characteristic physical and chemical indicators and the characteristic enzyme activity are: according to the change rules of the physical and chemical indicators and enzyme activity indicators of the high-temperature Daqu during the storage period, and the difference between the yellow qu and the black qu increases with the storage period; the screening criteria for the characteristic flavor compounds are: the VIP value of the orthogonal partial least squares discriminant analysis is greater than 1;
[0009] Step 3: Establish a stacking ensemble model: Use H2O's Stacking model integration strategy to use multiple optimized base learners as the first layer of the ensemble model, use the cross-validation prediction results of the first layer base learners as the input features of the second layer meta learner, and retrain with the training set response variables to obtain a complete stacking ensemble model;
[0010] Step 4, Daqu type identification: input the data of the high-temperature Daqu to be tested into the established stacking Daqu identification model to perform Daqu type identification.
[0011] In some embodiments, the high temperature Daqu includes Huangqu and Heiqu.
[0012] In some embodiments, the physical and chemical indicators include acidity, moisture content, starch content, amino nitrogen content, saccharifying power and liquefaction power;
[0013] The biological enzyme activity includes pectinase activity, tannase activity, acid / neutral protease activity, saccharifying enzyme activity, liquefying enzyme activity, lipase activity, phytase activity and cellulase activity.
[0014] In some embodiments, the characteristic microorganisms include bacteria genera and / or fungi genera microorganisms; the bacteria genera microorganisms include Kroppenstedtia, Lentibacillus, Pullulanibacillus, Saccharopolyspora, Staphylococcus; the fungi genera microorganisms include Aspergillus, Monascus, Paecilomyces, Thermoascus, Thermomyces; the characteristic physicochemical indexes include acidity, moisture content, saccharifying power, amino acid nitrogen content; the characteristic bioenzyme activities include the activities of pectinase, acid protease, neutral protease, lipase, liquefying enzyme; the characteristic flavor compounds include ethyl acetate, isovaleraldehyde, tetramethylpyrazine, isovaleric acid, ethyl phenylacetate.
[0015] In some embodiments, in step 2, after screening out the characteristic microorganisms, characteristic physicochemical indexes, characteristic bioenzyme activities and characteristic flavor compounds, it further includes an optimal variable screening step for the screened characteristic microorganisms, characteristic physicochemical indexes, characteristic bioenzyme activities and characteristic flavor compounds, sorting the important variables and further screening the top 5 indexes in the importance ranking as the modeling feature data;
[0016] In some embodiments, the optimal variable screening is carried out by using a random forest model, and the parameter settings are as follows:
[0017] set.seed(812)
[0018] train.rfcv<-rfcv(trainx=train_data[,-25],trainy=train_data$group,cv.fold=10,recursive=TRUE)
[0019] train.rfcv$n.var
[0020] train.rfcv$error.cv
[0021] train.re.rfcv<-replicate(5,rfcv(trainx=train_data[,-25],train_data$group,cv.fold=10,step=1.5),simplify=FALSE).
[0022] In some embodiments, the top 5 indicators in the importance ranking are Thermoascus, liquefying enzyme, Paecilomyces, neutral protease, and Saccharopolyspora.
[0023] In some embodiments, the base learners are random forest models, gradient boosting tree models, generalized linear models, and neural network models;
[0024] Preferably, the optimal parameter settings of the random forest model are as follows: ntrees = 100, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812;
[0025] The optimal parameter settings of the gradient boosting tree model are as follows: ntrees = 50, max_depth = 5, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812;
[0026] The optimal parameter settings of the generalized linear model are as follows: alpha = 0.5, lambda = 0.1, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812;
[0027] The optimal parameter settings of the neural network model are as follows: activation = "Rectifier", hidden = c(50, 50), epochs = 20, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812.
[0028] In some embodiments, in step 1, obtaining the relative abundance values, physicochemical indices, biological enzyme activities, and flavor compound contents of the microorganisms in the high-temperature Daqu includes a sampling step, and the sampling method is as follows: randomly select the high-temperature Daqu to be used stored for 0 - 8 months at different positions in the storage bin, knock off the four corners, the crust, and the core of each piece of Qu respectively, crush them, mix them, and then sieve them, and collect the Qu powder as the sample.
[0029] In some embodiments, in step 1, the relative abundance values of the microorganisms are obtained by high-throughput sequencing methods, and the flavor compound contents are determined by HS-SPME-GC-MS method.
[0030] In some embodiments, in step 3, in the step of establishing the stacking integration model, performance evaluation of the established stacking integration model is further included, and the accuracy rate, recall rate, F1 score, Kappa coefficient and Matthews correlation coefficient of the stacking integration model are calculated.
[0031] In a second aspect, the present invention provides an application of any one of the above discrimination methods in discriminating the categories of Daqu.
[0032] In summary, the present application includes at least one of the following beneficial technical effects:
[0033] By obtaining the relative abundance values, physicochemical indexes, biological enzyme activities and flavor compound contents of microorganisms in high-temperature Daqu as characteristic data, and further screening the characteristic data to obtain specific modeling characteristic data, a stacking integration model is established, and the stacking integration model can more accurately discriminate the types of high-temperature Daqu. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a result graph of sorting important variables in Example 1;
[0035] Figure 2 It is a heat map of the confusion matrix of the stacking integration model established in Example 1;
[0036] Figure 3 It is a result graph of the performance evaluation of the stacking integration model established in Example 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0037] The technical solutions of the present invention are further described below through specific embodiments. The specific embodiments do not represent limitations on the protection scope of the present invention. Some non-essential modifications and adjustments made by others based on the concept of the present invention still fall within the protection scope of the present invention.
[0038] Example 1
[0039] A discrimination method for high-temperature Daqu includes the following steps:
[0040] Step 1, obtaining characteristic data: obtaining the relative abundance values, physicochemical indexes, biological enzyme activities and flavor compound contents of microorganisms in high-temperature Daqu as characteristic data; including the following steps:
[0041] 1.1 Sampling: Randomly select 54 pieces of Aspergillus flavus and 54 pieces of Aspergillus niger stored for 0 - 8 months at different positions in the storage bin, a total of 108 pieces. Knock off the four corners, the surface layer, and the core of each piece of koji respectively, break them, mix them, and then sieve them. Collect the koji powder as a sample for standby. Since white koji can be judged manually with the naked eye based on its appearance, without the need to use other tools and methods for judgment, while Aspergillus flavus and Aspergillus niger may have relatively similar appearances, and it is easy to have errors in visual observation. Therefore, a more accurate judgment is made through the judgment method of the present invention.
[0042] 1.2 Determination of characteristic data: Use the koji powder of Aspergillus flavus and Aspergillus niger to measure the relative abundance values of microorganisms, physicochemical indexes, biological enzyme activities, and flavor compound contents respectively.
[0043] Among them, the relative abundance values of microorganisms are obtained by high-throughput sequencing. The specific obtaining method is as follows: Extract the DNA of microorganisms in the koji powder of Aspergillus flavus and Aspergillus niger and perform PCR amplification, and then perform high-throughput sequencing on it based on the Illlumina MiSeq platform to obtain the relative abundance values of various microorganisms in each sample of Aspergillus flavus and Aspergillus niger.
[0044] The specific measurement method of high-throughput sequencing is as follows: Use a soil DNA kit to extract the total genomic DNA in the samples of Aspergillus flavus and Aspergillus niger, and store the extracted total genomic DNA samples at -20°C; respectively use the forward primer: 515F (5'-GTGCCAGCMGCCGCGGTAA-3') and the reverse primer: 907R (5'-CCGTCAATTCMTTTRAGTTT-3') to perform PCR amplification on the V4 - V5 region of the bacterial 16S rRNA gene, and then use the forward primer ITS1F (5'-CTTGGTCATTTAGAGGAAGTAA-3') and the reverse primer ITS2R (5'-GCTGCGTTCTTCATCGATGC-3') to amplify the ITS1 region of fungi; Incorporate the sample-specific 7-bp barcode into the primer for multiplex sequencing; Purify and quantify the PCR amplicons using a kit; After quantification, combine equal amounts of amplicons and perform sequencing using the paired-end 2×250bp with the Illlumina MiSeq platform and MiSeq kit v3 located in Shanghai Sangon Biotech Co., Ltd. to obtain the relative abundance values of various microorganisms in each sample.
[0045] The physical and chemical indexes include acidity, moisture content, starch content, amino nitrogen content, saccharifying power and liquefying power. The acidity, moisture content, starch content, amino nitrogen content, saccharifying power and liquefying power of the koji powder are determined according to the methods in QB / T 4257-2011 General Analytical Methods for Brewing Daqu. Among them, the moisture content is expressed as a percentage; the acidity is the number of millimoles of 0.1 mol / L sodium hydroxide standard solution consumed by 10 g of absolutely dry Daqu, with the unit of mmol / 10 g; the starch content is the mass of starch contained in every 100 g of Daqu, with the unit of g / 100 g; the amino nitrogen content has the unit of g / kg; liquefying power: under the conditions of 35 °C and pH 4.6, the number of grams of starch that can be liquefied by 1 g of absolutely dry Daqu in 1 h is one liquefying power unit (U), expressed as (g / g·h); saccharifying power: under the conditions of 35 °C and pH 4.6, the number of milligrams of glucose produced by the conversion of soluble starch by 1 g of absolutely dry Daqu in 1 h is one saccharifying power unit (U), expressed as (mg / g·h).
[0046] The biological enzyme activities include pectinase activity, tannase activity, acidic / neutral protease activity, glucoamylase activity, liquefying enzyme activity, lipase activity, phytase activity and cellulase activity. The determination methods of the biological enzyme activities are as follows:
[0047] Accurately weigh 10.0 g of the koji powder of Aspergillus flavus and Aspergillus niger into a 100 mL volumetric flask, add ultrapure water to dissolve and make up to the mark, place it in a constant temperature water bath at 40 °C for 1 h, shake it every 15 min, centrifuge at 4000 rpm for 10 min, and take the supernatant for the determination of biocatalase activity; simultaneously determine the pectinase activity, tannase activity, acidic / neutral protease activity, glucoamylase activity, and liquefying enzyme activity in the koji powder by UPLC-HRMS. The specific determination method is as follows: For UPLC ultra-high performance liquid chromatography analysis, a HypersilGold C18 (150 mm × 2.1 mm, 1.9 μm) ultra-high pressure liquid chromatography column is selected, the column temperature is 35 °C, the injection volume is 5 μL, the flow rate is 0.2 mL / min, the analysis time is 12 min, 1 mmol / L ammonium acetate - 0.1% formic acid aqueous solution is used as the first mobile phase and methanol is used as the second mobile phase, and a preset gradient elution mode is adopted, from 0 min - 2.0 min - 3.0 min - 10.0 min - 10.1 min - 12.0 min, and the volume fraction of the first mobile phase is 80% - 80% - 40% - 100% - 100% - 100%. The ion source conditions of the HRMS high-resolution mass spectrometer are as follows: a heated electrospray ionization source is used, the sheath gas flow rate is 45 arb, the auxiliary gas flow rate is 15 arb, the spray voltage is -2.8 kV, the ion source temperature is 250 °C, and the capillary temperature is 300 °C. The scanning conditions of the RMS high-resolution mass spectrometer are as follows: a positive and negative ion simultaneous scanning acquisition mode of full scan / data-dependent secondary scan is adopted, the first scan range is m / z 50.0 - 400.0 Da, the first scan resolution is 70000, the automatic gain control for the number of ions entering the orbitrap (AGC target): 1e6, the maximum injection time (Maximum IT): 30 ms, the secondary scan resolution: 17500, AGC target: 5e4, Maximum IT: 50 ms, and the normalized collision energy (Normalized collision energies, NCE): 10 eV, 20 eV, 30 eV.
[0048] The titration method in GB / T 23535-2009 "Lipase Preparation" is used to determine the lipase activity; the spectrophotometric method is used to determine the phytase activity and cellulase activity respectively. The specific determination method is as follows: The phytase is determined with reference to GB / T 18634-2009 "Determination of Phytase Activity in Feed - Spectrophotometric Method", and the cellulase is determined with reference to NY / T 912-2020 "Determination of Cellulase Activity in Feed Additives - Spectrophotometric Method".
[0049] The content of flavor compounds was determined by HS-SPME-GC-MS. The specific determination method was as follows: 1.5 g of koji powder sample was placed into a 20 mL headspace vial, and then 5 mL of saturated NaCl solution was added. The vial was capped. A 120 μm DVB / CAR / PDMS extraction fiber head was used and inserted into the headspace vial that had been equilibrated at 50 °C for 10 min. Headspace adsorption was carried out for 30 min, and desorption was carried out at 250 °C for 5 min, followed by GC-MS separation and identification. Among them, the GC-MS detection conditions were as follows: DB-WAX chromatographic column (60 m × 0.25 mm, 0.25 μm); the carrier gas was He, with a flow rate of 1.0 mL / min. The programmed temperature rise was as follows: the initial temperature was 40 °C, held for 3 min, heated at a rate of 3 °C / min to 150 °C, held for 2 min, and finally heated at 7 °C / min to 230 °C and held for 5 min. The inlet temperature was 250 °C, and splitless injection was used. Electron impact ionization (EI) ion source, electron energy 70 eV, ion source temperature 230 °C, transfer line temperature 250 °C, quadrupole temperature 150 °C, full scan, mass range: 30 - 550, analyzed using the NIST 20 spectral library. The measured flavor compounds included 18 esters, 14 acids, 9 aldehydes, 9 pyrazines, 8 alcohols, 7 ketones, and 10 other types of compounds, as shown in the following table:
[0050] Flavor compounds
[0051]
[0052]
[0053] Step 2: Screen the modeling feature data: Based on the obtained feature data, screen characteristic microorganisms, characteristic physicochemical indexes, characteristic biological enzyme activities, and characteristic flavor compounds as the modeling feature data; it includes the following steps:
[0054] 2.1 Preliminary screening of variables
[0055] Among them, for the screening of the relative abundance values of microorganisms: Based on the relative abundance values of the microorganisms in high-temperature Daqu, at the genus level, the top 5 microorganisms with the highest relative abundance values of each microorganism were screened as characteristic microorganisms. Based on the above screening criteria, finally, Kroppenstedtia, Lentibacillus, Pullulanibacillus, Saccharopolyspora, Staphylococcus of the bacterial genus and Aspergillus, Monascus, Paecilomyces, Thermoascus, Thermomyces of the fungal genus were screened as characteristic microorganisms.
[0056] Screening for physical and chemical indicators and biological enzyme activities: According to the changing rules of various physical and chemical indicators and biological enzyme activity indicators of high-temperature Daqu during the storage period, and the increasing gap between Aspergillus flavus and Aspergillus niger as the storage period progresses, they are screened as characteristic quality indicators. Based on the above screening criteria, finally, acidity, moisture content, saccharifying power, and amino acid nitrogen content are screened as characteristic physical and chemical indicators; the activities of liquefying enzyme, pectinase, acidic protease, neutral protease, and lipase are screened as characteristic biological enzyme activities.
[0057] Screening for flavor compounds: Orthogonal partial least squares discriminant analysis (OPLS-DA) is performed on the flavor compound information in Aspergillus flavus and Aspergillus niger obtained by HS-SPME-GC-MS using SIMCA 14.1 software. Then, flavor compounds with VIP value > 1 are used as characteristic flavor compounds to screen the common characteristic flavor compounds in Aspergillus flavus and Aspergillus niger. Based on the above screening criteria, finally, ethyl acetate, isovaleraldehyde, tetramethylpyrazine, isovaleric acid, and ethyl phenylacetate are screened as characteristic flavor compounds.
[0058] 2.2. Optimal variable screening
[0059] Using the random forest model, the rfcv function is used for 10-fold cross-validation, and recursive feature elimination is used for cross-validation. Random forest cross-validation is repeated 5 times, and then the results preliminarily screened in 2.1 are subjected to processing such as extraction, integration, reshaping of the data structure, and data type conversion. The sorting results of important variables are as follows in the table and Figure 1 shown as follows:
[0060] Sorting of 24 important variables preliminarily screened
[0061]
[0062]
[0063] According to the above table and Figure 1 The results show that the top 5 variables have the best effect. At this time, the cross-validation error is the lowest. The top 5 indicators in terms of importance, Thermoascus, liquefying enzyme, Paecilomyces, neutral protease, and Saccharopolyspora, are used as the modeling feature data.
[0064] Step 3. Establish a stacking integrated model: Use the stacking model integration strategy of H2O to take multiple optimized base learners as the first layer of the integrated model, use the cross-validation prediction results of the first-layer base learners as the input features of the second-layer meta-learner, and combine with the response variable of the training set for retraining to obtain a complete stacking integrated model; including the following steps:
[0065] 3.1. Data preprocessing: Aspergillus flavus and Aspergillus niger were divided into two groups. The label of Aspergillus flavus was named Y, and the label of Aspergillus niger was named B. Using the "caret" package, the data was divided into a training set and a validation set in a 7:3 ratio according to the groups of Y and B. The training set was used for model training, and the validation set was used for validating the model prediction performance. The training set and the validation set were loaded, 5 optimal variables selected by random forest screening were chosen, and the sample grouping labels were extracted as the response variables, while the remaining feature variables were used for model training. The data was converted into an H2O object that supports distributed computing through the H2O platform.
[0066] 3.2. Basic model training: When selecting the base learner, the classification accuracy of the base learner and the differences between models need to be considered. Four basic models were constructed using Random Forest (RF), Gradient Boosting Machine (GBM), Generalized Linear Model (GLM), and Deep Learning (DL). Each model used 5-fold cross-validation, and at the same time, the prediction results of the cross-validation were retained to support subsequent stacking integration.
[0067] The random forest model was trained using h2o.randomForest. After multiple debuggings, the best parameter settings for the random forest model were as follows: ntrees = 100, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812;
[0068] The gradient boosting tree model was trained using h2o.gbm. After multiple debuggings, the best parameter settings for the gradient boosting tree model were as follows: ntrees = 50, max_depth = 5, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812;
[0069] The generalized linear model was trained using h2o.glm. After multiple debuggings, the best parameter settings for the generalized linear model were as follows: alpha = 0.5, lambda = 0.1, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812;
[0070] Train a neural network model using h2o.deeplearning. After multiple debuggings, the optimal parameter settings of the neural network model are as follows: activation = "Rectifier", hidden = c(50, 50), epochs = 20, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812.
[0071] 3.3. Construction of the stacking ensemble model: Use the Stacking model integration strategy of H2O to take the above four optimized base learners as the first layer of the ensemble model, and use the cross-validation prediction results of the first-layer base learners as the input features of the second-layer meta-learner. Retrain in combination with the response variable of the training set to obtain a complete model. The meta-learner selects a learner with a simple structure and generalization ability. In this embodiment, the meta-learner in the second layer selects the GLM model. The parameter settings are as follows:
[0072] stacked_model <- h2o.stackedEnsemble(x = predictors, y = response, training_fram = train.h2o, base_models = base_models, metalearner_algorithm = "glm", seed = 812)
[0073] summary(stacked_model)
[0074] The h2o.stackedEnsemble function of H2O is used to create a stacked model, specifying the feature list, response variable, dataset used to train the model, and the list of base learners of the model. The meta-learner is a generalized linear model (GLM), and the random number seed is set to 812.
[0075] 3.4. Performance evaluation of the stacking ensemble model:
[0076] Predict the stacked model on the validation set and output the predicted class and probability. The results are as follows:
[0077] Output of the prediction results
[0078]
[0079]
[0080] Calculate classification metrics such as model accuracy (Accuracy), recall (Recall), F1-score (F1-Score), Kappa coefficient (k), and Matthews correlation coefficient (MCC) based on the above data, and draw a heatmap of the confusion matrix, as Figure 2 shown. Comprehensively evaluate the performance metrics of the model to verify the robustness and generalization ability of the model. The samples are divided into true positives TP, false positives FP, true negatives TN, and false negatives FN, with Aspergillus flavus as the positive.
[0081] Accuracy refers to the proportion of correctly predicted samples among all predicted samples. It is one of the most intuitive model evaluation metrics. The calculation formula is as follows:
[0082]
[0083] Recall is also known as sensitivity (Sensitivity) or true positive rate (True Positive Rate, TPR). It measures the proportion of samples that are actually positive and are correctly predicted as positive by the model. The calculation formula is as follows:
[0084]
[0085] The F1-score is the harmonic mean of accuracy and recall. It comprehensively considers the accuracy and recall of the model and can more comprehensively evaluate the performance of the model. The calculation formula is as follows:
[0086] where
[0087] The Kappa coefficient is a statistic used to measure the classification consistency, mainly used to evaluate the degree of consistency between two classification results. It takes into account the consistency caused by accidental factors in the classification results and can more accurately reflect the true level of consistency. The calculation formula is as follows (P o represents the proportion of observed consistency, and P e represents the expected proportion of consistency):
[0088]
[0089] The Matthews correlation coefficient is a metric used to measure the classification quality in binary or multi-class classification problems. It takes into account the combination of the four classification results of true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN), and can comprehensively evaluate the performance of the model. The calculation formula is as follows:
[0090]
[0091] The results of the model performance evaluation are shown in the following table:
[0092] Metric Value Accuracy 1 Recall 1 F1Score 1 Kappa 1 MCC 1
[0093] The classification performance of the model is evaluated using the ROC curve and AUC. The ROC curve of the validation set is plotted to show the discrimination ability of the model. The ROC curve is a graphical tool for evaluating the performance of binary classification models. It takes the false positive rate (FPR) as the horizontal axis and the true positive rate (TPR), which is also called recall rate, as the vertical axis. AUC is the area under the ROC curve, which is a numerical metric for quantifying the performance of the model. The closer the ROC curve is to the upper left corner, the better the performance of the model; the closer the AUC is to 1, the better the performance of the model. The performance evaluation results of the stacking ensemble model established in this embodiment are as Figure 3 shown.
[0094] Step 4: Discrimination of Daqu types: Randomly select 5 pieces of Aspergillus flavus to be used stored for 0 - 8 months in the storage bin, denoted as Y, and 4 pieces of Aspergillus niger, denoted as B. Determine the relative abundance values of microorganisms, physicochemical indexes, biological enzyme activities, and flavor compound contents according to the method for obtaining characteristic data in Step 1. Screen the characteristic data in the same way as in Step 2, and then input them into the stacking ensemble model established in Step 3 for discrimination of Daqu types, and output the predicted category and probability.
[0095] The discrimination results are shown in the following table:
[0096] Prediction result output
[0097]
[0098]
[0099] It can be seen that using five characteristic data of Thermoascus, liquefying enzyme, Paecilomyces, neutral protease, and Saccharopolyspora as modeling features, the discrimination result accuracy output by the constructed stacking ensemble model is 100%, which is a method that can successfully discriminate Daqu types.
[0100] Example 2
[0101] A method for discriminating high-temperature Daqu. Based on Example 1, on the basis of Example 1, when screening and modeling feature data in Step 2, only the obtained feature data is subjected to preliminary variable screening. The 24 kinds of feature data obtained by preliminary screening, namely: the relative abundances of 10 microorganisms including Kroppenstedtia, Lentibacillus, Pullulanibacillus, Saccharopolyspora, Staphylococcus, Aspergillus, Monascus, Paecilomyces, Thermoascus, and Thermomyces, and 4 physicochemical indexes including acidity, moisture content, saccharifying power, and amino acid nitrogen content, and the activities of 5 biological enzymes including liquefying enzyme, pectinase, acidic protease, neutral protease, and lipase, and 5 flavor compounds including ethyl acetate, isovaleraldehyde, tetramethylpyrazine, isovaleric acid, and ethyl phenylacetate are used as modeling feature data and are brought into Step 3 to establish a stacking integrated model. The performance evaluation results of the stacking integrated model are shown in the following table:
[0102] Metric Value Accuracy 1 Recall 1 F1Score 1 Kappa 1 MCC 1
[0103] The performances of the models constructed in Example 2 are consistent with those in Example 1, and can also play a good judging effect on the types of high-temperature Daqu. However, too many modeling features are likely to cause overfitting. The increase in the number of features will increase the number of parameters of the model, making the model more complex. When there are too many features, there may be a high degree of correlation or collinearity between the features, resulting in the model fitting the training data too closely, reducing the generalization ability, and the running speed of the model will slow down when there are many features.
[0104] Proportion 1 explores the influence of 5 variables ranked in the top 5 in terms of important variable ranking on the performance of the stacking integrated model
[0105] A method for discriminating high-temperature Daqu. Based on Example 1, on the basis of Example 1, when screening and modeling feature data in Step 2, after initially screening 24 variables including Kroppenstedtia, Lentibacillus, Pullulanibacillus, Saccharopolyspora, Staphylococcus, Aspergillus, Monascus, Paecilomyces, Thermoascus, Thermomyces, acidity, moisture content, saccharifying power, amino acid nitrogen content, liquefying enzyme, pectinase, acid protease, neutral protease, lipase, ethyl acetate, isovaleraldehyde, tetramethylpyrazine, isovaleric acid, and ethyl phenylacetate, further perform optimal variable screening on this basis. Among the top 5 variables in the important variable ranking, further screen 4 as the modeling feature data to obtain 5 combinations of important variables. The first one is: Thermoascus, liquefying enzyme, Paecilomyces, neutral protease; the second one is: Thermoascus, liquefying enzyme, Paecilomyces, Saccharopolyspora; the third one is: Thermoascus, liquefying enzyme, neutral protease, Saccharopolyspora; the fourth one is: Thermoascus, Paecilomyces, neutral protease, Saccharopolyspora; the fifth one is: liquefying enzyme, Paecilomyces, neutral protease, Saccharopolyspora. Respectively use the important variables in the above 5 combinations as the modeling feature data and substitute them into Step 3 to establish 5 stacking ensemble models respectively. The performance evaluation results of the stacking ensemble models are shown in the following table:
[0106] Performance evaluation of models constructed with 4 variables
[0107]
[0108] In the above table, "-" represents the lack of this variable. For example, "-Saccharopolyspora" represents the performance evaluation result of the first group of stacking ensemble models lacking the Saccharopolyspora variable in this column. It can be seen from the results that if any one of the top 5 variables in the important variable ranking is missing, the performance of the constructed model is inferior to that of Example 1 in all aspects.
[0109] Comparative Example 2 explores the influence of sensory dimension scores on the performance of the stacking ensemble model
[0110] A method for discriminating high-temperature Daqu. Based on Example 1, on the basis of Example 1, when obtaining characteristic data in Step 1, in addition to the relative abundance values of microorganisms, physicochemical indexes, biological enzyme activities, and flavor compound contents, it also includes taking the sensory dimension score as characteristic data. The acquisition method of the sensory dimension score is as follows: Before crushing the Qu blocks taken when obtaining characteristic data in Step 1, professional production personnel in the workshop score the appearance (20 points), cross-section (30 points), aroma (30 points), dryness degree (20 points), etc. of the Qu blocks to obtain the sensory evaluation data of each Qu block.
[0111] Then, together with the top 5 indicators in terms of importance ranking obtained by optimal variable screening in Step 2, namely Thermoascus, liquefying enzyme, Paecilomyces, neutral protease, and Saccharopolyspora, they are used as modeling characteristic data and brought into the stacking ensemble model established in Step 3. The performance evaluation results of the stacking ensemble model are shown in the following table:
[0112] Constructing model performance evaluation by adding sensory dimension
[0113]
[0114]
[0115] The performance of the model constructed as a result is inferior to that of Example 1 in all aspects. This may be because the sensory dimension indexes are evaluated manually, with strong subjectivity and weak result accuracy, which easily affects the final discrimination result.
[0116] Comparative Example 3 explores the influence of the number of modeling characteristics on the performance of the stacking ensemble model
[0117] A method for discriminating high-temperature Daqu. Based on Example 1, on the basis of Example 1, when obtaining characteristic data in Step 2 and screening and modeling characteristic data in Step 2, 24 variables are obtained through preliminary variable screening, including the relative abundances of 10 microorganisms Kroppenstedtia, Lentibacillus, Pullulanibacillus, Saccharopolyspora, Staphylococcus, Aspergillus, Monascus, Paecilomyces, Thermoascus, Thermomyces, 4 physicochemical indexes including acidity, moisture content, saccharifying power, and amino acid nitrogen content, and the activities of 5 biological enzymes including liquefying enzyme, pectinase, acid protease, neutral protease, and lipase, as well as 5 flavor compounds including ethyl acetate, isovaleraldehyde, tetramethylpyrazine, isovaleric acid, and ethyl phenylacetate. The 4 single dimensions are respectively used as modeling characteristic data, that is, the relative abundances of microorganisms, physicochemical indexes, biological enzyme activities, and flavor compounds are respectively used as modeling characteristic data, and 4 stacking ensemble models are respectively established. The performance evaluation results of the stacking ensemble models are shown in the following table:
[0118] Performance evaluation of models constructed with single important variables
[0119]
[0120] The performance of each model constructed as a result is inferior to that of Example 1, indicating that the discrimination effect of establishing a stacking ensemble model with the relative abundances of microorganisms, physicochemical indexes, biological enzyme activities, and flavor compounds as modeling characteristic data respectively is not as accurate as that of establishing a stacking ensemble model with Thermoascus, liquefying enzyme, Paecilomyces, neutral protease, and Saccharopolyspora as modeling characteristic data together. The present invention selects multi-dimensional Daqu detection parameters as characteristic data, and screens and obtains the optimal variables as modeling characteristic data, so as to discriminate Daqu more comprehensively with better comprehensive quality.
[0121] Comparative Example 4 explores the influence of the modeling type on the discrimination effect of Daqu
[0122] A method for discriminating high-temperature Daqu. Based on Example 1, on the basis of Example 1, when obtaining characteristic data in Step 2 and screening modeling characteristic data in Step 2, after initially screening variables to obtain 24 variables including Kroppenstedtia, Lentibacillus, Pullulanibacillus, Saccharopolyspora, Staphylococcus, Aspergillus, Monascus, Paecilomyces, Thermoascus, Thermomyces and the activities of biological enzymes such as acidity, moisture content, saccharifying power, amino acid nitrogen content, liquefying enzyme, pectinase, acid protease, neutral protease, lipase, ethyl acetate, isovaleraldehyde, tetramethylpyrazine, isovaleric acid, ethyl phenylacetate, further perform optimal variable screening on this basis, and use the top 5 variables in the ranking of important variables, namely Thermoascus, liquefying enzyme, Paecilomyces, neutral protease, Saccharopolyspora as modeling characteristic data. Then use random forest to construct a discrimination model and evaluate the performance of the random forest model. Use the randomForest function to train a random forest model, specify the number of trees to be 500, and calculate the variable importance and the proximity between samples, determine the number of trees in the random forest, determine the parameters as mtry = 4, ntree = 150, the number of variables randomly selected when each node splits is 4, and construct 150 trees. The performance evaluation results of the random forest model are shown in the following table:
[0123] Performance evaluation of the random forest construction model
[0124] Metric Value Accuracy 0.91 Recall 0.81 F1Score 0.90 Kappa 0.81 MCC 0.83
[0125] The performance of each aspect of the constructed random forest model is inferior to that of Example 1, indicating that the Stacking model integration strategy of H2O is used to construct the model, integrating the advantages of multiple machine learning models, and can build a more comprehensive and accurate model. H2O uses a distributed computing architecture, can handle large-scale data sets, shows more excellent computing speed when processing massive data, supports multiple algorithms, and has strong ease of use and scalability.
[0126] It can be understood that the present invention is described through some embodiments. Those skilled in the art know that without departing from the spirit and scope of the present invention, various changes or equivalent replacements can be made to these features and embodiments. In addition, under the teaching of the present invention, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application belong to the scope protected by the present invention.
Claims
1. A method for discriminating high-temperature Daqu, characterized in that, The following steps are involved: Step 1, characteristic data acquisition: obtaining the relative abundance value of microorganisms, physical and chemical indicators, biological enzyme activity and flavor compound content of high-temperature Daqu as characteristic data; Step 2, screening modeling characteristic data: based on the acquired characteristic data, screening characteristic microorganisms, characteristic physical and chemical indicators, characteristic enzyme activities and characteristic flavor compounds as modeling characteristic data; Among them, the screening criteria for the characteristic microorganisms are: screening the top 5 in relative abundance value at the genus level; the screening criteria for the characteristic physical and chemical indicators and the characteristic enzyme activity are: according to the change rules of the physical and chemical indicators and enzyme activity indicators of the high-temperature Daqu during the storage period, and the difference between the yellow qu and the black qu increases with the storage period; the screening criteria for the characteristic flavor compounds are: the VIP value of the orthogonal partial least squares discriminant analysis is greater than 1; Step 3: Establish a stacking ensemble model: Use H2O's Stacking model integration strategy to use multiple optimized base learners as the first layer of the ensemble model, use the cross-validation prediction results of the first layer base learners as the input features of the second layer meta learner, and retrain with the training set response variables to obtain a complete stacking ensemble model; Step 4, Daqu type identification: input the data of the high-temperature Daqu to be tested into the established stacking Daqu identification model to perform Daqu type identification.
2. The discrimination method according to claim 1, wherein The high temperature Daqu includes yellow qu and black qu.
3. The discrimination method according to claim 1, wherein The physical and chemical indicators include acidity, moisture content, starch content, amino nitrogen content, saccharification power and liquefaction power; The biological enzyme activity includes pectinase activity, tannase activity, acid / neutral protease activity, saccharifying enzyme activity, liquefying enzyme activity, lipase activity, phytase activity and cellulase activity.
4. The discrimination method according to claim 1, wherein The characteristic microorganisms include bacterial and / or fungal microorganisms; the bacterial microorganisms include Kroppenstedtia, Lentibacillus, Pullulanibacillus, Saccharopolyspora, and Staphylococcus; the fungal microorganisms include Aspergillus, Monascus, Paecilomyces, Thermoascus, and Thermomyces; the characteristic physical and chemical indicators include acidity, moisture content, saccharifying power, and amino acid nitrogen content; the characteristic biological enzyme activities include the activities of pectinase, acid protease, neutral protease, lipase, and liquefaction enzyme; the characteristic flavor compounds include ethyl acetate, isovaleraldehyde, tetramethylpyrazine, isovaleric acid, and ethyl phenylacetate.
5. The discrimination method according to claim 1, wherein In the step 2, after the characteristic microorganisms, characteristic physical and chemical indicators, characteristic enzyme activities and characteristic flavor compounds are screened, the optimal variable screening step is also included for the screened characteristic microorganisms, characteristic physical and chemical indicators, characteristic enzyme activities and characteristic flavor compounds, and the important variables are sorted and the top 5 indicators in importance are further selected as modeling feature data; Preferably, the top 5 indicators in the importance ranking are Thermoascus, liquefying enzyme, Paecilomyces, neutral protease, and Saccharopolyspora.
6. The discrimination method according to claim 1, characterized in that, The base learners are random forest model, gradient boosting tree model, generalized linear model, and neural network model; Preferably, the optimal parameter settings of the random forest model are as follows: ntrees = 100, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812; The optimal parameter settings of the gradient boosting tree model are as follows: ntrees = 50, max_depth = 5, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812; The optimal parameter settings of the generalized linear model are as follows: alpha = 0.5, lambda = 0.1, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812; The optimal parameter settings of the neural network model are as follows: activation = "Rectifier", hidden = c(50, 50), epochs = 20, nfolds = 5, keep_cross_validation_predictions = TRUE, seed = 812.
7. The discrimination method according to claim 1, wherein In step 1, obtaining the relative abundance values, physicochemical indexes, biological enzyme activities, and flavor compound contents of the microorganisms in the high-temperature Daqu includes a sampling step, and the sampling method is as follows: randomly extract the high-temperature Daqu to be used stored for 0 - 8 months at different positions in the storage bin, respectively knock the four corners, the surface, and the core of each piece of Daqu, crush them, mix them, and then sieve them, and collect the Daqu powder as a sample.
8. The discrimination method according to claim 1, wherein In step 1, the relative abundance values of the microorganisms are obtained by high-throughput sequencing method, and the flavor compound contents are determined by HS-SPME-GC-MS method.
9. The discrimination method according to any one of claims 1-8, characterized in that, In step 3, the performance evaluation of the established stacking ensemble model is also included in the step of establishing the stacking ensemble model, and the accuracy rate, recall rate, F1 score, Kappa coefficient, and Matthews correlation coefficient of the stacking ensemble model are calculated.
10. Application of a discrimination method according to any one of claims 1 - 9 in discriminating the category of Daqu.
Citation Information
Patent Citations
A method for identifying the quality of large-batch koji based on a random forest model
CN109949863B
A method for identifying the quality of Daqu
CN111830079B
A method for identifying the type of high-temperature koji (a type of Chinese liquor).
CN113720790B