Method for detecting aging time of Chinese liquor based on GC-MS and machine learning model
Through the method of combining GC-MS and machine learning models, the key characteristic values related to aging time in liquor are identified, and the accuracy of liquor year identification is solved, and it is suitable for year identification of various liquor flavors.
Patent Information
- Application Number
- CN202310131160.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2043-02-17
AI Technical Summary
The prior art is difficult to accurately identify the aging time of liquor, resulting in the generation of counterfeit and inferior "agen" liquor, and lacks systematic testing methods.
GC-MS was used to qualitatively and quantify liquor compounds, calculate the reaction concentration quotient of the uniformity index and the reversible reaction of esterification and hydrolysis, and screen key characteristic values with machine learning model to establish a liquor year identification model.
It realizes accurate detection of the aging time of liquor, promotes in-depth understanding of the aging mechanism of liquor, and is suitable for the year identification of various liquor flavors.
Smart Images

Figure CN116110503B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of liquor identification, and particularly relates to a method for detecting the aging time of liquor based on GC-MS and a machine learning model. Background Art
[0002] Liquor is a traditional Chinese distilled spirit with a history of over 1,000 years and is very popular among consumers. The production volume in 2021 was 71.56 million liters. Liquor has ranked among the top five in the world's top 50 spirits for five consecutive years, and the total brand value in 2021 was approximately $90 billion (Brand Finance, 2021). There are 12 main flavor types of liquor, which have different production processes, such as different raw materials, fermenting agents, environmental conditions, and microorganisms, etc. These differences determine the different compositions and concentrations of flavor compounds, thus resulting in various aroma styles and quality grades (Jia et al., 2020).
[0003] The aging process has a positive impact on the sensory quality of fresh liquor, making its spicy and rough flavors become soft and harmonious. Therefore, like Western spirits such as wine, whisky, and brandy, aging is one of the key factors affecting the market price of liquor. The high profit of aged liquor has led to the emergence of some counterfeit and inferior "aged" liquors. Therefore, the age identification of liquor is of great significance for protecting consumers and liquor manufacturers with good reputations.
[0004] Currently, various analytical techniques are used for the identification of liquor, such as mid-infrared spectroscopy, time-resolved fluorescence, proton nuclear magnetic resonance spectroscopy (1H NMR), ultra-high performance liquid chromatography quadrupole orbitrap high-resolution mass spectrometry (UHPLC-Q-Orbitrap HRMS), gas chromatography-mass spectrometry (GC-MS), and so on. Among them, the GC-MS fingerprint of liquor has a large amount of information and can better reflect the sensory characteristics of liquor. It has been reported that more than 2,020 volatile flavor compounds have been detected, but not all compounds in the GC-MS fingerprint are related to the aging time of liquor. The relationship between the aging time and compounds is still in the exploratory stage. Therefore, it is necessary to identify the key differential compounds related to aging from the entire fingerprint dataset.
[0005] Previously, studies have reported on the variation patterns of compounds related to the aging of Chinese liquor. For example, Jia (2021) found that the contents of 10 esters in Fengxiang-flavor Chinese liquor (aged for 0 - 19 years) increased with the increase in aging time. After monitoring a batch of Laowuzeng Chinese liquor for one year, Zhu (2020) found that the aldehyde content showed an upward trend, while Tang (2020) found that with the aging (6, 12, 18, 24, and 30 months), the aldehyde content in Luzhou-flavor Chinese liquor showed a downward trend. Although there have been some studies on aged Chinese liquor, there is still no consistent conclusion on the variation patterns of compounds in Chinese liquor during the aging process. An important reason may be that the initial compound contents of various Chinese liquors are different, which may lead to inconsistent reaction directions and reaction rates of spontaneous reactions in the Chinese liquor system, resulting in inconsistent compound compositions after aging. Therefore, deeply exploring the aging mechanism of Chinese liquor and understanding the thermodynamics of spontaneous reactions related to aging helps to screen key aging compounds and identify the vintage year. Summary of the Invention
[0006] In order to more systematically and accurately detect the aging time of Chinese liquor, the present invention provides a method for detecting the aging time of Chinese liquor based on GC-MS and a machine learning model, which includes the following steps:
[0007] A. Qualitatively analyze the compounds in Chinese liquor samples with gradient aging times by using GC-MS to obtain a qualitative data set of compounds in each liquor sample, and quantitatively analyze the compounds in the qualitative data set to obtain a concentration data set of compounds in each liquor sample;
[0008] B. Calculate the evenness index of compounds in each Chinese liquor sample through chemometrics. The calculation formula is: where i is 1, 2, …… S, S is the total number of compounds in a certain Chinese liquor sample in step A, N is the total content of all compounds in this Chinese liquor sample, and N i is the content of a certain compound i in this Chinese liquor sample;
[0009] C. According to the qualitative and quantitative data sets determined in step A, select the combination of organic acids, alcohols, corresponding esters, and water that coexist in the esterification-hydrolysis reversible reaction in each Chinese liquor sample in step A, and establish a reversible reaction equilibrium equation: Calculate the reaction concentration quotient Qc of each esterification-hydrolysis reversible reaction in each Chinese liquor sample by taking the product of the powers of the concentrations of the products and dividing it by the product of the powers of the concentrations of the reactants. The calculation formula is: where C is the real-time concentration of each compound in the reversible reaction equilibrium equation;
[0010] D. The concentration dataset obtained in step A, the uniformity index obtained in step B, and the reaction concentration quotient obtained in step C are used as eigenvalues. An importance ranking is performed on the eigenvalues using a sorting algorithm. Starting from the eigenvalue with the highest ranking, one feature is added to the previous subset each time to form the next subset, resulting in (total number of eigenvalues × total number of sorting algorithms) input feature subsets of different sizes. Each feature subset is input into a machine learning model respectively, and the machine learning model is allowed to perform year discrimination to obtain (total number of eigenvalues × total number of sorting algorithms × total number of machine learning models) discrimination F1 values. All F1 value data are compared with a screening threshold, and M input feature subsets that are input into the machine learning model when F1 is greater than the screening threshold are retained. The intersection of these M input data subsets is taken to obtain the final key eigenvalues.
[0011] Among them, the above detection method further includes the following step: E. According to the key eigenvalues determined in step D, the key eigenvalues of the white liquor to be tested with an unknown aging time are detected and input into at least one of the machine learning models used in step D to determine its aging time.
[0012] Among them, in step A of the above detection method, when qualitatively analyzing the compounds in the liquor sample, compounds with a detection rate lower than 30% in the white liquor samples at gradient aging times are excluded according to the 30% rule, and only the remaining specific compounds are quantified.
[0013] Preferably, in step A of the above detection method, when qualitatively analyzing the compounds in the liquor sample, compounds with a detection rate lower than 50% in the white liquor samples at gradient aging times are excluded according to the 50% rule, and the remaining specific compounds are quantified.
[0014] Among them, in step A of the above detection method, an unsupervised hierarchical clustering algorithm is used to perform clustering analysis on the white liquor samples at gradient aging times.
[0015] Among them, in step D of the above detection method, the sorting algorithm is at least three of Information Gain, Information Gain Ratio, Gini Decrease, ANOVA, Chi-square test (χ 2 ), Feature Weighting Algorithm (ReliefF), and Fast Correlation-Based Filter (FCBF);
[0016] Preferably, in step D of the above detection method, the sorting algorithm is five kinds including Information Gain, Information Gain Ratio, ANOVA, Feature Weighting Algorithm, and Fast Correlation-Based Filter.
[0017] Among them, in the above detection method, in step D, the machine learning models are at least two of Random Forest, Support Vector Machine (SVM), Neural Network, K-Nearest Neighbors (KNN), Gradient Boosting, eXtreme Gradient Boosting, Linear Regression, Logistic Regression, AdaBoost, and Stochastic Gradient Descent.
[0018] Preferably, in the above detection method, in step D, the number of machine learning models is three, namely Random Forest, Support Vector Machine, and Neural Network.
[0019] In some embodiments of the present invention, in step A, the liquor samples with gradient aging time are selected from 0 to 11 years.
[0020] In some preferred embodiments of the present invention, in step A, the liquor samples with gradient aging time are divided into four groups: [0, 1) year, [1, 5) years, [5, 9) years, and [9, 11] years.
[0021] Among them, in the above detection method, in step A, headspace solid-phase microextraction, liquid-liquid microextraction, or liquid-liquid microextraction-BSTFA derivatization is used to pretreat the liquor samples.
[0022] Preferably, in the above detection method, in step A, the operation of headspace solid-phase microextraction is as follows: the liquor sample is diluted to a final ethanol content of 5-15 Vol%, NaCl is added, ethyl lactate and compounds with a retention time longer than acetic acid are both internal-standardized with 2-methylhexanoic acid, alcohols with a retention time shorter than acetic acid are internal-standardized with tert-amyl alcohol, and other compounds are internal-standardized with amyl acetate. Volatile compounds are extracted from the top space of the sample using an SPME fiber. The sample is pre-equilibrated at 45-60 °C for 0-10 min, and then extracted at 45-60 °C for 30-60 min for subsequent GC-MS determination.
[0023] Preferably, in the above detection method, in step A, the operation of liquid-liquid microextraction is as follows: three internal standards, namely amyl acetate, 2-methylhexanoic acid, and tert-amyl alcohol, and saturated sodium chloride solution are added to the liquor sample, and then solution A is added. After stirring well for more than 3 min, the upper organic phase is collected after static stratification. After the organic phase is concentrated by nitrogen blowing at room temperature, subsequent GC-MS determination is carried out; solution A is a mixed solvent with a volume ratio of anhydrous ether to pentane = 1:1.
[0024] Preferably, in the above detection method, in step A, the operation of liquid-liquid microextraction-BSTFA derivatization is as follows: Add nonadecanoic acid internal standard to the liquor sample, then add saturated sodium chloride solution, dilute to an ethanol concentration of about 10%, add solution A to the diluted system for extraction, vortex for more than 3 min, stand for more than 20 min, collect the organic layer, extract the remaining aqueous phase with acetonitrile and dichloromethane in sequence, combine the organic phases, add pyridine containing 2.5% (v / v) hydroxylamine hydrochloride to the organic phase, add BSTFA containing 1% TMCs, then vortex for more than 5 s, incubate in a metal bath at 45-85 °C for 1-5 h, after the reaction, centrifuge at 6000-12000 rpm for more than 3 min, take the supernatant, and perform subsequent GC-MS determination; the solution A is a mixed solvent of anhydrous ether:pentane with a volume ratio of 1:1.
[0025] Among them, in the above detection method, in step A, when headspace solid-phase microextraction is adopted, the GC-MS conditions are as follows: DB-WAX chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8-1.4 mL / min; electron impact mode: 70 eV; the transmission line temperature at the connection between the chromatographic column and the ion source is set at 200-230 °C; scanning mode, the scanning range is m / z 33-350; the inlet temperature is 250-300 °C, the desorption time is 5-15 min, and the split ratio is 2-20:1; temperature programming: starting from 50 °C, increase in stages to 235 °C and hold for more than 2 min.
[0026] Among them, in the above detection method, in step A, when liquid-liquid microextraction is adopted, the GC-MS conditions are as follows: DB-WAX chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8-1.4 mL / min; electron impact mode: 70 eV; the transmission line temperature at the connection between the chromatographic column and the ion source is set at 200-230 °C; scanning mode, the scanning range is m / z 33-350; the sample injection volume is 1 μL, the inlet temperature is 250-300 °C, the desorption time is 5-15 min, and the split ratio is 2-20:1; temperature programming: starting from 35 °C, increase in stages to 235 °C and hold for more than 10 min.
[0027] Among them, in the above detection method, in step A, when liquid-liquid microextraction-BSTFA derivatization is adopted, the GC-MS conditions are as follows: HP-5MS chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8-1.4 mL / min; electron impact mode: 70 eV; the transmission line temperature at the connection between the chromatographic column and the ion source is set at 200-230 °C; scanning mode, the scanning range is m / z 33-350; the sample injection volume is 1 μL, the inlet temperature is 250 °C, and the split ratio is 2-20:1; temperature programming: gradually increase from 65 °C to 280 °C and hold for more than 2 min.
[0028] Among them, in the above detection method, in step D, before the sorting algorithm performs importance sorting, the eigenvalue includes 108 compound concentrations in the compound concentration dataset in step A (see the compounds and their concentrations in Example 2 and Figure 1 ), 1 uniformity index of the compound in step B (see Example 3 and Figure 3 D, Table 4), and 19 reaction concentration quotients Qc in step C (see Table 5 specifically).
[0029] Among them, in the above detection method, in step D, the key eigenvalues include 36 compound concentrations screened from the compound concentration dataset in step A, 1 uniformity index of the compound in step B, and 4 Qc screened from the reaction concentration quotient Qc in step C (see Table 6 specifically).
[0030] Advantages of the present invention:
[0031] In this study, GC-MS was used for qualitative and quantitative analysis of the types and content changes of compounds in Baijiu, the uniformity of compounds in Baijiu was evaluated by calculating the uniformity index, the differences in the reaction concentration quotient and thermodynamic equilibrium constant of various reversible esterification reactions during the aging process of Baijiu were quantitatively analyzed, and the kinetic mechanism of the change in the acid and ester content in Baijiu was revealed; then, after combining the three datasets of compound concentration, uniformity index, and reaction concentration quotient, through the sorting algorithm and combined with the machine learning model, the key eigenvalues closely related to the identification of the aging time of Baijiu were screened out from them, thus establishing a Baijiu vintage identification model, and verifying the identification accuracy of the aging time of Baijiu, realizing the accurate detection of the aging time of Baijiu.
[0032] The present invention analyzes the flavor characteristics and content change rules of the screened vintage characteristics, promotes the in-depth understanding of the aging mechanism of Baijiu, and the proposed identification strategy and vintage identification method of vintage characteristics are not only applicable to Luzhou-flavor Baijiu, but also applicable to other Baijiu flavors or other spirits. Description of the Drawings
[0033] Figure 1 It is a heat map drawn after normalization of 54 wine samples and a result map of HCA clustering analysis based on Spearman similarity.
[0034] Figure 2 It is a result map of HCA clustering analysis of 54 wine samples based on Euclidean distance.
[0035] Figure 3Structural diagram for difference analysis of samples from different years; among them, A is the PCA analysis result diagram of Baijiu samples based on the concentrations of 108 compounds; B is the PCA analysis result diagram of Baijiu samples based on 41 key characteristic values; C is the total content diagram of various compounds in samples from different years; D is the uniformity index diagram of compound concentrations in samples from different years; E is the stacked bar chart of the contents of all compounds detected in samples from different years; F is the percentage chart of the contents of all compounds detected in samples from different years.
[0036] Figure 4 is the number of compounds detected in Baijiu with different aging times.
[0037] Figure 5 is the change rule diagram of acids or esters in wine samples from different years; among them, A is the total acid content and the content diagram of four main short-chain acids C2-C6; B is the content diagram of medium-chain acids C7-C 10 ; C is the content of long-chain acids C 12 -C 18 ; D is the total ethyl ester content and the content diagram of four main ethyl esters C4-C8; E is the content of medium-chain ethyl esters C9-C 12 ; F is the content of long-chain ethyl esters C 14 -C 20 ;
[0038] Figure 6 is the acid content, ethyl ester content and calculated Qc value diagram of each ethyl esterification reaction within the aging years; among them, A is ethyl hexanoate; B is ethyl octanoate; C is ethyl oleate; D is ethyl propionate; E is ethyl 4-methylvalerate; F is ethyl 3-methylbutyrate; G is the Qc calculated values diagram of 13 other esterification reactions at different ages; H is the Qc value diagram calculated from the research data of 53 aged Baijiu samples in the literature.
[0039] Figure 7 is the ethyl esterification reaction concentration quotient Qc corresponding to short-chain, medium-chain and long-chain acids and the acid and ethyl ester content diagram of 53 vintage Baijiu; among them, A is the ethyl esterification reaction concentration quotient Qc diagram corresponding to self-tested short-chain acids; B is the ethyl esterification reaction concentration quotient Qc corresponding to self-tested medium-chain acids; C is the ethyl esterification reaction concentration quotient Qc corresponding to self-tested long-chain acids; D is the acid content diagram of 53 vintage Baijiu obtained from research in the literature; E is the corresponding ethyl ester content diagram of 53 vintage Baijiu obtained from research in the literature.
[0040] Figure 8It is a process diagram of key feature selection based on a machine learning model; among them, A is a schematic diagram of the process of obtaining key features by combining 3 models with 5 sorting algorithms; B is a graph of the F1 value of a random forest using 5×128 feature subsets; C is a graph of the F1 value of a support vector machine using 5×128 feature subsets; D is a graph of the F1 value of a neural network using 5×128 feature subsets; E is a graph of obtaining 41 key feature values by taking the intersection of 9 sub-optimal subsets.
[0041] Figure 9 It is a schematic diagram of the decision tree of a random forest.
[0042] Figure 10 It is a schematic diagram of the principle of a support vector machine.
[0043] Figure 11 It is a schematic diagram of the principle of a neural network.
[0044] Figure 12 It is a graph of the change rule of the content of compounds detected only in aged Baijiu.
[0045] Figure 13 It is an analysis graph of the relationship between key feature values and Baijiu storage time and flavor characteristics; among them, A is a spearman correlation analysis graph of 41 key feature values and Baijiu storage time, 8 blue lines indicate negative correlation, and the remaining red lines indicate positive correlation. The thickness of the lines represents the absolute value of the correlation coefficient; B is a graph of the sensory evaluation description scores of samples of different age groups; C is the average value of each key feature value in four age groups, as well as the statistical results and identifications of the spearman correlation coefficient |r|. |r| greater than 0.9 is marked with **, and |r| in the range of 0.8 - 0.9 is marked with *. Specific implementation mode
[0046] Specifically, the method for detecting the aging time of Baijiu based on GC-MS and a machine learning model includes the following steps:
[0047] A. Qualitatively analyze the compounds in the Baijiu samples with gradient aging time by GC-MS to obtain a qualitative data set of compounds in each wine sample, and quantitatively analyze the compounds in the qualitative data set to obtain a concentration data set of compounds in each wine sample;
[0048] B. Calculate the evenness index of the compounds in each Baijiu sample through chemometrics. The calculation formula is: where i is 1, 2,..., S, S is the total number of compounds in a certain Baijiu sample in step A, N is the total content of all compounds in this Baijiu sample, and N i is the content of a certain compound i in this Baijiu sample;
[0049] C. Based on the qualitative and quantitative data sets determined in step A, select the combinations of organic acids, alcohols (including ethanol), corresponding esters, and water that coexist in the esterification-hydrolysis reversible reaction for each baijiu sample in step A, and establish a reversible reaction equilibrium equation: Calculate the reaction concentration quotient Qc for each esterification-hydrolysis reversible reaction in each baijiu sample by taking the product of the powers of the stoichiometric coefficients of the product concentrations and dividing it by the product of the powers of the stoichiometric coefficients of the reactant concentrations. The calculation formula is: where C is the real-time concentration of each compound (including ethanol and water) in the reversible reaction equilibrium equation;
[0050] D. The concentration data set obtained in step A, the uniformity index obtained in step B, and the reaction concentration quotient obtained in step C are the characteristic values. Use a sorting algorithm to sort the importance of the characteristic values. Start taking values from the characteristic value with the highest ranking. Each time, add one more characteristic to the previous subset to form the next subset, obtaining (total number of characteristic values × total number of sorting algorithms) input characteristic subsets of different sizes. Input each characteristic subset into a machine learning model respectively, and let the machine learning model perform year discrimination to obtain (total number of characteristic values × total number of sorting algorithms × total number of machine learning models) discrimination F1 values. Compare all F1 value data with the screening threshold, retain M input characteristic subsets that are input into the machine learning model when F1 is greater than the screening threshold, and take the intersection of these M input data subsets to obtain the final key characteristic values.
[0051] The present invention uses the above method to sort the qualitative and quantitative data sets, uniformity, and reaction concentration quotient through a sorting algorithm and combine with a machine learning model for discrimination, and screen out the key characteristic values closely related to the identification of the aging time of baijiu. By detecting the key characteristic values of the baijiu to be tested with an unknown aging time and inputting them into at least one of the machine learning models used in step D, its aging time can be determined. In step D, to ensure the accuracy of the discrimination result, more than one machine learning model is used. When detecting the baijiu to be tested with an unknown aging time, the key characteristic values of the baijiu to be tested with an unknown aging time can be input into one of the machine learning models as needed, generally the model with the highest accuracy. For example, in the embodiments of the present invention, three models are used simultaneously when screening the key characteristic values, but when identifying an unknown liquor sample, its key characteristic values can be input into one of the random forest, support vector machine, and neural network. After verification, the accuracy of the neural network machine learning model in the embodiments is the highest, so generally the neural network machine learning model is used.
[0052] In step A of the present invention, the qualitative and quantitative situations of water and ethanol in the liquor samples with gradient aging time are determined. Therefore, step A is to qualitatively and quantitatively analyze all other compounds in the liquor samples with gradient aging time except water and ethanol, and obtain the qualitative data set and concentration data set of all other compounds except water and ethanol in each liquor sample. These qualitative data sets and concentration data sets are used to calculate the uniformity index in step B, select and calculate the reaction concentration quotient in step C, and perform sorting and discrimination in step D. In step A, the unit of concentration in the concentration data set has no effect on the calculation of the uniformity index and the detection result of the aging time, but in the same set of experiments, a unified unit must be adopted; generally, the concentration is in mg / L, and it is converted to mol / L for calculating Qc when needed in subsequent steps.
[0053] In step A of the present invention, common methods in the art can be used to qualitatively analyze the compounds in the liquor samples, such as: comparing with the NIST17 database (MS), and the similarity should be greater than 80; or, comparing the retention index (RI) of the compound to be qualitatively analyzed with the retention index (RIs) of the standard product; or, comparing RI with the retention index (RI lit ) of the compound in the reference, and the difference in retention index should be less than 30 (X. Zhang et al., 2021). Through qualitative analysis, the structures of the compounds except water and ethanol in each liquor sample are determined.
[0054] In step A of the present invention, common methods in the art can be used to quantitatively analyze the compounds except water and ethanol in the liquor samples. For example: by establishing an internal standard calibration curve, with the abscissa and ordinate being the peak area ratio and mass concentration ratio of the target compound to be measured and the corresponding internal standard substance respectively, and calculating the compound concentration through the standard curve; or, using internal standard relative quantification; etc. There are many types of compounds in the liquor samples, and the appropriate quantitative method should be selected according to the situation of each compound to ensure the accuracy of the quantitative results.
[0055] In step A of the present invention, to reduce the calculation amount while ensuring the accuracy of the detection results, when qualitatively analyzing the compounds in the liquor samples, compounds with a detection rate lower than 30% in the liquor samples with gradient aging time can be excluded according to the 30% rule, and only the remaining specific compounds are quantitatively analyzed; preferably, compounds with a detection rate lower than 50% in the liquor samples with gradient aging time are excluded according to the 50% rule. Adopting the 30% or 50% rule can greatly reduce the data calculation amount. At this time, the qualitative data set and concentration data set obtained in step A change accordingly. Similarly, when calculating the uniformity index, selecting the esterification hydrolysis reversible reaction combination and calculating the reaction concentration quotient, sorting with the sorting algorithm, and discriminating with the machine learning model subsequently, it can be carried out according to the remaining specific compounds.
[0056] In step A of the present invention, the design of the liquor samples with gradient aging time can be accurate to the time interval unit of years, and even accurate to months when the sample size is sufficient. Those skilled in the art can design the liquor samples with gradient aging time by combining common sense, the standard liquor samples, and the liquor samples to be identified. In step A, using the unsupervised hierarchical clustering algorithm to perform clustering analysis on the liquor samples with gradient aging time is an optional step. The embodiments of the present invention prove that liquor samples with certain specific different aging times will show similar compound compositions. Therefore, when the number of liquor samples is sufficient, it can be determined whether to carry out unsupervised clustering based on the standard that the sample size in each determination time period is not less than 3. When carrying out unsupervised clustering, the liquor samples with gradient aging time are designed based on the criteria of continuous and non-overlapping time intervals.
[0057] In step B of the present invention, the evenness index of the compounds in each liquor sample is calculated by chemometrics. The calculation formula is: where i is 1, 2, …… S, S is the total number of compounds in a certain liquor sample in step A, N is the total content of all compounds in this liquor sample, and N i is the content of a certain compound i in this liquor sample. S can be determined according to the qualitative data set in step A, and N and N i can be determined according to the quantitative data set in step A; S, N, and N i will change according to the results of step A. If the 30% or 50% rule is not carried out, the three are the data of all other compounds except water and ethanol. If the 30% or 50% rule is carried out, the three are the remaining specific compounds except water and ethanol. As in the embodiments of the present invention, after qualitative analysis, 50% rule, and quantitative analysis, although the total number of specific compounds in the liquor samples of the four time periods is 108, specifically for a certain sample, the compounds contained in this liquor sample need to be selected from the 108 compounds. Therefore, the S of this liquor sample may be less than 108. The content unit has no influence on the calculation of the evenness index and the detection result of the aging time, but in the same group of experiments, a unified unit must be used; generally, mg / L is used as the content unit.
[0058] The following is an example by hypothesis to illustrate how the present invention calculates the evenness index:
[0059] For example, assume that in a certain liquor sample, except for water and ethanol, only ethyl octanoate 1 mg / L, octanoic acid 2 mg / L, ethyl oleate 3 mg / L, oleic acid 4 mg / L, ethyl acetate 5 mg / L, and hexanoic acid 6 mg / L are detected qualitatively and quantitatively. At this time, S is 6, N is the total content of ethyl octanoate, octanoic acid, ethyl oleate, oleic acid, ethyl acetate, and hexanoic acid in this liquor sample, which is 21 mg / L, and N 辛酸 is the content of octanoic acid in this liquor sample, which is 2 mg / L. Based on this, the evenness index of the compounds in this liquor sample is calculated.
[0060] For another example, assume that in a certain wine sample, besides water and ethanol, only ethyl octanoate, octanoic acid, ethyl oleate, oleic acid, ethyl acetate, hexanoic acid, lauric acid, and myristic acid are qualitatively detected. After applying the 50% rule, lauric acid and myristic acid are excluded, and the remaining specific compounds are quantified. It is detected that the content of ethyl octanoate is 1 mg / L, octanoic acid is 2 mg / L, ethyl oleate is 3 mg / L, oleic acid is 4 mg / L, ethyl acetate is 5 mg / L, and hexanoic acid is 6 mg / L. At this time, S is 6, and N is the total content of 21 mg / L of ethyl octanoate, octanoic acid, ethyl oleate, oleic acid, ethyl acetate, and hexanoic acid in the wine sample. N 辛酸 is the content of octanoic acid in the white wine sample, which is 2 mg / L. Based on this, the uniformity index of the compounds in the wine sample is calculated.
[0061] During the aging process, various physical and chemical reactions can occur, including oxidation, esterification, hydrolysis, volatilization, and dissolution. Besides water and ethanol (~98%) in white wine, esters and acids account for the vast majority (>90%) of the remaining compounds, and they play an important role in the flavor of white wine. Therefore, it is of practical significance to pay attention to the thermodynamic equilibrium of the esterification reaction. Theoretically, the Gibbs free energy of the system at any spontaneous phase transition point must be lower than its initial value (Meunier, Scalbert, & Thibault-Starzyk, 2015). Therefore, thermodynamics can be used to predict the specific reaction direction and rate, that is, to calculate the reaction concentration quotient Qc based on the ratio of reactants and products present in the system and compare it with the corresponding reaction equilibrium constant Kc. Revealing the chemical reaction kinetics during the aging process of white wine helps to deepen the scientific understanding of the aging process of white wine, helps to predict the compound composition and sensory quality of aged white wine, and helps to develop a method for identifying genuine and fake aged white wine.
[0062] In step C of the present invention, in the reversible reaction equilibrium equation, a…x, b…y are stoichiometric coefficients with the smallest integers, and in most cases, they are all 1. When calculating Qc, the real-time concentrations of each compound need to be unified into units of mol / L. If the system under study has reached thermodynamic equilibrium, the value of the reaction concentration quotient is equal to the reaction thermodynamic equilibrium constant Kc by definition.
[0063] Although the wine samples in the embodiments of the present invention all contain the same 19 esterification-hydrolysis reversible reactions and corresponding Qc values in the four time periods, in fact, the differences among the wine samples can more significantly distinguish the wine samples. Therefore, when selecting the combinations of organic acids, alcohols, corresponding esters, and water that coexist in the esterification-hydrolysis reversible reactions in each white liquor sample from the white liquor samples with gradient aging times in step A, there will be a situation where some wine samples have a certain esterification-hydrolysis reversible reaction, while some wine samples do not have this esterification-hydrolysis reversible reaction. For the wine samples without this esterification-hydrolysis reversible reaction, the Qc value is filled with 0. However, in order to reduce the calculation amount and ensure the accuracy of the discrimination result, it is also possible to only screen out the ethyl esterification-hydrolysis reversible reaction and its Qc value.
[0064] The following is an example by assumption to illustrate how to calculate Qc in the present invention:
[0065] For example, assume a set of wine samples with gradient aging times (two samples, both with an ethanol concentration of 52 Vol%, and water being 48%. After corresponding unit conversion during the calculation of Qc, the molar concentrations are 8.9 mol / L for ethanol and 26.7 mol / L for water). In wine sample 1, caprylic ethyl ester 1 mol / L, caprylic acid 2 mol / L, oleic ethyl ester 3 mol / L, oleic acid 4 mol / L, ethyl acetate 5 mol / L, hexanoic acid 6 mol / L, lauric acid, and myristic acid are qualitatively and quantitatively detected. This wine sample 1 has three esterification-hydrolysis reversible reaction equilibria. Among them, The Qc of caprylic ethyl ester = (1 mol / L 1 × 26.7 mol / L 1 ) / (2 mol / L 1 × 8.9 mol / L 1 ). The esterification-hydrolysis reversible reaction equilibria and Qc calculations of oleic ethyl ester and ethyl acetate are the same. In wine sample 2, caprylic acid 2 mg / L, oleic ethyl ester 3 mg / L, oleic acid 4 mg / L, ethyl acetate 5 mg / L, hexanoic acid 6 mg / L, lauric acid, and myristic acid are qualitatively and quantitatively detected. This wine sample 2 has two esterification-hydrolysis reversible reaction equilibria and does not have the esterification-hydrolysis reversible reaction equilibrium of caprylic ethyl ester. The Qc of caprylic ethyl ester is filled with 0, and the esterification-hydrolysis reversible reaction equilibria and Qc calculations of oleic ethyl ester and ethyl acetate are the same as those of wine sample 1.
[0066] Combining machine learning with sorting algorithms is a practical and simple method to solve the difficult problem of liquor age identification, and it has good effects on noise removal and clustering prediction for complex input data. The importance of features for discrimination, or the correlation between features, can be scored by sorting algorithms such as Information Gain (IG), Analysis of Variance (ANOVA), and Feature Weight Algorithm based on Dispersion or Correlation (Relief F). Then, many subsets with different features can be generated according to the sorting scores and evaluated by the model. Usually, the Wrapper method combined with the forward stepwise search method is used to generate subsets. Input each subset in turn, and when the discrimination accuracy of the model reaches the set threshold, an optimal feature subset will be obtained. The features in this optimal subset are usually determined as key features, and these features usually have stronger potential for year identification and will deepen the understanding of the aging mechanism.
[0067] In step D of the present invention, the sorting algorithm can select common sorting algorithms in the art, such as Information Gain, Information Gain Ratio, GiniDecrease, Analysis of Variance (ANOVA), Chi-Square Test (χ 2 )), Feature Weight Algorithm (ReliefF), Fast Correlation-Based Filter (FCBF), etc.; in order to ensure the accuracy of the discrimination result, generally at least three of them need to be selected for importance sorting. Preferably, in step D, the sorting algorithms are Information Gain, Information Gain Ratio, Analysis of Variance, Feature Weight Algorithm, and Fast Correlation-Based Filter, a total of five.
[0068] In step D of the present invention, the machine learning model can select common machine learning models in the art, such as Random Forest, Support Vector Machine (SVM), Neural Network, K-Nearest Neighbors (KNN), Gradient Boosting, eXtreme Gradient Boosting, Linear Regression, Logistic Regression, AdaBoost, Stochastic Gradient Descent, etc.; in order to ensure the accuracy of the discrimination result, generally at least two of them need to be selected for discrimination. Preferably, in step D, the number of machine learning models is Random Forest, Support Vector Machine, and Neural Network, a total of three.
[0069] In step D of the present invention, the closer the screening threshold is set to 1, the more concise the key feature values screened out are; generally, it is sufficient to set F1 greater than 0.9. However, in special cases, such as in the embodiments of the present invention, when the prediction capabilities of all models are very high, the model with the lowest F1 value is taken as the main consideration factor. Taking the embodiments of the present invention as an example, specifically, the following considerations are required: The prediction accuracies of the random forest and neural network are higher than that of SVM. Even for SVM, with its F1 > 0.9, it also shows great potential in predicting the age of Baijiu. Therefore, when designing the screening threshold, the F1 value of SVM is taken as an important consideration factor. The F1 value of SVM first rises, then stabilizes, and then drops, which meets the theoretical requirements of feature screening, that is, when the number of features is too small, the prediction accuracy of the model is relatively low, but when there are too many, it will introduce noise, resulting in a decrease in accuracy. By observing the F1 values corresponding to the ANOVA and FCBF sorting algorithms, it is found that when the number of feature values increases by more than 5, the F1 that satisfies the continuous rise or stability at a certain value is 0.946276. Taking its two significant figures, the screening threshold is set to 0.94. Therefore, if the F1 fluctuates violently (the number of times F1 is lower than 0.94 exceeds 3 times) after the first 32 (the first 25%) features of a certain sorting algorithm are input into a certain discrimination model, it is considered that the matching between this sorting algorithm and this model is poor, and the input feature subset (sub-optimal subset) is no longer obtained from this combination for subsequent union operations. If the F1 value of the subset composed of the first i features is greater than 0.94, while the F1 value of the first i + 1 subset is less than 0.94, then the subset composed of the first i features is defined as the sub-optimal subset (i ≥ 32). According to the above criteria, a total of 9 sub-optimal subsets are screened out, and their intersection is taken to obtain 41 key features.
[0070] In specific discrimination tests, those skilled in the art can appropriately adjust the screening threshold according to the sorting algorithm and machine learning model adopted to simultaneously ensure obtaining an appropriate number of key feature values and the accuracy of the discrimination (detection) results.
[0071] In the embodiments of the present invention, 54 samples of Baijiu with gradient aging times from 0 to 11 years are selected as Baijiu samples with gradient aging times, and the Baijiu samples with gradient aging times are divided into four groups: [0, 1) year, [1, 5) years, [5, 9) years, and [9, 11] years through cluster analysis.
[0072] In step A of the present invention, in order to achieve the accurate qualitative and quantitative analysis of various compounds in liquor, headspace solid-phase microextraction, liquid-liquid microextraction or liquid-liquid microextraction-BSTFA derivatization is used to pretreat the liquor samples. The operation of the headspace solid-phase microextraction is as follows: Dilute the liquor sample to a final ethanol content of 5-15 Vol%, add NaCl, ethyl lactate and compounds with retention times longer than acetic acid are all internal-standardized with 2-methylhexanoic acid, alcohols with retention times shorter than acetic acid are internal-standardized with tert-amyl alcohol, and other compounds are internal-standardized with amyl acetate. Use an SPME fiber to extract volatile compounds from the top space of the sample. The sample is pre-equilibrated at 45-60 °C for 0-10 min, and then extracted at 45-60 °C for 30-60 min for subsequent GC-MS determination. The operation of the liquid-liquid microextraction is as follows: Add three internal standards, namely amyl acetate, 2-methylhexanoic acid and tert-amyl alcohol, and saturated sodium chloride solution to the liquor sample, then add solution A, stir well for more than 3 min, collect the upper organic phase after static stratification, concentrate the organic phase by nitrogen blowing at room temperature, and then perform subsequent GC-MS determination; solution A is a mixed solvent of anhydrous ether:pentane with a volume ratio of 1:1. The operation of the liquid-liquid microextraction-BSTFA derivatization is as follows: Add heptadecanoic acid internal standard to the liquor sample, then add saturated sodium chloride solution, dilute to an ethanol concentration of about 10%, add solution A to the diluted system for extraction, vortex for more than 3 min, let stand for more than 20 min, collect the organic layer, extract the remaining aqueous phase with acetonitrile and dichloromethane in turn, combine the organic phases, add pyridine containing 2.5% v / v hydroxylamine hydrochloride to the organic phase, add BSTFA containing 1% TMCs, vortex for more than 5 s, incubate in a metal bath at 45-85 °C for 1-5 h, after the reaction is completed, centrifuge at 6000-12000 rpm for more than 3 min, take the supernatant for subsequent GC-MS determination; solution A is a mixed solvent of anhydrous ether:pentane with a volume ratio of 1:1.
[0073] In step A of the present invention, when headspace solid-phase microextraction is adopted, the GC-MS conditions are as follows: DB-WAX chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8-1.4 mL / min; electron impact mode: 70 eV; the transfer line temperature at the connection between the chromatographic column and the ion source is set at 200-230 °C; scanning mode, the scanning range is m / z 33-350; the inlet temperature is 250-300 °C, the desorption time is 5-15 min, and the split ratio is 2-20:1; temperature programming: starting from 50 °C, increasing in stages to 235 °C and holding for more than 2 min. When liquid-liquid microextraction is adopted, the GC-MS conditions are as follows: DB-WAX chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8-1.4 mL / min; electron impact mode: 70 eV; the transfer line temperature at the connection between the chromatographic column and the ion source is set at 200-230 °C; scanning mode, the scanning range is m / z 33-350; the sample injection volume is 1 μL, the inlet temperature is 250-300 °C, the desorption time is 5-15 min, and the split ratio is 2-20:1; temperature programming: starting from 35 °C, rising in stages to 235 °C and holding for more than 10 min. In step A, when liquid-liquid microextraction-BSTFA derivatization is adopted, the GC-MS conditions are as follows: HP-5MS chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8-1.4 mL / min; electron impact mode: 70 eV; the transfer line temperature at the connection between the chromatographic column and the ion source is set at 200-230 °C; scanning mode, the scanning range is m / z 33-350; the sample injection volume is 1 μL, the inlet temperature is 250 °C, the split ratio is 2-20:1; temperature programming: gradually heating from 65 °C to 280 °C and holding for more than 2 min.
[0074] In the embodiments of the present invention, by cooperating the target compound with a specific pretreatment process and GC-MS conditions, it is possible to ensure the accurate quantitative detection of the concentrations of up to 108 compounds in liquor; and the concentrations of each compound are used to calculate the evenness index Evenness index and Qc of the compounds in the liquor sample, which is beneficial to realizing the identification of the aging time of liquor and improving the accuracy of the identification results.
[0075] In step D of the present invention, as shown in the embodiments, when taking four groups of Luzhou-flavor liquors with [0,1) year, [1,5) years, [5,9) years, and [9,11] as samples, before the importance ranking by the sorting algorithm, the eigenvalues include the concentrations of 108 compounds in the concentration dataset of the compounds in step A, 1 evenness index of the compounds in step B, and 19 reaction concentration quotients Qc in step C.
[0076] In step D of the present invention, as shown in the examples, when taking four groups of Luzhou-flavor liquors with ages in the ranges of [0,1) year, [1,5) years, [5,9) years, and [9,11] years as samples, the key characteristic values include 36 compound concentrations screened from the compound concentration dataset in step A, 1 evenness index of the compounds in step B, and 4 Qc values screened from the reaction concentration quotient Qc in step C.
[0077] In the embodiments of the present invention, first, a gas chromatography-mass spectrometry (GC-MS) instrument is used to collect the compound content in premium Luzhou-flavor liquors. Subsequently, the evenness of the compounds in the liquor is calculated by chemometrics. Then, the esterification-hydrolysis reversible reactions in each liquor sample are screened out and the ethyl esterification reaction concentration quotient (Qc) is obtained. After combining the above three datasets, 41 characteristic values closely related to the age identification of the liquor are screened out from them by five ranking algorithms, namely Information Gain (IG), Information Gain ratio (IGR), Analysis of Variance (ANOVA), the feature weights algorithm (Relief F), and Fast correlation-based Filter (FCBF), in combination with three machine learning models, namely Random forest, Support Vector Machine (SVM), and Neural network, and the accuracy of their age identification is verified. Using the 41 characteristics, the age of the liquor sample can be identified by Random forest, Support Vector Machine, or Neural network. In particular, the accuracy of the Neural network can reach 100%.
[0078] It can be seen that the present invention deeply understands the aging mechanism of Chinese liquor and establishes a chemometric method for predicting the composition of flavor compounds after aging of liquor. The systematic research idea of the present invention, which combines volatile flavor analysis (GC-MS), real-time reaction concentration quotient Qc, compound evenness index, ranking algorithm, and machine learning, is not limited to Luzhou-flavor liquors with an aging time of 0 - 11 years, but also applicable to Luzhou-flavor liquors with a longer aging time, as well as the age identification of other-flavor liquors or other distilled spirits.
[0079] The present invention will be further described in detail below through examples, but the protection scope of the present invention is not limited to the scope of the described examples.
[0080] Materials and samples:
[0081] Sodium chloride, anhydrous ether, n-pentane, acetonitrile, anhydrous ethanol, and dichloromethane used for salting out or extraction were all from Shanghai Sinopharm Chemical Reagent Co., Ltd., China. 40 The retention indices (RIs) of the compounds were calculated using a mixture of n-alkanes (Sigma-Aldrich, Shanghai, China). The four internal standards, amyl acetate (IS1), 2-methylhexanoic acid (IS2), tert-amyl alcohol (IS3), and heptadecabutyric acid (IS4), and 98 standards (see Table 1 and Table S2) were purchased from Sigma-Aldrich or Aladdin Reagent Co., Ltd. (Shanghai, China) with a purity higher than 98%.
[0082] The authentic bottled aged liquor samples were provided by the preservation center of Sichuan Luzhou Laojiao Co., Ltd. All samples were premium liquor (Guojiao 1573) with the same alcohol content (52%, vol). The samples of different years (0-11 years) were well packaged and stored at room temperature.
[0083] Statistical analysis:
[0084] Each sample was analyzed three times. Analyses and visualizations were performed using R 3.6.2 (R Foundation for Statistical Computing, Vienna, Austria), GraphPad Prism 8.0.0 for Windows (GraphPad Software Inc., San Diego, CA, USA), Excel 2019 (Microsoft Corporation, USA), and Orange 14.1.2016 (University of Ljubljana, Republic of Slovenia). Similarity analysis (ANOSIM) was performed using R. Principal component analysis (PCA), hierarchical clustering (HCA), and heat maps were performed using SIMCA 14.1 (Umetrics, Sweden) and online websites to assess differences in compound concentrations in samples of the four wine age groups. Spearman correlations between compounds and wine storage time were calculated using R, and network diagrams were drawn using Cytoscape 3.8.2 (Shannon et al., 2003).
[0085] Example 1
[0086] Cluster analysis based on liquor aging time
[0087] In order to better and objectively understand the clustering of samples of different ages, this example is a sample of 54 samples (each sample aging time is shown in Figure 1 or Figure 2 The data on the horizontal axis are all clustered using an unsupervised hierarchical clustering algorithm (HCA). The results are as follows Figure 1 As shown. Based on Spearman ( Figure 1 ) and Euclidean(Figure 2 ) The HCA of Figure 3 A) was highly consistent with the clustering results. When 108 compounds were used as features, the samples could be divided into 4 groups. The results of PCA analysis ( Figure 1 ). Group I was the samples with an aging time of [0,1) years, and the samples with aging times of [1,5) years, [5,9) years, and [9,11) years were Group II, Group III, and Group IV, respectively. The ANOSIM test (Table 1) showed significant differences among the four groups.
[0088] Table 1 Analysis of the significance of differences in the compound composition of Chinese liquors of different ages by the ANOSIM method
[0089]
[0090] Similar age laws have also been presented in previous reports. Specifically, the material composition of Chinese liquor changes greatly within 1 - 3 years, the material compositions of the samples stored for 3 years and 4 years are similar, the samples stored for more than 5 years are quite different from the younger samples, while the samples stored for 5 years and 7 years are similar, and the samples of 9 years and 12 years show similar compound compositions (Dai et al., 2021; M. Xu, Yu, Ramaswamy, & Zhu, 2017; Zheng et al., 2021).
[0091] The research results of this example show that the storage time has an impact on the quality of Chinese liquor, and it is theoretically feasible to distinguish Chinese liquor samples with 4 groups of [0,1) years, [1,5) years, [5,9) years, and [9,11) years. Therefore, in this invention, the Chinese liquor stored for [0,1) years is used as new liquor, and the Chinese liquors of other years are aged Chinese liquors. The samples with aging times of [1,5) years, [5,9) years, and [9,11) years are Group II, Group III, and Group IV, respectively.
[0092] Example 2
[0093] 1. Pretreatment and GC-MS detection method:
[0094] (1). GC-MS combined with headspace solid-phase microextraction (HS-SPME):
[0095] The baijiu sample was diluted to a final ethanol content of 8% (6 mL), placed in a 20 mL glass bottle, and 2.0 g of NaCl was added. Ethyl lactate and compounds with a retention time longer than acetic acid were both internal-standardized with 2-methylhexanoic acid (10 μL, 14.00 g / L). Alcohols with a retention time shorter than acetic acid were internal-standardized with tert-amyl alcohol (10 μL, 8.05 g / L), and other compounds were internal-standardized with amyl acetate (10 μL, 10.33 g / L). Volatile compounds were extracted from the headspace of the sample using a PAL 3 autosampler (Zhejiang Alltech Analytical Instruments Co., Ltd., China) and an SPME fiber (80 μm thick, 10 mm long, DVB / C-WR / PDMS) (Agilent Technologies, Inc., USA). The sample was equilibrated at 60 °C for 5 min and then extracted at 60 °C for 40 min.
[0096] GC-MS: Analysis was performed using a 7890B gas chromatograph (GC), a 5977B mass spectrometer (MS) (Agilent Technologies, Inc., USA), and a DB-WAX (30 m × 0.25 mm × 0.25 μm) chromatographic column. The carrier gas was helium (purity 99.999%), and the flow rate was 1.2 mL / min. The MS was operated in electron impact (EI) mode (70 eV). The transfer line at the connection between the chromatographic column and the ion source required auxiliary heating, and the temperature was set to 230 °C. Detection was carried out in scan mode, with a range of m / z 33 - 350. The inlet temperature was 300 °C, the desorption time was 15 min, and the split ratio was 4:1. The temperature program started at 50 °C, held for 2 min, then increased to 145 °C at 3 °C / min and to 235 °C at 15 °C / min, and held for 8 min.
[0097] (2), GC-MS combined with liquid-liquid microextraction (LLME):
[0098] Transfer 4 mL of the baijiu sample to a 30 mL glass centrifuge tube, add 10 μL of each of the 3 internal standards (10.33 g / L amyl acetate, 14.00 g / L 2-methylhexanoic acid, and 10.22 g / L tert-amyl alcohol) and 14 mL of saturated sodium chloride solution. Then add 1.5 mL of solution A (anhydrous ether:pentane = 1:1, v / v), stir well for more than 3 min, and collect the upper organic phase after static stratification. The organic phase was concentrated to 250 μL by nitrogen blowing at room temperature.
[0099] GC-MS: Determination was carried out using an Agilent 7890B-5977B GC-MS. The chromatographic column was DB-WAX (30 m × 0.25 mm × 0.25 μm, Agilent Technologies). 1 μL of the concentrated organic phase was injected, and the split ratio was 4:1. The gas chromatography temperature program was as follows: hold at 35 °C for 0.5 min, increase to 50 °C at a rate of 10 °C / min and hold for 4 min, increase to 100 °C at a rate of 3 °C / min and hold for 3 min, increase to 240 °C at a rate of 3 °C / min and hold for 23 min. The carrier gas was helium at 1.4 mL / min (purity 99.999%). Other parameters were the same as in 1.(1).
[0100] (3), GC-MS combined with liquid-liquid microextraction (LLME)-BSTFA derivatization:
[0101] Transfer 4 mL of the liquor sample to a 30 mL glass sample vial, and add 100 μL of the internal standard (heptadecanoic acid at 0.0308 g / L). Then add 14 mL of saturated sodium chloride solution to dilute the liquor to about 10%. Add 1.5 mL of mixed solution A to the diluted sample for extraction, vortex for 3 min, and let stand for more than 20 min. Transfer the organic layer (containing acids) to a new test tube. The remaining aqueous phase was extracted successively with 1.5 mL of acetonitrile and 1.5 mL of dichloromethane. Take the organic layers, mix them with the previous organic layers, blow to dryness with nitrogen and then further derivatize. Add 50 μL of pyridine (containing 2.5% hydroxylamine hydrochloride, v / v) to dissolve the non-volatile compounds, add 200 μL of BSTFA (Bis(trimethylsilyl)trifluoroacetamide, containing 1% TMCs), cover tightly, and vortex for 5 s. Incubate in a metal bath at 55 °C for 5 h. After the reaction, centrifuge at 12,000 rpm for 3 min, transfer the supernatant to a sample vial for GC-MS determination.
[0102] GC-MS: The chromatographic column was HP-5MS (30 m × 0.25 mm × 0.25 μm, Agilent Technologies, USA). The sample injection volume was 1 μL, the inlet temperature was 250 °C, and the split ratio was 5:1. The gas chromatography temperature program was as follows: hold at 65 °C for 2 min, increase the temperature to 280 °C at a rate of 6 °C / min, and hold for 8 min. The carrier gas was helium at 1.0 mL / min (purity 99.999%). Other parameters were the same as in 1.(1).
[0103] 2. Detection limit, quantification limit and precision:
[0104] The limits of detection (LOD) and quantification (LOQ) are the concentrations corresponding to signal-to-noise ratios of 3 and 10, respectively. Based on the results of the mixed standard determination closest to the true concentration of the Chinese liquor, the intra-day precision was calculated by repeating the determination three times within one day, and the inter-day precision was calculated after repeating the determination on different three days. The precision was calculated using the mean value, standard deviation, and relative standard deviation (%) of the measured values.
[0105] 3. Qualitative analysis of compounds:
[0106] The identification of volatile compounds was based on two or three qualitative methods. One was to compare with the NIST17 database (MS), and the similarity should be greater than 80. One was to compare the retention index (RI) of the compound to be identified with that of the standard (RIs), and the other was to compare RI with the retention index of the compound in the reference (RI lit ), and the difference in retention index should be less than 30.
[0107] Specific operation: The samples were determined using various extraction and pretreatment methods such as HS-SPME, LLME, and LLME-BSTFA derivatization. Then, qualitative analysis was performed based on the NIST library matching degree and RI similarity of the compounds. The RI values of 98 compounds were referenced with the standards, and those of other compounds were referenced with the literature. A total of 234 compounds were preliminarily identified. Next, data processing was carried out using the modified "50% rule". When the detection rate of a compound in all samples was less than 50%, the compound was removed from the dataset. Therefore, a dataset of 108 compounds was obtained, as Figure 1 shown.
[0108] 4. Quantitative analysis of compounds:
[0109] A mixed standard solution was prepared by dissolving 91 standard compounds in a 52% ethanol solution prepared from chromatographically pure ethanol and ultrapure water, and was gradient-diluted into six concentration levels of mixed standard stock solutions. The content ratio of each compound should be close to that of the Chinese liquor sample. The method for determining the standard curve was exactly the same as that for sample detection. An internal standard calibration curve was established, with the abscissa and ordinate being the peak area ratio and mass concentration ratio of the target compound to be measured and the corresponding internal standard substance, respectively.
[0110] Due to the rarity of the aged Baijiu samples analyzed in this invention, the LLME with the minimum sample volume (4 mL) was used. In addition, the LLME-BSTFA derivatization method was adopted to determine the non-volatile acids in Baijiu. As is well known, water will affect the derivatization efficiency of BSTFA. In this study, LLME combined with nitrogen blowing and drying provided an anhydrous condition for BSTFA derivatization, which saved more time than water removal by rotary evaporation. The LLME-BSTFA-GC-MS method can accurately quantify ethyl laurate, ethyl myristate, ethyl palmitate, ethyl stearate, ethyl oleate, lauric acid, myristic acid, palmitic acid, stearic acid, linoleic acid, etc.
[0111] The standard curves of 73 compounds were established with a mixed standard as shown in Table 2, and the relative quantification of the other 35 compounds was carried out using internal standards as shown in Table 3.
[0112] Table 2 Quantitative standard curves of three pretreatment methods
[0113]
[0114]
[0115]
[0116] Note: a indicates using amyl acetate as the internal standard, b indicates using 2-methylhexanoic acid as the internal standard, c indicates using tert-amyl alcohol as the internal standard, d indicates using heptadecanoic acid as the internal standard, e indicates choosing HS-SPME for pretreatment, f indicates choosing LLME for pretreatment, g indicates choosing LLME-BSTFA for pretreatment, and h indicates that the RSD% is calculated from the standard product gradient closest to the true concentration of Baijiu.
[0117] Table 3 Internal standards corresponding to the relative quantification of compounds and qualitative methods
[0118]
[0119]
[0120] Note: a: IS1 is amyl acetate, IS2 is 2-methylhexanoic acid, IS3 is tert-amyl alcohol, and IS4 is heptadecanoic acid. b: MS, qualitative by comparing with the NIST spectral library; RI, qualitative by calculating the retention index of the compound, and the reference value can be queried through the website; S, qualitative by the standard product.
[0121] Example 3: Compound uniformity index of aged Baijiu
[0122] Overall, the total acid content increases with the extension of the aging time, while the total ester content shows the opposite trend (see Figure 3C). Esters, especially ethyl hexanoate, contribute to the pungency of Chinese liquor, while acids are beneficial for suppressing pungency. The increase in acids and the decrease in esters during the aging process may contribute to the reduction of pungency in aged Chinese liquor. In addition, the total amounts of compounds such as alcohols, aldehydes, pyrazines, terpenes, and ketones generally show an upward trend (see Figure 3 C). The amount of pyrazine compounds detected in samples aged over 5 years is significantly increased. These compositional changes may be due to a series of redox reactions during the aging process.
[0123] The number of compound species increases after 5 years of storage (see Figure 4 ). The species and contents of compounds in Chinese liquor vary with the aging time, and the uniformity of the compound composition of Chinese liquor also changes. To evaluate the uniformity, we adopted the Evenness index. In the present invention, the evenness index reflects the evenness of components in Chinese liquor from 0 (when there is only one compound in the system) to 1 (when the concentrations of multiple compounds are the same), and the closer the value is to 1, the more uniform it is.
[0124] The evenness index of compounds in Chinese liquor shows a linear positive correlation with the aging time (R 2 = 0.7012), increasing from 0.55 to 0.59 (see Figure 3 D and Table 4). The evenness of compounds shows an increasing trend with the aging of Chinese liquor. Therefore, the evenness index is used as an important input feature for vintage identification in subsequent analysis. The uniformity of Chinese liquor aging may be related to two factors: 1) a larger variety of flavor compounds, and 2) a smaller concentration difference between these compounds. The change in the structure of flavor compounds in aged Chinese liquor may be one of the reasons for the rich and harmonious sensory characteristics of aged Chinese liquor.
[0125] Table 4 Evenness index of compounds in Chinese liquor samples of different vintages
[0126] Month age (month) Age (year) Uniformity index Month age (month) Age (year) Uniformity index 4.5 0.38 0.555880 57.5 4.79 0.557221 7.0 0.58 0.550408 62.0 5.17 0.569318 10.0 0.83 0.552593 67.5 5.63 0.570448 20.0 1.67 0.558795 81.0 6.75 0.566001 20.5 1.71 0.557142 98.0 8.17 0.567247 32.0 2.67 0.572681 109.0 9.08 0.567247 35.5 2.96 0.566869 115.0 9.58 0.584498 43.5 3.63 0.570026 118.0 9.83 0.588906 45.0 3.75 0.577795 128.0 10.67 0.585145
[0127] Example 4: There is a correlation between the change patterns of acids or esters and their carbon chain lengths during the aging process
[0128] Except for the samples aged 1 - 5 years, the total content of the determined compounds generally shows an increasing trend with the increase of aging time ( Figure 3 E). Among the volatile compounds detected in Chinese liquor samples, the total contents of esters and acids account for about 95% ( Figure 3 F). In addition, acids with 2 - 18 carbons (C2 - C18) and their corresponding ethyl esters account for about 98% of the determined acids and esters. Therefore, thermodynamic analysis focusing on ethyl esters and acids may be helpful for understanding the aging mechanism.
[0129] The present invention analyzes the variation rules of esters and acids with different chain lengths. Esters and acids are classified according to the chain length of fatty acids. Acids and their corresponding esters with fatty acid chain lengths of 2 - 6, 7 - 10, and 12 - 18 carbons are regarded as short-chain, medium-chain, and long-chain, respectively.
[0130] The total acid content is slightly higher than the content of short-chain fatty acids (SCFAs), slightly decreases in the first 2 years of aging, then rises to the 6th year, and then decreases again ( Figure 5 A). Although the content of medium-chain fatty acids (MCFAs) fluctuates greatly between 3 - 6 years, there is an obvious upward trend overall ( Figure 5 B). The content of long-chain fatty acids (LCFAs) shows a sinusoidal fluctuation and increases slightly overall ( Figure 5 C). The total ethyl ester content ( Figure 5 D) shows a downward trend overall, especially in the first 4 years, and basically remains stable after 5 years. The contents of four short-chain fatty acid ethyl esters (including ethyl acetate, ethyl butyrate, ethyl hexanoate, and ethyl lactate), which were previously considered the most important flavor contributors, account for 98 - 99% of the total ethyl esters, and show a downward trend overall. While the contents of medium-chain fatty acid ethyl esters ( Figure 5 E) and long-chain fatty acid ethyl esters ( Figure 5 F) fluctuate greatly and increase overall with age. These results indicate that during the aging process, there is a correlation between the variation rules of the contents of acids and ethyl esters and their carbon chain lengths.
[0131] Therefore, the variation rules of esters and acids with different chain lengths are used as an important input feature for year identification in the subsequent analysis.
[0132] Example 4: Esterification kinetics is a key influencing factor for the change of acid and ester contents during the aging process
[0133] Ester and acid compounds in Chinese liquor are the main components of flavor compounds, and their contents change during the aging process. To better understand the aging mechanism, we studied the equilibrium of esterification and hydrolysis according to the reaction concentration quotient (Qc) of the esterification reaction. The difference between Qc and the thermodynamic equilibrium constant (Kc) can be used to judge the direction of each reversible esterification reaction. If the esterification reaction reaches thermodynamic equilibrium, the value of Qc is equal to the thermodynamic equilibrium constant (Kc) of the reaction. The Kc values of the esterification of ethyl acetate and ethyl butyrate are 4.7 and 2.28, respectively, and the Kc value of the esterification of branched-chain monocarboxylic acids and primary alcohols is about 4. In addition, in the esterification reaction of a given alcohol, the Kc value of esters with fatty acid lengths greater than ethyl butyrate is less than the Kc of ethyl butyrate, 2.28, because the steric hindrance of long-chain acids is larger and the reaction activation energy is higher. Therefore, the Kc of the esterification reaction of short-chain ethyl esters is estimated to be between 2.0 - 4.7, while the Kc of the esterification reaction of medium-chain and long-chain ethyl esters is estimated to be lower than 2.28.
[0134] When the esterification reaction Qc is greater than Kc, the reversible ethyl esterification reaction is mainly ester hydrolysis, and vice versa. In the newly brewed liquor stored for 0.38 years, the Qc and Kc values of the esterification reactions of several ethyl esters differ greatly, indicating that the reaction is far from reaching equilibrium( Figure 7 and Table 5). However, these Qc values tend to approach Kc (<4.7) during the aging process, indicating that the acid / ester conversion tends to reach an equilibrium state. For example, in the newly brewed liquor, the Qc values of ethyl hexanoate( Figure 6 A), ethyl octanoate( Figure 6 B), and ethyl oleate( Figure 6 C) are all higher than their Kc values. The Qc values of these esterification reactions decrease towards Kc within 5 years and then change slowly, reflecting that ester hydrolysis dominates during the aging process. On the contrary, the Qc values of ethyl propionate( Figure 6 D), ethyl 4-methylvalerate( Figure 6 E), and ethyl 3-methylbutyrate( Figure 6 F) are lower than Kc at the initial stage of aging and show an upward trend during the aging process, especially in the first 5 years of aging, indicating that esterification dominates. This causes the gradual accumulation and increase in the content of ethyl propionate, ethyl 3-methylbutyrate, and ethyl 4-methylvalerate during the aging process. Similarly, the Qc values of the other 13 ethyl esterification reactions all tend to approach Kc( Figure 6 G) during the aging process.
[0135] Figure 6 and Figure 7 In, the difference between the reaction quotient (Qc) and the equilibrium constant (Kc) determines the equilibrium position; the grey shading represents the range of observed Kc values.
[0136] In summary, although the changes in the content of each ester or acid during the aging process are related to their chain lengths, the 19 reversible ethyl esterification reactions all gradually tend to reach an equilibrium state with aging. In addition, the esters with a content advantage in the newly brewed liquor (such as ethyl hexanoate) are partially replaced by esters with a lower initial concentration (such as ethyl propionate) after aging. Generally speaking, driven by thermodynamics, a higher degree of uniformity and a more balanced material composition seem to be the general results of liquor aging.
[0137] To verify the above inferences and further explore the relationship between Qc and Kc during the aging process, the present invention conducted a Meta-analysis based on the research data of predecessors. Specifically, the original data of the concentration of ethanol, fatty acids, esters, and aging duration in the previously reported liquor aging studies were used to calculate the Qc of each ester. The re-analyzed samples included Luzhou-flavor liquor and Maotai-flavor liquor with different alcohol contents and quality grades. Figure 6 and Figure 7Among them, St is the abbreviation of Luzhou-flavor Baijiu, Sa is the abbreviation of Maotai-flavor Baijiu, PG is the abbreviation of premium-grade Baijiu, and FG is the abbreviation of first-grade Baijiu. For example, St_52_PG is premium-grade Luzhou-flavor Baijiu with an alcohol volume content of 52%. From Figure 6 H, Figure 7 It can be seen that although the concentrations of compounds in different Baijiu samples vary greatly, as the aging time prolongs, the calculated Qc values of 53 Baijiu samples almost all tend to the corresponding Kc values. These results indicate that according to this classical thermodynamic mechanism, the differences in the initial Qc of the esterification reaction may cause the reversible reaction to proceed in different reaction directions, resulting in inconsistent variation rules of the content of the same compound in different samples during aging. However, the rule that Qc approaches Kc is consistent. The distance between Qc and Kc determines whether the content of a specific acid or its ester increases or decreases during aging. A potential practical application of these findings is that the content of substances after aging can be predicted based on the distance between Qc and Kc in new Baijiu.
[0138] Table 5 Average values of 19 Qc of esterification reactions in samples of four age groups
[0139]
[0140]
[0141] Note: The values are mean ± standard deviation. Samples with different letters in the upper right corner of the values are significantly different (P < 0.05).
[0142] Example 5: Discrimination of the age of Baijiu based on eigenvalue screening engineering and machine learning
[0143] After Examples 1 to 4, the present invention used a total of 128 features for the identification of the age of Baijiu, including 108 compound concentrations (see Figure 1 , Tables 1 to 2), 1 uniformity index (see Figure 3 D), and 19 Qc values (see Figure 7 and Table 5). In order to remove noise, redundancy, and irrelevant data, eigenvalue screening was carried out. Five sorting algorithms were used to sort all features, and the wrapper method was used to input each feature set into 3 models respectively. Three models were used to evaluate the selected features to establish an age discrimination method ( Figure 8 A). Random forest ( Figure 9) is an effective multi - decision tree widely used in the field of Chinese liquor. To classify an input sample, it needs to be input into each tree for classification, and the classification results of several weak classifiers are voted to form a strong classifier. This is the basic idea of the random forest. To avoid overfitting problems, training data and sampling features are randomly selected. The best way to split the data is determined by the voting of internal nodes until the leaves of each tree meet the requirements for stopping splitting, and finally sample classification is achieved. Support vector machines and neural networks show great advantages in solving non - linear high - dimensional pattern recognition problems such as Chinese liquor GC - MS fingerprints. SVM uses a hyperplane to classify binary - labeled training data based on the geometric margin theory ( Figure 10 , Figure 10 in which the binary - labeled data are marked with different colors, and the hyperplane is in the middle of the five red data points). Different kernel functions can expand the classifier to solve non - linear classification problems by mapping input variables to a high - dimensional feature space. A neural network consists of node layers, and the node layers include an input layer, one or more hidden layers, and an output layer ( Figure 11 ). Each node is connected to another node and has associated weights and thresholds. If the output of any single node exceeds the specified threshold, the node will be activated and send the data to the next layer of the network. Finally, classification can be achieved.
[0144] As Figure 8 shown in A, first, the individual ranking algorithm is used to rank the features (Table 6), and then forward selection is carried out in turn (each time adding one feature to the previous subset to form the next subset), obtaining 5×128 input feature subsets of different sizes. The F1 value of the model using a certain feature subset is compared with 0.94 for evaluation: if the F1 value ≥ 0.94, then the next feature subset is evaluated and one more feature is added; when the F1 value < 0.94, the subset before this subset is selected as the sub - optimal subset. Theoretically, if the ranking algorithm matches the model well, F1 should show a downward trend or a trend of rising first and then falling as the number of input features increases. Therefore, if after the first 32 (the first 25%) features of a certain ranking algorithm are input into a certain discriminant model, the F1 fluctuates violently (the number of times the F1 is lower than 0.94 exceeds 3 times), it is considered that the ranking algorithm and the model do not match well, and no more input feature subsets are obtained from this combination for subsequent union operations (sub - optimal subsets). Figure 8 B - D show the F1 values of 3 models using 5×128 feature subsets. Finally, 9 sub - optimal subsets are screened out, and their intersection is taken to obtain 41 key features ( Figure 8 E).
[0145] Table 6 Ranking scores of optimal feature values under different ranking algorithms
[0146]
[0147]
[0148] The PCA results showed that the samples of the four age groups had better separation effects using 41 key features than using 108 features (compound concentrations). Figure 3 A and Figure 3 B). When using 41 key features as the input data set, the model had a high prediction accuracy for the age of Baijiu (Table 7), and the discrimination accuracy of the neural network reached 100%.
[0149] Table 7 Evaluation of the vintage identification effects of three models using different input data
[0150]
[0151]
[0152] Among the 41 key features, 6 features were highly correlated with the aging time (|r|>0.9), and the correlation coefficients |r| of 17 features were between 0.8 and 0.9 Figure 13 A and 13C). These 41 key features included 36 compound content values, 4 Qcs, and the uniformity index. The 36 main features were the concentrations of 14 esters, 6 acids, 5 alcohols, 3 aldehydes, 3 ketones, 2 pyrazines, 1 ether, 1 phenol, and 1 terpene.
[0153] To deeply understand the contribution of the 36 key compounds to the flavor characteristics of aged Baijiu, their OAVs were measured (Table 8), and 8 aroma descriptions of Luzhou-flavor Baijiu with different aging years were evaluated by sensory experts. Figure 13B). Among the four age groups, the OAVs of 17 compounds were continuously greater than or equal to, and the OAVs of 5 compounds were ≥1 in some age groups. Among them, ethyl hexanoate had the highest OAV (13478 - 25383), showing that the OAV was the highest in new liquor (Group I) and then decreased with aging. It has been reported that the content of ethyl hexanoate in Luzhou-flavor liquor decreases with aging, which is consistent with the research results of this invention. The second highest OAV was ethyl isovalerate (2455 - 6693), followed by ethyl octanoate (1681 - 2639) and ethyl valerate (1145 - 1873), all of which increased with age. In addition to the above four esters, ethyl isobutyrate (306 - 674), ethyl 4-methylvalerate (65 - 190), ethyl 2-methylbutyrate (50 - 144), and ethyl phenylacetate (2 - 4) were also determined, and their OAVs were all ≥1 and increased with aging. Therefore, although ethyl hexanoate decreased with aging, other ester compounds may increase and make a substitute contribution to the fruity aroma. In addition, there may be flavor interactions between compounds, resulting in higher fruity aroma scores for aged liquor in sensory analysis. In addition, two aldehydes with nut / fruit / sweet aroma, including 3-methylbutanal (716 - 1750) and furfural (799 - 850), increased with age. They may contribute to higher scores for the sweetness and nut aroma of aged liquor. In addition, four acids such as valeric acid (43 - 96), isovaleric acid (22 - 25), octanoic acid (4 - 16), and isobutyric acid (1 - 2) increased with aging, which was consistent with the sensory aroma characteristics, that is, the acid aroma score increased with the increase of aging time. In addition, except that 3-methylbutanal in Group IV was below the LOD, in the remaining three groups, its OAV was greater than 1. 3-Methylbutanal has an almond / malt aroma, and it decreased with aging, which was consistent with the result of the decrease in the grain aroma score during sensory evaluation.
[0154] Only longifolene (0 - 719), isophorone (0 - 9), acetoin (0 - 109), and 2,3,5-trimethylpyrazine (0 - 2) were detected in aged liquor ( Figure 12 ). Acetoin exhibited a buttery or creamy aroma, and 2,3,5-trimethylpyrazine had a baking aroma. Their appearance may be related to the Maillard reaction of sugars and amino residues during the aging process. Longifolene smelled of pine wood aroma and was an isomer of β-caryophyllene. Although both β-caryophyllene and isophorone that smelled like camphor were detected from Daqu and sorghum, the more detailed effects of longifolene and isophorone on the flavor of liquor need further study.
[0155] In summary, the six compounds that decreased during the aging process were mainly green, grain, and flower / fruit aromas, while the increased compounds were mainly beneficial to flower / fruit aroma, sweet aroma, acid aroma, baking / nut aroma, wood aroma, creamy aroma, etc. Generally speaking, the changes in the content or OAV of aroma compounds (Figure 13 C and Table 8) are basically consistent with the sensory evaluation profile of Chinese liquor( Figure 13 B). These 36 compounds are not only the year identification markers for aged Luzhou-flavor Chinese liquor, but also closely related to the aroma characteristics of aged Luzhou-flavor Chinese liquor.
[0156] In addition to the compound concentration, the other five key characteristics are the uniformity index and four Qc values. Among them, the Qc of ethyl oleate and the Qc of ethyl butyrate are negatively correlated with the aging time, while the Qc of ethyl isovalerate, the Qc of ethyl propionate and the uniformity index are positively correlated with the aging time( Figure 13 A). These results further confirm the above speculation that the trend of the esterification and hydrolysis reactions towards the thermodynamic equilibrium state plays an important driving role in promoting the change of compound content during the aging process of Chinese liquor, and the uniformity of the compound composition is a potential indicator parameter for the aging time of Chinese liquor.
[0157] Table 8 OAV values of 36 key compounds in Luzhou-flavor Chinese liquor of different ages
[0158]
[0159]
[0160] Note: a The OAVs are obtained by dividing the substance content by the odor threshold; b The odor threshold is obtained by referring to the literature; c The odor threshold is obtained from the literature (Liu & Sun, 2018); d The odor threshold is obtained from the literature (D. Zhao et al., 2018); e The odor threshold is obtained from the literature (Liu et al., 2021); f The odor threshold is obtained from the literature (Burdock, 2009); g The odor threshold is obtained from the literature (Vanderhaegen et al., 2004); h The odor threshold is obtained from the literature (L.J. van Gemert, 2011).
[0161] It can be seen from Examples 1 to 5 that the present invention uses HS-SPME, LLME and LLME-BSTFA combined with GC-MS to systematically analyze the compound content of high-end Luzhou-flavor Chinese liquor. By studying the relationships between the uniformity index, acid content, ethyl ester content, Qc value and aging time, the importance of esterification kinetics in the aging process of Chinese liquor is revealed. The self-measured acid and ester concentration data in this study and the data in the reference literature both show that the reversible esterification reaction in the Chinese liquor system tends to equilibrium with the increase of aging time, thus minimizing the free energy of the system, which is an important driving force for the transformation of flavor compounds during the aging process. This makes the composition of the aged Chinese liquor more stable and the flavor more harmonious and rich.
[0162] In addition, through unsupervised clustering analysis, it was found that aged Chinese liquor could be identified according to four age groups: [0,1) years, [1,5) years, [5,9) years, and [9,11) years. To improve the accuracy and discrimination efficiency of the model, 41 key features were screened out by combining 5 sorting algorithms with 3 models, including 4 Qc values, 1 uniformity index, and 36 compound concentrations. Using these 41 features, a method for identifying the age of Chinese liquor was successfully established. Using these 41 features, the neural network could identify the age of Chinese liquor samples with 100% accuracy.
[0163] In summary, the present invention has deeply understood the aging mechanism of Chinese liquor and established a chemometric method for predicting the composition of flavor compounds after aging of Chinese liquor. In addition, this systematic research idea combining volatile flavor analysis (GC-MS), real-time reaction concentration quotient Qc, compound uniformity index, ranking algorithm, and machine learning should be applicable to the age identification of other types of Chinese liquor and other distilled spirits.
Claims
1. A method for detecting the aging time of Chinese liquor based on GC-MS and machine learning models, characterized in that: It includes the following steps: A. Qualitatively analyze the compounds in the Baijiu samples with gradient aging time by GC-MS to obtain the qualitative data set of the compounds in each liquor sample, and quantitatively analyze the compounds in the qualitative data set to obtain the concentration data set of the compounds in each liquor sample; B. Calculate the evenness index of compounds in each Chinese liquor sample through chemometrics. The calculation formula is: Evenness index = , , where i is 1, 2,..., S, S is the total number of compounds in a certain Chinese liquor sample in step A, N is the total content of all compounds in this Chinese liquor sample, and N i is the content of a certain compound i in this Chinese liquor sample; C. According to the qualitative and concentration data sets determined in step A, select the combination of organic acids, alcohols, corresponding esters and water that coexist in the esterification hydrolysis reversible reaction in each Baijiu sample in step A, establish the reversible reaction equilibrium equation, and calculate the reaction concentration quotient Qc of each esterification hydrolysis reversible reaction in each Baijiu sample by taking the product of the powers of the stoichiometric coefficients of the products and dividing it by the product of the powers of the stoichiometric coefficients of the reactants; D. The concentration data set obtained in step A, the uniformity index obtained in step B, and the reaction concentration quotient obtained in step C are the characteristic values. Use the sorting algorithm to sort the characteristic values according to their importance. Start taking values from the most important characteristic value, and add one characteristic to the previous subset each time to form the next subset, obtaining the total number of characteristic values × the total number of sorting algorithms input feature subsets of different sizes. Input each feature subset into the machine learning model respectively, let the machine learning model perform year discrimination, obtain the total number of characteristic values × the total number of sorting algorithms × the total number of machine learning models discrimination F1 values, compare all F1 value data with the screening threshold, retain M input feature subsets that are input into the machine learning model when F1 is greater than the screening threshold, and take the intersection of the M input feature subsets to obtain the final key characteristic values; E. According to the key characteristic values determined in step D, detect the key characteristic values of the Baijiu to be tested with unknown aging time, and input them into at least one of the machine learning models used in step D to determine its aging time.
2. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 1, wherein: In step A, when qualitatively analyzing the compounds in the liquor sample, eliminate the compounds with a detection rate lower than 30% in the Baijiu samples with gradient aging time according to the 30% rule, and only quantitatively analyze the remaining specific compounds.
3. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 2, wherein: Eliminate the compounds with a detection rate lower than 50% in the Baijiu samples with gradient aging time according to the 50% rule.
4. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 1, characterized in that: In step A, use the unsupervised hierarchical clustering algorithm to perform clustering analysis on the Baijiu samples with gradient aging time.
5. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 1, characterized in that: In step D, the sorting algorithm is at least three of information gain, information gain rate, Gini index in descending order, analysis of variance, chi-square test, feature weight algorithm, and fast correlation filtering algorithm.
6. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 5, characterized in that: In step D, the sorting algorithm is five of information gain, information gain rate, analysis of variance, feature weight algorithm, and fast correlation filtering algorithm.
7. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 1, wherein: In step D, the machine learning model is at least two of random forest, support vector machine, neural network, nearest neighbor algorithm, gradient boosting, linear regression, logistic regression, adaptive boosting, and stochastic gradient descent.
8. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 7, wherein: In step D, the machine learning model is three of random forest, support vector machine, and neural network.
9. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 1, characterized in that: In step A, the Baijiu samples with gradient aging time are selected from 0 to 11 years.
10. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 9, characterized in that: The Baijiu samples with gradient aging time are divided into four groups: [0,1) years, [1,5) years, [5,9) years, and [9,11].
11. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 10, characterized in that: In step A, the liquor samples are pretreated by headspace solid-phase microextraction, liquid-liquid microextraction or liquid-liquid microextraction-BSTFA derivatization.
12. The method for detecting the aging time of Baijiu based on GC-MS and machine learning model according to claim 11, wherein: In step A, at least one of the following is satisfied: The operation of the headspace solid-phase microextraction is as follows: the liquor sample is diluted to a final ethanol content of 5-15 Vol%, NaCl is added, ethyl lactate and compounds with a retention time longer than acetic acid are both internal-standardized with 2-methylhexanoic acid, alcohols with a retention time shorter than acetic acid are internal-standardized with tert-amyl alcohol, and other compounds are internal-standardized with amyl acetate. The volatile compounds are extracted from the top space of the sample by using an SPME fiber. The sample is pre-equilibrated at 45-60 °C for 0-10 min, and then extracted at 45-60 °C for 30-60 min for subsequent GC-MS determination; The operation of the liquid-liquid microextraction is as follows: amyl acetate, 2-methylhexanoic acid and tert-amyl alcohol, three internal standards, and saturated sodium chloride solution are added to the liquor sample, and then solution A is added. After stirring well for more than 3 min, the upper organic phase is collected after static stratification. After the organic phase is concentrated by nitrogen blowing at room temperature, subsequent GC-MS determination is carried out; solution A is a mixed solvent with a volume ratio of anhydrous ether to pentane of 1:1; The operation of the liquid-liquid microextraction-BSTFA derivatization is as follows: heptadecanoic acid internal standard is added to the liquor sample, and then saturated sodium chloride solution is added and diluted to an ethanol concentration of 10%. Solution A is added to the diluted system for extraction, vortexed for more than 3 min, and left standing for more than 20 min. The organic layer is collected, and the remaining aqueous phase is extracted successively with acetonitrile and dichloromethane. The organic phases are combined. Pyridine containing 2.5% (v / v) hydroxylamine hydrochloride is added to the organic phase, and BSTFA containing 1% TMCs is added, followed by vortexing for more than 5 s. Incubate in a metal bath at 45-85 °C for 1-5 h. After the reaction is completed, centrifuge at 6000-12000 rpm for more than 3 min, and take the supernatant for subsequent GC-MS determination; solution A is a mixed solvent with a volume ratio of anhydrous ether to pentane of 1:
1.
13. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 12, wherein: In step A, at least one of the following is satisfied: When headspace solid-phase microextraction is adopted, the GC-MS conditions are as follows: DB-WAX chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8-1.4 mL / min; electron impact mode: 70 eV; the transfer line temperature at the connection between the chromatographic column and the ion source is set at 200-230 °C; scanning mode, the scanning range is m / z 33-350; the inlet temperature is 250-300 °C, the desorption time is 5-15 min, and the split ratio is 2-20:1; temperature programming: starting from 50 °C, increase in stages to 235 °C and hold for more than 2 min; When liquid-liquid microextraction is adopted, the GC-MS conditions are as follows: DB-WAX chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8 - 1.4 mL / min; electron impact mode: 70 eV; the transfer line temperature at the connection between the chromatographic column and the ion source is set at 200 - 230 °C; scanning mode, the scanning range is m / z 33 - 350; the sample injection volume is 1 μL, the inlet temperature is 250 - 300 °C, and the split ratio is 2 - 20:1; temperature programming: starting from 35 °C, it rises in stages to 235 °C and is held for more than 10 min; When liquid-liquid microextraction - BSTFA derivatization is adopted, the GC-MS conditions are as follows: HP-5MS chromatographic column; the carrier gas is helium with a purity of 99.999% at 0.8 - 1.4 mL / min; electron impact mode: 70 eV; the transfer line temperature at the connection between the chromatographic column and the ion source is set at 200 - 230 °C; scanning mode, the scanning range is m / z 33 - 350; the sample injection volume is 1 μL, the inlet temperature is 250 °C, and the split ratio is 2 - 20:1; temperature programming: gradually rising from 65 °C to 280 °C and held for more than 2 min.
14. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 10, characterized in that: In step D, before the sorting algorithm performs importance sorting, the eigenvalues include 108 compound concentrations in the compound concentration dataset in step A, 1 compound uniformity index in step B, and 19 reaction concentration quotients Qc in step C.
15. The method for detecting the aging time of Chinese liquor based on GC-MS and machine learning model according to claim 14, characterized in that: In step D, the key eigenvalues include 36 compound concentrations screened from the compound concentration dataset in step A, 1 compound uniformity index in step B, and 4 Qc screened from the reaction concentration quotient Qc in step C.
Citation Information
Patent Citations
Method for identifying storage time of white spirit by multiple linear stepwise regression
CN113203803A
Method for solid-phase microextraction and analysis, and a collector for this method
US20020098594A1