Method for identifying age of white spirit

By screening out 32 key characteristic components of sauce-flavored liquor, building a lightweight random forest model, the problem of standardization and quantitative evaluation of sauce-flavored liquor wine is solved, and the rapid and accurate age identification is achieved, reducing the detection cost and time.

CN120579078APending Publication Date: 2025-09-02MOUTAI INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510674568.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The existing technology is difficult to achieve standardization and quantitative evaluation of the wine age of sauce-flavored liquor. The traditional methods are greatly affected by individual sensory differences, and the detection cost is high, the efficiency is low, and the model is high, so it is impossible to effectively identify illegally added chemical substances or fake wine age.

Method used

Through cross-verification algorithm and variable importance analysis, 32 key characteristic components were screened out from 2,243 flavor substances, a lightweight random forest model was constructed, and the flavor data of liquor was obtained using gas chromatography-mass spectrometry combined technology, a mapping relationship between characteristic substance content and wine age was established, and 500 decision trees were generated for wine age prediction.

Benefits of technology

The rapid and accurate identification of liquor wine age was achieved, the detection variable was reduced by 98.6%, the model bag error was reduced to 11.11%, and the test set identification accuracy was improved to 89%, reducing the detection cost and time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579078A_ABST
    Figure CN120579078A_ABST
Patent Text Reader

Abstract

A method for identifying the age of white spirit comprises the following steps: S1, detecting Maotai-flavor white spirit standard samples of different ages by using a gas chromatography-mass spectrometry technology to obtain original detection data of flavor substances; s2, based on variable importance analysis of a random forest model, screening out key characteristic substances from the flavor substances, and carrying out normalization and centralization processing on the content of the screened characteristic substances; s3, constructing a random forest classification model by using an R language random Forest package, taking the content of the key characteristic substance as an input variable and taking the real age of the white spirit as an output variable; and S4, detecting the content of key characteristic substances in a white spirit sample to be detected, processing the key characteristic substances, inputting the processed characteristic data into the random forest classification model, and outputting a predicted spirit age. Compared with the prior art, the method has the advantages that the detection variable is reduced by 98.6%, the out-of-bag error (OOB) of the model is reduced from 55.56% to 11.11%, and the recognition accuracy of the test set is improved to 89%.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of liquor identification, and in particular to a method for identifying the age of liquor. Background Art

[0002] Maotai-flavor liquor, with its unique flavor and craftsmanship, holds a significant position in the liquor market. With rising consumption and expanding market size, the need for authenticity verification, a core indicator of liquor quality and value, has become increasingly urgent. However, the industry currently faces the following technical bottlenecks: Traditional authentication methods rely on the taster's sense of smell, taste, and vision, which are significantly affected by individual sensory differences, emotional state, and environmental factors, making standardization and quantitative assessment difficult. For example, different tasters may differ by 2-3 years on the age of the same five-year-old bottle. Furthermore, while traditional methods like gas chromatography-mass spectrometry (GC-MS) can analyze flavor compounds, they require testing hundreds or even thousands of components, taking hours. Furthermore, they cannot directly correlate components with liquor age, requiring manual database comparisons, resulting in low efficiency and an accuracy rate of less than 70%.

[0003] Some existing studies have attempted to construct wine age recognition systems using models such as partial least squares-discriminant analysis (PLS-DA). For example, one study developed a PLS-DA model by analyzing 2,243 flavor compounds. While the accuracy for identifying young base wines (less than 5 years) reached 83.3%, the overall accuracy was only 60%. Furthermore, the model required the detection of 774 flavor compounds, resulting in high model complexity and high detection costs in practical applications.

[0004] However, as market demand continues to grow, some unscrupulous businesses, seeking profit, have resorted to illegal methods such as adding chemicals or falsifying the age of liquor, severely damaging consumer rights and the healthy development of the Maotai-flavor liquor industry. Therefore, developing a method for rapidly and accurately identifying the authenticity of Maotai-flavor liquor's age is crucial for protecting consumer rights and promoting the healthy development of the industry. Summary of the Invention

[0005] This paper aims to provide a method for identifying the age of liquor. Using a cross-validation algorithm and variable importance analysis, 32 key characteristic components (such as 1,2-dimethoxybenzene and ethyl benzoate, all with MeanDecreaseAccuracy values ​​greater than 2.1) were screened from 2,243 flavor compounds and used to construct a lightweight random forest model. Compared with existing methods, this method reduces the number of test variables by 98.6%, lowers the model's out-of-bag error (OOB) from 55.56% to 11.11%, and improves test set recognition accuracy to 89%, achieving breakthroughs in both detection efficiency and accuracy.

[0006] A method for identifying the age of liquor, comprising the following steps:

[0007] S1. Obtaining sample flavor data: Using gas chromatography-mass spectrometry technology to test standard samples of Maotai-flavor liquor of different ages, obtaining original detection data of flavor substances;

[0008] S2. Feature screening and data preprocessing: Based on the variable importance analysis of the random forest model, key characteristic substances were screened from the flavor substances, and the content of the screened characteristic substances was normalized and centered;

[0009] S3. Constructing a random forest model: Using the R language randomForest package, with the content of key characteristic substances as input variables and the actual age of liquor as output variables, a random forest classification model was constructed;

[0010] S4. Wine age prediction: The content of key characteristic substances in the liquor sample to be tested is detected and normalized and centered, and the processed characteristic data is input into the random forest classification model to output the predicted wine age.

[0011] Optimally, the flavor substances include alcohols, esters, and ketones.

[0012] Optimized, the key characteristic substances include 1,2-dimethoxybenzene, ethyl benzoate, (3aS, 8aS)-6, 8a-dimethyl-3-(prop-2-ylidene)-1,2,3,3a,4,5,8,8a-octahydroazulene, and ethyl hexanoate.

[0013] The optimized screening process of the key characteristic substances includes: analyzing the relationship between the model error and the number of flavor substances based on the cross-validation curve, and determining that when the number of flavor substances is 32, the model error tends to a stable minimum value, and at this time the MeanDecreaseAccuracy value is greater than 2.1.

[0014] Optimized, 32 key characteristic substances were screened from flavor substances.

[0015] The relationship between the model error and the number of flavor substances was analyzed based on the cross-validation curve, and it was determined that the model error tended to a stable minimum value when the number of flavor substances was 32.

[0016] The model was optimized by generating 500 decision trees through sampling with replacement, with 47 variables randomly selected from each tree node, and the out-of-bag error (OOB) ≤ 11.11% was set as the optimization target.

[0017] Working principle and beneficial effects of the present invention:

[0018] The age of Maotai-flavor liquor is closely related to the composition, content, and interactions of its flavor substances. During storage, substances such as alcohols, esters, and ketones in the base liquor undergo chemical reactions such as oxidation, esterification, and hydrolysis, causing the flavor system to show regular changes over time: At the low-age stage (1-3 years), process differences (such as raw materials and fermentation conditions) and environmental factors dominate the flavor system. The flavors of base liquors from different distilleries vary significantly, and there are many types of characteristic substances with large fluctuations in abundance. At the high-age stage (more than 5 years), chemical reactions tend to balance, the composition of flavor substances converges to a "stable mode," and the similarity of base liquors from different sources increases significantly.

[0019] GC-MS detection was used to obtain the original data of 2243 flavor substances. Through cross-validation and variable importance analysis (MeanDecreaseAccuracy), 32 key substances with the strongest correlation with wine age were screened out (for example, the content of ester substances increases with wine age, and the content of alcohol substances first increases and then decreases), and a "characteristic substance content-wine age" mapping relationship was established.

[0020] 500 decision trees were generated through bootstrap sampling. Each tree randomly selected 47 variables (characteristic substances) at the node splitting to prevent a single variable from dominating the decision. Each tree independently predicted the sample age, and the final result was determined by majority voting to reduce the risk of overfitting of a single tree. The model's generalization ability was evaluated using the out-of-bag error (OOB). When OOB ≤ 11.11%, it indicated that the model's prediction reliability for unknown samples was over 90%.

[0021] This application uses a cross-validation algorithm and variable importance analysis to screen 32 key characteristic components (such as 1,2-dimethoxybenzene and ethyl benzoate, all with MeanDecreaseAccuracy values ​​> 2.1) from 2,243 flavor substances and construct a lightweight random forest model. Compared with existing technologies, this method reduces the number of detection variables by 98.6%, reduces the model's out-of-bag error (OOB) from 55.56% to 11.11%, and improves the recognition accuracy of the test set to 89%, achieving a double breakthrough in detection efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 This is the Venn diagram of flavor substances in Maotai-flavor liquors of different ages;

[0023] Figure 2 This is a heat map analysis of flavor substances in Maotai-flavor liquors of different ages (each sample is normalized according to its substance abundance);

[0024] Figure 3 This is the principal component analysis diagram of the flavor system of Maotai-flavor liquor of different ages;

[0025] Figure 4are the parameters of the partial least squares-discriminant analysis model (left: model score plot; right: variable scores in the first two axes of PLS-DA);

[0026] Figure 5 are the optimized partial least squares-discriminant analysis model parameters (left: model score plot; right: variable scores in the first two axes of PLS-DA);

[0027] Figure 6 This is the probability graph of the random forest model observation being judged correct;

[0028] Figure 7 is the cross validation curve graph;

[0029] Figure 8 A plot of the probability that an observation is correct for the optimized random forest model. DETAILED DESCRIPTION

[0030] The following is further described in detail through specific implementation methods:

[0031] The embodiment mainly completes data analysis and model construction based on R language. The main software and installation packages and their versions are shown in Table 1:

[0032] Table 1 - Software used and its versions

[0033]

[0034]

[0035] Using gas chromatography-mass spectrometry (GC-MS) analysis, 2243 flavor substances were detected in the sauce-flavor liquor, mainly alcohols, esters and ketones, indicating that the sauce-flavor liquor, as a product of solid-state mixed fermentation, has a rich and complex flavor substance structure.

[0036] By Venn analysis ( Figure 1 ) It can be seen that there are 405 substances that are common flavor components of base wines with an age of 1, 3, 5, 8, and 10 years, and their unique flavor substance types are 185, 204, 237, 197 and 219 respectively, indicating that the flavor system of base wine will undergo great changes during the storage process, and base wines of different ages will form significantly different flavor systems with the increase of storage time.

[0037] The flavor substance content in the sample was normalized and cluster analysis was performed, such as Figure 2Except for some samples that could not be accurately grouped together during the cluster analysis, most samples could be clustered and distinguished according to their origin and age, such as A01 (1-year-old base wine from Factory A), C01 (1-year-old base wine from Factory C), A03 (3-year-old base wine from Factory A), and B10 (10-year-old base wine from Factory B). This shows that base wines with similar ages have strong similarities in flavor systems.

[0038] However, from the above analysis, it can be seen that due to the characteristics of natural fermentation of sauce-flavor liquor, it has an extremely complex flavor substance system. Too many types of unnecessary flavor substances will reduce the differences between base liquors of different ages, making it difficult to distinguish their ages using simple cluster analysis. Therefore, it is necessary to identify the types of characteristic flavor substances based on further analysis of their changing patterns in order to realize the construction of the authenticity recognition model for the age of sauce-flavor liquor.

[0039] Similarity analysis of sauce-flavor liquors of different ages: The sauce-flavor liquor base liquors were grouped according to their age, and then the similarity between the two groups was examined using anosim analysis, as shown in Table 2. The results show that, first, there are significant differences between base liquors with shorter storage years (such as 1-3 years of age) and base liquors with longer storage years, such as 1 year of age and 5, 8 and 10 years of age, and 3 years of age and 8 and 10 years of age; second, there are no significant differences between base liquors with long-term storage (base liquors with a storage time of 5 years or more), and they have a strong similarity. In summary, the flavor system of the sauce-flavor liquor base liquor changes greatly in the early stage of storage, but when the storage time reaches 5 years, the degree of overall style change will decrease, which further explains Figure 2 There is a phenomenon that a large number of 5-, 8- and 10-year-old base wines are clustered together. See Table 2:

[0040] Table 2 - Pairwise Anosim analysis of Maotai-flavor liquors of different ages

[0041] group1 group2 distance R P_value 1 3 Bray-Curtis -0.01475 0.471 1 5 Bray-Curtis 0.180727 0.042 1 8 Bray-Curtis 0.290466 0.004 1 10 Bray-Curtis 0.260974 0.008 3 5 Bray-Curtis 0.070988 0.17 3 8 Bray-Curtis 0.150206 0.048 3 10 Bray-Curtis 0.162894 0.029 5 8 Bray-Curtis 0.002743 0.395 5 10 Bray-Curtis 0.049383 0.225 8 10 Bray-Curtis 0.024348 0.273

[0042] In order to further explore the effect of storage time on the flavor system of Maotai-flavor liquor, principal components analysis (PCA) was used to reduce the dimensionality of the complex flavor components in liquor. Figure 3As shown in the figure, it can be found that, firstly, the distribution of 1-year and 3-year-old base wines is relatively scattered, especially the base wines from different wineries have obvious differences in distribution, which is significantly different from the 5-, 8-, and 10-year-old base wines. Analysis shows that the influence of their flavor system is more due to factors such as process and storage conditions. Secondly, the distribution of 5-, 8-, and 10-year-old base wines on the PC1 axis is more concentrated. Even base wines from different wineries can be relatively concentrated when they are over 5 years old, indicating that the flavor composition of base wines is constantly changing towards a unified flavor composition pattern.

[0043] In summary, the flavor composition of the base liquor of Maotai-flavor baijiu continuously changes during storage, and these changes are directional. In the early stages of storage, factors such as process and environment may lead to significant differences between base liquors from different sources. However, as storage time increases, the impact of unstable environmental factors such as process, environment, and source will continue to weaken, and the base liquor flavor system will continue to change and converge towards a unified flavor composition pattern. Therefore, even the similarity between base liquors from different distilleries continues to increase, and there is even a higher degree of similarity between base liquors of different ages.

[0044] Construction and optimization of partial least squares-discriminant analysis model:

[0045] Unlike PCA, partial least squares-discriminant analysis (PLS-DA) is a supervised analytical method, classified as a model approach. It uses partial least squares regression to model the relationship between flavor compound content and base wine age. Data dimensionality reduction can better establish the relationship between samples and vintages. PLS-DA, available in the R package ropls, was used to model the relationship between flavor compound content and base wine age. R2X and R2Y represent the model's explanatory power for the X and Y matrices, respectively, while Q2 indicates the model's predictive power. Values ​​closer to 1 indicate a better model fit, and the more accurately the training set samples are classified into their original categories.

[0046] Depend on Figure 4 It can be seen that R2Y reached 0.982, indicating that the model explained 98.2% of the variance of the dependent variable matrix (base wine age), that is, the model can highly fit the base wine age. Figure 4 On the left, we can also see that base wines of different ages can be distinguished well. The Q2Y reached 0.62, which also shows that the model has good predictive ability.

[0047] However, since the model utilizes all flavor components, excessive variables can increase the risk of overfitting and complexity. Therefore, variable importance for the projection (VIP) is used to measure the influence and explanatory power of each flavor component's content on sample classification and discrimination, assisting in the selection of characteristic flavor components. A VIP value > 1 is typically used as the selection criterion.

[0048] After screening, the model obtained a data set containing 774 flavor substance variables (Appendix 1). The model was re-modeled based on the screened variables. The model evaluation was as follows: Figure 5 Although the R2Y of the optimized model decreased from 0.982 to 0.961, Q2Y increased to a certain extent, indicating that the predictive ability of the optimized model after streamlining the modeling variables has been improved to a certain extent.

[0049] Using the parameters provided by the ropls package, the original dataset was sorted by odd-even numbers to select the flavor composition of a subset of base wines as a test dataset. The model was then used to predict the ages of these wines and compared them with their actual ages to assess the model's performance. Table 3 shows the model's performance test results. The model's prediction accuracy was only 60%, demonstrating that it was able to predict the ages of younger base wines (less than five years old) well (with an accuracy of 83.3%). However, its prediction ability for older base wines remained significantly limited.

[0050] Table 3 Performance test of PLS-DA optimization model

[0051]

[0052] In summary, although the model based on partial least squares-discriminant analysis can make good predictions for young base wines, its overall prediction accuracy is only 60%, which is difficult to meet the application in production and life. Therefore, further consideration of model selection is urgently needed.

[0053] Construction and optimization of random forest discriminant analysis model

[0054] The random forest model was constructed using the randomForest package of R language to conduct further discriminant analysis on base wines of different ages, so as to improve the prediction accuracy of the base wine age.

[0055] A random forest model was constructed using the entire flavor composition. The randomForest() function randomly sampled 36 observation points from the training set with replacement, and 47 variables were randomly sampled at each node of each tree, generating 500 classic decision trees. The categories corresponding to the sample points not used in the tree generation can be estimated by the generated tree, and compared with their true categories to obtain the out-of-bag prediction error (OOB), which can be used to reflect the error rate of the classifier. The OOB value of this model was 55.56%, indicating low discrimination accuracy. Figure 6 It also shows a low judgment accuracy, so the original modeling data needs to be further screened and optimized.

[0056] The random forest model can be used to identify the characteristics of the importance of variables, and the cross-validation algorithm can be used to select flavor substances based on the cross-validation curve. Figure 7 ) shows the relationship between model error and the number of flavor components used for fitting. The error initially decreases significantly as the number of flavor components increases, but by the time the number of flavor components reaches approximately 32, the rate of decrease becomes less significant and even increases. Therefore, this study selected the 32 most important flavor components (Table 4) based on their MeanDecreaseAccuracy values ​​as the dataset for model construction.

[0057] Table 4 - List of the top 32 flavor substances screened

[0058]

[0059]

[0060] The random forest model was rebuilt using the dataset constructed from the top 32 flavor substances. Figure 8 As can be seen, the model's judgment accuracy has improved to a certain extent. Further examination of the model's OOB found that the model's OOB has dropped to 11.11%, indicating that the model has a low error rate, that is, the judgment accuracy has been significantly improved.

[0061] The optimized model was further tested on a test dataset consisting of nine samples. The model's performance is shown in Table 5. The test dataset contained the flavor composition of nine samples. The optimized model accurately predicted the true age of eight of these samples, achieving an accuracy rate of 89%. Further analysis of the incorrectly predicted samples revealed that while the model's prediction of the one-year-old sample was somewhat biased, its prediction of three years, the closest to the one-year-old, was accurate. This indicates no significant deviation, demonstrating the model's excellent performance.

[0062] Table 5 - Random Forest Optimization Model Performance Test

[0063]

[0064] In summary, the optimized random forest model has good predictive performance. Its data set construction requires the quantitative abundance of 32 flavor substances in the base wine flavor substance system, that is, only the content of 32 flavor substances is needed to construct the random forest model. It can not only significantly improve the accuracy of the model, but also effectively reduce the testing workload of testers and improve the timeliness of wine age judgment, thereby better meeting the needs of production and life.

[0065] The above is only an embodiment of the present invention, and the common knowledge such as the specific structure and characteristics of the scheme is not described in detail here. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention, and these will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A method for identifying the age of liquor, characterized in that: The steps include: S1. Obtaining sample flavor data: Using gas chromatography-mass spectrometry technology to test standard samples of Maotai-flavor liquor of different ages, obtaining original detection data of flavor substances; S2. Feature screening and data preprocessing: Based on the variable importance analysis of the random forest model, key characteristic substances were screened from the flavor substances, and the content of the screened characteristic substances was normalized and centered; S3. Constructing a random forest model: Using the R language randomForest package, with the content of key characteristic substances as input variables and the actual age of liquor as output variables, a random forest classification model was constructed; S4. Wine age prediction: The content of key characteristic substances in the liquor sample to be tested is detected and normalized and centered, and the processed characteristic data is input into the random forest classification model to output the predicted wine age.

2. The method for identifying the age of liquor according to claim 1, characterized in that: The flavor substances include alcohols, esters, and ketones.

3. The method for identifying the age of liquor according to claim 2, characterized in that: The key characteristic substances include 1,2-dimethoxybenzene, ethyl benzoate, (3aS, 8aS)-6, 8a-dimethyl-3-(prop-2-ylidene)-1,2,3,3a,4,5,8,8a-octahydroazulene, and ethyl hexanoate.

4. The method for identifying the age of liquor according to claim 3, characterized in that: The screening process of the key characteristic substances includes: analyzing the relationship between the model error and the number of flavor substances based on a cross-validation curve, and determining that when the number of flavor substances is 32, the model error tends to a stable minimum value, and at this time, the MeanDecreaseAccuracy value is greater than 2.

1.

5. The method for identifying the age of liquor according to claim 4, characterized in that: 32 key characteristic substances were screened from flavor substances.

6. The method for identifying the age of liquor according to claim 5, characterized in that: The relationship between the model error and the number of flavor substances was analyzed based on the cross-validation curve, and it was determined that the model error tended to a stable minimum value when the number of flavor substances was 32.

7. The method for identifying the age of liquor according to any one of claims 1 to 6, characterized in that: The model generates 500 decision trees by sampling with replacement, with 47 variables randomly selected from each tree node, and the out-of-bag error (OOB) ≤ 11.11% is used as the optimization target.