Method for predicting relative complexity index of white spirit flavor
Through quantitative analysis of liquor samples and synthetic chemistry machine learning methods, flavor substances representing aroma types were screened out, and a model was established and optimized, which solved the problem of accuracy in predicting liquor flavor complexity and achieved accurate prediction of liquor flavor complexity.
Patent Information
- Application Number
- CN202510986726.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-09-09
AI Technical Summary
Existing technologies cannot accurately reflect the complexity of liquor flavor, especially the difficulty in detecting low-content but highly active flavor substances, and fail to effectively optimize model parameters, resulting in low prediction accuracy.
By quantitatively analyzing the flavor substances in liquor samples, machine learning methods in the field of synthetic chemistry were used to screen out flavor substances representing the aroma. A machine learning model was established, and grid cross-validation was used to optimize the model hyperparameters to predict the relative complexity index of liquor flavor.
It achieves accurate prediction of the flavor complexity of liquor, provides data support for the impact of molecular complexity on flavor, and significantly improves prediction accuracy.
Smart Images

Figure CN120609954A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of liquor, and in particular to a method for predicting relative complexity indexes of liquor flavor. Background Art
[0002] Baijiu (Chinese liquor), a traditional Chinese distilled spirit, has its flavor quality as a key factor in determining its aroma, style, and quality grade. Baijiu's flavor profile is extremely complex, encompassing numerous trace components, some present in minute quantities but contributing significantly to the spirit's flavor. Furthermore, these components interact with each other. The traditional open, natural, multi-strain solid-state fermentation process results in a certain degree of flavor fluctuation in the base liquor. Therefore, controlling the quality stability of the finished liquor is a common challenge faced by baijiu companies.
[0003] Currently, baijiu flavor evaluation relies primarily on empirical sensory evaluation methods and conventional chromatographic skeleton component analysis. Sensory evaluation methods are subject to uncertainty and subjectivity. Conventional chromatographic skeleton component analysis, while capable of detecting some flavor compounds, presents significant challenges for detecting numerous components in complex systems, particularly low-content but highly active flavor compounds. Furthermore, it struggles to fully capture the complexity of flavor.
[0004] With the advancement of analytical technology, techniques such as gas chromatography-mass spectrometry (GC-MS) have been widely used to detect flavor compounds in baijiu. However, baijiu's composition is complex and diverse, especially the hundreds of esters, which have similar aromas and are prone to cross-talk during analysis. Furthermore, interference from high-dimensional, complex data and uncertainty in predictions remain challenges in flavoromics.
[0005] In recent years, with the rapid development of chemical analysis and machine learning techniques, flavor analysis methods based on substance content have gradually become a research hotspot. Although many studies have attempted to model the relationship between chemical composition and flavor, most methods rely on a small amount of flavor substance data or fail to effectively optimize model parameters, resulting in low prediction accuracy. Moreover, these methods often only consider the quantitative data of flavor substances, without fully considering the impact of the substance's molecular complexity on flavor expression, namely the flavor relative complexity index. The flavor relative complexity index has a strong positive correlation with the sensory evaluation results of baijiu (white spirits), and can more accurately reflect the flavor complexity of baijiu. Summary of the Invention
[0006] Technical problem solved by the present invention: The present invention provides a method for predicting the relative complexity index of liquor flavor, which solves the problem that the existing technology cannot accurately reflect the complexity of liquor flavor.
[0007] The present invention solves the above technical problems by adopting a technical solution: a method for predicting the relative complexity index of liquor flavor, comprising the following steps: S1. Quantitatively analyze flavor substances in multiple liquor samples of the same flavor type to obtain the types and corresponding contents of flavor substances; S2. Using machine learning methods in the field of synthetic chemistry, obtain five indices for various flavor substances in all liquor samples, and screen out flavor substances representing the aroma type based on the five indices; the five indices include synthetic complexity score, stoichiometry value, modified stoichiometry value, chemical space exploration value, and space score value; S3. Normalizing the contents of flavor substances representing the flavor type in all liquor samples by flavor substance type to obtain the normalized content of each flavor substance, and summing the normalized contents of each flavor substance in the same liquor sample to obtain a relative flavor complexity index for each liquor sample; S4. Establishing a machine learning model, wherein the machine learning model includes resampling, using flavor substances representing the aroma type and corresponding content data as input features, and using the flavor relative complexity index as a prediction target, training the machine learning model to obtain a flavor relative complexity index prediction model; S5. Optimizing hyperparameters of the flavor relative complexity index prediction model using a grid search cross-validation method to obtain an optimized flavor relative complexity index prediction model; S6. Inputting the flavor substances representing the aroma type and the corresponding content data of the liquor sample to be tested into the optimized liquor flavor relative complexity index prediction model to obtain the flavor relative complexity index of the liquor sample to be tested.
[0008] Furthermore, the flavor is one of twelve flavors, including: strong flavor, sauce flavor, light flavor, rice flavor, phoenix flavor, mixed flavor, Dong flavor, sesame flavor, special flavor, fermented soybean flavor, Laobaigan flavor and rich flavor.
[0009] Furthermore, quantitative analysis includes gas chromatography-tandem mass spectrometry, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry or flame ionization detection.
[0010] Furthermore, machine learning methods in the field of synthetic chemistry include synthetic complexity scoring, complexity index and space complexity score.
[0011] Furthermore, flavor substances for evaluating the relative complexity of liquor flavor are screened out based on the five indices, including: selecting flavor substances corresponding to the top 30% of each index value arranged from large to small as flavor substances for evaluating the relative complexity of liquor flavor, or selecting flavor substances with each index value exceeding the corresponding threshold as flavor substances for evaluating the relative complexity of liquor flavor.
[0012] Furthermore, the machine learning model is a random forest model, a linear regression model, a support vector machine model or a K-nearest neighbor model.
[0013] Beneficial effects of the present invention: The present invention provides a method for predicting the relative complexity index of liquor flavor, which obtains flavor substances and corresponding contents of multiple liquor samples of the same flavor type through quantitative analysis, obtains five indexes of various flavor substances in all liquor samples by adopting a machine learning method in the field of synthetic chemistry, and screens out flavor substances representing the flavor type, normalizes the contents of flavor substances representing the flavor type in all liquor samples to obtain the normalized content of each flavor substance, sums the normalized contents of each flavor substance in the same liquor samples to obtain the flavor relative complexity index of each liquor sample, and thereby constructs a flavor relative complexity index prediction model with the flavor substances and corresponding contents of the liquor samples as input and the flavor relative complexity index as output, and predicts the flavor relative complexity index of the liquor sample to be tested by the flavor relative complexity index prediction model, which solves the problem that the prior art cannot accurately reflect the flavor complexity of liquor and provides data support for the influence of molecular complexity on flavor. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 The present invention provides a flow chart of a method for predicting the relative complexity index of liquor flavor. DETAILED DESCRIPTION
[0015] The present invention addresses the problem that the existing technology does not consider the influence of molecular complexity on flavor and thus cannot accurately reflect the complexity of liquor flavor. A method for predicting the relative complexity index of liquor flavor is provided. Through quantitative analysis, the flavor substances and corresponding contents of multiple liquor samples of the same flavor type are obtained. A machine learning method in the field of synthetic chemistry is adopted to obtain five indices of various flavor substances in all liquor samples, and the flavor substances representing the flavor type are screened out. The contents of the flavor substances representing the flavor type in all liquor samples are normalized to obtain the normalized content of each flavor substance. The normalized content of each flavor substance in the same liquor samples is summed to obtain the flavor relative complexity index of each liquor sample. In this way, a flavor relative complexity index prediction model is constructed with the flavor substances and corresponding contents of the liquor samples as input and the flavor relative complexity index as output. The flavor relative complexity index prediction model predicts the flavor relative complexity index of the liquor sample to be tested.
[0016] like Figure 1 As shown, the present invention provides a method for predicting the relative complexity index of liquor flavor, comprising the following steps: S1. Quantitatively analyze the flavor substances in multiple liquor samples of the same flavor type to obtain the types and corresponding contents of the flavor substances.
[0017] Specifically, the aroma is one of twelve aroma types, including: strong aroma, sauce aroma, light aroma, rice aroma, phoenix aroma, mixed aroma, Dong aroma, sesame aroma, special aroma, fermented black bean aroma, Laobaigan aroma, and rich aroma. Quantitative analysis includes gas chromatography-tandem mass spectrometry, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry, or flame ionization detection.
[0018] S2. Using machine learning methods in the field of synthetic chemistry, five indices of various flavor substances in all liquor samples are obtained, and flavor substances representing the aroma are screened out based on the five indices; the five indices include synthetic complexity score value, stoichiometric value, modified stoichiometric value, chemical space exploration value and spatial score value.
[0019] Specifically, the machine learning methods in the field of synthetic chemistry include synthetic complexity scoring method, complexity index method and space complexity score method. The index SCScore is obtained by synthetic complexity scoring method, and the index CM is obtained by complexity index method. and CSE, and the index SPS is obtained by the spatial complexity score method to obtain five indices. The five indices are used to screen flavor substances representing the aroma type, including selecting the top 30% of flavor substances corresponding to each index value, arranged from largest to smallest, as flavor substances for evaluating the relative complexity of the liquor flavor, or selecting flavor substances with each index value exceeding the corresponding threshold as flavor substances for evaluating the relative complexity of the liquor flavor.
[0020] S3. Normalizing the contents of flavor substances representing the flavor type in all liquor samples by flavor substance type to obtain the normalized content of each flavor substance, and summing the normalized contents of each flavor substance in the same liquor sample to obtain a relative flavor complexity index for each liquor sample; Specifically, the content of flavor substances representing the aroma type is normalized according to the type of flavor substances, so that the content of each flavor substance in different samples is comparable, thereby establishing a correlation between the flavor substances and the corresponding content and the flavor relative complexity index.
[0021] Although the flavor relative complexity index of each liquor sample can be obtained from S1 to S3, if only the content of flavor substances representing the aroma type of a liquor is measured, the complexity score cannot be directly calculated due to lack of comparison. Therefore, the present invention also includes S4 to S6 to facilitate the prediction of the flavor relative complexity index of a liquor.
[0022] S4. Establishing a machine learning model. The machine learning model includes resampling, using flavor substances representing the aroma type and corresponding content data as input features, and using the flavor relative complexity index as a prediction target. The machine learning model is trained to obtain a flavor relative complexity index prediction model. Specifically, the machine learning model is a random forest model, a linear regression model, a support vector machine model, or a K-nearest neighbor model.
[0023] S5. Use a grid search cross-validation method to optimize the hyperparameters of the flavor relative complexity index prediction model to obtain an optimized flavor relative complexity index prediction model.
[0024] S6. Inputting the flavor substances representing the aroma type and the corresponding content data of the liquor sample to be tested into the optimized liquor flavor relative complexity index prediction model to obtain the flavor relative complexity index of the liquor sample to be tested.
[0025] Example: Taking 26 samples of Luzhou-flavor liquor as an example, the gas chromatography-tandem mass spectrometry (GC-MS / MS) method was used to quantitatively analyze and obtain the content data of their flavor substances, a total of 140 flavor substances. Then, the index SCScore was obtained by the synthetic complexity scoring method, and the index CM was obtained by the complexity index method. and CSE, the index SPS is obtained by the space complexity score method, as shown in Tables 1 to 4.
[0026] Table 1 Five indices of flavor substances corresponding to sequence numbers 1 to 36 in 26 Luzhou-flavor liquors
[0027] Table 2 Five indices of flavor substances corresponding to sequence numbers 37 to 72 in 26 Luzhou-flavor liquors
[0028] Table 3 Five indices of flavor substances corresponding to serial numbers 73 to 108 in 26 Luzhou-flavor liquors
[0029] Table 4 Five indices of flavor substances corresponding to serial numbers 109 to 140 in 26 Luzhou-flavor liquors
[0030] From the above 140 flavor substances, 64 flavor substances representing the strong-aroma type were screened out, and the contents of the 64 flavor substances representing the strong-aroma type were normalized according to the type of flavor substances, that is, the contents of the same flavor substances in 26 wine samples were normalized so that the contents of each flavor substance in different samples were comparable. Then the relative complexity index of the 26 wine samples was obtained by summing them up. The relative complexity index of the 26 wine samples is shown in Table 5.
[0031] Table 5 Relative complexity index of 26 wine samples
[0032] A random forest model was established, with 64 flavor substances representing the strong-flavor type and their corresponding contents as input data and the relative complexity index as output data. The random forest model was trained to obtain a prediction model for the relative complexity index of liquor flavor. Bootstrap resampling was performed on the data in the model. In this way, the training process obtained multiple different sub-datasets, thereby improving the stability and robustness of the model. The grid search cross-validation method (GridSearchCV) was used to optimize the model's hyperparameters to obtain the optimized flavor relative complexity index prediction model. The optimal hyperparameter combination is as follows:
[0033] n_estimators = 50; indicates that the number of trees in the random forest is 50, which is a moderate value that can strike a balance between model performance and computational efficiency.
[0034] max_depth = None; indicates that there is no limit to the depth of the tree, and the tree will grow until all leaf nodes are pure or the minimum number of samples is reached.
[0035] min_samples_split = 2; indicates that when splitting nodes, each node requires at least 2 samples to ensure the rationality of the split.
[0036] min_samples_leaf = 1; indicates that each leaf node requires at least one sample to avoid an overly complex tree structure.
[0037] max_features = 'sqrt'; indicates that at each split, the number of features considered is the square root of the total number of features, which helps reduce overfitting and improve model diversity.
[0038] The R² value of the optimized flavor relative complexity index prediction model was 0.98 and the mean square error (MSE) was 0.21.
[0039] Sensory verification: A correlation analysis was conducted between the flavor relative complexity index and the sensory complexity score. The resulting flavor relative complexity index was significantly positively correlated with the sensory complexity score. Therefore, the flavor relative complexity index can be used to reflect the flavor complexity of liquor, thereby providing data support for the impact of molecular complexity on flavor.
Claims
1. A method for predicting the relative complexity index of liquor flavor, characterized in that: The following steps are involved: S1. Quantitatively analyze flavor substances in multiple liquor samples of the same flavor type to obtain the types and corresponding contents of flavor substances; S2. Using machine learning methods in the field of synthetic chemistry, obtain five indices for various flavor substances in all liquor samples, and screen out flavor substances representing the aroma type based on the five indices; the five indices include synthetic complexity score, stoichiometry value, modified stoichiometry value, chemical space exploration value, and space score value; S3. Normalizing the contents of flavor substances representing the flavor type in all liquor samples by flavor substance type to obtain the normalized content of each flavor substance, and summing the normalized contents of each flavor substance in the same liquor sample to obtain a relative flavor complexity index for each liquor sample; S4. Establishing a machine learning model, wherein the machine learning model includes resampling, using flavor substances representing the aroma type and corresponding content data as input features, and using the flavor relative complexity index as a prediction target, training the machine learning model to obtain a flavor relative complexity index prediction model; S5. Optimizing hyperparameters of the flavor relative complexity index prediction model using a grid search cross-validation method to obtain an optimized flavor relative complexity index prediction model; S6. Inputting the flavor substances representing the aroma type and the corresponding content data of the liquor sample to be tested into the optimized liquor flavor relative complexity index prediction model to obtain the flavor relative complexity index of the liquor sample to be tested.
2. The method for predicting the relative complexity index of liquor flavor according to claim 1, characterized in that: The aroma is one of twelve aromas, including: strong aroma, sauce aroma, light aroma, rice aroma, phoenix aroma, mixed aroma, Dong aroma, sesame aroma, special aroma, fermented black bean aroma, Laobaigan aroma and rich aroma.
3. The method for predicting the relative complexity index of liquor flavor according to claim 1, characterized in that: Quantitative analysis includes gas chromatography-tandem mass spectrometry, gas chromatography-mass spectrometry, liquid chromatography-mass spectrometry, or flame ionization detection.
4. The method for predicting the relative complexity index of liquor flavor according to claim 1, characterized in that: Machine learning methods in the field of synthetic chemistry include synthetic complexity scoring, complexity index and space complexity score.
5. The method for predicting the relative complexity index of liquor flavor according to claim 1, characterized in that: The flavor substances representing the aroma type are screened out based on the five indexes, including: selecting the flavor substances corresponding to the top 30% of each index value arranged from large to small as the flavor substances for evaluating the relative complexity of the liquor flavor, or selecting the flavor substances whose each index value exceeds the corresponding threshold as the flavor substances for evaluating the relative complexity of the liquor flavor.
6. The method for predicting the relative complexity index of liquor flavor according to claim 1, characterized in that: The machine learning model is a random forest model, a linear regression model, a support vector machine model or a K-nearest neighbor model.