Method for predicting flue-cured tobacco note by fusing metabonomics with machine learning

By constructing an aroma score prediction model through metabolomics and machine learning, the subjectivity and data stability problems of traditional tobacco leaf aroma evaluation are solved, and rapid and accurate quantitative prediction of multiple aromas is achieved, supporting tobacco breeding and cigarette flavoring.

CN120685810APending Publication Date: 2025-09-23ZHENGZHOU TOBACCO RES INST OF CNTC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510847264.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional tobacco leaf aroma evaluation relies on manual sensory evaluation, which has problems of subjectivity and poor data stability. Existing technology makes it difficult to comprehensively predict aromas other than light fragrance and fruity aroma.

Method used

By collecting fresh tobacco leaf samples from different production areas and varieties, combining metabolomics and machine learning methods, a fragrance score prediction model was constructed. The SHAP method was used to interpret the model results and screen out the best prediction model.

Benefits of technology

It achieves rapid and accurate quantitative prediction of various tobacco leaf aromas, improves the efficiency and accuracy of evaluation, overcomes the limitations of traditional sensory evaluation, and provides a scientific basis to support tobacco breeding, cultivation and cigarette flavoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120685810A_ABST
    Figure CN120685810A_ABST
Patent Text Reader

Abstract

The invention discloses a method for predicting flue-cured tobacco aroma through metabonomics and machine learning, and belongs to the field of tobacco quality evaluation. According to the method, a note score prediction model is constructed through a machine learning method based on metabolic components of mature fresh tobacco leaves. The method comprises the following steps: sampling and measuring metabolome when tobacco leaves are mature in a field, baking and modulating samples in the same batch, and manually evaluating the note characteristics of the samples; performing correlation analysis on the metabolome data and the note score, and screening out characteristic metabolites; and a note score prediction model is further constructed through a machine learning method, and the contribution degree and the influence mode of different metabolites to note evaluation are explained and known in combination with the model. The method can predict the tobacco leaf note characteristics based on the fresh tobacco leaf metabolites, does not need to wait for tobacco leaf harvesting and modulation and expert evaluation, can obtain the tobacco leaf note characteristics in time, and has practical value for guiding tobacco breeding, tobacco leaf raw material allocation and guarantee of raw material supply.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for predicting the aroma of flue-cured tobacco by integrating metabolomics with machine learning, and belongs to the technical field of tobacco leaf quality evaluation. Background Art

[0002] In recent years, the rapid development of cigarette production technology has placed higher demands on the refined management of tobacco leaf raw materials. Cigarette aroma is the melody of the aroma expressed by the smoke and is an important component of the tobacco leaf's flavor characteristics. These include sweet, burnt, honey, burnt, and green aromas.

[0003] Traditional tobacco aroma evaluation relies on manual labor, requires high smoking skills, and suffers from poor data stability. The chemical composition of flue-cured tobacco leaves forms the material basis for the formation of tobacco aroma quality, and these chemical components are directly or indirectly derived from the accumulation of metabolites in fresh tobacco leaves. Therefore, exploring the intrinsic relationship between the sensory aroma of flue-cured tobacco leaves and the metabolites of fresh tobacco leaves, and constructing a metabolite-based aroma score prediction model, is of great significance for understanding the mechanism of tobacco aroma formation, guiding raw material production, and creating distinctive products.

[0004] A Chinese patent application with publication number CN202411774710 proposes a method for evaluating the light aroma of cigarette smoke. The method analyzes the aroma components of typical light-aroma cigarettes and non-light-aroma cigarettes, uses partial least squares discriminant analysis to obtain VIP values, and calculates the light aroma index Q. This method only predicts and evaluates light aroma and cannot obtain prediction results for other aromas. A Chinese patent application with publication number CN202411814044 proposes a method for evaluating the fruity aroma of cigarette smoke. The method captures the total particulate matter in the mainstream smoke of typical fruity and non-fruity cigarettes, performs GC-MS analysis, and calculates the fruity aroma index Q. However, this method primarily focuses on the evaluation of fruity cigarettes and has limited applicability to the evaluation of other aroma types. It also does not adequately consider the dynamic changes in aroma components. Summary of the Invention

[0005] The purpose of this paper is to propose a method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning. This method collects fresh tobacco leaf samples from different regions and varieties. Based on metabolomics analysis and aroma evaluation of flue-cured tobacco leaves, it uses machine learning to construct an aroma score prediction model. This model then screens aroma indices for predictability and further interprets those with relatively good aroma indices using the SHAP method.

[0006] The method of the present invention comprises the following steps:

[0007] (1) Collecting mature fresh tobacco leaf samples from different production areas and varieties;

[0008] (2) The collected fresh tobacco leaf samples were divided into two parts, one of which was used for metabolomics analysis;

[0009] (3) Another sample of fresh tobacco leaves was baked and cured, and the aroma index was manually evaluated after curing;

[0010] (4) Correlation analysis was performed between the metabolomics data and the aroma evaluation results to screen characteristic metabolites;

[0011] (5) Based on characteristic metabolites, different machine learning algorithms were compared to construct fragrance index prediction models and select the best prediction model;

[0012] (6) The prediction model is interpreted through the model interpretation method to reveal the contribution and influence of different metabolic components on the fragrance indicators.

[0013] Furthermore, in step (1), the sample collection includes selecting multiple representative tobacco varieties from multiple different production areas, and each variety is planted in each production area with no less than a certain number of plants to ensure the diversity and representativeness of the samples; in step (2), the metabolome determination can use liquid chromatography-mass spectrometry or gas chromatography-mass spectrometry technology to pre-treat the fresh tobacco leaf samples and extract metabolites, and calculate the average value of multiple biological replicates as the metabolome data to improve the reliability and stability of the data.

[0014] The beneficial effects of this technical solution are: by selecting multiple representative varieties from multiple producing areas and ensuring a sufficient number of plants for each variety, it effectively captures the genetic diversity and environmental variation of tobacco, avoids data bias, and improves the generalizability and application value of the model. Modern analytical techniques such as liquid chromatography-mass spectrometry (LC-MS) and gas chromatography-mass spectrometry (GC-MS) can widely detect primary and secondary metabolites in tobacco (such as sugars, organic acids, alkaloids, phenols, etc.), providing relatively comprehensive metabolome information.

[0015] Furthermore, in step (3), the fresh tobacco leaves are roasted and modulated using the local roasting process; the aroma evaluation is performed by a professional sensory evaluation team in accordance with relevant standards, and quantitative scores are given for various aroma indicators such as sweet aroma, burnt sweet aroma, honey sweet aroma, burnt aroma, green aroma, mellow sweet aroma, roasted aroma, hay aroma, woody aroma, and nutty aroma. The average of each result is taken, and the highest score of the indicator is set to a certain score.

[0016] The beneficial effects of this technical solution include: using local curing techniques for blending, it can truly reflect the actual processing conditions of tobacco from different producing areas, avoiding data bias caused by inconsistent curing processes. A professional sensory evaluation team, guided by standards, ensures the reliability and consistency of evaluation results and reduces subjective errors. Quantified scoring is performed for multiple aroma indicators (such as light sweet, burnt sweet, honey sweet, burnt aroma, and green aroma), covering key dimensions of tobacco aroma and providing a more comprehensive analysis.

[0017] Furthermore, in step (4), the association analysis first screens metabolites through Spearman rank correlation analysis, and then uses the random forest algorithm to calculate the feature importance to screen out metabolites that are closely related to the fragrance score.

[0018] Furthermore, in step (5), the different machine learning algorithms may include linear regression, K-nearest neighbor, support vector machine regression, random forest regression and gradient boosting decision tree, etc., and these algorithms are used to build models respectively, and then the root mean square error (RMSE), determination coefficient (R 2 ) and mean absolute error (MAE) and other evaluation indicators are used to evaluate the model prediction performance and select the prediction model.

[0019] The beneficial effects of this technical solution include: using feature selection methods to efficiently screen out characteristic metabolites significantly associated with aroma notes from a vast pool of metabolites, avoiding the subjectivity of manual screening and ensuring the reliability of the results. Comparing multiple machine learning algorithms and selecting the optimal prediction model can improve the accuracy of aroma score predictions.

[0020] Furthermore, in step (6), the model interpretation method can use methods such as SHAP local interpretation, feature importance or local dependence diagram to reveal the intrinsic relationship between metabolites and fragrance indicators.

[0021] The beneficial effects of this technical solution include: using SHAP (Shapley Additive Explanations) value analysis to quantify the contribution of each metabolite to aroma prediction from a game-theoretic perspective, making the "black box" machine learning model transparent and enhancing the credibility of the results. Combining feature importance ranking and local dependency graphs (PDP / ICE) to visually demonstrate the dose-effect relationship between key metabolites and aroma, consistent with biological logic.

[0022] In summary, the advantages of this invention over existing technologies are as follows: It utilizes metabolome analysis technology to comprehensively analyze metabolites in tobacco leaves, providing a richer material basis for aroma evaluation; it uses machine learning algorithms to construct a predictive model, enabling rapid and accurate quantitative prediction of aroma; and it uses model interpretation technology to deeply reveal the intrinsic relationship between metabolites and aroma formation, providing a scientific basis for tobacco breeding, cultivation, and cigarette flavoring. Furthermore, this invention can effectively overcome the subjectivity and limitations of traditional sensory evaluation, improve the efficiency and accuracy of aroma evaluation, and provide strong support for production and research and development in the tobacco industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a flow chart of the present invention for predicting tobacco aroma by integrating metabolomics with machine learning.

[0024] Figure 2 The data distribution of different fragrance ratings.

[0025] Figure 3 Metabolome analysis of different fresh tobacco leaf samples.

[0026] Figure 4 It is the number of characteristic metabolites of different fragrance indicators.

[0027] Figure 5 It is a scatter plot of the evaluation value and the predicted value.

[0028] Figure 6 This is the influence of monomeric characteristic metabolites on the fragrance prediction results. DETAILED DESCRIPTION

[0029] The present invention will be further described in detail below with reference to the accompanying drawings:

[0030] Example:

[0031] The present invention provides a method for predicting the aroma of flue-cured tobacco by integrating metabolomics with machine learning, and the flow chart thereof is as follows: Figure 1 As shown, the following steps are included:

[0032] 1. Collect samples

[0033] Twenty-one experimental accessions with excellent flue-curing performance and distinct aroma and flavor profiles were selected from the Hunan Tobacco Germplasm Bank (Table 1). These tobaccos originated from three production areas: Chenzhou, Hunan (a Nanling Hilly region producing a sweet and mellow aroma tobacco), Xiangxi, Hunan (Wuling and Qinba regions producing a mellow and sweet aroma tobacco), and Yuxi, Yunnan (a Southwest Plateau region producing a light and sweet aroma tobacco). At least 100 plants of each variety were planted in each region. Thirty leaves of each variety, near the middle of mature leaves, were collected at the time of mid-leave maturity. Fifteen of these leaves were removed from the main vein, wrapped in tin foil, and formed into three replicates (three biological replicates). These leaves were frozen in liquid nitrogen and used for metabolomics analysis. The remaining 15 leaves were labeled and hung in the center of the middle flue-curing barn with other varieties. They were then cured using the mid-leave curing process for the dominant local cultivar, K326.

[0034] Table 1 Tobacco varieties tested

[0035]

[0036] 2. Fragrance evaluation

[0037] The flue-cured tobacco leaves were evaluated with reference to the "Sensory Evaluation Method for Quality, Style and Characteristics of Flue-cured Tobacco Leaves" (YC / T530-2015). Nine experts with the qualification to conduct sensory evaluation of cigarette products were invited to form an evaluation team. The samples were evaluated on 10 aroma indicators, including sweet aroma, burnt sweet aroma, honey sweet aroma, burnt aroma, green aroma, mellow sweet aroma, baked aroma, hay aroma, wood aroma and nutty aroma. The average of each result was taken, and the highest score of the indicator was 5 points. A total of 63 valid samples were obtained.

[0038] Depend on Figure 2 As can be seen, tobacco leaves from different producing areas exhibit distinct aroma styles. Chenzhou tobacco leaves score highly for burnt-sweet, burnt, roasted, and nutty aromas, with burnt-sweet and burnt aromas mostly scoring between 1.5 and 2.0. Xiangxi tobacco leaves generally score between Chenzhou and Yuxi for all aromas, with only mellow sweetness and woody aroma scoring slightly higher than the other leaves. Yuxi tobacco leaves score much higher for light sweetness and green aroma than Chenzhou and Xiangxi. Yuxi's light sweetness scores range from 0.5 to 2.5, while the other two areas score 0. Yuxi's green aroma scores mostly range from 0.5 to 1.0, while the other two areas score below 0.5. In summary, Chenzhou tobacco leaves exhibit distinct burnt-sweet, burnt, roasted, and nutty aromas, while Yuxi tobacco leaves have more prominent light sweetness and green aromas.

[0039] Statistical analysis of all samples shows that (Table 2) the coefficients of variation of the sweet and nutty aromas are relatively large, both above 100%; followed by the honey sweet and burnt sweet aromas, which are 81% and 66% respectively; the coefficients of variation of the remaining aromas are all below 60%, among which the hay aroma is the smallest at only 3%, so the hay aroma indicator is removed in subsequent modeling.

[0040] Table 2 Standard deviation and coefficient of variation of tobacco leaf samples

[0041]

[0042] 3. Metabolome Analysis

[0043] Fresh tobacco leaf samples frozen in liquid nitrogen were vacuum freeze-dried and ground into powder. 20 mg of each sample was weighed and subjected to metabolite extraction and metabolome analysis using a derivatization-based GC-MS method. The average of three biological replicates was calculated as the metabolome for each fresh tobacco leaf variety for subsequent data analysis. A total of 131 metabolites were detected in the mature fresh tobacco leaf samples, including 37 organic acids, 30 sugars, 21 amino acids, 8 alkaloids, 5 steroids, 4 phenylpropanoids, 4 alcohols, 3 amines, 3 heterocyclics, and 16 other metabolites.

[0044] The results of principal component analysis showed that ( Figure 3 A), fresh tobacco samples have a certain separation trend according to origin. Through PLS-DA analysis, 60 metabolites with a large contribution to the origin grouping (VIP>1) were obtained. Among them, the content comparison of the 6 metabolites with the largest VIP in different origins is shown in Figure 2. Figure 3 B. The content of L-mimosine was higher in tobacco leaves from the Yuxi region, while the content of acetol, 2-hydroxyglutaric acid, ribonucleo-γ-lactone, and D-malic acid was higher in tobacco leaves from the Chenzhou region, and the content of pyruvic acid was higher in the western Hunan region.

[0045] 4. Feature Selection

[0046] A total of 131 metabolites were detected by GC-MS, and the number of features was higher than the sample size (63). In order to avoid overfitting of the model, a multi-stage feature selection strategy was used to screen the characteristic metabolites in the metabolomics data. In the feature selection process, the original metabolite concentration data was first normalized by Z-score to eliminate dimensional differences, and then FDR correction (p<0.05) was performed by Spearman rank correlation analysis and Benjamini-Hochberg method to preliminarily screen metabolites that were significantly correlated with the target aroma traits. On this basis, the Random Forest algorithm (RandomForest) was used to evaluate the importance of features, and the metabolites with the top 20% importance scores were retained as the final feature set. Figure 4 As can be seen, the Spearman-FDA correction method screened a relatively large number of characteristic metabolites for the fresh sweet, burnt sweet, burnt, and green aromas, all exceeding 60. However, the number of characteristic metabolites retained for the honey sweet, roasted, and woody aromas was relatively small, below 50, with only two for woody aroma. Further feature importance selection using random forests revealed that, with the exception of honey sweet, roasted, and woody aromas, which had fewer than 10 characteristic metabolites, the remaining aroma indices had 11 to 14 characteristic metabolites. Due to the small number of characteristic metabolites for woody aroma, woody aroma was not modeled in subsequent analyses.

[0047] 5. Model Construction

[0048] In order to explore the performance of different machine learning algorithms in predicting tobacco leaf aroma scores, this embodiment selects five widely used machine learning algorithms, namely linear regression, K-nearest neighbor, support vector machine regression, random forest and gradient boosting decision tree, for model construction and analysis. Due to the small sample size of data (63), it is divided into training set and test set at 8:2. The training set data (50) is modeled using 10-fold cross validation, and the test set (13) data is used to evaluate the model performance. The metabolomics data are preprocessed by standardization method. After 100 random iterations of training, R is selected. 2 The algorithm with the highest RMSE and the lowest MAE is the best algorithm. 2 ) and mean absolute error (MAE) are used to evaluate the model prediction performance. The smaller the MAE and RMSE values ​​are, the better the model prediction performance is; 2 The closer the value is to 1, the better the model fit is. MAE, RMSE and R 2 The calculation formula is as follows:

[0049]

[0050] Where: is the mean of all observations, y i is the i-th observation in the sample, is the predicted value corresponding to the i-th observation value, and n is the total number of samples.

[0051] As shown in Table 3, the model prediction performance of nutty aroma, burnt sweet aroma and mellow sweet aroma is better, and the R 2 All are above 0.85; the performance of the models for green, burnt, and sweet aromas are similar, and the test set R 2 The prediction performance of the honey sweet aroma and baking aroma detection models is poor, and the test set R 2 The prediction rates for nutty, sweet, and green aromas were only 0.53 and 0.44, respectively. RF was the best algorithm for predicting nutty, sweet, and green aromas, while SVR performed best for predicting burnt sweet aroma. Linear and KNN outperformed other algorithms for predicting burnt aroma, light sweet aroma, honey sweet aroma, and baked aroma.

[0052] Table 3 Best algorithm and model performance

[0053]

[0054] The comparison chart of the evaluation value and the predicted value shows that ( Figure 5), the prediction results of nutty aroma, caramel sweet aroma, and mellow sweet aroma are close to the true values, and the fitting line and diagonal basically coincide with each other; there are relatively more outliers for green aroma, caramel aroma, and light sweet aroma, and the deviation between the fitting line and the diagonal is large. The predicted value of light sweet aroma is mostly smaller than the true value; the predicted values ​​of honey sweet aroma and baking aroma deviate greatly from the true value, and the prediction results are mostly concentrated in a certain range, indicating that the model is difficult to fit the intrinsic relationship between sensory scores and characteristic metabolites.

[0055] 6. Model Interpretation

[0056] This example uses the SHAP (SHapley Additive exPlanations) method to interpret model prediction results. This method is based on the Shapley value in game theory and quantifies the contribution of each feature to the model prediction by distributing benefits fairly.

[0057] Figure 6 The influence of monomeric characteristic metabolites on the prediction results was demonstrated. Acetone had the greatest impact on the prediction results of the three aroma indices, and a positive correlation was observed. In addition to acetone, the nutty aroma was significantly affected by erythritol, 2-hydroxyglutaric acid, mucic acid, and galacturonic acid, all of which were positively correlated. Among the top seven metabolites with the greatest impact on the sweet aroma, except for acetone and ribonucleic acid-γ-lactone, the others were all organic acid metabolites: oleic acid, mucic acid, 2-hydroxyglutaric acid, 2-keto-L-gluconic acid, and D-malic acid, and all of which were positively correlated with the prediction results. The sweet aroma was most significantly affected by acetone, far exceeding the influence of other metabolites, indicating that the formation of the sweet aroma may be closely related to the acetone content.

Claims

1. A method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning, characterized by: Fresh tobacco leaf samples from different regions and varieties were collected. Based on metabolome analysis and aroma evaluation of flue-cured tobacco leaves, a machine learning approach was used to construct an aroma score prediction model. The predictability of aroma indices was screened, and the SHAP method was used to further interpret the aroma indices with relatively good results. The method includes the following steps: (1) Collecting mature fresh tobacco leaf samples from different production areas and varieties; (2) The collected fresh tobacco leaf samples were divided into two parts, one of which was used for metabolomics analysis; (3) Another sample of fresh tobacco leaves was baked and cured, and the aroma index was manually evaluated after curing; (4) Correlation analysis was performed between the metabolomics data and the aroma evaluation results to screen characteristic metabolites; (5) Based on characteristic metabolites, different machine learning algorithms were compared to construct fragrance index prediction models and select the best prediction model; (6) The prediction model is interpreted through the model interpretation method to reveal the contribution and influence of different metabolic components on the fragrance indicators.

2. The method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning according to claim 1, characterized in that: In step (1), the sample collection includes selecting multiple representative tobacco varieties from multiple different production areas, and each variety is planted in each production area with a minimum number of plants to ensure the diversity and representativeness of the samples.

3. The method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning according to claim 1, characterized in that: In step (2), the metabolome determination can be performed using liquid chromatography-mass spectrometry or gas chromatography-mass spectrometry technology. The fresh tobacco leaf samples are pretreated to extract metabolites, and the average value of multiple biological replicates is calculated as the metabolome data to improve the reliability and stability of the data.

4. The method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning according to claim 1, characterized in that: In step (3), the fresh tobacco leaves are roasted and conditioned using the local roasting process; the aroma evaluation is performed by a professional sensory evaluation team in accordance with relevant standards, and quantitative scoring is performed on a variety of aroma indicators such as sweet aroma, burnt sweet aroma, honey sweet aroma, burnt aroma, green aroma, mellow sweet aroma, roasted aroma, hay aroma, woody aroma, and nutty aroma. The average of each result is taken, and the highest score of the indicator is set to a certain score.

5. The method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning according to claim 1, characterized in that: In step (4), the association analysis first screens metabolites through Spearman rank correlation analysis, and then uses the random forest algorithm to calculate the feature importance to screen out metabolites that are closely related to the fragrance score.

6. The method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning according to claim 1, characterized in that: In step (5), the different machine learning algorithms include linear regression, K-nearest neighbor, support vector machine regression, random forest regression and gradient boosting decision tree, etc., and these algorithms are used to build models respectively, and then the root mean square error (RMSE), determination coefficient (R 2 ) and mean absolute error (MAE) evaluation indicators were used to evaluate the model prediction performance and select the best prediction model.

7. The method for predicting flue-cured tobacco aroma by integrating metabolomics with machine learning according to claim 1, characterized in that: In step (6), the model interpretation method can use SHAP local interpretation, supplemented by random forest feature importance ranking or local dependence graph method to reveal the intrinsic relationship between metabolites and fragrance indicators.

Citation Information

Patent Citations

  • Evaluation method for faint scent of cigarette smoke and application of evaluation method

    CN119291084A

  • Evaluation method for fruity fragrance of cigarette smoke and application of evaluation method

    CN119291086A