Method and device for predicting aroma type of flue-cured tobacco, electronic device and storage medium

By evaluating the aroma type and determining the chemical composition of flue-cured tobacco leaves, and combining separation screening, random forest importance screening, and correlation redundancy removal screening, a flue-cured tobacco aroma type prediction model was constructed. This model solved the problem of low accuracy in existing technologies and achieved efficient and stable aroma type prediction.

CN122288486APending Publication Date: 2026-06-26CHINA TOBACCO ZHEJIANG IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610404980.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-30
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in predicting the aroma type of flue-cured tobacco, making it difficult to capture the complex nonlinear relationship between aroma type and chemical components, and are easily affected by subjective factors.

Method used

By obtaining the aroma type evaluation results of flue-cured tobacco leaves, assigning label values, measuring chemical components and performing feature derivation, and using separation degree screening, random forest importance screening, correlation redundancy removal screening, and comprehensive cost consideration, a flue-cured tobacco aroma type prediction model is constructed, and a core indicator set is selected for prediction.

Benefits of technology

It improves the accuracy of predicting the aroma type of flue-cured tobacco, reduces detection costs, and enhances the objectivity and stability of prediction results, making it suitable for rapid screening and quality evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122288486A_ABST
    Figure CN122288486A_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, electronic device, and storage medium for predicting the aroma type of flue-cured tobacco. The method includes: acquiring and obtaining a label value for the target flue-cured tobacco leaf based on the evaluation results of its aroma type; measuring and derivation of the chemical components of the target flue-cured tobacco leaf to obtain a feature-derived dataset; sequentially filtering the feature-derived dataset using separation degree, random forest importance, relevance redundancy removal, and comprehensive cost considerations to obtain a final dataset; training a flue-cured tobacco aroma type classification prediction model using the final dataset as input and the label value as output; acquiring the content of each chemical component index in the final chemical component index dataset of the flue-cured tobacco leaf to be tested, inputting it into the flue-cured tobacco aroma type classification prediction model, and outputting the aroma type prediction result of the flue-cured tobacco leaf to be tested. By using separation degree, importance, and relevance filtering, the accuracy of the aroma type prediction model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of tobacco quality evaluation technology, and in particular to methods, apparatus, electronic devices and storage media for predicting the aroma type of flue-cured tobacco. Background Technology

[0002] The aroma of flue-cured tobacco is a core quality indicator that reflects the style characteristics of tobacco leaves. It directly determines the sensory quality and style positioning of cigarette products and is a key factor affecting consumer experience, brand recognition, and product market competitiveness. It has important guiding significance for cigarette formula design and raw material grading.

[0003] Currently, traditional methods of manual sensory evaluation are time-consuming and labor-intensive, and easily influenced by subjective factors, resulting in poor consistency and stability of classification results. Existing technologies typically use linear models (such as stepwise linear regression) to predict the aroma type of flue-cured tobacco, which struggles to capture the complex nonlinear relationship between aroma type and chemical components, and is prone to overfitting. This leads to the low accuracy of current flue-cured tobacco aroma type prediction.

[0004] There is currently no effective solution to the problem of low accuracy in related technologies. Summary of the Invention

[0005] This embodiment provides a method, apparatus, electronic device, and storage medium for predicting the aroma type of flue-cured tobacco, in order to solve the problem of low accuracy in related technologies.

[0006] In the first aspect, this embodiment provides a method for predicting the aroma type of flue-cured tobacco, including: obtaining the evaluation result of the aroma type of the target flue-cured tobacco leaf, and assigning a value to the target flue-cured tobacco leaf according to the evaluation result to obtain the label value of the target flue-cured tobacco leaf;

[0007] The chemical composition of the target flue-cured tobacco leaves was determined and its features were derived to obtain a feature-derived dataset;

[0008] The chemical composition indicators in the feature-derived dataset are sequentially filtered by separation degree, random forest importance, correlation redundancy removal, and comprehensive cost consideration to obtain the final chemical composition indicator dataset.

[0009] Using the final chemical composition index dataset as input and the label value of the target flue-cured tobacco leaf as output, a flue-cured tobacco aroma classification and prediction model is trained.

[0010] Obtain the content of each chemical component index in the final chemical component index dataset of the flue-cured tobacco leaf to be tested, input it into the flue-cured tobacco aroma type classification prediction model, and output the aroma type prediction result of the flue-cured tobacco leaf to be tested.

[0011] In some embodiments, the determination of the chemical composition of the target flue-cured tobacco leaves and the subsequent feature derivation to obtain a feature-derived dataset include:

[0012] The chemical composition of the target flue-cured tobacco leaves was determined to obtain the original chemical composition dataset;

[0013] Based on the original chemical composition dataset, feature derivation is performed through addition, subtraction, and division operations to obtain a feature-derived dataset.

[0014] In some embodiments, the chemical composition indicators in the feature-derived dataset are sequentially processed through separation filtering, random forest importance filtering, relevance redundancy removal filtering, and comprehensive cost considerations to obtain the final chemical composition indicator dataset, including:

[0015] The separation degree is used to filter the chemical composition indicators in the feature-derived dataset to obtain the first feature subset;

[0016] Using the first feature subset as input, a random forest classification model is constructed to perform the random forest importance screening, thereby obtaining the second feature subset;

[0017] The chemical composition indicators in the second feature subset are subjected to the aforementioned correlation redundancy removal screening to obtain the third feature subset;

[0018] The chemical composition indicators in the third feature subset are subjected to the comprehensive cost consideration to determine the final chemical composition indicator dataset.

[0019] In some embodiments, the separation filtering of chemical composition indicators in the feature-derived dataset to obtain a first feature subset includes:

[0020] Calculate the separation degree value of each chemical component index in the feature-derived dataset relative to the aroma type of different target flue-cured tobacco leaves, and retain the chemical component indexes with separation degree values ​​greater than a preset first threshold to obtain the first feature subset.

[0021] In some embodiments, the step of constructing a random forest classification model using the first feature subset as input to perform random forest importance filtering and obtain a second feature subset includes:

[0022] Using the first feature subset as input, construct a random forest classification model;

[0023] Calculate the contribution of each chemical component index in the second feature subset to the classification of flue-cured tobacco aroma type, obtain the feature importance of the chemical component index, and rank the chemical component index according to the feature importance;

[0024] Chemical composition indicators whose feature importance ranking is greater than a preset second threshold are retained to obtain the second feature subset.

[0025] In some embodiments, the process of performing the correlation redundancy removal screening on the chemical composition indicators in the second feature subset to obtain the third feature subset includes:

[0026] Calculate the correlation coefficients among all chemical component index pairs in the second feature subset;

[0027] For chemical component index pairs whose absolute values ​​of the correlation coefficients are greater than a preset third threshold, the chemical component indexes with high feature importance are retained to obtain a third feature subset.

[0028] In some embodiments, the step of performing the comprehensive cost consideration on the chemical composition indicators in the third feature subset to determine the final chemical composition indicator dataset includes:

[0029] Based on the detection cost of each chemical component index in the third feature subset and the importance of the features, the final chemical component index dataset is determined.

[0030] Secondly, this embodiment provides a device for predicting the aroma type of flue-cured tobacco, including: a dataset construction module, a feature selection module, and a model construction and application module; wherein:

[0031] The dataset construction module is used to obtain the evaluation results of the aroma type of the target flue-cured tobacco leaves, and assign values ​​to the target flue-cured tobacco leaves according to the evaluation results to obtain the label values ​​of the target flue-cured tobacco leaves; the chemical components of the target flue-cured tobacco leaves are measured and features are derived to obtain the feature-derived dataset;

[0032] The feature filtering module is used to sequentially filter the chemical composition indicators in the feature-derived dataset through separation degree filtering, random forest importance filtering, correlation redundancy removal filtering, and comprehensive cost consideration to obtain the final chemical composition indicator dataset.

[0033] The model building and application module is used to take the final chemical component index dataset as input and the label value of the target flue-cured tobacco leaf as output to train a flue-cured tobacco aroma type classification prediction model; obtain the content of each chemical component index in the final chemical component index dataset of the flue-cured tobacco leaf to be tested, input it into the flue-cured tobacco aroma type classification prediction model, and output the aroma type prediction result of the flue-cured tobacco leaf sample to be tested.

[0034] Thirdly, this embodiment provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the flue-cured tobacco aroma type prediction method described in the first aspect above.

[0035] Fourthly, this embodiment provides a storage medium storing a computer program that, when executed by a processor, implements the steps of the flue-cured tobacco aroma type prediction method described in the first aspect.

[0036] Compared with related technologies, this embodiment provides a method, apparatus, electronic device, and storage medium for predicting the aroma type of flue-cured tobacco. The method first obtains the evaluation results of the aroma type of the target flue-cured tobacco leaves and assigns a label value to the target tobacco leaves based on the evaluation results. Second, it measures the chemical components of the target tobacco leaves and performs feature derivation to obtain a feature-derived dataset. Next, it sequentially filters the chemical component indicators in the feature-derived dataset through separation degree screening, random forest importance screening, correlation redundancy removal screening, and comprehensive cost consideration to obtain a final chemical component indicator dataset. Further, it uses the final chemical component indicator dataset as input and the label value of the target flue-cured tobacco leaves as output to train a flue-cured tobacco aroma type classification prediction model. Finally, it obtains the content of each chemical component indicator in the final chemical component indicator dataset of the flue-cured tobacco leaves to be tested, inputs it into the flue-cured tobacco aroma type classification prediction model, and outputs the aroma type prediction result of the flue-cured tobacco leaves to be tested. It selects a core set of indicators from a massive amount of chemical component indicators through progressive screening, including separation degree screening, random forest importance screening, and correlation redundancy removal screening. Based on the core set of indicators, it constructs a flue-cured tobacco aroma type prediction model, which can improve the accuracy of the flue-cured tobacco aroma type prediction model.

[0037] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0038] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0039] Figure 1 This is a hardware structure block diagram of a terminal for a method for predicting the aroma type of flue-cured tobacco according to an embodiment of this application;

[0040] Figure 2 This is a flowchart of a method for predicting the aroma type of flue-cured tobacco according to an embodiment of this application;

[0041] Figure 3 This is a confusion matrix diagram of the classification results obtained by training a random forest model using a filtered feature dataset according to one embodiment of this application;

[0042] Figure 4 This is a feature importance score map of chemical component indicators retained after separation degree filtering in the original feature dataset of one embodiment of this application;

[0043] Figure 5 This is a graph showing the importance scores of the top 10 chemical component indicators in terms of feature importance in an additive composite feature dataset according to one embodiment of this application.

[0044] Figure 6 This is a graph showing the importance scores of the top 10 chemical component indicators in terms of feature importance in a subtractive composite feature dataset according to one embodiment of this application.

[0045] Figure 7 This is a graph showing the importance scores of the top 10 chemical component indicators in terms of feature importance in a division composite feature dataset according to one embodiment of this application;

[0046] Figure 8 This is a heat dissipation diagram showing the correlation coefficients of chemical composition index pairs in the original feature dataset in one embodiment of this application;

[0047] Figure 9 This is a heatmap of the correlation coefficients of chemical component index pairs in an additive composite feature dataset according to one embodiment of this application;

[0048] Figure 10 This is the correlation coefficient of chemical composition index pairs in the subtractive composite feature dataset in one embodiment of this application;

[0049] Figure 11 This is a heatmap of the correlation coefficients of chemical component index pairs in a division composite feature dataset according to one embodiment of this application;

[0050] Figure 12 This is a bar chart showing the average absolute SHAP values ​​of the feature importance of a flue-cured tobacco aroma type classification prediction model according to an embodiment of this application;

[0051] Figure 13 This application presents an embodiment of the influence of SHAP's sweet aroma characteristics on bee colony diagrams.

[0052] Figure 14 This application presents an embodiment of the influence of SHAP on bee colony diagrams based on the characteristics of honey sweetness.

[0053] Figure 15 This is an embodiment of the application based on the influence of SHAP's sweet aroma characteristics on bee colony diagrams;

[0054] Figure 16This is a scatter plot showing the characteristic dependence of rutin / xylitol in a sweet and refreshing flavor profile in one embodiment of this application;

[0055] Figure 17 This is a scatter plot showing the characteristic dependence of rutin / xylitol in a honey-sweet flavor profile in one embodiment of this application;

[0056] Figure 18 This is a scatter plot showing the characteristic dependence of rutin / xylitol in the sweet alcohol flavor profile in one embodiment of this application;

[0057] Figure 19 This is a scatter plot showing the characteristic dependence of total sugar + proline in a sweet and refreshing flavor in one embodiment of this application;

[0058] Figure 20 This is a scatter plot showing the characteristic dependence of total sugar + proline in a honey-sweet flavor in one embodiment of this application;

[0059] Figure 21 This is a scatter plot showing the characteristic dependence of total sugar + proline in the alcoholic sweet flavor profile in one embodiment of this application;

[0060] Figure 22 This is a structural block diagram of a flue-cured tobacco aroma type prediction device according to an embodiment of this application. Detailed Implementation

[0061] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0062] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.

[0063] The method embodiments provided in this example can be executed on a terminal, computer, or similar electronic device with a certain computing power. For example, it can run on a terminal. Figure 1 This is a hardware structure block diagram of a terminal for a method for predicting the aroma type of flue-cured tobacco according to an embodiment of this application. Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 and a memory 104 for storing data are also included. The processor 102 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.

[0064] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the flue-cured tobacco aroma type prediction method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer programs stored in the memory 104, thereby implementing the aforementioned method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0065] The transmission device 106 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0066] One embodiment of this application provides a method for predicting the aroma type of flue-cured tobacco. Figure 2 This is a flowchart of a method for predicting the aroma type of flue-cured tobacco according to an embodiment of this application, as shown below. Figure 2 As shown, the process includes the following steps:

[0067] Step S210: Obtain the evaluation result of the aroma type of the target flue-cured tobacco leaf, and assign a value to the target flue-cured tobacco leaf according to the evaluation result to obtain the label value of the target flue-cured tobacco leaf.

[0068] The aroma type of flue-cured tobacco leaves refers to the overall style characteristics of the volatile aroma components released during the combustion process, and is one of the important indicators for evaluating the quality of tobacco leaves. According to the People's Republic of China Tobacco Industry Standard YC / T 530, flue-cured tobacco leaves can be divided into three aroma types based on their aroma style characteristics: light and sweet aroma, honey-sweet aroma, and mellow and sweet aroma.

[0069] Specifically, to obtain the aroma type evaluation results of the experimental samples, a total of 600 flue-cured tobacco leaves from representative production areas in Sichuan Province over three consecutive years were selected as experimental samples. The aroma types of the experimental samples were evaluated according to the People's Republic of China Tobacco Industry Standard YC / T 530. Based on the evaluation results, the aroma type of each sample was determined and assigned values ​​of 0, 1, and 2, respectively, where 0 represents a light and sweet aroma, 1 represents a honey-sweet aroma, and 2 represents a mellow and sweet aroma. These values ​​were used for subsequent tobacco leaf quality analysis and the construction of a flue-cured tobacco aroma type classification prediction model.

[0070] Step S220: The chemical composition of the target flue-cured tobacco leaves is determined and features are derived to obtain a feature-derived dataset.

[0071] The chemical composition of flue-cured tobacco leaves refers to the intrinsic substances that affect their quality and aroma characteristics. It mainly includes conventional chemical indicators such as total sugar, reducing sugar, total nitrogen, nicotine, potassium, and chlorine, as well as components closely related to tobacco quality, such as starch, protein, and polyphenols. Specifically, the determination of total sugar, total alkaloids, total nitrogen, potassium, chlorine, starch, protein, polyphenols, single alkaloids, polybasic acids, and higher fatty acids follows the standards of the Tobacco Industry of the People's Republic of China. Free amino acid content is determined quantitatively using an automated amino acid analyzer after hydrolysis extraction and by external standard method. Sugar alcohols are determined qualitatively and quantitatively using gas chromatography-mass spectrometry (GC-MS) after derivatization treatment and by standard analysis. By determining these chemical components, basic data characterizing the intrinsic quality of tobacco leaves can be obtained.

[0072] Based on the obtained chemical composition determination results, to further explore the intrinsic relationship between chemical components and aroma types, feature derivation can be performed on the original chemical component indicators. Feature derivation refers to generating new feature variables by performing mathematical transformations, combination calculations, or ratio operations on the original chemical indicators, thereby enhancing the expressive power of the data and the predictive performance of the flue-cured tobacco aroma classification prediction model. Methods for feature derivation of chemical components in flue-cured tobacco leaves can include: calculating derived indicators such as the ratio of total sugar to nicotine, the ratio of total nitrogen to nicotine, and the potassium-chlorine ratio, as well as generating higher-order features through polynomial transformations and interactive feature construction.

[0073] Step S230: For the chemical composition indicators in the feature-derived dataset, the final chemical composition indicator dataset is obtained by sequentially filtering through separation degree, random forest importance, correlation redundancy removal, and comprehensive cost consideration.

[0074] After constructing the feature-derived dataset, to reduce data dimensionality and enhance feature interpretability, the chemical component indicators in the dataset underwent multiple rounds of screening. Features are variables characterizing the chemical composition of tobacco leaves. First, features with weak discriminative power were removed through separation screening, retaining those that effectively distinguish different aroma types. For example, in the original features, rutin, chlorogenic acid, neonicotinoids, and serine had separation values ​​exceeding the preset separation threshold, indicating significant distribution differences among sweet, honey-sweet, and mellow-sweet aroma samples, effectively distinguishing different aroma types and thus possessing strong discriminative power. Other features, such as potassium and chlorine, common chemical components, showed significant overlap in distribution across different aroma type samples, resulting in low separation values ​​and difficulty in effectively distinguishing different aroma types; these were removed during the separation screening stage. Next, random forest importance screening was used, ranking features based on their contribution to aroma type classification, retaining highly important features and removing redundant variables with minimal impact on aroma type discrimination. Next, a correlation redundancy removal screening process is performed, calculating the correlation coefficients between features and eliminating highly correlated features to avoid the impact of multicollinearity on the stability of subsequent data analysis. Finally, considering comprehensive cost factors such as detection cost, detection cycle, and operational complexity, priority is given to feature indicators that are easy to obtain and have lower detection costs.

[0075] Through the above four rounds of screening, features with good discriminative ability, high importance, low correlation and reasonable detection cost are selected from the original feature-derived dataset to form the final chemical composition index dataset.

[0076] Step S240: The final chemical composition index dataset is used as input, and the label value of the target flue-cured tobacco leaf is used as output to train a flue-cured tobacco aroma classification prediction model.

[0077] Specifically, the final chemical composition index dataset is a set of chemical composition indicators obtained through multiple rounds of screening. These indicators possess good discriminative power, high importance, low correlation, and reasonable detection costs, effectively characterizing the intrinsic chemical quality of flue-cured tobacco leaves. The label values ​​are quantitative results assigned based on the evaluation of tobacco leaf aroma types according to the People's Republic of China Tobacco Industry Standard YC / T 530, with 0, 1, and 2 representing light sweet aroma, honey sweet aroma, and mellow sweet aroma, respectively. Using chemical composition indicators as input and aroma type labels as output, a quantitative mapping relationship can be established between the intrinsic chemical composition of tobacco leaves and their extrinsic aroma style characteristics.

[0078] A tobacco aroma type classification and prediction model was constructed and trained. During training, the final chemical component index dataset and its corresponding label values ​​were first divided into a training set and a validation set. The training set was used for learning the tobacco aroma type classification and prediction model, while the validation set was used to evaluate its performance. Supervised learning algorithms, such as random forests, support vector machines, gradient boosting trees, or neural networks, were employed to learn the mapping relationship between input features and output labels. The tobacco aroma type classification and prediction model iteratively optimized its internal parameters to minimize the error between the prediction results and the true labels, gradually fitting the complex nonlinear relationship between chemical component indices and aroma type. During training, cross-validation and other methods were used to prevent overfitting and ensure the model had good generalization ability.

[0079] To verify the model performance, six machine learning algorithms were selected for performance evaluation: K-Nearest Neighbors, Random Forest, Decision Tree, Logistic Regression, Extreme Gradient Boosting, and Partial Least Squares Discriminant Analysis. Table 1 shows the performance evaluation results of the six machine learning algorithms in one embodiment of this application. In Table 1, accuracy is the core indicator. Decision Tree (87.1%) and Partial Least Squares Discriminant Analysis (82.8%) performed relatively poorly, while the accuracy of the other four models all exceeded 90%, with the Random Forest model achieving an accuracy of 93.5%. Among all models, Random Forest had the highest accuracy (94.1%), indicating that it still has strong classification accuracy under imbalanced data. Overall, the Random Forest model performs excellently on the dataset.

[0080] Table 1

[0081]

[0082] To further verify the impact of feature derivation and selection on model performance, random forest models were constructed based on three different datasets: the original dataset, the original dataset plus all derived features, and the selected dataset. Table 2 shows the evaluation metric scores of the random forest models constructed on the three different datasets in one embodiment of this application. The results in Table 2 show that, due to the excessive number of features, the accuracy of the model with the original dataset plus all derived features slightly decreased from 93.5% to 93.0% on the original dataset. After feature selection, the number of features was reduced from 51 to 7, and the model accuracy slightly decreased to 91.9%. However, the precision, recall, and F1 score remained at a good level, indicating that multiple rounds of feature selection can significantly reduce detection costs while maintaining good model performance.

[0083] Table 2

[0084]

[0085] Figure 3 This is a confusion matrix diagram of the classification results obtained by training a random forest model using a filtered feature dataset, according to one embodiment of this application. Figure 3 As can be seen, the confusion matrix, presented as a heatmap, shows the classification of samples of three aroma types (sweet, honey, and mellow) by the random forest model. The horizontal axis represents the aroma type predicted by the random forest model, the vertical axis represents the true aroma type of the sample, the values ​​on the diagonal represent the number of correctly classified samples, and the values ​​off-diagonally represent the number of misclassified samples. Figure 3 As can be seen, in the true 0 (sweet aroma) sample, 121 samples were correctly classified and 3 samples were misclassified as predicted 1 (honey sweet aroma); in the true 1 (honey sweet aroma) sample, 39 samples were correctly classified, 5 samples were misclassified as predicted 0 (sweet aroma), and 1 sample was misclassified as predicted 2 (mellow sweet aroma); in the true 2 (mellow sweet aroma) sample, 11 samples were correctly classified, 5 samples were misclassified as predicted 1 (honey sweet aroma), and 1 sample was misclassified as predicted 0 (sweet aroma).

[0086] The overall misclassification results are as follows: 3 true 0 (fresh and sweet aroma) samples were misclassified as predicted 1 (honey-sweet aroma); 6 true 1 (honey-sweet aroma) samples were misclassified, with 5 as predicted 0 (fresh and sweet aroma) and 1 as predicted 2 (mellow and sweet aroma); 6 true 2 (mellow and sweet aroma) samples were misclassified, with 5 as predicted 1 (honey-sweet aroma) and 1 as predicted 0 (fresh and sweet aroma). This misclassification may stem from the overlapping chemical characteristics between fresh and sweet aromas and honey-sweet aromas, as well as between honey-sweet aromas and mellow and sweet aromas, causing the random forest model to produce some bias when distinguishing these aroma types.

[0087] The final trained model for classifying and predicting the aroma type of flue-cured tobacco is capable of automatically identifying the aroma type of tobacco leaves based on their chemical composition indicators. This model, based on chemical composition data, avoids the subjective differences inherent in human evaluation and possesses good objectivity. After training, the model only requires inputting the chemical composition indicators of the tobacco leaves to quickly output aroma type prediction results, significantly improving detection efficiency and demonstrating high performance. Furthermore, the model is unaffected by subjective factors such as the experience and condition of the evaluators, resulting in good consistency in prediction results and ensuring the stability of the classification. In addition, the model can be applied to the rapid screening and quality evaluation of new batches of tobacco leaves, demonstrating high practical value and reusability.

[0088] Step S250: Obtain the content of each chemical component index in the final chemical component index dataset of the flue-cured tobacco leaves to be tested, input it into the flue-cured tobacco aroma type classification prediction model, and output the aroma type prediction result of the flue-cured tobacco leaves to be tested.

[0089] In the application phase of the flue-cured tobacco aroma classification prediction model, the chemical composition index data of the flue-cured tobacco leaves to be tested are first obtained. Specifically, the chemical composition of the flue-cured tobacco leaves to be tested is measured using the same measurement method as the training samples divided into the training set in the final chemical composition index dataset, to obtain the content data of each chemical component index in the final chemical composition index dataset. The final chemical composition index dataset is a set of feature indicators determined through multiple rounds of screening, containing chemical components that contribute significantly to aroma type identification, such as total sugar, nicotine, total nitrogen, potassium, chlorine, polyphenols, sugar alcohols, and free amino acids.

[0090] The acquired chemical component index data of the flue-cured tobacco leaves to be tested are organized and normalized according to the index order of the final chemical component index dataset to ensure that the data format is consistent with the input format during the training of the flue-cured tobacco aroma type classification prediction model. The processed chemical component index data is then input into the trained flue-cured tobacco aroma type classification prediction model, which automatically identifies the aroma type of the flue-cured tobacco leaves to be tested based on the learned mapping relationship between chemical component indexes and aroma types.

[0091] The flue-cured tobacco aroma classification prediction model outputs the predicted aroma type of the tested flue-cured tobacco leaves. The predicted result includes one of the following: sweet, honey-sweet, or mellow-sweet aroma. It can also simultaneously output the probability distribution or confidence score for each aroma type to characterize the reliability of the prediction results. This prediction result can be used to guide the quality evaluation, grading, raw material blending, and subsequent processing of flue-cured tobacco leaves.

[0092] Compared to related technologies that use linear models, such as stepwise linear regression, to predict the aroma type of flue-cured tobacco, linear models establish a linear mapping relationship between chemical component indicators and aroma type, enabling quantitative identification of aroma type. However, the relationship between flue-cured tobacco aroma type and chemical components often exhibits complex nonlinear characteristics, which linear models struggle to fully capture, resulting in limited model fitting ability and prediction accuracy. Furthermore, when there are many chemical component indicators and multicollinearity exists among them, linear models are prone to overfitting, meaning they perform well on the training set but have poor generalization ability on the test set, making it difficult to guarantee prediction performance in practical applications. This leads to the current problem of low accuracy in predicting the aroma type of flue-cured tobacco.

[0093] Steps S210 to S250 above involve: First, obtaining the evaluation results of the aroma type of the target flue-cured tobacco leaves, and assigning values ​​to the target flue-cured tobacco leaves based on the evaluation results to obtain the label values ​​of the target flue-cured tobacco leaves; second, measuring the chemical components of the target flue-cured tobacco leaves and performing feature derivation to obtain a feature derivation dataset; next, sequentially filtering the chemical component indicators in the feature derivation dataset through separation degree screening, random forest importance screening, correlation redundancy removal screening, and comprehensive cost consideration to obtain the final chemical component indicator dataset; further, using the final chemical component indicator dataset as input and the label values ​​of the target flue-cured tobacco leaves as output, training a flue-cured tobacco aroma type classification prediction model; finally, obtaining the content of each chemical component indicator in the final chemical component indicator dataset of the flue-cured tobacco leaves to be tested, inputting it into the flue-cured tobacco aroma type classification prediction model, and outputting the aroma type prediction results of the flue-cured tobacco leaves to be tested. It selects a core set of indicators from a massive amount of chemical component indicators through progressive screening, including separation degree screening, random forest importance screening, and correlation redundancy removal screening. Based on the core set of indicators, it constructs a flue-cured tobacco aroma type prediction model, which can improve the accuracy of the flue-cured tobacco aroma type prediction model.

[0094] Optionally, in one embodiment, the chemical composition of the target flue-cured tobacco leaves is measured and features are derived to obtain a feature-derived dataset, including: measuring the chemical composition of the target flue-cured tobacco leaves to obtain an original chemical composition dataset; and performing feature derivation based on the original chemical composition dataset through addition, subtraction, and division operations to obtain a feature-derived dataset.

[0095] Feature derivation datasets are obtained by performing addition, subtraction, and division operations on the original chemical composition dataset. Since the chemical characteristics of flue-cured tobacco are all continuous variables, addition, subtraction, and division operations are performed on the original chemical composition indicators (i.e., features) to ensure the interpretability of the relationships between variables. Specifically, additive feature derivation adds the measured values ​​of two or more original chemical composition indicators to obtain new composite features; subtractive feature derivation subtracts the measured values ​​of two original chemical composition indicators to obtain new composite features; and division feature derivation divides the measured values ​​of two original chemical composition indicators to obtain new composite features. Through these feature derivation operations, a large number of composite features are generated, which, together with the original chemical composition indicators, constitute the feature derivation dataset, used for subsequent feature selection and the construction of flue-cured tobacco aroma type classification prediction models.

[0096] For example, in one embodiment, the measured chemical components are used as the original dataset for feature derivation. Using 51 indicators, including common chemical components, alkaloids, polyphenols, sugar alcohols, and amino acids, as the original features, 1326 additive composite features, 1275 subtractive composite features, and 1275 divisive composite features are generated through feature derivation. These, together with the original 51 indicators, constitute a feature-derived dataset containing 3927 features. Table 3 is an example of a feature-derived dataset from one embodiment of this application.

[0097] Table 3

[0098]

[0099] By using addition, subtraction, and division operations to derive features, we can fully explore the interaction relationships between chemical composition indicators, enrich feature expression, enhance the representational ability of the dataset, and provide a more comprehensive information foundation for subsequent screening and modeling.

[0100] Furthermore, in one embodiment, the chemical composition indicators in the feature-derived dataset are sequentially subjected to separation degree screening, random forest importance screening, correlation redundancy removal screening, and comprehensive cost consideration to obtain the final chemical composition indicator dataset. This includes: performing separation degree screening on the chemical composition indicators in the feature-derived dataset to obtain a first feature subset; using the first feature subset as input, constructing a random forest classification model to perform random forest importance screening to obtain a second feature subset; performing correlation redundancy removal screening on the chemical composition indicators in the second feature subset to obtain a third feature subset; and performing comprehensive cost consideration on the chemical composition indicators in the third feature subset to determine the final chemical composition indicator dataset.

[0101] The chemical component indicators in the feature-derived dataset are screened for separation to obtain the first feature subset. Separation screening refers to screening based on the ability of each chemical component indicator to distinguish between samples of different aroma types. By calculating the mean difference of each indicator among different categories and using methods such as analysis of variance, the ability of each chemical component indicator to distinguish aroma types is evaluated. Chemical component indicators with weak distinguishing ability are eliminated, and those that can effectively distinguish different aroma types are retained. These chemical component indicators are used as features to form the first feature subset.

[0102] Using the first feature subset as input, a random forest classification model is constructed to perform random forest importance filtering, resulting in the second feature subset. Random forest importance filtering evaluates the importance of features in the first feature subset using the random forest algorithm. This is achieved by calculating the average contribution of each feature to classification accuracy within the random forest, or the average contribution of a feature to the reduction of impurity during decision tree node splits, thus obtaining an importance score for each feature. Features are then sorted from highest to lowest importance, retaining those with higher importance and removing redundant variables that have little impact on aroma type discrimination, thus forming the second feature subset.

[0103] The chemical component indicators in the second feature subset are subjected to correlation-based redundancy removal screening to obtain the third feature subset. Correlation-based redundancy removal screening involves calculating the correlation coefficients between features. For highly correlated feature pairs with absolute correlation coefficients exceeding a preset threshold, features with higher correlation to aroma type or higher importance scores are retained, while highly correlated redundant features are removed. This avoids the impact of multicollinearity on the stability of subsequent data analysis, thus forming the third feature subset.

[0104] A comprehensive cost consideration was conducted on the chemical component indicators in the third feature subset to determine the final chemical component indicator dataset. The comprehensive cost consideration refers to taking into account factors such as the detection cost, detection cycle, operational complexity, and instrument requirements of each chemical component indicator. Priority was given to features that are easy to obtain, have low detection costs, short detection cycles, and are easy to operate. While ensuring the model's prediction accuracy, features with reasonable detection costs and high practical application feasibility were selected to form the final chemical component indicator dataset.

[0105] Through the above four rounds of screening, features with good discriminative ability, high importance, low correlation and reasonable detection cost are selected from thousands of features in the feature-derived dataset to form the final chemical composition index dataset, which is used to build the subsequent flue-cured tobacco aroma classification and prediction model.

[0106] Through multiple rounds of progressive screening, including separation degree screening, random forest importance screening, correlation redundancy removal screening, and comprehensive cost consideration, it is possible to effectively reduce data dimensionality, eliminate redundant information, and improve feature interpretability while ensuring feature discrimination capability, and also take into account the economy and operability of practical applications.

[0107] In one embodiment, the separation degree screening of chemical component indicators in the feature-derived dataset to obtain a first feature subset includes: calculating the separation degree value of each chemical component indicator in the feature-derived dataset relative to the aroma type of different target flue-cured tobacco leaves, retaining chemical component indicators with a separation degree value greater than a preset first threshold, and obtaining the first feature subset.

[0108] Specifically, the separation score is used to measure the ability of each chemical component indicator to distinguish between samples of different aroma types. Using the aroma type label of the target flue-cured tobacco leaves as the classification basis, the aroma types include three categories: light sweet, honey sweet, and mellow sweet. For each feature in the feature-derived dataset, the distribution of each feature in the three aroma type samples is statistically analyzed. The mean difference, variance difference, or statistical test methods are used to evaluate the distinguishing ability of the chemical component indicator among different aroma types, thus obtaining the separation score of the chemical component indicator. Chemical component indicators with a separation score greater than a preset first threshold are retained, forming the first feature subset. The preset first threshold can be set according to the actual data distribution; it can be set as the F-statistic threshold or p-value threshold in the analysis of variance, or as a percentage threshold based on the distinguishing ability ranking. Through separation score screening, chemical component indicators with significant overlap in distribution and weak distinguishing ability among samples of different aroma types are removed, retaining features that can effectively distinguish different aroma types, forming the first feature subset.

[0109] In the separation screening, the separability index f is used to measure the ability of each chemical component index to distinguish between samples of different aroma types. The f value is calculated using the following formula: f = 1.18 × difference between adjacent peak center points / sum of half-peak widths of adjacent peaks. Here, the difference between adjacent peak center points represents the degree of difference in the mean values ​​of the chemical component index among samples of different aroma types, and the sum of half-peak widths of adjacent peaks represents the comprehensive dispersion of the distribution of the chemical component index within samples of each aroma type. The larger the f value, the stronger the ability of the chemical component index to distinguish between different aroma types, meaning the index is more effective in identifying different aroma types.

[0110] For example, separation analysis was performed on the four datasets derived from the features (original feature dataset, additive composite feature dataset, subtractive composite feature dataset, and division composite feature dataset). Table 4 shows the separation analysis results of the datasets derived from the features in one embodiment of this application. As can be seen from Table 4, in all datasets, the separation value f... 02 The number of features >0.8 is significantly greater than f. 01 >0.8 and f 12 >0.8 number of features, where f 02 Indicates the degree of separation between the light and sweet aroma type and the rich and sweet aroma type, f 01 Indicates the separation between the light and sweet fragrance type and the honey-sweet fragrance type, f 12 This indicates the degree of separation between honey-sweet and ethanol-sweet aroma types. This result shows that most characteristics are significantly more effective in distinguishing between light sweet and ethanol-sweet aroma types than between light sweet and honey-sweet, or honey-sweet and ethanol-sweet aroma types; that is, the differences in chemical composition between light sweet and ethanol-sweet aroma types are more significant.

[0111] Table 4

[0112]

[0113] In the original features, chemical components such as rutin, chlorogenic acid, neonicotinoids, and serine showed good separation ability, with resolution values ​​(f) all exceeding 0.8, indicating that these chemical components have good discriminative ability to distinguish different aroma types. However, the number of features with resolution values ​​greater than 0.8 in the three composite feature datasets (addition, subtraction, and division) is still relatively large, indicating that while the composite features generated through feature derivation further enrich the feature expression, they also introduce a large number of features. To reduce data dimensionality and improve the efficiency of subsequent flue-cured tobacco aroma type classification and prediction model construction, further feature screening is needed. Therefore, features with resolution values ​​(f) greater than 0.8 are retained for the next round of random forest importance analysis.

[0114] By using separation degree screening, chemical component indicators that have serious overlap in distribution and weak distinguishing ability among samples of different aroma types can be effectively removed, while chemical component indicators that make significant contributions to aroma type identification are retained. This reduces data dimensionality and redundant information, serving as data preparation for subsequent chemical component indicator screening and providing high-quality input feature subsets.

[0115] In one embodiment, a random forest classification model is constructed using a first feature subset as input to perform random forest importance screening to obtain a second feature subset. This includes: constructing a random forest classification model using the first feature subset as input; calculating the contribution of each chemical component index in the second feature subset to the classification of flue-cured tobacco aroma type, obtaining the feature importance of the chemical component index, and ranking the chemical component indexes according to their feature importance; retaining chemical component indexes whose feature importance ranking is greater than a preset second threshold to obtain the second feature subset.

[0116] A random forest classification model is constructed using the first feature subset as input. Specifically, the random forest classification model is an ensemble learning algorithm that constructs multiple decision trees and combines their voting results for classification prediction. The chemical composition indicators from the first feature subset are used as input features, and the aroma type label of the target flue-cured tobacco leaves is used as the output target. Hyperparameters such as the number of decision trees and maximum depth are set to train the random forest classification model. This model can effectively handle high-dimensional feature data and has good anti-overfitting ability.

[0117] The contribution of each chemical component index (i.e., feature) in the first feature subset to the classification of flue-cured tobacco aroma is calculated to obtain the feature importance of the chemical component index, and the chemical component indexes are ranked according to their feature importance. The calculation of feature importance can employ either a method based on impurity reduction or a method based on accuracy reduction. Specifically, the impurity reduction method calculates the average reduction in Gini impurity or information gain for each chemical component index during the splitting of all decision tree nodes in the random forest, using this as the importance score for that chemical component index. The accuracy reduction method randomly permutes the values ​​of the chemical component indexes and observes the degree of decrease in the prediction accuracy of the random forest classification model; a larger decrease indicates a more important feature. Using these methods, the feature importance score for each chemical component index is obtained, and the indexes are ranked from highest to lowest score.

[0118] The chemical composition indicators whose feature importance ranking is greater than a preset second threshold are retained to obtain the second feature subset. The preset second threshold can be set according to actual needs, for example, to retain the top 30% of features, or to retain features whose feature importance score is greater than a certain value.

[0119] For example, Figure 4 This is a feature importance score map of chemical component indicators retained after separation filtering from the original feature dataset of one embodiment of this application. Figure 4 As can be seen, the original feature dataset only retains four chemical component indicators: rutin, chlorogenic acid, neonicotinoids, and serine. Among them, rutin has the highest feature importance score, followed by chlorogenic acid, while neonicotinoids and serine have relatively lower feature importance scores. Figure 5 This is a score chart showing the importance scores of the top 10 chemical component indicators in terms of feature importance in an additive composite feature dataset according to one embodiment of this application. Figure 5 The results show that the top 10 most important features are, in descending order: total sugar + proline, total sugar + rutin, rutin + linoleic acid, rutin + mannitol, total sugar + total nitrogen, rutin + arginine, total sugar + protein, total sugar + neonicotinoids, rutin + serine, and rutin + glycerol. Additive features involving rutin account for more than half of these, further confirming the core role of rutin in aroma type identification. Simultaneously, additive features involving total sugar and multiple components also demonstrate high importance, indicating that sugars and their interactions with other components have a significant impact on aroma type identification. Figure 6 This is a score chart showing the importance scores of the top 10 chemical component indicators in terms of feature importance in a subtractive composite feature dataset according to one embodiment of this application. Figure 6The results show that the top 10 key characteristics include: total alkaloids-rutin, total sugars-starch, rutin-malonic acid, rutin-oxalic acid, valine-lysine, rutin-nicotine, total alkaloids-neonicotinine, phenylalanine-glutamic acid, and rutin-phenylalanine. Among these, subtractive characteristics involving rutin constitute a large proportion, and the difference between total alkaloids and rutin ranks first, indicating that the balance between rutin and alkaloids is crucial for aroma type identification. Furthermore, the differences between total sugars and starch, and between rutin and organic acids, also demonstrate high importance. Figure 7 This is a score chart showing the importance scores of the top 10 chemical component indicators in terms of feature importance in a division composite feature dataset according to one embodiment of this application. Figure 7 The results show that the top 10 key features are: rutin / xylitol, rutin / isoleucine, rutin / oxalic acid, rutin / valine, rutin / rhamnose, total alkaloids / neotoxin, proline / valine, rutin / mannitol, nicotine / neotoxin, and total alkaloids / rutin. Among these features, those involving rutin remain dominant, with ratios of rutin to sugar alcohols, amino acids, and organic acids all showing high importance. Simultaneously, the ratios of total alkaloids to neotoxin and nicotine to neotoxin also demonstrate high discriminative power, indicating that the proportional relationships between alkaloids also significantly contribute to aroma type identification.

[0120] By using random forest importance screening, we can quantitatively evaluate the actual contribution of chemical component indicators to the classification of flue-cured tobacco aroma types. We can remove chemical component indicators that contribute little to the identification of aroma types and have low importance, while retaining chemical component indicators that are strongly correlated with aroma types and have high discrimination ability. This will further reduce data dimensionality and redundant information, while enhancing the correlation between chemical component indicators and aroma types.

[0121] In one embodiment, the chemical component indicators in the second feature subset are subjected to correlation redundancy removal screening to obtain a third feature subset, including: calculating the correlation coefficient between all chemical component indicator pairs in the second feature subset; for chemical component indicator pairs whose absolute value of the correlation coefficient is greater than a preset third threshold, the chemical component indicators with high feature importance are retained to obtain the third feature subset.

[0122] Calculate the correlation coefficients between all chemical component indicator pairs in the second feature subset. Specifically, the Pearson correlation coefficient or Spearman rank correlation coefficient is used to measure the degree of linear correlation between each pair of chemical component indicators. For continuous variables that follow a normal distribution, the Pearson correlation coefficient can be used; for indicators that are not normally distributed or have a non-linear relationship, the Spearman rank correlation coefficient can be used. A correlation coefficient matrix is ​​obtained for all chemical component indicator pairs, where each element represents the strength of the correlation between the corresponding two chemical component indicators. The correlation coefficient ranges from [-1, 1], with the absolute value closer to 1 indicating a stronger correlation.

[0123] For chemical component index pairs whose absolute correlation coefficients exceed a preset third threshold, the chemical component indexes with higher feature importance are retained to obtain the third feature subset. The preset third threshold can be set according to actual needs, for example, to 0.9 or 0.95. When the absolute value of the correlation coefficient between two chemical component indexes exceeds this threshold, it indicates a high correlation between the two indexes, meaning they carry similar chemical information and are redundant.

[0124] During the screening process, for each pair of indicators whose absolute correlation coefficient exceeds a preset third threshold, the feature importance scores of the two indicators in the random forest importance screening are compared. Indicators with higher scores are retained, while those with lower scores are removed. Through this process, highly correlated redundant features are eliminated one by one, avoiding the impact of multicollinearity on the stability of subsequent data analysis and the prediction accuracy of the flue-cured tobacco aroma classification prediction model. This process is repeated until the absolute values ​​of the correlation coefficients between all retained chemical component indicator pairs do not exceed the preset third threshold, forming the third feature subset.

[0125] Figure 8 This is a heat map showing the correlation coefficients of chemical composition index pairs in the original feature dataset in one embodiment of this application. The feature importance is calculated based on the random forest algorithm. Figure 8 The results show that rutin has a correlation coefficient of 0.33 with chlorogenic acid, indicating a moderate positive correlation; a correlation coefficient of 0.45 with neonicotinoids, indicating a slightly higher-than-moderate positive correlation; and a correlation coefficient of 0.43 with serine, also indicating a moderate positive correlation. This suggests that rutin exhibits a certain degree of synergistic variation with chlorogenic acid, neonicotinoids, and serine, but the correlations are not highly correlated and can still be retained as independent characteristics.

[0126] Figure 9 This is a heatmap of the correlation coefficients of chemical component index pairs in an additive composite feature dataset from one embodiment of this application. Figure 9As can be seen, there are high correlations among some chemical component pairs in the additive composite feature dataset. For example, the correlation coefficient between rutin + serine and rutin + glycerol is 1, indicating a completely linear correlation between these two additive features, with highly redundant information. Furthermore, there are also high correlation coefficients (greater than 0.9) between rutin + mannitol and features such as total sugar + rutin and rutin + linoleic acid, indicating significant overlap in the chemical information carried by these features.

[0127] Figure 10 This is the correlation coefficient of chemical component index pairs in the subtractive composite feature dataset in one embodiment of this application. Figure 10 As can be seen, the correlation coefficients between the chemical component index pairs in the subtractive composite feature dataset exhibit both positive and negative differences. Some index pairs show a high degree of positive correlation; for example, the correlation coefficients between rutin-malonic acid and rutin-equisetine, rutin-oxalic acid, rutin-nicotinic acid, and rutin-phenylalanine are all close to 1, indicating a completely linear positive correlation between these subtractive features, with highly redundant information. Simultaneously, total alkaloids-rutin also show a high positive correlation (correlation coefficient greater than 0.8) with features such as rutin-malonic acid, rutin-equisetine, rutin-oxalic acid, and rutin-nicotinic acid. Furthermore, some index pairs show a negative correlation. For example, the correlation coefficient between total sugar-starch and valine-lysine is -0.62, showing a moderate negative correlation; total sugar-starch also shows a negative correlation (correlation coefficient around -0.5) with features such as rutin-malonic acid, rutin-equisetine, and rutin-oxalic acid, indicating an inverse relationship between these features. Whether positive or negative, all are considered highly correlated. Indicators with higher feature importance are retained, while redundant features that are highly correlated with them are removed.

[0128] Figure 11 This is a heatmap of the correlation coefficients of chemical component index pairs in a division composite feature dataset from one embodiment of this application. Figure 11As can be seen, some chemical component pairs in the division composite feature dataset exhibit high correlations. For example, the correlation coefficient between rutin / mannitol and rutin / rhamnose is 0.91, showing a high positive correlation; the correlation coefficient between rutin / xylitol and rutin / isoleucine is 0.86, also showing a strong positive correlation; and the correlation coefficient between total alkaloids / neotonicine and nicotine / neotonicine is 0.91, indicating a high linear positive correlation between these two ratio features. Meanwhile, some indicator pairs exhibit negative correlations. For example, the correlation coefficient between rutin / xylitol and proline / valine is -0.66, showing a moderate negative correlation; the correlation coefficient between rutin / oxalic acid and proline / valine is -0.55; and the correlation coefficient between rutin / isoleucine and nicotine / neotonicine is -0.63, indicating an inverse relationship between these division features. Similarly, regardless of whether the correlation is positive or negative, it is considered a high correlation. Indicators with higher feature importance are retained, while redundant features with high correlation are removed.

[0129] By using correlation-based redundancy removal screening, the linear correlation between chemical component indicators can be quantitatively assessed, and highly correlated (correlation coefficient absolute value exceeds a preset threshold) redundant features can be eliminated, while indicators with higher feature importance can be retained. This effectively eliminates the impact of multicollinearity on the stability of subsequent data analysis, reduces the dimensionality of chemical component indicators, and enhances the independence between chemical component indicators.

[0130] In addition, in one embodiment, the chemical component indicators in the third feature subset are comprehensively considered in terms of cost to determine the final chemical component indicator dataset, including: determining the final chemical component indicator dataset based on the detection cost and feature importance of each chemical component indicator in the third feature subset.

[0131] Specifically, comprehensive cost consideration refers to the process in practical applications of comprehensively considering factors such as detection cost, detection cycle, operational complexity, instrument and equipment requirements, and sample pretreatment difficulty for each chemical component indicator. While ensuring the prediction accuracy of the flue-cured tobacco aroma classification prediction model, priority is given to chemical component indicators that are easy to obtain, have low detection cost, short detection cycle, and simple operation. First, cost-related information for each chemical component indicator in the third feature subset is collected, including reagent costs, instrument and equipment costs, single detection time, sample pretreatment steps, and operator technical requirements. This information is then quantitatively evaluated to obtain a comprehensive cost score for each chemical component indicator. A lower cost score indicates lower detection cost, simpler operation, and higher detection efficiency for that chemical component indicator. Second, a comprehensive trade-off is made by combining the feature importance scores of each chemical component indicator in the random forest importance screening. For chemical component indicators with similar feature importance, those with lower comprehensive cost scores are prioritized; for key chemical component indicators with significantly higher feature importance, even if the detection cost is relatively high, they are still retained to ensure that the prediction accuracy of the flue-cured tobacco aroma classification prediction model is not affected.

[0132] Based on the above comprehensive cost considerations, chemical component indicators with both high feature importance and good detection economy and operability were selected from the third feature subset to form the final chemical component indicator dataset. This final chemical component indicator dataset not only retains key chemical component indicators closely related to the identification of flue-cured tobacco aroma types, but also takes into account detection costs and operational feasibility in practical applications, thereby improving the application value of the flue-cured tobacco aroma type classification and prediction model in actual production scenarios.

[0133] In one embodiment, the method for predicting the aroma type of flue-cured tobacco further includes interpretability analysis of the trained flue-cured tobacco aroma type classification prediction model. By employing Shapley Additive Explanations (SHAP) to perform interpretability analysis on the flue-cured tobacco aroma type classification prediction model, the average absolute SHAP value of each feature is quantified, and a bee colony diagram is used to visually demonstrate the positive and negative contribution direction and degree of influence of each feature on the aroma type of flue-cured tobacco. This reveals the intrinsic correlation between key chemical component indicators and the aroma type of flue-cured tobacco, enhancing the transparency and credibility of the model's decision-making process.

[0134] Specifically, in addition to evaluation indicators, transparency and interpretability are also important supplements to assessing the reliability of flue-cured tobacco aroma type classification prediction models. To enhance the interpretability of the model, the SHAP method is used for interpretive analysis of the flue-cured tobacco aroma type classification prediction model. SHAP is a game theory-based interpretive model framework that quantifies the importance of each feature in the decision-making process of the flue-cured tobacco aroma type classification prediction model by calculating the contribution value of each feature to the prediction result (i.e., the SHAP value). A positive SHAP value indicates that the feature has a positive driving effect on the prediction result, while a negative value indicates a negative inhibiting effect; the absolute value reflects the degree of influence. Figure 12 This is a bar chart showing the mean absolute SHAP values ​​of features in a flue-cured tobacco aroma type classification prediction model according to an embodiment of this application. Feature importance is quantified by the mean absolute SHAP value, and seven influential features were identified. The results show that in the random forest-based classification of flue-cured tobacco aroma types, the mean absolute SHAP value of a feature is positively correlated with its classification contribution. Among all features, rutin / xylitol has the highest mean absolute SHAP value and is the core feature for aroma type classification in the flue-cured tobacco aroma type classification prediction model. In addition, the mean absolute SHAP values ​​of six other features, including total sugar + proline, rutin, and chlorogenic acid, are also relatively high and are key auxiliary components for aroma type classification.

[0135] Figure 13 This is an embodiment of the application that uses the sweet aroma characteristics of SHAP to influence bee colony diagrams. Figure 13 In the diagram, the horizontal axis represents the SHAP value, reflecting the magnitude and direction of the feature's contribution to classification. A positive SHAP value indicates that the feature promotes the classification of the sample as sweet and refreshing, while a negative value indicates that it inhibits the classification. The vertical axis represents the feature variables, sorted from top to bottom according to feature importance. The color of the dots changes from blue to pink, representing feature value levels from low to high; blue indicates lower feature values, and pink indicates higher feature values. Figure 13 It can be seen that rutin / xylitol is the most important feature influencing the classification of sweet and refreshing aroma. Its SHAP value distribution is wide, and high eigenvalues ​​(pink dots) correspond to most positive SHAP values, indicating that when the rutin / xylitol content is high, this feature significantly promotes the model's classification as sweet and refreshing aroma. Total sugar + proline also shows a strong positive contribution, with the increase in eigenvalue correlated with positive SHAP values, indicating that the higher the content of this feature, the more the flue-cured tobacco aroma classification prediction model tends to classify the sample as sweet and refreshing aroma. The SHAP values ​​of rutin and chlorogenic acid are relatively concentrated in this type, and their contributions are relatively stable.

[0136] Figure 14 This is an embodiment of the present application that uses SHAP to influence bee colony diagrams based on the characteristics of honey sweetness. Figure 14In the diagram, the horizontal axis represents the SHAP value, reflecting the magnitude and direction of the feature's contribution to classification. A positive SHAP value indicates that the feature promotes the classification of the sample as honey-sweet, while a negative value indicates that it inhibits the classification. The vertical axis represents the feature variables, sorted from top to bottom according to feature importance. The color of the dots changes from blue to pink, representing feature value levels from low to high; blue indicates lower feature values, and pink indicates higher feature values. Figure 14 It can be seen that chlorogenic acid shows a strong positive contribution to the honey-sweet aroma classification. High content (pink dots) corresponds to most positive SHAP values, indicating that the higher the chlorogenic acid content, the more the flue-cured tobacco aroma classification prediction model tends to classify the sample as honey-sweet. Rutin / xylitol mainly shows a negative SHAP contribution in this type, and high eigenvalues ​​correspond to negative SHAP values, indicating that a high content of this feature will inhibit the model from classifying the sample as honey-sweet. The SHAP value distribution of total sugar + proline is relatively dispersed, with both positive and negative contributions, indicating that the influence of this feature on the honey-sweet aroma classification has a certain degree of uncertainty.

[0137] Figure 15 This is an embodiment of the application that uses SHAP to influence bee colony diagrams based on the characteristics of sweet and mellow aroma. Figure 15 In the diagram, the horizontal axis represents the SHAP value, reflecting the magnitude and direction of the feature's contribution to classification. A positive SHAP value indicates that the feature promotes the classification of the sample as "alcoholic and sweet," while a negative value indicates that it inhibits classification as "alcoholic and sweet." The vertical axis represents the feature variables, sorted from top to bottom according to feature importance. The color of the dots changes from blue to pink, representing feature value levels from low to high; blue indicates lower feature values, and pink indicates higher feature values. Figure 15 It can be seen that chlorogenic acid mainly contributes negative SHAP values ​​in this type, with high content (pink dots) corresponding to most negative SHAP values, indicating that high chlorogenic acid content inhibits the model from classifying samples as alcoholic sweet aroma type. Rutin / xylitol also mainly contributes negative SHAP values ​​in this type, with high eigenvalues ​​also corresponding to negative SHAP values, indicating that high content of this eigenvalue is unfavorable for the classification of alcoholic sweet aroma type. The SHAP value distribution of total sugar + proline is relatively concentrated, and the degree of contribution is relatively stable, with little impact on the classification of this type.

[0138] By calculating the mean absolute value (SHAP) of each feature, the importance of the features was quantified, and seven key features were identified, including rutin / xylitol, total sugar + proline, rutin, and chlorogenic acid. Further analysis using a bee colony diagram revealed the positive and negative contributions of each feature to the three aroma types: high rutin / xylitol content significantly promoted the sweet aroma type classification, while high chlorogenic acid content positively promoted the honey-sweet aroma type classification but inhibited the mellow-sweet aroma type classification. This interpretability analysis based on SHAP values ​​visually demonstrates the contribution and direction of influence of each chemical component index on aroma type classification, reveals the intrinsic correlation between key features and aroma types, enhances the transparency and credibility of the model's decision-making process, provides a reliable decision-making basis for the application of flue-cured tobacco aroma type classification prediction models in actual production, and offers scientific guidance for flue-cured tobacco aroma quality control and tobacco leaf grading.

[0139] In one embodiment, the method for predicting the aroma type of flue-cured tobacco further includes feature dependency analysis on the trained flue-cured tobacco aroma type classification prediction model. To explore the impact of key features on the model's classification performance and their numerical distribution across different aroma types, two features that significantly contribute to both global model interpretation and aroma type classification—rutin / xylitol and total sugar + proline—are selected. Feature dependency analysis is performed by generating scatter plots for each aroma type. Given that the tree-based model is insensitive to feature dimensions, all indicators are presented in their original form to more intuitively represent the feature intervals.

[0140] Figure 16 This is a scatter plot showing the characteristic dependence of rutin / xylitol in a sweet and refreshing flavor profile in one embodiment of this application. Figure 16 In the diagram, the X-axis represents the eigenvalue range of rutin / xylitol, and the Y-axis represents the corresponding SHAP value. A positive SHAP value indicates that the feature promotes the classification of the sample as a sweet and refreshing aroma, while a negative value indicates that it inhibits the classification as a sweet and refreshing aroma. Figure 16 It can be seen that rutin / xylitol has a significant positive contribution to the sweet aroma type, with a maximum SHAP value of 0.3. When the feature value exceeds 320, the SHAP value is mainly positive, indicating that when the content of this feature is high, it drives the flue-cured tobacco aroma type classification prediction model to classify the sample as sweet aroma type.

[0141] Figure 17 This is a scatter plot showing the characteristic dependence of rutin / xylitol in a honey-sweet flavor profile in one embodiment of this application. Figure 17 In the diagram, the X-axis represents the eigenvalue range of rutin / xylitol, and the Y-axis represents the corresponding SHAP value. A positive SHAP value indicates that the feature promotes the classification of the sample as honey-sweet, while a negative value indicates that it inhibits the classification as honey-sweet. Figure 17It can be seen that rutin / xylitol mainly has a negative contribution effect in the honey-sweet aroma type. When the eigenvalue is in the range of 250 to 320, the SHAP value distribution is relatively concentrated, indicating that this feature has a certain influence on the classification of honey-sweet aroma type in this range.

[0142] Figure 18 This is a scatter plot showing the characteristic dependence of rutin / xylitol in the sweet aroma profile of one embodiment of this application. Figure 18 In the diagram, the X-axis represents the eigenvalue range of rutin / xylitol, and the Y-axis represents the corresponding SHAP value. A positive SHAP value indicates that the feature promotes the classification of the sample as an alcoholic sweet aroma type, while a negative value indicates that it inhibits the classification. From Figure 18 It can be seen that rutin / xylitol mainly has a negative contribution to the sweet aroma type. When the eigenvalue is below 250, the SHAP value is mainly negative, indicating that when the content of this eigenvalue is low, the flue-cured tobacco aroma type classification prediction model tends to classify the sample as sweet aroma type.

[0143] Figure 19 This is a scatter plot showing the characteristic dependence of total sugar + proline in a sweet and refreshing flavor profile in one embodiment of this application. Figure 19 In the diagram, the X-axis represents the feature range of total sugar + proline, and the Y-axis represents the corresponding SHAP value. A positive SHAP value indicates that the feature promotes the classification of the sample as a sweet and refreshing flavor, while a negative value indicates that it inhibits the classification as a sweet and refreshing flavor. Figure 19 It can be seen that when the feature value exceeds 47, the corresponding SHAP value is mainly positive, indicating that when the feature content is high, it drives the flue-cured tobacco aroma classification prediction model to classify the sample as sweet aroma.

[0144] Figure 20 This is a scatter plot showing the characteristic dependence of total sugar + proline in the honey-sweet flavor in one embodiment of this application. Figure 20 In the diagram, the X-axis represents the feature range of total sugar + proline, and the Y-axis represents the corresponding SHAP value. A positive SHAP value indicates that the feature promotes the classification of the sample as honey-sweet, while a negative value indicates that it inhibits the classification as honey-sweet. Figure 20 It can be seen that when the eigenvalue is below 42, a positive SHAP value contribution occurs, indicating that a lower content of this eigenvalue is beneficial for the classification of honey-sweet aroma.

[0145] Figure 21 This is a scatter plot showing the characteristic dependence of total sugar + proline in the alcoholic sweet flavor profile in one embodiment of this application. Figure 21 In the diagram, the X-axis represents the feature range of total sugar + proline, and the Y-axis represents the corresponding SHAP value. A positive SHAP value indicates that the feature promotes the classification of the sample as a sweet and mellow flavor, while a negative value indicates that it inhibits the classification. From Figure 21It can be seen that the SHAP value of this feature has a relatively weak predictive correlation with the aroma of alcoholic sweetness. When the feature value is in the range of 45 to 47, the distribution of SHAP value tends to be alcoholic sweetness.

[0146] The feature dependency analysis described above can intuitively reveal the influence of key features on aroma type classification in different numerical ranges, clarify the feature value range corresponding to each aroma type, and provide quantitative basis and scientific guidance for the control of flue-cured tobacco aroma quality and tobacco leaf grading.

[0147] Figure 22 This is a structural block diagram of a flue-cured tobacco aroma type prediction device 50 according to an embodiment of this application, as shown below. Figure 22 As shown, the flue-cured tobacco aroma type prediction device includes: a dataset construction module 52, a feature selection module 54, and a model construction and application module 56; wherein: the dataset construction module 52 is used to obtain the evaluation results of the aroma type of the target flue-cured tobacco leaves, and assign values ​​to the target flue-cured tobacco leaves according to the evaluation results to obtain the label values ​​of the target flue-cured tobacco leaves; the chemical components of the target flue-cured tobacco leaves are measured and features are derived to obtain the feature-derived dataset; the feature selection module 54 is used to select the chemical component indicators in the feature-derived dataset through separation degree selection, random forest importance selection, correlation redundancy removal selection, and comprehensive cost consideration to obtain the final chemical component indicator dataset; the model construction and application module 56 is used to train the flue-cured tobacco aroma type classification prediction model by taking the final chemical component indicator dataset as input and the label values ​​of the target flue-cured tobacco leaves as output; the content of each chemical component indicator in the final chemical component indicator dataset of the flue-cured tobacco leaves to be tested is obtained, input into the flue-cured tobacco aroma type classification prediction model, and the aroma type prediction results of the flue-cured tobacco leaf samples to be tested are output.

[0148] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0149] This embodiment also provides an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.

[0150] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.

[0151] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:

[0152] S1, obtain the evaluation results of the aroma type of the target flue-cured tobacco leaves, and assign values ​​to the target flue-cured tobacco leaves according to the evaluation results to obtain the label value of the target flue-cured tobacco leaves;

[0153] S2, the chemical composition of the target flue-cured tobacco leaves is determined and features are derived to obtain a feature-derived dataset;

[0154] S3, for the chemical composition indicators in the feature-derived dataset, sequentially through separation degree screening, random forest importance screening, correlation redundancy removal screening, and comprehensive cost consideration, to obtain the final chemical composition indicator dataset;

[0155] S4. The final chemical composition index dataset is used as input, and the label value of the target flue-cured tobacco leaf is used as output to train a flue-cured tobacco aroma classification prediction model.

[0156] S5: Obtain the content of each chemical component index in the final chemical component index dataset of the flue-cured tobacco leaves to be tested, input it into the flue-cured tobacco aroma type classification prediction model, and output the aroma type prediction result of the flue-cured tobacco leaves to be tested.

[0157] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.

[0158] Furthermore, in conjunction with the flue-cured tobacco aroma type prediction method provided in the above embodiments, this embodiment can also provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the flue-cured tobacco aroma type prediction methods in the above embodiments.

[0159] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0160] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0161] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.

[0162] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0163] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.

Claims

1. A method for predicting the aroma type of flue-cured tobacco, characterized in that, include: The aroma type evaluation result of the target flue-cured tobacco leaf is obtained, and the target flue-cured tobacco leaf is assigned a value according to the evaluation result to obtain the label value of the target flue-cured tobacco leaf; The chemical composition of the target flue-cured tobacco leaves was determined and its features were derived to obtain a feature-derived dataset; The chemical composition indicators in the feature-derived dataset are sequentially filtered by separation degree, random forest importance, correlation redundancy removal, and comprehensive cost consideration to obtain the final chemical composition indicator dataset. Using the final chemical composition index dataset as input and the label value of the target flue-cured tobacco leaf as output, a flue-cured tobacco aroma classification and prediction model is trained. Obtain the content of each chemical component index in the final chemical component index dataset of the flue-cured tobacco leaf to be tested, input it into the flue-cured tobacco aroma type classification prediction model, and output the aroma type prediction result of the flue-cured tobacco leaf to be tested.

2. The method for predicting the aroma type of flue-cured tobacco according to claim 1, characterized in that, The determination of the chemical composition of the target flue-cured tobacco leaves and the subsequent feature derivation to obtain a feature-derived dataset include: The chemical composition of the target flue-cured tobacco leaves was determined to obtain the original chemical composition dataset; Based on the original chemical composition dataset, feature derivation is performed through addition, subtraction, and division operations to obtain a feature-derived dataset.

3. The method for predicting the aroma type of flue-cured tobacco according to claim 1, characterized in that, The chemical composition indicators in the feature-derived dataset are sequentially filtered through separation degree, random forest importance, relevance redundancy removal, and comprehensive cost considerations to obtain the final chemical composition indicator dataset, including: The separation degree is used to filter the chemical composition indicators in the feature-derived dataset to obtain the first feature subset; Using the first feature subset as input, a random forest classification model is constructed to perform the random forest importance screening, thereby obtaining the second feature subset; The chemical composition indicators in the second feature subset are subjected to the aforementioned correlation redundancy removal screening to obtain the third feature subset; The chemical composition indicators in the third feature subset are subjected to the comprehensive cost consideration to determine the final chemical composition indicator dataset.

4. The method for predicting the aroma type of flue-cured tobacco according to claim 3, characterized in that, The separation degree filtering of the chemical composition indicators in the feature-derived dataset to obtain the first feature subset includes: Calculate the separation degree value of each chemical component index in the feature-derived dataset relative to the aroma type of different target flue-cured tobacco leaves, and retain the chemical component indexes with separation degree values ​​greater than a preset first threshold to obtain the first feature subset.

5. The method for predicting the aroma type of flue-cured tobacco according to claim 3, characterized in that, The step of constructing a random forest classification model using the first feature subset as input to perform random forest importance filtering, and obtaining a second feature subset, includes: Using the first feature subset as input, construct a random forest classification model; Calculate the contribution of each chemical component index in the second feature subset to the classification of flue-cured tobacco aroma type, obtain the feature importance of the chemical component index, and rank the chemical component index according to the feature importance; Chemical composition indicators whose feature importance ranking is greater than a preset second threshold are retained to obtain the second feature subset.

6. The method for predicting the aroma type of flue-cured tobacco according to claim 5, characterized in that, The correlation redundancy removal screening of the chemical component indicators in the second feature subset yields a third feature subset, including: Calculate the correlation coefficients among all chemical component index pairs in the second feature subset; For chemical component index pairs whose absolute values ​​of the correlation coefficients are greater than a preset third threshold, the chemical component indexes with high feature importance are retained to obtain a third feature subset.

7. The method for predicting the aroma type of flue-cured tobacco according to claim 6, characterized in that, The process of performing a comprehensive cost assessment on the chemical component indicators in the third feature subset to determine the final chemical component indicator dataset includes: Based on the detection cost of each chemical component index in the third feature subset and the importance of the features, the final chemical component index dataset is determined.

8. A device for predicting the aroma type of flue-cured tobacco, characterized in that, include: The module comprises a dataset construction module, a feature selection module, and a model construction and application module; among which: The dataset construction module is used to obtain the evaluation results of the aroma type of the target flue-cured tobacco leaves, and assign values ​​to the target flue-cured tobacco leaves according to the evaluation results to obtain the label values ​​of the target flue-cured tobacco leaves; the chemical components of the target flue-cured tobacco leaves are measured and features are derived to obtain the feature-derived dataset; The feature filtering module is used to sequentially filter the chemical composition indicators in the feature-derived dataset through separation degree filtering, random forest importance filtering, correlation redundancy removal filtering, and comprehensive cost consideration to obtain the final chemical composition indicator dataset. The model building and application module is used to take the final chemical component index dataset as input and the label value of the target flue-cured tobacco leaf as output to train a flue-cured tobacco aroma type classification prediction model; obtain the content of each chemical component index in the final chemical component index dataset of the flue-cured tobacco leaf to be tested, input it into the flue-cured tobacco aroma type classification prediction model, and output the aroma type prediction result of the flue-cured tobacco leaf sample to be tested.

9. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the flue-cured tobacco aroma type prediction method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method for predicting the aroma type of flue-cured tobacco as described in any one of claims 1 to 7.