Cigarette smoke sensory quality prediction method based on multi-source data fusion

By using a multi-source data fusion method, various data features of cigarette smoke are obtained and screened, and a weighted fusion prediction model is constructed. This solves the problems of low prediction accuracy and poor adaptability in existing technologies, and enables accurate prediction and real-time control of the sensory quality of cigarette smoke.

CN122367265APending Publication Date: 2026-07-10CHINA TOBACCO HEBEI INDUSTRIAL CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA TOBACCO HEBEI INDUSTRIAL CO LTD
Filing Date
2026-04-17
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

Existing methods for predicting the sensory quality of cigarette smoke rely on a single data source, resulting in incomplete coverage, unscientific feature selection, ineffective fusion of multi-source data, low prediction accuracy, and weak generalization ability, which cannot meet the real-time prediction and control needs in the cigarette production process.

Method used

A multi-source data fusion method is adopted to obtain data on volatile, semi-volatile, and non-volatile components and physical parameters. Core features are selected through a dual-index screening strategy, and a weighted fusion algorithm is used to construct a prediction model to achieve accurate fusion and prediction of multi-source features.

Benefits of technology

It enables accurate and rapid prediction of the sensory quality of cigarette smoke, reduces the cost of smoking evaluation, improves the efficiency and precision of quality monitoring and control in the production process, and meets the needs of large-scale cigarette production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367265A_ABST
    Figure CN122367265A_ABST
Patent Text Reader

Abstract

The application discloses a cigarette smoke sensory quality prediction method based on multi-source data fusion, and relates to the field of tobacco processing and artificial intelligence. The method specifically comprises the following steps: multi-source data acquisition, multi-source data preprocessing and feature screening, construction of a multi-source fusion feature set, prediction model construction and optimization, and sensory quality prediction. The application eliminates the detection blind area through multi-source data fusion complementation, improves the prediction performance by combining feature screening and intelligent algorithms, is simple and convenient to operate, is accurate in prediction, has strong repeatability, can be directly adapted to online monitoring and prediction of sensory quality in the cigarette production process, provides reliable technical support for accurate regulation and control of cigarette quality, reduces the dependence cost of professional taste testers, and improves production efficiency and product consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the interdisciplinary field of tobacco processing and artificial intelligence, and specifically relates to a method for predicting the sensory quality of cigarette smoke based on multi-source data fusion, which is particularly suitable for online monitoring, prediction and precise control of the sensory quality of cigarette smoke during the cigarette production process. Background Technology

[0002] The sensory quality of cigarette smoke is a key indicator of the core competitiveness of cigarette products, directly determining the consumer smoking experience and market acceptance. Its main evaluation dimensions include comfort, irritation, and balance. Traditional sensory quality evaluation of cigarette smoke relies on the subjective judgment of professional smoke evaluators, which suffers from poor consistency in evaluation results, strong subjectivity, low efficiency, and high costs. Furthermore, it cannot achieve real-time prediction and control during the production process, making it difficult to adapt to the large-scale and precise demands of modern cigarette production.

[0003] With the development of detection and artificial intelligence technologies, existing technologies are gradually adopting a combination of single data sources and intelligent algorithms to predict sensory quality, such as building prediction models solely based on volatile component data or physical property data. However, the sensory quality of cigarette smoke is affected by multiple factors in synergy, and a single data source has limitations in terms of limited coverage and incomplete feature information, resulting in low accuracy and weak generalization ability of the prediction model, which cannot fully reflect the comprehensive impact of smoke components and physical properties on sensory quality.

[0004] Meanwhile, existing prediction methods lack scientific feature selection and data fusion strategies: on the one hand, the raw data is not effectively screened, and redundant features interfere with model performance, increase computational load, and reduce prediction accuracy; on the other hand, effective fusion of multi-source data is not achieved, failing to leverage the complementary advantages of each data source, resulting in low data utilization. Furthermore, existing model parameter settings are unclear, have poor feasibility, and insufficient prediction accuracy and stability, making them difficult to replace traditional professional sensory evaluation and unable to meet the actual needs of real-time prediction and precise control of sensory quality in cigarette production.

[0005] Therefore, there is an urgent need for an efficient and accurate prediction method based on multi-source data fusion to solve the technical problems of incomplete coverage of single data sources, unscientific feature selection, ineffective fusion of multi-source data, low prediction accuracy, weak generalization ability, and poor feasibility in existing technologies, so as to provide reliable technical support for quality control in cigarette production. Summary of the Invention

[0006] The purpose of this invention is to overcome the aforementioned defects in the existing technology and provide a method for predicting the sensory quality of cigarette smoke based on multi-source data fusion. This method aims to solve the technical problems of incomplete coverage of a single data source, unscientific feature selection, ineffective fusion of multi-source data, low prediction accuracy, weak generalization ability, and poor feasibility in existing prediction methods. It aims to achieve accurate and rapid prediction of the sensory quality of cigarette smoke, replace some functions of traditional professional smoking evaluation, reduce smoking evaluation costs, improve the efficiency and accuracy of quality monitoring and control in the cigarette production process, and ensure product quality consistency.

[0007] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:

[0008] A method for predicting the sensory quality of cigarette smoke based on multi-source data fusion is characterized by the following steps: Acquire data on volatile components, semi-volatile components, and non-volatile components of cigarette smoke, as well as physical parameters of the smoke; The acquired multi-source data were preprocessed, and feature parameters strongly correlated with the sensory quality of flue gas were selected based on a dual-index screening strategy to obtain the core feature subset of each data source. Based on the core feature subsets, the importance weights of each data source are determined, and a weighted fusion algorithm is used to fuse each core feature subset to construct a multi-source fusion feature set; Using the multi-source fusion feature set as input and the flue gas sensory quality score as output, a sensory quality prediction model is constructed. The optimal prediction model is obtained through model parameter optimization and training. The multi-source data and multi-source fusion feature set of the cigarette smoke to be predicted are obtained, and then input into the optimal prediction model to output the sensory quality prediction result of the cigarette smoke.

[0009] The beneficial effects of this invention are as follows: (1) Comprehensive coverage of multi-source data to eliminate prediction blind spots: For the first time, four core data sources of cigarette smoke, namely volatile, semi-volatile, non-volatile components and physical properties, are integrated to comprehensively cover the key factors affecting the sensory quality of cigarette smoke. This solves the defects of existing technologies, such as incomplete coverage of single data sources and insufficient feature information, and provides a solid data foundation for accurate prediction. The contribution of feature subsets of each data source to the prediction of sensory quality is objectively calculated by random forest model. Based on this, weighted fusion is performed to achieve synergistic complementarity of the four data sources, comprehensively reflect the chemical and physical properties of cigarette smoke, and ensure the comprehensiveness and accuracy of the prediction results.

[0010] (2) Dual-index feature screening to improve feature effectiveness: The dual-index screening strategy of “variance coefficient method + mutual information method” is adopted to ensure that the core features have sufficient discrimination and that they are strongly correlated with sensory quality. This effectively eliminates redundant features, reduces feature interference, reduces model computation, and improves model prediction accuracy. Compared with single feature screening methods, dual-index screening can more accurately screen out core effective features and avoid redundant features from interfering with model performance.

[0011] (3) The weighted fusion algorithm is scientific and gives full play to the complementary advantages of data: The weighted fusion algorithm based on the importance weight of features, combined with random forest to calculate the weight of each data source, realizes the accurate fusion of multi-source features, gives full play to the complementary advantages of various types of data, solves the problem of ineffective fusion of multi-source data and low data utilization in existing technologies, and further improves the stability and generalization ability of the prediction model; the objective weighting method avoids the bias of subjective weighting and ensures the optimal fusion effect of multi-source data.

[0012] (4) The model parameters are clear and the implementation is strong: the prediction model is constructed by using the random forest algorithm, the initial parameters and optimization range are clear, the parameters are optimized by the grid search method, the prediction accuracy and stability of the model are ensured, and the algorithm is mature and easy to implement, and can be directly adapted to laboratory scientific research and industrial online monitoring; the clear parameter settings reduce the difficulty of implementation of the method, and make it easy for those skilled in the art to quickly master and apply it.

[0013] (5) Replace some professional evaluations, reduce costs and improve efficiency: It can realize rapid and accurate prediction of the sensory quality of cigarette smoke, effectively replace some functions of traditional professional evaluations, solve the defects of traditional evaluations that are subjective, inefficient and costly, and can realize real-time prediction in the production process, improve the efficiency of cigarette quality control and ensure product consistency; no professional evaluation personnel are required, which greatly reduces labor costs, and the prediction speed is fast, which can meet the needs of large-scale cigarette production.

[0014] (6) Wide adaptability and strong practicality: It can be adapted to cigarette samples of different batches and different process parameters, supports dynamic model updates, and can meet the actual needs of online quality monitoring, prediction and precise control in large-scale cigarette production. It has strong practicality and promotion value; the model update function can adapt the method to changes in the production process, further improving the applicability and practicality of the method. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the method steps provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the multi-source data fusion logic in an embodiment of the present invention; Figure 3 This is a schematic diagram illustrating the fitting effect of the prediction model in an embodiment of the present invention; Figure 4 This is a schematic diagram of the system structure provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of an electronic device structure provided in an embodiment of the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Furthermore, the following description is for illustrative purposes and not for limitation, and sets forth specific details such as particular system architectures and techniques to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems and methods are omitted to avoid unnecessary detail that could obscure the description of the invention.

[0019] This invention first provides a method for predicting the sensory quality of cigarette smoke based on multi-source data fusion, referring to... Figure 1 ,include: Acquire data on volatile components, semi-volatile components, and non-volatile components of cigarette smoke, as well as physical parameters of the smoke; The acquired multi-source data were preprocessed, and feature parameters strongly correlated with the sensory quality of flue gas were selected based on a dual-index screening strategy to obtain the core feature subset of each data source. Based on the core feature subsets, the importance weights of each data source are determined, and a weighted fusion algorithm is used to fuse each core feature subset to construct a multi-source fusion feature set; Using the multi-source fusion feature set as input and the flue gas sensory quality score as output, a sensory quality prediction model is constructed. The optimal prediction model is obtained through model parameter optimization and training. The multi-source data and multi-source fusion feature set of the cigarette smoke to be predicted are obtained, and then input into the optimal prediction model to output the sensory quality prediction result of the cigarette smoke.

[0020] To fully disclose the method of this embodiment so that those skilled in the art can freely implement this solution, the embodiments of the present invention will be further explained in detail below in conjunction with specific technical details and experimental processes.

[0021] In this embodiment, to achieve the acquisition of multi-source data: Cigarette samples from different batches and with different process parameters were selected, and four core data sources were collected from the smoke of each cigarette to ensure comprehensive and representative data coverage. The specific collection method is as follows: (1) Volatile component data: Gas chromatography-mass spectrometry was used for detection. The chromatographic column was HP-5MS (30 m×0.25mm×0.25μm). The temperature program was 40°C for 3 min, then increased to 250°C at 5°C / min and held for 10 min. The injection port temperature was 250°C, the carrier gas was helium, the flow rate was 1.0 mL / min, the injection volume was 1 μL, and the split ratio was 10:1. The mass spectrometer detector voltage was 70 eV, the scanning range was m / z 35–450, and the content data of at least 80 volatile components (such as furfural, benzaldehyde, etc.) were obtained. Each sample was detected three times, and the average value was taken as the final data. (2) Semi-volatile component data: The data were detected using a two-dimensional gas chromatography-mass spectrometry system. The first dimension column was a DB-5MS (30 m × 0.25 mm × 0.25 μm), and the second dimension column was a DB-17MS (1.5 m × 0.25 mm × 0.25 μm). The temperature program was the same as that for the volatile component data. The content data of at least 50 semi-volatile components (such as 5-hydroxymethylfurfural, eugenol, etc.) were obtained. Each sample was tested three times, and the average value was taken as the final data. (3) Non-volatile component data: Liquid chromatography-tandem mass spectrometry was used for detection. The chromatographic column was C18 (150 mm × 2.1 mm × 1.8 μm), the mobile phase was methanol-0.1% formic acid aqueous solution, and the gradient elution program was as follows: 0–5 min methanol volume fraction increased from 30% to 60%, 5–15 min from 60% to 90%, 15–20 min maintained at 90%, and 20–25 min decreased to 30%; the flow rate was 0.3 mL / min, the injection volume was 5 μL, and the column temperature was 30°C; the mass spectrometry used an electrospray ionization source and positive ion mode to detect the content data of at least 30 non-volatile components (such as chlorogenic acid, caffeic acid, etc.). Each sample was detected three times, and the average value was taken as the final data. (4) Flue gas physical parameter data: The flue gas physical property tester was used to test the core parameters, including flue gas temperature (°C), flue gas pressure (kPa), flue gas density (kg / m³), and suction resistance (kPa). Each sample was tested three times, and the average value was taken as the final physical parameter data.

[0022] The selection of these four data sources is based on the inventors' in-depth analysis of the formation mechanism of smoke sensory quality: volatile components mainly determine the aroma quality and quantity; semi-volatile components affect the smoothness and fullness of the smoke; non-volatile components affect irritation by regulating acid-base balance; and physical parameters are directly related to vaping comfort. Experiments have verified that any single data source has a "prediction blind spot." However, due to the significant differences in dimensions, scale, and information density among the four data sources, they cannot be directly combined. Therefore, this embodiment successfully overcomes the technical bias of incompatibility between heterogeneous data sources through a three-step strategy of subsequent preprocessing, feature selection, and weighted fusion, providing a complete and solid data foundation for accurate prediction.

[0023] In this embodiment, to achieve multi-source data preprocessing and feature selection: The above four types of data sources are preprocessed to eliminate the impact of outliers, missing values, and dimensional differences on subsequent analysis. Then, a dual-indicator feature screening strategy is used to remove redundant features and retain core effective features. The specific steps are as follows: (1) Outlier removal: The Grubbs criterion (significance level α=0.05) is used to detect and remove outliers for individual feature parameters in each data source to ensure data reliability. The Grubbs criterion is an outlier detection method based on normal distribution, which is suitable for outlier removal of small sample data. It can accurately identify abnormal data that deviates from the normal range and avoid interference of abnormal data with subsequent model construction. (2) Missing value imputation: For missing values ​​that appear after outliers are removed, the K-nearest neighbor algorithm (K=5) is used for imputation. Based on the corresponding feature parameter values ​​of similar samples, the reasonableness of the imputation data is ensured. The K-nearest neighbor algorithm calculates the similarity between the missing value sample and other samples, and selects the mean of the corresponding features of the 5 most similar samples as the imputation value. Compared with the traditional mean imputation, it can better preserve the distribution characteristics of the data. (3) Standardization: The Z-score standardization method is used to standardize each type of data source after preprocessing, converting all feature parameters into standardized data with a mean of 0 and a standard deviation of 1, thus eliminating the difference in units. Since the feature parameters of the four types of data sources have different units (such as the unit of component content being μg / mL, and the unit of physical parameters being °C and kPa), the standardization process can avoid the impact of the difference in units on model training and ensure that each feature parameter has equal weight. (4) Feature screening: The dual-index screening strategy of "variance coefficient method + mutual information method" is adopted to screen core feature parameters that are strongly correlated with the sensory quality of flue gas and eliminate redundant features: ① The variation coefficient is ≥0.3 to ensure that the feature parameter has sufficient discrimination. The smaller the variation coefficient, the smaller the difference between different samples and the weaker the discrimination ability. It should be eliminated; ② The mutual information value is ≥0.5 to ensure that the feature parameter has a strong correlation with the sensory quality score. The larger the mutual information value, the higher the correlation between the feature and the sensory quality and the greater the contribution to the prediction result. Feature parameters that meet both criteria are included in the core feature subset, and the core feature subsets of volatile, semi-volatile, non-volatile components and physical parameters are obtained respectively.

[0024] In this embodiment, the present invention innovatively employs a tandem dual-indicator screening strategy of "coefficient of variation method + mutual information method," with the two methods exhibiting a significant synergistic screening effect. The synergistic effect lies in the following: the coefficient of variation (≥0.3) firstly filters out features with sufficient discriminative power among different samples from a statistical perspective, eliminating "mediocre features" with weak discriminative ability; then, from an informatics correlation perspective, it further filters out features strongly correlated with sensory quality scores, eliminating "irrelevant features" that, while having discriminative power, are irrelevant to the target. Compared to a single screening method, using only the coefficient of variation would retain highly fluctuating features unrelated to sensory quality, introducing noise; using only mutual information might select weakly discriminative features with low discriminative power but dependent on the overall distribution. The dual-indicator strategy of this embodiment forms a double barrier of "discriminative power" and "correlation," producing an unexpectedly precise purification effect.

[0025] In this embodiment, to construct a multi-source fusion feature set: Reference Figure 2 Based on the importance weights of each core feature subset, a weighted fusion algorithm is used to accurately fuse the four types of core feature subsets, constructing a multi-source fusion feature set. This fully leverages the complementary advantages of each data source and solves the problem of insufficient feature information from a single data source. The specific steps are as follows: (1) Weight calculation: Input the four core feature subsets into the initial random forest model (120 decision trees, minimum number of samples for node splitting 3, minimum number of samples for leaf 1), calculate the feature importance score of each feature subset, and determine the importance weight of each subset according to the score ratio. , , , And satisfy + + + =1; The random forest algorithm can calculate the contribution of features to the model's prediction results, obtain feature importance scores, and determine weights based on the score ratios. It can objectively reflect the degree of influence of each data source on sensory quality prediction and avoid the bias of subjective weighting. (2) Weighted fusion algorithm: Multi-source feature fusion is performed using the following formula: ,in For multi-source fusion feature sets, , , , The core feature subsets are volatile components, semi-volatile components, non-volatile components, and physical parameters. After fusion, the feature set is normalized to ensure that the data range is consistent, resulting in the final multi-source fused feature set. The normalization process can further eliminate the dimensional differences between the fused features and improve the model training efficiency.

[0026] In this embodiment, the four types of data sources work synergistically and complementarily to comprehensively reflect the chemical and physical characteristics of flue gas, ensuring the comprehensiveness and accuracy of the prediction results. "Complementarity" is reflected in the fact that the weight allocation objectively reflects the relative importance differences of different data sources in the prediction, avoiding information dilution caused by average fusion. For example, chemical components (the first three) have higher weights, directly characterizing the sensory material basis; while physical parameters have lower weights, they provide key environmental information on the release and delivery of chemical components. The two form a complementary explanatory framework of "chemical composition-physical delivery" in the model. "Synergy" is reflected in the fact that the final fused feature set F is not a simple concatenation, but rather achieves nonlinear superposition and enhancement of information from different dimensions through weights. This allows the model to make comprehensive judgments based on both the "quality and quantity" of chemical components and the "state and efficiency" of physical parameters, thereby comprehensively reflecting the flue gas quality and significantly improving the model's stability and generalization ability.

[0027] In this embodiment, to construct a sensory quality prediction model: Using a multi-source fusion feature set as input and professional sensory quality evaluation scores as output, a random forest sensory quality prediction model is constructed. Parameter optimization is used to improve the model's prediction accuracy and stability, ensuring that the model can meet the prediction needs of actual production. The specific steps are as follows: (1) Sensory quality evaluation score acquisition: A professional group of 6 evaluators with standardized training was formed. The 1-10 scoring method was used to score the comfort, irritation and harmony of cigarette smoke. Each cigarette was evaluated 3 times and the average value was taken as the individual score. The comprehensive score was calculated according to the weight of comfort (0.4), irritation (0.3) and harmony (0.3) as the core indicator of the model output. Standardized training can ensure that the evaluation standards of the evaluators are consistent and reduce subjective bias. Repeated evaluation and taking the average value can further improve the reliability of sensory evaluation scores. (2) Data set partitioning: The multi-source fusion feature set and the corresponding sensory quality evaluation comprehensive score are divided into a training set (for model training) and a validation set (for model validation) in a 7:3 ratio. The 7:3 partitioning ratio is the optimal ratio that balances the model training effect and the validation reliability. It can ensure that the training set has a sufficient number of samples so that the model can fully learn the relationship between features and output, and can also effectively test the generalization ability of the model through the validation set. (3) Initial model construction: Construct a random forest prediction model with the following initial parameters: 120 decision trees, 3 minimum number of samples for node splitting, and 1 minimum number of leaf samples. The random forest algorithm has the advantages of anti-overfitting, strong generalization ability, and ability to handle high-dimensional data. It is suitable for prediction tasks of multi-source fusion feature sets. The initial parameter settings are determined based on a large number of pre-experiments to ensure the initial performance of the model. (4) Model optimization: The model parameters are optimized using the grid search method. The optimization range is: 80-200 decision trees, 2-8 minimum number of samples for node splitting, and 1-4 minimum number of leaf samples. The optimal model parameters are determined with the highest prediction accuracy and the lowest prediction error on the validation set as the optimization objective. The grid search method can traverse all parameter combinations to find the optimal parameter configuration. Compared with manual parameter tuning, it is more efficient and more accurate. (5) Model validation: Input the validation set into the optimized prediction model, calculate the prediction accuracy and prediction error, and ensure that the prediction accuracy is ≥92% and the prediction error is ≤0.5 points (1-10 points) to obtain the optimal random forest sensory quality prediction model. Model validation can test the generalization ability and prediction accuracy of the model and ensure that the model can meet the accurate prediction needs in actual production.

[0028] In this embodiment, to achieve subsequent prediction of flue gas sensory quality: Multi-source data (volatile, semi-volatile, and non-volatile components, as well as physical parameters) of the cigarette smoke to be predicted are collected. The aforementioned preprocessing, feature selection, and multi-source fusion processes are repeated to obtain the multi-source fusion feature set to be predicted. This set is then input into the optimal prediction model, which outputs individual prediction scores for comfort, irritation, and harmony, as well as a comprehensive prediction score, thus completing the prediction of the cigarette smoke's sensory quality. This process is simple and fast, requires no professional smoke testers, and enables real-time prediction of cigarette smoke sensory quality.

[0029] In an optional embodiment, the prediction method further includes model updating: Every 50 new sets of multi-source data and corresponding sensory evaluation data are accumulated and added to the training set. The model is then retrained and the parameters are optimized. The optimal prediction model is updated to ensure the model's generalization ability and prediction accuracy, and to adapt to the prediction needs of different batches and different processes of cigarettes.

[0030] Since raw materials and processes may change during cigarette production, model updates can ensure that the model maintains high predictive performance, extend its lifespan, and improve the practicality of the method.

[0031] The above embodiments have provided a detailed explanation of the technical implementation process, core principles, and key parameter configuration of the spectral analysis method of the present invention. To further verify the effectiveness, stability, and technological advancement of the method in practical applications compared to existing technologies, the core technical steps of the present invention are further elaborated below with specific application examples. To simplify the text and avoid redundancy, simplified technical details are used in the following examples, and the omitted content can be referred to in the foregoing descriptions. Such simplified descriptions do not constitute any limitation on the scope of protection of the present invention.

[0032] This application example provides a method for predicting the sensory quality of cigarette smoke based on multi-source data fusion. The specific steps are as follows: Step 1: Multi-source data acquisition Ten batches of flue-cured cigarette samples were selected, with 10 cigarettes in each batch, totaling 100 cigarettes. Four types of data sources were collected from the smoke of each cigarette, as detailed below: (1) Volatile component data: Gas chromatography-mass spectrometry was used for detection. The detection conditions were set according to the technical plan. A total of 86 volatile components were detected. Each sample was tested three times and the average value was taken. (2) Semi-volatile component data: The semi-volatile components were detected by a two-dimensional gas chromatography-mass spectrometry system. The detection conditions were set according to the technical plan. A total of 54 semi-volatile components were detected. Each sample was tested three times and the average value was taken. (3) Non-volatile component data: Liquid chromatography-tandem mass spectrometry was used for detection. The detection conditions were set according to the technical plan. A total of 32 non-volatile components were detected. Each sample was tested three times and the average value was taken. (4) Flue gas physical parameter data: The core parameters were tested using a flue gas physical property tester: flue gas temperature (28.5±0.3°C), flue gas pressure (101.2±0.5 kPa), flue gas density (1.2±0.1 kg / m³), and suction resistance (1.8±0.2 kPa). Each sample was tested three times and the average value was taken.

[0033] Step 2: Multi-source data preprocessing and feature selection The above four types of data sources are preprocessed and feature-selected as follows: (1) Outlier removal: The Grubbs criterion (α=0.05) was used to remove outliers from various data sources. A total of 6 outliers were removed (3 volatile components, 2 semi-volatile components, and 1 physical parameter). (2) Missing value imputation: The K-nearest neighbor algorithm (K=5) is used to imput the 4 missing values ​​that appear after outlier removal to ensure data integrity; (3) Standardization: The Z-score standardization method is used to standardize all preprocessed data to eliminate dimensional differences; (4) Feature screening: The dual-index screening strategy of “variance coefficient ≥ 0.3 + mutual information value ≥ 0.5” was adopted to screen the core feature subsets: ① core feature subset of volatile components (F1): 25 types; ② core feature subset of semi-volatile components (F2): 18 types; ③ core feature subset of non-volatile components (F3): 15 types; ④ core feature subset of physical parameters (F4): 10 types; a total of 68 feature parameters in the four core feature subsets.

[0034] Step 3: Multi-source feature fusion (1) Weight calculation: Input F1, F2, F3, and F4 into the initial random forest model (120 decision trees, minimum number of samples for node splitting 3, minimum number of samples for leaf 1), calculate the feature importance score of each feature subset, and determine the weight according to the score ratio: =0.35、 =0.30、 =0.25、 =0.10; (2) Weighted fusion: The formula F=0.35F1+0.30F2+0.25F3+0.10F4 is used to perform weighted fusion on the four core feature subsets. After fusion, normalization is performed to obtain a multi-source fusion feature set of 100 samples (68 feature parameters for each sample).

[0035] Step 4: Predictive Model Construction and Optimization (1) Sensory evaluation score acquisition: Six professional evaluators conducted sensory evaluations on 100 cigarette samples, using a 1–10 point scoring method to evaluate comfort, irritation, and coordination. Each sample was evaluated three times, and the average value was taken as the individual score. The comprehensive score was calculated based on the weights of comfort (0.4), irritation (0.3), and coordination (0.3), and was used as the model output index. The comprehensive score range was 2.8–8.9 points. (2) Data set division: The multi-source fusion feature set of 100 samples and the corresponding comprehensive score are divided into a training set (70 samples) and a validation set (30 samples) in a ratio of 7:3. (3) Initial model construction: Construct a random forest prediction model with the following initial parameters: 120 decision trees, minimum number of samples for node splitting, and minimum number of leaf samples. (4) Model optimization: The grid search method was used to optimize the parameters. The optimization range was: 80-200 decision trees, 2-8 minimum number of samples for node splitting, and 1-4 minimum number of leaf samples. The optimal parameters after optimization were: 140 decision trees, 4 minimum number of samples for node splitting, and 2 minimum number of leaf samples. (5) Model Validation: Input the validation set into the optimal prediction model, calculate the model performance index, and refer to... Figure 3 Training set goodness of fit R 2 =0.90, validation set goodness of fit R 2 =0.85, prediction accuracy =91.0%, average prediction error =0.32 points, maximum prediction error =0.48 points, all of which meet the performance requirements, and the optimal random forest sensory quality prediction model is obtained.

[0036] Step 5, Sensory Quality Prediction Ten new batches of cigarette samples to be predicted were selected, and their smoke multi-source data were collected. The preprocessing, feature screening, and multi-source fusion processes in steps two and three were repeated to obtain the multi-source fusion feature set of the ten samples to be predicted. The feature set was then input into the optimal prediction model, and the individual prediction scores for comfort, irritation, and harmony, as well as the comprehensive prediction score, were output for each sample to be predicted. Some prediction results are shown in Table 1 below: Table 1. Predicted sensory quality results of cigarette samples to be predicted

[0037] Another embodiment of the present invention relates to a cigarette smoke sensory quality prediction system based on multi-source data fusion, referring to... Figure 4 ,include: The multi-source data acquisition unit is used to acquire data on volatile components, semi-volatile components, and non-volatile components of cigarette smoke, as well as physical parameters of the smoke. The preprocessing and feature selection unit is used to preprocess the acquired multi-source data and select feature parameters that are strongly correlated with the sensory quality of flue gas based on the dual-index selection strategy, thereby obtaining the core feature subset of each data source. A multi-source feature fusion unit is used to determine the importance weight of each data source based on the core feature subset, and to fuse each core feature subset using a weighted fusion algorithm to construct a multi-source fusion feature set; The prediction model construction and optimization unit is used to construct a sensory quality prediction model with the multi-source fusion feature set as input and the flue gas sensory quality score as output, and obtain the optimal prediction model through model parameter optimization and training. The sensory quality prediction unit acquires multi-source data and multi-source fusion feature sets of the cigarette smoke to be predicted, inputs them into the optimal prediction model, and outputs the sensory quality prediction results of the cigarette smoke to be predicted.

[0038] Figure 5 This is a schematic diagram of an electronic device 10 provided in another embodiment of the present invention. (See diagram below.) Figure 5 As shown, the electronic device 10 of this embodiment includes: a processor 11, a memory 12, and a computer program 13 stored in the memory 12 and executable on the processor 11, such as a program for a method of predicting the sensory quality of cigarette smoke based on multi-source data fusion. When the processor 11 executes the computer program 13, it implements the steps in the above embodiment of the method for predicting the sensory quality of cigarette smoke based on multi-source data fusion, for example... Figure 1 The steps are shown.

[0039] For example, computer program 13 may be divided into one or more modules / units, one or more of which are stored in memory 12 and executed by processor 11 to complete the present invention. One or more modules / units may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 13 in electronic device 10.

[0040] Electronic device 10 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Electronic device 10 may include, but is not limited to, a processor 11 and a memory 12. Those skilled in the art will understand that... Figure 5 This is merely an example of electronic device 10 and does not constitute a limitation on electronic device 10. It may include more or fewer components than shown, or combine certain components, or different components. For example, electronic device 10 may also include input / output devices, network access devices, buses, etc.

[0041] The processor 11 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0042] The memory 12 can be an internal storage unit of the electronic device 10, such as a hard disk or RAM of the electronic device 10. The memory 12 can also be an external storage device of the electronic device 10, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the electronic device 10. Furthermore, the memory 12 can include both internal and external storage units of the electronic device 10. The memory 12 is used to store computer programs and other programs and data required by the electronic device 10. The memory 12 can also be used to temporarily store data that has been output or will be output.

[0043] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0044] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments or adapt to the prior art.

[0045] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0046] In the embodiments provided by this invention, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0047] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0048] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0049] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.

[0050] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for predicting the sensory quality of cigarette smoke based on multi-source data fusion, characterized in that, Includes the following steps: Acquire data on volatile components, semi-volatile components, and non-volatile components of cigarette smoke, as well as physical parameters of the smoke; The acquired multi-source data were preprocessed, and feature parameters strongly correlated with the sensory quality of flue gas were selected based on a dual-index screening strategy to obtain the core feature subset of each data source. Based on the core feature subsets, the importance weights of each data source are determined, and a weighted fusion algorithm is used to fuse each core feature subset to construct a multi-source fusion feature set; Using the multi-source fusion feature set as input and the flue gas sensory quality score as output, a sensory quality prediction model is constructed. The optimal prediction model is obtained through model parameter optimization and training. The multi-source data and multi-source fusion feature set of the cigarette smoke to be predicted are obtained, and then input into the optimal prediction model to output the sensory quality prediction result of the cigarette smoke.

2. The method for predicting the sensory quality of cigarette smoke according to claim 1, characterized in that: To obtain the volatile component data: gas chromatography-mass spectrometry (GC-MS) was used for detection. The chromatographic column was HP-5MS, the temperature program was 40°C held for 3 min, then increased to 250°C at 5°C / min and held for 10 min, the injection port temperature was 250°C, the carrier gas was helium, the flow rate was 1.0 mL / min, the injection volume was 1 μL, the split ratio was 10:1, the mass spectrometer detector voltage was 70 eV, and the scan range was m / z 35–450. The content data of at least 80 volatile components were obtained. Each sample was detected three times, and the average value was taken as the final data. To obtain the semi-volatile component data: a two-dimensional gas chromatography-mass spectrometry (GC-MS) system was used for detection. The first dimension column was a DB-5MS column and the second dimension column was a DB-17MS column. The temperature program was the same as that for the volatile component data detection. The content data of at least 50 semi-volatile components were obtained. Each sample was detected three times, and the average value was taken as the final data. To obtain data on non-volatile components, liquid chromatography-tandem mass spectrometry (LC-MS / MS) was used. The chromatographic column was C18, and the mobile phase was methanol-0.1% formic acid aqueous solution. The gradient elution program was as follows: 0–5 min methanol volume fraction 30%–60%, 5–15 min 60%–90%, 15–20 min hold at 90%, 20–25 min reduce to 30%, flow rate 0.3 mL / min, injection volume 5 μL, column temperature 30°C, and electrospray ionization source in positive ion mode for mass spectrometry. The content data of at least 30 non-volatile components were obtained. Each sample was detected three times, and the average value was taken as the final data. To obtain flue gas physical parameter data: a flue gas physical property tester was used to test, including flue gas temperature, flue gas pressure, flue gas density, and suction resistance. Each parameter was tested three times, and the average value was taken as the final physical parameter data.

3. The method for predicting the sensory quality of cigarette smoke according to claim 1, characterized in that: Preprocessing of the acquired multi-source data includes: removing outliers using the Grubbs criterion, filling missing values ​​using the K-nearest neighbor algorithm, and standardizing the data using the Z-score standardization method.

4. The method for predicting the sensory quality of cigarette smoke according to claim 1, characterized in that: The dual-index screening strategy is based on the coefficient of variation and mutual information value to select feature parameters that are strongly correlated with the sensory quality of flue gas. Feature parameters that simultaneously meet the two criteria (coefficient of variation ≥ 0.3 and mutual information value ≥ 0.5) are included in the core feature subset.

5. The method for predicting the sensory quality of cigarette smoke according to claim 1, characterized in that: The specific formula for the weighted fusion algorithm is as follows: ,in For multi-source fusion feature sets, , , , These are the core feature subsets of volatile components, semi-volatile components, non-volatile components, and physical parameters. , , , These are the importance weights of the corresponding core feature subsets, and + + + =1; The importance weights are obtained through the random forest algorithm, specifically as follows: Each core feature subset is input into the initial random forest model, the feature importance score of each subset is calculated, and the importance weight is determined according to the score ratio.

6. The method for predicting the sensory quality of cigarette smoke according to claim 1, characterized in that: The sensory quality prediction model is a random forest model, and its initial parameters are: The number of decision trees is 100–150, the minimum number of samples for node splits is 2–5, and the minimum number of leaf samples is 1–2. The range of model parameters optimized by grid search is: the number of decision trees is 80–200, the minimum number of samples for node splits is 2–8, and the minimum number of leaf samples is 1–4, with the optimization objective being the highest prediction accuracy on the validation set. The ratio of the training set to the validation set is 7:

3.

7. The method for predicting the sensory quality of cigarette smoke according to claim 1, characterized in that: The sensory quality score of the smoke is obtained through professional sensory evaluation, which is conducted by a group of evaluators with smoking evaluation certificates. The evaluation method is 1-10 points, and the comfort, irritation and harmony are scored respectively. Each cigarette is evaluated three times and the average value is taken as the final sensory quality score. The scores for comfort, stimulation, and coordination are weighted at 0.4, 0.3, and 0.3, respectively, with the overall score serving as the core indicator for the model output.

8. A cigarette smoke sensory quality prediction system based on multi-source data fusion, characterized in that: The system is used to implement the method for predicting the sensory quality of cigarette smoke based on multi-source data fusion as described in any one of claims 1 to 7, including: The multi-source data acquisition unit is used to acquire data on volatile components, semi-volatile components, and non-volatile components of cigarette smoke, as well as physical parameters of the smoke. The preprocessing and feature selection unit is used to preprocess the acquired multi-source data and select feature parameters that are strongly correlated with the sensory quality of flue gas based on the dual-index selection strategy, thereby obtaining the core feature subset of each data source. A multi-source feature fusion unit is used to determine the importance weight of each data source based on the core feature subset, and to fuse each core feature subset using a weighted fusion algorithm to construct a multi-source fusion feature set; The prediction model construction and optimization unit is used to construct a sensory quality prediction model with the multi-source fusion feature set as input and the flue gas sensory quality score as output, and obtain the optimal prediction model through model parameter optimization and training. The sensory quality prediction unit acquires multi-source data and multi-source fusion feature sets of the cigarette smoke to be predicted, inputs them into the optimal prediction model, and outputs the sensory quality prediction results of the cigarette smoke to be predicted.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the steps of the method for predicting the sensory quality of cigarette smoke based on multi-source data fusion as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by the processor, it implements the steps of the method for predicting the sensory quality of cigarette smoke based on multi-source data fusion as described in any one of claims 1 to 7.