A method for simulating and formulating a cigarette leaf blend recipe based on tobacco leaf substitution
By constructing a tobacco leaf data information database and using linear discriminant analysis and classification training to generate a tobacco leaf type prediction model, optimizing tobacco leaf replacement and proportion, the problem of low maintenance efficiency of traditional cigarette formulas is solved, and an efficient and scientific cigarette leaf group formula design is achieved.
Patent Information
- Application Number
- CN202111126827.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-26
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-09-26
AI Technical Summary
The maintenance of traditional cigarette formulas relies on experts to evaluate the low efficiency, time-consuming and high cost, and the low data integration of cigarette companies affects the design efficiency.
Tobacco leaf data information database is constructed, and tobacco leaf type standard prediction model is generated using linear discriminant analysis and classification training. Tobacco leaf replacement is optimized through European-style distance and linear weighting calculations, and the ratio is adjusted to generate a similar style of cigarette leaf group formula.
The cigarette leaf group formula design based on non-artificial experience is realized, which improves the design efficiency and scientificity of the formula, reduces repetitive labor, and enhances production stability.
Smart Images

Figure CN115868656B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cigarette and cigarette product quality detection, in particular to a cigarette leaf group formula imitation design method based on tobacco leaf substitution. Background Art
[0002] Cigarette blends are the foundation and core of tobacco production for tobacco companies. Cigarette blends are composed of a variety of single-grade tobacco leaves blended in specific proportions. Typically, a finished cigarette blend is made up of 15 to 30 different single-ingredient tobacco leaves blended in specific proportions. During the cigarette production process, the cigarette blend plays a crucial role in maintaining the stable quality of cigarette brands. Traditionally, maintaining cigarette blends is a complex task, requiring sensory evaluation and measurement instruments to comprehensively assess whether the sensory and smoke performance of the updated cigarette blend meets the required standards. This typically requires extensive testing experience and extensive training and practical experience. The health and mental state of the testers also significantly influence the testing results. This reliance on expert testing and analytical testing is inefficient, time-consuming, and costly. Therefore, the development of intelligent cigarette blend maintenance technologies and methods is both necessary and urgent.
[0003] Over the years, cigarette companies have accumulated vast amounts of valuable basic and R&D data, but these data are often scattered, poorly integrated, poorly correlated, and inconsistent. In recent years, with the widespread application and development of machine learning, data mining, and artificial intelligence in various fields, these methods have been used to mine and process this data, unlocking its immense value. This improves understanding of the inherent laws of tobacco production, supports the design of tobacco leaf blends, and achieves digital and scientific approaches to cigarette product design. This has significant practical implications for cigarette companies in scientifically and efficiently designing and producing cigarette products, avoiding duplication of effort, improving work efficiency, and enhancing the stability of cigarette production, thereby increasing market competitiveness and promoting sustainable development.
[0004] The present invention aims to generate a leaf group formula list with a similar style by imitating the existing tobacco leaf group formula, thereby achieving tobacco leaf group formula design based on non-artificial experience and improving the efficiency of tobacco leaf group formula design. Summary of the Invention
[0005] The present invention provides a tobacco leaf substitution-based tobacco leaf replacement imitation design method, which aims to imitate the existing tobacco leaf group formula to generate a leaf group formula list with a similar style, thereby achieving tobacco leaf group formula design based on non-artificial experience and improving the efficiency of tobacco leaf group formula design.
[0006] A tobacco leaf group formulation imitation design method based on tobacco leaf substitution, comprising:
[0007] Construct a tobacco leaf data information database, which includes the production area, grade, classification, near-infrared spectrum and conventional chemical composition data of each tobacco leaf;
[0008] Based on the tobacco leaf data information database, linear discriminant analysis classification training is performed to obtain the tobacco leaf class label prediction model;
[0009] Select a target leaf group formula Q to be imitated, and use the tobacco leaf category prediction model to predict the category label of each tobacco leaf in Q to obtain the corresponding category label C;
[0010] In the tobacco leaf data information database, the infrared spectrum matrix of each tobacco leaf labeled C is used to calculate the Euclidean distance with the near-infrared spectrum matrix of the tobacco leaf in Q. The Euclidean distances are sorted from small to large, and the two tobacco leaves corresponding to the first two digits are selected as candidate replacements for the tobacco leaf in Q. They are recorded as the best replacement and the second best replacement, respectively.
[0011] For each tobacco leaf in the target leaf group formula Q to be imitated, the corresponding optimal alternative tobacco leaf is selected for replacement. If the optimal alternative tobacco leaf is also a component in Q, the suboptimal alternative tobacco leaf is selected for replacement to complete the generation of formula P.
[0012] The conventional chemical composition vectors of formula P and the target leaf group formula Q to be imitated are calculated based on linear weighting and the ratio is adjusted to minimize the difference in conventional chemical composition between formula P and the target leaf group formula Q to be imitated, thereby achieving ratio optimization of formula P.
[0013] Furthermore, the process of achieving the ratio optimization of the formula P includes:
[0014] For each tobacco leaf component in the distribution P, calculate the Euclidean distance d between its conventional chemical composition vector and the conventional chemical composition vector of the target leaf group formula Q to be imitated, and multiply the Euclidean distance d by the initial value of the tobacco leaf component ratio to obtain the influence factor r;
[0015] Sort by the size of the influencing factor r, and adjust the proportion of the tobacco leaf component with the largest influencing factor r so that the difference in conventional chemical composition between the adjusted formula P and the target leaf group formula Q to be imitated is minimized;
[0016] Recalculate the impact factor r of each tobacco component in the adjusted formula P and re-rank them;
[0017] If the reordered order is consistent with the previous order, the ratio optimization ends; otherwise, the above operation is repeated.
[0018] Furthermore, after achieving the ratio optimization of the formula P, the following steps are also included:
[0019] When the ratio of the formula P is optimized and the sum of the ratios of the tobacco leaf components is not equal to 100, it can be controlled by adjusting the proportion of the filler.
[0020] Furthermore, the filler includes cut stems, thin sheets, puffed tobacco or single-ingredient tobacco suitable for use as a filler.
[0021] Furthermore, in the process of generating the recipe P, if the optimal replacement tobacco leaf and the suboptimal replacement tobacco leaf are both components in the target leaf group recipe Q to be imitated, then based on the tobacco leaf type and grade of the tobacco leaf to be replaced in the target leaf group recipe Q to be imitated, a tobacco leaf that is consistent with the tobacco leaf type and grade of the tobacco leaf to be replaced and is not in the target leaf group recipe Q to be imitated is selected from the tobacco leaf data information database as the replacement tobacco leaf;
[0022] If there is no tobacco leaf that meets the conditions in the tobacco leaf data information database, the replacement of the tobacco leaf to be replaced will be abandoned.
[0023] Furthermore, the tobacco leaf types include light-flavor tobacco leaves, medium-flavor tobacco leaves, strong-flavor tobacco leaves, and imported tobacco leaves;
[0024] Before generating the recipe P, it also includes:
[0025] Conduct sensory evaluation on all tobacco leaves in the tobacco leaf database, and score the four aroma tones: fresh sweet aroma, honey sweet aroma, mellow sweet aroma, and burnt sweet aroma.
[0026] Tobacco leaf types are classified according to the scores of the four aroma flavors: in the evaluation results, the tobacco leaf with the highest score for light sweet aroma is defined as light aroma tobacco leaf; in the evaluation results, the tobacco leaf with the highest score for burnt sweet aroma is defined as strong aroma tobacco leaf; in the evaluation results, the tobacco leaf with the highest score for honey sweet aroma or mellow sweet aroma is defined as medium aroma tobacco leaf; tobacco leaves produced in foreign countries are defined as imported tobacco leaves.
[0027] Furthermore, the linear discriminant analysis classification training is performed based on the tobacco leaf data information database to obtain a tobacco leaf class label prediction model, which specifically includes:
[0028] Randomly select n near-infrared spectra of tobacco leaves of various categories and their corresponding class labels from the tobacco leaf data information database as a training set;
[0029] Based on the near-infrared spectra in the training set, a near-infrared spectrum matrix with a dimension of n×p is constructed, where p is the dimension of the near-infrared spectrum;
[0030] The near-infrared spectrum matrix is normalized to obtain a standardized near-infrared spectrum matrix X;
[0031] Compute the covariance matrix of the normalized near-infrared spectral matrix:
[0032] Calculate the eigenvalues and corresponding eigenvectors of the covariance matrix;
[0033] Sort the eigenvalues and their corresponding eigenvectors in descending order, and take the eigenvectors corresponding to the first k eigenvalues to form a dimensionality reduction matrix W with a dimension of p×k;
[0034] Use trainX=X*W as training data and the corresponding class labels as parameters to call Matlab's ClassificationDiscriminant.fit function to perform linear discriminant analysis classification training and obtain the tobacco leaf class label prediction model;
[0035] When using the tobacco leaf class label prediction model to predict the class label of each tobacco leaf in the target leaf group formula Q to be imitated, the near-infrared spectrum of each tobacco leaf in the target leaf group formula Q to be imitated is obtained and annotated to obtain testX, and the predict function of Matlab is called to obtain the predicted class label corresponding to each tobacco leaf.
[0036] Furthermore, the near-infrared spectrum matrix is normalized to obtain a normalized near-infrared spectrum matrix X, and the process includes:
[0037] The near-infrared spectrum matrix is Z-Score normalized, and each element is normalized as follows:
[0038]
[0039] in, is the standard deviation of the j-th variable, represents the average value of the jth variable, where the variable is absorbance, x ij Represents the element in the i-th row and j-th column of the near-infrared spectrum matrix;
[0040] The process of calculating the covariance matrix of the standardized near-infrared spectrum matrix includes:
[0041] Compute the mean vector of the normalized near-infrared spectral matrix:
[0042]
[0043] Among them, x i Represents the vector consisting of the i-th row of the normalized near-infrared spectrum matrix;
[0044] The covariance matrix S of the standardized near-infrared spectral matrix is calculated by the following formula:
[0045]
[0046] Furthermore, the classification label includes an integer part and a decimal part; the integer part represents the production area of the tobacco leaves; the decimal part has a value range of 1-3, corresponding to the upper leaves, middle leaves and lower leaves in the grade of tobacco leaves respectively; the form of tobacco leaves includes tobacco leaves, tobacco leaves or tobacco powder.
[0047] Furthermore, the near infrared spectrum of a single tobacco leaf was collected by a near infrared spectrometer, and the spectral range of the near infrared spectrometer was [12800 cm -1 , 3600cm -1 ] or [780nm, 2778nm]; the scanning speed range is: [1 round / second, 64 rounds / second], the number of scans per round range is: [1, 128], and the resolution range is: [2cm -1 , 64cm -1 ], the total number of collected data points ranges from: [1, 2592].
[0048] Beneficial effects
[0049] The present invention proposes a tobacco leaf group formula imitation design method based on tobacco leaf substitution. By constructing a tobacco leaf category prediction and optimization model through a mathematical intelligent algorithm, the existing tobacco leaf group formula is imitated to generate a leaf group formula list with a similar style, thereby achieving tobacco leaf group formula design based on non-artificial experience and improving the efficiency of tobacco leaf group formula design. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 This is a flow chart of a method for designing a tobacco leaf group recipe based on tobacco leaf substitution provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0052] To make the objectives, technical solutions, and advantages of the present invention more apparent, the technical solutions of the present invention will be described in detail below. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other implementations obtained by those of ordinary skill in the art without inventive effort are within the scope of protection of the present invention.
[0053] like Figure 1 As shown, an embodiment of the present invention provides a method for designing a tobacco leaf group formulation based on tobacco leaf substitution, comprising:
[0054] S1: Construct a tobacco leaf data database, which includes the production area, grade, classification, near-infrared spectrum, and conventional chemical composition data for each tobacco leaf. Conventional chemical composition includes the content of total sugar, reducing sugar, total nitrogen, total alkali, chlorine, and potassium. The combination of conventional chemical component content values constitutes a conventional chemical composition vector.
[0055] In this example, tobacco leaves are divided into 25 categories based on production area and grade. The raw materials for each category and their category labels are shown in the following table. The category label includes an integer and a decimal part. The integer part ranges from 1 to 9, representing different production areas of the tobacco leaves. The decimal part ranges from 1 to 3, corresponding to the upper, middle, and lower leaves of the tobacco leaves within the grade. When the integer part is 9, it indicates foreign tobacco leaves, and no decimal part is set.
[0056]
[0057] S2: Conduct a sensory evaluation of all tobacco leaves in the tobacco leaf database, scoring the four aroma characteristics of fresh sweet, honey sweet, mellow sweet, and burnt sweet. The more pronounced the aroma characteristics of the tobacco leaf, the higher the corresponding score. At the same time, combined with the regional information of the tobacco leaf production area, the tobacco leaves are classified according to the evaluation results.
[0058] The specific classifications are as follows: ① Tobacco leaves with the highest score for light sweet aroma in the evaluation results are defined as light aroma tobacco leaves. ② Tobacco leaves with the highest score for burnt sweet aroma in the evaluation results are defined as strong aroma tobacco leaves. ③ Tobacco leaves with the highest score for honey sweet aroma or mellow sweet aroma in the evaluation results are defined as medium aroma tobacco leaves. ④ Tobacco leaves produced in foreign countries are defined as imported tobacco leaves.
[0059] S3: Perform linear discriminant analysis and classification training based on the tobacco leaf data database to obtain a tobacco leaf class label prediction model. Specifically including:
[0060] S31: Randomly select n near-infrared spectra of tobacco leaves of various categories and corresponding class labels from the tobacco leaf data information database as a training set; in this embodiment, 74 tobacco leaves and corresponding near-infrared spectra shown in the following table are randomly selected from the tobacco leaf data information database as a training set for linear discriminant analysis training.
[0061] serial number Tobacco leaf name Classification serial number Tobacco leaf name Classification 1 Yunnan Lijiang-B2F 1.1 38 Jiangxi Ji'an-B2F 5.1 2 Yunnan Baoshan Longyang-B3F 1.1 39 B1F, Guiyang, Chenzhou, Hunan 5.1 3 Yunnan Qujing-B1F 1.1 40 C1F, Guiyang, Chenzhou, Hunan 5.2 4 Yunnan Honghe-C3F 1.2 41 Anhui Southern Anhui-C2F 5.2 5 Zhanyi, Qujing, Yunnan-C2F 1.2 42 Hunan Yongzhou Jiangyong-C2F 5.2 6 Dayao, Chuxiong, Yunnan-C2F 1.2 43 Jiangxi Ganzhou-C4F 5.3 7 Yunnan Wenshan-C4F 1.3 44 Anhui Wannan-X3F 5.3 8 Yunnan Lijiang-X2F 1.3 45 Anhui Southern Anhui-C4F 5.3 9 Kunming, Yunnan-C4F 1.3 46 B3F, Changting, Longyan, Fujian 6.1 10 Yuqing, Zunyi, Guizhou-B2F 2.1 47 Fujian Longyan-B3F 6.1 11 Guizhou Zunyi-B2F 2.1 48 Fujian Longyan-B2F 6.1 12 Guizhou Tongren-B4F 2.1 49 Fujian Liancheng-C3F 6.2 13 Guizhou Qiandongnan-C3F 2.2 50 Fujian Wuyishan-C3F 6.2 14 Guizhou Zunyi Yichuan-C3F 2.2 51 Fujian Longyan-C3F 6.2 15 Guizhou Zunyi Tongren-C3L 2.2 52 Fujian Longyan Changting-X3F 6.3 16 Guizhou Tongren-C4F 2.3 53 Fujian Liancheng-C4F 6.3 17 Guizhou Zunyi-X3F 2.3 54 Fujian Sanming-X2F 6.3 18 Guizhou Tongren-X2F 2.3 55 Shandong-B3F 7.1 19 Hubei Enshi-B2F 3.1 56 Shandong-B3F 7.1 20 Hunan Zhangjiajie-B2F 3.1 57 Shandong-B3F 7.1 21 Chongqing Wushan-B2F 3.1 58 Shandong Weifang-C4F 7.2 22 Chongqing Wushan-C3F 3.2 59 Shandong Weifang-C3F 7.2 23 Hunan Xiangxi Longshan-C2F 3.2 60 Shandong Weifang-C4F 7.3 24 Hunan Zhangjiajie-C2F 3.2 61 Shandong Linyi X3F 7.3 25 Hubei Enshi-X2F 3.3 62 Shandong Weifang-C4F 7.3 26 Hunan Xiangxi Longshan-X2F 3.3 63 Liaoning Chaoyang-B3F 8.1 27 Chongqing Wushan-C4F 3.3 64 Liaoning Chaoyang-B4F 8.1 28 Henan Luohe-B3F 4.1 65 Heilongjiang Mudanjiang-B2L 8.1 29 Henan Xuchang-B3F 4.1 66 Heilongjiang Mudanjiang-C3L 8.2 30 Henan Xuchang-B1F 4.1 67 Liaoning Chaoyang-C3F 8.2 31 Henan Shan County-C3F 4.2 68 Harbin, Heilongjiang-C3L 8.2 32 Nanyang, Henan-C3F 4.2 69 Liaoning Tieling-C4F 8.3 33 Shaanxi Shangluo-C3F 4.2 70 Liaoning Chaoyang-X3L 8.3 34 Nanyang, Henan--X2F 4.3 71 Liaoning Chaoyang-X3F 8.3 35 Henan Queshan-X3F 4.3 72 Brazil-BOA 9 36 Shaanxi Shangluo-X2F 4.3 73 Zimbabwe Book-L10 9 37 Anhui Southern Anhui-B3F 5.1 74 Zimbabwe Tianze-L1M 9
[0062] S32: Construct a near-infrared spectrum matrix of dimension n×p based on the near-infrared spectra in the training set, where p is the dimension of the near-infrared spectrum. In this embodiment, n is 74 and p is 1296.
[0063] S33: Normalize the near-infrared spectrum matrix to obtain the standardized near-infrared spectrum matrix X n×p . Specifically including:
[0064] The near-infrared spectrum matrix is Z-Score normalized, and each element is normalized as follows:
[0065]
[0066] in, is the standard deviation of the j-th variable, represents the average value of the jth variable, where the variable is absorbance, x ij Represents the element in the i-th row and j-th column of the near-infrared spectrum matrix.
[0067] S34: Calculate the covariance matrix S of the standardized near-infrared spectrum matrix. Specifically including:
[0068] Compute the mean vector of the normalized near-infrared spectral matrix:
[0069]
[0070] Among them, x i Represents the vector consisting of the i-th row of the normalized near-infrared spectrum matrix;
[0071] The covariance matrix S of the standardized near-infrared spectral matrix is calculated by the following formula:
[0072]
[0073] S35: Calculate the eigenvalue λ of the covariance matrix S according to the following formula i and the corresponding eigenvector v i .
[0074]
[0075] S36: The eigenvalue λ i and the corresponding eigenvector v i Perform descending sorting, and take the eigenvectors corresponding to the first k eigenvalues to form a dimensionality reduction matrix W with a dimension of p×k; in this embodiment, p is 48, so the dimension of the dimensionality reduction matrix W is 1296×48.
[0076] S37: Use trainX=X*W as training data, and use the corresponding class labels as parameters to call Matlab's ClassificationDiscriminant.fit function to perform linear discriminant analysis classification training to obtain a single tobacco producing area and part recognition model F: F=ClassificationDiscriminant.fit(trainX, trainid, 'discrimType', 'linear').
[0077] S4: Select a target leaf group formula Q to be imitated, and use the tobacco leaf class label prediction model to predict the class label of each tobacco leaf in Q to obtain the corresponding class label C. When using the tobacco leaf class label prediction model to predict the class label of each tobacco leaf in the target leaf group formula Q to be imitated, obtain the near-infrared spectrum of each tobacco leaf in the target leaf group formula Q and perform annotated processing to obtain testX. Then call Matlab's predict function to obtain the predicted class label C corresponding to each tobacco leaf: C = predict(F, testX*W).
[0078] S5: In the tobacco leaf data information database, the Euclidean distance is calculated between the infrared spectrum matrix of each tobacco leaf labeled as C and the near-infrared spectrum matrix of the tobacco leaf in Q. The Euclidean distances are sorted from small to large, and the two tobacco leaves corresponding to the first two digits are taken as candidate replacement tobacco leaves for the tobacco leaf in Q, and are recorded as the optimal replacement tobacco leaf and the suboptimal replacement tobacco leaf respectively.
[0079] S6: For each type of tobacco leaf in the target leaf group formula Q to be imitated, the corresponding optimal alternative tobacco leaf is selected for replacement. When the optimal alternative tobacco leaf is also a component in Q, the suboptimal alternative tobacco leaf is selected for replacement to complete the generation of formula P.
[0080] Preferably, in the process of generating formula P, if the optimal replacement tobacco leaf and the suboptimal replacement tobacco leaf are both components in the target leaf group formula Q to be imitated, then according to the tobacco leaf type and grade of the tobacco leaf to be replaced in the target leaf group formula Q to be imitated, tobacco leaves that are consistent with the tobacco leaf type and grade of the tobacco leaf to be replaced and are not in the target leaf group formula Q to be imitated are selected from the tobacco leaf data information database as replacement tobacco leaves; if there are no tobacco leaves that meet the conditions in the tobacco leaf data information database, the replacement of the tobacco leaf to be replaced is abandoned.
[0081] S7: Calculate the conventional chemical composition vectors of the formula P and the target leaf group formula Q to be imitated based on linear weighting and adjust the ratio to minimize the difference in conventional chemical composition between the formula P and the target leaf group formula Q to be imitated, thereby achieving the ratio optimization of the formula P. The specific process includes:
[0082] S71: For each tobacco leaf component in the allocation P, calculate the Euclidean distance d between its conventional chemical composition vector and the conventional chemical composition vector of the target leaf group formula Q to be imitated, and multiply the Euclidean distance d by the initial value of the tobacco leaf component ratio to obtain the influence factor r;
[0083] S72: sorting the tobacco leaf components according to the magnitude of the influencing factor r, and adjusting the proportion of the tobacco leaf components with the largest influencing factor r so that the difference in conventional chemical composition between the adjusted formula P and the target leaf group formula Q to be imitated is minimized;
[0084] S73: Recalculate the influence factor r of each tobacco component in the adjusted formula P and re-rank them;
[0085] S74: If the re-sorted order is consistent with the previous order, the ratio optimization is terminated; otherwise, steps S71 to S74 are repeated.
[0086] S8: When the total ratio of each tobacco leaf component after the optimized ratio of the recipe P is not equal to 100, the ratio of fillers is adjusted to control the ratio. In specific implementation, the fillers include cut stems, flakes, puffed shredded tobacco, or single-ingredient tobacco suitable for use as fillers.
[0087] During implementation, the near infrared spectrum of the single material smoke is collected by a near infrared spectrometer, and the spectral range of the near infrared spectrometer is: [12800cm -1 , 3600cm -1 ] or [780nm, 2778nm]; the scanning speed range is: [1 round / second, 64 rounds / second], the number of scans per round range is: [1, 128], and the resolution range is: [2cm -1 , 64cm -1 ], the total number of collected data points ranges from: [1, 2592].
[0088] It should be noted that the form of tobacco leaves includes but is not limited to tobacco leaves, tobacco leaves, or tobacco powder. The dimension of the infrared spectrum and the number of training set samples can be selected according to actual needs.
[0089] The following is an example of using the method of the present invention to carry out tobacco leaf group formulation imitation design.
[0090] First, we selected a recipe with hay and sweet aroma as the model. The recipe is as follows:
[0091]
[0092]
[0093] The tobacco leaf group formula is simulated and designed by the method of the present invention, and the generated formula is as follows:
[0094] Serial number Tobacco leaf name Proportion Category 1 2018-Lijiang, Yunnan-C2F 6 Light fragrance tobacco 2 2019-Ludian, Zhaotong, Yunnan-C2F 8 Light fragrance tobacco 3 2016-Lijiang, Yunnan-C1F 4 Light fragrance tobacco 4 2017-Yunnan Baoshan-C2F 8 Light fragrance tobacco 5 2017-Yunnan Pu'er Zhenyuan-C3F 2 Light fragrance tobacco 6 2017-Dali, Yunnan-C2F 5 Light fragrance tobacco 7 2016-Chuxiong, Yunnan-C1F 2 Light fragrance tobacco 8 2017-Changde, Hunan-C4F 3 Strong aroma tobacco 9 2017-Chongqing Wushan-C3F 2 Medium-flavor tobacco leaves 10 2017-Yunnan Baoshan-C3F 2 Light fragrance tobacco 11 2018-C3F, Jiangyong, Yongzhou, Hunan 2 Strong aroma tobacco 12 2016-Sichuan Liangshan-X2F 2 Light fragrance tobacco 13 2018-Longshan, Xiangxi, Hunan-C2F 4 Strong aroma tobacco 14 2016-Fujian Sanming-C4F 5 Light fragrance tobacco 15 2017-Sichuan Liangshan-C4F 5 Light fragrance tobacco 16 2019-Zimbabwe Tianze-L1OA 10 Imported tobacco leaves 17 2019-Zimbabwe Tianze-L1OC 10 Imported tobacco leaves 18 Filler 20 total 100
[0095] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0096] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A method for designing a tobacco leaf group formula based on tobacco leaf substitution, characterized in that: include: Construct a tobacco leaf data information database, which includes the production area, grade, classification, near-infrared spectrum and conventional chemical composition data of each tobacco leaf; Based on the tobacco leaf data information database, linear discriminant analysis classification training is performed to obtain the tobacco leaf class label prediction model; Select a target leaf group formula Q to be imitated, and use the tobacco leaf category prediction model to predict the category label of each tobacco leaf in Q to obtain the corresponding category label C; In the tobacco leaf data information database, the infrared spectrum matrix of each tobacco leaf labeled C is used to calculate the Euclidean distance with the near-infrared spectrum matrix of the tobacco leaf in Q. The Euclidean distances are sorted from small to large, and the two tobacco leaves corresponding to the first two digits are selected as candidate replacements for the tobacco leaf in Q. They are recorded as the best replacement and the second best replacement, respectively. For each tobacco leaf in the target leaf group formula Q to be imitated, the corresponding optimal alternative tobacco leaf is selected for replacement. If the optimal alternative tobacco leaf is also a component in Q, the suboptimal alternative tobacco leaf is selected for replacement to complete the generation of formula P. The conventional chemical composition vectors of the formula P and the target leaf group formula Q to be imitated are calculated based on linear weighting and the proportions are adjusted to minimize the difference in conventional chemical composition between the formula P and the target leaf group formula Q to be imitated, thereby achieving the proportion optimization of the formula P; The process of achieving the ratio optimization of the formula P includes: For each tobacco leaf component in recipe P, calculate the Euclidean distance d between its conventional chemical composition vector and the conventional chemical composition vector of the target leaf group recipe Q to be imitated, and multiply the Euclidean distance d by the initial value of the ratio of the tobacco leaf component to obtain the influence factor r; Sort by the size of the influencing factor r, and adjust the proportion of the tobacco leaf component with the largest influencing factor r so that the difference in conventional chemical composition between the adjusted formula P and the target leaf group formula Q to be imitated is minimized; Recalculate the impact factor r of each tobacco component in the adjusted formula P and re-rank them; If the reordered order is consistent with the previous order, the ratio optimization ends; otherwise, the above operation is repeated.
2. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to claim 1, characterized in that: After achieving the ratio optimization of formula P, it also includes: When the ratio of the formula P is optimized and the sum of the ratios of the tobacco leaf components is not equal to 100, it can be controlled by adjusting the proportion of the filler.
3. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to claim 2, characterized in that: The filler includes cut stems, thin sheets, puffed tobacco or single-ingredient tobacco suitable for use as filler.
4. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to claim 1, characterized in that: In the process of generating the recipe P, if the best alternative tobacco leaf and the second best alternative tobacco leaf are both components in the target leaf group recipe Q to be imitated, then according to the tobacco leaf type and grade of the tobacco leaf to be replaced in the target leaf group recipe Q to be imitated, a tobacco leaf that is consistent with the tobacco leaf type and grade of the tobacco leaf to be replaced and is not in the target leaf group recipe Q to be imitated is selected from the tobacco leaf data information database as the alternative tobacco leaf; If there is no tobacco leaf that meets the conditions in the tobacco leaf data information database, the replacement of the tobacco leaf to be replaced will be abandoned.
5. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to claim 4, characterized in that: The tobacco leaf types include light-flavor tobacco leaves, medium-flavor tobacco leaves, strong-flavor tobacco leaves, and imported tobacco leaves; Before generating the recipe P, it also includes: Conduct sensory evaluation on all tobacco leaves in the tobacco leaf database, and score the four aroma tones: fresh sweet aroma, honey sweet aroma, mellow sweet aroma, and burnt sweet aroma. Tobacco leaf types are classified according to the scores of the four aroma flavors: in the evaluation results, the tobacco leaf with the highest score for light sweet aroma is defined as light aroma tobacco leaf; in the evaluation results, the tobacco leaf with the highest score for burnt sweet aroma is defined as strong aroma tobacco leaf; in the evaluation results, the tobacco leaf with the highest score for honey sweet aroma or mellow sweet aroma is defined as medium aroma tobacco leaf; tobacco leaves produced in foreign countries are defined as imported tobacco leaves.
6. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to any one of claims 1 to 5, characterized in that: The linear discriminant analysis classification training is performed based on the tobacco leaf data information database to obtain a tobacco leaf class label prediction model, specifically including: Randomly select n near-infrared spectra of tobacco leaves of various categories and their corresponding class labels from the tobacco leaf data information database as a training set; Based on the near-infrared spectra in the training set, a near-infrared spectrum matrix with a dimension of n×p is constructed, where p is the dimension of the near-infrared spectrum; The near-infrared spectrum matrix is normalized to obtain a standardized near-infrared spectrum matrix X; Compute the covariance matrix of the normalized near-infrared spectral matrix: Calculate the eigenvalues and corresponding eigenvectors of the covariance matrix; Sort the eigenvalues and their corresponding eigenvectors in descending order, and take the eigenvectors corresponding to the first k eigenvalues to form a dimensionality reduction matrix W with a dimension of p×k; Use trainX=X*W as training data and the corresponding class labels as parameters to call Matlab's ClassificationDiscriminant.fit function to perform linear discriminant analysis classification training and obtain the tobacco leaf class label prediction model; When using the tobacco leaf class label prediction model to predict the class label of each tobacco leaf in the target leaf group formula Q to be imitated, the near-infrared spectrum of each tobacco leaf in the target leaf group formula Q to be imitated is obtained and annotated to obtain testX, and the predict function of Matlab is called to obtain the predicted class label corresponding to each tobacco leaf.
7. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to claim 6, characterized in that: The near-infrared spectrum matrix is normalized to obtain a normalized near-infrared spectrum matrix X, and the process includes: The near-infrared spectrum matrix is Z-Score normalized, and each element is normalized as follows: in, is the standard deviation of the j-th variable, represents the average value of the jth variable, where the variable is absorbance, x ij Represents the element in the i-th row and j-th column of the near-infrared spectrum matrix; The process of calculating the covariance matrix of the standardized near-infrared spectrum matrix includes: Compute the mean vector of the normalized near-infrared spectral matrix: Among them, x i Represents the vector consisting of the i-th row of the normalized near-infrared spectrum matrix; The covariance matrix S of the standardized near-infrared spectral matrix is calculated by the following formula:
8. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to claim 1, characterized in that: The classification mark includes an integer part and a decimal part; the integer part represents the production area of the tobacco leaves; the decimal part ranges from 1 to 3, corresponding to the upper leaves, middle leaves and lower leaves in the grade of tobacco leaves respectively; the form of tobacco leaves includes tobacco leaves, tobacco leaves or tobacco powder.
9. The method for designing tobacco leaf group recipes based on tobacco leaf substitution according to claim 1, characterized in that: The near infrared spectrum of a single tobacco leaf was collected by a near infrared spectrometer. The spectral range of the near infrared spectrometer is [12800 cm -1 , 3600cm -1 ] or [780nm, 2778nm]; the scanning speed range is: [1 round / second, 64 rounds / second], the number of scans per round range is: [1, 128], and the resolution range is: [2cm -1 , 64cm -1 ], the total number of collected data points ranges from: [1, 2592].
Citation Information
Patent Citations
Intelligent tobacco formulation method
CN102488309A
Alternative method for tobacco leaf and cigarette leaf group formula based on near infrared spectrum
CN109975238A
Cigarette mainstream smoke quality evaluation method
CN112881323A