Flue gas nicotine content prediction method based on LASSO and multiple linear regression

By combining LASSO with multiple linear regression, important variables were screened and a smoke nicotine content prediction model was established, which solved the problems of insufficient accuracy and generalization performance of nicotine content prediction in the existing technology, achieved efficient and accurate nicotine content prediction, and provided a scientific basis for cigarette formula design.

CN120656578APending Publication Date: 2025-09-16CHINA TOBACCO HENAN IND CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410293195.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

The accuracy and generalization performance of existing tobacco leaf nicotine content prediction models are insufficient, and traditional methods require a lot of manual intervention and prior knowledge, which increases the difficulty and cost of modeling.

Method used

The LASSO regression algorithm was used to screen important variables, and a multiple linear regression algorithm was combined to establish a prediction model for nicotine content in smoke. Representative samples were selected using the Kennard-Stone algorithm. The process parameters were controlled during the rolling process, and the chemical composition was determined using multiple detection methods to establish an accurate prediction model.

Benefits of technology

It improves the accuracy and efficiency of predicting nicotine content in tobacco leaves, provides an objective scientific basis, provides support for digital cigarette formula design, and reduces the cost of manual intervention and modeling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656578A_ABST
    Figure CN120656578A_ABST
Patent Text Reader

Abstract

The invention discloses a smoke nicotine content prediction method based on LASSO and multiple linear regression. The smoke nicotine content prediction method comprises the following steps: screening representative samples from all stock tobacco leaf samples; rolling the representative sample to obtain a rolled single tobacco sample; determining the nicotine content and the chemical component content of the rolled single-material tobacco sample; screening out a plurality of important variables from the nicotine content and the chemical component content by adopting an LASSO regression algorithm; establishing a nicotine prediction model by using the important variables and adopting a multiple linear regression algorithm; and predicting the nicotine content according to the nicotine prediction model. According to the smoke nicotine content prediction method based on LASSO and multiple linear regression, the nicotine content and chemical components of a representative sample are measured, an LASSO regression algorithm is used for screening important variables, a multiple linear regression algorithm is used for establishing a prediction model, an objective and accurate means is provided for prediction of the nicotine content of tobacco leaves, and the method is suitable for popularization and application. And a scientific basis is provided for digital formula design of cigarettes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of tobacco analysis, and in particular to a method for predicting nicotine content in smoke based on LASSO and multiple linear regression. Background Art

[0002] The nicotine in smoke comes from the tobacco leaves. During combustion, the nicotine in the tobacco leaves is vaporized and transferred into the smoke. Therefore, predicting the nicotine content in smoke from tobacco leaves is of great significance for cigarette formulation design.

[0003] Although numerous studies have attempted to develop predictive models for tobacco leaf nicotine content, the accuracy and generalization performance of these models are often limited by factors such as uneven data quality and imperfect feature selection. Furthermore, traditional modeling methods often require extensive manual intervention and prior knowledge, which undoubtedly increases the difficulty and cost of modeling.

[0004] Therefore, there is an urgent need for a method to predict nicotine content in smoke based on LASSO and multiple linear regression. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for predicting nicotine content in tobacco smoke based on LASSO and multiple linear regression to solve the problems in the above-mentioned prior art. The method can use the LASSO regression algorithm to screen important variables and use the multiple linear regression algorithm to establish a prediction model, thereby providing an objective and accurate means for predicting the nicotine content in tobacco leaves and providing a scientific basis for the digital formulation design of cigarettes.

[0006] The present invention provides a method for predicting nicotine content in smoke based on LASSO and multiple linear regression, which includes:

[0007] Select several representative samples from all tobacco leaf samples in stock;

[0008] Rolling each of the selected representative samples to obtain a number of rolled single-ingredient cigarette samples; and determining the nicotine content and chemical component content of the rolled single-ingredient cigarette samples corresponding to each of the representative samples;

[0009] LASSO regression algorithm was used to screen out several important variables from the nicotine content and chemical component content corresponding to each representative sample;

[0010] Using the selected important variables, a multiple linear regression algorithm is used to establish a nicotine prediction model;

[0011] The nicotine content in the smoke is predicted according to the nicotine prediction model.

[0012] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, the step of screening out several representative samples from all the tobacco leaf samples in stock specifically includes:

[0013] Conduct near infrared detection on all tobacco leaves in stock to obtain near infrared spectra of all tobacco leaves in stock;

[0014] Based on the near-infrared spectra of all stock tobacco leaves, the Kennard-Stone algorithm was used to screen out several representative samples from all stock tobacco leaf samples.

[0015] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, the method uses the Kennard-Stone algorithm based on the near-infrared spectra of all stock tobacco leaves to screen out several representative samples from all stock tobacco leaf samples, specifically including:

[0016] All near-infrared spectral samples of tobacco leaves in stock are regarded as candidate samples of the training set, and samples are selected from the candidate samples to enter the training set in turn;

[0017] Among all candidate samples, select the two samples with the farthest Euclidean distance to enter the training set;

[0018] For each remaining sample, calculate the Euclidean distance from the sample to each known sample in the training set;

[0019] Find the two samples closest to the selected samples in the training set, and select the one with the farther distance between them into the training set;

[0020] Repeat the steps of selecting samples to enter the training set until the number of samples in the training set reaches the preset requirement.

[0021] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, rolling each of the selected representative samples to obtain a plurality of rolled single-ingredient cigarette samples specifically includes:

[0022] Set the rolling processing parameters: drum speed 8r / min-12r / min, hot air temperature 100℃-120℃, drum wall temperature 110℃-130℃, outlet moisture content 12%-13%;

[0023] The representative samples were rolled according to the rolling processing parameters, wherein the control requirements for various indicators of the cigarettes are as follows: the filter rod uses a 24.1mm*100mm*2700Pa ordinary acetate filter rod, the cigarette paper uses a 26.5mm*28.5g / m2A70CU wood pulp cross-grain cigarette paper, the tipping paper uses a 64mm hot stamping tipping paper without perforations, the draw resistance control requirement is 1000Pa-1100Pa, the circumference control requirement is 24.1mm-24.5mm, the hardness control requirement is 58%-78%, the quality control requirement for a single cigarette is 0.82g-0.92g, the quality requirement for 20 cigarettes is 16.4g-18.4g, the length control requirement is 83.5mm-84.5mm, and the mass fraction control requirement for moisture content is 11.00%-13.00%.

[0024] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, the determining of the nicotine content and chemical component content of the rolled single-ingredient cigarette sample corresponding to each representative sample specifically includes:

[0025] Determining the nicotine content of the rolled single-ingredient cigarette samples corresponding to each of the representative samples using a smoking machine method or gas chromatography;

[0026] Determine the content of conventional chemical components of the rolled single-ingredient cigarette samples corresponding to each representative sample using a continuous flow method, wherein the conventional chemical components include reducing sugars, total alkaloids, total nitrogen, chlorine, potassium, and starch;

[0027] Determine the pH value of the rolled single-ingredient cigarette sample corresponding to each representative sample using a pH meter method;

[0028] Ion chromatography is used to determine the content of inorganic anions in the rolled single-ingredient cigarette samples corresponding to each representative sample, wherein the inorganic anions include sulfate and phosphate;

[0029] Determining the content of inorganic cations in the rolled single-ingredient cigarette samples corresponding to each representative sample using atomic absorption spectrometry, wherein the inorganic cations include magnesium ions and calcium ions;

[0030] Liquid chromatography was used to determine the content of polyphenols in the rolled single-ingredient cigarette samples corresponding to the representative samples, wherein the polyphenols included neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin, and rutin;

[0031] The GC-MS / MS method is used to determine the content of polybasic acids and higher fatty acids in the rolled single-ingredient cigarette samples corresponding to each representative sample, wherein the polybasic acids include oxalic acid, malonic acid, succinic acid, malic acid, citric acid and vanillic acid, and the higher fatty acids include myristic acid, hexadecanoic acid, linoleic acid, oleic acid + linolenic acid, octadecanoic acid and eicosanoic acid;

[0032] An amino acid analyzer is used to determine the content of free amino acids in the rolled single-ingredient cigarette samples corresponding to each representative sample, wherein the free amino acids include aspartic acid, threonine, serine, asparagine, glutamic acid, glutamine, glycine, alanine, valine, cystine, methionine, isoleucine, leucine, tyrosine, phenylalanine, 4-aminobutyric acid, lysine, histidine, tryptophan, arginine, and proline;

[0033] The content of Amadori compounds in the rolled single-ingredient cigarette samples corresponding to each representative sample was determined by HPLC-MS / MS, wherein the Amadori compounds include Glu-An, Fru-Amb, Fru-His, Fru-Pro, Fru-Val, Fru-Thr, Fru-Gly, Fru-Ala, Fru-Asn, Fru-Asp, Fru-Gln, Fru-Glu, Fru-Ile, Fru-Leu, Fru-Tyr, Fru-Phe and Fru-Trp;

[0034] Determine the content of dichloromethane extract of the rolled single-ingredient cigarette sample corresponding to each representative sample by using a filtration method;

[0035] Determining the solanesol content of the rolled single-ingredient tobacco samples corresponding to each representative sample using high performance liquid chromatography;

[0036] The content of neophytadiene in the rolled single-ingredient cigarette sample corresponding to each representative sample was determined by back-flushing gas chromatography.

[0037] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, the LASSO regression algorithm is used to screen out several important variables from the nicotine content and chemical component content corresponding to each representative sample, specifically including:

[0038] The objective function of the linear regression model for nicotine content is:

[0039] J(w)=∑(y-Vw) 2 +λ||w||1=∑(y-Vw) 2 +∑λ|w|=ESS(w)+λl1(w) (1)

[0040] Where y represents the nicotine content, which is the dependent variable, V=[v1,…,v p ] represents the content of all chemical composition indicators, which is the independent variable, p represents the number of chemical composition indicators, w=[w1,…,w p ] Trepresents the regression coefficient corresponding to each independent variable, λ||w||1 represents the penalty term of the L1 norm, ESS(w) represents the sum of squared errors of n groups of observation data corresponding to n representative samples, n represents the number of representative samples screened, and λl1(w) represents the penalty term;

[0041] Using the coordinate descent method, control the other regression parameters w unchanged, and for one of the w in the objective function J(w) q Find the partial derivative, q = 1, 2, ..., p, and so on,

[0042] Find the partial derivatives of the remaining regression parameters, and finally make each derivative function equal to 0, so as to obtain the objective function that reaches the global minimum.

[0043] at this time,

[0044] Taking partial derivative of formula (2), we get:

[0045]

[0046] in,

[0047] Since the penalty term is not differentiable, the following derivatives are used:

[0048]

[0049] Let the sum of the two partial derivatives of formula (3) and formula (4) be equal to 0, and we can get:

[0050]

[0051] From this we can get:

[0052]

[0053] All regression coefficients w j Remove the independent variables that are 0 and retain all regression coefficients w j Independent variables that are not 0 are considered as important variables that are screened out.

[0054] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, the method of using the LASSO regression algorithm to screen out several important variables from the nicotine content and chemical component content corresponding to each representative sample further includes:

[0055] The cross-validation method is used to determine the value of the penalty coefficient λ, which includes:

[0056] The dataset is divided into several parts, one of which is selected as the validation set each time, and the rest are used as the training set. The optimal λ value is selected by adjusting the value of λ and calculating the performance index of the model on the validation set.

[0057] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, the multiple linear regression algorithm is used to establish a nicotine prediction model using the selected important variables, which specifically includes:

[0058] The multiple linear regression analysis model is established as follows:

[0059]

[0060] Among them, y represents the dependent variable nicotine content, x1,…,x m Indicates the important variables screened out, β0, β1…, β m ,σ 2 are all related to x1,…,x m Unrelated unknown parameters, β0, β1…, β m is called the regression coefficient;

[0061] Obtain n independent observation data corresponding to n representative samples [b i ,a i1 ,…,a im ], where b i represents the observed value of y, [a i1 ,…,a im ] are x1,…,x m Observation value, i=1,…,n,n>m, according to formula (7):

[0062]

[0063] remember:

[0064]

[0065] ε=[ε1,…,ε n ] T , β=[β1,…,β m ] T ,

[0066] Formula (7) can be expressed as:

[0067]

[0068] Among them, E n represents the n-order identity matrix;

[0069] The least square method is used to calculate the parameters β0, β1…, βm Make an estimate, that is, choose the estimated value Make When , the error sum of squares is:

[0070]

[0071] Reach minimum;

[0072] To this end, let:

[0073]

[0074] After being sorted out and transformed into a normal system of equations, the matrix form is:

[0075] X T Xβ=X T Y (13)

[0076] When the matrix X has full column rank, X T X is a reversible matrix, and the solution of formula (12) is

[0077]

[0078] Will Substituting back into formula (7), we can get the estimated value of y, which is

[0079]

[0080] Formula (15) is the obtained multiple linear regression model.

[0081] In the above-mentioned method for predicting nicotine content in smoke based on LASSO and multiple linear regression, preferably, predicting the nicotine content in smoke according to the nicotine prediction model specifically includes:

[0082] Detecting the content of important variables of the rolled single-ingredient cigarette sample corresponding to the sample to be tested;

[0083] The contents of important variables of the sample to be tested are input into the nicotine prediction model to obtain the nicotine content in the smoke of the sample to be tested.

[0084] The method for predicting nicotine content in smoke based on LASSO and multiple linear regression as described above, wherein preferably, the method for predicting nicotine content in smoke based on LASSO and multiple linear regression further includes:

[0085] The prediction effect of the nicotine prediction model is evaluated, specifically including:

[0086] The prediction effect of the nicotine prediction model is evaluated using statistical indicators, wherein the statistical indicators include: at least one of mean square error, root mean square error and determination coefficient. The closer the determination coefficient is to 1, the smaller the mean square error and root mean square error are, indicating that the prediction accuracy of the nicotine prediction model is better.

[0087] The present invention provides a method for predicting nicotine content in tobacco gas based on LASSO and multiple linear regression. The method comprises the following steps: representative tobacco leaf samples are screened from all tobacco leaves in stock, and the screened tobacco leaves are rolled according to process requirements. The nicotine content and chemical composition of the rolled representative samples are then measured. The LASSO regression algorithm is then used to screen important variables from training set sample data, and a prediction model is established using the multiple linear regression algorithm. This method provides an objective and accurate means for predicting the nicotine content in tobacco leaves and a scientific basis for digital cigarette formulation design. The Kennard-Stone algorithm takes into account the diversity and representativeness of the samples, thereby ensuring that the selected samples have a wider range of quality and chemical composition distribution, thereby improving the efficiency and accuracy of modeling. The LASSO regression algorithm can screen out chemical components that have a greater impact on nicotine content as important variables, and has the advantages of controlling multicollinearity and reducing data noise, thereby helping to improve the accuracy and reliability of the prediction model. The prediction model is established using the multiple linear regression algorithm, thereby achieving accurate prediction of nicotine content. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be further described below with reference to the accompanying drawings, in which:

[0089] Figure 1 A flowchart of an embodiment of a method for predicting nicotine content in smoke based on LASSO and multiple linear regression provided by the present invention;

[0090] Figure 2 Schematic diagram of the cross-validated LASSO fitting MSE;

[0091] Figure 3 It is the LASSO fitting coefficient trajectory diagram;

[0092] Figure 4 Schematic diagram comparing the actual and predicted values ​​of nicotine. DETAILED DESCRIPTION

[0093] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. The description of the exemplary embodiments is merely illustrative and is in no way intended to limit the present disclosure, its application, or use. The present disclosure can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to make the present disclosure thorough and complete and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that unless otherwise specifically stated, the relative arrangement of parts and steps, the composition of materials, numerical expressions, and numerical values ​​set forth in these embodiments should be interpreted as being merely exemplary and not as limiting.

[0094] The terms "first," "second," and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are simply used to distinguish different parts. Terms such as "include" or "comprising" mean that the elements preceding the term include the elements listed after the term, and do not exclude the possibility of also including other elements. Terms such as "upper," "lower," and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.

[0095] In the present disclosure, when a specific component is described as being located between a first component and a second component, there may or may not be an intervening component between the specific component and the first component or the second component. When a specific component is described as being connected to another component, the specific component may be directly connected to the other component without an intervening component, or may not be directly connected to the other component but have an intervening component.

[0096] All terms (including technical or scientific terms) used in this disclosure have the same meaning as those understood by one of ordinary skill in the art to which this disclosure belongs, unless otherwise specifically defined. It should also be understood that terms defined in, for example, general dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or highly formal sense, unless explicitly defined herein.

[0097] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0098] like Figure 1 As shown, the method for predicting nicotine content in smoke based on LASSO and multiple linear regression provided in this embodiment includes the following steps in actual implementation:

[0099] Step S1: Select several representative samples from all tobacco leaf samples in stock.

[0100] In the present invention, all tobacco leaves in stock are scanned using near-infrared spectra. Based on the near-infrared spectra of all tobacco leaves in stock, the Kennard-Stone algorithm is used to select several representative samples from all tobacco leaves in stock for constructing a nicotine prediction model. In one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression, step S1 may specifically include:

[0101] Step S11: Perform near-infrared detection on all tobacco leaves in stock to obtain near-infrared spectra of all tobacco leaves in stock.

[0102] Step S12: Based on the near-infrared spectra of all the tobacco leaves in stock, a Kennard-Stone algorithm is used to screen out several representative samples from all the tobacco leaves in stock.

[0103] The Kennard-Stone algorithm is an algorithm for selecting samples. It can select a set of the most representative samples from a given data set. Its basic concept is to select samples by minimizing the distance between selected samples to ensure that the selected samples are as representative of the original data set as possible. In one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression of the present invention, step S12 may specifically include:

[0104] Step S121: All near-infrared spectrum samples of tobacco leaves in stock are regarded as candidate samples of the training set, and samples are selected from the candidate samples in turn to enter the training set.

[0105] Step S122: Among all candidate samples, select the two samples with the farthest Euclidean distance to enter the training set.

[0106] Step S123: For each remaining sample, calculate the Euclidean distance between the sample and each known sample in the training set.

[0107] Step S124: Find the two samples closest to the selected samples in the training set, and select the one with the farther distance between the two samples into the training set.

[0108] Step S125: Repeat the step of selecting samples to enter the training set until the number of samples in the training set reaches the preset requirement.

[0109] In one embodiment of the present invention, near-infrared spectra of all stock tobacco leaves are scanned, and the Kennard-Stone algorithm is used to screen out 120 representative samples from all stock tobacco leaf samples, covering multiple domestic and foreign production areas from 2013 to 2020, with nicotine content per unit tobacco ranging from 0.3 mg / g to 2.5 mg / g, for the construction of a nicotine prediction model.

[0110] Step S2: rolling the selected representative samples to obtain a number of rolled single-ingredient cigarette samples.

[0111] In one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression of the present invention, step S2 may specifically include:

[0112] Step S21, set the rolling processing parameters: drum speed is 8r / min-12r / min (for example, 10r / min), hot air temperature is 100℃-120℃ (for example, 110℃), drum wall temperature is 110℃-130℃ (for example, 120℃), outlet moisture content is 12%-13% (for example, 12.5%).

[0113] Step S22, rolling each of the representative samples according to the rolling processing parameters, wherein the control requirements for various indicators of the cigarette are as follows: the filter rod uses a 24.1mm*100mm*2700Pa ordinary acetate filter rod, the cigarette paper uses a 26.5mm*28.5g / m2A70CU wood pulp cross-grain cigarette paper, the tipping paper uses a 64mm hot stamping tipping paper without perforations, the draw resistance control requirement is 1000Pa-1100Pa, the circumference control requirement is 24.1mm-24.5mm, the hardness control requirement is 58%-78%, the single cigarette quality control requirement is 0.82g-0.92g, the 20 cigarette quality requirement is 16.4g-18.4g, the length control requirement is 83.5mm-84.5mm, and the moisture content mass fraction control requirement is 11.00%-13.00%.

[0114] Step S3: Determine the nicotine content and chemical component content of the rolled single-ingredient cigarette sample corresponding to each representative sample.

[0115] In one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression of the present invention, step S3 may specifically include:

[0116] Step S31: Determine the nicotine content of the rolled single-ingredient cigarette samples corresponding to each representative sample using a smoking machine method or gas chromatography.

[0117] Specifically, the nicotine content of the rolled single-ingredient cigarettes corresponding to the representative tobacco leaf samples was determined with reference to the relevant national standards "GB / T 19609-2004 Determination of total particulate matter and tar in cigarettes using a conventional analytical smoking machine" and "GB / T 23355-2009 Determination of nicotine in total particulate matter in cigarettes by gas chromatography."

[0118] Step S32: using a continuous flow method to determine the content of conventional chemical components of the rolled single-ingredient cigarette samples corresponding to each of the representative samples, wherein the conventional chemical components include reducing sugars, total alkaloids, total nitrogen, chlorine, potassium and starch.

[0119] Specifically, the determination of conventional chemical components (total alkaloids, reducing sugars, total sugars, total nitrogen, potassium, chlorine, and starch) refers to "YC / T 159-2002 Tobacco and Tobacco Products - Determination of Water-soluble Sugars - Continuous Flow Method", "YC / T 160-2002 Tobacco and Tobacco Products - Determination of Total Alkaloids - Continuous Flow Method", "YC / T 161-2002 Tobacco and Tobacco Products - Determination of Total Nitrogen - Continuous Flow Method", "YC / T 162-2011 Tobacco and Tobacco Products - Determination of Chlorine - Continuous Flow Method", "YC / T217-2007 Tobacco and Tobacco Products - Determination of Potassium - Continuous Flow Method", and "YC / T 216-2007 Tobacco and Tobacco Products - Determination of Starch - Continuous Flow Method".

[0120] Step S33: using a pH meter method to measure the pH value of the rolled single-ingredient cigarette sample corresponding to each representative sample.

[0121] Specifically, the pH value is determined with reference to "YC / T 222-2007 Tobacco and Tobacco Products - Determination of Tobacco pH Value".

[0122] Step S34: using ion chromatography to determine the content of inorganic anions in the rolled single-ingredient cigarette samples corresponding to the representative samples, wherein the inorganic anions include sulfate and phosphate.

[0123] Specifically, the determination of inorganic anions (sulfate, phosphate) refers to "YC / T 248-2008 Tobacco and tobacco products - Determination of inorganic anions - Ion chromatography method".

[0124] Step S35: using atomic absorption spectrometry to determine the content of inorganic cations in the rolled single-ingredient cigarette samples corresponding to the representative samples, wherein the inorganic cations include magnesium ions and calcium ions.

[0125] Specifically, the determination of inorganic cations (Mg, Ca) refers to "YC / T 174-2003 Determination of Calcium in Tobacco and Tobacco Products - Atomic Absorption Spectrometry" and "YC / T 175-2003 Determination of Magnesium in Tobacco and Tobacco Products - Atomic Absorption Spectrometry".

[0126] Step S36: using a liquid chromatograph to determine the content of polyphenols in the rolled single-ingredient cigarette samples corresponding to the representative samples, wherein the polyphenols include neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin and rutin.

[0127] Specifically, the determination of polyphenols (neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin, and rutin) refers to "YC / T 202-2006 Determination of polyphenol compounds chlorogenic acid, scopoletin, and rutin in tobacco and tobacco products."

[0128] Step S37, using GC-MS / MS to determine the content of polyacids and higher fatty acids in the rolled single-ingredient cigarette samples corresponding to each of the representative samples, wherein the polyacids include oxalic acid, malonic acid, succinic acid, malic acid, citric acid and vanillic acid, and the higher fatty acids include myristic acid, hexadecanoic acid, linoleic acid, oleic acid + linolenic acid, octadecanoic acid and eicosanoic acid.

[0129] Specifically, the determination of polybasic acids (oxalic acid, malonic acid, succinic acid, malic acid, citric acid, vanillic acid) and higher fatty acids (tetradecanoic acid, hexadecanoic acid, linoleic acid, oleic acid + linolenic acid, octadecanoic acid, eicosanoic acid) was carried out with reference to the paper "Simultaneous Analysis of 42 Organic Acids in Tobacco Leaves by GC-MS / MS Method".

[0130] Step S38: Determine the free amino acid content of the rolled single-ingredient cigarette sample corresponding to each of the representative samples using an amino acid analyzer, wherein the free amino acids include aspartic acid, threonine, serine, asparagine, glutamic acid, glutamine, glycine, alanine, valine, cystine, methionine, isoleucine, leucine, tyrosine, phenylalanine, 4-aminobutyric acid, lysine, histidine, tryptophan, arginine, and proline.

[0131] Specifically, the determination of free amino acids (aspartic acid, threonine, serine, asparagine, glutamic acid, glutamine, glycine, alanine, valine, cystine, methionine, isoleucine, leucine, tyrosine, phenylalanine, 4-aminobutyric acid, lysine, histidine, tryptophan, arginine, and proline) was performed according to the method described in YC / T 282-2009 Determination of Free Amino Acids in Tobacco Leaves - Amino Acid Analyzer Method.

[0132] Step S39: Determine the content of Amadori compounds in the rolled single-ingredient cigarette samples corresponding to each of the representative samples using HPLC-MS / MS, wherein the Amadori compounds include Glu-An, Fru-Amb, Fru-His, Fru-Pro, Fru-Val, Fru-Thr, Fru-Gly, Fru-Ala, Fru-Asn, Fru-Asp, Fru-Gln, Fru-Glu, Fru-Ile, Fru-Leu, Fru-Tyr, Fru-Phe, and Fru-Trp.

[0133] Specifically, the determination of Amadori compounds (Glu-An, Fru-Amb, Fru-His, Fru-Pro, Fru-Val, Fru-Thr, Fru-Gly, Fru-Ala, Fru-Asn, Fru-Asp, Fru-Gln, Fru-Glu, Fru-Ile, Fru-Leu, Fru-Tyr, Fru-Phe, Fru-Trp) was carried out with reference to the paper "Simultaneous Determination of 22 Amadori Compounds in Tobacco by HPLC-MS / MS".

[0134] Step S310: Determine the content of dichloromethane extract of the rolled single-ingredient cigarette sample corresponding to each representative sample by using a filtration method.

[0135] Specifically, the determination of dichloromethane extract refers to the patent "A filtration device for determining the content of dichloromethane extract of tobacco".

[0136] Step S311: using high performance liquid chromatography to determine the solanesol content of the rolled single-ingredient cigarette sample corresponding to each representative sample.

[0137] The determination of solanesol refers to "GB / T 31758-2015 Determination of solanesol in tobacco leaves and tobacco leaf extracts - High performance liquid chromatography method".

[0138] Step S312: using backflushing gas chromatography to determine the content of neophytadiene in the rolled single-ingredient cigarette sample corresponding to each representative sample.

[0139] Specifically, the determination of phytadiene refers to the paper "Detection and Analysis of the Content of Phytadiene in Tobacco Leaves from Different Ecological Zones by Backflush Gas Chromatography".

[0140] Step S4: using the LASSO regression algorithm to screen out several important variables from the nicotine content and chemical component content corresponding to each of the representative samples.

[0141] The selected important variables are used in the subsequent process of establishing a nicotine prediction model. In one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression of the present invention, step S4 may specifically include:

[0142] Step S41: Establishing the objective function of the linear regression model of nicotine content:

[0143] J(w)=∑(y-Vw) 2 +λ||w||1=∑(y-Vw) 2 +∑λ|w|=ESS(w)+λl1(w) (1)

[0144] Where y represents the nicotine content, which is the dependent variable, V=[v1,…,v p ] represents the content of all chemical composition indicators, which is the independent variable, p represents the number of chemical composition indicators, w=[w1,…,w p ] T represents the regression coefficient corresponding to each independent variable, λ||w||1 represents the penalty term of the L1 norm, ESS(w) represents the sum of squared errors of n groups of observation data corresponding to n representative samples, n represents the number of representative samples screened out, and λl1(w) represents the penalty term.

[0145] In the present invention, in order to ensure that the regression coefficient w can be calculated, the LASSO regression model adds an L1 norm penalty term to the objective function, which can reduce some unimportant regression coefficients to 0, thereby achieving the effect of eliminating variables.

[0146] Step S42: Using the coordinate descent method, control the other regression parameters w to be constant, and calculate one of the w in the objective function J(w) q Find the partial derivative, q = 1, 2, ..., p, and so on,

[0147] Find the partial derivatives of the remaining regression parameters, and finally make each derivative function equal to 0, so as to obtain the objective function that reaches the global minimum.

[0148] at this time,

[0149] Taking partial derivative of formula (2), we get:

[0150]

[0151] in,

[0152] Step S43: Since the penalty term is not differentiable, the following derivatives are used:

[0153]

[0154] Step S44: Let the sum of the two partial derivatives of formula (3) and formula (4) be equal to 0, and we can get:

[0155]

[0156] From this we can get:

[0157]

[0158] Step S45: All regression coefficients w j Remove the independent variables that are 0 and retain all regression coefficients w jIndependent variables that are not 0 are considered as important variables that are screened out.

[0159] Furthermore, in one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression of the present invention, step S4 may further include:

[0160] Step S46: Determine the value of the penalty coefficient λ using a cross-validation method.

[0161] Specifically, the dataset is divided into several parts, one of which is selected as the validation set each time, and the rest are used as the training set. The optimal λ value is selected by adjusting the value of λ and calculating the performance index of the model on the validation set.

[0162] In one embodiment of the present invention, the LASSO regression algorithm is used to screen variables, and the penalty coefficient λ corresponding to the minimum mean square error (MSE) is obtained by cross-validation, as shown in FIG. Figure 2 As shown. According to the selected optimal penalty coefficient λ, the independent variables whose estimated parameters are compressed to 0 can be eliminated, and the independent variables whose coefficients are not 0 can be retained, as shown in Figure 3 As shown in the figure, 14 independent variables were screened out through the LASSO regression algorithm, namely total plant alkaloids, sulfate, cryptochlorogenic acid, scopoletin, rutin, vanillic acid, serine, isoleucine, 4-aminobutyric acid, lysine, tryptophan, Fru-Pro, Fru-Gly, and Fru-Ile.

[0163] Step S5: using the selected important variables and adopting a multiple linear regression algorithm to establish a nicotine prediction model.

[0164] In one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression of the present invention, step S5 may specifically include:

[0165] Step S51: Establish a multiple linear regression analysis model:

[0166]

[0167] Among them, y represents the dependent variable nicotine content, x1,...,x m Indicates the important variables screened out, β0, β1…, β m ,σ 2 are all related to x1,...,x m Unrelated unknown parameters, β0, β1…, β m is called the regression coefficient;

[0168] Step S52: Obtain n independent observation data corresponding to n representative samples [b i ,a i1,…,a im ], where b i represents the observed value of y, [a i1 ,…,a im ] are x1,...,x m Observation value, i=1,…,n,n>m, according to formula (7):

[0169]

[0170] remember:

[0171]

[0172] ε=[ε1,…,ε n ] T , β=[β1,…,β m ] T ,

[0173] Formula (7) can be expressed as:

[0174]

[0175] Among them, E n represents the n-order identity matrix;

[0176] Step S53: Use the least square method to calculate the parameters β0, β1, β m Make an estimate, that is, choose the estimated value Make When , the error sum of squares is:

[0177]

[0178] Reach minimum;

[0179] To this end, let:

[0180]

[0181] After being sorted out and transformed into a normal system of equations, the matrix form is:

[0182] X T Xβ=X T Y (13)

[0183] When the matrix X has full column rank, X T X is a reversible matrix, and the solution of formula (12) is

[0184]

[0185] Step S54: Substituting back into formula (7), we can get the estimated value of y, which is

[0186]

[0187] Formula (15) is the obtained multiple linear regression model.

[0188] In one embodiment of the present invention, first define nicotine content as the dependent variable y, total alkaloid content as the independent variable x1, sulfate content as the independent variable x2, cryptochlorogenic acid content as the independent variable x3, scopoletin content as x4, rutin content as x5, vanillic acid content as the independent variable x6, serine content as the independent variable x7, isoleucine content as the independent variable x8, 4-aminobutyric acid content as the independent variable x9, lysine content as the independent variable x10. 10 , tryptophan content is the independent variable x 11 , Fru-Pro content is the independent variable x 12 , Fru-Gly content is the independent variable x 13 , Fru-Ile content is the independent variable x 14 ,

[0189] The prediction model was established using the multivariate linear regression algorithm, and the regression equation obtained was:

[0190] y=0.6567x1-0.0128x2-0.1866x3-0.5513x4-0.0217x5+1.6092x6+0.0009x7+0.0367x8-0.0010x9-0.0057x 10 -0.0008x 11 +0.000007x 12 +0.0097x 13 -0.0055x 14 +0.0771

[0191] Step S6: predicting the nicotine content in the smoke according to the nicotine prediction model.

[0192] In one embodiment of the method for predicting nicotine content in smoke based on LASSO and multiple linear regression of the present invention, step S6 may specifically include:

[0193] Step S61: detecting the contents of important variables of the rolled single-ingredient cigarette sample corresponding to the sample to be tested.

[0194] Step S62: Input the contents of important variables of the sample to be tested into the nicotine prediction model to obtain the nicotine content in the smoke of the sample to be tested.

[0195] Furthermore, in one embodiment of the present invention, the method for predicting nicotine content in smoke based on LASSO and multiple linear regression further includes:

[0196] Step S7: Evaluate the prediction effect of the nicotine prediction model.

[0197] Specifically, statistical indicators are used to evaluate the prediction effect of the nicotine prediction model.

[0198] The statistical indicators include mean square error (MSE), root mean square error (RMSE) and coefficient of determination (R 2 ), the closer the determination coefficient is to 1, the smaller the mean square error and root mean square error are, indicating that the prediction accuracy of the nicotine prediction model is better.

[0199] The F of the multivariate linear regression model obtained in step S5 is 69.5571, and the significance level is p=0.0000, indicating that the model is statistically significant. 2 =0.913, MSE=0.0148, RMSE=0.1215; R 2 =0.906, MSE=0.0113, RMSE=0.1067, indicating that the overall fitting effect of the model is good.

[0200] Furthermore, in some embodiments of the present invention, a regression analysis is performed on the actual values ​​of nicotine content in the test set and the predicted values, with the actual values ​​as the independent variable x and the predicted values ​​as the dependent vector y, to obtain the regression equation: y = 0.922x + 0.127. Figure 4 , which also shows that the fitting effect of the model is good.

[0201] The embodiment of the present invention provides a method for predicting nicotine content in flue gas based on LASSO and multiple linear regression. Representative tobacco samples are screened from all stock tobacco leaves, and the screened tobacco leaves are rolled according to process requirements. The nicotine content and chemical composition of the rolled representative samples are then measured. The LASSO regression algorithm is then used to screen important variables from the training set sample data, and a prediction model is established using the multiple linear regression algorithm. This provides an objective and accurate means for predicting the nicotine content in tobacco leaves and a scientific basis for the digital formulation design of cigarettes. The Kennard-Stone algorithm takes into account the diversity and representativeness of the samples, thereby ensuring that the selected samples have a wider range of quality and chemical composition distribution, thereby improving the efficiency and accuracy of modeling. The LASSO regression algorithm can screen chemical components that have a greater impact on nicotine content as important variables, and has the advantages of controlling multicollinearity and reducing data noise, which helps to improve the accuracy and reliability of the prediction model. By using the multiple linear regression algorithm to establish a prediction model, accurate prediction of nicotine content is achieved.

[0202] Thus far, various embodiments of the present disclosure have been described in detail. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Based on the above description, those skilled in the art can fully understand how to implement the technical solutions disclosed herein.

[0203] Although some specific embodiments of the present disclosure have been described in detail through examples, those skilled in the art will understand that the above examples are for illustration only and are not intended to limit the scope of the present disclosure. Those skilled in the art will understand that the above embodiments may be modified or some technical features may be replaced with equivalents without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. A method for predicting nicotine content in smoke based on LASSO and multiple linear regression, characterized in that: include: Select several representative samples from all tobacco leaf samples in stock; Rolling each of the selected representative samples to obtain a number of rolled single-ingredient cigarette samples; Determining the nicotine content and chemical component content of the rolled single-ingredient cigarette samples corresponding to each of the representative samples; LASSO regression algorithm was used to screen out several important variables from the nicotine content and chemical component content corresponding to each representative sample; Using the selected important variables, a multiple linear regression algorithm is used to establish a nicotine prediction model; The nicotine content in the smoke is predicted according to the nicotine prediction model.

2. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 1, characterized in that: The representative samples selected from all tobacco leaf samples in stock include: Conduct near infrared detection on all tobacco leaves in stock to obtain near infrared spectra of all tobacco leaves in stock; Based on the near-infrared spectra of all stock tobacco leaves, the Kennard-Stone algorithm was used to screen out several representative samples from all stock tobacco leaf samples.

3. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 1, characterized in that: Based on the near-infrared spectra of all tobacco leaves in stock, the Kennard-Stone algorithm was used to screen out several representative samples from all tobacco leaves in stock, including: All near-infrared spectral samples of tobacco leaves in stock are regarded as candidate samples of the training set, and samples are selected from the candidate samples to enter the training set in turn; Among all candidate samples, select the two samples with the farthest Euclidean distance to enter the training set; For each remaining sample, calculate the Euclidean distance from the sample to each known sample in the training set; Find the two samples closest to the selected samples in the training set, and select the one with the farther distance between them into the training set; Repeat the steps of selecting samples to enter the training set until the number of samples in the training set reaches the preset requirement.

4. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 1, wherein: The representative samples screened out are rolled to obtain a number of rolled single-ingredient cigarette samples, specifically comprising: Set the rolling processing parameters: drum speed 8r / min-12r / min, hot air temperature 100℃-120℃, drum wall temperature 110℃-130℃, outlet moisture content 12%-13%; The representative samples were rolled according to the rolling processing parameters, wherein the control requirements for various indicators of the cigarettes are as follows: the filter rod uses a 24.1mm*100mm*2700Pa ordinary acetate filter rod, the cigarette paper uses a 26.5mm*28.5g / m2A70CU wood pulp cross-grain cigarette paper, the tipping paper uses a 64mm hot stamping tipping paper without perforations, the draw resistance control requirement is 1000Pa-1100Pa, the circumference control requirement is 24.1mm-24.5mm, the hardness control requirement is 58%-78%, the quality control requirement for a single cigarette is 0.82g-0.92g, the quality requirement for 20 cigarettes is 16.4g-18.4g, the length control requirement is 83.5mm-84.5mm, and the mass fraction control requirement for moisture content is 11.00%-13.00%.

5. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 1, characterized in that: The determining of the nicotine content and chemical component content of the rolled single-ingredient cigarette samples corresponding to each of the representative samples specifically includes: Determining the nicotine content of the rolled single-ingredient cigarette samples corresponding to each of the representative samples using a smoking machine method or gas chromatography; Determine the content of conventional chemical components of the rolled single-ingredient cigarette samples corresponding to each representative sample using a continuous flow method, wherein the conventional chemical components include reducing sugars, total alkaloids, total nitrogen, chlorine, potassium, and starch; Determine the pH value of the rolled single-ingredient cigarette sample corresponding to each representative sample using a pH meter method; Ion chromatography is used to determine the content of inorganic anions in the rolled single-ingredient cigarette samples corresponding to each representative sample, wherein the inorganic anions include sulfate and phosphate; Determining the content of inorganic cations in the rolled single-ingredient cigarette samples corresponding to each representative sample using atomic absorption spectrometry, wherein the inorganic cations include magnesium ions and calcium ions; Liquid chromatography was used to determine the content of polyphenols in the rolled single-ingredient cigarette samples corresponding to the representative samples, wherein the polyphenols included neochlorogenic acid, chlorogenic acid, cryptochlorogenic acid, scopoletin, and rutin; The GC-MS / MS method is used to determine the content of polybasic acids and higher fatty acids in the rolled single-ingredient cigarette samples corresponding to each representative sample, wherein the polybasic acids include oxalic acid, malonic acid, succinic acid, malic acid, citric acid and vanillic acid, and the higher fatty acids include myristic acid, hexadecanoic acid, linoleic acid, oleic acid + linolenic acid, octadecanoic acid and eicosanoic acid; An amino acid analyzer is used to determine the content of free amino acids in the rolled single-ingredient cigarette samples corresponding to each representative sample, wherein the free amino acids include aspartic acid, threonine, serine, asparagine, glutamic acid, glutamine, glycine, alanine, valine, cystine, methionine, isoleucine, leucine, tyrosine, phenylalanine, 4-aminobutyric acid, lysine, histidine, tryptophan, arginine, and proline; The content of Amadori compounds in the rolled single-ingredient cigarette samples corresponding to each representative sample was determined by HPLC-MS / MS, wherein the Amadori compounds include Glu-An, Fru-Amb, Fru-His, Fru-Pro, Fru-Val, Fru-Thr, Fru-Gly, Fru-Ala, Fru-Asn, Fru-Asp, Fru-Gln, Fru-Glu, Fru-Ile, Fru-Leu, Fru-Tyr, Fru-Phe and Fru-Trp; Determine the content of dichloromethane extract of the rolled single-ingredient cigarette sample corresponding to each representative sample by using a filtration method; Determining the solanesol content of the rolled single-ingredient tobacco samples corresponding to each representative sample using high performance liquid chromatography; The content of neophytadiene in the rolled single-ingredient cigarette sample corresponding to each representative sample was determined by back-flushing gas chromatography.

6. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 1, characterized in that: The LASSO regression algorithm is used to screen out several important variables from the nicotine content and chemical component content corresponding to each representative sample, specifically including: The objective function of the linear regression model for nicotine content is: Where y represents the nicotine content, which is the dependent variable, V=[v1,…,v p ] represents the content of all chemical composition indicators, which is the independent variable, p represents the number of chemical composition indicators, w=[w1,…,w p ] T represents the regression coefficient corresponding to each independent variable, λ||w||1 represents the penalty term of the L1 norm, ESS(w) represents the sum of squared errors of n groups of observation data corresponding to n representative samples, n represents the number of representative samples screened, and λl1(w) represents the penalty term; Using the coordinate descent method, control the other regression parameters w unchanged, and for one of the w in the objective function J(w) q Find the partial derivative, q = 1, 2, ..., p, and so on, Find the partial derivatives of the remaining regression parameters, and finally make each derivative function equal to 0, so as to obtain the objective function that reaches the global minimum. at this time, Taking partial derivative of formula (2), we get: in, Since the penalty term is not differentiable, the following derivatives are used: Let the sum of the two partial derivatives of formula (3) and formula (4) be equal to 0, and we can get: From this we can get: All regression coefficients w j Remove the independent variables that are 0 and retain all regression coefficients w j Independent variables that are not 0 are considered as important variables that are screened out.

7. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 6, characterized in that: The LASSO regression algorithm is used to screen out several important variables from the nicotine content and chemical component content corresponding to each representative sample, and further includes: The cross-validation method is used to determine the value of the penalty coefficient λ, which includes: The dataset is divided into several parts, one of which is selected as the validation set each time, and the rest are used as the training set. The optimal λ value is selected by adjusting the value of λ and calculating the performance index of the model on the validation set.

8. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 1, characterized in that: The nicotine prediction model is established by using the screened important variables and adopting a multiple linear regression algorithm, specifically including: The multiple linear regression analysis model is established as follows: Among them, y represents the dependent variable nicotine content, x1,…,x m Indicates the important variables screened out, β0, β1…, β m ,σ 2 are all related to x1,…,x m Unrelated unknown parameters, β0, β1…, β m is called the regression coefficient; Obtain n independent observation data corresponding to n representative samples [b i ,a i1 ,…,a im ], where b i represents the observed value of y, [a i1 ,…,a im ] are x1,…,x m Observation value, i=1,…,n,n>m, according to formula (7): remember: ε=[ε1,…,ε n ] T ,β=[β1,…,β m ] T , Formula (7) can be expressed as: Among them, E n represents the n-order identity matrix; The least square method is used to calculate the parameters β0, β1…, β m Make an estimate, that is, choose the estimated value Make When , the sum of squared errors is: Reach minimum; To this end, let: After being sorted out and transformed into a normal system of equations, the matrix form is: X T Xβ=X T Y (13) When the matrix X has full column rank, X T X is a reversible matrix, and the solution of formula (12) is Will Substituting back into formula (7), we can get the estimated value of y, which is Formula (15) is the obtained multiple linear regression model.

9. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 8, characterized in that: The predicting of the nicotine content in the smoke according to the nicotine prediction model specifically includes: Detecting the content of important variables of the rolled single-ingredient cigarette sample corresponding to the sample to be tested; The contents of important variables of the sample to be tested are input into the nicotine prediction model to obtain the nicotine content in the smoke of the sample to be tested.

10. The method for predicting nicotine content in smoke based on LASSO and multiple linear regression according to claim 1, characterized in that: The method for predicting nicotine content in smoke based on LASSO and multiple linear regression also includes: The prediction effect of the nicotine prediction model is evaluated, specifically including: The prediction effect of the nicotine prediction model is evaluated using statistical indicators, wherein the statistical indicators include: at least one of mean square error, root mean square error and determination coefficient. The closer the determination coefficient is to 1, the smaller the mean square error and root mean square error are, indicating that the prediction accuracy of the nicotine prediction model is better.