Prediction model training method, prediction method and prediction system for predicting chemical component content of tobacco leaves

By constructing joint variables and using causal discovery methods to obtain causal indicators, establishing a deep learning model and using causal weights to predict, the problem of low prediction accuracy of chemical composition content in the existing technology is solved, and higher prediction accuracy and resource utilization efficiency are achieved.

CN119943185APending Publication Date: 2025-05-06ZHENGZHOU TOBACCO RES INST OF CNTC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411771845.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art predicts the chemical composition content of tobacco leaves with relatively low accuracy, which affects relationship mining and component prediction, resulting in waste of computing resources and reduced prediction accuracy.

Method used

By constructing joint variables, using the causal discovery method to obtain causal indicators between meteorological indicators and tobacco chemical components, establish a prediction model based on a deep learning model, and use causal weights as input to improve the accuracy of the prediction model.

Benefits of technology

It improves the accuracy of predicting chemical composition content of tobacco leaves, integrates the correlation relationship between meteorological indicators on chemical composition data, reduces waste of computing resources, and improves the prediction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943185A_ABST
    Figure CN119943185A_ABST
Patent Text Reader

Abstract

The invention relates to a prediction model training method, a prediction method and a prediction system for predicting the chemical component content of tobacco leaves, and belongs to the technical field of component content prediction. The method comprises the following steps that joint variables are constructed according to years and regions based on chemical component data and corresponding meteorological indexes of tobacco leaves of the same variety, two identical or different components are selected from the joint variables, causal discovery is carried out on the two components to obtain causal indexes between the two components, and the causal indexes of the two components are obtained; the causal indexes are at least used for reflecting the causal relationship between the chemical component data and the meteorological indexes; and establishing a deep learning model-based prediction model taking the meteorological indexes as input and the chemical component data as output based on the causal weight obtained based on the causal relationship. According to the method, the influence relationship of the meteorological indexes on the tobacco chemical component data is put into the tobacco chemical component prediction, so that the accuracy of the prediction result of the prediction model for predicting the tobacco chemical component content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a prediction model training method, a prediction method and a prediction system for predicting the content of chemical components in tobacco leaves, and belongs to the technical field of component content prediction. Background Art

[0002] The content of chemical components in tobacco leaves directly affects the quality of cigarettes, and the quality of cigarettes has a direct impact on their sales and popularity. The production and sales of cigarettes play an important role in the economy.

[0003] The main chemical components in tobacco leaves include total sugar, reducing sugar, total alkaloids, chlorine, potassium, total nitrogen, etc. Among them, total sugar and reducing sugar provide the necessary sweetness and soft taste for cigarettes, and the appropriate content can improve the comfort of smoking; total alkaloids determine the strength and stimulation of cigarettes. Too high total alkaloids may make smokers feel uncomfortable, and too low total alkaloids may make the smoke taste bland; chlorine content affects the combustion performance of cigarettes, and too high chlorine content will lead to poor combustion and produce bad odor; potassium helps to improve the combustibility and smoking quality of tobacco leaves; total nitrogen content is closely related to the aroma and concentration of smoke, and appropriate nitrogen helps to form a rich aroma. Therefore, if the influence relationship of chemical components in tobacco leaves is explored and the content of components is predicted through the collected meteorological indicators such as the maximum temperature, average temperature (average temperature), and day and night temperature difference, it will be of great significance to improve the quality of tobacco production and the stable production of tobacco growers.

[0004] When conducting data analysis between meteorological indicators and the content of chemical components in tobacco leaves, the existing methods for analyzing the content of chemical components in tobacco leaves are mainly divided into two parts: correlation analysis between meteorological indicators and the content of chemical components in tobacco leaves (correlation analysis, also known as influence relationship mining) and prediction of chemical component content based on meteorological indicators (prediction analysis).

[0005] When mining influencing relationships, most methods based on correlation analysis, such as correlation coefficient, cluster analysis, and similarity approximation, are used. Although these methods can mine the relationship between meteorological indicators and the content of chemical components in tobacco leaves, they cannot establish a clear regression expression, and cannot clarify the unidirectional influence of meteorological indicators on the chemical components of individual tobacco leaves. When predicting the content of chemical components in tobacco leaves, most methods use machine learning or deep learning to predict the content of tobacco leaf chemical components data through meteorological indicators. However, these methods mostly use meteorological indicators directly as input, and do not consider the influence of meteorological indicators on each chemical component during prediction, and directly perform prediction work.

[0006] In most cases, correlation analysis and prediction analysis are performed independently, so that in correlation analysis, the degree of influence between variables can be obtained, but the accurate quantitative change relationship between variables cannot be obtained; in prediction analysis, the content of chemical components in tobacco leaves can be predicted with the help of meteorological indicators, but the degree of influence of meteorological indicators is often not distinguished when predicting input, resulting in a waste of computing resources and a decrease in prediction accuracy.

[0007] Therefore, when conducting data analysis between meteorological indicators and the content of tobacco chemical components, the influence relationship mining and component prediction are independent, which makes the accuracy of the prediction results of tobacco chemical component content relatively low. In summary, how to realize the mining of influence relationships between variables and the component prediction combined with variable correlation in the analysis of chemical component content of tobacco leaves is a question worth considering. Summary of the invention

[0008] The purpose of the present invention is to provide a prediction model training method for predicting the content of chemical components in tobacco leaves, so as to solve the problem that the accuracy of the prediction results of the content of chemical components in tobacco leaves predicted by the existing prediction models is relatively low; and also to provide a prediction method and prediction system for predicting the content of chemical components in tobacco leaves, so as to solve the problem that the accuracy of the existing prediction results of the content of chemical components in tobacco leaves is relatively low.

[0009] To achieve the above object, the solution of the present invention includes:

[0010] A prediction model training method for predicting the content of chemical components in tobacco leaves of the present invention comprises the following steps:

[0011] The chemical composition data of tobacco leaves of the same variety and the corresponding meteorological indicators are used to construct joint variables according to year and region. Two identical or different components are randomly selected from the joint variables, and causal discovery is performed on the two components to obtain the causal indicators between the two components, and then the causal indicators between all components are obtained. The causal indicators are at least used to reflect the causal relationship between the chemical composition data and the meteorological indicators.

[0012] Based on the causal weights obtained based on the causal relationship, a prediction model based on a deep learning model is established with meteorological indicators as input and chemical composition data as output; the chemical composition data includes chemical composition and the corresponding chemical composition content.

[0013] Furthermore, meteorological indicators include average temperature, maximum temperature, minimum temperature, temperature difference between day and night, average relative humidity, rainfall and sunshine hours.

[0014] Furthermore, the joint variable is expressed as:

[0015]

[0016] In the formula, Z is the joint variable, Xi is the i-th meteorological index, Y j is the jth chemical composition data, n is the total number of samples, X n,i and Y n,j They are respectively the i-th meteorological index and the j-th chemical composition data of a certain region in a certain year.

[0017] Furthermore, the components are selected by combining the covariance matrix of each component within the variable to screen out components with high correlation;

[0018] The covariance matrix of each component in the joint variable is calculated by the correlation calculation method for the sample covariance of the joint variable. The covariance matrix of each component in the joint variable is used to reflect the correlation between the components in the joint variable. The sample covariance of the joint variable is obtained based on the mean of the joint variable and the joint variable.

[0019] Furthermore, correlation calculation methods include Graphical Lasso method, automatic correlation determination regression method, decision tree regression algorithm and feature selection graph neural network.

[0020] Furthermore, causal discovery is achieved using the IGCI method, additive noise model, CDS method or graph neural network.

[0021] Furthermore, the deep learning algorithms used in the deep learning model include BP neural network, convolutional neural network and recurrent neural network.

[0022] Furthermore, the loss function used by the prediction model is expressed as follows:

[0023]

[0024] In the formula, is the prediction result of the prediction model, Y k,j is the actual chemical composition data, j is one of the chemical components, m is the number of chemical components, and k is the number of samples.

[0025] A method for predicting the content of chemical components in tobacco leaves of the present invention comprises the following steps:

[0026] The meteorological indicators of the tobacco leaves to be tested are input into the prediction model to obtain the chemical composition data for predicting the tobacco leaves to be tested, so as to predict the chemical composition content of the tobacco leaves to be tested; the prediction model is trained using the prediction model training method for predicting the chemical composition content of tobacco leaves as mentioned above, and the variety of tobacco leaves in the prediction model training method is the same as the variety of tobacco leaves to be tested.

[0027] A prediction system for predicting the content of chemical components in tobacco leaves of the present invention comprises a processor, and the processor is used to execute a computer program to implement the steps of the prediction method for predicting the content of chemical components in tobacco leaves as described above.

[0028] Beneficial effects of the present invention:

[0029] The present invention is a pioneering invention, which provides a prediction model training method for predicting the content of chemical components in tobacco leaves. Taking into account the influence of meteorological indicators on the chemical component data of tobacco leaves, the influence relationship is put into the prediction of the chemical components of tobacco leaves. Specifically, based on the causal weight obtained by the causal relationship, a prediction model based on a deep learning model is established with meteorological indicators as input and chemical component data as output. The training of the prediction model incorporates the correlation between meteorological indicators and the chemical component data of tobacco leaves, so that the accuracy of the prediction results of the chemical component content of tobacco leaves predicted by the prediction model is improved. Among them, the causal relationship is at least the causal relationship between the chemical component data of tobacco leaves and the meteorological indicators. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flow chart of training a prediction model for predicting the content of chemical components in tobacco leaves;

[0031] Figure 2 It is a causal pointing diagram between meteorological variables and tobacco leaf chemical composition variables;

[0032] Figure 3 It is the structural diagram of the causal weighted BP model;

[0033] Figure 4 This is the framework diagram of the BP-causal weighted tobacco leaf chemical component content prediction model. DETAILED DESCRIPTION

[0034] The present invention uses a causal discovery method to mine the guiding relationship between meteorological indicators and various chemical components of tobacco leaves, and at the same time combines the results of the causal discovery as a weight mechanism into a prediction model based on deep learning, thereby realizing a deep reinforcement learning-causal weighted tobacco chemical component prediction model combined with a causal weight mechanism, so as to complete the causal-oriented discovery of meteorological variables (meteorological indicators) and chemical component content (chemical component data) during the analysis process, and add the correlation between variables in the prediction, thereby enriching data features, improving the prediction effect of tobacco leaf chemical components, and improving the accuracy of the prediction results of tobacco leaf chemical component content.

[0035] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments.

[0036] An embodiment of a prediction system for predicting the content of chemical components in tobacco leaves:

[0037] A prediction system for predicting the content of chemical components in tobacco leaves comprises a processor, wherein the processor is used for executing a computer program to implement the steps of a prediction method for predicting the content of chemical components in tobacco leaves.

[0038] A method for predicting the content of chemical components in tobacco leaves comprises the following steps: inputting meteorological indicators of the tobacco leaves to be tested into a prediction model to obtain chemical component data for predicting the content of chemical components in the tobacco leaves to be tested, so as to predict the content of chemical components in the tobacco leaves to be tested.

[0039] The prediction model is trained by using a prediction model training method for predicting the content of chemical components in tobacco leaves, and the variety of tobacco leaves in the prediction model training method is the same as the variety of tobacco leaves to be tested.

[0040] A prediction model training method for predicting the content of chemical components in tobacco leaves, such as Figure 1 As shown, the following steps are included:

[0041] The guiding relationship (causal relationship) between meteorological indicators and various chemical components of tobacco leaves is explored in the following ways:

[0042] The chemical composition data of the same variety of tobacco and the corresponding meteorological indicators are used to construct joint variables according to year and region. Two identical or different components are randomly selected from the joint variable, and causal discovery is performed on the two components to obtain the causal indicators between the two components, and then the causal indicators between all the components in the joint variable are obtained. The causal indicators are at least used to reflect the causal relationship between the chemical composition data and the meteorological indicators.

[0043] The results of causal discovery are combined as a weight mechanism into the prediction model based on deep learning. Specifically, a prediction model based on deep learning model is established with meteorological indicators as input and chemical composition data as output based on the causal weights obtained based on the causal relationship. The training of this prediction model incorporates the correlation between meteorological indicators and tobacco leaf chemical composition data, and can complete the causal-oriented discovery of meteorological variables (meteorological indicators) and chemical composition content (chemical composition data) during the analysis process, thereby improving the accuracy of the prediction results of the prediction model.

[0044] The chemical composition data includes chemical composition and corresponding chemical composition content.

[0045] Among them, the causal index can only reflect the causal relationship between chemical composition data and meteorological indicators.

[0046] Causal indicators can also be used to reflect the causal relationship between chemical composition data and meteorological indicators and the causal relationship between chemical composition data and chemical composition data.

[0047] Causal indicators can also be used to reflect the causal relationship between chemical composition data and meteorological indicators and the causal relationship between meteorological indicators.

[0048] Causal indicators can also be used to reflect the causal relationship between chemical composition data and meteorological indicators, the causal relationship between chemical composition data and chemical composition data, and the causal relationship between meteorological indicators and meteorological indicators.

[0049] Specifically, meteorological indicators include average temperature, maximum temperature, minimum temperature, temperature difference between day and night, average relative humidity, rainfall and sunshine hours.

[0050] As other implementations, the meteorological index includes average temperature, maximum temperature, minimum temperature, temperature difference between day and night, maximum humidity, minimum humidity, average humidity, relative humidity, rainfall and sunshine hours.

[0051] The meteorological index can be specifically selected according to its impact on the chemical composition of tobacco leaves.

[0052] Specifically, the chemical components of tobacco leaves include total sugars, reducing sugars, total alkaloids, chlorine, potassium and total nitrogen.

[0053] Specifically, the joint variable is represented as:

[0054]

[0055] In the formula, Z is the joint variable, X i is the i-th meteorological index, Y j is the jth chemical composition data, n is the total number of samples, X n,i and Y n,j They are respectively the i-th meteorological index and the j-th chemical composition data of a certain region in a certain year.

[0056] Specifically, the components are selected through the covariance matrix of each component in the joint variable to screen out components with high correlation, and the components with high correlation are used as components for causal discovery; the covariance matrix of each component in the joint variable is obtained by calculating the sample covariance of the joint variable through the correlation calculation method, and the covariance matrix of each component in the joint variable is used to reflect the correlation between each component in the joint variable, and the sample covariance of the joint variable is obtained according to the mean of the joint variable and the joint variable.

[0057] Specifically, the correlation calculation methods include the Graphical Lasso method, the automatic correlation determination regression method, the decision tree regression algorithm and the feature selection graph neural network.

[0058] Specifically, causal discovery is achieved using the IGCI method, additive noise model, CDS method or graph neural network.

[0059] Specifically, the deep learning algorithms used in the deep learning model include BP neural network, convolutional neural network and recurrent neural network.

[0060] Specifically, the loss function used by the prediction model is expressed as follows:

[0061]

[0062] In the formula, is the prediction result of the prediction model, Y k,j is the actual chemical composition data, j is one of the chemical components, m is the number of chemical components, and k is the number of samples.

[0063] Taking BP neural network as an example, the prediction model training method of the present invention is further explained.

[0064] In order to better consider the relationship between variables, the causal discovery results are transformed into a weight mechanism and added to the BP network prediction model. The hidden layer features are further processed to enhance the model's ability to predict the content of chemical components in tobacco leaves.

[0065] Specifically, the tobacco leaf chemical composition prediction model training method based on BP-causal weighting mainly includes the following steps:

[0066] 1) Collect tobacco leaf composition data and perform data preprocessing:

[0067] Step 1: Collect tobacco leaf composition data.

[0068] The data obtained include tobacco growing year, tobacco growing area, tobacco variety and tobacco chemical composition data.

[0069] The data on the chemical composition of tobacco leaves include chemical composition and the corresponding content of chemical composition. The chemical composition includes total sugar, reducing sugar, total alkaloids, chlorine, potassium and total nitrogen.

[0070] Among them, total sugar, reducing sugar, total alkaloids, chlorine, potassium, and total nitrogen were used as prediction target variables.

[0071] Step 2: Select the tobacco variety to be predicted.

[0072] Step 3: Use year and county (region) as key fields to construct groups, calculate the mean of the six chemical component variables of total sugar, reducing sugar, total alkaloids, chlorine, potassium, and total nitrogen in each group, and use the above variables as the variables to be predicted. j Indicates, j = 1, 2, ..., 6. If the amount of data collected is n, for j = 1, 2, ... 6, Y j It can be further expressed as Y j =(Y 1,j ,Y 2,j ,…,Y k,j ,…,Y n,j ) T.

[0073] 2) Collect meteorological data corresponding to tobacco leaf composition data and perform data preprocessing:

[0074] Step 1: Collect meteorological data corresponding to tobacco leaf composition data.

[0075] The data obtained include year, region and meteorological indicators. The meteorological indicators include average temperature, maximum temperature, minimum temperature, temperature difference between day and night, average relative humidity, rainfall, sunshine hours and other indicators.

[0076] Step 2: For the meteorological data obtained, grouping labels are also constructed based on years and counties. The mean values ​​of the seven meteorological variables, including average temperature, maximum temperature, minimum temperature, day-night temperature difference, average relative humidity, rainfall, and sunshine hours, are calculated for each county in each year as the independent variable data in the chemical composition prediction. i Indicates that i = 1, 2, ..., 7. If the amount of data collected is n, for n, for i = 1, 2, ..., 7, X i It can be further expressed as X i =(X 1,i ,X 2,i ,…,X k,i ,…,X n,i ) T .

[0077] 3) Conduct causal discovery on meteorological variables and tobacco leaf chemical composition variables, and explore the causal relationship between the variables:

[0078] Step 1: Match the tobacco leaf chemical composition data obtained in 1) and 2) with the meteorological indicators according to the year and district, and construct a joint variable Z in the following form:

[0079]

[0080] Step 2: First, calculate the mean μ of the joint variable Z obtained in Step 1) Z :

[0081] μ Z =(μ1μ2μ3μ4μ5μ6μ7μ8μ9μ 10 μ 11 μ 12 μ 13 )

[0082] in,

[0083] Then we can get the sample covariance S of the joint variable Z Z :

[0084]

[0085] Step 3: Based on sample covariance S Z , the following function can be minimized by the Graphical LASSO method:

[0086] -log‖∑ Z ‖+tr(S Z ∑ Z )+λ‖∑ Z ‖1

[0087] Get the covariance matrix ∑ of each component in the joint variable Z Z , in order to preliminarily obtain the correlation between the components within the joint variable.

[0088] Step 4: By using the IGCI method, we can get any two components Z in the joint variable k and Z l Causal indicators Where k,l∈[1,13].

[0089] Step 5: Through Step 4, the causal relationship between meteorological variables and the content of tobacco chemical components can be determined according to the IGCI method, and the causal adjacency matrix A between the corresponding variables can be obtained. Z :

[0090]

[0091] So far, we can get Figure 2 The causal relationship shown is directed to the diagram.

[0092] 4) Based on the causal mining results in 3), the causal weights are used to predict the content of chemical components in tobacco leaves:

[0093] Step 1: Convert the Y obtained in 1) j , j = 1, 2, ..., 6 can be rewritten as follows:

[0094]

[0095] Similarly, the X obtained in 2) i , i=1,2,…,7 can be rewritten as follows:

[0096]

[0097] Step 2: Extract A in Step 3) Z Normalize to get the weight transformation matrix W:

[0098]

[0099] in:

[0100]

[0101]

[0102] Step 3: Build Figure 3 The deep learning model shown.

[0103] In the model, for k = 1, 2, ..., n, when the input X k,· When , it first passes through a linear layer Linear1(·) to obtain the corresponding hidden layer feature H k :

[0104] H k =Linear1(X k,· )=(H k,1 ,H k,2 ,…,H k,13 ) 1×13

[0105] Then the hidden layer features H k Weighted

[0106]

[0107] The above After a linear layer Linear2(·), we can get X k,· The model prediction results

[0108]

[0109] Step 4: Construct the loss function Loss in the following form:

[0110]

[0111] The final prediction model is obtained by minimizing the above loss. The prediction model is the BP-causal weighted tobacco leaf chemical component content prediction model, and its framework structure is shown in the figure Figure 4 shown.

[0112] The present invention can realize the mining of causal relationships between variables while strengthening the feature prediction relationship between variables. The present invention is based on BP network, but it can also use methods such as convolutional neural network and recurrent neural network to change the position of action of causal weights to realize the overall model. Or it can be combined with other data fitting methods, such as NW (Nadaraya-Watson) estimation, decision tree regression, etc. Among them, NW estimation is mainly used to predict the value of continuous variables, which is a kind of non-parametric estimation method. On this basis, there are also local linear regression, local polynomial regression and its multivariate form expansion method; decision tree regression is a commonly used machine learning algorithm, which is mainly used to predict the value of continuous variables and is a commonly used machine learning algorithm.

[0113] Among them, the correlation calculation methods include ARD (automatic correlation determination regression), Decision Tree Regression (decision tree regression), FSGNN (Feature Selection Graph Neural Network), etc., and the causal relationship calculation methods include ANM (Additive Noise Model), CDS, GNN (Graph Neural Networks), etc.

[0114] The present invention uses the Graphical LASSO method and the IGCI method to realize the discovery of the causal relationship between meteorological indicators and various chemical components, and embeds the causal adjacency matrix obtained by the causal discovery into the BP network prediction as a weight matrix (causal weight matrix). The internal structure of the BP network is changed based on the causal weight matrix to obtain a BP-causal weighted tobacco leaf chemical component prediction model. While completing the prediction of the tobacco leaf chemical component content, the causal-oriented relationship between meteorological variables and tobacco leaf chemical component variables is explored, thereby enriching data features and improving the prediction effect of tobacco leaf chemical components.

[0115] Taking the absolute average error percentage as the evaluation index, the comparison results are shown in Table 1:

[0116] Table 1

[0117] method Total Sugar Reducing sugar Total plant alkaloids chlorine Potassium Total Nitrogen Normal BP 12.597 12.762 11.768 31.598 13.588 10.91 BP-Causal Weighting 11.975 11.444 11.765 28.855 13.26 10.871

[0118] In Table 1, "BP-Causal Weighting" is the evaluation index of the prediction results obtained by the BP-Causal Weighting tobacco chemical composition prediction model proposed in the present invention, and "Ordinary BP" represents the evaluation index of the prediction results of the model with the same neuron setting as "BP-Causal Weighting" but without the causal weighting mechanism. It can be seen that the effect of the present invention on the prediction of tobacco chemical composition is improved, which can prove the effectiveness of the present invention.

[0119] As a correlation calculation method, the core goal of the Graphical LASSO algorithm is to estimate a sparse inverse covariance matrix. This method has been widely used in variable selection problems in the field of high-dimensional statistics. The optimization goal of Graphical LASSO is as follows:

[0120]

[0121] where ∑ is the inverse covariance matrix to be estimated, S is the sample covariance matrix, tr(·) is the trace operator, ‖·‖1 is the matrix Chebyshev norm, which is the sum of the absolute values ​​of all elements in the matrix, and λ is the regularization parameter.

[0122] Information Geometric Causal Inference (IGCI) is a causal reasoning method based on the information orthogonality between variables. This method does not require prior knowledge and can directly realize causal reasoning from data.

[0123] In order to obtain the causal relationship between variables X and Y, the IGCI algorithm judges the causal asymmetry between the two variables through the feature that "the derivatives of the probability distribution p(X) and the mapping function Y=f(X) between the two variables are statistically independent", thereby inferring the causal relationship between the two variables. Generally speaking, the point set of X is mapped to the point set of variable Y through f. The irregularity of variable Y depends not only on the mapping f, but also on the irregularity of the point set X. The irregularity of Y is equal to the irregularity of X plus the irregularity of mapping f. The IGCI method constructs the following causal index C X→Y To determine the causal relationship between variables X and Y:

[0124] C X→Y =D(p X ∥ε X )-D(p Y ∥ε Y )=S(u)-S(v)+S(p Y )-S(p X )

[0125] Among them, ε X , ε Y represents the random sequence corresponding to X and Y, D(·∥·) is the relative entropy distance, and the relative entropy distance between p and q is defined as:

[0126]

[0127] S(·) is the differential entropy, which is a measure of the amount of information (complexity of change) of a random variable. The differential entropy of a random variable X with a probability density of P(X) is defined as:

[0128] S(x)=∫p(x)logp(x)dx

[0129] If C X→Y <0, then infer that X causes Y; if C X→Y >0, it is inferred that Y causes X. Therefore, when the IGCI method works, there is a unique conclusion for the causal relationship between X and Y.

[0130] An embodiment of a prediction model training method for predicting the content of chemical components in tobacco leaves:

[0131] A prediction model training method for predicting the content of chemical components in tobacco leaves has been described in detail in an embodiment of a prediction system for predicting the content of chemical components in tobacco leaves, and will not be repeated here.

[0132] An embodiment of a method for predicting the content of chemical components in tobacco leaves:

[0133] A method for predicting the content of chemical components in tobacco leaves has been described in detail in an embodiment of a prediction system for predicting the content of chemical components in tobacco leaves, and will not be repeated here.

Claims

1. A prediction model training method for predicting the content of chemical components in tobacco leaves, characterized in that: The steps include: The chemical composition data of tobacco leaves of the same variety and the corresponding meteorological indicators are used to construct a joint variable according to year and region, and two identical or different components are selected from the joint variable, and causal discovery is performed on the two components to obtain a causal indicator between the two components, and then the causal indicator between all components is obtained, and the causal indicator is at least used to reflect the causal relationship between the chemical composition data and the meteorological indicators; Based on the causal weight obtained from the causal relationship, a prediction model based on a deep learning model is established with meteorological indicators as input and chemical composition data as output; the chemical composition data includes chemical composition and corresponding chemical composition content.

2. The prediction model training method for predicting the content of chemical components in tobacco leaves according to claim 1, characterized in that: The meteorological indicators include average temperature, maximum temperature, minimum temperature, temperature difference between day and night, average relative humidity, rainfall and sunshine hours.

3. The prediction model training method for predicting the content of chemical components in tobacco leaves according to claim 1 or 2, characterized in that: The joint variable is represented as: In the formula, Z is the joint variable, X i is the i-th meteorological index, Y j is the jth chemical composition data, n is the total number of samples, X n,i and Y n,j They are respectively the i-th meteorological index and the j-th chemical composition data of a certain region in a certain year.

4. The prediction model training method for predicting the content of chemical components in tobacco leaves according to claim 1, characterized in that: The components are selected by combining the covariance matrix of each component in the variable to screen out components with high correlation; The covariance matrix of each component in the joint variable is obtained by calculating the sample covariance of the joint variable through a correlation calculation method. The covariance matrix of each component in the joint variable is used to reflect the correlation between each component in the joint variable. The sample covariance of the joint variable is obtained according to the mean of the joint variable and the joint variable.

5. The prediction model training method for predicting the content of chemical components in tobacco leaves according to claim 4, characterized in that: The correlation calculation methods include the Graphical Lasso method, the automatic correlation determination regression method, the decision tree regression algorithm and the feature selection graph neural network.

6. The prediction model training method for predicting the content of chemical components in tobacco leaves according to claim 1, characterized in that: The causal discovery is achieved by using an IGCI method, an additive noise model, a CDS method or a graph neural network.

7. The prediction model training method for predicting the content of chemical components in tobacco leaves according to claim 1, characterized in that: The deep learning algorithms adopted by the deep learning model include BP neural network, convolutional neural network and recurrent neural network.

8. The prediction model training method for predicting the content of chemical components in tobacco leaves according to claim 1, characterized in that: The expression of the loss function used in the prediction model is as follows: In the formula, is the prediction result of the prediction model, Y k,j is the actual chemical composition data, j is one of the chemical components, m is the number of chemical components, and k is the number of samples.

9. A method for predicting the content of chemical components in tobacco leaves, characterized in that: The steps include: The meteorological indicators of the tobacco leaves to be tested are input into the prediction model to obtain the chemical composition data for predicting the tobacco leaves to be tested, so as to predict the chemical composition content of the tobacco leaves to be tested; the prediction model is trained using the prediction model training method for predicting the chemical composition content of tobacco leaves as described in any one of claims 1 to 8, and the variety of tobacco leaves in the prediction model training method is the same as the variety of tobacco leaves to be tested.

10. A prediction system for predicting the content of chemical components in tobacco leaves, comprising a processor, characterized in that: The processor is used to execute a computer program to implement the steps of the method for predicting the content of chemical components in tobacco leaves as claimed in claim 9.