A prediction method for the occurrence of Chilo suppressalis in rice
By using a random forest model and grid search cross-validation method, combined with the importance of meteorological and environmental factors, the subjectivity and dependence problems of traditional rice pest prediction methods are solved, and the accurate prediction of rice borer is achieved, which improves the robustness and scientificity of the prediction.
Patent Information
- Application Number
- CN202410970666.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-07-19
AI Technical Summary
Traditional rice pest prediction methods are subjective and dependent, and it is difficult to effectively predict the nonlinear occurrence of rice borer, and it is prone to problems of poor performance and overfitting.
The random forest model was used to combine grid search and cross-validate to determine the input parameters and parameter thresholds of the model, predict the cumulative number of borer borer, and sort the importance of meteorological environmental factors.
The cumulative number prediction of the time series of rice borer borer has been achieved, the shortcomings of traditional prediction methods have been overcome, the accuracy and robustness of the prediction have been improved, and scientific basis for prevention and control policies have been provided.
Smart Images

Figure CN118916782B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of agricultural information technology, and more specifically, it relates to a method for predicting the occurrence of Chilo suppressalis in rice. Background Art
[0002] Monitoring and early warning of rice pests is a basic task in rice protection and a prerequisite for reasonably and scientifically guiding pest control. As one of the main pests in rice production, the larvae of Chilo suppressalis will bore into the stems, causing damage such as withered sheaths, withered hearts, and white panicles to the rice, seriously affecting the yield and quality of rice. Therefore, timely and accurately predicting the occurrence of Chilo suppressalis is of great guiding significance for ensuring food security and for the plant protection department to formulate effective control policies.
[0003] Traditional methods for predicting rice pests mainly rely on empirical prediction methods, experimental prediction methods, and mathematical statistics prediction methods. These methods have certain subjectivity, and in most cases, a simple linear relationship is constructed between the prediction factors and the pests. The prediction results are highly dependent on the input data. In fact, the relationship between the prediction factors and the pests is not a simple linear relationship but a highly non-linear one. In addition, the historical data of rice pests is often a typical small-sample data, and traditional prediction methods are prone to problems such as poor performance and overfitting.
[0004] The random forest method can identify key features through feature screening in small-sample analysis, and uses the structure of integrating multiple decision trees and randomness to reduce the risk of overfitting. It can avoid the deficiencies of traditional prediction methods, and at the same time can achieve non-linear prediction of the occurrence of Chilo suppressalis in rice, with strong robustness and anti-noise ability. However, the performance of the random forest model is sensitive to its parameter settings, and the model is composed of multiple integrated decision trees, which easily leads to difficult interpretation of the calculation process and results. Summary of the Invention
[0005] The purpose of the present invention is to address the deficiencies of the prior art and propose a method for predicting the occurrence of Chilo suppressalis in rice.
[0006] In a first aspect, a method for predicting the occurrence of Chilo suppressalis in rice is provided, including:
[0007] Step 1: Obtain the information on the number of Chilo suppressalis and the information on the growth and development dates of rice in the study area;
[0008] Step 2: Obtain the meteorological environment data during the growth and development period of rice and perform standardization processing on the meteorological environment data;
[0009] Step 3: Construct a data set based on the growth and development dates of rice, the information on the number of Chilo suppressalis, and the meteorological environment data, and divide the data set into a training set and a test set;
[0010] Step 4: Based on the random forest model, use grid search and cross-validation to determine the input parameters and parameter thresholds of the random forest model, and determine the optimal numerical combination of the model input parameters according to the evaluation metrics;
[0011] Step 5: Use the optimal input parameter combination of the random forest model to predict the cumulative number of Chilo suppressalis for the training set and the test set respectively, and output the predicted values of the cumulative number of Chilo suppressalis for the training set and the test set;
[0012] Step 6: Use the evaluation metrics to evaluate the training set and the test set respectively to obtain the goodness of fit of the model and the correlation degree between the predicted value and the actual value;
[0013] Step 7: Analyze the trend of the cumulative number of Chilo suppressalis using the test set to realize the prediction of the occurrence of Chilo suppressalis.
[0014] Step 8: Rank the meteorological environment factors affecting the cumulative number of Chilo suppressalis according to the contribution degree of the meteorological environment factors in the decision tree of the random forest model, and quantify the importance of the meteorological environment factors.
[0015] Preferably, in step 2, the meteorological environment data consists of several meteorological environment factors, and the meteorological environment factors include air temperature, air humidity, dew point temperature, soil temperature, soil moisture, wind speed, wind direction, atmospheric pressure, rainfall, light intensity, average temperature and relative humidity.
[0016] Preferably, in step 2, the standardization process of the meteorological environment data includes: through range standardization, eliminate the dimensional difference between different meteorological environment factors, so that each meteorological environment factor is at the same order of magnitude; the calculation formula for the standardization process is:
[0017]
[0018] where X' is the value after standardization, X is the original data value, X max is the maximum value in the data, X min is the minimum value in the data.
[0019] Preferably, in step 3, take the date of the rice growth and development period as the time series, the Chilo suppressalis quantity information as the dependent variable, and the meteorological environment data as the independent variable; the ratio of the training set to the test set is 7:3.
[0020] Preferably, step 4 includes:
[0021] Step 4.1: Determine the input parameters and parameter value ranges of the random forest model; the input parameters include the number of trees, the maximum depth of the trees, the minimum number of samples for node splitting, and the minimum number of samples for leaves;
[0022] Step 4.2: Use grid search to traverse each parameter combination in the parameter grid of Step 4.1;
[0023] Step 4.3: For each parameter combination, perform 3-fold cross-validation. The 3-fold cross-validation divides the dataset into 3 non-overlapping subsets. For each subset, use the current subset as the validation set and the other two subsets as the training sets. Repeat this process 3 times, each time selecting a different subset as the validation set;
[0024] Step 4.4: Take the average of the root mean square errors of the 3 validations as the performance evaluation metric for this parameter combination. The smaller the value, the better the prediction performance of the model. The formula for the root mean square error is:
[0025]
[0026] where n is the number of samples, y i is the i-th actual observation, is the i-th predicted value, and RMSE is the average of the squares of the prediction errors;
[0027] Step 4.5: According to the evaluation results of the cross-validation, select the combination with the smallest RMSE on the validation set as the best numerical combination of the input parameters of the random forest model.
[0028] Preferably, in Step 6, the evaluation metrics include root mean square error, mean absolute error, coefficient of determination, and Pearson correlation coefficient. The calculation formulas are:
[0029]
[0030] where n is the number of samples, y i is the i-th actual observation, is the i-th predicted value, and MAE measures the average absolute error between the predicted value and the actual value;
[0031]
[0032] where n is the number of samples, y i is the i-th actual observation, is the i-th predicted value, is the mean of the actual observations, and R 2 is the coefficient of determination, which ranges from 0 to 1. The closer it is to 1, the better the model fitting effect;
[0033]
[0034] where x i is the predicted value, y i is the actual observation, and They are the means of the predicted value and the actual observed value respectively, r is the Pearson correlation coefficient, and the value of r ranges from -1 to 1, where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation.
[0035] Preferably, in step 8, the contribution degree to the meteorological environment factors is calculated by the following formula:
[0036]
[0037] where FI(j) is the importance of feature j, T is the number of decision trees, N t is all the nodes in tree t, v(n) is the feature used to split node n, p(n) is the sample proportion passing through node n, and Δi(n,t) is the purity change on node n of i.
[0038] In a second aspect, a prediction system for the occurrence of Chilo suppressalis in rice is provided, which is used to execute the prediction method for the occurrence of Chilo suppressalis in rice according to any one of the first aspects, including:
[0039] A first acquisition module, which is used to acquire the information of Chilo suppressalis and rice information during the growth and development period of rice in the study area; the information of Chilo suppressalis includes the image information and quantity information of Chilo suppressalis;
[0040] A second acquisition module, which is used to acquire the meteorological environment data during the growth and development period of rice and perform standardization processing on the meteorological environment data;
[0041] A construction module, which is used to construct a data set according to the growth and development date of rice, the quantity information of Chilo suppressalis and the meteorological environment data, and divide the data set into a training set and a test set;
[0042] A determination module, which is used to determine the input parameters and parameter thresholds of the random forest model based on the random forest model, using grid search and cross-validation, and determine the best numerical combination of the model input parameters according to the evaluation index;
[0043] A prediction module, which is used to use the best input parameter combination of the random forest model to predict the cumulative quantity of Chilo suppressalis for the training set and the test set respectively, and output the predicted values of the cumulative quantity of Chilo suppressalis for the training set and the test set;
[0044] An evaluation module, which is used to evaluate the training set and the test set respectively using the evaluation index to obtain the goodness of fit of the model and the correlation degree between the predicted value and the actual value;
[0045] An analysis module, which is used to analyze the cumulative quantity trend of Chilo suppressalis using the test set to realize the prediction of the occurrence of Chilo suppressalis;
[0046] A sorting module, which is used to sort the meteorological environment factors affecting the cumulative number of Chilo suppressalis according to the contribution degree of the meteorological environment factors in the decision tree of the random forest model, so as to quantify the importance of the meteorological environment factors.
[0047] In a third aspect, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored executable program. When the executable program runs, it controls the device where the computer-readable storage medium is located to execute any one of the prediction methods for the occurrence of Chilo suppressalis in rice in the first aspect.
[0048] The beneficial effects of the present invention are as follows: By obtaining the data of Chilo suppressalis and meteorological environment data in the sex pheromone device, and performing range standardization on the meteorological environment data to eliminate the dimensional difference of different meteorological environment factors, so that each meteorological environment factor is at the same order of magnitude, which is conducive to analyzing the influence of meteorological environment factors on the occurrence of Chilo suppressalis. On this basis, the present invention predicts the cumulative number of Chilo suppressalis based on the random forest model, and introduces grid search and cross-validation to find the best numerical combination of the input parameters of the random forest, so as to solve the problem that the random forest model is sensitive to the setting of input parameters. In addition, the present invention ranks the importance of the meteorological environment factors affecting the cumulative number of Chilo suppressalis, providing a reference for formulating the prevention and control policy of Chilo suppressalis pests. Compared with the traditional prediction method of Chilo suppressalis, the prediction method for the occurrence of Chilo suppressalis in rice provided by the present invention can overcome the problems of traditional pest prediction relying on experience, low efficiency, high uncertainty, etc., and realize the prediction of the cumulative number of the time series of Chilo suppressalis in rice. Description of the Drawings
[0049] Figure 1a It is a flowchart of a prediction method for the occurrence of Chilo suppressalis in rice;
[0050] Figure 1b It is a schematic diagram of the data of Chilo suppressalis and meteorological environment data;
[0051] Figure 1c It is a schematic diagram of the standardization process;
[0052] Figure 1d It is a schematic diagram of adjusting the model parameters;
[0053] Figure 1e It is a schematic diagram of the experimental results;
[0054] Figure 2a It is a schematic diagram of the study area in the embodiment of the present invention;
[0055] Figure 2b It is an analysis diagram of the number of Chilo suppressalis in the embodiment of the present invention;
[0056] Figure 3This is a standardized processing diagram of meteorological environmental factors in Jiulongwang, the study area in the embodiment of the present invention;
[0057] Figure 4 This is a forecast diagram of the cumulative number of Chilo suppressalis in the study area in the embodiment of the present invention;
[0058] Figure 5 This is a trend chart of the cumulative number of Chilo suppressalis in the study area in the embodiment of the present invention;
[0059] Figure 6 This is a characteristic ranking diagram of meteorological environmental factors in an embodiment of the present invention. DETAILED DESCRIPTION
[0060] The present invention is further described below in conjunction with embodiments. The description of the following embodiments is only used to help understand the present invention. It should be noted that for ordinary persons in the art, without departing from the principle of the present invention, the present invention can also be modified in some ways, and these improvements and modifications also fall within the scope of protection of the claims of the present invention.
[0061] Embodiment 1:
[0062] As one of the main pests of rice, the larvae of the Chilo suppressalis feed by boring into the internal tissues of rice, resulting in symptoms such as dead tips, dead hearts, and white ears, which seriously affect rice yield and quality. The prediction of the Chilo suppressalis can provide information on the potential threats of pests, provide scientific data for the formulation of prevention and control measures, and reduce the damage of the Chilo suppressalis to rice. In this regard, Example 1 of the present application provides a prediction method for the occurrence of the Chilo suppressalis. When constructing a prediction model for the occurrence of the Chilo suppressalis based on random forests, a grid search method and cross-validation are used to systematically determine the optimal combination of input parameters and parameter values, and the meteorological environmental factors that affect the cumulative number of the Chilo suppressalis are ranked, thereby ensuring that the model can achieve accurate prediction performance.
[0063] like Figure 1a As shown, the embodiment of the present application sequentially obtains prediction results through data acquisition, data processing, modeling and optimization. Specifically, the method provided in the embodiment of the present application includes:
[0064] Step 1: Obtain information about the Chilo suppressalis and rice during the growth and development of rice in the study area; the Chilo suppressalis information includes Chilo suppressalis image information and Chilo suppressalis quantity information.
[0065] Specifically, information about the Chilo suppressalis is mainly obtained through the Top Digital Agriculture Cloud, which obtains IoT sex-induced detection data, such as Figure 1b As shown, it includes the image information and the number of Chilo suppressalis. Among them, sexual luring is carried out according to the habits of Chilo suppressalis using sexual attractants. Sexual luring photos are taken at 10:30 am every day, and the artificial intelligence AI recognition system is used to automatically identify and count the Chilo suppressalis.
[0066] Step 2: Obtain the meteorological environment data during the rice growth and development period, and perform standardization processing on the meteorological environment data.
[0067] The meteorological environment data is collected by the meteorological monitoring components set by the Top Digital Agriculture Cloud. The meteorological environment data includes hourly data and daily data, and the daily meteorological environment data is used in the research.
[0068] The meteorological environment data consists of several meteorological environment factors, and the meteorological environment factors include air temperature, air humidity, dew point temperature, soil temperature, soil moisture, wind speed, wind direction, atmospheric pressure, rainfall, light intensity, average temperature, and relative humidity.
[0069] The standardization processing of the meteorological environment data is as Figure 1c shown, including: through range standardization, eliminating the dimensional differences between different meteorological environment factors, so that each meteorological environment factor is at the same order of magnitude; the calculation formula for standardization processing is:
[0070]
[0071] where X’ is the standardized value, X is the original data value, X max is the maximum value in the data, and X min is the minimum value in the data.
[0072] Step 3: Construct a dataset based on the rice growth and development period date, the number information of Chilo suppressalis, and the meteorological environment data, and divide the dataset into a training set and a test set.
[0073] In Step 3, the rice growth and development period date is used as the time series, the number information of Chilo suppressalis is used as the dependent variable, and the meteorological environment data is used as the independent variable; the ratio of the training set to the test set is 7:3.
[0074] Step 4: Based on the random forest model, use grid search and cross-validation to determine the input parameters and parameter thresholds of the random forest model, and determine the optimal numerical combination of the model input parameters according to the evaluation index.
[0075] Step 4 includes:
[0076] Step 4.1: Determine the input parameters and parameter value ranges of the random forest model; the input parameters include the number of trees, the maximum depth of the trees, the minimum number of samples for node splitting, and the minimum number of samples for leaves; define the parameter values as follows;
[0077] Number of trees: Based on the number of trees in the decision tree ensemble method, they are 100, 200, and 400 trees respectively;
[0078] The maximum depth of the tree: The number of nodes on the longest path from the root node to the farthest leaf node is 10, 20, and 30 respectively;
[0079] The minimum number of samples for node splitting: The minimum number of samples that an internal node must contain before splitting is 2, 5, and 10 respectively;
[0080] The minimum number of samples in a leaf: The minimum number of samples that a leaf node must contain is 1, 2, and 4 respectively;
[0081] Step 4.2: Use grid search to traverse each parameter combination in the parameter grid of Step 4.1;
[0082] Step 4.3: For each parameter combination, perform 3-fold cross-validation. The 3-fold cross-validation is to divide the dataset into 3 non-overlapping subsets. For each subset, use the current subset as the validation set and the remaining two subsets as the training sets. Repeat this process 3 times, each time selecting a different subset as the validation set;
[0083] Step 4.4: Take the average of the root mean square errors of the 3 validations as the performance evaluation metric for this parameter combination. The smaller the value, the better the prediction performance of the model, that is, the smaller the difference between the predicted value and the actual value. The calculation formula for the root mean square error is:
[0084]
[0085] where n is the number of samples, y i is the i-th actual observation value, is the i-th predicted value, and RMSE is the average of the squares of the prediction errors;
[0086] Step 4.5: According to the evaluation results of the cross-validation, select the combination with the smallest RMSE on the validation set as the best numerical combination of the input parameters of the random forest model to achieve an accurate prediction effect of the random forest model on the validation set.
[0087] Step 5: Use the best input parameter combination of the random forest model to predict the cumulative number of Chilo suppressalis for the training set and the test set respectively, and output the predicted values of the cumulative number of Chilo suppressalis for the training set and the test set.
[0088] Step 6: Use the evaluation metrics to evaluate the training set and the test set respectively to obtain the goodness of fit of the model and the correlation degree between the predicted value and the actual value.
[0089] Step 7: Analyze the trend of the cumulative number of Chilo suppressalis using the test set to achieve the prediction of the occurrence of Chilo suppressalis.
[0090] Step 8: Sort the meteorological environment factors affecting the cumulative number of Chilo suppressalis according to the contribution degree of the meteorological environment factors in the decision trees of the random forest model, and quantify the importance of the meteorological environment factors.
[0091] In step 8, the calculation formula for the contribution degree of the meteorological environment factors is as follows:
[0092]
[0093] Where FI(j) is the importance of feature j, T is the number of decision trees, N t is all the nodes in tree t, v(n) is the feature used to split node n, p(n) is the sample proportion passing through node n, and Δi(n,t) is the purity change on node n of i.
[0094] Example 2:
[0095] Based on Example 1, Example 2 of the present application provides a more specific prediction method for the occurrence of Chilo suppressalis in rice, and its application in reality: Ningbo City, Zhejiang Province is a typical area with multiple rice planting patterns. Single-season rice and double-season rice are planted in this area, and the rice planting patterns are diverse and the planting structure is complex. A prediction method for the occurrence of Chilo suppressalis in rice is applied to Ningbo City, Zhejiang Province. The research area mainly includes three experimental fields, Jiulongwang, Modern Seed Industry Incubation Base, and Milang Farm, which are located in Zhenhai District, Yinzhou District, and Fenghua District of Ningbo City respectively. The random forest method is used to construct a cumulative number prediction model of the time series of Chilo suppressalis in rice by combining meteorological environment information and relevant pest information.
[0096] Specifically, the above method includes:
[0097] Step 1: As shown in Figure 2, obtain the information of Chilo suppressalis and rice during the growth and development period of rice in the research area; the information of Chilo suppressalis includes Chilo suppressalis image information and Chilo suppressalis quantity information.
[0098] Step 2: As Figure 3 shown, obtain the meteorological environment data during the growth and development period of rice, and perform standardization processing on the meteorological environment data.
[0099] Step 3: Construct a data set according to the rice growth and development period date, Chilo suppressalis quantity information, and meteorological environment data, and divide the data set into a training set and a test set.
[0100] Step 4: Based on the random forest model, use grid search and cross-validation to determine the input parameters and parameter thresholds of the random forest model, and determine the best numerical combination of the model input parameters according to the evaluation index.
[0101] Step 5: As Figure 4As shown in the figure, the best combination of input parameters of the random forest model is used to predict the cumulative number of Chilo suppressalis in the training set and the test set respectively, and the predicted values of the cumulative number of Chilo suppressalis in the training set and the test set are output.
[0102] Step 6, as Figure 1e shown in the figure, the training set and the test set are evaluated using evaluation indicators to obtain the goodness of fit of the model and the correlation degree between the predicted value and the actual value.
[0103] In Step 6, the evaluation indicators include root mean square error, mean absolute error, coefficient of determination and Pearson correlation coefficient, and the calculation formulas are as follows:
[0104]
[0105] where n is the number of samples, y i is the i-th actual observation value, is the i-th predicted value, and MAE measures the average absolute error between the predicted value and the actual value;
[0106]
[0107] where n is the number of samples, y i is the i-th actual observation value, is the i-th predicted value, is the mean of the actual observation values, and R 2 is the coefficient of determination, which ranges from 0 to 1, and the closer it is to 1, the better the model fitting effect;
[0108]
[0109] where x i is the predicted value, y i is the actual observation value, and are the means of the predicted value and the actual observation value respectively, and r is the Pearson correlation coefficient. The value of r ranges from -1 to 1, where 1 represents a perfect positive correlation, -1 represents a perfect negative correlation, and 0 represents no correlation.
[0110] Step 7, as Figure 5 shown in the figure, the test set is used to analyze the trend of the cumulative number of Chilo suppressalis to realize the prediction of the occurrence of Chilo suppressalis.
[0111] Step 8, as Figure 6 shown in the figure, according to the contribution degree of meteorological environment factors to the decision tree in the random forest model, the meteorological environment factors affecting the cumulative number of Chilo suppressalis are sorted to quantify the importance of meteorological environment factors.
[0112] In Step 8, the calculation formula for the contribution degree of the meteorological environment factors is:
[0113]
[0114] Among them, FI(j) is the importance of feature j, T is the number of decision trees, N t is all the nodes in tree t, v(n) is the feature used to split node n, p(n) is the proportion of samples passing through node n, and Δi(n,t) is the change in purity at node n of i.
[0115] It should be noted that the same or similar parts in this embodiment and Embodiment 1 can be referred to each other and will not be elaborated in this application.
[0116] Embodiment 3:
[0117] Based on Embodiment 1, Embodiment 3 of this application provides a prediction system for the occurrence of Chilo suppressalis in rice, which is used to execute the prediction method for the occurrence of Chilo suppressalis in rice described in any one of the first aspects, including:
[0118] A first acquisition module, which is used to acquire the Chilo suppressalis information and rice information during the growth and development period of rice in the study area; the Chilo suppressalis information includes Chilo suppressalis image information and Chilo suppressalis quantity information;
[0119] A second acquisition module, which is used to acquire the meteorological environment data during the growth and development period of rice and perform standardization processing on the meteorological environment data;
[0120] A construction module, which is used to construct a data set according to the growth and development date of rice, the Chilo suppressalis quantity information and the meteorological environment data, and divide the data set into a training set and a test set;
[0121] A determination module, which is used to determine the input parameters and parameter thresholds of the random forest model based on the random forest model, using grid search and cross-validation, and determine the best numerical combination of the model input parameters according to the evaluation index;
[0122] A prediction module, which is used to predict the cumulative quantity of Chilo suppressalis for the training set and the test set respectively using the best input parameter combination of the random forest model, and output the predicted values of the cumulative quantity of Chilo suppressalis for the training set and the test set;
[0123] An evaluation module, which is used to evaluate the training set and the test set respectively using the evaluation index to obtain the goodness of fit of the model and the correlation degree between the predicted value and the actual value;
[0124] An analysis module, which is used to analyze the trend of the cumulative quantity of Chilo suppressalis using the test set to realize the prediction of the occurrence of Chilo suppressalis;
[0125] A sorting module, configured to sort meteorological environmental factors affecting the cumulative number of Chilo suppressalis according to the contribution degree of the meteorological environmental factors to the decision trees in the random forest model, and quantify the importance of the meteorological environmental factors.
[0126] Specifically, the system provided in this embodiment is the system corresponding to the method provided in Embodiment 1. Therefore, for the parts that are the same or similar in this embodiment and Embodiment 1, reference can be made to each other, and details will not be described again in this application.
Claims
1. A method for predicting the occurrence of rice stem borer, characterized in that: include: Step 1, obtaining information on the number of Chilo suppressalis and the date of rice growth and development in the study area; Step 2, obtaining meteorological environment data during the growth and development of rice, and standardizing the meteorological environment data; in step 2, the meteorological environment data is composed of a number of meteorological environment factors, and the meteorological environment factors include air temperature, air humidity, dew point temperature, soil temperature, soil moisture, wind speed, wind direction, atmospheric pressure, rainfall, light intensity, average temperature and relative humidity; Step 3, constructing a data set according to the rice growth and development period date, the number information of the Chilo suppressalis and the meteorological environment data, and dividing the data set into a training set and a test set; Step 4: Based on the random forest model, grid search and cross validation are used to determine the input parameters and parameter thresholds of the random forest model, and the optimal numerical combination of the model input parameters is determined based on the evaluation indicators to solve the problem that the random forest model is sensitive to the input parameter settings; Step 4 includes: Step 4.1, determine the input parameters and parameter value range of the random forest model; the input parameters include the number of trees, the maximum depth of the tree, the number of samples for the minimum node segmentation, and the minimum number of samples for the leaves; Step 4.2: Use grid search to traverse each parameter combination in the parameter grid of step 4.1; Step 4.3: For each parameter combination, perform 3-fold cross validation, which is to divide the data set into 3 non-overlapping subsets. For each subset, use the current subset as the validation set and the other two subsets as the training set. Repeat this process 3 times, each time selecting a different subset as the validation set. Step 4.4: The average value of the root mean square error of the three validations is used as the performance evaluation index of the parameter combination. The smaller the value, the better the prediction performance of the model. The calculation formula of the root mean square error is: Where n is the number of samples, y i is the actual observation value of the ith is the i-th prediction value, and RMSE is the average of the squares of the prediction errors; Step 4.5: Based on the evaluation results of cross-validation, select the combination with the smallest RMSE on the validation set as the optimal numerical combination of the input parameters of the random forest model; Step 5, using the best input parameter combination of the random forest model to predict the cumulative number of Chilo suppressalis for the training set and the test set, and outputting the predicted values of the cumulative number of Chilo suppressalis for the training set and the test set; Step 6: Use the evaluation indicators to evaluate the training set and the test set respectively to obtain the goodness of fit of the model and the correlation between the predicted value and the actual value; Step 7: Analyze the cumulative number trend of Chilo suppressalis using the test set to predict the occurrence of Chilo suppressalis; Step 8: According to the contribution of meteorological environmental factors to the decision tree in the random forest model, rank the meteorological environmental factors that affect the cumulative number of Chilo suppressalis and quantify the importance of meteorological environmental factors.
2. The method for predicting the occurrence of rice stem borer according to claim 1, characterized in that: In step 2, the meteorological environment data is standardized, including: eliminating the dimensional differences between different meteorological environment factors through range standardization, so that each meteorological environment factor is at the same order of magnitude; the calculation formula for the standardized processing is: Among them, X' is the standardized value, X is the original data value, and X max is the maximum value in the data, X min is the minimum value in the data.
3. The method for predicting the occurrence of rice stem borer according to claim 2, characterized in that: In step 3, the date of rice growth and development period is used as the time series, the number of rice stem borers is used as the dependent variable, and the meteorological environment data is used as the independent variable; the ratio of the training set to the test set is 7:
3.
4. The method for predicting the occurrence of rice stem borer according to claim 3, characterized in that: In step 6, the evaluation indicators include root mean square error, mean absolute error, determination coefficient and Pearson correlation coefficient, and the calculation formula is: Where n is the number of samples, y i is the actual observation value of the ith is the i-th predicted value, and MAE measures the mean absolute error between the predicted value and the actual value; Where n is the number of samples, y i is the actual observation value of the ith is the ith predicted value, is the mean of the actual observations, R 2 is the coefficient of determination, which is between 0 and 1. The closer it is to 1, the better the model fit is; Among them, x i is the predicted value, y i is the actual observed value, and are the means of the predicted values and the actual observed values, respectively. r is the Pearson correlation coefficient. The value of r is between -1 and 1, where 1 indicates a perfect positive correlation, -1 indicates a perfect negative correlation, and 0 indicates no correlation.
5. The method for predicting the occurrence of rice stem borer according to claim 3, characterized in that: In step 8, the contribution to the meteorological environment factor is calculated as follows: Where FI(j) is the importance of feature j, T is the number of decision trees, and N t are all nodes in tree t, v(n) is the feature used to split node n, p(n) is the proportion of samples passing through node n, and Δi(n,t) is the purity change at node i n.
6. A prediction system for the occurrence of rice stem borer, characterized in that: The method for predicting the occurrence of rice stem borer according to any one of claims 1 to 5 comprises: The first acquisition module is used to acquire information about the Chilo suppressalis and rice during the growth and development of rice in the study area; the Chilo suppressalis information includes Chilo suppressalis image information and Chilo suppressalis quantity information; The second acquisition module is used to obtain meteorological and environmental data during the growth and development of rice, and to perform standardization on the meteorological and environmental data; A construction module is used to construct a data set according to the date of rice growth and development period, the number information of Chilo suppressalis and meteorological environment data, and divide the data set into a training set and a test set; A determination module is used to determine the input parameters and parameter thresholds of the random forest model based on the random forest model by using grid search and cross validation, and to determine the optimal numerical combination of the model input parameters according to the evaluation index; A prediction module is used to use the best input parameter combination of the random forest model to predict the cumulative number of Chilo suppressalis for the training set and the test set respectively, and output the predicted values of the cumulative number of Chilo suppressalis for the training set and the test set; The evaluation module is used to evaluate the training set and the test set using evaluation indicators to obtain the goodness of fit of the model and the correlation between the predicted value and the actual value; The analysis module is used to analyze the cumulative number trend of the Chilo suppressalis using the test set to predict the occurrence of the Chilo suppressalis; The sorting module is used to sort the meteorological and environmental factors that affect the cumulative number of Chilo suppressalis according to the contribution of meteorological and environmental factors to the decision tree in the random forest model, and to quantify the importance of meteorological and environmental factors.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the method for predicting the occurrence of rice stem borer according to any one of claims 1 to 5.
Citation Information
Patent Citations
A method for predicting concrete durability based on data mining and artificial intelligence algorithm
AU2020101854A4
A model and a method for predicting the occurrence of heat stroke based on machine learning
CN109359770A
Method for classifying soil environment quality of agricultural land
CN117541095A
Cited By
Cnaphalocrocis medinalis larva occurrence time prediction method, device and equipment
CN122413147A