Forestry pest and disease risk prediction method and system
Through multi-dimensional feature extraction, data normalization and LSTM combined with phase transformation methods, the insufficient capture of medium- and long-term dependencies in forestry pest prediction is solved, and the accuracy and explanatory nature of pest risk prediction is improved.
Patent Information
- Application Number
- CN202510373165.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When facing complex, multidimensional, nonlinear pest and disease data, existing forestry pest and disease prediction methods are difficult to effectively capture long-term and short-term dependencies, and lack deep feature extraction and normalization processing, resulting in insufficient prediction accuracy, lack of multidimensional feature synergy and model explanatory effect, affecting the prediction effect.
Multi-dimensional feature extraction, data normalization processing, long-term and short-term dependency modeling and phase transformation methods based on LSTM are used, and the long-range correlation modes in pest and disease data are identified, and prediction accuracy is improved through model evaluation and optimization.
It improves the accuracy and operability of pest risk prediction, can better capture the complexity and dynamic changes of pests and diseases, and provides more reference value for prediction results and explanatory.
Smart Images

Figure CN120256825A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pest risk prediction methods, and more specifically, to a forest pest risk prediction method and system. Background Art
[0002] With the continuous deterioration of global climate change and the ecological environment, the problem of forest pests has become increasingly serious, posing a huge threat to global forest resources and ecosystems. Traditional pest monitoring and control methods mainly rely on manual observation and expert experience. This method is not only time-consuming and laborious, but also often fails to respond in a timely manner when facing large-scale forest pests, resulting in the spread of pests beyond the controllable range. To solve this problem, in recent years, forest pest prediction methods based on data analysis and artificial intelligence have gradually attracted people's attention.
[0003] In the current prior art, time series analysis has been widely used in prediction models, especially for analyzing future pest risks based on historical pest data. However, existing time series prediction methods, such as traditional ARIMA models, support vector machines (SVM), etc., are difficult to effectively handle the complexity of pest data due to the lack of the ability to capture long-term and short-term dependencies in complex time series data. These methods mostly rely on fixed assumptions, such as the linear relationship or stability of data. Since forest pest data is affected by multiple factors such as climate, geography, and species distribution, it exhibits non-linear, long-range dependence, and high-dimensional characteristics, making it difficult for traditional models to accurately predict it.
[0004] In addition, existing forest pest prediction models usually lack in-depth feature extraction of pest data, especially the deficiencies in data normalization processing and feature weight allocation, which makes the model prone to deviation when dealing with feature data of different dimensions and difficult to balance the influence of different variables on the spread of pests. For example, there may be a huge difference in magnitude between climate data (such as temperature, humidity) and key parameters such as pest distribution and spread speed. Existing technologies often lead to model weight imbalance when dealing with such problems, thereby affecting the overall prediction result.
[0005] In the prior art, the processing of pest data usually relies on linear or single-scale data analysis methods, and this limitation is particularly obvious when facing pest data with strong seasonality and periodicity. Many existing models cannot effectively identify the spatio-temporal patterns of pest outbreaks, ignoring the potential correlations between data in different time periods, especially ignoring long-range dependence relationships, resulting in poor prediction effects when facing pest data with strong periodicity. This limitation mainly stems from the fact that existing methods cannot capture both short-term dynamic changes and long-term development trends when modeling complex time series data, thereby weakening the accuracy and practicality of the model in actual applications.
[0006] Regarding the above problems, although some studies have tried to introduce deep learning techniques, such as the LSTM model, to improve the processing ability of time series data, most existing models lack effective feature extraction, normalization processing, and model interpretability mechanisms, and still have great deficiencies in capturing the long-term and short-term dependencies of pest and disease data. For example, although LSTM can handle short-term and long-term time series dependencies, in existing technologies, it is usually not combined with effective phase transformation methods, resulting in limited performance in identifying long-range correlations in complex pest and disease data. Traditional methods rely solely on repeatedly training the model to improve prediction accuracy and are difficult to make breakthroughs in the complexity of time series data.
[0007] There is also a significant defect in the existing technology, that is, in the processing of multi-dimensional data, existing methods fail to reasonably solve the problems of synergy and complementarity between features of different dimensions, resulting in low efficiency of the model in processing high-dimensional pest and disease data and often ignoring the mutual correlations between features. For example, the interactions among features such as climate, geographical location, and insect distribution have not been fully utilized in existing models, resulting in the prediction model being unable to accurately identify key variables and their contributions to risk prediction when facing complex forest pest and disease scenarios, thus affecting the overall performance of the model.
[0008] In summary, the existing forest pest and disease prediction methods and systems generally have the following problems when facing the complexity, multi-dimensionality, and non-linearity of pest and disease data:
[0009] Existing time series analysis methods cannot effectively capture the long-term and short-term dependencies of complex pest and disease data, and the prediction accuracy is insufficient;
[0010] Traditional methods lack in-depth extraction and normalization processing of pest and disease data features, and the feature weight distribution is uneven, affecting the prediction performance of the model;
[0011] There is a lack of effective modeling methods for the synergy effect and mutual complementarity of features of different dimensions in pest and disease data, affecting the processing efficiency of high-dimensional data;
[0012] In terms of multi-variable correlation and interpretability, the existing technology lacks an in-depth model interpretation mechanism and is difficult to analyze the impact of key variables on the spread of pests and diseases. Summary of the Invention
[0013] In view of the above deficiencies in the prior art, the present invention proposes a method and system for predicting the risk of forestry pests and diseases. By combining multi-dimensional feature extraction of historical pest and disease data, data normalization, long short-term dependence modeling based on LSTM, and phase transformation methods, the present invention solves multiple problems in the prior art when dealing with complex time series data. The present invention can not only improve the accuracy of predicting the risk of pests and diseases, but also provide more operable prediction results through a multi-level feature processing and interpretation mechanism.
[0014] The present invention provides a method for predicting the risk of forestry pests and diseases, and the method includes the following steps:
[0015] Obtain historical forestry pest and disease data at a preset time granularity;
[0016] Extract features from the historical data to generate corresponding feature parameters;
[0017] Normalize the feature parameters to generate standardized feature data;
[0018] Based on the standardized feature data, randomly divide the training set, validation set, and test set according to a preset ratio;
[0019] Train the risk prediction model through the training set, and adjust the hyperparameters of the risk prediction model according to the validation set;
[0020] Model the long-range dependence in the time series based on the phase transformation method, specifically including obtaining the phase encoding of the time series, identifying long-range correlation patterns, constructing a dynamic dependence model, and integrating short-term and long-term transformations;
[0021] Use the test set to evaluate the trained risk prediction model;
[0022] Optimize the risk prediction model according to the model evaluation results, and finally determine the optimal model;
[0023] The step of obtaining historical forestry pest and disease data at a preset time granularity includes: obtaining national historical pest and disease data at a monthly time granularity, and the data includes pest and disease names, occurrence times, and geographical area information;
[0024] The phase transformation method includes the following steps:
[0025] First, obtain the phase encoding of the time series through a timing detector , and the calculation formula is as follows:
[0026] ;
[0027] Among them, the represents time The phase value at a moment, the is a time series in complex form, including an imaginary part and a real part Re ;
[0028] Then, calculate the statistical dependence of the time series , and its calculation formula is:
[0029] ;
[0030] Among them, the is the statistical dependence, the is the total number of time steps of the time series, represents the change rate of phase encoding;
[0031] Next, identify the long-range correlation pattern through convolution operation, and the convolution calculation formula is:
[0032] ;
[0033] Among them, the is the convolution kernel weight, the is the convolution kernel size, the is the historical value of phase encoding;
[0034] Finally, construct a dynamic dependence model, and the dynamic dependence model is based on the following formula:
[0035] ;
[0036] Among them, the is the state value at time , the and are system parameters, the is a non-linear excitation term.
[0037] Preferably, the feature extraction step includes extracting geographical location, climate characteristics, insect distribution, agricultural pest and disease conditions, historical data of forestry pest and disease prediction, and relevant information on forestry and agricultural pests and diseases.
[0038] Preferably, the normalization process is carried out through the following formula:
[0039] ;
[0040] Among them, the represents the original value of the feature parameter, the and the respectively represent the minimum and maximum values of the feature parameter, and the is the normalized feature value.
[0041] Preferably, the risk prediction model is trained by a long short-term memory network (LSTM), and the loss function of the model is defined as:
[0042] ;
[0043] wherein, the is the actual value of the th sample, the is the predicted value, the are the parameters of the model, the is the regularization coefficient, and the is the number of samples.
[0044] Preferably, the model evaluation step includes: evaluating the performance of the risk prediction model through a test set, and the performance evaluation is based on the following mean square error (MSE) and root mean square error (RMSE) formulas:
[0045] ;
[0046] ;
[0047] wherein, the are the actual values in the test set, the are the predicted values, and the is the number of samples.
[0048] Preferably, the model interpretation index evaluation includes calculating the variance of the weight score matrix to obtain a variance score matrix, and the calculation formula is:
[0049] ;
[0050] wherein, the represents the variance of the weight score matrix, the is the th weight score, the is the mean of the weight scores, and the is the number of weights.
[0051] A forest pest and disease risk prediction system for executing the method includes:
[0052] A data acquisition unit for acquiring historical forest pest and disease data with a preset time granularity;
[0053] A feature extraction unit for extracting features from the historical data to generate feature parameters;
[0054] A normalization processing unit for performing normalization processing on the feature parameters;
[0055] A dataset division unit for dividing a training set, a validation set, and a test set according to the normalized feature parameters;
[0056] A model training unit for training the risk prediction model with the training set and adjusting the hyperparameters of the model according to the validation set;
[0057] A phase transformation unit for identifying long-range dependencies in time series;
[0058] A model evaluation unit for evaluating the risk prediction model using the test set.
[0059] Preferably, the model training unit includes:
[0060] An LSTM model construction module for constructing an LSTM model;
[0061] A training module for performing data training based on the training set and adjusting the weights and thresholds of the model through the backpropagation algorithm;
[0062] A hyperparameter optimization module for optimizing the hyperparameters of the model based on the validation set.
[0063] The beneficial effects of the present invention are reflected in the following aspects:
[0064] First, by introducing the phase transformation method, the present invention effectively solves the deficiency of existing time series methods in dealing with long-range dependencies. Phase transformation can capture the periodic patterns in the spread of pests and diseases through the phase encoding analysis of time series data, and combined with the long short-term dependence modeling ability of the LSTM network, enabling the model to handle both short-term fluctuations and long-term trends simultaneously. The introduction of this method not only makes up for the limitations of LSTM in dealing with long-range dependencies but also can capture the hidden patterns in the data through phase transformation, thereby improving the prediction ability of the model.
[0065] Second, through normalization processing, the present invention reasonably solves the problem of model weight imbalance caused by the dimensional difference between different features. Normalization processing not only enables the model to process different feature data on the same scale but also can optimize the feature weights through multiple training iterations, avoiding biases in the model when dealing with key features such as climate and insect distribution. This processing method makes the model more robust in multi-variable data.
[0066] In addition, the multi-dimensional feature extraction and collaborative modeling method proposed by the present invention can make full use of the mutual correlation of multi-dimensional features such as climate, geography, and insect distribution to achieve the complementary, superimposed, and collaborative effects among features. The collaborative effect of different features in the model can enhance the ability to capture the complexity and dynamic changes of pests and diseases, and help the model better identify the importance of key variables in the risk of pests and diseases.
[0067] Finally, the present invention solves the problem of "black box" of the model in the prior art through the model interpretability mechanism. By interpreting and analyzing the model weights, users can clearly understand the key features and variables on which the model bases its risk prediction, providing a more valuable reference basis for the management and decision-making of forestry pests and diseases.
[0068] In summary, the present invention makes up for many deficiencies in the prior art through the capture of long-term dependencies in time series data, the deep extraction and normalization of multi-dimensional features, and the model interpretability mechanism. The system can effectively cope with the complexity and variability of pest and disease data, greatly improving the accuracy and operability of forestry pest and disease risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a logic block diagram of the system of the present invention.
[0070] Figure 2 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following specifically describes the specific implementation manner, structure, features, and effects of a forestry pest and disease risk prediction method proposed according to the present invention in conjunction with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art belonging to the technical field of the present invention.
[0073] The following specifically describes the specific solution of a forestry pest and disease risk prediction method provided by the present invention in conjunction with the accompanying drawings.
[0074] Please refer to Figure 1-2, the present invention relates to the scenario application of a forestry pest and disease risk prediction method and system, aiming to improve the accuracy and efficiency of risk prediction through in-depth analysis of forestry pest and disease data. The invention unfolds around this application scenario and involves data acquisition, feature extraction, normalization processing, training of a long short-term memory (LSTM) model, application of a phase transformation method, and model evaluation and optimization.
[0075] First, the present invention provides a forestry pest and disease risk prediction method, including a series of steps, specifically: obtaining historical forestry pest and disease data at a preset time granularity; extracting features from the historical data according to preset data dimensions and generating feature parameters; subsequently performing normalization processing on these feature parameters. Preferably, the normalized feature data is randomly divided into a training set, a validation set, and a test set according to a preset ratio. The risk prediction model is trained using the training set, and the hyperparameters of the model are optimized using the validation set. In addition, in the present invention, based on the phase transformation method, long-range dependencies in the time series are modeled, specifically including obtaining the phase encoding of the time series, identifying long-range correlation patterns, constructing a dynamic dependence model, and integrating short-term and long-term transformations. After the model training is completed, the trained model is evaluated using the test set, and the model is further optimized according to the evaluation results to ensure that the model performance meets the preset requirements.
[0076] In a specific application scenario, a large-scale spread of pine wilt disease has occurred in the pine trees in a certain forest area. To predict the pest and disease risk in the next period of time, the system first obtains the pest and disease data for each month in the past 5 years from relevant data platforms, including the name of the pest and disease, the occurrence time, and the location area. After obtaining this historical data, the system will perform feature extraction for each pest and disease data point, such as extracting the climate conditions (temperature, humidity) of each region, the types of pests, the spread speed of the pest and disease, etc. Through feature extraction, we can analyze the correlation between climate conditions and pest and diseases more deeply. The beneficial effect of this step is that it can establish a comprehensive and accurate historical data foundation for pest and diseases, thus providing a reliable input source for subsequent model prediction. These data can not only reflect the time law of pest and diseases, but also reveal the external environmental factors for the spread of pest and diseases through geographical location, climate characteristics, etc.
[0077] Preferably, in an embodiment of the present invention, forestry pest and disease data nationwide is obtained with a monthly preset time granularity, including pest and disease names, occurrence times, and their corresponding geographical area information. By using this time granularity, the seasonal and regional differences of pests and diseases can be reflected more accurately. The acquisition of this data can rely on existing relevant standards in the agricultural field, providing a basis for subsequent feature extraction and modeling. Suppose in a certain area, the risk of pests and diseases is the highest during the summer (from June to August) every year. Therefore, the system will obtain pest and disease data for this period on a monthly basis, and these data include the specific time points of pest and disease occurrence and the spreading areas. Through this monthly acquisition granularity, the seasonal patterns of pests and diseases can be analyzed, and corresponding predictions can be made according to different months. Using data with a monthly granularity can capture the seasonal changes of pests and diseases more precisely, helping the model identify high-risk and low-risk periods of pest and disease outbreaks, thereby improving the timeliness and accuracy of the prediction model.
[0078] Preferably, in an embodiment of the present invention, in the feature extraction step, the system processes historical data according to preset data dimensions. Preferably, in an embodiment of the present invention, the extracted feature parameters include geographical location, climate characteristics, insect distribution, historical conditions of agricultural pests and diseases, etc. For example, areas with dense insect distribution may have a higher risk of pests and diseases, and climate characteristics (such as temperature and humidity) play an important role in the occurrence and spread of pests and diseases. Therefore, the extraction of these features is crucial for accurately predicting pest and disease risks.
[0079] Suppose the temperature and humidity in a certain area change greatly, and the system will extract these climate characteristics. Further combined with the insect distribution, for example, a certain leaf-eating insect is concentrated in the low-altitude area, and this information is also input into the model as a feature parameter. Through multi-dimensional feature extraction, the system can consider the influence of different environmental factors on the spread of pests and diseases, thereby generating more targeted risk predictions. Through multi-dimensional feature extraction, the complex factors behind the spread of pests and diseases can be fully considered. These features can reveal the internal laws of pest and disease occurrence, especially in the influence of climate, geographical factors, insect distribution, etc. on the spread of pests and diseases, improving the model's ability to identify future risks.
[0080] Preferably, in an embodiment of the present invention, normalization processing is essential for the standardization of feature parameters. In the present invention, common normalization processing methods are used, as shown in the following formula:
[0081] ;
[0082] wherein, the represents the original value of the feature parameter, the and the respectively represent the minimum and maximum values of the characteristic parameters, and the is the normalized eigenvalue. Through normalization, the differences in dimensions between features can be eliminated, making the gradient descent in the model training process more stable, thereby improving the convergence efficiency of the model.
[0083] Suppose the temperature data range obtained by the system is from 10°C to 40°C, and the humidity data is from 0% to 100%. To avoid affecting the results due to differences in data magnitudes during model training, these features need to be normalized. After normalization, both temperature and humidity are converted to the range of 0 to 1, so that all features can be placed on the same scale, enabling the model to process different features evenly.
[0084] Normalization processing can improve the convergence speed of the model and avoid the weight imbalance of the model during training due to excessive differences in feature data. At the same time, it also helps to prevent the model from being biased towards different features and improves the prediction accuracy.
[0085] Preferably, in an embodiment of the present invention, the normalized characteristic parameters are randomly divided into a training set, a validation set, and a test set according to the ratio of 6:2:2. The training set is used for training the risk prediction model, the validation set is used for adjusting hyperparameters, and the test set is used for the final evaluation of the model. This way of dividing the dataset ensures that the model can maintain good generalization ability during the training process and avoids overfitting. Through reasonable data division, the overfitting problem can be effectively avoided, ensuring that the model can not only perform well on the training data but also maintain a high prediction ability on unknown data. This division method can improve the robustness and generalization of the model.
[0086] Preferably, in an embodiment of the present invention, the training process of the risk prediction model relies on a long short-term memory (LSTM) network. By using the LSTM network in the present invention, the long-term and short-term dependencies in time series data can be effectively captured. Preferably, in an embodiment of the present invention, the loss function of the model is as follows:
[0087] ;
[0088] wherein, the is the actual value of the th sample, the is the predicted value, the are the parameters of the model, the is the regularization coefficient, which is used to prevent the model from overfitting, and the is the number of samples. Through the backpropagation algorithm, the gradient of the loss function is calculated and fed back into the LSTM model to adjust the model's weight and threshold parameters. After multiple trainings, an optimal LSTM model is finally generated.
[0089] In a specific scenario, the actual incidence rate of pests and diseases is , and the result predicted by the model is . The difference between the two is calculated through the loss function and fed back to the model for weight adjustment. During the training process, the system repeatedly adjusts the model's weights to gradually reduce the value of the loss function, thereby improving the prediction accuracy. By using the LSTM model, the system can effectively capture the long-term and short-term dependencies of the occurrence of pests and diseases in the time series, enabling the model to still have a high prediction accuracy when facing complex time series data. At the same time, the addition of the regularization term can avoid overfitting of the model and improve the generalization ability.
[0090] During the outbreak of pine wilt disease in a certain forest area, through the phase transformation network, the system can identify the periodic changes of pest and disease data in different time periods. For example, the outbreak period of pine wood nematodes is usually highly correlated with the local temperature and humidity. Through phase transformation, the system can analyze the phase encoding of these external influencing factors (such as climate data) and capture the key time points and their periodic changes of the occurrence of pests and diseases. The phase transformation network uses the following formula to calculate the phase encoding of the time series:
[0091] ;
[0092] wherein, the represents the phase value at time , and the is a time series in complex number form, including the imaginary part and the real part Re . In this way, the key phase information can be extracted from the time series data and then used to model the long-range dependencies in the sequence.
[0093] Next, the steps for calculating the statistical dependence degree are proposed. Preferably, in an embodiment of the present invention, the statistical dependence degree is calculated through the following formula:
[0094] ;
[0095] wherein, the is the statistical dependence degree, and the is the total number of time steps of the time series; Represents the change rate of phase encoding. This dependency can be used to measure the long-range dependency in time series data, providing a basis for subsequent convolutional operations to identify patterns.
[0096] The phase transformation method can capture hidden patterns in time series. For example, in the pest outbreak cycle, phase transformation can reveal the seasonal changes and long-term development trends of pests. By calculating the phase encoding and dependency, the system can identify long-range correlations in pest data for more accurate prediction. The phase transformation method can effectively capture long-range dependencies in pest occurrence data, helping the system better identify potential outbreak cycles. This method can perform in-depth pattern recognition on complex time series data, enhancing the model's prediction ability for the long-term development trends of pests.
[0097] The identified phase patterns can be further extracted through convolutional operations. By convolutional operations, long-range correlation patterns are identified, and the convolutional calculation formula is:
[0098] ;
[0099] where, the is the convolutional kernel weight, the is the convolutional kernel size, and the is the historical value of phase encoding; through convolutional operations, the system can further extract the key features of phase patterns and identify potential patterns of pest spread. Through the phase transformation network, the system can effectively capture long-range dependency relationships in time series, especially for pest data with strong periodicity and complex time series, and can identify potential pest outbreak rules. Convolutional operations further enhance the ability to extract phase patterns, enabling the system to not only identify short-term dynamic changes but also capture long-term time series correlations. Compared with existing linear time series analysis methods, this method can more accurately predict the outbreak period and spread trend of pests, thereby improving the accuracy of risk prediction.
[0100] During the pest spread process, external factors (such as climate fluctuations, geographical environment) often show complex non-linear changes. At this time, simple linear models are difficult to accurately predict the development trajectory of pests. Therefore, the present invention performs dynamic modeling on time series data through a chaotic system unit to capture the non-linearity and uncertainty in the pest spread process.
[0101] The chaotic system unit constructs a dynamic dependency model using chaos theory, and the dynamic dependency model is based on the following formula:
[0102] ;
[0103] where, the is time The state value at a moment, where the and are system parameters, and the is a non - linear excitation term. In this equation, the chaotic system can simulate the randomness and complex interactions in the process of pest and disease spread by introducing non - linear excitation. For example, the non - linear impact of climate change on the spread of pests and diseases. By introducing the chaotic system, the system can better model and predict the spread path and impact range of pests and diseases.
[0104] Through the chaotic system unit, the present invention can capture the complex dynamic characteristics in the process of pest and disease spread, especially the spread of pests and diseases that are sensitive to changes in external conditions. For example, the impact of temperature and humidity changes on pine wilt disease is highly non - linear, and it is difficult for traditional linear models to accurately predict. The chaotic system can effectively simulate these complex external impacts through non - linear equations, thereby improving the model's prediction ability for pest and disease spread.
[0105] In addition, the combination of the chaotic system and the phase transformation network realizes the multi - dimensional identification of pest and disease spread patterns. The phase transformation captures the periodicity and long - range correlations in the time series, while the chaotic system solves the complex dynamic problems that are difficult to handle by traditional linear models by introducing non - linear excitation. This method not only makes up for the deficiencies of traditional prediction methods but also improves the processing ability for complex time - series data through synergy.
[0106] Preferably, in an embodiment of the present invention, the evaluation of the model is carried out through a test set. Preferably, in an embodiment of the present invention, the evaluation indicators include the mean square error (MSE) and the root mean square error (RMSE), and their calculation formulas are as follows:
[0107] ;
[0108] ;
[0109] where the is the actual value in the test set, the is the predicted value, and the is the number of samples. By calculating these two error indicators, the prediction performance of the model can be quantitatively evaluated to ensure the generalization ability of the model on the test set.
[0110] For example, when the system predicts that the incidence rate of pests and diseases in a certain month is 80%, while the actual incidence rate is 85%, the system will calculate the error through MSE and RMSE and feedback it to the model for adjustment.
[0111] Through the calculation of MSE and RMSE, the system can accurately measure the performance of the model on the test set to ensure that the model can handle risk prediction in actual scenarios.
[0112] Finally, in an embodiment of the present invention, a method for evaluating the model interpretability index is proposed. Preferably, in an embodiment of the present invention, the model interpretation index is realized by calculating the variance of the weight score matrix as follows:
[0113] ;
[0114] wherein, the represents the variance of the weight score matrix, the is the th weight score, the is the mean value of the weight scores, and the is the number of weights. In this way, the influence of the weights of each feature in the model on the prediction result can be evaluated, thereby improving the interpretability of the model.
[0115] By interpreting the weight scores, the system can identify which features have the greatest impact on the prediction of pest and disease risks. For example, certain climate features may be more important for the spread of pests and diseases than other factors. By interpreting these weights, the system can provide more guiding prediction results for forestry managers. The model interpretation index can improve the transparency of the system and help users understand how the model makes predictions. This is crucial for the interpretability of the model, especially when specific decisions need to be made, as it can more clearly identify which factors play a key role in the prediction of pests and diseases.
[0116] The present invention also relates to a forestry pest and disease risk prediction system, which is mainly applied to the acquisition, analysis, and prediction of historical data on forestry pests and diseases. The system forms a complete solution for forestry pest and disease risk prediction through the collaborative work of multiple modules, including functional modules such as data acquisition, feature extraction, normalization processing, dataset division, model training, phase transformation modeling, and model evaluation.
[0117] The forestry pest and disease risk prediction system of the present invention can be divided into multiple units, each unit undertaking different functions and collaborating with each other to complete the entire prediction process. First, the system includes a data acquisition unit 1, whose main function is to obtain data from the historical records of forestry pests and diseases. These data are usually obtained at a preset time granularity, preferably on a monthly basis, covering the historical data of national pests and diseases. The data acquisition unit 1 can be docked with an agricultural data platform or a forest monitoring database to ensure the accuracy and timeliness of the data.
[0118] Next, the feature extraction unit 2 in the system is responsible for performing multi-dimensional feature extraction on the acquired historical data. Preferably, the features include geographical location, climate conditions, insect distribution, historical spread of forestry pests and diseases, etc. The feature extraction unit 2 automatically extracts representative features from the historical data according to preset rules or machine learning-based methods and generates corresponding feature parameters. For example, in a certain area, climate data such as temperature and humidity have a greater impact on the occurrence of pests and diseases, and the extraction of these features will directly affect the prediction accuracy of the model.
[0119] To ensure that the model can process feature data with different dimensions, the present invention provides a normalization processing unit 3 for normalizing the extracted feature parameters. The normalization processing unit 3 can effectively eliminate the influence of data dimensions, enabling different features to be trained in the same scale for the model.
[0120] The system further includes a dataset division unit 4 for dividing the training set, validation set, and test set according to the normalized feature parameters in a preset ratio. Preferably, the dataset division unit 4 can divide the data into the training set, validation set, and test set in a ratio of 6:2:2. Through this division method, the model can be fully trained, and at the same time, the validation set and test set can ensure the generalization ability of the model on different datasets.
[0121] In the model training stage, the model training unit 5 in the system trains the risk prediction model with the training set and adjusts the hyperparameters of the model according to the validation set. Preferably, in an embodiment of the present invention, the model training unit 5 includes an LSTM model construction module 51, a training module 52, and a hyperparameter optimization module 53. The LSTM model construction module 51 is used to construct a long short-term memory (LSTM) model. The training module 52 trains the model through the backpropagation algorithm and adjusts the weights and thresholds of the model to ensure that the model can capture long-term and short-term dependencies in the time series. The hyperparameter optimization module 53 automatically adjusts the hyperparameters of the model based on the results of the validation set, enabling the model to maintain the best prediction performance in different scenarios.
[0122] In addition, the system further includes a phase transformation unit 6 for identifying long-range dependencies in the time series. The phase transformation unit 6 obtains the phase encoding of the time series through a timing detector and analyzes the time series using the phase transformation method. Through the phase encoding, the system can identify long-range correlation patterns in the pest and disease data, especially in seasonal pests and diseases or pests and diseases that break out under certain specific conditions, and can effectively identify their potential spread trends.
[0123] After completing model training and dependency recognition, the system uses the model evaluation unit 7 to evaluate the trained risk prediction model. The model evaluation unit 7 evaluates the performance of the model through the test set. Preferably, evaluation metrics such as mean squared error (MSE) and root mean squared error (RMSE) are used to judge the prediction accuracy of the model. For example, if the actual incidence rate of a certain pest and disease in the test set is 70%, and the incidence rate predicted by the model is 65%, then the error value obtained by evaluating the model can be used for further optimization of the subsequent model.
[0124] Through the close cooperation of the above-mentioned multiple functional units, the forest pest and disease risk prediction system of the present invention realizes the accurate prediction of forest pest and diseases. First of all, the data acquisition unit 1 ensures that the system can obtain high-quality historical data and can flexibly adjust the data collection scope according to different time granularities. The feature extraction unit 2 ensures that the model can capture the impact of different environmental variables on pest and diseases through multi-dimensional data extraction, thereby improving the performance of the prediction model.
[0125] The normalization processing unit 3 ensures that the feature data can be fully utilized by the model at the same scale by normalizing the feature data with different dimensions. The reasonable partitioning scheme of the data set partitioning unit 4 ensures the generalization ability of the model, and at the same time the partitioning of the training set and the validation set ensures that the model can maintain consistent performance on different data sets.
[0126] The introduction of the LSTM model in the model training unit 5 solves the problem of difficult to capture long-term and short-term dependencies in time series data. The phase transformation unit 6 further improves the pattern recognition ability of complex pest and disease data by identifying long-range dependencies in time series. Finally, the model evaluation unit 7 ensures that the model can be applied in actual scenarios through strict testing and evaluation, providing a reliable prediction basis for forestry management.
[0127] A forest pest and disease risk prediction system of the present invention combines the advantages of multiple functional units and provides an efficient and accurate pest and disease risk prediction solution through a series of processes such as data acquisition, feature extraction, normalization processing, model training, phase transformation, and model evaluation. Each module in the system can effectively cope with the complexity of pest and disease transmission through reasonable design and optimization, providing strong support for the prevention and control and management of forest pest and diseases.
[0128] The present invention provides an innovative forest pest and disease risk prediction method and system, which forms a complete process from the acquisition of historical data, feature extraction, model training, the introduction of phase transformation methods, to the evaluation and optimization of the final model. By combining prior knowledge and algorithm innovation in the field, the present invention can effectively improve the accuracy of pest and disease risk prediction and provide more powerful decision-making support for forestry management.
[0129] The embodiments described in this specification are all preferred embodiments of the present invention, and do not limit the scope of implementation of the present invention. Therefore, all equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.
Claims
1. A method for predicting the risk of forestry pests and diseases, characterized in that, The method includes the following steps: Obtain historical data of forestry pests and diseases at a preset time granularity; Extract features from the historical data to generate corresponding feature parameters; Perform normalization processing on the feature parameters to generate standardized feature data; Based on the standardized feature data, randomly divide the training set, validation set, and test set according to a preset ratio; Train the risk prediction model through the training set and adjust the hyperparameters of the risk prediction model according to the validation set; Model the long-range dependence in the time series based on the phase transformation method, specifically including obtaining the phase encoding of the time series, identifying long-range correlation patterns, constructing a dynamic dependence model, and integrating short-term and long-term transformations; Use the test set to evaluate the trained risk prediction model; Optimize the risk prediction model according to the model evaluation results and finally determine the optimal model; The step of obtaining historical data of forestry pests and diseases at a preset time granularity includes: obtaining national historical data of pests and diseases at a monthly time granularity, and the data includes pest and disease names, occurrence times, and geographical area information; The phase transformation method includes the following steps: First, obtain the phase encoding of the time series through a timing detector , and the calculation formula is as follows: ; Among them, the represents the phase value at the time instant, and the is a complex-valued time series that includes an imaginary part and a real part Re ; Then, calculate the statistical dependence of the time series , and its calculation formula is: ; Among them, the is the statistical dependence, and the is the total number of time steps of the time series, represents the change rate of phase encoding; Next, identify long-range correlation patterns through convolution operations, and the convolution calculation formula is: ; Among them, the is the convolution kernel weight, the is the convolution kernel size, and the is the historical value of phase encoding; Finally, construct a dynamic dependence model, and the dynamic dependence model is based on the following formula: ; Among them, the is the state value at time , the and are system parameters, and the is a non-linear excitation term.
2. The forestry pest and disease risk prediction method according to claim 1, characterized in that, The feature extraction step includes extracting geographical location, climate characteristics, insect distribution, agricultural pest and disease conditions, historical data of forestry pest and disease prediction, and relevant information on forestry and agricultural pests and diseases.
3. A method for predicting the risk of forestry pests and diseases according to claim 1, characterized in that, The normalization processing is performed through the following formula: ; Among them, the represents the original value of the characteristic parameter, the and the respectively represent the minimum value and the maximum value of the characteristic parameter, and the is the normalized characteristic value.
4. A forestry pest and disease risk prediction method according to claim 1, characterized in that, The risk prediction model is trained through a long short-term memory network LSTM, and the loss function of the model is defined as: ; Among them, the is the actual value of the th sample, the is the predicted value, the are the parameters of the model, the is the regularization coefficient, and the is the number of samples.
5. A forestry pest and disease risk prediction method according to claim 1, characterized in that The model evaluation step includes: evaluating the performance of the risk prediction model through the test set, and the performance evaluation is based on the following mean square error MSE and root mean square error RMSE formulas: ; ; Among them, the is the actual value in the test set, the is the predicted value, and the is the number of samples.
6. The forestry pest and disease risk prediction method according to claim 5, characterized in that The evaluation of model interpretation metrics includes calculating the variance of the weight score matrix to obtain the variation score matrix, and the calculation formula is as follows: ; Among them, the represents the variance of the weight score matrix, the is the -th weight score, the is the mean of the weight scores, and the is the number of weights.
7. A forestry pest and disease risk prediction system for implementing the method according to any one of claims 1-6, characterized in that, Includes: A data acquisition unit for obtaining historical data of forestry pests and diseases at a preset time granularity; A feature extraction unit for extracting features from the historical data to generate feature parameters; A normalization processing unit for performing normalization processing on the feature parameters; A data set division unit for dividing the training set, validation set, and test set according to the normalized feature parameters; A model training unit for training the risk prediction model through the training set and adjusting the hyperparameters of the model according to the validation set; A phase transformation unit for identifying long-range dependence in the time series; A model evaluation unit for evaluating the risk prediction model using the test set.
8. The forestry pest and disease risk prediction system according to claim 7, characterized in that The model training unit includes: An LSTM model construction module for constructing an LSTM model; A training module for performing data training based on the training set and adjusting the weights and thresholds of the model through the backpropagation algorithm; A hyperparameter optimization module for optimizing the hyperparameters of the model based on the validation set.
Citation Information
Cited By
Forestry disease and pest risk prediction method and system based on big data
CN121352501A