Forest pest and disease damage prediction system based on data analysis
By generating a spatiotemporal standardized data set and adaptive feature weight matrix, dynamically adjusting the feature weights of the pest and disease prediction system, the problem of instability in pest and disease prediction in the existing technology is solved, and high-precision and spatially accurate pest and disease prediction and risk assessment are achieved.
Patent Information
- Application Number
- CN202510354490.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing forest pest prediction system has unstable distribution of feature weights under multiple time periods and multiple environmental conditions, and has failed to fully adapt to the differences in pest and disease characteristic weights, resulting in low prediction accuracy, uneven resource allocation, and ineffective control of the risk of pest and disease spread.
The spatiotemporal and standardized data set is generated through data preprocessing, the feature weights are dynamically adjusted, and the feature weight values of pest samples and the adaptive feature weight weight matrix are combined, key feature vectors are extracted, the weight error change trend is dynamically calculated, and the pest dynamic prediction results are generated and the risk area is evaluated.
It achieves high accuracy and time and space accuracy of pest and disease prediction, dynamically adjusts feature weights, reduces interference from irrelevant factors, provides scientific resource allocation guidance, and improves prevention and control effects.
Smart Images

Figure CN120296495A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pest control, and in particular to a forest pest prediction system based on data analysis. Background Art
[0002] The field of pest control technology is a crucial link in agricultural and forestry production. It mainly monitors, prevents and controls the threats of diseases and pests to crops and trees through physical, chemical, biological and modern information technology. The development of this technical field includes the identification of pests and diseases, the study of their transmission mechanisms, the application and innovation of prevention and control methods, and the promotion of intelligent and automated monitoring and prediction systems. With the advancement of science and technology, especially the introduction of big data, artificial intelligence and remote sensing technology, pest control technology has gradually developed in the direction of high efficiency, precision and environmental protection, providing important support for the sustainable development of agriculture and forestry.
[0003] Among them, the forest pest and disease prediction system refers to the prediction and early warning of the occurrence and spread of forest diseases and pests through technical means such as data analysis, model building and sensor monitoring. The main purpose of this system is to help forestry workers take prevention and control measures in advance, reduce the harm of pests and diseases to forest resources, improve the efficiency and effectiveness of forest health management, reduce economic losses, and promote the stability and sustainable development of forestry ecosystems.
[0004] The fixed weight allocation of existing technologies ignores the dynamic changes of errors and fails to fully adapt to the differences in the weights of pest and disease characteristics under multiple time periods and environmental conditions, resulting in unstable prediction performance. The feature recognition link relies too much on limited data dimensions and fails to deeply extract key features. The noise interference is obvious, which reduces the accuracy of pest and disease prediction. In addition, the existing methods are relatively rough in the division of pest and disease risk areas and lack quantification of spatiotemporal distribution, resulting in uneven resource allocation, failure to effectively control the risk of pest and disease spread, and affecting the prevention and control effect. Summary of the invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a forest pest prediction system based on data analysis.
[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: A forest pest prediction system based on data analysis comprises:
[0007] The data preprocessing module extracts the temperature value, humidity value, soil moisture content, and tree growth index value in the time series based on meteorological data, soil condition data, and tree growth status data, fills in missing values, removes outliers, and generates a spatiotemporal standardized data set;
[0008] The dynamic weight allocation module analyzes the error distribution of meteorological values, soil values, tree index values and pest and disease events based on the spatiotemporal standardized data set, quantifies the impact of errors on feature weights, iteratively updates feature weight values in multiple time periods, stores the update results, and generates an adaptive feature weight matrix;
[0009] The pest feature recognition module extracts the meteorological characteristic values, soil characteristic values, and tree index values of the pest samples based on the spatiotemporal standardized data set and the adaptive characteristic weight matrix, quantifies the correlation between multiple characteristic values and the occurrence of pests and diseases, and selects key characteristic vectors of pests and diseases according to the weights;
[0010] The prediction model training module selects meteorological characteristic values, soil characteristic values, and tree index values within a time period based on the key characteristic vectors of pests and diseases and the adaptive characteristic weight matrix, gradually calculates the changing trend of weight errors, stores the corresponding results, selects the optimal weight configuration as input data, and generates dynamic prediction results of pests and diseases;
[0011] The prediction result output module extracts the prediction data and actual occurrence data of pests and diseases for each time period based on the dynamic prediction results of pests and diseases, calculates the prediction error and integrates the distribution trend, divides the spatiotemporal distribution of risk areas, and generates a pest and disease risk assessment value.
[0012] The spatiotemporal standardized data set includes temperature values, humidity values, soil moisture content, and tree growth index values; the adaptive feature weight matrix includes meteorological value weights, soil value weights, and tree index value weights; the key feature vectors of pests and diseases include meteorological feature values, soil feature values, and tree index values; the dynamic prediction results of pests and diseases include weight error change trends, optimal weight configurations, and predicted output data; the risk assessment values of pests and diseases include prediction errors, distribution trends, and spatiotemporal distribution of risk areas.
[0013] As a further solution of the present invention, the step of acquiring the spatiotemporal standardized data set is specifically:
[0014] Extract temperature and humidity values from meteorological data, soil moisture content from soil condition data, and tree growth index values from tree growth status data, screen the data parameters at each time point and space point, remove abnormal parameters that exceed a reasonable threshold range, and obtain a parameter set after removing the abnormalities;
[0015] Calling the parameter set after removing the anomalies, identifying the missing value intervals of multiple parameters in the time series, performing weighted average calculations on the data of the temperature value, humidity value, soil moisture content, and tree growth index value before and after the time point and space point by interpolation method, filling in the missing values, and generating a filled parameter set;
[0016] Based on the completed parameter set, standardize the temperature value, humidity value, soil moisture content, and tree growth index value respectively according to the time axis and spatial point coordinates, using the formula:
[0017]
[0018] Generate a spatio-temporal standardized data set;
[0019] Among them, Z i,j represents the standardized value at the i-th time point and the j-th spatial point, X i,j represents the original parameter value at the i-th time point and the j-th spatial point, μ j is the mean value of all time points at the j-th spatial point, σ j is the standard deviation of all time points at the j-th spatial point.
[0020] As a further solution of the present invention, the steps for obtaining the adaptive feature weight matrix are specifically as follows:
[0021] Extract meteorological values, soil values, and tree index values from the spatio-temporal standardized data set, combine with pest event records, calculate the error values between multi-feature parameters and pest events, group them by time period, and obtain an error distribution parameter set within the time period;
[0022] According to the error distribution parameter set within the time period, analyze the influence of the error value on the feature weight, calculate the weight adjustment factor for each feature, using the formula:
[0023]
[0024] Generate an initial feature weight vector;
[0025] Among them, W i represents the weight value of the i-th feature, E i represents the error value of the i-th feature, F i represents the influence factor of the i-th feature, and n is the total number of feature parameters;
[0026] Use the initial feature weight vector, according to the error distribution trend of multiple time periods, refer to the weight value of the previous time period and the current error data to iteratively update the feature weight, and use the weighted average method to update the weight value to generate an adaptive feature weight matrix.
[0027] As a further solution of the present invention, the steps for obtaining the key pest feature vector are specifically as follows:
[0028] Based on the spatio-temporal standardized data set, extract the corresponding meteorological feature values, soil feature values, and tree index values, and generate a pest sample feature data set for the specified pest samples;
[0029] Using the pest and disease sample feature dataset, combined with the adaptive feature weight matrix, calculate the weighted scores of the association between multiple eigenvalues and the occurrence of pests and diseases, using the formula:
[0030]
[0031] Generate a feature correlation degree score vector;
[0032] Among them, C i represents the correlation degree score of the i-th feature, W j represents the weight value of the j-th feature, V i,j represents the value of the j-th feature in the i-th sample, μ j is the mean value of all sample values of the j-th feature, used for data standardization, σ j is the standard deviation of the j-th feature, used for data standardization processing, and n is the total number of features;
[0033] According to the feature correlation degree score vector, select the feature with the highest correlation degree score, and verify that the feature is the key feature vector of pests and diseases.
[0034] As a further solution of the present invention, the steps for obtaining the dynamic prediction result of pests and diseases are specifically as follows:
[0035] Call the key feature vector of pests and diseases and the adaptive feature weight matrix, select meteorological feature values, soil feature values, and tree index values from the current spatio-temporal dataset to generate a feature set;
[0036] Conduct a weight error analysis on the feature set, calculate the change trend of the weight error of multiple eigenvalues, and update the weights, using the formula:
[0037]
[0038] Generate a weight update result;
[0039] Among them, W i,t+1 represents the weight of the i-th feature in the next time period, W i,t represents the weight of the i-th feature in the current time period t, which is the original value before weight update, ΔE i,t represents the change of weight error, E max is the maximum error within the time period, and η is the adjustment coefficient to ensure the smoothness of weight adjustment;
[0040] According to the weight update result, select the optimal weight configuration, store the weights as the input data of the prediction model, and generate the optimal weight configuration;
[0041] Using the optimal weight configuration, input it into the pest and disease prediction model, calculate the future occurrence of pests and diseases, and generate a dynamic prediction result of pests and diseases by synthesizing the data analysis results of multiple time periods.
[0042] As a further solution of the present invention, the steps for obtaining the pest and disease risk assessment value are specifically as follows:
[0043] Based on the dynamic prediction result of pests and diseases, extract the prediction data for each time period, combine it with the corresponding actual occurrence data, calculate the error between the prediction and the actual, and generate prediction error data for multiple time periods;
[0044] Analyze the prediction error data for multiple time periods, calculate the statistical distribution of the errors, including the average error and the standard deviation, using the formula:
[0045]
[0046] Generate an error trend analysis result;
[0047] where E trend represents the error trend value, P i represents the predicted value for the i-th time period, A i represents the actual value for the i-th time period, and n is the number of samples;
[0048] Using the error trend analysis result, divide the risk regions, apply a clustering algorithm to identify high-risk and low-risk regions, and generate a risk region classification result;
[0049] Based on the comprehensive risk region classification result, calculate the pest and disease risk assessment value for each region, adjust the risk assessment model according to the risk level and the error trend, and generate the pest and disease risk assessment value.
[0050] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0051] In the present invention, by analyzing the error distribution to dynamically adjust the feature weights, the influence of each feature on the pest and disease events is accurately quantified, and the self-adaptability of weight allocation is improved. Combining multi-feature correlation analysis, key features highly related to pests and diseases are screened, interference from irrelevant factors is reduced, and the quality of model input is enhanced. Dynamically calculate the changing trend of weight errors to form an optimal weight configuration in time series, and achieve high-precision prediction of the occurrence trend of pests and diseases. Based on the comparison between the prediction data and the actual occurrence data, the risk distribution trend is quantified and the risk regions are divided, providing a scientific spatio-temporal assessment result for pest and disease prevention and control, and assisting in the precise allocation of prevention and control resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is the system flow chart of the present invention;
[0053] Figure 2Flow chart of the acquisition steps of the spatio-temporal standardized data set of the present invention;
[0054] Figure 3 Flow chart of the acquisition steps of the adaptive feature weight matrix of the present invention;
[0055] Figure 4 Flow chart of the acquisition steps of the key feature vectors of pests and diseases of the present invention;
[0056] Figure 5 Flow chart of the acquisition steps of the dynamic prediction results of pests and diseases of the present invention;
[0057] Figure 6 Flow chart of the acquisition steps of the risk assessment value of pests and diseases of the present invention. Detailed implementation manners
[0058] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0059] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "a plurality of" is two or more unless otherwise specifically defined.
[0060] Embodiment 1
[0061] Please refer to Figure 1 , a forest pest and disease prediction system based on data analysis includes:
[0062] The data preprocessing module extracts temperature values, humidity values, soil moisture content, and tree growth index values in time series based on meteorological data, soil condition data, and tree growth status data, fills in missing values and eliminates outliers, and generates a spatio-temporal standardized data set;
[0063] The dynamic weight allocation module analyzes the error distribution of meteorological values, soil values, and tree index values with pest and disease events based on the spatio-temporal standardized data set, quantifies the influence of errors on feature weights, iteratively updates the feature weight values in multiple time periods, stores the update results, and generates an adaptive feature weight matrix;
[0064] The pest and disease feature recognition module extracts the meteorological characteristic values, soil characteristic values, and tree index values of pest and disease samples based on the spatiotemporal standardized data set and the adaptive characteristic weight matrix, quantifies the correlation between multiple characteristic values and the occurrence of pests and diseases, and selects the key characteristic vectors of pests and diseases according to the weights;
[0065] The prediction model training module is based on the key feature vectors of pests and diseases and the adaptive feature weight matrix. It selects the meteorological feature values, soil feature values, and tree index values within a time period, gradually calculates the changing trend of the weight error, stores the corresponding results, selects the optimal weight configuration as input data, and generates dynamic prediction results of pests and diseases.
[0066] The prediction result output module is based on the dynamic prediction results of pests and diseases, extracts the prediction data and actual occurrence data of pests and diseases for each time period, calculates the prediction error and integrates the distribution trend, divides the temporal and spatial distribution of risk areas, and generates a pest and disease risk assessment value.
[0067] The spatiotemporal standardized data set includes temperature values, humidity values, soil moisture content, and tree growth index values. The adaptive feature weight matrix includes meteorological value weights, soil value weights, and tree index value weights. The key feature vectors of pests and diseases include meteorological feature values, soil feature values, and tree index values. The dynamic prediction results of pests and diseases include the weight error change trend, optimal weight configuration, and prediction output data. The risk assessment values of pests and diseases include prediction errors, distribution trends, and spatiotemporal distribution of risk areas.
[0068] See also Figure 2 ,The specific steps for obtaining the spatiotemporal standardized dataset are:
[0069] Extract temperature and humidity values from meteorological data, soil moisture content from soil condition data, and tree growth index values from tree growth status data, screen the data parameters at each time point and space point, remove abnormal parameters that exceed a reasonable threshold range, and obtain a parameter set after removing the abnormalities;
[0070] In the meteorological data obtained through long-term monitoring at the weather station, the original temperature and humidity values collected by the sensors need to undergo preliminary screening to identify and eliminate reading errors or abnormal values. This process relies on an automated system, which will mark and eliminate data that exceeds a preset threshold (for example, the temperature exceeds 45°C or is below -10°C). In addition, for humidity values, the system will also eliminate any extreme readings below 5% or above 95%, ensuring that the resulting parameter set accurately reflects the actual climate conditions, and using this data set as the basis for subsequent processing.
[0071] Call the parameter set after excluding anomalies, identify the missing value intervals of multiple parameters in the time series, and perform weighted average operations on the data of the time points before and after and spatial points of temperature values, humidity values, soil moisture content, and tree growth index values through interpolation to complete the missing values and generate a completed parameter set;
[0072] Using the parameter set after preliminary screening, further identify the data missing problems therein, especially the data loss caused by equipment failures that may occur during continuous monitoring. By comparing the data of adjacent time points, identify continuous or discontinuous missing data points, and use linear interpolation to estimate the missing temperature values and humidity values. For example, if the data between 10:00 and 12:00 is missing, the data at 9:00 and 13:00 will be used to estimate the data between 10:00 and 12:00. This not only improves the integrity of the data but also maintains the coherence of the data, providing a more reliable data basis for subsequent analysis.
[0073] Based on the completed parameter set, standardize the temperature values, humidity values, soil moisture content, and tree growth index values respectively according to the time axis and spatial point coordinates, using the formula:
[0074]
[0075] Generate a spatio-temporal standardized data set;
[0076] Among them, Z i,j represents the standardized value at the i-th time point and the j-th spatial point, X i,j represents the original parameter value at the i-th time point and the j-th spatial point, μ j is the mean value of all time points at the j-th spatial point, and σ j is the standard deviation of all time points at the j-th spatial point.
[0077] Formula:
[0078]
[0079] The advantage of the formula is that through standardization processing, the data deviation caused by environmental differences between different geographical locations can be eliminated, enabling data from different locations to be compared and analyzed under the same standard.
[0080] Detailed explanation of the formula and the derivation process of formula calculation:
[0081] Suppose the temperature value X i,j measured by a weather station on a specific day is 15 °C, the average temperature μ j at this location is 10 °C, and the standard deviation σ j is 5 °C. According to the standardization formula calculation, the obtained standardized value is:
[0082]
[0083] The result shows that the temperature of the weather station on this day is 1 standard deviation higher than the average temperature of this location, providing a quantitative index for facilitating the analysis of how the climate conditions at a specific time at this location compare with historical data.
[0084] Please refer to Figure 3 , and the steps for obtaining the adaptive feature weight matrix are specifically as follows:
[0085] Extract meteorological values, soil values, and tree index values from the spatio-temporal standardized dataset, combine with pest and disease event records, calculate the error values between multiple feature parameters and pest and disease events, group by time period, and obtain the error distribution parameter set within the time period;
[0086] Extract meteorological values, soil values, and tree index values from the spatio-temporal standardized dataset. These data respectively represent the climate conditions, soil moisture, and growth status of trees within the region. This is to ensure that the information in the dataset can comprehensively reflect various environmental factors affecting plant growth. By combining pest and disease event records, the core of this step is to conduct a correlation analysis between environmental data and actual pest and disease events. The purpose is to find out how environmental factors affect the occurrence of pest and diseases, and which specific environmental conditions are more likely to lead to an increase in pest and diseases. Calculate the error for each feature parameter, which involves statistical analysis. By calculating the difference between the actual pest and disease events and the pest and disease events predicted by environmental parameters, determine the error value. This calculation is crucial for subsequent weight adjustment because only by accurately understanding the influence of each environmental factor on pest and disease prediction can the weights in the model be effectively adjusted to improve the prediction accuracy. Obtain the error distribution parameter set within each time period. These parameter sets will serve as the basis for adjusting the feature weights in the prediction model and provide a basis for further data processing and analysis.
[0087] According to the error distribution parameter set within the time period, analyze the influence of the error value on the feature weight, calculate the weight adjustment factor for each feature, and use the formula:
[0088]
[0089] Generate the initial feature weight vector;
[0090] Among them, W i represents the weight value of the i-th feature, E i represents the error value of the i-th feature, F i represents the influence factor of the i-th feature, and n is the total number of feature parameters;
[0091] Formula:
[0092]
[0093] The advantage of the formula is that it reduces the impact of errors through exponential weights, thus focusing more on those features with smaller errors in pest and disease events. This can reduce the sensitivity of the model to noise and enhance the response sensitivity of the model to actual effective features.
[0094] Detailed explanation of the formula and the derivation process of formula calculation:
[0095] Suppose there are three features with error values of E1 = 2, E2 = 3, and E3 = 1 respectively, and the influence factors are set as F1 = 0.5, F2 = 1.5, and F3 = 1. First, calculate the sum of the exponential negatives of the errors:
[0096]
[0097] Then calculate the weights for each feature respectively:
[0098]
[0099] The result shows that Feature 3 has obtained the largest weight because it has the smallest error with the pest and disease event. This indicates that the model will rely more on this feature for the prediction of pest and disease events. Through this weight adjustment, the model can more accurately identify and respond to those environmental factors that are more critical for pest and disease prediction.
[0100] Using the initial feature weight vector, according to the error distribution trend in multiple time periods, referring to the weight value of the previous time period and the current error data, iteratively update the feature weights, and use the weighted average method to update the weight values to generate an adaptive feature weight matrix.
[0101] Using the initial feature weight vector, the key to this step is to effectively iteratively update the weight vector to reflect the dynamic changes of environmental factors and pest and disease patterns over time. According to the error distribution trend in multiple time periods, this requires analyzing the error data within each time period to identify the changing trend of the errors. This analysis helps to understand which feature weights should be increased and which should be decreased, so that the model can adapt to the current pest and disease occurrence situation; by iteratively updating the feature weights, considering the weight value of the previous time period and the new error data in each iteration, such an iterative method can ensure that the model maintains its prediction accuracy and sensitivity in a long time series. Using the weighted average method to update the weight values, this method utilizes all available historical data to ensure that each update is based on the most comprehensive data perspective. Generate an adaptive feature weight matrix, which not only reflects the current weight status of each feature but also indicates the possible future adjustment directions, thus providing continuous optimization and adaptability for the pest and disease prediction model. This is especially important for the dynamically changing agricultural environment, ensuring that the model can flexibly respond to sudden pest and disease events and environmental changes.
[0102] Please refer to Figure 4 , and the steps for obtaining the key feature vectors of pests and diseases are specifically as follows:
[0103] Based on the spatio-temporal standardized dataset, extract the corresponding meteorological feature values, soil feature values, and tree index values, and generate a pest and disease sample feature dataset for the specified pest and disease samples;
[0104] When extracting the feature data of specific pest and disease samples, it is first necessary to call relevant data from the spatio-temporal standardized dataset. This dataset includes meteorological feature values, soil feature values, and tree index values. These data have been standardized to ensure data consistency and comparability. By precisely matching the recording time points and locations of pest and disease samples, the feature data corresponding to the samples are accurately extracted. This process ensures the precise correspondence and effectiveness of the data. Using a data integration tool, these feature values are integrated according to the samples to form a pest and disease sample feature dataset, which will be directly used for subsequent feature correlation analysis to support the acquisition of an accurate pest and disease prediction model.
[0105] Using the pest and disease sample feature dataset, combined with the adaptive feature weight matrix, calculate the weighted scores of the association between multiple feature values and the occurrence of pests and diseases, using the formula:
[0106]
[0107] Generate a feature correlation score vector;
[0108] Among them, C i represents the correlation score of the i-th feature, W j represents the weight value of the j-th feature, V i,j represents the value of the j-th feature in the i-th sample, μ j is the mean value of all sample values of the j-th feature for data standardization, σ j is the standard deviation of the j-th feature for data standardization, and n is the total number of features;
[0109] Formula:
[0110]
[0111] The benefit of the formula is that by considering the feature weights and the standardization of feature values, it enhances the accuracy and applicability of the model in identifying the associated features of pests and diseases.
[0112] Detailed explanation of the formula and the derivation process of formula calculation:
[0113] Set W j as the weight value of feature j, V i,j as the actual measured value of feature j in sample i, μj and σ j are the overall sample mean and standard deviation of feature j, respectively. Here, it is assumed that the data has been collected, and μ of each feature is obtained through data preprocessing j and σ j . Then, the standardized scores of the eigenvalues are calculated, and the weighted sum is performed by applying the feature weights. Finally, an association score C is obtained for each sample i . For example, for a certain feature, if its weight W j = 0.3, the mean μ j = 50, the standard deviation σ j = 10, and it is assumed that the actual value V of this feature in a sample i,j = 55, then the calculation process of its association score is as follows:
[0114]
[0115] This result shows that by considering the weights and standardized eigenvalues, the contribution degree of features to the occurrence of diseases and pests can be quantified
[0116] According to the feature association score vector, select the feature with the highest association score, and verify that the feature is the key feature vector of diseases and pests
[0117] When determining the key feature vector of diseases and pests, the key lies in using the feature association scores, which reflect the degree of association between each feature and the occurrence of diseases and pests. By setting the threshold of the association scores, select the features with scores higher than this threshold. These features are regarded as key features, and these features are integrated into a vector, that is, the key feature vector of diseases and pests. This vector will be used to establish a disease and pest prediction model. The determination of the key feature vector is based on the actual data analysis results, ensuring the practicability of the model and the accuracy of prediction. Through the actual data-driven method, the key features that can best reflect the impact of diseases and pests are determined, improving the efficiency and effect of disease and pest management
[0118] Please refer to Figure 5 , and the specific steps for obtaining the dynamic prediction results of diseases and pests are as follows:
[0119] Call the key feature vector of diseases and pests and the adaptive feature weight matrix, select the meteorological feature values, soil feature values, and tree index values from the current spatio-temporal dataset to generate a feature set
[0120] Based on the key feature vectors of pests and diseases and the adaptive feature weight matrix, accurately extract relevant eigenvalue from detailed meteorological, soil, and tree growth datasets, systematically analyze the contribution degree and influence of each feature on pests and diseases, use advanced data processing technologies to ensure the accuracy and integrity of the data, and generate a specific feature set by integrating these data through computer algorithms. This feature set will be directly used for subsequent weight error analysis and model input, ensuring the efficiency of the data processing process and the control of data quality.
[0121] Conduct weight error analysis on the feature set, calculate the change trend of the multi-eigenvalue weight error, update the weights, using the formula:
[0122]
[0123] Generate the weight update result;
[0124] Among them, W i,t+1 represents the weight of the i-th feature in the next time period, and W i,t represents the weight of the i-th feature in the current time period t, which is the original value before weight update. ΔE i,t represents the change in weight error, and E max is the maximum error within the time period, and η is the adjustment coefficient to ensure the smoothness of weight adjustment;
[0125] Formula:
[0126]
[0127] The benefit of the formula is that it reflects the real-time importance of different features in the prediction model by dynamically adjusting the weights, and this method enhances the adaptability and prediction accuracy of the model.
[0128] Detailed explanation of the formula and the derivation process of formula calculation:
[0129] Suppose in a specific application, the weight of a certain feature in the time period t is 0.85, the learning rate η is set to 0.05, the error change rate of this feature is 0.1, and the maximum error within this time period is 0.2. According to the formula, the calculation process is as follows:
[0130]
[0131] W i,t+1 = 0.85·(1 - 0.05·0.5)
[0132] W i,t+1 = 0.85·0.975 = 0.82875;
[0133] The results show that considering the error change rate of the current feature, its weight has been slightly reduced, which reflects the change trend of this feature in prediction. And in this way, the model can be dynamically adjusted to adapt to the changing data features.
[0134] According to the weight update results, select the weight configuration with the best performance, store the weights as the input data of the prediction model, and generate the optimal weight configuration.
[0135] Using the weight update results, finely adjust the model. Through multiple iterative calculations, optimize the weight configuration of each feature, store this data for modeling use, so as to generate a more accurate weight configuration file. This process involves advanced mathematical calculations and data analysis techniques to ensure that the selected weight configuration minimizes the prediction error, determine the weight configuration with the best performance, and record these configurations in detail for future pest and disease dynamic prediction models to ensure that the model can reflect the actual situation and accurately predict the occurrence of pests and diseases.
[0136] Input the optimal weight configuration into the pest and disease prediction model, calculate the future occurrence of pests and diseases, and generate the pest and disease dynamic prediction results by integrating the data analysis results of multiple time periods.
[0137] Combined with the optimal weight configuration, conduct a comprehensive analysis through the pest and disease prediction model. The model calculates the possible occurrence of pests and diseases according to the input optimized weights. This calculation process includes complex algorithm operations and multi-layer data processing to generate the dynamic prediction results of pests and diseases. These results are of great significance for actual agricultural production, providing a scientific basis and preventive measures to help managers make more accurate decisions and generate the pest and disease dynamic prediction results.
[0138] Please refer to Figure 6 , and the specific steps for obtaining the pest and disease risk assessment value are as follows:
[0139] Based on the pest and disease dynamic prediction results, extract the prediction data for each time period, combine with the corresponding actual occurrence data, calculate the error between the prediction and the actual, and generate the prediction error data for multiple time periods.
[0140] The difference between the prediction data and the actual occurrence data is carefully calculated and presented in the form of prediction error data. Through further analysis of these data, the accuracy of the prediction model is verified, and data support is provided for the adjustment of the model. The calculation of these data not only depends on simple difference calculations but also includes the verification of data validity and integrity to ensure the reliability of the analysis results. Through this process, the prediction model is optimized, thereby improving the prediction accuracy of future pest and disease occurrences.
[0141] Analyze the predicted error data for multiple time periods, calculate the statistical distribution of the errors, including the mean error and standard deviation, using the formula:
[0142]
[0143] Generate the error trend analysis result;
[0144] Where, E trend represents the error trend value, P i represents the predicted value for the i-th time period, A i represents the actual value for the i-th time period, and n is the number of samples;
[0145] Formula:
[0146]
[0147] The advantage of the formula is that it provides a robust method for quantifying the prediction error by calculating the average of the squared differences between the predicted and actual values. This method reduces the impact of accidental errors on the final result, thus providing a more stable error trend.
[0148] Detailed explanation of the formula and the formula calculation derivation process:
[0149] Set the number of samples n = 100 (representing data for 100 time periods), the predicted value P i and the actual value A i are obtained from the daily pest occurrence records and corresponding prediction records for the past 100 days. For example, if the predicted incidence rate on a certain day is 30%, and the actual record is 28%, then for this day, P i = 30, A i = 28. Substitute into the formula for calculation:
[0150]
[0151] The result shows that the predicted error for each data point on average is 4, which reflects the deviation degree between the prediction model and the actual situation, and helps to evaluate and adjust the accuracy of the prediction model.
[0152] Utilize the error trend analysis result to divide the risk areas, apply the clustering algorithm to identify the high-risk and low-risk areas, and generate the risk area classification result;
[0153] Based on the analysis results of the error trend, high-risk and low-risk areas are identified. Using clustering algorithms, regional markings of the geographical information system are carried out according to the risk levels. This process not only involves the collection and classification of data, but also includes spatial analysis and pattern recognition of the data. By geocoding, the risk data is combined with specific geographical locations to achieve the visualization of risks, enabling decision-makers to intuitively identify areas with concentrated risks, optimize resource allocation, and intervention measures.
[0154] Based on the comprehensive classification results of risk areas, calculate the pest and disease risk assessment value for each area. According to the risk level and error trend, adjust the risk assessment model to generate the pest and disease risk assessment value.
[0155] Based on the comprehensive classification results of risk areas, combined with geographical and previous data, calculate the pest and disease risk assessment value for each area. This calculation is not only based on the risk division in the previous steps, but also takes into account the occurrence frequency and severity of pests and diseases in each area. Through comprehensive analysis, determine the risk level of each area, providing decision-making support for subsequent resource allocation and prevention measures. This process ensures the pertinence and efficiency of pest and disease management, thus minimizing potential losses.
[0156] The above is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A forest pest prediction system based on data analysis, characterized in that, The system includes: Based on meteorological data, soil condition data, and tree growth status data, the data preprocessing module extracts temperature values, humidity values, soil moisture content, and tree growth index values under time series, fills in missing values, eliminates outliers, and generates a spatio-temporal standardized dataset. Based on the spatio-temporal standardized dataset, the dynamic weight allocation module analyzes the error distribution of meteorological values, soil values, and tree index values with respect to pest and disease events, quantifies the impact of errors on feature weights, iteratively updates the feature weight values within multiple time periods, stores the updated results, and generates an adaptive feature weight matrix. Based on the spatio-temporal standardized dataset and the adaptive feature weight matrix, the pest and disease feature recognition module extracts meteorological feature values, soil feature values, and tree index values of pest and disease samples, quantifies the correlation between multi-feature values and the occurrence of pest and diseases, and selects key pest and disease feature vectors according to weights. Based on the key pest and disease feature vectors and the adaptive feature weight matrix, the prediction model training module selects meteorological feature values, soil feature values, and tree index values within a time period, gradually calculates the changing trend of weight errors, stores the corresponding results, screens the optimal weight configuration as input data, and generates dynamic pest and disease prediction results. Based on the dynamic pest and disease prediction results, the prediction result output module extracts the predicted data and the actual occurrence data of pest and diseases for each time period, calculates the prediction error, integrates the distribution trend, divides the spatio-temporal distribution of risk areas, and generates a pest and disease risk assessment value.
2. The forest pest prediction system based on data analysis according to claim 1, wherein, The spatio-temporal standardized dataset includes temperature values, humidity values, soil moisture content, and tree growth index values. The adaptive feature weight matrix includes meteorological value weights, soil value weights, and tree index value weights. The key pest and disease feature vectors include meteorological feature values, soil feature values, and tree index values. The dynamic pest and disease prediction results include the changing trend of weight errors, the optimal weight configuration, and the predicted output data. The pest and disease risk assessment value includes the prediction error, the distribution trend, and the spatio-temporal distribution of risk areas.
3. The forest pest prediction system based on data analysis according to claim 2, wherein The specific steps for obtaining the spatio-temporal standardized dataset are as follows: Extract the temperature values and humidity values from meteorological data, the soil moisture content from soil condition data, and the tree growth index values from tree growth status data. Screen the data parameters for each time point and spatial point, and eliminate the abnormal parameters in the data that exceed the reasonable threshold range to obtain a parameter set after eliminating outliers. Call the parameter set after eliminating outliers, identify the missing value intervals of multiple parameters in the time series, and perform weighted average operations on the data of the front and back time points and spatial points of the temperature values, humidity values, soil moisture content, and tree growth index values through interpolation to fill in the missing values and generate a parameter set after filling in the missing values. Based on the parameter set after filling in the missing values, perform standardized processing on the temperature values, humidity values, soil moisture content, and tree growth index values respectively according to the time axis and spatial point coordinates, using the formula: Generate a spatio-temporal standardized dataset. Among them, Z i,j represents the standardized value at the j-th spatial point at the i-th time point, and X i,j represents the original parameter value at the j-th spatial point at the i-th time point, μ j is the mean value of all time points at the j-th spatial point, and σ j is the standard deviation of all time points at the j-th spatial point.
4. The forest pest prediction system based on data analysis according to claim 3, characterized in that The specific steps for obtaining the adaptive feature weight matrix are as follows: Extract meteorological values, soil values, and tree index values from the spatio-temporal standardized dataset, combine with pest and disease event records, calculate the error values between multi-characteristic parameters and pest and disease events, group by time period, and obtain the error distribution parameter set within the time period; According to the error distribution parameter set within the time period, analyze the influence of the error value on the characteristic weight, calculate the weight adjustment factor for each characteristic, and use the formula: Generate an initial characteristic weight vector; Among them, W i represents the weight value of the i-th feature, E i represents the error value of the i-th feature, F i represents the influence factor of the i-th feature, and n is the total number of feature parameters; Use the initial characteristic weight vector, according to the error distribution trend of multiple time periods, refer to the weight value of the previous time period and the current error data to iteratively update the characteristic weight, and use the weighted average method to update the weight value to generate an adaptive characteristic weight matrix.
5. The forest pest prediction system based on data analysis according to claim 4, characterized in that, The specific steps for obtaining the key pest and disease characteristic vector are as follows: Based on the spatio-temporal standardized dataset, extract the corresponding meteorological characteristic values, soil characteristic values, and tree index values, and generate a pest and disease sample characteristic dataset for the specified pest and disease samples; Using the pest and disease sample characteristic dataset, combine with the adaptive characteristic weight matrix, calculate the weighted scores of the association between multi-characteristic values and the occurrence of pests and diseases, and use the formula: Generate a characteristic correlation score vector; Among them, C i represents the correlation score of the i-th feature, W j represents the weight value of the j-th feature, V i,j represents the value of the j-th feature in the i-th sample, μ j is the mean of all sample values of the j-th feature, used for data standardization, σ j is the standard deviation of the j-th feature, used for data standardization, and n is the total number of features; According to the characteristic correlation score vector, select the characteristic with the highest correlation score, and verify that the characteristic is the key pest and disease characteristic vector.
6. The forest pest prediction system based on data analysis according to claim 5, characterized in that The specific steps for obtaining the dynamic prediction result of pests and diseases are as follows: Call the key pest and disease characteristic vector and the adaptive characteristic weight matrix, select meteorological characteristic values, soil characteristic values, and tree index values from the current spatio-temporal dataset to generate a characteristic set; Conduct a weight error analysis on the characteristic set, calculate the change trend of the weight error of multi-characteristic values, and update the weight, using the formula: Generate a weight update result; Among them, W i,t+1 represents the weight of the i-th feature in the next time period, and W i,t represents the weight of the i-th feature in the current time period t, which is the original value before weight update. ΔE i,t represents the change in weight error, and E max is the maximum error within the time period, and η is the adjustment coefficient to ensure the smoothness of weight adjustment; According to the weight update result, select the optimal weight configuration, store the weight as the input data of the prediction model, and generate the optimal weight configuration; Use the optimal weight configuration, input it into the pest and disease prediction model, calculate the future occurrence situation of pests and diseases, and comprehensively analyze the data results of multiple time periods to generate a dynamic prediction result of pests and diseases.
7. The forest pest prediction system based on data analysis according to claim 6, wherein, The specific steps for obtaining the pest and disease risk assessment value are as follows: Based on the dynamic prediction result of pests and diseases, extract the prediction data for each time period, combine with the corresponding actual occurrence data, calculate the error between the prediction and the actual, and generate the prediction error data for multiple time periods; Analyze the prediction error data for multiple time periods, calculate the statistical distribution of the error, including the average error and the standard deviation, using the formula: Generate an error trend analysis result; Among them, E trend represents the error trend value, P i represents the predicted value in the i-th time period, A i represents the actual value in the i-th time period, and n is the number of samples; Use the error trend analysis result to divide the risk areas, apply the clustering algorithm to identify high-risk and low-risk areas, and generate a risk area classification result; Comprehensively consider the risk area classification result, calculate the pest and disease risk assessment value for each area, adjust the risk assessment model according to the risk level and the error trend, and generate the pest and disease risk assessment value.
Citation Information
Cited By
Landscaping plant disease and insect pest distribution data statistical method
CN121210540A
Forestry disease and pest risk prediction method and system based on big data
CN121352501A