Forestry pest and disease prediction method and system based on artificial intelligence
By obtaining dynamic monitoring data in forest areas, performing feature extraction and pest prediction model analysis, and generating a control optimization instruction set, it solves the problems of insufficient data for pest prediction and lack of control strategies in the existing technology, and achieves efficient, precise control and resource optimization of forestry pests and diseases.
Patent Information
- Application Number
- CN202510759296.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-06-09
AI Technical Summary
The existing forestry pest and disease prediction technologies have problems such as few monitoring points, single data types, weak data processing capabilities, low generalization capabilities of prediction models, and inability to provide specific control strategies and resource scheduling solutions, resulting in low prevention and control efficiency, high cost and poor effect.
By obtaining the dynamic monitoring data set of the target forest area, feature extraction is performed, forestry environment feature collection is generated, and abnormal parameters are analyzed using the pre-trained pest prediction model, and a set of pest prediction parameters are generated. Combined with the prevention and control priority strategy, forest area prevention and control optimization instruction set is generated to realize the full process automation and intelligent resource scheduling.
It has achieved comprehensive, accurate and efficient prediction and control guidance on pest and diseases, improved prevention and control efficiency and effectiveness, provided scientific decision-making support, and improved resource utilization efficiency.
Smart Images

Figure CN120278346B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based forestry pest and disease prediction method and system. Background Art
[0002] Forestry is of great significance to ecological balance and resource supply, but pests and diseases are important factors threatening its healthy development. Traditional prevention and control relies on manual labor and experience, which is inefficient, difficult to detect early, misses the best opportunity, is costly and has poor results. There is an urgent need to accurately predict pests and diseases.
[0003] There are many deficiencies in the existing forestry pest and disease prediction technology. In terms of data collection, there are few monitoring points, limited scope, and single data types, mostly common environmental parameters. There is a lack of data related to vegetation physiology and the environment and pests and diseases, which affects the accuracy of predictions. When processing data, the feature extraction and analysis capabilities are weak, and surface features are mostly extracted. It is difficult to dig deep information and cannot capture early signals and potential patterns. Prediction models are mostly based on traditional statistics or simple machine learning algorithms, which are difficult to process complex forestry environmental data and rely on large amounts of historical data. Due to the occasional outbreak of pests and diseases, the integrity and accuracy of the data are difficult to guarantee, and the model generalization ability and prediction credibility are low. In addition, existing technologies focus on prediction results and neglect prevention and control decision support. They are unable to provide specific prevention and control strategies and resource scheduling plans, which increases the difficulty and uncertainty of prevention and control, resulting in waste of resources and poor results. Summary of the Invention
[0004] In view of the above-mentioned problems, in combination with the first aspect of the present invention, an embodiment of the present invention provides a forestry pest and disease prediction method based on artificial intelligence, the method comprising:
[0005] Acquire a dynamic monitoring data set of a target forest area, wherein the dynamic monitoring data set includes an environmental parameter sequence and a vegetation physiological parameter sequence collected by multiple monitoring nodes, wherein each monitoring node corresponds to a spatial coordinate region;
[0006] Extracting features from the dynamic monitoring data set to obtain a forestry environment feature set for each monitoring node, wherein the forestry environment feature set includes vegetation growth state features, environmental fluctuation correlation features, and potential identification features of pests and diseases;
[0007] Based on the pre-trained pest and disease prediction model, the forestry environment feature set is subjected to abnormal parameter analysis processing to generate a pest and disease prediction parameter set for the monitoring node, wherein the pest and disease prediction parameter set is used to characterize the probability of occurrence of pests and diseases and the level of the impact range;
[0008] Performing a dynamic prediction matching operation based on the pest and disease prediction parameter set to determine a pest and disease distribution heat map of the target forest area and generate a pest and disease control priority strategy;
[0009] The pest distribution heat map is integrated with the pest control priority strategy to generate a forest area control optimization instruction set, and the forest area control optimization instruction set is fed back to the forestry monitoring platform to trigger the control resource scheduling operation.
[0010] On the other hand, an embodiment of the present invention also provides an artificial intelligence-based forestry pest and disease prediction system, including a processor and a machine-readable storage medium, wherein the machine-readable storage medium is connected to the processor, the machine-readable storage medium is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the machine-readable storage medium to implement the above method.
[0011] Based on the above aspects, the embodiments of the present invention realize comprehensive, accurate and efficient prediction and prevention guidance of the pest and disease situation in the target forest area. First, a dynamic monitoring data set of the target forest area is obtained, covering the environmental parameter sequence and vegetation physiological parameter sequence of multiple monitoring nodes, and each monitoring node corresponds to a specific spatial coordinate area. When the dynamic monitoring data set is feature extracted to obtain a forestry environment feature set, not only attention is paid to the vegetation growth state characteristics, but also in-depth exploration of environmental fluctuation correlation characteristics and potential identification characteristics of pests and diseases, and a comprehensive portrayal of the forestry environment conditions from multiple dimensions. Based on the pre-trained pest and disease prediction model, the forestry environment feature set is subjected to abnormal parameter analysis processing to generate a pest and disease prediction parameter set. The pest and disease prediction parameter set can accurately characterize the probability of occurrence of pests and diseases and the level of the impact range. The model algorithm is used to perform in-depth analysis and anomaly identification of the features, so that pest and disease prediction is no longer limited to simple qualitative judgment, but achieves quantitative and accurate evaluation. Furthermore, dynamic prediction and matching operations are performed based on the set of pest and disease prediction parameters to determine the pest and disease distribution heat map of the target forest area and generate a pest and disease control priority strategy. The discrete prediction parameters are converted into an intuitive distribution heat map, which clearly shows the distribution of pests and diseases in the forest area. At the same time, combined with the prevention and control priority strategy, prevention and control work can be targeted, improving prevention and control efficiency and effectiveness. Finally, the pest and disease distribution heat map is integrated with the pest and disease control priority strategy to generate a forest area prevention and control optimization instruction set, which is fed back to the forestry monitoring platform to trigger the prevention and control resource scheduling operation. This realizes the automation and intelligence of the entire process from prediction to prevention and control decision-making to resource scheduling. This not only improves the accuracy and timeliness of forestry pest and disease prediction, but also provides scientific and precise decision-making support for forestry prevention and control work, effectively improving the overall level of forestry pest and disease control and resource utilization efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a schematic diagram of the execution flow of the forestry pest and disease prediction method based on artificial intelligence provided by an embodiment of the present invention.
[0013] Figure 2 Schematic diagram of an artificial intelligence-based forestry pest prediction system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0014] The present invention will be described in detail below with reference to the accompanying drawings. Figure 1 This is a flow chart of an artificial intelligence-based forestry pest prediction method provided by an embodiment of the present invention. The artificial intelligence-based forestry pest prediction method is introduced in detail below.
[0015] Step S110: Acquire a dynamic monitoring data set of the target forest area, wherein the dynamic monitoring data set includes an environmental parameter sequence and a vegetation physiological parameter sequence collected by multiple monitoring nodes, wherein each monitoring node corresponds to a spatial coordinate area.
[0016] In this embodiment, in order to achieve effective prediction of pests and diseases in the target forest area, it is first necessary to obtain a dynamic monitoring data set of the target forest area. In the target forest area, a plurality of monitoring nodes are pre-arranged. These monitoring nodes are distributed in different positions of the forest area, and each monitoring node corresponds to a specific spatial coordinate area. The monitoring nodes will continuously collect environmental parameter sequences and vegetation physiological parameter sequences. The environmental parameter sequence includes data such as temperature, humidity, light intensity, and soil pH, while the vegetation physiological parameter sequence includes information such as chlorophyll content, trunk water saturation, and canopy density. For example, in a target forest area with a larger area, 100 monitoring nodes are evenly distributed, and each monitoring node collects environmental parameters and vegetation physiological parameters every hour. For example, at monitoring node A, the temperature sequence collected over the course of a day is [20°C, 22°C, 23°C, …, 21°C], the humidity sequence is [60%, 62%, 61%, …, 63%], and the chlorophyll content sequence is [0.5 mg / g, 0.52 mg / g, 0.51 mg / g, …, 0.53 mg / g]. By collecting this data, a dynamic monitoring data set containing data from multiple monitoring nodes is formed.
[0017] Step S120: extracting features from the dynamic monitoring data set to obtain a forestry environment feature set for each monitoring node, wherein the forestry environment feature set includes vegetation growth state features, environmental fluctuation correlation features, and potential identification features of pests and diseases.
[0018] In this embodiment, a feature extraction operation is performed on the acquired dynamic monitoring data set to obtain a forestry environment feature set for each monitoring node. This process can be further divided into the following sub-steps.
[0019] Step S121: aligning the environmental parameter sequence in time dimension to generate a standardized environmental parameter sequence.
[0020] In this embodiment, since there may be certain differences in the time when different monitoring nodes collect data, in order to facilitate subsequent analysis and processing, it is necessary to align the environmental parameter sequence in the time dimension. For example, the time intervals for monitoring node A and monitoring node B to collect data may not be completely consistent. Monitoring node A collects data every 1 hour, while monitoring node B collects data every 1.5 hours. Through time dimension alignment, the data of the two monitoring nodes are unified to the same time scale. The specific approach is to use the shorter time interval as a benchmark and perform interpolation processing on the data collected at a longer time interval. Assuming that the time interval is 1 hour, for the data collected by monitoring node B every 1.5 hours, the data value at 1 hour is calculated by linear interpolation. After such processing, a standardized environmental parameter sequence is generated, making the environmental parameter data of different monitoring nodes comparable in time dimension.
[0021] Step S122: calling a preset vegetation physiological characteristic encoder to perform physiological state analysis on the vegetation physiological parameter sequence to generate a vegetation growth state feature vector, wherein the vegetation growth state feature vector includes the chlorophyll content change rate, trunk water saturation and canopy density attenuation coefficient.
[0022] Then, the preset vegetation physiological characteristic encoder is called to process the vegetation physiological parameter sequence. The vegetation physiological characteristic encoder is a trained model that can parse the physiological status information of the vegetation based on the input vegetation physiological parameter sequence. For example, for the chlorophyll content sequence, the chlorophyll content change rate is obtained by calculating the ratio of the chlorophyll content difference between adjacent time points and the chlorophyll content at the previous moment. Assuming that within a certain time period, the chlorophyll content sequence of monitoring node A is [0.5mg / g, 0.52mg / g, 0.51mg / g], the chlorophyll content change rate of the first time interval is (0.52-0.5) / 0.5=0.04. Similarly, for the trunk water saturation and canopy density data, after processing by the encoder, the trunk water saturation and canopy density attenuation coefficients are obtained. Finally, these features are combined into a vegetation growth state feature vector, such as [0.04, 0.8, 0.02], where 0.04 represents the chlorophyll content change rate, 0.8 represents the trunk water saturation, and 0.02 represents the canopy density attenuation coefficient.
[0023] Step S123: performing environmental correlation analysis on the standardized environmental parameter sequence to extract environmental fluctuation correlation features, wherein the environmental fluctuation correlation features include the temperature and humidity coordinated change gradient, soil pH offset, and light intensity cumulative deviation value.
[0024] Furthermore, the standardized environmental parameter sequence is subjected to environmental correlation analysis to extract environmental fluctuation correlation features. This process can be carried out according to the following specific steps.
[0025] Step S1231: Acquire the temperature sequence, humidity sequence, and light intensity sequence in the standardized environmental parameter sequence.
[0026] In this embodiment, the temperature sequence, humidity sequence, and light intensity sequence are extracted from the standardized environmental parameter sequence. For example, in the standardized environmental parameter sequence of monitoring node A, the temperature sequence is [20°C, 22°C, 23°C, 21°C], the humidity sequence is [60%, 62%, 61%, 63%], and the light intensity sequence is [1000lux, 1100lux, 1050lux, 1080lux].
[0027] Step S1232: Calculate the dynamic covariance matrix between the temperature sequence and the humidity sequence, and determine the temperature-humidity coordinated change gradient, wherein the temperature-humidity coordinated change gradient is quantified by the covariance change rate within the sliding time window.
[0028] In this embodiment, a sliding time window method can be used, assuming that the window size is 3 time points. For the temperature sequence [20°C, 22°C, 23°C, 21°C] and the humidity sequence [60%, 62%, 61%, 63%], the temperature data in the first window is [20°C, 22°C, 23°C], and the humidity data is [60%, 62%, 61%]. The covariance of the two sequences is calculated. Then, the window is slid back one time point and the covariance in the second window is calculated. By comparing the covariance values of adjacent windows, the covariance change rate is obtained to quantify the temperature and humidity coordinated change gradient. For example, if the covariance of the first window is 0.5 and the covariance of the second window is 0.6, the covariance change rate is (0.6-0.5) / 0.5=0.2, which is the temperature and humidity coordinated change gradient in this time period.
[0029] Step S1233: performing trend fitting processing on the soil pH data in the standardized environmental parameter sequence to extract the soil pH offset, where the soil pH offset is determined by the standard deviation multiple of the current pH value and the historical reference value.
[0030] In this embodiment, first, a historical baseline value is established. This historical baseline value can be the average value of soil pH over a period of time. For example, the average soil pH value of monitoring node A over the past month is 7.0. Then, the current soil pH data is fitted to obtain the current pH value. Assuming the current pH value is 7.2, the difference between the current value and the historical baseline value is calculated, that is, 7.2-7.0=0.2. Then, the standard deviation of the historical data is calculated, assuming the standard deviation is 0.1. The soil pH offset is 0.2 / 0.1=2 times the standard deviation.
[0031] Step S1234: performing cumulative deviation calculation on the illumination intensity sequence to generate a cumulative deviation value of illumination intensity, where the cumulative deviation value is the cumulative difference between the measured illumination intensity value and the theoretical value within a continuous time interval.
[0032] In this embodiment, the theoretical light intensity value can be calculated based on factors such as the local geographical location and time. For example, based on the geographical location and time of monitoring node A, the theoretical light intensity sequence is calculated to be [980lux, 1020lux, 1060lux, 1040lux], while the measured light intensity sequence is [1000lux, 1100lux, 1050lux, 1080lux]. The difference between the measured value and the theoretical value at each time point is calculated as (1000-980) = 20lux, (1100-1020) = 80lux, (1050-1060) = -10lux, and (1080-1040) = 40lux. These differences are then accumulated to obtain a cumulative deviation value of light intensity of 20+80-10+40 = 130lux.
[0033] Step S1235: normalize the temperature and humidity coordinated change gradient, soil pH offset, and light intensity cumulative deviation value to obtain the environmental fluctuation correlation feature.
[0034] In this embodiment, the purpose of normalization is to map these eigenvalues into a unified range to facilitate subsequent analysis and processing. For example, using the minimum-maximum normalization method, assuming that the temperature-humidity synergistic variation gradient ranges from [0, 1], the soil pH offset ranges from [0, 5], and the light intensity cumulative deviation ranges from [0, 200]. The temperature-humidity synergistic variation gradient of 0.2, the soil pH offset of 2, and the light intensity cumulative deviation of 130 are normalized to obtain the normalized environmental fluctuation correlation eigenvector, such as [0.2, 0.4, 0.65].
[0035] Step S124: performing abnormal identification fusion processing on the vegetation growth state feature vector and the environmental fluctuation correlation feature to generate potential identification features of pests and diseases, wherein the potential identification features of pests and diseases include an estimated value of insect egg distribution density, a fungal infection diffusion rate, and an intensity of disease symptom manifestation.
[0036] Next, the vegetation growth state feature vector and the environmental fluctuation correlation feature are fused to generate potential pest and disease identification features. This process can be carried out in the following specific steps.
[0037] For example, step S1241: discretize and segment the chlorophyll content change rate, trunk water saturation and canopy density attenuation coefficient in the vegetation growth state characteristic vector to generate chlorophyll content change rate intervals, trunk water saturation intervals and canopy density attenuation coefficient intervals.
[0038] In this embodiment, each feature in the vegetation growth state feature vector is discretized and segmented. For example, for the chlorophyll content change rate, its value range is divided into multiple intervals, such as [0, 0.02), [0.02, 0.05), [0.05, +∞). Assuming that the chlorophyll content change rate of monitoring node A is 0.04, it falls within the interval [0.02, 0.05). Similarly, the trunk water saturation and canopy density attenuation coefficient are similarly segmented to obtain the trunk water saturation interval and canopy density attenuation coefficient interval.
[0039] Step S1242: determining an estimated value of the insect egg distribution density according to a mapping relationship between the temperature and humidity coordinated change gradient in the environmental fluctuation correlation feature and the chlorophyll content change rate interval.
[0040] In this embodiment, a mapping table can be pre-established to record estimated values of insect egg distribution density for different combinations of temperature-humidity synergistic gradients and chlorophyll content change rate intervals. For example, when the temperature-humidity synergistic gradient is 0.2 and the chlorophyll content change rate interval is [0.02, 0.05), by querying the mapping table, the estimated insect egg distribution density is 50 eggs / m².
[0041] Step S1243: Based on the soil pH offset and the light intensity cumulative deviation value in the environmental fluctuation correlation feature, combined with a predefined fungal growth threshold range, a fungal infection diffusion rate is generated.
[0042] In this embodiment, first, the suitable soil pH range and light intensity range for fungal growth are determined. For example, the suitable soil pH range for fungal growth is [6.5, 7.5], and the suitable light intensity range is [800lux, 1200lux]. Then, the current environment is judged to be suitable for fungal growth based on the soil pH offset and the light intensity cumulative deviation value. If the soil pH offset is small and the light intensity cumulative deviation value is also within a certain range, it is considered that the environment is suitable for fungal growth, and the fungal infection diffusion rate is calculated according to a pre-established model. Assuming that the soil pH offset is 2 times the standard deviation and the light intensity cumulative deviation value is 130lux, the fungal infection diffusion rate calculated by the model is 0.1m² / day.
[0043] Step S1244: Determine the intensity of disease symptoms based on the ratio of overlapping areas between the trunk water saturation interval and the canopy density attenuation coefficient interval.
[0044] In this embodiment, the length of the overlapping area of the two intervals is first calculated, and then divided by the total length of the two intervals to obtain the overlapping area ratio. For example, if the trunk water saturation interval is [0.7, 0.9] and the canopy density attenuation coefficient interval is [0.01, 0.03], the overlapping area length is 0, and the total length is (0.9-0.7)+(0.03-0.01)=0.22, then the overlapping area ratio is 0 / 0.22=0. According to the pre-established mapping relationship, when the overlapping area ratio is 0, the intensity of disease symptoms is weak.
[0045] Step S1245: performing normalized weighted concatenation on the estimated value of the insect egg distribution density, the fungal infection diffusion rate, and the intensity of the disease symptoms to generate the potential identification features of the pests and diseases.
[0046] In this example, the estimated egg density, fungal infection spread rate, and disease symptom intensity are first normalized and mapped to the range [0, 1]. For example, an estimated egg density of 50 eggs / m² is normalized to 0.5; a fungal infection spread rate of 0.1 m² / day is normalized to 0.3; and weak disease symptom intensity is represented by 0.1. Next, weights are assigned to different features based on their importance. For example, the estimated egg density is weighted 0.5, the fungal infection spread rate is weighted 0.3, and the disease symptom intensity is weighted 0.2. The normalized feature values are multiplied by the weights and then concatenated to obtain a potential pest and disease identification feature vector, such as [0.5×0.5, 0.3×0.3, 0.2×0.1] = [0.25, 0.09, 0.02].
[0047] Step S125: Integrate the vegetation growth state characteristics, the environmental fluctuation correlation characteristics, and the potential identification characteristics of pests and diseases to form the forestry environment characteristic set.
[0048] In this embodiment, for example, the vegetation growth state feature vector is [0.04, 0.8, 0.02], the environmental fluctuation association feature vector is [0.2, 0.4, 0.65], and the pest and disease potential identification feature vector is [0.25, 0.09, 0.02]. They are spliced together to form a forestry environment feature set, such as [0.04, 0.8, 0.02, 0.2, 0.4, 0.65, 0.25, 0.09, 0.02].
[0049] Step S130: Based on the pre-trained pest and disease prediction model, the forestry environment feature set is subjected to abnormal parameter analysis processing to generate a pest and disease prediction parameter set of the monitoring node, wherein the pest and disease prediction parameter set is used to characterize the probability of occurrence of pests and diseases and the level of impact range.
[0050] In this embodiment, a pre-trained pest prediction model is used to process the forestry environment feature set to generate a pest prediction parameter set for the monitoring node. This process can be further divided into the following sub-steps.
[0051] Step S131: inputting the forestry environment feature set into the feature coding layer of the pest and disease prediction model to generate a high-dimensional fusion feature vector.
[0052] In this embodiment, the feature encoding layer is an integral part of the model that converts low-dimensional input features into high-dimensional features to better capture the information in the data. For example, if the forestry environment feature set is [0.04, 0.8, 0.02, 0.2, 0.4, 0.65, 0.25, 0.09, 0.02], after processing by the feature encoding layer, a high-dimensional fused feature vector is generated, such as [0.1, 0.2, 0.3, ..., 0.05]. The dimension of this vector may be much higher than the dimension of the input forestry environment feature set.
[0053] Step S132: Calculate the deviation coefficient between the high-dimensional fusion feature vector and the pre-stored healthy vegetation feature template through the anomaly detection layer of the pest prediction model, where the deviation coefficient is determined by the weighted sum of the Euclidean distance and the cosine similarity.
[0054] Next, the anomaly detection layer of the pest and disease prediction model calculates the deviation coefficient between the high-dimensional fusion feature vector and the pre-stored healthy vegetation feature template. The specific steps are as follows.
[0055] Step S1321: Obtain the standard chlorophyll content, standard trunk water saturation, and standard canopy density in the healthy vegetation feature template.
[0056] In this embodiment, the standard chlorophyll content, standard trunk water saturation, and standard canopy density are obtained from a pre-stored healthy vegetation feature template. For example, the standard chlorophyll content recorded in the healthy vegetation feature template is 0.6 mg / g, the standard trunk water saturation is 0.9, and the standard canopy density is 0.8.
[0057] Step S1322: Calculate the first absolute value of the difference between the chlorophyll content change rate in the high-dimensional fusion feature vector and the standard chlorophyll content, and the second absolute value of the difference between the trunk water saturation and the standard trunk water saturation.
[0058] In this example, the chlorophyll content change rate and trunk water saturation information are extracted from the high-dimensional fused feature vector. Assume that the element value corresponding to the chlorophyll content change rate in the high-dimensional fused feature vector is 0.04, and the element value corresponding to the trunk water saturation is 0.8. The absolute value of the first difference is calculated as |0.04 - 0.6| = 0.56, and the absolute value of the second difference is |0.8 - 0.9| = 0.1.
[0059] Step S1323: normalize the first absolute value of the difference and the second absolute value of the difference to generate a primary deviation index.
[0060] Normalize the first and second absolute differences. Assuming the chlorophyll content change rate ranges from [0, 0.1] and the trunk water saturation ranges from [0, 1]. Normalize the first and second absolute differences, 0.56, to 1 (the maximum value), since 0.56 exceeds the chlorophyll content change rate range of [0, 0.1]. The second absolute difference, 0.1, remains within the trunk water saturation range of [0, 1] and is therefore still 0.1 after normalization. Then, take the weighted average of these two normalized values. Assuming the weight of the chlorophyll content change rate is 0.6 and the weight of the trunk water saturation is 0.4, the primary deviation index is 1 × 0.6 + 0.1 × 0.4 = 0.64.
[0061] Step S1324: extracting the canopy density attenuation coefficient from the high-dimensional fusion feature vector, calculating the inverse of its ratio to the standard canopy density, and generating a secondary deviation index.
[0062] In this example, assume that the element value corresponding to the canopy density attenuation coefficient in the high-dimensional fused feature vector is 0.02, and the standard canopy density is 0.8. The calculated ratio of the canopy density attenuation coefficient to the standard canopy density is 0.02 / 0.8 = 0.025, and the inverse of this ratio is 1 / 0.025 = 40. This 40 is the secondary deviation index.
[0063] Step S1325: linearly combine the primary deviation index and the secondary deviation index to obtain a comprehensive deviation coefficient, wherein the combination weight of the linear combination is dynamically adjusted according to the importance of features in the historical pest and disease data.
[0064] Consider a linear combination of the primary deviation index of 0.64 and the secondary deviation index of 40. Assume that, based on feature importance analysis of historical pest and disease data, the primary deviation index is weighted at 0.3 and the secondary deviation index is weighted at 0.7. The combined deviation coefficient is 0.64 × 0.3 + 40 × 0.7 = 0.192 + 28 = 28.192.
[0065] Step S133: matching a preset pest and disease type mapping table according to the deviation coefficient to determine the pest and disease occurrence probability, wherein the pest and disease occurrence probability is exponentially related to the deviation coefficient.
[0066] In this embodiment, the preset pest type mapping table records the pest occurrence probability corresponding to different deviation coefficient ranges. For example, when the deviation coefficient is between 0-10, the pest occurrence probability is 10%; when the deviation coefficient is between 10-20, the pest occurrence probability is 30%; when the deviation coefficient is between 20-30, the pest occurrence probability is 60%; and when the deviation coefficient is greater than 30, the pest occurrence probability is 90%. Since the comprehensive deviation coefficient is 28.192, which falls within the range of 20-30, the pest occurrence probability for this monitoring node is determined to be 60%.
[0067] Step S134: calling the impact range calculation module of the pest prediction model to calculate the pest impact range level based on the spatial correlation parameters in the high-dimensional fusion feature vector, where the pest impact range level is determined by the feature similarity chain propagation rate of adjacent monitoring nodes.
[0068] In this embodiment, first, spatial correlation parameters are extracted from the high-dimensional fusion feature vector. These spatial correlation parameters may include information such as the spatial coordinates of the monitoring node and the distance to the adjacent monitoring nodes. Assuming that there are three adjacent monitoring nodes B, C, and D around the monitoring node A, the feature similarity between the monitoring node A and the adjacent monitoring nodes is calculated. The feature similarity can be obtained by calculating the cosine similarity between the high-dimensional fusion feature vectors. For example, the high-dimensional fusion feature vector of monitoring node A is [0.1, 0.2, 0.3, ..., 0.05], and the high-dimensional fusion feature vector of monitoring node B is [0.12, 0.22, 0.28, ..., 0.06], and the cosine similarity between them is calculated to be 0.9. Similarly, the feature similarities of monitoring node A and monitoring nodes C and D are calculated to be 0.8 and 0.7 respectively.
[0069] Then, the chain propagation rate is determined based on the feature similarity. Assuming the feature similarity is greater than 0.8, the chain propagation rate is one adjacent monitoring node per day; between 0.6 and 0.8, the chain propagation rate is one adjacent monitoring node every two days; and when the feature similarity is less than 0.6, the chain propagation rate is one adjacent monitoring node every three days. For monitoring node A and its adjacent monitoring nodes, the chain propagation rate with node B is one adjacent monitoring node per day, with node C it is one adjacent monitoring node every two days, and with node D it is one adjacent monitoring node every two days.
[0070] Based on the chain transmission rate and the distribution of adjacent monitoring nodes, the spread of pests and diseases within a certain period of time is predicted. Assuming a one-week time period, node B will be infected within a week, and nodes C and D have a 50% chance of being infected within a week. Based on these predictions, the impact range of the pest and disease is determined. If the impact range is small, affecting only a few adjacent monitoring nodes, the impact range is low; if the impact range is large, involving many adjacent monitoring nodes, the impact range is high. In this example, the impact range of the pest and disease is medium.
[0071] Step S135: combining and packaging the pest occurrence probability and impact range level to generate the pest prediction parameter set.
[0072] The calculated pest and disease occurrence probability of 60% and the impact range level of "medium" are combined and encapsulated to generate a pest and disease prediction parameter set. For example, it can be expressed as [60%, "medium"]. This pest and disease prediction parameter set is used to represent the pest and disease occurrence probability and impact range level of the monitoring node.
[0073] Step S140: performing a dynamic prediction matching operation based on the pest and disease prediction parameter set, determining a pest and disease distribution heat map of the target forest area, and generating a pest and disease control priority strategy.
[0074] This step requires dynamic prediction and matching operations based on the pest and disease prediction parameter set to determine the pest and disease distribution heat map of the target forest area and generate a pest and disease control priority strategy. This process can be broken down into the following sub-steps.
[0075] Step S141: extracting the probability of occurrence of pests and diseases and the level of impact range from the pest and disease prediction parameter set.
[0076] In this embodiment, the pest and disease occurrence probability and impact level are extracted from the previously generated pest and disease prediction parameter set. For example, for the pest and disease prediction parameter set [60%, "Medium"] for monitoring node A, the pest and disease occurrence probability of 60% and the impact level of "Medium" are extracted. This extraction operation is performed on the pest and disease prediction parameter sets of all monitoring nodes in the target forest area to obtain the pest and disease occurrence probability and impact level data for each monitoring node.
[0077] Step S142: constructing a three-dimensional geographic information grid based on the spatial coordinate area of the monitoring node, and mapping the probability of pest and disease occurrence of each grid node to the corresponding spatial coordinate.
[0078] A 3D geographic information grid is constructed based on the spatial coordinates of the monitoring nodes. First, the geographic scope of the target forest area is determined. For example, the target forest area is 1000 meters long, 800 meters wide, and 50 meters high (accounting for certain vertical spacing, such as tree height). This area is divided into 3D grids, for example, each grid is 10 meters x 10 meters x 5 meters.
[0079] Next, map the spatial coordinates of each monitoring node to a grid node in the three-dimensional geographic information grid. For example, if monitoring node A has spatial coordinates (200, 300, 10), it corresponds to the grid node in row 20, column 30, and layer 2 of the three-dimensional geographic information grid. Map the 60% probability of pest and disease occurrence for that monitoring node to this corresponding grid node. Repeat this mapping operation for all monitoring nodes, ensuring that each grid node has a corresponding probability of pest and disease occurrence.
[0080] Step S143: performing spatial interpolation processing on the discrete probability of occurrence of pests and diseases based on the Kriging interpolation algorithm to generate a continuous probability distribution surface.
[0081] In this step, the Kriging interpolation algorithm is used to perform spatial interpolation on the discrete probability of occurrence of pests and diseases. The specific steps are as follows.
[0082] Step S1431: Generate a spatial distance matrix between monitoring nodes based on the spatial coordinates of the monitoring nodes.
[0083] In this embodiment, the spatial distance between any two monitoring nodes is calculated based on their spatial coordinates. Assuming that the spatial coordinates of monitoring node A are (200, 300, 10) and the spatial coordinates of monitoring node B are (210, 310, 12), the distance between them is calculated using a spatial distance calculation formula (such as the three-dimensional Euclidean distance formula) as: square root of ((210 - 200) squared + (310 - 300) squared + (12 - 10) squared) = square root of (100 + 100 + 4) = square root of 204 ≈ 14.28 meters.
[0084] This distance calculation is performed on all monitoring nodes to obtain a spatial distance matrix between monitoring nodes. For example, there are three monitoring nodes A, B, and C, and the spatial distance matrix is:
[0085] [0, 14.28, 20.5],
[0086] [14.28, 0, 18.3],
[0087] [20.5, 18.3, 0]
[0088] Step S1432: Calculate the semivariogram model parameters according to the spatial distance matrix and construct a semivariogram model.
[0089] The semivariogram describes the variability of spatial data. The parameters of the semivariogram model are calculated based on the spatial distance matrix and the corresponding pest and disease probability data. First, the monitoring nodes are grouped by distance and the average of the squared differences in pest and disease probability for each monitoring node within each group is calculated. For example, for monitoring nodes within a distance range of 0-10 meters, the average of the squared differences in pest and disease probability is calculated.
[0090] Next, a semivariogram model is fitted based on these calculated results. Common semivariogram models include spherical and exponential models. Assuming a spherical model is used, the model parameters, such as the nugget value, sill value, and range, are obtained through fitting. Assume that the fitted spherical model parameters are: nugget value of 0.05, sill value of 0.3, and range of 50 meters. Based on these parameters, a semivariogram model is constructed.
[0091] Step S1433: Calculate the interpolation weight coefficient between the target grid point and the monitoring node according to the semivariogram model parameters and the spatial distance matrix.
[0092] For each target grid point in the 3D geographic information grid, calculate the spatial distance between it and all monitoring nodes. Then, calculate the interpolation weight coefficient based on the semivariogram model and these distances. Assume that the coordinates of target grid point G are (205, 305, 11), its distance from monitoring node A is 5 meters, its distance from monitoring node B is 12 meters, and its distance from monitoring node C is 22 meters.
[0093] Substitute these distances into the semivariogram model to obtain the corresponding semivariogram values. Then, based on the principles of kriging interpolation, solve a set of linear equations to calculate the interpolation weight coefficients between the target grid point G and the monitoring nodes A, B, and C. Assume that the weight coefficients obtained are 0.4, 0.3, and 0.3, respectively.
[0094] Step S1434: performing weighted summation on the pest and disease occurrence probabilities of the monitoring nodes according to the interpolation weight coefficients to generate an interpolation probability value of the target grid point.
[0095] In this example, assume that the probability of pests and diseases occurring at monitoring node A is 60%, the probability of pests and diseases occurring at monitoring node B is 55%, and the probability of pests and diseases occurring at monitoring node C is 65%. The interpolated probability value of the target grid point G is 0.4 × 60% + 0.3 × 55% + 0.3 × 65% = 0.4 × 0.6 + 0.3 × 0.55 + 0.3 × 0.65 = 0.24 + 0.165 + 0.195 = 0.6.
[0096] Step S1435: traverse all target grid points in the three-dimensional geographic information grid to generate a continuous probability distribution surface covering the target forest area.
[0097] In this embodiment, by traversing each target grid point, an interpolated probability value is obtained for each grid point. These interpolated probability values are visualized within the three-dimensional geographic information grid, generating a continuous probability distribution surface covering the target forest area. This surface intuitively displays the probability of pests and diseases occurring at different locations within the target forest area.
[0098] Step S1436: performing Gaussian filtering on the continuous probability distribution surface to eliminate local mutation noise of the interpolation probability value, and outputting a smoothed continuous probability distribution surface.
[0099] To eliminate possible localized noise mutations in a continuous probability distribution surface, perform Gaussian filtering. Gaussian filtering is a linear smoothing filter that smooths surfaces through convolution. Choose an appropriate Gaussian kernel size and standard deviation, for example, a 3×3×3 kernel size and a standard deviation of 1.
[0100] For each grid point in the continuous probability distribution surface, the interpolation probability values of the grid points in its neighborhood are weighted averaged according to the weight of the Gaussian kernel. For example, for the target grid point G, there are 26 grid points in its neighborhood (in three-dimensional space). The interpolation probability values of these 26 grid points are weighted averaged according to the weight of the Gaussian kernel to obtain the smoothed interpolation probability value. This process is repeated for all grid points, and the smoothed continuous probability distribution surface is output.
[0101] Step S144: performing regional segmentation on the continuous probability distribution surface according to the influence range level, dividing it into a plurality of partitions, and assigning a color gradient identifier to each partition.
[0102] In this embodiment, the continuous probability distribution surface is segmented into regions based on the influence level of each monitoring node. For regions with a "low" influence level, relatively small regions are marked on the continuous probability distribution surface; for regions with a "medium" influence level, medium-sized regions are marked; and for regions with a "high" influence level, larger regions are marked.
[0103] For example, for monitoring node A, whose influence range is "medium," a medium-sized region is divided on the continuous probability distribution surface with the grid point corresponding to the monitoring node as the center. This region division operation is performed on all monitoring nodes to obtain multiple partitions.
[0104] Then, assign a color gradient to each zone. Different colors can be used to represent different ranges of pest and disease probability. For example, areas with a pest and disease probability of 0-20% are represented by green, areas with a pest and disease probability of 20%-40% are represented by yellow, areas with a pest and disease probability of 40%-60% are represented by orange, areas with a pest and disease probability of 60%-80% are represented by red, and areas with a pest and disease probability of 80%-100% are represented by dark red. This color gradient allows you to intuitively identify the pest and disease probability of each zone.
[0105] Step S145: superimposing the color gradient mark and the three-dimensional geographic information grid to generate the pest distribution heat map, wherein the heat intensity is proportional to the product of the pest occurrence probability and the impact range level.
[0106] In this embodiment, the color gradient identifier of each partition is mapped to the corresponding area in the three-dimensional geographic information grid, so that the entire three-dimensional geographic information grid presents a distribution of different colors, forming a heat map of pest and disease distribution.
[0107] Thermal intensity is proportional to the product of the probability of pest and disease occurrence and the impact area level. For example, for a subarea with a "medium" impact area level and a 60% probability of pest and disease occurrence, its thermal intensity is 60% x medium (assuming "medium" corresponds to 0.5) = 0.3. Based on this thermal intensity, the color intensity of the subarea on the heat map is adjusted to make the color more vivid, highlighting the severity of the pest and disease in that area.
[0108] Step S146: Generate a pest control priority strategy, which specifically includes the following sub-steps.
[0109] Step S1461: Obtain a significant clustering region coordinate set and corresponding pest type identifiers in the pest distribution heat map.
[0110] In this embodiment, a significant clustering area refers to an area where the probability of occurrence of pests and diseases is high and concentrated. By setting a pest and disease occurrence probability threshold, for example, 70%, areas with a pest and disease occurrence probability greater than 70% are marked as significant clustering areas.
[0111] Obtain the coordinates of these significant clusters. For example, there are three significant clusters with coordinates [(200-220, 300-320, 0-10), (400-420, 500-520, 0-10), (600-620, 700-720, 0-10)]. At the same time, based on previous pest and disease prediction results, obtain the corresponding pest and disease type identifiers for these areas, such as "pine wilt disease" and "poplar anthracnose."
[0112] Step S1462: Match the control resource requirement list corresponding to each type of pest and disease according to the historical control record database. The control resource requirement list includes the type of pesticide, the number of equipment, and the manpower allocation.
[0113] Query the historical pest control record database, which records the pest control resource information used in the past. For each identified pest type, match it with the corresponding pest control resource requirement list.
[0114] For example, historical control records for pine wilt disease indicate the need for nematicides, 10 sprayers, 5 drills, and 20 specialized control personnel. For poplar anthracnose, carbendazim, 5 sprayers, and 10 control personnel are required. Thus, a corresponding list of control resource requirements was established for each pest and disease type in a significant cluster.
[0115] Step S1463: Calculating the urgency score of each region in the significant cluster region coordinate set, specifically including the following sub-steps.
[0116] Step S14631: Obtain the pest and disease occurrence probability P and impact range level R in the pest and disease prediction parameter set.
[0117] From the previously generated set of pest and disease prediction parameters, obtain the corresponding pest and disease occurrence probability P and impact range level R for each significant cluster area. For example, for the first significant cluster area, its pest and disease prediction parameter set is [80%, "high"], then the pest and disease occurrence probability P is obtained as 80%, and the impact range level R is obtained as "high" (assuming the value corresponding to "high" is 0.8).
[0118] Step S14632: query the vegetation economic value database of the target forest area and extract the economic value coefficient V of the corresponding area. The economic value coefficient is calculated comprehensively based on the vegetation type, tree age and market unit price.
[0119] The target forest area's vegetation economic value database is queried, which records the vegetation economic value information for different areas within the target forest area. Based on the coordinates of the significant clustered areas, the economic value coefficient V of the corresponding area is extracted.
[0120] The economic value coefficient (V) is calculated based on the vegetation type, tree age, and market unit price. For example, if a certain area is primarily populated by pine trees, each 20 years old, with a market unit price of 1,000 yuan per cubic meter, the economic value coefficient (V) for this area would be 0.6, calculated using a pre-defined method.
[0121] Step S14633: Calculate the initial urgency score using the formula S=P×R×V.
[0122] The obtained probability of pest occurrence P, impact level R, and economic value coefficient V are used to calculate the initial urgency score S. Using the data from the previous example, the probability of pest occurrence P is 80%, or 0.8, the corresponding value of impact level R is 0.8, and the economic value coefficient V is 0.6. Therefore, the initial urgency score S is equal to 0.8 multiplied by 0.8 and then multiplied by 0.6, resulting in a calculated S of 0.384.
[0123] Step S14634: If the current inventory of prevention and control resources is lower than the demand threshold, the initial urgency score is increased according to a preset ratio to obtain a revised urgency score.
[0124] Query the current inventory of pest control resources and compare it to the demand threshold. The demand threshold is a resource quantity standard set based on historical pest control experience and current pest and disease conditions. For example, for pine wilt disease control, the demand threshold for the nematicide agent is 1000 liters. The current inventory is 800 liters, which is below the demand threshold.
[0125] The preset increase percentage is determined based on the degree of resource shortage and the severity of the pest or disease. Assuming the preset increase percentage is 20%, the revised urgency score is equal to the initial urgency score of 0.384 multiplied by (1 + 20%), or 0.384 multiplied by 1.2, resulting in a revised urgency score of 0.4608.
[0126] Step S14635: normalize the corrected urgency score to obtain a final urgency score and write it into the prevention and control task priority queue.
[0127] Normalization is used to map the corrected urgency scores to a uniform range for easier comparison and ranking. Assuming the normalized range is [0, 1], we first find the maximum and minimum corrected urgency scores for all significant clusters.
[0128] Assume that the revised urgency scores for all regions are 0.4608, 0.35, 0.52, and so on, with a minimum of 0.3 and a maximum of 0.55. For the revised urgency score of 0.4608 calculated above, its normalized final urgency score is (0.4608 - 0.3) divided by (0.55 - 0.3), or 0.1608 divided by 0.25, resulting in a final urgency score of 0.6432. This final urgency score is written to the prevention and control task priority queue.
[0129] Step S1464: Sort the significant clustered area coordinate sets from high to low according to the urgency scores to generate a prevention and control task priority queue.
[0130] Sort the coordinates of the significant clusters based on their final urgency score. Place regions with higher final urgency scores first, and regions with lower scores last. For example, after sorting, the priority queue for prevention and control tasks might be [(600-620, 700-720, 0-10), (200-220, 300-320, 0-10), (400-420, 500-520, 0-10)], indicating that the first region has the highest priority, and so on.
[0131] Step S1465: Generate the pest control priority strategy including time nodes, resource allocation scheme and execution path according to the control task priority queue and control resource requirement list.
[0132] Based on the priority queue of prevention and control tasks and the corresponding prevention and control resource demand list for each area, specific time nodes, resource allocation plans and execution paths are formulated.
[0133] For the highest-priority area, such as the coordinates (600-620, 700-720, 0-10), assuming pine wilt disease is the problem, the control resource requirements list requires nematicides, 10 sprayers, 5 drills, and 20 specialized control personnel. Control work is scheduled to begin within the next two days, representing the timeline.
[0134] Regarding resource allocation, the appropriate amount of pesticides, equipment, and personnel will be deployed to the area from existing resource inventories. For example, if the current inventory of nematicide is 800 liters, 500 liters will be allocated to the area; 15 sprayers will be allocated, 10; and 8 drills will be allocated, 5. Regarding personnel, 20 existing pest control personnel will be deployed to the area.
[0135] The execution path is the route that pest control personnel and equipment take from the resource storage point to the target area. Based on the geographic information and traffic conditions of the target forest area, an optimal execution path is planned, for example, starting from the resource storage point, passing through specific roads and passages, and finally reaching the target area.
[0136] For other areas with lower priority, the same method is used to determine the time nodes, resource allocation plans and implementation paths in turn according to their priority order and the list of prevention and control resource requirements, and finally form a complete pest and disease control priority strategy.
[0137] Step S150: Fusing the pest distribution heat map with the pest control priority strategy to generate a forest area control optimization instruction set, and feeding the forest area control optimization instruction set back to the forestry monitoring platform to trigger a control resource scheduling operation.
[0138] This step requires integrating the previously generated pest and disease distribution heat map and the pest and disease control priority strategy to generate a forest area prevention and control optimization instruction set, and then feeding the instruction set back to the forestry monitoring platform.
[0139] First, a fusion process is performed to integrate the pest distribution information in the pest distribution heat map with the time nodes, resource allocation plan, and execution path of the pest control priority strategy. For example, in the pest distribution heat map, the corresponding control time node and the type and amount of resources required for each significant cluster area are marked.
[0140] For the area with coordinates (600-620, 700-720, 0-10), mark the corresponding location in the heat map as requiring 500 liters of nematicide, 10 sprayers, 5 drills, and 20 specialized pest control personnel for the next two days. Repeat this marking and integration process for all significant clusters.
[0141] After fusion processing, an optimized forest prevention instruction set is generated. This instruction set contains detailed prevention information, such as the prevention time, required resources, and execution path for each area.
[0142] The optimized forest control instruction set is then fed back to the forestry monitoring platform. Upon receiving the instruction set, the platform triggers resource scheduling based on the information contained therein. Based on the resource allocation plan, the platform allocates the appropriate pesticides, equipment, and manpower from its inventory and arranges for personnel to transport these resources to the target areas along the execution path. Simultaneously, the platform monitors the progress of control efforts in real time to ensure they proceed smoothly and on schedule.
[0143] For example, the platform will arrange vehicles to transport prepared pesticides and equipment to the target area and notify relevant pest control personnel to arrive at the designated location on time. During the control process, through monitoring equipment installed in the forest area and feedback from pest control personnel, the platform can obtain real-time information on the progress of the control work, such as whether the pesticides are sprayed properly and whether the equipment is functioning properly. If any problems arise, the platform can promptly adjust the control strategy to ensure effective pest control.
[0144] Based on the above steps, starting from obtaining the dynamic monitoring data set of the target forest area, through feature extraction, pest and disease prediction, generation of pest and disease distribution heat map and prevention and control priority strategy, this information is finally integrated to generate the forest area prevention and control optimization instruction set and fed back to the forestry monitoring platform, realizing the effective prediction of pests and diseases in the target forest area and the rational scheduling of prevention and control resources, thereby improving the efficiency and effectiveness of forestry pest and disease control.
[0145] Furthermore, for example, the method may also include a pre-training step of a pest and disease prediction model, as follows.
[0146] Step S210: Acquire a historical monitoring sample data set, which includes multiple healthy sample data subsets and pest and disease sample data subsets. Each sample data subset contains an environmental parameter sequence, a vegetation physiological parameter sequence, and a marked pest and disease occurrence status label.
[0147] In order to pre-train the pest and disease prediction model, it is first necessary to obtain a historical monitoring sample data set, which contains multiple healthy sample data subsets and pest and disease sample data subsets. Each sample data subset contains an environmental parameter sequence, a vegetation physiological parameter sequence, and an annotated pest and disease occurrence status label.
[0148] This embodiment can collect historical monitoring data from multiple different forest areas, covering different geographical environments, climate conditions, and vegetation types, to ensure data diversity and representativeness. The environmental parameter sequence includes data such as temperature, humidity, light intensity, and soil pH, which reflect the environmental conditions of the forest area. The vegetation physiological parameter sequence includes information such as chlorophyll content, trunk water saturation, and canopy density, which are closely related to the growth status of the vegetation.
[0149] Each sample data subset also needs to be labeled with the pest and disease occurrence status. This label can be determined through field surveys, expert assessments, or long-term monitoring records. For example, if a field survey reveals the presence of pests and diseases in a sample data subset, the data can be labeled as "mild pest and disease," "moderate pest and disease," or "severe pest and disease" based on the severity and scope of the pest and disease.
[0150] Step S220: performing time window interception and dimension unit normalization processing on the historical monitoring sample data set to generate a standardized sample data set.
[0151] After obtaining the historical monitoring sample data set, it is necessary to perform time window interception and dimension unit normalization on it to generate a standardized sample data set.
[0152] Because environmental and plant physiological parameters are time-varying sequence data, time windowing is necessary to better capture the data's characteristics. A set time window size is selected, such as a one-week window. For each sample data subset, divide it into multiple time windows in chronological order.
[0153] Assuming that a sample data subset contains one month of monitoring data, it can be divided into four time windows with a time window of one week. The data in each time window contains the environmental parameter sequence and vegetation physiological parameter sequence within that time period.
[0154] Different environmental parameters and vegetation physiological parameters have different dimensions and ranges. To eliminate the impact of dimensions, data normalization is required. Common normalization methods include min-max normalization and Z-score normalization.
[0155] Taking minimum-maximum normalization as an example, for each environmental parameter and vegetation physiological parameter, the minimum and maximum values in the entire historical monitoring sample data set are found. Then, for each data point, the minimum value is subtracted and the result is divided by the difference between the maximum and minimum values to obtain the normalized data.
[0156] For example, for the temperature parameter, in the historical monitoring sample data set, the minimum value is 10°C and the maximum value is 30°C. For a temperature data point of 20°C, the normalized value is (20-10) / (30-10)=0.5.
[0157] By extracting the time window and normalizing the dimensional units, a standardized sample data set is generated, making the data between different sample data subsets comparable, which facilitates subsequent model training.
[0158] Step S230: Input the standardized sample data set into the feature coding layer of the initial pest and disease prediction model for iterative training to generate a high-dimensional feature vector set, wherein the feature coding layer extracts spatiotemporal correlation features through the parallel structure of convolutional neural network and long short-term memory network.
[0159] The standardized sample data set is input into the feature encoding layer of the initial pest and disease prediction model for iterative training to generate a set of high-dimensional feature vectors. The feature encoding layer extracts spatiotemporal correlation features through a parallel structure of a convolutional neural network (CNN) and a long short-term memory network (LSTM).
[0160] CNN is mainly used to extract spatial features of data. The environmental parameter sequence and vegetation physiological parameter sequence in the standardized sample data set are regarded as two-dimensional matrix data, where one dimension represents time and the other dimension represents different parameters.
[0161] In a CNN, multiple convolution kernels are used to perform convolution operations on the data, extracting features at different scales. Each convolution kernel can be thought of as a small filter that slides over the data to compute the convolution result. By combining multiple convolutional and pooling layers, high-level features of the data can be gradually extracted.
[0162] For example, a 3×3 convolution kernel is used to perform a convolution operation on the data to obtain a new feature map. The feature map is then downsampled through a pooling layer (such as max pooling) to reduce the dimensionality of the data while retaining important feature information.
[0163] LSTM is mainly used to process sequence data and can capture the temporal dependencies of data. The environmental parameter sequences and vegetation physiological parameter sequences in the standardized sample data set are input into the LSTM in chronological order.
[0164] LSTM controls the flow of information through a gating mechanism, including input gate, forget gate, and output gate. The input gate determines whether new information should be added to the cell state, the forget gate determines which information should be forgotten, and the output gate determines which parts of the cell state should be output.
[0165] Through iterative training of LSTM, the time series characteristics of the data can be learned, such as the changing trends of vegetation physiological parameters over time.
[0166] The outputs of CNN and LSTM are concatenated to obtain a high-dimensional feature vector. This high-dimensional feature vector contains the spatial and temporal features of the data and can more comprehensively describe the sample data.
[0167] During the iterative training process, the parameters of the CNN and LSTM are continuously adjusted so that the generated high-dimensional feature vectors can better reflect the characteristics of the sample data. After multiple iterations, a set of high-dimensional feature vectors is finally generated.
[0168] Step S240: Filter the high-dimensional feature vectors corresponding to the healthy sample data subset from the high-dimensional feature vector set, calculate the mean vector and covariance matrix of each feature dimension, generate a standard healthy feature vector and store it in the healthy vegetation feature template library.
[0169] The high-dimensional feature vectors corresponding to the healthy sample data subset are screened out from the generated high-dimensional feature vector set. These healthy sample data subsets are the sample data marked as free of pests and diseases in step S210.
[0170] After filtering out the high-dimensional feature vectors corresponding to the healthy sample data subset, they are further processed to generate standard healthy feature vectors.
[0171] For the high-dimensional feature vectors corresponding to the screened healthy sample data subset, the mean vector and covariance matrix of each feature dimension are calculated.
[0172] The mean vector represents the average value of each feature dimension in the healthy sample data. For a high-dimensional feature vector, assuming there are n feature dimensions, calculate the average value of each feature dimension in all healthy sample data to obtain an n-dimensional mean vector.
[0173] The covariance matrix describes the correlation between different feature dimensions. Calculating the covariance between any two feature dimensions yields an n×n covariance matrix.
[0174] Based on the calculated mean vector and covariance matrix, a standard healthy feature vector is generated. The standard healthy feature vector can be regarded as the representative feature of the healthy sample data.
[0175] The generated standard healthy feature vector is stored in the healthy vegetation feature template library and used as a reference template for subsequent anomaly detection. In the actual pest and disease prediction process, the difference between the high-dimensional feature vector of the input data and the standard healthy feature vector is compared to determine whether there is an abnormal pest and disease situation.
[0176] Step S250: Based on the subset of pest and disease sample data in the high-dimensional feature vector set, calculate the multi-dimensional deviation coefficient of each high-dimensional feature vector and the standard healthy feature vector, and optimize the combined weight of the Euclidean distance and cosine similarity of the anomaly detection layer through the back propagation algorithm.
[0177] Based on the subset of pest and disease sample data in the high-dimensional feature vector set, the multidimensional deviation coefficient is calculated for each high-dimensional feature vector and the standard healthy feature vector. The multidimensional deviation coefficient is used to measure the degree of difference between the pest and disease sample data and the healthy sample data.
[0178] For the high-dimensional feature vector corresponding to each pest and disease sample data subset, the Euclidean distance and cosine similarity between it and the standard health feature vector are calculated.
[0179] Euclidean distance measures the distance between two vectors in space. A larger distance indicates a greater difference between the two vectors. Cosine similarity measures the directional similarity between two vectors. A similarity closer to 1 indicates a more similar direction.
[0180] The multi-dimensional deviation coefficient is obtained by weighting the Euclidean distance and cosine similarity. The weights can be optimized using the back-propagation algorithm.
[0181] The backpropagation algorithm is an optimization algorithm used to train neural networks. It continuously adjusts weights to minimize the error between the model's predictions and the true labels. During this process, an error function is calculated based on the multidimensional deviation coefficient and the annotated pest and disease status labels. Then, by backpropagating the error, the combined weights of the Euclidean distance and cosine similarity are adjusted, ensuring that the multidimensional deviation coefficient more accurately reflects the differences between pest and disease sample data and healthy sample data.
[0182] Step S260: Based on the marked pest and disease occurrence status labels and the preset true value of the impact range level, supervised training is performed on the impact range calculation module of the pest and disease prediction model to generate a chain propagation rate matching function of the spatial correlation parameters.
[0183] Based on the annotated pest and disease occurrence status labels and the preset true value of the impact range level, the impact range calculation module of the pest and disease prediction model is supervised and trained to generate a chain transmission rate matching function of the spatial correlation parameters.
[0184] Spatial correlation parameters are extracted from the high-dimensional feature vector set. These spatial correlation parameters may include the spatial coordinates of the monitoring node, the distance to adjacent monitoring nodes, etc. Spatial correlation parameters are used to describe the spread of pests and diseases in space.
[0185] The labeled pest and disease occurrence status and the preset true impact range level are input into the impact range calculation module as supervisory signals. The impact range calculation module learns from these supervisory signals and adjusts its own parameters to generate a chain propagation rate matching function for the spatial correlation parameters.
[0186] The chain transmission rate matching function is used to describe the transmission rate of pests and diseases between different spatial locations. Through training, the impact range calculation module can accurately predict the impact range level of pests and diseases based on spatial correlation parameters.
[0187] Step S270: adjusting the hyperparameters of the feature encoding layer, anomaly detection layer, and impact range calculation module through cross-validation until the model prediction accuracy reaches a convergence threshold, and outputting a trained pest and disease prediction model.
[0188] Through cross-validation, the hyperparameters of the feature encoding layer, anomaly detection layer, and impact range calculation module are adjusted until the model prediction accuracy reaches the convergence threshold, and the trained pest and disease prediction model is output.
[0189] For example, the standardized sample data set is divided into a training set and a validation set. For example, the data is divided into a training set and a validation set according to the ratio of 80% and 20%.
[0190] The model is trained using the training set and evaluated using the validation set. Through multiple cross-validations, the hyperparameters of the feature encoding layer, anomaly detection layer, and impact range calculation module, such as the learning rate, convolution kernel size, and number of hidden layer neurons, are continuously adjusted.
[0191] In each cross-validation step, the model's prediction accuracy is calculated. Prediction accuracy can be measured by comparing the model's predictions with the true labels, for example using metrics like precision, recall, or F1 score.
[0192] The hyperparameters are continuously adjusted until the model's prediction accuracy reaches the convergence threshold. The convergence threshold is a pre-set accuracy value. When the model's prediction accuracy reaches or exceeds this threshold, the model training is considered to have converged.
[0193] When the model training converges, the trained pest and disease prediction model is output. This model can be used for actual forest pest and disease prediction. By inputting new monitoring data, it outputs the probability of pest and disease occurrence and the level of impact, providing decision support for forest pest and disease prevention.
[0194] During the actual pre-training process, it's important to ensure consistent dimensionality and matching feature dimensions during feature comparison and parameter calculation. For example, when calculating Euclidean distance and cosine similarity, ensure that the dimensions of the high-dimensional feature vector and the standard healthy feature vector are consistent. Furthermore, when adjusting hyperparameters, appropriate choices should be made based on the model's performance and the characteristics of the data to improve the model's predictive accuracy and generalization capabilities.
[0195] Figure 2A schematic diagram illustrates exemplary hardware and software components of an artificial intelligence-based forestry pest prediction system 100 that can implement the concepts of the present application, as provided in some embodiments of the present application. For example, a processor 120 can be used in the artificial intelligence-based forestry pest prediction system 100 to perform the functions described in the present application.
[0196] The artificial intelligence-based forestry pest prediction system 100 can be a general-purpose server or a special-purpose server, both of which can be used to implement the artificial intelligence-based forestry pest prediction method of the present application. Although only one server is shown in this application, for convenience, the functions described in this application can be implemented in a distributed manner on multiple similar platforms to balance the processing load.
[0197] For example, the forestry pest and disease prediction system 100 based on artificial intelligence can include a network port 110 connected to a network, one or more processors 120 for executing program instructions, a communication bus 130, and storage media 140 in different forms, such as a disk, ROM, or RAM, or any combination thereof. Exemplarily, the forestry pest and disease prediction system 100 based on artificial intelligence can also include program instructions stored in ROM, RAM, or other types of non-transitory storage media, or any combination thereof. The method of the present application can be implemented according to these program instructions. The forestry pest and disease prediction system 100 based on artificial intelligence also includes an I / O interface 150 between the computer and other input and output devices.
[0198] For ease of explanation, only one processor is described in the artificial intelligence-based forestry pest and disease prediction system 100. However, it should be noted that the artificial intelligence-based forestry pest and disease prediction system 100 in the present application may also include multiple processors, so the steps performed by one processor described in the present application may also be performed jointly or individually by multiple processors. For example, if the processor of the artificial intelligence-based forestry pest and disease prediction system 100 executes step A and step B, it should be understood that step A and step B may also be performed jointly by two different processors or individually in one processor. For example, the first processor executes step A, the second processor executes step B, or the first processor and the second processor execute steps A and B together.
[0199] In addition, an embodiment of the present invention further provides a readable storage medium, in which computer-executable instructions are preset. When a processor executes the computer-executable instructions, the above-mentioned artificial intelligence-based forestry pest and disease prediction method is implemented.
[0200] It should be noted that in order to simplify the description of the present invention and thus help understand one or more embodiments of the invention, in the foregoing description of the embodiments of the present invention, multiple features are sometimes combined into one embodiment, figure or description thereof.
Claims
1. A forestry pest and disease prediction method based on artificial intelligence, characterized in that: The method comprises: Acquire a dynamic monitoring data set of a target forest area, wherein the dynamic monitoring data set includes an environmental parameter sequence and a vegetation physiological parameter sequence collected by multiple monitoring nodes, wherein each monitoring node corresponds to a spatial coordinate region; Extracting features from the dynamic monitoring data set to obtain a forestry environment feature set for each monitoring node, wherein the forestry environment feature set includes vegetation growth state features, environmental fluctuation correlation features, and potential identification features of pests and diseases; Based on the pre-trained pest and disease prediction model, the forestry environment feature set is subjected to abnormal parameter analysis processing to generate a pest and disease prediction parameter set for the monitoring node, wherein the pest and disease prediction parameter set is used to characterize the probability of occurrence of pests and diseases and the level of the impact range; Performing a dynamic prediction matching operation based on the pest and disease prediction parameter set to determine a pest and disease distribution heat map of the target forest area and generate a pest and disease control priority strategy; The pest distribution heat map is integrated with the pest control priority strategy to generate a forest area control optimization instruction set, and the forest area control optimization instruction set is fed back to the forestry monitoring platform to trigger a control resource scheduling operation; The feature extraction of the dynamic monitoring data set to obtain a forestry environment feature set of each monitoring node includes: Performing time dimension alignment on the environmental parameter sequence to generate a standardized environmental parameter sequence; Calling a preset vegetation physiological characteristic encoder to perform physiological state analysis on the vegetation physiological parameter sequence to generate a vegetation growth state feature vector, wherein the vegetation growth state feature vector includes a chlorophyll content change rate, a trunk water saturation, and a canopy density attenuation coefficient; Performing environmental correlation analysis on the standardized environmental parameter sequence to extract environmental fluctuation correlation features, wherein the environmental fluctuation correlation features include the temperature and humidity coordinated change gradient, soil pH offset, and light intensity cumulative deviation value; Performing anomaly identification fusion processing on the vegetation growth state feature vector and the environmental fluctuation correlation feature to generate potential identification features of pests and diseases, wherein the potential identification features of pests and diseases include an estimated value of insect egg distribution density, a fungal infection diffusion rate, and an intensity of disease symptom manifestation; Integrating the vegetation growth state characteristics, the environmental fluctuation correlation characteristics, and the potential identification characteristics of pests and diseases to form the forestry environment characteristic set; The abnormal identification fusion processing of the vegetation growth state feature vector and the environmental fluctuation correlation feature to generate potential identification features of pests and diseases includes: Discretizing and segmenting the vegetation growth state characteristic vector to generate a chlorophyll content change rate interval, a trunk water saturation interval, and a canopy density attenuation coefficient interval; Determining the suitability level for hatching eggs based on the temperature and humidity coordinated change gradient in the environmental fluctuation correlation characteristics, and calculating an estimated value of the egg distribution density based on the chlorophyll content change rate interval; Based on the soil pH offset and the cumulative light intensity deviation value, a fungal growth probability model is constructed to output a fungal infection and diffusion rate, wherein the fungal infection and diffusion rate is calculated by a linear combination coefficient of the pH offset and the light intensity deviation; Determining the intensity of disease symptom manifestation based on the cross-validation results of the trunk water saturation interval and the canopy density attenuation coefficient interval, wherein the intensity of disease symptom manifestation is inversely proportional to the product of water saturation and canopy density; Performing weighted concatenation of the estimated value of the insect egg distribution density, the fungal infection diffusion rate, and the intensity of the disease symptom manifestation to generate the potential identification features of the pest; The strategy for generating pest control priority includes: Obtaining a set of significant clustered area coordinates and corresponding pest type identifiers in the pest distribution heat map; Matching a control resource requirement list corresponding to each pest type based on a historical control record database, the control resource requirement list includes the type of pesticide, the amount of equipment, and the manpower allocation; Calculate the urgency score of each area in the significant clustering area coordinate set, where the urgency score is determined by multiplying the probability of occurrence of pests and diseases, the level of affected area, and the economic value of vegetation; Sort the coordinate sets of significant clustered areas according to their urgency scores from high to low to generate a priority queue for prevention and control tasks; Generate the pest control priority strategy including time nodes, resource allocation plan and execution path according to the control task priority queue and control resource demand list; The step of calculating the urgency score of each region in the significant cluster region coordinate set includes: Obtaining the pest and disease occurrence probability P and impact range level R from the pest and disease prediction parameter set; Query the vegetation economic value database of the target forest area and extract the economic value coefficient V of the corresponding area, which is calculated based on the vegetation type, tree age and market unit price; Calculate the initial urgency score using the formula S=P×R×V; If the current inventory of prevention and control resources is lower than the demand threshold, the initial urgency score is increased by a preset ratio to obtain a revised urgency score; The corrected urgency score is normalized to obtain a final urgency score and written into the prevention and control task priority queue.
2. The artificial intelligence-based forestry pest prediction method according to claim 1, characterized in that: The performing environmental correlation analysis on the standardized environmental parameter sequence to extract environmental fluctuation correlation features includes: Acquiring a temperature sequence, a humidity sequence, and a light intensity sequence from the standardized environmental parameter sequence; Calculating a dynamic covariance matrix between the temperature sequence and the humidity sequence to determine a temperature-humidity coordinated change gradient, wherein the temperature-humidity coordinated change gradient is quantified by a covariance change rate within a sliding time window; Performing trend fitting processing on the soil pH data in the standardized environmental parameter sequence to extract the soil pH offset, where the soil pH offset is determined by a multiple of the standard deviation between the current pH value and the historical benchmark value; Performing cumulative deviation calculation on the light intensity sequence to generate a light intensity cumulative deviation value, wherein the cumulative deviation value is the cumulative difference between the measured light intensity value and the theoretical light intensity value in a continuous time interval; The temperature and humidity coordinated change gradient, soil pH offset and light intensity cumulative deviation value are normalized to obtain the environmental fluctuation correlation characteristics.
3. The artificial intelligence-based forestry pest prediction method according to claim 1, characterized in that: The pre-trained pest and disease prediction model performs abnormal parameter analysis on the forestry environment feature set to generate the pest and disease prediction parameter set of the monitoring node, including: Inputting the forestry environment feature set into the feature coding layer of the pest and disease prediction model to generate a high-dimensional fusion feature vector; Calculating, through the anomaly detection layer of the pest and disease prediction model, a deviation coefficient between the high-dimensional fusion feature vector and a pre-stored healthy vegetation feature template, wherein the deviation coefficient is determined by a weighted sum of Euclidean distance and cosine similarity; According to the deviation coefficient, a preset pest and disease type mapping table is matched to determine the probability of pest and disease occurrence, wherein the probability of pest and disease occurrence is exponentially related to the deviation coefficient; Invoking the impact range calculation module of the pest prediction model to calculate the pest impact range level based on the spatial correlation parameters in the high-dimensional fusion feature vector, where the pest impact range level is determined by the feature similarity chain propagation rate of adjacent monitoring nodes; The probability of occurrence of the pests and diseases and the level of the impact range are combined and packaged to generate the pest and disease prediction parameter set.
4. The artificial intelligence-based forestry pest prediction method according to claim 3, characterized in that: The calculating of the deviation coefficient between the high-dimensional fusion feature vector and a pre-stored healthy vegetation feature template includes: Obtain the standard chlorophyll content, standard trunk water saturation and standard canopy density in the healthy vegetation feature template; Calculating a first absolute value of a difference between a chlorophyll content change rate in the high-dimensional fusion feature vector and the standard chlorophyll content, and a second absolute value of a difference between a trunk water saturation and the standard trunk water saturation; Normalizing the first difference absolute value and the second difference absolute value to generate a primary deviation; Extracting the canopy density attenuation coefficient from the high-dimensional fusion feature vector, calculating the inverse of its ratio to the standard canopy density, and generating a secondary deviation index; The primary deviation index and the secondary deviation index are linearly combined to obtain a comprehensive deviation coefficient, wherein the combination weight of the linear combination is dynamically adjusted according to the importance of features in historical pest and disease data.
5. The artificial intelligence-based forestry pest prediction method according to claim 1, characterized in that: The performing of a dynamic prediction matching operation based on the pest and disease prediction parameter set to determine a pest and disease distribution heat map of the target forest area includes: Extracting the probability of occurrence of pests and diseases and the level of impact range from the pest and disease prediction parameter set; Constructing a three-dimensional geographic information grid based on the spatial coordinate area of the monitoring node, and mapping the probability of pest and disease occurrence of each grid node to the corresponding spatial coordinate; Based on the Kriging interpolation algorithm, the discrete probability of occurrence of pests and diseases is spatially interpolated to generate a continuous probability distribution surface; Performing regional segmentation on the continuous probability distribution surface according to the impact range level, dividing it into multiple partitions, and assigning a color gradient label to each partition; The color gradient logo is superimposed on the three-dimensional geographic information grid to generate the pest and disease distribution thermal map, in which the thermal intensity is proportional to the product of the pest and disease occurrence probability and the impact range level.
6. The artificial intelligence-based forestry pest prediction method according to claim 5, characterized in that: The Kriging interpolation algorithm is used to perform spatial interpolation processing on the discrete probability of occurrence of pests and diseases to generate a continuous probability distribution surface, including: generating a spatial distance matrix between monitoring nodes based on the spatial coordinates of the monitoring nodes; Calculating semivariogram model parameters according to the spatial distance matrix and constructing a semivariogram model; Calculating interpolation weight coefficients between target grid points and monitoring nodes based on the semivariogram model parameters and the spatial distance matrix; Performing weighted summation on the pest and disease occurrence probabilities of the monitoring nodes according to the interpolation weight coefficients to generate an interpolation probability value of the target grid point; Traversing all target grid points in the three-dimensional geographic information grid to generate a continuous probability distribution surface covering the target forest area; Gaussian filtering is performed on the continuous probability distribution surface to eliminate local mutation noise of the interpolation probability value, and a smoothed continuous probability distribution surface is output.
7. An artificial intelligence-based forestry pest and disease prediction system, characterized in that: It includes a processor and a memory, the memory is connected to the processor, the memory is used to store programs, instructions or codes, and the processor is used to execute the programs, instructions or codes in the memory to implement the artificial intelligence-based forestry pest prediction method described in any one of claims 1 to 6.
Citation Information
Patent Citations
Rice disease and pest early warning method and system based on image recognition
CN119251568A
Deep learning-based farmland disease and pest monitoring and prevention method, system and device
CN119904328A