Crop disease and pest data prediction method and system
By obtaining environmental parameters of crop-growing areas, identifying key nodes of transmission paths, generating risk boundary distribution maps and evaluating synchronization indicators of environmental factors, the deficiencies in real-time and accuracy of traditional pest and disease prediction methods are addressed, and efficient pest and disease early warning and dynamic risk analysis are achieved.
Patent Information
- Application Number
- CN202510902022.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-09-05
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional crop disease and pest prediction methods rely on meteorological monitoring data and historical records, which makes it difficult to cope with sudden disease and pest emergencies and complex climate changes. The prediction accuracy is limited by data missing and processing methods, and real-time and accurate early warning cannot be achieved.
By obtaining the environmental parameters of the planting area, identifying the key nodes of the propagation path, generating the risk boundary distribution map, statistically analyzing the species diffusion frequency data, evaluating the synchronization index of environmental factors, screening the synchronization deviation nodes, generating the environmental factor synchronization deviation prediction data set, and using the cluster analysis method to predict dynamic changes.
It has achieved more detailed pest and disease warning and risk analysis, can efficiently identify the spread range and intensity of species, improve the real-time and accuracy of predictions, and provide more effective agricultural management decision support.
Smart Images

Figure CN120598129A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of agricultural data analysis, and in particular to a method and system for predicting crop disease and insect pest data. Background Art
[0002] The field of agricultural data analysis involves the collection, analysis, and processing of various types of data in agricultural production. Its purpose is to support improving agricultural production efficiency, optimizing agricultural management decisions, providing early warning of agricultural disasters, and promoting sustainable agricultural development. Core issues in this field include crop growth monitoring, soil quality analysis, meteorological data analysis, and pest and disease prediction. Widely used technologies include sensor data acquisition, remote sensing, data mining, and machine learning. Agricultural data analysis integrates and analyzes data through systematic technical means, helping farmers and agricultural managers make more accurate decisions, increase crop yields, and effectively control pests and diseases.
[0003] Traditional crop pest and disease data prediction methods involve collecting environmental data, meteorological data, and pest and disease records from crop-growing areas, and then using statistical or machine learning methods to analyze and predict the likelihood of crop pest and disease outbreaks. Traditional methods rely on meteorological monitoring data and historical pest and disease records to predict the risk of pest and disease outbreaks through regression or statistical models. However, these methods are limited in their effectiveness in responding to sudden pest and disease outbreaks and complex climate changes, and their prediction accuracy is significantly affected by missing data and limitations in processing methods. Traditional methods primarily use empirical formulas and historical data to make rough predictions of pest and disease outbreaks, making it difficult to provide real-time, accurate early warnings.
[0004] Existing technologies for pest and disease forecasting rely too heavily on meteorological monitoring data and historical records, using regression or statistical models for extrapolation. This approach is limited in its effectiveness when responding to sudden pest and disease outbreaks or complex climate change. Furthermore, due to data gaps and limitations in processing methods, it is difficult to provide accurate and real-time early warnings. For example, traditional methods struggle to accurately predict the outbreak of emerging pests and diseases, and are unable to promptly respond to potential risks posed by climate change. This leads to delayed early warnings, impacting farmers' prevention and control measures and yields. Traditional methods rely too heavily on historical data and ignore the complex interactions between environmental factors, resulting in insufficient flexibility in forecasts and an inability to adapt to changing climatic conditions and geographic environments. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings of the prior art and to propose a crop disease and insect pest data prediction method and system.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a method for predicting crop disease and insect pest data, comprising the following steps:
[0007] S1: Obtain temperature, humidity, and soil moisture distribution information from the environmental parameters of the planting area, perform time series interval segmentation, locate the mutation points of environmental parameters and determine the locations of key nodes in the propagation path, determine the risk range of species diffusion, and generate a risk boundary distribution map;
[0008] S2: Based on the risk boundary distribution map, read the key node sequence of the propagation path, identify the temperature and humidity change direction vectors, perform propagation intensity counting on the vectors and aggregate them by node, count the temperature and humidity change frequencies, and obtain species diffusion frequency data;
[0009] S3: extracting the propagation path distribution information based on the species diffusion frequency data, combining it with the node response time distribution, identifying the environmental factor correlation strength index, and generating the propagation path environmental factor synchronization index;
[0010] S4: Based on the synchronization index of the environmental factors of the propagation path, a set of nodes whose synchronization deviates from the median value is screened, the corresponding differences between the node intervals and the response time intervals are analyzed, and an environmental factor synchronization deviation prediction data set is generated.
[0011] As a further solution of the present invention, the risk boundary distribution map includes the location distribution of mutation points, the distribution of key nodes in the propagation path, and the risk range of species diffusion. The species diffusion frequency data includes the frequency of propagation node changes, the number of vector changes, and the characteristics of the propagation direction. The synchronization index of the propagation path environmental factors includes propagation distribution information, response time distribution, and environmental factor correlation strength index. The environmental factor synchronization deviation prediction data set includes node time distribution, node interval mean, and corresponding difference distribution of the propagation path.
[0012] As a further solution of the present invention, the steps of obtaining the risk boundary distribution map are specifically as follows:
[0013] S111: Obtaining temperature, humidity, and soil moisture distribution information from the environmental parameters of the planting area, performing time series interval segmentation on the continuous data, identifying mutation points within the interval, extracting points where the two points before and after the current data point show opposite changes as mutation points, and obtaining a mutation point distribution position sequence;
[0014] S112: Based on the mutation point distribution position sequence, for each mutation point and adjacent data sequence, calculate the continuous change and determine whether it is less than zero, mark the key nodes of the propagation path, combine the mutation points to form a pairing structure, and obtain the species diffusion risk range sequence;
[0015] S113: Based on the species diffusion risk range sequence, group the transmission paths by starting nodes, calculate the start and end changes, transmission intervals and average transmission rates of each group of transmission paths, integrate the boundary features of the transmission paths, and generate a risk boundary distribution map.
[0016] As a further solution of the present invention, the steps for obtaining the species diffusion frequency data are specifically as follows:
[0017] S211: Based on the risk boundary distribution map, read the propagation path subsequence within the corresponding time range, identify the index sequence of each propagation path in ascending time order, perform reverse sorting on the temperature and humidity data sequence, and obtain a temperature and humidity difference vector sequence;
[0018] S212: Extract the change signs of adjacent elements in the temperature and humidity difference vector sequence, if the difference is less than zero, it is considered a propagation event, and count the number of propagations of each segment according to the index segment to obtain a propagation count sequence;
[0019] S213: Based on the propagation count sequence and in combination with the time span information of each propagation path, the propagation frequency per unit time in the propagation path is calculated, the corresponding frequency of the propagation path is counted, and species diffusion frequency data is established.
[0020] As a further solution of the present invention, the step of obtaining the synchronization index of the propagation path environmental factor is specifically as follows:
[0021] S311: Based on the species diffusion frequency data, mapping and aggregating by propagation path number, extracting the associated propagation path number and propagation frequency, grouping and calculating the distribution of each propagation path during the propagation process, and generating a propagation path distribution sequence;
[0022] S312: Based on the propagation path distribution sequence, extract the propagation path response time series, calculate the fluctuation of the distribution information corresponding to the time point, evaluate the time distribution fluctuation degree of each propagation path response sequence, and generate the propagation path response time distribution sequence;
[0023] S313: Based on the response time distribution sequence of the propagation path and the distribution of short-time environmental factors of the propagation path, statistical energy concentration and standard deviation are performed to evaluate the correlation structure of environmental factors and obtain the synchronization index of the propagation path environmental factors.
[0024] As a further solution of the present invention, the steps for obtaining the environmental factor synchronization deviation prediction data set are specifically as follows:
[0025] S411: Based on the synchronization index of the propagation path environmental factor, identify the synchronization value set according to the propagation path number, extract the quartiles and set the deviation threshold, filter the propagation path index values, and output the abnormal propagation path number set;
[0026] S412: Extracting and normalizing the difference between consecutive node intervals of the propagation path based on the set of abnormal propagation path numbers, calculating the ratio of the corresponding time of the propagation path to the node misalignment, aggregating and averaging the misalignment ratio sequence of each propagation path, and generating a propagation path alignment deviation mean sequence;
[0027] S413: Based on the propagation path alignment deviation mean sequence, sort the average alignment deviation by propagation path number, identify the corresponding deviation between the propagation path sampling time and interval, use the deviation as the discrete trajectory point of the propagation path, and generate an environmental factor synchronization deviation prediction data set.
[0028] As a further embodiment of the present invention, the method further comprises step S5:
[0029] S5: performing cluster analysis on the reorganized node density distribution based on the environmental factor synchronization deviation prediction dataset, calculating the node density gradient in the cluster area, performing a difference operation with the baseline density and performing normalization processing to generate a planting area dynamic change prediction dataset;
[0030] The planting area dynamic change prediction data set includes time window clustering areas, node density gradient change rate, and normalized difference area distribution records.
[0031] As a further solution of the present invention, the steps for acquiring the planting area dynamic change prediction dataset are specifically as follows:
[0032] S511: reorganize and sort all nodes on the time axis according to the environmental factor synchronization deviation prediction data set, analyze the change trend of node density, calculate the node change rate per unit time, and obtain the sliding window node density gradient;
[0033] S512: Based on the sliding window node density gradient, extract the node density within the time window and perform a difference operation with the reference density, perform normalization processing on the difference data, uniformly scale it to a standard range, and obtain a cluster density normalized difference;
[0034] S513: Arrange the density offset values in the order of the time window center points according to the cluster density normalization difference, aggregate the time windows in each cluster area and superimpose the normalization results, reconstruct a complete and sequentially consistent node sequence structure, and establish a planting area dynamic change prediction dataset.
[0035] The crop disease and insect pest data prediction system is used to implement the above-mentioned crop disease and insect pest data prediction method, and the system includes:
[0036] The risk extraction module obtains the temperature, humidity, and soil moisture distribution information of the planting area environmental parameters, detects the mutation points of the time series interval as the key nodes of the propagation path, determines the continuous decreasing trend of the propagation path, and generates a risk boundary distribution map;
[0037] The propagation calculation module calculates the number of temperature and humidity change signs based on the inverted data sequence in the risk boundary distribution map, aggregates the propagation intensity value, and generates species diffusion frequency data;
[0038] The synchronization evaluation module aggregates distribution information by propagation path based on the species diffusion frequency data, identifies the environmental factor correlation strength index in combination with the propagation path response time distribution, and generates a propagation path environmental factor synchronization index;
[0039] The node analysis module screens abnormal propagation paths based on the synchronization index of the propagation path environmental factors, extracts time series and node interval series, calculates the corresponding difference means, and generates an environmental factor synchronization deviation prediction data set;
[0040] The density normalization module is based on the environmental factor synchronization deviation prediction data set, counts the changes in the number of nodes in the time window, calculates the density gradient difference and performs normalization processing to generate a planting area dynamic change prediction data set.
[0041] Compared with the prior art, the advantages and positive effects of the present invention are:
[0042] The present invention achieves more detailed pest and disease early warning and risk analysis by accurately capturing changes in environmental parameters and monitoring key nodes of the propagation path in real time. This allows for more efficient identification of the range and intensity of species spread, and thus accurate prediction of pest and disease occurrence. By analyzing environmental factors and inter-node response times, it is possible to effectively capture and predict pest and disease occurrence trends in different regions, avoiding the limitations of traditional methods that rely solely on meteorological data and historical records. This allows early warnings to be more than just retrospective predictions, but rather possesses a higher level of dynamic response capability. Cluster analysis is used to accurately assess and normalize changes in node density, improving the accuracy of predictions of dynamic changes in planting areas, making risk assessments more real-time and accurate, and providing more effective decision support for agricultural managers. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic diagram of the workflow of the present invention;
[0044] Figure 2 This is a flow chart for obtaining the risk boundary distribution map in the present invention;
[0045] Figure 3 This is a flow chart for obtaining species diffusion frequency data in the present invention;
[0046] Figure 4 This is a flow chart for obtaining the synchronization index of the propagation path environmental factor in the present invention;
[0047] Figure 5 This is a flow chart for obtaining a synchronous deviation prediction data set of environmental factors in the present invention;
[0048] Figure 6 This is a flow chart for obtaining a dataset for predicting dynamic changes in planting areas in the present invention. DETAILED DESCRIPTION
[0049] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0050] In the description of the present invention, it should be understood that the terms "length," "width," "up," "down," "front," "back," "left," "right," "vertical," "horizontal," "top," "bottom," "inside," "outside," and the like, indicating positions or relationships, are based on the positions or relationships shown in the accompanying drawings and are intended only to facilitate the description of the present invention and simplify the description. They do not indicate or imply that the devices or elements referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limiting the present invention. Furthermore, in the description of the present invention, "plurality" means two or more, unless otherwise expressly and specifically defined.
[0051] Example 1
[0052] See also Figure 1 The present invention provides a technical solution: a method for predicting crop disease and insect pest data, comprising the following steps:
[0053] S1: Obtain temperature, humidity, and soil moisture distribution information from the environmental parameters of the planting area, perform time series interval segmentation, locate the mutation points of environmental parameters and determine the locations of key nodes in the propagation path, determine the risk range of species diffusion, and generate a risk boundary distribution map;
[0054] S2: Based on the risk boundary distribution map, read the key node sequence of the propagation path, identify the temperature and humidity change direction vectors, count the propagation intensity of the vectors and aggregate them by node, count the temperature and humidity change frequencies, and obtain species diffusion frequency data;
[0055] S3: Extract the propagation path distribution information based on the species diffusion frequency data, combine it with the node response time distribution, identify the environmental factor correlation strength index, and generate the propagation path environmental factor synchronization index;
[0056] S4: Based on the synchronization index of the environmental factors of the propagation path, the node set with synchronization deviation from the median value is screened, the corresponding difference between the node interval and the response time interval is analyzed, and the environmental factor synchronization deviation prediction dataset is generated;
[0057] S5: Based on the environmental factor synchronization deviation prediction dataset, cluster analysis is performed on the reorganized node density distribution, the node density gradient in the cluster area is statistically calculated, and the difference operation is performed with the baseline density and normalized to generate a planting area dynamic change prediction dataset.
[0058] The risk boundary distribution map includes the location distribution of mutation points, the distribution of key nodes in the propagation path, and the risk range of species diffusion. The species diffusion frequency data includes the frequency of propagation node changes, the number of vector changes, and the characteristics of the propagation direction. The synchronization indicators of environmental factors in the propagation path include propagation distribution information, response time distribution, and environmental factor correlation strength indicators. The environmental factor synchronization deviation prediction data set includes node time distribution, node interval mean, and corresponding difference distribution of the propagation path. The planting area dynamic change prediction data set includes time window clustering areas, node density gradient change rate, and normalized difference area distribution records.
[0059] See also Figure 2 , the steps to obtain the risk boundary distribution map are as follows:
[0060] S111: Obtaining temperature, humidity, and soil moisture distribution information from the environmental parameters of the planting area, performing time series interval segmentation on the continuous data, identifying mutation points within the interval, extracting points where the two points before and after the current data point show opposite changes as mutation points, and obtaining a mutation point distribution position sequence;
[0061] When obtaining the temperature, humidity and soil moisture distribution information of the environmental parameters of the planting area, data is automatically collected every hour from 100 wireless sensor nodes evenly distributed in a smart greenhouse. For example, from 0:00 on June 1 to 0:00 on June 2, sensor node A records a temperature of 25.1°C, a humidity of 70%, and a soil moisture content of 35% at 0:00, and a temperature of 25.3°C, a humidity of 70.5%, and a soil moisture content of 35.2% at 1:00, and so on. When performing time series interval segmentation on continuous data, 24 hours is a time series interval. For example, the data from 0:00 to 23:59 on June 1 is taken as an interval. When identifying the mutation points in the interval, the change trends of temperature, humidity and soil moisture content of adjacent data points are compared. For example, the temperature of sensor node B is 25.1°C, 70% and 35% at 0:00 in a certain time interval. After the temperature rises from 28℃ to 29℃, it drops to 28.5℃ in the next hour. After the humidity drops from 60% to 58%, it rises to 59% in the next hour. After the soil moisture content rises from 30% to 31%, it drops to 30.5% in the next hour. At this time, the point where the temperature, humidity and soil moisture content all change in the opposite direction is marked as a mutation point. The points before and after the current data point where the two points change in the opposite direction are extracted as mutation points. For example, if the temperature of the current data point is 28.5℃, the previous point is 29℃, and the next point is 28.8℃, and the humidity is 59%, the previous point is 58%, and the next point is 59.5%, and the soil moisture content is 30.5%, the previous point is 31%, and the next point is 30.8%, then the current data point will be identified as a mutation point, and the mutation point distribution position sequence will be obtained.
[0062] S112: Based on the mutation point distribution position sequence, for each mutation point and adjacent data sequence, calculate the continuous change and determine whether it is less than zero, mark the key nodes of the propagation path, combine the mutation points to form a pairing structure, and obtain the species diffusion risk range sequence;
[0063] According to the mutation point distribution position sequence, for example, the mutation point sequence is [P1, P2, P3, P4], where P1 represents the position of the first mutation point (for example, the data collected by a sensor at a specific time point), and P2 represents the position of the second mutation point. For each mutation point and the adjacent data sequence, for example, for P1 and the subsequent continuous data sequence (for example, 10 consecutive data points after P1 occurs), calculate the continuous change and determine whether it is less than zero. For example, if the temperature drops by 5 units continuously, the humidity drops by 3 units continuously, and the soil moisture content drops by 2 units continuously, and all change values are less than zero, the continuous change sequence will be marked as the key node of the propagation path. For example, the node where the temperature, humidity, and soil moisture content drop at the same time is marked as the key node. Combined with the mutation point, the continuous change sequence is the key node of the propagation path. When change points form a pairing structure, each mutation point is paired with one or more key node sequences. For example, P1 is paired with the continuously decreasing key node sequence C1, forming a correspondence between the start and propagation path of a diffusion event. If, after the mutation point, the temperature of sensor node C continuously decreases from 27°C to 25°C (a change of -2°C), the humidity continuously decreases from 65% to 60% (a change of -5%), and the soil moisture content continuously decreases from 32% to 30% (a change of -2%), and the change values are all less than zero, then the nodes in the decreasing process are marked as key nodes, and finally a species diffusion risk range sequence is obtained. For example, the risk range sequence contains {[P1, C1], [P2, C2], ...}, where each element represents a diffusion event and its propagation path.
[0064] S113: Based on the species diffusion risk range sequence, group the transmission paths by starting nodes, calculate the start and end changes, transmission intervals, and average transmission rates of each transmission path, integrate the boundary characteristics of the transmission paths, and generate a risk boundary distribution map;
[0065] Based on the species diffusion risk range sequence, for example, the risk range sequence includes {[P1, C1], [P2, C2], [P3, C3]}. When grouping by the starting node of the propagation path, the paths are grouped according to the starting mutation point of each propagation path. For example, all propagation paths starting from P1 are grouped into one group, and those starting from P2 are grouped into another group. If P1 and P2 are mutations at different sensor points at the same time, the paths are grouped into the same group. If P1, P2, and P3 all occur at different starting time points and at different sensor points, different groups will be formed. For example, all propagation paths with the mutation point detected by sensor A at 0:00 on June 1 as the starting node are grouped into the first group, and all propagation paths with the mutation point detected by sensor B at 1:00 on June 1 as the starting node are grouped into the second group. When calculating the start and end changes, propagation interval, and average propagation rate of each group of propagation paths, for each propagation path in each group, the temperature from the starting mutation point to the propagation end point (the last point of the key node sequence) is calculated. The total change in temperature, humidity, and soil moisture content. The propagation interval refers to the length of time from the starting point to the end point. For example, if the propagation path starts at 0:00 on June 1 and ends at 12:00 on June 1, the propagation interval is 12 hours. The average propagation rate is the total change divided by the propagation interval. For example, if the total temperature change is -5°C and the propagation interval is 10 hours, the average propagation rate is -0.5°C / hour. When integrating the boundary characteristics of the propagation path, the calculated start and end changes, propagation interval, and average propagation rate are used as the boundary characteristics of the propagation path. For example, for the first group of propagation paths, the temperature change range is [-5°C, -2°C], the humidity change range is [-10%, -5%], the soil moisture content change range is [-3%, -1%], and the average propagation rate range is [-0.5°C / hour, -0.2°C / hour]. Finally, a risk boundary distribution map is generated, which visualizes the starting point, propagation path length, propagation direction, and the range of environmental parameters for different species dispersal events.
[0066] See also Figure 3 , the specific steps for obtaining species diffusion frequency data are:
[0067] S211: Based on the risk boundary distribution map, read the propagation path subsequence within the corresponding time range, identify the index sequence of each propagation path in ascending time order, perform reverse sorting on the temperature and humidity data sequences, and obtain the temperature and humidity difference vector sequence;
[0068] Based on the risk boundary distribution map, for example, the risk boundary distribution map shows the risk areas and propagation paths for the spread of multiple species, and the propagation path subsequence within the corresponding time range is read. For example, all propagation path data from June 1 to June 7 are selected. The path subsequence contains the detailed environmental parameter changes of each propagation path. When identifying the index sequence of each propagation path in ascending time order, the selected propagation path subsequences are sorted in ascending order according to their occurrence time, and a unique index number is assigned to each path. For example, path 1 occurs at 0:00 on June 1, path 2 occurs at 1:00 on June 1, and so on. When performing reverse sorting of the temperature and humidity data sequences, the temperature and humidity data sequences corresponding to each propagation path are obtained. For example, for path 1, its temperature data sequence is [28℃, 27.5℃, 27℃ , 26.5℃], the humidity data sequence is [60%, 59.5%, 59%, 58.5%], and then the sequence is sorted in reverse order. For example, the temperature reverse sequence is [26.5℃, 27℃, 27.5℃, 28℃], and the humidity reverse sequence is [58.5%, 59%, 59.5%, 60%]. When obtaining the temperature and humidity difference vector sequences, the difference between adjacent data points after reverse sorting is calculated. For example, for the temperature reverse sequence, the difference vector is [27℃-26.5℃, 27.5℃-27℃, 28℃-27.5℃]=[0.5℃, 0.5℃, 0.5℃], and for the humidity reverse sequence, the difference vector is [59%-58.5%, 59.5%-59%, 60%-59.5%]=[0.5%, 0.5%, 0.5%].
[0069] S212: Extract the change signs of adjacent elements in the temperature and humidity difference vector sequence. If the change signs are less than zero, it is considered a propagation event. Count the number of propagations in each segment according to the index segment to obtain a propagation count sequence.
[0070] According to the temperature and humidity difference vector sequence, for example, the temperature difference vector sequence is [0.5℃, 0.5℃, 0.5℃], and the humidity difference vector sequence is [0.5%, 0.5%, 0.5%]. When extracting the change signs of adjacent elements in the difference, for each element of the temperature difference vector, if the difference is less than zero, it is recorded as a negative sign, otherwise it is a positive sign. For example, if an element of the difference vector is -0.2℃, its change sign is negative. If it is less than zero, it is regarded as a propagation event. When judging the extracted change sign, when there is an element less than zero in the temperature difference vector or the humidity difference vector, it means that the environmental conditions are moving towards a favorable direction. The direction of species diffusion changes. For example, when the temperature changes from high to low, the difference is negative, which is regarded as a propagation event. When counting the number of propagations in each segment by index segment, the number of propagation events occurring in each index segment (that is, in each propagation path) is counted according to the previously identified index sequence. For example, for index segment 1 (path 1), if its temperature difference vector has 2 elements less than zero and its humidity difference vector has 1 element less than zero, then the number of propagations in this segment is 3, and finally the propagation count sequence is obtained. For example, the propagation count sequence is [3, 2, 4], which means that there are 3 propagation events in path 1, 2 in path 2, and 4 in path 3.
[0071] S213: Based on the propagation count sequence and the time span information of each propagation path, the formula is used:
[0072] ;
[0073] Calculate the propagation frequency per unit time within the propagation path, count the corresponding frequencies of the propagation path, and establish species diffusion frequency data;
[0074] in, Represents the propagation frequency per unit time within the propagation path, represents the propagation time of the i-th propagation path, represents the time span of the i-th propagation path, represents the total time span of the propagation path, Represents the total number of propagation paths;
[0075] According to the propagation count sequence, for example, the propagation count sequence is [3, 2, 4], combined with the time span information of each propagation path, for example, the time span of path 1 The time span of Path 2 is 10 hours The time span of path 3 is 8 hours The total time span of the propagation path is 12 hours. To determine the time span of all propagation paths, sum the time spans of all propagation paths. For example, if there are three propagation paths with time spans of 10 hours, 8 hours, and 12 hours respectively, the total time span is hours, of which Represents the frequency of spread per unit time within the propagation path, that is, the frequency of species dispersal events within a specific time period;
[0076] Representative The propagation time of a propagation path segment refers to the sum of the time when all propagation events occur in the propagation path. For example, if there are three propagation events in the first propagation path segment, which occur at the second hour, the fifth hour, and the eighth hour respectively, then hours. If the duration of a propagation event is 0.5 hours, then the propagation time is 0.5×3=1.5 hours. Here, it refers to the time point when each propagation event occurs. Here, the number of propagation events is used to represent the propagation time. Therefore, The value of should be the number of propagation events. For example, if there are 3 propagation events in the first propagation path, then , No. The propagation time of a segment propagation path refers to the number of actual propagation events that occur in the segment path. For example, if Second-rate, Second-rate, Second-rate;
[0077] Representative The time span of a segment propagation path refers to the total time length from the beginning to the end of the path, for example, Hour, Hour, Hour;
[0078] represents the total time span of the propagation path, that is, the sum of the time spans of all propagation paths, for example, Hour;
[0079] represents the total number of propagation paths, e.g. ;
[0080] Using the formula Calculate the propagation frequency per unit time within the propagation path. This formula is designed to quantify the frequency of species diffusion. Its core idea is to calculate the difference between each propagation event and the corresponding time span, and divide the sum of the differences by the total time span to obtain the average propagation frequency. Specifically, the numerator Indicates the The absolute difference between the number of propagation events along a propagation path and the time span of the path, Indicates that all The absolute difference of each propagation path is summed up. This sum reflects the total deviation between the propagation event frequency and time span of all propagation paths. Represents the total time span of the propagation path, divide the numerator by The total deviation can be normalized to obtain the average propagation frequency per unit time;
[0081] For example, for path 1, ;
[0082] For path 2, ;
[0083] For path 3, ;
[0084] but ;
[0085] When counting the frequencies corresponding to the propagation paths, the calculated frequency values are associated with the corresponding propagation paths, and finally the species diffusion frequency data are established, which contains the unique identifier of each propagation path and its corresponding propagation frequency per unit time.
[0086] See also Figure 4 ,The specific steps for obtaining the synchronization index of the propagation path environmental factors are as follows:
[0087] S311: Based on the species diffusion frequency data, the propagation path number is mapped and aggregated, the associated propagation path number and propagation frequency are extracted, the distribution of each propagation path in the propagation process is calculated in groups, and a propagation path distribution sequence is generated;
[0088] According to the species diffusion frequency data, for example, the species diffusion frequency data includes a path number and a corresponding diffusion frequency value, as shown in Table 1. When mapping and aggregating by the diffusion path number, the records with the same diffusion path number are aggregated. For example, if the path number "Path001" has multiple frequency records, the records are grouped into the "Path001" group. When extracting the associated diffusion path number and diffusion frequency, each unique diffusion path number and the diffusion frequency value associated with it are extracted from the aggregated data. For example, "Path001" and its frequency 0.7, "Path002" and its frequency 0.6 are extracted from Table 1 and grouped. When calculating the distribution of each propagation path during the propagation process, the propagation frequency value of each propagation path (i.e., each group) is statistically analyzed to calculate its distribution characteristics. For example, the frequency distribution of "Path001" at different time points or under different environmental conditions is calculated. If the frequency of Path001 in the morning is 0.7, in the afternoon is 0.6, and in the evening is 0.8, then its frequency distribution is calculated. Statistics such as standard deviation and mean can be used to describe the distribution. For example, the frequency mean and standard deviation of Path001 are calculated to finally generate a propagation path distribution sequence, which contains the number of each propagation path and the distribution statistics of its propagation frequency.
[0089] Table 1: Species dispersal frequency data
[0090]
[0091] As shown in Table 1 , the diffusion frequency data of some species are listed, and the data are used for subsequent aggregation and analysis.
[0092] S312: Based on the propagation path distribution sequence, extract the propagation path response time series, calculate the fluctuation of the distribution information corresponding to the time point, evaluate the time distribution fluctuation degree of each propagation path response sequence, and generate the propagation path response time distribution series;
[0093] Based on the propagation path distribution sequence, for example, the propagation path distribution sequence includes the frequency distribution statistics of each propagation path. When extracting the propagation path response time series, the species response time series is extracted from the detailed data of each propagation path. For example, if the propagation frequency of Path001 changes at different time points, the time points and their corresponding frequency values are extracted to form a response time series. For example, the response time series of Path001 is [(T1, F1), (T2, F2), (T3, F3)], where T represents the time point and F represents the frequency at the time point. When calculating the fluctuation of the distribution information corresponding to the time point, for each propagation path response time series, the degree of fluctuation of its frequency value at different time points is calculated. For example, the variance or standard deviation of the frequency F of Path001 in the time period from T1 to T3 is calculated. When evaluating the degree of fluctuation of the time distribution of the response sequence of each propagation path, a statistical method is used to evaluate this fluctuation. For example, if the frequency value of Path001 fluctuates widely, it is considered that its time distribution fluctuation degree is high. If the frequency value is relatively stable, the fluctuation degree is low. Finally, a propagation path response time distribution sequence is generated. This sequence contains the response time series of each propagation path and the evaluation results of its fluctuation degree. For example, the response time distribution sequence of Path001 is (Path001_Response_Series, Volatility_Measure_Path001).
[0094] S313: Based on the response time distribution sequence of the propagation path, combined with the distribution of short-term environmental factors along the propagation path, the energy concentration and standard deviation are statistically analyzed to evaluate the correlation structure of environmental factors using the formula:
[0095] ;
[0096] Obtain the synchronization index of the environmental factors of the propagation path;
[0097] in, Represents the synchronization index of environmental factors in the propagation path, represents the response time of the kth propagation path, represents the average value of the propagation path response time distribution, represents the short-term environmental factor of the k-th propagation path, Represents the average value of the short-term environmental factor distribution along the propagation path, represents the total number of propagation paths, represents the weighted index associated with the propagation path response time, Represents the weighted index associated with the short-term environmental factors of the propagation path;
[0098] According to the propagation path response time distribution sequence, for example, the propagation path response time distribution sequence includes the response time series of each propagation path and its fluctuation degree evaluation result. For example, the response time series of Path001 is , combined with the distribution of short-term environmental factors of the propagation path, for example, the distribution of short-term environmental factors (such as temperature and humidity) corresponding to Path001 at the response time point is: , Among them, when calculating the energy concentration and standard deviation, the energy concentration refers to the concentration degree of the environmental factor data within a certain range. For example, for temperature data, if most of the temperature values are concentrated around 25℃, the energy concentration is high. The standard deviation reflects the degree of dispersion of the environmental factor data. For example, if the temperature standard deviation is 1℃, it means that the temperature fluctuation is small. When evaluating the correlation structure of environmental factors, the correlation between the environmental factors and the propagation path response time is judged by observing the energy concentration and standard deviation. If the environmental factor fluctuation is small and the energy concentration is high, while the propagation path response fluctuates greatly, it indicates that the correlation between the environmental factor and the response is weak, otherwise the correlation is strong. Among them, It represents the synchronization index of environmental factors along the transmission path, reflecting the degree of consistency between species diffusion response and changes in environmental factors;
[0099] Representative The response time of the propagation path, where Refers to a specific point in time For example, if the response frequency of a propagation path in the first hour is 0.7 and the response frequency in the second hour is 0.65, then ;
[0100] Represents the average value of the response time distribution of the propagation path, which is the average value of all response times The result of taking the arithmetic mean, for example, ;
[0101] Representative The short-term environmental factors of a propagation path refer to the factors that Corresponding time point On the other hand, the measured values of environmental factors, for example, when When the corresponding environmental factors ;
[0102] Represents the average value of the distribution of short-term environmental factors in the propagation path, which is the average value of all short-term environmental factors The result of taking the arithmetic mean, for example, ;
[0103] represents the total number of propagation paths, in this case, ;
[0104] It represents the weighted index associated with the propagation path response time. Its setting purpose is to adjust the contribution of the response time term to the synchronization index. For example, If it is set to 2, it means that the impact of response time fluctuation on synchronization index is a square relationship. If the response time fluctuation is large, the synchronization index will be significantly increased. The setting reference is, by comparing different The impact of the value on the prediction accuracy in practical applications, for example, through multiple experiments, setting Adjust between 1 and 3. When the prediction accuracy reaches the highest, the final choice ;
[0105] It represents the weighted index associated with the short-term environmental factor of the propagation path. Its setting purpose is to adjust the contribution of the short-term environmental factor item to the synchronization index. For example, If it is set to 1, it means that the impact of environmental factor fluctuations on the synchronization index is linear. If the environmental factor fluctuates greatly, the synchronization index will increase linearly. The setting reference is, by comparing different The impact of the value on the prediction accuracy in practical applications, for example, through multiple experiments, setting Adjust between 0.5 and 2. When the prediction accuracy reaches the highest, the final choice ;
[0106] The system uses the formula The synchronization index of the environmental factors of the propagation path is obtained. This formula measures the synchronization between the response time and the environmental factors by calculating the degree to which they deviate from their average values, taking the weighted sum and then taking the square root. Indicates the The absolute difference between the response time of each propagation path and the average response time The power reflects the contribution of response time fluctuation to the synchronization index. Indicates the The absolute difference between the short-term environmental factor of each propagation path and the average value of the environmental factor The power reflects the contribution of environmental factor fluctuations to the synchronization index. After adding these two items, Indicates that all The weighted difference of each propagation path is summed up. This sum reflects the total response time and environmental factor fluctuations of all propagation paths, and then divided by The synchronization index is obtained by averaging and finally taking the square root to ensure its dimensional consistency with the original data. Assume , , , , , ,but:
[0107] k=1: ;
[0108] k=2: ;
[0109] k=3: ;
[0110] ;
[0111] The results show that there is a certain degree of synchronization between the response time of the propagation path and the environmental factors. If the S value is smaller, the synchronization is higher, and the changes in environmental factors are more consistent with the response of species diffusion. Conversely, the larger the S value, the lower the synchronization, which helps to screen abnormal propagation paths in subsequent steps. The benefit of the formula is that by introducing the weighted index of response time and environmental factors, the contribution of different factors to the synchronization index can be flexibly adjusted according to actual conditions, thereby more accurately evaluating the correlation between species diffusion and the environment, and providing a more refined measurement for the identification of abnormal propagation paths.
[0112] See also Figure 5 ,The steps for obtaining the environmental factor synchronization deviation prediction dataset are as follows:
[0113] S411: Based on the synchronization index of the propagation path environmental factor, identify the synchronization value set according to the propagation path number, extract the quartiles and set the deviation threshold, filter the propagation path index values, and output the abnormal propagation path number set;
[0114] Based on the synchronization index of the propagation path environmental factor, for example, the synchronization index list is [0.5788, 0.62, 0.55, 0.70, 0.59]. When identifying the synchronization value set by the propagation path number, each propagation path number is associated with its corresponding synchronization index value. For example, the synchronization index of Path001 is 0.5788, and the synchronization index of Path002 is 0.62. When extracting the quartiles and setting the deviation threshold, all synchronization index values are sorted, for example, [0.55, 0 .5788, 0.59, 0.62, 0.70], then calculate the first quartile (Q1), median (Q2) and third quartile (Q3), for example, Q1 is 0.5788, Q2 is 0.59, Q3 is 0.62, then set the deviation threshold, which is the limit used to judge whether the synchronization index is abnormal. Its setting reference is the interquartile range (IQR), IQR=Q3-Q1, for example, IQR=0.62-0.5788=0.0412, and the judgment of the abnormal value is that it falls within the interval Therefore, the deviation threshold is set to , the lower threshold is , the upper threshold is When filtering the propagation path index value, the synchronization index value of each propagation path is traversed and compared with the set deviation threshold. For example, if the synchronization index value is less than 0.517 or greater than 0.6818, the propagation path is considered abnormal. For example, the synchronization index 0.70 is greater than 0.681818, so Path004 is filtered as abnormal. The numbers of all the filtered abnormal propagation paths are collected to finally obtain the abnormal propagation path number set. For example, the abnormal propagation path number set is {Path004}.
[0115] S412: Based on the abnormal propagation path number set, extract the difference between consecutive node intervals of the propagation path and standardize it, identify the ratio of the corresponding time of the propagation path to the node dislocation, and use the formula:
[0116] ;
[0117] Aggregate and average the propagation path misalignment ratio series to generate the propagation path alignment deviation mean series;
[0118] in, represents the average value of the propagation path misalignment ratio series, represents the corresponding time of the mth propagation path, Represents the node misalignment time difference of the mth propagation path, Represents the interval difference between consecutive nodes in the mth propagation path, Represents the total number of propagation paths;
[0119] According to the abnormal propagation path number set, for example, the abnormal propagation path number set is {Path004}, the system extracts the difference between the intervals of consecutive nodes in the propagation path and standardizes it. For Path004, the time interval difference between its consecutive nodes is, for example, in a certain propagation path, data of 10 nodes are collected continuously, and the time intervals are: 1 hour from node 1 to node 2, 1.5 hours from node 2 to node 3, and 1 hour from node 3 to node 4. The time difference between adjacent nodes is calculated, for example, [1, 1.5, 1, ...], and then the difference is standardized, for example, by Z-score standardization ( ), scale the difference to a standard range. Assume that after standardization, the difference sequence is When identifying the ratio of the corresponding time of the propagation path to the node misalignment, the system determines the corresponding time of each abnormal propagation path. For example, the overall propagation time of Path004 is 24 hours. At the same time, the system identifies the node misalignment time difference in the path. The node misalignment time difference refers to the deviation between the actual recorded node time and the expected time. For example, if the expected node should appear at the 3rd hour, but actually appears at the 3.5th hour, the node misalignment time difference is 0.5 hours. The system uses the formula Aggregate and average the propagation path misalignment ratio series, where Represents the average value of the propagation path misalignment ratio series, which is used to quantify the average level of node misalignment in the propagation path; Representative For example, if the overall propagation time of the first abnormal propagation path is 10 hours, then Hour; Representative The node misalignment time difference of a propagation path refers to the deviation between the actual node time and the expected node time. For example, if a node of the first abnormal propagation path is expected to appear at the second hour and actually appears at the 2.5th hour, then the node's Hours, here refers to the sum or average of the time differences of multiple nodes on a single propagation path. For example, if Path004 has 5 nodes and the time differences of each node are 0.1h, 0.2h, 0.15h, 0.05h, and 0.1h respectively, then Hour; Representative The interval difference of consecutive nodes in a propagation path refers to the time interval between adjacent nodes. For example, the interval difference sequence of consecutive nodes in Path004 is [1h, 1.5h, 1h], and its mean is Hours, here refers to the standardized difference between consecutive node intervals, for example, from the above The system can sum and average the values, or for each Take a representative difference, for simplicity, assume is the mean of the normalized sequence of consecutive node interval differences, for example, Hour; represents the total number of propagation paths, e.g. The benefit of the formula is that by considering the corresponding time, node misalignment time difference and node interval difference, it can comprehensively evaluate the time deviation in the propagation path, thereby more accurately quantifying the synchronization deviation of environmental factors. Assuming that the corresponding time of Path004 is hours, node misalignment time difference Hours, difference between consecutive node intervals hours, then Finally, a propagation path alignment deviation mean sequence is generated, which contains the average value of the misalignment ratio of each abnormal propagation path. For example, the propagation path alignment deviation mean sequence is [8.466].
[0120] S413: Based on the propagation path alignment deviation mean sequence, sort the average alignment deviation by propagation path number, identify the corresponding deviation between the propagation path sampling time and interval, use the deviation as the discrete trajectory point of the propagation path, and generate the environmental factor synchronization deviation prediction data set;
[0121] Based on the propagation path alignment deviation mean sequence, for example, the propagation path alignment deviation mean sequence is [8.466]. When sorting the average alignment deviation by the propagation path number, the alignment deviation mean of each propagation path is associated with its corresponding propagation path number and sorted. For example, if there is an additional path, there will be [Path004:8.466, Path005:7.921]. When identifying the corresponding deviation between the propagation path sampling time and interval, the sampling time of each propagation path is compared with the preset ideal sampling interval to identify the deviation. For example, if Path004 The ideal sampling interval is 1 hour, but the actual sampling time interval has a deviation of 0.1 hours. The deviation is used as the discrete trajectory point of the propagation path. For example, for Path004, its alignment deviation average is 8.466, and the time deviation at a certain sampling point is 0.1 hours. This information together constitutes the discrete trajectory point of the propagation path. The discrete trajectory points reflect the deviation of the synchronization of environmental factors. Finally, an environmental factor synchronization deviation prediction dataset is generated. This dataset contains the number of the abnormal propagation path, its average alignment deviation, and the time deviation at each sampling point, providing input for subsequent dynamic change prediction.
[0122] See also Figure 6 ,The steps for obtaining the dataset for dynamic change prediction of ,planting areas are as follows:
[0123] S511: Based on the environmental factor synchronization deviation prediction data set, all nodes are reorganized and sorted on the time axis, the change trend of node density is analyzed, the node change rate per unit time is calculated, and the sliding window node density gradient is obtained;
[0124] According to the synchronization deviation prediction data set of environmental factors, for example, the data set contains the number of abnormal propagation paths, the average alignment deviation, and the time deviation of each sampling point. When all nodes are reorganized and sorted on the time axis, all nodes in the propagation paths (including nodes in normal and abnormal paths) are uniformly sorted according to the time points of their occurrence. For example, all nodes from 0:00 on June 1 to 23:59 on June 7 are arranged in chronological order. When analyzing the changing trend of node density, multiple time windows are set on the time axis (for example, each window is 1 hour), and statistics are collected for each time window. The number of nodes contained in the sliding window is used to analyze the changing trend of node density. For example, from 9 to 10 a.m., the node density increases from 100 / hour to 120 / hour. When counting the node change rate per unit time, the system calculates the change rate of the number of nodes in each time window relative to the previous window. For example, if the node density in the previous hour is 100 and the current hour is 120, the change rate is (120-100) / 100=0.2. Finally, the sliding window node density gradient is obtained, which represents the rate and direction of node density change in different time windows.
[0125] S512: Based on the sliding window node density gradient, extract the node density within the time window and perform a difference operation with the reference density, perform normalization processing on the difference data, uniformly scale it to a standard range, and obtain a normalized difference of cluster density;
[0126] Based on the sliding window node density gradient, for example, the sliding window node density gradient shows the rate and direction of change of node density in different time windows. When extracting the node density in the time window and performing the difference operation with the benchmark density, the current node density value is first extracted from each time window. For example, in a 1-hour time window, the node density is 110 / hour, and then the difference operation is performed with the preset benchmark density. The benchmark density refers to the average node density of the planting area under normal and stable conditions. It is used to measure the degree of deviation of the current node density. Its setting reference is to perform statistical analysis on the node density data of the planting area in the past year and calculate its average value and standard deviation. For example, historical data shows that the average node density of the greenhouse is 90 / hour and the standard deviation is 5 / hour. The benchmark density is set to the historical average value of 90 / hour, and a reasonable fluctuation range is considered to be set so that the density differences of different time windows can be compared on the same scale. Finally, the normalized difference of cluster density is obtained. This difference reflects the degree of deviation of the node density in each time window from the benchmark density and is standardized for subsequent processing.
[0127] S513: Arrange the density offset values in the order of the time window center points based on the cluster density normalization difference, aggregate the time windows within each cluster area, and superimpose the normalization results to reconstruct a complete and sequentially consistent node sequence structure, thereby establishing a planting area dynamic change prediction dataset;
[0128] According to the cluster density normalized difference, for example, the cluster density normalized difference is [0.83, 0.75, 0.90, 0.68, ...], when the density offset values are arranged in the order of the time window center point, each normalized difference is associated with the center point time of its corresponding time window and arranged in time order. For example, if the center point of the first time window is 0:30 on June 1, and the second is 1:30 on June 1, the density offset values are arranged in this order. When the time windows in each cluster area are aggregated and the normalized results are superimposed, the system divides multiple adjacent or related time windows into Different clustering areas, for example, the east and west areas of the greenhouse are divided into different clustering areas, and then the normalized results of the time windows in each clustering area are superimposed or averaged to obtain the overall density offset of the area. When reconstructing a complete and sequentially consistent node sequence structure, the superposition results of each clustering area are reintegrated to form a complete and temporally consistent node sequence. For example, the density offset data of the east and west areas are merged by time, and finally a planting area dynamic change prediction dataset is established. This dataset contains the node density offset information of all clustering areas in the planting area at different time points, providing a basis for predicting the dynamic changes of the planting area.
[0129] The crop disease and insect pest data prediction system is used to implement the above-mentioned crop disease and insect pest data prediction method, and the system includes:
[0130] The risk extraction module obtains the temperature, humidity, and soil moisture distribution information of the planting area environmental parameters, detects the mutation points of the time series interval as the key nodes of the propagation path, determines the continuous decreasing trend of the propagation path, and generates a risk boundary distribution map;
[0131] The propagation calculation module calculates the number of temperature and humidity sign changes based on the inverted data sequence in the risk boundary distribution map, aggregates the propagation intensity values, and generates species diffusion frequency data;
[0132] The synchronization assessment module aggregates distribution information by propagation path based on species diffusion frequency data, identifies the correlation strength index of environmental factors in combination with the response time distribution of propagation paths, and generates the synchronization index of environmental factors in propagation paths;
[0133] The node analysis module screens abnormal propagation paths based on the synchronization index of environmental factors in the propagation path, extracts time series and node interval series, calculates the corresponding difference means, and generates an environmental factor synchronization deviation prediction dataset;
[0134] The density normalization module is based on the environmental factor synchronization deviation prediction dataset, statistics the changes in the number of nodes within the time window, calculates the density gradient difference and performs normalization processing to generate a planting area dynamic change prediction dataset.
[0135] The above are merely preferred embodiments of the present invention and do not limit the present invention in any other form. Any technician familiar with the profession may use the technical content disclosed above to change or modify it into an equivalent embodiment with equivalent changes and apply it to other fields. However, any simple modification, equivalent change and modification made to the above embodiment based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A method for predicting crop disease and insect pest data, characterized in that: The following steps are involved: S1: Obtain temperature, humidity, and soil moisture distribution information from the environmental parameters of the planting area, perform time series interval segmentation, locate the mutation points of environmental parameters and determine the locations of key nodes in the propagation path, determine the risk range of species diffusion, and generate a risk boundary distribution map; S2: Based on the risk boundary distribution map, read the key node sequence of the propagation path, identify the temperature and humidity change direction vectors, perform propagation intensity counting on the vectors and aggregate them by node, count the temperature and humidity change frequencies, and obtain species diffusion frequency data; S3: extracting the propagation path distribution information based on the species diffusion frequency data, combining it with the node response time distribution, identifying the environmental factor correlation strength index, and generating the propagation path environmental factor synchronization index; S4: Based on the synchronization index of the environmental factors of the propagation path, a set of nodes whose synchronization deviates from the median value is screened, the corresponding differences between the node intervals and the response time intervals are analyzed, and an environmental factor synchronization deviation prediction data set is generated.
2. The crop disease and insect pest data prediction method according to claim 1, characterized in that: The risk boundary distribution map includes the location distribution of mutation points, the distribution of key nodes in the propagation path, and the species diffusion risk range. The species diffusion frequency data includes the frequency of propagation node changes, the number of vector changes, and the propagation direction characteristics. The propagation path environmental factor synchronization index includes propagation distribution information, response time distribution, and environmental factor correlation strength index. The environmental factor synchronization deviation prediction data set includes node time distribution, node interval mean, and corresponding difference distribution of propagation paths.
3. The crop disease and insect pest data prediction method according to claim 1, characterized in that: The steps for obtaining the risk boundary distribution map are specifically as follows: S111: Obtaining temperature, humidity, and soil moisture distribution information from the environmental parameters of the planting area, performing time series interval segmentation on the continuous data, identifying mutation points within the interval, extracting points where the two points before and after the current data point show opposite changes as mutation points, and obtaining a mutation point distribution position sequence; S112: Based on the mutation point distribution position sequence, for each mutation point and adjacent data sequence, calculate the continuous change and determine whether it is less than zero, mark the key nodes of the propagation path, combine the mutation points to form a pairing structure, and obtain the species diffusion risk range sequence; S113: Based on the species diffusion risk range sequence, group the transmission paths by starting nodes, calculate the start and end changes, transmission intervals and average transmission rates of each group of transmission paths, integrate the boundary features of the transmission paths, and generate a risk boundary distribution map.
4. The crop disease and insect pest data prediction method according to claim 3, characterized in that: The steps for obtaining the species diffusion frequency data are specifically as follows: S211: Based on the risk boundary distribution map, read the propagation path subsequence within the corresponding time range, identify the index sequence of each propagation path in ascending time order, perform reverse sorting on the temperature and humidity data sequence, and obtain a temperature and humidity difference vector sequence; S212: Extract the change signs of adjacent elements in the temperature and humidity difference vector sequence, if the difference is less than zero, it is considered a propagation event, and count the number of propagations of each segment according to the index segment to obtain a propagation count sequence; S213: Based on the propagation count sequence and in combination with the time span information of each propagation path, the propagation frequency per unit time in the propagation path is calculated, the corresponding frequency of the propagation path is counted, and species diffusion frequency data is established.
5. The crop disease and insect pest data prediction method according to claim 4, characterized in that: The steps for obtaining the synchronization index of the transmission path environmental factors are specifically as follows: S311: Based on the species diffusion frequency data, mapping and aggregating by propagation path number, extracting the associated propagation path number and propagation frequency, grouping and calculating the distribution of each propagation path during the propagation process, and generating a propagation path distribution sequence; S312: Based on the propagation path distribution sequence, extract the propagation path response time series, calculate the fluctuation of the distribution information corresponding to the time point, evaluate the time distribution fluctuation degree of each propagation path response sequence, and generate the propagation path response time distribution sequence; S313: Based on the response time distribution sequence of the propagation path and the distribution of short-time environmental factors of the propagation path, statistical energy concentration and standard deviation are performed to evaluate the correlation structure of environmental factors and obtain the synchronization index of the propagation path environmental factors.
6. The crop disease and insect pest data prediction method according to claim 5, characterized in that: The steps for obtaining the environmental factor synchronization deviation prediction data set are specifically as follows: S411: Based on the synchronization index of the propagation path environmental factor, identify the synchronization value set according to the propagation path number, extract the quartiles and set the deviation threshold, filter the propagation path index values, and output the abnormal propagation path number set; S412: Extracting and normalizing the difference between consecutive node intervals of the propagation path based on the set of abnormal propagation path numbers, calculating the ratio of the corresponding time of the propagation path to the node misalignment, aggregating and averaging the misalignment ratio sequence of each propagation path, and generating a propagation path alignment deviation mean sequence; S413: Based on the propagation path alignment deviation mean sequence, sort the average alignment deviation by propagation path number, identify the corresponding deviation between the propagation path sampling time and interval, use the deviation as the discrete trajectory point of the propagation path, and generate an environmental factor synchronization deviation prediction data set.
7. The crop disease and insect pest data prediction method according to claim 1, characterized in that: The method further comprises step S5: S5: performing cluster analysis on the reorganized node density distribution based on the environmental factor synchronization deviation prediction dataset, calculating the node density gradient in the cluster area, performing a difference operation with the baseline density and performing normalization processing to generate a planting area dynamic change prediction dataset; The planting area dynamic change prediction data set includes time window clustering areas, node density gradient change rate, and normalized difference area distribution records.
8. The crop disease and insect pest data prediction method according to claim 7, characterized in that: The steps for obtaining the planting area dynamic change prediction dataset are as follows: S511: reorganize and sort all nodes on the time axis according to the environmental factor synchronization deviation prediction data set, analyze the change trend of node density, calculate the node change rate per unit time, and obtain the sliding window node density gradient; S512: Based on the sliding window node density gradient, extract the node density within the time window and perform a difference operation with the reference density, perform normalization processing on the difference data, uniformly scale it to a standard range, and obtain a cluster density normalized difference; S513: Arrange the density offset values in the order of the time window center points according to the cluster density normalization difference, aggregate the time windows in each cluster area and superimpose the normalization results, reconstruct a complete and sequentially consistent node sequence structure, and establish a planting area dynamic change prediction dataset.
9. The crop disease and insect pest data prediction system is characterized by: The system is used to implement the crop disease and insect pest data prediction method according to any one of claims 1 to 8, and the system comprises: The risk extraction module obtains the temperature, humidity, and soil moisture distribution information of the planting area environmental parameters, detects the mutation points of the time series interval as the key nodes of the propagation path, determines the continuous decreasing trend of the propagation path, and generates a risk boundary distribution map; The propagation calculation module calculates the number of temperature and humidity change signs based on the inverted data sequence in the risk boundary distribution map, aggregates the propagation intensity value, and generates species diffusion frequency data; The synchronization evaluation module aggregates distribution information by propagation path based on the species diffusion frequency data, identifies the environmental factor correlation strength index in combination with the propagation path response time distribution, and generates a propagation path environmental factor synchronization index; The node analysis module screens abnormal propagation paths based on the synchronization index of the propagation path environmental factors, extracts time series and node interval series, calculates the corresponding difference means, and generates an environmental factor synchronization deviation prediction data set; The density normalization module is based on the environmental factor synchronization deviation prediction data set, counts the changes in the number of nodes in the time window, calculates the density gradient difference and performs normalization processing to generate a planting area dynamic change prediction data set.
Citation Information
Cited By
Disease resistance identification combined potato germplasm resource evaluation method
CN121562972A