A tobacco black shank fine intelligent early warning method and system

By constructing a hierarchical monitoring network and a dynamic Bayesian network model that integrates multi-source data, combined with a long short-term memory network for black shank early warning, the problems of limited monitoring coverage and poor early warning timeliness in existing technologies have been solved. This has enabled efficient and accurate disease early warning and resource allocation, reduced prevention and control costs, and adapted to different environmental changes.

CN120765067BActive Publication Date: 2026-04-17ZHAOTONG CO LTD YUNNAN TOBACCO CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHAOTONG CO LTD YUNNAN TOBACCO CO LTD
Filing Date
2025-07-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing black shank early warning methods suffer from limited monitoring coverage, poor timeliness of early warning, weak targeted control measures, and a lack of multi-source data fusion, causal relationship modeling, and adaptive optimization capabilities. This results in low early warning accuracy, frequent false alarms and missed alarms, unreasonable resource allocation, and difficulty in meeting the needs of precision agriculture.

Method used

A hierarchical monitoring network is constructed, risk classification is performed using hotspot analysis, causal inference is performed using a dynamic Bayesian network model that combines multi-source data fusion, a survival index grid is generated using a maximum entropy model, short-term outbreak probability is predicted using a long short-term memory network, resource scheduling is optimized using a greedy heuristic vehicle routing algorithm, and data security and traceability management are implemented to achieve adaptive optimization of the system.

Benefits of technology

It significantly improves monitoring efficiency and prediction accuracy, reduces disease incidence and control costs, enhances resource utilization efficiency, ensures data security and system reliability, and adapts to environmental changes in different years and regions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765067B_ABST
    Figure CN120765067B_ABST
Patent Text Reader

Abstract

This invention discloses a refined intelligent early warning method and system for tobacco black shank disease, belonging to the field of intelligent early warning technology for agricultural diseases. It constructs a hierarchical monitoring network of basic grid and high-risk subgrids, and uses hotspot analysis for risk classification. A dynamic Bayesian network model is used to perform causal inference on meteorological factors, soil factors, and planting management measures, quantifying the average treatment effect of each driving factor on disease occurrence. A suitability index grid is generated based on a maximum entropy model, and a comprehensive economic risk value index is calculated by combining tobacco yield and purchase price data. A long short-term memory network model is used to infer the outbreak probability grid within the next seven days. A weighted fusion calculation of the comprehensive scheduling score is performed, and a greedy heuristic vehicle routing algorithm is used to generate a control resource scheduling list, which can significantly improve the accuracy of early warning and reduce control costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent early warning technology for agricultural diseases, specifically a refined intelligent early warning method and system for tobacco black shank disease. Background Technology

[0002] Black shank is one of the most devastating soil-borne diseases in tobacco production, caused by Phytophthora infestans, and occurs in all major tobacco-producing regions worldwide. This disease is characterized by rapid spread, high mortality, and difficulty in control. Once an outbreak occurs, it often leads to large-scale yield reductions or even total crop failure in tobacco fields, seriously threatening the sustainable development of the tobacco industry. As the world's largest tobacco producer, with a tobacco planting area exceeding ten million mu (approximately 6.67 million hectares), my country's effective control of black shank is of great significance to ensuring the security of the national tobacco industry.

[0003] Traditional blackleg control relies mainly on manual field inspections and experience-based judgment, which suffers from limited monitoring coverage, poor early warning timeliness, and a lack of targeted control measures. With the development of modern information technology, scholars both domestically and internationally have conducted extensive research in the field of agricultural disease early warning. Existing technologies mainly include statistical prediction models based on meteorological data, disease monitoring technologies based on remote sensing imagery, and decision support systems based on expert systems.

[0004] Existing technologies, by establishing statistical relationship models between disease occurrence and environmental factors, or by using remote sensing technology to monitor changes in crop health status, have improved the scientific rigor and accuracy of disease early warning to some extent. However, they still have certain limitations, such as: the monitoring network layout lacks scientific rigor, often employing a uniform distribution method that fails to highlight key monitoring areas in high-risk regions; data acquisition methods are limited, mainly relying on meteorological data from fixed observation stations, lacking real-time monitoring of key factors such as soil physicochemical properties and planting management measures; prediction models are mostly empirical models based on correlation analysis, lacking a deep understanding of disease occurrence mechanisms and making it difficult to reveal the causal relationships between various driving factors; early warning results are mainly presented in the form of disease incidence probability or risk level, lacking organic integration with prevention and control decisions, and failing to guide the optimal allocation of prevention and control resources; data management and quality control systems are imperfect, with problems such as difficulty in ensuring data authenticity and difficulty in tracing historical data; and early warning systems are mostly static models, lacking continuous learning and adaptive optimization capabilities, making it difficult to adapt to environmental changes in different years and regions.

[0005] The aforementioned technical deficiencies greatly limit the practicality and effectiveness of existing black shank early warning methods, resulting in low accuracy, frequent false alarms and missed alarms, and unreasonable allocation of prevention and control resources, making it difficult to meet the actual needs of modern tobacco precision agriculture development.

[0006] Therefore, there is an urgent need for a refined intelligent early warning method and system for tobacco black shank disease that can achieve multi-source data fusion, causal relationship modeling, short-term accurate prediction, intelligent scheduling decision-making, data security management, and system adaptive optimization. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of the prior art and propose a refined intelligent early warning method and system for tobacco black shank disease to solve the above-mentioned problems.

[0008] The objective of this invention is achieved through the following technical solution: a refined intelligent early warning method for tobacco black shank disease, comprising the following steps:

[0009] S1. Monitoring Network Construction and Data Acquisition: A hierarchical monitoring network is constructed through a basic grid and high-risk subgrids. Hotspot analysis is used for risk classification, and field monitoring data, meteorological remote sensing data, soil physicochemical data, and planting planning data are acquired and processed.

[0010] S2. Multi-source data fusion and causal relationship modeling: A dynamic Bayesian network model is used to make causal inferences on meteorological factors, soil factors and planting management measures, and to quantify the average treatment effect of each driving factor on the occurrence of diseases.

[0011] S3. Long-term economic risk assessment: Based on historical disease locations and multi-year average environmental data, a black shank suitability index grid is generated using the maximum entropy model. The suitability index grid is then multiplied pixel by pixel with the tobacco production grid and the local purchase price grid to obtain the comprehensive economic risk value index grid. The suitability index grid is then dynamically corrected using the average treatment effect.

[0012] S4. Short-term outbreak probability prediction: Based on recent daily environmental grids and planting management characteristics, a long short-term memory network model with monthly incremental training and expert review and optimization is used to infer the daily outbreak probability grid of black shank within the next seven days.

[0013] S5. Intelligent scheduling decision: Based on the comprehensive economic risk value index grid and the short-term outbreak probability grid, a comprehensive scheduling score for high-risk fields is generated through weighted fusion calculation, and fields to be dealt with are selected based on the score.

[0014] S6. Prevention and Control Resource Scheduling: For fields to be treated, a greedy heuristic vehicle routing algorithm is used to generate a scheduling list containing travel routes and operation sequences for prevention and control resource equipment.

[0015] S7. Data Security and Traceability Management: Core data is managed using hash fingerprints through audit triggers, and monthly offline archiving tasks are executed.

[0016] S8. Continuous System Optimization: Through human-machine collaborative learning mechanisms and model performance monitoring, the system achieves adaptive optimization and iterative upgrades.

[0017] Step S1, the construction of the monitoring network, includes: creating a basic grid covering the entire tobacco field boundary through a geographic information system platform, generating a unique identifier for each grid; using a hotspot analysis tool to classify the risk of black shank disease diagnosis coordinates over the years, defining significant hotspot grids as high-risk grids and subdividing them into sub-grids; setting fixed monitoring points in ordinary grids and adding supplementary monitoring points in high-risk sub-grids.

[0018] The acquisition of soil physicochemical data in step S1 includes: performing annual baseline spectral mapping, acquiring reflectance curves using a portable spectrometer, and establishing a spectral inversion model using partial least squares regression in conjunction with laboratory physicochemical calibration; performing quarterly real-time measurements, with technicians using handheld multi-parameter soil pens to measure the monitoring points; and performing monthly random sampling calibration, randomly selecting monitoring points for laboratory chemical analysis, and performing regression correction on the measurement data of the handheld multi-parameter soil pens based on the analysis results.

[0019] The causal relationship modeling in step S2 includes: using a constraint algorithm for structure learning, using a scoring function to run an equivalent search to determine the structure of the directed acyclic graph; estimating the conditional probability table for discrete nodes using priors, and estimating the continuous nodes using a Gaussian conditional linear model; evaluating the model performance through cross-validation, and outputting the log-likelihood, area under the curve, and skill score indicators.

[0020] The comprehensive scheduling score in step S5 is calculated using the following formula: Comprehensive scheduling score = First weight × Normalized comprehensive economic risk value index + Second weight × Maximum outbreak probability grid value within the next seven days. The comprehensive scheduling score is then judged based on a preset threshold to select fields to be dealt with.

[0021] A refined intelligent early warning system for tobacco black shank disease includes: a network construction module configured to construct a hierarchical monitoring network through a basic grid and high-risk subgrids, and to perform risk classification using hotspot analysis; a data acquisition and processing module for acquiring and processing field monitoring data, meteorological remote sensing data, soil physicochemical data, and planting planning data; a causal inference module configured to run a dynamic Bayesian network model to perform causal inference on driving factors to quantify the average treatment effect; and a risk assessment module configured to run a maximum entropy model to generate a suitability index grid, and to combine it with tobacco yield and purchase price data to generate a comprehensive economic risk value index grid, while simultaneously using the average treatment effect to further analyze the risk. The system includes: a dynamic correction module; a short-term prediction module configured to run a long short-term memory network model that undergoes monthly incremental training and expert review and optimization, generating an outbreak probability grid for the next seven days; an intelligent decision-making module configured to calculate a comprehensive scheduling score based on the comprehensive economic risk value index grid and the outbreak probability grid using a weighted fusion formula, and to screen out high-risk fields; a resource scheduling module configured to run a greedy heuristic vehicle routing algorithm for high-risk fields, generating a scheduling list for prevention and control equipment; a data security and traceability module configured to manage the hash fingerprint of core data through audit triggers and perform monthly offline archiving tasks; and a system optimization module configured to achieve adaptive system optimization through a human-machine collaborative learning mechanism and performance monitoring.

[0022] The data acquisition and processing module includes: a spectral data interface and a built-in partial least squares regression model for retrieving soil physicochemical properties;

[0023] A handheld device interface is provided for connecting a handheld multi-parameter soil pen to receive real-time measurement data.

[0024] The laboratory data interface is used to import monthly laboratory analysis results and is configured with a calibration subroutine for regression correction of handheld multi-parameter soil pen measurement data.

[0025] The causal inference module includes:

[0026] The structure learning submodule is configured to determine the dynamic Bayesian network structure using constraint algorithms and scoring functions; the parameter estimation submodule is configured to perform conditional probability table estimation for discrete nodes and Gaussian conditional linear modeling for continuous nodes; and the causal quantification submodule is configured to calculate the average treatment effect of each driving factor.

[0027] The short-term prediction module includes a human-machine collaborative training submodule, which is configured as follows:

[0028] Automatically select samples whose predicted probabilities from the Long Short-Term Memory Network model fall within a preset uncertainty range;

[0029] Generate a checklist containing sample images and location information for plant protection experts to manually annotate;

[0030] The expert annotation results will be used as high-weight samples and included in the next round of monthly incremental training dataset.

[0031] The data security and traceability module includes: an audit trigger submodule, configured to automatically generate hash fingerprints when core data such as field monitoring data, model version information, and scheduling lists are changed at the database level; a verification submodule, configured to run a verification script to recalculate the hash and compare it with the stored fingerprint, recording and alarming when inconsistencies are found; and an offline archiving submodule, which connects to a storage drive that can be written to once and read multiple times to perform monthly offline archiving tasks for core data and audit logs.

[0032] The beneficial effects of this invention are:

[0033] By combining hotspot analysis technology with a hierarchical monitoring network, and through the layered construction of a basic grid and high-risk subgrids, accurate identification and key monitoring of high-risk areas were achieved under limited monitoring resources, significantly improving monitoring efficiency. The rapid detection technology for soil physicochemical parameters, combining portable spectrometers with partial least squares regression models, broke through the technical bottlenecks of long analysis cycles and high costs associated with traditional laboratory methods, greatly improving detection timeliness and significantly reducing detection costs, providing technical support for large-scale real-time monitoring. The dynamic Bayesian network model, through constraint algorithms and scoring function optimization, achieved deep causal inference of meteorological factors, soil factors, and planting management measures. It can quantify the average treatment effect of each driving factor on disease occurrence, providing clear causal evidence for scientific prevention and control decisions, and effectively overcoming the subjectivity and uncertainty of traditional experience-based judgments.

[0034] A long short-term memory network model, optimized through monthly incremental training and expert review, combined with a human-machine collaborative learning mechanism, achieves accurate prediction of the probability of blackleg outbreaks within the next seven days, with a prediction accuracy significantly superior to existing early warning methods. By weighted fusion calculation of the comprehensive economic risk value index and short-term outbreak probability, the generated comprehensive scheduling score simultaneously considers economic losses and disease risk, making the identification of high-risk fields more scientific and reasonable, and effectively reducing false alarm and false negative rates. The greedy heuristic vehicle path algorithm established in this invention can generate optimal travel routes and operation sequences for prevention and control resource equipment, significantly improving resource utilization efficiency and effectively shortening prevention and control operation time.

[0035] By employing audit triggers to manage core data using hash fingerprints, combined with monthly offline archiving tasks and write-once-read-many storage technology, a complete data security and traceability management system has been constructed. This effectively ensures data integrity and provides a reliable basis for agricultural insurance claims and liability tracing. The verification script designed in this invention can monitor data consistency in real time and immediately record alarms when anomalies are detected, significantly improving data security.

[0036] This invention can significantly reduce the incidence and control costs of black shank disease. Through precise early warning and scientific scheduling, it effectively reduces the disease incidence, significantly decreases control costs, and significantly increases the proportion of high-quality tobacco leaves. By employing intelligent resource allocation and precise control measures, this invention greatly improves the utilization efficiency of control resources, effectively reduces control costs per unit area, and achieves excellent input-output benefits.

[0037] The technical solution adopted in this invention features high standardization, flexible equipment configuration, and low maintenance costs, making it easy to promote and apply in tobacco-producing areas of different sizes and regions. The human-machine collaborative learning mechanism established in this invention enables the system to continuously improve and optimize itself in practical applications, effectively adapting to environmental changes and disease evolution patterns in different years and regions.

[0038] This invention helps promote the transformation and upgrading of traditional agriculture to smart agriculture, effectively improving the technological content and management level of agricultural production. Through precise early warning and scientific prevention and control, it can effectively reduce pesticide use, significantly reduce environmental pollution, and actively promote green agricultural development. The technical solution provided by this invention helps stabilize tobacco farmers' income, effectively improve the quality of tobacco products, and significantly enhance the international competitiveness of my country's tobacco industry. The technical system and methodology established by this invention have strong versatility and, with appropriate adjustments, can be applied to disease early warning and prevention and control of other crops, showing broad prospects for promotion. Attached Figure Description

[0039] Figure 1 The process of this invention Figure 1 ;

[0040] Figure 2 The process of this invention Figure 2 ;

[0041] Figure 3 This is a system architecture diagram of the present invention. Detailed Implementation

[0042] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0043] It should be noted that the directional concepts of "left", "right", "up", "down", "front", "back", "inner", and "outer" in the following scheme are all relative directions, and will not be listed one by one here.

[0044] Example 1:

[0045] like Figures 1 to 3 As shown, this embodiment provides a basic implementation plan for a refined intelligent early warning method for black shank disease in tobacco. Through technologies such as constructing a hierarchical monitoring network, multi-source data acquisition and processing, soil spectral detection, and long-term economic risk assessment, it achieves basic early warning functions for black shank disease. This embodiment is particularly suitable for grassroots agricultural technology extension agencies such as county-level plant protection stations.

[0046] A tiered monitoring network and risk classification were implemented. A basic grid of 2km × 2km squares was created in the geographic information system platform to ensure complete coverage of the vector boundaries of all tobacco fields. Historical black shank disease diagnosis coordinate data were imported into the system, and spatial statistical analysis was performed using the Getis-Ord Gi* hotspot analysis tool to generate a grid-level Z-value distribution. The formula for hotspot analysis is as follows:

[0047]

[0048] in: Let i be the Getis-Ord Gi* statistic; The attribute value of grid j (number of defect points); The spatial weights between grid i and j; The mean of all grid attribute values; The standard deviation of the attribute values; This represents the total number of grid cells.

[0049] By setting a significance threshold Significant hotspot grids are defined as high-risk grids. Within each high-risk grid, subdivisions of 1 km x 1 km are created to form a tiered monitoring network.

[0050] To improve the sensitivity of the monitoring network to spatiotemporal changes in disease incidence, this embodiment introduces an adaptive hotspot dynamic weighting algorithm. This algorithm dynamically adjusts the weights of monitoring points based on historical disease frequency, seasonal fluctuations, and spatial proximity effects.

[0051] The core formula of the algorithm is:

[0052]

[0053] in:

[0054] To monitor the dynamic weight of point $i$ at time $t$, Use the base weight (default is 1.0). As a time-sensitive factor, Spatial influence factor, Persistent decay factor

[0055] Formula for calculating time sensitivity factor:

[0056]

[0057] in: The seasonal intensity coefficient is 0.3-0.8. The time decay coefficient is 0.01-0.05. This is the historical peak date for the incidence of the disease.

[0058] Spatial impact factor is based on the disease history in the neighborhood:

[0059]

[0060] in: For point The neighborhood set; For neighborhood points The historical severity of the disease; Spatial distance; This is the spatial influence radius parameter.

[0061] The persistent decay factor reflects the pathogenesis memory effect:

[0062]

[0063] in: This refers to the time interval since the last onset of illness; The memory decay period is typically 180 days.

[0064] A standard grid is equipped with one fixed monitoring point; a high-risk subgrid is equipped with an additional supplementary monitoring point to achieve higher monitoring coverage.

[0065] The soil spectral detection and inversion model adopts a three-tiered strategy combining annual baseline spectral mapping, quarterly real-time measurements, and monthly random sampling calibration. Annual baseline spectral mapping utilizes a portable visible, near-infrared, and mid-infrared spectrometer to perform area scanning, with a sampling depth of 0-20 cm, and at least one spectral sampling point randomly selected for each 1 km grid.

[0066] Laboratory calibration was performed according to national standard methods to determine key indicators such as pH, volumetric water content, organic matter content, and cation exchange capacity. After sequentially performing Savitzky-Golay smoothing, standard normal transformation, and first derivative processing on the spectral curves, a partial least squares regression (PLSR) method was used to establish an inversion model.

[0067] The standard normal transformation formula is:

[0068]

[0069] The partial least squares regression model is based on linear equations:

[0070] The evaluation index uses the coefficient of determination from 10-fold cross-validation. and root mean square error :

[0071]

[0072] Requires all target attributes and It meets the accuracy requirements of soil testing standards in the agricultural industry.

[0073] Quarterly real-time measurements were performed using a handheld multi-parameter soil pen. Measurements were taken every 7 days at permanent monitoring sites and every 15 days at mobile monitoring sites. For monthly sampling and calibration, at least 5% of the monitoring sites were randomly selected for re-collection of soil samples, which were then sent to the laboratory for chemical analysis of the same indicators. The regression equation is as follows:

[0074]

[0075] If the regression slope deviates Or intercept deviation The system prompts you to recalibrate the device.

[0076] Long-term economic risk assessment and MaxEnt modeling were employed. Based on historical disease locations and multi-year average environmental data, a maximum entropy (MaxEnt) model was used to generate a black shank suitability index raster. Positive point data consisted of unique samples obtained from monitoring networks and historical data. Background samples were randomly selected from the study area at a ratio of 1.5 times the number of existing samples.

[0077] Predictors were selected from eight layers: multi-year average temperature, precipitation, temperature seasonality, precipitation seasonality, topographic humidity index, soil pH, soil organic matter, and mid-year normalized vegetation index. Any highly correlated pairs with an absolute correlation coefficient greater than 0.8 were removed.

[0078] The sample segmentation uses a spatial k-means clustering method with 70% training and 30% validation, outputting the training AUC, validation AUC, and true skill statistic (TSS):

[0079]

[0080] Model performance requirement verification .

[0081] To integrate the potential map with economic decision-making, the system performs a pixel-by-pixel multiplication operation on the susceptibility index grid, the county-level tobacco production grid, and the local purchase price grid to obtain the Comprehensive Economic Risk Value Index (CRI):

[0082]

[0083] in: The comprehensive economic risk value index (¥ / km²); The suitability index (dimensionless, 0-1); Tobacco production (kg / ha); The price is the local purchase price (¥ / kg).

[0084] The system uses Jenks' natural breakpoint method to divide the comprehensive economic risk value index into 5 levels: very low, low, medium, high, and very high.

[0085] Data security and hash fingerprint management: Add fingerprint, creation timestamp, and last update timestamp fields to core tables; use a unified audit trigger to immediately calculate row-level SHA-256 hash values ​​after each insert or update.

[0086]

[0087] The system runs a verification script during off-peak hours each day to recalculate the hash value of core table records added or modified the previous day and compare it with the stored fingerprint.

[0088]

[0089] At the end of each month, the core and audit tables' data for that month are automatically exported as a compressed file. The compressed file is then written to a Write-once-Read-many (WORM) medium, and a packet-level SHA-256 digest is calculated.

[0090]

[0091] Workflow

[0092] This embodiment establishes a basic grid and high-risk sub-grids covering the entire area through a geographic information system platform to form a hierarchical monitoring network; it regularly collects field disease samples, soil physicochemical parameters, meteorological remote sensing data, and planting planning information at each monitoring point; it uses a dynamic Bayesian network model to make causal inferences on various driving factors; it constructs a maximum entropy suitability model based on historical data and environmental factors to generate a long-term spatial potential assessment; and it combines the suitability index with economic value to form a comprehensive risk assessment result.

[0093] This embodiment can achieve significant results, including reducing the incidence of diseases by no less than 15%, reducing prevention and control costs by no less than 10%, and increasing the proportion of high-quality tobacco leaves by no less than 8%.

[0094] Example 2:

[0095] like Figures 1 to 3 As shown, this embodiment, based on Embodiment 1, further integrates causal relationship modeling and short-term outbreak probability prediction functions. It employs intelligent algorithms such as dynamic Bayesian networks and long short-term memory networks to achieve accurate prediction and causal analysis of black shank disease. This embodiment is particularly suitable for municipal-level plant protection agencies with strong technical capabilities.

[0096] Multi-source data fusion and causal relationship modeling employ dynamic Bayesian network structure learning.

[0097] Based on the monitoring network in Example 1, a monthly offline-updated dynamic Bayesian network was established to quantify the causal contributions of meteorological moisture factors, soil factors, and planting management practices to short-term outbreaks of blackleg. Each month, the following data were extracted: daily average precipitation, temperature, and relative humidity grids; the latest pH and volumetric moisture content grids for the current season; variety resistance scores, growth stage codes, and management intensity indices for the current planting grids; and the monthly average disease incidence rate summarized by grid identifier.

[0098] Standardized for continuous variables using standard scores:

[0099] Structure learning is performed using a constrained algorithm, and a greedy equivalence search is run on the candidate skeleton using the Bayesian information criterion as the scoring function.

[0100] Where: L is the likelihood function; k is the number of model parameters; and n is the number of samples.

[0101] For consecutive nodes, the conditional mean and covariance are estimated using a Gaussian conditional linear model:

[0102]

[0103] Causal quantification and model evaluation were conducted, and intervention calculations were performed on 72-hour cumulative precipitation, soil moisture, and management intensity index to calculate the average treatment effect.

[0104] do-calculus implementation:

[0105]

[0106] The average treatment effect inferred from the model is directly weighted pixel-by-pixel with the maximum entropy fitness index to generate a comprehensive risk index:

[0107]

[0108] in: To normalize weights .

[0109] To improve the reliability of causal inference results, this embodiment proposes a multivariate causal relationship reliability assessment algorithm. This algorithm comprehensively evaluates the reliability of causal relationships from three dimensions: statistical significance, biological plausibility, and spatiotemporal consistency.

[0110] The formula for calculating the overall reliability index of the algorithm is:

[0111]

[0112] in: The causal relationship confidence index between variable i and variable j; Weighting coefficients .

[0113] The statistical significance score is based on the edge weights of the Bayesian network:

[0114]

[0115] in: The coefficient represents the causal relationship. is the standard error; n is the number of samples; p is the number of parameters.

[0116] Biological rationality score combined with expert knowledge:

[0117]

[0118] in: Rate the experts (0-1); For observation time lag; This is the theoretically optimal time delay; This represents the prior probability.

[0119] Spatiotemporal consistency score measures the stability of a relationship:

[0120]

[0121] in: Spatial variance; For time variance; This is the stability threshold.

[0122] Based on the reliability index, the system automatically filters causal relationships with high credibility:

[0123]

[0124] in: This is the reliability threshold (0.7 is recommended).

[0125] For short-term outbreak probability prediction, an LSTM model architecture and training approach was adopted, employing a technical solution of monthly offline incremental training, sliding window daily inference, and regular expert review. The model input consists of daily average meteorological grid, the latest soil moisture grid, planting management feature grid, and the disease rate from the previous observation round. The output is a daily blackleg outbreak probability grid for the next 7 days.

[0126] LSTM cell state update formula:

[0127]

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] in: Forgotten Gate; For input gates; For output gate; In cellular state; It is in a hidden state.

[0134] Incremental training is performed at midnight on the 1st of each month, using samples added in the last 30 days for iterative training. The loss function used is binary cross-entropy.

[0135]

[0136] The human-machine collaborative training mechanism uses the uncertainty of the prediction probability to filter samples that require expert annotation:

[0137]

[0138] Where: p is the model prediction probability; U is the uncertainty, the smaller the value, the more uncertain the model is.

[0139] The system automatically filters samples whose predicted probabilities of the long short-term memory network model are within a preset uncertainty range, generates a checklist containing sample images and location information, which is then manually labeled by plant protection experts. The expert labeling results are then included as high-weight samples in the next round of monthly incremental training dataset.

[0140] Intelligent scheduling decision-making and comprehensive scheduling score calculation: Based on the comprehensive economic risk value index grid and the short-term outbreak probability grid, a comprehensive scheduling score for high-risk fields is generated through weighted fusion calculation.

[0141] First, based on the results of S5 multi-source data fusion and causal relationship modeling, the enhanced comprehensive risk index is calculated:

[0142]

[0143] in: To normalize weights , The average treatment effects are precipitation, soil moisture, and management intensity, respectively.

[0144] Then, the final comprehensive scheduling score is calculated by combining the short-term outbreak probability:

[0145]

[0146] in: First weight and second weight ; This is the normalized enhanced composite risk index; This represents the grid value representing the maximum probability of an outbreak within the next 7 days.

[0147] Threshold filtering rules:

[0148]

[0149] System performance evaluation

[0150] Set strict quality control thresholds:

[0151] Bayesian network model: AUC ≥ 0.75

[0152] LSTM prediction model: AUC ≥ 0.80, Brier Score ≤ 0.20

[0153] Causal effect drift: relative change ≤ ±25%

[0154] This embodiment uses dynamic Bayesian network causal inference to gain a deeper understanding of the causal relationship between various environmental factors and management measures on disease occurrence; it employs long short-term memory network for short-term probability prediction, which significantly improves the timeliness and accuracy of early warning; the human-machine collaborative training mechanism ensures that the model can continuously learn and optimize; and the multi-weight fusion decision-making mechanism comprehensively considers economic losses and disease probability, making resource allocation more scientific and rational.

[0155] Compared with traditional methods, this embodiment improves the prediction accuracy by no less than 20%; through short-term accurate prediction, the early warning period can be extended to 7 days; and the prevention and control cost is further reduced by no less than 15%.

[0156] Example 3:

[0157] like Figures 1 to 3 As shown, this embodiment, based on embodiment 2, further includes advanced functions such as intelligent resource scheduling, a greedy heuristic vehicle routing algorithm, job execution monitoring, data security and traceability management, and system performance monitoring, forming a complete closed-loop system for refined intelligent early warning and prevention of black shank disease. This embodiment is particularly suitable for high-tech institutions such as provincial agricultural management departments.

[0158] To achieve dynamic and balanced allocation of equipment resources, this embodiment proposes an intelligent load balancing scheduling algorithm. This algorithm comprehensively considers differences in equipment capabilities, task urgency, and regional workload to achieve optimal resource allocation.

[0159] The core objective function of the algorithm is:

[0160]

[0161] in: For load imbalance; Total number of devices; For equipment The workload; For equipment Processing capacity; This represents the system's average utilization rate. For device weights; This is the delay penalty coefficient; For the task Priority; For the task The delay time.

[0162] Calculation of system average utilization rate:

[0163]

[0164] Formula for assessing equipment dynamic capabilities:

[0165]

[0166] in: For basic equipment capabilities; This is the aging degradation coefficient; The service life of the equipment; This is a real-time state correction factor; These are weather-related factors.

[0167] Dynamic task priority adjustment strategy:

[0168]

[0169] in: Basic priority; This represents the time urgency coefficient. For task submission time; The deadline; This is a risk level adjustment factor.

[0170] Load redistribution decision function:

[0171]

[0172] in: The amount by which the load balance is improved; The redistribution threshold; This serves as a feasibility constraint.

[0173] System load warning mechanism:

[0174]

[0175] Based on the comprehensive scheduling score calculation in Example 2, a Greedy-VRP algorithm is further designed to generate the optimal travel route and operation sequence for the control resource equipment.

[0176] The core idea of ​​the algorithm is to gradually construct a feasible solution using a greedy strategy, selecting the optimal local decision for the current state at each step. Given m control devices and n fields to be treated, the objective function is:

[0177]

[0178] in: Let be the travel time from field i to field j; For the time allotted for the assignment; This is a binary decision variable, indicating whether device k travels from field i to field j.

[0179] The constraints include:

[0180] Equipment capacity constraints: ,in For the operational needs of field i, For the capacity of device $k$

[0181] Field access constraints: To ensure that each field is served by only one device.

[0182] Path continuity constraints:

[0183] The implementation of the greedy algorithm follows the greedy-VRP optimization process in design scheme S8:

[0184] Initialization: All devices start from the dispatch center.

[0185] For each device, select the nearest high-scoring field that has not been visited.

[0186] Update device location and remaining capacity

[0187] Repeat steps 2-3 until all fields are allocated or equipment capacity is exhausted.

[0188] Equipment returned to dispatch center

[0189] The algorithm's time complexity is Where N is the number of fields and M is the number of equipment.

[0190] Based on the comprehensive scheduling score in Example 2, further consideration is given to equipment accessibility and operation time windows:

[0191] First, calculate the basic integrated scheduling score:

[0192]

[0193] Then add accessibility and time window coefficients:

[0194]

[0195] in: The reachability coefficient is (0-1). The time window coefficient is (0-1).

[0196] Job execution monitoring system real-time status monitoring equipment

[0197] Establish a real-time equipment status monitoring mechanism to track the location, operation progress, and equipment status of prevention and control equipment throughout the entire process. Monitoring indicators include:

[0198] Location information: GPS coordinates, altitude, movement speed

[0199] Operating status: Remaining pesticide solution, spraying pressure, and operating area.

[0200] Equipment health: Battery level, engine status, fault codes

[0201] Equipment status data is uploaded every 30 seconds, and the system calculates job coverage in real time.

[0202]

[0203] Real-time evaluation of work quality based on equipment trajectory data and operational parameters:

[0204]

[0205] in:

[0206]

[0207]

[0208]

[0209] Building upon the hash fingerprint management in Example 1, a hierarchical data security system is further established:

[0210] Core business data: monitoring raw records, model prediction results, scheduling execution logs

[0211] Row-level hash fingerprints are calculated using the SHA-256 algorithm:

[0212]

[0213] Establish a multi-layered backup mechanism, including local backup, cloud backup, and WORM media storage.

[0214] Implement blockchain-based link verification to ensure data time-series integrity.

[0215] Operational audit data: User operation logs, system operation logs, device status logs

[0216] Fast integrity verification using the MD5 algorithm

[0217] Establish an operational traceability chain to record the complete process of data creation, modification, and deletion.

[0218] Data integrity verification adopts the same verification method as design scheme S9:

[0219]

[0220] Chained hash verification uses a time-series connection method:

[0221]

[0222] in: Let be the hash value of the nth record; $Data_n$ is the hash value of the previous record; $Data_n$ is the content of the current record. For timestamps.

[0223] Offline archiving and disaster recovery: Establish a three-tier offline archiving system.

[0224] Daily backup: Automatically backs up core data to local RAID storage daily.

[0225] Monthly archiving: Core data is compressed and archived to WORM media monthly.

[0226] Grade-level data archiving: All data will be archived to an off-site disaster recovery center each year.

[0227] Archived data integrity verification:

[0228]

[0229] Archive integrity check:

[0230]

[0231] System performance monitoring and optimization employs an adaptive optimization algorithm.

[0232] Establish a system performance adaptive optimization mechanism to dynamically adjust key parameters based on actual operating results:

[0233] Prediction accuracy monitoring:

[0234]

[0235] If the accuracy rate is below the set threshold for 7 consecutive days, the model retraining mechanism is triggered.

[0236] Resource utilization optimization:

[0237]

[0238] Based on historical utilization data, dynamically adjust equipment scheduling strategies.

[0239] Establish a human-machine collaborative learning effectiveness evaluation system:

[0240] Consistency of expert annotations:

[0241]

[0242] Improvement in learning effectiveness:

[0243]

[0244] Application Cases

[0245] The following example, using a tobacco-growing area in a certain city, demonstrates the practical application effect of this embodiment. This city, as a major tobacco-producing area, has a planting area exceeding 150,000 mu (approximately 10,000 hectares) and an annual tobacco leaf production of about 350,000 dan (approximately 16,500 tons). Its complex and varied terrain makes black shank disease characterized by its scattered distribution, difficulty in prediction, and high control costs.

[0246] Launched in March 2024, this project was jointly implemented by a tobacco monopoly bureau and a municipal agriculture and rural affairs bureau, covering three key tobacco-producing counties. It established a three-tiered technical architecture consisting of a provincial command center, county-level execution terminals, and a farmer participation interface. The provincial command center deployed a high-performance server cluster with 64 CPU cores, 256GB of RAM, and 100TB of storage capacity, responsible for dynamic Bayesian network model training, multi-source data fusion, comprehensive analysis, and decision support across the province. The county-level execution terminals were equipped with industrial-grade servers, UAV ground stations, mobile command vehicles, and field monitoring equipment, responsible for the construction of the monitoring network, data collection, resource scheduling, and execution monitoring within their respective counties. A smart tobacco app was developed for growers, providing services such as disease early warning push notifications, prevention and control guidance, and operation scheduling.

[0247] The system implementation was carried out in four phases. Phase 1 involved system deployment and data accumulation, including the establishment of 126 fixed monitoring points, creating a 2km×2km grid monitoring system covering the main tobacco-growing areas of the three counties, installing 45 automatic weather stations, and establishing a database of black shank incidence over the past four years. Phase 2 focused on model training and parameter optimization, integrating five years of historical meteorological remote sensing data and three years of remote sensing imagery data to create a comprehensive dataset containing 18 environmental variables and 12 management variables. The area under the curve of the dynamic Bayesian network model reached 0.87, and the 7-day prediction accuracy of the LSTM model reached 83.5%. Phase 3 involved trial operation and effect verification. The intelligent prevention and control system was officially launched on August 1st, covering 68,000 mu of tobacco fields in the three counties. The incidence of black shank decreased from 8.3% to 6.8%, the timeliness of prevention and control increased from 65.2% to 89.7%, and equipment utilization increased by 82.4%. The fourth phase involves system optimization and promotion preparation, including three months of data integrity verification, and a second round of model parameter tuning based on actual operating data, which improved the prediction accuracy to 86.2%.

[0248] In terms of application effectiveness, the accuracy rate of short-term early warning has increased from 55% to 86.2%, the spatial early warning accuracy has improved from the county level to the 2km grid level, and the early warning lead time has been extended from 1-2 days to 7 days. Regarding prevention and control efficiency, equipment operation efficiency has increased by 76.8%, the timeliness rate of prevention and control has reached 89.7%, and the utilization rate of pesticides has increased by 23.5%. In terms of system reliability, the system's continuous operation stability has reached 99.2%, data integrity has reached 99.7%, and the equipment failure rate has decreased to 3.1%.

[0249] From an economic perspective, the project reduced yield losses due to diseases by approximately 2,100 tons, saving 75.6 million yuan in economic losses; reduced prevention and control costs by 37 yuan per mu, resulting in total cost savings of 2.516 million yuan across the three counties; improved tobacco leaf quality, increasing the proportion of Grade 1 tobacco by 8.7%, and generating approximately 12 million yuan in additional income; the total project investment was 18.5 million yuan, with an annual return on investment of 473%. Indirect benefits included a 19.0% reduction in pesticide use, the training of 96 technical personnel, and the creation of a technological demonstration effect.

[0250] This embodiment, through an intelligent resource scheduling system, can significantly improve the utilization efficiency and operational quality of prevention and control equipment; through a greedy heuristic vehicle routing algorithm, it achieves optimal allocation of prevention and control resources; through comprehensive operation execution monitoring, it ensures the timely and effective implementation of prevention and control measures; through enhanced data security and traceability management, it guarantees the integrity and reliability of system data; and through an adaptive optimization mechanism, it enables the system to continuously improve and optimize.

[0251] Compared with traditional prevention and control methods, this embodiment can improve resource utilization efficiency by no less than 25%, shorten operation time by no less than 30%, and further reduce prevention and control costs by no less than 20%, while ensuring that data security and system reliability meet enterprise-level application standards.

[0252] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A refined intelligent early warning method for tobacco black shank disease, characterized in that, Includes the following steps: S1. Monitoring Network Construction and Data Acquisition: A hierarchical monitoring network is constructed through a basic grid and high-risk subgrids. Hotspot analysis is used for risk classification, and field monitoring data, meteorological remote sensing data, soil physicochemical data, and planting planning data are acquired and processed. S2. Multi-source data fusion and causal relationship modeling: A dynamic Bayesian network model is used to make causal inferences on meteorological factors, soil factors and planting management measures. A monthly offline updated dynamic Bayesian network is established to quantify the causal contribution of meteorological moisture factors, soil factors and planting management measures to the short-term outbreak of blackleg. The monthly average precipitation, temperature and relative humidity grids, the latest pH and volumetric water content grids of the current season, the variety resistance score, growth stage code and management intensity index of the current planting grid, and the monthly average disease rate summarized by grid identifier are extracted to quantify the average treatment effect of each driving factor on the occurrence of the disease. S3. Long-term economic risk assessment: Based on historical disease locations and multi-year average environmental data, a black shank suitability index grid is generated using a maximum entropy model. The suitability index grid is then multiplied pixel by pixel with the tobacco production grid and the local purchase price grid to obtain a comprehensive economic risk value index grid. The average treatment effect is then used to dynamically correct the suitability index grid. S4. Short-term outbreak probability prediction: Based on recent daily environmental grids and planting management characteristics, a long short-term memory network model with monthly incremental training and expert review and optimization is adopted. The model input consists of daily average meteorological grids, the latest soil moisture grids, planting management characteristic grids, and the disease rate of the previous observation round. The output is the daily black shank outbreak probability grid for the next 7 days, and the daily black shank outbreak probability grid for the next 7 days is inferred. S5. Intelligent scheduling decision: Based on the comprehensive economic risk value index grid and the short-term outbreak probability grid, a comprehensive scheduling score for high-risk fields is generated through weighted fusion calculation, and fields to be dealt with are selected based on the score. S6. Prevention and Control Resource Scheduling: For the fields to be treated, a greedy heuristic vehicle path algorithm is used to generate a scheduling list containing travel routes and operation sequences for prevention and control resource equipment. S7. Data Security and Traceability Management: Core data is managed using hash fingerprints through audit triggers, and monthly offline archiving tasks are executed. S8. Continuous System Optimization: Through human-machine collaborative learning mechanisms and model performance monitoring, the system achieves adaptive optimization and iterative upgrades.

2. The method according to claim 1, characterized in that, The monitoring network construction in step S1 includes: creating a basic grid covering the entire tobacco field boundary through a geographic information system platform, generating a unique identifier for each grid; using a hotspot analysis tool to classify the risk of black shank disease diagnosis coordinates over the years, defining significant hotspot grids as high-risk grids and subdividing them into sub-grids; setting fixed monitoring points in ordinary grids, and adding supplementary monitoring points in the high-risk sub-grids.

3. The method according to claim 1, characterized in that, The acquisition of soil physicochemical data in step S1 includes: performing annual baseline spectral mapping, acquiring reflectance curves using a portable spectrometer, and establishing a spectral inversion model using partial least squares regression in conjunction with laboratory physicochemical calibration; performing quarterly real-time measurements, with technicians using handheld multi-parameter soil pens to measure the monitoring points; and performing monthly random sampling calibration, randomly selecting monitoring points for laboratory chemical analysis, and performing regression correction on the measurement data of the handheld multi-parameter soil pens based on the analysis results.

4. The method according to claim 1, characterized in that, The causal relationship modeling in step S2 includes: using a constraint algorithm for structure learning, using the Bayesian information criterion as the scoring function to run a greedy equivalence search on the candidate skeleton to determine the directed acyclic graph structure of the dynamic Bayesian network; estimating the conditional probability table for discrete nodes using priors, and estimating the conditional mean and covariance for continuous nodes using a Gaussian conditional linear model; evaluating the model performance through cross-validation, and outputting the log-likelihood, area under the curve, and skill scoring index.

5. The method according to claim 1, characterized in that, The comprehensive scheduling score in step S5 is calculated using the following formula: Comprehensive scheduling score = First weight × Normalized comprehensive economic risk value index + Second weight × Maximum outbreak probability grid value within the next seven days. The comprehensive scheduling score is then judged based on a preset threshold to select the fields to be disposed of.

6. A refined intelligent early warning system for tobacco black shank disease, used to implement the method described in any one of claims 1 to 5, characterized in that, Includes a network building module that communicates with the processor and memory, configured to build a hierarchical monitoring network through a base grid and high-risk subgrids, and to perform risk classification using hotspot analysis; The data acquisition and processing module is used to acquire and process field monitoring data, meteorological remote sensing data, soil physicochemical data, and planting planning data. The causal inference module is configured to run a dynamic Bayesian network model to perform causal inference on driving factors in order to quantify the average treatment effect. The risk assessment module is configured to run a maximum entropy model to generate a suitability index grid, and combine it with tobacco production and purchase price data to generate a comprehensive economic risk value index grid, while using the average processing effect for dynamic correction; the short-term prediction module is configured to run a long short-term memory network model with monthly incremental training and expert review and optimization, and infer to generate an outbreak probability grid for the next seven days. The intelligent decision-making module is configured to calculate a comprehensive scheduling score based on the comprehensive economic risk value index grid and the outbreak probability grid using a weighted fusion formula, and then screen out high-risk fields. The resource scheduling module is configured to run a greedy heuristic vehicle routing algorithm for the high-risk fields to generate a scheduling list for the prevention and control equipment; the data security and traceability module is configured to manage the hash fingerprint of core data through audit triggers and execute monthly offline archiving tasks; the system optimization module is configured to achieve system adaptive optimization through human-machine collaborative learning mechanism and performance monitoring.

7. The system according to claim 6, characterized in that, The data acquisition and processing module includes: a spectral data interface and a built-in partial least squares regression model for retrieving soil physicochemical properties; A handheld device interface is provided for connecting a handheld multi-parameter soil pen to receive real-time measurement data. The laboratory data interface is used to import monthly laboratory analysis results and is configured with a calibration subroutine for regression correction of the handheld multi-parameter soil pen measurement data.

8. The system according to claim 6, characterized in that, The causal inference module includes: The structure learning submodule is configured to determine the dynamic Bayesian network structure using constraint algorithms and scoring functions; the parameter estimation submodule is configured to perform conditional probability table estimation for discrete nodes and Gaussian conditional linear modeling for continuous nodes; and the causal quantification submodule is configured to calculate the average treatment effect of each driving factor.

9. The system according to claim 6, characterized in that, The short-term prediction module includes a human-machine collaborative training submodule, which is configured as follows: Automatically select samples whose predicted probabilities from the Long Short-Term Memory Network model fall within a preset uncertainty range; Generate a checklist containing sample images and location information for plant protection experts to manually annotate; The expert annotation results will be used as high-weight samples and included in the next round of monthly incremental training dataset.

10. The system according to claim 6, characterized in that, The data security and traceability module includes: an audit trigger submodule, configured to automatically generate hash fingerprints when core data of field monitoring data, model version information, and scheduling lists are changed at the database level; a verification submodule, configured to run a verification script to recalculate the hash and compare it with the stored fingerprint, and record and alarm when inconsistencies are found; and an offline archiving submodule, which connects to a storage drive that can be written to once and read multiple times, and performs monthly offline archiving tasks for core data and audit logs.

Citation Information

Patent Citations

  • Tobacco main pest and disease damage prediction method based on big data

    CN110837926A

  • Vegetable disease and pest outbreak prevalence prediction method, medium and system

    CN117852726A