Tobacco pesticide residue dynamic prediction method and system fusing deep learning and mechanism model
By integrating deep learning with mechanism models, a tobacco pesticide residue prediction method was developed to solve the problem of insufficient prediction across varieties, plots, and climate conditions, and to achieve dynamic, explainable, and stable pesticide residue monitoring and decision support.
Patent Information
- Application Number
- CN202511191586.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-25
- Publication Date
- 2025-10-17
AI Technical Summary
Existing tobacco pesticide residue prediction methods have insufficient generalization capabilities across varieties, plots, and climate conditions, lack real-time online assimilation and updating capabilities, cannot effectively achieve dynamic monitoring and decision support, and lack the interpretability and uncertainty characterization of prediction results.
By integrating deep learning with mechanism models, we acquire multi-source data associated with the target plots and perform consistent processing to establish a mechanism model of crop-environment coupling. This model is then coupled with the deep learning model, and physical consistency constraints and online assimilation mechanisms are introduced to achieve dynamic prediction and adaptive calibration.
The accuracy and stability of the prediction are improved, and it can maintain interpretability and transferability under different conditions, output uncertainty intervals and exceedance probabilities, and support planting management and quality control decisions.
Smart Images

Figure CN120808942A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of deep learning, in particular to a tobacco pesticide residue dynamic prediction method and system fusing deep learning and mechanism model. BACKGROUND
[0002] In the process of tobacco planting, various pesticides are often used, and the residual level directly affects product safety and compliance of international trade. The current industry mainstream still relies on laboratory detection means, such as gas chromatography-mass spectrometry (GC-MS) and liquid chromatography-tandem mass spectrometry (LC-MS / MS), to evaluate pesticide residues by regular sampling inspection. To assist prediction, empirical first-order kinetic half-life models or simple statistical regression models are usually used to post-assess and trend extrapolate the residual amount. In recent years, there have also been studies trying to combine environmental factors (such as temperature, humidity, rainfall, and light) with pesticide application records, but most of them remain at the level of offline batch processing and static formula, relying on the simple superposition of empirical parameters, and it is difficult to achieve high timeliness and daily-level dynamic prediction under continuous monitoring conditions in the field.
[0003] With the popularity of Internet of Things sensors, remote sensing technology, and numerical weather prediction, multi-source environmental and agronomic data can be obtained in real time. At the same time, mechanism models describing the fate of crops and compounds are constantly developing. Such models, based on the principle of mass conservation and kinetic equations, can describe the absorption, transport, and metabolism of pesticides in the crop-environment system in detail. The future direction of development is gradually shifting towards "data and mechanism fusion": providing an interpretable prior framework through mechanism models, and combining deep learning models to describe non-linear processes that are difficult to parameterize, such as rainfall-induced erosion, exchange between leaf surface and internal tissue, differences between varieties, and disturbances of field management measures. Fusion modeling combined with online calibration and uncertainty quantification can not only improve the robustness of prediction, but also calculate the over-limit probability of the maximum residue limit (MRL) and the early warning time window, thereby serving the dual needs of field management and market supervision.
[0004] However, the existing method still has obvious defects. On the one hand, the deep learning model purely relying on data performs poorly in out-of-sample scenarios, and has insufficient generalization ability in the face of variety differences, plot differences or climate variations. On the other hand, the purely mechanism model is highly sensitive to some difficult-to-measure environmental and physiological parameters, such as leaf cuticle thickness, stomatal opening rate and inter-tissue distribution coefficient. Not only is the parameter acquisition cost high, but also the stability is difficult to maintain under the change of season and management conditions. In addition, the existing method generally lacks three capabilities: (1) unable to realize real-time online assimilation and update when new observation data arrives; (2) lack of constraints based on physical laws, such as the constraint that the residual amount should not rise under the condition of no new pesticide application; (3) the prediction output is usually limited to point prediction, lacking of unified uncertainty characterization of interval prediction and over-limit probability. These defects make the reliability of the existing method insufficient in dynamic decision-making and compliance determination. SUMMARY
[0005] In order to overcome the deficiencies of the prior art, the purpose of the present application is to provide a tobacco pesticide residue dynamic prediction method and system fusing deep learning and mechanism model, through deep coupling of mechanism model and deep learning model, joint design of online assimilation and physical consistency constraints, to improve the prediction accuracy while ensuring the result to be interpretable, transferable and consistent with the actual law, thereby providing an efficient, reliable and scalable technical solution for dynamic monitoring and scientific decision-making of tobacco pesticide residues.
[0006] To achieve the above purpose, the present application provides the following scheme: A tobacco pesticide residue dynamic prediction method fusing deep learning and mechanism model, comprising: Obtaining pesticide application event data, environmental driving data, agronomic management data, plant phenotype data and historical residue observation data associated with a target plot, and performing time alignment, missing completion, anomaly rejection and dimension unification on the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotype data and the historical residue observation data to form consistent data; Establishing a mechanism model coupled with crops and environment based on mass conservation and dynamics relationship, and generating a mechanism prior trajectory by time-varying simulation using the consistent data; Building a deep learning model, taking the consistent data and the mechanism prior trajectory as joint input, coupling the deep learning model and the mechanism model through parameter sharing or residual mapping to represent the nonlinear effects not fully described by the mechanism model, and obtaining a target coupled model; training the target coupling model based on a joint loss; the joint loss at least includes a fitting term for historical residual observations and a physically consistent constraint term, the physically consistent constraint includes that the predicted residual does not increase within a time window without new pesticide application and mass conservation across compartments; when a new residual observation arrives, using a recursive update method to online assimilate and adaptively calibrate the state and identifiable parameters of the trained target coupling model; outputting residual time series prediction and uncertainty interval at the granularity of harvest batch or leaf position, and calculating over-limit probability, compliance confidence and compliance time window based on maximum residual limit, for decision-making of planting management or quality control system.
[0007] Preferably, pesticide application event data, environmental driving data, agronomic management data, plant phenotype data and historical residual observation data associated with the target plot are obtained, and the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotype data and the historical residual observation data are time-aligned, missing data completed, outliers removed and dimension unified to form consistent data, including: aligning the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotype data and the historical residual observation data according to a unified time stamp, and generating an interpolated time label for data without a time stamp; missing data items are filled in by interpolation based on adjacent time slices, matching under similar environmental conditions or statistical mean; outliers beyond a reasonable threshold range or not conforming to physical laws are detected and corrected by rule removal or alternative filling; data of different sources and different units are normalized according to a unified dimension to obtain the consistent data with consistent structure.
[0008] Preferably, a mechanism model of crop and environment coupling is established based on mass conservation and kinetic relationship, and the consistent data is used for time-varying simulation to generate mechanism prior trajectory, including: a compartment structure including leaf surface compartment, mesophyll compartment and soil surface compartment is established, and leaf surface residual mass, mesophyll residual mass and soil surface residual mass are taken as state variables, and absorption flux, leaching flux, volatilization flux and degradation flux are set between compartments; air temperature, rainfall, relative humidity, irradiance, pesticide dose and pesticide application time in the consistent data are taken as external driving and initial conditions, and pesticide input is given to the leaf surface compartment at the pesticide application time and zero input at the non-pesticide application time; the net change rate of each compartment is described by mass conservation and kinetic relationship to form the following state equation group: ; ; ; where, is the leaf surface residual mass; is the leaf tissue residual mass; is the soil surface residual mass; is the application input; is the leaf surface to leaf tissue absorption coefficient; is the stomata modulation factor as a function of relative humidity; is the leaching coefficient determined by rainfall intensity or amount; is the volatilization coefficient; is the leaf tissue to other parts translocation coefficient; , , are the degradation coefficients of leaf surface, leaf tissue and soil surface, respectively; is the leaching coefficient affected by water content; is the temperature; is the relative humidity; is the rainfall intensity or cumulative rainfall amount; is the soil volumetric water content; to reduce the empirical weight and improve cross-scene transferability, the absorption, leaching, volatilization and degradation processes are determined by using a function form with physical meaning; a one-step method with non-negativity preservation is used for time-varying simulation, and mass conservation check is performed for irreversible loss, and the update formula is: ; ; where, is the residual mass of the ith compartment at time t; Δt is the time step; is the net flux function of the ith compartment; is the external driving vector; is the uncorrected update amount; is the sum of the mass of each compartment; is the allowed total amount obtained by adding the application input to the mass at the previous time and subtracting the irreversible loss; is the total amount of irreversible loss caused by volatilization and degradation; ε is a very small positive number for numerical stability; is the external input to the leaf surface compartment at the application time, the value is determined by the application dose, application method and application time; the compartment state is converted into the residual prediction of the detectable part by using a linear observable mapping, and the mechanism prior trajectory is output: ; where, is the observable residual vector at time t; C is the observation mapping matrix determined by the correspondence between sampling depth and leaf position; is the compartment state vector; the mechanism prior trajectory is the combination of the observable residual vector at each time.
[0009] Preferably, to reduce empirical weights and improve cross-scene transferability, the absorption, leaching, volatilization and degradation processes are determined using function forms with physical meanings, including: The absorption coefficient jointly modulated by stomata and cuticle adopts the following expression: ; Wherein, is the reference absorption coefficient; is the wax thickness sensitivity coefficient; is the equivalent thickness of leaf cuticle wax; is the sensitivity coefficient of stomata opening and closing to relative humidity; is the reference relative humidity of stomata half-opening; Considering the stripping effect of rainfall kinetic energy on the adherend, the leaching coefficient is defined as: ; Wherein, is the sensitivity coefficient to rainfall kinetic energy; is the cumulative rainfall kinetic energy measure within the time window; is the rainfall intensity of the th time slice is the corresponding time slice length; The volatilization and degradation processes adopt temperature-driven Arrhenius-type coefficients: ; Wherein, is the reference volatilization coefficient; is the apparent activation energy of volatilization; R is the gas constant; is the reference temperature; is the reference degradation coefficient of compartment x; is the degradation apparent activation energy of compartment x ; x Take the leaf surface or mesophyll or soil surface layer; The soil water content driven leaching coefficient adopts a power index type seepage relationship: ; Wherein, is the reference leaching coefficient; m is the response index of soil moisture to leaching process.
[0010] Preferably, a deep learning model is constructed, the uniform data and the mechanism prior trajectory are used as joint inputs, and the deep learning model is coupled with the mechanism model through parameter sharing or residual mapping to characterize the nonlinear effects that are not fully described by the mechanism model, thereby obtaining a target coupling model, including: Constructing a deep learning model that takes the uniformized data and the mechanism prior trajectory as joint input, wherein the input form includes a time series, an environmental factor series, and a residual trajectory predicted by the mechanism; A feature extraction unit is provided in the deep learning model to extract temporal features and environment-dependent features from the uniform data and the mechanism prior trajectory, respectively, and perform feature alignment in a unified latent space; By sharing parameters, some parameters are reused in the feature extraction unit to maintain consistent representation capabilities across different input sources. By means of residual mapping, a difference learning module is introduced between the mechanism prior trajectory and the output of the deep learning model, so that the deep learning model can focus on learning the nonlinear deviations that are not fully characterized by the mechanism model; The parameter sharing mechanism is combined with the residual mapping module to form the target coupling model.
[0011] Preferably, the calculation formula of the joint loss is: ;
[0012] in, For the target coupling model at time Output observable residual predictions; It is a historical residual observation; It is a time index collection with no new pesticide application; For the moment the total residual mass of the model summed across compartments; For the moment The total mass allowed by the law of conservation of mass is obtained by adding the total mass at the previous moment to the current amount of injection and deducting the irreversible loss; is the scale factor obtained by robust estimation based on the observation sequence; is the scale factor of the total mass dimension; is the Huber function of robust fitting; is the constant of Huber threshold; is the weight coefficient of the constraint of “no increase when no new pesticides are applied”; is the weight coefficient of the mass conservation constraint.
[0013] Preferably, each of the weight coefficients is automatically normalized to a single constant according to the data scale.
[0014] Preferably, upon arrival of new residual observations, an online assimilation and adaptive calibration of the trained state and identifiable parameters of the target coupled model is performed using a recursive update method, including: receiving new residual observation data and comparing the observation data with the output of the target coupled model at the same time to obtain observation errors; inputting the observation errors into a recursive update unit and correcting the internal state of the target coupled model according to a preset state transition relationship, so that the updated state fits the real observation; on the basis of the corrected state, identifying identifiable parameters that have a significant impact on the prediction of the target coupled model, and adjusting the values of the identifiable parameters according to the error size and direction; performing consistency test on the updated state and identifiable parameters to ensure that the physical constraint conditions are met and that the residual amount is not physically increased; using the corrected state and parameters as new initial conditions for dynamic prediction in the next period to achieve adaptive calibration of the model under different environmental and management conditions.
[0015] Preferably, the residual time series prediction and uncertainty interval are output according to the harvest batch or leaf position granularity, and the over-limit probability, compliance confidence and compliance time window are calculated based on the maximum residual limit to provide decision support for planting management or quality control systems, including: aggregating the residual prediction results output by the target coupled model at each prediction time according to the harvest batch or leaf position granularity to form a residual time series; based on distribution parameter regression, quantile prediction or uncertainty estimation method, generating upper and lower limit intervals for each time of the residual time series to represent the uncertainty of the prediction; comparing the residual time series and uncertainty interval with the maximum residual limit specified in the target area to obtain the over-limit judgment at each time; calculate the over-limit probability by counting the proportion of residual prediction values higher than the maximum residual limit, and calculate the compliance confidence under the regulatory threshold; According to the downward trend of the residual prediction value with time, determine the time point when the residual prediction value first and continuously falls below the maximum residual limit, define as the start of the compliance time window, and output to the planting management or quality control system for decision support of harvesting, processing or warehousing.
[0016] A tobacco pesticide residue dynamic prediction method combining deep learning and mechanism model, including: a data acquisition and unification unit configured to acquire pesticide application event data, environmental driving data, agronomic management data, plant phenotype data, and historical residue observation data associated with a target plot, and to perform time alignment, missing data completion, outlier removal, and dimension unification on the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotype data, and the historical residue observation data to form unified data; a mechanism modeling and simulation unit configured to establish a mechanism model coupled with crops and environment based on mass conservation and kinetic relations, and to perform time-varying simulation on the unified data to generate mechanism prior trajectories; a deep learning coupled modeling unit configured to construct a deep learning model, to use the unified data and the mechanism prior trajectories as joint inputs, and to couple the deep learning model and the mechanism model through parameter sharing or residual mapping to represent nonlinear influences not sufficiently depicted by the mechanism model, so as to obtain a target coupled model; a joint loss training unit configured to train the target coupled model based on a joint loss; the joint loss at least includes a fitting term for historical residue observations and a physical consistency constraint term; the physical consistency constraint includes that the predicted residue does not increase within a time window without new pesticide application, and mass conservation across compartments; an online assimilation and adaptive calibration unit configured to, when new residue observations arrive, perform online assimilation and adaptive calibration on the state and identifiable parameters of the trained target coupled model through a recursive update method; a prediction output and risk determination unit configured to output residue time series prediction and uncertainty interval at a harvest batch or leaf position granularity, and to calculate over-limit probability, compliance confidence, and compliance time window based on maximum residue limit, so as to provide decision-making for planting management or quality control systems.
[0017] According to the specific embodiments of the present application, the following technical effects are disclosed: Compared with a model that simply relies on data driving, the present application explicitly introduces physical processes and agronomic management priors through the fusion of mechanism models and deep learning models, thereby significantly improving the prediction generalization ability across varieties, plots, and climate conditions, and avoiding the defects of traditional methods in poor performance in out-of-sample scenarios.
[0018] By introducing physical constraints and data-driven corrections in joint modeling, the present application effectively alleviates the high sensitivity of pure mechanism models to difficult-to-measure parameters (such as leaf cuticle wax thickness and inter-tissue distribution coefficients), and reduces the parameter identification cost and the prediction instability caused by seasonal drift.
[0019] The online assimilation mechanism can recursively correct the model state and parameters when new observation data arrives, realize dynamic updating and adaptive calibration of the model, and thus ensure the continuous accuracy of the prediction result in long-term application, instead of relying on a static model with fixed parameters.
[0020] The application ensures that the prediction result conforms to the actual physical law by embedding the physical consistency constraints of "no increase in residual under no new application condition" and "mass conservation across compartments" in the joint loss function, and avoids the "numerical drift" or "non-physical increment" that may occur in conventional deep learning models.
[0021] The application not only outputs point prediction, but also generates uncertainty interval through distribution parameter estimation or interval prediction, and further calculates the over-limit probability and compliance confidence, to provide more comprehensive risk information for supervision and decision-making, while the traditional method can only give single numerical prediction.
[0022] The application can directly provide decision basis for harvesting, processing and warehousing by predicting the time series of residual amount and determining the compliance time window, help the grower to reasonably arrange the harvest period, reduce the economic loss and trade barriers caused by non-compliance, and has significant practical value. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 The method flowchart provided for the embodiments of the present application is provided. Figure 2 The system structure schematic diagram provided for the embodiments of the present application is provided. DETAILED DESCRIPTION
[0025] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0026] The purpose of the present application is to provide a tobacco pesticide residue dynamic prediction method and system combining deep learning and mechanism model, which significantly improves the accuracy and practicability of tobacco pesticide residue prediction.
[0027] In order to make the above objectives, characteristics and advantages of the present application more apparent, comprehensible and easier to be understood, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0028] Figure 1 The method flowchart provided by the embodiments of the present application is shown in Figure 1 The present application provides a tobacco pesticide residue dynamic prediction method combining deep learning and mechanism model, which comprises the following steps: Step 100: Obtain pesticide application event data, environmental driving data, agronomic management data, plant phenotype data and historical residue observation data associated with the target plot, and perform time alignment, missing completion, abnormality elimination and dimension unification on the pesticide application event data, environmental driving data, agronomic management data, plant phenotype data and historical residue observation data to form consistent data; Step 200: Establish a mechanism model coupled with crops and environment based on the mass conservation and kinetic relationship, and use the consistent data to generate a mechanism prior trajectory through time-varying simulation; Step 300: Construct a deep learning model, use the consistent data and the mechanism prior trajectory as joint input, and couple the deep learning model and the mechanism model through parameter sharing or residual mapping to represent the nonlinear influence not fully described by the mechanism model, so as to obtain a target coupled model; Step 400: Train the target coupled model based on a joint loss; the joint loss at least includes a fitting term for historical residue observation and a physical consistency constraint term, the physical consistency constraint includes that the predicted residue does not increase within a time window without new pesticide application and the mass conservation across compartments; Step 500: When new residue observation arrives, use a recursive updating method to perform online assimilation and adaptive calibration on the state and identifiable parameters of the trained target coupled model; Step 600: Output residue time series prediction and uncertainty interval at harvest batch or leaf position granularity, and calculate over-limit probability, compliance confidence and compliance time window based on maximum residue limit, so as to provide decision-making for planting management or quality control system.
[0029] In a preferred embodiment of the present application, step 100 first collects multi-source data associated with the target plot, including pesticide application event data, environmental driving data, agronomic management data, plant phenotype data, and historical residue observation data. The pesticide application event data is used to characterize the dosage, mode, and time of pesticide application at different time points; the environmental driving data includes continuous environmental factors such as temperature, humidity, rainfall, wind speed, irradiance, and soil moisture content; the agronomic management data includes operation information such as irrigation, mulching measures, and harvesting arrangements; the plant phenotype data is used to reflect the appearance and physiological state of tobacco at different growth stages, such as leaf vigor index and variety characteristics; the historical residue observation data is the residue concentration data obtained through laboratory detection or field rapid detection. The above multi-source data is aligned according to a unified timestamp. For static or discrete data without a timestamp, the corresponding time marker is generated by interpolation method to ensure consistency of different source data in the same time dimension.
[0030] In the aligned data, some records may have null values due to sensor failure or detection missing. The present application preferably uses three types of methods for missing completion: one is interpolation calculation using adjacent time slices, suitable for short-term missing of environmental factors; two is matching filling under similar environmental conditions, suitable for local missing of event data such as harvesting or pesticide application; three is to replace with statistical mean or historical typical value, suitable for occasional missing of observation data. After completing the missing completion process, all data are subjected to anomaly detection. When the data exceeds the reasonable threshold range or does not conform to the physical law, it is removed by rules or replaced by adjacent values to ensure the authenticity and reasonableness of the data input.
[0031] To further ensure that different source data can be processed by mechanism models and deep learning models, the present application unifies different units and dimensions of data after data cleaning. Specifically, the environmental data such as temperature, humidity, and rainfall are standardized to a unified dimension, and the pesticide dosage and detection concentration are normalized to relative values or standardized values to eliminate the interference of dimension difference on model training and prediction. After time alignment, missing completion, anomaly removal, and dimension unification, the data is called "consistent data". Among them, the consistent data refers to the input data stream formed after the multi-source data is processed by the above steps, which has consistency and comparability in time dimension and dimension dimension. This data stream can be directly used as the input of mechanism simulation and deep learning modeling to ensure the stability and interpretability of the prediction results.
[0032] The embodiment first establishes a compartment structure coupled with crop and environment, including at least leaf surface compartment, leaf mesophyll compartment and soil surface layer compartment, taking leaf surface residual mass, leaf mesophyll residual mass and soil surface residual mass as state variables. Absorption flux, leaching flux, volatilization flux and degradation flux are set between the compartments to describe the transfer and consumption of matter between the compartments and between the compartments and the environment. External driving and initial conditions come from the homogenization data, including hourly or daily air temperature, relative humidity, rainfall or rainfall intensity, irradiance intensity, pesticide dose and pesticide application time. At the pesticide application time, the input matching the pesticide dose is injected into the leaf surface compartment, and no input is injected at the non-pesticide application time. The simulation time step can be selected as one hour or one day, and the initial state can be set as non-zero for the leaf surface compartment and the rest of the compartments according to the experience or historical measured value at the last pesticide application time; if a new season or a new field is entered, the initial state of all compartments is set to zero or the background value is detected.
[0033] The embodiment adopts mass conservation and kinetic relationship to describe the net change rate of each compartment, that is, the increase of each compartment in a time step is equal to the sum of the relevant input and the transfer amount from other compartments minus the transfer amount to other compartments or the environment and irreversible loss. In order to reduce the empirical weight and enhance the transferability, the flux function adopts a setting with clear physical meaning: the absorption flux is affected by the opening and closing of stomata and the wax of leaf epidermis, the opening and closing of stomata is modulated by relative humidity, and the thicker the wax, the weaker the absorption; the leaching flux increases with the increase of rainfall kinetic energy, which can be reflected by the cumulative rainfall and its intensity; the volatilization and degradation processes accelerate with the increase of temperature, which is embodied by temperature-sensitive coefficients; the soil leaching process is enhanced with the increase of soil volume water content, which is embodied by the power index relationship of water content. For example, under the condition of common flue-cured tobacco varieties, the reference absorption intensity can be taken as zero point eight per hour, and when the equivalent thickness of leaf epidermis wax is two to five microns, the absorption modulation coefficient corresponds to about zero point six to zero point four; under the condition of medium to heavy rain, if the rainfall intensity is ten millimeters per hour for one hour, the leaf surface leaching ratio can reach about thirty percent; when the reference temperature is twenty-five degrees Celsius, the volatilization and degradation intensity can be increased by about twenty to forty percent with the increase of ten degrees Celsius; when the soil volume water content is zero point two five, the leaching intensity can be set as zero point zero two per hour. The above values are embodiment parameters, and actual application can select them within a given range according to varieties, soil types and local standards.
[0034] The embodiment adopts a non-negative preserving one-step update method for time-varying simulation: at each time step, the uncorrected state quantity of the next step is calculated according to the state quantity of the previous step and each flux, and if a negative value occurs, it is clamped to zero. Then, the mass conservation check is performed: according to the total mass at the last time, the amount of input at the current time and the irreversible loss, the total mass allowed at the current time is calculated, and the uncorrected state quantity is scaled to the allowed total mass in proportion to avoid total mass drift caused by numerical cumulative error. In order to convert the internal compartment state into the observable output corresponding to the detection link, the embodiment adopts a linear observable mapping: according to the sampling depth and the leaf position distribution, a set of non-negative weights whose sum is one is configured for each detection site, and the leaf surface and mesophyll state are linearly combined according to the weight to obtain the observable residual prediction of the site; for example, for the upper leaf position, the leaf surface weight can be set to 0.7 and the mesophyll weight to 0.3, and for the middle leaf position, the leaf surface weight can be set to 0.5 and the mesophyll weight to 0.5. Collecting the observable residual predictions of all detection sites in time sequence, the mechanism prior trajectory is obtained, the time resolution of which is consistent with the simulation step, which can be directly used for subsequent deep learning coupling and online assimilation.
[0035] Specifically, the specific parameters in the embodiment are explained as follows: (1) Compartment structure: refers to dividing the crop and environment system into a plurality of physically meaningful storage units, including leaf surface, mesophyll and soil surface layer, each unit corresponding to a time-varying residual amount; (2) Flux: refers to the material transfer or loss between compartments or between compartments and the environment per unit time, including absorption, leaching, volatilization and degradation; (3) Consistent data: refers to the multi-source input data stream formed after time alignment, missing data completion, abnormal data removal and dimension unification, which can be directly used for mechanism simulation and model training; (4) Linear observable mapping: refers to linearly combining the internal compartment state with a set of non-negative weights to obtain the observable residual prediction corresponding to the actual sampling position, the weights are derived from the sampling depth and leaf position distribution, and the sum of each weight is one; (5) Mechanism prior trajectory: refers to the residual prediction sequence of each detection position obtained by mechanism simulation under given external driving and initial conditions. Parameter sources and examples: pesticide application dose and application time come from field pesticide application records; air temperature, relative humidity, rainfall amount or intensity, and irradiance intensity come from plot weather stations or nearby weather service platforms; soil volume water content comes from soil moisture sensors or sampling and testing; leaf surface wax equivalent thickness can be determined by experiment or using variety statistical values, preferably two to five microns; the relative humidity corresponding to the half-opened stomata can be taken as fifty-five percent to sixty-five percent; the reference absorption intensity is preferably in the range of zero point zero three to zero point one two per hour, and the default is zero point zero eight per hour; the soil reference leaching intensity is preferably in the range of zero point zero one to zero point zero three per hour, and the water content index is preferably in the range of one to two; the reference temperature is twenty-five degrees Celsius, and the temperature sensitive relationship can be selected according to local standards or registered data; the rainfall kinetic energy sensitive coefficient can be set according to the statistical interval of moderate rain to heavy rain, so that ten to twenty millimeters of rainfall per hour corresponds to a leaching ratio of about twenty percent to fifty percent. The numerical interval given in this embodiment is a directly implementable value range, which can ensure stable simulation, mass conservation and output observability.
[0036] Further, the embodiment first calculates the absorption intensity at each simulation time step. The absorption process is modulated by the stomatal opening and the leaf cuticular wax. The calculation sequence is as follows: firstly, a stomatal modulation factor is obtained according to the relative humidity, and a saturation curve that monotonically increases and has an upper limit of one is selected, so that when the relative humidity is at the reference value of the half-opened stomata, the modulation factor is about 0.5, and when the relative humidity is low, the factor is close to zero, and when the relative humidity is high, the factor gradually approaches one. Secondly, a wax modulation factor is obtained according to the equivalent thickness of the leaf cuticular wax, and a curve that decreases with the increase of the wax thickness and has a lower limit of not less than 0.3 is selected, which is used to represent the inhibition of the wax barrier to absorption. Thirdly, the absorption intensity at the time step is obtained by multiplying the reference absorption intensity and the two modulation factors, and is used for the transfer of substances from the leaf surface to the leaf mesophyll. In terms of parameter setting, the reference absorption intensity can be 0.08 per hour; the reference relative humidity for half-opened stomata can be 60%; when the relative humidity increases from 50% to 70%, the stomatal modulation factor increases from about 0.4 to about 0.6 to 0.7; when the equivalent thickness of the wax increases from 2 microns to 5 microns, the wax modulation factor decreases from about 0.6 to about 0.4. In order to ensure numerical stability, the embodiment sets a non-negative lower limit of zero and a reasonable upper limit of 0.2 per hour for the absorption intensity, and truncates the values outside the range.
[0037] The embodiment calculates the leaf surface leaching intensity at each simulation time step to depict the stripping effect of rainfall on the attached residues. The calculation sequence is as follows: within a preset time window (for example, nearly six hours or nearly one day), the rainfall intensity of each time slice is nonlinearly transformed, multiplied by the length of the time slice, and summed to obtain an accumulated rainfall kinetic energy index; then the index is mapped to a saturation type stripping coefficient between zero and one, and the small value stage is approximately linearly responsive, and gradually tends to saturation after the index increases, to reflect the upper limit of the leaching effect under heavy rainfall. In terms of parameter setting, the sensitivity coefficient can be selected between 0.02 and 0.06, so that one hour of 10 mm of rainfall corresponds to a leaf surface leaching ratio of about 30%; in the case of strong rainfall of 20 mm per hour for two hours, the leaching ratio can reach 50% to 60%; and the leaching ratio corresponding to light rain (for example, 2 mm per hour for one hour) is generally less than 5%. In order to avoid excessive stripping leading to non-physical negative values, the embodiment uses non-negative preservation in numerical updating and uniformly constrains the total amount in mass conservation correction.
[0038] The Arrhenius temperature sensitivity law is used to calculate the volatilization and degradation intensity, and twenty-five degrees Celsius is selected as the reference temperature to give the multiple adjustment of intensity due to temperature change: when the temperature rises from twenty-five degrees Celsius to thirty-five degrees Celsius, the volatilization intensity is multiplied by about one to three to five, and the degradation intensity is multiplied by about one to two to four; when the temperature drops to fifteen degrees Celsius, the volatilization intensity and the degradation intensity are multiplied by about zero point seven to zero point eight and zero point seven five to zero point eight five, respectively. The reference volatilization intensity can be taken as zero point zero five per hour, and the reference degradation intensity of leaf surface and mesophyll can be taken as zero point zero three and zero point zero two per hour, respectively. The reference degradation intensity of the soil surface can be taken as zero point zero one to zero point zero two per hour. The soil leaching intensity is realized by a power index relationship driven by water content: when the soil volume water content is zero point two five, the leaching intensity is taken as zero point zero two per hour; when the water content rises to zero point three five, the leaching intensity is increased to zero point zero three to zero point zero four per hour; when the water content drops to zero point one five, the leaching intensity is reduced to zero point zero one to zero point zero fifteen per hour. In order to ensure cross-scene migration, reasonable upper and lower limits are set for all intensity quantities, and quality conservation check is performed after updating.
[0039] Specifically, the parameter interpretation and explanation of the present embodiment are as follows: One, stomata modulation factor: a normalized quantity driven by relative humidity to depict the influence of stomata opening and closing on leaf surface absorption process, zero means absorption almost closed, one means absorption reaches the upper limit; two, wax equivalent thickness: the equivalent thickness converted from the comprehensive resistance of leaf epidermis wax layer to diffusion, which can be obtained by microscopic measurement or variety statistical data; three, cumulative rainfall kinetic index: a scalar obtained by nonlinear transformation of rainfall intensity in each time slice, multiplied by the time length and summed up in a preset time window, used to represent the scouring ability of rainfall on the adherend; four, Arrhenius type temperature sensitivity law: an empirical law that takes the reference temperature as the benchmark, the volatility and degradation increase exponentially with the increase of temperature, and decrease exponentially with the decrease of temperature, which can be realized in engineering by the multiple relationship of reference temperature and temperature rise or fall; five, power index type percolation relationship: an empirical relationship that the soil leaching intensity monotonically increases with the volume water content according to the power index. Parameter sources and examples: relative humidity, air temperature and rainfall intensity come from the plot meteorological station or the nearest meteorological platform; the wax equivalent thickness can be measured by experiment or reference to the statistical value of the same variety, preferably two to five microns; the reference relative humidity of stomata half opening can be taken as fifty-five percent to sixty-five percent; the reference absorption intensity is preferably in the range of zero point zero three to zero point twelve per hour, and the default is zero point zero eight per hour; the sensitivity coefficient can be selected between zero point zero two and zero point zero six according to the rainfall characteristics of the region, so that the leaching ratio under the condition of medium to heavy rain is in the engineering acceptable interval of twenty percent to sixty percent; the reference volatility and degradation intensity is selected according to the registered data or public literature, and the upper limit is not more than zero point two per hour; the soil volume water content is obtained by soil moisture sensor or sampling test, and the power index can be selected between one and two to match the local soil texture. The above values and ranges are directly implementable engineering parameters, which meet the requirements of numerical stability and physical rationality.
[0040] Specifically, the step 300 of the embodiment first constructs an input channel of the deep learning model, and takes the homogenized data and the mechanism prior trajectory as joint input. The joint input is composed of two parts: one is the time series of environmental factors and management information, including air temperature, relative humidity, rainfall, irradiance, soil moisture content, pesticide dosage, pesticide application method, harvesting batch identification, and plant phenotype index; the other is the time series of the mechanism prior trajectory, i.e., the mechanism prediction residual curve obtained according to the leaf position or batch granularity. The two parts of data use the same time step and window length, for example, sliding sampling with one day as the step and thirty days as the window length, and then performing standardization and missing data completion on each channel before splicing the input. In order to enable the model to distinguish information from different sources, the embodiment appends a source identification bit to each input sequence, and aligns the time stamp uniformly and ensures the one-to-one correspondence between batches and leaf positions during the data loading stage. The main body of the model adopts a time series encoding structure, which is preferably composed of a plurality of layers of attention units and feedforward units stacked alternately, and the number of layers can be six, and the number of single-layer hidden units can be one hundred and twenty-eight, and residual connection and layer normalization are enabled in each layer to balance long sequence dependence and training stability.
[0041] The embodiment sets a feature extraction unit in the deep learning model, which includes two input branches and an alignment branch. The two input branches are respectively oriented to the homogenized data and the mechanism prior trajectory, and each extracts time series features and environment-dependent features; the alignment branch is used to map the features of the two branches to a unified hidden space, so that the representations of different sources have consistent scales and distributions. In order to reduce the complexity of the model and improve the cross-scene transferability, the embodiment adopts a parameter sharing manner for the first several layers of the two input branches, i.e., reusing the same set of convolution kernels, attention projection, and normalization statistics, so as to maintain consistent bottom-level representation ability among different sources; the number of shared layers can be two to three, and then one to two non-shared layers are added respectively to retain source specificity. The alignment branch realizes alignment through gated fusion and channel rescaling: first, the features of the two branches are spliced, then a lightweight gating unit is used to learn the source weight, and finally a channel rescaling unit is used to amplify high-correlation channels and suppress low-correlation channels, outputting joint representations in a unified hidden space for use by the downstream residual mapping module.
[0042] The embodiment couples the deep learning model and the mechanism model by residual mapping to obtain a target coupling model. The residual mapping is realized by a difference learning module: the module takes the aligned joint representation and the local segment of the mechanism prior trajectory at the same time step as input, and outputs a nonlinear deviation term that is not fully depicted by the mechanism model; the module structure adopts a lightweight configuration of two to three layers of feedforward network superimposed with gate units, to highlight the representation of nonlinear disturbance without overfitting the basic trend. In the inference stage, the deviation term output by the difference learning module is combined with the mechanism prior trajectory at the corresponding time step to obtain the final residual prediction sequence; in the training stage, the difference learning module, together with the aforementioned feature extraction unit and feature alignment branch, forms an end-to-end coupling structure, and is optimized with the joint loss to focus on the mechanism blind area while ensuring physical consistency. Thus, the parameter sharing mechanism provides a consistent foundation across sources, the residual mapping module focuses on learning nonlinear deviations, and the combination of the two constitutes the target coupling model, which can directly predict different leaf positions or different batches dynamically.
[0043] Specifically, the parameter interpretation and explanation of the embodiment are as follows: I. Joint input: refers to the double-channel or multi-channel input formed after aligning the consistent data and the mechanism prior trajectory at the same time step and window length, used to simultaneously carry environmental driving, management information and mechanism prediction information; II. Mechanism prior trajectory: refers to the residual prediction curve output by the mechanism model over time under given external driving and initial conditions, as the prior baseline for deep learning coupling; III. Hidden space: refers to the public representation space inside the deep learning model for carrying abstract features, with a fixed dimension, facilitating the fusion of features from different sources at the same scale; IV. Parameter sharing: refers to the use of the same set of trainable parameters by two or more input branches in several preposition layers, to reduce repeated learning and improve generalization and robustness; V. Residual mapping and difference learning module: refers to a structure module that takes the mechanism prior trajectory as the baseline, learns the deviation between the baseline and the true trend, takes the deviation as the output, and combines it with the baseline to obtain the final prediction. The above terms are functional designations used in the specification, aiming to define the module functions and data flow, but not to limit the specific implementation language and underlying library.
[0044] Step 400 of the embodiment is trained using a combined loss function comprising three parts: one is a fitting term for measuring the deviation of the model from the historical residual observations at each time point, preferably using a robust loss function with the ability to resist outliers; the second is a consistency constraint that does not increase when there is no new administration, which is used to punish the situation where the predicted residual is physically increased compared to the last time in the time window when no new administration event occurs, and a small tolerance is allowed to absorb detection noise and numerical errors; the third is a mass conservation constraint across compartments, which is used to punish the deviation between the total residual amount at any time and the allowed total amount, which is obtained by adding the input amount of the current period to the total residual amount of the last time and deducting the irreversible loss (volatilization and degradation) of the current period. To reduce the burden of manual parameter tuning, the three parts are scaled in this embodiment: first, calculate the initial magnitude of the three parts on a batch of calibration samples, then automatically scale the three weights to similar contribution, and then fix them as a single constant for the whole process training. Numerical example: the threshold constant of the robust loss function is one; the "non-increasing" tolerance is 0.5 micrograms per kilogram; the relative tolerance of mass conservation is one percent of the allowed total amount; the three weights are automatically normalized to fall between one and two, and the default is one.
[0045] The joint input is divided into training samples by uniform time step and window length in this embodiment, preferably with one day as the step and thirty days as the window for sliding sample building; each sample contains time series of environmental and management information, time series of mechanism prior trajectory, and historical residual observations aligned with the end of the window. Small batch optimization is used during training, and the batch size can be thirty-two; the optimizer uses an adaptive moment estimation method, with an initial learning rate of one thousandth, and a learning rate that decays by cosine or step according to the number of rounds; to avoid gradient explosion, threshold clipping is used with a threshold of one to two; to prevent overfitting, an early stopping strategy is enabled, and training is stopped when the validation set loss does not improve significantly for ten consecutive rounds; to obtain stable uncertainty estimates, the parameters except the output layer are frozen at the end of training, and the output layer is fine-tuned for two to three rounds to make the distribution of predicted residuals as close as possible to the assumptions of the robust loss function. Numerical example: total training rounds fifty to eighty; early stopping patience ten; learning rate halved at twenty rounds.
[0046] The embodiment synchronously calculates two types of physical consistency constraints in each step of training iteration. For the "no increase without new application" constraint, first generate a time index set according to the application record, compare the predicted residues of the adjacent two time points in the set, and if the increase of the latter exceeds the tolerance, the excess part is counted into the penalty; for the mass conservation constraint, first sum to get the model total residual amount of each compartment at this time, then get the allowed total amount according to the total residual amount at the last time, the current input amount and the irreversible loss estimate value, and the deviation of the two after dimensionless by the scale factor is counted into the penalty. In order to ensure numerical stability, the embodiment imposes non-negative constraints on the compartment residual amount at each time step, and performs conservation check after each update: if the total residual amount is slightly higher than the allowed total amount, scale the compartments proportionally to return to the allowed total amount range; if it is slightly lower than the allowed total amount and the difference is within the statistical fluctuation range of irreversible loss, do not adjust it. Numerical example: irreversible loss is composed of two parts of volatilization and degradation, which can be estimated by the mechanism flux given above. Under the condition of twenty-five degrees Celsius and moderate sunshine, the irreversible loss of the hour level generally does not exceed two percent of the total amount; when the conservation deviation exceeds one percent of the allowed total amount, the scaling correction is triggered.
[0047] Specifically, the parameter interpretation and explanation of the embodiment are as follows: I. Joint loss: refers to the training target obtained by weighted sum of fitting term, "no increase without new application" constraint and mass conservation constraint; II. Robust loss function: refers to the loss form whose penalty growth rate is lower than square for large deviation samples, used to resist the influence of abnormal points; III. Scale factor: refers to the value that maps errors of different dimensions or different orders of magnitude to a comparable range, the scale factor of the fitting part is preferably the value obtained by multiplying the absolute median of the observation error by 1.426, and the scale factor of the conservation part is preferably the robust scale of the allowed total amount during the training period; IV. Time index set of no new application: refers to the set of time steps in which no new application event occurs in the application record; V. Allowed total amount: refers to the total residual amount that is physically allowed to exist at a certain time, which is determined by the total residual amount at the last time, the current period input amount and the current period irreversible loss; VI. Weight automatic normalization: refers to the method of first evaluating the initial magnitude of each loss with calibration samples, and then automatically scaling the weights to make their contribution similar. Parameter source and example: historical residual observation comes from laboratory detection or field rapid detection; application record comes from land management system; irreversible loss and total residual amount are obtained from the compartment flux and state of the model; the threshold constant of robust loss is one zero; the weights of the two constraints are defaulted to one after automatic normalization, and can be fine-tuned in the range of zero point five to two to adapt to different varieties and land blocks if necessary. The above values and ranges are directly implementable engineering parameters, meeting the requirements of training stability and physical reasonableness.
[0048] Step 500 of the present embodiment, after a new residual observation arrives, is matched by the data access module with the model output according to the observation time, preferably in daily time steps, and the multi-leaf or multi-batch observations are respectively matched to the model output of the same granularity. To avoid the interference of outliers, quality control is performed first: including value range verification (such as limiting residual concentration to zero to ten thousand micrograms per kilogram), duplicate record deduplication, sample temperature and humidity consistency check, and instrument batch drift correction (preferably using a contemporaneous control sample to correct the coefficient). Then, according to the field sampling record, the observation is mapped to the observation channel of the model (such as the upper leaf, middle leaf, and lower leaf), and compared with the model output at the same time to obtain the observation error. To improve robustness, the present embodiment makes a hierarchical discrimination on the error: when the error is less than the empirical tolerance (such as 0.5 micrograms per kilogram), it is recorded as noise; when the error is between the tolerance and the diagnostic threshold (such as one to three times the tolerance), it enters the recursive update; when the error exceeds the diagnostic threshold and at the same time triggers the abnormal flag of weather or application record, the update is suspended and a review prompt is issued. The default configuration is: daily step, one to three day update window, observation delay allowed not more than two days.
[0049] The present embodiment sets the state transition relationship and the parameter identifiable list, and at each assimilation, a forward prediction is first performed to obtain the prior state at the current time, and then the observation error generated in the last segment is input into the recursive update unit for state correction, so that the updated state is consistent with the true observation. To suppress overfitting and parameter drift, only a small number of parameters with high sensitivity and good stability are adaptively adjusted, preferably limited to one to three, such as the absorption intensity reference of leaf to mesophyll, the rainfall stripping sensitivity coefficient, and the volatilization intensity reference; the remaining parameters remain frozen or are slowly updated with a small amplitude. The parameter boundary is given by the registration data and mechanism prior, for example, the absorption intensity reference is limited to 0.003 to 0.012 per hour, the rainfall stripping sensitivity coefficient is limited to 0.02 to 0.06, and the volatilization intensity reference is limited to 0.003 to 0.008 per hour. Recursive update adopts threshold gating and stepwise step: when the error is one to two times the tolerance and the diagnostic threshold, a small step correction is used; when it reaches two to three times, a medium step correction is used and the parameter slow-release strategy is triggered (the last adjustment direction and amplitude are attenuated for this update); when it exceeds three times, parameter update is suspended and only state correction is performed and recorded as a sample requiring manual verification. The default configuration is: updated once a day, forgetting factor 0.98, and single parameter change not more than 10% of the current value.
[0050] After the state and parameter update, the embodiment immediately performs consistency check, including non-negativity check, non-increase in the time window without new application, mass conservation check across compartments, and leaf position combination rationality check. If a negative value is found in an individual compartment, it is clamped to zero; if the predicted residue rises from the previous time and the amplitude exceeds the tolerance in the time window without new application record, the excess part is allocated back to the leaf surface or mesophyll reversible flux to eliminate the non-physical increase; if the total amount deviates from the allowed total amount by more than the allowed error (e.g. one percent of the allowed total amount), each compartment is scaled proportionally within the conservation range. After all the checks, the updated set is submitted: the corrected state and parameters are written to the model cache as the initial conditions for the next period prediction, and an audit log is generated to record the error, step, threshold trigger, and parameter change of this correction for traceability and risk control. The default deployment strategy is: batch assimilation at night is performed at 00:05 every day, and urgent assimilation triggered by on-site rapid detection data is completed within one hour after receiving the data; when the threshold is not triggered for three consecutive assimilations, the parameter learning rate is automatically reduced; when the highest threshold is triggered for two consecutive times, the parameters are automatically frozen and only state assimilation is performed until manual review is completed.
[0051] Specifically, the parameter interpretation and explanation of the embodiment are as follows: I. Recursive update unit: refers to the function module for gradually correcting the prior state of the model when new observations arrive, including error threshold discrimination, state correction, parameter fine-tuning, and boundary constraint; II. Identifiable parameters: refer to model parameters that have a significant impact on the prediction results under the current data density and noise level, and can converge to stable values through assimilation, such as absorption intensity benchmark, rainfall stripping sensitivity coefficient, and volatilization intensity benchmark; III. State transition relationship: refers to the evolution rule between adjacent time steps, which gives the prior state in combination with mechanism flux and external driving; IV. Consistency check: refers to the physical and numerical rationality check performed on the assimilated state and parameters, including non-negativity, non-increase in the time window without new application, and mass conservation; V. Threshold and tolerance: refers to two or more levels of numerical boundaries used for error classification management, used to distinguish between noise, assimilable samples, and samples that need to be reviewed. Parameter sources and examples: observation values come from laboratory tests or field rapid detection; application records and harvest records come from land management systems; environmental driving comes from land weather stations; default tolerance is 0.5 micrograms per kilogram, diagnostic threshold is three times the tolerance, forgetting factor is 0.98, parameter change upper limit is 10%, and conservation allowed error is 1%. The above value range is an engineering implementable configuration, which can be adjusted within the range described in the application file according to varieties, land, and instrument performance without changing the essence of the invention.
[0052] Step 600 of the embodiment collects the residual prediction results output by the target coupling model at each prediction time according to the "batch mapping table" or the "leaf position mapping table": when taking the harvesting batch as the granularity, the predicted values of different leaf positions or the same block at the same time are aggregated according to the batch number, and the aggregation is preferably performed in a quality weighting or area weighting manner (the quality weight comes from the actual harvesting weight or the planned harvesting weight, and the area weight comes from the plot operation area); when taking the leaf position as the granularity, the predicted values of the corresponding channels are directly read according to the upper leaf, the middle leaf, the lower leaf or the finer leaf order number. After aggregation, an equidistant residual time sequence is formed, and the time step is preferably one day, or can be set to one hour or three days according to the detection frequency. In order to improve the robustness of the sequence, the embodiment performs missing data filling and outlier processing (for example, replacing isolated spikes with the median of a sliding window) on the aggregation results, and then writes them into the "prediction cache area". The cache area records the time stamp, batch or leaf position identifier, aggregation method, weight source and version number, so as to trace back.
[0053] The embodiment generates upper and lower limit intervals for each prediction time to represent uncertainty. The implementation path includes three alternative or combinable ways: one is a distribution parameter regression method, which simultaneously learns the distribution parameters such as prediction value and variance in the training stage, and directly outputs the mean and interval in the inference stage; the second is a quantile prediction method, which targets different quantiles in the training stage, and returns the interval of the corresponding upper and lower quantiles in the inference stage; the third is an empirical calibration method, which constructs an interval with controllable coverage according to the historical residual distribution on the validation set, and updates it seasonally. In order to ensure engineering usability, the embodiment defaults to output an interval with a coverage rate of ninety percent, while retaining a ninety-five percent interval for risk review; the coverage rate is calibrated through a hold-out set or a near-month rolling sample, so that the actual coverage rate deviation does not exceed two percentage points. Numerical example: in the spring sample, the center value of daily prediction of a batch is 800 micrograms per kilogram, and the corresponding ninety percent interval is 650 micrograms per kilogram to 950 micrograms per kilogram; when entering the rainy season, the interval width automatically expands by about ten percent to twenty percent with the increase of rainfall uncertainty.
[0054] The embodiment compares the residual time sequence and its uncertainty interval with the maximum residual limit of the target area at each time point, and obtains an over-limit judgment mark. The over-limit probability is obtained by counting the proportion of the prediction distribution higher than the maximum residual limit at each time point, and the compliance confidence is defined as one minus the over-limit probability; when using quantile prediction or empirical calibration, the proportion is directly estimated according to the positional relationship between the interval and the threshold. To avoid misjudgment caused by single-point fluctuation, the embodiment defines the start of the "compliance time window" as the first time point at which the prediction central value and the upper boundary of the uncertainty interval are both lower than the maximum residual limit, and continue to meet at least three consecutive time steps (preferably three days, and can also be set to two to five days according to local standards). The system generates a "compliance suggestion" when the condition is met, including the recommended harvesting date range, the current over-limit probability, the compliance confidence, the key driving factor summary (such as recent rainfall and temperature), and pushes it to the planting management or quality control system; when the over-limit probability exceeds the warning threshold (default is thirty percent), a "risk warning" is generated, accompanied by the next management suggestion (such as postponing harvesting, strengthening ventilation and drying, arranging re-inspection).
[0055] Specifically, the parameter explanations and descriptions of the embodiment are as follows: I. Batch mapping table: refers to a table used to map plots, harvesting dates and leaf position information into harvesting batch numbers, which is derived from a production management system; II. Prediction cache area: refers to a storage structure used to temporarily store aggregated prediction results and their metadata, which is used for auditing and tracing; III. Uncertainty interval: refers to an upper and lower limit pair used to describe the prediction uncertainty range under a given coverage rate, with coverage rate values commonly being ninety percent or ninety-five percent; IV. Over-limit probability: refers to the probability that the prediction distribution falls above the maximum residual limit at a certain time; V. Compliance confidence: refers to the probability that the prediction distribution falls below the maximum residual limit at a certain time; VI. Compliance time window: refers to a continuous time period starting from the first time point at which the prediction value and the upper boundary of the interval are both lower than the maximum residual limit and continue to meet; VII. Maximum residual limit: refers to the maximum allowable residual value of a specific active ingredient in tobacco according to national or regional regulations, with units usually being micrograms per kilogram or milligrams per kilogram. Parameter source examples: the maximum residual limit comes from a regulatory standard library; the batch and leaf position weights come from harvesting records and plot areas; the coverage rate is by default ninety percent and can be configured between eighty percent and ninety-five percent; the compliance judgment continuous step number is by default three days and can be configured between two and five days; the warning threshold is by default thirty percent and can be adjusted between twenty percent and fifty percent. The above value ranges are engineering implementation configurations, which allow adjustment according to plot characteristics and regulatory requirements without changing the essence of the invention.
[0056] Corresponding to the above method, as shown in Figure 2 the embodiment also provides a tobacco pesticide residue dynamic prediction method combining deep learning and mechanism models, comprising: a data acquisition and unification unit configured to acquire pesticide application event data, environmental driving data, agronomic management data, plant phenotype data, and historical residue observation data associated with a target plot, and to perform time alignment, missing value completion, outlier removal, and dimension unification on the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotype data, and the historical residue observation data to form unified data; a mechanism modeling and simulation unit configured to establish a mechanism model coupled with crops and environment based on mass conservation and kinetic relations, and to perform time-varying simulation on the unified data to generate mechanism prior trajectories; a deep learning coupled modeling unit configured to construct a deep learning model, to take the unified data and the mechanism prior trajectories as joint inputs, and to couple the deep learning model and the mechanism model through parameter sharing or residual mapping to represent nonlinear influences not sufficiently depicted by the mechanism model, so as to obtain a target coupled model; a joint loss training unit configured to train the target coupled model based on a joint loss; the joint loss at least includes a fitting term for historical residue observations and a physical consistency constraint term; the physical consistency constraint includes that the predicted residue does not increase within a time window without new pesticide application and that mass conservation is conserved across compartments; an online assimilation and adaptive calibration unit configured to, when a new residue observation arrives, perform online assimilation and adaptive calibration on the state and identifiable parameters of the trained target coupled model through a recursive update method; a prediction output and risk determination unit configured to output residue time series prediction and uncertainty interval at a harvest batch or leaf position granularity, and to calculate over-limit probability, compliance confidence, and compliance time window based on maximum residue limit, so as to provide decision-making for planting management or quality control systems.
[0057] The embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other.
[0058] The principles and implementation manners of the present application are described by using specific examples in the specification. The above description of the embodiments is only for the purpose of helping to understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will be changed. In conclusion, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A dynamic prediction method for tobacco pesticide residues that integrates deep learning and mechanism models, characterized by: include: Acquire pesticide application event data, environmental driving data, agronomic management data, plant phenotypic data, and historical residue observation data associated with the target plot, and perform time alignment, missing complementation, anomaly elimination, and dimension unification on the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotypic data, and the historical residue observation data to form consistent data; Establishing a mechanism model for crop-environment coupling based on mass conservation and kinetic relationships, and using the uniformed data to perform time-varying simulations to generate mechanism prior trajectories; Constructing a deep learning model, taking the uniform data and the mechanism prior trajectory as joint input, coupling the deep learning model with the mechanism model through parameter sharing or residual mapping to characterize nonlinear effects not fully described by the mechanism model, and obtaining a target coupling model; Training the target coupling model based on a joint loss; the joint loss includes at least a fitting term for historical residue observations and a physical consistency constraint term, the physical consistency constraint including predicting no increase in residue within a time window without new application and conservation of mass across compartments; When new residual observations arrive, a recursive update method is used to perform online assimilation and adaptive calibration of the state and identifiable parameters of the trained target coupling model; Output residue time series predictions and uncertainty intervals by harvest batch or leaf size, and calculate the probability of exceeding the limit, compliance confidence level and compliance time window based on the maximum residue limit for decision-making in planting management or quality control systems.
2. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism model according to claim 1, characterized in that: Acquire the pesticide application event data, environmental driving data, agronomic management data, plant phenotypic data, and historical residue observation data associated with the target plot, and perform time alignment, missing complementation, anomaly elimination, and dimension unification on the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotypic data, and the historical residue observation data to form consistent data, including: Aligning the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotypic data, and the historical residue observation data according to a unified timestamp, and generating an interpolated time stamp for data that does not have a timestamp; Missing data items are filled by interpolation based on adjacent time slices, matching of similar environmental conditions, or statistical mean filling; Perform anomaly detection on data that exceeds reasonable thresholds or does not conform to physical laws, and correct them through rule elimination or substitution filling; The data from different sources and different unit systems are normalized according to a unified dimension to obtain the consistent data with a consistent structure.
3. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism model according to claim 1, characterized in that: A mechanism model for crop-environment coupling is established based on mass conservation and dynamics. Time-varying simulations are performed using the standardized data to generate a priori mechanism trajectories, including: A compartment structure consisting of leaf compartment, mesophyll compartment and soil surface compartment was established, with the leaf residue mass, mesophyll residue mass and soil surface residue mass as state variables, respectively. Absorption flux, leaching flux, volatilization flux and degradation flux were set between the compartments. Using the temperature, rainfall, relative humidity, irradiance, pesticide dosage and application time in the consistent data as external driving and initial conditions, the leaf compartment is given a pesticide input at the time of application and a zero input at the time of non-application. The net rate of change of each compartment is described by mass conservation and kinetic relationship, forming the following set of state equations: ; ; ; in, is the residual mass on the leaf surface; is the mass of mesophyll residue; is the residual mass of the soil surface; Input for pesticide application; is the absorption coefficient from leaf surface to mesophyll; is the stomatal modulation factor that varies with relative humidity; is the elution coefficient determined by rainfall intensity or rainfall amount; is the volatility coefficient; is the transport coefficient from mesophyll to other parts; 、 、 are the degradation coefficients of leaf surface, mesophyll and soil surface, respectively; is the leaching coefficient affected by moisture content; is temperature; is the relative humidity; is the rainfall intensity or accumulated rainfall; is the soil volume moisture content; In order to reduce the empirical weight and improve cross-scenario transferability, a physically meaningful function form is used to determine the absorption, leaching, volatilization and degradation processes; A one-step method with non-negativity preservation is used for time-varying simulation, and mass conservation is checked for irreversible losses. The update formula is: ; ; in, is the residual mass of the i-th compartment at time t; Δt is the time step; is the net flux function of the i-th compartment; is the external driving vector; is the uncorrected update amount; Sum the masses of each compartment; The allowable total amount is obtained by adding the mass at the previous moment to the pesticide input minus the irreversible loss; is the total amount of irreversible loss caused by volatilization and degradation; ε is a very small positive number used for numerical stability; It is the external input to the foliar compartment at the time of application, and its value is determined by the dosage, application method and application time. The compartmental states are converted into detectable residue predictions using a linear observable mapping and the mechanistic prior trajectory is output: ; in, is the observable residual vector at time t; C is the observation mapping matrix determined by the corresponding relationship between sampling depth and leaf position; is the compartment state vector; the mechanism prior trajectory is the combination of the observable residual vectors at each moment.
4. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism model according to claim 3, characterized in that: To reduce the empirical weight and improve cross-scenario transferability, a physically meaningful function form is used to determine the absorption, leaching, volatilization and degradation processes, including: The absorption coefficient modulated by stomata and stratum corneum is expressed as follows: ; in, is the reference absorption coefficient; is the wax thickness sensitivity coefficient; is the equivalent thickness of leaf epidermal wax; is the sensitivity coefficient of stomatal opening and closing to relative humidity; The reference relative humidity is when the stomata are half open; Considering the peeling effect of rainfall kinetic energy on the attachment, the elution coefficient is defined as: ; in, is the sensitivity coefficient to rainfall kinetic energy; is a measure of the cumulative rainfall kinetic energy within the time window; For the Rainfall intensity in a time slice is the length of the corresponding time slice; The volatilization and degradation processes use temperature-driven Arrhenius-type coefficients: ; in, is the base volatility coefficient; is the apparent activation energy of volatilization; R is the gas constant; is the reference temperature; is the baseline degradation coefficient of compartment x; For compartments x Apparent activation energy of degradation; x Take the leaf surface or mesophyll or soil surface; The leaching coefficient driven by soil moisture content adopts a power exponential seepage relationship: ; in, is the baseline leaching coefficient; m It is the response index of soil moisture to leaching process.
5. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism model according to claim 1, characterized in that: Construct a deep learning model, take the consistent data and the mechanism prior trajectory as joint input, couple the deep learning model with the mechanism model through parameter sharing or residual mapping to characterize the nonlinear effects not fully described by the mechanism model, and obtain the target coupling model, including: Constructing a deep learning model that takes the uniformized data and the mechanism prior trajectory as joint input, wherein the input form includes a time series, an environmental factor series, and a residual trajectory predicted by the mechanism; A feature extraction unit is provided in the deep learning model to extract temporal features and environment-dependent features from the uniform data and the mechanism prior trajectory, respectively, and perform feature alignment in a unified latent space; By sharing parameters, some parameters are reused in the feature extraction unit to maintain consistent representation capabilities across different input sources. By means of residual mapping, a difference learning module is introduced between the mechanism prior trajectory and the output of the deep learning model, so that the deep learning model can focus on learning the nonlinear deviations that are not fully characterized by the mechanism model; The parameter sharing mechanism is combined with the residual mapping module to form the target coupling model.
6. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism model according to claim 1, characterized in that: The calculation formula of the joint loss is: ; ; in, For the target coupling model at time Output observable residual predictions; It is a historical residual observation; It is a time index collection with no new pesticide application; For the moment the total residual mass of the model summed across compartments; For the moment The total mass allowed by the law of conservation of mass is obtained by adding the total mass at the previous moment to the current amount of injection and deducting the irreversible loss; is the scale factor obtained by robust estimation based on the observation sequence; is the scale factor of the total mass dimension; is the Huber function of robust fitting; is the constant of Huber threshold; is the weight coefficient of the constraint of “no increase when no new pesticides are applied”; is the weight coefficient of the mass conservation constraint.
7. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism model according to claim 6, characterized in that: Each weight coefficient is automatically normalized to a single constant according to the data scale.
8. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism model according to claim 1, characterized in that: When new residual observations arrive, a recursive update method is used to perform online assimilation and adaptive calibration of the state and identifiable parameters of the trained target coupling model, including: receiving new residual observation data, and comparing the observation data with the output result of the target coupling model at the same time to obtain an observation error; The observation error is input into a recursive updating unit, and the internal state of the target coupling model is corrected according to a preset state transition relationship so that the updated state fits the real observation; Based on the correction state, identifying identifiable parameters that significantly affect the prediction of the target coupling model, and adjusting the values of the identifiable parameters according to the magnitude and direction of the error; Perform consistency checks on the updated state and identifiable parameters to ensure that physical constraints are met and that no unphysical increase in residual volume occurs; The corrected states and parameters are used as new initial conditions for dynamic prediction in the next period to achieve adaptive calibration of the model under different environmental and management conditions.
9. The method for dynamic prediction of tobacco pesticide residues by integrating deep learning and mechanism modeling according to claim 1, characterized in that: Output residue time series predictions and uncertainty intervals by harvest batch or leaf size, and calculate the probability of exceeding the limit, confidence level of compliance, and time window for compliance based on the maximum residue limit to provide support for plantation management or quality control system decision-making, including: Aggregating the residue prediction results output by the target coupling model at each prediction moment according to harvest batch or leaf position granularity to form a residue time series; Based on distribution parameter regression, quantile forecasting or uncertainty estimation methods, generating upper and lower bounds for each moment of the residual time series to characterize the uncertainty of the forecast; Comparing the residue time series and uncertainty intervals with the maximum residue limits specified in the target area to obtain the limit-exceeding judgment at each time; By counting the proportion of predicted residue values that exceed the maximum residue limit, the probability of exceeding the limit is calculated, and the confidence level of compliance under the regulatory threshold is inferred from this; Based on the downward trend of the predicted residue value over time, the time point when the predicted residue value is first and continuously lower than the maximum residue limit is determined, which is defined as the start of the compliance time window and output to the planting management or quality control system for decision support for harvesting, processing or warehousing.
10. A dynamic prediction method for tobacco pesticide residues that integrates deep learning and mechanism models, characterized in that: include: A data acquisition and consistency unit is used to acquire pesticide application event data, environmental driving data, agronomic management data, plant phenotypic data, and historical residue observation data associated with the target plot, and to perform time alignment, missing complementation, anomaly elimination, and dimensional unification on the pesticide application event data, the environmental driving data, the agronomic management data, the plant phenotypic data, and the historical residue observation data to form consistent data; a mechanism modeling and simulation unit for establishing a mechanism model of crop-environment coupling based on mass conservation and kinetic relationships, and performing time-varying simulation using the uniformed data to generate a mechanism prior trajectory; A deep learning coupling modeling unit is used to construct a deep learning model, take the consistent data and the mechanism prior trajectory as joint input, couple the deep learning model with the mechanism model through parameter sharing or residual mapping, so as to characterize the nonlinear effects not fully characterized by the mechanism model, and obtain a target coupling model; a joint loss training unit for training the target coupling model based on a joint loss; the joint loss includes at least a fitting term for historical residue observations and a physical consistency constraint term, the physical consistency constraint including predicting no increase in residue within a time window without new application of pesticides and conservation of mass across compartments; An online assimilation and adaptive calibration unit, configured to perform online assimilation and adaptive calibration of the state and identifiable parameters of the trained target coupling model using a recursive update method when new residual observations arrive; The prediction output and risk assessment unit is used to output residue time series predictions and uncertainty intervals by harvest batch or leaf position particle size, and calculate the probability of exceeding the limit, compliance confidence level and compliance time window based on the maximum residue limit for decision-making in planting management or quality control systems.