Intelligent crop yield forecasting method based on deep learning and data assimilation

By combining deep learning and data assimilation methods with physical mechanisms and data-driven models, the accuracy and robustness of crop yield forecasts are improved. This solves the problem of independent use of physical and data models in existing technologies, and provides autonomous optimization and scientific decision support.

CN121809776APending Publication Date: 2026-04-07张静
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-12
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

In existing crop yield forecasting technologies, physical mechanism models are unable to characterize the complex nonlinear environment-crop response relationship, data-driven models lack physical constraints, resulting in unreasonable forecast results and uncertainty. The uncertainty of these forecasts lacks systematic management, and the models lack self-optimization capabilities.

Method used

We construct an intelligent crop yield forecasting method based on deep learning and data assimilation. Through hybrid growth engine, uncertainty guidance, multi-source information fusion and feedback optimization steps, we achieve deep integration of physical mechanisms and data intelligence, autonomously accumulate knowledge and continuously optimize.

Benefits of technology

It improves the accuracy, robustness, and practicality of yield forecasts, ensures that forecast results are within agronomically reasonable ranges, reduces risks under unforeseen scenarios, and provides scientific decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121809776A_ABST
    Figure CN121809776A_ABST
Patent Text Reader

Abstract

The invention discloses a crop yield intelligent forecasting method based on deep learning and data assimilation, and belongs to the technical field of intelligent agriculture and agricultural information. According to the method, a mixed growth engine containing differentiable rule constraints is constructed as a digital twin body core, and an analog state is calibrated through dynamic data assimilation; furthermore, the uncertainty is actively evaluated and predicted, and a directional experiment is performed in a virtual environment to conclude causal knowledge; then, real-time data, historical knowledge and expert experience are fused through a multi-source evidence reasoning framework, and a probabilistic yield prediction and agronomic plan is generated; finally, rule base self-correction, model parameter self-updating and self-optimization evolution are achieved through growth season feedback data. According to the method, the key problems that a physical model and a data model are separated, and the system passively predicts and cannot be evolved are solved, and the prediction accuracy, reliability and self-adaptive capability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart agriculture and agricultural information technology, specifically to a method for intelligent crop yield forecasting based on deep learning and data assimilation. Background Technology

[0002] Accurate crop yield forecasts are crucial for national food security early warning, agricultural policy formulation, and farmer production management. Existing yield forecasting technologies primarily rely on the following two types of methods: The first category consists of physical models based on crop growth mechanisms, such as WOFOST and DSSAT. These models are based on the principles of energy and matter balance, simulating key processes such as photosynthesis, respiration, and bioassay in crops through a series of biophysical equations. Their advantage lies in the transparency of the processes and the clarity of the mechanisms. However, these models typically contain a large number of parameters that require local calibration, making them extremely sensitive to the quality and completeness of input data (such as soil parameters and management practices). Furthermore, mechanistic models have limited ability to characterize the complex nonlinear relationships of crop responses under extreme climate events (such as combined high temperature and drought), and their high computational complexity makes large-scale, real-time operational implementation difficult.

[0003] The second category is data-driven models based on statistics and machine learning. With the development of remote sensing big data and artificial intelligence technologies, these methods predict yields by mining statistical relationships between historical yields and multi-source environmental factors (meteorology, remote sensing spectroscopy, soil, etc.). Deep learning models, such as convolutional neural networks (CNNs) and recurrent neural networks (RNNs), have demonstrated powerful nonlinear fitting capabilities in this field. However, purely data-driven "black box" models lack physical interpretability, and their predictions sometimes violate basic agricultural common sense and physical laws (e.g., predicting increased biomass under no-light conditions). Furthermore, these models heavily rely on the coverage of the training data; when encountering novel climate-management combinations not present in the training data, the reliability and stability of their extrapolation predictions decrease significantly.

[0004] To combine the advantages of the two methods mentioned above, existing technologies have made some attempts, mainly including: using machine learning methods to invert or calibrate key parameters of the mechanistic model; and using the simulation results of the mechanistic model as features to input into the statistical prediction model. However, these methods are essentially "post-processing" or "serialized," with physical constraints only acting on the model's input or output, and not deeply embedded in the core learning and reasoning process of the data-driven model, thus failing to solve the problem of the physical irrationality of the prediction results.

[0005] Therefore, a technological solution is needed that can achieve a deep integration of physical mechanisms and data intelligence, and endow the forecasting system with the ability to proactively recognize uncertainties, autonomously accumulate knowledge, and continuously optimize itself, thereby improving the accuracy, robustness, and practicality of yield forecasts. Summary of the Invention

[0006] To overcome the aforementioned shortcomings of existing technologies, embodiments of the present invention provide a smart crop yield forecasting method based on deep learning and data assimilation. Specifically, this addresses the problems mentioned in the background art: physical mechanism-based growth models struggle to accurately characterize complex nonlinear environment-crop response relationships; while purely data-driven deep learning models may generate predictions that violate fundamental agricultural physics. These two methods are typically used independently or simply in series, resulting in a failure to integrate physical constraints with the data learning process. Furthermore, existing solutions lack a systematic quantification and proactive mitigation mechanism for model prediction uncertainties, and the models lack a mechanism for autonomously updating parameters using new data, hindering continuous performance evolution.

[0007] To address the aforementioned problems, this invention provides an intelligent crop yield forecasting method based on deep learning and data assimilation. Its core lies in constructing a system that sequentially executes the steps of 'growth simulation and assimilation', 'knowledge discovery guided by uncertainty', 'multi-source information fusion reasoning', and 'feedback-based model optimization', specifically including the following steps: S1. Constructing and Dynamically Calibrating a Growth Simulator. A hybrid growth engine integrating neural network units and rule-based constraint units is established as the core of the crop growth digital twin. This engine receives time-series environmental data and iteratively generates a growth state sequence. By introducing real-time observation data, the engine parameters are dynamically assimilated to continuously approximate the simulated state to real-world observations.

[0008] S2. Uncertainty-Driven Knowledge Discovery. Evaluate the predictive consistency of the hybrid growth engine under different environmental scenarios, quantifying and locating spatiotemporal regions with high predictive uncertainty. Within these regions, automatically design and execute multiple sets of virtual growth experiments. By analyzing "environmental intervention-yield response" data, extract qualitative or semi-quantitative correlation rules between environmental factors and yield, forming an interpretable knowledge base.

[0009] S3. Multi-source information fusion reasoning and decision-making. An evidence-based reasoning framework is constructed, simultaneously inputting real-time assimilation status, historical association rule knowledge, and external expert experience. This framework performs credibility assessment, conflict resolution, and dynamic weighted fusion of multi-source evidence, ultimately outputting yield forecasts expressed as probability intervals and corresponding agronomic management plans, while quantifying their uncertainty.

[0010] S4. Model Self-Iteration and Optimization. After the growth cycle ends, attribution analysis is performed on the system based on the difference between the final prediction and the measured results. The focus is on auditing the triggering effectiveness of the rule constraint units and correcting erroneous triggering rules; and using new data to enhance the training of neural network units, thereby enabling the system to learn autonomously and iterate its performance from practical applications.

[0011] S5. Pre-sowing strategy pre-assessment. The core processes of S1 to S3 above are applied to the pre-sowing planning stage. By simulating the long-term performance of different varieties and cultivation schemes under various possible climatic scenarios, a comprehensive evaluation and ranking of multiple objectives (such as average yield, yield stability, and risk resistance) is conducted to provide data-driven optimal solutions for sowing decisions.

[0012] Furthermore, in S1, the hybrid growth engine operates through a differentiable rule embedding mechanism: the neural network unit outputs a preliminary prediction, and the rule constraint unit calculates in parallel its conformity with preset agricultural rules; the system dynamically adjusts the weights between the preliminary prediction and the rule-guided correction term based on the conformity, thereby realizing the soft constraint of physical laws on data prediction.

[0013] Furthermore, in S2, the uncertainty quantification is achieved by running a set of growth simulators, and cognitive weaknesses are identified by statistically analyzing the prediction variance of key reproductive period state variables; the virtual experiment is conducted using a space-filling design sampling within a parameter space consisting of high-uncertainty reproductive periods and their key environmental factors.

[0014] Furthermore, in S3, the operation of the evidence reasoning framework includes: activating relevant rule knowledge based on matching degree; calculating the comprehensive weight of each piece of evidence by combining the historical confidence of the rules, the arbitration result of the current observation on conflicting evidence, and the external information bias; and finally integrating the weighted outputs of each piece of evidence to form a comprehensive prediction.

[0015] Furthermore, in S4, the rule correction is achieved by analyzing the historical trigger log of the rule: comparing the state evolution direction of the simulation and observation after the rule is triggered, and statistically analyzing the "suspicious triggering" rate of the rule; for rules with a high suspicious triggering rate, analyzing the common environmental characteristics of their false triggering scenarios, and then refining or adjusting their logical premises.

[0016] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects: 1. By embedding differentiable rules and using dynamic fusion mechanisms in the hybrid growth engine, a deep integration of data-driven models and domain physical knowledge is achieved, ensuring that the prediction results are always within the range of agronomical rationality and enhancing the reliability of the model in unseen scenarios.

[0017] 2. Through virtual experiments that quantify and guide uncertainty, the system can proactively detect and compensate for its own cognitive blind spots, automatically generate and accumulate "condition-outcome" agronomical knowledge to guide the system's own reasoning, and reduce the prediction risk at key decision points.

[0018] 3. Through the multi-source evidence fusion reasoning framework, it effectively integrates real-time data, historical patterns and expert judgment, and can handle the incompleteness and conflict of information, outputting predictions and suggestions with probability assessment and causal explanation, thereby improving the scientificity and credibility of decision-making.

[0019] 4. A closed-loop optimization mechanism based on actual production feedback was established, enabling the model to continuously diagnose its own defects, correct the rule base and update the neural network during use, solving the problem of traditional model solidification and realizing the system's self-iteration and long-term improvement.

[0020] 5. By extending forecasting capabilities to the sowing planning stage, and through multi-scenario simulation and risk assessment, forward-looking optimization basis is provided for variety selection and cultivation strategy formulation, realizing the extension from mid-production forecasting to pre-production planning and improving practicality. Attached Figure Description

[0021] Figure 1 This is a flowchart of the intelligent crop yield forecasting system of the present invention.

[0022] Figure 2 This is a flowchart illustrating the collaborative operation of the hybrid growth engine of the present invention.

[0023] Figure 3 This is a flowchart of the virtual experiment exploration and knowledge discovery branch of the present invention.

[0024] Figure 4 This is a flowchart of the multi-source evidence reasoning and decision generation branch process of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] Example 1 As attached Figures 1 to 4 The intelligent crop yield forecasting method based on deep learning and data assimilation, as shown below, has the following specific implementation details: Step 1: Construct a digital twin of crop growth and perform dynamic data assimilation The goal of this step is to construct a digital twin of crop growth. The core dynamic simulation capability of the digital twin is provided by a hybrid growth engine, which includes a neural network unit and a rule constraint unit.

[0027] The specific implementation process is as follows: First, data preparation and simulator initialization are performed. Historical time-series data of the target area are acquired. This historical time-series data includes environmental driving data, soil data, crop parameters, and corresponding crop growth status observation data.

[0028] The environmental driving data is daily meteorological data, specifically including daily maximum temperature. Daily minimum temperature Daily precipitation Total solar radiation .

[0029] The soil data includes the soil volumetric water content at the time of sowing. .

[0030] The crop parameters include the photoperiod sensitivity parameters of the crop variety and the accumulated temperature requirements during key growth periods.

[0031] The growth status observation data includes periodically acquired measured values ​​of leaf area index (LAI) and aboveground biomass (AGB).

[0032] Initialize the hybrid growth engine. The neural network unit is constructed as a feedforward neural network, whose input layer receives the crop state variables (including) from the end of the previous day. and ) and the environmental driving variables of the day (including , , and The output layer outputs predictions of the daily state variable increments, including... and .

[0033] The rule constraint unit contains pre-set agricultural knowledge rules expressed in the form of logical conditions.

[0034] For example, one rule is defined as: if the daily average temperature Below 0°C, the daily biomass increase The predicted value is not greater than 0, where .

[0035] Another rule is defined as: leaf area index The predicted value is in the range Inside, among which This is the theoretical value of the maximum leaf area index set according to the variety.

[0036] Secondly, the hybrid growth engine is pre-trained based on historical data. The historical environment-driven sequence is input into the initialized hybrid growth engine to obtain the predicted growth state sequence. The weight parameters of the neural network units are then optimized by minimizing an objective function.

[0037] The objective function The definition is as follows: .in, and They represent Simulated state vector and observed state vector at time t. This represents the total number of time steps in the historical data. This represents the squared error of a vector. For indicator functions, if in If any rule is triggered during the time-lapse simulation, then ,otherwise . This is the regularization coefficient, used to balance prediction accuracy and rule compliance; its value can be determined based on cross-validation.

[0038] Finally, dynamic data assimilation is performed during the growing season. Starting from the sowing day, the following operations are performed daily: the simulated state from the previous day and the environmentally driven data of the current day are input into a pre-trained hybrid growth engine to obtain the simulated state for that day. This is done when the state observation data for the day is obtained. At that time, calculate the simulated state. With observation status The error between them.

[0039] Based on this error, a gradient descent algorithm is used to update the current parameters of the hybrid growth engine, making subsequent simulation predictions closer to actual observations. This assimilation process is repeated throughout the growing season. This step outputs a real-time calibrated digital twin of crop growth, providing core simulation capabilities for subsequent steps.

[0040] As a preferred embodiment, the specific collaborative operation process of the hybrid growth engine in daily simulation prediction is as follows: The neural network unit calculates based on the input and outputs a vector of preliminary state predictions without rule constraints. This includes the , Prediction of variables such as inequalities.

[0041] Rule-constrained units for vectors Conduct a compliance check. Calculate the degree of violation for the stated daily average temperature rule. .

[0042] For the LAI range rule, calculate the degree of violation. .

[0043] Calculate the total violation level at the current time step. .

[0044] according to Calculate dynamic fusion weights The calculation formula is: .in The preset sensitivity coefficient is used to adjust the impact of rule violation on the fusion weights. The intensity of the impact. The higher the value, the more serious the consequences of even minor rule violations. Significantly reduced. Its value can be adjusted according to the strictness of the model's adherence to physical rules, typically ranging from 0.5 to 5.0. For example, in one implementation, it is taken as... .

[0045] Generate a correction vector based on the violated rule. The purpose of the correction vector is to pull the predicted values ​​that violate the rules back into the compliant range. For example, if Then a will be generated Value direction Correction amount for direction adjustment; if and Then a reduced The correction amount.

[0046] Calculate the final output simulation state: .in This is a preset correction step size coefficient used to control the magnitude of a single correction and avoid overcorrection. A typical value range is 0.01 to 0.5. For example, in one implementation, it is... .

[0047] when hour, The final output equals the initial prediction. ; when hour, The output is a weighted combination of the preliminary prediction and the correction terms.

[0048] Step Two: Exploring Virtual Experiments Guided by Uncertainty The goal of this step is to evaluate the stability of crop growth digital twin predictions, proactively identify areas of high uncertainty in its cognition, and establish correlation rules between environmental factors and yield response by performing pre-set experiments in the digital twin.

[0049] The specific implementation process is as follows: First, uncertainty is quantified. Using the crop growth digital twin constructed in step one, the input... The group comprises representative historical environment-driven sequences, obtained by resampling historical data. For example, integers greater than 10 For several predefined key growth stages, such as flowering and grain-filling, we collect the predicted values ​​of all key state variables at the end of each stage, such as the number of ears per unit area (SN) at the end of flowering. We then calculate the value of this variable at the end of each stage. Variance in the simulation .

[0050] Secondly, identify regions of high uncertainty. Set a variance threshold. The threshold One method to determine this is to calculate the median of the variance sequence for all key reproductive periods and all analyzed state variables. This will satisfy... The reproductive period was marked as a high-uncertainty period. Further analysis was conducted on the main environmental drivers that led to increased prediction variance during these periods, and the typical ranges of these factors within these periods were recorded.

[0051] As a clearer example, Table 1 below shows some of the results of uncertainty quantification analysis on winter wheat in a certain region: Table 1: Examples of High Uncertainty Region Identification Note: The data in Table 1 are for illustrative purposes only. The units of variance and threshold are the squares of the corresponding state variable units.

[0052] As shown in Table 1, the prediction variances for both the jointing stage and the flowering stage exceeded the threshold. (120.5), therefore it was marked by the system as a period of high uncertainty, and its corresponding key environmental factors and sensitive range were identified. Thus, for example, "flowering period, and daily maximum temperature" is used to define this period. The range of values ​​is "It is defined as a specific region of high uncertainty and serves as the target for subsequent virtual experiments."

[0053] Then, a virtual experiment was designed and executed. For an identified region of high uncertainty, it was assumed that it was caused by… Define one key environmental factor. Within the dimensional environmental factor value space, a Latin hypercube sampling method is used to systematically generate... By using a combination of parameters as the experimental group, this method can cover the parameter space well with fewer experiments. The value of needs to balance exploration efficiency and computational cost, for example .

[0054] Simultaneously, a set of environmental parameters representing the average climatic conditions or conventional agronomic management conditions of the region was set as a control group. While keeping the growth simulator parameters constant, simulations were driven by the environmental parameters of both the experimental and control groups, and the final yield of each simulation was recorded. .

[0055] Finally, the association patterns were summarized. The yields of each experimental group were calculated. Yield relative to control group rate of change: For all Perform cluster analysis, such as using the K-means algorithm, to divide them into several categories such as "significantly increased production", "normal production", and "significantly decreased production".

[0056] For the "significant production reduction" category, we analyzed the environmental driving data corresponding to all experimental groups falling into this category. For each environmental factor, we statistically analyzed its value distribution in the experimental groups of this category and identified the value ranges that occurred more than a preset proportion (e.g., 70%).

[0057] For example, it might be found that when "the highest daily temperature reaches 3 consecutive days during the flowering period" At this point, most experimental groups experienced significant yield reductions. This pattern of "environmental condition-yield response" is extracted as a correlation and stored in a knowledge base. This step, through active exploration, outputs a knowledge base on the relationship between environmental stress and yield loss.

[0058] Step 3: Prediction and Decision Generation Based on Multi-Source Evidence This step constructs a reasoning framework to collaboratively utilize real-time state estimation, historical knowledge patterns, and external information to generate probabilistic yield forecasts and corresponding agronomic recommendations.

[0059] The specific implementation process is as follows: First, evidence input and matching are performed. The reasoning framework receives three types of input: first, the association pattern knowledge base established in step two; second, the optimal estimate of the current crop growth status from the dynamic assimilation output in step one; and third, external information input in natural language (such as "current soil moisture is low" input by agricultural technicians).

[0060] The framework compares the current growth state and its environmental context with the preconditions of each associated pattern in the knowledge base. These preconditions are typically composed of multiple sub-conditions linked by a logical AND relationship. Matching degree. The calculation represents the proportion of the number of sub-conditions that satisfy the current state and the environment out of the total number of sub-conditions. Set activation threshold ,For example All matching degrees The pattern is activated, forming a candidate evidence set.

[0061] Secondly, evidence weights are dynamically allocated. Each activated mode... It has a basic weight This weight is initialized and periodically updated based on its historical prediction accuracy.

[0062] The reasoning framework detects whether there are conflicting pattern pairs in the candidate evidence set. For pattern pairs with a matching difference less than a set tolerance (e.g., 0.1) and conflicting conclusions, the system calls higher-precision or more real-time specialized observation data (such as water content sensor data for a specific soil layer) to re-evaluate the specific environmental factors on which the conflicting patterns depend.

[0063] Based on the results of this special assessment, we determine which model's preconditions are more consistent with current observations, and thus assign it a higher provisional weight. and correspondingly reduce the other party's Simultaneously, it parses the natural language information from external input and, through a predefined keyword-rule mapping table, transforms it into adjustment amounts for the weights of specific patterns. (For example, inputting "drought" will increase the weight of patterns related to water stress.)

[0064] Ultimately, the pattern Overall weight Calculated in the following way: .in These are weighting coefficients, and These coefficients are used to balance historical performance, current conflict resolution, and external experience. They are optimized through validation on historical data; one example value is... .

[0065] Then, weighted fusion prediction is performed. Each activated mode... Output the probability distribution of output change indicated by its conclusion. This distribution can be represented by the mean. and standard deviation The description approximates a normal distribution. The inference framework distributes the output of all activated modes according to their combined weights. By performing a weighted mixture, we obtain the comprehensive probability distribution regarding future output changes: .

[0066] Finally, forecast results and recommendations are generated. Based on the comprehensive probability distribution... and historical production baseline (e.g., average output over the previous three years), calculate the forecast range for future output. For example, calculate... of quantiles and quantiles The production forecast range is then... At the same time, the patterns with the highest comprehensive weights are extracted, and their associated agronomic control measures (such as "it is recommended to implement sprinkler irrigation to cool down under high temperature warning during flowering period") are output as decision-making suggestions.

[0067] Step 4: System closed-loop optimization based on prediction-observation feedback After a growing season ends, this step uses the final measured yield data to retrospectively evaluate and autonomously correct the system, enabling iterative evolution of the model.

[0068] The specific implementation process is as follows: First, error calculation and attribution analysis are performed to obtain the final actual output. Compared with the final production forecast generated in step three (e.g., the midpoint of the predicted interval) are compared to calculate the prediction error. Simultaneously, the detailed simulation log recorded throughout the entire growing season is replayed, containing the moment each rule constraint unit is triggered. Triggered rule identifier And a snapshot of the environment and state at the time of triggering.

[0069] Secondly, the rule base is evaluated and revised. This involves reviewing and revising each rule trigger record in the logs. Check a preset time window after triggering. (For example Within a day, check whether the direction of change of the simulated state is consistent with the direction of change of the actual observed state. If the directions are inconsistent, mark the trigger as a suspicious trigger.

[0070] Statistics for each rule Number of times marked as suspicious triggers throughout the growing season and its total number of triggers Calculate its suspected trigger rate. .

[0071] Set a threshold for suspicious trigger rate This threshold is used to filter potentially unreliable rules and can be set according to the requirements for model stability, typically ranging from 0.2 to 0.4. For example, in one implementation, it is set to... For all that satisfy The rule was used to determine it as a high-suspection problem rule.

[0072] For each rule with a high probability of being a problem, extract the set of environment and state data corresponding to all its suspected triggering cases. .analyze The statistical characteristics or common patterns of the data. For example, if... The solar radiation values ​​in most cases If all values ​​are below a certain level, it indicates that the original rule may fail under low-light conditions. Based on this analysis, the preconditions of the rule are refined, for example, by adding constraint variables or adjusting the threshold.

[0073] The specific revision is as follows: A sub-condition related to the common characteristic is added to the original rule's premise. For example, the rule "If..." ,but "If in Most cases meet Then it can be corrected to "if" and ,but Among them, the newly added threshold parameter According to middle The distribution (e.g., taking the 80th percentile) is used to determine this.

[0074] Finally, the neural network units are retrained. The newly acquired "environment-driven-state observation" data pairs from this growing season are merged into the historical training dataset. Using a loss function similar in form to the pre-training stage in step one, which includes a rule-triggered penalty term, the neural network units in the hybrid growth engine are incrementally trained (fine-tuned) to update their network weight parameters, enabling them to better internalize the revised physical rule constraints while absorbing new data.

[0075] Step 5: Pre-sowing forecasting and decision support applications This step applies the aforementioned methods to the planning stage before sowing, providing decision support based on multi-scenario simulation for variety selection and cultivation program development.

[0076] The specific implementation process is as follows: First, input forward-looking scenarios and alternative solutions. Input future growing season... A series of possible climate scenario sequences, For example, an integer greater than 5. These sequences are typically derived from ensemble forecasts of climate models or analyses of historical similar years. Input Genetic characteristic parameters of each candidate variety. Input Different sowing management plans are available, each plan including parameters such as sowing date, planting density, and base fertilizer application rate.

[0077] Secondly, run parallel scenario simulations. For each climate scenario... ,variety Management Plan Combinations Initialize an independent digital twin instance. On this instance, run the complete simulation and prediction process from step one to step three to obtain a corresponding output forecast. This process can be performed in parallel to improve efficiency.

[0078] Then, a comprehensive evaluation based on multiple indicators is conducted. This applies to each combination of variety and cultivation plan. Summarizing all of them Yield forecasts under various climate scenarios Calculate several evaluation metrics for this combination, including: average expected yield. Standard deviation of output Used to measure production stability and risk production. For example, it can be defined as the 10th quantile of the production series, representing the minimum production level under adverse conditions.

[0079] To visually illustrate the evaluation process, Table 2 below lists the indicator results after simulated evaluation, using three candidate solutions as examples: Table 2: Examples of Simulation Results for Pre-Sowing Program Evaluation Note: The comprehensive score in Table 2 is calculated based on preset weights (e.g., average output weight 0.4, stability weight 0.4, risk output weight 0.2), and the score has been standardized.

[0080] Finally, decision recommendations are generated. These recommendations are based on the decision-makers' assessment of each evaluation indicator (…). , , The preference weights are determined, and a multi-attribute decision-making method (such as weighted summation) is used to evaluate all preferences. The combinations are comprehensively evaluated and ranked. As shown in Table 2, Option 2 (Variety B, timely sowing, medium density) has the highest overall score because it has better stability (lower standard deviation) and higher risk yield (better guaranteed return) while having a higher average yield. The system will output a list of such top-ranked candidate combinations as recommended options and provide a detailed analysis report for each combination as a scientific basis for sowing decisions.

[0081] The above description is merely a preferred embodiment of this application, but the scope of protection of this application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in this application, based on the technical solution and inventive concept of this application, should be included within the scope of protection of this application.

Claims

1. A smart crop yield forecasting method based on deep learning and data assimilation, characterized in that, Includes the following steps: S1. Construct a digital twin of crop growth and perform dynamic data assimilation. The digital twin of crop growth is implemented through a hybrid growth engine, which includes a neural network unit and a rule constraint unit. Input time-series environmental data into the hybrid growth engine to iteratively generate a growth state simulation sequence. During the generation of the growth state simulation sequence, compare the real-time observed state data with the corresponding simulated state, and dynamically adjust the parameters of the hybrid growth engine based on the comparison differences. S2. Uncertainty-guided virtual experiment exploration, specifically: based on the differences in the prediction of key growth period states by the hybrid growth engine under different environmental driving sequences, identify regions with high prediction uncertainty; for the identified regions, automatically design and execute multiple sets of virtual growth experiments with different input parameters in the crop growth digital twin, and record the yield results of each experiment; from the yield results of all experiments, summarize the correlation pattern between environmental conditions and yield response. S3. Prediction and decision generation based on multi-source evidence: Specifically, construct a reasoning framework, taking the association pattern obtained in step S2, the current growth state estimate obtained after assimilation in step S1, and externally inputted empirical information as inputs to the reasoning framework; the reasoning framework performs consistency verification on various types of input evidence and assigns weights, performs reasoning based on the weighted fused evidence, and outputs a future yield range prediction and corresponding agronomic control suggestions. S4. System closed-loop optimization based on prediction-observation feedback, specifically: after completing the forecast of a growth cycle, the optimization process is triggered based on the difference between the final yield prediction value and the actual observed value; the optimization process includes: analyzing the correlation between the difference and the triggering record of the rule constraint unit in step S1, and correcting the rule constraint unit accordingly. And update the parameters of the neural network units using the newly added observation data; S5. Pre-sowing prediction and decision support application: Before sowing, based on long-term climate prediction data, multiple candidate variety parameters, and multiple sowing management schemes, the processes of steps S1 to S3 are run in parallel to simulate the yield performance of each scheme under different climate scenarios; based on the simulation results, the schemes are comprehensively evaluated and ranked by multiple indicators, and sowing decision suggestions are output.

2. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 1, characterized in that, In step S1, the hybrid growth engine generates the final simulation state for each time step by performing the following operations: The neural network unit outputs a preliminary state prediction value based on the input previous state and current environmental data; The rule constraint unit performs a compliance check on the preliminary state prediction value based on preset agricultural knowledge rules and calculates the degree of violation. Calculate a dynamic fusion weight based on the calculated degree of violation; Based on the degree of violation, a state correction item is generated; The preliminary state prediction value and the state correction term are weighted and fused according to the dynamic fusion weight to obtain the final simulated state.

3. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 2, characterized in that, The dynamic fusion weight is calculated as follows: the value of the dynamic fusion weight is negatively correlated with the value of the violation degree; when the violation degree is zero, the dynamic fusion weight makes the final simulated state equal to the preliminary state prediction value; when the violation degree is greater than zero, the dynamic fusion weight makes the state correction term affect the final simulated state.

4. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 1, characterized in that, The identification of regions with high prediction uncertainty in step S2 is achieved in the following way: Using the hybrid growth engine, parallel simulations were performed under multiple different environment-driven sequences to obtain multiple sets of simulation results. For multiple preset key reproductive periods, the variance of the same state variable at the end of each set of simulation results is calculated. The reproductive periods with calculated variances exceeding a preset threshold, along with the corresponding ranges of values ​​for the main environmental factors leading to high variances, are collectively marked as regions with high prediction uncertainty.

5. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 4, characterized in that, The automatic design and execution of the virtual growth experiment in step S2 specifically involves: generating multiple sets of environmental parameter combinations within the parameter space defined by the marked growth period and the range of values ​​of the main environmental factors, using the Latin hypercube sampling method as the experimental group; setting a set of baseline environmental parameters as the control group; and running the simulation in the crop growth digital twin using the parameters of the experimental group and the control group, respectively.

6. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 5, characterized in that, The inductive association pattern described in step S2 is achieved in the following way: Calculate the rate of change of simulated yield for each experimental group relative to simulated yield for the control group; Cluster analysis was performed on the rate of change in yield for all experimental groups to form multiple categories; For a selected category, analyze the distribution of environmental parameter values ​​for all experimental groups falling into that category, and extract common environmental characteristics; The common environmental characteristics are associated with the corresponding categories of output changes to form an association pattern.

7. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 1, characterized in that, In step S3, the inference framework performs consistency checks and assigns weights. Includes the following operations: The current growth status and its environmental conditions are matched with the preconditions of each association pattern, and association patterns with a matching degree exceeding a preset threshold are activated. Assign a base weight to each activated association pattern; Check if there are any conflicting conclusions between the activated association patterns. If there is a conflict, call the special observation data to arbitrate the conflicting parties and assign a temporary weight to each party. Analyze the empirical information from external inputs and transform it into adjustments to the weights of specific association patterns; Calculate the overall weight of each activated association pattern based on the base weight, temporary weight, and adjustment amount.

8. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 7, characterized in that, Step S3, which involves performing reasoning based on the weighted fusion of evidence, includes: Each activated association pattern outputs a probability distribution of output variation; The probability distributions of all activated association modes are weighted and mixed according to their respective comprehensive weights to obtain a comprehensive probability distribution of output change. Based on the comprehensive probability distribution of output changes and the pre-obtained historical output baseline, the prediction range for future output is calculated.

9. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 1, characterized in that, Step S4, which involves analyzing the correlation between differences and rule-triggered records and correcting rule constraint units, specifically includes: The simulation log of the entire growth cycle is replayed, and the log records the time and rule identifier of each time the rule constraint unit is triggered; For each trigger record, check whether the direction of change of the simulated state is consistent with the direction of change of the actual observed state within the preset time window after the trigger, and mark the records with inconsistent directions as suspicious triggers; Count the frequency at which each rule is marked as a suspicious trigger; Rules whose suspected trigger frequency exceeds a preset threshold are identified as rules that need to be corrected. Analyze the environmental and state data corresponding to all suspicious trigger records for each rule to be corrected to identify common characteristics; Based on the common features identified, the preconditions of the rule to be corrected are modified.

10. The intelligent crop yield forecasting method based on deep learning and data assimilation according to claim 1, characterized in that, The multi-indicator comprehensive evaluation and ranking in step S5 is as follows: For each combination of candidate varieties and sowing management schemes, calculate the average expected value, standard deviation, and risk quantile value of the simulated yield under all climate scenarios; calculate the comprehensive score of each combination according to the decision-maker's preset weights for these three indicators; and rank all combinations in descending order according to the comprehensive score.