A crop real-time intelligent irrigation decision-making method coupling agent model and reinforcement learning

CN122597103APending Publication Date: 2026-08-18NORTHEAST AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610773719.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-01
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0003]然而,作物生长模型通常具有参数数量多、计算过程复杂、逐日迭代计算量大等特点,在区域高分辨率空间栅格尺度条件下,面对实时气象变化、土壤墒情动态变化及多区域同步灌溉决策需求时,存在计算效率低、响应速度慢的问题,难以满足实时动态灌溉决策对快速预测与快速迭代的要求

Benefits of technology

[0040] This invention proposes a real-time intelligent irrigation decision-making method that couples a crop growth surrogate model with reinforcement learning. Utilizing an LSTM surrogate model incorporating physical constraints, it achieves high-precision and rapid temporal inference of crop growth status and root zone soil moisture, replacing the cumbersome iterative calculations of traditional mechanistic models. By deeply integrating the surrogate model with the SAC reinforcement learning algorithm, it achieves daily dynamic optimization and closed-loop iterative control of irrigation strategies around the goals of high yield and water conservation. This method can quickly adapt to real-time changes in weather, field soil moisture, and regional water supply conditions, effectively overcoming the technical shortcomings of traditional crop growth models, such as low computational efficiency and inability to meet real-time dynamic irrigation decision-making requirements. This invention can fully consider the spatial heterogeneity of regional agricultural water and soil resources, achieving grid-scale refined and temporally ordered intelligent irrigation optimization scheduling, providing a reliable and efficient theoretical basis and technical support for precision irrigation management of field crops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122597103A_ABST
    Figure CN122597103A_ABST
Patent Text Reader

Abstract

This invention discloses a real-time intelligent irrigation decision-making method for crops that couples a surrogate model with reinforcement learning, belonging to the field of agricultural irrigation technology. First, multi-source basic data of the study area are collected and decision sub-regions are divided. The PCSE crop growth model is localized and calibrated using Sobol global sensitivity analysis and particle swarm optimization algorithm, constructing a daily high-resolution database of crop growth status and soil moisture for the region from 2000 to 2020. Second, an LSTM deep learning surrogate model embedded with soil water balance physical constraints is constructed to achieve rapid and high-precision prediction of root zone soil moisture content and aboveground crop biomass for the next day. Finally, the surrogate model is embedded into a soft actor-commentator (SAC) reinforcement learning decision-making framework, defining the state space, action space, and multi-objective reward function, establishing a daily coupling iterative mechanism between the surrogate model and the reinforcement learning model, and outputting the optimal irrigation regime at the grid scale daily level, constrained by crop growth status, soil moisture, and regional water supply conditions. This invention solves the problems of low computational efficiency and difficulty in meeting the needs of real-time dynamic irrigation decision-making in traditional crop growth models, enabling high-resolution, rapid-response, and refined intelligent irrigation management for the region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of agricultural irrigation technology, and in particular to a real-time intelligent crop irrigation decision-making method that couples a proxy model with reinforcement learning. Background Technology

[0002] Currently, most regional-scale crop irrigation regime optimization research is based on the coupling of crop growth models and optimization algorithms, using methods such as multi-objective optimization, genetic algorithms, or dynamic programming to optimize irrigation regimes. These methods can effectively describe crop growth processes, soil moisture migration processes, and crop response mechanisms to changes in meteorological conditions, demonstrating high reliability in terms of irrigation decision-making accuracy.

[0003] However, crop growth models typically have a large number of parameters, complex calculation processes, and large daily iteration computational loads. Under the condition of regional high-resolution spatial grid scale, when facing the needs of real-time weather changes, dynamic changes in soil moisture, and synchronous irrigation decisions in multiple regions, they suffer from low computational efficiency and slow response speed, making it difficult to meet the requirements of real-time dynamic irrigation decision-making for rapid prediction and rapid iteration.

[0004] As the demands for dynamic response capabilities in regional real-time irrigation management continue to increase, existing irrigation decision-making methods based on the feedback and iteration of crop growth models and optimization algorithms are insufficient to meet the rapid dynamic decision-making needs under real-time changes in weather, soil moisture, and regional water supply conditions. Therefore, there is an urgent need for a regional real-time intelligent irrigation method that can replace the complex calculation process of traditional crop growth models. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention proposes a real-time intelligent irrigation decision-making method for crops that couples a surrogate model with reinforcement learning. This method first calibrates crop growth model parameters and constructs a long-term time-series dataset. Then, it builds a crop growth surrogate model that integrates physical constraints using an LSTM network, replacing the traditional model to rapidly extrapolate growth status and soil moisture. Next, the surrogate model is embedded into the SAC reinforcement learning framework to establish a closed-loop decision-making iteration mechanism. Combining multiple constraints such as crop growth thresholds, soil moisture, and total water supply, it dynamically optimizes daily irrigation amounts, efficiently adapting to real-time changes in weather and soil moisture, balancing yield increase and water conservation needs, and meeting the application requirements for rapid intelligent irrigation decision-making at the regional grid scale.

[0006] To solve the technical problem, the technical solution of the present invention is as follows:

[0007] A real-time intelligent irrigation decision-making method for crops that couples a proxy model with reinforcement learning, the method comprising:

[0008] S1: Collect crop planting distribution data, daily meteorological driving data, soil data, field management data and historical irrigation data within the study area; divide decision sub-regions based on differences in crop planting distribution, irrigation boundaries and soil types, and construct a daily spatial raster driving database covering each sub-region;

[0009] S2: Input the daily spatial grid-driven database into the crop growth model, and after parameter sensitivity analysis and localization calibration, generate a multi-year daily crop growth status and root zone soil moisture content database for each sub-region;

[0010] S3: Train a long short-term memory network model using the multi-year daily crop growth status and root zone soil moisture content database to construct a surrogate model for crop growth process; wherein, daily meteorological driving data, crop water requirement and irrigation amount are used as input variables, and the root zone soil moisture content and crop aboveground biomass of the next day are used as output targets. The model is trained by combining the mean squared error loss function and the physical constraints of soil water balance to obtain the surrogate model for crop growth process in each sub-region.

[0011] S4: Embed the crop growth process surrogate model into a reinforcement learning decision framework to construct a real-time intelligent irrigation decision system; input real-time meteorological data, soil moisture data, crop growth status data, and historical irrigation data into the reinforcement learning model to construct a daily irrigation decision state space, using daily irrigation volume as the action output variable, and input the irrigation decision results into the crop growth process surrogate model in real time to predict the changes in crop growth indicators and root zone soil moisture content for the next day, and feed the prediction results back to the reinforcement learning decision environment to complete the state update, realizing the daily coupling iteration of the surrogate model and the reinforcement learning model; during the adjustment of the irrigation system, use crop growth status, root zone soil moisture content, and regional water supply conditions as constraints to implement dynamic correction of the irrigation strategy; couple the crop growth process surrogate model with the soft actor-critic reinforcement learning algorithm for multi-environment training and strategy iterative optimization to obtain the daily real-time optimal irrigation system for each sub-region.

[0012] Furthermore, the specific process of parameter sensitivity analysis and localized calibration described in step S2 is as follows:

[0013] S21: The Sobol global sensitivity analysis method is used to perform sensitivity analysis on the input parameters of the crop growth model, calculate the first-order sensitivity index, second-order sensitivity index and global sensitivity index of each parameter, sort and select sensitive parameters according to the sensitivity index, and set non-sensitive parameters as default values.

[0014] S22: The combined particle swarm optimization algorithm and distributed crop growth model aim to minimize the error between the measured and simulated values ​​of crop growth indicators and root zone soil moisture content. The sensitive parameters at the grid scale of each sub-region are automatically calibrated and localized to obtain crop parameter sets and soil parameter sets adapted to each sub-region, and a regional high-resolution spatially distributed crop growth model is constructed.

[0015] Furthermore, the objective function for the localized calibration in step S22 is:

[0016]

[0017] In the formula, and These are the measured and simulated values ​​of the crop leaf area index, respectively. and These are the measured and simulated values ​​of dry matter accumulation in the aboveground parts of crops, respectively. and These are the measured and simulated values ​​of soil moisture content in the root zone, respectively. and These are the measured and simulated values ​​of crop yield, respectively. , , , These are the weighting coefficients for each indicator; The sampling frequency was determined to measure crop leaf area index, root zone soil moisture content, and aboveground dry matter accumulation. The number of times crop yield was sampled.

[0018] Furthermore, in step S3, the training process combines the mean squared error loss function and the physical constraints of soil water balance, employing a joint loss function, the formula of which is:

[0019]

[0020] Among them, the mean square error basic loss term for:

[0021]

[0022] Physical constraints of soil water balance for:

[0023]

[0024] In the formula, For the joint loss function, These are the weighting coefficients for the physical constraint terms. This represents the number of training samples; and The predicted and actual values ​​of soil volumetric moisture content in the root zone are shown below, respectively. and These are the predicted and actual values ​​of crop aboveground biomass for the following day; This represents the initial volumetric water content of the root zone soil on that day. This represents the rainfall for the day. This represents the amount of irrigation water used that day. This represents the actual evapotranspiration of the crop on that day. This represents the deep seepage volume in the root zone on that day. This represents the surface runoff for that day.

[0025] Furthermore, the input data of the long short-term memory network model described in step S3 is preprocessed using the min-max normalization method before training, mapping each index data to the [0,1] interval to eliminate the influence of dimensional differences on model training.

[0026] Furthermore, the state variables in the daily irrigation decision state space mentioned in step S4 include: the daily aboveground dry matter accumulation of the crop, the daily root zone soil volumetric moisture content, the 7-day weather forecast data, the daily regional water availability, and the current growth stage of the crop; the action space of the action output variable is a continuous action space, and the daily irrigation amount satisfies... ,in This is the maximum quota for a single irrigation.

[0027] Furthermore, the reward function of the reinforcement learning model described in step S4 is a multi-objective comprehensive reward function that takes into account both high crop yield and water-saving irrigation, and the formula is:

[0028]

[0029] In the formula, For the first Instant reward value on the decision-making day; The daily crop growth and yield gain is calculated from the daily increase in aboveground crop biomass. Conversion coefficient of biomass to economic output Calculated; To enforce penalties for violations, penalties are triggered when the root zone soil volumetric moisture content exceeds the upper or lower limits allowed for crop growth or when the cumulative irrigation amount exceeds the cumulative available water volume in the region as of that day. , These are the weight coefficients for each sub-item.

[0030] Furthermore, the daily coupling iteration mechanism between the surrogate model and the reinforcement learning model in step S4 is as follows:

[0031] In the On the decision-making day, the reinforcement learning agent receives the current environmental state. Based on the current strategy network, output irrigation decision actions. ;Will The current meteorological data and initial soil moisture content are input into the crop growth process surrogate model to predict the root zone soil moisture content for the next day. With crop aboveground biomass Calculate the daily instant reward based on the prediction results and reward function. The environmental status for the next day is obtained by combining the prediction results with real-time observation data from the following day. ; Transition sample The data is stored in the experience replay pool for updating the network parameters of the soft actor-critic reinforcement learning algorithm. The above process is repeated to complete the daily decision-making iteration throughout the entire reproductive period, realizing the real-time coupling and state closed-loop update of the agent model and the reinforcement learning model.

[0032] Furthermore, the constraints on the irrigation decision in step S4 include:

[0033] Root zone soil moisture content constraints: ,in For the first Soil volumetric moisture content in the Sungen area The crop wilting coefficient. It refers to the soil field water holding capacity;

[0034] Irrigation water constraints: ,and ,in For the first Daily irrigation volume per unit area The total number of days in the entire growth period of the crop. This refers to the total water supply per unit area throughout the entire reproductive period.

[0035] Aboveground dry matter constraints: ,in This refers to the actual aboveground dry matter weight of the crop. This is the preset minimum aboveground dry matter quality assurance threshold.

[0036] Furthermore, during the training process of the soft actor-critic reinforcement learning algorithm described in step S4, the target Q-network weights are updated using a soft update method, and the update formula is:

[0037]

[0038] In the formula, For the target Q network weights, For double-Q network weights, The soft update coefficient is used; at the same time, an entropy regularization term and an adaptive temperature coefficient are introduced to automatically adjust the temperature parameters by minimizing the temperature loss function, thus exploring and utilizing the balancing strategy.

[0039] This application has the following advantages:

[0040] This invention proposes a real-time intelligent irrigation decision-making method that couples a crop growth surrogate model with reinforcement learning. Utilizing an LSTM surrogate model incorporating physical constraints, it achieves high-precision and rapid temporal inference of crop growth status and root zone soil moisture, replacing the cumbersome iterative calculations of traditional mechanistic models. By deeply integrating the surrogate model with the SAC reinforcement learning algorithm, it achieves daily dynamic optimization and closed-loop iterative control of irrigation strategies around the goals of high yield and water conservation. This method can quickly adapt to real-time changes in weather, field soil moisture, and regional water supply conditions, effectively overcoming the technical shortcomings of traditional crop growth models, such as low computational efficiency and inability to meet real-time dynamic irrigation decision-making requirements. This invention can fully consider the spatial heterogeneity of regional agricultural water and soil resources, achieving grid-scale refined and temporally ordered intelligent irrigation optimization scheduling, providing a reliable and efficient theoretical basis and technical support for precision irrigation management of field crops. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This application provides a technical roadmap for a real-time intelligent crop irrigation decision-making method that couples a proxy model with reinforcement learning. Detailed Implementation

[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0044] Example 1:

[0045] like Figure 1 As shown, this embodiment provides a real-time intelligent irrigation decision-making method that couples a surrogate model of crop growth process with a reinforcement learning model, including the following steps:

[0046] S1: Basic data collection and irrigation decision-making area delineation;

[0047] Data on crop planting distribution, daily meteorological driving data, soil data, field management data, and historical irrigation data were collected within the study area. The meteorological driving data included solar radiation, maximum temperature, minimum temperature, precipitation, wind speed, and relative humidity. The soil attribute data included soil texture parameters, field water holding capacity, saturated water content, and root activity layer depth parameters. Based on the differences in crop planting distribution, irrigation boundaries, and soil types within the study area, decision sub-regions were divided, and a daily spatial raster-driven database covering each sub-region from 2000 to 2020 was constructed.

[0048] S2: Construct a long-term, high-resolution regional crop growth status database;

[0049] The model-driven data is input into the crop growth model to identify and optimize the sensitive crop and soil parameters that affect crop growth dynamics and soil moisture migration processes in each sub-region. The key parameters of the model are calibrated locally by combining the measured crop growth indicators and root zone soil moisture observation data of each sub-region, and a high-resolution spatially distributed PCSE growth model is constructed. Based on the daily spatial raster-driven database of each sub-region from 2000 to 2020 and the localized PCSE crop growth model, a multi-year daily crop growth status and root zone soil moisture database of each sub-region is generated.

[0050] S3: Construct a surrogate model of crop growth process based on LSTM;

[0051] The multi-year daily crop growth status and root zone soil moisture content database of each sub-region are input into a Long Short-Term Memory (LSTM) network model to construct a surrogate model of crop growth process at the regional spatial grid scale. Using daily meteorological data, crop water requirement, and irrigation amount as input variables, and the root zone soil moisture content and crop aboveground biomass of the next day as output targets, a dynamic mapping relationship of crop growth status is established. The model parameters are iteratively trained and jointly optimized by combining standardized preprocessing methods, mean squared error loss function, and physical constraints of soil water balance to obtain a high-resolution spatial grid scale surrogate model for rapid prediction of crop growth indicators and root zone soil moisture content based on crop growth process in each sub-region.

[0052] S4: Real-time irrigation decision-making by coupling crop growth surrogate model and reinforcement learning model;

[0053] By embedding crop growth surrogate models of each sub-region into a reinforcement learning decision-making framework, a real-time intelligent irrigation decision-making system for each sub-region's spatial grid scale is constructed. Real-time meteorological data, soil moisture data, crop growth status data, and historical irrigation data are input into the reinforcement learning model to construct a daily irrigation decision state space for the region. The daily irrigation amount is used as the action output variable, and the irrigation decision results are input into the crop growth surrogate model in real time to predict the changes in crop growth indicators and root zone soil moisture content for the next day. The prediction results are fed back to the reinforcement learning decision-making environment to complete the state update, realizing the daily coupling and iteration of the surrogate model and the reinforcement learning model. During the dynamic adjustment of the irrigation regime, the irrigation strategy is dynamically modified using crop growth status, root zone soil moisture content, and regional water supply conditions as constraints. The surrogate model and the soft actor-critic reinforcement learning algorithm (SAC) are coupled to carry out multi-environment training and strategy iterative optimization to obtain the daily real-time optimal irrigation regime for each sub-region at a high-resolution spatial grid scale.

[0054] For example, the specific steps of the crop growth model parameter sensitivity analysis, PCSE model localization calibration, and regional long-term high-resolution crop growth status database construction described in step S2 are as follows:

[0055] S21: Parameter Sensitivity Analysis and Sensitive Parameter Screening

[0056] The Sobol global sensitivity analysis method was used to conduct sensitivity analysis on the input parameters of the PCSE model, and the first-order sensitivity index, second-order sensitivity index and global sensitivity index of each parameter to the target variable were calculated. According to the sensitivity index ranking, the sensitive parameters that have a significant impact on crop growth dynamics and soil moisture migration process were screened out, and the non-sensitive parameters were set as the model default value. Only the sensitive parameters were subjected to subsequent local calibration.

[0057] Furthermore, the specific steps of Sobol in S21 are as follows:

[0058] S211: Determine the input parameters and reasonable value ranges for the PCSE crop growth model. The input parameters include crop variety genetic parameters, soil hydraulic properties parameters, and field management parameters. Determine the target variables for model calibration. The target variables are the crop leaf area index, aboveground dry matter accumulation, and root zone soil volumetric water content output in the time series, as well as the crop final yield output in the non-time series.

[0059] S212: Sampling of input variables:

[0060] Latin hypercube sampling is used to ensure that each variable is uniformly distributed within its range. A set of sample points needs to be generated for each input variable; these sample points will be used in the model's computation.

[0061] S213: Model Evaluation

[0062] The sample data generated by the Latin hypercube sampling method is output, the model is run to calculate the output results, and this process is repeated until enough sample output points are obtained to perform sensitivity analysis, so as to ensure that the sample size is large enough (several thousand samples) to ensure the accuracy of the analysis.

[0063] S214: Calculate variance decomposition:

[0064] The Sobol method evaluates the impact of input variables through variance decomposition. Specifically, the output variance can be represented as the sum of the variance contributions of each input variable:

[0065]

[0066] Among them, V i V represents the contribution of a single input variable to the output variance. ij This represents the contribution of the interaction effect of the two input variables to the output variance, where k is the number of input parameters and Var(Y) is the total output variance of the model.

[0067] S215: Calculate the Sobol sensitivity index:

[0068] The Sobol index is a quantitative indicator used to measure the contribution of each input parameter to the model output; the first-order sensitivity index, second-order sensitivity index, and full-order sensitivity index are calculated using programming software.

[0069] First-order sensitivity index (S i The variance denoted by represents the contribution of a single input variable to the output variance. It can be understood as the degree of influence of that input variable on the output result without considering other interactions. The calculation formula is as follows:

[0070]

[0071] S i It is a first-order sensitivity index, V i Var(Y) represents the contribution of a single input variable to the output variance, and Var(Y) is the total output variance of the model.

[0072] Second-order sensitivity index (S ij The variance denoted by represents the contribution of the interaction between two input variables to the output variance. The calculation formula is as follows:

[0073]

[0074] S ij It is a second-order sensitivity index, V ij This represents the contribution of the interaction effect of two input variables to the output variance.

[0075] Global Sensitivity Index (ST) i This represents the total contribution of a given input variable to the output, including the first-order effect of that variable and its interaction effects with other input variables. The calculation formula is as follows:

[0076]

[0077] ST i It is the global sensitivity index, V -i This represents the variance of all parameters except the i-th parameter.

[0078] S216: Results Analysis and Interpretation

[0079] Based on the calculated sensitivity indices, the sensitivity of each parameter to crop yield, leaf area index and aboveground dry matter accumulation is analyzed, and parameter calibration is performed on input parameters with larger sensitivity indices.

[0080] S22: Localized calibration of PCSE model parameters:

[0081] By combining the particle swarm optimization (PSO) algorithm with the distributed PCSE crop growth model, automated calibration and adjustment of the selected sensitive parameters are carried out. With the optimization objective of minimizing the error between the measured and simulated values ​​of crop growth indicators and root zone soil moisture content, the localized calibration of model parameters at the grid scale of each sub-region is completed, and a set of crop and soil parameters adapted to the study area is obtained, thus constructing a high-resolution spatially distributed PCSE growth model for the region.

[0082] Furthermore, the specific steps of PSO in S22 are as follows:

[0083] S221: Initialization: Set the initial particle swarm size, define the maximum number of algorithm iterations, and set the initial values ​​for particle position and velocity. Each particle represents a set of model parameter combinations; its position is the model parameter combination, and its velocity is the step size for updating that particle.

[0084] S222: Based on the objective function, evaluate the fitness of each particle, that is, minimize the error between the measured and simulated values ​​of crop growth indicators. The fitness value is used to measure the quality of the solution of that particle.

[0085] S223: Update speed and position: Based on the particle's historical best position (individual best) and the global best position in the swarm, adjust the particle's speed and position using the particle swarm optimization update formula;

[0086] S224: Local and Global Optimal Updates: After updating its velocity and position, each particle recalculates its fitness value and compares it with its historical best position. If the current fitness is better, the particle's best position is updated. The particle with the best fitness across the entire population is selected as the global optimum through fitness evaluation.

[0087] S225: Evaluation: The updated particles are evaluated for fitness based on the objective function, and their performance in the optimization objective is calculated. Fitness evaluation is used to determine whether to update the individual optimal position and global optimal position of the particles;

[0088] S226: Iterative update: Repeat steps 3 to 5 until the maximum number of iterations or the fitness value converges.

[0089] S227: Termination condition: The algorithm terminates when the maximum number of iterations is reached or the fitness converges.

[0090] S228: Output result: Output the optimal solution set, which is the optimal combination of model parameters that minimizes the error between the simulated and measured values.

[0091] According to the implementation method, the objective function for PCSE model localization calibration in step S22 is as follows:

[0092]

[0093] In the formula, and These are the measured and simulated values ​​of the crop leaf area index, respectively. and These are the measured and simulated values ​​of dry matter accumulation in the aboveground parts of crops, respectively. and These are the measured and simulated values ​​of soil moisture content in the root zone, respectively. and These are the measured and simulated values ​​of crop yield, respectively. , , , These are the weighting coefficients for each indicator; n represents the number of samplings for the measured crop leaf area index, root zone soil moisture content, and aboveground dry matter accumulation; and M represents the number of samplings for crop yield.

[0094] S23: Construction of a regional long-term high-resolution crop growth status database;

[0095] Daily meteorological driving data, soil data, and field management data at the raster scale for each sub-region of the study area from 2000 to 2020 were input into the calibrated spatially distributed PCSE growth model to conduct long-sequence daily simulations. Daily crop growth status data (including leaf area index, aboveground dry matter accumulation, and crop growth period) and root zone soil moisture data of each raster were extracted from the simulation output to construct a daily crop growth status and soil moisture database covering the entire raster of the study area from 2000 to 2020, providing a basic dataset for subsequent surrogate model training.

[0096] According to some implementation methods, the construction, training, and optimization of the LSTM-based crop growth process surrogate model in step S3 are as follows:

[0097] S31: Determine the input and output variables of the surrogate model and the dataset partitioning:

[0098] The input variables and output targets of the LSTM surrogate model are defined as follows: the input variables are daily meteorological driving data (solar radiation, maximum temperature, minimum temperature, precipitation, wind speed, relative humidity) at the grid scale of the study area, the daily crop water requirement, and the daily irrigation amount; the output targets are the root zone soil volumetric water content and the aboveground biomass of crops the next day; the constructed daily dataset from 2000 to 2020 is divided into two parts, with the data from 2000 to 2017 used as the model training set and the data from 2018 to 2020 used as the model validation and test sets.

[0099] S32: Data standardization preprocessing:

[0100] The raw data of the input variables and output targets are preprocessed using standardization. The min-max normalization method is used to map all data to the [0,1] interval to eliminate the influence of differences in the scale of different indicators on model training. The normalization formula is as follows:

[0101]

[0102] In the formula, X norm The data is normalized, and X is the original data. min X is the minimum value of the data series of this indicator. max This represents the maximum value of the data sequence for this indicator.

[0103] S33: LSTM proxy model network structure construction:

[0104] A multi-input, multi-output LSTM deep learning network structure was constructed, consisting of an input layer, an LSTM hidden layer, a fully connected layer, and an output layer. The input layer receives temporally sequenced input variables, the LSTM hidden layer extracts temporal features of crop growth and soil moisture migration, the fully connected layer performs feature mapping and dimension transformation, and the output layer outputs predicted values ​​of the next day's root zone soil volumetric water content and crop aboveground biomass increment. Hyperparameters such as the model's temporal input step size, the number of hidden layer neurons, and the dropout ratio were set to suppress model overfitting.

[0105] S34: Model Loss Function Construction and Physical Constraint Embedding

[0106] We construct a basic loss function with mean squared error (MSE) as the core, and embed soil water balance physical constraints to build a joint loss function. This achieves the integration and optimization of model data-driven and physical mechanisms, and avoids the model from producing prediction results that violate hydrophysical laws.

[0107] The joint loss function formula for the LSTM surrogate model in step S34 is as follows:

[0108]

[0109] In the formula, Loss is the joint loss function of the model. MSE Loss is the basic loss term for mean squared error. WSB This is a physical constraint term for soil water balance. These are the weighting coefficients for the physical constraint terms.

[0110] The formula for calculating the mean squared error basic loss term is as follows:

[0111]

[0112] In the formula, N is the number of training samples, and SM i,pred and SM i,true These represent the model-predicted and actual values ​​of the root zone soil volumetric water content for the i-th sample on the following day; TAGP i,pred and TAGP i,true These are the model-predicted and actual values ​​of the aboveground biomass of the crop on the next day for the i-th sample, respectively.

[0113] The formula for calculating the physical constraints of soil water balance is:

[0114]

[0115] In the formula, SM i,in Let P be the initial root zone soil volumetric water content of the i-th sample on that day. i I represents the daily rainfall. i ET represents the daily irrigation volume.a,i D represents the actual evapotranspiration of the crop on that day. i R represents the deep seepage volume in the root zone on that day. i This represents the surface runoff for that day.

[0116] S35: Iterative Training and Hyperparameter Optimization of LSTM Surrogate Model

[0117] The Adam algorithm is used to iteratively update the model parameters, setting training hyperparameters such as the initial learning rate, batch size, and maximum number of training epochs. Forward propagation calculations are performed on the training set, and the prediction error is calculated using the joint loss function. Then, the network weights and bias parameters are updated using the backpropagation algorithm. The model fitting effect is monitored in real time based on the validation set, and an early stopping mechanism is used to suppress overfitting. Training is terminated when the validation set loss does not decrease for several consecutive epochs. The grid search method is combined to optimize the model hyperparameters and obtain the optimal hyperparameter combination.

[0118] S36: Model Accuracy Evaluation and Generalization Validation

[0119] The accuracy of the trained LSTM surrogate model was evaluated using test set data, with the coefficient of determination R being selected. 2 The root mean square error (RMSE) and mean absolute error (MAE) are used as model accuracy evaluation indicators to verify the model's prediction accuracy for crop growth status and root zone soil moisture content. Through sample testing across multiple sub-regions, growth stages, and meteorological years, the model's spatial and temporal generalization are verified, ultimately obtaining a high-precision surrogate model of the crop growth process at the grid scale for each sub-region. According to the implementation method, the calculation formula for the model accuracy evaluation indicators in step S36 is as follows:

[0120] Coefficient of determination:

[0121]

[0122] Root mean square error:

[0123]

[0124] Mean absolute error:

[0125]

[0126] In the formula, For the sample true value, ȳ is the model's predicted value, ȳ is the average of the true value sequence, and N is the number of test samples.

[0127] For example, the construction of the crop growth surrogate model coupled with the SAC reinforcement learning model and the optimization of real-time irrigation decisions described in step S4 are as follows:

[0128] S41: Building a Reinforcement Learning Decision-Making Environment

[0129] A digital environment for irrigation decision-making is constructed based on a crop growth agent model. The state space, action space, reward function, and state transition mechanism of the decision-making process are defined to realize the digital mapping of irrigation decision-making scenarios.

[0130] S42: State space construction:

[0131] Define the daily state space of the reinforcement learning model. The state variables are real-time observational data that characterize crop growth, soil moisture, meteorological conditions, and water supply constraints. Specifically, these include: daily aboveground dry matter accumulation, daily root zone soil volumetric water content, 7-day weather forecast data, daily regional water availability, and the current growth stage of the crop. After standardizing the above-mentioned real-time collected data, it is input into the reinforcement learning model to form the daily irrigation decision state space S. t , where t is the decision date sequence.

[0132] S43: Action Space Construction:

[0133] The action output space of the reinforcement learning model is defined, and the action variable is the irrigation amount per unit area At on the decision day. Combining the irrigation engineering design standards of the study area, the water demand characteristics of the crop growth period and the regional water supply capacity, upper and lower limits of irrigation amount are set, i.e. At∈[0,Amax], where Amax is the maximum quota for a single irrigation. The action space is a continuous action space, which is adapted to the continuous decision-making characteristics of the SAC algorithm to realize the fine and continuous control of daily irrigation amount.

[0134] S44: Reward Function Construction:

[0135] A multi-objective comprehensive reward function that balances high crop yield and water-saving irrigation is constructed as the core guide for iterative optimization of reinforcement learning strategies. The reward function calculation formula is as follows:

[0136]

[0137] In the formula, The immediate reward value on decision day t. The crop yield for that day. To constrain violations of penalties, , These are the weight coefficients for each sub-item.

[0138] The formula for calculating the production gain component is as follows:

[0139]

[0140] In the formula, This represents the increase in aboveground biomass of the crop on day t. This is the conversion coefficient from biomass to economic output.

[0141] The constraint violation penalty item is used to penalize violations of soil moisture constraints, water supply constraints, and excessive irrigation. The calculation formula is as follows:

[0142]

[0143] In the formula, The maximum penalty value, Let be the soil volumetric water content in the root zone on day t. , These are the lower and upper limits of the permissible root zone soil moisture content during the crop's growth period;

[0144] For cumulative irrigation volume, This represents the cumulative available water supply for the region up to day t.

[0145] S45: Specific execution steps of the SAC reinforcement learning algorithm:

[0146] S451: Initialization: Set the hyperparameters of the SAC algorithm, including the experience replay pool capacity, batch size, soft update coefficient, discount factor, learning rate, and maximum number of training rounds; initialize the policy network (actor network), double Q network (critic network), and target Q network, set the initial weights and bias parameters of each network, and synchronize the weights of the target Q network with the weights of the double Q network.

[0147] S452: Environmental Interaction and Sample Collection: In each training round, starting from the initial state of the crop growth period, the agent outputs irrigation actions based on the current policy network, interacts with the decision environment of the embedded agent model, completes the daily state transitions throughout the entire growth period, collects state transition samples and stores them in the experience replay pool.

[0148] S453: Sample Sampling and Network Training: Once the number of samples in the experience replay pool reaches the batch size requirement, a batch of state transition samples is randomly sampled from the replay pool and used to train and update the policy network and the double Q network respectively. The double Q network parameters are updated by minimizing the Q network loss function, and the policy network parameters are updated by maximizing the expected soft Q value of the policy network. At the same time, the reparameterization technique is used to handle the gradient propagation problem of continuous action sampling.

[0149] S454: Target Network Soft Update: The weight parameters of the target Q-network are updated using a soft update method. The update formula is as follows:

[0150]

[0151] In the formula, For the target Q network weights, For double-Q network weights, This is the soft update coefficient.

[0152] S455: Entropy Regularization and Adaptive Temperature Coefficient Update: The exploration and utilization of the balancing strategy by introducing an entropy regularization term, setting an adaptive temperature coefficient, and automatically adjusting the temperature parameter by minimizing the temperature loss function to ensure the stochastic exploration and convergence stability of the strategy.

[0153] S456: Iteration Termination Judgment: Repeat steps S462 to S465 until the preset maximum number of training rounds is reached, or the cumulative reward value of the policy network converges to a stable range for multiple consecutive rounds, then terminate the algorithm training.

[0154] S457: Optimal Strategy Output: After training, the converged optimal strategy network is output, which is the real-time irrigation decision optimal strategy model adapted to the grid scale of each sub-region of the study area.

[0155] S46: Daily coupling and iterative mechanism between the surrogate model and the reinforcement learning model:

[0156] The trained LSTM crop growth agent model is embedded into the SAC reinforcement learning decision framework as the core module for state transition in the decision environment, and a daily coupled iterative mechanism is constructed.

[0157] According to the implementation method, the specific process of constructing the daily coupling iteration mechanism in step S45 is as follows:

[0158] S461: On decision day t, the reinforcement learning agent receives the current environmental state S. t Based on the current strategy network, output irrigation decision action A. t (Daily irrigation volume);

[0159] S462: Decision Action A t With the current state S t Meteorological data and initial soil moisture content are input into the LSTM crop growth surrogate model to predict the root zone soil moisture content SM the following day. t+1 TAGP of crop aboveground biomass t ;

[0160] S463: Calculate the daily instant reward R based on the prediction results and reward function. t Simultaneously, by combining the prediction results with real-time observation data from the following day, the environmental state S for the next day is updated. t+1 ;

[0161] S464: Store the (St, At, Rt, St+1) state transition samples into the empirical replay pool for network parameter updates in the SAC algorithm;

[0162] S465: Repeat steps S451 to S454 to complete the daily decision-making iteration throughout the entire reproductive period, realizing the real-time coupling and state closed-loop update of the surrogate model and the reinforcement learning model.

[0163] S47: Constraints for Real-Time Irrigation Decisions

[0164] Root zone soil moisture content constraints:

[0165]

[0166] In the formula, Let be the soil volumetric moisture content in the crop root zone on day t. , These are the lower and upper limits of the root zone soil moisture content corresponding to the crop's growth stage on day t, where the lower limit is the crop wilting coefficient and the upper limit is the soil field holding capacity.

[0167] Irrigation water constraints:

[0168]

[0169]

[0170] In the formula, Let be the irrigation amount per unit area on day t. Let T be the maximum irrigation quota for a single irrigation during the crop's growth stage on day t, where T is the total number of days in the crop's entire growth period. This refers to the total water supply per unit area throughout the entire growth period of the crop.

[0171] Aboveground dry matter constraints:

[0172] TAGP≥TAGPmin

[0173] In the formula, TAGP represents the actual aboveground dry matter mass per unit area of ​​the crop, and TAGPmin represents the minimum aboveground dry matter mass guarantee threshold set for the study area.

[0174] S48: Reads meteorological monitoring data, soil moisture data and field management data of each grid in real time, inputs them into the decision model and outputs the daily optimal irrigation amount. When there are sudden changes in regional water supply conditions and weather forecasts, the decision state space is updated based on the latest real-time data, and the strategy iteration and irrigation decision correction are completed again to realize real-time intelligent irrigation decision-making that is rolled out daily throughout the entire growth period.

[0175] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0176] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A real-time intelligent irrigation decision-making method for crops that couples a proxy model with reinforcement learning, characterized in that, The method includes: S1: Collect crop planting distribution data, daily meteorological driving data, soil data, field management data and historical irrigation data within the study area; divide decision sub-regions based on differences in crop planting distribution, irrigation boundaries and soil types, and construct a daily spatial raster driving database covering each sub-region; S2: Input the daily spatial grid-driven database into the crop growth model, and after parameter sensitivity analysis and localization calibration, generate a multi-year daily crop growth status and root zone soil moisture content database for each sub-region; S3: Train a long short-term memory network model using the multi-year daily crop growth status and root zone soil moisture content database to construct a surrogate model for crop growth process; wherein, daily meteorological driving data, crop water requirement and irrigation amount are used as input variables, and the root zone soil moisture content and crop aboveground biomass of the next day are used as output targets. The model is trained by combining the mean squared error loss function and the physical constraints of soil water balance to obtain the surrogate model for crop growth process in each sub-region. S4: Embed the crop growth process surrogate model into a reinforcement learning decision framework to construct a real-time intelligent irrigation decision system; input real-time meteorological data, soil moisture data, crop growth status data, and historical irrigation data into the reinforcement learning model to construct a daily irrigation decision state space, using daily irrigation volume as the action output variable, and input the irrigation decision results into the crop growth process surrogate model in real time to predict the changes in crop growth indicators and root zone soil moisture content for the next day, and feed the prediction results back to the reinforcement learning decision environment to complete the state update, realizing the daily coupling iteration of the surrogate model and the reinforcement learning model; during the adjustment of the irrigation system, use crop growth status, root zone soil moisture content, and regional water supply conditions as constraints to implement dynamic correction of the irrigation strategy; couple the crop growth process surrogate model with the soft actor-critic reinforcement learning algorithm for multi-environment training and strategy iterative optimization to obtain the daily real-time optimal irrigation system for each sub-region.

2. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, The specific process of parameter sensitivity analysis and localized calibration in step S2 is as follows: S21: The Sobol global sensitivity analysis method is used to perform sensitivity analysis on the input parameters of the crop growth model, calculate the first-order sensitivity index, second-order sensitivity index and global sensitivity index of each parameter, sort and select sensitive parameters according to the sensitivity index, and set non-sensitive parameters as default values. S22: The combined particle swarm optimization algorithm and distributed crop growth model aim to minimize the error between the measured and simulated values ​​of crop growth indicators and root zone soil moisture content. The sensitive parameters at the grid scale of each sub-region are automatically calibrated and localized to obtain crop parameter sets and soil parameter sets adapted to each sub-region, and a regional high-resolution spatially distributed crop growth model is constructed.

3. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 2, characterized in that, The objective function for the localized calibration in step S22 is: , In the formula, and These are the measured and simulated values ​​of the crop leaf area index, respectively. and These are the measured and simulated values ​​of dry matter accumulation in the aboveground parts of crops, respectively. and These are the measured and simulated values ​​of soil moisture content in the root zone, respectively. and These are the measured and simulated values ​​of crop yield, respectively. , , , These are the weighting coefficients for each indicator; The sampling frequency was determined to measure crop leaf area index, root zone soil moisture content, and aboveground dry matter accumulation. The number of times crop yield was sampled.

4. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, Step S3 describes training by combining the mean squared error loss function and the physical constraints of soil water balance, using a joint loss function, the formula of which is: , Among them, the mean square error basic loss term for: , Physical constraints of soil water balance for: , In the formula, For the joint loss function, These are the weighting coefficients for the physical constraint terms. This represents the number of training samples; and The predicted and actual values ​​of soil volumetric moisture content in the root zone are shown below, respectively. and These are the predicted and actual values ​​of crop aboveground biomass for the following day; This represents the initial volumetric water content of the root zone soil on that day. This represents the rainfall for the day. This represents the amount of irrigation water used that day. This represents the actual evapotranspiration of the crop on that day. This represents the deep seepage volume in the root zone on that day. This represents the surface runoff for that day.

5. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, The input data of the long short-term memory network model described in step S3 is preprocessed by the min-max normalization method before training, which maps each index data to the [0,1] interval to eliminate the influence of dimensional differences on model training.

6. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, The state variables in the daily irrigation decision state space mentioned in step S4 include: the daily aboveground dry matter accumulation of the crop, the daily root zone soil volumetric moisture content, the 7-day weather forecast data, the daily regional water availability, and the current growth stage of the crop; the action space of the action output variable is a continuous action space, and the daily irrigation amount satisfies... ,in This is the maximum quota for a single irrigation.

7. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, The reward function of the reinforcement learning model described in step S4 is a multi-objective comprehensive reward function that takes into account both high crop yield and water-saving irrigation, and the formula is: , In the formula, For the first Instant reward value on the decision-making day; The daily crop growth and yield gain is calculated from the daily increase in aboveground crop biomass. Conversion coefficient of biomass to economic output Calculated; To enforce penalties for violations, penalties are triggered when the root zone soil volumetric moisture content exceeds the upper or lower limits allowed for crop growth or when the cumulative irrigation amount exceeds the cumulative available water volume in the region as of that day. , These are the weight coefficients for each sub-item.

8. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, The daily coupling iteration mechanism between the surrogate model and the reinforcement learning model mentioned in step S4 is as follows: In the On the decision-making day, the reinforcement learning agent receives the current environmental state. Based on the current strategy network, output irrigation decision actions. ;Will The current meteorological data and initial soil moisture content are input into the crop growth process surrogate model to predict the root zone soil moisture content for the next day. With crop aboveground biomass Calculate the daily instant reward based on the prediction results and reward function. The environmental status for the next day is obtained by combining the prediction results with real-time observation data from the following day. ; Transition sample The data is stored in the experience replay pool for updating the network parameters of the soft actor-critic reinforcement learning algorithm. The above process is repeated to complete the daily decision-making iteration throughout the entire reproductive period, realizing the real-time coupling and state closed-loop update of the agent model and the reinforcement learning model.

9. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, The constraints for the irrigation decision in step S4 include: Root zone soil moisture content constraints: ,in For the first Soil volumetric moisture content in the Sungen area This is the crop wilting coefficient. It refers to the soil field water holding capacity; Irrigation water constraints: ,and ,in For the first Daily irrigation volume per unit area The total number of days in the entire growth period of the crop. This refers to the total water supply per unit area throughout the entire reproductive period. Aboveground dry matter constraints: ,in This refers to the actual aboveground dry matter weight of the crop. This is the preset minimum aboveground dry matter quality assurance threshold.

10. The real-time intelligent irrigation decision-making method based on the surrogate model and reinforcement learning model of the crop growth process according to claim 1, characterized in that, In step S4, during the training of the soft actor-critic reinforcement learning algorithm, the target Q-network weights are updated using a soft update method. The update formula is as follows: , In the formula, For the target Q network weights, For double-Q network weights, The soft update coefficient is used; at the same time, an entropy regularization term and an adaptive temperature coefficient are introduced to automatically adjust the temperature parameters by minimizing the temperature loss function, thus exploring and utilizing the balancing strategy.