Decision-making methods, systems, equipment and media for paddy field irrigation and drainage

By constructing a simulated paddy field growth environment and transforming it into a Markov decision process, the policy network and value network of the agent are trained, solving the problem of optimizing the entire growth period in paddy field irrigation and drainage management, and achieving efficient water resource utilization and improved decision quality.

CN122089113APending Publication Date: 2026-05-26WUHAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV
Filing Date
2026-04-23
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve long-term optimization throughout the entire growth cycle in paddy field irrigation and drainage management. Furthermore, methods based on multi-objective optimization and reinforcement learning are limited by the action space and can only achieve short-term benefit optimization.

Method used

By acquiring historical environmental parameters of the target paddy field, a paddy field growth simulation environment is constructed. The irrigation and drainage decision-making process is transformed into a Markov decision process. A reinforcement learning environment is constructed to train the agent's policy network and value network. The trained policy network is then used to execute irrigation and drainage decisions, thereby achieving dynamic control of the irrigation and drainage process.

Benefits of technology

It improves water resource utilization efficiency, reduces the probability of over-irrigation or under-drainage, and enhances long-term decision modeling capabilities and decision quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089113A_ABST
    Figure CN122089113A_ABST
Patent Text Reader

Abstract

This application relates to the field of agricultural irrigation control technology, and in particular to a method, system, equipment, and medium for paddy field irrigation and drainage decision-making. The method includes: acquiring historical environmental parameters of the target paddy field; constructing a paddy field growth simulation environment based on the historical environmental parameters; transforming the paddy field irrigation and drainage decision-making process concept into a Markov decision process within the paddy field growth simulation environment, and constructing a reinforcement learning environment through the Markov decision process; training the agent's policy network and value network within the paddy field growth simulation environment; updating the network parameters of the value network through the reinforcement learning environment during training, and updating the network parameters of the policy network through the updated value network; and executing the target paddy field irrigation and drainage decisions using the trained policy network. This solves the problems of related technologies that focus on static or short-term decision analysis at local time scales, model only single irrigation actions, have limited action space, and can only achieve short-term benefit optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of agricultural irrigation control technology, and in particular to a method, system, equipment and medium for making decisions on irrigation and drainage in paddy fields. Background Technology

[0002] Rice irrigation and drainage management directly affects water resource utilization efficiency and crop yield, making precise regulation essential. Traditional methods rely heavily on manual experience, which struggles to respond promptly to changes in weather and field water conditions, resulting in yield losses and water waste. Therefore, developing intelligent optimization for paddy field irrigation and drainage is of great significance.

[0003] In related technologies, multi-objective optimization-based methods focus on static or short-term decision analysis at local time scales, and can only reflect short-term regulatory needs; reinforcement learning-based methods only model single irrigation actions, have limited action space, and lack the ability to optimize long-term benefits throughout the entire growth period. Summary of the Invention

[0004] This application provides a method, system, equipment, and medium for paddy field irrigation and drainage decision-making, in order to solve the problems of related technologies that focus on static or short-term decision analysis at local time scales, only model a single irrigation action, have limited action space, and can only achieve short-term benefit optimization.

[0005] The first aspect of this application provides a method for making irrigation and drainage decisions for paddy fields, comprising the following steps: obtaining historical environmental parameters of a target paddy field; constructing a paddy field growth simulation environment based on the historical environmental parameters; transforming the concept of the paddy field irrigation and drainage decision-making process into a Markov decision process in the paddy field growth simulation environment, and constructing a reinforcement learning environment through the Markov decision process; training the agent's policy network and value network in the paddy field growth simulation environment, and during the training process, updating the network parameters of the value network through the reinforcement learning environment, and updating the network parameters of the policy network through the updated value network; and executing the irrigation and drainage decisions for the target paddy field using the trained policy network.

[0006] Optionally, in one embodiment of this application, the reinforcement learning environment includes a state space, an action space, a transition function, a reward function, and a discount factor. The environmental state vector of the state space includes the forecast rainfall sequence for the next preset number of days, the current water depth, the lower limit of the suitable water depth, the upper limit of the suitable water depth, the upper limit of rainwater storage, and water depth threshold parameters related to flood tolerance. The action space is constructed based on the decision variables of the paddy field irrigation and drainage decision cycle. The reward function includes immediate rewards and round rewards. Immediate rewards are used to evaluate rainfall utilization, irrigation and drainage volume, and the impact on crop yield within a single decision cycle. Round rewards are used to penalize the number of non-zero irrigation and drainage decisions throughout the entire growth period. The transfer function expression is:

[0007] in, For the irrigation cycle number The initial water depth of the sky, For the irrigation cycle number The initial water depth of the sky, For the first The amount of irrigation and drainage decisions per day and To determine the amount of irrigation and drainage based solely on factors other than rainfall, evapotranspiration, and seepage. The resulting changes in the water layer.

[0008] Optionally, in one embodiment of this application, the expression for the instant reward is:

[0009] in, For the first Instant rewards during the decision-making cycle. For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. For the first The decision-making cycle is related to rewards and penalties for drought or flood-induced yield reduction. For the first The decision-making cycle is related to the water consumption penalty associated with the current irrigation and drainage volume; The expression for round reward is:

[0010] in, For the total reward function, The penalty coefficient is... The total number of irrigation and drainage actions within a round. This represents the number of decision steps within a complete decision-making cycle.

[0011] Optionally, in one embodiment of this application, updating the network parameters of the value network through a reinforcement learning environment includes: inputting the environmental state vector of the state space into the agent; inputting the irrigation and drainage decision quantity output by the agent into the action space; executing the irrigation and drainage decision quantity through the action space; inputting the execution result of the action space into the reinforcement learning environment; updating the environmental state vector through the reinforcement learning environment; inputting the updated environmental state vector into the agent; updating the network parameters of the value network through the updated environmental state vector; and updating the network parameters of the policy network through the updated value network.

[0012] Optionally, in one embodiment of this application, before inputting the irrigation and drainage decision quantity output by the agent into the action space, the method further includes: extracting the current water layer depth and water layer threshold parameters from the environmental state vector; and correcting the irrigation and drainage decision quantity based on the current water layer depth and water layer threshold parameters.

[0013] Optionally, in one embodiment of this application, the paddy field growth simulation environment includes a paddy field water balance model, a rice water production function, and a rice yield reduction model due to waterlogging, wherein the expression of the paddy field water balance model is:

[0014] in, For the irrigation cycle number The initial water depth of the day; For the irrigation cycle number t The initial water depth of the day; For the first Forecast rainfall for the day; To determine the first action based on the decision-making action Daily irrigation and drainage decision-making volume; For the first Forecast of crop water requirements for the day; For the first Forecast field seepage volume for the day; The expression for the rice yield reduction model due to waterlogging is:

[0015] in, For rice yield reduction due to flooding; For the first The depth of the water layer in the field at that time; For rice plant height; To withstand deep flooding; These are empirical parameters. These are empirical parameters; The expression for the rice water production function is:

[0016] in, For the first Crop yield indicators during the decision-making cycle; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage Forecast evapotranspiration for the day; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage The potential evapotranspiration of the day; This is a water sensitivity index determined based on the current growth stage of rice.

[0017] Optionally, in one embodiment of this application, the irrigation and drainage decisions for the target paddy field are executed using a trained policy network, including: at the beginning of each irrigation and drainage decision cycle, collecting real-time environmental parameters of the target paddy field; constructing a state vector based on the real-time environmental parameters, inputting the state vector into a trained and converged policy network, outputting the irrigation and drainage decision quantity corresponding to the current decision cycle through the policy network, generating decision instructions based on the irrigation and drainage decision quantity to control at least one action of the irrigation device and the drainage device; at the end of the current decision cycle, collecting actual rainfall, evapotranspiration, and water depth to update the state vector for the next decision cycle, until the end of the rice growth period in the target paddy field is reached.

[0018] A second aspect of this application provides a paddy field irrigation and drainage decision-making system, comprising: an acquisition module for acquiring historical environmental parameters of a target paddy field; a construction module for constructing a paddy field growth simulation environment based on the historical environmental parameters; a transformation module for transforming the concept of the paddy field irrigation and drainage decision-making process into a Markov decision process in the paddy field growth simulation environment, and constructing a reinforcement learning environment through the Markov decision process; an update module for training the agent's policy network and value network in the paddy field growth simulation environment, updating the network parameters of the value network through the reinforcement learning environment during the training process, and updating the network parameters of the policy network through the updated value network; and an execution module for executing the irrigation and drainage decision of the target paddy field using the trained policy network.

[0019] Optionally, in one embodiment of this application, the reinforcement learning environment includes a state space, an action space, a transition function, a reward function, and a discount factor. The environmental state vector of the state space includes the forecast rainfall sequence for the next preset number of days, the current water depth, the lower limit of the suitable water depth, the upper limit of the suitable water depth, the upper limit of rainwater storage, and water depth threshold parameters related to flood tolerance. The action space is constructed based on the decision variables of the paddy field irrigation and drainage decision cycle. The reward function includes immediate rewards and round rewards. Immediate rewards are used to evaluate rainfall utilization, irrigation and drainage volume, and the impact on crop yield within a single decision cycle. Round rewards are used to penalize the number of non-zero irrigation and drainage decisions throughout the entire growth period. The transfer function expression is:

[0020] in, For the first The initial water depth of the sky, For the first The initial water depth of the sky, For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. and To determine the amount of irrigation and drainage based solely on factors other than rainfall, evapotranspiration, and seepage. The resulting changes in the water layer.

[0021] Optionally, in one embodiment of this application, the expression for the instant reward is:

[0022] in, For the first Instant rewards during the decision-making cycle. For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. For the first The decision-making cycle is related to rewards and penalties for drought or flood-induced yield reduction. For the first The decision-making cycle is related to the water consumption penalty associated with the current irrigation and drainage volume; The expression for round reward is:

[0023] in, For the total reward function, The penalty coefficient is... The total number of irrigation and drainage actions within a round. This represents the number of decision steps within a complete decision-making cycle.

[0024] Optionally, in one embodiment of this application, the update module is further configured to input the environmental state vector of the state space into the agent; input the irrigation and drainage decision quantity output by the agent into the action space; execute the irrigation and drainage decision quantity through the action space; input the execution result of the action space into the reinforcement learning environment; update the environmental state vector through the reinforcement learning environment; input the updated environmental state vector into the agent; update the network parameters of the value network through the updated environmental state vector; and update the network parameters of the policy network through the updated value network.

[0025] Optionally, in one embodiment of this application, the method further includes: before inputting the irrigation and drainage decision quantity output by the agent into the action space, the extraction module is used to extract the current water layer depth and water layer threshold parameters from the environmental state vector; and the irrigation and drainage decision quantity is corrected according to the current water layer depth and water layer threshold parameters.

[0026] Optionally, in one embodiment of this application, the paddy field growth simulation environment includes a paddy field water balance model, a rice water production function, and a rice yield reduction model due to waterlogging, wherein the expression of the paddy field water balance model is:

[0027] in, For the irrigation cycle number The initial water depth of the day; For the irrigation cycle number t The initial water depth of the day; For the first Forecast rainfall for the day; To determine the first action based on the decision-making action Daily irrigation and drainage decision-making volume; For the first Forecast of crop water requirements for the day; For the first Forecast field seepage volume for the day; The expression for the rice yield reduction model due to waterlogging is:

[0028] in, For rice yield reduction due to flooding; For the first The depth of the water layer in the field at that time; For rice plant height; To withstand deep flooding; These are empirical parameters. These are empirical parameters; The expression for the rice water production function is:

[0029] in, For the first Crop yield indicators during the decision-making cycle; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage Forecast evapotranspiration for the day; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage The potential evapotranspiration of the day; This is a water sensitivity index determined based on the current growth stage of rice.

[0030] Optionally, in one embodiment of this application, the execution module is further configured to: collect real-time environmental parameters of the target paddy field at the beginning of each irrigation and drainage decision cycle; construct a state vector based on the real-time environmental parameters; input the state vector into a trained and converged policy network; output the irrigation and drainage decision quantity corresponding to the current decision cycle through the policy network; generate decision instructions based on the irrigation and drainage decision quantity to control at least one action of the irrigation device and the drainage device; and collect the actual rainfall, evapotranspiration and water depth at the end of the current decision cycle to update the state vector of the next decision cycle until the end of the rice growth period of the target paddy field is reached.

[0031] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor executes the program to implement the paddy field irrigation and drainage decision-making method as described in the above embodiments.

[0032] A fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the paddy field irrigation and drainage decision-making method as described above.

[0033] Therefore, this application has the following beneficial effects: First, historical environmental parameters of the target paddy field are obtained to provide basic data support for model construction. Second, a paddy field growth simulation environment is constructed based on the historical environmental parameters to reproduce the dynamic changes of the paddy field under different meteorological and water conditions, thereby improving the realism and usability of the training environment. Then, the irrigation and drainage decision-making process is modeled as a Markov decision process in the paddy field growth simulation environment, and a reinforcement learning environment is constructed to transform the irrigation and drainage decisions from static optimization to a sequential decision optimization problem, thereby enhancing the long-term decision modeling capability. Next, a policy network and a value network are trained in the simulation environment. The value network provides feedback constraints for policy updates, realizing the synergistic optimization of the policy network and the value network, thereby improving the policy convergence stability and decision quality. Finally, the trained policy network is used to execute the irrigation and drainage decisions of the target paddy field, realizing the dynamic control of the irrigation and drainage process, thereby improving water resource utilization efficiency and reducing the probability of over-irrigation or under-drainage. This solves the problems of related technologies that focus on static or short-term decision analysis at local time scales, model only a single irrigation action, have limited action space, and can only achieve short-term benefit optimization.

[0034] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0035] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a paddy field irrigation and drainage decision-making method according to an embodiment of this application; Figure 2 This is an example diagram of an intelligent agent structure according to an embodiment of this application; Figure 3 A flowchart of the functional modules of the paddy field irrigation and drainage decision-making method according to an embodiment of this application; Figure 4 This is an example diagram illustrating the logical flow of the paddy field irrigation and drainage decision-making method according to an embodiment of this application; Figure 5 This is an example diagram of a paddy field irrigation and drainage decision-making device according to an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0036] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0037] The following description, with reference to the accompanying drawings, illustrates a method, system, device, and medium for making decisions regarding irrigation and drainage in paddy fields according to embodiments of this application. Addressing the problems mentioned in the background section, this application provides a method for making decisions regarding irrigation and drainage in paddy fields. In this method, firstly, historical environmental parameters of the target paddy field are acquired to provide basic data support for model construction. Secondly, a paddy field growth simulation environment is constructed based on the historical environmental parameters to reproduce the dynamic changes of the paddy field under different meteorological and water conditions, thereby improving the realism and usability of the training environment. Then, the irrigation and drainage decision-making process is modeled as a Markov decision process within the paddy field growth simulation environment, and a reinforcement learning environment is constructed to transform the irrigation and drainage decision from static optimization into a sequential decision optimization problem, thereby enhancing the long-term decision modeling capability. Next, a policy network and a value network are trained in the simulation environment. The value network provides feedback constraints for policy updates, achieving collaborative optimization between the policy network and the value network, thereby improving policy convergence stability and decision quality. Finally, the trained policy network is used to execute the irrigation and drainage decisions for the target paddy field, achieving dynamic control of the irrigation and drainage process, thereby improving water resource utilization efficiency and reducing the probability of over-irrigation or under-drainage. This solves the problems of related technologies that focus on static or short-term decision analysis at local time scales, model only a single irrigation action, have limited action space, and can only achieve short-term benefit optimization.

[0038] Specifically, Figure 1 This is a flowchart illustrating a paddy field irrigation and drainage decision-making method provided in an embodiment of this application.

[0039] like Figure 1 As shown, the paddy field irrigation and drainage decision-making method includes the following steps: In step S101, historical environmental parameters of the target paddy field are obtained.

[0040] The target paddy field is the actual paddy field that requires irrigation and drainage control; the historical environmental parameters are a set of data used to characterize the hydrological, meteorological and crop growth status of the target paddy field during historical periods, and in this application, they represent the basic input data for constructing the paddy field growth simulation environment and training the reinforcement learning environment.

[0041] It is understood that the embodiments of this application obtain historical environmental parameters of the target paddy field to provide a real data foundation for constructing a high-precision paddy field growth simulation environment, thereby accurately depicting the field water volume change pattern and crop growth process; at the same time, it provides representative training samples for reinforcement learning models, improving the stability and convergence efficiency of policy learning.

[0042] Specifically, in this application embodiment, historical environmental parameters of the target paddy field used to construct a paddy field growth simulation environment are collected or read. The historical environmental parameters are used to characterize the current water conditions, crop water consumption characteristics and future meteorological processes, and include at least: the current water depth of the paddy field, the crop growth and development stage, the crop coefficient, measured meteorological data and weather forecast data within the next preset number of days.

[0043] In step S102, a rice paddy growth simulation environment is constructed based on historical environmental parameters.

[0044] Among them, the paddy field growth simulation environment is a virtual environment constructed based on historical environmental parameters to simulate changes in paddy field moisture and crop growth process.

[0045] Understandably, constructing a rice paddy growth simulation environment based on historical environmental parameters enables the reinforcement learning training process to be carried out under controllable and repeatable simulation conditions, thereby avoiding the costs and risks of trial and error in real rice paddies. At the same time, it can simulate water volume changes and crop response processes under different meteorological conditions and growth stages, thereby improving the stability of the learned irrigation and drainage strategies.

[0046] In step S103, the concept of paddy field irrigation and drainage decision process is transformed into Markov decision process in the paddy field growth simulation environment, and a reinforcement learning environment is constructed through Markov decision process.

[0047] In this application, Markov decision process is a mathematical model used to describe sequential decision problems. It represents the abstraction of paddy field irrigation and drainage process into a decision framework consisting of states, actions, state transitions, and rewards. Reinforcement learning environment is an interactive operating environment that obtains state transition and reward feedback. In this application, it represents an interactive simulation method based on Markov decision process for training irrigation and drainage strategies.

[0048] Understandably, by formalizing the paddy field irrigation and drainage decision-making process as a Markov decision process, the original continuous time-series decision problem, which relied on empirical rules and was difficult to quantify, is transformed into a standard mathematical framework consisting of states, actions, state transition probabilities, and reward functions. This enables a unified description of paddy field moisture changes, crop growth status, and external meteorological disturbances, significantly reducing the complexity of problem modeling. Furthermore, by constructing a reinforcement learning environment, the embodiments of this application can continuously interact with the environment in a simulated paddy field growth environment, obtaining immediate and long-term reward feedback based on different irrigation and drainage actions, thereby optimizing long-term cumulative benefits rather than just optimizing single-step decisions. At the same time, this application can complete strategy training without the need for a large number of real field experiments, reducing experimental costs and risks, and improving the model's adaptability and generalization ability under different climatic conditions and different growth stages, thereby improving the overall intelligence level and decision-making accuracy of irrigation and drainage regulation.

[0049] In one embodiment of this application, the reinforcement learning environment includes a state space, an action space, a transition function, a reward function, and a discount factor. The environmental state vector of the state space includes the forecast rainfall sequence for a preset number of days in the future, the current water depth, the lower limit of the suitable water depth, the upper limit of the suitable water depth, the upper limit of rainwater storage, and water depth threshold parameters related to flood tolerance. The action space is constructed based on the decision variables of the paddy field irrigation and drainage decision cycle. The reward function includes immediate rewards and round rewards. Immediate rewards are used to evaluate rainfall utilization, irrigation and drainage volume, and the impact on crop yield within a single decision cycle. Round rewards are used to penalize the number of non-zero irrigation and drainage decisions throughout the entire growth period. The transfer function expression is:

[0050] in, For the first The initial water depth of the sky, For the first The initial water depth of the sky, For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. and To determine the amount of irrigation and drainage based solely on factors other than rainfall, evapotranspiration, and seepage. The resulting changes in the water layer.

[0051] In this application, the state space is a mathematical representation of the current state of the environment, specifically a set of environmental state vectors composed of the predicted rainfall sequence within a preset number of days, the current water depth, the lower limit of the suitable water depth, the upper limit of the suitable water depth, the upper limit of rainwater storage, and water layer threshold parameters related to flood-resistant water depth. The action space is the set of all decision actions that the agent can execute in a given state, specifically a set of irrigation or drainage operations constructed based on the decision variables of the paddy field irrigation and drainage decision cycle. The transition function is a mapping function describing the change of state with actions, specifically representing the root... The state evolution rule for updating the paddy field water layer state is based on current irrigation and drainage actions and external rainfall input; the reward function is a feedback function used to evaluate the quality of actions, and in this application, it represents an evaluation mechanism composed of immediate rewards and round rewards. The immediate reward is used to assess the rainfall utilization rate, irrigation and drainage volume, and impact on crop yield within a single decision cycle, while the round reward is used to penalize the number of non-zero irrigation and drainage decisions within the entire growth period; the discount factor is a weighting coefficient used to measure the importance of future rewards relative to current rewards, and in this application, it represents a decay weighting parameter used to balance short-term and long-term benefits.

[0052] It is understandable that by standardizing the modeling of the reinforcement learning environment, the paddy field irrigation and drainage decision-making problem is systematically divided into five core components: state space, action space, transition function, reward function, and discount factor. This gives the complex agricultural water resource regulation problem a unified and computable mathematical expression. By introducing a state space that includes future forecast rainfall sequences and water layer threshold parameters, the environment can fully characterize rainfall forecast information and field water constraints. By constructing an action space based on the decision cycle, the irrigation and drainage behaviors in this embodiment can be expressed in a unified discrete decision form, reducing control complexity. By setting a reward function composed of immediate and round rewards, it is possible to optimize the effect of single-step water regulation and constrain unnecessary irrigation and drainage behaviors throughout the entire growth period, thereby achieving synergy between local and global optimization. At the same time, by combining the discount factor to weight future returns, the strategy learning can take into account both short-term response and long-term yield returns, ultimately improving the stability, economy, and adaptability of irrigation and drainage decisions.

[0053] In this embodiment of the application, the process of constructing the state space and action space in the process of transforming the concept of paddy field irrigation and drainage decision-making into Markov decision-making is as follows: First, construct the environment state vector. The environmental state vector includes at least the forecast rainfall sequence for the next preset number of days and the current water depth. Lower limit of suitable water depth Suitable upper limit of water depth , upper limit of rain storage and flood resistance depth The relevant water layer threshold parameters enable this application to simultaneously perceive future rainfall scenarios, current moisture conditions, and safe water layer control boundaries when making decisions. The water layer threshold parameters can be set according to different rice cultivation types and growth stages.

[0054] In one embodiment of this application, the reinforcement learning environment can be represented as:

[0055] Where S is the state space, A is the action space, P is the transition function, R is the reward function, and γ is the discount factor.

[0056] In one embodiment of this application, the environment state vector can be represented as:

[0057] in, For the current irrigation decision cycle The cumulative forecast rainfall for the next 1-7 days in the region. For the current irrigation decision cycle The depth of the water layer on the field surface, To suit the lower limit of water depth, To suit the upper limit of water depth, This is the upper limit of rainwater storage. To ensure rice can withstand flooding, the aforementioned water depth threshold parameters can be preset based on local rice varieties and water management standards for different growth stages.

[0058] Secondly, the decision variables for the paddy field irrigation and drainage decision cycle are defined as discrete water quantity options, constituting the action space. Define continuous actions. ∈[ , As a decision-making factor for irrigation and drainage, among which >0 indicates the amount of irrigation water. <0 indicates the amount of water discharged. =0 indicates no irrigation and no drainage; the range of irrigation and drainage values ​​can be pre-defined within the allowable range based on local water conservancy project conditions and field characteristics to ensure the physical feasibility of the action and facilitate project implementation.

[0059] Suppose that when using a proximal policy optimization algorithm to process discrete actions, the action space can be defined as:

[0060] Positive values ​​represent irrigation, negative values ​​represent drainage, and zero values ​​represent neither irrigation nor drainage. This represents discrete irrigation and drainage volumes. The thickness gradually increases from 100 mm to 100 mm, with each increment being 10 mm.

[0061] In one embodiment of this application, the expression for the instant reward is:

[0062] in, For the first Instant rewards during the decision-making cycle. For the first The basic reward of the decision-making cycle. For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. For the first The decision-making cycle is related to rewards and penalties for drought or flood-induced yield reduction. For the first The decision-making cycle is related to the water consumption penalty associated with the current irrigation and drainage volume; The expression for round reward is:

[0063] in, For the total reward function, For the first Instant rewards during the decision-making cycle. The penalty coefficient is... The total number of irrigation and drainage actions within a round. This represents the number of decision steps within a complete decision-making cycle.

[0064] It is understood that the embodiments of this application, by constructing an immediate reward in a product form, can achieve joint constraints on multiple factors such as rainfall utilization rate, risk of reduced production due to drought or flooding, and irrigation and drainage water consumption, so that a synergistic adjustment relationship is formed between the sub-reward items; by constructing a round reward, the single-step decision benefit and the cumulative performance throughout the entire growth period are uniformly optimized, and a penalty term for the total number of irrigation and drainage actions is introduced to effectively suppress frequent ineffective operations.

[0065] In this embodiment of the application, the process of constructing the reward function in the process of transforming the concept of paddy field irrigation and drainage decision-making into Markov decision-making is as follows: Specifically, in this embodiment, the reward function is divided into two parts: immediate reward and round reward. The immediate reward is used to evaluate the decision quality within a single decision cycle, while the round reward is used to measure the overall performance of irrigation and drainage behavior throughout the entire growth period.

[0066] For example, basic reward It can be given by the following formula:

[0067] For example, rainfall utilization reward The calculation can be performed based on the utilization of rainfall within a preset number of days (7 days in this example). The expression is:

[0068] in, In order to make decisions during the irrigation cycle The cumulative actual crop evapotranspiration for the next 7 days after the irrigation and drainage decision is implemented, based on the default management model. In order to make decisions during the irrigation cycle The cumulative deep seepage over the next 7 days following the implementation of irrigation and drainage decisions, under the default management model. In the irrigation decision cycle The cumulative rainfall over the next 7 days after the irrigation and drainage decision is implemented, according to the default management model.

[0069] Specifically, if no drainage incident occurs within the next 7 days, then A value close to 1 indicates that rainfall is being fully utilized; if numerous drainage events cause some rainfall to be discharged, then... Reduce rainfall as punishment for wasted rain.

[0070] Production bonus This is used to reflect the impact of the current decision on rice yield, taking into account both drought-induced yield reduction and flood-induced yield reduction scenarios. These scenarios will be explained in subsequent embodiments in conjunction with irrigation and drainage decision scenarios.

[0071] To balance the relationship between irrigation and drainage volume and the number of irrigation and drainage decisions, this application introduces a round reward to constrain the number of non-zero irrigation and drainage decisions within a complete growth period. The round reward can be expressed as:

[0072] in, For the total reward function, The number of decision steps within a complete decision-making cycle. For the first Instant rewards for each step This is the penalty coefficient, which can be chosen based on experience, for example, 5; This refers to the total number of non-zero actions (i.e., irrigation / drainage) within a round. By introducing round rewards, the embodiments of this application can effectively suppress the tendency of agents to obtain high immediate rewards through "multiple small-volume irrigation and drainage," achieving a comprehensive balance between water consumption, irrigation and drainage frequency, and yield loss on a full-lifecycle scale.

[0073] In step S104, the agent's policy network and value network are trained in a rice paddy growth simulation environment. During the training process, the network parameters of the value network are updated through a reinforcement learning environment, and the network parameters of the policy network are updated through the updated value network.

[0074] In this application, the agent is the entity that interacts with the environment and makes decisions during reinforcement learning. It represents a decision-making learning model composed of a policy network and a value network. The policy network generates irrigation or drainage actions, the value network evaluates the long-term reward of a state or state-action relationship, and the policy network is optimized and updated based on feedback from the value network. This includes, but is not limited to, proximal policy optimization algorithms. The policy network is a function model that outputs the probability of each action chosen by the agent in a given state. In this application, it represents a neural network that generates the probability distribution of irrigation or drainage decisions based on the input of the paddy field environment state. The value network is a function model that evaluates the expected long-term cumulative reward of the current state or state-action relationship. In this application, it represents a neural network that predicts the long-term returns of paddy field irrigation and drainage decisions throughout the entire growing season.

[0075] It is understandable that by introducing a collaborative training mechanism between the policy network and the value network in the rice paddy growth simulation environment, the reinforcement learning process can achieve closed-loop optimization of "decision generation - effect evaluation - parameter update". Among them, the value network estimates the long-term reward of the state or state-action, so that the policy update no longer depends on single-step immediate feedback, and enhances the optimization of long-term yield and water resource utilization efficiency. By continuously providing state transition and reward feedback through the reinforcement learning environment, the value network parameters are continuously corrected and gradually approach the payoff distribution in the real environment, thereby further improving the accuracy of policy network updates.

[0076] In one embodiment of this application, before inputting the irrigation and drainage decision quantity output by the intelligent agent into the action space, the method further includes: extracting the current water layer depth and water layer threshold parameters from the environmental state vector; and correcting the irrigation and drainage decision quantity based on the current water layer depth and water layer threshold parameters.

[0077] It is understood that, before the agent outputs irrigation and drainage decision quantities and inputs them into the action space, the embodiments of this application introduce a correction mechanism based on the current water layer state and water layer threshold parameters, so that the decision results can be dynamically matched with the actual hydrological constraints in the field, avoiding unreasonable irrigation or drainage actions generated by the strategy network in extreme or boundary states; in particular, by using the current water layer depth and threshold parameters such as the upper and lower limits of suitable water layer depth, the upper limit of rainwater storage, and the flood-resistant water depth to constrain and correct the original decision quantities, it can effectively ensure that the action output is always within the physically feasible range.

[0078] Specifically, the transfer function Used to describe the quantity of irrigation and drainage decisions made. The evolution process of the state in the subsequent embodiments of this application can be used to define the ideal water layer change relationship, i.e., the transfer function expression:

[0079] in, For the irrigation cycle number The initial water depth of the sky, For the irrigation cycle number The initial water depth of the sky, For the first The amount of irrigation and drainage decisions per day and To determine the amount of irrigation and drainage based solely on factors other than rainfall, evapotranspiration, and seepage. The resulting changes in the water layer.

[0080] Furthermore, during the training period of this application embodiment, in order to avoid the agent from making actions that significantly contradict traditional irrigation and drainage experience, the irrigation and drainage decision quantities output by the executing agent are... At that time, it can be based on the current environment state vector The actions are corrected to obtain the corrected irrigation and drainage decision quantities. Its expression is:

[0081] Furthermore, embodiments of this application obtain corrected irrigation and drainage decision quantities. Subsequently, based on actual meteorological factors and crop water consumption processes, the field water depth is updated to obtain the actual water depth for the next decision-making cycle. The expression is:

[0082] in, Irrigation cycle The actual water depth of the sky, Irrigation cycle The actual water depth of the sky, Irrigation cycle The actual rainfall within the day, Irrigation cycle Actual crop evapotranspiration within a day Irrigation cycle The amount of deep seepage within a day, when a water layer exists in the field. An empirical constant can be taken, which can be approximated as 0 when there is no water layer in the field.

[0083] It can be calculated using the following formula:

[0084] in, This refers to crop coefficients related to the growth stage. This is the soil moisture stress coefficient. For reference crop evapotranspiration, It can be based on the water deficit in the root zone. Total effective water volume and easily accessible water volume Perform segmented calculations:

[0085]

[0086]

[0087] in, Field holding capacity The wilting coefficient, The depth of the root layer. The water content is designed for easy utilization.

[0088] Furthermore, the environmental state vector for the next irrigation and drainage decision cycle. Based on the updated water depth The rainfall forecast sequence after forward shift, as well as the updated water layer threshold and reproductive stage information, are reconstructed.

[0089] In one embodiment of this application, the rice paddy growth simulation environment includes a rice paddy water balance model, a rice water production function, and a rice flooding-induced yield reduction model, wherein, The expression for the paddy field water balance model is:

[0090] in, For the irrigation cycle number The initial water depth of the day; For the irrigation cycle number t The initial water depth of the day; For the first Forecast rainfall for the day; To determine the first action based on the decision-making action Daily irrigation and drainage decision-making volume; For the first Forecast of crop water requirements for the day; For the first Forecast field seepage volume for the day; The expression for the rice yield reduction model due to waterlogging is:

[0091] in, For the function of rice yield reduction due to waterlogging, For the first The depth of the water layer in the fields at that time For rice plant height, To withstand deep floodwaters, These are empirical parameters. These are empirical parameters; The expression for the rice water production function is:

[0092] in, For the first Crop yield indicators during the decision-making cycle; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage Forecast evapotranspiration for the day; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage The potential evapotranspiration of the day; This is a water sensitivity index determined based on the current growth stage of rice.

[0093] Among them, the paddy field water balance model is a mathematical model used to describe the relationship between paddy field water depth and time. In this application, it represents the state update equation for calculating the dynamic changes of paddy field water layer based on water expenditure factors such as rainfall, irrigation drainage, crop evapotranspiration, and field seepage. The rice flooding yield reduction model is a functional model used to describe the relationship between field flooding depth and rice yield. In this application, it represents a yield reduction function that quantitatively assesses yield loss caused by flooding based on the relationship between water layer depth and rice plant height and flood tolerance depth. The rice water production function is a functional model used to describe the relationship between crop evapotranspiration level and yield formation. In this application, it represents a yield evaluation function that comprehensively characterizes the degree of water stress and yield impact at different growth stages based on the ratio of actual evapotranspiration to potential evapotranspiration.

[0094] It is understandable that by constructing a paddy field water balance model, the embodiments of this application can quantitatively describe the dynamic evolution of paddy field water depth under the influence of multiple factors such as rainfall, irrigation and drainage, crop evapotranspiration, and field seepage. By introducing a rice flood-induced yield reduction model, the yield loss caused by water depth exceeding the flood tolerance threshold is explicitly quantified, enabling the decision-making process to directly perceive flood risk and form a constrained optimization mechanism. Furthermore, by constructing a rice water production function, the difference between actual and potential evapotranspiration at different growth stages is transformed into yield evaluation indicators, enabling this application to characterize the stage-specific impact of water stress on crop growth and achieve synergistic optimization between water use efficiency and yield formation.

[0095] In the context of irrigation decision-making, this embodiment of the application quantifies the impact of water deficit on yield. An improved Jensen water production function model can be used to establish the relationship between water deficit and yield. Therefore, the expression for the rice water production function in this embodiment of the application is as follows:

[0096] in, For the first Crop yield indicators during the decision-making cycle; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage Forecast evapotranspiration for the day; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage The potential evapotranspiration of the day; The water sensitivity index is determined based on the current growth stage of rice; Based on rice Moisture sensitivity index determined for each reproductive stage.

[0097] Specifically, in the irrigation decision scenario, the first Crop yield indicators during the decision-making cycle A result below 1 indicates that the current irrigation strategy may cause crops to suffer a certain degree of water stress in later growth stages, thus adversely affecting the final yield; conversely, if the result is above 1, it indicates that the current irrigation strategy may cause crops to suffer a certain degree of water stress in later growth stages, thus adversely affecting the final yield. A value of 1 indicates that the decision effectively maintains the crop's water requirements and does not pose a significant threat to yield.

[0098] In the context of drainage decision-making, to reflect the risk of "damage and reduced yield due to heavy rain caused by failure to pre-drain in time," this application uses a rice flood-induced yield reduction model to represent the impact of water depth and water accumulation duration on yield. The expression for the rice flood-induced yield reduction model is as follows:

[0099]

[0100] in, For the function of rice yield reduction due to waterlogging, Let t be the field water depth. For rice plant height, To withstand deep floodwaters, For the number of days since transplanting, These are empirical parameters. These are empirical parameters, and their values ​​are affected by the rice growth stage, as shown in Table 1, which presents the values ​​for different growth stages. The value of .

[0101] Table 1

[0102] Furthermore, in the drainage decision-making scenario Expressed by the following formula:

[0103] in, For the drainage decision scenario, the first Crop yield indicators during the decision-making cycle. In order to make decisions during the irrigation cycle The cumulative rice yield reduction due to flooding in the 7 days following the implementation of the drainage decision. Depending on whether the current situation is irrigation or drainage, the appropriate option can be selected. or ;like A value less than 1 indicates that the crop will suffer from waterlogging in the near future, resulting in yield loss; otherwise... A value of 1 indicates that the drainage decision has no adverse effect on crop yield.

[0104] To limit the amount of water needed for a single irrigation or drainage operation, encourage water conservation, and avoid extreme practices, embodiments of this application, exemplarily, reward irrigation or drainage volume. Defined by the following formula:

[0105] in, The preset maximum irrigation or drainage volume (100 mm is acceptable). The larger | is, the more The smaller the value, the better, thus guiding the agent to select the smallest possible value while still meeting the crop's water requirements.

[0106] In one embodiment of this application, updating the network parameters of the value network through a reinforcement learning environment includes: inputting the environmental state vector of the state space into the agent; inputting the irrigation and drainage decision quantity output by the agent into the action space; executing the irrigation and drainage decision quantity through the action space; inputting the execution result of the action space into the reinforcement learning environment; updating the environmental state vector through the reinforcement learning environment; inputting the updated environmental state vector into the agent; updating the network parameters of the value network through the updated environmental state vector; and updating the network parameters of the policy network through the updated value network.

[0107] It is understood that by inputting the environmental state vector into the agent and driving the policy network to generate irrigation and drainage decision quantities, the embodiments of this application can make adaptive decisions based on real-time hydrological information, thereby improving the responsiveness to complex weather and field changes. By executing decisions in the action space and receiving updated state information from the reinforcement learning environment, the embodiments of this application can realistically depict the impact of decisions on water layer evolution, thereby enhancing the causal consistency of environmental modeling. Furthermore, by iteratively updating the value network parameters using the updated environmental state vector, the embodiments of this application can improve the accuracy of value estimation. At the same time, by back-optimizing the policy network parameters based on the updated value network, the policy network can obtain more accurate long-term benefit guidance, thereby improving the global optimization capability of irrigation and drainage strategies and their generalization performance under different growth stages and climatic conditions.

[0108] The intelligent agent structure in this application embodiment is as follows: Figure 2 As shown, specifically, the embodiments of this application employ a proximal policy optimization algorithm based on a policy-value architecture to implement policy learning: The policy network uses the environment state vector Given the input, output the probability distribution of actions. ,in, Here are the policy parameters. The policy network updates its parameters using the policy gradient method, aiming to maximize the expected cumulative reward. The corresponding update form can be written as:

[0109] in, For strategy parameters; The gradient of the objective function with respect to the parameters; The objective function is... For expectations; This represents the probability distribution of actions. The advantage function measures the merits of a particular action relative to the average decision.

[0110] The value network uses the environment state vector s t Given input, output state value estimate ,in, Let be the value parameters. The value network can be trained by minimizing the time difference error, and its loss function is:

[0111] in, The loss function of the value network, This is the discount factor (e.g., 0.99). For state value estimation.

[0112] Furthermore, to reduce the variance of policy updates and balance bias and stability, embodiments of this application introduce a generalized dominance estimation method to calculate the dominance function, which is defined as follows:

[0113] in, For a moment t The timing difference error; The discount factor (e.g., 0.99); The generalized dominance estimation coefficient (e.g., 0.90) is used to balance the estimation bias and variance.

[0114] Furthermore, to prevent excessively large policy update steps from causing training instability, an agent pruning objective function is introduced during policy optimization. Specifically, the agent reuses samples through importance sampling, defining the old policy as... The current strategy is The importance sampling ratio can then be expressed as:

[0115] Furthermore, the objective function of the pruning strategy is defined as follows:

[0116] in, Used to Limited to [ Within the interval, This is the clipping factor (for example, a value of 0.2).

[0117] Furthermore, to encourage a degree of exploratory nature in the strategy, an entropy regularization term is introduced in one embodiment, and the entropy reward can be expressed as:

[0118] in, The entropy function of the policy distribution is used to encourage higher uncertainty in policy output, thereby avoiding early entrapment in local optima.

[0119] Furthermore, the mean squared error loss of the value network can be expressed as:

[0120] in, The target value is calculated using the generalized advantage estimation method.

[0121] Furthermore, combining the above-mentioned policy loss, value loss, and entropy reward to form the agent's total loss function, the expression of which is:

[0122] in, The weighting coefficients for the loss of the value function (e.g., a value of 1 is acceptable); This is the coefficient of the entropy regularization term (for example, its value range is 0.01-0.05).

[0123] Furthermore, during training, the agent interacts in multiple rounds within a simulated rice paddy irrigation and drainage environment, with each sampling yielding a state transition sample. Calculate the corresponding advantage function With target value Then, the total loss function is processed in batch form. Perform gradient ascent and gradient descent operations to update the policy network parameters respectively. With value network parameters This continues until the accumulated rewards converge.

[0124] Preferably, to prevent gradient explosion or training oscillations, the gradient norms of the policy network and the value network are clipped to ensure... ,in The maximum gradient norm is used (typically between 0.5 and 1.0). Finally, after multiple rounds of training iterations, a converged irrigation and drainage strategy model is obtained, which is used to output the optimal continuous irrigation and drainage decisions in actual paddy field management.

[0125] In step S105, the trained policy network is used to make irrigation and drainage decisions for the target paddy field.

[0126] Understandably, this application can automatically generate irrigation and drainage control quantities based on the real-time environmental conditions of the target paddy field, thereby improving the automation level and response speed of decision-making. At the same time, since the policy network has been fully trained in the paddy field growth simulation environment and reinforcement learning framework, its decision results can comprehensively consider the influence of multiple factors such as rainfall, evapotranspiration and water layer threshold constraints, which helps to improve water resource utilization efficiency, reduce unnecessary irrigation and drainage times, and reduce the risk of drought and flooding.

[0127] In one embodiment of this application, the irrigation and drainage decisions for a target paddy field are executed using a trained policy network, including: at the beginning of each irrigation and drainage decision cycle, real-time environmental parameters of the target paddy field are collected; a state vector is constructed based on the real-time environmental parameters, the state vector is input into the trained and converged policy network, the policy network outputs the irrigation and drainage decision quantity corresponding to the current decision cycle, and a decision instruction is generated based on the irrigation and drainage decision quantity to control at least one action of the irrigation device and the drainage device; at the end of the current decision cycle, the actual rainfall, evapotranspiration and water depth are collected to update the state vector for the next decision cycle, until the end of the rice growth period in the target paddy field is reached.

[0128] Among them, real-time environmental parameters are the instantaneous observation data of the target paddy field within the current decision-making cycle, and in this application, they represent real-time collected data used to describe field hydrological and meteorological conditions; decision instructions are control signals used to drive the action of the execution device, and in this application, they represent the set of operation instructions for controlling the irrigation or drainage device obtained by converting the irrigation and drainage decision quantities; irrigation device is an execution device used to replenish water to the paddy field, and in this application, it represents a water supply facility system used to perform irrigation actions; drainage device is an execution device used to drain excess water from the paddy field, and in this application, it represents a drainage control facility system used to perform drainage actions; the end point of the rice growth period is the time node at which the rice growth cycle ends, and in this application, it represents the time limit for the termination of the irrigation and drainage decision execution process.

[0129] It is understood that, by establishing a closed-loop execution mechanism based on real-time state acquisition and policy network inference, the embodiments of this application enable the trained policy network to achieve continuous automated irrigation and drainage control in the target paddy field; by periodically updating the actual rainfall, evapotranspiration and water depth information, the embodiments of this application can continuously reflect changes in the field state, thereby improving the real-time nature of decision-making, adaptability and overall water regulation effect.

[0130] During actual operation, at the beginning of each irrigation and drainage decision cycle, an environmental state vector is constructed by collecting real-time environmental parameters. Input it into the trained and converged policy network, and it outputs the irrigation and drainage decision quantity corresponding to the current decision cycle. The irrigation and drainage decision quantities are converted into decision instructions for irrigation devices or drainage devices, and at least one action of the irrigation device and drainage device is controlled. Before executing the irrigation and drainage decision quantities corresponding to the current decision cycle, the irrigation and drainage decision quantities can be modified based on the current water layer depth and water layer threshold parameters to meet the appropriate water layer constraints and rainwater storage upper limit constraints.

[0131] At the end of the current decision cycle, the water layer state and related variables are updated based on parameters such as actual rainfall, evapotranspiration, and water layer depth, and the state vector for the next decision cycle is updated. ; If the rice growth period of the target paddy field has not yet reached its end, the above process will continue in the next decision cycle. When the field is judged to be in the drying period based on the crop growth stage, the irrigation and drainage actions will be suspended, and the water layer status will only be updated based on the natural drying process. The irrigation and drainage decisions will be resumed after the drying period ends. The entire growth period irrigation and drainage strategy sequence will be output until the rice growth period ends.

[0132] The paddy field irrigation and drainage decision-making method of this application embodiment can be divided into the following modules according to their functions, such as... Figure 3 As shown: 1. Data acquisition module, used to collect environmental parameters of the target paddy field in the current decision-making period and obtain weather forecast data for the next preset number of days; environmental parameters include the current water depth of the paddy field, crop coefficient, measured meteorological data, and information such as rainfall, temperature, and evapotranspiration in the future forecast period.

[0133] 2. Environmental modeling module, used to construct paddy field water balance model, rice water production function and rice waterlogging reduction model based on collected environmental parameters, forming a paddy field growth simulation environment that can simulate the evolution of field water layer, evapotranspiration process and yield response.

[0134] 3. The policy learning module is used in the offline training phase to use the rice paddy growth simulation environment as the reinforcement learning environment. It uses a proximal policy optimization algorithm based on the policy-value architecture to train the policy network and the value network, optimize the policy parameters and value function parameters, and obtain a converged policy network.

[0135] The strategy learning module may include: A policy network is used to output a continuous action distribution based on the input state vector and sample irrigation and drainage decisions; Value networks are used to estimate the value function in the current state and calculate the advantage function. The loss calculation unit is used to construct the total loss function based on the pruning policy loss, value loss, and entropy regularization term, and to update the parameters of the policy network and value network through gradients to ensure the stability of training and the generalization ability of the policy.

[0136] 4. The online decision-making module is used to receive the environmental state vector of the current decision cycle during the actual operation phase, input it into the strategy model obtained through offline training, and output the corresponding continuous irrigation and drainage decision.

[0137] 5. Control execution module, used to convert decision quantities into control commands for irrigation or drainage devices and execute them.

[0138] 6. The state update and cycle determination module is used to update the initial state vector of the next decision cycle after each decision cycle ends, based on the measured rainfall, evapotranspiration, water depth and other parameters, the water balance relationship and the transfer function, and determine whether the end of the rice growth period or the drying period has been reached.

[0139] The cycle determination module can determine whether the cycle has ended by monitoring the rice growth period. When the decision cycle reaches the end of the rice growth period, the cycle is terminated. When the rice is detected to have entered the drying period, irrigation and drainage operations are suspended. After the drying period ends, the next decision process is restarted, realizing the coordinated coupling of intelligent decision-making and agronomic processes.

[0140] 7. The results output module is used to output the optimal irrigation and drainage strategy sequence for the entire growth period of the target paddy field, and generate corresponding results data such as irrigation water consumption, rainfall utilization rate and yield indicators, so as to provide a reference for agricultural water resource scheduling and management.

[0141] In summary, the specific process of the paddy field irrigation and drainage decision-making method in this application embodiment is as follows: Figure 4 As shown: Logically, it can be divided into an offline training phase and an online decision-making phase: In the offline training phase, historical environmental state parameters and weather forecast data for the complete decision-making cycle of the target paddy field are first acquired. These historical environmental state parameters include field water depth, crop growth stage, crop coefficient, and measured meteorological data. Based on this data, the irrigation and drainage decision-making process is formalized as a Markov decision process within the paddy field growth simulation environment. The state space, action space, state transition function, and reward function are defined to achieve mathematical modeling of the paddy field hydrological regulation problem. On this basis, a paddy field water balance model, a rice water production function, and a rice flood-induced yield reduction model are constructed to characterize water layer change patterns, crop water use characteristics, and the impact of flooding on yield, forming a complete paddy field growth simulation environment.

[0142] Furthermore, the simulated rice paddy growth environment is used as a reinforcement learning environment. The agent learns policies through cyclical interactions involving states, actions, rewards, and state updates, enabling it to gradually learn collaborative decision-making strategies for irrigation and drainage under different weather conditions. During policy optimization, a proximal policy optimization algorithm is used to train the policy network and value network offline to maximize the cumulative reward throughout the rice growth period. When the convergence condition is met, training ends and the process enters the online decision-making phase.

[0143] During the online decision-making phase, the current environmental state parameters are first acquired and a state vector is constructed. This vector is then input into the converged policy network, outputting the irrigation or drainage decision for the current decision cycle. This decision is then corrected, and the irrigation or drainage operation is executed. If a field drying period exists, it is determined whether the field is currently in this phase. If so, natural drying is implemented; otherwise, intelligent decision-making is executed. At the end of each decision cycle, measured data such as actual rainfall, evapotranspiration, and field water depth are collected, and the initial state vector for the next decision cycle is updated based on the water balance relationship and the state transition function.

[0144] Finally, determine whether the rice growth period has ended. If it has not ended, repeat the above online decision-making process. If it has ended, output the irrigation and drainage execution strategy sequence corresponding to each decision cycle throughout the entire growth period.

[0145] According to the paddy field irrigation and drainage decision-making method proposed in this application, firstly, historical environmental parameters of the target paddy field are obtained to provide basic data support for model construction; secondly, a paddy field growth simulation environment is constructed based on the historical environmental parameters to reproduce the dynamic changes of the paddy field under different meteorological and water conditions, thereby improving the realism and usability of the training environment; then, the irrigation and drainage decision-making process is modeled as a Markov decision process in the paddy field growth simulation environment, and a reinforcement learning environment is constructed to transform the irrigation and drainage decision from static optimization to a sequential decision optimization problem, thereby enhancing the long-term decision modeling capability; then, a policy network and a value network are trained in the simulation environment, and the value network provides feedback constraints for policy updates, realizing the synergistic optimization of the policy network and the value network, thereby improving the policy convergence stability and decision quality; finally, the trained policy network is used to execute the irrigation and drainage decisions of the target paddy field, realizing the dynamic control of the irrigation and drainage process, thereby improving water resource utilization efficiency and reducing the probability of over-irrigation or under-drainage. This solves the problems of related technologies that focus on static or short-term decision analysis at local time scales, model only a single irrigation action, have limited action space, and can only achieve short-term benefit optimization.

[0146] Next, the paddy field irrigation and drainage decision system proposed according to the embodiments of this application is described with reference to the accompanying drawings.

[0147] Figure 5 This is a block diagram of a paddy field irrigation and drainage decision system according to an embodiment of this application.

[0148] like Figure 5 As shown, the paddy field irrigation and drainage decision system 10 includes: an acquisition module 100, a construction module 200, a conversion module 300, an update module 400, and an execution module 500.

[0149] The system comprises: an acquisition module 100 for acquiring historical environmental parameters of the target paddy field; a construction module 200 for constructing a paddy field growth simulation environment based on the historical environmental parameters; a transformation module 300 for transforming the concept of paddy field irrigation and drainage decision-making process into a Markov decision process within the paddy field growth simulation environment, and constructing a reinforcement learning environment through the Markov decision process; an update module 400 for training the agent's policy network and value network within the paddy field growth simulation environment, updating the network parameters of the value network through the reinforcement learning environment during training, and updating the network parameters of the policy network through the updated value network; and an execution module 500 for executing the irrigation and drainage decisions of the target paddy field using the trained policy network.

[0150] In one embodiment of this application, the reinforcement learning environment includes a state space, an action space, a transition function, a reward function, and a discount factor. The environmental state vector of the state space includes the forecast rainfall sequence for a preset number of days in the future, the current water depth, the lower limit of the suitable water depth, the upper limit of the suitable water depth, the upper limit of rainwater storage, and water depth threshold parameters related to flood tolerance. The action space is constructed based on the decision variables of the paddy field irrigation and drainage decision cycle. The reward function includes immediate rewards and round rewards. Immediate rewards are used to evaluate rainfall utilization, irrigation and drainage volume, and the impact on crop yield within a single decision cycle. Round rewards are used to penalize the number of non-zero irrigation and drainage decisions throughout the entire growth period. The transfer function expression is: in, Irrigation cycle The actual water depth of the sky, Irrigation cycle The actual rainfall within the day, This is an irrigation and drainage decision based on the current water layer depth and water layer threshold parameters. Irrigation cycle Actual crop evapotranspiration within a day Irrigation cycle The amount of deep seepage within a day, when a water layer exists in the field. An empirical constant can be taken, which can be approximated as 0 when there is no water layer in the field.

[0151] in, For the first The initial water depth of the sky, For the first The initial water depth of the sky, For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. and To determine the amount of irrigation and drainage based solely on factors other than rainfall, evapotranspiration, and seepage. The resulting changes in the water layer.

[0152] In one embodiment of this application, the expression for the instant reward is:

[0153] in, For the first Instant rewards during the decision-making cycle. For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. For the first The decision-making cycle is related to rewards and penalties for drought or flood-induced yield reduction. For the first The decision-making cycle is related to the water consumption penalty associated with the current irrigation and drainage volume; The expression for round reward is:

[0154] in, For the total reward function, The penalty coefficient is... The total number of irrigation and drainage actions within a round. This represents the number of decision steps within a complete decision-making cycle.

[0155] In one embodiment of this application, the update module 400 is further configured to input the environmental state vector of the state space into the agent; input the irrigation and drainage decision quantity output by the agent into the action space; execute the irrigation and drainage decision quantity through the action space; input the execution result of the action space into the reinforcement learning environment; update the environmental state vector through the reinforcement learning environment; input the updated environmental state vector into the agent; update the network parameters of the value network through the updated environmental state vector; and update the network parameters of the policy network through the updated value network.

[0156] In one embodiment of this application, the method further includes: an extraction module for extracting the current water depth and water layer threshold parameters from the environmental state vector before inputting the irrigation and drainage decision quantity output by the agent into the action space; and correcting the irrigation and drainage decision quantity based on the current water depth and water layer threshold parameters.

[0157] In one embodiment of this application, the paddy field growth simulation environment includes a paddy field water balance model, a rice water production function, and a rice yield reduction model due to waterlogging. The expression for the paddy field water balance model is as follows:

[0158] in, For the irrigation cycle number The initial water depth of the day; For the irrigation cycle numbert The initial water depth of the day; For the first Forecast rainfall for the day; To determine the first action based on the decision-making action Daily irrigation and drainage decision-making volume; For the first Forecast of crop water requirements for the day; For the first Forecast field seepage volume for the day; The expression for the rice yield reduction model due to waterlogging is:

[0159] in, For rice yield reduction due to flooding; For the first The depth of the water layer in the field at that time; For rice plant height; To withstand deep flooding; These are empirical parameters. These are empirical parameters; The expression for the rice water production function is:

[0160] in, For the first Crop yield indicators during the decision-making cycle; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage Forecast evapotranspiration for the day; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage The potential evapotranspiration of the day; This is a water sensitivity index determined based on the current growth stage of rice.

[0161] In one embodiment of this application, the execution module 500 is further configured to: collect real-time environmental parameters of the target paddy field at the beginning of each irrigation and drainage decision cycle; construct a state vector based on the real-time environmental parameters; input the state vector into a trained and converged policy network; output the irrigation and drainage decision quantity corresponding to the current decision cycle through the policy network; generate decision instructions based on the irrigation and drainage decision quantity to control at least one action of the irrigation device and the drainage device; and collect the actual rainfall, evapotranspiration and water depth at the end of the current decision cycle to update the state vector of the next decision cycle until the end of the rice growth period of the target paddy field is reached.

[0162] One embodiment of this application can be deployed on an intelligent agricultural control terminal with a data acquisition interface and a control port, or it can be integrated into a cloud-based farmland water management platform. It can collect meteorological and hydrological data through a wireless sensor network and call the trained strategy model in real time to make optimal irrigation and drainage decisions.

[0163] It should be noted that the foregoing explanation of the implementation method for paddy field irrigation and drainage decision-making also applies to the paddy field irrigation and drainage decision-making system of this embodiment, and will not be repeated here.

[0164] According to the paddy field irrigation and drainage decision-making system proposed in this application, firstly, historical environmental parameters of the target paddy field are acquired to provide basic data support for model construction; secondly, a paddy field growth simulation environment is constructed based on the historical environmental parameters to reproduce the dynamic changes of the paddy field under different meteorological and water conditions, thereby improving the realism and usability of the training environment; then, the irrigation and drainage decision-making process is modeled as a Markov decision process in the paddy field growth simulation environment, and a reinforcement learning environment is constructed to transform the irrigation and drainage decision from static optimization to a sequential decision optimization problem, thereby enhancing the long-term decision modeling capability; then, a policy network and a value network are trained in the simulation environment, and the value network provides feedback constraints for policy updates, realizing the synergistic optimization of the policy network and the value network, thereby improving the policy convergence stability and decision quality; finally, the trained policy network is used to execute the irrigation and drainage decisions of the target paddy field, realizing the dynamic control of the irrigation and drainage process, thereby improving water resource utilization efficiency and reducing the probability of over-irrigation or under-drainage. This solves the problems of related technologies that focus on static or short-term decision analysis at local time scales, model only a single irrigation action, have limited action space, and can only achieve short-term benefit optimization.

[0165] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.

[0166] When the processor 602 executes the program, it implements the paddy field irrigation and drainage decision-making method provided in the above embodiments.

[0167] Furthermore, electronic devices also include: Communication interface 603 is used for communication between memory 601 and processor 602.

[0168] The memory 601 is used to store computer programs that can run on the processor 602.

[0169] The memory 601 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.

[0170] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0171] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.

[0172] The processor 602 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of this application.

[0173] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described paddy field irrigation and drainage decision-making method.

[0174] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0175] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0176] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0177] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.

[0178] Those skilled in the art will understand that all or part of the steps of the methods implementing the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0179] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.

Claims

1. A method for making decisions on irrigation and drainage in paddy fields, characterized in that, Includes the following steps: Obtain historical environmental parameters of the target paddy field; A simulated environment for paddy field growth was constructed based on the historical environmental parameters. In a rice paddy growth simulation environment, the concept of the rice paddy irrigation and drainage decision process is transformed into a Markov decision process, and a reinforcement learning environment is constructed through the Markov decision process. The agent's policy network and value network are trained in the rice paddy growth simulation environment. During the training process, the network parameters of the value network are updated through the reinforcement learning environment, and the network parameters of the policy network are updated through the updated value network. The trained policy network is used to execute irrigation and drainage decisions for the target paddy field.

2. The paddy field irrigation and drainage decision-making method according to claim 1, characterized in that, The reinforcement learning environment includes a state space, an action space, a transition function, a reward function, and a discount factor, wherein... The environmental state vector of the state space includes the forecast rainfall sequence within a preset number of days in the future, the current water layer depth, the lower limit of the suitable water layer depth, the upper limit of the suitable water layer depth, the upper limit of rainwater storage, and water layer threshold parameters related to flood-resistant depth. The action space is constructed based on the decision variables of the paddy field irrigation and drainage decision cycle; The reward function includes an immediate reward and a round reward. The immediate reward is used to evaluate rainfall utilization, irrigation drainage and its impact on crop yield within a single decision cycle. The round reward is used to penalize the number of non-zero irrigation drainage decisions during the entire growth period. The expression for the transfer function is: in, For the first The initial water depth of the sky, For the first The initial water depth of the sky, For the first The amount of irrigation and drainage decisions per day.

3. The paddy field irrigation and drainage decision-making method according to claim 2, characterized in that, The expression for the instant reward is: in, For the first Instant rewards during the decision-making cycle. For the first Rainfall utilization rate reward and penalty items in the decision-making cycle. For the first The decision-making cycle is related to rewards and penalties for drought or flood-induced yield reduction. For the first The decision-making cycle is related to the water consumption penalty associated with the current irrigation and drainage volume; The expression for the round reward is: in, For the total reward function, The penalty coefficient is... The total number of irrigation and drainage actions within a round. This represents the number of decision steps within a complete decision-making cycle.

4. The paddy field irrigation and drainage decision-making method according to claim 1, characterized in that, The step of updating the network parameters of the value network through a reinforcement learning environment includes: The environmental state vector of the state space is input into the agent; The irrigation and drainage decision output by the agent is input into the action space, the irrigation and drainage decision is executed through the action space, and the execution result of the action space is input into the reinforcement learning environment. The environment state vector is updated through the reinforcement learning environment, and the updated environment state vector is input into the agent. The network parameters of the value network are updated through the updated environment state vector, and the network parameters of the policy network are updated through the updated value network.

5. The paddy field irrigation and drainage decision-making method according to claim 4, characterized in that, Before inputting the irrigation and drainage decision values ​​output by the agent into the action space, the method further includes: Extract the current water depth and water threshold parameters from the environmental state vector; The irrigation and drainage decision quantity is adjusted based on the current water layer depth and the water layer threshold parameter.

6. The paddy field irrigation and drainage decision-making method according to claim 1, characterized in that, The simulated environment for rice growth includes a rice paddy water balance model, a rice water production function, and a rice yield reduction model due to waterlogging. The expression for the paddy field water balance model is as follows: in, For the first The initial water depth of the day; For the first The initial water depth of the day; For the first Forecast rainfall for the day; To determine the first action based on the decision-making action Daily irrigation and drainage decision-making volume; For the first Forecast of crop water requirements for the day; For the first Forecast field seepage volume for the day; The expression for the rice yield reduction model due to waterlogging is: in, For rice yield reduction due to flooding; For the first The depth of the water layer in the field at that time; For rice plant height; To withstand deep flooding; These are empirical parameters. These are empirical parameters; The expression for the rice water production function is as follows: in, For the first Crop yield indicators during the decision-making cycle; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage Forecast evapotranspiration for the day; In the first After the decision-making cycle executes the irrigation decision, the first Future within each reproductive stage The potential evapotranspiration of the day; This is a water sensitivity index determined based on the current growth stage of rice.

7. The paddy field irrigation and drainage decision-making method according to claim 1, characterized in that, The process of using a trained policy network to execute irrigation and drainage decisions for the target paddy field includes: Real-time environmental parameters of the target paddy field are collected at the beginning of each irrigation and drainage decision cycle; A state vector is constructed based on the real-time environmental parameters. The state vector is input into a trained and converged policy network. The policy network outputs the irrigation and drainage decision quantity corresponding to the current decision cycle. A decision instruction is generated based on the irrigation and drainage decision quantity to control at least one action of the irrigation device and the drainage device. At the end of the current decision cycle, the actual rainfall, evapotranspiration and water depth are collected to update the state vector for the next decision cycle, until the end of the rice growth period in the target paddy field is reached.

8. A paddy field irrigation and drainage decision-making system, characterized in that, include: The acquisition module is used to obtain historical environmental parameters of the target paddy field; The construction module is used to construct a rice paddy growth simulation environment based on the historical environmental parameters; The transformation module is used to transform the concept of paddy field irrigation and drainage decision process into a Markov decision process in a paddy field growth simulation environment, and to build a reinforcement learning environment through the Markov decision process. The update module is used to train the policy network and value network of the agent in the rice field growth simulation environment. During the training process, the network parameters of the value network are updated through the reinforcement learning environment, and the network parameters of the policy network are updated through the updated value network. The execution module is used to execute irrigation and drainage decisions for the target paddy field using the trained policy network.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the paddy field irrigation and drainage decision method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the paddy field irrigation and drainage decision-making method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Method and system for dynamically optimizing irrigation and drainage strategy of rice field

    CN120258244A

  • Cache optimization method for cell-free multi-input multi-output environment based on VPPO algorithm

    CN120957192A

  • Rice field irrigation and drainage decision-making method and system based on human-in-loop reinforcement learning

    CN121541477A

  • Rice irrigation online learning forecasting method and system

    CN121766815A

  • Irrigation decision-making method and system based on crop model and deep reinforcement learning

    CN121788287A