Irrigation pipe network system layout optimization method, apparatus, medium, and product
By training deep Q-network models in virtual and real irrigation network systems, the application challenges of model-free control algorithms in irrigation network systems were solved, achieving stable optimization and improved economic benefits of the system.
Patent Information
- Application Number
- PCT/CN2024/103517
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-14
- Filing Date
- 2024-07-04
- Publication Date
- 2025-11-20
AI Technical Summary
Existing technologies make it difficult to effectively apply model-free control algorithms in irrigation rotation network systems, leading to abnormal system operation or damage. Furthermore, the abnormal data output by sensors cannot be used for learning and training, resulting in a lack of adaptability and accuracy.
By establishing a virtual irrigation network system to pre-train the deep Q-network model, and then retraining it in conjunction with a real irrigation network system, the deep Q-network model is optimized to adapt to the actual system and achieve layout optimization.
The model-free control algorithm was effectively applied to the irrigation network system, which improved the system's adaptability and optimization effect, ensured stable system operation, and improved economic benefits.
Smart Images

Figure CN2024103517_20112025_PF_FP_ABST
Abstract
Description
An irrigation pipe network system layout optimization method, device, medium and product
[0001] The present application claims priority to the Chinese patent application No. 2024106001552, filed on May 14, 2024, and entitled "An irrigation pipe network system layout optimization method, device, medium and product", the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to the field of agricultural intelligent irrigation technology, and in particular to an irrigation pipe network system layout optimization method, device, medium and product. BACKGROUND
[0003] In the field of agricultural irrigation technology, intelligent algorithms represented by model predictive control (MPC) are usually used to design and optimize the irrigation wheel irrigation pipe network system. However, this method still has the following shortcomings: due to the complexity of the wheel irrigation pipe network system control, a lot of historical information needs to be collected to build an accurate pipe network model; it is not easy to develop a simplified but accurate enough pipe network model, especially when the model is not accurate enough in describing the operation of the wheel irrigation pipe network system, the performance control may deviate from the expectation; in addition, most of the model-based control methods need to be adjusted according to the actual wheel irrigation pipe network system, and the self-adaptability is not strong. The above shortcomings of the model predictive control method promote the development of model-free control methods. However, on the one hand, since the output of the model-free control algorithm is uncontrollable and random before learning and training, it may cause abnormal operation or damage of the wheel irrigation pipe network system, so the model-free control algorithm cannot be directly connected to the wheel irrigation pipe network system; on the other hand, the learning and training process of the model-free control algorithm needs to be based on the interaction data of the algorithm and the wheel irrigation pipe network system, and without connecting to the wheel irrigation pipe network system, the data cannot be obtained; on the other hand, the data output by the sensors of the actual wheel irrigation pipe network system may be abnormal, which cannot be directly used for learning and training of the model-free control algorithm. Therefore, how to apply the model-free control algorithm to optimize the wheel irrigation pipe network system is a problem that needs to be solved at present.
[0004] SUMMARY
[0005] The purpose of the present application is to provide an irrigation pipe network system layout optimization method, device, medium and product, which pre-trains and re-trains a deep Q network model by using a virtual wheel irrigation pipe network system and a real wheel irrigation pipe network system, and optimizes the layout of the wheel irrigation pipe network system based on the trained deep Q network model.
[0006] To achieve the above-mentioned purpose, the present application provides the following solutions:
[0007] A kind of irrigation pipe network system layout optimization method, the method comprises:
[0008] S10: the system state of target round irrigation pipe network system current time is acquired.
[0009] S20: the system state of target round irrigation pipe network system current time is input into the depth Q network model trained, and the control action of current time is obtained;Wherein, the determination process of the depth Q network model trained is: according to virtual round irrigation pipe network system, the depth Q network model is pre-trained, and the depth Q network model of preliminary training is obtained;According to real round irrigation pipe network system, the depth Q network model of preliminary training is retrained, and the depth Q network model trained is obtained.
[0010] S30: according to the control action of current time, target round irrigation pipe network system is optimized.
[0011] Wherein, the system state of target round irrigation pipe network system current time is determined according to the real round irrigation pipe network system optimized after last time.
[0012] The application also discloses a computer system, comprising: memory, processor and computer program stored on the memory and executable on the processor, the processor executes the computer program to realize the steps of the irrigation pipe network system layout optimization method.
[0013] The application also discloses a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the steps of the irrigation pipe network system layout optimization method.
[0014] The application also discloses a computer program product, comprising a computer program, which is executed by a processor to realize the steps of the irrigation pipe network system layout optimization method.
[0015] According to the specific embodiments of the application, the following technical effects are disclosed:
[0016] In the face of how to apply model-free control algorithm to optimize irrigation pipe network system, the application firstly pre-trains the depth Q network model according to the established virtual round irrigation pipe network system, obtains the depth Q network model of preliminary training, and then connects it to the real round irrigation pipe network system, so as to avoid directly connecting the untrained depth Q network model to the real round irrigation pipe network system, which causes system abnormality or damage due to uncontrollable output action;Then, the depth Q network model of preliminary training is retrained in the real round irrigation pipe network system, and secondary learning is carried out according to the real round irrigation pipe network system, so as to enhance the applicability of the depth Q network model;Further, the depth Q network model trained is used to optimize the round irrigation pipe network system.
[0017] Drawings
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description only represent some of the embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative labor based on these drawings also belong to the protection scope of the present application.
[0019] Fig. 1 is a schematic diagram of the irrigation pipe network system layout optimization method provided by the embodiment of the present application;
[0020] Fig. 2 is a schematic diagram of the Matlab-EPANET joint simulation based on DQN provided by the embodiment of the present application;
[0021] Fig. 3 is a schematic diagram of the program directory created in Matlab provided by the embodiment of the present application;
[0022] Fig. 4 is a schematic diagram of the DQN-RL framework provided by the embodiment of the present application;
[0023] Fig. 5 is a schematic diagram of the neural network framework in the deep Q network model provided by the embodiment of the present application;
[0024] Fig. 6 is a schematic diagram of the iteration process based on DQN provided by the embodiment of the present application;
[0025] Fig. 7 is a schematic diagram of the irrigation wheel irrigation pipe network system simulation model based on DQN provided by the embodiment of the present application;
[0026] Fig. 8 is a schematic diagram of the internal structure of the computer device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0027] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only represent some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor also belong to the protection scope of the present application.
[0028] Since the mid-1990s, China's irrigation technology has evolved from the initial demonstration stage to the large-scale development stage. In this process, as the control area of the irrigation system expands, the head loss of the system also increases. When optimizing the irrigation pipe network system in large-scale irrigation areas with relatively uniform ground slope, the rotation irrigation working method is often used to address the challenge of low-cost application of field crop drip irrigation. However, in the academic field at home and abroad, there is relatively little research on rotation irrigation grouping optimization, and a comprehensive system pipe network optimization design scheme has not yet been formed. At the same time, in the actual design field, there is also a lack of high-reliability pipe network optimization algorithms.
[0029] In the field of agricultural irrigation technology, intelligent algorithms represented by Model Predictive Control (MPC) provide new assistance for rotation irrigation pipe network system design and optimization. However, the complexity of rotation irrigation pipe network system control requires people to collect more historical information to build an accurate pipe network model. However, it is not easy to develop a simplified but sufficiently accurate pipe network model, especially when the model is not accurate enough in describing the operation of the rotation irrigation pipe network system, the control of performance may deviate from the expected. In addition, most of these model-based control methods need to be adjusted according to the actual rotation irrigation pipe network system, resulting in weak adaptability. In this regard, the shortcomings of model-based control methods promote the development of model-free control methods. The RL (Reinforcement Learning) algorithm, as a classic model-free algorithm, provides an improved methodology for solving this problem.
[0030] Considering the advantages of RL in adaptability and its interactive learning ability in learning control preferences, this algorithm shows significant potential in dealing with complex systems, especially in rotation irrigation pipe network system design and optimization applications. For example: (1) simplicity: model-free methods are usually simpler because they do not require explicit modeling of the environment's dynamics. They learn a policy or value function directly from interactions with the environment. (2) flexibility: model-free RL is usually more flexible and robust to environmental changes because it does not rely on a model that may not be accurate or complete. (3) sample efficiency: although model-free methods may have lower sample efficiency than model-based methods, they are usually simpler to implement and perform well in environments where it is difficult to accurately model the dynamics, which is suitable for use in rotation irrigation pipe network system optimization control. (4) computational efficiency: model-free methods can be more computationally efficient because they avoid the potentially costly step of learning an accurate model of the environment. (5) applicability: model-free RL can be applied to a wider range of problems, especially when the environment is complex or difficult to accurately model. Therefore, it is necessary to further explore RL to coordinate the agent system participating in rotation irrigation pipe network optimization.
[0031] However, on the one hand, since the output of the model-free control algorithm is uncontrollable and random before learning and training, it may cause abnormal operation or damage of the wheel irrigation pipe network system, so the model-free control algorithm cannot be directly connected to the wheel irrigation pipe network system; on the other hand, the learning and training process of the model-free control algorithm needs to be based on the interaction data of the algorithm and the wheel irrigation pipe network system, and cannot be obtained without connecting to the wheel irrigation pipe network system; on the other hand, the data output by the actual wheel irrigation pipe network system sensor may be abnormal and cannot be directly used for learning and training of the model-free control algorithm. Therefore, how to apply the model-free control algorithm to optimize the irrigation pipe network system is a problem that needs to be solved at present.
[0032] The purpose of the present application is to provide an irrigation pipe network system layout optimization method, device, medium and product, which pre-trains and re-trains a deep Q network model based on a virtual wheel irrigation pipe network system and a real wheel irrigation pipe network system, so as to optimize the layout of the wheel irrigation pipe network system based on the trained deep Q network model.
[0033] In order to make the above-mentioned purpose, characteristics and advantages of the present application more obvious and easy to understand, the present application will be further described in detail below in combination with the drawings and specific embodiments.
[0034] Embodiment 1
[0035] Referring to FIG. 1, the present embodiment provides an irrigation pipe network system layout optimization method, which comprises:
[0036] S10: obtaining the system state of the target wheel irrigation pipe network system at the current time.
[0037] S20: inputting the system state of the target wheel irrigation pipe network system at the current time into the trained deep Q network model to obtain the control action at the current time; wherein the determination process of the trained deep Q network model is: pre-training the deep Q network model according to the virtual wheel irrigation pipe network system to obtain a preliminarily trained deep Q network model; re-training the preliminarily trained deep Q network model according to the real wheel irrigation pipe network system to obtain the trained deep Q network model.
[0038] S30: optimizing the target wheel irrigation pipe network system according to the control action at the current time.
[0039] The system state of the target wheel irrigation pipe network system at the current time is determined according to the real wheel irrigation pipe network system after optimization at the previous time.
[0040] The deep Q network model is pre-trained according to the virtual wheel irrigation pipe network system to obtain a preliminarily trained deep Q network model, which specifically comprises:
[0041] training the deep Q network model according to data in the experience pool to obtain a preliminarily trained deep Q network model;
[0042] The data in the experience pool is from an interaction process of the deep Q network model and the virtual rotating irrigation pipe network system; the deep Q network model and the virtual rotating irrigation pipe network system are respectively built in Matlab software; and the virtual rotating irrigation pipe network system is simulated in EPANET software.
[0043] The application sets key parameters of the irrigation pipe network system in the Matlab environment, such as spatial coordinates of the irrigation emitter node, coordinates of the water source point (water source position), coordinates of the water outlet point (water outlet pile), and positions of the pressure regulator, rated flow and lift parameters of the water pump, irrigation quota, length and outer diameter of the pipe section, and roughness coefficient of the pipe, etc. By creating a program directory coupled with EPANET, a comprehensive simulation model is constructed, and the system pressure of the pipe node and the dynamic change of the water flow in the pipe section at different time points are monitored and analyzed by using the EPANET software. The specific process of building the simulation model of the rotating irrigation pipe network system by EPANET and Matlab software is shown in Fig. 2.
[0044] By using the above parameters, the application establishes a program system closely coupled with the EPANET software, thereby constructing a composite simulation model. The model applies the analysis capability of EPANET to continuously track and evaluate the pressure performance of each node in the system at different time points and the change of the water flow (such as flow and flow rate) in the pipe section.
[0045] The coupling of EPANET and Matlab is realized by the interactive interface tool EPANET-Matlab Toolkit. The System toolbox in Matlab provides a standardized interface (EPANET-Matlab Toolkit) for model exchange and collaborative simulation, which facilitates efficient data structure access and integration between EPANET and Matlab. This architecture supports the definition and calculation process of variables of the reinforcement learning algorithm of the irrigation pipe network model in the Matlab environment, thereby facilitating the implementation of the optimization control strategy of the irrigation pipe network system.
[0046] The building and running of the virtual rotating irrigation pipe network system specifically include:
[0047] Due to the complexity of the pipe network information of the micro-irrigation system, the Matlab 2016 software can be used to edit the input file in the EPANET software, including the node coordinates of each irrigation emitter in the pipe network, the water source point coordinates, the water outlet pile coordinates, the pipe section length, and the pipe roughness coefficient, etc. The program directory created by Matlab is shown in Fig. 3, and the main meanings of the programs are as follows:
[0048] * setvalve.m: This program is mainly used for setting the coordinate position, connection node and coefficient of pressure regulator.
[0049] * setsource.m: This program is mainly used for setting the coordinate position of water source node.
[0050] * setreactions.m: This program is mainly used for setting the flow coefficient of emitter node.
[0051] * setpump.m: This program is mainly used for setting the coordinate position of water pump node and pump-pressure curve coefficient.
[0052] * setpipes.m: This program is used to set the pipe parameters for generating pipe network, mainly including pipe label, starting node coordinate, end node coordinate, pipe length, pipe diameter and pipe friction coefficient, etc.
[0053] * setjunction.m: This program is mainly used for setting the parameters of each node.
[0054] * sethead.m: This program is mainly used to generate information of water source node.
[0055] * setemitters.m: This program is mainly used to set the node flow coefficient of emitter.
[0056] * setcoord.m: This program is mainly used to calculate the coordinates and elevations of emitter node, branch pipe, water outlet pile and water source node.
[0057] * output.m: This program is mainly used to output pipe segment information and node information to the result file, and calculate the corresponding indicators.
[0058] * inpmake.m: This program is mainly used to collect the calculation data of each program and generate the inp file required by EPANET.
[0059] * getinp.m: This program is mainly used to read the inp file, separately extract the required pipe network coordinate and pipe segment data file and write it into a new file.
[0060] * application.m: This program is the total program, which is used to unify the calculation of each program and set the parameters of pipe network area, shape, diameter and water head, adjust and optimize the parameters to simulate the layout of micro-irrigation pipe network.
[0061] The edited input file (generated.inp) is imported into EPANET, and simulation is run to realize batch calculation of each node and each pipe in the simulated pipe network, and to achieve the purpose of the model application, thereby providing technical support for optimization of the hydraulic performance of the micro-irrigation pipe network system.
[0062] In reinforcement learning, the policy is usually represented by the symbol π, which is determined by the following formula and defined as a mapping from the system state to the action (π: S→A) ; π(a|s)=P(A=a t |S=S t ).
[0063] Wherein, π(a|s) represents the conditional probability of taking a specific action a in the current state s, P represents the conditional probability of outputting the control action, A represents the control action, and S refers to the current state. On this basis, the evaluation of the policy depends on the expected value of the cumulative reward. The cumulative reward can be estimated by the experience trajectory, as shown in the following formula,
[0064] Wherein, G t represents the cumulative reward, r represents the current reward, T represents the termination time, that is, the total length of the expected reward or the length of the sequence, t represents the current time step, k represents the kth time step after the current time step t, and γ represents the discount factor, which is usually a constant between 0 and 1. The cumulative reward is the sum of the current reward multiplied by the corresponding discount factor γ, and when γ≥1, the cumulative reward does not converge.
[0065] The state value function refers to the expected cumulative reward obtained by selecting the action according to the given policy π after starting from the current state S t , as shown in the following formula. The state value function can evaluate the good or bad degree of the state at this moment,
[0066] In the case of knowing the current state S t and the current action a t , the expected cumulative reward G t obtained by selecting the action according to the given policy π is called the state-action value function under the current state and action, as shown in the following formula. The state-action value function can be used to evaluate the good or bad degree of the action a t selected under the current state S t ,
[0067] The core goal of reinforcement learning is to output an optimal control policy that maximizes cumulative rewards through a series of iterations and updates. In this process, policy improvement usually relies on an advantage function, which can be determined by the difference between a state value function and a state-action value function, as shown in the following equation, A π (s,a)=Q π (s,a)-v π (s)。
[0068] The advantage function is used to evaluate the superiority of a given action relative to the average action value. If the advantage function A π (s,a) is greater than 0 in the current state, it means that taking the output action in the current state will result in a higher average reward value than taking all possible actions in that state. That is, if the advantage function is less than 0, it means that the reward value of the current output action is lower than the average reward value.
[0069] We use the advantage function to evaluate the pros and cons of the policy, which combines the advantages of state value and state-action value, and can distinguish the relative value of different actions in a specific state.
[0070] Traditional Q-learning reinforcement learning algorithms use tables to store Q values, which are suitable for small discrete state spaces and action spaces. However, irrigation pipe network systems usually have large-scale and continuous state and action spaces, which may cause dimensionality explosion in practical applications. Since deep neural networks have strong descriptive capabilities for complex nonlinear state information, RL can use deep neural networks to estimate Q values. Deep Q-network (DQN) is a RL algorithm that combines deep learning and reinforcement learning. In the DQN algorithm, a deep neural network is used to approximate the Q value of each action in the current state, which can greatly improve the storage and calculation efficiency of actions.
[0071] The DQN algorithm is similar to the traditional Q-learning algorithm framework, but it adds a deep learning network to approximate the Q function. In the DQN algorithm, the target Q value update is as follows, Q target =r+γmaxQ(s,a;θ)。
[0072] where Q target is the calculated target Q value used to train the current Q network; r is the received immediate reward; γ represents the discount factor; Q(s,a; θ) is the Q value of the next state and all possible actions using the weight θ of the target Q network.
[0073] In the training process of the deep Q network model, the current Q network and the target Q network are initialized with random weights, and a balance is achieved between exploration and utilization based on the experience replay mechanism. The DQN agent (deep Q network model) accumulates transition experiences and iteratively updates the Q network parameters using experience data to reduce the deviation between the predicted value and the target value.
[0074] The DQN framework (i.e., deep Q network model) built mainly consists of two parts, namely the control loop and the learning loop, as shown in FIG. 4. The control loop is responsible for real-time interaction with the environment (such as a virtual wheel irrigation pipe network system), while the learning loop is responsible for learning and optimizing the control strategy from the interaction. The specific process of the DQN algorithm is as follows:
[0075] 1) Real-time interaction (control loop): Monitor the environment state and output real-time control actions using the greedy strategy, and store the results (in the experience pool).
[0076] 2) Batch learning (learning loop): Randomly sample from the experience pool, calculate the current Q value, and update the current Q network weight.
[0077] 3) Regular synchronization (learning loop): Regularly copy the current Q value network weight to the target Q network to avoid self-divergence problems in the learning process.
[0078] The neural network in the deep Q network model includes an input layer, two hidden layers, and an output layer, as shown in FIG. 5. The neural network is divided into a current Q network and a target Q network.
[0079] The specific process of the DQN algorithm can also be as follows:
[0080] 1. Initialization:
[0081] Initialize the policy (current) network and its weight parameters.
[0082] Initialize the target network and its weight parameters, and the weight parameters of the policy network and the target network are initially the same.
[0083] Initialize the experience replay pool (i.e., experience pool) D as an empty set.
[0084] 2. Experience collection:
[0085] For each round, initialize the initial state s of the environment.
[0086] For each step, select an action a from the policy network according to the ∈-greedy strategy (∈ represents the probability):
[0087] Randomly select an action with probability ∈;
[0088] Select the action corresponding to the larger Q value with probability 1-∈.
[0089] 3. Perform action:
[0090] Perform action a in the environment, observe reward r and new state s′.
[0091] 4. Store experience:
[0092] Store experience (s, a, r, s′) in experience replay pool D.
[0093] 5. Learn:
[0094] Randomly sample a mini-batch of training data from experience replay pool D.
[0095] For each experience (s, a, r, s′) in the mini-batch, compute target Q value.
[0096] Compute loss function from current Q value and target Q value.
[0097] Perform gradient descent step to update weight parameters θ of policy network to minimize loss function.
[0098] 6. Target network update:
[0099] Every few steps, copy weight parameters of policy network to target network, updating its weight parameters.
[0100] 7. Iterative optimization:
[0101] Repeat steps 2 through 6 until termination condition is met, such as network convergence or predetermined number of episodes reached.
[0102] The interaction process between the deep Q network model and the virtual rotating irrigation pipe network system includes: the deep Q network model obtains the system state of the virtual rotating irrigation pipe network system, adopts the greedy strategy to output the control action, and calculates the corresponding reward value; the virtual rotating irrigation pipe network system performs layout optimization according to the control action; the deep Q network model then obtains the system state of the virtual rotating irrigation pipe network system after layout optimization, and so on; the system state, control action and reward value and other data obtained in the interaction process are stored in the experience pool.
[0103] Referring to FIG. 6, the specific steps based on Matlab-EPANET joint simulation are as follows:
[0104] The irrigation rotating irrigation pipe network model completes the initial simulation in EPANET, and the result is output as the system state to the Matlab environment; a round of iteration of the algorithm is performed by using Matlab, and the obtained control action is fed back to EPANET through the System toolbox to update the state of the irrigation rotating irrigation pipe network model; the whole process is repeated until the simulation is completed, and the optimal irrigation rotating irrigation pipe network control strategy is output.
[0105] Specifically: in the initial iteration phase, the DQN agent obtains the system state S from the irrigation rotation canal network model in EPANET t , and transmits the control action a t ; this process is repeated until the algorithm model converges. The iteration period is mainly limited by the length of the EPANET simulation.
[0106] In order to standardize the iteration period, the time step of EPANET simulation and DQN agent iteration is set to 10 seconds, that is, the time increment Δt = 10 seconds.
[0107] The joint simulation environment constructed by the present application can effectively solve the inherent constraint problem of the built-in control logic of EPANET software: in most studies, EPANET software is applied to the construction of an irrigation pipe network system simulation model. However, EPANET has limitations in integrating intelligent control algorithms for irrigation systems, especially when integrating advanced deep reinforcement learning algorithms such as DQN. However, the System toolbox in Matlab provides a standardized interface for model exchange and collaborative simulation, facilitating efficient data structure access and integration between EPANET and Matlab. This architecture supports the definition and calculation of variables for the reinforcement learning algorithm for the irrigation rotation canal network model in the Matlab environment, thereby facilitating the implementation of an optimal control strategy for the irrigation pipe network system. Therefore, the present application realizes the layout optimization of the irrigation pipe network system by building a Matlab-EPANET joint simulation platform based on RL. The developed Matlab-EPANET joint simulation model makes it easier to apply advanced control algorithms such as DQN to the built-in simulation environment of EPANET. In addition, the developed joint simulation model evaluates the performance of the control strategy in the OpenAI Gym (a tool in Matlab) environment, which is easier to debug and has better scalability. This embodiment uses the RL algorithm, and the joint simulation platform realizes the optimization of economic benefits while ensuring hydraulic performance through continuous interaction with the environment and iterative learning, to determine the optimal irrigation pipe network layout method.
[0108] The limitations of EPANET are specifically embodied in the following aspects: EPANET can achieve efficient hydraulic simulation of the water supply network, but its editing ability and post-processing information processing ability have obvious shortcomings. For micro-irrigation system design managers, it is necessary to further expand the pipe network hydraulic simulation according to specific needs, and through the help of the advanced programming environment of MATLAB data processing and analysis, the data can be connected with EPANET, and the water supply network can be studied more deeply, especially when combining with advanced deep reinforcement learning algorithms such as DQN, which is the most obvious and significant limitation, and this is also the reason why the joint simulation model is built.
[0109] According to the real rotating irrigation pipe network system, the preliminarily trained deep Q network model is retrained to obtain a trained deep Q network model, specifically including:
[0110] The preliminarily trained deep Q network model is trained according to the data in the experience pool to obtain a trained deep Q network model; wherein the data in the experience pool comes from the interaction process of the preliminarily trained deep Q network model and the real rotating irrigation pipe network system.
[0111] Through the second learning and training according to the actual rotating irrigation pipe network system, i.e. the real rotating irrigation pipe network system, the applicability of the preliminarily trained deep Q network model can be enhanced. For example, in the actual rotating irrigation pipe network system, the sensor may be aged for a long time, resulting in a shift in the measured value. When the control action is output to the rotating irrigation pipe network system according to this measured value, a reward value or system state different from the optimal strategy will be obtained. However, based on the internal reward and punishment mechanism of the deep Q network model, after the second learning and training, a deep Q network model suitable for the actual rotating irrigation pipe network system will be obtained. The output of this deep Q network model is the optimal control action suitable for the actual rotating irrigation pipe network system. When the sensor has a shifted measured value is obtained again, a control action different from the control action before the second training (i.e. different from the control action output after pre-training) will be output. When this control action is output to the rotating irrigation pipe network system, a reward value or system state corresponding to the optimal strategy will be obtained.
[0112] Referring to FIG. 7, the system state includes the scale area of the irrigation pipe network system, the rated flow of the water pump, the working pressure downstream of the pressure regulator, the ground slope of the pipe (I1), the nominal outer diameter of the pipe (Dm, Ds, and Dl of the main pipe, branch pipe, and capillary pipe, respectively, in millimeters), and the design flow of the sprinkler under a working pressure of 0.1 MPa (qd, in liters / hour).
[0113] The DQN framework is suitable for the optimal control of irrigation pipe network systems, aiming to optimize the pipe network layout on the basis of ensuring the hydraulic design specifications to enhance the economic benefits of the system. Generally, for an irrigation pipe network system considering RL, the key steps include describing the system state, designing the control action, and constructing the reward function. In this embodiment, six state variables closely related to the optimization of irrigation pipe network layout are identified for use in the reinforcement learning process through correlation analysis screening.
[0114] The control action includes the number of branch pipes, the arrangement mode of the capillary pipes (one-way, two-way, or ring arrangement), and whether a pressure regulator is used in the system, as shown in FIG. 7.
[0115] The control action is used for the DQN algorithm to make optimal control decisions. The optimal irrigation pipe network layout is the core goal, so the number of branch pipes, the arrangement mode of the capillary pipes (one-way, two-way, or ring arrangement), and whether a pressure regulator is used in the system are used as the reinforcement learning control action.
[0116] The reward value calculation formula of the deep Q network model is as follows: Reward = -μ1QV - μ2EU - μ3CT.
[0117] Wherein, Reward is the reward value, QV is the flow deviation rate, EU is the irrigation uniformity index, CT is the annual cost of the irrigation system, and μ1, μ2, and μ3 represent the corresponding weight coefficients, respectively.
[0118] In the above-mentioned reward function framework, the comprehensive performance index is composed of the following three core evaluation factors: (1) flow deviation rate; (2) irrigation uniformity; and (3) annual cost of the irrigation system. In the above-mentioned comprehensive reward function, μ1, μ2, and μ3 represent the weight coefficients of the above-mentioned evaluation factors in the optimization strategy, reflecting the importance of each performance index in the control decision.
[0119] In view of the fact that the hydraulic performance and economic benefits constitute the core indicators of irrigation system optimization, this embodiment selects the flow deviation rate and the irrigation uniformity index of the irrigation system unit pipe network as the key variables for constructing the reward function; for refined economic benefit evaluation, this embodiment includes the annual cost of the irrigation system in the calculation of the reward function, as shown in FIG. 7.
[0120] The flow deviation rate QV includes the design flow deviation rate within the irrigation unit and the design flow deviation rate within the rotation irrigation group.
[0121] According to the “Technical Standards for Irrigation Engineering” (GB / T 50485-2020), the design flow deviation rate qv(%) of the sprinkler is selected as the hydraulic design index within the irrigation unit and the rotation irrigation group, which needs to meet the following conditions:
[0122] wherein: q v-sb is the design flow rate deviation in the irrigation unit, %; q max-sb is the maximum flow rate of the emitters in the irrigation unit, L / h; q min-sb is the minimum flow rate of the emitters in the irrigation unit, L / h; q d-sb is the design flow rate of the emitters in the irrigation unit, L / h; q d-rg is the design flow rate of the emitters in the wheel irrigation group, L / h; q v-rg is the design flow rate deviation in the wheel irrigation group, %; q max-sb is the maximum flow rate of the emitters in the wheel irrigation group, L / h; q min-sb is the minimum flow rate of the emitters in the wheel irrigation group, L / h.
[0123] The irrigation uniformity index EU includes the Keller uniformity coefficient of the irrigation unit and the Keller uniformity coefficient of the wheel irrigation group. That is, the Keller uniformity coefficient EU (%) that can comprehensively reflect the combined effects of the hydraulic deviation, the manufacturing deviation, and the combination effect is selected to evaluate the hydraulic performance in the irrigation unit and the wheel irrigation group:
[0124] wherein: EU sb is the Keller uniformity coefficient of the irrigation unit, %; is the average flow rate of the emitters in the irrigation unit, L / h; CV m is the emitter manufacturing deviation coefficient; N is the number of emitters per plant; EU rg is the Keller uniformity coefficient of the wheel irrigation group, %; is the average flow rate of the emitters in the wheel irrigation group, L / h.
[0125] The annual cost CT of the irrigation system includes the initial investment cost of the pipe network system, the maintenance and operation expenses of the system, the energy consumption cost (i.e., the power cost related to energy consumption), and the irrigation cost.
[0126] In order to further accurately evaluate the economic benefits, the annual cost of the wheel irrigation system is introduced into the reward function in this embodiment. The annual cost of the wheel irrigation system is calculated in four aspects: the annual cost (CT) of the wheel irrigation group includes the initial investment cost (Ca) of the pipe network system, the maintenance and operation expenses (Cm) of the system, the power cost (Ce) related to energy consumption, and the irrigation cost (Cw). The above costs are all calculated as the annual cost per unit area (yuan / (hm2·a)). When calculating the annual cost per unit area of the wheel irrigation group, the working life (m) and the annual interest rate (i) of each component of the system are considered, and the annual cost per unit area is calculated as follows: C inp = L l P l + L mP m +L n P n +N pi P pi +N wi P wi +N vi P vi +N ai P ai +N si P si .
[0127] where CRF is the capital recovery factor, 1 / a; C inp is the investment cost of each component of the rotating irrigation group, yuan; L l , L m and L n are the lengths of the main pipe, branch pipe and rough pipe, respectively, m; P l , P m and P n are the unit prices of the main pipe, branch pipe and rough pipe, respectively, yuan / m; N pi , N wi , N vi , N ai , N si are the numbers of water pumps, water outlet piles, pressure regulating valves, bypasses and plugs, respectively, pieces; P pi , P wi , P vi , P ai , P si are the unit prices of water pumps, water outlet piles, pressure regulating valves, bypasses and plugs, respectively, yuan / piece. With reference to the current irrigation situation in the northwest region of China, the service life of the main pipe, water pump and water outlet pile is 15 years, the service life of the branch pipe and pressure regulating valve is 5 years, and the service life of the rough pipe, bypass and plug is 1 year. The maintenance cost Cm of the irrigation pipe network is 5% of the investment of the system pipe network.
[0128] The energy consumption cost of the irrigation system is calculated according to the power consumption when the irrigation water is lifted to a specific working pressure under different working conditions:
[0129] where N p is the working power of the water pump, kW; O t is the annual irrigation time to meet the irrigation demand, h / a; En c is the electricity cost, yuan / kWh; Q s is the outlet flow of the water pump, m 3 / s; H0 is the working lift of the water pump, m; E p is the efficiency of the water pump, which can be taken as 0.65 in general; and R g is the rough irrigation quota, m3 / (hm 2 ·a)。
[0130] To ensure that all positions in the rotation irrigation group are fully irrigated, R g is calculated by the following formula:
[0131] In the formula: R n is the net irrigation rate, m 3 / (hm 2 ·a)。
[0132] The water fee is calculated according to the gross irrigation rate (R g ) considering the uniformity of irrigation and the water price (P w , yuan / m 3 ): C w =R g P w .
[0133] Based on the existing research results, the specific problems of rotation irrigation group division are analyzed in depth, and the influence of factors such as the number of rotation irrigation group units, the configuration mode of pipe network and whether to use pressure regulator on the flow deviation rate, irrigation uniformity and annual economic cost is analyzed in detail. On this basis, a Matlab-EPANET irrigation pipe network system co-simulation model integrated with an efficient RL framework is developed, and through its advanced intelligent irrigation pipe network layout optimization method, it brings an innovative perspective for the research and development of large-scale irrigation systems. The purpose is to ensure that the irrigation pipe network system meets the hydraulic design specification, and to determine the maximum economic benefit of the irrigation system.
[0134] In order to make the application of the present application easier to understand, the application of the present application is further introduced through the following cases:
[0135] An irrigation pipe network system in an agricultural area is composed of main pipes and branch pipes. The main pipes run through the entire farmland center, and the branch pipes extend from the main pipes to each irrigation area. The original design of such a system may not take into account the special water demand of some plots, or some areas may have insufficient water pressure during peak periods. For example, the system may be a uniformly distributed grid structure, with pipes of the same diameter in each grid, but in fact, due to the topography or geology of some areas, more water may be needed.
[0136] The main problems before optimization are:
[0137] ① Water pressure inequality: During peak hours, some areas far from the water source have significantly insufficient water pressure. ② Resource waste: Due to the suboptimal design of the pipe network, some areas are over-supplied with water, while others are under-supplied. ③ Energy consumption issues: Due to the unreasonable system design, more energy is needed to pump water to remote areas.
[0138] First, by monitoring the system status in real-time, such as water flow, water pressure, and pipe usage, the data is input into a trained deep Q-network model. This model can predict the optimal control action based on the current state, such as adjusting the valve opening degree of certain pipes, reconfiguring the water pressure, or increasing or decreasing the water supply of certain pipe sections.
[0139] Advantages of the optimized structure and layout:
[0140] ① Reorganization of pipe network layout: Based on the model's recommendations, the original uniform grid structure is adjusted, and for areas with high water demand, the pipe diameter is increased or parallel pipes are added to improve water supply capacity. ② Dynamic water pressure adjustment: The output of the pump station is adjusted by the intelligent control system when needed to optimize the water pressure distribution and ensure that each area can obtain appropriate water pressure. ③ Energy saving and emission reduction: The optimized system reduces the running time of the pump station and reduces unnecessary water pressure output, achieving the effect of energy saving and emission reduction.
[0141] Through optimization, the irrigation efficiency of this agricultural area is significantly improved, water resource utilization is more reasonable, energy consumption is reduced, and the impact on the environment is also reduced. Each area has been allocated more reasonable water resources according to actual demand, improving the growth conditions and yield of crops. Through such an example, it can be seen that the optimization method based on DQN not only improves the physical layout of the irrigation system, but also improves its operating efficiency and the rationality of resource use.
[0142] Example 2
[0143] A computer system comprising: a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the steps of the irrigation pipe network system layout optimization method of embodiment 1.
[0144] Example 3
[0145] A computer-readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the steps of the irrigation pipe network system layout optimization method of embodiment 1.
[0146] Example 4
[0147] A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the irrigation pipe network system layout optimization method of embodiment 1.
[0148] Embodiment 5
[0149] A computer device can be a database, and an internal structure diagram thereof can be as shown in FIG. 8. The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is used to store transactions to be processed. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals through network connection. The computer program is executed by the processor to implement the irrigation pipe network system layout optimization method described in embodiment 1.
[0150] It should be noted that the object information (including but not limited to object device information, object personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the object or fully authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0151] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The database involved in the embodiments provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0152] In summary, the present application has the following advantages:
[0153] 1) The present application avoids the situation that the untrained deep Q network model is directly connected to the real rotating irrigation pipe network system, which may cause system abnormalities or damage due to uncontrollable model output actions, and solves the problem that the deep Q network model cannot be directly trained according to the real rotating irrigation pipe network system (which may output abnormal data), and provides a method for using a model-free control algorithm to optimize the layout of the real rotating irrigation pipe network system; on the other hand, by retraining the pre-trained deep Q network model according to the real rotating irrigation pipe network system, the deep Q network model is enhanced for the applicability of the real rotating irrigation pipe network system, so even if the data output by the real rotating irrigation pipe network system is abnormal or has errors, the deep Q network model can output the optimal control action suitable for the real rotating irrigation pipe network system.
[0154] 2) The intelligent irrigation pipe network system layout optimization method based on the deep Q network reinforcement learning algorithm provided by the present application comprehensively considers the system scale characteristics and the actual deviation of the terrain, and integrates all the capillary tubes and branch pipes of the units in the rotating irrigation system into the calculation framework of the hydraulic model. By calculating and analyzing the flow deviation rate of the irrigation pipe network, the irrigation uniformity coefficient and the economic indicators of the system, the layout of the pipe network system is optimized to improve the overall irrigation efficiency and economic benefit of the irrigation system.
[0155] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the various embodiments can be referred to each other.
[0156] The principles and implementation modes of the present application are described by applying specific examples in this paper, and the above description of the embodiments is only used to help understand the device and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In view of the above, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for layout optimization of an irrigation pipe network system, characterized in that, The method comprises: acquiring a system state of a target wheel irrigation pipe network system at a current time point; inputting the system state of the target wheel irrigation pipe network system at the current time point into a trained deep Q network model to obtain a control action at the current time point; wherein, a determination process of the trained deep Q network model comprises: pre-training a deep Q network model according to a virtual wheel irrigation pipe network system to obtain a preliminarily trained deep Q network model; and re-training the preliminarily trained deep Q network model according to an actual wheel irrigation pipe network system to obtain the trained deep Q network model; optimizing the target wheel irrigation pipe network system according to the control action at the current time point; wherein, the system state of the target wheel irrigation pipe network system at the current time point is determined according to an actual wheel irrigation pipe network system optimized at a previous time point.
2. The method of claim 1, wherein, The pre-training of the deep Q network model according to the virtual wheel irrigation pipe network system to obtain the preliminarily trained deep Q network model specifically comprises: training the deep Q network model according to data in an experience pool to obtain the preliminarily trained deep Q network model; wherein, the data in the experience pool comes from an interaction process of the deep Q network model and the virtual wheel irrigation pipe network system; the deep Q network model and the virtual wheel irrigation pipe network system are respectively built in Matlab software; and the virtual wheel irrigation pipe network system is simulated in EPANET software.
3. The method of claim 2, wherein, The re-training of the preliminarily trained deep Q network model according to the actual wheel irrigation pipe network system to obtain the trained deep Q network model specifically comprises: training the preliminarily trained deep Q network model according to data in an experience pool to obtain the trained deep Q network model; wherein, the data in the experience pool comes from an interaction process of the preliminarily trained deep Q network model and the actual wheel irrigation pipe network system.
4. The method of claim 1, wherein, The system state comprises a dimension area of the irrigation pipe network system, a rated flow of a water pump, a working pressure downstream of a pressure regulator, a ground slope, a nominal outer diameter of a pipe, and a design flow of a sprinkler under a working pressure of 0.1 MPa.
5. The method of claim 1, wherein, The control action comprises a number of branch pipes, an arrangement mode of capillary tubes, and whether to use a pressure regulator.
6. The method of claim 1, wherein, A reward value calculation formula of the deep Q network model is as follows: Reward = -μ1QV - μ2EU - μ3CT; wherein, Reward is the reward value, QV is a flow deviation rate, EU is an irrigation uniformity index, CT is an annual cost of the irrigation system, and μ1, μ2, and μ3 respectively represent corresponding weight coefficients. A neural network in the deep Q network model comprises an input layer, two hidden layers, and an output layer; and the neural network is divided into a current Q network and a target Q network.
7. The method of claim 1, wherein, A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the steps of the irrigation pipe network system layout optimization method in any one of claims 1-7.
8. A computer system comprising: The computer program is executed by the processor to implement the steps of the irrigation pipe network system layout optimization method in any one of claims 1-7.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, 10. A computer program product comprising a computer program, characterized in that, The computer program, which is executed by a processor, implements the steps of the method for optimizing the layout of an irrigation pipe network system according to any one of claims 1-7.
Citation Information
Patent Citations
Irrigation control method and device, electronic equipment and storage medium
CN115530054A
Network model method and device, data processing method and device and electronic equipment
CN115983346A
Intelligent optimization method and system for water resource allocation of water conservancy project
CN117575245A
Method and device for optimisation of collection & processing of soil moisture data to be used in automatic instructions generation for optimal irrigation in agriculture
EP4173477A1
Irrigation control with deep reinforcement learning and smart scheduling
US20220248616A1
Cited By
Intelligent remote control method and system for irrigation electric valve
CN121325568A
Wetland group dynamic water distribution method, device and equipment based on multi-modal perception
CN121365853A
Method and device for determining operation scheduling scheme of micro-service system and computer equipment
CN121501462A
PCM grid layout method and system
CN121615707A