A reinforcement learning-based method for pump and valve joint control in fluid transmission and distribution pipelines

By constructing a pump and valve control model and a digital twin system based on Markov decision processes, and combining it with deep reinforcement learning using the Actor-Critic framework, the problems of regulation lag and low efficiency in fluid transmission and distribution network control systems were solved, achieving collaborative optimization and high-efficiency energy-saving control of the pump and valve system.

CN120871646BActive Publication Date: 2026-01-06HANGZHOU ZETA TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511407346.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-01-06
Estimated Expiration
2045-09-29

AI Technical Summary

Technical Problem

When faced with complex nonlinear systems, frequent load fluctuations, and multivariable coupling, existing fluid transmission and distribution network control systems exhibit hysteresis, low efficiency, and difficulty in self-adaptation due to traditional control strategies. Furthermore, the application of reinforcement learning in this scenario suffers from high state dimensionality, poor action continuity, and safety control boundary issues.

Method used

A pump and valve control model based on Markov decision process is constructed. Combined with a digital twin system, deep reinforcement learning is carried out using the Actor-Critic framework. By introducing a reward function and meta-learning mechanism, the pump and valve joint control is realized. It has online learning and policy update capabilities, and the digital twin system is used for real-time data feedback and control command output.

Benefits of technology

It achieves coordinated optimization control of pump and valve systems, supports online learning and strategy updates, adapts to dynamic load changes, significantly saves energy and reduces consumption, enhances the robustness and safety of control, and is suitable for complex industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120871646B_ABST
    Figure CN120871646B_ABST
Patent Text Reader

Abstract

The present application relates to fluid distribution and process control technology, and aims to provide a fluid distribution network pump-valve joint debugging control method based on reinforcement learning. The method comprises the following steps: constructing a pump-valve control model based on Markov decision process, realizing safety guarantee of three-level time scale by introducing a reward function; carrying out deep reinforcement learning based on an Actor-Critic framework, continuously interacting with the running environment of the fluid distribution network without prior model, making the model learn and continuously train online until the strategy converges; collecting the state variables of the network in real time and simultaneously inputting them into the digital twin system and the pump-valve control model, realizing pump-valve joint debugging and collaborative control. The present application can realize collaborative optimization control of the pump-valve system, avoid coupling imbalance caused by separate control; support online learning and strategy updating, and has the ability to adapt to dynamic load changes; has a significant energy saving effect, and has strong adaptability to model uncertainty, disturbance and delay.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of fluid transmission and distribution and process control technology, specifically relating to a method for applying reinforcement learning to pump and valve joint control in fluid transmission and distribution pipeline networks, applicable to complex pipeline systems in industries such as water supply, heating, petrochemicals, and metallurgy. Background Technology

[0002] In existing fluid distribution network control systems, pumps and valves, as key actuators, directly affect the overall system's pressure stability and energy efficiency through their coordinated regulation capabilities. Common control strategies currently include proportional-integral control, fuzzy control, and manual setting methods based on empirical rules. While these methods possess some engineering adaptability, they often exhibit problems such as hysteresis, low efficiency, or difficulty in self-adaptation when facing complex nonlinear systems, frequent load fluctuations, and industrial scenarios with multivariate coupling. Furthermore, traditional methods rely on expert experience, making it difficult to achieve strategy self-updating and broad transferability.

[0003] With the development of deep learning and intelligent control theory, reinforcement learning has attracted attention due to its ability to learn policies automatically without prior models, and has been gradually introduced into fields such as energy system scheduling, intelligent transportation, and industrial process control. However, in the scenario of pump and valve linkage control in fluid transmission and distribution pipeline networks, the practical application of reinforcement learning methods is still in its early stages, facing technical challenges such as high state dimensionality, continuous actions, poor policy stability, and safety control boundaries.

[0004] Therefore, there is an urgent need for an intelligent control method that can integrate system modeling and autonomous optimization capabilities to meet the higher demands of modern industry for energy saving, high efficiency and stable operation. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a method for joint control of pumps and valves in fluid transmission and distribution pipelines based on reinforcement learning.

[0006] To solve the above-mentioned technical problems, the present invention adopts the following solution:

[0007] A method for integrated control of pumps and valves in a fluid transmission and distribution pipeline network based on reinforcement learning is provided. This method includes:

[0008] (1) Construct a pump and valve control model based on Markov decision process, using key operating state variables in the fluid transmission and distribution network and hydraulic feature vectors calculated in real time by the digital twin system as inputs, and pump frequency control and valve opening adjustment as joint control actions; by introducing a reward function that includes pressure supply error, energy consumption penalty and action smoothness, a three-level time scale safety guarantee is achieved;

[0009] (2) Deep reinforcement learning is carried out based on the Actor-Critic framework. The Actor network is used to generate control actions, and the Critic network is used to evaluate the quality of the strategy. The MAML framework for pre-training general initialization is embedded in the Critic network. Without the need for a prior model, the model continuously interacts with the operating environment of the fluid transmission and distribution network, so that the model can learn online and continuously train until the strategy converges, and obtain a model with the optimal control strategy.

[0010] (3) Deploy the digital twin system on a cloud server and deploy the trained pump and valve control model as a controller in the fluid transmission and distribution network; collect the state variables of the network in real time and use them as inputs to the digital twin system and the pump and valve control model; the digital twin system feeds back the hydraulic feature vector calculated in real time to the pump and valve control model, which outputs the control commands for pump frequency and valve opening after calculation, so as to realize the joint commissioning and coordinated control of pump and valve.

[0011] As a preferred embodiment of the present invention, the Markov decision process includes four elements: a state space, an action space, a state transition function, and a reward function. The state space takes the operating state variables in the fluid distribution network and the hydraulic characteristic vectors calculated in real-time by digital twins as inputs, including: pump frequency, valve opening, pressure and flow rate at key nodes, as well as dynamic impedance coefficients, pump similarity parabolic coefficients, valve flow characteristic coefficients, and the equivalent elastic cavity of the network. The action space is a continuous action space, including adjustment values ​​for pump frequency and valve opening. The state transition function is determined by the hydraulic model of the fluid distribution network and execution feedback. The reward function is defined as an immediate reward, used to encourage stable pressure supply, reduce energy consumption, and control drastic changes in action.

[0012] As a preferred embodiment of the present invention, the state space is represented as follows:

[0013] ;

[0014] in, It is a state-space function; For key nodes in fluid transmission and distribution pipeline networks Pressure at any given moment, MPa; Main pipeline Flow rate at any given time, m³ / h; For pumps The operating frequency at any given time, in Hz; for The valve opening at any given time, %; ΔH / ΔQ is the dynamic impedance coefficient. The coefficients of the similarity parabola for the pump; This refers to the valve's flow characteristic coefficient. It is the equivalent elastic cavity of the pipeline network.

[0015] As a preferred embodiment of the present invention, the action space is represented as follows:

[0016] ;

[0017] in, For the action space function, corresponding to Valve opening at any given time, % for The operating frequency of the pump at any given time, in Hz; The valve opening degree at the next moment, % This indicates that two levels of regularization are added to the output.

[0018] As a preferred embodiment of the present invention, the state transition function is expressed as:

[0019] ;

[0020] in, express The state at any given moment; This represents the defined transfer function. express The state at any given moment; express Valve opening at any given time, % This indicates environmental disturbances or model errors.

[0021] As a preferred embodiment of the present invention, the reward function is expressed as follows:

[0022] ;

[0023] in, for Momentary rewards; for The pressure of constant time; For target pressure; Energy consumption per unit time; This refers to the frequency change of the pump. This represents the change in valve opening. , , These are the weighting coefficients for the corresponding parameters; accordingly, This represents the pressure supply error. Represents energy consumption penalty, It represents the smoothness of movement.

[0024] As a preferred embodiment of the present invention, the Actor-Critic framework is optimized using the TD3 algorithm or the DDPG algorithm to achieve a target network soft update and policy delayed update mechanism, while introducing Ornstein-Uhlenbeck noise to improve exploration efficiency.

[0025] As a preferred embodiment of the present invention, in the Actor-Critic framework:

[0026] Actor networks are based on state-space functions. generate Valve opening at all times The control instructions; its objective function is:

[0027] ;

[0028] in, Represent the objective function; Represents the parameters of the Actor network; For mathematical expectation; The action value function is calculated by the Critic network; for The state at any given moment; For Actor networks;

[0029] Critic networks estimate value functions To evaluate the merits of the current strategy, the following loss function is used:

[0030] ;

[0031] in, Represent the objective function; These represent the parameters of the Critic network; For mathematical expectation; for Momentary rewards; Discount factor; The action value function; for The state at any given moment; for Valve opening at any given time, % for The state at any given moment; for The valve opening at any given time, %.

[0032] As a preferred embodiment of the present invention, the real-time acquisition of the pipeline network status variables is implemented through Modbus / TCP or 4-20mA industrial protocol; the control commands for pump frequency and valve opening are received and executed by the pump frequency converter and the electric valve actuator, respectively.

[0033] As a preferred embodiment of the present invention, the controller is deployed in a PLC or edge gateway of the fluid transmission and distribution network and connected to a cloud server through a communication module with networking capabilities.

[0034] Description of the invention principle:

[0035] 1. A Markov Decision Process (MDP) is a mathematical model used to describe how an agent makes decisions in an environment with Markov properties. It influences the environment by performing actions and adjusts its policy based on feedback (rewards) to achieve long-term goals. The Actor-Critic model mainly consists of two neural networks: an Actor (policy network) and a Critic (value function network). It combines the ideas of policy and value, balancing sample efficiency and learning stability, and is an important architecture in reinforcement learning. However, in the scenario of pump and valve linkage control in fluid transmission and distribution pipeline networks, if the Markov Decision Process (MDP) and Actor-Critic framework are applied in a textbook manner using conventional deep learning and intelligent control methods, the following shortcomings usually occur: directly treating pump frequency and valve opening as continuous actions leads to oscillations or divergence in the high-dimensional action space training; the pipeline network has large hydraulic inertia and significant delay, and traditional TD errors cannot accurately reflect the real physical causality, resulting in slow convergence and poor robustness; the reward function only considers pressure error and energy consumption squares, lacking engineering safety constraints such as water hammer, cavitation, and mechanical life, making it impossible to directly implement online; after training, the model becomes fixed and cannot adapt to seasonal drift of topology or load, requiring manual periodic retraining.

[0036] 2. This application departs from the conventional thinking of those skilled in the art and proposes a novel solution to the aforementioned technical application dilemmas. By constructing a system model based on a Markov decision process, key operating state variables in the pipeline network (such as pressure, flow rate, pump frequency, and valve position) are used as state inputs. The pump's variable frequency control signal and valve opening adjustment are used as joint control actions. Actor-Critic structures (such as TD3 and DDPG) in deep reinforcement learning are used for policy training. Targeted state representation, continuous action generation strategies, and energy consumption-oriented reward functions are designed, enabling the control system to learn the optimal joint control strategy online through continuous interaction with the environment without the need for a prior model.

[0037] 3. The following innovative practices are key features of the joint control strategy of this invention:

[0038] (1) Physically guided construction of hybrid state space

[0039] In addition to conventional data from sensors based on "pressure-flow-pump frequency-valve position," hydraulic characteristic vectors calculated in real time by digital twins are introduced; these include: dynamic impedance coefficient ΔH / ΔQ, pump similarity parabolic coefficient, etc. Valve flow characteristic coefficient Equivalent elastic cavity of pipeline network By embedding prior physical quantities into the state space, the effective dimensionality is significantly reduced and sample efficiency is improved.

[0040] (2) Action safety regularization

[0041] Two levels of engineering constraints are applied to the continuous action output of the Actor network: a Butterworth low-pass filter is used to limit the instantaneous rate of change of the pump frequency command and suppress water hammer.

[0042] (3) The reward function ensures safety across three time scales.

[0043] The reward function consists of three parts: millisecond-level safety barrier, second-level action smoothness, and minute-level energy efficiency. The weights are adaptively updated according to the valve wear model to ensure that dangerous areas can be avoided in the early stages of training.

[0044] (4) Meta-learning-hot migration online update mechanism

[0045] The Model-Agnostic Meta-Learning (MAML) framework is embedded in the Critic network to achieve general pre-training initialization; during the runtime phase, only the most recent 2 hours of data and 5 to 10 steps of gradient fine-tuning are needed to complete minute-level hot transfer without manual retraining.

[0046] Compared with the prior art, the beneficial effects of the present invention are:

[0047] 1. The present invention has the following significant technical advantages: it can realize the coordinated optimization control of pump-valve system and avoid the coupling mismatch caused by individual control; it supports online learning and strategy update and has the ability to adapt to dynamic load changes; it has significant energy saving and consumption reduction effect, and guides the control system to run towards the path of minimum energy consumption; it enhances control robustness and has strong adaptability to model uncertainty, disturbance and delay.

[0048] 2. The pump and valve joint control model based on reinforcement learning in this invention has high deployability and can be embedded in edge gateways or PLC devices with industrial interfaces to realize real-time status acquisition interfaces (such as 4-20mA pressure sensors, flow meter Modbus readings), control signal output interfaces (Modbus RTU / TCP protocol downlink to VFD and electric valves), can execute breakpoint recovery and model hot update mechanisms, and can support online learning and cloud strategy distribution.

[0049] 3. This invention is applicable to the following typical industrial scenarios: industrial water supply / cooling systems such as thermal power plants, chemical plants, and steel plants; urban centralized water supply and heating systems; long-distance oil / natural gas pipeline systems; blast furnace cooling water systems in the metallurgical industry; energy consumption optimization systems in intelligent buildings; and any scenario that requires joint adjustment of fluid dynamic parameters (pressure / flow rate). Attached Figure Description

[0050] Figure 1 This is a schematic diagram of the system operation of the pump and valve control model in this invention.

[0051] Figure 2 This is a flowchart for the closed-loop control of pump and valve linkage.

[0052] Figure 3 This is a schematic diagram of the Actor-Critic network operation.

[0053] Figure 4 This is a schematic diagram illustrating how the pressure control error changes during training.

[0054] Figure 5 This is a schematic diagram showing how energy consumption per unit changes with the number of training steps. Detailed Implementation

[0055] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the specific implementation of the present invention will be described below with reference to the accompanying drawings.

[0056] The reinforcement learning-based pump and valve joint control method for fluid transmission and distribution pipelines in this invention includes the following steps:

[0057] 1. Construct a pump and valve control model based on Markov decision process. The model takes the key operating state variables in the fluid transmission and distribution network and the hydraulic feature vector calculated in real time by the digital twin system as inputs, and takes the pump frequency control and valve opening adjustment as joint control actions. By introducing a reward function that includes pressure supply error, energy consumption penalty and action smoothness, a three-level time scale safety guarantee is achieved.

[0058] This invention abstracts the pump and valve control system as a Markov Decision Process (MDP), mainly comprising four elements: State Space, Action Space, State Transition Function, and Reward Function. The State Space takes the operating state variables of the fluid distribution network and the hydraulic characteristic vectors calculated in real-time by digital twins as inputs, including: pump frequency, valve opening, pressure and flow rate at key nodes, as well as dynamic impedance coefficients, pump similarity parabolic coefficients, valve flow characteristic coefficients, and the equivalent elastic cavity of the network. The Action Space is a continuous action space, including the adjustment values ​​of pump frequency and valve opening. The State Transition Function is determined by the hydraulic model of the fluid distribution network and execution feedback. The Reward Function is defined as an immediate reward, used to encourage stable pressure supply, reduce energy consumption, and control drastic changes in action.

[0059] (1) The state space is represented as:

[0060] ;

[0061] in, It is a state-space function; For key nodes in fluid transmission and distribution pipeline networks Pressure at any given time (e.g., main network or terminal pressure), MPa; Main pipeline Flow rate at time (i.e., the current time, the same below), m³ / h; For pumps The operating frequency at any given time, in Hz; for The valve opening at any given time, %.

[0062] This invention, based on a digital twin system and utilizing data from conventional sensors based on "pressure-flow-pump frequency-valve position" for real-time calculation, obtains four hydraulic characteristic vectors as prior physical quantities, including: dynamic impedance coefficient ΔH / ΔQ, pump similarity parabola coefficient, etc. Valve flow characteristic coefficient Equivalent elastic cavity of pipeline network Embedding these prior physical quantities into the state space can significantly reduce the effective dimensionality and improve sample efficiency.

[0063] The calculation method for the four hydraulic characteristic vectors described in this invention can be found in published literature, such as "(Coulbeck, Bryan (1977) Optimisation and modelling techniques in dynamic control of water distribution systems. PhD thesis, University of Sheffield.", or "Zhang Mingyuan, Jing Sirui, and Li Guojun (eds.), Advanced Engineering Fluid Mechanics, Xi'an Jiaotong University Press, 2006". The digital twin system used in this invention and its specific solution process can be found in published literature, such as "Lu Jianfeng, Zhang Hao, and Zhao Rongyong. Digital Twin Technology and Engineering Practice—Model- and Data-Driven Intelligent Systems. 2022". Since these contents are not part of the innovative technology of this invention, they will not be elaborated further.

[0064] (2) The action space is represented as:

[0065] ;

[0066] in, For the action space function, corresponding to Valve opening at any given time, % for The operating frequency of the pump at any given moment (i.e., the next moment, the same below), in Hz; The valve opening degree at the next moment, % This indicates that two levels of regularization are added to the output.

[0067] This section covers action safety regularization, specifically applying engineering constraints to the continuous actions output by the Actor network. In practical implementation, a Butterworth low-pass filter can be used to limit the instantaneous rate of change of the pump frequency command, thereby suppressing water hammer.

[0068] Because it is a continuous motion space, this setting meets the industrial site's need for fine-grained control.

[0069] (3) The state transition function is expressed as:

[0070] ;

[0071] in, express The state at any given moment; This represents the defined transfer function. express The state at any given moment; express Valve opening at any given time, % This indicates environmental disturbances or model errors.

[0072] This setting is used to reflect the randomness of the system.

[0073] (4) The reward function is expressed as:

[0074] ;

[0075] in, for Momentary rewards; for The pressure of constant time; For target pressure; Energy consumption per unit time; This refers to the frequency change of the pump. This represents the change in valve opening. , , These are the weighting coefficients for the corresponding parameters; accordingly, This represents the supply pressure error, used to achieve a millisecond-level safety barrier. This represents an energy consumption penalty, used to achieve minute-level energy efficiency. Represents motion smoothness, used to achieve second-level motion smoothness.

[0076] 2. Deep reinforcement learning is performed based on the Actor-Critic framework. The Actor network is used to generate control actions, and the Critic network is used to evaluate the quality of the strategy. The Critic network embeds the MAML framework for pre-training general initialization. Without the need for a prior model, it continuously interacts with the operating environment of the fluid transmission and distribution network, enabling the model to learn online and continuously train until the strategy converges, thus obtaining a model with the optimal control strategy.

[0077] In this invention, the Actor-Critic framework employs either the TD3 or DDPG algorithm for optimization, achieving policy optimization by maximizing the value of actions within the current Critic network. While utilizing a target network soft update and policy delayed update mechanism, Ornstein-Uhlenbeck noise is incorporated to enhance exploration capabilities, ensuring convergence stability and training efficiency. Specifically,

[0078] Actor networks are based on state-space functions. Generate the current valve opening. The control instructions; its objective function is:

[0079] ;

[0080] in, Represent the objective function; Represents the parameters of the Actor network; For mathematical expectation; The action value function is calculated by the Critic network; for The state at any given moment; This is an Actor network.

[0081] Critic networks estimate value functions To evaluate the merits of the current strategy, the following loss function is used:

[0082] ;

[0083] in, Represent the objective function; These represent the parameters of the Critic network; For mathematical expectation; for Momentary rewards; Discount factor; The action value function; for The state at any given moment; for Valve opening at any given time, % for The state at any given moment; for The valve opening at any given time, %.

[0084] This invention proposes embedding a Model-Agnostic Meta-Learning (MAML) framework into the Critic network for general initialization during pre-training. General initialization refers to pre-training the initial parameters of the Critic network through meta-learning, enabling it to quickly adapt to the dynamic characteristics of different fluid distribution networks (such as impedance changes, pump and valve parameter differences, load fluctuations, etc.).

[0085] Based on this design, the present invention can obtain a model with the optimal control strategy by continuously interacting with the operating environment of the fluid transmission and distribution network without the need for a prior model.

[0086] 3. Deploy the digital twin system on a cloud server and deploy the trained pump and valve control model as a controller in the fluid transmission and distribution network; collect the state variables of the network in real time and use them as inputs to both the digital twin system and the pump and valve control model; the digital twin system feeds back the hydraulic feature vectors calculated in real time to the pump and valve control model, which then outputs control commands for pump frequency and valve opening after calculation, thereby realizing pump and valve joint commissioning and coordinated control.

[0087] In this invention, the digital twin system leverages real-time data interaction and analysis to achieve dynamic simulation, prediction, and optimization of the entire lifecycle of physical entities. This places high demands on the hardware system; therefore, this invention deploys the digital twin system on a cloud server. The trained pump and valve control model, acting as a controller, is deployed in the PLC or edge gateway of the fluid distribution network. The controller connects to the cloud server via a network-enabled communication module, enabling data interaction with the digital twin system, online fine-tuning and learning, cloud model updates, and breakpoint recovery, thereby enhancing the intelligence and adaptability of the control system. Real-time acquisition of the network's state variables is achieved through Modbus / TCP or the 4-20mA industrial protocol. Control commands for pump frequency and valve opening are received and executed by the pump's frequency converter and the electric valve actuator, respectively.

[0088] Figure 1 The diagram illustrates the system operation of a pump and valve control model. The "environment" in the diagram refers to the critical operating states within the fluid distribution network, serving as inputs to the Markov decision process in the controller. After calculation, the controller outputs the pump frequency and valve opening, which are then executed by frequency converters and electric valve actuators within the fluid distribution network. The adjusted system operating state is then re-inputted to the controller after signal acquisition.

[0089] Figure 2 The flowchart of the pump-valve integrated closed-loop control is shown below:

[0090] (1) Initialize the system;

[0091] (2) Collect key operating status data in the current fluid transmission and distribution network;

[0092] (3) Input the state variables into the Actor-Critic network and output the current valve opening. Control commands;

[0093] (4) The controller sends instructions to the pump's frequency converter and electric valve actuator;

[0094] (5) Pipeline operation and feedback New state of mind and rewards Used to train and update the Critic-Actor network;

[0095] (6) Repeat the above process to achieve adaptive closed-loop control.

[0096] Figure 3 The diagram shows the operation of the Actor-Critic network.

[0097] Below is a specific verification example:

[0098] To further verify the effectiveness of the reinforcement learning-based pump and valve joint control method for fluid transmission and distribution pipelines proposed in this invention, the applicant constructed a simulation platform based on Matlab. Figure 1 , 3 The Critic-Actor network and pump-valve control model shown are used as references for the deployment of the digital twin system in the publicly available literature "Lu Jianfeng, Zhang Hao, Zhao Rongyong, Digital Twin Technology and Engineering Practice - Model and Data-Driven Intelligent Systems, 2022".

[0099] The control objective is to stabilize the main pressure and flow of the water supply network. The parameter settings and procedures are as follows:

[0100] (1) Environment and parameter settings

[0101] Initial state: The system is initialized to a medium-pressure, low-flow state (pressure 0.3MPa, flow rate 110 m³ / h).

[0102] Target pressure setting: 0.35 MPa; Target flow rate setting: 120 m³ / h.

[0103] Actor-Critic network structure: The TD3 algorithm is used. Both the policy network and the evaluation network are two-layer fully connected neural networks with 64 neurons in each layer and ReLU activation function.

[0104] Exploration strategy: Add Ornstein-Uhlenbeck process noise to perturb the motion;

[0105] Training rounds: 100,000 steps;

[0106] Reward function weight parameters: = 1.0 (pressure supply error); = 0.01 (energy consumption penalty); = 0.1 (motion smoothness).

[0107] (2) Performance indicators and results

[0108] The system control error converged rapidly within 30,000 steps, and the final control error stabilized at ±0.01 MPa; the valve adjustment process was smooth and there were no drastic fluctuations; the system response time was shortened from 12 seconds under traditional PID control to 7 seconds; the energy consumption assessment model showed that the total energy consumption decreased by about 13%.

[0109] Figure 4 The data shows that the pressure control error gradually decreases and converges during the training process; Figure 5 The figure shows the trend of unit energy consumption gradually decreasing with the number of training steps. From the data changes in the figure, it can be seen that the pump-valve integrated control method of the present invention has the following significant advantages:

[0110] (1) High control accuracy: The pressure control error converges rapidly during training (within 30,000 steps) and eventually stabilizes at ±0.01 MPa. Figure 4 As shown in the figure, it is significantly better than traditional control methods.

[0111] (2) Strong robustness and stability: The valve opening adjustment process is smooth and there are no violent fluctuations, which effectively avoids water hammer effect and ensures the safety of pipeline network.

[0112] (3) Fast dynamic response: The system response time is reduced from 12 seconds in traditional PID control to 7 seconds, improving the regulation efficiency by 42%.

[0113] (4) Significant energy-saving effect: Energy consumption per unit decreases continuously with the number of training steps ( Figure 5 As shown in the figure, the total energy consumption was reduced by about 13%, which verifies the effectiveness of the energy consumption penalty term in the reward function.

[0114] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A fluid distribution network pump-valve co-simulation control method based on reinforcement learning, characterized in that, The application relates to a pump-valve control model based on a Markov decision process, which takes key operating state variables in a fluid distribution network and a hydraulic characteristic vector calculated by a digital twin system as inputs, takes frequency control of a pump and adjustment of a valve opening as joint control actions, and realizes safety guarantee in three time scales by introducing a reward function containing a supply pressure error, an energy consumption penalty and action smoothness. The application also relates to deep reinforcement learning based on an Actor-Critic framework, which generates control actions by using an Actor network and evaluates the advantages and disadvantages of a strategy by using a Critic network; the Critic network is embedded with a MAML framework for pre-training general initialization, and the model is learned and trained online by continuously interacting with the operating environment of the fluid distribution network without prior models, so that the model with an optimal control strategy is obtained. The target function of the Actor-Critic framework is: The application further relates to a digital twin system deployed in a cloud server and a trained pump-valve control model deployed in the fluid distribution network as a controller; state variables of the network are collected in real time and used as inputs of the digital twin system and the pump-valve control model; the digital twin system feeds back a hydraulic characteristic vector calculated in real time to the pump-valve control model, and the pump-valve control model outputs control instructions of pump frequency and valve opening after operation, so that pump-valve joint debugging and collaborative control are realized. Actor network based on state space function Generate Valve opening at the moment Control command; The Markov decision process includes four elements, namely a state space, an action space, a state transition function and a reward function; the state space takes operating state variables in the fluid distribution network and a hydraulic characteristic vector calculated by the digital twin in real time as inputs, and includes pump frequency, valve opening, pressure and flow of key nodes, dynamic impedance coefficient, pump similar parabolic coefficient, valve flow characteristic coefficient and pipe network equivalent elastic cavity; the action space is a continuous action space, and includes adjustment values of pump frequency and valve opening; the state transition function is determined by a hydraulic model of the fluid distribution network and execution feedback; and the reward function is defined as an immediate reward for encouraging stable supply pressure, reducing energy consumption and controlling action violent change. ; wherein, represents the objective function; represents the parameters of the Actor network; is the mathematical expectation; is the action value function, computed by the Critic network; is the state at time instant; the Actor network; The critic network evaluates the current policy by estimating the value function which is expressed as the following loss function: ; wherein, represents an objective function; represents parameters of the Critic network; is a mathematical expectation; is a reward at time instant; is a discount factor; is an action value function; is a state at time instant; is a valve opening at time instant; is a state at time instant; is a valve opening at time instant; The Actor-Critic framework adopts a TD3 algorithm or a DDPG algorithm for optimization, realizes a target network soft update and a policy delay update mechanism, and introduces Ornstein-Uhlenbeck noise to improve exploration efficiency.

2. The method of claim 1, wherein, The state variables of the network are collected in real time by Modbus / TCP or 4-20mA industrial protocols; and the control instructions of the pump frequency and the valve opening are received and executed by a frequency converter of the pump and an electric valve actuator respectively.

3. The method of claim 2, wherein, The state space is represented as: ; in, It is a state-space function; For key nodes in fluid transmission and distribution pipeline networks Pressure at any given moment, MPa; Main pipeline Flow rate at any given time, m³ / h; For pumps The operating frequency at any given time, in Hz; for The valve opening at any given time, %; ΔH / ΔQ is the dynamic impedance coefficient. The coefficients of the similarity parabola for the pump; This refers to the valve's flow characteristic coefficient. It is the equivalent elastic cavity of the pipeline network.

4. The method of claim 2, wherein, The action space is represented as: ; wherein, is the action space function corresponding to is the valve opening at time instant is the is the pump operating frequency at time instant is the valve opening at next time instant indicates that two levels of regularization are added at the output.

5. The method of claim 2, wherein, The state transition function is represented as: ; wherein represents the state at time instant represents a defined transfer function, represents the state at time instant represents the valve opening at time instant represents an environmental disturbance or model error.

6. The method of claim 2, wherein, The reward function is represented as: ; wherein, is a reward at the moment; is a stress at the moment; is a power consumption per unit time; is a change in opening of the valve; , , are weight coefficients of the corresponding parameters, respectively, and represents a supply pressure error, represents a power consumption penalty, represents an action smoothness.

7. The method of claim 1, wherein, The controller is deployed in a PLC or an edge gateway of the fluid distribution network and connected to the cloud server through a communication module with networking function.

8. The method of claim 1, wherein, ​ 9. The method of claim 1, wherein, ​

Citation Information

Patent Citations

  • Energy-saving dispatching method and system for pump station water supply system, electronic equipment and storage medium

    CN115526504A

  • Tobacco enterprise intelligent production composite scheduling method and system

    CN118295353A