Pump control method and control system for a water treatment plant

The integration of simulation and Reinforcement Learning techniques in pump control systems addresses inefficiencies in water treatment plants by optimizing pump operation and reducing energy consumption through real-time adaptation to dynamic inflows, improving efficiency and safety.

WO2026062308A1PCT designated stage Publication Date: 2026-03-26ACCIONA AGUA SAU
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2026-03-26

AI Technical Summary

Technical Problem

Existing pump control systems in water treatment plants face inefficiencies due to reliance on historical data for Reinforcement Learning (RL) models, which do not cover the entire state space, leading to errors and suboptimal pump efficiency, especially when dealing with unknown and dynamic water inflows.

Method used

A method combining simulation and Reinforcement Learning (RL) techniques, using neural networks and predefined heuristics to optimize pump operation, ensuring safe and efficient control by predicting water flow and preventing cavitation, with a control system that includes a graphical interface for real-time monitoring and adjustment.

Benefits of technology

The method enhances pump efficiency and safety by dynamically adapting to changing conditions, reducing energy consumption and maintaining stable tank levels, outperforming traditional PI controllers in real-world scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure ES2025070512_26032026_PF_FP_ABST
    Figure ES2025070512_26032026_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to a method for controlling pumps in a water treatment plant comprising a first tank and a second tank located at a higher level, comprising: obtaining a set of configuration data for the pumps and the plant; obtaining estimates of inlet water flow from historical data and a change in frequency; simulating the behaviour of the pumps by determining an operating point, efficiency and power consumed; optimising the operating point by applying reinforcement learning (RL) techniques, using neural networks trained with simulation values and / or historical data, using, as hyperparameters, error metrics of the level of the first tank and frequency variation after an action; and applying monitoring techniques to prevent the first tank from emptying or overflowing, and to ensure that the operating point of the pumps is safe.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] PUMP CONTROL METHOD AND CONTROL SYSTEM FOR A WATER TREATMENT PLANT DESCRIPTION OBJECT OF THE INVENTION The invention focuses on pumping systems and more particularly on methods for controlling said pumping systems. The object of the present invention is a pump control method in a water treatment plant that allows for optimizing pump performance based on an unknown and dynamic variable.the inflow of water. The invention also relates to a control system incorporating the method of the invention for a treatment plant comprising a set of pumps. BACKGROUND OF THE INVENTION Studies for optimizing energy consumption in pumping systems have been carried out in the prior art. The greedy strategy currently deployed in many plants uses a series of pumps in parallel to maintain a constant water level in the storage tank. Control is achieved by regulating the variable frequency of the pumps and activating or deactivating them based on the water level in the tank, which depends on the inflow of water. There are works that address the control of multiple pumps with variable frequency speeds using Reinforcement Learning (RL). In these works, the control is learned from a plant model generated from historical data. In this approach,Errors in model learning contribute to the optimization of Reinforcement Learning (RL). However, these solutions only consider the application of a Proximal Policy Optimization (PPO) algorithm. On the other hand, Reinforcement Learning (RL) algorithms generally operate through trial and error and need to be able to explore the entire space of possible combinations to determine the optimal strategy. Learning a model from historical data has the limitation that the collected data will only cover a subset of the state space, since, due to security and operational constraints, it is not possible to record system data in every situation. Therefore,In these solutions, errors increase, decreasing the precision in optimizing pump efficiency. DESCRIPTION OF THE INVENTION The invention relates to a pump control method for pumping in a water treatment plant. The treatment plant comprises a first tank and a second tank, the second tank being located at a higher level than the first tank. The method of the invention comprises the steps of: - obtaining a configuration data set for the pump assembly, the treatment plant tanks, and a set of pipes connecting the tanks; - obtaining estimates of the incoming water flow, which is unknown and varies over time;Based on historical plant data and a desired change in frequency; - perform a simulation of the pump behavior for pumping water from the first tank to the second tank by following these steps: ● determine a system curve relating the static head of the pumps to the flow using flow estimation equations, such as Manning's, as a function of the level of the first tank; ● obtain a characteristic curve for each pump at all frequencies, from the characteristic curve at a specific frequency using affinity laws; ● determine an operating point for the pumps as the intersection point between the system curve and the characteristic curves, determining a pumped water flow; ● determine a pump efficiency at this operating point and a power consumption; - determine an optimal operating point by following these steps: ● apply Reinforcement Learning (RL) techniques,using neural networks trained with simulation values ​​and / or historical plant data, using as hyperparameters: error metrics, between historical data and simulation values, of the first tank level and pump frequency variation after an action, as a function of the incoming water flow, the first tank level, and the pump status; and ● applying monitoring techniques for the first tank level and incoming water flow, preventing the first tank from emptying or overflowing, and ensuring that the pump operating point is above the NPSH (Net Positive Suction Head) curve; - updating the first tank level. In particular, in the stage of obtaining incoming water flow estimates, data on the first tank level and pump operating frequency can be used. The incoming water flow is a stochastic value whose estimation and prediction are inherently complex. On the other hand,In the stage of applying monitoring techniques, predefined heuristics and strategies can preferably be used to correct the pump configuration and bring them to a safe state. This allows for safe operation by preventing cavitation in the pumps. Preferably, the metric defined for calculating the hyperparameters can be the sum of a water tank level error metric and a frequency variation error metric of the form: where, where is a level value of the first tank defined by the simulation, is the working frequency of the pumps at an instant, is the working frequency of the pumps at an instant after t. Additionally, the following formula can be applied during the tank level update stage: where L is the tank level, V is the volume of water in the tank, V total the total volume of the tank, Q in the inflow, QoutOutflow, ∆t the time interval between two consecutive measurements. Preferably, the setup data includes: - System Curve: the first and second term coefficients of the system curve, which describe how the system resistance varies as a function of the water flow; - Tank Volume and Surface Area and Maximum Static Head of the system; - Characteristic Curve of each pump at a fixed frequency; and - Efficiency Surface of the pumps at discrete frequency intervals. Historical data may include: - the level of the first tank; - the operating frequencies of the pumps, which, together with the level, allow estimation of the inflow; and - the power consumed by the pumps, which allows verification of the simulator's operation. Thus, the simulation can obtain, among other things: - Tank Level: as it receives the incoming wastewater flow.- System Curve: based on tank level and other factors to reflect real-time operating conditions. - Pump Curves at Different Frequencies (Hz): model pump behavior at varying speeds. - Outlet Water Flow: the intersection of the pump curves and the system curve to estimate real-time outlet water flow. - Pump Efficiency and Power: estimates the efficiency and power of each pump based on its current operation and conditions. This efficiency can be calculated for a continuous frequency range using multivariable regression. This regression allows estimating pump efficiency based on multiple variables, such as pump speed and system conditions, which is essential for evaluating pump performance and efficiency under unsampled conditions.Through the simulation stage, the system can make mistakes and adjust its strategies before being deployed in the actual plant, ensuring more reliable and safer operation. Furthermore, simulation allows the system to adapt dynamically to changing conditions, thus improving its ability to optimize pump operation in real time. Preferably, the simulation can be performed using the following relationships and expressions that govern the operation of the plant and its pumps. In the stage of determining a system curve, the following expression can be used: where is the head. The highest ethical standard is the head tank level at time t, ybyc are coefficients of the system curve polynomial obtained for flow values ​​determined by the following Manning equation: where n is Manning's coefficient, A is the pipe area, R is the hydraulic radius, and S is the slope energy gradient. These relationships allow the determination of a dynamic head loss. ௗ as a function of the flow rate (Q). For its part, in the stage of obtaining a characteristic curve of the pumps, the following expression can be used: where a, b, c, and d are coefficients of the characteristic curve polynomial. To obtain a characteristic curve for each pump at all frequencies, affinity laws of the form N can be used, where N is the rotational speed. of a pump motor (in rpm), Q are the flows, H the heads and P the powers, and where the rotation speed is obtained: where are face engine characteristics. The stage of determining an operating point can be carried out using the following expressions: For the case where there is only one pump, it is equal to in t. Pump efficiency depends on the pump's frequency and flow rate. Efficiency can be determined using previously obtained characteristic data at a specific frequency or frequencies (provided by the manufacturer) and using a multivariable regression model to determine an efficiency surface at any given frequency and flow rate. Subsequently, the power consumption is calculated, which can be determined from the efficiency using a motor power curve or using the following expression: On the other hand, the pumping well may comprise multiple pumps located in parallel with separate or shared piping. If the piping is separate, then the following holds true: the outlet flow is equal to the system flow. If the pipes are shared, then: In this case, the operating point is obtained through an iterative process by increasing the system head from 0 until the system curve intersects the characteristic curves, estimating for each system head value the total water flow, as the sum of the pump flows: and looking for the intersection with the system curve for which the total flow differs from the flow of the system curve by a value less than a required accuracy: Regarding Reinforcement Learning (RL), neural networks can be trained continuously with new data measured at the plant and changes in the inflow distribution of water; or, alternatively, in a loop, recording new inflow data to update the simulation and determining a new policy within that simulation, ensuring it is safe and efficient. Reinforcement Learning (RL) can be applied in this case using off-policy techniques, which focus on learning from past experiences, selected from Q-learning and G-learning, or on-policy techniques, which focus on learning from the ongoing interaction, selected from SARSA and Actor-Critical.Preferably, Reinforcement Learning (RL) can be defined as a Markov decision process in which a state space S, an action space A, a transition function space P (which stores the probability of transitioning from state s to state s by means of action a), and a reward function for action a are defined. In this application, the state space (S) defines possible values ​​for the tank level (L), the frequencies of each pump, and the change in tank level between two steps. de tiempo Preferably, the frequencies take discrete values ​​in the range [1, 50] Hz, the tank level values ​​are continuous in the range [0, 100], and the tank level change values ​​are continuous in the range [-100, 100]. The action space (A) can be defined as a frequency variation applied to each pump. Preferably, the frequency variation is applied using the values ​​{-1, 0, 1}, and the frequency is kept in the range [0, 50]. The transition function can be defined as: The reward function can be defined, in this case, as a penalty term; if the tank level is within an acceptable range, it is the negative sum of the bomb power, and if not, it is a constant value -c, which is the maximum penalty for violating the rule that the first tank level is too high or too low: where is the power of the pump By default, there is a constant penalty for the maximum power consumption value. pump In the stage of determining an optimal operating point, an estimate can be used. of the optimal action-value function that determines the future reward of action a, of the form: where r is the learning area and the action that gives a maximum value of Q is selected: Preferably, a greedy-scanning strategy can be used to randomly select the action and prevent local minima. More preferably, neural networks can use a Deep Q Network (DQN) technique that defines an action-value function as ^^ ఏ with parameters ^^ and updates them using a stochastic gradient that descends with the square of the TD-error loss, in the form: where the gradient is: where D is a dataset of transitions, is a copy of updated to a shorter timescale, and ^^ is a discount factor that reduces the impact of future rewards. Furthermore, Reinforcement Learning (RL) can also comprise an experience replay technique and the use of two networks with the same structure: a Q-network and a target network. The experience replay technique involves storing transitions in memory and sampling subsets of transitions. In these implementations, a sample error (e) can be defined as the distance between Q(s, a) and a target. being T and make error a priority for each sample, in the following way: being a small positive constant and where It controls the relative difference between a high and low error, and where Reinforcement Learning (RL) is performed with transitions according to this probability: The target network technique involves the use of two networks of the same structure, but with different weights, ^^ for the network- for the target network, such that the Q-network is updated regularly while the target network is updated to the parameters of the Q-network every C time steps. More preferably, the neural networks can use a Double Deep Q Network (DDQN) technique with a function: and where the Q-network is used to select the action and the target network is used to evaluate the action. Alternatively, the method of the invention can use a PPO technique for neural networks that approximates the advantage function, dependent on advantage parameters, and the policy, dependent on policy parameters, using neural networks. In this case, in iteration k, the policy parameters are updated as follows: where are the policy parameters without updating, is the objective associated with the policy Under the unupdated policy parameters, the Kullback-Leibler divergence between stock policies The astics controls the influence of the KL term. Therefore, the new parameter vector k+1 is calculated by maximizing the objective while at the same time remaining close to the old policy. The PPO target is defined as the policy's trimmed gradient: d where (st,at) are the state and action in time is the estimated advantage associated with the old policy the state-action pair (st,at) according to the current parameters is a clipping function. In this approximation, for a given state represents a probability distribution over the possible actions for the current parameter vector θ, and the action to be taken is selected by randomly drawing an action from the probability distribution. The invention also relates to a control system for a water treatment plant comprising: - a set of pumps; - a first tank; and - a second tank, located at a higher level than the first tank; wherein the control system comprises: - a pump control module connected to the pumps and configured to carry out the method of the invention described. Preferably, the control system may further comprise a graphical interface configured to display the system curve, the pump characteristic curves and their operating point at any given time, a tank level value, an inlet water flow value, the frequency, efficiency, and power of the pumps, and an outlet flow value.BRIEF DESCRIPTION OF THE DRAWINGS To complement the description provided and to aid in a better understanding of the features of the invention, according to a preferred embodiment thereof, a set of diagrams is attached, in which, for illustrative and non-limiting purposes, the following is represented: Figure 1 shows an example embodiment of the water treatment plant of the invention. Figure 2 shows the pump characteristic curve at 50 Hz and other frequencies calculated using affinity laws. Figure 3 shows a system curve and the pump characteristic curves at two frequencies. Figure 4 shows a calculated efficiency surface for the pump. Figure 5 shows efficiency curves at different frequencies. Figure 6 shows an example of what would be displayed in a graphical interface. Figure 7 shows the reward function for three different approximations over time.Figure 8 shows the tank level for three different approximations over time. Figure 9 shows the power consumed for three different approximations over time. Figure 10 shows the evolution of the frequency obtained for three different approximations over time. Figure 11 shows the efficiency obtained for three different approximations over time. Figure 12 shows the evolution of the water inlet flow in a real plant. Figure 13 shows the evolution of the water inlet flow for use case 1 of the invention's method with RL, trained with a PPO policy. Figure 14 shows the evolution of the tank level for use case 1 of the invention's method with RL, trained with a PPO policy. Figure 15 shows the evolution of the power for use case 1 of the invention's method with RL, trained with a PPO policy.Figure 16 shows the evolution of specific consumption for use case 1 of the invention method with RL, trained with a PPO policy. Figure 17 shows the evolution of the water inlet flow for use case 2 of the invention method with RL, trained with a PPO policy. Figure 18 shows the evolution of the tank level for use case 2 of the invention method with RL, trained with a PPO policy. Figure 19 shows the evolution of power for use case 2 of the invention method with RL, trained with a PPO policy. Figure 20 shows the evolution of specific consumption for use case 2 of the invention method with RL, trained with a PPO policy. Figure 21 shows the evolution of the number of active pumps for use case 2 of the invention method with RL, trained with a PPO policy.Figure 22 shows the evolution of the incoming water flow for use case 3 of the invention's method with RL, trained with a PPO policy. Figure 23 shows the evolution of the tank level for use case 3 of the invention's method with RL, trained with a PPO policy. Figure 24 shows the evolution of the power for use case 3 of the invention's method with RL, trained with a PPO policy. Figure 25 shows the evolution of the specific consumption for use case 3 of the invention's method with RL, trained with a PPO policy. Figure 26 shows the evolution of the number of active pumps for use case 3 of the invention's method with RL, trained with a PPO policy. PREFERRED EMBODIMENTS OF THE INVENTION The invention relates to a control method for pumps in a water treatment plant that incorporates a simulation stage and an optimization stage using reinforcement learning.Regarding the simulation stage, the goal is for the obtained values ​​to approximate the operation of a pump in a real plant. To incorporate Reinforcement Learning (RL) algorithms, the aim is to gather extensive experience data to calculate a near-optimal policy. This experience involves simulating transitions in the form (S, A, R, S'), where S is the current state, A is the action chosen by the current decision strategy in state S, R is the reward received by the system for applying action A in state S, and S' is the new state resulting from applying action A in state S. Depending on the problem size, the number of transitions required to calculate a near-optimal policy can be on the order of several million. Therefore, accumulating experience from the operation of a real pump is highly inefficient, as the learning process would be excessively slow.Instead, the Reinforcement Learning (RL) algorithm in the present invention is trained in a simulator that determines the reward R and the next state S′ corresponding to a given state S and a chosen action A. It is important that the simulator's response time be as fast as possible, as this means the system can accumulate experience more quickly, and consequently, the Reinforcement Learning (RL) algorithms can also be trained more quickly. Thus, in the present invention, a first tank that collects incoming wastewater and a centrifugal pump that pumps water from the first tank to a second tank in a water treatment plant are simulated. The simulation of a single pump is described below and then extended to the simulation of multiple pumps. First, a time interval for the simulation is defined.In this case, the pump operation has been simulated at fixed 1-second intervals due to the rate of change in the tank level, allowing for precise plant control. This time interval is a configurable hyperparameter of the simulator. Every second, the system can choose a new action to perform, which in turn determines the pump's current operating frequency. The current tank level is then updated based on the current incoming flow and the pump's water flow rate corresponding to the operating frequency. The current tank level is approximated at each time step t. Specifically, the tank level ^^. ௧ at time t depends on the tank level at the previous moment, and is defined as: where It is the variable that describes the flow of water entering the tank, the flow of water from the pump from the first tank to the second tank that of the plant, ∆t is the simulation time interval, V total is the total volume of the tank. The simulator repeatedly updates the level d the tank at each time step. To model the incoming flow of wastewater The approach is to use historical data recorded from a real wastewater treatment plant, for example, a plant in La Almunia, Spain, or in Burgos, Spain. The simulation of the outflow is described below. The system curve describes the relationship between the head (H) and the water flow (Q) of the system. The head (H) is the difference in height between the level of the first tank and the water level in the plant, i.e., in the second tank. The water level in the plant, in the second tank, is constant in this case, so the head (H) depends directly on the level of the first tank (L). The system curve, in this case, is estimated by measuring the head (H) at different water flows (Q). Given the plant design in Figure 1, the dynamic head loss H is estimated. d for some points (Q, H) d ) using the following Manning equation: where n is Manning's coefficient, A is the cross-sectional area of ​​the pipe, R is the hydraulic radius, and S is the energy gradient of the slope. From this equation, the following points are obtained: From these, the system's curve can be estimated as a second-degree polynomial: where is the maximum static head, is the tank level head at time t, and the constants b and c are the coefficients of the polynomial that can be estimated by solving the following linear system of equations: To simulate the operation of a pump, the pump's characteristic curve is analyzed, deriving its operating point and estimating its efficiency and energy consumption. The pump's characteristic curve describes the relationship between the pump's head (H) and water flow rate (Q). It can be estimated by measuring the head (H) at different flow rates (Q). This data is usually provided by the pump manufacturer in the form of a table or graph for a fixed frequency (f). The characteristic curve can be estimated as a third-degree polynomial of the form: Using a third-degree polynomial allows for a high degree of accuracy while maintaining simplicity in calculating its roots. Given the points (Q, H) measured from the manufacturer's table, the coefficients a, b, c, and d can be estimated by solving the linear system: Thus, the pump characteristic curve can be obtained at the frequency used to measure the points (Q, H). However, to estimate the pump characteristic curve at any frequency (f), the following affinity laws can be used: where N1 and N2 are the rotational speeds of a pump motor in rpm, Q1 and Q2 are the flow rates, H1 and H2 are the heads, and P1 and P2 are the corresponding powers. The motor's rotational speed is obtained from the frequency (f) as follows: donde These are characteristics of the motor. Next, the coefficients a, b, c, d can be derived from the pump characteristic curve for any frequency f, by solving the following system of equations:

[0002] The operating point is the point of intersection between the system curve and the pump characteristic curve. In the case of a single pump, the operating point can be derived by finding the root of the polynomial: Where: It is the flow at the operating point. The operating point of head H op It can be derived simply from the pump curve as: In the case of a single pump, the total water that flows from the first tank to the second tank of the plant at time t is: Now we proceed to estimate the efficiency of the pump given the frequency (f) and the water flow (Q) f ). Efficiency can be estimated using efficiency curves / tables provided by the pump manufacturer. As shown in Figure 5, efficiency is a function of frequency (f) and water flow rate (Q). f ). The missing values ​​of These can be filled using a multivariable regression model to fit a surface to the data previously provided by the manufacturer. The fitted model can then be used to estimate efficiency. ^ for any frequency (f) and water flow (Q) f Next, the power (P) of the pump is estimated given the water flow (Q). f ) to the frequency (f) and the head (H) f This can be done using efficiency the motor's power curve. To do this, we first estimate the operating point for the pump at the frequency (f). Additionally, it is possible to estimate the electrical power ( ) of the pump at the frequency (f) as: The following describes how to extend the simulation to multiple pumps in parallel, with separate or shared piping. When simulating multiple pumps in parallel with separate piping, the following applies: where the superscript system represents the pump system, and the subscripts P1 and P2 represent individual pumps. In this case, the system flow is derived by summing the individual operating flow of each pump. Finally, the total flow from the first tank to the plant is: Conversely, when simulating multiple pumps in parallel with shared pipes, the following holds true: where the superscript system represents the pump system, and the subscripts P1 and P2 represent individual pumps. In this case, the total flow is derived. The system curve and the characteristic curve of each pump are used to find the operating point of the pump system. This is done through an iterative procedure that starts from H = 0 and increases H until the system curve and the characteristic curve of each pump intersect. For each head H, the total water flow for the given head (H) is estimated as: and the intersection with the system curve is determined for which where ε ∈ [0, 1] is a desired precision. Finally, we have: where ^^ ∗ is a feasible head for which In order to facilitate human interaction with the simulator, the system of the invention may comprise a graphical interface that displays the following components: • System curve graph and pump characteristic curves: A graph that displays the system curve and pump curves for each pump as the system evolves over time. The graph is updated at each simulation time interval to reflect the current state of the system. The total water flow is also represented (^^ ௧ ^௨௧From the first tank to the main plant, or second tank. • Tank level: The tank level is displayed with a bar that increases or decreases depending on the current tank level. • Water inlet: The water flowing into the tank is displayed and updated based on the current incoming water flow. • Pump status: The frequency, efficiency, and power of each pump are displayed. • Water outlet: The water outlet from the tank is displayed and updated based on the current system status. The graphical interface is implemented adaptively and can display any number of pumps. Figure 6 shows a screenshot of an example of the graphical interface. Regarding the optimization process, the PI controller used in the state of the art in a wastewater treatment plant is initially used as the reference control strategy.The PI controller is a single-input, single-output (SISO) controller that uses the error between the setpoint and the current value of the controlled variable to calculate the control signal. The control signal is calculated as the sum of the proportional and integral terms. The proportional term (P) is proportional to the error, while the integral term (I) is proportional to the integral of the error. In this case, since the simulator operates in discrete time intervals of 1 second, a discrete version of the PI controller has been implemented, given by the following equation: where u(t) is the control signal, e(t) is the error, is the proportional gain and ^^ is ^ the overall gain. The setpoint in this case is the desired water level in the tank, set to 60% of the first tank's capacity. The controlled variable is the current water level in the tank. Therefore, the error is: and the control signal u(t) is the frequency change that will be applied to the pump. Thus, the PI controller has two hyperparameters: the proportional gain ^^ ^ and the overall gain ^^ ^ To adjust these hyperparameters, a grid search is performed by evaluating the controller's performance in the simulator. This evaluation of controller performance is carried out by combining the following metrics: where ^^ ௧ is the water level in the tank at time ty T is the total number of time steps. where ^^^^ ௧ is the pump frequency at time t, and T is the total number of time steps. The final metric is the sum of the two previous metrics: This metric is used to select the best hyperparameters that minimize the tank level error and keep the pump frequency as stable as possible (i.e., the pump should not turn on and off too frequently, and the frequency should not change too often). This PI controller hyperparameter selection procedure can be used for an existing PI controller in a wastewater treatment plant without implementing Reinforcement Learning (RL). Regarding Reinforcement Learning (RL), in this case, the decision problem is modeled as a Markov Decision Process (MDP), that is, a tuple M = {S, A, P, R}, where: • S is the state space, containing all possible states the system can be in. • A is the action space, containing all possible actions the agent can perform when interacting with the system. This is the transition dynamics. Where Δ(S) is the probability simplex on S, that is, the set of all probability distributions on S. For each state s, an action a, and a subsequent state s′, P(s, a, s′) indicates the probability of reaching state s′ after performing action a in state s. Sometimes the conditional probability of transitioning to state s′ is represented as P(s′|s, a). is the reward function. The value R(s, a, s′) gives the amount of "reward" associated with the transition to state s′ when action a is performed from state s. Different algorithms have been implemented for Reinforcement Learning (RL): Proximal Policy Optimization (PPO), Q-learning, and Deep Q-learning (DQN). The Proximal Policy Optimization (PPO) algorithm is a gradient policy-type algorithm that approximates the value of the function for each state-action pair. In contrast, other algorithms require storing these values, so they are not useful when the action and state spaces are very large. The Proximal Policy Optimization (PPO) algorithm approximates the advantage function, which depends on advantage parameters, and the policy, which depends on policy parameters, using neural networks. In iteration k, the policy parameters are updated as follows: where ^^ ^ These are the outdated policy parameters, It is the objective associated with the policy ^^ ఏ under the unupdated policy parameters The Kullback-Leibler divergence between stochastic policies controls the influence of the KL term. Therefore, the new vector of pa k+1 parameters are calculated by maximizing the objective while at the same time remaining close to the old policy The objective of PPO is defined as the policy's trimmed gradient: do nde (s t ,to t ) are the state and the action in time is the estimated advantage associated with the old policy the state-action pair (s t ,to t ) according to the current parameters is a trimming function. In PPO, the policy is stochastic, and for a given state s, πθ(^|s) represents a probability distribution over the possible actions for the current parameter vector θ. In this case, action selection consists of randomly drawing an action from the probability distribution ^^ఏ^^ |^^^. A common problem in Reinforcement Learning (RL) is the balance between exploration and exploitation, that is, the balance between exploring the environment to find better policies and exploiting the current policy to maximize the expected reward. In PPO, a stochastic policy is used to ensure that the agent explores the environment, and preferably, an entropy term can be added to the goal to encourage exploration.Q-learning is the most popular Reinforcement Learning (RL) algorithm, although it requires that the state space S and action space A be discrete and small enough to represent the policy in a lookup table. DQN is a powerful deep Reinforcement Learning (RL) algorithm that, unlike Q-learning, can handle continuous state spaces using function approximation in the form of a neural network. Q-learning maintains an estimate. of the optimal action-value function. Specifically, for each state s and action a, It estimates how much future reward the system will receive by applying action a in a state s. Given a transition (s, a, r, s′) recorded during the execution of the system, the algorithm updates the estimate a) as follows: where ^^ ∈ ^0, 1^ is a learning rate. Given a state s, the strategy for action selection is represented by a policy that simply selects the action with the maximum Q value: However, during learning, it is sometimes necessary to select a random action to avoid local minima. The simplest and most common exploration strategy is called ε-greedy action selection, in which the exploration policy selects a random action with probability ε and the action with the maximum Q-value with probability 1 − ε. On the other hand, the Deep Q Network (DQN) combines Q-learning with Deep Neural Network function approximation to approximate the optimal action-value function. The parameterized action-value function of the neural network is defined as ^^ ఏ with the parameters θ. DQN updates the parameters θ by performing a stochastic gradient descent with the square of the error loss TD: where the gradient is: In this case, D is a set of transition data (s, a, r, s′) that is called the replay memory It's a copy of ^^ ఏupdated to a slower timescale. In this case, the neural network training encompasses two additional mechanisms: experiment replay and a target network. ● Experiment replay: The idea is to store transitions (s, a, s′, r) in a buffer D and uniformly sample small batches of transitions from D. ● Target network: The target network is used to mitigate the fact that the loss is being calculated based on a moving target that changes as the learning process unfolds. Therefore, two networks with the same structure but different weights are used: θ for the Q-network and θ− for the target network. The Q-network is periodically updated according to the loss function explained earlier, while the target network is updated by copying the parameters of the Q-network to the target network. every C time periods. Therefore, the target network weights remain frozen for time periods: C. The target network or it is updated much less frequently than θ to increase the stability of the learning. Additionally, the Double DQN (DDQN) strategy has been proposed as an improvement over DQN by addressing the problem of overestimating the values ​​of Q. In this case, the update rule is: The DDQN update rule uses the Q network to select the stock and the target network to evaluate the stock. This decoupling of the selection and evaluation processes helps reduce overestimation and leads to more stable and accurate stock value estimates. Furthermore, an improved technique can be implemented: the Prioritized Experience Reproduction (PER) technique, which improves how experience is sampled. The main idea is that it is preferable to learn from transitions (s, a, s′, r) that do not fit well into the current approximation of the Q value. In this case, an error e of a sample (s, a, s′, r) can be defined as the distance between Q(s, a) and its target T(s, a, r, s′): where T(s, a, r, s′) in the case of DDQN would be: This error then becomes a priority for each sample (s, a, s′, r): where ε is a small positive constant that ensures no transition has zero priority. The parameter θ (0 ≤ 1) controls the relative difference between high and low error, i.e., it determines how much prioritization is used. With θ = 0, the uniform case would be obtained. Priority translates into the probability of being chosen for repetition. A sample i has a probability of being chosen during the repetition of the experiment determined by: At each learning step, the algorithm samples a batch of transitions using this probability distribution and trains the network on that batch. The implementation of pump control using the Reinforcement Learning (RL) algorithms explained above is described below. In Reinforcement Learning (RL), the goal is to maximize the sum of rewards from a given state. That is, given a trajectory, we wish to maximize ∑where γ ∈ (0, 1) is a factor of discount that exponentially reduces the im agreement of future rewards. As before, Reinforcement Learning (RL) is modeled as a Markov decision process M = ^S, A, P, R^. Thus, the state space S of the problem is defined as the set of all possible values ​​of the tank level L measured as a percentage, the current pump frequencies and the change in tank level between two consecutive time steps The pump frequencies are discrete variables that can take values ​​in the domain {0 Hz, 1 Hz, ..., 50 Hz}, while the tank level L and the change in tank level ΔL are continuous variables that can take real values ​​in the intervals [0, 100] and [−100, 100], respectively. Therefore, any state in S is defined by n+2 variables corresponding to the discrete frequencies of the n pumps and the two continuous variables, namely, the tank level and its change in level between two time steps. Thus, a state s t ∈ S is represented as a tuple: One of the main consequences of factoring states with real variables is that the size of the state space becomes infinite. Thus, the complexity of the problem can be reduced by discretizing the factors, also allowing the use of tabular methods such as Q-learning in a discrete and finite state space. The tank level L is discretized simply by taking the integer part of its value. The pumping frequencies are already discrete variables. The change in tank level ΔL is discretized by rounding it to two decimal places and multiplying by 100. Therefore, the final discretized representation of a state is: The use of tabular methods is limited to small state spaces, as some memory must be allocated to each combination of variables and their evaluation. Another option is to deal with continuous states and use universal function approximators, that is, Neural Networks (NNs), which accept real numbers as input to approximate the desired function. One of the problems when learning with NNs is that the weights can be greatly affected by the domain of the input variables. The solution to this problem is to normalize all the input data so that each factor has the same domain, and therefore the same relevance when learning the neural network's weights. In this case, all variables have a known domain, so it is possible to normalize as follows: In order to respect physical constraints, the action space A is defined as the frequency variation (Δf) that will be applied to each of the n pumps, that is,. More specifically, for each pump p i the action is limited to i to be discrete and take values ​​in the set {−1Hz, 0Hz, +1Hz}. Therefore, the size of the action space is |A| = 3n, where each action is represented as: The transition function is denoted as that which returns the probability of a tuple es decir, is the probability of reaching state S when applying action A in state S at time t. Thus, the next state is defined as: where each adds −1, 0 or 1 to the i-th pump, keeping the final frequency within the range [0, 50]; and the difference in the tank level is due to the difference between the incoming water flow where the latter can only be controlled through system design. Finally, the reward function is defined as a penalty term, that is, a cost. If the water level in the tank is within an acceptable range, the reward is the negative sum of each pump's power. Otherwise, the reward is a negative constant (−c) representing the maximum penalty for violating the rules of making the tank level too high or too low: where π is the pump power; ε = 5 by default; and the penalty is a constant. where denotes the maximum possible energy consumption of the pump p i In the experiments performed, this value was for all pumps. Alternatively, in problems where the problem size is too large, a PPO or DQN type algorithm can be used. Additionally, the method of the invention could also be used with a PI controller, as used in the prior art, but determining, by means of the method of the invention, the number of pumps to activate / deactivate to control the tank level. The optimal number of pumps to activate to minimize energy consumption can be determined by simulation as described in the method of the invention. Also, the method of the invention could comprise the integration of exploration strategies within the RL algorithm to capture a long-horizon dependency that justifies short-term costs with long-term benefits. Additionally, Predictive Models can be integrated to estimate the water inlet data and feed it into the RL algorithm.This enables proactive decision-making and further optimizes system performance. The results of evaluating tabular Q-learning (Tab-Q-Learning) and DDQN with prioritized experience replay (DDQN-PE) strategies against the state-of-the-art PI controller are then presented. The performance of the three controllers was compared in a real-world scenario where the simulator was configured to simulate the wastewater treatment plant of the city of La Almunia, as shown in Figure 1. A single-pump system and a tank with a capacity of 103.75 m³ were considered. 3With a maximum static head of 9.2 m, the system curve can be derived from the plant representation in Figure 1. In this case, the actual plant is equipped with four pumps: two for normal operation and two backups in case of exceptional events. The pump curves are derived from the manufacturer's documentation and are shown in Figure 2. The water inlet flow is simulated using real data recorded at the plant, as shown in Figure 12. The three algorithms were trained and evaluated in the simulator over 4 million time steps, representing approximately 46 days. Various metrics were recorded to assess the controllers' performance. Figure 7 shows the reward collected by the controllers over time. As can be seen initially, Tab-Q-learning and DDQN-PE explore the environment, performing worse than the PI controller in terms of reward. However, as time passes, the Reinforcement Learning (RL) algorithms learn to exploit the environment, and the reward increases. After two million time steps (approximately 23 days), the Reinforcement Learning (RL) controllers outperform the PI controller. It is further shown that both DDQN-PE and Tab-Q-learning operate similarly.Figure 8 shows the tank level over time. Monitoring the tank level allows for the interpretation of the policy learned by the Reinforcement Learning (RL) algorithms. As can be seen in Figure 8, the PI controller maintains the tank level at a constant 60%, while the RL algorithms learn to maintain the tank at a higher level of 93%. As the tank level rises, the system curve descends; consequently, the pump can move the same amount of water at a lower frequency. This advantage is reflected in Figure 9, where the Reinforcement Learning (RL) algorithms are operating at a lower frequency than the PI controller. Finally, the same effect is reflected in Figure 10, where the Reinforcement Learning (RL) algorithms are operating at a lower power than the PI controller.Figure 11 shows the pump's efficiency over time. It reveals that, after the initial exploration phase, the Reinforcement Learning (RL) algorithms learn to operate the pump with greater efficiency than the PI controller. Table I lists all the hyperparameters used to configure the algorithms in the experiments. Table I. Hyperparameters. Based on the results obtained with DDQN-PER and Tab-Q-learning, it can be deduced that the optimal strategy in the described simulator is to remain close to the upper limit of the tank level. Compared to the PI controller, this generates a significant decrease in pump power, resulting in a corresponding energy cost saving of 2.3 kW / h on average. In the case where a PPO-type algorithm is used, with a single pump, the decision problem is modeled as a Markov Decision Process (MDP) with the following data: composed of three continuous variables. tank level percentage.pump frequency. pump power.^^: composed of three discrete actions^^ ^ increased frequency of bo mba at 1 Hz. ^^ ଶ Reduction of the pump frequency by 1 Hz. ^^ ଷMaintain a constant pump frequency. Any action to increase or decrease the pump frequency will only have effect if it is maintained within the pump's operating range: [39, 50] Hz. ℛ defined as: Where ^^ ௧ It's the tank level as a percentage over time. Thanks ^^ ^^This is the pump's power. Therefore, restrictions have been imposed on the tank level (between 30 and 70), the number of times per hour the pumps are activated / deactivated, and the pump's operating frequency (between 39 and 50 Hz), which is applied directly to the action space. Imposing restrictions allows for balancing system safety with optimizing the performance of the RL algorithm: • Frequency restrictions: These prevent the algorithm from discovering policies that operate outside the prescribed frequency limits. • Tank level restrictions: These impose restrictions on the algorithm's search space, according to safety limits, to obtain optimal solutions. • Pump activation restrictions: Restrictions that allow only one activation or deactivation at a given time mitigate rapid changes that could destabilize the system.In this case, the results mainly refer to: • Power (kW) consumed per hour. • Specific consumption (kWh / m). 3Energy per unit of water. Integrated over time, these metrics allow for the determination of energy consumption metrics used in the plant's daily operation, such as the energy consumption ratio per cubic meter of treated water (kWh / m³). Additionally, for each use case, the tank level, water inflow, and average number of active pumps are displayed. These metrics are shown with a 100-hour moving average smoothing to reduce noise in the data and make the graphs more readable. Shaded registers represent the actual metric values, while solid lines represent the smoothed moving average. RL represents the Reinforcement Learning (RL) controller, trained with the PPO algorithm, and PI represents the proportional-integral controller. The performance of the RL agent is compared to the PI controller at 60% and 69% of the tank level.The PI-60% controller is the one typically used in the water treatment plant, while the PI-69% controller is the same PI controller with the optimal tank level suggested by the RL agent. The PI-69% allows you to see the improvement of the RL agent over the PI controller when the same tank level is set. The optimal tank level is the level that minimizes power consumption without violating the constraint of maintaining the tank level between 70% and 30%. The RL agent automatically determines the optimal tank level that minimizes power consumption without violating the tank level constraints. Table II: Results of Use Case 1: La Almunia. As can be seen in Table II, the RL agent outperforms the PI controller at 60% of the tank level. The RL agent reduces power by 9.65% and decreases specific consumption by 9.65%. On the other hand, when comparing the RL agent with the PI controller at 69% of the tank level, the RL agent performs similarly with a power reduction of 0.23% and a specific consumption reduction of 0.23%. The inlet water flow is shown in Figure 13, the tank level in Figure 14, the power in Figure 15, and the specific consumption in Figure 16. The RL agent reduces power by setting a higher tank level, allowing the pumps to work less. When the same operating tank level is set for the PI controller (PI 69%), the RL agent and the PI 69% perform similarly. As the number of pumps increases, the model changes.In this case, the same plant is simulated with two pumps in parallel, which can be activated or deactivated according to demand. In this case, the model comprises the following data: composed of five continuous variables (ℒ, ^^௭^భ, ^^௭^మ, ∑ே^ ^^^ , ^^): ℒ ∈ ^0, 100^ percentage of tank level. ^^^௭^భ, ^^௭^మ^ frequency of the pumps. sum of the pump powers. ^^^^ ^0, 1^ is a normalized variable that represents the time that has passed since the last time the pump was activated. ^^: composed of nine discrete actions (^^^, … , ^^ଽ): activate an extra pump. ^^. ଶ defuse a bomb. ^^ ଷ Maintain a constant pump frequency. ^^ ସ Increase the pump frequency by 1 Hz. ହ Increase the pump frequency by 2 Hz in 1 Hz increments. ^ Pump frequency reduction of 1 Hz at a time. ^ Pump frequency reduction 2 in 1 Hz. ^^ ଼Increase the frequency of all pumps by 1 Hz. ଽ Reduction of the frequency of all pumps by 1 Hz. Any action to increase or decrease the pump frequency will only have effect if it is maintained within the pump's operating range: [39, 50] Hz and if the pump is active. ℛ: defined as: Where ^^^^ ^0, 1^ is reset each time a bomb is activated and ^^1 ^^ ^0, 1^ is a weight that controls the importance of the bomb's power in the reward function. Table III: Resu Results of Use Case 2: La Almunia v2. Table III shows that the RL agent outperforms the PI controller at 60% of the tank level. The RL agent reduces power by 14.16% and decreases specific consumption by 14.74%. On the other hand, when comparing the RL agent with the PI controller at 69% of the tank level, the power reduction is 5.72% and the specific consumption reduction is 6.25%. Also shown are: the inlet water flow (Figure 17), the tank level (Figure 18), the power (Figure 19), the specific consumption (Figure 20), and the number of active pumps (Figure 21). Once again, the RL agent reduces power by setting a higher tank level, allowing the pumps to operate less. By setting the same operating tank level for the PI controller, in this case, the RL agent also outperforms the PI controller with the optimal tank level (PI 69%).This is because, as shown in Figure 12, the RL agent uses fewer pumps than the PI controller, allowing the RL agent to reduce power consumption. Finally, tests were carried out in Burgos, with a plant equipped with nine pumps, five of which operate in parallel during normal plant operation, pumping water into a single pipe. Up to four of the five pumps can be switched on or off. In this case, the model comprises twelve continuous variables. ^^: composed of fifteen discrete shares ^^ ^ Activate an extra bomb. ^^ ଶ defuse a bomb. ^^ ଷ Maintain a constant pump frequency. ^^ ସ Increase the pump frequency by 1 Hz increments. … ^^ ଼ Increase the pump frequency by 5 Hz in 1 Hz increments. ଽ Reduction of the pump frequency by 1 Hz increments. … ^^ ^ଷ Pump frequency reduction of 5 in 1 Hz. ^^^ସ Increase the frequency of all pumps by 1 Hz. ^ହ Reduction of the frequency of all pumps by 1 Hz. Any action to increase or decrease the pump frequency will only have effect if it is maintained within the pump's operating range: [39, 50] Hz and if the pump is active. ℛ: defined as: Table IV: Results of Use Case 3: Burgos. Table IV shows that the RL agent outperforms the PI controller at 60% of the tank level. The RL agent reduces power by 6.01% and decreases specific consumption by 6.55%. On the other hand, when comparing the RL agent with the PI controller at 69% of the tank level, a power reduction of 3.08% and a reduction in specific consumption of 3.85% are obtained. The inlet water flow is shown in Figure 22, the tank level in Figure 23, the power in Figure 24, the specific consumption in Figure 25, and the number of active pumps in Figure 26. The RL agent reduces power by establishing a higher tank level, allowing the pumps to operate less. When the same operating tank level is set for the PI controller, PI 69%, the RL agent also outperforms the PI controller with the optimal tank level (PI 69%).Thus, in all the use cases shown, the Reinforcement Learning (RL) algorithm is able to learn a policy that outperforms the current strategy (PS). In the first use case, energy consumption savings of up to 33.84 kWh / day on average are achieved; in the second use case, savings of up to 74.88 kWh / day are obtained; and in the third use case, savings of up to 300 kWh / day are achieved. Table V: KPIs of the invention system

Claims

AMENDED CLAIMS received by the International Bureau on March 6, 2026 (06.03.2026) 1. A pump control method for pumping in a water treatment plant comprising a first tank and a second tank, located at a higher level with respect to the first tank, comprising the steps of: obtaining a configuration data set of the pump assembly of the treatment plant tanks and of a set of pipes connecting the tanks; obtaining estimates of the incoming water flow into the first tank, unknown and time-varying, from historical plant data and a desired change in frequency; performing a simulation of the pump behavior for pumping water from the first tank to the second tank by means of the steps of: • Determine a system curve that relates the static pump head to the flow using flow estimation equations, such as Manning's equation as a function of the level of the first tank; • obtain a characteristic curve for each pump for all frequencies, from the characteristic curve at a specific frequency using affinity laws; • determine an operating point for the pumps as the point of intersection between the system curve and the characteristic curves, determining a pumped water flow; • Determine the pump efficiency and power consumption at that operating point; determine an optimal operating point through the following steps: • Apply Reinforcement Learning (RL) techniques, using neural networks trained with simulation values ​​and / or historical plant data, where a metric defined to calculate the hyperparameters is the sum of a water tank level error metric and a frequency variation error metric of the form: where where setpoint is a level value of the first tank defined by simulation, Hz t is the operating frequency of the pumps at a given moment ty Hz t+1is the operating frequency of the pumps at a later time at; and where in Reinforcement Learning (RL) is defined as a Markov decision process in which a state space S, an action space A, a transition function space P are defined, which stores the probability of moving to a state s' from a state s by means of an action a, and a reward function of the action a, and where the state space (S) defines possible values ​​of tank level (L), frequencies of each of the pumps and change of tank level between 2 time steps where the stock space (A) is defined as a frequency variation that is applied to each pump; and the transition function is defined as: and where the frequencies take discrete values ​​in the range [1, 50] Hz, the tank level values ​​are continuous in the range [0, 100] and the tank level change values ​​are continuous in the range [-100, 100], and the frequency variation is applied using the values: (-1, 0, 1}, and keeping the frequency in the range [0, 50]; and where the reward function is defined as a penalty term, if the tank level is in an acceptable range, it is the negative sum of the pump power and if not, it is a constant value -c which is the maximum penalty: where the default pump power p yc = is a penalty constant te para el valor máximo de consumo de poder de la pu p t ; • Apply monitoring techniques for the first tank level and incoming water flow, preventing the first tank from emptying or overflowing, and ensure that the pump operating point is above the NPSH (Net Positive Suction Head) curve; update the first tank level by applying the following formulas: and where the stage of determining an optimal operating point is carried out by updating an estimate Q of the optimal action-value function that determines the future reward of action a, in the form: where a is the learning rate and the action that gives a maximum value of Q is selected:

2. A control method according to claim 1, wherein the configuration data comprises: System Curve: the first and second term coefficients of the system curve, which describe how the resistance of the system varies as a function of the water flow; Tank Volume and Maximum Static Head of the system; Characteristic curve of each pump at a fixed frequency; and Efficiency Surface of pumps at discrete frequency intervals.

3. A control method according to claim 1, wherein the historical data comprise: the level of the first tank; the working frequencies of the pumps; and the power consumed by the pumps.

4. A control method according to claim 1, wherein the step of determining a system curve is performed using the expression: where is the maximum static head, is the tank level head at t, ybyc are coefficients of the system curve polynomial obtained for flow values ​​determined by the following Manning equation: where n is the Manning coefficient, A is the pipe area, R is the hydraulic radius, and S is the slope energy gradient.

5. A control method according to claim 1, wherein the step of obtaining a characteristic curve of the pumps is performed using the expression: where a, b, c, d are coefficients of the characteristic curve polynomial.

6. A control method according to claim 1, wherein the step of obtaining a characteristic curve for each pump for all frequencies is performed using affinity laws of the form: where N is the rotational speed of a pump motor, Q are the flows, H represents the heads and P the powers, and where the rotation speed is obtained: where glide refers to engine characteristics.

7. A control method according to claim 1, wherein the step of determining an operating point is carried out by means of the expressions:

8. A control method according to claim 1, wherein the efficiency depends on the frequency and flow of the pump, and is determined using previously obtained characteristic data at a specific frequency or frequencies and using a multivariable regression model to determine an efficiency surface at any frequency and flow.

9. A control method according to claim 1, wherein the power consumed is determined from the efficiency using a motor power curve or using the expression:

10. A control method according to claim 1, wherein the treatment plant comprises multiple pumps located in parallel with separate or shared pipes.

11. A control method according to claim 10, wherein the pipes are separated and wherein:

12. A control method according to claim 10, wherein the pipes are shared and wherein: where the operating point is obtained through an iterative process by increasing the system head from 0 until the system curve intersects the characteristic curves, estimating for each system head value the total water flow, as the sum of the pump flows: and looking for the intersection with the system curve for which the total flow differs from the flow of the system curve by a value less than a required accuracy:

13. A control method according to claim 1, wherein the Reinforcement Learning (RL) neural networks are continuously trained with new data measured at the plant and changes in the inlet water flow distribution.

14. A control method according to claim 1, the Reinforcement Learning (RL) neural networks are trained in a loop, recording new inlet water flow data to update the simulation and determine in said simulation a new policy, ensuring that it is safe and efficient.

15. A control method according to claim 1, wherein Reinforcement Learning (RL) is applied using off-policy techniques, which focus on learning from past experiences, selected from Q-learning and G-learning, or on-policy techniques, which focus on learning from the ongoing interaction, selected from SARSA and Actor-Critic.

16. A control method according to claim 1, wherein an α-greedy exploration strategy is used to randomly select the action and prevent local minima.

17. A control method according to claim 1, wherein the neural networks use a Deep Q Network (DQN) technique that defines an action-value function as Qe with parameters 6 and updates it by means of a stochastic gradient that descends with the square of the TD-error loss, in the form: where D is a set of transition data It is a copy of Q e updated to a shorter timescale, yy is a discount factor that reduces the impact of future rewards.

18. A control method according to claim 17, wherein Reinforcement Learning (RL) further comprises a technique for reproducing experiences which involves storing transitions in a memory and sampling subsets of transitions.

19. A control method according to claim 17, wherein Reinforcement Learning (RL) further comprises defining a sample error (e) as the distance between Q(s, a) and a target T(s, a, r, s'): T being: and error becomes a priority for each sample, in the following way: being a small positive constant and where It controls the relative difference between a high and low error, and where Reinforcement Learning (RL) is performed with transitions according to this probability:

20. A control method according to claim 17, wherein Reinforcement Learning (RL) further comprises a target network technique comprising the use of two networks of the same structure but with different weights, 6 for the Q-network and 0 for the target network, such that the Q-network is updated regularly while the target network is updated to the parameters of the Q-network every C time steps.

21. A control method according to claim 20, wherein the neural networks use a double Deep Q Network (DDQN) technique with a function: and where the Q-network is used to select the action and the target network is used to evaluate the action.

22. A control method according to claim 21, wherein a Proximal Policy Optimization (PRO) type algorithm is used, which approximates for each state-action pair the value of the advantage function dependent on advantage parameters a policy dependent on policy parameters 6, using neural networks; where in iteration k, the policy parameters are updated as follows: where ft are the unupdated policy parameters, is the target associated with politics under the unupdated policy parameters K is the Kullback-Leibler divergence between stochastic policies controls the influence of the KL term; and where the new parameter vector k+1 is calculated by maximizing the objective where the PPO target is defined as the trimmed policy gradient: where (s t ,to t ) are the state and action over time is the estimated advantage associated with the old policy the state-action pair (s t ,tot ) according to the current parameters is a function of cutout.

23. Control method according to claim 22, wherein rre( |s) represents a probability distribution over the possible actions for the current parameter vector 0, and wherein the selection of actions consists of randomly extracting an action from the probability distribution rre(- |s).

24. Control system for a water treatment plant comprising: a set of pumps; a first tank; and a second tank, located at a higher level with respect to the first tank; wherein the control system comprises: a pump control module connected to the pumps and configured to carry out the method according to any of claims 1 to 23.

25. Control system according to claim 24, wherein the system further comprises a graphical interface configured to display the system curve, the pump characteristic curves, and their operating point at any given time, a tank level value, an inlet water flow value, the frequency, efficiency and power of the pumps, and an outlet flow value.

Citation Information

Patent Citations

  • Method and device for controlling a wastewater tank pumping system

    EP3690758A1

  • Cavitation detection in a process plant

    US20020123856A1

  • Digital model based configurable plant optimization and operation

    US20230213922A1