Cascade pump station collaborative scheduling method, device and equipment based on artificial intelligence and medium thereof
By constructing a lightweight hydraulic transient model and a reinforcement learning agent, we have achieved active control over hydraulic transients in the cascade pumping station system, solving the equipment damage and safety risks caused by pressure fluctuations in traditional models, and improving system safety and energy efficiency.
Patent Information
- Application Number
- CN202511840077.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-02-13
AI Technical Summary
Traditional AI scheduling models have failed to effectively handle hydraulic transients in cascade pumping station systems, leading to equipment fatigue damage and system safety risks caused by pressure fluctuations, and failing to achieve synergistic optimization of safety and efficiency.
A lightweight one-dimensional nonlinear hydraulic transient model is constructed and combined with a reinforcement learning agent. The agent is trained through state space, action space and reward function to generate a pre-trained scheduling strategy, predict pressure oscillations in real time and generate compensatory scheduling instructions, so as to achieve active control of hydraulic transients.
It improves the control accuracy and equipment life of the cascade pumping station system, reduces operating energy consumption and maintenance costs, and enhances the system's adaptability to complex working conditions.
Smart Images

Figure CN121523048A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of automatic control of water conservancy projects, and particularly relates to a cascade pumping station collaborative scheduling method, device and equipment based on artificial intelligence and a medium thereof. BACKGROUND
[0002] In the operation and scheduling of the cascade pumping station water conveyance system, artificial intelligence technology has been widely applied to energy efficiency optimization and water distribution. Traditional AI scheduling models are usually based on steady-state or quasi-steady-state assumptions, simplify the hydraulic process into a static hydraulic balance relationship, and mainly focus on macro objectives such as pump station combination optimization, speed regulation energy saving, etc. Although these models have improved the system operation efficiency to some extent, their core defect is that they completely ignore the inherent dynamic characteristics of the hydraulic transient process. In actual engineering, the start-stop, valve regulation or speed regulation operation of the upstream pumping station will excite pressure waves (i.e. water hammer phenomenon), which propagate along the pipeline at the speed of sound and cause complex reflection, superposition and resonance phenomena when encountering downstream pumping stations, valves or pipeline diameter changes. The traditional AI model simply regards this water hammer effect as a constraint condition that needs to be avoided, and usually adopts a passive defense strategy of setting a large safety margin, such as limiting the speed regulation frequency, avoiding rapid start-stop, etc. Although this conservative strategy ensures system safety, it seriously sacrifices the flexibility and optimization space of scheduling.
[0003] More seriously, when the AI scheduling strategy pursues extreme energy efficiency through fine and frequent speed regulation, these small operations may become new vibration sources that excite pressure oscillation. Due to the lack of perception and modeling ability of the hydraulic transient dynamic process, the traditional scheduling system cannot predict the nonlinear oscillation risk that may be caused by its own decision, forming a vicious cycle of "optimization-oscillation-instability". This continuous pressure fluctuation not only causes cumulative fatigue damage to the pipeline, pumps and valves, but also may lead to out-of-control system pressure, threatening the safe operation of the entire water conveyance system.
[0004] Although existing research has recognized this problem, the solutions have obvious limitations. On the one hand, high-precision hydraulic transient simulation models are complex and time-consuming to calculate, which cannot meet the real-time requirements of online scheduling; on the other hand, existing AI methods such as reinforcement learning lack quantitative consideration of pressure oscillation dynamic characteristics in the design of reward functions, which makes the agent unable to learn effective oscillation suppression strategies. This fragmentation of model and control makes the system have to make a difficult choice between "ensuring safety at the expense of efficiency" or "pursuing efficiency at the risk of safety", and cannot achieve the collaborative optimization of safety and efficiency.
[0005] Therefore, there is an urgent need for an innovative method that can deeply integrate the hydraulic transient dynamic characteristics into the AI decision framework, and fundamentally solve the coupling problem between scheduling behavior and system dynamic response. SUMMARY
[0006] Therefore, it is necessary to provide an artificial intelligence-based collaborative scheduling method, device, equipment, and medium for cascade pumping stations to address the aforementioned technical problems.
[0007] Firstly, this application provides an artificial intelligence-based collaborative scheduling method for cascade pumping stations, including:
[0008] S1. Based on the pipeline geometric parameters, pump station equipment parameters and historical SCADA data of the cascade pump station system, a lightweight one-dimensional nonlinear hydraulic transient model is constructed to obtain the hydraulic transient model.
[0009] S2. Based on the real-time status data of the cascade pumping station system and the predicted output of the hydraulic transient model, a reinforcement learning agent is obtained by constructing a state space, action space, and reward function for reinforcement learning.
[0010] S3. Using historical system operation data of the cascade pumping station system, simulate state transition and reward calculation in a simulation environment to train the reinforcement learning agent and output the pre-trained scheduling strategy.
[0011] S4. Input the real-time system status of the cascade pumping station system into the hydraulic transient model to predict pressure oscillations; output scheduling instructions to cope with pressure oscillations based on the scheduling strategy, and execute pumping station control according to the scheduling instructions.
[0012] Secondly, this application also provides an artificial intelligence-based cascade pumping station coordinated scheduling device for implementing the method described in the first aspect, the device comprising:
[0013] The hydraulic dynamic modeling module is used to construct a lightweight one-dimensional nonlinear hydraulic transient model based on the pipeline geometric parameters, pump station equipment parameters and historical SCADA data of the cascade pump station system, thus obtaining the hydraulic transient model.
[0014] The strategy space construction module is used to predict the output based on the real-time status data of the cascade pumping station system and the hydraulic transient model. By constructing the state space, action space and reward function for reinforcement learning, a reinforcement learning agent is obtained.
[0015] The agent training and optimization module is used to train the reinforcement learning agent by using historical system operation data of the cascade pumping station system, simulating state transitions and reward calculations in a simulation environment, and outputting a pre-trained scheduling strategy.
[0016] The scheduling instruction execution module is used to input the real-time system status of the cascade pumping station system into the hydraulic transient model to predict pressure oscillations; based on the scheduling strategy, it outputs scheduling instructions to deal with pressure oscillations and executes pumping station control according to the scheduling instructions.
[0017] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement an artificial intelligence-based collaborative scheduling method for cascade pumping stations as described in the first aspect.
[0018] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an artificial intelligence-based coordinated scheduling method for cascade pumping stations as described in the first aspect.
[0019] The aforementioned method, device, equipment, and medium for coordinated scheduling of cascade pumping stations based on artificial intelligence achieves dynamic prediction of pressure wave propagation by constructing a lightweight hydraulic transient model. Based on the model prediction results and real-time system status, a reinforcement learning agent is constructed. Using historical operating data, a scheduling strategy capable of actively identifying and responding to pressure oscillations is trained. Finally, in the online control stage, the real-time status is input into the model to predict pressure oscillations, and compensatory scheduling instructions are generated based on the trained strategy. This achieves a fundamental shift from passively avoiding hydraulic transients to actively managing them. As a result, while ensuring the safe and stable operation of the system, control accuracy and equipment lifespan are significantly improved, operating energy consumption and maintenance costs are effectively reduced, and the system's adaptability to complex operating conditions is enhanced. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 A flowchart illustrating an artificial intelligence-based collaborative scheduling method for cascade pumping stations provided by this invention;
[0022] Figure 2 This is a schematic diagram illustrating the process of constructing a reinforcement learning agent in one optional embodiment of the present invention;
[0023] Figure 3 This is a schematic diagram of the structure of a cascade pumping station collaborative scheduling device based on artificial intelligence, provided by the present invention. Detailed Implementation
[0024] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0025] refer to Figure 1 The document presents a flowchart illustrating an artificial intelligence-based collaborative scheduling method for cascade pumping stations, which includes the following steps:
[0026] S1. Based on the pipeline geometric parameters, pump station equipment parameters and historical SCADA data of the cascade pump station system, a lightweight one-dimensional nonlinear hydraulic transient model is constructed to obtain the hydraulic transient model.
[0027] Specifically, pipeline geometric parameters include the pipeline's inner diameter, wall thickness, laying length, pipe material roughness, and terrain elevation characteristics along the pipeline. These parameters directly determine the propagation path and resistance characteristics of water flow within the pipeline. Pump station equipment parameters include the main pump's head-flow characteristics, efficiency-flow characteristics, motor speed range, and valve flow coefficients and opening / closing characteristics. These parameters reflect the inherent performance boundaries of the equipment. Historical SCADA data covers continuous monitoring data such as pump station inlet and outlet pressures, flow rates, pump speeds, valve openings, and water tank levels during past operations, providing a data foundation for model parameter calibration and verification.
[0028] The model is built upon one-dimensional unsteady flow theory, using Saint-Venant's equations as the mathematical foundation for describing hydraulic transient processes. Its continuity equation is expressed as follows:
[0029]
[0030] The momentum equation is expressed as follows:
[0031]
[0032] In the above two formulas, Represents the water flow rate within the pipe. Represents the spatial coordinates along the pipe's axial direction. Represents the cross-sectional area of the pipe. Represents the piezometric head. Represents time, Represents gravitational acceleration. Represents the friction coefficient. This represents the pipe's inner diameter. To achieve a lightweight model design, the Saint-Venant equations are discretized using the finite volume method combined with the implicit Crank-Nicolson scheme. This discretization method significantly reduces computational complexity while ensuring computational stability. Spatially, the pipe is discretized using a staggered grid, alternating pressure and flow nodes to improve the ability to capture abrupt changes in hydraulic parameters. The time step is determined based on the Courant number constraint to ensure the numerical stability of the flow fluctuation propagation process.
[0033] To address the error problem caused by the linearization of the friction term in traditional models, the model retains the friction coefficient along the path. The nonlinear characterization is calculated in real time using the Colebrook-White formula. The value, whose expression is:
[0034]
[0035] In the formula, Represents the absolute roughness of the pipe. Representing the Reynolds number, this formula dynamically adjusts the friction coefficient according to the actual flow state (laminar or turbulent) to ensure the accuracy of friction loss calculation. After the model is built, its reliability is ensured through static and dynamic verification. Static verification verifies the model's accuracy under stable operating conditions by comparing the consistency between the model output and actual monitoring data under steady-state conditions. Dynamic verification, on the other hand, simulates transient conditions such as pump start-up and shutdown, and valve adjustment, comparing the differences between the pressure wave propagation law and peak pressure predicted by the model and the actual data to ensure the model accurately portrays hydraulic transient processes.
[0036] S2. Based on the real-time status data of the cascade pumping station system and the predicted output of the hydraulic transient model, a reinforcement learning agent is obtained by constructing a state space, action space, and reward function for reinforcement learning.
[0037] Specifically, the state space construction comprehensively covers the static characteristics and dynamic responses of the system, ensuring that the agent can fully perceive the current state and future trends of the scheduling environment. The input data for the state space includes two parts: first, real-time state data of the cascade pumping station system, covering the inlet and outlet pressures, flow rates, pump speeds, valve openings, and water levels in the pools of each pumping station. This data directly reflects the current operating conditions of the system; second, the predictive output of the hydraulic transient model, including predicted pressure values, pressure change rates, and water hammer risk levels at key sections of each pipeline at multiple future time points. This data supports the agent in predicting hydraulic transient risks in advance. To eliminate the impact of differences in data dimensions on agent training, all data in the state space needs to be normalized, mapping the data to a fixed interval to ensure a balanced weighting of the influence of each dimension of data on the training process.
[0038] The definition of the action space combines the mechanical constraints and scheduling requirements of equipment operation to ensure that the actions output by the intelligent agent have engineering feasibility and practical operational value. The action space includes three core types of actions: first, pump speed adjustment actions, where the adjustment range must strictly adhere to the motor's speed range and single adjustment rate limits to avoid motor overload or hydraulic shock due to sudden speed changes; second, valve opening adjustment actions, where the adjustment step size and rate must match the valve's opening and closing mechanical characteristics to prevent water hammer effects caused by valve disc impact or sudden flow changes; and third, pump station start-up and shutdown actions, which must meet the water level constraints of the pool and downstream flow demand constraints to avoid system flow interruption or overload due to erroneous pump station start-up and shutdown. The dimensions of the action space are determined by the number of adjustable devices in the cascade pump station system. Each adjustable device corresponds to one or more action dimensions, forming a structured action space matrix.
[0039] The reward function employs a multi-objective weighted summation design, balancing system safety, energy efficiency, and hydraulic stability to avoid performance imbalances caused by single-objective optimization. The expression for the reward function is:
[0040]
[0041] In the formula, Represents the total reward value. , , , These represent the weighting coefficients for safety rewards, energy efficiency rewards, oscillation suppression rewards, and penalties, respectively. The weighting coefficients are determined through offline grid search and performance verification to ensure that each objective occupies a reasonable weight in the reward function. The safety bonus is calculated based on the deviation between the pressure value predicted by the hydraulic transient model and the safe pressure range. The closer the pressure is to the median of the safe range, the greater the bonus value. This represents the energy efficiency bonus, calculated based on the ratio of the system's total energy consumption to the historical best energy consumption. The lower the energy consumption, the greater the bonus value. The reward for oscillation suppression is calculated based on the amplitude and frequency of pressure fluctuations. The smaller the fluctuation and the lower the frequency, the greater the reward value. This represents a penalty item. When the system experiences abnormal operating conditions such as pressure exceeding the threshold, equipment overload, or insufficient flow, a corresponding penalty value is assigned based on the degree of abnormality to guide the agent to avoid dangerous actions.
[0042] S3. Using historical system operation data of the cascade pumping station system, simulate state transitions and reward calculations in a simulation environment to train the reinforcement learning agent and output a pre-trained scheduling strategy.
[0043] Specifically, the first step is to screen and preprocess historical operational data to provide high-quality training data support for the simulation environment. The historical data screening process involves selecting continuous and complete operational cycle data, covering system operating states under different seasons, flow demands, and equipment conditions, ensuring the diversity and representativeness of the training data. The preprocessing process includes outlier removal and missing value imputation. Statistical methods are used to identify and remove abnormal data such as sudden pressure changes, negative flow rates, and out-of-range rotational speeds to avoid interference with agent training. For missing values encountered during data acquisition, a time-series prediction model is used to imput them based on normal data from preceding and following times, ensuring the continuity of the data sequence. The preprocessed historical data is then divided into training, validation, and test sets according to the time series, used for agent parameter training, hyperparameter tuning, and performance verification, respectively.
[0044] The construction of the simulation environment is the core of realizing offline training of the intelligent agent. Its function is to simulate the state transition process and reward feedback of a cascade pumping station system under different scheduling actions. The simulation environment uses a lightweight hydraulic transient model as its core computing module. It takes the actions output by the reinforcement learning agent (such as pump speed adjustment and valve opening adjustment) as input, calculates the new system state (such as changes in pressure and flow) after the action is executed through the hydraulic transient model, and calculates the reward value corresponding to the action according to the preset reward function, forming a closed-loop feedback of "action-state-reward". To improve the realism of the simulation environment, actual engineering factors such as the dynamic response delay of the equipment and sensor measurement noise are incorporated into the environment to ensure that the training scheduling strategy can adapt to the operating characteristics of the real system.
[0045] The training process of the agent adopts a training framework combining temporal difference learning and experience replay. The specific steps are as follows: First, initialize the agent's network parameters (such as the weights and biases of the deep neural network) and the experience replay pool; In each training cycle, the agent obtains the current system state from the state space, selects an action based on the exploration strategy (such as the ε-greedy strategy), and inputs it into the simulation environment; After the simulation environment executes the action, it returns a new state and reward value, and stores the experience tuple of "current state-action-reward-new state" into the experience replay pool; When the experience replay pool accumulates a certain amount of experience, randomly sample batches of experience to update the agent's network parameters, and optimize the network by minimizing the loss function between the target Q value and the predicted Q value; During the training process, the exploration rate of the exploration strategy is dynamically adjusted, and the exploration rate is gradually reduced as the number of training iterations increases, so that the agent gradually shifts from "exploration" to "utilizing" the learned optimal strategy.
[0046] During training, explicit termination conditions and performance evaluation metrics are set to ensure the effectiveness and convergence of agent training. Termination conditions include reaching a preset upper limit for the number of training iterations, the average reward value on the validation set remaining stable for multiple consecutive periods (i.e., reward value fluctuations below a preset threshold), and the safety rate and energy efficiency metrics of the agent's output scheduling policy on the test set meeting preset standards. Performance evaluation metrics include the frequency of system stress exceeding thresholds, average energy consumption level, and stress fluctuation amplitude. These metrics are periodically calculated to monitor the agent's training progress and performance improvement. When the agent training meets the termination conditions, the network parameters and scheduling policy at this point are saved to form a pre-trained scheduling policy model for subsequent real-time scheduling.
[0047] S4. Input the real-time system status of the cascade pumping station system into the hydraulic transient model to predict pressure oscillations; output scheduling instructions to cope with pressure oscillations based on the scheduling strategy, and execute pumping station control according to the scheduling instructions.
[0048] Specifically, the first step is to collect and preprocess real-time system status data to provide timely and accurate input information for scheduling decisions. Data acquisition is achieved through the SCADA monitoring network of the cascade pumping station system. The collected data includes inlet and outlet pressures, flow rates, pump speeds, valve openings, motor currents, water levels in the reservoirs, and pressure data at key pipe sections for each pumping station. The acquisition frequency is determined based on the response speed of hydraulic transients to ensure timely capture of dynamic changes in the system status. The collected real-time data undergoes real-time preprocessing, including data validity verification (e.g., determining whether the data is within a reasonable range), real-time outlier removal (e.g., identifying instantaneous impulse noise using the sliding window averaging method), and data smoothing (e.g., using Kalman filtering to eliminate sensor measurement noise). The preprocessed real-time data must meet the real-time requirements of scheduling decisions, and data latency must be controlled within a preset range.
[0049] The pre-processed real-time system status is input into a lightweight hydraulic transient model for real-time prediction and risk assessment of pressure oscillations. The hydraulic transient model simulates the system's hydraulic dynamic response over a short period based on real-time status data, outputting pressure prediction curves for key pipeline sections, the time of pressure peak occurrence, and the amplitude of pressure fluctuations. Water hammer risk is assessed based on the prediction results, categorizing risks into no risk, low risk, medium risk, and high risk. Assessment criteria include the deviation between the predicted pressure value and the safe pressure range, the comparison between the pressure change rate and the equipment's tolerance limit, and the duration of pressure fluctuations. The risk assessment results will serve as a crucial basis for generating dispatch instructions; under high-risk conditions, priority will be given to ensuring system safety, while under medium- and low-risk conditions, energy efficiency optimization can be considered.
[0050] Based on real-time system status, prediction results from hydraulic transient models, and pre-trained scheduling strategies, scheduling instructions for the current operating conditions are generated. The generation process of scheduling instructions is as follows: First, the real-time system status and pressure prediction results are input into the pre-trained scheduling strategy model. The model outputs the optimal action plan according to its built-in decision logic (such as the Q-value selection mechanism of a deep Q-network). This plan includes the target values for pump speed adjustment, valve opening adjustment, or pump start / stop instructions for each pump station. Then, combined with the actual operating constraints of the current system (such as the current value and adjustment range of pump speed, and the current value and opening / closing rate limit of valve opening), the feasibility of the optimal action plan output by the model is verified and adjusted to ensure that the generated scheduling instructions are within the allowable range of the equipment's mechanical performance. Finally, the adjusted action plan is converted into specific control instructions, such as frequency adjustment instructions for pump inverters, opening control instructions for valve actuators, and start / stop control signals for pump stations. The format of the control instructions must match the communication protocol of the field control equipment to ensure that they can be correctly parsed and executed by the equipment.
[0051] Dispatch commands are issued to the field control units of each pumping station through the control network of the cascade pumping station system. The field control units then drive the actuators (such as frequency converters, valve actuators, and pumping station control cabinets) to execute the dispatch commands. During command execution, changes in the system state are monitored in real time. Data such as pressure, flow rate, and rotational speed after execution are collected through the SCADA network and compared with the prediction results of the hydraulic transient model to evaluate the effectiveness of the dispatch command execution. If the deviation between the actual system state and the prediction results exceeds a preset threshold, it indicates that the current dispatch strategy may not be able to adapt to real-time changes in the system (such as unforeseen hydraulic disturbances). In this case, a dynamic adjustment mechanism needs to be activated, feeding back the deviation data to the online fine-tuning module of the reinforcement learning agent. Based on the real-time deviation data, the dispatch strategy is updated with small parameters, generating a corrected dispatch command, which is then reissued for execution until the system state returns to a safe and efficient operating range. Through this closed-loop control process of "prediction-decision-execution-feedback-adjustment," real-time dynamic collaborative dispatching of the cascade pumping station system is achieved, continuously ensuring the system's safety and energy efficiency performance.
[0052] The aforementioned AI-based collaborative scheduling method for cascade pumping stations achieves dynamic prediction of pressure wave propagation by constructing a lightweight hydraulic transient model. Based on the model's prediction results and real-time system status, a reinforcement learning agent is built. Using historical operating data, a scheduling strategy capable of proactively identifying and responding to pressure oscillations is trained. Finally, during the online control phase, the real-time status is input into the model for pressure oscillation prediction, and compensatory scheduling instructions are generated based on the trained strategy. This achieves a fundamental shift from passively avoiding hydraulic transients to actively managing them. As a result, while ensuring the safe and stable operation of the system, control accuracy and equipment lifespan are significantly improved, operating energy consumption and maintenance costs are effectively reduced, and the system's adaptability to complex operating conditions is enhanced.
[0053] refer to Figure 2 In one optional embodiment, based on the real-time state data of the cascade pumping station system and the predicted output of the hydraulic transient model, a reinforcement learning agent is obtained by constructing a state space, action space, and reward function for reinforcement learning, including the following steps:
[0054] S11. Clean and timestamp-align the flow rate, pressure, pump speed, and valve opening data of each pump station in the real-time status data to obtain the cleaned system status dataset. Input the cleaned system status dataset into the hydraulic transient model, calculate the predicted value of pressure wave propagation, extract the pressure oscillation amplitude and dominant frequency features, and obtain the transient pressure feature vector.
[0055] Specifically, the first step is to clean the real-time status data of each pump station, including flow rate, pressure, pump speed, and valve opening. This cleaning process must eliminate outliers and noise that may have been introduced during data acquisition. For flow rate data, a sliding window median method is used to identify abrupt changes exceeding the normal operating range. The window length is set according to the data acquisition frequency (e.g., 5 data points when the acquisition frequency is 1Hz). If the deviation of a data point within the window from the median exceeds a preset proportion (e.g., 20%), it is considered an outlier and replaced with linear interpolation of adjacent normal data points. For pressure data, instantaneous pulse noise is detected using the first-order difference method. If the difference between the pressure data at a certain moment and the data at the previous moment exceeds the normal hydraulic fluctuation range of the system, a moving average method is used for smoothing. For pump speed and valve opening data, since both are directly driven by equipment control signals, data that significantly deviate from the equipment's mechanical range (e.g., pump speed exceeding the rated speed range, valve opening less than 0 or greater than 100%) must be removed to ensure the data conforms to the actual operating boundaries of the equipment.
[0056] After data cleaning, timestamp alignment is performed. Since the acquisition trigger times of monitoring sensors at different pump stations may vary slightly (e.g., millisecond delays), directly using the raw data would lead to misalignment of different parameters in the time dimension, affecting the model's accurate judgment of the system state. Timestamp alignment employs linear interpolation, using a unified system clock signal as a reference to map the acquisition time of all parameters to preset time series nodes (e.g., one node every 0.1 seconds). For data not acquired at a node, the interpolation result is calculated based on the values and time intervals of the two adjacent valid data points, ensuring that all parameters have corresponding data at the same time node. This ultimately results in a cleaned system state dataset with consistent time dimensions and meeting data quality standards.
[0057] The system state dataset after cleaning is input into the constructed lightweight one-dimensional nonlinear hydraulic transient model, and the predicted value of pressure wave propagation is obtained through model calculation. The core of pressure wave propagation prediction is based on the numerical solution of the Saint-Venant equations. The model simulates the pressure propagation process in each section of the pipeline over a future period (e.g., the next 10 seconds) based on input parameters such as real-time flow rate, pump speed, and valve opening, and outputs the pressure time series prediction curves for key sections of each pipeline (e.g., pump station inlet / outlet, pipeline midpoint). Based on these prediction curves, two core transient pressure characteristics, pressure oscillation amplitude and dominant frequency, are further extracted: the pressure oscillation amplitude is calculated using the peak-to-valley difference half-method, and its formula is:
[0058]
[0059] In the formula, Represents the amplitude of pressure oscillation. This represents the maximum value of the pressure within the predicted time series. This represents the minimum pressure within the predicted time series. This indicator directly reflects the strength of pressure fluctuations; the stronger the fluctuation, the higher the risk of water hammer the system faces.
[0060] The dominant frequency was calculated using a Fast Fourier Transform (FFT) to perform frequency domain analysis on the pressure time series. The formula is as follows:
[0061]
[0062] In the formula, Represents the dominant frequency. The frequency domain amplitude function represents the pressure time series obtained after Fast Fourier Transform. Representing the angular frequency, the `arg max` function is used to find the frequency corresponding to the maximum amplitude in the frequency domain. This indicator reflects the main periodic characteristics of pressure oscillations; an excessively high dominant frequency may increase the risk of pipeline resonance. The pressure oscillation amplitudes and dominant frequencies of each pipeline's key cross-sections are arranged in a preset order to form a transient pressure feature vector. This vector, together with the cleaned system state dataset, constitutes the basic input to the state space.
[0063] S12. Combine the macroscopic variables and transient pressure feature vectors in the system state dataset after cleaning to define the state space of the reinforcement learning agent and obtain the state vector; where the state vector includes the flow rate, pressure, pump speed, valve opening, pressure oscillation amplitude, dominant frequency and timestamp of each pump station.
[0064] Specifically, the macroscopic variables in the system status dataset directly reflect the overall operating conditions of the cascade pumping station system, including the real-time flow rate of each pumping station (reflecting water delivery capacity), inlet and outlet pressure (reflecting pipeline pressure status), pump speed (reflecting equipment operating intensity), and valve opening (reflecting flow regulation status). These variables are the core basis for the intelligent agent to perceive the current macroscopic operating status of the system. The absence of any variable may lead to an incomplete judgment of the system status by the intelligent agent (for example, the lack of pump speed data will prevent the intelligent agent from knowing the current output of the equipment, and thus from accurately assessing the impact of speed regulation actions).
[0065] The aforementioned macroscopic variables are combined with the transient pressure feature vector obtained in S11, and a timestamp parameter is introduced to jointly constitute the dimension of the state space. The introduction of the timestamp parameter aims to establish temporal correlations between states, enabling the agent to identify the evolutionary patterns of the system state over time (e.g., the increasing trend of pressure oscillation amplitude at a pumping station over time may indicate the accumulation of water hammer risk), avoiding the agent's decision-making based solely on static data at a single moment. The specific form of the state vector can be expressed as: ,in Represents the state vector. This represents the number of pumping stations in a cascade pumping station system. Representing the Real-time flow rate of each pumping station Representing the The inlet pressure of each pumping station, Representing the The outlet pressure of each pumping station, Representing the Pump speed of each pumping station Representing the The opening degree of the outlet valve of each pumping station Representing the Pressure oscillation amplitude at key sections of the inlet and outlet pipelines of each pumping station Representing the The dominant frequency of pressure oscillation at key cross-sections of the inlet and outlet pipelines of each pumping station. Represents the current timestamp.
[0066] The definition of the state space must ensure that it covers the key dimensions of system operation, while avoiding dimensional redundancy (such as avoiding the repeated introduction of auxiliary parameters that are highly linearly correlated with the core parameters), in order to balance the integrity of the state space and the computational efficiency of the agent. After constructing the state vector, the data of all dimensions are normalized. The min-max normalization method is used to map the data to the interval [0,1], and the formula is as follows:
[0067]
[0068] In the formula, Represents the normalized data. Represents the original data. This represents the minimum value of this dimension of data during its historical operation. This represents the maximum value of the data in this dimension throughout the historical process. The purpose of normalization is to eliminate the weight imbalance caused by differences in the units of measurement of data in different dimensions, and to ensure that the impact of data in each dimension on the agent's decision-making is reasonably reflected.
[0069] S13. Based on the physical constraints of pump station speed regulation and valve opening, define a continuous action space to obtain the action vector; wherein, the action vector includes the speed regulation increment and valve opening increment of each pump station.
[0070] Specifically, the reason for choosing incremental rather than absolute form is that: absolute form may cause the agent to output actions that exceed the real-time adjustment capability of the equipment (for example, if the current pump speed of a pumping station is 1000 r / min, if the action directly outputs an absolute target value of 1500 r / min, it may be necessary to adjust by 500 r / min in a single operation, which far exceeds the maximum speed regulation rate allowed by the motor), while incremental form can ensure that the action conforms to the mechanical characteristics constraints of the equipment by limiting the single adjustment range.
[0071] Before defining the action space, the physical constraints of pump station speed regulation and valve opening must be clearly defined. These constraints originate from the technical parameters provided by the equipment manufacturer and engineering practice experience. For pump station speed regulation increments, the core constraints include the maximum single-instance speed regulation increment and the maximum speed regulation rate per unit time. The maximum single-instance speed regulation increment is determined by the motor's torque characteristics. If the single-instance speed regulation increment is too large, it will cause a sudden increase in the motor stator current, exceeding the rated current limit and triggering overload protection. The maximum speed regulation rate per unit time is determined by the pump's hydraulic characteristics. An excessively fast speed regulation rate will cause a sudden change in the flow rate in the pipeline, triggering a strong water hammer effect. Based on the above constraints, the speed regulation increment range for each pump station can be determined as follows: ,in This represents the maximum speed adjustment increment in a single operation. Its value must meet the following requirements: the motor's single speed adjustment current does not exceed 1.2 times the rated current, and the speed adjustment rate per unit time does not exceed the maximum allowable rate of the equipment (e.g., 50 r / min·s).
[0072] For valve opening increments, the core constraints include the maximum single-cycle opening increment and the valve opening / closing rate. The maximum single-cycle opening increment is determined by the valve's sealing characteristics and structural strength; an excessively large increment can cause severe impact between the valve disc and seat, shortening the valve's lifespan. The valve opening / closing rate is determined by the pipeline's hydraulic stability requirements; an excessively rapid rate can lead to sudden changes in flow rate within the pipeline, causing pressure fluctuations. Based on these constraints, the opening increment range for each pump station outlet valve can be determined as follows: ,in This represents the maximum opening increment in a single operation. Its value must meet the following requirements: the impact force of a single valve opening and closing operation does not exceed the design limit, and the opening and closing rate does not exceed the maximum allowable rate in the project (e.g., 5% / s).
[0073] Based on the above physical constraints, the continuous action space of a reinforcement learning agent can be defined as a multi-dimensional continuous space composed of all pump station speed increments and valve opening increments, with the corresponding action vector form as follows:
[0074]
[0075] In the formula, Represents the action vector. This represents the number of pumping stations in a cascade pumping station system. Representing the Speed increase of each pumping station Representing the The opening increment of the outlet valve of each pump station. Each motion dimension must meet the corresponding physical constraints, i.e. , .
[0076] The reason for using a continuous rather than discrete action space is that the scheduling needs of a cascade pumping station system are continuous (e.g., requiring fine-tuning of flow based on actual water demand). A discrete action space would result in insufficient adjustment precision, making fine-grained scheduling impossible. A continuous action space, on the other hand, provides an unlimited number of action options, allowing the agent to output adjustments that better match actual needs. To ensure that the agent's output actions always conform to physical constraints, a constraint check is added before action execution: if the agent's output action increment exceeds a preset range, it is automatically truncated to the nearest constraint boundary value (e.g., the agent's output...). ,and Then the actual execution (Adjusted to 50 r / min) to further ensure the engineering feasibility of the action.
[0077] S14. The total reward value is obtained by calculating the weighted sum of the energy efficiency reward, water volume reward, and pressure oscillation penalty.
[0078] Specifically, by constructing a weighted combination model of energy efficiency reward, water consumption reward, and pressure oscillation penalty, multi-dimensional feedback on the scheduling behavior of the reinforcement learning agent is achieved, guiding the agent to balance energy efficiency optimization and water demand satisfaction under the premise of safe operation. The total reward value is calculated using a weighted summation formula, the specific expression of which is:
[0079]
[0080] In the above formula, Representing the total reward value, it is the core indicator for the agent to judge the quality of scheduling actions; , , These represent the weighting coefficients for energy efficiency rewards, water volume rewards, and pressure oscillation penalties, respectively. The sum of these three is 1. The weighting values are determined through offline simulation experiments and sensitivity analysis and need to be dynamically adjusted according to the operational priority of the cascade pumping station system (e.g., the weight of the water volume reward needs to be increased during the dry season, and the weight of the pressure oscillation penalty needs to be increased during the equipment aging period). This represents an energy efficiency award, used to evaluate the energy utilization efficiency of dispatching actions; This represents a water volume reward item, used to assess the degree to which scheduling actions meet water demand; This represents a pressure oscillation penalty term, used to suppress pressure fluctuation behavior that may cause safety risks.
[0081] Energy Efficiency Bonus The calculation is based on the ratio of the system's actual total energy consumption to the optimal baseline energy consumption, expressed as:
[0082]
[0083] In the formula, The optimal baseline energy consumption of the system under the current water demand is calculated by the hydraulic transient model and equipment efficiency curve. Specifically, it is the total energy consumption of each pumping station operating at the efficiency peak point when the current total water supply is met. This represents the total actual energy consumption of the system after the scheduling action is executed. It is calculated as the sum of the energy consumption of each pump station. The energy consumption expression for a single pump station is: (in The density of water, It is the acceleration due to gravity. For the pump station flow rate, For the pump station head, The pump station operating efficiency is obtained by interpolation of the pump's HQ-η characteristic curve. The closer the actual total energy consumption is to the optimal baseline energy consumption, the better. The closer to 1, the higher the energy efficiency bonus; if the actual total energy consumption is higher than the optimal baseline energy consumption, If the value is less than 1, the energy efficiency reward is reduced, thereby incentivizing the agent to choose a low-energy scheduling scheme.
[0084] Water volume bonus The calculation is based on the deviation between the actual water supply and the demand for water supply, and the expression is:
[0085]
[0086] In the formula, The total actual water supply of the system after the scheduling action is executed is the sum of the outlet flow of each pump station (collected in real time by flow sensors or calculated by a hydraulic transient model). This represents the system's water demand for the current period, determined based on the production, domestic, or ecological water use plans of the downstream water-using areas. This is a deviation coefficient used to adjust the impact of water volume deviation on the reward value. For example, its value should be chosen to ensure that the water volume deviation is less than 5%. The value should be no less than 0.9 to encourage agents to prioritize water demand; when the water deviation exceeds 15%, A value below 0.3 creates a clear reward and penalty system to prevent water supply from deviating significantly from demand.
[0087] Pressure oscillation penalty item The calculation is based on the pressure oscillation amplitude and dominant frequency extracted from S11, and the expression is:
[0088]
[0089] In the formula, This represents the pressure oscillation amplitude, which is the difference between the peak and trough values in the pressure fluctuation curve (predicted by the hydraulic transient model or calculated from real-time pressure data). This represents the dominant frequency of pressure oscillations, which is the frequency component with the highest energy proportion in pressure fluctuations (obtained through fast Fourier transform analysis of pressure time series). This is a penalty coefficient, and its value needs to be determined in conjunction with the tolerance capacity of the pipeline and equipment. Exceeding the safe oscillation amplitude threshold or When approaching the natural frequency of the pipe, Automatically increases the penalty intensity to prevent pipeline fatigue damage or resonance risks caused by pressure oscillations. When there is no significant pressure oscillation ( Less than the threshold and When the frequency is below the low-frequency threshold, When the value approaches zero, the penalty effect disappears; as the amplitude and frequency of the oscillation increase, The linear increase suppresses scheduling actions that might cause oscillations by reducing the total reward value.
[0090] S15. Construct a reinforcement learning agent based on the state vector, action vector, and total reward value.
[0091] Specifically, the S15 stage uses the state vector defined in S12, the action vector defined in S13, and the total reward value calculated in S14 as the core elements. By building the network architecture, designing the training mechanism and interaction logic, the reinforcement learning agent is constructed, enabling it to output the optimal scheduling action based on the system state.
[0092] First, the intelligent agent network architecture is designed, adopting an actor-critic dual-network architecture. This architecture can simultaneously achieve action generation and value evaluation, improving decision-making accuracy and training stability. The "actor network" is responsible for generating scheduling actions that conform to action space constraints based on the input state vector. The network uses a fully connected neural network structure, with the input layer dimension consistent with the state vector dimension (i.e., the sum of the number of variables in the state vector). There are 2-3 hidden layers, with the number of neurons in each layer being 1.5-2 times the input layer dimension. The ReLU activation function is used to enhance nonlinear fitting ability. The output layer dimension is consistent with the action vector dimension (i.e., the sum of the number of speed increments and valve opening increments for each pump station). The Tanh activation function is used in the output layer to map the output action values to a preset action constraint range (e.g., speed increments are mapped to...). Valve opening increment mapped to This ensures that the generated action vectors satisfy the physical constraints defined in S13.
[0093] The "critic network" is responsible for evaluating the long-term value (i.e., the expected total reward in the future) of an action based on the input state vector and action vector. The network also uses a fully connected neural network structure. The input layer dimension is the sum of the state vector dimension and the action vector dimension. The hidden layer structure is consistent with the actor network, and the output layer is a single neuron, with the output value being the action value estimate. (in Represents the state vector. (Representing the action vector), used to guide the parameter updates of the actor network. To avoid overfitting, dropout layers are added to the hidden layers of both networks, with dropout probabilities set to 0.1-0.2. Batch normalization is also used to process the input data, accelerating network training convergence.
[0094] Secondly, a training interaction mechanism for the agent is designed, employing the Temporal Difference (TD) algorithm to update network parameters. During training, the agent interacts with the simulation environment (constructed based on historical operating data and a hydraulic transient model). At each time step, the agent obtains the current state vector from the environment. Generate motion vectors through actor networks ,Will After the environment is input, the environment performs an action and returns the next state vector. Total reward value of the current step ; Empirical tuples The data is stored in the experience replay pool and updated using a first-in, first-out (FIFO) strategy to ensure the timeliness of the experience data.
[0095] When the number of experience tuples stored in the experience replay pool reaches a preset threshold (e.g., 1000), batch sampling training begins. A batch of experiences (batch size set to 32-64) is randomly drawn from the experience replay pool, and the commentator network is used to calculate the value of the current action. Value of the target action The calculation of the target action value uses the time-series difference error formula:
[0096]
[0097] In the formula, This is a discount factor used to weigh the importance of current rewards against future rewards. The closer the value is to 1, the more the agent focuses on long-term rewards. For state-based The next action vector is generated. The commentator network parameters are updated by minimizing the mean squared error (MSE) between the current action value and the target action value; the loss function expression is:
[0098]
[0099] In the formula, The number of empirical tuples in the batch sampling. This serves as the index for experience tuples. Simultaneously, a policy gradient algorithm is used to update the actor network parameters, aiming to maximize action value. Gradient calculation is based on the action value output by the commentator network, and the loss function expression is:
[0100]
[0101] The gradient is calculated using the backpropagation algorithm, and the weight parameters of the two networks are updated using the Adam optimizer. The number of training iterations is determined based on the convergence status. For example, when the average total reward value on the validation set fluctuates by less than 5% for 100 consecutive iterations, the network training is considered to have converged.
[0102] Finally, the online decision-making logic of the agent is constructed, completing the overall construction of the agent. After training convergence, the agent outputs scheduling actions during actual operation through the following process: first, it receives real-time collected and preprocessed system state data, and generates a state vector conforming to the definition of S12. ;Will Input the actor network to generate initial motion vectors ;right Physical constraint verification is performed. If a certain motion component (such as the speed regulation increment of a pumping station) exceeds the constraint range defined in S13, it is adjusted to the nearest constraint boundary (if it exceeds the maximum speed regulation increment, the maximum increment value is taken) to obtain the final motion vector. ;Will The commands are converted into actual scheduling instructions (such as frequency adjustment values for pump inverters and opening adjustment values for valve actuators) and sent to field equipment. Simultaneously, the agent receives real-time system feedback (new state vector and total reward value) after the actions are executed. Through an online fine-tuning mechanism (updating network parameters using small-batch experience, reducing the learning rate to 1 / 10 of the training phase), the network is continuously optimized to ensure the agent maintains decision-making performance even when system conditions change. Ultimately, this results in a reinforcement learning agent with dynamic adaptability and multi-objective optimization capabilities.
[0103] In one optional embodiment, the total reward value is obtained by calculating a weighted sum of the energy efficiency reward, water volume reward, and pressure oscillation penalty, including the following steps:
[0104] S21. Based on the flow rate, pressure, and efficiency parameters of the pumping stations, calculate the power consumption of each pumping station using the power consumption calculation formula to obtain the power consumption value of each pumping station; sum and negate the power consumption values of each pumping station to obtain the energy efficiency bonus item; the power consumption calculation formula is:
[0105]
[0106] in, This represents the power consumption value of pump station i. This indicates the density of water. Represents gravitational acceleration. This represents the flow rate value of pump station i. Indicates the pressure head of pump station i. This represents the efficiency of pump station i.
[0107] Specifically, this step focuses on the energy utilization efficiency of pump station operation. By calculating the actual power consumption of each pump station, an energy efficiency bonus is obtained to guide scheduling behavior towards low energy consumption optimization. The calculation process uses the pump station's flow rate, pressure, and efficiency as core input parameters: First, for each pump station, a specific power consumption calculation formula is used to obtain the power consumption value of a single pump station. This formula accurately reflects the energy loss during the process of converting electrical energy into hydraulic energy. Its expression is: In the formula Represents a single pumping station The power consumption value reflects the energy consumption level of the pump station during operation; The density of water is an inherent physical property of water. It represents the acceleration due to gravity, reflecting the effect of gravity on the work done by the water flow; Representative pump station The actual flow rate reflects the water conveyance capacity of the pumping station; Representative pump station The pressure head reflects the energy required by the pumping station to lift the water flow. Representative pump station The operating efficiency reflects the effectiveness of energy conversion at the pumping station.
[0108] After obtaining the power consumption value of each pump station Then, the power consumption values of all pump stations are summed to obtain the total power consumption of the cascade pump station system. Since the energy efficiency bonus must be negatively correlated with energy consumption—that is, the lower the total power consumption, the higher the bonus value—the total power consumption is negatively evaluated to obtain the final energy efficiency bonus. This computational method enables intelligent agents to prioritize scheduling schemes with lower total power consumption when making decisions, thereby optimizing system energy efficiency.
[0109] S22. Based on the total flow and planned flow of the cascade pumping station system, calculate the absolute deviation between the actual flow and the planned flow; take the negative value of the absolute deviation to obtain the water volume bonus; the calculation formula for the water volume bonus is:
[0110]
[0111] in, This indicates a water volume reward item. Indicates the number of pumping stations. This represents the flow rate of pump station i. This indicates the planned flow rate.
[0112] Specifically, this step focuses on the completion rate of water conveyance tasks in the cascade pumping station system. By comparing the deviation between the actual total flow and the planned flow, water volume rewards are generated to ensure that scheduling actions meet the preset water conveyance requirements. First, two core flow indicators are defined: one is the actual total flow of the cascade pumping station system, which is determined by analyzing the individual flow of all pumping stations. Summing yields, i.e. (in (1) The total number of pumping stations in the cascade pumping station system; 2) The planned flow rate. This value is predetermined based on the production, domestic or ecological water demand of the downstream water-using area and is the water transfer target that the system needs to achieve.
[0113] During the calculation, first determine the absolute deviation between the actual total flow and the planned flow, i.e. This absolute deviation reflects the degree of deviation between the actual water delivery and the planned target—the larger the deviation, the lower the completion rate of the water delivery task. To ensure a positive correlation between the reward items and the task completion rate, the absolute deviation is negatively evaluated to obtain the water volume reward item. The calculation formula is as follows: When the actual total flow is exactly the same as the planned flow, the absolute deviation is 0 and the water reward is 0 (the reward is the highest at this time). The more the actual total flow deviates from the planned flow, the greater the absolute deviation and the smaller the negative value of the water reward. This incentivizes the agent to adjust the scheduling action so that the actual water delivery of the system is as close as possible to the planned flow.
[0114] S23. Based on the pressure oscillation amplitude and dominant frequency in the transient pressure eigenvector, the pressure oscillation penalty term is calculated through linear combination; the formula for calculating the pressure oscillation penalty term is:
[0115]
[0116] in, This indicates a penalty term for pressure oscillations. and Indicates the scaling factor; Indicates the amplitude of pressure oscillation; Indicates the dominant frequency.
[0117] Specifically, this step addresses the pressure safety risks in the operation of a cascade pumping station system. Based on key parameters in the transient pressure feature vector, it calculates a pressure oscillation penalty term to suppress pressure fluctuations that may cause equipment damage or system instability. The core parameters required for the calculation are derived from the previously extracted transient pressure feature vector, specifically the pressure oscillation amplitude. and dominant frequency :in It reflects the strength of pressure fluctuations, that is, the difference between the pressure peak and the trough. It reflects the frequency of pressure fluctuations, that is, the main frequency of periodic pressure changes.
[0118] The pressure oscillation penalty term is obtained by linearly combining these two parameters, and its calculation formula is as follows: In the formula This represents the penalty for pressure oscillation; the larger the value, the higher the safety risk caused by pressure oscillation. and This is a scaling factor used to adjust the weights of the pressure oscillation amplitude and dominant frequency on the penalty term, respectively. If the system is more sensitive to pressure fluctuations (e.g., due to severe pipe aging), it can be increased. The value of can be increased if the system is more prone to resonance from high-frequency fluctuations (such as long straight pipe sections). The value of is determined through this linear combination, allowing the penalty term to fully reflect the safety risk of pressure oscillations. or As the penalty term increases, the agent is guided to avoid scheduling actions that exacerbate stress oscillations.
[0119] S24. The energy efficiency bonus, water consumption bonus, and pressure oscillation penalty are weighted and summed to generate the total bonus value; the formula for calculating the total bonus value is:
[0120]
[0121] in, This represents the total reward value. Indicates energy efficiency bonus items, This represents the weighting coefficient.
[0122] Specifically, this step integrates the energy efficiency reward, water volume reward, and pressure oscillation penalty, and generates a total reward value through weighted summation, providing a unified evaluation standard for the reinforcement learning agent's decision-making. The total reward value needs to comprehensively balance the system's energy efficiency, water delivery task completion rate, and pressure safety; therefore, the importance of each component needs to be adjusted using weighting coefficients during calculation. The calculation formula is as follows:
[0123]
[0124] In the formula, The total reward value is the core basis for the agent to judge the quality of scheduling actions; The energy efficiency bonus calculated for S21 The water volume bonus calculated for S22 The pressure oscillation penalty term calculated for S23; This is a weighting coefficient used to adjust the influence of the pressure oscillation penalty term on the total reward value. It can be increased when the system is in a high-safety-risk scenario (such as water transfer during flood season). The value of this parameter increases the pressure safety weight, prioritizing system stability; when the system is in normal operating conditions, it can be appropriately reduced. While prioritizing safety, the focus is more on energy efficiency and water consumption targets.
[0125] The total reward value calculated using this formula can fully reflect the comprehensive performance of scheduling actions under multi-objective conditions: if a certain scheduling action can reduce the total power consumption ( (increased), approaching planned flow ( Increase) and suppress pressure oscillations ( (Decrease), then the total reward value The efficiency will increase significantly, and the agent will be more inclined to choose such actions to achieve multi-objective collaborative optimization of the cascade pumping station system.
[0126] In one optional embodiment, a lightweight one-dimensional nonlinear hydraulic transient model is constructed based on the pipeline geometry parameters, pump station equipment parameters, and historical SCADA data of the cascade pumping station system, resulting in the hydraulic transient model, including the following steps:
[0127] S31. Based on the pipeline geometric parameters, pump station equipment parameters, and historical SCADA data of the cascade pumping station system, and combined with the one-dimensional unsteady flow control equations, the continuity equation and momentum equation are derived to obtain the basic physical model; the continuity equation is:
[0128]
[0129] The momentum equation is:
[0130]
[0131] in, Represents the cross-sectional area of the pipe. Indicates flow rate. Indicates time, Indicates the distance along the pipe. Represents gravitational acceleration. Indicates pressure head, This represents the Darcy-Weisbach friction coefficient. Indicates the pipe diameter.
[0132] Specifically, this step uses the core foundational data of the cascade pumping station system as support, and combines it with the one-dimensional unsteady flow control equations to derive the continuity and momentum equations describing the hydraulic transient process, thus constructing the physical basis of the model. Pipeline geometric parameters (such as pipe diameter and wall thickness) directly determine the pipe's cross-sectional area. Calculation ( , The parameters of the pump station equipment (such as the pump's head-flow characteristics and valve regulation characteristics) are used to set the boundary conditions of the subsequent equations, while historical SCADA data (such as past pressure and flow time series data) provide a reference for verifying the rationality of the equations.
[0133] The derivation revolves around the one-dimensional unsteady flow governing equation, which is the core theoretical basis for characterizing the unsteady motion of water flow within a pipe. Among these equations, the continuity equation reflects the law of conservation of water mass, meaning that the rate of change of water mass within the pipe equals the rate of change of flow rate along the pipe's axial direction, and its expression is:
[0134]
[0135] In the formula, Represents the cross-sectional area of the pipe, which varies with the pipe diameter. Changes (if the pipe is of constant diameter) It is a constant value; if it is a reducing pipe, along (Change of direction) It represents the water flow rate within the pipe, reflecting the amount of water passing through a certain cross section per unit time. It represents time and depicts the dynamic changes in hydraulic parameters; This represents the distance along the pipe's axial direction, used to locate different cross-sections of the pipe. The equation shows that as the cross-sectional area of a pipe increases over time, the flow rate along the pipe's axial direction decreases accordingly, and vice versa, ensuring the conservation of water flow mass.
[0136] The momentum equation reflects the law of conservation of water flow, taking into account the effects of inertial force, gravity, friction, and other factors on water flow. Its expression is:
[0137]
[0138] In the formula, It represents the acceleration due to gravity and is a key parameter affecting the effect of gravity. Represents pressure head, reflecting the sum of pressure energy and positional energy of water flow at the pipe cross-section; Represents the Darcy-Weisbach friction coefficient, used to quantify the frictional resistance of the pipe's inner wall to water flow (its value is related to the pipe's roughness and the water flow Reynolds number, and can be determined through empirical formulas or experimental data). Represents the pipe diameter, and These factors collectively affect the flow capacity and frictional resistance of the water. Each term in the equation corresponds to the rate of change of the local inertial force of the water flow (…). ), convective rate of change of inertial force ( ), Gravity term ( ) and frictional resistance term ( By analyzing the balance relationships among these terms, the momentum change pattern of water flow in unsteady motion can be fully characterized.
[0139] S32. Linearize the basic physical model and use the method of characteristics to transform the partial differential equations into ordinary differential equations to obtain the linearized model.
[0140] Specifically, this step addresses the complexity of the fundamental physical model by using linearization and the method of characteristics to transform the partial differential equations, which are difficult to solve directly, into ordinary differential equations that are easier to compute. This lays the foundation for subsequent model lightweighting and real-time computation.
[0141] First, the fundamental physical model is linearized. This is because the continuity and momentum equations in the fundamental physical model contain nonlinear terms (such as...). , Directly solving this problem faces issues of high computational cost and poor convergence, making it difficult to meet the real-time requirements of online scheduling. The core idea of linearization is to perform a Taylor expansion of the nonlinear terms in the equation around a certain steady-state operating point (selecting common steady-state conditions from historical SCADA data, such as those corresponding to rated flow and design pressure), ignoring higher-order terms and retaining only first-order linear terms. For example, for... Let the flow rate at the steady-state operating point be... Cross-sectional area is The Taylor expansion is approximately equal to This approximation transforms the nonlinear relationships in the equations into linear ones, reducing the difficulty of solving the equations. Simultaneously, the linearization process must ensure that the error is controlled within a preset range (e.g., the deviation between the linearized model's predicted value and the original nonlinear model's predicted value is less than 5%), avoiding oversimplification that could lead to a loss of model accuracy.
[0142] The linearized partial differential equations are then transformed into ordinary differential equations using the method of characteristics. The core principle of the method of characteristics is to find a set of characteristic lines that transform the partial differential equations into ordinary differential equations along these lines. The specific steps are as follows: based on the linearized continuity and momentum equations, eliminate the time or spatial partial derivatives to derive the characteristic line equations (describing the variation of hydraulic parameters along a specific direction) and the compatibility equations (ordinary differential equations that hold along the characteristic lines). For example, by simultaneously solving the two linearized equations, two characteristic line directions can be obtained (corresponding to the downstream and upstream propagation directions of the pressure wave, respectively). Along each characteristic line direction, the flow rate... With pressure head The changes satisfy a specific ordinary differential relation, namely the compatibility equation. Through this transformation, what was originally required in the two-dimensional spacetime domain ( The partial differential equations solved in the domain are transformed into ordinary differential equations solved on one-dimensional characteristic lines, which greatly simplifies the calculation process while preserving the core characteristics of pressure wave propagation in hydraulic transients (such as propagation direction and propagation speed), ensuring that the physical meaning of the model is not lost. Finally, through linearization and the method of characteristics, a linearized model that combines accuracy and computability is obtained.
[0143] S33. By dividing the pipeline into multiple grid points and using an implicit difference scheme to solve for pressure and head, the linearized model is discretized to obtain a discretized model.
[0144] Specifically, this step discretizes the linearized model in both space and time, transforming continuous hydraulic parameter changes into numerical calculations at discrete grid points, thus enabling the computer implementation of the model.
[0145] First, the pipeline is meshed. The mesh density is determined based on the actual pipeline length and the gradient of hydraulic parameter changes (e.g., areas with drastic hydraulic changes such as pipe diameter changes and pump station inlets / outlets). For straight pipe sections with gradual hydraulic parameter changes, the mesh spacing can be appropriately increased (e.g., one mesh point every 100-200m); for areas with drastic hydraulic changes (e.g., before and after valves, at pipe reducers), the mesh spacing needs to be decreased (e.g., one mesh point every 20-50m) to ensure accurate capture of local hydraulic transients. During mesh generation, the entire pipeline network of the cascade pump station system (including connecting pipes between pump stations and inlet / outlet pipes within pump stations) is uniformly divided into a continuous sequence of mesh nodes, with each mesh node corresponding to a spatial location. ( , (Total number of grid nodes), and record the pipe parameters at each node (such as diameter). Cross-sectional area coefficient of friction This forms a discrete spatial computational domain.
[0146] Subsequently, an implicit difference scheme is used to discretize the ordinary differential equations of the linearized model in time. The core advantage of the implicit difference scheme lies in its good stability and the ability to use a larger time step (determined based on the pressure wave propagation velocity and grid spacing, typically taking a value that can be taken as...). , (For the propagation speed of the pressure wave), balancing computational efficiency and stability. Specifically, for each time step... (For example, taking 0.01-0.1s), the time domain is divided into discrete time nodes. ( And assume that within each time step, hydraulic parameters (such as flow rate) Pressure head The changes in pressure head satisfy specific difference relationships (such as first-order backward difference, second-order central difference). For example, in time step partial derivatives It can be approximated as , Indicates the first The grid node at the ... Pressure head value at each time step This represents the value at the previous time step.
[0147] By combining spatially discrete grid nodes with time-discrete difference schemes and substituting them into the ordinary differential equations of the linearized model, we can obtain the hydraulic parameters of each grid node at the current time step. , The system of linear algebraic equations is used to solve the system of equations. By solving this system of equations (using numerical methods such as Gaussian elimination and conjugate gradient method), the pressure and head values of each grid node at each time step can be obtained, realizing the discretization simulation of the hydraulic transient process, and finally forming a discretized model.
[0148] S34. Based on the pressure and flow readings in historical SCADA data, the discretized model is calibrated using the least squares method to obtain the calibrated hydraulic transient model.
[0149] Specifically, this step uses historical SCADA data to calibrate the parameters of the discretized model, reducing the deviation between the model's predicted values and the actual system operating data, and improving the model's accuracy and reliability.
[0150] First, historical SCADA data is screened and preprocessed. Pressure and flow data covering different operating conditions (including steady-state and transient conditions) are selected from the historical SCADA database. Steady-state data (such as pressure and flow readings during long-term stable operation of a pumping station) are used to calibrate the steady-state error of the model, while transient data (such as pressure fluctuation data during pump start-up and shutdown, and valve adjustment) are used to calibrate the model's simulation accuracy for transient processes. The preprocessing process includes: removing outliers (such as sudden pressure increases or negative flow rates caused by sensor malfunctions), smoothing high-frequency noise using a moving average method (to avoid noise interference with calibration results), and aligning the timestamps of the pressure and flow data with the model's time step (to ensure time matching between data and model calculation nodes), ultimately forming the dataset for calibration. ,in , The first The monitoring point corresponding to the grid node is at the [number]th [node]. Actual pressure head and flow rate data at each time step.
[0151] The least squares method was then used for model calibration. The core objective of calibration was to adjust key parameters in the discretized model (such as the Darcy-Weisbach friction coefficient). Pipe cross-sectional area Correction factor, pressure wave propagation speed (Correction factors, etc.) to minimize the sum of squared residuals between the model predictions and the actual data. Let the predicted pressure head and flow rate be respectively... , ( The vector of parameters to be calibrated, such as , for Correction factor, for (assuming the correction coefficients are used), then the objective function of the least squares method is:
[0152]
[0153] In the formula, The weighting coefficient for the flow residual (determined based on the accuracy of the pressure and flow monitoring data; for example, if the pressure sensor has higher accuracy). A value of 0.5-0.8 can be used. The objective function is solved using numerical optimization algorithms (such as gradient descent or L-BFGS). The minimum value is used to obtain the optimal calibration parameters. .
[0154] After calibration, the model accuracy is verified: the calibrated model is applied to a test dataset that was not calibrated (separated from historical SCADA data), and the root mean square error (RMSE) and mean absolute percentage error (MAPE) between the model's predicted values and the actual data are calculated. For example, if the RMSE of the pressure head is ≤0.02MPa, the RMSE of the flow rate is ≤0.05m³ / s, and the MAPE is ≤5%, the model calibration is considered successful, and the final hydraulic transient model is obtained. If the accuracy requirements are not met, the data is re-screened or the range of parameters to be calibrated is adjusted, and the calibration process is repeated until the model accuracy meets the standards.
[0155] In one optional embodiment, historical system operation data of the cascade pumping station system is used to train a reinforcement learning agent by simulating state transitions and reward calculations in a simulation environment, and outputting a pre-trained scheduling strategy, including the following steps:
[0156] S41. Based on the hydraulic transient model and pump station system dynamics, a simulation environment is constructed for simulating state transitions and reward calculations.
[0157] Specifically, this step relies on the hydraulic transient model and the dynamic characteristics of the pumping station system to build a simulation environment capable of accurately simulating the state transition process and reward calculation logic of a cascade pumping station system, providing a training platform for reinforcement learning agents that closely approximates real-world operating scenarios. The construction of the simulation environment requires the integration of two core models: first, the previously established hydraulic transient model, which simulates the dynamic changes in hydraulic parameters caused by scheduling actions. For example, when the agent outputs actions to adjust pump speed or valve opening, the hydraulic transient model can calculate the pressure and flow rate changes at each pipe section after the action is executed, providing hydraulic-dimensional parameter support for state transition; second, the pumping station system dynamic model, which characterizes the dynamic response characteristics of the equipment itself, including the time delay for adjusting pump speed from the current value to the target value, the rate constraint of valve opening changes, and the dynamic matching relationship between motor output power and speed. These characteristics directly affect the timeliness and accuracy of system state changes after action execution and need to coordinate with the output of the hydraulic transient model to jointly complete the state transition simulation.
[0158] In the state transition simulation logic design, the simulation environment adopts the process of "action input - multi-model collaborative calculation - new state output". After receiving the action vector output by the reinforcement learning agent, the action is first decomposed into specific operation instructions for each pump station (such as the speed adjustment increment of a pump station, the opening increment of a valve), and the pump station system dynamic model is combined to determine whether the instructions meet the physical constraints of the equipment. If they do, the equipment execution trajectory is generated (such as the curve of speed change over time, the curve of opening change over time). Then, the execution trajectory is input into the hydraulic transient model, and the hydraulic parameters such as pressure, flow rate, and head loss of the system at different time points are calculated by the model. Finally, the state parameters corresponding to the equipment execution trajectory (such as real-time speed and real-time opening) are integrated with the hydraulic parameters output by the hydraulic transient model to form the new state of the system after the action is executed, thus completing one state transition simulation.
[0159] The design of the reward calculation module is consistent with the previously defined logic for calculating the total reward value. This module extracts key parameters required for reward calculation from the state transition process. For example, it obtains the flow rate, pressure, and efficiency parameters of each pump station from the new state to calculate the energy efficiency reward, obtains the deviation between the total system flow and the planned flow to calculate the water volume reward, and extracts the pressure oscillation amplitude and dominant frequency from the hydraulic transient model output to calculate the pressure oscillation penalty. Then, it obtains the reward value corresponding to the action through weighted summation. The reward value and "initial state-action-new state" are stored together as empirical data to provide feedback for subsequent agent training.
[0160] S42. Based on the preset network parameter configuration, the network parameters and experience replay buffer of the reinforcement learning agent are initialized to obtain the initialized agent.
[0161] Specifically, this step lays the foundation for the training of the reinforcement learning agent by pre-configuring and initializing the network parameters and experience replay buffer, ensuring a stable start and efficient progress in the training process. The initialization of network parameters must be carried out according to the pre-configured network architecture. If the agent uses an actor-critic dual-network architecture, the parameters of the actor network and the critic network are initialized separately: the actor network is responsible for generating actions based on the state, and its parameters include the weights and biases of neurons in each layer. Random initialization (such as Xavier initialization) should be used during initialization to ensure that the network output is within a reasonable range in the early stages of training, avoiding training divergence due to extreme parameter values; the critic network is responsible for evaluating the value of actions, and its parameter initialization method is consistent with that of the actor network. It is also necessary to ensure that the initialization parameters of the two networks are independent of each other to avoid initial correlation affecting training stability.
[0162] The function of the experience replay buffer is to store experience tuples generated during the interaction between the agent and the simulation environment. Each experience tuple contains four types of information: initial state, executed action, corresponding reward value, and new state after transition. This information is the core data source for subsequent batch training. The maximum storage capacity of the buffer is defined during initialization, based on training data requirements and computational resources. Data storage rules are also set. A first-in, first-out (FIFO) strategy is adopted; when the buffer reaches its maximum capacity, the oldest stored experience tuple is automatically deleted to ensure the timeliness of the data in the buffer and avoid interference from outdated data on training results. Furthermore, the buffer's index pointer and data count variable must be initialized for subsequent experience tuple storage and retrieval operations, ensuring the orderly interaction of data.
[0163] S43. Sample the initial state from the historical system operation data; based on the initialized agent, obtain the output action by inputting the initial state into the simulation environment; based on the output action, simulate the state transition and reward calculation through the simulation environment to obtain the updated system state and the corresponding reward value.
[0164] Specifically, this step extracts the initial state from historical system operation data, driving the reinforcement learning agent to interact with the simulation environment for the first time, completing the initial action output, state transition simulation, and reward calculation, generating the first batch of experience data required for training. The sampling of the initial state follows the principle of "covering multiple operating conditions and ensuring data validity." Data segments containing different operating scenarios are selected from historical system operation data, including water demand scenarios in different seasons, operating scenarios with different equipment combinations, and different hydraulic load scenarios, to avoid the sampling data being concentrated in a single operating condition, which could lead to insufficient generalization ability of the agent. At the same time, the validity of the sampled raw data is verified, removing segments with outliers (such as pressure exceeding the safe range or negative flow) or missing data, ensuring that the sampled initial state accurately reflects the parameter distribution during normal system operation, providing reliable starting conditions for subsequent interactions.
[0165] Once the initial state is determined, it is input into the pre-initialized agent. The agent generates output actions based on its current policy. At this stage, the agent is in the early training phase, and the policy is primarily exploratory, typically employing an ε-greedy exploration strategy. This strategy involves randomly selecting an action to explore the action space with a certain probability, and generating what is considered the optimal action based on the current network parameters with the remaining probability. This strategy helps the agent fully explore the impact of different actions on the system state in the early training phase, avoiding getting trapped in local optima. The generated output actions need to be converted into a format recognizable by the simulation environment, i.e., a combination of speed increments and valve opening increments consistent with the action vector definition, before being input into the simulation environment.
[0166] Upon receiving the output action, the simulation environment initiates the state transition simulation process. Combining the pump station system dynamics model and the hydraulic transient model, it calculates the changes in various system parameters after the action is executed, obtaining the updated system state. This state includes the rotational speed, valve opening, flow rate, and pressure of each pump station after the action is executed, as well as the hydraulic parameters of each pipeline cross-section. Simultaneously, the reward calculation module is activated. Based on the parameters in the updated system state, it generates the corresponding reward value according to the preset total reward value calculation logic. The total reward value is calculated using a pre-defined formula, the expression of which is:
[0167]
[0168] In the formula, This represents the reward value corresponding to this action. Represents energy efficiency awards. Represents water volume reward items, This represents the penalty term for pressure oscillation. , , These represent the weight coefficients of the three sub-items, used to balance the importance of different optimization objectives. After calculating the updated system state and the corresponding reward value, the "initial state - output action - reward value - updated system state" is combined into an experience tuple and stored in the experience replay buffer to complete one interaction process.
[0169] S44. Using the proximal policy optimization algorithm, based on the updated system state and the corresponding reward value, the initial agent is trained in batches through experience replay. When the reward value of the initial agent converges or reaches the preset maximum number of iterations, the training process is terminated and the trained scheduling policy is output.
[0170] Specifically, this step employs the proximal policy optimization algorithm, using batches of experience data in the experience replay buffer to train the initial agent. Network parameters are continuously updated iteratively until the agent's reward value converges or reaches a preset number of iterations, ultimately outputting a pre-trained scheduling policy with stable scheduling capabilities. The core advantage of the proximal policy optimization algorithm lies in its "pruning" mechanism for policy updates, limiting the magnitude of each update and preventing training instability due to policy mutations. Its core objective function expression is:
[0171]
[0172] In the formula, The objective function representing the policy update, The parameters representing the current policy network, These represent the parameters of the policy network before the update. Representing the old strategy Calculate the expectation over the generated empirical distribution. The ratio of the probability of an action under the new and old strategies (i.e.) ,in For the current policy in state Select action The probability, (The probability of choosing the same action under the same conditions in the old strategy). Representative advantage function (used to measure action) The advantage over average movement, namely ,in As a discount factor, For the critics network in the state (value estimate below) The clipping factor (used to limit) The range of values should be determined to avoid excessively large or small values that could lead to excessively large policy update fluctuations.
[0173] The batch training process follows a cyclical flow of "empirical sampling - parameter update - convergence judgment". First, a batch of empirical tuples is randomly sampled from the empirical replay buffer (batch size is determined based on computational resource configuration). For each empirical tuple, the dominance function is calculated. The sampled state, action, and advantage functions are then input into the actor network, combined with the objective function. The policy gradient is calculated, and the actor network parameters are updated using the gradient descent algorithm. Simultaneously, the state and objective value function are input into the critic network, and the critic network parameters are updated by minimizing the value estimation error. After each parameter update, the current policy network parameters are assigned to the old policy parameters. This is in preparation for the next update.
[0174] During training, convergence is determined by analyzing the trend of reward value changes. The average reward value of the agent in multiple consecutive rounds of interaction in the simulation environment is periodically calculated. If the fluctuation range of this average reward value over multiple iterations is less than a preset threshold, it indicates that the agent has found a stable optimal policy, and training has reached convergence. If the average reward value has not converged but the preset maximum number of iterations has been reached, the training process must be terminated to avoid overfitting due to overtraining. After training is terminated, the final policy network parameters (or action decision rules generated based on these parameters) are saved as a pre-trained scheduling policy. This policy can receive real-time system status during actual operation and directly output scheduling instructions that meet safety and efficiency goals.
[0175] The aforementioned AI-based collaborative scheduling method for cascade pumping stations achieves dynamic prediction of pressure wave propagation by constructing a lightweight hydraulic transient model. Based on the model's prediction results and real-time system status, a reinforcement learning agent is built. Using historical operating data, a scheduling strategy capable of proactively identifying and responding to pressure oscillations is trained. Finally, during the online control phase, the real-time status is input into the model for pressure oscillation prediction, and compensatory scheduling instructions are generated based on the trained strategy. This achieves a fundamental shift from passively avoiding hydraulic transients to actively managing them. As a result, while ensuring the safe and stable operation of the system, control accuracy and equipment lifespan are significantly improved, operating energy consumption and maintenance costs are effectively reduced, and the system's adaptability to complex operating conditions is enhanced.
[0176] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0177] Based on the same inventive concept, this application also provides an apparatus for implementing the aforementioned AI-based cascade pumping station collaborative scheduling method. The solution provided by this apparatus is similar to the implementation described in the above method; therefore, the specific limitations of one or more AI-based cascade pumping station collaborative scheduling apparatus embodiments provided below can be found in the above-described limitations of the AI-based cascade pumping station collaborative scheduling method, and will not be repeated here.
[0178] In one exemplary embodiment, such as Figure 3 As shown, an artificial intelligence-based cascade pumping station collaborative scheduling device 30 is provided to implement the methods in the above-described method embodiments. The device includes:
[0179] The hydraulic dynamic modeling module 31 is used to construct a lightweight one-dimensional nonlinear hydraulic transient model based on the pipeline geometric parameters, pump station equipment parameters and historical SCADA data of the cascade pump station system, and obtain the hydraulic transient model.
[0180] The strategy space construction module 32 is used to predict the output based on the real-time state data of the cascade pumping station system and the hydraulic transient model. By constructing the state space, action space and reward function for reinforcement learning, a reinforcement learning agent is obtained.
[0181] The agent training and optimization module 33 is used to train the reinforcement learning agent by using historical system operation data of the cascade pumping station system, simulating state transitions and reward calculations in a simulation environment, and outputting a pre-trained scheduling strategy.
[0182] The scheduling instruction execution module 34 is used to input the real-time system status of the cascade pumping station system into the hydraulic transient model to predict pressure oscillations; output scheduling instructions to cope with pressure oscillations based on the scheduling strategy; and execute pumping station control according to the scheduling instructions.
[0183] Embodiments of this application also provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the aforementioned method embodiments.
[0184] Embodiments of this application also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-described method embodiments.
[0185] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0186] The above-described embodiments are merely illustrative of several implementation methods of the embodiments of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the embodiments of this application, and these modifications and improvements all fall within the protection scope of the embodiments of this application.
Claims
1. An artificial intelligence-based step-by-step pump station collaborative scheduling method, characterized in that, The method comprises: S1, based on the pipeline geometric parameters, pump station equipment parameters and historical SCADA data of the cascade pump station system, constructing a lightweight one-dimensional nonlinear hydraulic transient model to obtain a hydraulic transient model; S2, based on the real-time state data of the cascade pump station system and the prediction output of the hydraulic transient model, constructing a state space, an action space and a reward function for reinforcement learning to obtain a reinforcement learning agent; S3, using the historical system operation data of the cascade pump station system, simulating state transition and reward calculation through a simulation environment, training the reinforcement learning agent, and outputting a pre-trained scheduling strategy; S4, inputting the real-time system state of the cascade pump station system into the hydraulic transient model to predict pressure oscillation, outputting a scheduling instruction for coping with the pressure oscillation based on the scheduling strategy, and executing pump station control according to the scheduling instruction.
2. The method of claim 1, wherein, The reinforcement learning agent obtained based on the real-time state data of the cascade pump station system and the prediction output of the hydraulic transient model by constructing a state space, an action space and a reward function for reinforcement learning comprises: S11, performing cleaning and timestamp alignment processing on the flow, pressure, pump speed and valve opening data of each pump station in the real-time state data to obtain a cleaned system state data set; inputting the cleaned system state data set into the hydraulic transient model, extracting pressure oscillation amplitude and dominant frequency characteristics by calculating the prediction value of pressure wave propagation, and obtaining a transient pressure feature vector; S12, combining the macro variables in the cleaned system state data set and the transient pressure feature vector to define the state space of the reinforcement learning agent, and obtaining a state vector; wherein the state vector comprises the flow, pressure, pump speed, valve opening, pressure oscillation amplitude, dominant frequency and timestamp of each pump station; S13, defining a continuous action space based on the physical constraints of pump speed and valve opening to obtain an action vector; wherein the action vector comprises the speed increment and valve opening increment of each pump station; S14, obtaining a total reward value by calculating the weighted sum of the energy efficiency reward item, the water quantity reward item and the pressure oscillation penalty item; S15, constructing the reinforcement learning agent according to the state vector, the action vector and the total reward value.
3. The method of claim 2, wherein, The total reward value obtained by calculating the weighted sum of the energy efficiency reward item, the water quantity reward item and the pressure oscillation penalty item comprises: S21, calculating the power consumption of each pump station based on the flow, pressure and efficiency parameters of the pump station by a power consumption calculation formula to obtain the pump station power consumption value of each pump station; taking the negative sum of the pump station power consumption values of each pump station to obtain the energy efficiency reward item; the power consumption calculation formula is: wherein, represents the power consumption value of the pump station i, represents the water density, represents the gravitational acceleration, represents the flow value of the pump station i, represents the pressure head of the pump station i, represents the efficiency of the pump station i; S22, calculating the absolute deviation of actual flow from planned flow based on the total flow and planned flow of the cascade pump station system; taking the negative value of the absolute deviation to obtain the water quantity reward item; the calculation formula of the water quantity reward item is: wherein, represents the water quantity reward term, represents the number of pump stations, represents the flow rate of pump station i, represents the planned flow rate; S23, calculating the pressure oscillation penalty item by linear combination based on the pressure oscillation amplitude and dominant frequency in the transient pressure feature vector; the calculation formula of the pressure oscillation penalty item is: wherein, denotes a pressure oscillation penalty term, and denotes a scaling factor; denotes a pressure oscillation amplitude; denotes a dominant frequency; S24, the total reward value is generated by weighted sum of the energy efficiency reward item, the water quantity reward item and the pressure oscillation penalty item; the calculation formula of the total reward value is: wherein, represents the total reward value, represents the energy efficiency reward term, represents a weight coefficient.
4. The method of claim 1, wherein, The pipeline geometric parameters, pump station equipment parameters and historical SCADA data of the cascade pump station system are used to construct a lightweight one-dimensional nonlinear hydraulic transient model to obtain a hydraulic transient model, including: S31, based on the pipeline geometric parameters, pump station equipment parameters and historical SCADA data of the cascade pump station system, the continuity equation and momentum equation are derived by combining the one-dimensional unsteady flow control equation to obtain a basic physical model; the continuity equation is: The momentum equation is: wherein, denotes the pipe cross-sectional area, denotes the flow rate, denotes the time, denotes the distance along the pipe, denotes the gravitational acceleration, denotes the pressure head, denotes the Darcy-Weisbach friction factor, denotes the pipe diameter; S32, the basic physical model is linearized and the partial differential equation is converted into an ordinary differential equation by using the method of characteristics to obtain a linearized model; S33, the linearized model is discretized by dividing the pipeline into multiple grid points and using an implicit difference format to solve the pressure and water head to obtain a discretized model; S34, based on the pressure and flow readings in the historical SCADA data, the discretized model is calibrated by the least square method to obtain the calibrated hydraulic transient model.
5. The method according to any one of claims 1 to 4, characterized in that, The historical system operation data of the cascade pump station system is used to train the reinforcement learning agent by simulating state transition and reward calculation in a simulation environment, and output a pre-trained scheduling strategy, including: S41, based on the hydraulic transient model and pump station system dynamics, a simulation environment for simulating state transition and reward calculation is constructed; S42, based on the preset network parameter configuration, the reinforcement learning agent is initialized in network parameters and experience replay buffer to obtain an initialized agent; S43, an initial state is sampled from the historical system operation data; based on the initialized agent, the initial state is input into the simulation environment to obtain an output action; based on the output action, the simulation environment simulates state transition and reward calculation to obtain an updated system state and a corresponding reward value; S44, using a proximal policy optimization algorithm, the initialized agent is batch trained based on the updated system state and the corresponding reward value through experience replay; when the reward value of the initialized agent converges or reaches a preset maximum iteration number, the training process is terminated, and a trained scheduling strategy is output.
6. An artificial intelligence-based cascade pumping station collaborative scheduling device for implementing the method of any one of claims 1 to 5, characterized in that, The device comprises: A hydraulic dynamic modeling module is configured to construct a lightweight one-dimensional nonlinear hydraulic transient model based on pipeline geometric parameters, pump station equipment parameters and historical SCADA data of a cascade pump station system to obtain a hydraulic transient model; A policy space construction module is configured to obtain a reinforcement learning agent by constructing a state space, an action space and a reward function for reinforcement learning based on real-time state data of the cascade pump station system and predicted output of the hydraulic transient model; An agent training optimization module is configured to train the reinforcement learning agent by simulating state transition and reward calculation in a simulation environment using historical system operation data of the cascade pump station system, and output a pre-trained scheduling strategy. The scheduling instruction execution module is configured to input a real-time system state of the cascade pump station system into the hydraulic transient model to predict pressure oscillation, output a scheduling instruction for coping with the pressure oscillation based on the scheduling strategy, and execute pump station control according to the scheduling instruction. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The computer program, when executed by the processor, implements the method of any one of claims 1 to 5.
8. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the method of any one of claims 1 to 5.