Intelligent regulation and control method for direct connection type pipe network pressure-superposed water supply equipment

By constructing a water supply knowledge graph and a digital twin hydraulic model, and combining it with a working condition causal graph model for counterfactual reasoning and reinforcement learning, the control strategy of the direct-connection pipeline superimposed pressure water supply equipment is optimized. This solves the safety hazards of existing equipment under rapid start-up and shutdown and hydraulic transient risks, and achieves safe and adaptive optimized water supply control.

CN121680201AInactive Publication Date: 2026-03-17YUNNAN NANFANG INTELLIGENT EQUIP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511877994.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-17
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The control logic of existing direct-connection pipeline superimposed pressure water supply equipment is difficult to identify and avoid the risks of hydraulic transients such as negative pressure and water hammer caused by rapid start-up and shutdown, large frequency changes or rapid opening and closing of valves, which may lead to safety hazards such as overpressure in local pipe sections, pipe bursts or backflow pollution, and lacks adaptability and reliability.

Method used

A knowledge graph and digital twin hydraulic model for water supply are constructed. The attributes of entity nodes are updated through real-time monitoring data. The digital twin hydraulic model is established and simulation calculations are performed. Combined with the working condition causal graph model, counterfactual reasoning and reinforcement learning are carried out to construct safety constraints and policy priors, and optimize the control strategy to reduce safety risks and energy consumption.

Benefits of technology

It significantly reduces water supply security risks, improves operational safety and reliability, reduces energy consumption and mechanical shock, extends equipment life, and has adaptive optimization capabilities, enhancing human-machine collaborative decision-making capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121680201A_ABST
    Figure CN121680201A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of municipal water supply, in particular to an intelligent regulation and control method for direct connection type pipe network pressure-superposed water supply equipment, and the method comprises the steps: extracting the structure and operation feature information of the pressure-superposed water supply equipment and a pipe network connected with the pressure-superposed water supply equipment through a water supply knowledge graph, and building a digital twin hydraulic model; performing simulation calculation and evaluation on hydraulic response of the candidate regulation and control actions in a preset prediction time domain in each regulation and control period, and constructing a first safety constraint based on a negative pressure risk index and a water attack risk index; meanwhile, anti-fact reasoning is conducted on candidate regulation and control actions on a working condition causal graph model, a second safety constraint and strategy prior are constructed in combination with a predefined accident chain, and a safety action domain is limited from physical simulation and working condition causal double-layer constraint; regulation and control actions possibly causing safety incidents such as negative pressure, water hammer or pipe explosion can be recognized and eliminated in advance before control issuing, and compared with a traditional control mode only based on in-situ pressure regulation or an empirical threshold value, the water supply safety risk can be remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of municipal water supply technology, and in particular to an intelligent control method for a direct-connection pipeline superimposed pressure water supply equipment. Background Technology

[0002] With the continuous expansion of urban water supply networks and the increase in the number of high-rise buildings, municipal water supply pressure often struggles to simultaneously meet the water needs of users in low-rise areas, remote areas, and high-rise buildings under fluctuating operating conditions throughout the day. To avoid the need for large-capacity elevated water tanks or towers, and to reduce secondary pollution and civil engineering investment, direct-connection network booster pump systems are increasingly being widely adopted in this field. These systems use variable frequency pump sets to boost pressure on top of the municipal water supply network, directly delivering water to the building's internal water supply network.

[0003] However, the control logic of existing direct-connection pipeline booster pump water supply equipment is mostly based on single-point pressure feedback and simple logic control. The control strategy often relies on the instantaneous pressure values ​​of a small number of monitoring points to set a single target pressure or simple segmented settings. It is difficult to identify and avoid the risks of hydraulic transients such as negative pressure and water hammer caused by rapid start-up and shutdown, large frequency changes, or rapid opening and closing of valves. This can easily lead to safety hazards such as local pipe section overpressure, pipe bursts, or backflow pollution. In addition, the existing control parameters mostly rely on the experience of commissioning personnel for repeated tuning. There is a lack of a mechanism that can automatically correct with the adjustment of pipeline structure, changes in water demand, and equipment aging. It is difficult to adapt to the gradual changes in operating conditions during long-term operation, resulting in problems such as excessively high constant pressure settings or frequent start-up and shutdown, which increases energy consumption and mechanical shock to equipment.

[0004] Therefore, an intelligent control method for direct-connection pipeline superimposed pressure water supply equipment is proposed. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, this invention provides an intelligent control method for a direct-connection pipeline superimposed pressure water supply equipment.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: an intelligent control method for a direct-connection superimposed pressure water supply equipment, comprising the following steps: S1, constructing a water supply knowledge graph, introducing water supply scenario elements as entity nodes in the water supply knowledge graph, defining the association relationships and operational constraints between entity nodes as graph edge relationships, updating and instantiating the attributes of entity nodes based on real-time monitoring data during operation of the superimposed pressure water supply equipment and its connected pipe network, obtaining a water supply operating condition scenario subgraph representing the current water supply status; S2, based on the structure and operation reflected in the water supply knowledge graph... Feature information is used to establish a corresponding digital twin hydraulic model, and the digital twin hydraulic model is continuously calibrated using hydraulic state monitoring data. In each control cycle, the calibrated digital twin hydraulic model is called to simulate the hydraulic response of candidate control actions within a preset prediction time domain. According to preset safety control conditions, the set of actions that meet the safety control conditions is determined from the candidate control actions, constituting the first safety constraint based on the digital twin; S3, a causal graph model of operating conditions is constructed in the water supply knowledge graph, and the nodes representing the operating mode and water supply safety risk are associated through causal edges during the control process. Combining the subgraph of the water supply operating condition scenario, counterfactual reasoning is performed on the causal graph model of the operating condition for each candidate control action. When the reasoning result meets the preset risk triggering conditions, the candidate control action is adjusted. Based on the current water supply operating condition scenario, similar historical scenarios are retrieved from the water supply knowledge graph, and the corresponding policy prior information is used to set the relevant parameters of the reinforcement learning policy network, forming a second security constraint and policy prior based on the knowledge graph; S4. Within the safety action domain that satisfies the first and second security constraints, the operating status characteristics reflecting the operating status of the booster pump water supply equipment and its connected pipe network are used as... The state variables used in reinforcement learning are used as action variables to represent regulatory behavior. A reward function with comprehensive operational indicators as the optimization objective is constructed. Reinforcement learning training and policy updates are carried out in the digital twin hydraulic model. The parameters of the reinforcement learning policy network are adjusted according to the operational evaluation indicators between the digital twin prediction results and the actual operational performance. When the operational evaluation indicators exceed the preset range, the policy protection mechanism is triggered, and the control actions after the above safety assessment are sent to the booster pump. The operational feedback information generated during the execution of control is written back to the water supply knowledge graph and the digital twin hydraulic model.

[0007] As a preferred technical solution of the present invention, updating the attributes of the entity nodes refers to refreshing the attribute values ​​of the entity nodes corresponding to the monitored objects in the water supply knowledge graph at each preset time interval during operation, based on real-time monitoring data; instantiating the attributes of the entity nodes refers to selecting, after completing the above attribute update, entity nodes corresponding to the water supply scenario elements corresponding to the current operating state and their graph edge relationships from the water supply knowledge graph according to the time information corresponding to the current time or the current control cycle, and combining the updated attribute values, the selected entity nodes and the graph edge relationships together.

[0008] As a preferred technical solution of the present invention, the digital twin hydraulic model is modeled based on pipeline nodes and hydraulic components. During continuous calibration, in each control cycle or each preset time window, the model parameters are updated several times using the collected hydraulic state monitoring data. During each update, the parameter vector is corrected according to the sensitive direction of the loss function on the model parameters, so that the loss function value defined based on the pressure deviation of each monitoring location gradually decreases. When the change of the loss function between two adjacent updates is lower than the preset convergence threshold, or the number of updates reaches the preset upper limit, the parameter calibration of this round ends, and the parameter vector obtained at this time is used as the parameter of the current version of the digital twin hydraulic model. The safety control conditions include simulating the hydraulic response of each candidate control action in the digital twin hydraulic model in the preset prediction time domain, calculating the corresponding negative pressure risk index and water hammer risk index, and making step-by-step judgments in the prediction time domain. When the negative pressure risk index or water hammer risk index exceeds the preset safety threshold at any prediction time, the corresponding candidate control action is marked as not meeting the safety control conditions and is eliminated.

[0009] As a preferred technical solution of the present invention, the counterfactual reasoning refers to classifying candidate control actions into corresponding operating modes while keeping other conditions unchanged in the current water supply operating condition scenario subgraph, and selecting the corresponding operating mode node as the starting node for counterfactual reasoning in the operating condition causal graph model. Following the direction of the causal edges in the operating condition causal graph model, the possible changes in operating conditions and abnormal operating situations caused by the candidate control actions are gradually deduced within a limited number of reasoning steps, and it is checked whether water supply safety risk event nodes will be triggered during the deduction process. The predefined accident chain is a set of causal paths predetermined according to operating procedures and historical accident cases, which connect operating mode nodes sequentially to water supply safety risk event nodes through one or more operating condition and abnormal event relationships. When the counterfactual reasoning result indicates that, under the assumption of executing a certain candidate control action, there exists a path along the above predefined accident chain from the current operating mode... When a node reaches the causal path of a target water supply safety risk event node, it is determined that the candidate control action will evolve along a predefined accident chain, thereby marking the candidate control action as having an unacceptable safety risk, which is used to reject or modify it in the safety constraints. The acquisition of policy prior information includes graph embedding calculation on the subgraph of water supply operating conditions to obtain feature vectors that represent the characteristics of the current operating condition. The feature vectors are then compared with the historical operating condition vectors pre-stored in the water supply knowledge graph. The similarity is measured by cosine similarity, and the ratio of the inner product of the current operating condition vector and the product of their magnitudes is used as the similarity value. Several historical operating conditions are selected from large to small based on the similarity value to form a set of similar operating conditions. The historical control strategy parameters, reward weights and exploration strategies corresponding to the set of similar operating conditions are extracted as the policy prior information of the reinforcement learning policy network.

[0010] As a preferred embodiment of the present invention, the operational status characteristics include characteristics reflecting the water supply pressure distribution, characteristics reflecting the water load level, characteristics reflecting the consistency between the digital twin hydraulic model and the on-site operation, and characteristics reflecting the equipment operating conditions. The method for obtaining the characteristics reflecting the consistency between the digital twin hydraulic model and the on-site operation is as follows: at each moment, several pressure monitoring locations are selected for comparison, and the measured pressure at each monitoring location and the simulated pressure calculated by the digital twin hydraulic model are obtained respectively. The average difference between the simulated pressure and the measured pressure at each monitoring location is calculated, and the average difference is used as the average model pressure deviation at the corresponding moment, characterizing the consistency between the digital twin hydraulic model and the on-site operation. Consistency among on-site operational data; comprehensive operational indicators include multiple evaluation indicators corresponding to pressure service quality, water supply energy efficiency, water supply safety risks, and equipment lifespan; the immediate reward value of reinforcement learning at each evaluation moment is calculated by weighting the evaluation items corresponding to water supply pressure deviation and pressure fluctuation, the evaluation items corresponding to unit water supply energy consumption, the evaluation items corresponding to negative pressure risk and water hammer risk, and the lifespan evaluation items corresponding to equipment start-up and shutdown times and operating condition change amplitude according to each weight coefficient to obtain the reward value at the evaluation moment. Each weight coefficient reflects the relative importance of pressure service quality, water supply energy efficiency, water supply safety risks, and equipment lifespan in the comprehensive operational indicators; Reinforcement learning training and strategy updating refers to using the operational state characteristics reflecting the operation of the booster pump system and its connected pipe network as input state variables within the safe action domain that satisfies the first and second safety constraints. The control variables, consisting of frequency setpoints, pump start-up or shutdown commands, and adjustments to the opening of key pressure regulating components, are used as output action variables. A digital twin hydraulic model is used to simulate the pressure and flow changes at each monitoring location after executing each candidate control action in the prediction time domain. The corresponding reward value is calculated based on the reward function to form sample data between the state variables, control variables, reward values, and the operational state characteristics at the next moment. Based on this sample data... The parameters of the reinforcement learning policy network are iteratively adjusted using existing reinforcement learning methods to improve the overall operational performance under the desired meaning, thereby completing the reinforcement learning training. As training progresses, the updated policy network can output a new combination of control variables under the constraints of given operational state characteristics and safety action domain. The control variable combination obtained from the latest policy network is used as the updated control strategy, and after passing a safety assessment, it is applied as the current version of the control strategy to the booster pump water supply equipment. The operational assessment indicators include at least one indicator for quantifying the deviation between the digital twin hydraulic model and the actual operation, and at least one indicator for quantifying the level of safety risk.During reinforcement learning training, the learning rate and policy update step size of the reinforcement learning policy network are adjusted according to the model bias index in the performance evaluation metrics. When the model bias index is low, the learning rate and policy update step size are kept near the pre-set baseline learning rate and policy update step size. When the model bias index increases, the learning rate and policy update step size are increased proportionally within a preset allowable range.

[0011] Compared with the prior art, the beneficial effects that this invention can achieve are: 1. This invention extracts structural and operational characteristic information of booster pumping equipment and its connected pipe networks from a water supply knowledge graph, establishes a digital twin hydraulic model, and simulates and evaluates the hydraulic response of candidate control actions within a preset prediction time domain in each control cycle, constructing a first safety constraint based on negative pressure risk indicators and water hammer risk indicators. Simultaneously, counterfactual reasoning is performed on candidate control actions on the operating condition causal graph model, and a second safety constraint and strategy prior are constructed by combining a predefined accident chain. By limiting the safety action domain from both physical simulation and operating condition causal double-layer constraints, control actions that may cause safety events such as negative pressure, water hammer, or pipe bursts can be identified and eliminated in advance before control is issued. Compared with traditional control methods that are only based on local pressure regulation or empirical thresholds, this invention can significantly reduce water supply safety risks and improve the safety and reliability of booster pumping operation.

[0012] 2. Within the safe operating domain, this invention uses the operating state characteristics reflecting the operation of the booster water supply system as reinforcement learning state variables, and uses the frequency setpoint of the variable frequency pump set, the pump set start-up or stop command, and the opening adjustment of key pressure regulating components as control variables. It constructs a reward function with pressure service quality, water supply energy efficiency, water supply safety risk, and equipment lifespan consumption as comprehensive operating indicators. Reinforcement learning training and strategy updates are performed in a digital twin hydraulic model, enabling the control strategy to automatically balance factors such as user-end pressure deviation and fluctuation, unit water supply energy consumption, and equipment start-up and shutdown frequency while meeting safety constraints. Compared to fixed setpoints or simple segmented control strategies, this can reduce ineffective high-pressure operation and frequent start-ups and shutdowns, reduce energy consumption and mechanical shock, and extend the service life of key components such as pump sets and valves.

[0013] 3. This invention continuously iterates and calibrates the digital twin hydraulic model by introducing hydraulic state monitoring data to reduce the deviation between simulated and measured pressures. It also adaptively adjusts the learning rate and policy update step size of reinforcement learning using model deviation indices. When deviation or risk indicators exceed limits, a policy rollback and freeze update mechanism is triggered. Simultaneously, monitoring data, abnormal events, and their processing results during operation are written back to the water supply knowledge graph and the digital twin hydraulic model to update the water supply operating condition scenario subgraph and model parameters. As a result, the control strategy can continuously learn and adaptively optimize with adjustments to the pipeline structure, changes in water load, and aging of equipment performance. Compared with traditional control parameters that remain unchanged for a long time after being tuned, this invention has better long-term robustness and adaptability.

[0014] 4. This invention uses a water supply knowledge graph to uniformly represent the structural topology, operational constraints, water usage conditions, and abnormal events of booster pumping equipment and its connected pipe networks. It also constructs a cause-effect graph model of operating conditions and an accident chain, reusing historical operating condition scenarios and control strategies as policy priors in reinforcement learning. On the one hand, this makes control decisions not only rely on real-time data but also make full use of existing operation and maintenance experience and accident knowledge. On the other hand, the graph structure can intuitively show the causal path between a certain control action and potential risk events such as negative pressure and water hammer. Compared with the completely black-box pure data-driven control method, it has better interpretability, which is conducive to operation and maintenance personnel to understand and verify control strategies and improve human-machine collaborative decision-making capabilities. Attached Figure Description

[0015] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0016] To make the technical means, creative features, objectives, and effects of this invention easier to understand, the invention is further described below with reference to specific embodiments. However, the following embodiments are merely preferred embodiments of this invention and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments described herein without creative effort are all within the protection scope of this invention.

[0017] Example: Figure 1 As shown, an intelligent control method for a direct-connection pipeline superimposed pressure water supply system includes the following steps: S1. Construct a water supply knowledge graph, introduce water supply scenario elements as entity nodes in the water supply knowledge graph, and define the association and operation constraints between entity nodes as graph edge relationships. During the operation of the superimposed water supply equipment and its connected pipeline network, update and instantiate the attributes of entity nodes based on real-time monitoring data during the operation process to obtain a water supply operating condition scenario subgraph representing the current water supply status.

[0018] It should be noted that the elements of the water supply scenario include the pipe sections, pipe network nodes, and actuators, pressure monitoring points, and flow monitoring points arranged on the pipes in the municipal water supply network; the inlet interface, outlet interface, bypass interface, pump set, and pressure stabilizing unit in the booster pump system; the vertical main pipes, floor branch pipes, and terminal water points in the building's internal water supply network; different types of water users; and the operating condition elements and event elements used to describe water usage conditions and abnormal operation events.

[0019] The water supply knowledge graph models the elements of the above-mentioned water supply scenario by representing them as different types of entity nodes. The water supply knowledge graph can be implemented using conventional modeling methods of existing knowledge graph technology. Those skilled in the art can select specific data structures and storage methods based on existing knowledge graph construction methods, and this invention does not limit them in this regard.

[0020] In practice, static information of water supply scenario elements, such as pipe diameter, length, elevation parameters, and rated operating parameters of pump sets and pressure stabilizing units, can be imported from existing pipeline design data, as-built drawings, geographic information systems (GIS), or asset management systems. Dynamic information of water supply scenario elements is obtained through on-site monitoring instruments and automation systems and written into the attributes of the corresponding entity nodes.

[0021] In addition, the relationships between entity nodes include topological relationships to represent the connectivity of nodes and pipe segments in the pipeline network, hydraulic relationships to represent the interaction between pressure and flow, and relationships between operating conditions and abnormal events to represent the connection between different water use conditions and operational anomalies; operating constraints include operating constraints such as pump start-up and shutdown, frequency regulation, and valve opening and closing restrictions.

[0022] Topological relationships can be generated based on pipeline design drawings and GIS pipeline data. Hydraulic relationships can be configured by combining hydraulic calculation results or historical operation analysis results. Operating conditions and abnormal event relationships, as well as operating constraints, can be set according to operating procedures, scheduling strategies, and historical event records, or dynamically modified during operation.

[0023] In addition, real-time monitoring data includes pressure data from pressure monitoring points, flow data from flow monitoring points, operating frequency and start / stop status data of pump sets, opening data of actuator valves, and alarm and alarm information recorded by the control system.

[0024] Real-time monitoring data is collected by pressure sensors, flow meters, and other field instruments located at appropriate positions, and transmitted to the upper-level monitoring system or edge control device via fieldbus or industrial Ethernet through local programmable controllers (PLCs), remote terminal units (RTUs), or monitoring and data acquisition systems (SCADA) to the control method of this invention for access and processing. Pump unit operating frequency and start / stop status data, as well as valve opening data, are collected by frequency converters, PLCs, or remote terminals connected to the pump unit and valves. Alarm and alarm information is automatically generated and written into the corresponding data records by the control system when it detects that operating parameters exceed limits or equipment failures.

[0025] Finally, updating the attributes of entity nodes refers to refreshing the attribute values ​​of entity nodes corresponding to the monitored objects in the water supply knowledge graph at each preset time interval during operation, based on real-time monitoring data. For example, the pressure attribute of the entity node corresponding to the pressure monitoring point is updated to the currently collected pressure value; the flow attribute of the entity node corresponding to the flow monitoring point is updated to the currently collected flow value; the operating frequency attribute and start / stop status attribute of the entity node corresponding to the pump group are updated to the current operating status fed back by the frequency converter; and the opening attribute of the entity node corresponding to the execution valve is updated to the current valve position feedback opening percentage. Through the above attribute updates, the attribute values ​​stored by each entity node in the water supply knowledge graph can reflect the operating status of the booster pump and its connected pipe network at the current moment.

[0026] Instantiating the attributes of entity nodes refers to, after completing the above attribute updates, selecting entity nodes and their graph edge relationships corresponding to the water supply scenario elements corresponding to the current operating state from the water supply knowledge graph based on the time information corresponding to the current time or the current control cycle. For example, topology filtering based on key boundaries and control locations, correlation filtering based on water demand and monitoring layout, and operating condition correlation filtering combined with operating state thresholds. The updated attribute values, selected entity nodes, and graph edge relationships are then combined to form a water supply operating condition scenario subgraph with the current attribute values.

[0027] S2. Based on the structural and operational characteristics of the superimposed pressure water supply equipment and its connected pipe network reflected in the water supply knowledge graph, a corresponding digital twin hydraulic model is established. The digital twin hydraulic model is continuously calibrated using hydraulic state monitoring data. In each control cycle, the calibrated digital twin hydraulic model is called to simulate and evaluate the hydraulic response of candidate control actions in the preset prediction time domain. According to the preset safety control conditions, the set of actions that meet the safety control conditions is determined from the candidate control actions, so that the set of actions constitutes the first safety constraint based on the digital twin.

[0028] It should be noted that the structural and operational characteristic information includes the topology of the booster pump and its connected network nodes and pipe sections recorded in the water supply knowledge graph, the geometric parameters and elevation parameters of each pipe section, the pump configuration and pressure stabilizing unit connection method of the booster pump, as well as the rated operating parameters of the pump and pressure stabilizing unit pre-entered as attributes of the water supply scenario. The structural and operational characteristic information is imported from existing engineering data during the initialization phase of the water supply knowledge graph.

[0029] The digital twin hydraulic model is modeled based on pipeline nodes and hydraulic components. During continuous calibration, the model parameters are updated several times in each control cycle or each preset time window using the collected hydraulic state monitoring data. Each time the parameter vector is updated, it is corrected according to the sensitive direction of the loss function on the model parameters, so that the value of the loss function defined based on the pressure deviation of each monitoring location gradually decreases. When the change of the loss function between two adjacent updates is lower than the preset convergence threshold, or the number of updates reaches the preset upper limit, the parameter calibration ends and the parameter vector obtained at this time is used as the parameters of the current version of the digital twin hydraulic model.

[0030] Specifically, hydraulic elements are used to represent hydraulic components such as pipes, pump sets, valves, and pressure stabilizing units that connect network nodes; parameter updates satisfy the following relationship: .

[0031] in, For the first The parameter vector at the next iteration This is the step size coefficient. The loss function is defined based on the pressure deviation at each monitoring location. The gradient of the loss function with respect to the parameters is defined by a preset convergence threshold, which limits the allowable variation of the loss function between two adjacent iterations. A preset upper limit limits the maximum number of iterations allowed in a single parameter calibration process. The specific values ​​of both can be set by those skilled in the art through debugging, based on model accuracy requirements, convergence speed requirements, and available computing resources. The step size coefficient is also defined. The parameters can be selected or adjusted in segments according to the model's convergence speed and stability requirements, so as to balance the convergence speed and numerical stability.

[0032] loss function The formula is as follows: .

[0033] in, The number of pressure monitoring locations involved in the calibration. This represents the number of sample times collected within the calibration time window. For in the parameter vector When taking the current value, the first value calculated by the digital twin hydraulic model is... The pressure monitoring location is at the first Simulated stress at a given moment, The measured pressure at the same pressure monitoring location at the same time. For the first The weighting coefficients corresponding to each pressure monitoring location are used to reflect the importance of different monitoring locations in the overall calibration. This is a normalization factor used to scale the loss function to an order of magnitude independent of the number of monitoring points and the number of samples.

[0034] In addition, the hydraulic condition monitoring data includes online pressure data at the connection points of the booster pump and its connecting pipe network, online pressure data at the outlet of the booster pump, online pressure data at key nodes inside the building, and online flow data corresponding to the above pressure monitoring locations. The hydraulic condition monitoring data is collected by pressure monitoring points and flow monitoring points located at the corresponding locations. The measurement signals from each pressure monitoring point and flow monitoring point are collected by the local control unit or monitoring system and then uploaded to form the hydraulic condition monitoring data.

[0035] In addition, the control cycle refers to the time interval used to execute a control decision. It can be a fixed control cycle set in the monitoring system for the booster pump and its connected pipe network, or a time window coordinated with the sampling cycle.

[0036] Simulation calculations and evaluations of the hydraulic response of candidate control actions within a preset prediction time domain are performed. Within the current control cycle, the current operating state of the booster pump and its connected pipe network is determined based on hydraulic state monitoring data. This current operating state is used as the initial state of the digital twin hydraulic model, and the control variables corresponding to the candidate control actions are used as the control inputs of the digital twin hydraulic model. Starting from the current moment, the simulated pressure and simulated flow rate at each monitoring location are calculated step-by-step along the preset prediction time domain at preset time intervals. Based on this, the pressure changes at each monitoring location within adjacent time intervals are obtained to reflect the hydraulic transient response. Here, a candidate control action refers to a... Within the control cycle, under the premise of meeting the operational constraints of the booster pump system and its connected pipe network, one or more alternative combinations of control variables such as pump frequency setting, number of pumps in operation, and adjustment of key pressure regulating valve opening are given. Each combination of control variables constitutes a candidate control action. The prediction time domain refers to a simulated time range that is extrapolated backward from the moment the candidate control action is executed in the current control cycle. It is used to evaluate the impact of the candidate control action on the pressure and flow at each monitoring location in the digital twin hydraulic model. Its length can be set by those skilled in the art according to the hydraulic response characteristics and safety assessment requirements of the booster pump system and its connected pipe network.

[0037] Finally, safety control conditions include negative pressure risk indicators and water hammer risk indicators calculated based on simulation results from a digital twin hydraulic model. The negative pressure risk indicators... Calculate according to the following formula: .

[0038] in, For a moment Negative pressure risk indicators For the first Monitoring location at time The measured or simulated pressure, For the first The minimum allowable pressure threshold for the monitoring location.

[0039] Water hammer risk indicators Calculate according to the following formula: .

[0040] in, For a moment Water strike risk indicators For the first Monitor the pressure value at the previous moment. To monitor the maximum allowable pressure variation limit at the location, when and When all parameters do not exceed the preset safety threshold, the corresponding candidate control action is determined to meet the safety control conditions; the preset safety threshold refers to the threshold used to limit the negative pressure risk indicators. and water hammer risk indicators The upper limit of the allowable range is used to determine whether the candidate control action meets the safety control conditions.

[0041] For each candidate control action, a hydraulic response simulation in the preset prediction time domain is performed in the digital twin hydraulic model, and the corresponding negative pressure risk index is calculated. and water hammer risk indicators Furthermore, the system makes step-by-step judgments in the prediction time domain. When the negative pressure risk index or water hammer risk index exceeds the preset safety threshold at any prediction time, the corresponding candidate control action is marked as not meeting the safety control conditions and is eliminated. The remaining candidate control actions constitute the first safety constraint based on digital twin.

[0042] S3. Construct a condition-cause graph model in the water supply knowledge graph to describe the relationship between operating conditions and risks. Connect the nodes representing operating modes and water supply safety risks through causal edges. During the control process, combine the water supply operating condition scenario subgraph and perform counterfactual reasoning on the condition-cause graph model for each candidate control action to determine whether it will evolve along a predefined accident chain if the candidate control action is assumed to be executed. When the reasoning result meets the preset risk triggering conditions, adjust the candidate control action. Based on the current water supply operating condition scenario, retrieve similar historical scenarios from the water supply knowledge graph and use the corresponding policy prior information to set the relevant parameters of the reinforcement learning policy network, forming a second safety constraint and policy prior based on the knowledge graph.

[0043] It should be noted that the operating mode nodes in the operating condition cause-effect graph model include nodes representing high flow rate operating conditions, rapid valve closure operating conditions, significant frequency reduction operating conditions, significant frequency increase operating conditions, and high pressure differential operating conditions; the water supply safety risk nodes include nodes representing negative pressure risk, water hammer risk, and pipe burst risk; and the causal edges are used to represent the relationships that trigger the corresponding water supply safety risks under different operating modes.

[0044] Furthermore, counterfactual reasoning is performed on the operating condition causal graph model for each candidate control action. This involves classifying the candidate control action into the corresponding operating mode while keeping other conditions unchanged in the current water supply operating condition scenario subgraph. The operating mode node corresponding to the operating mode is selected as the starting node for counterfactual reasoning in the operating condition causal graph model. According to the direction of the causal edges in the operating condition causal graph model, the possible changes in operating conditions and abnormal operating situations caused by the candidate control action are deduced step by step within a limited number of reasoning steps. It is also checked whether the water supply safety risk event node will be triggered during the deduction process. The predefined accident chain is a set of causal paths that are predetermined based on the operating procedures and historical accident cases, and are sequentially connected from the operating mode node to the water supply safety risk event node through one or more operating conditions and abnormal event relationships. When the counterfactual reasoning results show that, under the assumption that a certain candidate control action is executed, there exists a causal path from the current operating mode node to the target water supply safety risk event node along the above predefined accident chain, it is determined that the candidate control action will evolve along the predefined accident chain, thereby marking the candidate control action as having an unacceptable safety risk, which is used to reject or modify it in the safety constraints.

[0045] In addition, the risk triggering condition refers to the path search in the working condition causal graph model, starting from the current water supply working condition scenario subgraph and the node corresponding to the candidate control action, and searching along the causal edge within a preset number of inference steps. When there is a path from the operating mode node to the water supply safety risk node through one or more causal edges and the corresponding triggering condition meets the preset threshold set by the technicians, the candidate control action is determined to meet the risk triggering condition, and the candidate control action is modified or replaced with a predefined safety control action template.

[0046] Furthermore, the acquisition of policy prior information includes graph embedding calculations on the subgraph of water supply operating conditions to obtain feature vectors representing the characteristics of the current operating condition. The feature vectors are then compared with the historical operating condition vectors pre-stored in the water supply knowledge graph for similarity calculation. The similarity can be measured using cosine similarity, i.e., the ratio of the inner product of the current operating condition vector and the historical operating condition vector to the product of their magnitudes is used as the similarity value. Several historical operating conditions are selected from large to small based on the similarity values ​​to form a set of similar operating conditions. The historical control strategy parameters, reward weights, and exploration strategies corresponding to the set of similar operating conditions are extracted and used as policy prior information for the reinforcement learning policy network to set the initial parameters and related hyperparameters of the policy network.

[0047] Specifically, similarity is calculated using the following formula: .

[0048] in, The similarity value. For vector dot product, and Let be the vector magnitude.

[0049] S4. Within the safe action domain that satisfies the first and second safety constraints, the operating state characteristics reflecting the operation status of the booster pumping equipment and its connected pipe network are used as the state variables for reinforcement learning, and the control variables used to characterize the control behavior are used as the action variables for reinforcement learning. A reward function with the comprehensive operation index as the optimization objective is constructed. Reinforcement learning training and strategy updates are carried out in the digital twin hydraulic model. Based on the operation evaluation index between the digital twin prediction results and the actual operation performance, the parameters of the reinforcement learning strategy network are adaptively adjusted. When the operation evaluation index exceeds the preset range, the strategy protection mechanism is triggered, and the control actions after the above safety evaluation are sent to the booster pumping equipment to realize intelligent control of the booster pumping operation status.

[0050] It should be noted that the operational status characteristics include features reflecting the distribution of water supply pressure, features reflecting the water load level, features reflecting the consistency between the digital twin hydraulic model and the on-site operation, and features reflecting the operating conditions of the equipment. The method for obtaining the features reflecting the consistency between the digital twin hydraulic model and the on-site operation is as follows: at each moment, several pressure monitoring locations are selected for comparison, and the measured pressure at each monitoring location and the simulated pressure calculated by the digital twin hydraulic model are obtained respectively. The average difference between the simulated pressure and the measured pressure at each monitoring location is calculated, and the average difference is used as the average model pressure deviation at the corresponding moment to characterize the consistency between the digital twin hydraulic model and the on-site operating data. The number of monitoring locations participating in the comparison can be determined according to the layout of the monitoring points.

[0051] Specifically, the characteristics reflecting the consistency between the digital twin hydraulic model and the field operation include the model deviation index calculated according to the following formula. : .

[0052] in, For a moment The average model pressure deviation, The number of monitoring locations included in the comparison. For the first Monitoring location at time The measured pressure, For digital twin hydraulic models at time Simulated pressure.

[0053] In addition, the control variables include frequency setpoints and input / output commands to represent the target operating status of each variable frequency pump group, as well as opening adjustment amounts to represent the operating status of key pressure regulating components. The control variables can only be output as action quantities for reinforcement learning after being screened by the first and second safety constraints.

[0054] In addition, the comprehensive operation indicators include multiple evaluation indicators corresponding to pressure service quality, water supply energy efficiency, water supply safety risk, and equipment lifespan. The calculation of the immediate reward value of reinforcement learning at each evaluation moment involves weighting the evaluation items corresponding to water supply pressure deviation and pressure fluctuation, the evaluation items corresponding to unit water supply energy consumption, the evaluation items corresponding to negative pressure risk and water hammer risk, and the lifespan evaluation items corresponding to equipment start-up and shutdown frequency and operating condition change amplitude according to pre-set weight coefficients to obtain the reward value at the evaluation moment. Among them, each weight coefficient is used to reflect the relative importance of pressure service quality, water supply energy efficiency, water supply safety risk, and equipment lifespan in the comprehensive operation indicators.

[0055] Specifically, the reward function of reinforcement learning The calculation includes the following forms: .

[0056] in, As an evaluation item corresponding to water supply pressure deviation and fluctuation, for several key user pressure monitoring locations, the actual pressure of each monitoring location is compared with its corresponding target reference pressure in each evaluation time interval. When the actual pressure is close to the reference pressure and within the allowable deviation range, a higher evaluation value is given. When the pressure deviation increases or exceeds the allowable deviation range, the evaluation value is reduced accordingly. As an evaluation item corresponding to the unit water supply energy consumption, in each evaluation time interval, the motor power of each operating pump group of the booster water supply equipment and the water supply flow at the equipment outlet are obtained, and the energy consumption level corresponding to the unit flow is calculated, that is, the unit water supply energy consumption. Then, the unit water supply energy consumption is compared with the reference unit energy consumption determined according to the design conditions or historical energy-saving operation data. When the current unit water supply energy consumption is less than or close to the reference value, the energy-saving effect is considered to be good and a higher evaluation value is given. When the unit water supply energy consumption is significantly higher than the reference value, the energy consumption is considered to be too high and the evaluation value is reduced by a certain proportion. As evaluation items corresponding to negative pressure risk and water hammer risk, a digital twin hydraulic model is used to simulate the hydraulic response of candidate control actions in the prediction time domain. The negative pressure and water hammer risk indicators are obtained based on the calculation of whether negative pressure or water hammer peaks exceeding the preset pressure limit occur at each monitoring location during the time period. For life evaluation items corresponding to the number of equipment start-ups and shutdowns and the magnitude of changes in operating conditions, within each evaluation time window, the number of start-ups and shutdowns for each pump unit, as well as the magnitude and frequency of frequency changes during the corresponding frequency adjustment process, are statistically analyzed. Based on equipment technical data or operation and maintenance experience, a recommended maximum start-up and shutdown frequency can be set for the number of pump unit start-ups and shutdowns, and a recommended range can be set for the frequency adjustment magnitude and the number of adjustments. Based on these, a reference number of start-ups and shutdowns and a reference total frequency change are given, which are then used in the calculation... When the actual number of start-stop cycles in the current window is compared with the reference number of start-stop cycles, if the number of start-stop cycles is close to or lower than the reference value, it is considered to have little impact on lifespan. If the number of start-stop cycles is significantly higher than the reference value, the lifespan evaluation value is reduced by the excess ratio. , , , These are the corresponding weighting coefficients.

[0057] Furthermore, reinforcement learning training and strategy updating refer to using the operational state characteristics reflecting the operation of the booster pump system and its connected pipe network as the state input for reinforcement learning within the safe action domain that satisfies the first and second safety constraints. The control variables, consisting of frequency setpoints, pump group activation or deactivation commands, and adjustments to the opening of key pressure regulating components, are used as the action output for reinforcement learning. A digital twin hydraulic model is used to simulate the pressure and flow changes at each monitoring location after executing each candidate control action in the prediction time domain. The corresponding reward value is calculated based on the reward function to form sample data between the state variables, control variables, reward values, and the operational state characteristics at the next moment. Using sample data, existing reinforcement learning methods are employed to iteratively adjust the parameters of the reinforcement learning policy network. For example, reinforcement learning algorithms suitable for continuous action spaces, such as Deep Deterministic Policy Gradient (DDPG) and Proximal Policy Optimization (PPO), can be selected to improve the overall operating performance of the policy network in the desired sense, thereby completing the reinforcement learning training. As training progresses, the updated policy network can output a new combination of control variables under the constraints of given operating state characteristics and safe action domain. The combination of control variables obtained based on the latest policy network is used as the updated control strategy, and after passing a safety assessment, it is applied as the current version of the control strategy to the booster pump water supply equipment.

[0058] In addition, the operational evaluation indicators include at least one indicator for quantifying the deviation between the digital twin hydraulic model and the actual operation, and at least one indicator for quantifying the level of safety risk. During reinforcement learning training, the learning rate and policy update step size of the reinforcement learning policy network can be adaptively adjusted according to the model deviation indicator in the operational evaluation indicators: when the model deviation indicator is small, the learning rate and policy update step size are kept near the pre-set base learning rate and base policy update step size; when the model deviation indicator increases, the learning rate and policy update step size are appropriately increased proportionally within the preset allowable range, so that the policy network can complete parameter correction more quickly when the model deviation is large.

[0059] Specifically, adaptive adjustment satisfies: .

[0060] in, and Update the step size based on the currently used learning rate and strategy. and The step size for updating the base learning rate and base policy. For adjustment coefficients, This is a model bias index.

[0061] Finally, the strategy protection mechanism includes the following: within a preset evaluation window, if the model deviation index or security risk-related index in the performance evaluation indicators continuously exceeds the corresponding threshold, or exceeds a preset number of times within the evaluation window, the current reinforcement learning strategy version will be rolled back to the most recent historical strategy version that passed the security evaluation. After the rollback, the online update of the strategy network parameters will be temporarily frozen, and offline training and verification will only be allowed in the digital twin hydraulic model. When the offline training results meet the security control conditions and the comprehensive performance indicators are better than the current strategy, the freeze will be lifted and online updates will resume.

[0062] S5. Write back the operational feedback information generated during the execution control process to the water supply knowledge graph and digital twin hydraulic model to update the knowledge information and model parameters corresponding to the operating scenario, so as to realize the continuous self-learning and adaptive optimization of the intelligent control strategy of the superimposed pressure water supply equipment.

[0063] It should be noted that the operational feedback information includes monitoring data collected during operation, recorded abnormal events and their occurrence conditions, control actions taken and their execution results, and information on manual intervention. The operational feedback information is used to update the entity nodes corresponding to the operating conditions and their relationships in the water supply knowledge graph, and to perform incremental parameter calibration in the digital twin hydraulic model, so as to improve the model's fitting accuracy to the actual operating behavior of the booster pump equipment, and support the continuous self-learning and adaptive optimization of the intelligent control strategy.

[0064] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.

Claims

1. An intelligent control method for a direct connection type pipe network pressure boosting water supply device, characterized by, The method comprises the following steps: S1, constructing a water supply knowledge graph, introducing water supply scene elements as entity nodes in the water supply knowledge graph, and defining the association relationship and operation constraint between the entity nodes as graph edge relationship, updating the attributes of the entity nodes based on real-time monitoring data in the operation process during the operation process of the superimposed water supply equipment and its connected pipe network, and obtaining a water supply working condition scene subgraph representing the current water supply state; S2, based on the structure and operation characteristic information reflected by the water supply knowledge graph, a corresponding digital twin hydraulic model is established, and the digital twin hydraulic model is continuously calibrated using hydraulic state monitoring data, in each control period, the calibrated digital twin hydraulic model is called to simulate and calculate the hydraulic response of the candidate control action in the preset prediction time domain, and the action set satisfying the safety control condition is determined from the candidate control action according to the preset safety control condition, to form the first safety constraint based on the digital twin; S3, constructing a working condition causal graph model in the water supply knowledge graph, associating nodes representing operation modes and water supply safety risks through causal edges, in the control process, combining the water supply working condition scene subgraph, performing counterfactual reasoning on each candidate control action on the working condition causal graph model, adjusting the candidate control action when the reasoning result meets the preset risk triggering condition, and based on the current water supply working condition scene, retrieving similar historical scenes from the water supply knowledge graph, and using the corresponding strategy prior information to set the related parameters of the reinforcement learning strategy network, forming the second safety constraint and strategy prior based on the knowledge graph; S4, within the safety action domain defined by the first safety constraint and the second safety constraint, the operating state characteristics reflecting the operating conditions of the superimposed water supply equipment and its connected pipe network are used as the state quantity of reinforcement learning, the control variables used to represent the control behavior are used as the action quantity of reinforcement learning, a reward function with the comprehensive operation index as the optimization objective is constructed, reinforcement learning training and strategy updating are performed in the digital twin hydraulic model, and the parameters of the reinforcement learning strategy network are adjusted according to the operation evaluation index between the digital twin prediction result and the actual operation performance, the strategy protection mechanism is triggered when the operation evaluation index exceeds the preset range, and the control action after the above safety evaluation is issued to the superimposed water supply equipment; S5, writing the operation feedback information generated in the execution control process back to the water supply knowledge graph and the digital twin hydraulic model. 2.The intelligent control method of a direct connection pipe network pressure boosting water supply device according to claim 1, characterized in that, The attribute updating of the entity node refers to that, in the operation process, at each preset time interval, the attribute value of the entity node corresponding to the monitoring object in the water supply knowledge graph is refreshed according to the real-time monitoring data; The attribute instantiation of the entity node refers to that, after the attribute updating is completed, the entity nodes and graph edge relationships corresponding to the water supply scene elements corresponding to the current operation state are selected from the water supply knowledge graph according to the time information corresponding to the current time or the current control period, and the updated attribute value, the selected entity nodes and graph edge relationships are combined together. 3.The intelligent control method of a direct connection pipe network pressure boosting water supply device according to claim 2, characterized in that, The digital twin hydraulic model is modeled based on pipe network nodes and hydraulic elements, and during continuous calibration, the hydraulic state monitoring data collected is used to perform several small-step updates on the model parameters in each control cycle or each preset time window, and the parameter vector is corrected according to the sensitive direction of the loss function on the model parameters each time, so that the loss function value defined based on the pressure deviation of each monitoring position gradually decreases, and when the change amount of the loss function between adjacent two updates is less than a preset convergence threshold, or the number of updates reaches a preset upper limit, the current round of parameter calibration is ended, and the parameter vector obtained at this time is taken as the parameter of the current version of the digital twin hydraulic model.

4. The intelligent control method of a direct connection pipe network superposition water supply equipment according to claim 3, characterized in that, The safety control condition includes hydraulic response simulation of each candidate control action in a preset prediction time domain in the digital twin hydraulic model, calculation of corresponding negative pressure risk indicators and water hammer risk indicators, and step-by-step judgment on the prediction time domain, and when the negative pressure risk indicator or the water hammer risk indicator at any prediction time exceeds a preset safety threshold, the corresponding candidate control action is marked as not meeting the safety control condition and is removed.

5. The intelligent control method of a direct connection pipe network superposition water supply equipment according to claim 4, characterized in that, The counterfactual reasoning refers to classifying the candidate control action into the corresponding operation mode while keeping other conditions in the current water supply working condition scene subgraph unchanged, selecting the corresponding operation mode node in the working condition causal graph model as the starting node of the counterfactual reasoning, and gradually inferring the working condition changes and operation abnormal conditions that may be caused by the candidate control action within a limited number of reasoning steps according to the direction of the causal edges in the working condition causal graph model, and checking whether the supply water safety risk event node will be triggered in the inference process; The pre-defined accident chain is a set of causal paths determined in advance according to the operation rules and historical accident cases, and connected from the operation mode node to the supply water safety risk event node through one or more working condition and abnormal event relationships; When the counterfactual reasoning result shows that there is a causal path from the current operation mode node to the target supply water safety risk event node along the pre-defined accident chain under the assumption of executing a certain candidate control action, it is determined that the candidate control action will evolve along the pre-defined accident chain, and the candidate control action is marked as having unacceptable safety risks, which is used for rejection or modification in the safety constraint.

6. The intelligent control method of a direct connection pipe network superposition water supply equipment according to claim 5, characterized in that, The acquisition of the strategy prior information includes graph embedding calculation on the water supply working condition scene subgraph to obtain a feature vector for representing the current working condition characteristics, and similarity calculation between the feature vector and each historical working condition vector pre-stored in the water supply knowledge graph, the similarity is measured by using cosine similarity, the ratio of the inner product of the current working condition vector and the historical working condition vector to the product of their module lengths is taken as the similarity value, and a number of historical working condition scenes are selected from large to small according to the similarity value to form a similar working condition set, and the historical control strategy parameters, reward weights and exploration strategies corresponding to the similar working condition set are extracted as the strategy prior information of the reinforcement learning strategy network.

7. The intelligent control method of a direct connection pipe network superposition water supply equipment according to claim 6, characterized in that, The operation state features include features reflecting the supply water pressure distribution condition, features reflecting the water load level, features reflecting the consistency between the digital twin hydraulic model and the field operation, and features reflecting the equipment operation condition; The feature reflecting the consistency between the digital twin hydraulic model and the field operation is obtained by selecting a plurality of pressure monitoring positions for comparison at each time, respectively obtaining the measured pressure and the simulation pressure calculated by the digital twin hydraulic model at each monitoring position, calculating the average value of the difference between the simulation pressure and the measured pressure at each monitoring position, and taking the average difference value as the average model pressure deviation at the corresponding time, thereby representing the consistency between the digital twin hydraulic model and the field operation data. 8.The intelligent control method of a direct connection pipe network pressure boosting water supply device according to claim 7, characterized in that, The comprehensive operation index includes a plurality of evaluation indexes corresponding to the pressure service quality, the water supply energy efficiency, the water supply safety risk and the equipment operation life. The instant reward value of the reinforcement learning is calculated at each evaluation time by weighting and combining the evaluation items corresponding to the water supply pressure deviation and the pressure fluctuation, the evaluation item corresponding to the unit water supply energy consumption, the evaluation items corresponding to the negative pressure risk and the water hammer risk, and the life evaluation item corresponding to the number of equipment start-stop and the working condition change amplitude according to the weight coefficients, thereby obtaining the reward value at the evaluation time, and each weight coefficient represents the relative importance of the pressure service quality, the water supply energy efficiency, the water supply safety risk and the equipment life in the comprehensive operation index. 9.The intelligent control method of a direct connection pipe network pressure boosting water supply device according to claim 8, characterized in that, The reinforcement learning training and policy updating are performed within the safety action domain defined by the first safety constraint and the second safety constraint, the operation state features reflecting the operation conditions of the cascade pressure water supply equipment and the connected pipe network are taken as the state quantity input of the reinforcement learning, the control variables composed of the frequency setting value, the pump group input or exit instruction and the key pressure regulating element opening adjustment amount are taken as the action quantity output of the reinforcement learning, the digital twin hydraulic model is used to simulate the pressure and flow changes at each monitoring position after the execution of each candidate control action in the prediction time domain, and the corresponding reward value is calculated according to the reward function, thereby forming sample data between the state quantity, the control variable, the reward value and the next time operation state feature; Based on the sample data, the existing reinforcement learning method is used to iteratively adjust the parameters of the reinforcement learning policy network, so that the policy network improves the comprehensive operation index in the expected sense, thereby completing the reinforcement learning training; With the training, the updated policy network can output a new control variable combination under the given operation state feature and the constraint of the safety action domain, the control variable combination obtained according to the latest policy network is taken as the control strategy after the policy update, and it is applied to the cascade pressure water supply equipment as the current version of the control strategy after the safety evaluation.

10. The intelligent control method of a direct connection pipe network superposition water supply equipment according to claim 9, characterized in that, The operation evaluation index includes at least one index for quantifying the deviation between the digital twin hydraulic model and the actual operation, and at least one index for quantifying the safety risk level. In the reinforcement learning training process, according to the model bias index in the running evaluation index, the learning rate and the policy update step of the reinforcement learning strategy network are adjusted, when the model bias index is small, the learning rate and the policy update step are kept near the pre-set basic learning rate and basic policy update step, when the model bias index increases, the learning rate and the policy update step are proportionally increased within the pre-set allowable range.

Citation Information

Cited By

  • Port water resource whole-process intelligent regulation and control method and system adaptive to dynamic working conditions

    CN122219707A