Water supply network multi-target toughness maintenance method fusing graph convolution reinforcement learning

Through the multi-objective toughness maintenance method of fusion graph convolution reinforcement learning, combined with GCN and DRL, the recovery sequence of the water supply pipeline network is optimized, and the problems of single-dimensional evaluation and local optimal bias in the existing technology are solved, and the multi-objective coordinated optimization and toughness improvement of the water supply pipeline network are achieved.

CN120146709AInactive Publication Date: 2025-06-13SOUTHWEST JIAOTONG UNIV

Patent Information

Application Number
CN202510624467.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing water supply pipeline maintenance methods have problems such as single-dimensional evaluation, local optimal bias and high-dimensional discrete space solution, which is difficult to effectively improve the recovery degree of the pipeline and reduce the recovery time.

Method used

The multi-objective toughness maintenance method of fusion graph convolution reinforcement learning is adopted. By building a multi-objective toughness evaluation framework, combining graph convolution network (GCN) and deep reinforcement learning (DRL), the pipeline recovery sequence is optimized, and the recovery degree of the water supply network is synergistically improved and the recovery time is shortened.

Benefits of technology

Multi-dimensional collaborative optimization has been achieved, which significantly improves the toughness of the water supply pipeline network. The generated recovery sequence toughness index has been increased by 28.0% to 108.8% compared with traditional methods, and the computing efficiency of the model in complex environments is improved through transfer learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146709A_ABST
    Figure CN120146709A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of water supply pipe networks, and provides a water supply pipe network multi-target toughness maintenance method based on fusion graph convolution reinforcement learning, which comprises the following steps: S1, building a data set; s2, constructing a degradation damage module; s3, building a performance evaluation module; s4, calculating transportation time; s5, calculating a recovery degree; s6, calculating a toughness evaluation index; s7, calculating a weight; s8, an M2DRL model is built; s9, training the model; and S10, generating a pipeline recovery sequence through the trained model, and performing evaluation based on a toughness evaluation index. By optimizing the pipeline recovery sequence, the recovery degree of the water supply pipe network can be synergistically improved, the recovery time is shortened, and the system toughness is finally enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of water supply networks, and in particular, to a multi-objective resilience maintenance method for water supply networks integrating graph convolution and reinforcement learning. Background Technique

[0002] Water distribution networks (WDNs) are the core components of urban water supply systems, and their failures may cause serious damage to urban safety, economy, and public health. However, affected by factors such as material aging and environmental loads, the deterioration problem of WDNs is becoming increasingly prominent, and effective maintenance strategies are urgently needed to ensure their long-term stable operation. In recent years, the research focus has gradually shifted from traditional reliability-based protection methods to resilience-based protection methods. The resilience of infrastructure systems is defined as the ability to effectively reduce risks and quickly restore services with minimal harm to the public, including recovery ability, absorption ability, and adaptive ability. Among them, the recovery ability is very crucial in ensuring the water supply safety of communities and has become the primary concern of water supply companies. Specifically, the recovery ability can be conceptualized by the degree of recovery and the recovery time. Among them, the degree of recovery is usually quantified by hydraulic and water quality indicators. At the same time, in actual work, pipeline maintenance is not carried out on all pipelines at once, but in a certain order. Therefore, how to improve the recovery degree of WDNs while minimizing the recovery time, so as to obtain the optimal recovery sequence of WDNs and enhance the resilience of WDNs has become a current research difficulty.

[0003] In recent years, many studies have proposed resilience-based WDN recovery decision-making methods. The existing methods can be classified into three categories: (1) priority ranking-based methods; (2) greedy search algorithms; (3) metaheuristic algorithms. Among them, the priority ranking-based methods determine the recovery sequence by quantifying the criticality indicators of pipelines, avoiding complex hydraulic simulation processes, and thus having significant computational advantages. However, such methods overly rely on expert judgments or preset rules with strong subjectivity and are difficult to obtain the optimal recovery sequence. Greedy search algorithms select operation steps through an iterative local search mechanism, achieving a balance between timeliness and solution quality. However, this method has the risk of premature convergence and is prone to falling into local optimal solutions.

[0004] For the global optimization requirements, meta-heuristic algorithms characterized by intelligent optimization have become a research hotspot, including genetic algorithms (GA), ant colony optimization (ACO), tabu search (TS), and deep reinforcement learning (DRL), etc. Some scholars have used multi-objective genetic algorithms (MOGA) to reduce risks, control costs, and meet the strategic policies of water companies. At the same time, some scholars have combined GA, ACO, and TS algorithms to propose a multi-objective resilience-driven restoration planning model, which provides an optimization solution for the rapid restoration of WDN by optimizing restoration time, cost, and resilience level. In the field of DRL, some scholars have explored the applications of online and offline DRL in WDN maintenance planning, demonstrating the advantages of DRL in dealing with large-scale complex state space problems. Some scholars have proposed a model based on graph convolutional network (GCN) and DRL to optimize the repair decisions of damaged WDN. By optimizing a single hydraulic index, the restoration ability and resilience of the system are improved. In terms of multi-index optimization, some scholars have also used DRL and GCN to optimize the pipeline maintenance timing, achieving the balance of two indexes of cost and reliability at the same time. These studies not only illustrate the powerful capabilities of meta-heuristic algorithms in global optimization but also provide new ideas and methods for solving complex optimization problems in practical engineering.

[0005] However, the existing research methods for WDN resilience based on meta-heuristic algorithms mainly have three disadvantages: First, a multi-dimensional resilience evaluation system of hydraulics-water quality-time has not been established, and the existing maintenance decisions are limited to single hydraulic index evaluation, resulting in doubts about the actual utility of the restoration plan; Second, the topological characteristics of the pipe network are not fully mined, leading to a local optimum bias in the maintenance sequence and limited improvement in the overall resilience of the pipe network; Third, the computational efficiency of solving high-dimensional discrete spaces is low, and the generalization ability is insufficient due to the preset damage scenarios. Summary of the Invention

[0006] The content of the present invention is to provide a multi-objective resilience maintenance method for water supply pipe networks integrating graph convolutional reinforcement learning, which can synergistically improve the restoration degree of the water supply pipe network and shorten the restoration time by optimizing the pipeline restoration sequence, and ultimately enhance the system resilience.

[0007] A multi-objective resilience maintenance method for water supply pipe networks integrating graph convolutional reinforcement learning according to the present invention includes the following steps: S1: Build a data set; The data set includes the topological structure of the WDN, the age, pipe material, burial depth, and soil type information of each pipeline, obtaining the data under the normal state of the WDN; S2: Build a degradation damage module; The degradation damage module is based on the risk value of each pipeline in the WDN , convert the data in the normal state of the WDN into data in the deteriorated state to provide data support for the subsequent recovery stage; S3: Build a performance evaluation module; After determining the pipeline to be maintained, generate a recovery sequence based on the recovery method, and the toughness improvement effect of the sequence is quantified through the performance evaluation module and the transportation time result; specifically in the performance evaluation module, to obtain the global performance of the WDN, it is necessary to first calculate the evaluation weight and comprehensive performance of each pipeline connection point in the WDN; S4: Calculate the transportation time; S5: Calculate the recovery degree; Recovery degree is usually quantified by the area method of the gray area under the recovery curve. The faster the performance recovers, the larger the area under the recovery curve, and the corresponding is higher; S6: Calculate the toughness evaluation index; The toughness evaluation index includes the recovery time and two parts; the quantification of the recovery time is completed through S4, and the quantification of the recovery degree is completed through S5; S7: Calculate the weight; Simulate the recovery effects of a large number of different pipeline recovery sequences, integrate the obtained transportation time results and recovery degree results, and thus use the entropy weight method to assign weights to the reward values; S8: Build an M2DRL model; M2DRL is a DRL model integrating GCN. At each time step, the agent obtains the current state based on the graph representation of the WDN , executes an action , and after that, the environmental state evolves into , and obtains an immediate reward for policy optimization; the implementation of the model first requires constructing the graph structure of the WDN as the input of the agent, and secondly clearly defining the state, action, and reward value in the interaction process; S9: Train the model; S10: Generate a pipeline recovery sequence through the trained model and evaluate it based on the toughness evaluation index.

[0008] Preferably, the specific steps of S2 are as follows: S21: To determine the of each pipeline, it is necessary to first quantify the failure probability and the failure consequence . Among them, is determined by comparing the maximum stress changing with time on the pipeline with the tensile failure strength based on the pipeline material, and the formula is: ; In the formula, is the service life of the pipeline, is the pipeline diameter; is the traffic load; is the soil elastic modulus; is the pipeline elastic modulus; is the unit soil weight; is the pipeline burial depth; is the time-varying wall thickness of the pipeline, is the internal water pressure of the pipeline; is the lateral earth pressure coefficient; S22: Calculate the of the pipeline using a fuzzy inference system; S23: Through the calculated and , obtain the of each pipeline; The quantization formula is: ; S24: Calculate the impact of pipeline deterioration on performance; By constructing a leakage point, simulate the degree of pipeline deterioration with the leakage flow rate, thereby affecting the hydraulic and water quality changes at the connection point. The formula is: ; In the formula, is the leakage flow rate; is the leakage area; is the flow coefficient; Use the WNTRSimulator and EpanetSimulator provided by the open-source Python package WNTR to simulate the hydraulic changes and water quality changes after adding a leakage point respectively, so as to obtain the water pressure value and water quality calculation coefficient after leakage for the performance evaluation model; S25: Screen the pipelines to be maintained; Assume that when the of the pipeline exceeds the given threshold , then take maintenance measures and put this part of the pipeline into the set of pipelines to be maintained ;

[0009] Preferably, the specific steps of S3 are as follows: S31: Calculate the evaluation weight; The evaluation weight of the connection point represents the importance degree of this point in the WDN, which is determined by its failure consequence . The formula is: ; In the formula, is the connection point fault consequence; is the total number of connection points; is determined by the average value of the fault consequences of the pipes connected to the connection point , and the formula is: ; In the formula, is the fault consequence of the pipe ; The value range is (0, 1]; S32: Calculate the hydraulic performance; In the pressure-driven hydraulic model, the hydraulic performance is defined as the ability of the connection point to meet the water demand of users, and is quantified by pressure adequacy; when , the water pressure meets the designed water supply demand, and the formula is: ; In the formula, is the hydraulic performance of the connection point at the moment; is the water pressure of the connection point ; is the water pressure at which the water supply stops; is the minimum water pressure to meet the water supply demand; S33: Calculate the water quality performance; In the WDN, maintaining an appropriate amount of residual chlorine is the key standard to ensure the drinking water safety of the water supply system. The residual chlorine concentration is used as an evaluation index for the water quality performance of the WDN ; In the WDN, the attenuation process of the residual chlorine concentration conforms to the first-order reaction kinetics, and the general model is: ; In the formula, is the residual chlorine concentration of the connection point at the moment; is the initial chlorine concentration of the connection point ; is the time; is the total chlorine attenuation coefficient; S34: Calculate the global performance; Using the evaluation weights of each connection point (the calculation results of S31) and the comprehensive performance (hydraulic performance and water quality performance, which are the calculation results of S32 and S33), calculate the global performance of the WDN at the moment, including the global hydraulic performance HP and the global water quality performance , and the formula is: 。

[0010] Preferably, in S4, the recovery time is measured by the sum of the transportation times between pipelines during maintenance work, and the formula is: ; In the formula, H is the total transportation time; is the transportation time at the moment of is the total step length of the recovery time, and the value is the number of pipelines to be maintained; is a decision variable, indicating that it takes 1 when maintaining pipeline at the moment of and 0 otherwise; is the transportation time between pipeline and pipeline ; is the set of pipelines to be maintained; among them, calculating needs to be obtained based on OpenStreetMap and ArcGIS to create a start-end matrix.

[0011] Preferably, the specific calculation steps of S5 are as follows: Based on the quantification schematic diagram of the global performance, the degree of hydraulic recovery and the degree of water quality recovery can be quantified, and the formula is: ; In the formula, and respectively represent the start and end times of the recovery process, and each time step in the middle represents a pipeline being maintained.

[0012] Preferably, the specific steps of S6 are as follows: Obtain the recovery time and the degree of recovery, calculate the resilience index , and transform it into a multi-objective optimization problem. The formula is: ; In the formula, and respectively represent the after normalization, and , corresponding to the weights and respectively, and the weight values are calculated by the entropy weight method; is used to evaluate the effects of different pipeline recovery sequences. The larger the value, the stronger the resilience improvement ability of the corresponding pipeline recovery sequence.

[0013] Preferably, the specific steps of S8 are as follows: S81: Create a graph structure; The WDN graph structure is represented as , where is a set of nodes, and each node represents a pipeline; is a set of edges, and each edge represents the connection relationship between two pipelines; is an adjacency matrix with a dimension of , where each element , indicates whether there is a connection between node and , ; S82: Define states; Each node in the graph structure corresponds to a feature vector , where is the pipeline ID; is the pipeline length; is the pipeline diameter; is the pipeline material; is the pipeline risk value; is the pipeline hydraulic performance; is the pipeline water quality performance; is the pipeline maintenance status, with the value to be maintained being 0 and the value of having been maintained being 1; at each time step, and will be updated according to the maintenance actions taken; S83: Define actions; Based on the state of the WDN, the agent selects a pipeline from the set of pipelines that meet the maintenance conditions at time for maintenance, which is represented as , thus forming the current action . At the same time, the maintenance actions at all times need to be equal to the set of pipelines to be maintained , and the formula is: ; In the formula, represents the action at the first time; represents the action at the last time; S84: Set rewards; The reward value takes the resilience index as the only optimization goal, that is: ; In the formula, is the immediate reward.

[0014] Preferably, the specific steps of S9 are as follows: S91: Initialize parameters: Set the initial number of steps , the current training cycle , the initial exploration rate , the exploration rate decay coefficient and the total number of training cycles ; Identify the state of the WDN , the maintenance status of all pipelines is set to 0, and the time step ; S92: Calculate the dynamic exploration rate: Update the current exploration rate through the decay formula , so that the agent tries different maintenance methods frequently in the initial stage and gradually reduces randomness in the later stage, focusing on the optimal strategy; S93: Action selection: Adopt the greedy strategy to select the maintenance action, randomly select a pipeline with probability , or select the pipeline with the largest value according to the agent's choice with probability ; S94: State update and reward calculation: Modify the of the selected pipeline to 1, calculate the new state , and calculate the immediate reward through formula (19) ; S95: Value update and network training: The model updates the network parameters by combining the value with the maximum future value and comparing it with the current value; As the training progresses, the agent can gradually judge the true value of the action and achieve autonomous learning of the WDN maintenance decision; S96: Loop control: When all pipelines have completed maintenance, that is, complete one , reset the WDN state and start a new round of training. When the number of training times reaches , the training ends.

[0015] The beneficial effects of the present invention are as follows: (1) Construct a multi-objective resilience evaluation framework to achieve multi-dimensional collaborative optimization. The present invention proposes a multi-objective resilience evaluation framework that integrates hydraulic performance, water quality performance, and transportation time, breaking through the limitations of traditional single-dimensional evaluation. By assigning weights to each dimension through the entropy weight method, this framework can comprehensively quantify the recovery degree and recovery time of the water supply network, providing a scientific basis for maintenance decisions and achieving multi-objective collaborative optimization.

[0016] (2) Integrate GCN and DRL to enhance the global optimization ability in complex networks. To make up for the deficiencies of traditional DRL in dealing with graph-structured data, the present invention combines the graph convolutional network (GCN) with deep reinforcement learning (DRL) to construct the M2DRL model. GCN encodes the topological structure and dynamic node features of the water supply network, and the generated node embedding vectors are used as the state input of DRL, enabling it to accurately perceive the global state, so as to achieve more effective global policy search in the high-dimensional discrete action space, and significantly enhance the global optimization ability of the model.

[0017] (3) To solve the problems of model adaptability and computational efficiency in different damage scenarios, the present invention introduces a transfer learning strategy. Through parameter sharing and fine-tuning of the pre-trained model, the model can quickly adapt to new scenarios, avoid the high computational cost of training from scratch, and enhance the cross-scenario adaptability and computational efficiency of the model. Experimental results show that transfer learning significantly reduces the computational time and improves the engineering practicability of the model in complex environments. Description of the Drawings

[0018] Figure 1 A multi-objective resilience maintenance method for water supply networks integrating graph convolutional reinforcement learning in the embodiment; Figure 2 A schematic diagram of the performance of WDN in the embodiment considering before and after deterioration and during the maintenance process; Figure 3 A schematic diagram of the interaction between the DQN agent and the environment in the embodiment; Figure 4 A schematic diagram of the architecture and training flow chart of the M2DRL model in the embodiment; Figure 5 A schematic diagram of the CA1 network in the embodiment; Figure 6 A failure probability curve of four types of pipelines over time in the embodiment; Figure 7(a) is a schematic diagram of the failure probability distribution in the risk estimation based on the risk equation in the embodiment; Figure 7(b) is a schematic diagram of the failure consequence distribution in the risk estimation based on the risk equation in the embodiment; Figure 7(c) is a schematic diagram of the risk distribution in the risk estimation based on the risk equation in the embodiment; Figure 8 A schematic diagram of the node weight distribution in the embodiment; Figure 9(a) is a schematic diagram of the degree of hydraulic recovery during the training process in the embodiment; Figure 9(b) is a schematic diagram of the degree of water quality recovery during the training process in the embodiment; Figure 9(c) is a schematic diagram of the total transportation time during the training process in the embodiment; Figure 9(d) is a schematic diagram of the toughness index during the training process in the embodiment; Figure 10(a) is a schematic diagram of the overall hydraulic performance of different restoration methods in the embodiment; Figure 10(b) is a schematic diagram of the overall water quality performance of different restoration methods in the embodiment; Figure 10(c) is a schematic diagram of the transportation time of different restoration methods in the embodiment; Figure 11(a) is a schematic diagram of the distribution of pipes to be maintained in Scenario 1 (83 pipes) in the embodiment; Figure 11(b) is a schematic diagram of the distribution of pipes to be maintained in Scenario 2 (52 pipes) in the embodiment; Figure 11(c) is a schematic diagram of the distribution of pipes to be maintained in Scenario 3 (62 pipes) in the embodiment; Figure 12 It is a schematic diagram of the comparison of RI values of different restoration methods in Scenarios 1-3 in the embodiment; Figure 13(a) is a schematic diagram of the distribution of pipes to be maintained in Scenario 4 (37 pipes) in the embodiment; Figure 13(b) is a schematic diagram of the distribution of pipes to be maintained in Scenario 5 (29 pipes) in the embodiment; Figure 13(c) is a schematic diagram of the distribution of pipes to be maintained in Scenario 6 (24 pipes) in the embodiment. Detailed implementation manners

[0019] To further understand the content of the present invention, the present invention will be described in detail in combination with the accompanying drawings and embodiments. It should be understood that the embodiments are only for explaining the present invention rather than limiting it.

[0020] Embodiment As Figure 1 shown, this embodiment provides a multi-objective toughness maintenance method for water supply pipe networks integrating graph convolution and reinforcement learning, which includes the following steps: S1: Build a data set; The data set includes the topological structure of the WDN, the age, pipe material, burial depth and soil type information of each pipe, and obtains the data in the normal state of the WDN.

[0021] S2: Build a deterioration damage module; The deterioration damage module converts the data in the normal state of the WDN into data in the deteriorated state based on the risk value of each pipe in the WDN , and provides data support for the subsequent restoration stage.

[0022] The specific steps of S2 are as follows: S21: To determine the of each pipe, it is necessary to first quantify the failure probability and the consequences of failure . Among them, by comparing the maximum stress that changes over time on the pipeline with the tensile failure strength based on the pipeline material to determine, the formula is: ; In the formula, is the service life of the pipeline, is the pipeline diameter; is the traffic load; is the soil elastic modulus; is the pipeline elastic modulus; is the unit soil weight; is the pipeline burial depth; is the time-varying wall thickness of the pipeline, is the internal water pressure of the pipeline; is the lateral earth pressure coefficient; S22: Use a fuzzy inference system (existing technology) to calculate the ; S23: Through the calculated and , obtain the for each pipeline; The quantification formula is: ; S24: Calculate the impact of pipeline deterioration on performance; By constructing a leakage point, use the leakage flow rate to simulate the degree of pipeline deterioration, thereby affecting the hydraulic and water quality changes at the connection point. The formula is: ; In the formula, is the leakage flow rate; is the leakage area; is the flow coefficient; Use the WNTRSimulator and EpanetSimulator provided by the open-source Python package WNTR to simulate the hydraulic changes and water quality changes after adding a leakage point respectively, so as to obtain the water pressure value and water quality calculation coefficient after leakage, which are used in the performance evaluation model.

[0023] S25: Screen the pipelines to be maintained; Assume that when the of the pipeline exceeds the given threshold , then take maintenance measures and put this part of the pipeline into the set of pipelines to be maintained . Among them, the specific should be determined by the water utility company according to its budget, resources, and repair policy.

[0024] S3: Build a performance evaluation module; After determining the pipeline to be maintained, a recovery sequence is generated based on the recovery method, and the toughness improvement effect of the sequence is quantified through the performance evaluation module and the transportation time result; specifically, in the performance evaluation module, to obtain the global performance of the WDN, it is necessary to first calculate the evaluation weight and comprehensive performance of each pipeline connection point in the WDN.

[0025] The specific steps of S3 are as follows: S31: Calculate the evaluation weight;

[0026] S31: Calculate the evaluation weight; Connection point The evaluation weight of represents the importance degree of this point in the WDN, which is determined by its failure consequence The formula is: ; In the formula, is the failure consequence of the connection point ; is the total number of connection points; is determined by the average value of the failure consequences of the root pipelines connected to the connection point The formula is: ; ; In the formula, is the failure consequence of the pipeline ; The value range is (0, 1]; S32: Calculate the hydraulic performance (one of the comprehensive performances); In the pressure-driven hydraulic model, the hydraulic performance is defined as the ability of the connection point to meet the water demand of users, and is quantified by pressure adequacy; when , the water pressure meets the designed water supply demand, and the formula is: ; In the formula, is the hydraulic performance of the connection point at the moment; is the water pressure of the connection point ; is the water pressure when the water supply stops; is the minimum water pressure to meet the water supply demand. and are set to 0 and 30 respectively.

[0027] S33: Calculate the water quality performance (one of the comprehensive performances); In a WDN, maintaining an appropriate amount of residual chlorine is a key criterion to ensure the safety of drinking water in the water supply system. The residual chlorine concentration serves as an evaluation index for the water quality performance of the WDN. In the WDN, the attenuation process of the residual chlorine concentration conforms to the first-order reaction kinetics, and the general model is: ; In the formula, is the residual chlorine concentration at the connection point at time; is the initial chlorine concentration at the connection point ; is the time; is the total chlorine attenuation coefficient; S34: Calculate the global performance; Using the evaluation weights of each connection point (the calculation result of S31) and the comprehensive performance (hydraulic performance and water quality performance, which are the calculation results of S32 and S33), calculate the global performance of the WDN at time, including the global hydraulic performance HP and the global water quality performance , and the formula is: .

[0028] S4: Calculate the transportation time; In S4, the focus of this embodiment is to obtain the optimal restoration timing of the pipeline. Therefore, the restoration time is measured by the sum of the transportation times between pipelines during the maintenance work, and the formula is: ; In the formula, H is the total transportation time; is the transportation time at time; is the total step length of the restoration time, and its value is the number of pipelines to be maintained; is a decision variable, which takes 1 when maintaining the pipeline at time , and 0 otherwise; is the transportation time between pipeline and pipeline ; is the set of pipelines to be maintained; among them, calculating

[0029] S5: Calculate RD ; Based on the quantification schematic diagram of the global performance (such as Figure 2As shown), it can quantify the degree of hydraulic restoration and the degree of water quality restoration , and the formula is: ; In the formula, and respectively represent the start and end times of the restoration process, and each time step in the middle represents the maintenance of a pipeline.

[0030] S6: Calculate the resilience evaluation index; The resilience evaluation index includes the restoration time and two parts; the quantification of the restoration time is completed through S4, and the quantification of the restoration degree is completed through S5. Obtain the restoration time and restoration degree, and calculate the resilience index , which is transformed into a multi-objective optimization problem, and the formula is: ; In the formula, and respectively represent the after normalization, and , corresponding to the weights and respectively, and the weight values are calculated by the entropy weight method; is used to evaluate the effects of different pipeline restoration sequences. The larger the value, the stronger the resilience improvement ability of the corresponding pipeline restoration sequence.

[0031] S7: Calculate the weights; Simulate the restoration effects of a large number of different pipeline restoration sequences, integrate the obtained transportation time results and restoration degree results, and thus use the entropy weight method to assign weights to the reward values.

[0032] S8: Build the M2DRL model; M2DRL is a DRL model integrating GCN, Figure 3 which shows the interaction process between the agent and the environment. At each time step, the agent obtains the current state based on the graph representation of the WDN , executes the action and then the environmental state evolves into , and obtains the immediate reward for policy optimization; the implementation of the model first requires constructing the graph structure of the WDN as the input of the agent, and secondly clearly defining the state, action and reward values in the interaction process.

[0033] The specific steps of S8 are as follows: S81: Create the graph structure; The WDN graph structure is represented as where is a set of nodes, and each node represents a pipeline; is a set of edges, and each edge represents the connection relationship between two pipelines; is an adjacency matrix with a dimension of where each element , indicates whether there is a connection between nodes and , ; S82: Define the state; Each node in the graph structure corresponds to a feature vector , where is the pipeline ID; is the pipeline length; is the pipeline diameter; is the pipeline material; is the pipeline risk value; is the pipeline hydraulic performance; is the pipeline water quality performance; is the pipeline maintenance status, with the value to be maintained being 0 and the value of having been maintained being 1; at each time step, and will be updated according to the maintenance actions taken; S83: Define the action; Based on the state of the WDN, the agent selects a pipeline from the set of pipelines that meet the maintenance conditions at time for maintenance, denoted as , thus forming the current action . At the same time, the maintenance actions at all times need to be equal to the set of pipelines to be maintained , and the formula is: ; In the formula, represents the action at the first time; represents the action at the last time.

[0034] S84: Set the reward; The goal of this embodiment is to find the optimal recovery sequence that can maximize the resilience value of the WDN. Therefore, the reward value uses the resilience index as the only optimization goal, that is: ; In the formula, is the immediate reward.

[0035] S9: Train the model; The specific steps of S9 are as follows: S91: Initialize parameters: Set the initial number of steps , the current training cycle , the initial value of the exploration rate , the exploration rate decay coefficient and the total number of training cycles ; Identify the state of the WDN , and set the maintenance status of all pipelines to 0, and the time step ; S92: Calculate the dynamic exploration rate: Update the current exploration rate through the decay formula , so that the agent tries different maintenance methods frequently in the initial stage and gradually reduces randomness in the later stage, focusing on the optimal strategy.

[0036] S93: Action selection: Adopt the greedy strategy to select the maintenance action, randomly select a pipeline with probability , or select the pipeline with the largest value according to the agent's choice with probability .

[0037] S94: State update and reward calculation: Modify the of the selected pipeline to 1, calculate the new state , and calculate the immediate reward through formula (19) .

[0038] S95: value update and network training: The model updates the network parameters by comparing the combination of with the future maximum value and the current value; As the training progresses, the agent can gradually judge the true value of the action and achieve autonomous learning of the WDN maintenance decision.

[0039] S96: Loop control: When all pipelines have completed maintenance ( are all 1), that is, complete one , reset the WDN state and start a new round of training. When the number of training times reaches , the training ends.

[0040] Figure 4 On the right is the input and specific network structure of the agent. The agent uses the GCN module to extract environmental features and predicts the Q value through the FCN module to guide the maintenance decision. In addition, in this embodiment, the network parameters are updated by the Adam optimizer, and the experience replay technology is used to improve the training stability.

[0041] S10: Generate a pipeline restoration sequence through the trained model and evaluate it based on the resilience evaluation index.

[0042] Experimental analysis I. Data preparation (1) Dataset construction The present invention uses the WDN research database publicly available at the University of Kentucky to analyze a certain WDN located in California, named the CA1 network, as Figure 5 shown. The WDN includes 111 connection points, 126 pipelines, and 1 elevated storage tank. The CA1 pipe network covers an area of 4 square kilometers, and the total length of the pipelines is 17.90 kilometers. Among them, the pipe diameter range of CA1 is 8 - 16 inches. Therefore, the pipe diameter range considered in this embodiment is 150 - 400 mm.

[0043] (2) Deterioration damage results Fault probability analysis results: Given that most pipeline materials in the United States are made of cast iron (CI), it is assumed that the pipeline material in this case is CI. To ensure consistency with the CA1 data, four pipeline types are adopted in this embodiment (Type 1: D = 150 mm; Type 2: D = 200 mm; Type 3: D = 300 mm; Type 4: D = 400 mm) for fault probability analysis. Considering that most CI pipes were produced before 1948, the age of the pipelines is defined as a random number uniformly distributed between 70 and 100 years. Using the Monte Carlo simulation method with 100,000 sampling points, the time-varying fault probability curve of CI pipelines is obtained as Figure 6 .

[0044] Risk analysis results: Combine the fault probability with the fault consequence to quantify the risk of WDN pipelines. In terms of fault probability, according to Figure 6 the fault probability of each pipeline is evaluated, and the calculation results are shown in Figure 7(a). In terms of fault consequence, a fuzzy hierarchical inference system is used to calculate the comprehensive consequence of each pipeline, and the result is in the range of 0 - 1, as shown in Figure 7(b). Figure 7(c) is the risk value calculated by Formula 2, where the risk value of Pipeline 123 is the highest, being 0.535, and the risk value of Pipeline 226 is the lowest, being 0.169. By maintaining the pipelines with high risk values, the fault probability of the pipelines is reduced, thereby reducing the risk.

[0045] Pipelines to be maintained: When the risk threshold is set to 0.3, the risk values of 83 out of 126 pipelines in the WDN exceed this threshold. In this embodiment, these 83 pipelines are included in the set of pipelines to be maintained and maintained using different restoration methods.

[0046] For each pipeline of the WDN is used for leakage simulation to obtain the performance results of the pipeline. Meanwhile, the node weight results of the WDN are as Figure 8 shown. According to the performance results and weight values of the pipeline, before the maintenance starts in the deteriorated state of the WDN, HP is approximately 0.6874, is approximately 0.1775.

[0047] (3) Entropy weight method for weighting To determine the importance of different variables in the resilience index, the present invention randomly generates 1,000 pipeline restoration sequences, and based on these sequences, 1,000 sets of maintenance results are obtained, including , and . The entropy weight method is used to quantitatively analyze the weights of these variables. The results show that HRD , QRD and H have weights of approximately 52%, 25%, and 23% respectively, as shown in Table 1. Among them, has the highest weight, indicating its main role in contributing to the system resilience during the pipeline restoration process; and have relatively small weights, but still provide important decision-making bases for optimizing the restoration sequence.

[0048] Table 1 Table of weight calculation results based on the entropy weight method

[0049] II. Model training The key parameters used during the M2DRL training process are shown in Table 2. In the case of 83 pipelines being damaged, the total number of training cycles is 500, indicating that the model is updated 41,500 times during the training process. The decay coefficient is set to 8,800. Therefore, at the end of the training, the exploration rate will be close to 0 (approximately 0.00895).

[0050] Table 2 Table of key parameters during the M2DRL model training process

[0051] Figures 9(a), 9(b), 9(c), and 9(d) show the , , and of the WDN during 500 rounds of training, as well as the corresponding value trends, and the Savitzky-Golay filter is used to smooth the change trends. Because Among the calculations the weight value is the largest, so optimizing during the training process becomes the key task of the intelligent agent, showing relatively high stability. The experimental results show that within the first 300 rounds of training, due to the control parameter is still large, most maintenance operations are randomly selected. As the training progresses, the control parameter gradually decreases. The intelligent agent mainly makes decisions based on GCN and obtains a recovery plan with a high resilience index. In the later stage, there are still fluctuations in various parameters, which are mainly attributed to the inherent randomness of the neural network and the complexity of the high-dimensional state space.

[0052] III. Comparative Analysis of Multiple Methods To evaluate the resilience improvement effect of the pipeline recovery sequence generated by the proposed M2DRL model (named M1), this embodiment compares it with five of the most common traditional decision-making methods (named M2 - M6), including a method based on genetic algorithm (GA) (M2), two methods based on greedy search, including the static importance method (M3) and the dynamic importance method (M4), a risk-value-based prioritization method (M5), and a diameter-based prioritization method (M6). The experimental results are shown in Table 3.

[0053] Table 3 Recovery Effects and Computational Time Tables of Different Recovery Methods

[0054] The results show that the resilience index of the M2DRL model (M1) is improved by 28.0% - 108.8% compared with the traditional methods (M2 - M6), verifying the effectiveness of the multi-objective M2DRL model in pipeline maintenance decision-making. Specifically in terms of various indicators, the of M1 reaches 1.000, significantly better than other models, 39.1% more than the sub-optimal method M6 (0.719); is close to the theoretical optimal value (0.993), with only a 0.7% gap compared with 1.000 of M5; Because its weight is the lowest, the value is 0.181. In contrast, the of M3 has the highest value, but its extreme strategy results in only 0.628, indicating that M1 can effectively avoid local optimal traps. Although the computational time of M1 is slightly higher than that of some traditional methods, its advantages in comprehensive performance make it a better choice.

[0055] Among M2 to M6, for M2 based on GA, the (0.385) performed the worst in the experiment, reflecting the limitations of GA in the problem of WDN restoration in this high-dimensional discrete space (including 83! restoration schemes); although the greedy search methods (M3-M4) have the advantage of computational speed (0.5-1h), they only consider the immediate reward at each step and are limited by the local optimal solution; the computational efficiency advantage (<1min) of the sorting-based methods (M5-M6) is in sharp contrast to their performance defects (0.488-0.584), and there is also a single-objective bias. Therefore, traditional methods may show certain advantages in some specific performances, but they are still inferior to M1 in terms of comprehensive indicators.

[0056] Specifically, in each time step, different methods are related to , and The restoration trajectories of are shown in Figs. 10(a), 10(b) and 10(c). In Fig. 10(a), assuming that HP reaching 0.8 is the hydraulic performance safety threshold, M1 recovers the fastest, only requiring 43 time steps, M6 is sub-optimal (44 steps), and the time consumption of M2-M5 is between 54-61 steps. In Fig. 10(b), the qualified threshold of residual chlorine concentration is usually set at 0.2 mg / L. M5 performs the best, only requiring 18 time steps to meet the standard, M1 is sub-optimal (19 steps), and the other methods are between 23-25 steps. In Fig. 10(c), at the end of the maintenance, the total transportation time of M3 is the shortest, only 47.98 min, the total transportation time of M5 is the longest, 311.74 min, and the time consumption of the other methods is between 146.44-269.99 min. This shows that the M2DRL model can achieve multi-objective collaborative optimization and has important engineering value in the daily maintenance work scenario.

[0057] IV. Stability Analysis To further verify the stability of the prediction effect of the M2DRL model, in this embodiment, by adjusting different risk thresholds different numbers of pipes that need to be maintained are obtained, named Scenario 2 and Scenario 3 respectively, which are set to 0.35 and 0.31 respectively, resulting in 52 and 72 pipes that need to be maintained respectively. The distributions of the pipes that need to be maintained in the three scenarios (all pipes covered in the training) are shown in Figs. 11(a), 11(b) and 11(c).

[0058] Use M1-M6 to restore the WDNs of Scenarios 1-3, and compare the final RI Specifically, as Figure 12The results show that the M1 model outperforms traditional decision-making methods in all degradation scenarios. Therefore, the global optimization decision of the proposed M2DRL model can effectively improve the system resilience and adapt to diverse degradation damage scenarios, showing good stability.

[0059] V. Transfer Learning Traditional meta-heuristic algorithms have high computational resource requirements and are limited to preset damage scenarios. When new damaged pipes appear in the WDN, the algorithm needs to recalculate the WDN. In this regard, M2DRL can adopt the method of transfer learning to directly fine-tune on the existing model to enhance the adaptability to new scenarios, thereby improving the computational efficiency. To verify the advantages of transfer learning, three degradation scenarios (Scenarios 4-6) were created in this embodiment, and 37, 29, and 24 pipes to be maintained were randomly selected, as shown in Figures 13(a), 13(b), and 13(c).

[0060] In Scenarios 4-6, the comparison results of different restoration methods are shown in Table 4. M1 still obtains the highest RI value in all scenarios, and as the maintenance scale increases, M1 has stronger solution-solving ability in the high-dimensional discrete space compared to other methods. RI The value is higher. In terms of computational time, taking Scenario 4 as an example, the computational time required also drops from 5.9 h (control experiment result) to 3.4 min, and the computational efficiency is significantly improved.

[0061] Table 4 Values and computational time table of different restoration methods in Scenarios 4-6 RI value and computational time

[0062] VI. Feasibility and Beneficial Effects This embodiment proposes a multi-objective resilience maintenance method for water supply networks that integrates graph convolution and reinforcement learning. Specifically, a multi-objective maintenance decision-making model (M2DRL) is built, aiming to synergistically improve the recovery degree of the WDN and shorten the recovery time by optimizing the pipeline recovery sequence, ultimately enhancing the system resilience. First, a multi-objective resilience evaluation framework including hydraulic performance, water quality performance, and transportation time is constructed to quantify the comprehensive impact of the recovery degree and recovery time. Second, the GCN is used to encode the network topology structure, and combined with the DRL framework to achieve global policy search in the high-dimensional discrete action space, and an optimal maintenance sequence is generated with the maximization of the resilience index as the goal. Taking the CA1 system in the United States as a case, the following effects can be achieved: (1) The resilience index of the recovery sequence generated by the M2DRL model is improved by 28.0% to 108.8% compared with five traditional methods, which can ensure the best index performance while effectively avoiding the local optimum trap, verifying the effectiveness of multi-objective collaborative optimization; (2) In different risk threshold scenarios 1-3, the model resilience index is higher than that of traditional methods, and the performance is stable; (3) In the new scenarios 4-6, through the sharing and fine-tuning of pre-trained model parameters in transfer learning, the calculation time is reduced from 5.9 hours to 3.4 minutes, achieving fast adaptation across scenarios. Therefore, the M2DRL model solves the problem of single-dimensionality through a multi-objective evaluation framework, and at the same time has the ability of global optimization in complex water supply network maintenance decision-making, with stable performance and fast generalization ability, which can provide an effective tool for improving the resilience of urban water supply systems.

[0063] The above schematically describes the present invention and its implementation manners. This description is not restrictive. What is shown in the drawings is only one of the implementation manners of the present invention, and the actual structure is not limited thereto. Therefore, if those of ordinary skill in the art are inspired by it and design similar structural manners and embodiments to this technical solution without creative work under the premise of not departing from the purpose of the present invention, they shall fall within the protection scope of the present invention.

Claims

1. A multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning, characterized by: The following steps are involved: S1: Build the dataset; The data set includes the topological structure of the WDN, the age, pipe material, burial depth and soil type information of each pipeline, and obtains the data of the WDN in normal state; S2: Build degradation damage module; The degradation damage module is based on the risk value of each pipeline in the WDN. , converting the WDN normal state data into degraded state data to provide data support for the subsequent recovery phase; S3: Build performance evaluation module; After the pipeline to be maintained is determined, a recovery sequence is generated based on the recovery method. The resilience improvement effect of the sequence is quantified through the performance evaluation module and the transportation time results. Specifically, in the performance evaluation module, obtaining the global performance of the WDN requires first calculating the evaluation weight and comprehensive performance of each pipeline connection point in the WDN. S4: Calculate the transportation time; S5: Calculate the degree of recovery; Degree of recovery The quantification of is usually done by the gray area method under the recovery curve. The faster the performance recovery, the larger the area under the recovery curve. the higher; S6: Calculate toughness evaluation index; Resilience evaluation indicators include recovery time and Two parts; the quantification of recovery time is completed by S4, and the quantification of recovery degree is completed by S5; S7: Calculate weights; The restoration effects of a large number of different pipeline restoration sequences are simulated, and the obtained transportation time results and restoration degree results are integrated to weight the reward values ​​using the entropy weight method; S8: Build the M2DRL model; M2DRL is a DRL model that integrates GCN. At each time step, the agent obtains the current state based on the graph representation of WDN. , perform the action After that, the environmental state evolves to , and receive instant rewards Used for strategy optimization; the implementation of the model first requires building the WDN graph structure as the input of the agent, and then clearly defining the state, action and reward value during the interaction process; S9: training model; S10: Generate pipeline recovery sequences through the trained model and evaluate them based on resilience evaluation indicators.

2. According to claim 1, a multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning is characterized by: The specific steps of S2 are as follows: S21: To determine the , we need to quantify the failure probability first and consequences of failure ;in, By comparing the maximum stress on the pipe over time Based on the tensile strength of the pipe material To determine, the formula is: ; In the formula, is the service life of the pipeline, is the pipe diameter; For traffic load; is the elastic modulus of soil; is the elastic modulus of the pipe; is the unit soil weight; The buried depth of the pipeline; is the time-varying wall thickness of the pipe, is the water pressure inside the pipe; is the lateral earth pressure coefficient; S22: Calculation pipeline using fuzzy inference system ; S23: calculated and , get the value of each pipeline ; The quantification formula is: ; S24: Calculate the impact of pipeline degradation on performance; By constructing leakage points, the leakage flow is used to simulate the degree of pipeline deterioration, thereby affecting the hydraulic and water quality changes at the connection point. The formula is: ; In the formula, is the leakage flow; is the leakage area; is the flow coefficient; WNTRSimulator and EpanetSimulator provided by the open source Python package WNTR are used to simulate the hydraulic changes and water quality changes after adding the leakage point, so as to obtain the water pressure value and water quality calculation coefficient after the leakage, which are used for the performance evaluation model; S25: Screening pipelines to be maintained; Assume that when the pipeline Exceeds a given threshold , then take maintenance measures and put this part of the pipeline into the pipeline collection to be maintained middle.

3. According to claim 2, a multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning is characterized by: The specific steps of S3 are as follows: S31: Calculate evaluation weight; Connect the dots Evaluation weight Indicates the importance of the point in the WDN, based on the consequences of its failure OK, the formula is: ; In the formula, For connection point Consequences of failure; is the total number of connection points; By connecting the points Connected The average value of the failure consequence of the root pipeline is determined by the formula: ; In the formula, For pipeline Consequences of failure; The value range is (0,1]; S32: Calculate hydraulic performance; In the pressure-driven hydraulic model, hydraulic performance Defined as the ability of the connection point to meet the water needs of users, quantified by pressure adequacy; When , the water pressure meets the design water supply demand, the formula is: ; In the formula, For connection point exist Hydraulic performance at all times; For connection point Water pressure; The water pressure for water supply stop; Minimum water pressure to meet water supply demand; S33: Calculate water quality performance; In WDN, maintaining an appropriate amount of residual chlorine is a key criterion for ensuring the safety of drinking water in the water supply system. The residual chlorine concentration is a water quality performance indicator of WDN. Evaluation index; In WDN, the decay process of residual chlorine concentration conforms to the first-order reaction kinetics, and the general model is: ; In the formula, For connection point exist Residual chlorine concentration at the time; For connection point Initial chlorine concentration; For time; is the total chlorine attenuation coefficient; S34: Calculate global performance; The WDN is calculated using the evaluation weights, hydraulic performance and water quality performance of each connection point. Global performance at the time, including global hydraulic performance HP and global water quality performance , the formula is: 。 4. According to claim 3, a multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning is characterized by: In S4, the recovery time is measured by the sum of the transportation time between pipelines during maintenance work, and the formula is: ; In the formula, H is the total transportation time; for The transit time of the moment; is the total step length of the recovery time, and its value is the number of pipelines to be maintained; is the decision variable, indicating that at time For pipeline The value is 1 when maintenance is in progress, otherwise it is 0; For pipeline and pipeline The transportation time between is the set of pipelines to be maintained; among them, calculation It is necessary to create a start-point-destination matrix based on OpenStreetMap and ArcGIS.

5. According to claim 4, a multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning is characterized in that: The specific calculation steps of S5 are as follows: Based on global performance Quantitative diagram to quantify the degree of hydraulic recovery and water quality recovery , the formula is: ; In the formula, and They represent the start and end time of the recovery process respectively, and each time step in between represents a pipeline being maintained.

6. According to claim 5, a multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning is characterized by: The specific calculation steps of S6 are as follows: Get the recovery time and degree of recovery, and calculate the resilience index , transformed into a multi-objective optimization problem, the formula is: ; In the formula, and Respectively represent the normalized , and , corresponding to the weights and , the weight value is calculated by the entropy weight method; Used to evaluate the effects of different pipeline recovery sequences. The larger the value, the stronger the resilience improvement ability of the corresponding pipeline recovery sequence.

7. According to claim 6, a multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning is characterized by: The specific steps of S8 are as follows: S81: Create graph structure; The WDN graph structure is represented as ,in is a node set, each node Represents a pipeline; is an edge set, each edge Indicates the connection relationship between two pipelines; The dimension is The adjacency matrix of , indicating a node and Is there a connection between them? ; S82: define status; Each node in the graph Corresponding to a feature vector ,in, is the pipeline ID; is the length of the pipeline; is the pipe diameter; is the pipe material; is the pipeline risk value; is the pipeline hydraulic performance; For pipeline water quality performance; is the pipeline maintenance status, the waiting maintenance value is 0, and the maintained value is 1; at each time step, and Will be updated based on maintenance actions taken; S83: define actions; Based on the state of WDN, the agent Always select the pipelines that meet the maintenance conditions Select a pipeline Maintenance is performed, represented by , thus forming the current action At the same time, the maintenance actions at all times need to be equal to the set of pipelines to be maintained , the formula is: ; In the formula, It indicates the action at the first moment; Indicates the action at the last moment; S84: Set rewards; Reward Value Resilience Index As the only optimization goal, that is: ; In the formula, For immediate rewards.

8. According to claim 7, a multi-objective resilience maintenance method for water supply network integrating graph convolutional reinforcement learning is characterized by: The specific steps for S9 are as follows: S91: Initialization parameters: Set the initial number of steps , current training cycle , initial value of exploration rate , exploration rate decay coefficient and the total number of training cycles ; Identify the state of WDN , the maintenance status of all pipelines Set to 0, time step ; S92: Calculate dynamic exploration rate: Update the current exploration rate through the decay formula , which makes the agent frequently try different maintenance methods in the early stage, and gradually reduce randomness in the later stage to focus on the optimal strategy; S93: Action selection: Adopt The greedy strategy selects maintenance actions with probability Select pipelines randomly, or with probability According to the agent selection The pipe with the largest value; S94: Status update and reward calculation: Modify to 1 and calculate the new state , and calculate the immediate reward by formula (19) ; S95: Value update and network training: The model combines The biggest future value , and the current Compare values ​​to update Network parameters; as training progresses, the agent gradually becomes able to determine the true nature of the action value, achieving autonomous learning of WDN maintenance decisions; S96: Cycle control: When all pipelines are maintained, a cycle is completed. , reset the WDN state and start a new round of training. When the number of training times reaches The training ends when.

Citation Information

Patent Citations

  • Method for calculating cascade robustness of water supply pipe network of dynamic emergency recovery mechanism

    CN107066666A

  • Water supply pipe network optimization method and system based on ordered binary decision diagram

    CN114676536A

  • Urban water supply pipe network anti-seismic toughness comprehensive calculation method considering water quality and water quantity indexes

    CN116011232A

  • Systems and methods to facilitate decision making for utility networks

    US20240112150A1

  • Method and System for Generating a Resilience Analysis of a Real-world System

    US20250077746A1

Cited By

  • HPLC (high performance liquid chromatography) mobile phase gradient intelligent proportioning online purification system

    CN121049436A

  • HPLC mobile phase gradient intelligent proportioning online purification system

    CN121049436B