Heating pipe network diameter optimization method and medium based on graph attention and reinforcement learning

CN122818883APending Publication Date: 2026-09-25HUADIAN ZHENGZHOU MECHANICAL DESIGN INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610692693.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,当前的热网管径设计参照工程指南和设计手册进行,往往只涉及简单的估算和查表,无法在建设成本、运行成本、供需匹配、速度限制等多维目标下实现全局优化

Benefits of technology

[0019]由上述技术方案可知,本发明的基于图注意力与强化学习的供热管网管径优化方法,利用图注意力网络提取热网的高维非线性特征,利用强化学习实现不同规模热网案例的快速和泛化求解。该方法展现了强大的搜索-学习-泛化能力,能够在陌生的案例下快速得到全面超越现有方法的优质解,拓宽了强化学习在能源领域的应用范围。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122818883A_ABST
    Figure CN122818883A_ABST
Patent Text Reader

Abstract

The application discloses a heat supply pipe network pipe diameter optimization method and medium based on graph attention and reinforcement learning, converts a pipe diameter optimization problem into a Markov decision process, defines an agent state containing node features, pipe features and graph structure features, takes single pipe pipe diameter selection as an action, and adopts a mask mechanism to constrain a decision sequence; a heat network hydraulic model is built based on a lumped parameter method, a reinforcement learning environment containing initialization, state transition and a composite reward function is constructed; a strategy network and a value network are built by using a graph attention network and a full connection neural network to form a reinforcement learning agent; heat network instance data enhancement is realized through topological transformation, pipe length sampling and load sampling, and an agent is trained by using a PPO algorithm; and a to-be-decided case is input into the trained agent to obtain a globally optimal pipe diameter combination through pipe-by-pipe decision-making. The application can significantly reduce heat supply pipe network construction and operation costs, improve load satisfaction, and avoid flow rate over-limiting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of urban heating technology, specifically to a method and medium for optimizing the diameter of heating pipeline networks based on graph attention and reinforcement learning. Background Technology

[0002] Heating in northern cities is a major issue affecting people's livelihoods, and the pipe diameter of the heating network is a key factor determining the construction and operation costs of the urban heating system, the energy consumption of the heating network, and the satisfaction of residents with public services. Optimizing the design of the heating network pipe diameter can not only reduce investment and operating costs for enterprises but also improve heating quality and increase user satisfaction. However, current heating network pipe diameter designs often rely on engineering guidelines and design manuals, involving only simple estimations and table lookups, failing to achieve global optimization across multiple dimensions such as construction costs, operating costs, supply and demand matching, and speed limitations. Summary of the Invention

[0003] The present invention proposes a method for optimizing the pipe diameter of heating pipe networks based on graph attention and reinforcement learning, which can at least solve one of the technical problems in the background art.

[0004] To achieve the above objectives, the present invention adopts the following technical solution: A method for optimizing the pipe diameter of a heating network based on graph attention and reinforcement learning is implemented through computer equipment. The method comprises the following steps: S1, modeling the pipe diameter optimization problem of the heating network as a Markov decision process, defining the state and action of the reinforcement learning agent: the state is composed of node features, pipe features and graph structure features; the action is to select the pipe diameter for a single pipe; and setting a masking mechanism to ensure that each pipe is selected only once in the same decision round. S2. A hydraulic simulation module for a heating network is built based on the lumped parameter method. The network topology and heat source flow rate are used as inputs, and the pressure and flow rate of each node are output. A reinforcement learning environment is built based on this hydraulic model, including environment initialization, state transition and reward function. The reward function comprehensively considers construction cost, operating cost, load non-compliance penalty and flow velocity over-limit penalty. S3. Construct a reinforcement learning agent that includes a policy network and a value network: The policy network consists of a graph attention network, a fully connected layer, and a softmax layer, and outputs the probability distribution of pipe actions; the value network consists of a graph attention network, a fully connected layer, and an additive pooling layer, and outputs the state value. S4. Based on real heating network cases, multiple sets of heating network instances are generated through topology transformation, pipe length sampling, and load sampling to enhance training data; S5. The PPO algorithm is used to enable the agent to interact with the environment in the enhanced hot network instance for training, and to complete the convergence of the policy network and the value network. S6. Input the pipeline network to be optimized into the trained agent, make decisions on the pipe diameter for each pipeline, and obtain the global optimal pipe diameter combination scheme.

[0005] In the above technical solution, step S1 further comprises: S11, the state s of the reinforcement learning agent consists of node features X, pipeline features E, and graph structure features Edge_index.

[0006]

[0007] The node characteristics include: node type, node pressure, distance from the heat source node, total degree, total downstream load, and downstream load satisfaction rate; the pipe characteristics include: pipe length, pipe diameter, flow rate, flow velocity, number of downstream nodes, flow rate percentage, average pipe diameter of adjacent pipes, downstream load, and downstream load satisfaction rate; the graph structure characteristics are represented by an edge list, which is used to describe the first and last nodes of each pipe.

[0008] S12, each action of the reinforcement learning agent is to select the diameter of a pipe. Assuming there are p possible diameters, each action should satisfy:

[0009] in, a Refers to actions, d1,..., d p Refers to all options for pipe diameter; S13. Set an appropriate masking mechanism so that once the reinforcement learning agent selects a pipe, that pipe cannot be changed again in the same round, thus making the number of actions in a round equal to the number of pipes.

[0010] In the above technical solution, step S2 further comprises: S21. Construct a hydraulic simulation module for the heating network. This module is based on the lumped parameter method, which treats the hydraulic transmission process as the combined action of water resistance elements and water pressure source elements. The inputs to the hydraulic simulation module are the heating network topology information and the heat source flow rate, and the outputs are the flow rate and pressure of each node.

[0011] S22, Based on the above hydraulic simulation module, a reinforcement learning environment is constructed, including an environment initialization function, an environment transition function, and a reward function.

[0012] S23, the environment initialization function takes the heat network topology information and heat source flow rate as input, randomly initializes a pipe diameter combination for the current heat network instance, and solves the flow rate and pressure information of the current pipe diameter combination through the above hydraulic simulation module.

[0013] S24, the environment transfer function generates the state under the latest pipe diameter combination by receiving the agent's action, namely the pipe diameter selection.

[0014] S25, the reward function consists of three parts: building cost CC, operating cost OC, load non-compliance penalty LUP, and overspeed penalty VVP, as shown in the following formula:

[0015] In the formula: The cur subscript refers to the indicators corresponding to the current pipe diameter combination; The prev subscript refers to the indicators corresponding to the previous pipe diameter combination; W, a, b are coefficients; m is the number of pipes in the heating network; vi is the water flow velocity in the i-th pipe; k is the number of load nodes in the heating network; Preal,i is the actual heating load of the i-th pipe; Pneed,i is the required heating load of the i-th tube; Furthermore, step S3 specifically includes: S31, a policy network is constructed using a graph attention network module, a fully connected neural network module, and a softmax function; S32, a value network is constructed using a graph attention network module, a fully connected neural network module, and a summation pooling module; S33, Construct a reinforcement learning agent based on the above policy network and value network.

[0016] Furthermore, step S4 specifically includes: S41, select multiple real-world hot network cases of different sizes to ensure the diversity of training and testing datasets; S42 achieves data enhancement by changing the topology of the heating network through swapping the positions of heat source nodes and load nodes and sampling pipe lengths, and by changing the operating mode of the heating network through load sampling.

[0017] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.

[0018] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.

[0019] As can be seen from the above technical solution, the heating network pipe diameter optimization method based on graph attention and reinforcement learning of the present invention utilizes graph attention networks to extract high-dimensional nonlinear features of the heating network and uses reinforcement learning to achieve fast and generalized solutions for heating network cases of different scales. This method demonstrates powerful search-learning-generalization capabilities, and can quickly obtain high-quality solutions that comprehensively surpass existing methods even in unfamiliar cases, thus broadening the application scope of reinforcement learning in the energy field.

[0020] Specifically, this invention first determines the state and actions of the reinforcement learning agent, thereby correctly transforming the pipe diameter optimization problem into a Markov decision process. Then, a hydraulic model of the heating network is built to construct the reinforcement learning environment, defining the environment initialization function, environment transition function, and reward function. Based on the characteristics of heating network pipe diameter optimization, a graph attention network and a fully connected neural network are used to construct the agent. Multiple effective heating network training instances are generated based on actual heating network cases to achieve data augmentation for the agent. The agent interacts with the environment in the generated heating network training instances to continuously improve its performance. Finally, the cases to be decided are input into the trained agent, which can make decisions pipe by pipe, ultimately obtaining the globally optimal pipe diameter combination scheme. The ultimately trained policy network can utilize the learned experience to demonstrate significantly better performance than traditional pipe diameter design methods in new heating network case test sets. The obtained pipe diameter configuration scheme reduces economic costs by an average of 31.4% and improves load satisfaction by 47.1% compared to traditional methods, achieving both improved economic efficiency and increased resident heating satisfaction.

[0021] In summary, this invention provides a heat network pipe diameter optimization planning method based on reinforcement learning. This method addresses the numerous complex features with nonlinear relationships in the heat network by adding graph neural network modules to the policy network and value network, and utilizes a reinforcement learning framework to train the policy network and value network. The ultimately trained policy network leverages learned experience to demonstrate significantly better performance than traditional pipe diameter design methods in new heat network case test sets. The resulting pipe diameter configuration scheme reduces economic costs by an average of 31.4% and improves load satisfaction by 47.1% compared to traditional methods, while completely avoiding excessively high flow velocities. This achieves both improved economic efficiency and increased resident heating satisfaction. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of a reinforcement learning agent network structure according to an embodiment of the present invention; Figure 2 This is a schematic diagram illustrating the generation of a heating network instance according to an embodiment of the present invention; Figure 3 This is a schematic diagram of a reinforcement learning training framework according to an embodiment of the present invention; Figure 4The flowchart illustrates a method for optimizing pipe diameter in urban heating systems that integrates graph attention mechanism and reinforcement learning, as provided in this invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments.

[0024] like Figure 4 As shown in the figure, the heating network pipe diameter optimization method based on graph attention and reinforcement learning described in this embodiment is implemented by a computer device through the following steps: S1 determines the state and actions of the reinforcement learning agent, thereby transforming the pipe diameter optimization problem into a Markov decision process, which can then be solved using a reinforcement learning framework. The state includes node features, pipe features, and graph structure features; the action is the candidate pipe diameter.

[0025] S2, a hydraulic model of the heating network is built to realize the construction of the reinforcement learning environment; the hydraulic model takes the heating network topology and user load information as input, and can output the pressure and flow rate of each node; the reinforcement learning environment includes an environment initialization function, an environment transition function, and a reward function; S3 utilizes graph attention networks and fully connected neural networks to construct an intelligent agent, while the state and actions determined by S1 determine the input and output dimensions of the network.

[0026] S4, based on actual heating network cases, generates multiple effective heating network instances through three methods: topology transformation, pipe length sampling, and load sampling, thereby enhancing the training data.

[0027] S5 trains the agent built in S3 to interact with the environment in the heat network instance generated in S4, and improves and converges the agent's policy network and value network based on the PPO algorithm.

[0028] S6 uses the case of the hot network to be decided to build a new environment, allowing the agent trained in S5 to interact with the new environment, and finally reasoning to obtain the globally optimal pipe diameter combination scheme.

[0029] The following is a detailed explanation: Step S1: Determine the state of the reinforcement learning agent. s ,action a .

[0030] Further, in step S11, the state of the reinforcement learning agent... s It consists of node features X, pipeline features E, and graph structure features Edge_index, such as Figure 2As shown.

[0031]

[0032] The node characteristics include: node type, node pressure, distance from the heat source node, total degree, total downstream load, and downstream load satisfaction rate; the pipe characteristics include: pipe length, pipe diameter, flow rate, flow velocity, number of downstream nodes, flow rate percentage, average pipe diameter of adjacent pipes, downstream load, and downstream load satisfaction rate; the graph structure characteristics are represented by an edge list, which is used to describe the first and last nodes of each pipe.

[0033] Further, in step S12, each action of the reinforcement learning agent is to select the diameter of a pipe. There are a total of 8 pipe diameter options, and each action should satisfy:

[0034] Further, in step S13, a suitable masking mechanism is set so that once the reinforcement learning agent selects a pipe, that pipe cannot be changed again in the same round, thus making the number of actions in one round equal to the number of pipes.

[0035] Step S2: Build a hydraulic model of the heating network to realize the construction of the reinforcement learning environment. Further, in step S21, a hydraulic simulation module for the heating network is constructed. This module is constructed based on the lumped parameter method, which treats the hydraulic transmission process as the combined action of water resistance elements and water pressure source elements.

[0036] Further, in step S22, a reinforcement learning environment is constructed based on the above-mentioned hydraulic simulation module, including an environment initialization function, an environment transfer function, and a reward function.

[0037] Further, in step S23, the environment initialization function takes the input of the heating network topology information and the heat source flow rate, and randomly initializes a pipe diameter combination for the current heating network instance, and solves the flow rate and pressure information of the current pipe diameter combination through the above-mentioned hydraulic simulation module.

[0038] Further, in step S24, the environment transfer function generates the state under the latest pipe diameter combination by receiving the agent's action, namely pipe diameter selection.

[0039] Further, in step S25, the reward function is determined by the construction cost. CC Operating costs OC Penalty for unmet load requirements LUP Speeding penalties VVP It consists of three parts, and the formula is as follows:

[0040] In the formula: curThe subscript refers to the indicators corresponding to the current pipe diameter combination; prev The subscript refers to the indicators corresponding to the previous pipe diameter combination; W, a, b are coefficients, where W is 1e6, a is 2e7, and b is 1e7; m This refers to the number of pipes in the heating network; v i It is the first i Water flow rate in the root canal; k This refers to the number of load nodes in the heating network; P real,i It is the first i The actual heating load of the root canal; P need,i It is the first i The heating load required for root canals; The formulas for calculating construction costs and operating costs are as follows:

[0041] e () is a function representing the construction cost per unit length of water pipe. ; d It is the pipe diameter; L i It is the first i The length of the root canal; P pump It is the operating power of the water pump; C ele This is the unit price of electricity; T This refers to the pump running time; S4, Generating Heating Network Instances Based on a Real-World Heating Network Case. This case study generates 35 training cases and 7 test cases based on a 24-node primary heating network in southern China. The generation method is as follows: Figure 3 As shown.

[0042] In S5, the agent constructed in S3 is trained to interact with the environment in the heat network instance generated in S4, and the agent's policy network and value network are improved and converged based on the PPO algorithm. Table 1 shows the hyperparameter settings of the PPO algorithm.

[0043] Table 1. Hyperparameter settings for the PPO algorithm

[0044] In step S6, a new environment is constructed using the case of the heating network to be decided, allowing the agent trained in step S5 to interact with the new environment, ultimately obtaining the globally optimal pipe diameter combination scheme. Table 2 shows a comparison of the target values ​​obtained by the proposed method and the traditional method. It can be seen that the proposed method is inferior to the traditional method in terms of economic cost, load non-compliance, and overspeed, demonstrating the superiority of the proposed method.

[0045] Table 2 Reward values ​​for different actions

[0046] In another aspect, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method described above.

[0047] In another aspect, the present invention also discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the method described above.

[0048] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute any of the heating pipe network diameter optimization methods based on graph attention and reinforcement learning in the above embodiments.

[0049] It is understood that the systems, devices, and storage media provided in the embodiments of the present invention correspond to the methods provided in the embodiments of the present invention, and the explanations, examples, and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.

[0050] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).

[0051] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0052] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0053] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for optimizing the diameter of heating pipe networks based on graph attention and reinforcement learning, characterized in that, Includes the following steps, S1. Model the pipe diameter optimization problem of the heating network as a Markov decision process, and define the state and action of the reinforcement learning agent: the state is composed of node features, pipe features and graph structure features; the action is to select the pipe diameter for a single pipe; set a masking mechanism so that each pipe is selected only once in the same decision round. S2. A hydraulic simulation module for a heating network is built based on the lumped parameter method. The network topology and heat source flow rate are used as inputs, and the pressure and flow rate of each node are output. A reinforcement learning environment is built based on this hydraulic model, including environment initialization, state transition and reward function. The reward function comprehensively considers construction cost, operating cost, load non-compliance penalty and flow velocity over-limit penalty. S3. Construct a reinforcement learning agent that includes a policy network and a value network: The policy network consists of a graph attention network, a fully connected layer, and a softmax layer, and outputs the probability distribution of pipe actions; the value network consists of a graph attention network, a fully connected layer, and an additive pooling layer, and outputs the state value. S4. Based on real heating network cases, multiple sets of heating network instances are generated through topology transformation, pipe length sampling, and load sampling to enhance training data; S5. The PPO algorithm is used to enable the agent to interact with the environment in the enhanced hot network instance for training, and to complete the convergence of the policy network and the value network. S6. Input the pipeline network to be optimized into the trained agent, make decisions on the pipe diameter for each pipeline, and obtain the global optimal pipe diameter combination scheme.

2. The method for optimizing the diameter of heating pipe networks based on graph attention and reinforcement learning according to claim 1, characterized in that: S1 specifically includes, S11, The state of the reinforcement learning agent s It consists of node features X, pipeline features E, and graph structure features Edge_index; The node characteristics include: node type, node pressure, distance from the heat source node, total degree, total downstream load, and downstream load satisfaction rate. Pipeline characteristics include: pipeline length, pipeline diameter, flow rate, flow velocity, number of downstream nodes, flow rate percentage, average diameter of adjacent pipelines, downstream load, and downstream load fulfillment rate. The graph structure features are represented by a list of edges, which describes the first and last nodes of each pipe. S12. Each action of the reinforcement learning agent is like selecting the diameter of a pipe. Assume there are a total of... p Given a choice, each action should satisfy: in, a Refers to actions, d1,..., d p Refers to all options for pipe diameter; S13. Set up a masking mechanism so that once the reinforcement learning agent selects a pipe, that pipe cannot be changed again in the same round, thus making the number of actions in a round equal to the number of pipes.

3. The method for optimizing the diameter of heating pipe networks based on graph attention and reinforcement learning according to claim 1, characterized in that: S2 specifically includes, S21. Construct a hydraulic simulation module for the heating network. This module is based on the lumped parameter method, which treats the hydraulic transmission process as the combined action of water resistance elements and water pressure source elements. The input of the hydraulic simulation module is the heating network topology information and heat source flow rate, and the output is the flow rate and pressure of each node. S22, Based on the above hydraulic simulation module, a reinforcement learning environment is constructed, including an environment initialization function, an environment transition function, and a reward function; S23, the environment initialization function takes the input of the heating network topology information and heat source flow rate, and randomly initializes a pipe diameter combination for the current heating network instance, and solves the flow rate and pressure information of the current pipe diameter combination through the above hydraulic simulation module; S24, the environment transfer function generates the state under the latest pipe diameter combination by receiving the agent's action, namely pipe diameter selection; S25, the reward function is determined by construction cost. CC Operating costs OC Penalty for unmet load requirements LUP Speeding penalties VVP It consists of three parts, and the formula is as follows: In the formula: cur The subscript refers to the indicators corresponding to the current pipe diameter combination; prev The subscript refers to the indicators corresponding to the previous pipe diameter combination; W , a , b It is a coefficient; m This refers to the number of pipes in the heating network; v i It is the first i Water flow rate in the root canal; k This refers to the number of load nodes in the heating network; P real,i It is the first i The actual heating load of the root canal; P need,i It is the first i The heating load required for root canal treatment.

4. The method for optimizing the diameter of heating pipe networks based on graph attention and reinforcement learning according to claim 1, characterized in that: S3 specifically includes, S31, a policy network is constructed using a graph attention network module, a fully connected neural network module, and a softmax function; S32, a value network is constructed using a graph attention network module, a fully connected neural network module, and a summation pooling module; S33, Construct a reinforcement learning agent based on the above policy network and value network.

5. The method for optimizing the diameter of heating pipe networks based on graph attention and reinforcement learning according to claim 4, characterized in that: S4 specifically refers to: S41, select multiple real-world hot network cases of different sizes to ensure the diversity of training and testing datasets; S42 achieves data enhancement by changing the topology of the heating network through swapping the positions of heat source nodes and load nodes and sampling pipe lengths, and by changing the operating mode of the heating network through load sampling.

6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor performs the steps of the method as described in any one of claims 1 to 5.