Optimal dispatching method for electric heating gas comprehensive energy system, terminal device and storage medium
Patent Information
- Application Number
- CN202310332234.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2043-03-30
AI Technical Summary
这些求解工作大多集中在近似求解、非线性求解等方面,不可避免地要面对算法复杂度高、系统状态变化需重新求解的问题,在面向大规模系统时那以实现快速相应,且无法保证达到全局最优解
[0019] This invention adopts the above technical solution and proposes a SAC reinforcement learning optimization scheduling algorithm based on a GNN architecture. Compared with methods based on an MLP architecture, the utilization of system topology information leads to a faster convergence speed, making it more advantageous in the optimization scheduling of integrated energy systems (electricity, heat, and gas).
Smart Images

Figure CN116451947B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy optimization scheduling, and in particular to a method, terminal equipment and storage medium for optimizing scheduling of integrated electric, heat and gas energy systems. Background Technology
[0002] Optimal scheduling of integrated electric, thermal, and gas energy systems is one of the core problems in multi-energy flow analysis, and it is of great significance for energy conservation, emission reduction, and the full utilization of energy. Current research mainly focuses on optimal power flow solutions for electric, thermal, or electrical energy systems. In recent years, there has also been a surge in modeling and solving electric, thermal, and gas systems. These solutions mostly focus on approximate and nonlinear solutions, inevitably facing problems such as high algorithm complexity and the need to recalculate solutions when system states change. They also struggle to achieve rapid response for large-scale systems and cannot guarantee a globally optimal solution. Furthermore, the widespread deployment of new energy power plants and the uncertainty of their output bring new challenges to optimization solutions. Summary of the Invention
[0003] To address the aforementioned problems, this invention proposes an optimized scheduling method, terminal equipment, and storage medium for an integrated electric, thermal, and gas energy system.
[0004] The specific plan is as follows:
[0005] A method for optimizing the scheduling of an integrated electric, thermal, and gas energy system includes the following steps:
[0006] S1: The integrated energy system of electricity, heat, and gas is abstractly modeled as a state diagram. In the power system, the node characteristics of the equipment are represented by the electrical load of the equipment, and the edge characteristics are represented by the susceptance and conductance between the two equipment corresponding to the two nodes. The node characteristics of the heat system are represented by the heat load of the equipment, and the edge characteristics are represented by the length of the pipeline branch between the two equipment corresponding to the two nodes and the pipeline mass flow rate. The node characteristics of the natural gas system are represented by the gas load of the equipment, and the variable characteristics are represented by the pipeline length between the two equipment corresponding to the two nodes and the pipeline constant.
[0007] S2: Collect historical state diagrams of the integrated electric, heat, and gas energy system to form a training set;
[0008] S3: Construct a maximum entropy reinforcement learning model based on graph neural network for the optimal scheduling of integrated electric, thermal and gas energy systems. The model changes the original multilayer perceptron network to a graph neural network, and sets the action space, state space and reward function of the model based on the integrated electric, thermal and gas energy system. The model is trained using a training set.
[0009] S4: Obtain the output results of the integrated electric, thermal, and gas energy system through the trained model.
[0010] Furthermore, in the action space of the reinforcement learning model, the active power output of the thermal power plant, the electrical output of the CHP unit, the thermal power output of the CHP unit, the absorption coefficient of the wind power plant, and the gas supply of the gas supply station are used as action variables, and the value range of each action variable is set.
[0011] Furthermore, the system state in the state space of the reinforcement learning model is represented by the node features and edge features of the graph.
[0012] Furthermore, the formula for calculating the reward function of a reinforcement learning model is as follows:
[0013]
[0014] Where, r t F represents the reward value at time t. t To represent the operating cost at time t, λ i Let |L| represent the penalty factor corresponding to the i-th constraint, where i represents the constraint number. i | represents the absolute value of the difference between the establishment of the i-th constraint and the establishment of the constraint. When the constraint is established, |L i | is 0; when the constraint is not met, |L i | represents the minimum absolute value of the difference between the boundary conditions and the boundary conditions.
[0015] Furthermore, the operating cost of the integrated electric, heat, and gas energy system is the sum of the operating costs of the thermal power plant, the operating costs of the CHP unit, and the cost of supplying natural gas to the gas supply station.
[0016] Furthermore, the reinforcement learning model uses an attention mechanism in its graph neural network to aggregate node information and obtain node representations.
[0017] An optimized scheduling terminal device for an integrated electric, thermal, and gas energy system includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in the embodiments of the present invention.
[0018] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described above in the embodiments of the present invention.
[0019] This invention adopts the above technical solution and proposes a SAC reinforcement learning optimization scheduling algorithm based on a GNN architecture. Compared with methods based on an MLP architecture, the utilization of system topology information leads to a faster convergence speed, making it more advantageous in the optimization scheduling of integrated energy systems (electricity, heat, and gas). Attached Figure Description
[0020] Figure 1The diagram shown is a flowchart of Embodiment 1 of the present invention.
[0021] Figure 2 The diagram shows a pipeline structure of a compressor driven by a gas turbine in this embodiment.
[0022] Figure 3 The diagram shown is a schematic of the algorithm framework of the reinforcement learning model in this embodiment.
[0023] Figure 4 The diagram shown is a schematic of the Actor network structure of the graph neural network and the multilayer perceptron in this embodiment.
[0024] Figure 5 The diagram shown is a structural schematic of the integrated electric, thermal, and gas energy system in this embodiment.
[0025] Figure 6 The diagram shown illustrates the comparison results between GNN and MLP in this embodiment.
[0026] Figure 7 The diagram shown is a schematic representation of the output results of the simulated experimental system in this embodiment. Detailed Implementation
[0027] To further illustrate the various embodiments, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, primarily used to illustrate the embodiments, and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementations and the advantages of the present invention.
[0028] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.
[0029] Example 1:
[0030] This invention provides an optimized scheduling method for an integrated electric, thermal, and gas energy system, such as... Figure 1 As shown, the method includes the following steps:
[0031] S1: The combined electric, heat, and gas energy system is abstractly modeled as a state diagram.
[0032] (I) Environmental Model
[0033] The environmental model is an integrated energy system model of electricity, heat and gas that interacts with intelligent agents, and includes an electric system, a thermal system, a natural gas system and a coupling system.
[0034] 1. Power System Model
[0035] The power flow equations for an AC power system are:
[0036]
[0037] In the above formula, P i Q i U represents the active power and reactive power injected at node i, respectively. i G represents the voltage magnitude at node i. ij B represents the conductance between node i and node j. ij θ represents the susceptance between node i and node j. ij =θ i -θ j This represents the phase angle difference between node i and node j.
[0038] 2. Thermodynamic System Model
[0039] 1) Hydraulic model
[0040] The hydraulic model consists of the flow continuity equation and the loop pressure equation.
[0041]
[0042] In the above formula, A is the node-branch correlation matrix, m is the pipeline mass flow rate vector, and m q Inject flow vectors into nodes, where B is the loop-branch correlation matrix, and h f This represents the head loss vector, which is related to the pipe's damping coefficient and mass flow rate.
[0043] 2) Thermodynamic model
[0044] The thermal model includes nodal power model, pipe temperature drop equation, and medium mixing equation.
[0045]
[0046] In the above formula, H i C represents the thermal power of node i. p m is the specific heat capacity of water. q,i T represents the injected traffic of node ... in i ,T i out T represents the supply water temperature and return water temperature of node i, respectively. j,i T represents the water temperature at end j of pipe branch i→j. i,j T represents the water temperature at end i. e Where L is the ambient temperature, λ is the thermal conductivity, and L is the external temperature. ij Let m represent the length of pipe branch ij. For the medium mixing equation, m k,i T represents the mass flow rate from node k to node i. k Let |n| represent the water temperature at node k when it flows to node i. i | represents the total number of nodes that flow to node i.
[0047] 3. Natural Gas System Model
[0048] Natural gas system pipeline flow rate f ij With node pressure p i The relationship between them is:
[0049]
[0050] In the above formula, κ ij This is the pipeline constant.
[0051] Because natural gas requires a certain number of compressors to ensure stable transportation during transmission, a pipeline model that includes compressors, such as... Figure 2 As shown.
[0052] For pipelines containing gas compressors, the model is as follows:
[0053]
[0054] In the above formula, f c p represents the gas consumption of the gas turbine. j ,p i ,p o ,p k These represent the pressures at the four nodes in the diagram, k. c =p o / p i f is the compression ratio. io T is the air flow rate through the compressor. gas Let q be the temperature of the natural gas. gas α represents the calorific value of natural gas, and α is the polytropic index.
[0055] 4. Coupling links in the electric heating gas system
[0056] For an integrated power, heat, and gas energy system, this embodiment considers a gas-fired CHP unit capable of simultaneously generating electricity and providing heat to meet the load demands of both the power and heating systems. The polygonal output model corresponding to the CHP unit is as follows:
[0057]
[0058] In the above formula, P CHP H CHP These represent the electrical output and thermal output of the CHP unit, respectively. These represent the upper and lower limits of the electrical output power of the CHP unit, respectively. α1, α2, and α3 represent the upper and lower limits of the thermal output power of the CHP unit, respectively, and are polygonal region coefficients.
[0059] (II) Objective Function
[0060] The goal of optimizing the scheduling of an integrated energy system combining electricity, heat, and gas is to minimize operating costs without violating constraints.
[0061] 1) Operating costs of thermal power plants
[0062]
[0063] In the above formula, T represents the total running time, and |N P | P represents the number of thermal power plants. i,t Let αi represent the active power output of thermal power plant i at time t, and α0, α1, α2 be the parameters of the consumption characteristic curve of the thermal power unit.
[0064] 2) Operating costs of CHP units
[0065]
[0066] In the above formula, |N CHP | represents the number of CHP units, P i CHP H i CHP Let μ0, μ1, μ2, μ3, μ4, and μ5 represent the electrical and thermal output of CHP unit i at time t, respectively, and μ0, μ1, μ2, μ3, μ4, and μ5 are the consumption characteristic curve parameters of CHP unit.
[0067] 3) Natural gas cost
[0068]
[0069] In the above formula, |N G |C represents the number of gas supply stations. gas f represents the unit price of natural gas. i,t This represents the gas supply volume at time t of the gas supply station.
[0070] 4) Objective function
[0071] minF t =F 1,t +F 2,t +F 3,t (10)
[0072] In the above formula, F t This represents the total operating cost of the integrated electric, thermal, and gas energy system at time t.
[0073] (III) Constraints
[0074] 1) Safety constraints
[0075] The safety constraints that must be met for the stable operation of an integrated electric, thermal, and gas energy system include: voltage constraints, phase angle difference constraints, and line transmission constraints for the power system; node temperature and pipeline flow constraints for the thermal system; and node pressure constraints for the natural gas system.
[0076]
[0077] In the above formula, U i,min U i,max These are the upper and lower limits of the voltage amplitude at node i, respectively. P is the upper limit of the phase angle difference. l T is the upper limit of line transmission power. i,max ,T i,min Let m be the upper and lower limits of the water supply temperature at node i, respectively. ij,max ,m ij,min Let p be the upper and lower limits of the water supply flow rate of pipe ij, respectively. i,min ,p i,max These are the upper and lower pressure limits for node i, respectively.
[0078] 2) Climbing constraints
[0079] The ramp constraint refers to the fact that the difference between actions at different times cannot exceed a certain range, including: upper and lower limits for power ramp and upper and lower limits for thermal ramp.
[0080]
[0081] (iv) State Diagram
[0082] In this embodiment, the power system, heating system, and natural gas system are abstractly modeled as graphs G(V,E) consisting of nodes and edges. Here, V represents a node in the system, and E represents an edge in the system.
[0083] S2: Collect historical state diagrams of the integrated electric, heat, and gas energy system to form a training set.
[0084] S3: Construct a maximum entropy reinforcement learning model based on graph neural network for the optimal scheduling of integrated electric, thermal, and gas energy systems. The original multilayer perceptron network is changed to a graph neural network, and the action space, state space, and reward function of the model are set based on the integrated electric, thermal, and gas energy system. The model is trained using a training set.
[0085] Reinforcement learning differs from supervised and unsupervised learning. Its core idea is to learn through exploration, adjust the policy based on feedback, and ultimately obtain the optimal solution in the current environment. Specifically, the agent obtains the current state s from the environment and outputs an action a, which affects the environment and yields a corresponding reward r(s,a). The agent adjusts its policy and learns network parameters based on the reward value, aiming to obtain the maximum cumulative reward.
[0086] The Actor-Critic (AC) algorithm, a widely used reinforcement learning algorithm, consists of two relatively independent but interactive networks: the Actor and the Critic. The Actor communicates through the policy network π. θ (a|s) gives the action a in the current state s. Critic, based on state s and reward r, uses the value network Q. β Provide an evaluation of the current strategy, its basic framework as follows: Figure 3 As shown, θ and β are the parameters of the neural network to be learned.
[0087] To enhance the algorithm's exploration capabilities, the SAC (Maximum Entropy Reinforcement Learning) algorithm adds an entropy term to the AC framework, aiming to maximize both the reward and the entropy value.
[0088]
[0089] The entropy term in the above formula Let π(a|s) represent the entropy value of policy π(a|s) in state s, where α is the temperature coefficient that adjusts the weight of the entropy term in the cumulative reward. The purpose of increasing the entropy term is to randomize the policy, making the probability distribution of each action as dispersed as possible, resulting in a larger entropy. This ensures the randomness of the exploration, expands the exploration range, and avoids getting trapped in local optima.
[0090] When evaluating the strategy, the Soft Q-value function and the modified Bellman operator are used.
[0091]
[0092]
[0093] In the above formula, γ is a discount factor used to adjust the agent's emphasis on short-term and long-term rewards, V(s t ) is a state value function used to evaluate the current state.
[0094] Value Network Q β With policy network π θ Update by minimizing the Bellman residual and minimizing the KL divergence respectively:
[0095]
[0096]
[0097] In the above formula, Z β (·) is a function that normalizes the distribution.
[0098] Graph-structured data can be defined as G = (V, E), where V represents the set of nodes in the system, and E represents the set of edges in the system. In the training of graph neural network models, node updates and representations are achieved through information exchange and aggregation between nodes.
[0099]
[0100] In the above formula This represents the vector representation of node i after passing through the k-th layer of the neural network. Let γ and φ represent the neighboring nodes of node i, and let e represent different differentiable functions. i,j Let be the eigenvectors of the edges.
[0101] A major limitation of graph neural network models in graph representation learning is that when too many network layers are stacked, the training performance deteriorates due to the convergence of node representations, a phenomenon commonly referred to as oversmoothing. To better utilize information in the graph and prevent oversmoothing, this paper employs an attention mechanism to aggregate node information to obtain node representations, specifically:
[0102]
[0103] In the above formula, W represents the neural network parameter matrix, which is used to perform linear transformation on the node features, and α i,j Attention coefficient:
[0104]
[0105] In the above formula, vector a is the parameter vector of the attention network, and W e It is a parameter matrix that performs a linear transformation on the edge information, GELU is the activation function, and || is the vector concatenation operator.
[0106] This embodiment employs a maximum entropy reinforcement learning model based on graph neural networks, where the Actor network structure is compared to the Actor network structure of traditional reinforcement learning models, for example... Figure 4 As shown.
[0107] like Figure 4 As shown on the left, the Actor network, based on a graph neural network structure, takes the state graph G(V,E) at the current time t as input. After passing through k layers of the graph neural network, with GELU as the activation function of each layer, the outputs are the mean μ and the logarithm of the variance lnσ for each action. After performing an exponential transformation on lnσ, a normal distribution N(μ,σ) is obtained. 2 After sampling and adding noise, the values in the range (-1,1) are obtained through the Tanh layer, and then linearly mapped to the action range to obtain the actual output value.
[0108] When constructing a maximum entropy reinforcement learning model, it is also necessary to define the action space, state space, and reward function.
[0109] 1) Action Space
[0110] The actions output by the intelligent agent include: active power output of thermal power plants, electrical and thermal power output of CHP units, wind power absorption coefficient, and gas supply of natural gas supply stations.
[0111]
[0112] In the above formula, α i Let α be the absorption coefficient of the wind power station. i P i W This refers to the grid-connected power of wind power. The corresponding operating ranges are as follows:
[0113]
[0114] In the above formula These represent the upper and lower limits of the active power output of thermal power plant i, respectively. These are the upper and lower limits of the electrical output of the CHP unit i, respectively. These represent the upper and lower limits of the thermal power output of the CHP unit, respectively. i,min ,f i,max These represent the upper and lower limits of the gas supply volume of gas supply station i within a certain period of time.
[0115] 2) State Space
[0116] The integrated energy system of electricity, heat and gas is modeled as a graph G(V,E), and the system state is reflected by the node features and edge features in the graph.
[0117] The node characteristics and edge characteristics of the power system are as follows:
[0118]
[0119] In the above formula, P i L Let be the electrical load of node i.
[0120] The node characteristics and edge characteristics of the thermal system are as follows:
[0121]
[0122] In the above formula Let be the heat load of node i.
[0123] The node characteristics and edge characteristics of the natural gas system are as follows:
[0124]
[0125] In the above formula, f i Lis the gas load at node i, l ij is the length of the pipeline between nodes i and j.
[0126] 3) Reward function
[0127] The reward value includes the system operation cost and the penalty for constraint violation. Since the goal pursues minimization while reinforcement learning pursues maximization of return, it is necessary to take a negative value of the reward function:
[0128]
[0129] In the above formula, r t represents the reward value at time t, F t is the operation cost at time t shown in formula (10), λ i is the penalty factor corresponding to the i-th constraint. To ensure that the training results meet the constraint conditions, a relatively large penalty factor is generally set. i represents the serial number of the constraints listed in formulas (11) and (12), |L i | represents the absolute value of the satisfaction difference of the i-th constraint. When the constraint is satisfied, |L i | is 0; when the constraint is not satisfied, |L i | is the minimum value among the absolute values of the differences from the boundary conditions. When the constraint includes upper and lower limit boundary conditions, the minimum value among the absolute values of the differences from the upper limit and the lower limit is taken (for example, a<X<b, when |X-a|<|X-b|, then |L i |=|X-a|; when |X-a|>|X-b|, then |L i |=|X-b|); when the constraint only includes an upper limit boundary condition or a lower limit boundary condition, the absolute value of the difference from the upper limit or the absolute value of the difference from the lower limit is taken.
[0130] S4: Obtain the output result of the integrated electricity-heat-gas energy system through the trained model.
[0131] Experimental verification and analysis
[0132] As shown in Figure 5 , this example adopts a modified and fused 6-6-6 node integrated electricity-heat-gas energy system example
[17] to verify the training efficiency of the proposed algorithm. The black part represents the power system, the red part represents the thermal system, and the blue part represents the natural gas system. Among them, Bus 2 is connected to a thermal power station; Bus 6 is connected to a wind power station; CHP1 is a coal-fired combined heat and power unit, and CHP2 is a gas-fired combined heat and power unit.
[0133] In this experiment, comparative tests were conducted on GNN-based and MLP-based architectures with the same number of network layers and neurons. The Actor network had 3 layers with 96 neurons in each layer. The Critic network also had 3 layers with 96 neurons in each layer, using GELU as the activation function and a pool size of 24000. The Adam optimizer was used to automatically adjust the learning rate.
[0134] from Figure 6 As can be seen, the reward curve based on the GNN architecture converges after 2000 training rounds, and compared with the reward curve based on the MLP architecture, the GNN architecture converges faster. This indicates that the algorithm model based on the GNN architecture utilizes edge information, resulting in a larger exploration space and faster training speed.
[0135] After training, the policy network obtains system output based on load, and the 24-hour scheduling results are as follows: Figure 7 As shown.
[0136] Figure 7 The system output results for each time period show that the total output basically matches the load curve. CHP1, with its larger installed capacity and lower cost coefficient, undertakes more output tasks. During peak daytime electricity consumption, it exhibits a slow ramp-up characteristic, effectively meeting the increased actual electricity demand. As nighttime temperatures drop, the heat load increases, and CHP1's heat output also increases accordingly. More wind power generation occurs at night, leading to a corresponding increase in grid-connected power, with a wind power absorption rate exceeding 90%. The output of thermal power plants remains relatively stable, with the power gap being filled by CHP2.
[0137] This invention proposes a reinforcement learning model based on a graph neural network architecture. By modeling the integrated energy system as a graph and feeding it into a Soft Actor-Critic reinforcement learning model based on a graph neural network architecture, faster exploration and learning are achieved, ultimately leading to an optimized scheduling scheme. Compared to MLP-based reinforcement learning models, this method can achieve convergence results faster while satisfying safety constraints.
[0138] Example 2:
[0139] The present invention also provides an optimized scheduling terminal device for an integrated electric, thermal, and gas energy system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method embodiment described above in Embodiment 1 of the present invention.
[0140] Furthermore, as an executable solution, the integrated electric, heat, and gas energy system optimization and scheduling terminal device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The integrated electric, heat, and gas energy system optimization and scheduling terminal device may include, but is not limited to, a processor and a memory. Those skilled in the art will understand that the above-described composition of the integrated electric, heat, and gas energy system optimization and scheduling terminal device is merely an example and does not constitute a limitation on the device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the integrated electric, heat, and gas energy system optimization and scheduling terminal device may also include input / output devices, network access devices, buses, etc., and this embodiment of the invention does not limit this.
[0141] Furthermore, as an executable solution, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. This processor is the control center of the integrated electric, thermal, and gas energy system optimization and scheduling terminal equipment, connecting various parts of the terminal equipment via various interfaces and lines.
[0142] The memory can be used to store the computer programs and / or modules. The processor, by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory, realizes various functions of the integrated electric, heat, and gas energy system optimization scheduling terminal equipment. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory and may also include non-volatile memory, such as hard disk, RAM, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0143] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the methods described in the embodiments of the present invention.
[0144] If the modules / units integrated in the optimized scheduling terminal equipment of the aforementioned integrated electric, heat, and gas energy system are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc.
[0145] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.
Claims
1. A method for optimal scheduling of an integrated energy system of electric heating and gas, characterized in that, Includes the following steps: S1: The integrated energy system of electricity, heat and gas is abstractly modeled as a state diagram, in which the electrical system equipment represents the node characteristics through the electrical load of the equipment, and the edge characteristics are represented by the susceptance and conductance between the two equipment corresponding to the two nodes. In a thermal system, node characteristics are represented by the heat load of the equipment, and edge characteristics are represented by the length of the pipeline branch between the two devices corresponding to the two nodes and the pipeline mass flow rate. In a natural gas system, node characteristics are represented by the gas load of the equipment, and edge characteristics are represented by the pipeline length between the two devices corresponding to the two nodes and the pipeline constant. S2: Collect historical state diagrams of the integrated electric, heat, and gas energy system to form a training set; S3: Construct a maximum entropy reinforcement learning model based on graph neural network for the optimal scheduling of integrated electric, thermal and gas energy systems. The model changes the original multilayer perceptron network to a graph neural network, and sets the action space, state space and reward function of the model based on the integrated electric, thermal and gas energy system. The model is trained using a training set. S4: Obtain the output results of the integrated electric, thermal, and gas energy system through the trained model.
2. The optimized scheduling method for an integrated electric, thermal, and gas energy system according to claim 1, characterized in that: In the action space of the reinforcement learning model, the active power output of the thermal power plant, the electrical output of the CHP unit, the thermal power output of the CHP unit, the absorption coefficient of the wind power plant, and the gas supply of the gas supply station are used as action variables, and the value range of each action variable is set.
3. The optimized scheduling method for an integrated electric, thermal, and gas energy system according to claim 1, characterized in that: In the state space of a reinforcement learning model, the system state is represented by the node and edge features of the graph.
4. The optimized scheduling method for an integrated electric, thermal, and gas energy system according to claim 1, characterized in that: In the graph neural network of reinforcement learning models, an attention mechanism is used to aggregate node information to obtain node representations.
5. A terminal device for optimizing and dispatching an integrated electric, heat, and gas energy system, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 4.
6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Energy routing optimization method
CN113132232A
Optimized operation method and system of electric-thermal combined system, equipment and medium
CN113780688A