Optimization scheduling methods, terminal equipment and storage media for combined electric and thermal energy systems
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-30
- Publication Date
- 2026-08-11
AI Technical Summary
但过高的复杂度以及每当系统状态变化就需要重新求解使得这些方法在面对大规模问题时难以实现快速响应
[0021]本发明采用如上技术方案,提出了一种基于GNN架构的强化学习优化调度方法,与基于MLP架构的方法相比,系统拓扑信息的利用带来了更大的探索空间与更快的收敛速度,在电热联合能源系统优化调度工作中更具优势。
Smart Images

Figure CN116362504B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of energy optimization scheduling, and in particular to a method, terminal equipment and storage medium for optimizing scheduling of combined electric and thermal energy systems. Background Technology
[0002] With the increasing demand for energy in society and the growing contradiction between energy conservation and emission reduction, how to make full use of new energy sources to reduce the use of traditional energy sources and achieve the goals of reducing operating costs and energy conservation and emission reduction has become an urgent problem to be solved.
[0003] The development of the energy internet has provided a guarantee for the complementarity and conversion of multiple energy flows, realizing the full utilization of energy. Among these, multi-energy flow scheduling and coupling are key to the efficient operation of integrated energy systems. Due to numerous nonlinear constraints in energy systems, multi-energy flow optimization scheduling, as a non-convex optimization problem, is difficult to find a globally optimal solution. Traditional solutions to this problem have mostly focused on approximate and nonlinear solutions, and intelligent algorithms such as particle swarm optimization have emerged. However, the high complexity and the need to re-solve the problem whenever the system state changes make these methods difficult to respond quickly to large-scale problems. With the popularization of renewable energy sources such as photovoltaics and wind power, the volatility and uncertainty of their output bring new challenges to optimization scheduling. Summary of the Invention
[0004] To address the aforementioned problems, this invention proposes an optimized scheduling method, terminal equipment, and storage medium for a combined electric and thermal energy system.
[0005] The specific plan is as follows:
[0006] A method for optimizing the scheduling of a combined electric and thermal energy system includes the following steps:
[0007] S1: The combined electric and thermal energy system is abstractly modeled as a state diagram, where the electrical system equipment represents the node characteristics through the electrical load of the equipment, and the edge characteristics are represented by the susceptance and conductance between the two equipment corresponding to the two nodes; the thermal system equipment represents the node characteristics through the thermal load of the equipment, and the edge characteristics are represented by the length of the pipe branch and the pipe mass flow rate between the two equipment corresponding to the two nodes.
[0008] S2: Collect historical state diagrams of the combined electric and thermal energy system to form a training set;
[0009] S3: Construct a reinforcement learning model for the optimal scheduling of a combined electric and thermal energy system. Change the multilayer perceptron network in the reinforcement learning model to a graph neural network. Define the action space, state space, and reward function of the reinforcement learning model. Use the maximum entropy of the value distribution as the algorithm objective of reinforcement learning. Train the reinforcement learning model using the training set.
[0010] S4: Obtain the power output of the combined electrothermal energy system through the trained reinforcement learning model.
[0011] Furthermore, in the action space of the reinforcement learning model, the active power output of the thermal power plant, the electrical output of the CHP unit, the thermal power output of the CHP unit, the thermal power output of the heating station, and the absorption coefficient of the wind power plant are used as action variables, and the range of values for each action variable is set.
[0012] Furthermore, the system state in the state space of the reinforcement learning model is represented by the node features and edge features of the graph.
[0013] Furthermore, the formula for calculating the reward function of a reinforcement learning model is as follows:
[0014]
[0015] Where, r t F represents the reward at time t. t Let λ represent the operating cost at time t, i represent the constraint number, and λ represent the operating cost at time t. i Let |L| represent the penalty factor corresponding to the i-th constraint. i | represents the absolute value of the difference between the establishment of the i-th constraint and the establishment of the constraint. When the constraint is established, |L i | is 0; when the constraint is not met, |L i | represents the minimum absolute value of the difference between the boundary conditions and the boundary conditions.
[0016] Furthermore, the combined heat and power (CHP) energy system consists of a thermal power plant, a heating plant, a CHP unit, and a wind power plant. Its operating cost is the sum of the operating costs of the thermal power plant, the heating plant, the CHP unit, and the wind curtailment cost of the wind power plant.
[0017] Furthermore, the reinforcement learning model uses an attention mechanism in its graph neural network to aggregate node information and obtain node representations.
[0018] Furthermore, in the reinforcement learning model, instead of directly calculating the expected value of the soft Q function reward, the model models the value distribution function of the soft Q function reward and learns the soft Q function reward based on the Bellman operator.
[0019] An optimized scheduling terminal device for an electric and thermal combined energy system includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in the embodiments of the present invention.
[0020] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described above in the embodiments of the present invention.
[0021] This invention adopts the above technical solution and proposes a reinforcement learning optimization scheduling method based on GNN architecture. Compared with the method based on MLP architecture, the utilization of system topology information brings a larger exploration space and a faster convergence speed, which is more advantageous in the optimization scheduling of combined electric and thermal energy systems. Attached Figure Description
[0022] Figure 1 The diagram shown is a flowchart of Embodiment 1 of the present invention.
[0023] Figure 2 The diagram shown is a schematic of the reinforcement learning model algorithm framework in this embodiment.
[0024] Figure 3 The diagram shown is a schematic of the Actor network structure in this embodiment.
[0025] Figure 4 The diagram shown is a schematic of the combined electrothermal energy system in this embodiment.
[0026] Figure 5 The diagram shown is a comparison of GNN and MLP results in this embodiment.
[0027] Figure 6 The diagram shown is a schematic representation of the power system output in this embodiment.
[0028] Figure 7 The diagram shown is a schematic representation of the power output of the thermal system in this embodiment. Detailed Implementation
[0029] To further illustrate the various embodiments, the present invention provides accompanying drawings. These drawings are part of the disclosure of the present invention, primarily used to illustrate the embodiments, and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementations and the advantages of the present invention.
[0030] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments.
[0031] Example 1:
[0032] This invention provides an optimized scheduling method for a combined electric and thermal energy system, such as... Figure 1 As shown, the method includes the following steps:
[0033] S1: Abstract the combined electric and thermal energy system into a state diagram.
[0034] I. Model of Combined Electricity and Heat Energy System
[0035] The electrothermal combined energy system constructed in this embodiment includes a power system model, a thermal system model, and an electrothermal system coupling link.
[0036] 1. Power System Model
[0037] The power flow equations for an AC power system are:
[0038]
[0039] Where: N P Let P represent the set of nodes in a power system. i Q i U represents the active power and reactive power injected at node i, respectively. i G represents the voltage magnitude at node i. ij B represents the conductance between node i and node j. ij θ represents the susceptance between node i and node j. ij =θ i -θ j This represents the phase angle difference between node i and node j.
[0040] 2. Thermodynamic System Model
[0041] Since heat conduction requires a medium, this embodiment selects the most commonly used hydraulic medium and divides the thermal system into a hydraulic model and a thermal model.
[0042] 1) Hydraulic model
[0043] The hydraulic model consists of the flow continuity equation and the loop pressure equation:
[0044]
[0045] In the formula: A represents the node-branch correlation matrix, m represents the pipeline mass flow rate vector, m q Let B represent the node injection flow vector, and let h represent the loop-branch correlation matrix. f This represents the head loss vector, which is related to the pipe's damping coefficient and mass flow rate.
[0046] 2) Thermodynamic model
[0047] The thermal model includes nodal power equations, pipe temperature drop equations, and medium mixing equations:
[0048]
[0049] Where: N H H represents the set of nodes in a thermal system. iC represents the thermal power of node i. p m is the specific heat capacity of water. q,i T represents the injected flow at node i. i in ,T i out T represents the inlet and outlet water temperatures at node i, respectively. j,i T represents the water temperature at end j of pipe branch ij. i,j This represents the water temperature at end i of pipe branch ij, T. e Where L is the ambient temperature, λ is the thermal conductivity, and L is the external temperature. ij T represents the length of pipe branch ij. i The water temperature at node i is represented by m. ik Represents the mass flow rate between node k and node i, |n i | represents the total number of nodes that flow to node i.
[0050] 3. Coupling links in the electrothermal system
[0051] For the coupling link of the combined power and heat energy system, this embodiment considers a CHP cogeneration unit that can generate electricity and supply heat simultaneously to meet the load demand of the power system and the heat system.
[0052] Common models for combined heat and power (CHP) units include polygonal models and linear models using a fixed heat-to-electricity ratio. In this embodiment, a more flexible extraction-condensing unit is selected, and its corresponding polygonal model is as follows:
[0053]
[0054] In the formula: P CHP H CHP These represent the electrical output and thermal output of the CHP unit, respectively. These represent the upper and lower limits of the electrical output power of the CHP unit, respectively. α1, α2, and α3 represent the upper and lower limits of the thermal output power of the CHP unit, respectively, and are polygonal region coefficients.
[0055] II. Objective Function
[0056] For the task of optimizing the scheduling of the combined power and heat system, this embodiment aims to minimize operating costs and absorb as much renewable energy output as possible.
[0057] 1) Operating costs of thermal power plants
[0058]
[0059] In the formula: F 1,t |N represents the operating cost of all thermal power plants at time t. P | P represents the number of thermal power plants.i,t Let αi represent the active power output of thermal power plant i at time t, and α0, α1, α2 be the parameters of the consumption characteristic curve of the thermal power unit.
[0060] 2) Operating costs of heating stations
[0061]
[0062] In the formula: F 2,t |N represents the operating cost of all heating stations at time t. H H represents the number of heating stations. i,t β0, β1, and β2 represent the heat output of heating station i at time t, and β0, β1, and β2 are the parameters of the consumption characteristic curve of the heating station.
[0063] 3) Operating costs of CHP units
[0064]
[0065] In the formula: F 3,t Represents the operating cost of all CHP units at time t, |N CHP | This represents the number of CHP units. denoted as electrical output and thermal output of CHP unit i at time t, respectively, and μ0~μ5 are the consumption characteristic curve parameters of CHP unit.
[0066] 4) Cost of wind curtailment
[0067]
[0068] In the formula: F 4,t Let |N| be the cost of wind curtailment for all wind power plants at time t. W | represents the number of wind farms, α i The wind power output absorption coefficient is the coefficient of performance. Let i be the output of wind power station i at time t, corresponding to For grid-connected wind power, C w This represents the cost coefficient for wind curtailment.
[0069] 5) Objective function
[0070] minF t =F 1,t +F 2,t +F 3,t +F 4,t (9)
[0071] In the formula, F t This represents the total operating cost of the combined electric and thermal energy system at time t.
[0072] III. Constraints
[0073] 1) Power balance constraints
[0074]
[0075] In the formula: Let represent the output of the conventional unit, CHP unit, and renewable energy power plant at node i at time t, respectively. Let represent the thermal output of the conventional unit and the CHP unit at node i at time t, respectively. These are electrical load and thermal load, respectively.
[0076] 2) Safety constraints
[0077] The stable operation of a combined heat and power system requires meeting certain safety constraints, including voltage constraints, phase angle difference constraints, and line transmission constraints for the power network, and node temperature and pipeline flow constraints for the thermal network.
[0078]
[0079] In the formula: U i,min U i,max These are the upper and lower limits of the voltage amplitude at node i, respectively. P is the upper limit of the phase angle difference. l T represents the upper limit of power transmission capacity of power lines. i,max ,T i,min Let m be the upper and lower limits of the water supply temperature at node i, respectively. ij,max ,m ij,min These represent the upper and lower limits of the water supply flow rate for pipeline ij, respectively.
[0080] IV. State Diagram
[0081] Since power and heat networks have a natural graph structure, they can be abstractly modeled as a graph G(V,E) consisting of nodes and edges, without considering internal device information. Here, V represents a node in the system, and E represents an edge in the system.
[0082] S2: Collect historical state diagrams of the combined electric and thermal energy system to form a training set.
[0083] S3: Construct a reinforcement learning model for the optimal scheduling of a combined electric and thermal energy system. Change the fully connected neural network in the reinforcement learning model to a graph neural network. Define the action space, state space, and reward function of the reinforcement learning model. Use the maximum entropy of the value distribution as the algorithm objective of reinforcement learning. Train the reinforcement learning model using a training set.
[0084] Reinforcement learning involves an agent learning through exploration to find the optimal solution to a problem in the current environment. The agent acquires the current state `s` and outputs an action `a`, which affects the environment and yields a corresponding reward `r`. The agent learns network parameters based on the reward values and continuously adjusts its output policy to maximize the cumulative reward.
[0085]
[0086] In the formula: G is the cumulative return value, γ∈[0,1] is the discount rate, which is used to adjust the agent's weights on short-term and long-term returns.
[0087] The Actor-Critic algorithm is a reinforcement learning method that combines policy gradient and temporal difference learning. Its basic architecture is as follows: Figure 2 As shown. Where Actor refers to the policy network π θ (a|s), which means learning a strategy to obtain the highest possible reward, Critic is the value network V φ (s) is used to estimate the current policy and output the evaluation value. Therefore, the Actor-Critic algorithm can update parameters in one step, without waiting for the environment time to end before updating the network. The policy network π in this algorithm framework... θ (s,a) and value network V φ In (s), θ and φ are functions to be learned, which need to be learned during the training process. In each update step, the Actor adjusts the current environment state s. t Output action a t and receive an immediate reward r(s) t ,a t ,s t+1 ). Critic calculates the actual return based on the environment and the score r+γV under previous standards. φ (s t+1 The Critic adjusts their scoring criteria to make their ratings more closely reflect the true rewards of the environment. The Actor, in turn, adjusts their strategy based on the Critic's ratings. θ .
[0088] The maximum entropy Actor-Critic algorithm for value distribution used in this embodiment aims to:
[0089]
[0090] In the formula, Let π(a|s) represent the entropy value of policy π(a|s) in state s. Compared to the Actor-Critic algorithm, the purpose of adding the entropy term is to randomize the policy, that is, to distribute the probability of each action as widely as possible, rather than concentrating it on a single action. This ensures the randomness of policy learning, maximizes the scope of exploration, and avoids falling into local optima.
[0091] To evaluate policy π, a soft Q-value function is defined and the Bellman operator is used.
[0092]
[0093] The goal of strategy improvement is to find a new strategy π. new A better strategy than the current one results in a higher expected return, and the policy network updates its learning based on maximizing the soft Q value:
[0094]
[0095] To avoid overestimating the Q-value during learning and thus reducing policy performance, this algorithm no longer directly calculates the soft reward Z. π The expected value Q of (s,a) π (s,a) was modeled instead of soft reward Z. π Distribution of (s,a): It is called the value distribution function, and the soft reward Z is learned based on the Bellman operator. π (s,a):
[0096]
[0097] In the formula, r ~ R(·|s,a), s t+1 ~p,a t+1 ~π. Symbol This indicates that the random variables on both sides have the same probability distribution. Assume... Follows distribution Update parameters by minimizing the distribution distance:
[0098]
[0099] In the formula, d is a distance function that measures the distance between two distributions, commonly known as KL divergence.
[0100] Compared to MLPs (Multilayer Perceptrons) that do not utilize topological information, graph neural network models can transmit information between nodes based on the connections between them. To better utilize the information in the graph, this embodiment employs an attention mechanism to aggregate node information to obtain node representations:
[0101]
[0102] In the formula Let W represent the vector representation of node i in the k-th layer of the neural network, and let W represent the neural network parameter matrix, to perform a linear transformation on the node features. Let α represent the neighboring nodes of node i. i,j Attention coefficient:
[0103]
[0104] In the formula, vector a is the parameter vector of the attention network, and W e It is the parameter matrix for linear transformation of the edge information, e i,j Let be the feature vector of the edge, GELU be the activation function, and || be the vector concatenation operator.
[0105] Solving for the optimal scheduling strategy of a combined electric and thermal energy system requires defining the action, state, and reward functions for the problem.
[0106] 1) Action Space
[0107] The active power output of thermal power plants, the electrical and thermal power output of CHP units, the thermal power output of heating stations, and the absorption coefficient of wind power plants are used as action variables:
[0108]
[0109] The corresponding action range is as follows:
[0110]
[0111] In the formula: These represent the upper and lower limits of the active power output of thermal power plant i, respectively. These are the upper and lower limits of the electrical output of the CHP unit i, respectively. These represent the upper and lower limits of the thermal power output of the CHP unit i, respectively. These are the upper and lower limits of the heat output of heating station i, respectively.
[0112] 2) State Space
[0113] Since the system is modeled as a graph G(V,E), the system state is reflected through node and edge characteristics. The equipment in a combined electric and thermal energy system is divided into power system equipment and thermal system equipment.
[0114] The node characteristics and edge characteristics of the power system are as follows:
[0115]
[0116] In the formula: P i L Let be the electrical load of node i.
[0117] The node features and edge features of the thermal system are as follows:
[0118]
[0119] In the formula: is the heat load of node i. L ij represents the length of pipeline branch ij, m ij represents the pipeline mass flow rate between node i and node j.
[0120] 3) Reward function
[0121] The calculation of the reward value includes the system operation cost and the penalty for violating constraints. Since the operation cost aims to be minimized while reinforcement learning aims to maximize the reward, the reward needs to be taken as a negative value:
[0122]
[0123] In the formula: rt represents the reward value at time t, F t is the operation cost at time t shown in formula (9), λ i is the penalty factor corresponding to the i-th constraint, i represents the serial number of the constraints listed in formulas (10) and (11), |L i | represents the absolute value of the difference in the establishment of the i-th constraint. When the constraint is established, |L i | is 0; when the constraint is not established, |L i | is the minimum value among the absolute values of the differences from the boundary conditions. For example, when the constraint is an equation in formula (10), when the constraint is not established, |L i | is the absolute value of the difference between both sides of the equation; for example, when the constraint is an inequality in formula (11), when the constraint is not established, if the upper and lower limit boundary conditions are included, the minimum value among the absolute values of the differences from the upper and lower limits is taken (if there is a < X < b, when |X - a| < |X - b|, then |L i | = |X - a|; when |X - a| > |X - b|, then |L i | = |X - b|); if only the upper limit boundary condition or the lower limit boundary condition is included, the absolute value of the difference from the upper limit or the absolute value of the difference from the lower limit is taken.
[0124] In this embodiment, the network architecture of the Actor is as Figure 3 shown. The model input is the state graph G(V t , E t ) at the current time t. After passing through k layers of graph neural networks, the representation h t,k is obtained. The activation function of each layer is GELU, and the mean μ and the logarithm of the variance lnσ of each action are output. After performing an exponential transformation on lnσ, the normal distribution N(μ, σ2 After sampling and adding noise, the values in (-1,1) are obtained through the Tanh layer, and finally mapped to the action range listed in formula (21) to obtain the actual dispatch output value.
[0125] The learning of the value network Critic needs to be combined with the output actions, and the node characteristics of the power system at this time. Nodal characteristics of thermal systems They are respectively:
[0126]
[0127] The graph G(V) in this state t ',E t The soft Q value is obtained by feeding it into the value network.
[0128] S4: Obtain the power output of the combined electrothermal energy system through the trained reinforcement learning model.
[0129] Experimental verification analysis
[0130] This embodiment uses, as follows: Figure 4 The combined heat and power (CHP) system shown is analyzed as a case study, consisting of an IEEE-33 bus grid and a 32-node heating network in Bali, to analyze the optimization effect of reinforcement learning. G1 and G2 are two thermal power plants; W is a wind power plant; GB1 and GB2 are two heating stations; and CHP is a combined heat and power (CHP) unit.
[0131] In this embodiment, comparative experiments were conducted with the same number of network layers and neurons. The Actor network had 4 layers, with 128, 64, 32, and 32 neurons per layer. The Critic network had 5 layers, with 128, 64, 64, 32, and 32 neurons per layer. The activation function for each layer was GELU, and the experience pool size was 500,000. The Adam optimizer was used to automatically adjust the learning rate, with an adjustment range of 5 × 10⁻⁶. -4 ~5×10 -6 .
[0132] from Figure 5 As can be seen, the reward curve based on the GNN architecture converges after 3000 training rounds. Compared with the reward curve based on the MLP architecture, GNN starts to rise and converge faster, and the reward value is larger, meaning the operating cost is lower. This indicates that the algorithm model based on the GNN architecture utilizes edge information, resulting in a larger exploration space and faster training speed.
[0133] After training, the policy network can derive system output actions based on system load, resulting in 24-hour scheduling results. Figure 6 and Figure 7 As shown.
[0134] Figure 6 The power system output results for each time period show that the total output basically matches the load curve. Thermal power units 1 and 2, due to their larger installed capacity, undertook more power generation tasks and exhibited a slow ramp-up characteristic during peak daytime electricity demand, thus meeting the increased actual electricity demand. The load gap during peak periods was filled by the CHP units, which operated at minimum output power during other periods. Wind power generation increased at night, and the corresponding grid-connected power also increased accordingly, with the wind power absorption coefficient consistently remaining around 97%. Figure 7 The heat output results shown indicate that the total output basically matches the load curve, and it can gradually ramp up to meet demand during peak nighttime load periods. Due to the upper limit of output, the output differences between the heat sources are small, which can effectively reduce heat loss during transmission.
[0135] This invention proposes a value distribution maximum entropy Actor-Critic reinforcement learning optimization scheduling method based on a GNN architecture. This method can fully utilize the system's topological structure information to achieve more effective exploration and learning. Compared with methods based on an MLP architecture, this embodiment utilizes system topological information, resulting in a larger exploration space and faster convergence speed, making it more advantageous in the optimization scheduling of combined electric and thermal energy systems.
[0136] Example 2:
[0137] The present invention also provides an optimized scheduling terminal device for an electric and thermal combined energy system, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps in the method embodiment described above in Embodiment 1 of the present invention.
[0138] Furthermore, as an executable solution, the power-heat combined energy system optimization and scheduling terminal device can be a computing device such as a desktop computer, laptop, handheld computer, or cloud server. The power-heat combined energy system optimization and scheduling terminal device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that the above-described composition of the power-heat combined energy system optimization and scheduling terminal device is merely an example and does not constitute a limitation on the power-heat combined energy system optimization and scheduling terminal device. It may include more or fewer components than described above, or combine certain components, or different components. For example, the power-heat combined energy system optimization and scheduling terminal device may also include input / output devices, network access devices, buses, etc., and this embodiment of the invention does not limit this.
[0139] Furthermore, as an executable solution, the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. This processor is the control center of the combined electric and thermal energy system optimization scheduling terminal equipment, connecting various parts of the terminal equipment via various interfaces and lines.
[0140] The memory can be used to store the computer programs and / or modules. The processor, by running or executing the computer programs and / or modules stored in the memory and calling the data stored in the memory, realizes various functions of the optimized scheduling terminal equipment of the combined electric and thermal energy system. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a function; the data storage area may store data created based on the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0141] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the method described in the embodiments of the present invention.
[0142] If the modules / units integrated in the optimized scheduling terminal equipment of the aforementioned combined electric and thermal energy system are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), and software distribution media, etc.
[0143] Although the invention has been specifically shown and described in conjunction with preferred embodiments, those skilled in the art should understand that various changes in form and detail may be made to the invention without departing from the spirit and scope of the invention as defined in the appended claims, all of which shall be within the scope of protection of the invention.
Claims
1. A method for optimal scheduling of a combined electric and thermal energy system, characterized in that, Includes the following steps: S1: The combined electric and thermal energy system is abstractly modeled as a state diagram, in which the electrical system equipment represents the node characteristics through the electrical load of the equipment, and the edge characteristics are represented by the susceptance and conductance between the two equipment corresponding to the two nodes. In a thermal system, the characteristics of nodes are represented by the heat load of the equipment, and the characteristics of edges are represented by the length of the pipe branch between the two devices corresponding to the two nodes and the mass flow rate of the pipe. S2: Collect historical state diagrams of the combined electric and thermal energy system to form a training set; S3: Construct a reinforcement learning model for the optimal scheduling of a combined electric and thermal energy system. Change the multilayer perceptron network in the reinforcement learning model to a graph neural network, and define the action space, state space, and reward function of the reinforcement learning model. Use the maximum entropy of the value distribution as the algorithm objective of reinforcement learning, and train the reinforcement learning model using a training set. In the reinforcement learning model, instead of directly calculating the expected value of the soft Q function reward, model the value distribution function of the soft Q function reward and learn the soft Q function reward based on the Bellman operator. S4: Obtain the power output of the combined electrothermal energy system through the trained reinforcement learning model.
2. The method for optimal scheduling of a combined electric and thermal energy system according to claim 1, characterized in that: In the action space of the reinforcement learning model, the active power output of the thermal power plant, the electrical power output of the CHP unit, the thermal power output of the CHP unit, the thermal power output of the heating station, and the absorption coefficient of the wind power plant are used as action variables, and the range of values for each action variable is set.
3. The method for optimal scheduling of a combined electric and thermal energy system according to claim 1, characterized in that: In the state space of a reinforcement learning model, the system state is represented by the node and edge features of the graph.
4. The method for optimal scheduling of a combined electric and thermal energy system according to claim 1, characterized in that: The formula for calculating the reward function of a reinforcement learning model is: in, This represents the reward value at time t. This represents the operating cost at time t. i Indicates the constraint number. Indicates the first i The penalty factor corresponding to each constraint Indicates the first i The absolute value of the difference between the fulfillment of the constraints; when the constraints are fulfilled. =0; when the constraint is not met, It is the minimum absolute value of the difference between the boundary conditions and the boundary conditions.
5. The method for optimal scheduling of a combined electric and thermal energy system according to claim 4, characterized in that: The combined heat and power (CHP) energy system consists of a thermal power plant, a heating plant, a CHP unit, and a wind power plant. Its operating cost is the sum of the operating costs of the thermal power plant, the heating plant, the CHP unit, and the wind curtailment cost of the wind power plant.
6. The method for optimal scheduling of a combined electric and thermal energy system according to claim 1, characterized in that: In the graph neural network of reinforcement learning models, an attention mechanism is used to aggregate node information to obtain node representations.
7. A terminal device for optimized scheduling of a combined electric and thermal energy system, characterized in that: It includes a processor, a memory, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 6.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Optimized operation method and system of electric-thermal combined system, equipment and medium
CN113780688A
Method, system and device for economic dispatching decision-making of electric power system and medium
CN114358520A