Optimization methods, devices, and electronic equipment for electric heating control strategies

CN120470774BActive Publication Date: 2026-09-01STATE GRID INNER MONGOLIA EASTERN ELECTRIC POWER CO LTD TONGLIAO POWER SUPPLY CO +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510549107.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2026-09-01
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

[0005]本发明提供一种电采暖调控策略的优化方法、装置和电子设备,用以解决现有技术中导致确定的调控优化策略的准确性降低的缺陷

Benefits of technology

[0053]本发明提供的电采暖调控策略的优化方法、装置和电子设备,基于目标区域内的配电网结构,构建电网拓扑图,其中,电网拓扑图中的节点用于表征配电网结构中的母线,电网拓扑图中的边用于表征与边连接的两个节点之间的耦合度,且耦合度是基于两个节点各自对应配电台区的资源密度、居住面积比例、室内温度、热阻、电采暖集中度和电气距离等指标确定的,基于耦合度阈值和各边对应的耦合度,对电网拓扑图进行划分,得到至少两个电网子图,针对各电网子图,获取电网子图内各目标节点在当前时刻的净负荷功率和各目标节点对应的配电台区在当前时刻的目标室内温度,并将各净负荷功率和各目标室内温度输入电采暖优化调控网络中,得到电采暖优化调控网络输出的目标子区域内的电采暖调控优化策略。可知,本发明能够基于每个节点对应的配电台区的资源密度、居住面积比例、室内温度、热阻、电采暖集中度和电气距离等多维度指标对整个目标区域的电网拓扑图进行划分,使得划分得到的各电网子图对应的子区域内用户的温度需求类似,进而针对各电网子图,基于电网子图内各目标节点的净负荷功率和各目标节点对应的配电台区的目标室内温度,确定目标子区域内的电采暖调控优化策略,实现了对每个目标子区域的单独调控,从而提高了确定的调控优化策略的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120470774B_ABST
    Figure CN120470774B_ABST
Patent Text Reader

Abstract

This invention provides an optimization method, apparatus, and electronic device for electric heating control strategies, relating to the field of power system technology. The method includes: constructing a power grid topology map of a target area, where nodes in the topology map represent buses in the distribution network structure, and edges represent the coupling degree between two nodes, determined based on the index values ​​of the respective indices of the two nodes in an index system; dividing the power grid topology map based on the coupling degrees corresponding to each edge; and inputting the net load power and target indoor temperature of each target node within the sub-map at the current moment into an electric heating optimization control network to obtain an electric heating control optimization strategy. This invention can divide the target area into control zones based on multi-dimensional indices corresponding to the distribution sub-regions of each node, ensuring similar temperature requirements for users within the sub-regions of the resulting power grid sub-maps, and then determining the electric heating control optimization strategy for each sub-map, thus improving the accuracy of the determined control optimization strategy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system technology, and in particular to an optimization method, apparatus, and electronic equipment for electric heating control strategies. Background Technology

[0002] As a crucial component of the new power system, the intermittency and volatility of new energy sources pose significant challenges to the power grid's source-load balance after grid connection. Due to factors such as the power grid's frequent "heat-driven power generation" operation in winter and the high cost of large-scale energy storage equipment, the regulation capacity of both "source" and "storage" is insufficient. Therefore, tapping the load-side regulation potential is key to improving the absorption of new energy and promoting the safe and stable operation of the power grid. As an important distributed flexible resource on the load side, electric heating accounts for a high proportion of the load during the winter heating season and shows a trend of massive development. Optimizing and regulating electric heating can help promote the safe and stable operation of the power grid and improve its flexibility and economy.

[0003] In related technologies, the control and optimization strategy for electric heating is usually determined based on the electrical information of all busbars under the distribution network structure of the entire region.

[0004] However, in the aforementioned technologies, if the temperature requirements of different sub-regions within the entire region are different, the accuracy of the determined control optimization strategy will be reduced. Summary of the Invention

[0005] This invention provides a method, apparatus, and electronic device for optimizing electric heating control strategies, in order to address the shortcomings of existing technologies that lead to reduced accuracy of determined control optimization strategies.

[0006] This invention provides an optimization method for electric heating control strategy, comprising the following steps.

[0007] Based on the distribution network structure within the target area, a power grid topology diagram is constructed. The nodes in the power grid topology diagram are used to represent the busbars in the distribution network structure, and the edges in the power grid topology diagram are used to represent the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in the index system. The indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node.

[0008] Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology is divided to obtain at least two power grid sub-graphs.

[0009] For each of the aforementioned power grid sub-graphs, obtain the net load power of each target node within the power grid sub-graph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time;

[0010] The net load power and the target indoor temperature are input into the electric heating optimization and control network to obtain the electric heating control optimization strategy in the target sub-region output by the electric heating optimization and control network. The electric heating control optimization strategy includes the electric heating power in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region.

[0011] According to the present invention, an optimization method for electric heating control strategy is provided, wherein the power grid topology is divided based on a coupling degree threshold and the coupling degree corresponding to each edge to obtain at least two power grid sub-graphs, including:

[0012] Delete edges with coupling degrees less than the coupling degree threshold from the power grid topology graph to obtain the power grid topology graph after deletion;

[0013] In the deleted power grid topology graph, identify all connected subgraphs;

[0014] Each of the connected subgraphs is defined as the power grid subgraph.

[0015] According to the present invention, an optimization method for electric heating control strategy is provided, the method further includes:

[0016] For each of the aforementioned indicators, the information entropy of the indicator is determined based on formula (1):

[0017] (1)

[0018] in, Indicates the first The value of each indicator, Indicates the first Information entropy of each indicator Indicates the total number of indicators;

[0019] Based on the information entropy of the indicator, the weight of the indicator is determined using formula (2):

[0020] (2)

[0021] in, Indicates the first The weight of each indicator;

[0022] Based on the respective index values ​​of the two nodes and the weights of each index, the coupling degree between the two nodes is determined using formula (3):

[0023] (3)

[0024] in, Represents a node With nodes The degree of coupling between them Represents a node The The value of each indicator, Represents a node The The value of each indicator.

[0025] According to the optimization method of electric heating control strategy provided by the present invention, the electric heating optimization control network is trained based on the following method:

[0026] The net load power of each sample node in the sample area at the time of acquisition and the indoor temperature of the distribution substation corresponding to each sample node at the time of acquisition are obtained. The sample node is used to characterize the bus in the sample distribution network structure in the sample area.

[0027] The net load power and indoor temperature of each sample at the acquisition time are used as state space information and input into the actor network to obtain the predicted action value of the acquisition time output by the actor network. The predicted action value includes the predicted power of electric heating in the sample sub-region and the predicted charging and discharging power of the energy storage system in the sample sub-region.

[0028] The state space information and the predicted action value are input into the critic network to obtain the value estimate output by the critic network;

[0029] With the goal of maximizing value estimation, the parameters of the actor network are adjusted until the convergence condition is met, and the final actor network is determined as the electric heating optimization and control network.

[0030] According to the present invention, an optimization method for electric heating control strategy is provided, wherein the actor network includes a long short-term memory network and a fully connected network;

[0031] The net load power and indoor temperature of each sample at the acquisition time are input as state-space information into the actor network to obtain the predicted action value at the acquisition time output by the actor network, including:

[0032] The state space information at the acquisition time and the hidden state information at the previous acquisition time are respectively input into the forget gate, input gate, memory update gate and output gate of the long short-term memory network to obtain the first result output by the forget gate, the second result output by the input gate, the third result output by the memory update gate and the fourth result output by the output gate.

[0033] Based on the first result and the second result, the third result of the memory update gate is updated to obtain the updated third result;

[0034] Based on the fourth result and the updated third result, the hidden state information at the acquisition time is determined;

[0035] The hidden state information at the acquisition time is input into the fully connected network to obtain the predicted action value output by the fully connected network.

[0036] According to the present invention, an optimization method for an electric heating control strategy includes adjusting the parameters of the actor network with the objective of maximizing value estimation until a convergence condition is met, and determining the final actor network as the electric heating optimization control network. The method comprises:

[0037] Based on the predicted action value and the action value label corresponding to the state space information, a loss function is constructed;

[0038] With the goal of maximizing value estimation, the weight matrices and bias terms of the forget gate, the input gate, the memory update gate and the output gate are adjusted based on the loss function until the convergence condition is met, and the final actor network is determined as the electric heating optimization and control network.

[0039] According to the optimization method of an electric heating control strategy provided by the present invention, the step of inputting the state space information and the predicted action value into a commentator network to obtain the value estimate output by the commentator network includes:

[0040] The state space information and the predicted action value are input into the critic network, and the critic network determines the value estimate based on the reward value.

[0041] The reward value is determined based on formula (4):

[0042] (4)

[0043] in, Indicates the reward value. Indicates the operating cost of the power distribution network. This indicates a deviation in the absorption of new energy sources. This indicates the weight corresponding to the operating cost of the distribution network. This indicates the weight corresponding to the deviation in the absorption of new energy sources. ,

[0044] , Indicates the first The cost of purchasing electricity from the upper-level power grid at the time of data collection. Indicates the first The regulation cost of the energy storage system at the time of data acquisition. Indicates the first The control cost of electric heating at the time of data collection. Indicates the first The first acquisition time The penalty cost of each constraint condition. Represents a set of constraints. Indicates the first The power output of the power distribution network at the time of data collection. Indicates the first Load difference at the time of data acquisition Indicates the total number of nodes. This represents the total duration of all data collection moments.

[0045] The present invention also provides an optimization device for electric heating control strategy, comprising:

[0046] The construction unit is used to construct a power grid topology map based on the distribution network structure within the target area. The nodes in the power grid topology map are used to represent the buses in the distribution network structure, and the edges in the power grid topology map are used to represent the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in the index system. The indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node.

[0047] A partitioning unit is used to partition the power grid topology based on a coupling degree threshold and the coupling degree corresponding to each edge, to obtain at least two power grid subgraphs;

[0048] The acquisition unit is used to acquire, for each of the power grid subgraphs, the net load power of each target node in the power grid subgraph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time;

[0049] The control unit is used to input the net load power and the target indoor temperature into the electric heating optimization control network to obtain the electric heating control optimization strategy in the target sub-region output by the electric heating optimization control network. The electric heating control optimization strategy includes the power of the electric heating load in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region.

[0050] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement an optimization method for any of the electric heating control strategies described above.

[0051] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements an optimization method for the electric heating control strategy as described above.

[0052] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements an optimization method for electric heating control strategies as described above.

[0053] The present invention provides an optimization method, apparatus, and electronic device for electric heating control strategies. Based on the distribution network structure within a target area, a power grid topology diagram is constructed. Nodes in the topology diagram represent buses in the distribution network structure, and edges represent the coupling degree between two nodes connected to the edges. The coupling degree is determined based on indicators such as resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution sub-area corresponding to each node. Based on the coupling degree threshold and the coupling degree corresponding to each edge, the topology diagram is divided into at least two sub-graphs. For each sub-graph, the net load power of each target node at the current time and the target indoor temperature of the corresponding distribution sub-area are obtained. These net load powers and target indoor temperatures are then input into an electric heating optimization control network to obtain the electric heating control optimization strategy for the target sub-area output by the network. As can be seen, the present invention can divide the power grid topology of the entire target area based on multi-dimensional indicators such as resource density, living area ratio, indoor temperature, thermal resistance, concentration of electric heating, and electrical distance of the distribution sub-area corresponding to each node. This makes the temperature requirements of users in the sub-areas corresponding to the divided power grid sub-maps similar. Then, for each power grid sub-map, based on the net load power of each target node in the power grid sub-map and the target indoor temperature of the distribution sub-area corresponding to each target node, the electric heating control optimization strategy in the target sub-area is determined, realizing individual control of each target sub-area, thereby improving the accuracy of the determined control optimization strategy. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0055] Figure 1 This is one of the flowcharts illustrating the optimization method of electric heating control strategy provided in the embodiments of the present invention.

[0056] Figure 2 This is the second flowchart illustrating the optimization method of the electric heating control strategy provided in this embodiment of the invention.

[0057] Figure 3This is a schematic diagram of the structure of the optimization device for electric heating control strategy provided in an embodiment of the present invention.

[0058] Figure 4 This is a schematic diagram of the physical structure of the electronic device provided by the present invention. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0060] The following is combined Figure 1 and Figure 2 This invention describes an optimization method for an electric heating control strategy. The execution entity of this optimization method can be an electronic device such as a terminal, tablet computer, or computer, or it can be an optimization device for the electric heating control strategy installed in the electronic device. This optimization device can be implemented through software, hardware, or a combination of both.

[0061] Figure 1 This is one of the flowcharts illustrating the optimization method for electric heating control strategy provided in this embodiment of the invention, such as... Figure 1 As shown, the optimization method for this electric heating control strategy includes the following steps:

[0062] Step 101: Based on the distribution network structure within the target area, construct a power grid topology diagram. The nodes in the power grid topology diagram are used to represent the buses in the distribution network structure, and the edges in the power grid topology diagram are used to represent the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in the index system. The indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node.

[0063] The busbar is used to concentrate the electrical energy from power sources such as generators, transformers, and transmission lines. The target area can be a specific town or a block, etc., and this invention does not limit it.

[0064] For example, the distribution network structure within the target area is obtained, and each bus in the distribution network structure is taken as a node in the power grid topology graph. When there is a connection between two buses, the nodes corresponding to the two buses are connected to form an edge in the power grid topology graph. The index values ​​of each indicator of the distribution substation corresponding to each node are obtained. For two nodes with an edge, the coupling degree between the two nodes is calculated based on the index values ​​of all indicators of the two nodes. The coupling degree is used to characterize the electrical coupling strength between the two nodes and is used as an attribute of the corresponding edge. Finally, the power grid topology graph is obtained. The power grid topology graph can be an undirected graph.

[0065] Specifically, the resource density value of the distribution radio zone corresponding to each node can be calculated based on the following formula (5):

[0066]

[0067] in, Represents a node The corresponding value of the resource density index for the distribution radio area. Represents a node The total resource output of the corresponding distribution area can be understood as the total economic output. Represents a node The total area of ​​the corresponding distribution radio area.

[0068] The residential area ratio of each distribution zone can be calculated based on the following formula (6):

[0069]

[0070] in, Represents a node The value of the indicator corresponding to the residential area ratio of the distribution area. Represents a node The total resource output of the corresponding distribution area can be understood as the total economic output. Represents a node The corresponding residential area area of ​​the distribution station zone, Represents a node The total area of ​​the corresponding distribution radio area.

[0071] The indoor temperature value of the distribution area corresponding to each node can be calculated based on the following formula (7):

[0072]

[0073] in, Represents a node The corresponding value of the indoor temperature index in the distribution area. Represents a node The number of buildings within the corresponding distribution area, Represents a node The corresponding distribution radio station area The indoor temperature of a building.

[0074] The thermal resistance of the distribution area corresponding to each node can be calculated based on the following formula (8):

[0075]

[0076] in, Represents a node The corresponding value of the thermal resistance index for the distribution area. Represents a node The corresponding distribution radio station area The thermal resistance of a building , Indicates the thickness of the building's material layers. This indicates the thermal conductivity of the material used in the material layer.

[0077] The concentration of electric heating in the distribution zone corresponding to each node can be calculated based on the following formula (9):

[0078] (9)

[0079] in, Represents a node The value of the indicator "centralization of electric heating in the corresponding distribution area". Represents a node The corresponding electric heating load power of the distribution area Represents a node The total load power of the corresponding distribution area.

[0080] The electrical distance to the distribution zone corresponding to each node can be calculated based on the following formula (10):

[0081] (10)

[0082] in, Represents a node The value of the corresponding electrical distance indicator for the distribution transformer area. Represents a node The number of nodes connected in the power grid topology diagram Represents a node and the first connection The impedance value of the lines between nodes.

[0083] Step 102: Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology is divided to obtain at least two power grid sub-graphs.

[0084] Optionally, edges with coupling degrees less than the coupling degree threshold are deleted from the power grid topology graph to obtain a power grid topology graph after deletion; all connected subgraphs are determined in the power grid topology graph after deletion; and each of the connected subgraphs is determined as the power grid subgraph.

[0085] For example, dividing the power distribution network into zones based on the electricity consumption characteristics of electric heating helps optimize electric heating control strategies. Traditional methods cannot simultaneously consider the user's thermal comfort needs (temperature requirements) and electrical characteristics. Therefore, a comprehensive thermal zone division is performed based on graph theory. When obtaining the coupling degree of each edge in the power grid topology graph, the coupling degree of each edge is compared with a coupling degree threshold. Edges with coupling degrees less than the coupling degree threshold are selected from all comparison results and deleted from the power grid topology graph, resulting in a deleted power grid topology graph. Connected subgraphs are then separated from the deleted power grid topology graph, and each connected subgraph is treated as an independent power grid subgraph. It is understood that the coupling degree of edges within the same power grid subgraph is greater than or equal to the coupling degree threshold, while the coupling degree between different power grid subgraphs is less than the coupling degree threshold. This invention integrates multiple dimensions of indicators, such as resource density, residential area ratio, indoor temperature, thermal resistance, concentration of electric heating, and electrical distance, of the distribution substation corresponding to the node to construct a power grid topology map. It breaks through the limitations of traditional methods that only consider a single electrical factor, and realizes a distribution network area division that is more in line with the user's thermal comfort needs and the electrical characteristics of electric heating, laying the foundation for subsequent refined regulation.

[0086] Step 103: For each of the power grid sub-graphs, obtain the net load power of each target node in the power grid sub-graph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time.

[0087] For example, for each power grid sub-graph, the net load power of each target node in the power grid sub-graph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time are obtained. Here, the net load power is used to reflect the source-load balance of the target node at the current time, and the target indoor temperature can be the average indoor temperature of all buildings in the distribution sub-area corresponding to the target node at the current time. The specific methods for obtaining the net load power and the target indoor temperature can be referred to relevant technologies, and will not be elaborated here.

[0088] Step 104: Input the net load power and the target indoor temperature of each item into the electric heating optimization and control network to obtain the electric heating control optimization strategy of the target sub-region output by the electric heating optimization and control network. The electric heating control optimization strategy includes the power of the electric heating load in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region.

[0089] For example, the net load power of each target node in the power grid subgraph at the current moment and the target indoor temperature of the corresponding distribution sub-area at the current moment are input as state space information into a pre-trained electric heating optimization and control network. The electric heating optimization and control network performs feature analysis on the net load power of each target node in the power grid subgraph at the current moment and the target indoor temperature of the corresponding distribution sub-area at the current moment. Finally, the electric heating optimization and control network outputs the electric heating control optimization strategy for the target sub-region. The electric heating control optimization strategy includes the power of the electric heating load in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region. This enables the energy storage system in the target sub-region to operate based on the calculated charging and discharging power, and enables the electric heating in the target sub-region to operate based on the calculated power, thereby ensuring the safe and stable operation of the power grid in the target sub-region and improving the flexibility and economy of power grid operation.

[0090] The present invention provides an optimization method for electric heating control strategies. Based on the distribution network structure within the target area, a power grid topology diagram is constructed. Nodes in the power grid topology diagram represent buses in the distribution network structure, and edges represent the coupling degree between two nodes connected to the edges. The coupling degree is determined based on indicators such as resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution sub-regions corresponding to each node. Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology diagram is divided into at least two sub-graphs. For each sub-graph, the net load power of each target node at the current time and the target indoor temperature of the corresponding distribution sub-region are obtained. These net load powers and target indoor temperatures are then input into the electric heating optimization control network to obtain the electric heating control optimization strategy for the target sub-region output by the electric heating optimization control network. As can be seen, the present invention can divide the power grid topology of the entire target area based on multi-dimensional indicators such as resource density, living area ratio, indoor temperature, thermal resistance, concentration of electric heating, and electrical distance of the distribution sub-area corresponding to each node. This makes the temperature requirements of users in the sub-areas corresponding to the divided power grid sub-maps similar. Then, for each power grid sub-map, based on the net load power of each target node in the power grid sub-map and the target indoor temperature of the distribution sub-area corresponding to each target node, the electric heating control optimization strategy in the target sub-area is determined, realizing individual control of each target sub-area, thereby improving the accuracy of the determined control optimization strategy.

[0091] In one embodiment, the coupling degree of each side is calculated as follows:

[0092] For each of the aforementioned indicators, the information entropy of the indicator is determined based on formula (1):

[0093] (1)

[0094] in, Indicates the first The value of each indicator, Indicates the first Information entropy of each indicator Indicates the total number of indicators;

[0095] Based on the information entropy of the indicator, the weight of the indicator is determined using formula (2):

[0096] (2)

[0097] in, Indicates the first The weight of each indicator;

[0098] Based on the respective index values ​​of the two nodes and the weights of each index, the coupling degree between the two nodes is determined using formula (3):

[0099] (3)

[0100] in, Represents a node With nodes The degree of coupling between them Represents a node The The value of each indicator, Represents a node The The value of each indicator.

[0101] For example, Table 1 lists all the indicators included in the indicator system. As shown in Table 1, these include urban indicators, thermal environment indicators, and electricity indicators. Urban indicators include the resource density of the distribution sub-area corresponding to the node and the residential area ratio of the sub-area. Thermal environment indicators include the indoor temperature of the distribution sub-area corresponding to the node and the thermal resistance of the sub-area. Electricity indicators include the concentration of electric heating in the distribution sub-area corresponding to the node and the electrical distance to the sub-area. The data for each indicator are normalized to obtain the normalized value of each indicator, which is the indicator value for each individual indicator. Indicates the first The value of the indicator will be the first indicator. Substituting the value of the first indicator into the formula (1) above, we can calculate the value of the second indicator. The information entropy of the first indicator, and then the information entropy of the second indicator. Substituting the information entropy of the first indicator into the above formula (2), we can calculate the first... The weights of each indicator can be calculated for each node using the same method. The The indicator value and node of each indicator Weights of each metric, nodes The The indicator value and node of each indicator The weight of each indicator is substituted into the above formula (3) to calculate the node. With nodes The coupling degree between nodes can be calculated using the same method, for every two nodes with an edge.

[0102] Table 1

[0103]

[0104] It should be noted that the indicator system may include other indicators, and this invention does not limit this.

[0105] In this embodiment, the coupling degree between the two nodes is determined based on their respective index values ​​and the weights of each index, thus achieving automatic determination of the coupling degree.

[0106] In one embodiment, Figure 2 This is a schematic diagram of the training method for the electric heating optimization and control network provided in an embodiment of the present invention, as shown below. Figure 2 As shown, the electric heating optimization and control network is trained in the following way:

[0107] Step 201: Obtain the net load power of each sample node in the sample area at the time of acquisition and the indoor temperature of the distribution substation corresponding to each sample node at the time of acquisition. The sample node is used to characterize the bus in the sample distribution network structure in the sample area.

[0108] The sample area can be the target area or other areas using electric heating; this invention does not limit this.

[0109] For example, a sample power grid topology map of the sample area is constructed in a manner similar to step 101 above, and the sample power grid topology map is divided in a manner similar to step 102 above to obtain multiple sample power grid sub-maps of the sample power grid topology map. The sample net load power of each sample node in the sample power grid sub-map at the time of acquisition and the sample indoor temperature of the distribution sub-area corresponding to each sample node at the time of acquisition are obtained. Here, the sample net load power is used to reflect the source-load balance of the sample node at the time of acquisition, and the sample indoor temperature can be the average indoor temperature of all buildings in the distribution sub-area corresponding to the sample node at the time of acquisition. The specific methods for obtaining the sample net load power and sample indoor temperature can be referred to relevant technologies, and will not be elaborated here.

[0110] Step 202: Input the net load power of each sample and the indoor temperature of each sample at the acquisition time as state space information into the actor network to obtain the predicted action value of the acquisition time output by the actor network. The predicted action value includes the predicted power of electric heating in the sample sub-region and the predicted charging and discharging power of the energy storage system in the sample sub-region.

[0111] The actor network, within the Deep Deterministic Policy Gradient (DDPG) framework, is a policy network that primarily outputs an action based on the current environment's state space information. It receives state space information as input, calculates the action value in the action space through forward propagation of the neural network, and outputs the action value. The state space is the set of variables used by the agent to perceive the environmental state. In the electric heating optimization and control scenario, the agent can act as a controller in the power system, using the net load power of each node at different times and the indoor temperature of the corresponding distribution substation as state space information to reflect the overall current state of the power system within the target area. The action space is the set of all possible actions the agent can take. In the electric heating optimization and control scenario, the power of electric heating and the charging and discharging power of the energy storage system are used as action space variables, i.e., action values.

[0112] For example, the net load power of each sample node in the sample power grid subgraph at the time of acquisition and the indoor temperature of the corresponding distribution sub-area of ​​each sample node are used as state space information and input into the actor network. The actor network performs feature analysis on the net load power of each sample node in the sample power grid subgraph at the time of acquisition and the indoor temperature of the corresponding distribution sub-area of ​​each sample node. Finally, the predicted action value of the sample sub-region calculated by the actor network through forward propagation is obtained. The predicted action value includes the predicted power of electric heating in the sample sub-region and the predicted charging and discharging power of the energy storage system in the sample sub-region.

[0113] Step 203: Input the state space information and the predicted action value into the critic network to obtain the value estimate output by the critic network.

[0114] The critic network evaluates the value of the action values ​​output by the actor network in the current state. It receives state space information and action values ​​as input, and the neural network calculates the value estimate (Q-value) of the state-action pair, which is the value estimate corresponding to the state space information and the predicted action value.

[0115] For example, the state space information and corresponding predicted action value at the acquisition time are input into the critic network to obtain the value estimate of the acquisition time output by the critic network. The state space information at the acquisition time, the predicted action value at the acquisition time, the reward value at the acquisition time, and the state space information of the next time step at the acquisition time are used as a quadruple. The quadruple is stored as experience data in the experience replay pool. During training, a batch of experience data is randomly selected from the experience replay pool to update the parameters of the actor network and the critic network. This can break the correlation of data and improve learning efficiency.

[0116] Step 204: With the goal of maximizing value estimation, adjust the parameters of the actor network until the convergence condition is met, and determine the final actor network as the electric heating optimization and control network.

[0117] For example, in the scenario of electric heating optimization and control, based on the actor network continuously interacting with the environment, with the goal of maximizing the value estimate (i.e., long-term cumulative reward) of the critic network's output, the parameters of the actor network (such as the weights and biases of the fully connected network) are dynamically adjusted using the policy gradient method. The specific process is as follows: In each round of training, the actor network generates the corresponding predicted action value based on the current state space information. The critic network evaluates the Q value of the predicted action value and feeds back the gradient signal. The actor network gradually improves the action value through the gradient ascent optimization strategy. During the training process, the learning process is stabilized by combining random sampling of the experience replay pool and soft update technology of the target network until the preset convergence condition is met (such as the average reward fluctuation being less than 2% for 100 consecutive rounds or the number of training rounds exceeding 1000). At this time, the training is terminated and the parameters of the actor network are frozen. The finally converged actor network is deployed as the electric heating optimization and control network.

[0118] It should be noted that the target networks refer to the target actor network and the target critic network, used to increase the stability and convergence of training. The target actor network has the same structure as the actor network, and the target critic network has the same structure as the critic network. Parameters are updated via soft updates every preset number of training steps. The target actor network provides a relatively stable target action distribution, and the target critic network provides a relatively stable target Q-value estimate.

[0119] In this embodiment, an electric heating optimization and control network is trained based on an actor network and a critic network. Addressing the challenges of high uncertainty and difficulty in fine-grained modeling of novel power systems, the trained deep deterministic policy gradient algorithm framework is used as the electric heating optimization and control network. This network directly interacts with the environment to solve for the optimal electric heating control strategy, enabling it to autonomously output the optimal strategy based on environmental conditions and improving the accuracy of the optimization.

[0120] In one embodiment, the actor network includes a long short-term memory network and a fully connected network; step 202 above inputs the net load power of each sample and the indoor temperature of each sample at the acquisition time as state space information into the actor network to obtain the predicted action value of the acquisition time output by the actor network, which can be implemented in the following way:

[0121] The state space information at the acquisition time and the hidden state information of the previous acquisition time are respectively input into the forget gate, input gate, memory update gate, and output gate of the Long Short-Term Memory network to obtain a first result output by the forget gate, a second result output by the input gate, a third result output by the memory update gate, and a fourth result output by the output gate. Based on the first result and the second result, the third result of the memory update gate is updated to obtain an updated third result. Based on the fourth result and the updated third result, the hidden state information at the acquisition time is determined. The hidden state information at the acquisition time is input into the fully connected network to obtain the predicted action value output by the fully connected network.

[0122] For example, the state space contains node data at multiple time points, resulting in high variable dimensionality and information redundancy, which increases the risk of overfitting and the curse of dimensionality. Considering the significant temporal nature of state space information, the neural network of the actor network in related technologies is replaced with a long short-term memory network. The long short-term memory network extracts features from the input state space information to obtain a temporal feature vector. The extracted temporal feature vector is then flattened and input into a fully connected network to learn spatial relationships.

[0123] Specifically, the hidden state information at the acquisition time can be obtained through the following formulas (11) to (16):

[0124]

[0125] in, Indicates the acquisition time of the forget gate output. The first result Indicates the acquisition time of the input gate output. The second result, Indicates the acquisition time of the memory update gate output. The third result, Indicates the acquisition time of the output gate output. The fourth result, The weight matrix represents the forget gate. The bias term representing the forget gate. This represents the weight matrix of the input gate. This represents the bias term of the input gate. This represents the weight matrix of the memory update gate. This represents the bias term of the memory update gate. This represents the weight matrix of the memory output gate. This represents the bias term of the output gate. This indicates that the hidden state information from the previous moment will be displayed. and state-space information at the time of acquisition Splicing in the horizontal direction. This represents the sigmoid function. Represents the hyperbolic tangent function. Indicates the previous moment As a result, This indicates the third result after the update. The hidden state information at the acquisition time is used as the extracted time feature vector. That is, the hidden state information at the acquisition time is input into the fully connected network to obtain the predicted action value output by the fully connected network.

[0126] It should be noted that before training the electric heating optimization and control network, the sample data needs to be preprocessed and normalized. This involves cleaning the sample net load power and sample indoor temperature, removing outliers and missing data to ensure data quality. Considering the spatiotemporal coupling characteristics of the information in each state space, the sample data is arranged with rows representing node numbers and columns representing time. Each row of data is considered as a time series, corresponding to the state change of a node over time. Normalization is performed on each row of data to obtain normalized data. The normalized sample net load power and normalized sample indoor temperature are then input into the Long Short-Term Memory (LSTM) network to obtain the extracted time feature vector, reducing the information redundancy of the input data.

[0127] It should be noted that the data format for inputting the actor network can be (batch_size, num_nodes, num_timesteps, num_variables), where batch_size represents the number of samples randomly drawn from the experience replay pool each time the parameters are updated, num_nodes represents the total number of sample nodes, num_timesteps represents the total duration of all acquisition times, and num_variables represents the stacking of sample net load power and sample room temperature along the feature dimension to adapt to the input requirements of long short-term memory networks in deep reinforcement learning frameworks.

[0128] It should be noted that when the actor network includes a long short-term memory network and a fully connected network, it is also necessary to determine key parameters such as the number of layers and hidden units of the long short-term memory network and the fully connected network, set the input dimension to adapt to the state space information, and set the output dimension to adapt to the action value. These will not be elaborated here.

[0129] In this embodiment, a Long Short-Term Memory (LSTM) network is introduced into the deep deterministic policy gradient algorithm framework to effectively address the high dimensionality and temporal complexity of state-space information and reduce the redundancy of the network input data. Specifically, this invention, considering the spatiotemporal uncertainties of electric heating and the differences in user thermal comfort requirements, divides the distribution network into target areas based on the power grid topology. Within different sub-regions, a deep deterministic policy gradient algorithm using LSM networks is employed to obtain optimized electric heating control strategies, thereby improving the renewable energy absorption capacity of the distribution network and reducing its operating costs.

[0130] In one embodiment, step 204 above aims to maximize value estimation by adjusting the parameters of the actor network until convergence is achieved. The resulting actor network is then determined as the electric heating optimization and control network. This can be implemented in the following ways:

[0131] Based on the predicted action values ​​and the action value labels corresponding to the state space information, a loss function is constructed. With the goal of maximizing value estimation, the weight matrices and bias terms of the forget gate, the input gate, the memory update gate, and the output gate are adjusted according to the loss function until the convergence condition is met. The final actor network is then determined as the electric heating optimization and control network.

[0132] For example, when the predicted action value at the acquisition time is obtained, a loss function is constructed based on the predicted action value at the acquisition time and the action value label corresponding to the state space information at the acquisition time. Specifically, the loss function can be calculated based on the relative squared error. The performance of the actor network under the current parameter settings is quantified through the loss function, that is, the difference between the predicted action value output by the actor network and the expected action value label is quantified. The weight matrices and bias terms of the forget gate, input gate, memory update gate and output gate are adjusted through backpropagation. The specific adjustment principle of the weight matrix and bias term is: once the gradient of all parameters is obtained, the weight matrix and bias term can be updated in the direction of gradient descent. That is, if an increase in a certain weight matrix leads to an increase in loss, then the weight matrix should be decreased, and vice versa. This is repeated iteratively, and each iteration will try to further reduce the loss until the performance of the actor network reaches the preset level or no longer improves significantly. This indicates that the convergence condition has been met, and the final actor network is determined as the electric heating optimization and control network.

[0133] It's important to note that the step size used during the update process is determined by the learning rate. Choosing an appropriate learning rate is crucial because it determines the speed and stability of the weight matrix and bias term updates. An excessively large learning rate may lead to unstable training or failure to converge, while an excessively small learning rate may make the training process very slow. Therefore, it is necessary to set an appropriate learning rate during training.

[0134] It should be noted that, in addition to gradient descent, optimization algorithms such as Adam and RMSprop can also be used. These algorithms can adaptively adjust the learning rate and take into account the historical information of previous gradients to better guide the update process of the weight matrix and bias terms. These will not be elaborated on here.

[0135] In this embodiment, a loss function is constructed based on the predicted action value and the action value label corresponding to the state space information. With the goal of maximizing value estimation, the weight matrices and bias terms of the forget gate, input gate, memory update gate, and output gate are adjusted based on the loss function until the convergence condition is met. The final actor network is determined as the electric heating optimization and control network, so that the trained electric heating optimization and control network can autonomously output the optimal electric heating control optimization strategy according to the environmental state, thereby improving the accuracy of electric heating control strategy optimization.

[0136] In one embodiment, step 203 above inputs the state space information and the predicted action value into the critic network to obtain the value estimate output by the critic network, which can be specifically implemented in the following way:

[0137] The state space information and the predicted action value are input into the critic network, and the critic network determines the value estimate based on the reward value.

[0138] The reward value is determined based on formula (4):

[0139] (4)

[0140] in, Indicates the reward value. Indicates the operating cost of the power distribution network. This indicates a deviation in the absorption of new energy sources. This indicates the weight corresponding to the operating cost of the distribution network. This indicates the weight corresponding to the deviation in the absorption of new energy sources. ,

[0141] , Indicates the first The cost of purchasing electricity from the upper-level power grid at the time of data collection. Indicates the first The regulation cost of the energy storage system at the time of data acquisition. Indicates the first The control cost of electric heating at the time of data collection. Indicates the first The first acquisition time The penalty cost of each constraint condition. Represents a set of constraints. Indicates the first The power output of the power distribution network at the time of data collection. Indicates the first Load difference at the time of data acquisition Indicates the total number of nodes. This represents the total duration of all data collection moments.

[0142] For example, when the critic network calculates the value estimate at the acquisition time, it needs to calculate based on the reward value at the acquisition time. The specific calculation can be referred to in related technologies, and will not be repeated here. In addition, the reward function is a function used to evaluate the quality of the action taken by the agent. Based on the effect of the agent's action in the environment, a reward value is returned, which is used to guide the learning process of the agent. In the scenario of electric heating optimization and control, since the actor network in DDPG will learn how to maximize the reward, the negative of the weighted average of the distribution network operating cost and the new energy consumption deviation is used as the reward value. The weights of the distribution network operating cost and the new energy consumption deviation can be determined by the principal component analysis method. The specific calculation of the reward value is shown in the above formula (4).

[0143] In this embodiment, the state space information and predicted action values ​​are input into the critic network, and the critic network determines the value estimate based on the reward value, thereby improving the accuracy of the value estimate calculation.

[0144] The following describes the optimization device for the electric heating control strategy provided by the present invention. The optimization device for the electric heating control strategy described below and the optimization method for the electric heating control strategy described above can be referred to in correspondence.

[0145] Figure 3 This is a schematic diagram of the structure of the optimization device for electric heating control strategy provided in an embodiment of the present invention, as shown below. Figure 3 As shown, the optimization device 300 for the electric heating control strategy includes a construction unit 301, a division unit 302, an acquisition unit 303, and a control unit 304; wherein:

[0146] The construction unit 301 is used to construct a power grid topology map based on the distribution network structure in the target area. The nodes in the power grid topology map are used to represent the buses in the distribution network structure, and the edges in the power grid topology map are used to represent the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in the index system. The indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration and electrical distance of the distribution substation corresponding to the node.

[0147] The partitioning unit 302 is used to partition the power grid topology based on the coupling degree threshold and the coupling degree corresponding to each edge, to obtain at least two power grid sub-graphs.

[0148] The acquisition unit 303 is used to acquire, for each of the power grid subgraphs, the net load power of each target node in the power grid subgraph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time;

[0149] The control unit 304 is used to input the net load power and the target indoor temperature into the electric heating optimization control network to obtain the electric heating control optimization strategy in the target sub-region output by the electric heating optimization control network. The electric heating control optimization strategy includes the power of the electric heating load in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region.

[0150] The electric heating control strategy optimization device provided by this invention constructs a power grid topology map based on the distribution network structure within the target area. Nodes in the power grid topology map represent buses in the distribution network structure, and edges represent the coupling degree between two nodes connected to the edges. The coupling degree is determined based on indicators such as resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution sub-area corresponding to each node. Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology map is divided into at least two power grid sub-maps. For each power grid sub-map, the net load power of each target node at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time are obtained. These net load powers and target indoor temperatures are then input into the electric heating optimization control network to obtain the electric heating control optimization strategy for the target sub-area output by the electric heating optimization control network. As can be seen, the present invention can divide the power grid topology of the entire target area based on multi-dimensional indicators such as resource density, living area ratio, indoor temperature, thermal resistance, concentration of electric heating, and electrical distance of the distribution sub-area corresponding to each node. This makes the temperature requirements of users in the sub-areas corresponding to the divided power grid sub-maps similar. Then, for each power grid sub-map, based on the net load power of each target node in the power grid sub-map and the target indoor temperature of the distribution sub-area corresponding to each target node, the electric heating control optimization strategy in the target sub-area is determined, realizing individual control of each target sub-area, thereby improving the accuracy of the determined control optimization strategy.

[0151] Based on any of the above embodiments, the partitioning unit 302 is specifically used for:

[0152] Delete edges with coupling degrees less than the coupling degree threshold from the power grid topology graph to obtain the power grid topology graph after deletion;

[0153] In the deleted power grid topology graph, identify all connected subgraphs;

[0154] Each of the connected subgraphs is defined as the power grid subgraph.

[0155] Based on any of the above embodiments, the electric heating control strategy optimization device 300 further includes:

[0156] The first determining unit is used to determine the information entropy of each indicator based on formula (1):

[0157] (1)

[0158] in, Indicates the first The value of each indicator, Indicates the first Information entropy of each indicator Indicates the total number of indicators;

[0159] The second determining unit is used to determine the weight of the indicator based on the information entropy of the indicator using formula (2):

[0160] (2)

[0161] in, Indicates the first The weight of each indicator;

[0162] The third determining unit is used to determine the coupling degree between the two nodes based on their respective index values ​​and the weights of each index, using formula (3):

[0163] (3)

[0164] in, Represents a node With nodes The degree of coupling between them Represents a node The The value of each indicator, Represents a node The The value of each indicator.

[0165] Based on any of the above embodiments, the electric heating optimization and control network is trained in the following manner:

[0166] The net load power of each sample node in the sample area at the time of acquisition and the indoor temperature of the distribution substation corresponding to each sample node at the time of acquisition are obtained. The sample node is used to characterize the bus in the sample distribution network structure in the sample area.

[0167] The net load power and indoor temperature of each sample at the acquisition time are used as state space information and input into the actor network to obtain the predicted action value of the acquisition time output by the actor network. The predicted action value includes the predicted power of electric heating in the sample sub-region and the predicted charging and discharging power of the energy storage system in the sample sub-region.

[0168] The state space information and the predicted action value are input into the critic network to obtain the value estimate output by the critic network;

[0169] With the goal of maximizing value estimation, the parameters of the actor network are adjusted until the convergence condition is met, and the final actor network is determined as the electric heating optimization and control network.

[0170] Based on any of the above embodiments, the actor network includes a long short-term memory network and a fully connected network;

[0171] The step of inputting the net load power and indoor temperature of each sample at the acquisition time as state-space information into the actor network to obtain the predicted action value at the acquisition time output by the actor network includes:

[0172] The state space information at the acquisition time and the hidden state information at the previous acquisition time are respectively input into the forget gate, input gate, memory update gate and output gate of the long short-term memory network to obtain the first result output by the forget gate, the second result output by the input gate, the third result output by the memory update gate and the fourth result output by the output gate.

[0173] Based on the first result and the second result, the third result of the memory update gate is updated to obtain the updated third result;

[0174] Based on the fourth result and the updated third result, the hidden state information at the acquisition time is determined;

[0175] The hidden state information at the acquisition time is input into the fully connected network to obtain the predicted action value output by the fully connected network.

[0176] Based on any of the above embodiments, the step of adjusting the parameters of the actor network with the objective of maximizing value estimation until the convergence condition is met, and determining the final actor network as the electric heating optimization and control network, includes:

[0177] Based on the predicted action value and the action value label corresponding to the state space information, a loss function is constructed;

[0178] With the goal of maximizing value estimation, the weight matrices and bias terms of the forget gate, the input gate, the memory update gate and the output gate are adjusted based on the loss function until the convergence condition is met, and the final actor network is determined as the electric heating optimization and control network.

[0179] Based on any of the above embodiments, the step of inputting the state space information and the predicted action value into the critic network to obtain the value estimate output by the critic network includes:

[0180] The state space information and the predicted action value are input into the critic network, and the critic network determines the value estimate based on the reward value.

[0181] The reward value is determined based on formula (4):

[0182] (4)

[0183] in, Indicates the reward value. Indicates the operating cost of the power distribution network. This indicates a deviation in the absorption of new energy sources. This indicates the weight corresponding to the operating cost of the distribution network. This indicates the weight corresponding to the deviation in the absorption of new energy sources. ,

[0184] , Indicates the first The cost of purchasing electricity from the upper-level power grid at the time of data collection. Indicates the first The regulation cost of the energy storage system at the time of data acquisition. Indicates the first The control cost of electric heating at the time of data collection. Indicates the first The first acquisition time The penalty cost of each constraint condition. Represents a set of constraints. Indicates the first The power output of the power distribution network at the time of data collection. Indicates the first Load difference at the time of data acquisition Indicates the total number of nodes. This represents the total duration of all data collection moments.

[0185] Figure 4 This is a schematic diagram of the physical structure of the electronic device provided in the embodiments of the present invention, such as... Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute an optimization method for electric heating control strategy. The method includes: constructing a power grid topology map based on the distribution network structure in the target area, wherein the nodes in the power grid topology map are used to represent the buses in the distribution network structure, and the edges in the power grid topology map are used to represent the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in an index system, wherein the indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node.

[0186] Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology is divided to obtain at least two power grid sub-graphs.

[0187] For each of the aforementioned power grid sub-graphs, obtain the net load power of each target node within the power grid sub-graph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time;

[0188] The net load power and the target indoor temperature are input into the electric heating optimization and control network to obtain the electric heating control optimization strategy for the target sub-region output by the electric heating optimization and control network. The electric heating control optimization strategy includes the electric heating power in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region.

[0189] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0190] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the optimization method of the electric heating control strategy provided by the above methods. The method includes: constructing a power grid topology map based on the distribution network structure in the target area. The nodes in the power grid topology map are used to represent the buses in the distribution network structure. The edges in the power grid topology map are used to represent the coupling degree between two nodes connected to the edges. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in the index system. The indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node.

[0191] Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology is divided to obtain at least two power grid sub-graphs.

[0192] For each of the aforementioned power grid sub-graphs, obtain the net load power of each target node within the power grid sub-graph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time;

[0193] The net load power and the target indoor temperature are input into the electric heating optimization and control network to obtain the electric heating control optimization strategy for the target sub-region output by the electric heating optimization and control network. The electric heating control optimization strategy includes the electric heating power in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region.

[0194] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an optimization method for performing the electric heating control strategies provided by the above methods. The method includes: constructing a power grid topology diagram based on the distribution network structure within a target area, wherein nodes in the power grid topology diagram are used to characterize buses in the distribution network structure, and edges in the power grid topology diagram are used to characterize the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in an index system, wherein the indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node.

[0195] Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology is divided to obtain at least two power grid sub-graphs.

[0196] For each of the aforementioned power grid sub-graphs, obtain the net load power of each target node within the power grid sub-graph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time;

[0197] The net load power and the target indoor temperature are input into the electric heating optimization and control network to obtain the electric heating control optimization strategy for the target sub-region output by the electric heating optimization and control network. The electric heating control optimization strategy includes the electric heating power in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region.

[0198] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0199] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0200] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An optimization method for electric heating control strategy, characterized in that, include: Based on the distribution network structure within the target area, a power grid topology diagram is constructed. Nodes in the power grid topology diagram are used to represent the busbars in the distribution network structure, and edges in the power grid topology diagram are used to represent the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in the index system. The indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node. Based on the coupling degree threshold and the coupling degree corresponding to each edge, the power grid topology is divided to obtain at least two power grid sub-graphs. For each of the aforementioned power grid sub-graphs, obtain the net load power of each target node within the power grid sub-graph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time; The net load power and the target indoor temperature are input into the electric heating optimization and control network to obtain the electric heating control optimization strategy in the target sub-region output by the electric heating optimization and control network. The electric heating control optimization strategy includes the electric heating power in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region. The method further includes: For each of the aforementioned indicators, the information entropy of the indicator is determined based on formula (1): (1) in, Indicates the first The value of each indicator, Indicates the first Information entropy of each indicator Indicates the total number of indicators; Based on the information entropy of the indicator, the weight of the indicator is determined using formula (2): (2) in, Indicates the first The weight of each indicator; Based on the respective index values ​​of the two nodes and the weights of each index, the coupling degree between the two nodes is determined using formula (3): (3) in, Represents a node With nodes The degree of coupling between them Represents a node The The value of each indicator, Represents a node The The value of each indicator.

2. The optimization method for electric heating control strategy according to claim 1, characterized in that, The power grid topology is divided based on a coupling degree threshold and the coupling degree corresponding to each edge to obtain at least two power grid sub-graphs, including: Delete edges with coupling degrees less than the coupling degree threshold from the power grid topology graph to obtain the power grid topology graph after deletion; In the deleted power grid topology graph, identify all connected subgraphs; Each of the connected subgraphs is defined as the power grid subgraph.

3. The optimization method for electric heating control strategy according to claim 1, characterized in that, The electric heating optimization and control network was trained using the following method: The net load power of each sample node in the sample area at the time of acquisition and the indoor temperature of the distribution substation corresponding to each sample node at the time of acquisition are obtained. The sample node is used to characterize the bus in the sample distribution network structure in the sample area. The net load power and indoor temperature of each sample at the acquisition time are used as state space information and input into the actor network to obtain the predicted action value of the acquisition time output by the actor network. The predicted action value includes the predicted power of electric heating in the sample sub-region and the predicted charging and discharging power of the energy storage system in the sample sub-region. The state space information and the predicted action value are input into the critic network to obtain the value estimate output by the critic network; With the goal of maximizing value estimation, the parameters of the actor network are adjusted until the convergence condition is met, and the final actor network is determined as the electric heating optimization and control network.

4. The optimization method for electric heating control strategy according to claim 3, characterized in that, The actor network includes a long short-term memory network and a fully connected network; The net load power and indoor temperature of each sample at the acquisition time are input as state-space information into the actor network to obtain the predicted action value at the acquisition time output by the actor network, including: The state space information at the acquisition time and the hidden state information at the previous acquisition time are respectively input into the forget gate, input gate, memory update gate and output gate of the long short-term memory network to obtain the first result output by the forget gate, the second result output by the input gate, the third result output by the memory update gate and the fourth result output by the output gate. Based on the first result and the second result, the third result of the memory update gate is updated to obtain the updated third result; Based on the fourth result and the updated third result, the hidden state information at the acquisition time is determined; The hidden state information at the acquisition time is input into the fully connected network to obtain the predicted action value output by the fully connected network.

5. The optimization method for electric heating control strategy according to claim 4, characterized in that, The process of adjusting the parameters of the actor network with the goal of maximizing value estimation until convergence is achieved, and determining the final actor network as the electric heating optimization and control network, includes: Based on the predicted action value and the action value label corresponding to the state space information, a loss function is constructed; With the goal of maximizing value estimation, the weight matrices and bias terms of the forget gate, the input gate, the memory update gate and the output gate are adjusted based on the loss function until the convergence condition is met, and the final actor network is determined as the electric heating optimization and control network.

6. The optimization method for electric heating control strategy according to claim 3, characterized in that, The step of inputting the state space information and the predicted action value into the critic network to obtain the value estimate output by the critic network includes: The state space information and the predicted action value are input into the critic network, and the critic network determines the value estimate based on the reward value. The reward value is determined based on formula (4): (4) in, Indicates the reward value. Indicates the operating cost of the power distribution network. This indicates a deviation in the absorption of new energy sources. This indicates the weight corresponding to the operating cost of the distribution network. This indicates the weight corresponding to the deviation in the absorption of new energy sources. , , Indicates the first The cost of purchasing electricity from the upper-level power grid at the time of data collection. Indicates the first The regulation cost of the energy storage system at the time of data acquisition. Indicates the first The control cost of electric heating at the time of data collection. Indicates the first The first acquisition time The penalty cost of each constraint condition. Represents a set of constraints. Indicates the first The power output of the power distribution network at the time of data collection. Indicates the first Load difference at the time of data acquisition Indicates the total number of nodes. This represents the total duration of all data collection moments.

7. An optimization device for electric heating control strategy, characterized in that, include: A construction unit is used to construct a power grid topology map based on the distribution network structure within a target area. Nodes in the power grid topology map represent buses in the distribution network structure, and edges in the power grid topology map represent the coupling degree between two nodes connected to the edge. The coupling degree is determined based on the index values ​​of the respective indicators of the two nodes in an index system. The indicators include the resource density, residential area ratio, indoor temperature, thermal resistance, electric heating concentration, and electrical distance of the distribution substation corresponding to the node. A partitioning unit is used to partition the power grid topology based on a coupling degree threshold and the coupling degree corresponding to each edge, to obtain at least two power grid subgraphs; The acquisition unit is used to acquire, for each of the power grid subgraphs, the net load power of each target node in the power grid subgraph at the current time and the target indoor temperature of the distribution sub-area corresponding to each target node at the current time; The control unit is used to input the net load power and the target indoor temperature into the electric heating optimization control network to obtain the electric heating control optimization strategy in the target sub-region output by the electric heating optimization control network. The electric heating control optimization strategy includes the power of the electric heating load in the target sub-region and the charging and discharging power of the energy storage system in the target sub-region. The device further includes: The first determining unit is used to determine the information entropy of each indicator based on formula (1): (1) in, Indicates the first The value of each indicator, Indicates the first Information entropy of each indicator Indicates the total number of indicators; The second determining unit is used to determine the weight of the indicator based on the information entropy of the indicator using formula (2): (2) in, Indicates the first The weight of each indicator; The third determining unit is used to determine the coupling degree between the two nodes based on their respective index values ​​and the weights of each index, using formula (3): (3) in, Represents a node With nodes The degree of coupling between them Represents a node The The value of each indicator, Represents a node The The value of each indicator.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the optimization method of the electric heating control strategy as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the optimization method of the electric heating control strategy as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Node coupling degree analysis method based on injection currents

    CN104716646A

  • High-proportion new energy power grid island division method considering frequency constraint

    CN119651563A