Active power distribution network energy storage partition control network training method and device and control method
By setting up energy storage zones in the distribution network and optimizing the scheduling strategy using policy networks and evaluation networks, the problem of unstable operation of energy storage devices under different temperature conditions is solved, thereby improving the scheduling efficiency and security of the distribution network.
Patent Information
- Application Number
- CN202511517295.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-23
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2045-10-23
AI Technical Summary
In existing technologies, the temperature impact of energy storage devices in distributed power distribution networks is not adequately considered, leading to unstable operation of energy storage devices in different seasons, which can easily cause thermal runaway or capacity reduction, affecting the dispatch efficiency and safety of the power distribution network.
By setting up multiple energy storage zones in the distribution network, using policy networks and evaluation networks for training, and considering the temperature-related constraints of energy storage devices, the scheduling strategy is optimized, including parameter updates for action evaluation and target evaluation networks, thereby improving the accuracy of action evaluation.
It enables optimized scheduling of energy storage devices under different temperature conditions, improves the stability and security of the power distribution network, reduces the risk of thermal runaway of energy storage devices, and enhances the utilization efficiency of energy storage devices.
Smart Images

Figure CN120999641B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power distribution network dispatching technology, and more specifically, to a power distribution network control network training method, device, and control method for a power distribution network. Background Technology
[0002] With the development of technology, power distribution networks are gradually evolving into distributed structures. As the proportion of photovoltaic equipment in power distribution networks increases and fluctuating loads such as electric vehicle charging stations are connected on a large scale, power distribution networks are gradually shifting towards an active distribution model. However, with this transformation, the control and scheduling methods used in related technologies for power distribution networks are no longer applicable. Summary of the Invention
[0003] In view of this, this application provides a control network training method, device, and control method for a power distribution network.
[0004] One aspect of this application provides a method for training a control network of a power distribution network, the power distribution network including multiple energy storage zones, each energy storage zone being equipped with energy storage devices. The method includes: for each energy storage zone, inputting training data, the training data including a first state, a first action taken in the first state, a reward corresponding to the first action, and a second state achieved by taking the first action in the first state, wherein the first state and the second state each include the ambient temperature of the energy storage devices, and the reward represents the impact of the first action on the power distribution network, including temperature-related impacts on the energy storage devices; and using a policy network, determining a target zone control action based on the first state. The target zoning control action is subject to the operational constraints of the distribution network, including temperature-related constraints for the energy storage device. A predicted action value is determined using an action evaluation network, representing the expected cumulative reward when the target zoning control action is performed in the first state. A target action value is determined using a target evaluation network, representing the expected cumulative reward when the action begins in the second state. Based on the predicted action value and the target action value, the parameters of the strategy network, the action evaluation network, and the target evaluation network are updated to control the energy storage zoning according to the updated strategy network.
[0005] Another aspect of this application provides a control method for a power distribution network, the power distribution network including multiple energy storage zones, the control method comprising: for each energy storage zone, obtaining the current state of the energy storage zone; determining a current zone control action based on the current state according to an updated policy network obtained using the control network training method described above; and controlling the energy storage zone based on the current zone control action.
[0006] Another aspect of this application provides a control network training device for a power distribution network, the power distribution network including multiple energy storage zones, each energy storage zone being equipped with energy storage devices. The control network training device includes: an input module for inputting training data for each energy storage zone, the training data including a first state, a first action taken in the first state, a reward corresponding to the first action, and a second state achieved by taking the first action in the first state, wherein the first state and the second state each include the ambient temperature of the energy storage devices, and the reward represents the impact of the first action on the power distribution network, including temperature-related effects on the energy storage devices; and a first determination module for determining a target zone control action based on the first state using a policy network, wherein... The aforementioned target zoning control action is subject to the operational constraints of the aforementioned distribution network, including temperature-related constraints for the aforementioned energy storage devices; the second determination module is used to determine the predicted action value using an action evaluation network, the predicted action value representing the expected cumulative reward when the aforementioned target zoning control action is performed in the aforementioned first state; the third determination module is used to determine the target action value using a target evaluation network, the target action value representing the expected cumulative reward when the action is performed starting from the aforementioned second state; the update module is used to update the parameters of the aforementioned strategy network, the parameters of the aforementioned action evaluation network, and the parameters of the aforementioned target evaluation network based on the aforementioned predicted action value and the aforementioned target action value, so as to control the aforementioned energy storage zoning according to the updated strategy network.
[0007] According to embodiments of this application, a method for training the control network of a distribution network is provided for training each energy storage zone within the distribution network, which can refine the control and scheduling of the distribution network. Furthermore, the impact of temperature on energy storage devices is considered during the training process, further improving the optimal scheduling of energy storage devices by the distribution network. By using two evaluation networks—an action evaluation network and a target evaluation network—to evaluate the actions executed in the first state and the actions executed in the second state respectively, the results of the target evaluation network can be used as constraints for the action evaluation network, improving the accuracy of the action evaluation network. Attached Figure Description
[0008] The above and other objects, features and advantages of this application will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0009] Figure 1 An exemplary system architecture for a control network training method applicable to a power distribution network, according to an embodiment of this application, is shown.
[0010] Figure 2 A flowchart of a control network training method for a power distribution network according to an embodiment of this application is shown.
[0011] Figure 3 A schematic diagram illustrating the relationship between the number of clusters and cluster partitioning evaluation indicators according to an embodiment of this application is shown.
[0012] Figure 4 A schematic diagram showing the relationship between temperature and capacity of an energy storage device according to an embodiment of this application is provided.
[0013] Figure 5 A schematic diagram of different states of an energy storage device according to an embodiment of this application is shown.
[0014] Figure 6 A schematic diagram of the scheduling of an energy storage device according to an embodiment of this application is shown.
[0015] Figure 7 The relationship between cumulative rewards and time steps according to embodiments of this application is illustrated.
[0016] Figure 8 A schematic block diagram of a control network training method for a power distribution network according to an embodiment of this application is shown.
[0017] Figure 9 A flowchart of a control method for a power distribution network according to an embodiment of this application is shown.
[0018] Figure 10 A schematic diagram of an energy storage zone of a distribution network according to an embodiment of this application is shown in one example.
[0019] Figure 11 A schematic diagram of the photovoltaic output and load power factor of a distribution network according to an embodiment of this application is shown in one example.
[0020] Figure 12 The diagram shows temperature curves of a distribution network according to an embodiment of this application during a summer surge in photovoltaic power and during a winter overload, in one example.
[0021] Figure 13 A graph showing the energy consumption of an energy storage device in a distribution network according to an embodiment of this application at different temperatures is shown in one example.
[0022] Figure 14 A schematic diagram illustrating the relevant changes in the distribution network under the condition of a sudden increase in photovoltaic power in the distribution network in Example 1 of this application embodiment is shown.
[0023] Figure 15 A schematic diagram illustrating the relevant changes in the distribution network under the condition of a sudden increase in photovoltaic power in the distribution network in Example 2 of this application embodiment is shown.
[0024] Figure 16 A schematic diagram illustrating the relevant changes in the distribution network under the condition of a sudden increase in photovoltaic power in the distribution network in Example 3 of this application embodiment is shown.
[0025] Figure 17 A schematic diagram illustrating the relevant changes in the distribution network in the case of a sudden increase in photovoltaic power in the distribution network in Example 4 of this application embodiment is shown.
[0026] Figure 18 A schematic diagram illustrating the relevant changes in the distribution network under load changes in Example 5 of this application embodiment is shown.
[0027] Figure 19 A schematic diagram illustrating the relevant changes in the distribution network in the case of a sudden increase in photovoltaic power in the distribution network in Example Six of this application embodiment is shown.
[0028] Figure 20 A block diagram of a control network training device for a power distribution network according to an embodiment of this application is shown. Detailed Implementation
[0029] The embodiments of this application will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of this application. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of this application for ease of explanation. However, it will be apparent that one or more embodiments may be implemented without these specific details. Furthermore, descriptions of well-known structures and technologies are omitted in the following description to avoid unnecessarily obscuring the concepts of this application.
[0030] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0031] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0032] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0033] Energy storage devices in power distribution networks can effectively manage the randomness, volatility, and intermittency caused by distributed resources due to their flexibility. They can also alleviate pressure on the distribution network through charging and discharging strategies during heavy loads. However, in some regions, the temperature difference between seasons is significant, reaching 40℃~50℃ in summer and -15℃~0℃ in winter. In summer, the high ambient temperature, combined with the heat generated by charging and discharging of energy storage devices, can easily lead to thermal runaway. In winter, the low ambient temperature can cause energy storage devices to malfunction or have reduced charging and discharging capacity, making it difficult for them to fully discharge and alleviate overload conditions during heavy loads. It is necessary to consider the impact of ambient temperature changes on energy storage devices. On the one hand, this ensures the normal operation of energy storage devices and extends their lifespan; on the other hand, energy storage devices used in power distribution networks may experience dangerous accidents such as thermal runaway.
[0034] Research and discussion on optimizing the scheduling of energy storage devices to improve renewable energy power generation are extensive. They can be mainly divided into three aspects: 1. Control methods for energy storage devices; 2. Sharing strategies for energy storage devices; 3. Modeling of energy storage devices. Control methods for energy storage devices can include mathematical methods, reinforcement learning, heuristic algorithms, etc. Sharing strategies for energy storage devices include the application of hybrid energy storage devices and the sharing of energy storage devices. Modeling of energy storage devices includes the entire life cycle of energy storage devices, degradation costs, operating costs, peak-valley arbitrage, etc. However, related technologies do not adequately consider the impact of temperature on the operational safety of energy storage devices.
[0035] Figure 1 An exemplary system architecture for applying a control network training method for power distribution networks according to embodiments of this application is shown. It should be noted that... Figure 1 The examples shown are merely examples of system architectures that can be applied to the embodiments of this application, in order to help those skilled in the art understand the technical content of this application, but do not mean that the embodiments of this application cannot be used in other devices, systems, environments or scenarios.
[0036] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 serves as a medium for providing communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0037] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 via the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social media platform software, etc. (for example only).
[0038] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0039] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (this is just an example). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0040] It should be noted that the power distribution network control network training method provided in this application embodiment can generally be executed by server 105. Correspondingly, the power distribution network control network training device provided in this application embodiment can generally be located in server 105. The power distribution network control network training method provided in this application embodiment can also be executed by a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Correspondingly, the power distribution network control network training device provided in this application embodiment can also be located in a server or server cluster that is different from server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or server 105. Alternatively, the power distribution network control network training method provided in this application embodiment can also be executed by the first terminal device 101, the second terminal device 102, or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102, or the third terminal device 103. Accordingly, the power distribution network control network training device provided in this application embodiment can also be set in the first terminal device 101, the second terminal device 102 or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102 or the third terminal device 103.
[0041] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0042] Figure 2 A flowchart of a control network training method for a power distribution network according to an embodiment of this application is shown.
[0043] like Figure 2 As shown, the distribution network includes multiple energy storage zones, each equipped with energy storage devices. The control network training method includes operations S210 to S250 for each energy storage zone.
[0044] In operation S210, input training data.
[0045] The training data includes a first state, a first action taken in the first state, a reward corresponding to the first action, and a second state achieved by taking the first action in the first state. The first state and the second state each include the ambient temperature of the energy storage device. The reward represents the impact of the first action on the distribution network and includes the temperature-related impact on the energy storage device.
[0046] When operating S220, the policy network is used to determine the target partition control action based on the first state.
[0047] The target zone control actions are subject to the operating constraints of the distribution network, including temperature-related constraints for energy storage devices.
[0048] In operating S230, the action evaluation network is used to determine the value of the predicted action.
[0049] Predicted action value represents the expected cumulative reward when the target partition control action is performed in the first state.
[0050] In operating S240, the target evaluation network is used to determine the value of the target action.
[0051] The target action value represents the expected cumulative reward when the action is performed starting from the second state.
[0052] In operation of S250, the parameters of the policy network, the action evaluation network, and the target evaluation network are updated based on the predicted action value and the target action value, so as to control the energy storage zoning according to the updated policy network.
[0053] According to an embodiment of this application, the training data can be historical data of the power distribution network. The first state can be the operating state of the energy storage zone.
[0054] According to embodiments of this application, the policy network can be a network for selecting actions. A first state can be a state of abnormal distribution network operation. The policy network can be used to determine multiple actions that can restore normal distribution network operation based on the first state, and then determine a target zone control action from among these actions. The target zone control action can be randomly determined, or actions that control only the target zone can be preferentially selected. Alternatively, the first state can be a state of normal distribution network operation. The policy network can be used to determine multiple actions that can make the distribution network operation more stable based on the first state, and then determine a target zone action from among these actions.
[0055] According to embodiments of this application, the action evaluation network can be a network used to evaluate target partition control actions performed in a first state, and the target evaluation network can be a network used to evaluate actions performed in a second state.
[0056] According to embodiments of this application, a method for training the control network of a distribution network is provided for training each energy storage zone within the distribution network, which can refine the control and scheduling of the distribution network. Furthermore, the impact of temperature on energy storage devices is considered during the training process, further improving the optimal scheduling of energy storage devices by the distribution network. By using two evaluation networks—an action evaluation network and a target evaluation network—to evaluate the actions executed in the first state and the actions executed in the second state respectively, the results of the target evaluation network can be used as constraints for the action evaluation network, improving the accuracy of the action evaluation network.
[0057] According to an embodiment of this application, the control network training method may further include: determining the electrical distance between multiple nodes of the distribution network based on the power grid data of the distribution network, wherein the energy storage device is installed at at least one node; and determining the energy storage zone of the distribution network based on the electrical distance between the multiple nodes.
[0058] According to embodiments of this application, the power grid data of the distribution network may include active power data, phase angle data, reactive power data, and voltage data, and the electrical distance between multiple nodes can be determined through the power grid data.
[0059] A modified equation can be used to replace the traditional power flow calculation, as shown in the following formula.
[0060] (1)
[0061] (2)
[0062] (3)
[0063] (4)
[0064] (5)
[0065] in, , , , These represent the active power data, reactive power data, phase angle data, and voltage data of the distribution network, respectively. , , , These are the matrix elements of the Jacobian matrix. This represents the voltage active power sensitivity matrix. This represents the voltage reactive power sensitivity matrix. This indicates the overall impact of the power change at node j on node i. Represents a node Incorporate the effect of active power on the voltage at node i. This indicates that node j adds active power to node j. Voltage influence, This indicates the effect of adding reactive power to node i on the voltage at node i. This indicates the effect of adding reactive power to node j on the voltage at node i. and This represents the electrical distance between node i and node j. This represents the symmetry and normalization process performed on the electrical distance between nodes i and j based on n nodes.
[0066] Energy storage zones can be divided using the K-means clustering algorithm. The K-means clustering algorithm can... The nodes are divided into K disjoint clusters. For example, the node set is... ,in It is If the feature vectors are dimensional, then the K-means clustering algorithm can be expressed as the following formula.
[0067] (6)
[0068] (7)
[0069] (8)
[0070] in, The objective function for partitioning the cluster, Clustering The central node, Represents a node With the central node The distance between them Represents a cluster The number of nodes in the array.
[0071] However, the K-means clustering algorithm is prone to getting trapped in local optima. Therefore, we can specify the center node of the initial cluster, define the node index based on the density around the center node of the initial cluster, and select the center node from the nodes with high node indices.
[0072] Calculate the distance between the central node and other nodes. As shown in the following formula.
[0073] (9)
[0074] Then calculate the nodal index, which can be... Sort the elements in the array in descending order from largest to smallest, and set the parameters. Select the first Each element is used as a node. node index Node index The more surrounding nodes there are, the smaller the node index.
[0075] Then select the node index, and you can set the parameters. ,Will Sort in ascending order and select the first... Using a threshold element, if the selected node's index is lower than the threshold, the node is called a high-index node. A smaller index indicates more surrounding nodes, thus increasing the probability that the node will become the central node, thereby identifying the high-index node group. This avoids the situation where the central node in the cluster appears on an isolated edge node.
[0076] Then, the node corresponding to the highest node index is determined as the central node of the cluster. The central node of the first cluster can be called... .
[0077] Select from the remaining nodes There is one central node. A node farthest from the central node can be selected as the next central node, as shown in the following formula.
[0078] (10)
[0079] The sum of squared errors (SSE) between nodes within the cluster and the central node of the cluster can be used as an evaluation metric for cluster partitioning, expressed as the following formula.
[0080] (11)
[0081] in, Represents a node With the central node distance, SSE indicates the number of nodes in the cluster; a smaller SSE indicates a more compact cluster partition. The relationship between the number of clusters and SSE can be expressed as follows: Figure 3 As shown.
[0082] Figure 3 A schematic diagram illustrating the relationship between the number of clusters and cluster partitioning evaluation indicators according to an embodiment of this application is shown.
[0083] like Figure 3 As shown, when the number of clusters X is less than 6, the larger the number of clusters, the smaller the cluster partitioning evaluation index. When the number of clusters X is greater than 6, the larger the number of clusters, the more stable the cluster partitioning evaluation index becomes. Therefore, it can be concluded that a larger number of clusters is not necessarily better.
[0084] Figure 4 A schematic diagram showing the relationship between temperature and capacity of an energy storage device according to an embodiment of this application is provided. Figure 5 A schematic diagram of different states of an energy storage device according to an embodiment of this application is shown.
[0085] like Figure 4 and Figure 5 As shown, at temperatures below 20℃, the capacity of the energy storage device gradually decreases as the temperature drops, indicating a low-temperature, low-capacity state. At temperatures above 30℃, the capacity gradually decreases as the temperature increases, placing the device in a high-temperature, flammable state. However, at temperatures between 20 and 30℃, the capacity of the energy storage device remains relatively stable, representing a normal, healthy state. Therefore, maintaining a stable temperature for the energy storage device can ensure its capacity.
[0086] For energy storage devices, temperature changes are determined by the internal temperature changes of the energy storage device and external heat exchange. In some regions, the temperature difference between summer (30℃-40℃) and winter (-15℃-0℃) is large. In summer, heat dissipation is difficult and thermal runaway of energy storage devices is likely to occur. In winter, the temperature is low, and the battery management system of the energy storage device will limit charging below 0℃. Low temperature cycling leads to irreversible capacity loss of the energy storage device. Therefore, it is necessary to consider the impact of external temperature and its own charging and discharging on the energy storage device. According to the law of conservation of energy, the dynamic energy balance of the energy storage device can be derived, which is expressed by the following formula.
[0087] (12)
[0088] in, This represents the equivalent mass of the components in an energy storage device that primarily affect temperature changes. Indicates specific heat capacity. It is the heat balance differential equation approximated by the difference, where, This indicates that the temperature is rising. This indicates that the temperature is decreasing. This indicates that the temperature is in a stable state. This represents the heat power applied or removed to the energy storage device to maintain a constant temperature, where positive values indicate heating and negative values indicate cooling. This represents the heat exchange coefficient of an energy storage device due to environmental changes. The overall heat exchange coefficient, Indicates the effective heat transfer area outside the energy storage device. This indicates the internal temperature of the energy storage device. Indicates the ambient temperature of the energy storage device. This indicates the constant temperature of the energy storage device. Indicates time, Indicates the time step. This represents the internal temperature of the energy storage device at time t. This indicates that at time t, elapsed The internal temperature of the energy storage device after the time step.
[0089] According to embodiments of this application, the control network training method may further include: determining the operating cost of the distribution network based on the grid power purchase cost, energy storage operation cost, energy storage lifetime loss cost, and temperature cost; determining the voltage deviation of the distribution network based on the number of nodes and the voltage standard value; and determining a reward based on the operating cost and the voltage deviation, wherein the energy storage lifetime loss cost and temperature cost are related to temperature.
[0090] The objective function for optimizing energy storage zones in a distribution network can be a function that minimizes the weighted sum of voltage deviation and operating costs. A smaller voltage deviation indicates a more stable distribution network operation, which can effectively slow down the aging of energy storage devices and improve the network's carrying capacity. Excessively high or low voltage will increase additional losses. The objective function is constructed as follows:
[0091] (13)
[0092] in, Describe the objective function. Represents the voltage deviation sub-function. This represents the sub-function for operating costs. This represents the weighting coefficients of the voltage deviation sub-function. This represents the weighting coefficient of the operating cost sub-function. and These can be dynamic weighting coefficients. In this case, it indicates a greater emphasis on the power quality of the distribution network. In this case, it indicates a greater emphasis on the economic efficiency of the distribution network operation.
[0093] The voltage deviation sub-objective function calculates the degree to which the voltage of the distribution network deviates from the rated value, and is expressed by the following formula.
[0094] (14)
[0095] in, Indicates the number of nodes in the distribution network. Indicates the control cycle. This represents the voltage deviation at node i at time t. This indicates the standard voltage value, for example, it could be 0.93pu-1.07pu.
[0096] The operating cost sub-function can be expressed as the following formula.
[0097] (15)
[0098] in, This indicates the cost of purchasing electricity from the power grid. Indicates the operating cost of energy storage. This indicates the cost of energy storage lifespan depletion. This indicates the cost of temperature.
[0099] The cost of purchasing electricity from the power grid is expressed by the following formula.
[0100] (16)
[0101] in, This represents the electricity price at time t. This indicates the power purchased from the power grid. For time step.
[0102] The operating cost of energy storage is expressed by the following formula.
[0103] (17)
[0104] in, Indicates the number of energy storage devices. This represents the operation and maintenance cost of the s-th energy storage device. This represents the charging and discharging power of the s-th energy storage device at time t, with charging being positive and discharging being negative.
[0105] The energy storage life loss cost can be the life loss of lithium batteries, which can be obtained using the raindrop counting method and nonlinear life loss model, as shown in the following formula.
[0106] (18)
[0107] (19)
[0108] (20)
[0109] (twenty one)
[0110] in, This represents the investment cost of the s-th energy storage device. This represents the theoretical lifespan cycle number of the s-th energy storage device. This represents the depth of discharge of the s-th energy storage device at time t. This represents the nonlinear relationship between the lifetime loss and depth of discharge of the s-th energy storage device at time t. Indicates the influence factor of temperature. This represents the external temperature of the energy storage device at time t. Indicates the standard temperature value. This represents the exponential coefficient between the lifetime loss and depth of discharge of the s-th energy storage device at time t. Represents the nonlinear coefficients. This represents the state of charge of the energy storage device at time t. This represents the state of charge of the energy storage device at time t-1.
[0111] The temperature cost is expressed by the following formula.
[0112] (twenty two)
[0113] (twenty three)
[0114] in, This represents the operating temperature of the s-th energy storage device at time t. This indicates that the s-th energy storage device is at a temperature of The additional aging rate, expressed as %h, can be calculated using the Arrhenius equation. This represents the rated capacity of the s-th energy storage device. This represents the unit replacement cost of the s-th energy storage device. Indicates the baseline aging rate. Indicates activation energy. is the gas constant.
[0115] According to embodiments of this application, the temperature cost considers the heat loss of the energy storage device in maintaining a constant temperature to cope with changes in external temperature and the heat generated by its own charging and discharging. It also considers the cost of a nonlinear energy storage device loss model based on raindrop counting and temperature-related costs. While ensuring the temperature safety of the energy storage device, the management of the energy storage device is optimized, guaranteeing the discharge capacity and power output of the battery.
[0116] According to embodiments of this application, temperature costs are related to the additional aging of energy storage devices based on temperature.
[0117] Figure 6 A schematic diagram of the scheduling of an energy storage device according to an embodiment of this application is shown.
[0118] like Figure 6 As shown, the horizontal axis represents the load power, and the vertical axis represents the power of the energy storage device and the photovoltaic device. When the load power of the distribution network is low, the energy storage device can charge. When the load is high, the energy storage device can discharge. When the power of the photovoltaic device increases and meets the load operation requirements, the energy storage device can charge.
[0119] According to embodiments of this application, the operating constraints may further include at least one of voltage constraints, power balance constraints, energy storage operation constraints, energy storage temperature control constraints, and photovoltaic output constraints. The voltage constraints characterize the voltage range of multiple nodes in the distribution network, the power balance constraints characterize the power range of multiple nodes in the distribution network, the energy storage operation constraints characterize the power range of the energy storage device under charging or discharging conditions, the energy storage temperature control constraints characterize the power range of the energy storage device under heating or cooling conditions, and the photovoltaic output constraints characterize the power range provided by the photovoltaic devices in the distribution network. The temperature-related constraints include the energy storage temperature control constraints.
[0120] The voltage constraint condition can be expressed as the following formula.
[0121] (twenty four)
[0122] in, This represents the voltage at node i at time t. , These represent the lower voltage limit and the upper voltage limit, respectively. This represents the maximum capacity of bus ij. This represents the capacity of bus ij at time t.
[0123] The power balance constraint can be expressed as the following formula.
[0124] (25)
[0125] in, , This represents the active and reactive power injected by node i at time t. , This represents the conductance and susceptance of bus ij at time t. This represents the voltage phase angle between nodes i and j at time t. This represents the total number of nodes.
[0126] To ensure that the thermal power required by the energy storage device is within the allowable physical range, the heating or cooling power of the energy storage device has upper and lower limits. Therefore, the energy storage temperature control constraint can be expressed as the following formula.
[0127] (26)
[0128] in, , These represent the minimum and maximum values of the energy storage heating or cooling power constraints, respectively.
[0129] According to embodiments of this application, heating or cooling can be performed to maintain the energy storage device at a predetermined temperature.
[0130] The operating constraints of energy storage devices satisfy the following formula.
[0131] (27)
[0132] in, , , Let represent the minimum power, power at time t, and maximum power of the j-th energy storage device during charging and discharging, respectively. , These represent the charging and discharging states of the j-th energy storage device at time t, respectively. In the discharge state , This represents the storage capacity of the j-th energy storage device at time t. , These represent the charging efficiency and discharging efficiency of the energy storage device, respectively. This represents the state of charge of the j-th energy storage device at time t.
[0133] The photovoltaic output constraint can be expressed by the following formula.
[0134] (28)
[0135] in, This represents the photovoltaic output of the photovoltaic equipment at time t. , These represent the minimum and maximum photovoltaic output, respectively.
[0136] According to embodiments of this application, the training method of this application can be combined with Markov cooperative game modeling.
[0137] The distribution network can be divided into zones, with an agent set up in each energy storage zone. The agent controls the action of the energy storage equipment. The agent makes actions by locally observing the internal state of the energy storage zone and cooperating to minimize the objective function. Therefore, it can be modeled as a Markov cooperative game, expressed as the following formula.
[0138] (29)
[0139] in, Representing the state space, Indicates an action, Represents the reward function, This represents the discount factor.
[0140] In the process of optimizing the scheduling of energy storage devices for distribution networks, the state space can define various operating states of the distribution network. The state space can be defined using the active power demand of the distribution network load, the active power output of photovoltaic systems, and the state of charge of the energy storage devices, and can be expressed by the following formula.
[0141] (30)
[0142] in, Let represent the state of the agent at time t. Represents time t Active power demand of regional load, Represents time t The active power output of the photovoltaic equipment in the region Represents time t State of charge of regional energy storage devices.
[0143] Actions can be any possible operations for a given state, used to improve the power acceptance capacity of photovoltaic equipment in the distribution network, the aggregation and regulation of energy storage equipment, and the economic efficiency of operation. They can be expressed as the following formula.
[0144] (31)
[0145] (32)
[0146] in, Represents time t The operation of the regional energy storage equipment meets all operational constraints. Represents time t The charging and discharging power of regional energy storage equipment This indicates the storage capacity of the energy storage device.
[0147] The reward that can be obtained by taking an action can be expressed by the following formula.
[0148] (33)
[0149] in, The reward at time t, Let the objective function be [function name], and if the operation of the distribution network satisfies the constraints, then [function name]... It is 0 otherwise it is 1. It is the penalty vector.
[0150] Within a decision-making cycle, the agent can determine the actions of energy storage devices based on the current state, interact with the environment to perform power flow calculations, and calculate rewards accordingly. Then it proceeds to the next time step until a predetermined number of steps is reached and the process terminates. The cumulative reward is as follows: Figure 7 As shown.
[0151] Figure 7 The relationship between cumulative rewards and time steps according to embodiments of this application is illustrated.
[0152] like Figure 7 As shown, the number of training steps can be the number of steps pushed down from the first state. As the number of training steps increases to 1000, the reward gradually stabilizes.
[0153] The cumulative maximum reward is expressed in the form of expectation, which is the predicted action value, and can be represented by the following formula.
[0154] (34)
[0155] in, Let 's' represent the predicted action value, 's' represent the state, and 'a' represent the action. Expressing expectations, Indicates the control cycle.
[0156] According to an embodiment of this application, updating the parameters of the policy network, the action evaluation network, and the target evaluation network based on the predicted action value and the target action value includes: determining the policy network loss value of the policy network and the action evaluation network loss value of the action evaluation network based on the predicted action value and the target action value; updating the parameters of the policy network and the action evaluation network based on the policy network loss value and the action evaluation network loss value; and softly updating the parameters of the target evaluation network based on the updated parameters of the action evaluation network.
[0157] According to embodiments of this application, the policy network and the action evaluation network can each have their own loss function. The policy network loss value of the policy network and the action evaluation network loss value of the action evaluation network can be determined using the loss function based on the predicted action value and the target action value.
[0158] Then, based on the loss values of the policy network and the action evaluation network, the parameters of the policy network and the action evaluation network are updated. The target evaluation network and the action evaluation network can have the same structure. Therefore, the parameters of the target evaluation network can be updated based on the updated parameters of the action evaluation network. This update is called a soft update.
[0159] According to embodiments of this application, determining the policy network loss value of the policy network and the action evaluation network loss value of the action evaluation network based on the predicted action value and the target action value may include: determining the policy network loss value of the policy network based on the predicted action value; and determining the action evaluation network loss value of the action evaluation network based on the predicted action evaluation value and the target action value.
[0160] The loss function of the policy network can be expressed as the following formula.
[0161] (35)
[0162] in, The loss function of the policy network is represented. Indicates that at time t, the state is in the first state. Execute the first action Expectations Indicates that at time t, the state is in the first state. Execute the first action The reward The regularization coefficient representing entropy. For policy entropy, This represents the policy distribution function in the first state.
[0163] The specific update of the policy network is as follows:
[0164] (36)
[0165] in, This represents the loss function of the policy network after substituting the network parameters into the policy network. This represents the function value of the action evaluation network, with the first state as the input. First action and action evaluation network parameters , For the network parameters of the policy network, It's the learning rate. This is the gradient of the loss function.
[0166] The specific update formula for the action evaluation network is as follows.
[0167] (37)
[0168] in, This represents the loss function of the action evaluation network. This represents the reward obtained from performing the first action in the first state and the expectation in the second state. This represents the predicted output value of the action evaluation network. This represents the expected action to be performed at time t+1. This represents the function value of the target evaluation network, with the second state as the input. Second action and target evaluation network parameters , express The first derivative, This is the experience pool, used to store parameters during the training process.
[0169] Since the voltage deviation subfunction and the operation subfunction have different units, they can be handled by normalizing the vector, as shown in the following formula.
[0170] (38)
[0171] in, , Let represent the maximum and minimum values of the reward value of the voltage deviation sub-function at time t, respectively. This represents the reward value of the voltage deviation sub-function at time t. This represents the reward value of the operating subfunction at time t. and Let t represent the maximum and minimum reward values of the sub-function and the operation function, respectively.
[0172] Figure 8 A schematic block diagram of a control network training method for a power distribution network according to an embodiment of this application is shown.
[0173] like Figure 8 As shown, training data is input into the experience pool, and data is extracted from the experience pool. Random sampling yields the first state. The first action taken in the first state Rewards corresponding to the first action and the second state achieved by taking the first action in the first state. Using a policy network, the target partition control action is determined based on the first state. The predicted action value, i.e., the expected cumulative reward for executing the target partition control action in the first state, is obtained using an action evaluation network. The target action value, i.e., the cumulative reward for performing the action starting from the second state, is obtained using a target evaluation network. The expected value is determined by the policy network. Based on the predicted action value, the target action value, and the respective loss functions of the three networks, the parameters of the policy network, the action evaluation network, and the target evaluation network are updated. Training can be performed offline, and then the updated policy network is executed online in the power distribution network. The agent determines the actions of the power distribution network based on its state through the policy network, thereby controlling the power distribution network according to these actions.
[0174] According to embodiments of this application, the application can be divided into three parts: distribution network partitioning, offline pre-training, and online training execution. First, the distribution network is partitioned using the K-means clustering method. Second, pre-training is performed based on training data using a distribution network control network training method. Finally, the updated policy network and the real-time operating status of the distribution network are used to optimize and schedule energy storage devices in real time, achieving optimal control while reducing training time and ensuring grid security.
[0175] According to embodiments of this application, in the event of a sudden surge in photovoltaic power and regional overload, different energy storage zones can utilize energy storage devices for regional autonomy or mutual assistance, mitigating the impact of uncertainties on the distribution network. During topology changes, online training can be performed based on the network obtained through offline training, reducing the sample requirements for online interactive learning, accelerating convergence, and increasing real-time response capabilities.
[0176] Figure 9 A flowchart of a control method for a power distribution network according to an embodiment of this application is shown.
[0177] like Figure 9 As shown, the method includes operations S910 to S930.
[0178] When operating the S910, the current status of each energy storage zone is obtained.
[0179] When operating the S920, the current partition control action is determined based on the current state, according to the updated policy network obtained by using the control network training method.
[0180] When operating the S930, control the energy storage zone based on the current zone control action.
[0181] With the development of energy storage technology in power distribution networks, the use of energy storage devices to mitigate fluctuating loads has been proposed. However, the operation of energy storage devices is affected by natural conditions. In summer, high ambient temperatures combined with the heat generated by the charging and discharging of energy storage devices can easily lead to thermal runaway. In winter, low temperatures significantly reduce the capacity of energy storage devices, preventing full utilization of their resources. The usable capacity of energy storage devices varies under different temperatures. Figure 4 As shown in the figure. In power system simulation, most studies focus on peak shaving and valley filling and the entire life cycle of energy storage devices, while lacking research that considers the impact of external temperature on energy storage devices.
[0182] According to an embodiment of this application, data from a regional power distribution network is used to verify a power distribution network control network training method.
[0183] Figure 10 A schematic diagram of an energy storage zone of a distribution network according to an embodiment of this application is shown in one example.
[0184] like Figure 10As shown, the distribution network has 33 nodes, divided into 6 energy storage zones: Zone 1, Zone 2, Zone 3, Zone 4, Zone 5, and Zone 6. To ensure efficient response of energy storage devices across zones, a particle swarm optimization algorithm is used for optimal configuration. Energy storage devices are added to nodes 3, 6, 9, 16, 21, and 31, with capacities of 0.6MW, 0.8MW, 1MW, 1MW, 0.9MW, and 1MW respectively. Photovoltaic equipment is added to node 9. The load and photovoltaic equipment are normalized to match the data requirements of the simulation example to train the agent.
[0185] The following analysis uses six case studies, with the variables and parameters for each case shown in Tables 1 and 2.
[0186] Table 1
[0187]
[0188] Table 2
[0189]
[0190] Figure 11 A schematic diagram of the photovoltaic output and load power factor of a distribution network according to an embodiment of this application is shown in one example. Figure 12 The diagram shows temperature curves of a distribution network according to an embodiment of this application during a summer surge in photovoltaic power and during a winter overload, in one example. Figure 13 A graph showing the energy consumption of an energy storage device in a distribution network according to an embodiment of this application at different temperatures is shown in one example.
[0191] like Figures 11-13 As shown, using typical summer days and typical winter days as experimental days, the relationship between the power factor of photovoltaic output and load, the temperature changes during photovoltaic surges and winter load overloads, and the energy consumption changes of energy storage devices at different temperatures were obtained.
[0192] Figure 14 A schematic diagram illustrating the relevant changes in the distribution network under the condition of a sudden increase in photovoltaic power in the distribution network in Example 1 of this application embodiment is shown.
[0193] like Figure 14 As shown, the power supplied by the photovoltaic equipment at node 9 is increased until a voltage over-limit occurs in region 5, at which point the power supplied by the photovoltaic equipment is 4MW. Among these, Figure 14 (a) is a graph showing the voltage changes at node 9, region 5, and the total node. Figure 14 (b) is a graph showing the variation from region 1 to region 6. From... Figure 14 From (a), it can be seen that the voltage over-limit period is from 10:00 to 14:00. Figure 14As can be seen in (b), at the same time that the voltage of region 5 exceeds the limit, the voltage of region 6 also exceeds the limit. The reason is that the photovoltaic power generation suddenly increases and region 5 cannot fully absorb it. Therefore, it is transmitted to the end of the distribution network. Region 6 at the end cannot consume the remaining photovoltaic power, and at this time, power flow reversal will occur.
[0194] Figure 15 A schematic diagram illustrating the relevant changes in the distribution network under the condition of a sudden increase in photovoltaic power in the distribution network in Example 2 of this application embodiment is shown.
[0195] like Figure 15 As shown, Case 2, based on Case 1, adds the energy loss required for the energy storage device to maintain its own temperature. Among these, Figure 15 (a) is a graph showing the voltage changes at node 9, region 5, and the total node. Figure 15 (b) is a graph showing the variation from region 1 to region 6. (and) Figure 14 Compared to (a), Figure 15 In (a), node 9 and region 5 experienced more severe voltage limit exceedances between 10:00 and 12:00 compared to when no energy storage was used, and a brief period of voltage non-exceedance occurred between 12:00 and 14:00, because the energy storage device consumed some of the photovoltaic power during this time. Figure 14 Compared to (b), Figure 15 In (b), brief voltage dips occurred in regions 4, 5, and 6 between 1:00-4:00 and 21:00-23:00. This was due to energy consumption by the energy storage devices and increased load during the 21:00-23:00 period, which led to the voltage dips.
[0196] Figure 16 A schematic diagram illustrating the relevant changes in the distribution network under the condition of a sudden increase in photovoltaic power in the distribution network in Example 3 of this application embodiment is shown.
[0197] like Figure 16 As shown, Case 3, based on Case 2, adds the scheduling of energy storage devices in Region 5. Among them, Figure 16 (a) is a graph showing the voltage changes at node 9, region 5, and the total node. Figure 16 (b) is a graph showing the changes from region 1 to region 6. Figure 16 (c) is a graph showing the voltage changes at each node. Figure 16 (d) is a graph showing the energy dispatch changes of the energy storage devices in region 5 over 24 hours. For example... Figure 16As shown in (d), the device charges when photovoltaic power increases rapidly and discharges when the load power is high. Blue indicates the change in the state of charge (SOC) of the energy storage device, and red indicates the energy storage strategy (ESS). The energy storage device is negative when charging and positive when discharging.
[0198] from Figure 16 of (a) Figure 16 As can be seen from (b), compared with Case 2 without energy storage equipment scheduling, the time when the voltage exceeds the upper limit is significantly shortened, but there is still a period of time when the voltage exceeds the upper limit; the ability to improve the voltage below the lower limit is weak, indicating that the autonomy of the energy storage equipment in Region 5 is limited. Therefore, the energy storage equipment in Region 6 can be called upon to participate in regulation.
[0199] Figure 17 A schematic diagram illustrating the relevant changes in the distribution network in the case of a sudden increase in photovoltaic power in the distribution network in Example 4 of this application embodiment is shown.
[0200] like Figure 17 As shown, Case 4, based on Case 3, adds the scheduling of energy storage devices in Region 6. Among them, Figure 17 (a) is a graph showing the voltage changes at node 9, region 5, and the total node. Figure 17 (b) is a graph showing the changes from region 1 to region 6. Figure 17 (c) is a graph showing the voltage changes at each node. Figure 17 (d) is a graph showing the energy dispatch changes of energy storage devices in regions 5 and 6 over 24 hours. Figure 17 (a) and Figure 17 As can be seen from (b), through the combined regulation of the energy storage devices in regions 5 and 6, the voltage in each time period and region is significantly improved. Figure 17 As can be seen from (c), the voltage changes at each node are all within the acceptable voltage range. Figure 17 As can be seen from (d), the energy storage devices in region 5 charge from 11:00 to 14:00, while the energy storage devices in region 6 charge from 12:00 to 13:00 and discharge at 11:00 and 14:00. The reason for the different strategies is that the photovoltaic surge occurs between 12:00 and 1:00. The charging of upstream region 5 causes downstream region 6 to be unable to meet the load requirements. Therefore, the energy storage devices in region 6 discharge to maintain the power balance of their own region.
[0201] Figure 18 A schematic diagram illustrating the relevant changes in the distribution network under load changes in Example 5 of this application embodiment is shown.
[0202] like Figure 18As shown, Case 5 simulates an extremely cold, low-temperature scenario without photovoltaic power. Based on the existing load power factor, a 0.5MW load is added to node 9 from 18:00 to 21:00 until the voltage at node 9 and the voltage in region 5 both cross the lower limit. Among these... Figure 18 (a) is a graph showing the voltage changes at node 9, region 5, and the total node. Figure 18 (b) is a graph showing the changes from region 1 to region 6. Figure 18 (c) shows the voltage changes at each node. Since region 4 is the end region and the energy storage equipment needs to maintain optimal operation due to the low ambient temperature, it will exceed the limit between 18:00 and 24:00.
[0203] Figure 19 A schematic diagram illustrating the relevant changes in the distribution network in the case of a sudden increase in photovoltaic power in the distribution network in Example Six of this application embodiment is shown.
[0204] like Figure 19 As shown, Case Six, based on Case Five, adds energy storage devices from Regions 4, 5, and 6 to participate in regulation. Among them, Figure 19 (a) is a graph showing the voltage changes at node 9, region 5, and the total node. Figure 19 (b) is a graph showing the changes from region 1 to region 6. Figure 19 (c) is a graph showing the voltage changes at each node. Figure 19 (d) is a graph showing the energy dispatch changes of energy storage devices in regions 4, 5 and 6 over 24 hours.
[0205] and Figure 18 (a) comparison, Figure 19 In (a), the voltage can be kept within a reasonable range by the action of three energy storage devices. Figure 18 (b) comparison, Figure 19 Although the voltage in each region of (b) is within a reasonable range, Figure 19 (b) The voltage fluctuation is significantly higher than Figure 18 (b) is due to voltage fluctuations caused by the charging and discharging of the energy storage device.
[0206] The changes in energy storage equipment in regions 4, 5 and 6 are as follows: Figure 19 As shown in (d), the energy storage device discharges between 19:00 and 21:00. During this period, the discharge capacity of region 4 is the smallest because the overload node is at node 9. The reason for the voltage exceeding the limit in region 4 is that it is located at the end of the line, resulting in an excessively high power factor. Therefore, only a small portion of the discharge is needed to ensure the safe voltage threshold. Regions 5 and 6 have a deeper discharge depth because node 9 is an overload node. In addition, region 6 is located at the end of the line, so more capacity needs to be released to ensure the safe voltage threshold.
[0207] The voltage offsets and operating costs for Case 1 through Case 6 are shown in Table 3 below.
[0208] Table 3
[0209]
[0210] Table 3 shows that the voltage deviation and operating cost in Case 2 are higher than in Case 1. This is because the energy consumed by the added energy storage device due to temperature changes is added in the form of energy storage, and the increased operating cost is due to the need to purchase more electricity. In Cases 3 and 4, both voltage deviation and operating costs are reduced because the energy storage device makes the load more stable. During periods of rapid photovoltaic growth, the energy storage device stores energy and releases it when the load is high, thus reducing the operating cost of the distribution network. Cases 5 and 6 show the difference between adding and not adding energy storage devices when the regional load is overloaded. The voltage deviation is significantly reduced, and the operating cost is lower because the energy storage device charges during low load periods (1:00-4:00) and discharges during high load periods, utilizing peak-valley arbitrage to reduce operating costs.
[0211] Figure 20 A block diagram of a control network training device for a power distribution network according to an embodiment of this application is shown.
[0212] like Figure 20 As shown, the power distribution network includes multiple energy storage zones, each of which is equipped with energy storage devices. The control network training device 2000 of the power distribution network includes an input module 2010, a first determination module 2020, a second determination module 2030, a third determination module 2040, and an update module 2050.
[0213] Input module 2010 is used to input training data for each energy storage zone. The training data includes a first state, a first action taken in the first state, a reward corresponding to the first action, and a second state achieved by taking the first action in the first state. The first and second states each include the ambient temperature of the energy storage device. The reward represents the impact of the first action on the distribution network and includes temperature-related effects on the energy storage device. First determination module 2020 is used to determine the target zone control action based on the first state using a policy network. The target zone control action is constrained by the operating constraints of the distribution network, including temperature-related constraints on the energy storage device. Second determination module 2030 is used to determine the predicted action value using an action evaluation network. The predicted action value represents the expected cumulative reward when the target zone control action is performed in the first state. Third determination module 2040 is used to determine the target action value using a target evaluation network. The target action value represents the expected cumulative reward when the action is performed starting from the second state. The update module 2050 is used to update the parameters of the policy network, the action evaluation network, and the target evaluation network based on the predicted action value and the target action value, so as to control the energy storage zoning according to the updated policy network.
[0214] Any one or more of the modules, submodules, units, and subunits according to the embodiments of this application, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to the embodiments of this application can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to the embodiments of this application can be at least partially implemented as hardware circuits, such as field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), systems-on-a-chip, systems-on-a-substrate, systems-on-package, application-specific integrated circuits (ASICs), or implemented by hardware or firmware in any other reasonable manner by integrating or packaging circuits, or implemented in any one of software, hardware, and firmware, or in a suitable combination of any of these. Alternatively, one or more of the modules, submodules, units, and subunits according to the embodiments of this application can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0215] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of this application can be combined and / or combined in various ways, even if such combinations are not explicitly described in this application. In particular, without departing from the spirit and teachings of this application, the features described in the various embodiments of this application can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of this application.
[0216] The embodiments of this application have been described above. However, these embodiments are merely illustrative and not intended to limit the scope of this application. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of this application, those skilled in the art can make various substitutions and modifications, all of which should fall within the scope of this application.
Claims
1. An active power distribution network energy storage zone control network training method, characterized in that, The power distribution network comprises a plurality of energy storage partitions, and energy storage devices are arranged in the energy storage partitions. Input training data, the training data comprising a first state, a first action taken in the first state, a reward corresponding to the first action, and a second state achieved by taking the first action in the first state, wherein the first state and the second state each comprise an ambient temperature of the energy storage device, and the reward represents an influence of the first action on the power distribution network and comprises a temperature-related influence on the energy storage device; Determine a target partition control action based on the first state using a policy network, wherein the target partition control action is constrained by an operation constraint condition of the power distribution network, and the operation constraint condition comprises a temperature-related constraint condition for the energy storage device; Determine a predicted action value using an action evaluation network, wherein the predicted action value represents an expectation for cumulative reward in the case of executing the target partition control action in the first state; Determine a target action value using a target evaluation network, wherein the target action value represents an expectation for cumulative reward in the case of taking actions from the second state; Update parameters of the policy network, parameters of the action evaluation network, and parameters of the target evaluation network according to the predicted action value and the target action value, so as to control the energy storage partitions according to the updated policy network; The updating of the parameters of the policy network, the parameters of the action evaluation network, and the parameters of the target evaluation network according to the predicted action value and the target action value comprises: Determine a policy network loss value of the policy network and an action evaluation network loss value of the action evaluation network according to the predicted action value and the target action value; Update the parameters of the policy network and the parameters of the action evaluation network according to the policy network loss value and the action evaluation network loss value; Softly update the parameters of the target evaluation network based on the updated parameters of the action evaluation network; The control network training method further comprises: Determine an operation cost of the power distribution network according to a grid purchase cost, an energy storage operation cost, an energy storage life consumption cost, and a temperature cost; Determine a voltage deviation of the power distribution network according to a number of nodes of the power distribution network and a voltage standard value; Determine the reward according to the operation cost and the voltage deviation, wherein the energy storage life consumption cost and the temperature cost are related to temperature; The control network training method further comprises: Determine electrical distances between a plurality of nodes of the power distribution network according to grid data of the power distribution network, wherein the energy storage devices are arranged at at least one node; Determine energy storage partitions of the power distribution network according to the electrical distances between the plurality of nodes.
2. The control network training method of claim 1, wherein, The determination of the policy network loss value of the policy network and the action evaluation network loss value of the action evaluation network according to the predicted action value and the target action value comprises: determine the policy network loss value of the policy network according to the predicted action value; determine the action evaluation network loss value of the action evaluation network according to the predicted action evaluation value and the target action value.
3. The control network training method of claim 1, wherein, The temperature cost is related to additional aging of the energy storage device based on temperature.
4. The control network training method of claim 1, wherein, The operation constraint condition further includes at least one constraint condition in a voltage constraint condition, a power balance constraint condition, an energy storage operation constraint condition, an energy storage temperature control constraint condition, and a photovoltaic output constraint condition; The voltage constraint condition characterizes a voltage range of a plurality of nodes in the power distribution network. The power balance constraint condition characterizes a power range of a plurality of nodes in the power distribution network. The energy storage operation constraint condition characterizes a power range of the energy storage device in the case of charging or discharging. The energy storage temperature control constraint condition characterizes a power range of the energy storage device in the case of heating or refrigeration. The photovoltaic output constraint condition characterizes a power range provided by a photovoltaic device in the power distribution network. The temperature-related constraint condition includes the energy storage temperature control constraint condition.
5. The control network training method of claim 4, wherein, The heating or refrigeration is performed to maintain the energy storage device at a predetermined temperature.
6. A control method of a power distribution network, characterized by, The power distribution network includes a plurality of energy storage sub-zones, and the control method of the power distribution network includes: obtaining a current state of the energy storage sub-zone; determining a current sub-zone control action based on the current state by using an updated policy network obtained by using the control network training method in any one of claims 1-5; controlling the energy storage sub-zone based on the current sub-zone control action.
7. An active power distribution grid energy storage zone control network training device, characterized by, The active power distribution network energy storage sub-zone control network training device is applied to the control network training method in any one of claims 1-5, the power distribution network includes a plurality of energy storage sub-zones, and an energy storage device is arranged in each energy storage sub-zone. The control network training device includes: an input module configured to input training data for each energy storage sub-zone, the training data including a first state, a first action taken in the first state, a reward corresponding to the first action, and a second state achieved by taking the first action in the first state, wherein the first state and the second state each include an ambient temperature of the energy storage device, and the reward represents an influence of the first action on the power distribution network and includes a temperature-related influence on the energy storage device; a first determination module configured to determine a target sub-zone control action based on the first state by using a policy network, wherein the target sub-zone control action is constrained by an operation constraint condition of the power distribution network, and the operation constraint condition includes a temperature-related constraint condition for the energy storage device; a second determination module configured to determine a predicted action value by using an action evaluation network, the predicted action value representing an expectation of cumulative reward in the case of executing the target sub-zone control action in the first state; a third determination module configured to determine a target action value by using a target evaluation network, the target action value representing an expectation of cumulative reward in the case of performing an action from the second state; and a fourth determination module configured to determine a policy network loss value of the policy network according to the predicted action value. An updating module is configured to update the parameters of the policy network, the parameters of the action evaluation network and the parameters of the target evaluation network according to the predicted action value and the target action value, so as to control the energy storage partition according to the updated policy network.
Citation Information
Patent Citations
Collaborative optimization scheduling method for power distribution and micro-grid and related device
CN119787370A
Micro-grid energy management method and system based on deep reinforcement learning
CN120109917A