Active distribution network cluster dynamic division method and system based on deep reinforcement learning
By applying deep reinforcement learning algorithms in the distribution network to construct and solve the CMDP model, the efficiency and accuracy problems of traditional methods under high proportion distributed photovoltaic access are solved, and an efficient, real-time and accurate solution for dynamic division of distribution network clusters is realized.
Patent Information
- Application Number
- CN202411666509.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-11-21
AI Technical Summary
In the face of high proportion of distributed photovoltaic access, traditional distribution network cluster division method is difficult to effectively deal with dynamic changes and uncertainties in system state, resulting in model limitations, high computational complexity and poor real-time performance.
Using a deep reinforcement learning method, a CMDP model for dynamic division of distribution network clusters is constructed, and solved through the AC-DRL framework. The Lagrangian relaxation method is used to transform it into a non-constraint problem. Combined with Gumbel-Softmax function and dual Q clipping learning technology, a deployable CMDP model is trained.
It significantly improves the efficiency and accuracy of the distribution network cluster division results, can quickly respond to changes in the distribution network operating status, reduces the problems of poor computing complexity and real-time performance, and adapts to the drastic changes in the system state caused by high proportion of distributed photovoltaics.
Smart Images

Figure CN119154412B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power distribution networks, and in particular to a method and system for dynamically dividing active power distribution network clusters based on deep reinforcement learning. Background Art
[0002] As a key link between power transmission and users, the safe and reliable operation of the distribution network is of great importance to the stability of the power system and user experience. The current distribution system still adopts a unified and centralized management method. Faced with the current situation of a high proportion of distributed power sources, the disadvantages of this management method are gradually emerging. Therefore, it is necessary to adopt a cluster management method for active distribution networks containing a high proportion of distributed photovoltaics, which also provides another new research direction for handling a large number of dispersed distributed power sources.
[0003] Traditional distribution network clustering methods are mostly based on static models and mathematical optimization algorithms. These methods usually assume that system parameters and external conditions are constant or vary within a small range, such as linear programming, mixed integer programming, and heuristic algorithms. However, the actual operating environment of the distribution network is complex and changeable, especially when a high proportion of distributed photovoltaic access is connected, the photovoltaic output and load demand are highly dynamic and uncertain. Traditional optimization methods have the following shortcomings in dealing with these complex dynamic changes: 1) Model limitations: Traditional clustering methods rely on accurate mathematical models to describe the system, but the various uncertainties and complex relationships in the actual distribution network are difficult to represent with accurate mathematical models. 2) High computational complexity: When dealing with large-scale, multi-variable distribution network systems, the computational complexity of traditional clustering methods increases exponentially, and it is difficult to obtain the optimal solution in a short time. 3) Poor real-time performance: Due to the long calculation time, traditional clustering methods are difficult to achieve rapid response and dynamic adjustment to changes in the operating status of the distribution network. Summary of the invention
[0004] In order to solve the deficiencies in the prior art, the present invention provides a method for dynamically partitioning high-proportion distributed photovoltaic access distribution network clusters based on a deep security reinforcement learning algorithm. This method aims to solve several key problems in the existing cluster partitioning technology, including high requirements for distribution network model parameter accuracy, long decision-making time, and inability to effectively cope with drastic changes in system state caused by high-proportion distributed photovoltaics. Through this method, the efficiency and accuracy of the distribution network cluster partitioning results can be significantly improved.
[0005] The present invention adopts the following technical solution.
[0006] A first aspect of the present invention provides a method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning, comprising the following steps:
[0007] Step 1: Construct a CMDP model for dynamic partitioning of distribution network clusters;
[0008] Step 2: Use the Lagrangian relaxation method to transform the CMDP model constructed in step 1 into an unconstrained problem, and construct an AC-DRL-based solution framework to solve the CMDP model;
[0009] Step 3: Train the AC-DRL-based solution framework constructed in step 2 to obtain a deployable CMDP model;
[0010] Step 4: Deploy the deployable CMDP model obtained in step 3 to the distribution network dispatching cloud platform, and provide cluster division results based on the real-time observation status of the distribution network.
[0011] Preferably, in step 1, the CMDP model for dynamic partitioning of distribution network clusters is expressed by the following formula:
[0012] (10)
[0013] Where:
[0014] is the discount factor, is the cost limit, represents the optimal strategy;
[0015] for The agent state space contains the distribution network state information at all times;
[0016] for The action space of the agent as the result of the dynamic clustering of the distribution network at all times;
[0017] is the agent’s reward function;
[0018] is the cost function of the agent.
[0019] Preferably, At this moment, the agent state space containing the distribution network state information is also expressed by the following formula:
[0020] (1)
[0021] Where:
[0022] Respectively At time , the active and reactive load demand vectors of the distribution network;
[0023] Respectively The active and reactive output vectors of photovoltaics in the distribution network at all times;
[0024] The agent action space as a result of the dynamic division of the distribution network into dynamic clusters is also expressed by the following formula:
[0025] (2)
[0026] Where:
[0027] is a discrete integer variable, indicating the The cluster to which the node belongs is time;
[0028] is the number of distribution network nodes, is the number of clusters to be divided.
[0029] Preferably, At this moment, the agent's reward function is expressed as follows:
[0030] (8)
[0031] Where:
[0032] Respectively At this moment, the system's intra-group aggregation, inter-group sparsity, and power balance.
[0033] Preferably, the intra-cluster aggregation, inter-cluster sparsity and power balance are expressed by the following formula:
[0034] (3)
[0035] (4)
[0036] (5)
[0037] (6)
[0038] (7)
[0039] Where:
[0040] Used to evaluate the coupling degree between nodes within the cluster;
[0041] Used to evaluate the coupling degree between nodes in the cluster;
[0042] Represents a distribution network node and nodes Electrical distance;
[0043] For Node The set of internal node numbers of the sub-area to which it belongs;
[0044] It is the reactive power balance index;
[0045] c is the number of clusters;
[0046] For cluster The reactive balance degree, The maximum reactive power supply within the cluster, including the reactive power provided by the node reactive power compensation device and the reactive power that can be provided by some inverters. is the reactive power demand value within the cluster;
[0047] For cluster The degree of active balance, is the maximum value of active power supply within the cluster, is the required value of active power within the cluster.
[0048] Preferably, At this moment, the cost function of the agent is also expressed as follows:
[0049] (9)
[0050] Where:
[0051] Represents a cluster The number of isolated islands in the
[0052] Preferably, in step 2, the Gumbel-Softmax function is used to generate the sampling probability of the action, and the corresponding discrete action is selected accordingly, which is expressed as the following formula:
[0053] (12)
[0054] Where:
[0055] π is the strategy, which is formalized as a neural network whose input is the state and output is the partition number of each node;
[0056] Is The generated Gumbel noise is used to introduce randomness into the sampling process;
[0057] For Node The probability vector of belonging to which cluster;
[0058] is a vector of H random variables uniformly distributed between 0 and 1;
[0059] is a hyperparameter;
[0060] is the softmax function.
[0061] Preferably, in step 2, in the process of training the action network and the evaluation network, the double Q clipping learning technology is introduced, and the strategy loss function based on the double Q clipping learning is expressed by the following formula:
[0062] (13)
[0063] Where:
[0064] Indicates that the agent is in state according to the action network The sampled actions, φ is the neural network parameter of the discrete policy network; and They are the reward value function and the cost value function respectively.
[0065] Preferably, step 3 comprises:
[0066] ① Initialization phase: First, initialize the parameters of the Gumbel-Softmax-based policy network, cost evaluation network, and reward evaluation network; create an empty experience playback buffer D to store the experience data of the interaction between the agent and the distribution network environment;
[0067] ② Experience collection: By executing the actions output by the action network in the distribution network environment, experience data is collected, including: distribution network status, cluster division results, reward value, cost value, which can be expressed as , each set of experience data is stored in a buffer;
[0068] ③Data sampling: randomly extract a batch of experience data from the playback buffer for subsequent network training;
[0069] ④ Evaluation network update: Use the cost evaluation network, reward evaluation network and policy network to estimate the target cost and reward Q value of each sampled experience; calculate the mean square error loss between the target cost, reward Q value and actual Q value, and update the parameters θ and θ of the reward and cost evaluation networks accordingly. ;
[0070] ⑤ Policy network update: Apply the gradient ascent method to update the parameter φ of the discrete policy network to optimize the action selection of the policy network, that is, the clustering result;
[0071] ⑥ Soft update target network: Use soft update method to gradually adjust the parameters of target reward and cost evaluation network and , to ensure the stability of network parameter updates;
[0072] ⑦ Update the dual variable: fix the trained cost evaluation network and policy network parameters, use the gradient ascent method to update the dual variable, and ensure that it is greater than or equal to 0;
[0073] ⑧Training iteration: Repeat steps 2 to 7 until the set training end condition is reached.
[0074] The second aspect of the present invention provides a high-proportion distributed photovoltaic access distribution network cluster dynamic division system based on secure deep reinforcement learning, and runs the active distribution network cluster dynamic division method based on deep reinforcement learning, including:
[0075] Data acquisition module, used to input the real-time observation status of the distribution network;
[0076] CMDP module, with built-in CMDP model for dynamic partitioning of distribution network clusters;
[0077] Solver module, with built-in AC-DRL framework, for solving CMDP models;
[0078] The cluster division result output module is used to provide cluster division results according to the real-time observation status of the distribution network.
[0079] Compared with the prior art, the beneficial effects of the present invention include at least:
[0080] 1) Compared with traditional clustering methods (particle swarm optimization, genetic algorithm, etc.), the method proposed in this invention makes decisions in a data-driven and model-free manner, which is beneficial in practice in view of the difficulty of accurately estimating system parameters in real time.
[0081] 2) The proposed method uses neural networks to extract features and optimize strategies. Compared with traditional clustering methods, it greatly reduces the computational complexity in practical applications and can make extremely fast decisions even for large-scale systems, making it more capable of responding to fluctuations in distribution network status in real time.
[0082] 3) The proposed method is a dynamic cluster division method for high-penetration photovoltaic access distribution networks. Compared with the traditional static cluster division method that only focuses on a single time section, it can dynamically adjust the cluster division results according to the operating status of the distribution network (node load demand and photovoltaic output), and can effectively adapt to the dynamic changes of photovoltaic output and network topology. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] Figure 1 Dynamic partitioning process of active distribution network clusters based on deep reinforcement learning; DETAILED DESCRIPTION
[0084] In order to make the purpose, technical scheme and advantages of the present invention clearer, the technical scheme of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the spirit of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work belong to the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein should be the common meanings understood by people with general skills in the field to which the present invention belongs.
[0085] Compared with existing technologies, deep reinforcement learning algorithm, as an artificial intelligence method combining deep learning and reinforcement learning, has the following significant advantages: 1) Strong adaptability: Deep reinforcement learning algorithm can adapt to different system states and external conditions through continuous interaction and learning with the environment, without the need to predefine precise mathematical models. 2) Efficient processing of complexity: Deep reinforcement learning can handle high-dimensional input data and complex system dynamics, and perform feature extraction and strategy optimization through neural networks, greatly reducing computational complexity. 3) Real-time decision-making ability: Deep reinforcement learning can respond to system changes in real time, make decisions quickly and dynamically adjust the cluster division strategy of the distribution network through online learning and updating. Therefore, the dynamic distribution network cluster division method based on deep reinforcement learning algorithm can effectively overcome the shortcomings of traditional optimization methods and provide a more flexible, efficient and intelligent solution for distribution networks with a high proportion of distributed photovoltaic access.
[0086] like Figure 1 As shown, embodiment 1 of the present invention provides a method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning, comprising the following steps:
[0087] Step 1: Construct a CMDP (Constrained Markov decision processes) model for dynamic partitioning of distribution network clusters.
[0088] It is worth noting that Safe Deep Reinforcement Learning (SDRL) is an extension of classic DRL, adding a cost function to consider the safety of the strategy. Based on the CMDP model, the goal of SDRL is to maximize the long-term expected return while satisfying the long-term expected cost constraint. The present invention first models the cluster dynamic partitioning problem of the distribution network as CMDP. In a preferred but non-restrictive implementation, the construction of CMDP for cluster dynamic partitioning mainly includes the following four key elements:
[0089] State Space: is the state space of the intelligent agent, which contains all the state information in the actual distribution network, including: the active power and reactive power of the load, and the active and reactive power of photovoltaics. At this moment, the observed state of the agent is expressed by the following formula:
[0090] (1)
[0091] Where:
[0092] They represent the active and reactive load demand vectors of the distribution network at time t, respectively, and their dimensions are determined by the number of nodes in the distribution network;
[0093] They represent the active and reactive output vectors of photovoltaics in the distribution network at time t, respectively, and their dimensions are determined by the number of photovoltaic configurations in the distribution network.
[0094] Action Space: is the action space of the intelligent agent, which is the cluster division result of the distribution network. Assume that the number of nodes in the distribution network is , the number of clusters to be divided is Then the action of the agent is expressed by the following formula:
[0095] (2)
[0096] Where:
[0097] is a discrete integer variable, indicating the The cluster to which the node belongs is time.
[0098] Reward function: is the reward function of the agent, which is the feedback given by the environment after the action is performed. The reward function with a penalty term is ineffective in enhancing the safety of the policy and will have a negative impact on the convergence of the critic network. In the present invention, the CMDP model decouples the reward and cost functions to more accurately characterize the goals and safety constraints, thereby promoting stable and efficient training.
[0099] Given that cluster division needs to satisfy the characteristics of high coupling of nodes within a cluster and low coupling of nodes between clusters, we first define two indicators: intra-cluster aggregation and inter-cluster sparsity. They are expressed as:
[0100] (3)
[0101] (4)
[0102] Where:
[0103] Used to evaluate the coupling degree between nodes within the cluster. The larger it is, the higher the coupling degree between nodes in the cluster is;
[0104] Used to evaluate the coupling degree between nodes in the cluster. The larger it is, the lower the coupling degree of nodes between sub-regions is;
[0105] represents the electrical distance between nodes i and j in the distribution network;
[0106] is the set of internal node numbers of the sub-region to which node i belongs.
[0107] In addition, in order to improve the level of cluster autonomy, a power balance evaluation index including reactive power and active power balance is constructed, which is specifically expressed as follows:
[0108] (5)
[0109] (6)
[0110] (7)
[0111] Where:
[0112] It is the reactive power balance index;
[0113] c is the number of clusters;
[0114] For cluster The reactive balance degree, The maximum reactive power supply within the cluster, including the reactive power provided by the node reactive power compensation device and the reactive power that can be provided by some inverters. is the reactive power demand value within the cluster;
[0115] For cluster The degree of active balance, is the maximum value of active power supply within the cluster, is the required value of active power within the cluster.
[0116] In order to facilitate further solution, it is necessary to normalize the intra-group aggregation, inter-group sparsity and power balance indicators of different dimensions. The weighting method based on sensitivity analysis is adopted to study the influence of different weights on the objective function value by changing the weight values one by one, and finally select the appropriate weight. , to empower the three evaluation indicators. t At this moment, the cluster division evaluation index (i.e., reward function) of the distribution network can be expressed as:
[0117] (8)
[0118] Where:
[0119] Respectively At this moment, the system's intra-group aggregation, inter-group sparsity, and power balance.
[0120] Cost function: Considering that the cluster division is intended to facilitate the regulation of distribution network operation, when a single node forms a cluster, the actual regulation workload may increase. Therefore, when performing cluster division, constraints need to be added to determine whether there are isolated nodes in each cluster after the action given by the agent. Therefore, At this moment, the cost function of CMDP can be expressed as:
[0121] (9)
[0122] Where:
[0123] Represents a cluster The number of isolated islands in the
[0124] According to the established state, action, reward and cost function, the goal of CMDP is to maximize the long-term cumulative reward function value while the cost function value satisfies the constraints, which can be expressed as:
[0125] (10)
[0126] Where:
[0127] is the discount factor, is the cost limit, Represents the optimal strategy.
[0128] Step 2: Construct a secure DRL algorithm to solve the CMDP. Specifically, for the established CMDP problem, a primal-dual DRL method is adopted to solve it. The original problem defined in formula (10) is transformed using the Lagrangian relaxation method and can be restated as an unconstrained problem, expressed as follows:
[0129] (11)
[0130] Where:
[0131] is the Lagrange multiplier,
[0132] represents the optimal strategy,
[0133] represents the optimal Lagrange multiplier.
[0134] The primal-dual DRL method based on safe deep reinforcement learning emphasizes satisfying constraints and obtaining future rewards by introducing dual variables. During the training process, the present invention adopts the actor-critic (AC) DRL framework and uses an iterative primal-dual method to update the primal policy π and the dual variable λ in sequence. In this framework, the agent consists of an actor and two evaluators, namely a reward evaluator and a cost evaluator. The actor uses policy gradients to determine the optimal policy, while the reward evaluator approximates the expected reward by capturing the relationship between state, action, and reward. Similarly, the cost evaluator approximates the expected constraint cost by considering the relationship between state, action, and constraint violation.
[0135] The standard AC algorithm is mainly designed for the optimization problem of continuous decision variables, and it is difficult to directly solve discrete optimization problems such as cluster division. Therefore, as one of the outstanding substantive features of the present invention, the present invention improves the policy network and adopts the Gumbel-Softmax sampling technology.
[0136] In the policy network, the present invention uses the Gumbel-Softmax function to generate the sampling probability of the action and selects the corresponding discrete action accordingly. Gumbel-Softmax provides a differentiable strategy to approximate discrete distribution sampling, which is achieved by introducing Gumbel noise and adjusting the temperature parameter. This design allows discrete variables to be effectively processed when performing gradient descent optimization, which is particularly suitable for deep reinforcement learning (DRL) models involving backpropagation. The specific action selection process is based on the Gumbel-Softmax mechanism and is expressed as the following formula:
[0137] (12)
[0138] Where:
[0139] π is the strategy, which is formalized as a neural network whose input is the state and output is the partition number of each node;
[0140] Is The generated Gumbel noise is used to introduce randomness into the sampling process;
[0141] For Node The probability vector of belonging to which cluster;
[0142] is a vector of H random variables uniformly distributed between 0 and 1;
[0143] is a hyperparameter;
[0144] is the softmax function.
[0145] In order to reduce the overestimation bias of the evaluator network and enhance the stability and convergence of the algorithm, the double Q clipping learning technology is introduced in the process of training the action network and the evaluation network. First, the algorithm randomly samples a set of experience points from the buffer D. , and fix the Lagrange multiplier , train the action network to learn the optimal strategy that maximizes the Lagrangian function. Therefore, the strategy loss function based on double Q clipping learning is expressed as follows:
[0146] (13)
[0147] Where:
[0148] Indicates that the agent is in state according to the action network The sampled actions, φ is the neural network parameter of the discrete policy network; and Respectively i For reward and cost, dual Q clipping learning uses two evaluation networks and introduces a minimum clipping operation when calculating parameter updates to avoid extreme Q values affecting the update process.
[0149] The parameters of the main actor network are updated by stochastic gradient ascent, expressed as follows:
[0150] (14)
[0151] Where:
[0152] represents the learning rate,
[0153] is the loss function about The first derivative of .
[0154] The reward and cost evaluation network consists of a deep neural network with fully connected layers. The trainable parameters of the reward evaluation network are The parameters of the cost evaluation network are expressed as The reward evaluation network takes the observed state and action as input and generates the reward Q value of the state-action pair, which is recorded as The main reward critic network is trained using a loss function defined as the mean squared error between the predicted value (the output of the main reward critic network) and the target value (the immediate reward plus the output of the target critic network), expressed as follows:
[0155] (15)
[0156] Where:
[0157] is the loss function of the reward evaluation network;
[0158] is the loss function about The first derivative of ;
[0159] represents the target reward evaluation network;
[0160] Parameters of the evaluation network for the target reward;
[0161] Soft update coefficients for evaluating network parameters for target reward.
[0162] Given the same input as the reward critic, the cost critic network produces the cost Q-value of the state-action pair, expressed as As above, the loss function of the cost critic network is expressed as follows:
[0163] (16)
[0164] Where:
[0165] Evaluate the network for target cost;
[0166] Evaluate the parameters of the network for the target cost;
[0167] is the soft update coefficient of the target cost critic network parameters;
[0168] is the loss function about The first derivative of .
[0169] In the case of a fixed action network and cost evaluation network, the agent's dual variables λ The gradient ascent method can be used for updating, which can be expressed as follows:
[0170] (17)
[0171] Where:
[0172] To ensure that the dual variables .
[0173] Step 3: Train the AC-DRL-based solution framework constructed in step 2 to obtain a deployable CMDP model.
[0174] In a preferred but non-limiting embodiment of the present invention, the training process of the proposed secure AC algorithm can be divided into the following core steps:
[0175] ① Initialization phase: First, initialize the parameters of the Gumbel-Softmax-based policy network, cost evaluation network, and reward evaluation network. Create an empty experience playback buffer D to store the experience data of the interaction between the agent and the distribution network environment.
[0176] ② Experience collection: Collect experience data by executing the actions output by the action network in the distribution network environment. It mainly includes the distribution network status, cluster division results, reward value, and cost value, which can be expressed as , each set of experience data is stored in a buffer.
[0177] ③Data sampling: Randomly extract a batch of experience data from the playback buffer for subsequent network training.
[0178] ④ Evaluation network update: Use the cost evaluation network, reward evaluation network, and policy network to estimate the target cost and reward Q value for each sampled experience. Calculate the mean square error loss between the target cost, reward Q value, and actual Q value, and update the parameters θ and θ of the reward and cost evaluation networks accordingly. .
[0179] ⑤ Policy network update: Apply the gradient ascent method to update the parameter φ of the discrete policy network to optimize the action selection of the policy network, that is, the cluster division result.
[0180] ⑥ Soft update target network: Use soft update method to gradually adjust the parameters of target reward and cost evaluation network and , to ensure the stability of network parameter updates.
[0181] ⑦ Update the dual variable: fix the trained cost evaluation network and policy network parameters, use the gradient ascent method to update the dual variable, and ensure that it is greater than or equal to 0.
[0182] ⑧Training iteration: Repeat steps 2 to 7 until the set number of training times is reached.
[0183] Step 4: Online decision making and execution, deploy the deployable CMDP model obtained in step 3 to the distribution network dispatching cloud platform, and provide cluster division results based on the real-time observation status of the distribution network.
[0184] Specifically, after the algorithm training is completed, the parameters of the action network are fixed and deployed to the distribution network scheduling cloud platform. In the actual decision-making process, the algorithm only provides cluster division results based on the real-time observation status of the distribution network, without relying on complex calculations. This method enables the cluster division results to quickly respond to changes in the operating status of the distribution network, improves real-time performance and reliability, and reduces the impact of distributed photovoltaic and load uncertainties on the division results.
[0185] Embodiment 2 of the present invention provides a high-proportion distributed photovoltaic access distribution network cluster dynamic division system based on secure deep reinforcement learning, which runs the active distribution network cluster dynamic division method based on deep reinforcement learning according to embodiment 1, including:
[0186] Data acquisition module, used to input the real-time observation status of the distribution network;
[0187] CMDP module, with built-in CMDP model for dynamic partitioning of distribution network clusters;
[0188] Solver module, with built-in AC-DRL framework, for solving CMDP models;
[0189] The cluster division result output module is used to provide cluster division results according to the real-time observation status of the distribution network.
[0190] In order to more clearly introduce the outstanding essential features of the present invention and the significant progress it brings to the prior art, the specific scheme of the present invention is described below with a specific example.
[0191] The best way to implement the whole method is as follows:
[0192] The following takes the IEEE 33-node distribution system as an example to illustrate the method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning described in the present invention, wherein the system includes 32 load nodes, 9 photovoltaic power generation systems, and the number of clusters to be divided is set to 5.
[0193] (1) First, we build a CMDP for dynamic cluster partitioning. It mainly includes the following four key elements:
[0194] State Space: is the state space of the agent, which contains all the state information in the environment (i.e., the actual distribution network), including the active power and reactive power of the load, and the active and reactive output of photovoltaics. At time t, the observed state of the agent can be defined as:
[0195] (1)
[0196] in, They represent the active and reactive load demand vectors of the 33-bus system at time t, respectively, which are two 32×1 vectors. They represent the active and reactive outputs of photovoltaic power in the distribution network at time t, respectively, and are two 9×1 vectors.
[0197] Action Space: is the action space of the intelligent agent, that is, the cluster division result of the distribution network. At time T, the action of the intelligent agent can be defined as:
[0198] (2)
[0199] in, is a discrete integer variable, indicating the cluster to which the Nth node belongs at time t. In this embodiment, M=5, N=33.
[0200] Reward function: R is the reward function of the agent, which is the feedback given by the environment after the action is performed. Reward functions with penalty terms are ineffective in enhancing policy security and have a negative impact on the convergence of the critic network. In this method, the CMDP model decouples the reward and cost functions to more accurately characterize the goals and safety constraints, thereby promoting stable and efficient training. Given that cluster division needs to meet the characteristics of high coupling of nodes within the group and low coupling of nodes between groups, two indicators are first defined: intra-group aggregation and inter-group sparsity. They are expressed as:
[0201] (3)
[0202] (4)
[0203] In the formula, Used to evaluate the coupling degree between nodes within the cluster. The larger the value, the higher the coupling degree between nodes in the cluster. Used to evaluate the coupling degree between nodes in the cluster. The larger it is, the lower the coupling degree of nodes between sub-regions. Represents the electrical distance between nodes i and j in the distribution network. is the set of node numbers in the sub-area to which node i belongs. In the IEEE 33-node power distribution system, N=33.
[0204] In addition, in order to improve the level of cluster autonomy, a power balance evaluation index including reactive power and active power balance is constructed, which is specifically expressed as follows:
[0205] (5)
[0206] (6)
[0207] (7)
[0208] Where: is the reactive power balance index; c is the number of clusters; is the reactive balance degree of cluster i; The maximum value of reactive power supply within the cluster, including reactive power provided by the node reactive power compensation device and reactive power that can be provided by some inverters; is the reactive power demand value within the cluster. is the active power balance degree of cluster i; The maximum value of active power supply within the cluster; is the required value of active power within the cluster. In this embodiment, M=5.
[0209] In order to facilitate further solution, it is necessary to normalize the intra-group aggregation, inter-group sparsity and power balance indicators of different dimensions. The weighting method based on sensitivity analysis is adopted to study the influence of different weights on the objective function value by changing the weight values one by one, and finally select the appropriate weight. , to weight the three evaluation indicators. At time t, the cluster division evaluation indicator (i.e., reward function) of the distribution network can be expressed as:
[0210] (8)
[0211] In the formula, They represent the system's intra-group aggregation, inter-group sparsity, and power balance at time t. In this embodiment, They are 0.3, 0.3 and 0.4 respectively.
[0212] Cost function: Considering that the cluster division is intended to facilitate the regulation of distribution network operation, when a single node forms a cluster, the actual regulation workload may increase. Therefore, when performing cluster division, constraints need to be added to determine whether there are isolated nodes in each cluster after the action given by the agent. Therefore, at time t, the cost function of CMDP can be expressed as:
[0213] (9)
[0214] In the formula, Indicates the number of islands in cluster i.
[0215] According to the established state, action, reward and cost function, the goal of CMDP is to maximize the long-term cumulative reward function value while the cost function value satisfies the constraints, which can be expressed as:
[0216] (10)
[0217] In the formula, is the discount factor, is the cost limit, In this embodiment, =0.99, d=0,
[0218] (2) Construct a secure DRL algorithm to solve CMDP:
[0219] First, the original problem defined in (10) is transformed using the Lagrangian relaxation method and reformulated as an unconstrained problem:
[0220] (11)
[0221] Where: is the Lagrange multiplier, represents the optimal strategy, represents the optimal Lagrange multiplier.
[0222] During training, an iterative primal-dual approach is used to sequentially update the parameters of the primal policy π and the dual variable λ. In this framework, the agent consists of a policy network and two critic networks, a reward critic network and a cost critic network. The policy network uses policy gradients to determine the optimal policy, while the reward critic network approximates the expected reward by capturing the relationship between states, actions, and rewards. Similarly, the cost critic network approximates the expected constraint cost by considering the relationship between states, actions, and constraint violations.
[0223] In the policy network, there is an input layer, two hidden layers and an output layer. The number of neurons in the input layer corresponds to the dimension of the distribution network state, and the number of neurons in the output layer corresponds to the possible cluster number of each node. The output layer is activated by the Gumbel-Softmax function as follows:
[0224] (12)
[0225] Where: π is the strategy, which is formalized as a neural network. Its input is the state, and the output is the partition number of each node; Is The generated Gumbel noise is used to introduce randomness into the sampling process; is the probability vector of which cluster node i belongs to; is a vector of H random variables uniformly distributed between 0 and 1; is a hyperparameter; is the softmax function.
[0226] The feasibility and optimality of the actions output by the strategy network are evaluated, and a cost evaluation network and a reward evaluation network are constructed. The two networks have the same structure, including an input layer, three hidden layers and an output layer. The number of neurons in the input layer is equal to the sum of the distribution network state dimension and the action dimension, and the number of neurons in the output layer is 1, which outputs the cost value and reward value .
[0227] To train the action network and the evaluation network, the algorithm randomly samples a set of experience points from the buffer D. , and fix the Lagrangian multiplier λ, train the policy network to maximize the optimal policy of the Lagrangian function. Therefore, the policy loss function is defined as:
[0228] (13)
[0229] Where: represents the action sampled by the agent according to the action network in state , and φ are the neural network parameters of the discrete policy network;
[0230] Update the parameters of the main actor network through stochastic gradient ascent:
[0231] (14)
[0232] Wherein, represents that the learning rate is set to 0.0001, is the loss function with respect to the first derivative.
[0233] The trainable parameters of the cost evaluation network are represented by , and the parameters of the target evaluation network are represented by . The reward evaluation network takes the observed state and action as inputs and generates the reward Q value of the state-action pair, denoted as . The main reward evaluation network is trained using a loss function defined as the mean squared error between the predicted value (the output of the main reward critic network) and the target value (the immediate reward plus the output of the target critic network):
[0234] (15)
[0235] Wherein, J(θ) is the loss function of the reward evaluation network; ∇θ is the first derivative of the loss function J(θ) with respect to θ. represents the target reward evaluation network, are the parameters of the target reward evaluation network. is the soft update coefficient of the parameters of the target reward evaluation network.
[0236] Under the same input as the reward critic, the cost critic network generates the cost Q value of the state-action pair, denoted as . Similarly, the loss function of the cost critic network can be defined as:
[0237] (16)
[0238] Where: is the target cost evaluation network; are the parameters of the target cost evaluation network; is the soft update coefficient; is the normalized reward function vector; is the loss function with respect to the first derivative.
[0239] In the case of fixed action network and cost evaluation network, the agent's dual variable λ can be updated using the gradient ascent method, which can be expressed as:
[0240] (17)
[0241] In the formula, To ensure that the dual variables .
[0242] In summary, the training process of the proposed secure AC algorithm can be divided into the following core steps:
[0243] ① Initialization phase: First, initialize the parameters of the Gumbel-Softmax-based policy network, cost evaluation network, and reward evaluation network. Create an empty experience playback buffer D to store the experience data of the interaction between the agent and the distribution network environment.
[0244] ② Experience collection: Collect experience data by executing the actions output by the action network in the distribution network environment. It mainly includes the distribution network status, cluster division results, reward value, and cost value, which can be expressed as , each set of experience data is stored in a buffer.
[0245] ③Data sampling: Randomly extract a batch of experience data from the playback buffer for subsequent network training.
[0246] ④ Evaluation network update: Use the cost evaluation network, reward evaluation network, and policy network to estimate the target cost and reward Q value for each sampled experience. Calculate the mean square error loss between the target cost, reward Q value, and actual Q value, and update the parameters θ and θ of the reward and cost evaluation networks accordingly. .
[0247] ⑤ Policy network update: Apply the gradient ascent method to update the parameter φ of the discrete policy network to optimize the action selection of the policy network, that is, the cluster division result.
[0248] ⑥ Soft update target network: Use soft update method to gradually adjust the parameters of target reward and cost evaluation network and , to ensure the stability of network parameter updates.
[0249] ⑦ Update the dual variable: fix the trained cost evaluation network and policy network parameters, use the gradient ascent method to update the dual variable, and ensure that it is greater than or equal to 0.
[0250] ⑧Training iteration: Repeat steps 2 to 7 until the set number of training times is reached.
[0251] Finally, the trained cluster division strategy is deployed online through the action network to realize the implementation of the proposed method in the distribution system.
[0252] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning, characterized in that: The following steps are involved: Step 1: Construct a CMDP model for dynamic partitioning of distribution network clusters, which is expressed as follows: (1) Where: is the discount factor, is the cost limit, represents the optimal strategy; for The agent state space contains the distribution network state information at all times; for The action space of the agent as the result of the dynamic clustering of the distribution network at all times; is the reward function of the agent; is the cost function of the agent; At this moment, the cost function of the agent is also expressed as follows: (2) Where: Represents a cluster The number of isolated islands in the Step 2: Use the Lagrangian relaxation method to transform the CMDP model constructed in step 1 into an unconstrained problem, and build a solution framework based on AC-DRL to solve the CMDP model; use the Gumbel-Softmax function to generate the sampling probability of the action, and select the corresponding discrete action based on it, which is expressed as the following formula: (3) Where: π is the strategy, which is formalized as a neural network whose input is the state and output is the partition number of each node; Is The generated Gumbel noise is used to introduce randomness into the sampling process; For Node The probability vector of belonging to which cluster; is a vector of H random variables uniformly distributed between 0 and 1; is a hyperparameter; is a softmax function. In the process of training the action network and the evaluation network, the double Q clipping learning technology is introduced. The strategy loss function based on the double Q clipping learning is expressed as follows: (4) Where: Indicates that the agent is in state according to the action network The sampled actions, φ is the neural network parameter of the discrete policy network; and They are the reward value function and the cost value function respectively; Step 3: Train the AC-DRL-based solution framework constructed in step 2. After training, fix the parameters of the action network to obtain a deployable CMDP model. Step 4: Deploy the deployable CMDP model obtained in step 3 to the distribution network dispatching cloud platform, and provide cluster division results only based on the real-time observation status of the distribution network.
2. The method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning according to claim 1 is characterized in that: exist At this moment, the agent state space containing the distribution network state information is also expressed by the following formula: (5) Where: Respectively At time , the active and reactive load demand vectors of the distribution network; Respectively The active and reactive output vectors of photovoltaics in the distribution network at all times; The agent action space as a result of the dynamic division of the distribution network into dynamic clusters is also expressed by the following formula: (6) Where: is a discrete integer variable, indicating the The cluster to which the node belongs is time; is the number of distribution network nodes, is the number of clusters to be divided.
3. The method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning according to claim 1, characterized in that: exist At this moment, the agent's reward function is expressed as follows: (7) Where: Respectively At this moment, the system's intra-group aggregation, inter-group sparsity, and power balance.
4. The method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning according to claim 3 is characterized in that: The intra-cluster aggregation, inter-cluster sparsity, and power balance are also expressed by the following formula: (8) (9) (10) (11) (12) Where: Used to evaluate the coupling degree between nodes within the cluster; Used to evaluate the coupling degree between nodes in the cluster; Represents a distribution network node and nodes Electrical distance; For Node The set of internal node numbers of the sub-area to which it belongs; It is the reactive power balance index; c is the number of clusters; For cluster The reactive balance degree, The maximum reactive power supply within the cluster, including the reactive power provided by the node reactive power compensation device and the reactive power that can be provided by some inverters. is the reactive power demand value within the cluster; For cluster The degree of active balance, is the maximum value of active power supply within the cluster, is the required value of active power within the cluster.
5. The method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning according to claim 4 is characterized in that: Step 3 includes: ① Initialization phase: First, initialize the parameters of the Gumbel-Softmax-based policy network, cost evaluation network, and reward evaluation network; create an empty experience playback buffer , used to store the experience data of the interaction between the intelligent agent and the distribution network environment; ② Experience collection: By executing the actions output by the action network in the distribution network environment, experience data is collected, including: distribution network status, cluster division results, reward value, cost value, which can be expressed as , each set of experience data is stored in a buffer; ③Data sampling: randomly extract a batch of experience data from the playback buffer for subsequent network training; ④ Evaluation network update: Use the cost evaluation network, reward evaluation network and policy network to estimate the target cost and reward Q value of each sampled experience; calculate the mean square error loss between the target cost, reward Q value and actual Q value, and update the parameters θ and θ of the reward and cost evaluation networks accordingly. ; ⑤ Policy network update: Apply the gradient ascent method to update the parameter φ of the discrete policy network to optimize the action selection of the policy network, that is, the clustering result; ⑥ Soft update target network: Use soft update method to gradually adjust the parameters of target reward and cost evaluation network and , to ensure the stability of network parameter updates; ⑦ Update the dual variable: fix the trained cost evaluation network and policy network parameters, use the gradient ascent method to update the dual variable, and ensure that it is greater than or equal to 0; ⑧Training iteration: Repeat steps 2 to 7 until the set training end condition is reached.
6. A system for dynamic partitioning of active distribution network clusters based on deep reinforcement learning, running the method for dynamic partitioning of active distribution network clusters based on deep reinforcement learning according to any one of claims 1-5, characterized in that: include: Data acquisition module, used to input the real-time observation status of the distribution network; CMDP module, with built-in CMDP model for dynamic partitioning of distribution network clusters; Solver module, with built-in AC-DRL framework, for solving CMDP models; The cluster division result output module is used to provide cluster division results according to the real-time observation status of the distribution network.
Citation Information
Patent Citations
Active power distribution network multi-region division optimization method based on MOEA / D
CN114417566A
Power distribution network voltage reactive power control method and system based on safety reinforcement learning algorithm
CN116760047A
Power grid active scheduling intelligent decision-making method and system based on Lagrange relaxation
CN117254468A
Railway BIM data edge caching method based on multi-agent reinforcement learning
CN117473616A
Distributed photovoltaic optimization scheduling strategy method, device and system based on GA-MADRL-PPO combination
CN118157143A