Reactive voltage distributed control method, server, controller and storage medium
Through the control architecture parallel to online federated learning and distributed control, a multi-agent deep reinforcement learning algorithm is deployed, which solves the problem that reactive voltage control in massive DG scenarios in the existing technology requires a large amount of communication resources and privacy leakage, and achieves a low communication and high privacy protection reactive voltage control effect.
Patent Information
- Application Number
- CN202210873424.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The existing reactive voltage control method based on deep reinforcement learning of multi-agents requires a large number of communication resources and leaks the privacy of each VPP, and cannot take into account the global model-free adaptive optimization capabilities and local fast control capabilities.
Adopting a control architecture parallel to online federated learning and distributed control, multi-agent deep reinforcement learning algorithm is deployed, marginal variable exchange and local learning are performed on local controllers, and the server is centralized to coordinate the learning process of each local controller to realize the reactive voltage control of the power grid.
By reducing communication volume, saving communication costs and protecting the privacy of each subject, low communication and high privacy protection reactive voltage control is achieved, with global model-free adaptive optimization capabilities and local fast control capabilities.
Smart Images

Figure CN115377993B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power system operation and control, and in particular relates to a reactive voltage distributed control method, a server, a controller and a storage medium. Background Art
[0002] Large-scale distributed generation (DG) has become an important part of the new power system. If there is a lack of effective coordinated control, the grid connection of massive DG will seriously affect the safe operation of the system. DG is usually connected to the distribution network by power electronic converters. It has fast control performance. Combined with new technologies such as artificial intelligence and 5G communication, it can become a flexible resource for the control of the receiving power grid, actively participate in the control of the power grid, and support the construction of new power systems. In this context, the value of various advanced operation control methods of distribution networks with high DG penetration is becoming increasingly prominent. Among them, reactive voltage control has always been an important component of the energy management system of the distribution network. It can eliminate safety hazards such as voltage over-limit of the distribution network and improve the operation efficiency of the distribution network by rationally configuring various reactive resources in the distribution network.
[0003] Traditional reactive voltage control methods often adopt a model-based centralized optimization architecture, that is, the state information of the entire distribution network is collected at the distribution network control center, mathematical optimization is performed through the pre-established distribution network model, and the optimized control instructions are issued to each controlled device. However, this model-based centralized optimization architecture often faces the following two problems in practical applications, which limits the application of model-based centralized reactive voltage control methods. On the one hand, the centralized architecture has key problems such as single point failure, high communication and computing burden, serious impact of communication delay, and difficulty in protecting the privacy of each subject. This is because in a high-penetration distribution network, there are many controlled DGs, the network structure is complex, and the communication conditions are difficult to support fast DG control. In engineering practice, virtual power plants (VPPs) are often used to aggregate DGs to form multiple subjects within the distribution network. Not only do the subjects need to be coordinated and controlled, but the privacy of each subject also needs to be guaranteed. Therefore, a large number of distributed reactive voltage control methods have emerged in the industry, and local controllers are set on each VPP to achieve reactive voltage control under point-to-point communication conditions or no communication conditions. Compared with centralized methods, distributed methods often have weaker requirements on communication conditions, faster control speed, and are suitable for large-scale DG control. On the other hand, massive DGs are incorporated into the distribution network, and their models are difficult to maintain accurately and in a timely manner. This is because the huge scale of equipment and frequent changes lead to huge maintenance costs for the distribution network model. In fact, theoretical parameters are basically used with low accuracy. Since the ideal model of the distribution network is difficult to obtain, traditional model-based optimization control methods cannot guarantee the control effect, and often the control instructions are far away from the optimal point and the power grid operates in a suboptimal state. Therefore, a number of model-free reactive voltage optimization methods based on deep reinforcement learning have emerged in the industry, which can adaptively obtain the optimal reactive voltage control strategy through online learning of the massive data generated by the control process.
[0004] Combining the improvements in the above two directions, a class of reactive voltage control methods based on multi-agent deep reinforcement learning has emerged in the industry. Through the "centralized learning-decentralized control" approach, it can combine the global optimization capability of model-free adaptation and the high-speed and flexible local control capability of decentralized control, thereby improving the practicality of reactive voltage control in massive DG scenarios. However, the existing reactive voltage control methods based on multi-agent deep reinforcement learning require that the operating data in each local controller be uploaded to a centralized server for learning, and the learned strategies be distributed to each local controller. This process not only consumes a large amount of communication resources, but also causes the privacy of each VPP to be leaked.
[0005] Therefore, it is necessary to propose a low-communication, high-privacy protection control method that combines global model-free adaptive optimization capabilities and local fast control capabilities while using only a small amount of communication bandwidth, saving communication costs and protecting the privacy of each subject. Summary of the invention
[0006] In view of the above problems, the present invention provides a reactive voltage distributed control method, which adopts a control architecture of online federated learning and distributed control in parallel, including:
[0007] Deploy multi-agent deep reinforcement learning algorithms on the local controller;
[0008] Through the marginal variable exchange between the local controller and the centralized server, the local controller learns the control strategy of each local controller;
[0009] Each local controller performs grid reactive power and voltage control based on control strategies and local measurement data;
[0010] The centralized server is used to coordinate the learning process of the intelligent agents on each local controller.
[0011] Furthermore, the following steps are performed on each regional local controller:
[0012] Framework construction steps: Construct a multi-agent interactive training framework based on Markov game;
[0013] Network deployment steps: deploy the neural network;
[0014] Local control steps: parallel execution of reactive power and voltage control;
[0015] Local exchange step: executing marginal variable exchange with the centralized server in parallel, including obtaining downlink marginal variables from the centralized server and sending uplink marginal variables to the centralized server;
[0016] Local learning step: Perform local learning in parallel based on the downstream marginal variables;
[0017] The downlink marginal variables are obtained by the centralized server after centralized learning based on the collected uplink marginal variables;
[0018] The local control step, the local exchange step and the local learning step are repeatedly performed in parallel.
[0019] Further, the upstream marginal variable includes scalar data calculated by the local controller;
[0020] The downstream margin variables include data used to calculate gradient information.
[0021] Furthermore, the centralized server is used to establish a distributed reactive power voltage optimization model, including determining reactive power voltage control targets in each region:
[0022] The distribution network is divided into |A| areas according to the actual installation of local controllers. The area label set is denoted by A. A local controller is set in each area to quickly obtain the measurement data of the area. The local controller can communicate with the centralized server. Among them, the reactive voltage control target of each area is to minimize the network loss of the whole network under the condition of satisfying the voltage constraint:
[0023]
[0024] Among them, L opt is the distributed reactive power voltage optimization target, P loss is the total network loss, Represent the active power injection and voltage amplitude of the jth node in the ith region, V, are the lower and upper limits of the voltage amplitude respectively; q S,i represents the reactive power output of all static VAR compensators SVC in the i-th region, q S,1 represents the reactive power output of all SVCs in the first zone, q S,|A| represents the reactive power output of all SVCs in the |A|th region, in represents the reactive power output of the jth SVC in the ith zone, They represent the lower and upper limits of the reactive power output of the jth SVC in the i-th region respectively; q G,i represents the reactive power output of all DGs in the ith area, q G,1 represents the reactive power output of all DGs in the first area, q G,|A| represents the reactive power output of all DGs in the |A|th area, where represents the reactive power output of the jth DG in the ith region, is the installed capacity and active power output of the jth DG in the ith region, N i is the set of nodes in the ith region, is the set of DG of the ith region, is the set of SVCs in the i-th region.
[0025] Furthermore, a multi-agent interactive training framework based on Markov game is constructed, including:
[0026] Corresponding to the measurement data of each regional system, construct the observation variables o of each region i As shown in (1.2):
[0027]
[0028] represents the reactive power injection of the jth node in the i-th region;
[0029] Corresponding to the reactive power and voltage optimization objectives of each region, the unified feedback variable r of each region is constructed as shown in (1.3):
[0030]
[0031] Among them, C V is the penalty coefficient of voltage constraint; function [] + is defined as [x] + =max(0,x);
[0032] Decomposing the unified feedback variable by region, we can obtain:
[0033]
[0034] Among them, R i is the decomposed variable of r in the ith region, i.e., the local feedback variable of the ith region, which can be calculated through the local measurement data of the ith region;
[0035] Corresponding to the reactive power of controllable flexibility resources, construct the action variables a of each area i As shown in (1.5):
[0036]
[0037] in, represents the reactive power output of the kth SVC in the ith zone, Represents the reactive power output of the jth DG in the ith zone. The controllable flexibility resources include SVC and DG.
[0038] Furthermore, the neural network is deployed on each local controller in each region, including:
[0039] Deploy the neural network on the local controller of each region and initialize it. For any i-th region, the following steps are included:
[0040] Define the first neural network c of the i-th region i For an input (o i ,a i ), a fully connected neural network that outputs a single scalar value, denoted by c i The network parameters are The corresponding freezing parameters are The neural network corresponding to the frozen parameters is recorded as Random Initialization and The first neural network is used to estimate the expected reward;
[0041] Define the second neural network e of the i-th region i For an input (o i ,a i ) is a fully connected neural network that outputs a single scalar value, denoted by e i The network parameters are The corresponding freezing parameters are The neural network corresponding to the frozen parameters is recorded as Random Initialization and The second neural network is used to evaluate entropy;
[0042] Define the third neural network And the fourth neural network in, and All are input o i Output k i dimensional vector policy neural network, where k i is the action vector a i Dimensions; and They each have an independent output layer, while sharing the same neural network input layer and hidden layer, including several hidden layers; and All network parameters are θ i .
[0043] Furthermore, the method further comprises:
[0044] Initialize the discrete time variables, set the current time t = 0, and the actual time interval between two iterations is Δt, which is determined according to the actual measurement data and command control speed of the local controller;
[0045] Initialize each local controller's local cache experience database D i ,i∈A, for any i-th region, Initialize temporary variables
[0046] Furthermore, each regional local controller executes the marginal variable exchange step in parallel, including:
[0047] Obtain measurement data from the measurement device of the regional power grid and update the corresponding observation variable o i ;
[0048] According to the local policy neural network and Update the action variable a corresponding to this iteration i :
[0049]
[0050] ξ i represents the i-th random variable, N represents the normal distribution, I i Indicates k i -dimensional unit vector;
[0051] The local controller will a i Send it to the corresponding local DG and / or SVC;
[0052] Each regional local controller performs the following marginal variable exchange steps at time t in parallel:
[0053] According to the current time, randomly in D i The sample is drawn from i ,a i ,R i ,o′ i}, and calculate the upward marginal variable:
[0054]
[0055] Among them, o′ i for o i The sample of a′ i and The calculation method is:
[0056]
[0057]
[0058] The upward marginal variable u i Upload to the centralized server via the communication network.
[0059] Furthermore, each local controller collects the downlink marginal variables sent by the centralized server and executes the learning steps in parallel, including:
[0060] According to the downside marginal variable d i The following calculation is performed with local data:
[0061]
[0062]
[0063]
[0064]
[0065] in, represents the target value of the entropy estimation function of the ith region, π i represents the policy function of the ith region, represents the loss function of the neural network e, γ represents the reinforcement learning discount value, α represents the penalty coefficient, represents the gradient of the first neural network c in the i-th region, is the total network parameters θ of the ith region i The gradient of It means to find the expectation of all observed variables and random variables in actions; It means to find the expectation of o,a,r,o' on the data set, o=[o1,…,o n ] T represents the observed variables of all regions, a=[a1,…,a n ] T represents the action variable of all regions, and o′1 is the sample observation variable of all regions;
[0066]
[0067] in, represents the inferred action variable of the entire region, c i represents the neural network of the i-th region;
[0068] Denotes the first neural network c of neural network f for region i i The partial derivatives of the values; Denotes the loss function L φ Partial derivatives of the neural network f;
[0069] Update Parameters With θ i :
[0070]
[0071]
[0072]
[0073] Among them, η θ represents the learning rate of the policy neural network, L θ represents the loss function of the policy neural network parameters, η Q is a hyperparameter representing the learning rate of the value function;
[0074] Update frozen parameters and
[0075]
[0076]
[0077] Furthermore, the downstream marginal variables are obtained by the centralized server collecting upstream marginal variables of each region and performing centralized learning:
[0078] Collect the upstream marginal variables u uploaded by each local controller i ;
[0079] According to u i Information calculation in:
[0080]
[0081] Among them, y represents the target value of the value function, γ represents the discount value of reinforcement learning, represents the neural network corresponding to the frozen parameters of the first region, (o′1, a′1) represents the input of the neural network corresponding to the frozen parameters of the first region, where o′1 is the sample observation variable of the first region, and a′1 is the sample action variable of the first region; represents the neural network corresponding to the frozen parameters of the nth region, (o′ n ,a′ n ) represents the input of the neural network corresponding to the frozen parameters of the nth region, where o′ n is the sample observation variable of the nth region, a′ n is the sample action variable of the nth region; the neural network f is a fully connected neural network with an input of |A| scalars and an output of 1 scalar value, and the network parameter of f is φ f , the corresponding freezing parameter is The neural network corresponding to the frozen parameters is recorded as
[0082]
[0083] Among them, L φ represents the loss function of the neural network parameter φ, represents the network parameters of the first neural network c1 of the first region, The first neural network c represents the nth region n The network parameters, φ f Gradient Calculated by the following formula:
[0084]
[0085] Update network parameters φ f :
[0086]
[0087] Step 5-4: Update the freeze parameters
[0088]
[0089] Where η is a hyperparameter, representing the freezing coefficient;
[0090] Form the downlink marginal variable d sent to each local controller i ;
[0091] Each downside marginal variable d i It is sent to the corresponding local controller through the communication network.
[0092] The present invention also provides a reactive voltage distributed control method, which adopts a control architecture of online federated learning and distributed control in parallel, including:
[0093] The learning process of the agents on each local controller is coordinated by a centralized server;
[0094] A multi-agent deep reinforcement learning algorithm is deployed on the local controller;
[0095] The centralized server exchanges marginal variables with the local controllers so that the local controllers learn the control strategies of the local controllers and perform reactive power voltage control of the power grid based on the control strategies and local measurement data.
[0096] Further, the method comprises:
[0097] Model determination steps: Establish a distributed reactive power voltage optimization model;
[0098] Network deployment steps: deploy the neural network on a centralized server;
[0099] Centralized learning steps: The centralized server collects the uplink marginal variables of each region, performs centralized learning, and sends the downlink marginal quantities to each local controller;
[0100] The downstream marginal variables are used by each local controller to perform local learning in parallel to obtain a control strategy;
[0101] Among them, the centralized learning step is repeatedly performed.
[0102] Further, the upstream marginal variable includes scalar data calculated by the local controller;
[0103] The downstream margin variables include data used to calculate gradient information.
[0104] Furthermore, establishing a distributed reactive power voltage optimization model includes determining the reactive power voltage control target of each region:
[0105] The distribution network is divided into |A| areas according to the actual installation of local controllers. The area label set is denoted by A. A local controller is set in each area to quickly obtain the measurement data of the area. The local controller can communicate with the centralized server. Among them, the reactive voltage control target of each area is to minimize the network loss of the whole network under the condition of satisfying the voltage constraint:
[0106]
[0107] Among them, L opt is the distributed reactive power voltage optimization target, P loss is the total network loss, Represent the active power injection and voltage amplitude of the jth node in the ith region, V, are the lower and upper limits of the voltage amplitude respectively; q S,i represents the reactive power output of all static VAR compensators SVC in the i-th region, q S,1 represents the reactive power output of all SVCs in the first zone, q S,|A| represents the reactive power output of all SVCs in the |A|th region, in represents the reactive power output of the jth SVC in the ith zone, They represent the lower and upper limits of the reactive power output of the jth SVC in the i-th region respectively; q G,i represents the reactive power output of all DGs in the ith area, q G,1 represents the reactive power output of all DGs in the first area, q G,|A| represents the reactive power output of all DGs in the |A|th area, where represents the reactive power output of the jth DG in the ith region, is the installed capacity and active power output of the jth DG in the ith region, N i is the set of nodes in the ith region, is the set of DG of the ith region, is the set of SVCs in the i-th region.
[0108] Furthermore, the local controller has a multi-agent interactive training framework based on Markov game formed in the following manner:
[0109] Corresponding to the measurement data of each regional system, construct the observation variables o of each region i As shown in (1.2):
[0110]
[0111] represents the reactive power injection of the jth node in the i-th region;
[0112] Corresponding to the reactive power and voltage optimization objectives of each region, the unified feedback variable r of each region is constructed as shown in (1.3):
[0113]
[0114] Among them, C V is the penalty coefficient of voltage constraint; function [] + is defined as [x] + =max(0,x);
[0115] Decomposing the unified feedback variable by region, we can obtain:
[0116]
[0117] Among them, R i is the decomposed variable of r in the ith region, i.e., the local feedback variable of the ith region, which can be calculated through the local measurement data of the ith region;
[0118] Corresponding to the reactive power of controllable flexibility resources, construct the action variables a of each area i As shown in (1.5):
[0119]
[0120] in, represents the reactive power output of the kth SVC in the ith zone, Represents the reactive power output of the jth DG in the ith zone. The controllable flexibility resources include SVC and DG.
[0121] Further, the neural network is deployed and initialized on a centralized server;
[0122] Define the neural network f as a fully connected neural network with an input of |A| scalars and an output of 1 scalar value; let the network parameter of f be φ f , the corresponding freezing parameter is The neural network corresponding to the frozen parameters is recorded as Randomly initialize φ f and
[0123] Furthermore, the centralized server collects the uplink marginal variables of each region and performs the learning steps and sends the downlink marginal quantities to each local controller, including:
[0124] Collect the upstream marginal variables u uploaded by each local controller i ;
[0125] According to ui Information calculation in:
[0126]
[0127] Among them, y represents the target value of the value function, γ represents the discount value of reinforcement learning, represents the neural network corresponding to the frozen parameters of the first region, (o′1, a′1) represents the input of the neural network corresponding to the frozen parameters of the first region, where o′1 is the sample observation variable of the first region, and a′1 is the sample action variable of the first region; represents the neural network corresponding to the frozen parameters of the nth region, (o′ n ,a′ n ) represents the input of the neural network corresponding to the frozen parameters of the nth region, where o′ n is the sample observation variable of the nth region, a′ n is the sample action variable of the nth region;
[0128]
[0129] Among them, L φ represents the loss function of the neural network parameter φ, represents the network parameters of the first neural network c1 of the first region, The first neural network c represents the nth region n The network parameters, Indicates the expectation of o, a, r, o' on the data set; φ f Gradient Calculated by the following formula:
[0130]
[0131] Update network parameters φ f :
[0132]
[0133] where η Q is a hyperparameter representing the learning rate of the value function;
[0134] Step 5-4: Update the freeze parameters
[0135]
[0136] Where η is a hyperparameter, representing the freezing coefficient;
[0137] Form the downlink marginal variable d sent to each local controller i :
[0138]
[0139] Where o=[o1,…,o n ] T represents the observed variables of all regions, a=[a1,…,a n ] T represents the action variables of the entire region, represents the inferred action variable of the entire region, c i represents the neural network of the i-th region;
[0140] Denotes the first neural network c of neural network f for region i i The partial derivatives of the values; Denotes the loss function L φ Partial derivatives of the neural network f;
[0141] Each downside marginal variable d i It is sent to the corresponding local controller through the communication network.
[0142] The present invention provides a controller for reactive voltage distributed control, wherein the controller is a local controller in a control architecture capable of applying online federated learning and distributed control in parallel;
[0143] The distribution network is divided into several controlled areas, each of which includes multiple nodes;
[0144] The node of each controlled area is connected to the local controller of the controlled area for communication, and the collected measurement data is uploaded to the local controller of the controlled area;
[0145] Each controller is in communication with a centralized server, which coordinates the learning process of the agents on each local controller;
[0146] Deploy multi-agent deep reinforcement learning algorithms on the local controller;
[0147] The local controller exchanges marginal variables with the centralized server, and the local controller learns the control strategy of each local controller;
[0148] Each local controller performs grid reactive power and voltage control based on the control strategy and local measurement data.
[0149] The present invention provides a server for reactive voltage distributed control, wherein the server is a centralized server in a control architecture capable of applying online federated learning and distributed control in parallel;
[0150] The centralized server is used to coordinate the learning process of the agents on each local controller;
[0151] The distribution network is divided into several controlled areas, each of which includes multiple nodes; the nodes of each controlled area are connected to the local controller of the controlled area in communication, and the collected measurement data are uploaded to the local controller of the controlled area;
[0152] The centralized server is in communication connection with the local controller and exchanges marginal variables with the local controller;
[0153] The local controller deploys a multi-agent deep reinforcement learning algorithm, which can learn the control strategy of each local controller based on the exchanged marginal variables, and perform grid reactive voltage control based on the control strategy and local measurement data.
[0154] The present invention provides a computer storage medium having computer executable instructions stored thereon. When executed by a computing device, the computer executable instructions enable the computing device to execute the above-mentioned reactive voltage distributed control method capable of running on a local controller.
[0155] The present invention provides a computer storage medium having computer executable instructions stored thereon. When executed by a computing device, the computer executable instructions enable the computing device to execute the above-mentioned reactive voltage distributed control method that can be run on a centralized server.
[0156] The present invention adopts a distributed architecture, and can achieve approximately optimal reactive voltage control without the need for an accurate model of the controlled distribution network. Each local controller can quickly acquire and execute the control strategy by exchanging marginal variables and performing local learning with a centralized server, thereby ensuring high-speed local control. The marginal variable exchange method has low communication volume and strong privacy protection, which can save communication bandwidth and protect the privacy of each controlled area. At the same time, compared with the distributed learning method that relies on the P2P (peer-to-peer) architecture, the present invention makes full use of the centralized coordination architecture of the power grid and has higher training efficiency. Other features and advantages of the present invention will be described in the subsequent description, and part of them will become apparent from the description, or can be understood through the implementation of the present invention. The objects and other advantages of the present invention can be achieved and obtained through the structures indicated in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0157] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0158] Figure 1 A flow chart of a reactive voltage distributed control method according to an embodiment of the present invention is shown;
[0159] Figure 2 A flow chart of another reactive voltage distributed control method according to an embodiment of the present invention is shown;
[0160] Figure 3 A schematic diagram of the structure of a reactive voltage distributed control system according to an embodiment of the present invention is shown;
[0161] Figure 4 A schematic structural diagram of a computing device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0162] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0163] The embodiment of the present invention proposes a method for distributed control of reactive voltage, which combines a federated learning architecture with multi-agent reinforcement learning, adopts a control architecture that combines online federated learning with distributed control in parallel, deploys an efficient multi-agent deep reinforcement learning algorithm, and coordinates the learning process of agents on each local controller through a centralized server; the control strategy of each controller is learned through asynchronous exchange of marginal variables between the local controller and the centralized server, and each controller performs reactive voltage control of the power grid based on the local control strategy and local measurement data.
[0164] Exemplarily, the distribution network includes a plurality of controlled areas, each of which is provided with a controller, and each controller can communicate with a cloud server of a regional power grid control center to exchange marginal variables. Each controlled area includes a plurality of nodes, and includes an SVC node with a static VAR compensator SVC and / or a DG node with a renewable energy generation device DG. The nodes in each controlled area are connected by branches to form a distribution network. Adjacent controlled areas have a communication connection relationship, for example, communication connection is performed through a controller.
[0165] The reactive voltage distributed control method of the embodiment of the present invention can be applied to the centralized server and / or local controller in the above-mentioned distribution network. Implementing the method of the embodiment of the present invention on a centralized server or a local controller alone, or implementing the method of the embodiment of the present invention on a centralized server and a local controller at the same time, all fall within the protection scope of the present invention.
[0166] An aspect of an embodiment of the present invention provides a reactive voltage distributed control method capable of running on a local controller, wherein the method adopts a control architecture in which online federated learning and distributed control are carried out in parallel, such as Figure 1 As shown, the method includes: deploying a multi-agent deep reinforcement learning algorithm on a local controller; the local controller learns the control strategy of each local controller through the exchange of marginal variables between the local controller and the centralized server; each local controller performs reactive power voltage control of the power grid based on the control strategy and local measurement data; wherein the centralized server is used to coordinate the learning process of the agents on each local controller.
[0167] Specifically, the method includes executing the following steps on the local controller of each region:
[0168] Framework construction steps: Construct a multi-agent interactive training framework based on Markov game;
[0169] Network deployment steps: deploy the neural network;
[0170] Local control steps: parallel execution of reactive power and voltage control;
[0171] Local exchange step: parallel execution of marginal variable exchange with the centralized server, including obtaining downstream marginal variables from the centralized server and sending upstream marginal variables to the centralized server;
[0172] Local learning step: Perform local learning in parallel based on the downstream marginal variables;
[0173] Among them, the downlink marginal variables are obtained by the centralized server after performing centralized learning based on the collected uplink marginal variables;
[0174] The local control step, the local exchange step and the local learning step are repeatedly performed in parallel.
[0175] Another aspect of the embodiment of the present invention provides a reactive voltage distributed control method that can be run on a centralized server, using a control architecture that parallels online federated learning and distributed control, such as Figure 2 As shown, the method includes: coordinating the learning process of the intelligent agents on each local controller through a centralized server; deploying a multi-agent deep reinforcement learning algorithm on the local controller; the centralized server exchanges marginal variables with the local controller so that the local controller learns the control strategy of each local controller, and performs reactive voltage control of the power grid based on the control strategy and local measurement data.
[0176] Specifically, the method comprises the following steps:
[0177] Model determination steps: Establish a distributed reactive power voltage optimization model;
[0178] Network deployment steps: deploy the neural network on a centralized server;
[0179] Centralized learning steps: The centralized server collects the uplink marginal variables of each region, performs centralized learning, and sends the downlink marginal quantities to each local controller;
[0180] The downstream marginal variables are used by each local controller to perform local learning in parallel to obtain the control strategy;
[0181] Among them, the centralized learning step is repeatedly performed.
[0182] Compared with the P2P distributed learning method, this method can achieve higher training efficiency. In federated learning, each controller does not use the local observation variables (o i ), action variables (a i ) and control strategy (neural network function and ) is uploaded to the centralized server, but the scalar data calculated by the local controller (agent) is uploaded to the centralized server. The centralized server transmits the data used to calculate the gradient information (specifically, the partial derivative used to calculate the gradient information) to the local controller. During the learning phase, the local controller performs updates on its own, which not only utilizes global information but also protects individual privacy. Among them, a local controller can be regarded as an agent.
[0183] The embodiment of the present invention adopts an efficient multi-agent deep reinforcement learning algorithm, which has both global model-free adaptive optimization capability and local fast control capability, thereby realizing optimal control of the reactive voltage of the power grid under incomplete model scenarios; through the organic combination of federated learning, the control data and strategy information of each controller are only processed locally, which not only avoids a large amount of data exchange and saves communication costs, but also protects the privacy of each subject; and thus has low communication requirements, high privacy protection, high efficiency, and high flexibility, thereby greatly improving the voltage quality of the power grid and reducing the network loss of the power grid operation under poor communication conditions and poor model conditions.
[0184] The local controller of the embodiment of the present invention can perform rapid control according to the real-time local measurement data, giving full play to the speed advantage of distributed control, and is particularly suitable for distribution networks with a large number of high-speed DG resources; on the other hand, the present invention does not need to expose the local measurement data and control strategies of each intelligent agent, and only needs to exchange a small number of marginal variables, with low communication bandwidth requirements, which can save communication costs and protect the privacy of each subject, and is particularly suitable for use in distribution networks with multiple VPP access. The control method of the embodiment of the present invention also combines the flexible and efficient advantages of distributed control with the global optimization capabilities of multi-agent reinforcement learning, which not only avoids the high risk of single-point failure and large control instruction delay caused by centralized control, but also saves the high cost of repeated maintenance of accurate models, and avoids the reduction of control effect caused by rough models; at the same time, it reduces the requirements for communication bandwidth, saves communication costs and protects the privacy of each subject, and is suitable for large-scale promotion.
[0185] The following takes the cooperation between the local controller and the centralized server as an example to explain in detail the reactive power voltage distributed control method applied on the local controller and the centralized server.
[0186] Step 1:
[0187] Model determination step: determine the reactive power and voltage control targets of each controlled area on the centralized server and establish a distributed reactive power and voltage optimization model;
[0188] Framework construction steps: Based on the distributed reactive power and voltage optimization model, a multi-agent interactive training framework based on Markov game is constructed on the local controller, including:
[0189] Step 1-1, establish a distributed reactive voltage optimization model: The distribution network is divided into |A| areas according to the actual controller installation situation, and the area label set is denoted by A. Each area is installed with a controller (i.e., a controlled area controller or a local controller) to quickly obtain the measurement data of the area. The controller has a certain computing power and can communicate with the centralized server (i.e., the cloud server of the regional power grid control center). Establishing a distributed reactive voltage optimization model includes determining the reactive voltage control target of each area, where the reactive voltage control target of each area is to minimize the network loss of the entire network while satisfying the voltage constraint:
[0190]
[0191] Among them, L opt is the distributed reactive power voltage optimization target, P loss is the total network loss, represents the active power injection and voltage amplitude of the jth node in the i-th region, V, are the lower and upper limits of the voltage amplitude respectively; q S,i represents the reactive power output (q) of all SVCs (Static Var Compensator) in the ith region S,1 represents the reactive power output of all SVCs in the first zone, q S,|A| represents the reactive power output of all SVCs in the |A|th region), where represents the reactive power output of the jth SVC in the ith zone, represent the lower and upper limits of reactive power output of the jth SVC in the ith region, respectively; q G,i represents the reactive power output of all DGs in the ith region, ( q G,1 It represents the reactive power output of all DGs in the first area when i=1. q G,|A| represents the reactive power output of all DGs in the |A|th area), where represents the reactive power output of the jth DG in the ith region, is the installed capacity and active power output of the jth DG in the i-th region. i is the set of nodes in the ith region, is the set of DG of the ith region, is the set of SVCs in the ith region. In addition, the distribution network in the controlled region also satisfies the power flow constraint, and the power flow constraint equation is well known to those skilled in the art and is omitted here.
[0192] Construct a multi-agent interactive training framework based on Markov game, including:
[0193] Step 1-2: Construct the observed variables o for each region according to the measurement data of each regional system i As shown in (1.2).
[0194]
[0195] represents the reactive power injection of the jth node in the i-th region;
[0196] In the embodiments of the present invention, unless otherwise specified, the same symbols represent the same physical parameters and are not described in detail.
[0197] Step 1-3: Corresponding to the reactive power and voltage optimization objectives of each region, construct the unified feedback variable r of each region as shown in (1.3).
[0198]
[0199] Among them, C Vis the penalty coefficient of voltage constraint, and for example, the value is 1; function [] + is defined as [x] + =max(0,x).
[0200] Decomposing the unified feedback variable by region, we can obtain:
[0201]
[0202] Among them, R i is the decomposed variable of r in the ith region, that is, the local feedback variable of the ith region, which can be calculated through the local measurement data of the ith region.
[0203] Step 1-4: Construct the action variable a of each area corresponding to the reactive power of the controllable flexibility resource i As shown in (1.5).
[0204]
[0205] in, represents the reactive power output of the kth SVC in the ith zone. represents the reactive power output of the jth DG in the ith region. The controllable flexibility resources include SVC and DG.
[0206] Step 2: Deploy neural networks on the centralized server and local controllers in each region, and initialize each neural network and related control process variables, specifically a series of variables used for the control process described in steps 2-1 to 2-4;
[0207] Step 2-1: Deploy the neural network on a centralized server and initialize it;
[0208] Define neural network f as a fully connected neural network with an input of |A| scalars and an output of 1 scalar value, including several hidden layers (for example, 2 hidden layers), each hidden layer contains several neurons (for example, 512 neurons), and the activation function is the ReLU function. Let the network parameter of f be φ f , the corresponding freezing parameter is The neural network corresponding to the frozen parameters is recorded as Randomly initialize φ f and
[0209] Step 2-2: Deploy and initialize the neural network on the local controller of each region. Taking the i-th region as an example, the steps include:
[0210] Step 2-2-1: Define the first neural network c of the i-th region i For an input (o i ,ai ), a fully connected neural network that outputs a single scalar value, including several hidden layers (for example, 2 hidden layers), each hidden layer contains several neurons (for example, 512 neurons), and the activation function is the ReLU function. i The network parameters are The corresponding freezing parameters are The neural network corresponding to the frozen parameters is recorded as Random Initialization and The first neural network is used to estimate the expected reward.
[0211] Step 2-2-2: Define the second neural network e of the i-th region i For an input (o i ,a i ) is a fully connected neural network that outputs a single scalar value, including several hidden layers (for example, 2 hidden layers), each hidden layer contains several neurons (for example, 512 neurons), and the activation function is the ReLU function. i The network parameters are The corresponding freezing parameters are The neural network corresponding to the frozen parameters is recorded as Random Initialization and The second neural network is used to estimate entropy.
[0212] Preferably, the first neural network and the second neural network have the same structure.
[0213] Step 2-2-3: Define the third neural network And the fourth neural network in, and All are input o i Output k i dimensional vector neural network, where k i is the action vector a i Dimension. and Each has an independent output layer, while sharing the same neural network input layer and hidden layer, including several hidden layers (for example, 2 hidden layers), each hidden layer contains several neurons (for example, 512 neurons), and the activation function is the ReLU function. and All network parameters are θ i . Randomly initialize θ i .
[0214] Step 2-3: Initialize the discrete time variable, set the current time t = 0, and the actual time interval between two iterations is Δt, which is determined according to the actual measurement data and command control speed of the local controller;
[0215] Step 2-4: Initialize each controller's local cache experience database D i ,i∈A, taking the i-th region as an example, Initialize temporary variables
[0216] Step 3: Each local controller in the i-th region executes the following control steps at time t, in parallel without interfering with each other:
[0217] Step 3-1: Obtain measurement data from the measurement device of the regional power grid and update the corresponding observation variable o i ;
[0218] Among them, the measuring device includes voltage sensors and current sensors arranged at each node in the area. Through the current sensors and voltage sensors of each node, the current signal and voltage signal of each node can be obtained, and then the active power injection, reactive power injection and voltage amplitude of each node can be obtained.
[0219] Step 3-2: According to the local policy neural network and Update the action variable a corresponding to this iteration i :
[0220]
[0221] t represents the current time, ξ i represents the i-th random variable, N represents the normal distribution, I i Indicates k i -dimensional unit vector.
[0222] Step 3-3: The controller will a i Send it to the corresponding local DG and / or SVC;
[0223] Step 3-4: If TMP i If it is not an empty set, the observed variables at the current time (the time node of the current iteration) and the observed variables, action variables, and local feedback variables at the previous iteration are stored in D i Middle, i.e. D i ←D i ∪{o i ∪TMP i};
[0224] Step 3-5: According to o i Calculate R i , and calculate TMPi ={o i ,a i ,R i};
[0225] Step 4: Each regional local controller performs the following marginal variable exchange steps at time t, in parallel without interfering with each other:
[0226] Step 4-1: According to the current time, randomly select D i The sample is drawn from i ,a i ,R i ,o′ i}, and calculate the upward marginal variable:
[0227]
[0228] Among them, o′ i for o i The sample of a′ i and The calculation method is:
[0229]
[0230]
[0231] Step 4-2: Set the upward marginal variable u i Upload to the centralized server via the communication network.
[0232] Step 5: The centralized server performs the following learning steps (which may be performed in parallel with the local controller in the current iteration, for example, the centralized server also performs the following steps at time t):
[0233] Step 5-1: Collect the upstream marginal variables u uploaded by each local controller i ;
[0234] Step 5-2: According to u i Information calculation in:
[0235]
[0236] Among them, y represents the target value of the value function, γ represents the discount value of reinforcement learning (the typical value is 0.99), represents the neural network corresponding to the frozen parameters of the first region, (o′1, a′1) represents the input of the neural network corresponding to the frozen parameters of the first region, where o′1 is the sample observation variable of the first region, and a′1 is the sample action variable of the first region; represents the neural network corresponding to the frozen parameters of the nth region, (o′ n,a′ n ) represents the input of the neural network corresponding to the frozen parameters of the nth region, where o′ n is the sample observation variable of the nth region, a′ n is the sample action variable of the nth region;
[0237]
[0238] Among them, L φ Represents the parameters of the neural network The loss function is represents the network parameters of the first neural network c1 of the first region, The first neural network c represents the nth region n The network parameters, Indicates the expectation of o, a, r, o' on the data set; φ f Gradient Calculated by the following formula:
[0239]
[0240] Step 5-3: Update network parameters φ f :
[0241]
[0242] where η Q is a hyperparameter representing the value function learning rate. By way of example, its value is 0.001.
[0243] Step 5-4: Update the freeze parameters
[0244]
[0245] Wherein η is a hyperparameter representing the freezing coefficient, and illustratively, its value is 0.995.
[0246] Step 5-5: Form the downlink marginal variable d sent to each local controller i , where the partial derivatives can be directly obtained using a commercial neural network automatic differentiator, such as TensorFlow or PyTorch frameworks:
[0247]
[0248] Where o=[o1,…,o n ] T represents the observed variables of all regions, a=[a1,…,a n ] T represents the action variables of the entire region, represents the inferred action variable of the entire region, c i represents the neural network of the i-th region; Denotes the first neural network c of neural network f for region i i The partial derivatives of the values; Denotes the loss function L φ Partial derivatives of the neural network f.
[0249] Step 5-6: Substitute each downside marginal variable d i Send it to the corresponding local controller through the communication network;
[0250] Step 6: The local controller in each region i receives the d sent by the centralized server at time t i Then execute the following control steps in parallel without interfering with each other:
[0251] Step 6-1: Based on the downside marginal variable d i The following calculation is performed with local data:
[0252]
[0253]
[0254]
[0255]
[0256] in, represents the target value of the entropy estimation function of the ith region, π i represents the policy function of the ith region, represents the loss function of the neural network e, γ represents the reinforcement learning discount value, α represents the penalty coefficient, represents the gradient of the first neural network c in the i-th region, It means to find the expectation of all observed variables and random variables in actions.
[0257] Step 6-2: Update parameters With θ i :
[0258]
[0259]
[0260]
[0261] Among them, η θ represents the learning rate of the policy neural network, L θThe loss function representing the policy neural network parameters;
[0262] Step 6-3: Update the freeze parameters and
[0263]
[0264]
[0265] Step 7: t=t+1, in the next iteration, return to step 3 and repeat steps 3, 4, 5, and 6.
[0266] Based on the same inventive concept, an embodiment of the present invention also provides a controller for distributed control of reactive voltage, which is a local controller in a control architecture that can apply online federated learning and distributed control in parallel; the distribution network is divided into a number of controlled areas, each of which includes multiple nodes; the nodes of each controlled area are communicated with the local controller of the controlled area, and the collected measurement data are uploaded to the local controller of the controlled area; each controller is communicated with a centralized server, and the centralized server coordinates the learning process of the intelligent agents on each local controller; a multi-agent deep reinforcement learning algorithm is deployed on the local controller; the local controller exchanges marginal variables with the centralized server, and the local controller learns and obtains the control strategy of each local controller; each local controller performs reactive voltage control of the power grid based on the control strategy and local measurement data.
[0267] Based on the same inventive concept, an embodiment of the present invention also provides a server for reactive voltage distributed control, which is a centralized server in a control architecture that can apply online federated learning and distributed control in parallel. The centralized server is used to coordinate the learning process of the agents on each local controller; wherein the distribution network is divided into a number of controlled areas, each of which includes a plurality of nodes; the nodes of each controlled area are connected in communication with the local controller of the controlled area, and the collected measurement data are uploaded to the local controller of the controlled area; the centralized server is connected in communication with the local controller, and exchanges marginal variables with the local controller; the local controller deploys a multi-agent deep reinforcement learning algorithm, which can learn the control strategy of each local controller based on the exchanged marginal variables, and execute reactive voltage control of the power grid based on the control strategy and local measurement data.
[0268] The above centralized server and local controller can work together when applied to form a reactive voltage distributed control system. Figure 3As shown, the system includes a regional power grid control center cloud server (i.e., centralized server), a controlled regional power grid controller, i.e., a local controller (referred to as controller). The distribution network is divided into several controlled areas, each of which includes multiple nodes. A node is a unit connected to the distribution network that contains energy-consuming equipment or energy storage equipment.
[0269] A measuring device is provided at each node. The power grid is also provided with one of a static VAR compensator (SVC) and a renewable energy generation equipment (DG) at some or all of its nodes according to its own actual situation. Specifically, the controlled power grid node may be a node provided with only a measuring device, a node provided with a measuring device and a DG, or a node provided with a measuring device and an SVC. The measuring device is used to collect measurement data of the node. The measuring device includes: a voltage measuring device, a current measuring device, and a power measuring device. The current measuring device is used to measure the current of each node, the voltage measuring device is used to measure the voltage of each node, and the power measuring device is used to measure active power and reactive power. The measuring device may use a voltage sensor or a current sensor, but is not limited to this.
[0270] The node of each controlled area is connected to the controller of the controlled area and uploads the collected measurement data to the controller of the controlled area. The controller controls the actions of the SVC node and the DG node by sending control signals to them.
[0271] Each controller is connected to the centralized server in communication, such as through a communication network. The centralized server and the controller exchange marginal variables. Specifically, the centralized server sends downlink marginal variables (d1...dN) to the controller, and the controller uploads uplink marginal variables (u1...uN) to the centralized server.
[0272] Neural networks are deployed on both the centralized server and the controller. The neural network is used to implement a distributed reactive voltage optimization model, which is established based on the reactive voltage control target of each controlled area. Specifically, a computer system capable of deploying a neural network is provided on both the centralized server and the controller, including a processor and a memory. Specifically, a local learning module is deployed on the controller, and the neural network is deployed on the local learning module, and continuously learns in each iteration to generate a control strategy. The local learning module includes a collection of software and hardware capable of executing neural network deployment, initialization, and learning.
[0273] The centralized server and the controller can be used to execute the reactive power voltage distributed control method of the above-mentioned embodiment, and the specific execution method of each step will not be repeated here.
[0274] The control method of the embodiment of the present invention adopts a federated learning architecture, in which the centralized server only performs coordination, and each controller controls data and policy information only locally, which not only avoids a large amount of data exchange and saves communication costs, but also protects the privacy of each subject.
[0275] The control method of the embodiment of the present invention does not require the establishment of an accurate distribution network model in advance. It can perform model-free optimization and obtain the optimal control strategy of each intelligent agent through the accumulation and learning of control process data. It can optimize the reactive power distribution of the distribution network in the scenario of incomplete model, eliminate voltage over-limit and ensure efficient operation.
[0276] The control method of the embodiment of the present invention has each controller running completely in parallel, and the main amount of calculation is concentrated locally in each controller. As the number of connected controllers increases, the amount of calculation on the centralized server increases slowly. The newly connected controller can complete the local calculation, so the computing resources can be fully utilized, and it has good compatibility with the problem scale and computing conditions.
[0277] When executing control operations, each regional controller in the present invention does not need to communicate with a centralized server or other controllers, but can quickly generate control instructions based on the stored strategy neural network, efficiently utilize high-speed flexibility resources, and improve the efficiency of reactive voltage control.
[0278] Based on the same inventive concept, an embodiment of the present invention also provides a computer storage medium having computer executable instructions stored thereon. When executed by a computing device, the computer executable instructions enable the computing device to execute the above-mentioned reactive voltage distributed control method, including a reactive voltage distributed control method that can be run on a local controller or a reactive voltage distributed control method that can be run on a centralized server.
[0279] For example, Figure 4 As shown, the computing device includes: at least one processor, Figure 4 A processor is taken as an example; and a memory (memory), and may also include a communication interface (Communication Interface) and a bus. Among them, the processor, the communication interface, and the memory can communicate with each other through a data connection. The communication interface can be used for information transmission. The processor can call the logic instructions in the memory to execute the reactive voltage distributed control method of the above embodiment.
[0280] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0281] The memory, as a computer-readable storage medium, can be used to store software programs and computer executable programs, such as program instructions / modules corresponding to the method in the embodiment of the present invention. The processor executes functional applications and data processing by running the software programs, instructions and modules stored in the memory, that is, implementing the reactive voltage distributed control method in the above method embodiment.
[0282] The memory may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required for at least one function; the data storage area may store data created according to the use of the terminal device, etc. In addition, the memory may include a high-speed random access memory and may also include a non-volatile memory.
[0283] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A reactive voltage distributed control method, characterized in that: The control architecture adopts online federated learning and distributed control in parallel, including: Deploy multi-agent deep reinforcement learning algorithms on the local controller; Through the marginal variable exchange between the local controller and the centralized server, the local controller learns the control strategy of each local controller; Each local controller performs grid reactive power and voltage control based on control strategies and local measurement data; Wherein, the centralized server is used to coordinate the learning process of the agents on each local controller; Perform the following steps on each regional local controller: Framework construction steps: Construct a multi-agent interactive training framework based on Markov game; Network deployment steps: deploy the neural network; Local control steps: parallel execution of reactive power and voltage control; Local exchange step: executing marginal variable exchange with the centralized server in parallel, including obtaining downlink marginal variables from the centralized server and sending uplink marginal variables to the centralized server; Local learning step: Perform local learning in parallel based on the downstream marginal variables; The downlink marginal variables are obtained by the centralized server after centralized learning based on the collected uplink marginal variables; The local control step, the local exchange step and the local learning step are repeatedly performed in parallel; The upstream marginal variables include scalar data calculated by the local controller; The downstream marginal variable includes data for calculating gradient information; The centralized server is used to establish a distributed reactive power voltage optimization model, including determining the reactive power voltage control targets of each region: The distribution network is divided into |A| areas according to the actual installation of local controllers. The area label set is denoted by A. A local controller is set in each area to quickly obtain the measurement data of the area. The local controller can communicate with the centralized server. Among them, the reactive voltage control target of each area is to minimize the network loss of the whole network under the condition of satisfying the voltage constraint: Among them, L opt is the distributed reactive power voltage optimization target, P loss is the total network loss, Represent the active power injection and voltage amplitude of the jth node in the ith region, V, are the lower and upper limits of the voltage amplitude respectively; q S,i represents the reactive power output of all static VAR compensators SVC in the i-th region, q S,1 represents the reactive power output of all SVCs in the first zone, q S,|A| represents the reactive power output of all SVCs in the |A|th region, in represents the reactive power output of the jth SVC in the ith zone, They represent the lower and upper limits of the reactive power output of the jth SVC in the i-th region respectively; q G,i represents the reactive power output of all DGs in the ith area, q G,1 represents the reactive power output of all DGs in the first area, q G,|A| represents the reactive power output of all DGs in the |A|th area, where represents the reactive power output of the jth DG in the ith region, is the installed capacity and active power output of the jth DG in the ith region, N i is the set of nodes in the ith region, is the set of DG of the ith region, is the set of SVCs in the i-th region.
2. The reactive voltage distributed control method according to claim 1, characterized in that: Construct a multi-agent interactive training framework based on Markov game, including: Corresponding to the measurement data of each regional system, construct the observation variables o of each region i As shown in (1.2): represents the reactive power injection of the jth node in the i-th region; Corresponding to the reactive power and voltage optimization objectives of each region, the unified feedback variable r of each region is constructed as shown in (1.3): Among them, C V is the penalty coefficient of voltage constraint; function [] + is defined as [x] + =max(0,x); Decomposing the unified feedback variable by region, we can obtain: Among them, R i is the decomposed variable of r in the ith region, i.e., the local feedback variable of the ith region, which can be calculated through the local measurement data of the ith region; Corresponding to the reactive power of controllable flexibility resources, construct the action variables a of each area i As shown in (1.5): in, represents the reactive power output of the kth SVC in the ith zone, represents the reactive power output of the jth DG in the ith zone, and the controllable flexibility resources include SVC and DG.
3. The reactive voltage distributed control method according to claim 2, characterized in that: Deploy neural networks on local controllers in each region separately, including: Deploy the neural network on the local controller of each region and initialize it. For any i-th region, the following steps are included: Define the first neural network c of the i-th region i For an input (o i ,a i ), a fully connected neural network that outputs a single scalar value, denoted by c i The network parameters are The corresponding freezing parameters are The neural network corresponding to the frozen parameters is recorded as Random Initialization and The first neural network is used to estimate the expected reward; Define the second neural network e of the i-th region i For an input (o i ,a i ) is a fully connected neural network that outputs a single scalar value, denoted by e i The network parameters are The corresponding freezing parameters are The neural network corresponding to the frozen parameters is recorded as Random Initialization and The second neural network is used to evaluate entropy; Define the third neural network And the fourth neural network in, and All are input o i Output k i dimensional vector policy neural network, where k i is the action vector a i Dimensions; and They each have an independent output layer, while sharing the same neural network input layer and hidden layer, including several hidden layers; and All network parameters are θ i .
4. The reactive power voltage distributed control method according to claim 3, characterized in that: Also includes: Initialize the discrete time variables, set the current time t = 0, and the actual time interval between two iterations is Δt, which is determined according to the actual measurement data and command control speed of the local controller; Initialize each local controller's local cache experience database D i ,i∈A, for any i-th region, Initialize temporary variables 5. The reactive voltage distributed control method according to claim 4, characterized in that: Each regional local controller executes the marginal variable exchange steps in parallel, including: Obtain measurement data from the measurement device of the regional power grid and update the corresponding observation variable o i ; According to the local policy neural network and Update the action variable a corresponding to this iteration i : ξ i represents the i-th random variable, N represents the normal distribution, I i Indicates k i -dimensional unit vector; The local controller will a i Send it to the corresponding local DG and / or SVC; Each regional local controller performs the following marginal variable exchange steps at time t in parallel: According to the current time, randomly in D i The sample is drawn from i ,a i ,R i ,o′ i }, and calculate the upward marginal variable: Among them, o′ i for o i The sample of a′ i and The calculation method is: The upward marginal variable u i Upload to the centralized server via the communication network.
6. A reactive voltage distributed control method, characterized in that: The control architecture adopts online federated learning and distributed control in parallel, including: The learning process of the agents on each local controller is coordinated by a centralized server; A multi-agent deep reinforcement learning algorithm is deployed on the local controller; The centralized server exchanges marginal variables with the local controllers so that the local controllers learn the control strategies of the local controllers and perform reactive power voltage control of the power grid based on the control strategies and local measurement data; Methods include: Model determination steps: Establish a distributed reactive power voltage optimization model; Network deployment steps: deploy the neural network on a centralized server; Centralized learning steps: The centralized server collects the upstream marginal variables of each region, performs centralized learning, and sends the downstream marginal variables to each local controller; The downstream marginal variables are used by each local controller to perform local learning in parallel to obtain a control strategy; Among them, the centralized learning step is repeatedly performed; The upstream marginal variables include scalar data calculated by the local controller; The downstream marginal variable includes data for calculating gradient information; Among them, establishing a distributed reactive power voltage optimization model includes determining the reactive power voltage control target of each area: The distribution network is divided into |A| areas according to the actual installation of local controllers. The area label set is denoted by A. A local controller is set in each area to quickly obtain the measurement data of the area. The local controller can communicate with the centralized server. Among them, the reactive voltage control target of each area is to minimize the network loss of the whole network under the condition of satisfying the voltage constraint: Among them, L opt is the distributed reactive power voltage optimization target, P loss is the total network loss, represent the active power injection and voltage amplitude of the jth node in the ith region, respectively. V , are the lower and upper limits of the voltage amplitude respectively; q S,i represents the reactive power output of all static VAR compensators SVC in the i-th region, q S,1 represents the reactive power output of all SVCs in the first zone, q S,|A| represents the reactive power output of all SVCs in the |A|th region, in represents the reactive power output of the jth SVC in the ith zone, They represent the lower and upper limits of the reactive power output of the jth SVC in the i-th region respectively; q G,i represents the reactive power output of all DGs in the ith area, q G,1 represents the reactive power output of all DGs in the first area, q G,|A| represents the reactive power output of all DGs in the |A|th area, where represents the reactive power output of the jth DG in the ith region, is the installed capacity and active power output of the jth DG in the ith region, N i is the set of nodes in the ith region, is the set of DG of the ith region, is the set of SVCs in the i-th region.
7. The reactive power voltage distributed control method according to claim 6, characterized in that: The local controller has a multi-agent interactive training framework based on Markov game formed in the following manner: Corresponding to the measurement data of each regional system, construct the observation variables o of each region i As shown in (1.2): represents the reactive power injection of the jth node in the i-th region; Corresponding to the reactive power and voltage optimization objectives of each region, the unified feedback variable r of each region is constructed as shown in (1.3): Among them, C V is the penalty coefficient of voltage constraint; function [] + is defined as [x] + =max(0,x); Decomposing the unified feedback variable by region, we can obtain: Among them, R i is the decomposed variable of r in the ith region, i.e., the local feedback variable of the ith region, which can be calculated through the local measurement data of the ith region; Corresponding to the reactive power of controllable flexibility resources, construct the action variables a of each area i As shown in (1.5): in, represents the reactive power output of the kth SVC in the ith zone, represents the reactive power output of the jth DG in the ith zone, and the controllable flexibility resources include SVC and DG.
8. The reactive power voltage distributed control method according to claim 7, characterized in that: Deploy the neural network on a centralized server and initialize it; Define the neural network f as a fully connected neural network that inputs |A| scalars and outputs 1 scalar value; let the network parameter of f be φ f , the corresponding freezing parameter is The neural network corresponding to the frozen parameters is recorded as Randomly initialize φ f and 9. The reactive power voltage distributed control method according to claim 8, characterized in that: The centralized server collects the uplink marginal variables of each region and performs the learning steps and sends the downlink marginal quantities to each local controller, including: Collect the upstream marginal variables u uploaded by each local controller i ; According to u i Information calculation in: Among them, y represents the target value of the value function, γ represents the discount value of reinforcement learning, represents the neural network corresponding to the frozen parameters of the first region, (o′1, a′1) represents the input of the neural network corresponding to the frozen parameters of the first region, where o′1 is the sample observation variable of the first region, and a′1 is the sample action variable of the first region; represents the neural network corresponding to the frozen parameters of the nth region, (o′ n ,a′ n ) represents the input of the neural network corresponding to the frozen parameters of the nth region, where o′ n is the sample observation variable of the nth region, a′ n is the sample action variable of the nth region; Among them, L φ represents the loss function of the neural network parameter φ, represents the network parameters of the first neural network c1 of the first region, The first neural network c represents the nth region n The network parameters, Indicates the expectation of o, a, r, o' on the data set; φ f Gradient Calculated by the following formula: Update network parameters φ f : where η Q is a hyperparameter representing the learning rate of the value function; Step 5-4: Update the freeze parameters Where η is a hyperparameter, representing the freezing coefficient; Form the downlink marginal variable d sent to each local controller i : Where o=[o1,…,o n ] T represents the observed variables of all regions, a=[a1,…,a n ] T represents the action variables of the entire region, represents the inferred action variable of the entire region, c i represents the neural network of the i-th region; Denotes the first neural network c of neural network f for region i i The partial derivatives of the values; Denotes the loss function L φ Partial derivatives of the neural network f; Each downside marginal variable d i It is sent to the corresponding local controller through the communication network.
10. A computer storage medium having computer executable instructions stored thereon, which, when executed by a computing device, cause the computing device to perform the method of any one of claims 1-5.
11. A computer storage medium having computer executable instructions stored thereon, which, when executed by a computing device, cause the computing device to perform the method of any one of claims 6-9.
Citation Information
Patent Citations
Power grid reactive voltage distributed control method and system
CN111799808A
Distributed photovoltaic power station system and reactive self-balancing control method
CN112653178A