A self-organizing group decision-making method and system for regional power grids based on multi-objective phased learning

Through the multi-objective phased learning of regional power grid self-organized group decision-making method, the self-organized group action volume is used to perform grid interaction and neural network parameters are optimized, and the model complexity and power fluctuation problems of distributed resources are solved, and the intelligent interaction of flexible resources and intelligent decision-making of power grids are realized.

CN119742868BActive Publication Date: 2025-09-02CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411840889.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-09-02
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

The existing centralized optimization of provincial power grids is difficult to cope with the model complexity and power fluctuation of distributed flexible resources. Especially in regional power grids where the proportion of distributed new energy increases, power balance and regulation of resources have become challenges.

Method used

The regional power grid self-organized group decision-making method is adopted with multi-objective staged learning. By establishing a group strategy network, the staged neural network parameter optimization is carried out to realize the independent intelligent decision-making of resource groups, and the power grid interaction and trend calculation is used to optimize neural network parameters to improve the level of intelligent interaction.

Benefits of technology

It realizes intelligent interaction of massive flexible resources, reduces model complexity and power fluctuation, improves the learning efficiency and convergence of neural networks, and is suitable for intelligent decision-making of regional power grids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119742868B_ABST
    Figure CN119742868B_ABST
Patent Text Reader

Abstract

A multi-objective, phased learning, regional power grid self-organizing group decision-making method and system includes establishing several group strategy networks using regional power grid resources. The group strategy networks perform forward calculations based on their own observations, outputting group action quantities under the current environmental state. Group power is output based on the group action quantities, and the power injected by nodes connected to the group is accounted for in the group power. Single-round reward values ​​are calculated based on the environmental state and actions before and after the decision, and the single-round decision experience is stored in an experience pool. Neural network parameters are optimized, the loss function value of the value network is calculated, and convergence is determined based on the change in the loss function value until all objective function items are accounted for and the loss function converges. The group strategy network performs forward calculations and outputs the group action quantity under the current environmental state. Group resources perform operations based on the group action quantity, causing flexible resource power to be output to grid-connected nodes. This invention enables autonomous and intelligent decision-making by resource groups.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of source-grid-load-storage coordinated dispatching, and specifically relates to a regional power grid self-organizing group decision-making method and system with multi-objective phased learning. Background Art

[0002] With the orderly development of the construction of new power systems, distributed flexible resources are rich in variety, high in penetration, and large in adjustable capacity. The granularity of adjustable loads has been reduced from megawatts to kilowatts. Taking a certain regional power grid as an example, the source-grid-load-storage control system is connected to more than 100,000 diverse micro-loads, including a university, a company's headquarters base, a heat pump energy station, and a photovoltaic storage and charging building. In addition, the types of flexible resources are diverse, including energy storage, non-industrial air conditioning, commercial buildings, residential loads, electric vehicles, electric boilers, industrial parks, microgrids, etc., which poses a huge challenge to the power balance of the regional power grid taking into account new energy, adjustable loads and other resources.

[0003] As the proportion of distributed renewable energy in regional power grids increases, regional power grids are transitioning from one-way energy distribution to a two-way balance of power generation and consumption. Existing centralized optimization of provincial power grids struggles to cope with the model complexity and power fluctuations of large groups. Summary of the Invention

[0004] The purpose of the present invention is to address the problems in the above-mentioned prior art and provide a multi-objective phased learning regional power grid self-organizing group decision-making method and system, which takes the regional power grid resource group as the scheduling object, adopts phased neural network parameter optimization, realizes autonomous intelligent decision-making of the resource group, and effectively improves the level of intelligent interaction of massive flexible resources.

[0005] In order to achieve the above object, the present invention has the following technical solutions:

[0006] First, a multi-objective phased learning method for regional power grid self-organization group decision-making is provided, including:

[0007] Several group strategy networks are established through regional power grid resources. The i-th group strategy network performs forward calculation based on its own observations and outputs the current environment state S t The group action amount a under t,i , according to the group action amount a t,i Output group power p i,t , the power injected by the node connected to the group is included in the group power p i,t , using grid interactive simulation to calculate power flow, output the active power p of the power system node bus , reactive power q bus and voltage v bus ; Based on the environmental state S before and after the decision t and S t+1 and action at,i Calculate the single-round reward value r t , the single-round decision experience (S t ,r t ,a t,i ,S t+1 ) is stored in the experience pool;

[0008] Optimize the neural network parameters, extract experience data from the experience pool, and set the network value according to (S t ,r t ,S t+1 )Calculate the network gradient; According to (S t ,a t,i ) Calculate the full-cycle benefit valuation Q of the group action in the current environment through the value network t , based on the full cycle return valuation Q t and group action amount a t,i Calculate the gradient of the swarm strategy network; calculate the loss function value of the value network, and determine whether it has converged based on the change in the loss function value. If it has converged, add a new stage objective function item, otherwise continue to update the value network and the swarm strategy network; until all objective function items are taken into account and the loss function converges;

[0009] The swarm strategy network performs forward calculations based on the power grid measurement information and user-side measurement information, and outputs the current environmental state S t The group action amount a under t,i , the regional power grid group resources are based on the group action amount a t,i Execute operations to enable flexible resource power to be exported to regional power grid connection nodes.

[0010] As a preferred solution, the step of establishing a plurality of group strategy networks using regional power grid resources includes:

[0011] Self-organizing group observation space design:

[0012]

[0013] Self-organizing group action space design:

[0014] a t,i =(ΔT ac ,ΔP disc ,ΔP charg ,ΔP pv ,ΔP wd )

[0015] Design of multi-objective reward function for self-organizing groups:

[0016] r=ω blc r blc +ω cst rcst +ω line r line +ω v r v +ω ne r ne +ω err r err

[0017] Where, obs t,i is the observation space of the ith agent in time period t, is the mean room temperature within the group, T env is the ambient temperature, P se is the energy storage power, P pv is the photovoltaic power, P wd is the wind power, P bus Inject power into the node, V bus is the node voltage; a t,i is the agent action space, ΔT ac Temperature adjustment, ΔP disc is the discharge power adjustment amount, ΔP charg is the charging power adjustment amount, ΔP pv is the photovoltaic power regulation, ΔP wd is the wind power regulation value.

[0018] As a preferred solution, the multi-objective phased learning regional power grid self-organizing group decision-making method further includes multi-agent reinforcement learning of several group strategy networks established through regional power grid resources, including:

[0019] Multi-agent strategy network design;

[0020] Multi-agent target strategy network design;

[0021] Multi-agent global value network design;

[0022] Multi-agent global objective value network design;

[0023] Update the multi-agent global value network and the target global value network;

[0024] Perform multi-agent policy network and target policy network updates;

[0025] Add stage goals and end training after stage-by-stage training converges.

[0026] As a preferred solution, the steps of designing the multi-agent strategy network include:

[0027] Input layer dimension calculation:

[0028] NP in,i =N obs,i

[0029] Output layer dimension calculation:

[0030] NP out,i =N a,i

[0031] Hidden layer dimension calculation:

[0032]

[0033] Where NP in,i is the dimension of the policy network input layer, N obs,i is the observation vector dimension, NP out,i is the dimension of the policy network output layer, N a,i is the action vector dimension, NP h,i is the hidden layer dimension of the policy network, N s is the number of training samples, α is the diversity adjustment parameter;

[0034] NP in,i ≤NP h,i ≤NP out,i

[0035] NP h,i ≤2*NP in,i

[0036] NP obtained based on the hidden layer dimension calculation formula h,i , bring in the above two dimensional conditions determined by the input and output layer dimensions, when NP h,i If the conditions are met, the original value can be used; otherwise, the boundary value of the above conditions can be used;

[0037] Set the number of hidden layers to be greater than 2. By comparing the results, increase the number of layers for underfitting and reduce the number of layers for overfitting. Overfitting is judged based on the fact that the reward obtained by decision-making using the training set data reaches the upper limit, but the reward obtained by decision-making using the test set data reaches the lower limit. Underfitting is judged based on the fact that the reward values ​​obtained by decision-making on both the training set and test set data reach the lower limit.

[0038] As a preferred solution, the steps of designing the multi-agent global value network include:

[0039] Input layer dimension calculation:

[0040] NQ in,i =N obs,i +N a,i

[0041] Output layer dimension calculation:

[0042] NQ out,i =1

[0043] Hidden layer dimension calculation:

[0044]

[0045] NQ in,i ≤NQ h,i ≤NQ out,i

[0046] NQ h,i ≤2*NQ in,i

[0047] Where NQ in,i Input dimension for the value network, NQ out,i is the output dimension of the value network, NQ h,i is the hidden layer dimension of the network.

[0048] As a preferred solution, the step of updating the multi-agent global value network and the target global value network includes:

[0049] Target global value network θ tar Initialize with a copy of the global value network θ;

[0050] The global value network θ is in the current target global value network θ tar Add random disturbance on the basis of

[0051] Extract multiple experiences (S t ,r t ,a t,i ,S t+1 ) to update the target global value network, including:

[0052] Calculate the action in period t+1:

[0053]

[0054] Where a i,t+1 is the strategy network at time t+1 The decision result of π(·) represents the forward calculation result of the policy network;

[0055] Calculate the target value network estimation:

[0056] V t+1 =V θ (S t+1 ,a t+1 )

[0057] Among them, V t+1 is the target global value network θ at time t+1 tar The valuation result of V θ (·) represents the forward calculation result of the global value network θ;

[0058] Profit calculation:

[0059] V t =r t +γV t+1

[0060] Where V t is the total global revenue of multi-agent decision-making under the action and state at time t calculated according to the Bellman equation, r t is the reward function value at time t, and γ is the discount coefficient of future benefits;

[0061] Action estimation calculation:

[0062] Q t =Q θ (a i,t ,S t )

[0063] Where Q t is the estimated value of the global value network θ for the global benefit of multi-agent decision-making at the action and state at time t, Q θ (·) is the forward calculation result of the global value network θ;

[0064] Loss function calculation:

[0065]

[0066] Among them, J Q (θ) represents the loss function of the global value network θ;

[0067] Gradient calculation:

[0068]

[0069] Where, Represents the target global value network θ tar The gradient, Represents the gradient of all parameters of the global value network θ to the valuation, Calculated using the back-propagation method;

[0070] Target value network update:

[0071]

[0072] Among them, λ Q is the target global value network θ tar The learning rate.

[0073] As a preferred solution, the steps of updating the multi-agent policy network and the target policy network include:

[0074] Calculate action estimates:

[0075] Take the global value network θ to estimate the global benefit Q of the multi-agent decision at the action and state at time t t ;

[0076] Calculate the gradient of the action with respect to the loss function:

[0077] The policy network loss function takes the estimate of the global value network θ as:

[0078]

[0079] Where, is the forward calculation result of the global value network;

[0080] Therefore, the gradient of the action with respect to the valuation is expressed as:

[0081]

[0082] Where, is the gradient of the action of the i-th agent with respect to the global value network valuation;

[0083] Calculate the gradient of the policy network parameters with respect to the action:

[0084]

[0085] Where, is the gradient of the policy network parameters of the i-th agent in time period t with respect to the action;

[0086] The gradient of the policy network parameters to the action is calculated by the policy network back propagation method;

[0087] Calculate the gradient of the policy network parameters with respect to the estimate:

[0088]

[0089] Where, The gradient of the policy network of the i-th agent to the valuation is equal to the product of the gradient of the policy network parameters to the action and the gradient of the action to the valuation network in the same period of the corresponding agent;

[0090] Targeted Policy Network Updates:

[0091]

[0092] Where, φ tar is the target policy network parameter, and λ is the target policy network learning rate.

[0093] As a preferred solution, the step of adding a stage target and ending the training after the staged training converges includes:

[0094] Determine whether the loss function converges:

[0095]

[0096] Calculate whether the change in the loss function satisfies the convergence condition ε π ;

[0097] If not satisfied, continue to perform the policy network and global value network θ and the corresponding target global value network θ under the current reward function tar Parameter update;

[0098] Weighting coefficients and stage mapping:

[0099] Phase 1: Adding Balance Rewards blc r blc and cost reward ω cst r cst , so that the self-organizing group regulation can meet the regional power grid power balance constraints under the minimum cost target;

[0100] Phase 2: Adding line voltage bonus ω line r line and node voltage reward ω v r v , so that the regulation target of the self-organizing group meets the safety constraints of power grid operation;

[0101] Phase 3: Adding new energy consumption incentives ne r ne and group response bias reward ω err r err , making the self-organized group adjustment target conducive to the consumption of new energy and the completion of group resources;

[0102] The objective function is adjusted in stages. The weighted coefficient of the objective function is initialized to 0. After each stage of training reaches the convergence condition, the weighted coefficient is modified to 1 in sequence until the training ends after the last stage of training converges.

[0103] As a preferred solution, the swarm strategy network performs forward calculation based on the power grid measurement information and the user side measurement information, and outputs the current environment state S t The group action amount a under t,i In the step, the self-organizing group i is based on the electricity quantity and resource measurement statistics obs observed by the group. t,i Perform forward calculation of the strategy network and adjust the group action amount a t,i :

[0104]

[0105] Where, Represents the forward calculation result of the policy network;

[0106] The group resources of the regional power grid are based on the group action amount a t,i In the step of executing the operation to output the power of the flexible resources to the regional power grid connection node, the internal resources of the regional power grid group perform charging and discharging, temperature regulation and distributed new energy power control according to the published group action strategy.

[0107] Secondly, a multi-objective phased learning regional power grid self-organizing group decision-making system is provided, including:

[0108] The decision-making process establishment module is used to establish several group strategy networks through regional power grid resources. The i-th group strategy network performs forward calculation based on its own observations and outputs the current environment state S t The group action amount a under t,i , according to the group action amount a t,i Output group power p i,t , the power injected by the node connected to the group is included in the group power p i,t , using grid interactive simulation to calculate power flow, output the active power p of the power system node bus , reactive power q bus and voltage v bus ; Based on the environmental state S before and after the decision t and S t+1 and action a t,i Calculate the single-round reward value r t , the single-round decision experience (S t ,r t ,a t,i ,S t+1 ) is stored in the experience pool;

[0109] Neural network parameter optimization module is used to optimize neural network parameters, extract experience data from the experience pool, and calculate the network parameters according to (S t ,r t ,S t+1 )Calculate the network gradient; According to (S t ,a t,i ) Calculate the full-cycle benefit valuation Q of the group action in the current environment through the value network t , based on the full cycle return valuation Q t and group action amount a t,i Calculate the gradient of the swarm strategy network; calculate the loss function value of the value network, and determine whether it has converged based on the change in the loss function value. If it has converged, add a new stage objective function item, otherwise continue to update the value network and the swarm strategy network; until all objective function items are taken into account and the loss function converges;

[0110] The decision output module is used by the group strategy network to perform forward calculations based on the power grid measurement information and user-side measurement information, and output the current environment state S t The group action amount a under t,i , the regional power grid group resources are based on the group action amount a t,i Execute operations to enable flexible resource power to be exported to regional power grid connection nodes.

[0111] As a preferred solution, the steps of the decision-making process establishment module establishing a plurality of group strategy networks through regional power grid resources include:

[0112] Self-organizing group observation space design:

[0113]

[0114] Self-organizing group action space design:

[0115] a t,i =(ΔT ac ,ΔP disc ,ΔP charg ,ΔP pv ,ΔP wd )

[0116] Design of multi-objective reward function for self-organizing groups:

[0117] r=ω blc r blc +ω cst r cst +ω line r line +ω v r v +ω ne r ne +ω err r err

[0118] Where, obs t,i is the observation space of the ith agent in time period t, is the mean room temperature within the group, T env is the ambient temperature, P se is the energy storage power, P pv is the photovoltaic power, P wd is the wind power, P bus Inject power into the node, V bus is the node voltage; a t,i is the agent action space, ΔT ac Temperature adjustment, ΔP disc is the discharge power adjustment amount, ΔP charg is the charging power adjustment amount, ΔP pvis the photovoltaic power regulation, ΔP wd is the wind power regulation;

[0119] Multi-agent reinforcement learning is performed on several swarm strategy networks established through regional power grid resources, including:

[0120] Multi-agent strategy network design;

[0121] Multi-agent target strategy network design;

[0122] Multi-agent global value network design;

[0123] Multi-agent global objective value network design;

[0124] Update the multi-agent global value network and the target global value network;

[0125] Perform multi-agent policy network and target policy network updates;

[0126] Add stage goals and end training after stage-by-stage training converges.

[0127] As a preferred solution, the decision output module uses the self-organizing group i to calculate the electricity quantity and resource measurement statistical information obs observed by the group. t,i Perform forward calculation of the strategy network and adjust the group action amount a t,i :

[0128]

[0129] Where, Represents the forward calculation result of the policy network;

[0130] The internal resources of the regional power grid group perform charging and discharging, temperature regulation, and distributed new energy power control based on the published group action strategy.

[0131] In a third aspect, an electronic device is provided, comprising a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the multi-objective phased learning regional power grid self-organizing group decision-making method.

[0132] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the regional power grid self-organizing group decision-making method of multi-objective staged learning is implemented.

[0133] Compared with the prior art, the first aspect of the present invention has at least the following beneficial effects:

[0134] A self-organizing group multi-agent decision-making intelligent agent training framework has been constructed to achieve autonomous intelligent decision-making of resource groups and realize intelligent interaction between sources, networks, loads and storage. A self-organizing group is a massive resource aggregate. The aggregate is formed through spontaneous interaction, and the group allows users to join and exit when access conditions are met. Massive device-level resources are aggregated layer by layer to form an adjustable resource group. The present invention realizes the coordinated interaction of sources, networks, loads and storage through a centralized decision-making-distributed autonomous hybrid mode, reducing the model complexity and power volatility. The present invention is suitable for intelligent decision-making of massive flexible resource groups in regional power grids, including but not limited to interactive regulation of sources, networks, loads and storage in regional power grid zones, etc., which can effectively improve the level of intelligent interaction of massive flexible resources.

[0135] Furthermore, an embodiment of the present invention performs multi-agent reinforcement learning on several group strategy networks established through regional power grid resources. The proposed training method adopts a staged neural network parameter optimization technology to improve the security of intelligent decision-making, as well as the learning efficiency and convergence of the neural network.

[0136] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0137] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0138] Figure 1 The training phase architecture diagram of the regional power grid self-organizing group decision-making method with multi-objective phased learning of the present invention;

[0139] Figure 2 Schematic diagram of the intelligent decision-making process of the regional power grid self-organizing group decision-making method with multi-objective phased learning according to the present invention;

[0140] Figure 3 The multi-agent global value network and target global value network update flow chart of the present invention;

[0141] Figure 4 The multi-agent strategy network and target strategy network update flow chart of the present invention. DETAILED DESCRIPTION

[0142] In the following description, specific details such as specific system structures and techniques are provided for purposes of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present application with unnecessary detail.

[0143] See also Figure 1 and Figure 2 The embodiment of the present invention proposes a multi-objective phased learning regional power grid self-organizing group decision-making method, which is mainly divided into two parts: an intelligent decision-making process and a neural network parameter optimization process, specifically including:

[0144] S1. Establish several group strategy networks through regional power grid resources. The i-th group strategy network performs forward calculation based on its own observations and outputs the current environment state S t The group action amount a under t,i , according to the group action amount a t,i Output group power p i,t , the power injected by the node connected to the group is included in the group power p i,t , using grid interactive simulation to calculate power flow, output the active power p of the power system node bus , reactive power q bus and voltage v bus ; Based on the environmental state S before and after the decision t and S t+1 and action a t,i Calculate the single-round reward value r t , the single-round decision experience (S t ,r t ,a t,i ,S t+1 ) is stored in the experience pool;

[0145] S2, optimize the neural network parameters, extract experience data from the experience pool, and value the network according to (S t ,r t ,S t+1 ) calculates the network gradient by the least mean square root; according to (S t ,a t,i ) Calculate the full-cycle benefit valuation Q of the group action in the current environment through the value network t , based on the full cycle return valuation Q t and group action amount a t,iCalculate the gradient of the swarm strategy network; calculate the loss function value of the value network, and determine whether it has converged based on the change in the loss function value. If it has converged, add a new stage objective function item, otherwise continue to update the value network and the swarm strategy network; until all objective function items are taken into account and the loss function converges;

[0146] S3, the group strategy network performs forward calculation based on the power grid measurement information and user-side measurement information, and outputs the current environment state S t The group action amount a under t,i , the regional power grid group resources are based on the group action amount a t,i Execute operations to enable flexible resource power to be exported to regional power grid connection nodes.

[0147] Forward computation is a fundamental process in neural networks, passing input data through each layer of the network until an output is obtained. Aggregating massive amounts of device-level resources layer by layer to form a scalable resource group, and enabling the coordinated interaction of sources, grids, loads, and storage through a hybrid model of centralized decision-making and distributed autonomy has become a key task in the construction of new power systems.

[0148] In one possible implementation, the steps of establishing a plurality of group strategy networks using regional power grid resources include:

[0149] Self-organizing group observation space design:

[0150]

[0151] Self-organizing group action space design:

[0152] a t,i =(ΔT ac ,ΔP disc ,ΔP charg ,ΔP pv ,ΔP wd )

[0153] Design of multi-objective reward function for self-organizing groups:

[0154] r=ω blc r blc +ω cst r cst +ω line r line +ω v r v +ω ne r ne +ω err r err

[0155] Where, obs t,i is the observation space of the ith agent in time period t, is the mean room temperature within the group, T env is the ambient temperature, P se is the energy storage power, P pv is the photovoltaic power, P wd is the wind power, P bus Inject power into the node, V bus is the node voltage; a t,i is the agent action space, ΔT ac Temperature adjustment, ΔP disc is the discharge power adjustment amount, ΔP charg is the charging power adjustment amount, ΔP pv is the photovoltaic power regulation, ΔP wd is the wind power regulation value.

[0156] In one possible implementation, multi-agent reinforcement learning is performed on several swarm strategy networks established using regional power grid resources, including:

[0157] Multi-agent strategy network design;

[0158] Multi-agent target strategy network design;

[0159] Multi-agent global value network design;

[0160] Multi-agent global objective value network design;

[0161] Update the multi-agent global value network and the target global value network;

[0162] Perform multi-agent policy network and target policy network updates;

[0163] Add stage goals and end training after stage-by-stage training converges.

[0164] In one possible implementation, the steps of designing the multi-agent strategy network include:

[0165] Input layer dimension calculation:

[0166] NP in,i =N obs,i

[0167] Output layer dimension calculation:

[0168] NP out,i =N a,i

[0169] Hidden layer dimension calculation:

[0170]

[0171] Where NP in,iis the dimension of the policy network input layer, N obs,i is the observation vector dimension, NP out,i is the dimension of the policy network output layer, N a,i is the action vector dimension, NP h,i is the hidden layer dimension of the policy network, N s is the number of training samples, and α is the diversity adjustment parameter. The number of samples determines the model complexity of the agent, but the diversity of the samples affects the actual information content of the samples. To address this, a diversity adjustment parameter α is added, generally ranging from 2 to 10. In general, the more training samples there are, the higher the diversity, and the more neurons that can support the neural network.

[0172] NP in,i ≤NP h,i ≤NP out,i

[0173] NP h,i ≤2*NP in,i

[0174] The more complex the decision neural network is, the better it is. An overly complex neural network will make it difficult to converge in training or cause overfitting. h,i , bring in the above two dimensional conditions determined by the input and output layer dimensions, when NP h,i If the conditions are met, the original value can be used; otherwise, the boundary value of the above conditions can be used.

[0175] The decision-making problem involves complex mapping relationships and a large number of decision boundaries. The number of hidden layers is set to be greater than 2. By comparing the results, the number of layers is increased for underfitting and decreased for overfitting. Among them, the basis for judging overfitting is that the reward for making decisions using the training set data is extremely high, but the reward for making decisions using the test set is low. The basis for judging underfitting is that the reward values ​​for making decisions on both the training set and the test set data are low.

[0176] In one possible implementation, the multi-agent target policy network design is the same as the multi-agent policy network design.

[0177] In one possible implementation, the steps of designing a multi-agent global value network include:

[0178] Input layer dimension calculation:

[0179] NQ in,i =N obs,i +N a,i

[0180] Output layer dimension calculation:

[0181] NQ out,i =1

[0182] Hidden layer dimension calculation:

[0183]

[0184] NQ in,i ≤NQ h,i ≤NQ out,i

[0185] NQ h,i ≤2*NQ in,i

[0186] Where NQ in,i Input dimension for the value network, NQ out,i is the output dimension of the value network, NQ h,i is the hidden layer dimension of the network.

[0187] In one possible implementation, the multi-agent global objective value network design is the same as the multi-agent global value network design.

[0188] See also Figure 3 In one possible implementation, the step of updating the multi-agent global value network and the target global value network uses one global value network (referred to as value network) and one target global value network (referred to as target value network).

[0189] The specific steps include:

[0190] Target global value network θ tar Initialize with a copy of the global value network θ;

[0191] The global value network θ is in the current target global value network θ tar Add random disturbance on the basis of

[0192] Extract multiple experiences (S t ,r t ,a t,i ,S t+1 ) to update the target global value network, including:

[0193] Calculate the action in period t+1:

[0194]

[0195] Where a i,t+1 is the strategy network at time t+1 The decision result of π(·) represents the forward calculation result of the policy network;

[0196] Calculate the target value network estimation:

[0197] V t+1 =Vθ (S t+1 ,a t+1 )

[0198] Among them, V t+1 is the target global value network θ at time t+1 tar The valuation result of V θ (·) represents the forward calculation result of the global value network θ;

[0199] Profit calculation:

[0200] V t =r t γV t+1

[0201] Where V t is the total global revenue of multi-agent decision-making under the action and state at time t calculated according to the Bellman equation, r t is the reward function value at time t, and γ is the discount coefficient of future benefits;

[0202] Action estimation calculation:

[0203] Q t =Q θ (a i,t ,S t )

[0204] Where Q t is the estimated value of the global value network θ for the global benefit of multi-agent decision-making at the action and state at time t, Q θ (·) is the forward calculation result of the global value network θ;

[0205] Loss function calculation:

[0206]

[0207] Among them, J Q (θ) represents the loss function of the global value network θ;

[0208] Gradient calculation:

[0209]

[0210] Where, Represents the target global value network θ tar The gradient, Represents the gradient of all parameters of the global value network θ to the valuation, It is calculated using the back-propagation method; back-propagation is a common method for training artificial neural networks. It is an algorithm used in conjunction with optimization methods such as gradient descent. The back-propagation algorithm uses the chain rule to calculate the gradient from the output layer of the network backward along the connections of the network layer by layer.

[0211] Target value network update:

[0212]

[0213] Among them, λ Q is the target global value network θ tar The learning rate.

[0214] See also Figure 4 In one possible implementation, the step of updating the multi-agent policy network and the target policy network includes:

[0215] Calculate action estimates:

[0216] Take the estimated value Q of the global benefit of the global value network θ for the multi-agent decision-making at time t and state in the previous step t ;

[0217] Calculate the gradient of the action with respect to the loss function:

[0218] The policy network loss function takes the estimate of the global value network θ as:

[0219]

[0220] Where, is the forward calculation result of the global value network;

[0221] Therefore, the gradient of the action with respect to the valuation is expressed as:

[0222]

[0223] Where, is the gradient of the action of the i-th agent with respect to the global value network valuation;

[0224] Calculate the gradient of the policy network parameters with respect to the action:

[0225]

[0226] Where, is the gradient of the policy network parameters of the i-th agent in time period t with respect to the action;

[0227] The gradient of the policy network parameters to the action is calculated by the policy network back propagation method;

[0228] Calculate the gradient of the policy network parameters with respect to the estimate:

[0229]

[0230] Where, The gradient of the policy network of the i-th agent to the valuation is equal to the product of the gradient of the policy network parameters to the action and the gradient of the action to the valuation network in the same period of the corresponding agent;

[0231] Targeted Policy Network Updates:

[0232]

[0233] Where, φ tar is the target policy network parameter, λ φ is the target policy network learning rate.

[0234] In a possible implementation, the step of adding a stage target and ending the training after the staged training converges includes:

[0235] Determine whether the loss function converges:

[0236]

[0237] Calculate whether the change in the loss function satisfies the convergence condition ε π ;

[0238] If not satisfied, continue to perform the policy network and global value network θ and the corresponding target global value network θ under the current reward function tar Parameter update;

[0239] Weighting coefficients and stage mapping:

[0240] Phase 1: Adding Balance Rewards blc r blc and cost reward ω cst r cst , so that the self-organizing group regulation can meet the regional power grid power balance constraints under the minimum cost target;

[0241] Phase 2: Adding line voltage bonus ω line r line and node voltage reward ω v r v , so that the regulation target of the self-organizing group meets the safety constraints of power grid operation;

[0242] Phase 3: Adding new energy consumption incentives ne r ne and group response bias reward ω err r err, making the self-organized group adjustment target conducive to the consumption of new energy and the completion of group resources;

[0243] The objective function is adjusted in stages. The weighted coefficient of the objective function is initialized to 0. After each stage of training reaches the convergence condition, the weighted coefficient is modified to 1 in sequence until the training ends after the last stage of training converges.

[0244] The self-organizing group makes autonomous decisions. The group strategy network performs forward calculations based on the power grid measurement information and user-side measurement information, and outputs the current environmental state S t The group action amount a under t,i In the step, the self-organizing group i is based on the electricity quantity and resource measurement statistics obs observed by the group. t,i Perform forward calculation of the strategy network and adjust the group action amount a t,i :

[0245]

[0246] Where, Represents the forward calculation result of the policy network;

[0247] Self-organizing flexible resource response, the regional power grid group resources according to the group action amount a t,i In the step of executing the operation to output the power of the flexible resources to the regional power grid connection node, the internal resources of the regional power grid group perform charging and discharging, temperature regulation and distributed new energy power control according to the published group action strategy.

[0248] The multi-objective, phased learning method for regional power grid self-organizing group decision-making in this embodiment of the present invention establishes a training framework for self-organizing group multi-agent decision-making agents, enabling autonomous and intelligent decision-making within resource groups and intelligent interaction between sources, grids, loads, and storage. The proposed training method utilizes phased neural network parameter optimization technology to improve neural network learning efficiency and convergence.

[0249] Another embodiment of the present invention further provides a regional power grid self-organizing group decision-making system with multi-objective phased learning, comprising:

[0250] The decision-making process establishment module is used to establish several group strategy networks through regional power grid resources. The i-th group strategy network performs forward calculation based on its own observations and outputs the current environment state S t The group action amount a under t,i , according to the group action amount a t,i Output group power p i,t , the power injected by the node connected to the group is included in the group power p i,t , using grid interactive simulation to calculate power flow, output the active power p of the power system node bus , reactive power qbus and voltage v bus ; Based on the environmental state S before and after the decision t and S t+1 and action a t,i Calculate the single-round reward value r t , the single-round decision experience (S t ,r t ,a t,i ,S t+1 ) is stored in the experience pool;

[0251] Neural network parameter optimization module is used to optimize neural network parameters, extract experience data from the experience pool, and calculate the network parameters according to (S t ,r t ,S t+1 ) calculates the network gradient by the least mean square root; according to (S t ,a t,i ) Calculate the full-cycle benefit valuation Q of the group action in the current environment through the value network t , based on the full cycle return valuation Q t and group action amount a t,i Calculate the gradient of the swarm strategy network; calculate the loss function value of the value network, and determine whether it has converged based on the change in the loss function value. If it has converged, add a new stage objective function item, otherwise continue to update the value network and the swarm strategy network; until all objective function items are taken into account and the loss function converges;

[0252] The decision output module is used by the group strategy network to perform forward calculations based on the power grid measurement information and user-side measurement information, and output the current environment state S t The group action amount a under t,i , the regional power grid group resources are based on the group action amount a t,i Execute operations to enable flexible resource power to be exported to regional power grid connection nodes.

[0253] In a possible implementation, the step of the decision-making process establishment module establishing a plurality of group strategy networks using regional power grid resources includes:

[0254] Self-organizing group observation space design:

[0255]

[0256] Self-organizing group action space design:

[0257] a t,i =(ΔT ac ,ΔP disc ,ΔP charg ,ΔP pv ,ΔP wd )

[0258] Design of multi-objective reward function for self-organizing groups:

[0259] r=ω blc r blc +ω cst r cst +ω line r line +ω v r v +ω ne r ne +ω err r err

[0260] Where, obs t,i is the observation space of the ith agent in time period t, is the mean room temperature within the group, T env is the ambient temperature, P se is the energy storage power, P pv is the photovoltaic power, P wd is the wind power, P bus Inject power into the node, V bus is the node voltage; a t,i is the agent action space, ΔT ac Temperature adjustment, ΔP disc is the discharge power adjustment amount, ΔP charg is the charging power adjustment amount, ΔP pv is the photovoltaic power regulation, ΔP wd is the wind power regulation;

[0261] Multi-agent reinforcement learning is performed on several swarm strategy networks established through regional power grid resources, including:

[0262] Multi-agent strategy network design;

[0263] Multi-agent target strategy network design;

[0264] Multi-agent global value network design;

[0265] Multi-agent global objective value network design;

[0266] Update the multi-agent global value network and the target global value network;

[0267] Perform multi-agent policy network and target policy network updates;

[0268] Add stage goals and end training after stage-by-stage training converges.

[0269] In a possible implementation, the self-organizing group makes autonomous decisions, and the decision output module outputs the statistical information obs of the self-organizing group i based on the electrical quantity and resource measurement observed by the group. t,i Perform forward calculation of the strategy network and adjust the group action amount a t,i :

[0270]

[0271] Where, Represents the forward calculation result of the policy network;

[0272] Self-organizing flexible resource response, the internal resources of the regional power grid group perform charging and discharging, temperature regulation and distributed new energy power control based on the published group action strategy.

[0273] Another embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the multi-objective phased learning regional power grid self-organizing group decision-making method.

[0274] Another embodiment of the present invention further provides a computer-readable storage medium, which stores at least one instruction. When the at least one instruction is executed by a processor, it implements the regional power grid self-organizing group decision-making method with multi-objective staged learning.

[0275] The computer program includes computer program code, which may be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device, medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code. It should be noted that the content contained in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electric carrier signals and telecommunication signals. For ease of explanation, the above content only shows the part related to the embodiment of the present invention. For specific technical details not disclosed, please refer to the method part of the embodiment of the present invention. The computer-readable storage medium is non-transitory and can be stored in a storage device formed by various electronic devices, and can implement the execution process recorded in the method of the embodiment of the present invention.

[0276] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0277] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems) and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0278] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0279] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0280] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.

Claims

1. A multi-objective phased learning regional power grid self-organizing group decision-making method, characterized by: include: Several group strategy networks are established through regional power grid resources. The i-th group strategy network performs forward calculation based on its own observations and outputs the current environment state S t The group action amount a under t,i , according to the group action amount a t,i Output group power p i,t , the power injected by the node connected to the group is included in the group power p i,t ; Based on the environmental state S before and after the decision t and S t+1 and action a t,i Calculate the single-round reward value r t , the single-round decision experience (S t ,r t ,a t,i ,S t+1 ) is stored in the experience pool; Optimize the neural network parameters, extract experience data from the experience pool, and set the network value according to (S t ,r t ,S t+1 )Calculate the network gradient; According to (S t ,a t,i ) Calculate the full-cycle benefit valuation Q of the group action in the current environment through the value network t , based on the full cycle return valuation Q t and group action amount a t,i Calculate the gradient of the swarm strategy network; calculate the loss function value of the value network, and determine whether it has converged based on the change in the loss function value, until all objective function items are taken into account and the loss function converges; The swarm strategy network performs forward calculations and outputs the current environment state S t The group action amount a under t,i , the regional power grid group resources are based on the group action amount a t,i Execute operations to enable flexible resource power to be exported to regional power grid connection nodes; The steps of establishing a plurality of group strategy networks by using regional power grid resources include: Self-organizing group observation space design: Self-organizing group action space design: a t,i =(ΔT ac ,ΔP disc ,ΔP charg ,ΔP pv ,ΔP wd ) Design of multi-objective reward function for self-organizing groups: r=ω blc r blc +oh cst r cst +oh line r line +oh v r v +oh ne r ne +oh err r err Where, obs t,i is the observation space of the ith agent in time period t, is the mean room temperature within the group, T env is the ambient temperature, P se is the energy storage power, P pv is the photovoltaic power, P wd is the wind power, P bus Inject power into the node, V bus is the node voltage; a t,i is the agent action space, ΔT ac Temperature adjustment, ΔP disc is the discharge power adjustment amount, ΔP charg is the charging power adjustment amount, ΔP pv is the photovoltaic power regulation, ΔP wd is the wind power regulation; r blc To balance the reward value, ω blc To balance the reward weight; r cst is the cost reward value, ω cst is the cost reward weight; r line is the line voltage bonus value, ω line is the line voltage reward weight; r v is the node voltage reward value, ω v is the node voltage reward weight; r ne is the new energy consumption reward value, ω ne is the reward weight for new energy consumption; r err is the group response deviation reward value, ω err Reward weights for group response bias.

2. The regional power grid self-organizing group decision-making method with multi-objective phased learning according to claim 1 is characterized in that: It also includes multi-agent reinforcement learning of several swarm strategy networks established through regional power grid resources, including: Multi-agent strategy network design; Multi-agent target strategy network design; Multi-agent global value network design; Multi-agent global objective value network design; Update the multi-agent global value network and the target global value network; Perform multi-agent policy network and target policy network updates; Add stage goals and end training after stage-by-stage training converges.

3. The regional power grid self-organizing group decision-making method with multi-objective phased learning according to claim 2 is characterized in that: The steps of designing the multi-agent strategy network include: Input layer dimension calculation: NP in,i =N obs,i Output layer dimension calculation: NP out,i =N a,i Hidden layer dimension calculation: Where NP in,i is the dimension of the policy network input layer, N obs,i is the observation vector dimension, NP out,i is the dimension of the policy network output layer, N a,i is the action vector dimension, NP h,i is the hidden layer dimension of the policy network, N s is the number of training samples, α is the diversity adjustment parameter; E.G in,i ≤NP h,i ≤NP out,i NP h,i ≤2*NP in,i NP obtained based on the hidden layer dimension calculation formula h,i , bring in the above two dimensional conditions determined by the input and output layer dimensions, when NP h,i If the conditions are met, the original value can be used; otherwise, the boundary value of the above conditions can be used; Set the number of hidden layers to be greater than 2. By comparing the results, increase the number of layers for underfitting and reduce the number of layers for overfitting. Overfitting is judged based on the fact that the reward obtained by decision-making using the training set data reaches the upper limit, but the reward obtained by decision-making using the test set data reaches the lower limit. Underfitting is judged based on the fact that the reward values ​​obtained by decision-making on both the training set and test set data reach the lower limit.

4. The regional power grid self-organizing group decision-making method with multi-objective phased learning according to claim 2 is characterized in that: The steps of designing the multi-agent global value network include: Input layer dimension calculation: NQ in,i =N obs,i +N a,i Output layer dimension calculation: NQ out,i =1 Hidden layer dimension calculation: NQ in,i ≤NQ h,i ≤NQ out,i NQ h,i ≤2*NQ in,i Where NQ in,i Input dimension for the value network, NQ out,i is the output dimension of the value network, NQ h,i is the hidden layer dimension of the network.

5. The regional power grid self-organizing group decision-making method with multi-objective phased learning according to claim 2 is characterized in that: The steps of updating the multi-agent global value network and the target global value network include: Target global value network θ tar Initialize with a copy of the global value network θ; The global value network θ is in the current target global value network θ tar Add random disturbance on the basis of Extract multiple experiences (S t ,r t ,a t,i ,S t+1 ) to update the target global value network, including: Calculate the action in period t+1: Where a i,t+1 is the strategy network at time t+1 The decision result of π(·) represents the forward calculation result of the policy network; Calculate the target value network estimation: In t+1 =V θ (With t+1 ,and t+1 ) Among them, V t+1 is the target global value network θ at time t+1 tar The valuation result of V θ (·) represents the forward calculation result of the global value network θ; Profit calculation: 5 t =r t +γV t+1 Where V t is the total global benefit of multi-agent decision-making under the action and state at time t, r t is the reward function value at time t, and γ is the discount coefficient of future benefits; Action estimation calculation: Q t =Q θ (a i,t ,S t ) Where Q t is the estimated value of the global value network θ for the global benefit of multi-agent decision-making at the action and state at time t, Q θ (·) is the forward calculation result of the global value network θ; Loss function calculation: Among them, J Q (θ) represents the loss function of the global value network θ; Gradient calculation: Where, Represents the target global value network θ tar The gradient, Represents the gradient of all parameters of the global value network θ to the valuation, Calculated using the back-propagation method; Target value network update: Among them, λ Q is the target global value network θ tar The learning rate.

6. The regional power grid self-organizing group decision-making method with multi-objective phased learning according to claim 5 is characterized in that: The steps of updating the multi-agent policy network and the target policy network include: Calculate action estimates: Take the global value network θ to estimate the global benefit Q of the multi-agent decision at the action and state at time t t ; Calculate the gradient of the action with respect to the loss function: The policy network loss function takes the estimate of the global value network θ as: Jπ i =Q θ (S t ,a i ) Where Jπ i is the forward calculation result of the global value network; Therefore, the gradient of the action with respect to the valuation is expressed as: Where, is the gradient of the action of the i-th agent with respect to the global value network valuation; Calculate the gradient of the policy network parameters with respect to the action: Where, is the gradient of the policy network parameters of the i-th agent to the action in period t; The gradient of the policy network parameters to the action is calculated by the policy network back propagation method; Calculate the gradient of the policy network parameters with respect to the estimate: Where, The gradient of the policy network of the i-th agent to the valuation is equal to the product of the gradient of the policy network parameters to the action and the gradient of the action to the valuation network in the same period of the corresponding agent; Targeted Policy Network Updates: Where, φ tar is the target policy network parameter, λ φ is the target policy network learning rate.

7. The regional power grid self-organizing group decision-making method with multi-objective phased learning according to claim 2 is characterized in that: The steps of increasing the stage target and ending the training after the staged training converges include: Determine whether the loss function converges: Calculate whether the change in the loss function satisfies the convergence condition ε π ; If not satisfied, continue to perform the policy network and global value network θ and the corresponding target global value network θ under the current reward function tar Parameter update; Weighting coefficients and stage mapping: Phase 1: Adding Balance Rewards blc r blc and cost reward ω cst r cst , so that the self-organizing group regulation can meet the regional power grid power balance constraints under the minimum cost target; Phase 2: Adding line voltage bonus ω line r line and node voltage reward ω v r v , so that the regulation target of the self-organizing group meets the safety constraints of power grid operation; The third stage: adding new energy consumption incentives ne r ne and group response bias reward ω err r err , making the self-organized group adjustment target conducive to the consumption of new energy and the completion of group resources; The objective function is adjusted in stages. The weighted coefficient of the objective function is initialized to 0. After each stage of training reaches the convergence condition, the weighted coefficient is modified to 1 in sequence until the training ends after the last stage of training converges.

8. The regional power grid self-organizing group decision-making method with multi-objective phased learning according to claim 1 is characterized in that: In the step of calculating the loss function value of the value network and determining whether convergence is achieved based on the change in the loss function value, if convergence is achieved, a new stage objective function item is added; otherwise, the value network and the swarm strategy network are continued to be updated; In the step of forward calculation of the group strategy network, the group strategy network performs forward calculation based on the power grid measurement information and the user side measurement information, and outputs the current environment state S t The group action amount a under t,i ; The self-organizing group i is based on the electricity quantity and resource measurement statistics obs observed by the group t,i Perform forward calculation of the strategy network and adjust the group action amount a t,i : Where, Represents the forward calculation result of the policy network; The group resources of the regional power grid are based on the group action amount a t,i In the step of executing the operation to output the power of the flexible resources to the regional power grid connection node, the internal resources of the regional power grid group perform charging and discharging, temperature regulation and distributed new energy power control according to the published group action strategy.

9. A multi-objective phased learning regional power grid self-organizing group decision-making system, characterized by: include: The decision-making process establishment module is used to establish several group strategy networks through regional power grid resources. The i-th group strategy network performs forward calculation based on its own observations and outputs the current environment state S t The group action amount a under t,i , according to the group action amount a t,i Output group power p i,t , the power injected by the node connected to the group is included in the group power p i,t ; Based on the environmental state S before and after the decision t and S t+1 and action a t,i Calculate the single-round reward value r t , the single-round decision experience (S t ,r t ,a t,i ,S t+1 ) is stored in the experience pool; Neural network parameter optimization module is used to optimize neural network parameters, extract experience data from the experience pool, and calculate the network parameters according to (S t ,r t ,S t+1 )Calculate the network gradient; According to (S t ,a t,i ) Calculate the full-cycle benefit valuation Q of the group action in the current environment through the value network t , based on the full cycle return valuation Q t and group action amount a t,i Calculate the gradient of the swarm strategy network; calculate the loss function value of the value network, and determine whether it has converged based on the change in the loss function value, until all objective function items are taken into account and the loss function converges; The decision output module is used by the group strategy network to perform forward calculations based on the power grid measurement information and user-side measurement information, and output the current environment state S t The group action amount a under t,i , the regional power grid group resources are based on the group action amount a t,i Execute operations to enable flexible resource power to be exported to regional power grid connection nodes; The steps of the decision-making process establishment module establishing a plurality of group strategy networks through regional power grid resources include: Self-organizing group observation space design: Self-organizing group action space design: a t,i =(ΔT ac ,ΔP disc ,ΔP charg ,ΔP pv ,ΔP wd ) Design of multi-objective reward function for self-organizing groups: r=ω blc r blc +oh cst r cst +oh line r line +oh v r v +oh ne r ne +oh err r err Where, obs t,i is the observation space of the ith agent in period t, T ac is the mean room temperature within the group, T env is the ambient temperature, P se is the energy storage power, P pv is the photovoltaic power, P wd is the wind power, P bus Inject power into the node, V bus is the node voltage; a t,i is the agent action space, ΔT ac Temperature adjustment, ΔP disc is the discharge power adjustment amount, ΔP charg is the charging power adjustment amount, ΔP pv is the photovoltaic power regulation, ΔP wd is the wind power regulation; r blc To balance the reward value, ω blc To balance the reward weight; r cst is the cost reward value, ω cst is the cost reward weight; r line is the line voltage bonus value, ω line is the line voltage reward weight; r v is the node voltage reward value, ω v is the node voltage reward weight; r ne is the new energy consumption reward value, ω ne is the reward weight for new energy consumption; r err is the group response deviation reward value, ω err Reward weights for group response bias.

10. The regional power grid self-organizing group decision-making system with multi-objective phased learning according to claim 9 is characterized in that: The decision-making process establishment module performs multi-agent reinforcement learning on several group strategy networks established through regional power grid resources, including: Multi-agent strategy network design; Multi-agent target strategy network design; Multi-agent global value network design; Multi-agent global objective value network design; Update the multi-agent global value network and the target global value network; Perform multi-agent policy network and target policy network updates; Add stage goals and end training after stage-by-stage training converges.

11. The regional power grid self-organizing group decision-making system with multi-objective phased learning according to claim 9 is characterized in that: The decision output module converts the self-organizing group i into the electricity quantity and resource measurement statistical information obs observed by the group t,i Perform forward calculation of the strategy network and adjust the group action amount a t,i : Where, Represents the forward calculation result of the policy network; The internal resources of the regional power grid group perform charging and discharging, temperature regulation, and distributed new energy power control based on the published group action strategy.

12. An electronic device, characterized in that: The system comprises a processor and a memory, wherein the processor is configured to execute a computer program stored in the memory to implement the regional power grid self-organizing group decision-making method with multi-objective staged learning as claimed in any one of claims 1 to 8.

13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, the regional power grid self-organizing group decision-making method with multi-objective staged learning as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Power grid multi-section power automatic control method based on distributed multi-agent reinforcement learning

    CN112615379A

  • Temperature control load cluster characteristic analysis method, system and device and readable storage medium

    CN115437255A