Control system, method, device and power system for a microgrid

CN122660041APending Publication Date: 2026-08-28SUNGROW POWER SUPPLY (NANJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510221482.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

然而常用的控制系统在处理各个微电网的部分故障时会依赖主机的集中决策,主机一旦故障会严重影响电力系统的稳定性和运行安全

Benefits of technology

[0037] This application discloses a microgrid control system comprising multiple independently operating intelligent agents. Each agent is connected to a corresponding microgrid and is used to acquire the state parameters of the corresponding microgrid under the current scenario. Each agent is equipped with a pre-built and offline trained policy network, which outputs control actions based on the state parameters of the corresponding microgrid. These control actions are used to adjust the state of the corresponding microgrid under the current scenario. The microgrid control system provided in this application adopts a decentralized distributed architecture, utilizing multiple distributed agents to perform private observation and individual decision-making on their respective microgrids. This effectively handles fault conditions. Furthermore, the multiple agents do not communicate with each other, and the microgrid's decision-making process does not rely on a host computer or require global observation information to participate in the decision-making process. This not only greatly reduces communication pressure but also effectively improves the stability and operational safety of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122660041A_ABST
    Figure CN122660041A_ABST
Patent Text Reader

Abstract

The application discloses a micro-grid control system, method, device and power system, and relates to the technical field of power systems. The system comprises a plurality of agents that operate independently of each other, each agent being connected to a corresponding micro-grid and being configured to obtain state parameters of the corresponding micro-grid in a current scenario, wherein each agent is provided with a pre-constructed and offline-trained strategy network, the strategy network being configured to output a control action based on the state parameters of the corresponding micro-grid, the control action being used to adjust the state of the corresponding micro-grid in the current scenario. The application adopts a decentralized distributed architecture, and utilizes a plurality of distributed agents to respectively perform private observation and independent decision-making on corresponding micro-grids, so that fault conditions can be effectively handled, the decision-making process of the micro-grid does not need to rely on an upper host, and global observation information is not needed to participate in decision-making. As a result, not only can communication pressure be greatly reduced, but also the stability and operation safety of the power system can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power system technology, specifically to a control system, method, device, and power system for a microgrid. Background Technology

[0002] With the increasing penetration rate of distributed energy resources in the power grid and the continuous development of energy storage technology, new power systems that coordinate and interact across the power generation, grid, load, and storage systems are also constantly evolving. Setting up integrated microgrids is an effective way to promote regional energy structure transformation. A microgrid is a small-scale power generation and distribution system composed of distributed power sources, energy storage devices, electrical loads, monitoring and protection devices, etc. Microgrids can be connected to the main power grid, and their coordinated control plays a crucial role in promoting power system balance and ensuring reliable power supply.

[0003] With the high proportion of microgrids integrated into the power grid, the stability of the power system faces significant challenges. To address these safety and stability issues, power systems typically employ control systems to handle faults arising from microgrids. However, commonly used control systems rely on centralized decision-making by the main control unit when dealing with partial faults in individual microgrids. A failure of this main control unit can severely impact the stability and operational safety of the power system. Summary of the Invention

[0004] The embodiments of this application provide a control system, method, device, and power system for a microgrid to effectively handle microgrid fault conditions and improve the stability and operational safety of the power system.

[0005] To address the aforementioned technical problems, embodiments of this application disclose the following technical solutions:

[0006] Firstly, a control system for a microgrid is provided, comprising:

[0007] At least two intelligent agents, which operate independently of each other, and each intelligent agent is connected to the corresponding microgrid to obtain the state parameters of the corresponding microgrid in the current scenario;

[0008] Each of the intelligent agents is equipped with a pre-built and offline trained policy network. The policy network is used to output control actions based on the state parameters of the corresponding microgrid. The control actions are used to adjust the state of the corresponding microgrid in the current scenario.

[0009] In some embodiments, the control actions output by each policy network are configured as the optimal joint actions in the current scenario. The optimal joint actions are used to characterize the joint actions with the best control effect, and the joint actions are used to characterize the sum of the control actions corresponding to all the microgrids.

[0010] In some embodiments, the intelligent agent is located at the grid connection point corresponding to the microgrid.

[0011] In some embodiments, the status parameters include grid connection point electrical quantity information, switch position information within the microgrid, and controllable resource power information connected within the microgrid.

[0012] In some embodiments, the controllable resource includes adjustable energy, adjustable load, and controllable load, wherein the adjustable load is used to characterize a load with adjustable power, the controllable load is used to characterize a load that can be disconnected or connected, and the control action includes at least one of switching on / off, disconnecting the controllable load, adjusting the power of the adjustable load, and adjusting the power of the adjustable energy.

[0013] In a second aspect, a control method for a microgrid is provided, applied to an intelligent agent in a control system of a microgrid as described in any of the first aspects, the method comprising:

[0014] Obtain the state parameters of the microgrid connected in the current scenario;

[0015] Input the state parameters of the microgrid in the current scenario into the policy network, and obtain the control actions output by the policy network;

[0016] The state of the microgrid connected in the current scenario is adjusted based on the control action.

[0017] In some embodiments, the step of obtaining the policy network includes:

[0018] Construct the policy network and initialize the parameters of the policy network;

[0019] Determine the state parameters of each microgrid under each simulation scenario;

[0020] Based on the state parameters of the microgrid in each of the simulated scenarios, the policy network corresponding to the microgrid is trained by a deterministic gradient policy to determine the control action corresponding to the microgrid in each of the simulated scenarios.

[0021] A target function is constructed to evaluate the control effect of executing joint actions in each of the simulation scenarios, and the joint actions are used to characterize the sum of control actions corresponding to all the microgrids.

[0022] Based on the state parameters of all microgrids in each simulation scenario, the control actions corresponding to all microgrids, and the optimal solution of the objective function in each simulation scenario, the parameters of each policy network are optimized to obtain each trained policy network.

[0023] In some embodiments, the state parameters of the microgrid in the current scenario are input into a policy network to obtain the control actions output by the policy network, including:

[0024] If the current scenario is any of the simulated scenarios, the control action corresponding to the microgrid in the simulated scenario is determined as the control action output by the strategy network;

[0025] If the current scene does not belong to any of the simulated scenes, determine the target scene that is closest to the current scene from each of the simulated scenes;

[0026] Based on the Euclidean distance between the state parameters of the microgrid in the current scenario and the state parameters of the microgrid in the target scenario, and the control action corresponding to the microgrid in the target scenario, the control action output by the policy network in the current scenario is determined.

[0027] In some embodiments, the construction of the objective function includes:

[0028] Based on the impact of each microgrid on the power system, the timeliness coefficient of the joint action, the control action corresponding to each microgrid, and the state parameters of each microgrid in the simulation scenario, the objective function is constructed. The timeliness coefficient of the joint action is used to characterize the impact of the joint action on the power system.

[0029] In some embodiments, the step of obtaining the joint action timeliness coefficient includes:

[0030] Determine the total time consumption of the combined actions in the simulated scenario;

[0031] Based on the preset functional relationship between the total time consumption and the timeliness coefficient of the joint action, the timeliness coefficient of the joint action in the simulated scenario is determined.

[0032] Thirdly, a microgrid control device is provided, comprising:

[0033] A memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements a microgrid control method as described in any of the second aspects.

[0034] Fourthly, a power system is provided, comprising:

[0035] At least two microgrids, each microgrid comprising a new energy generation unit, an energy storage unit, and a load unit;

[0036] At least two intelligent agents, which operate independently of each other, each intelligent agent corresponds one-to-one with the microgrid, and each intelligent agent is connected to the new energy power generation unit, the energy storage unit and the load unit respectively. The intelligent agents are used to adopt the microgrid control method as described in any of the second aspects, or include the microgrid control equipment as described in the third aspect.

[0037] This application discloses a microgrid control system comprising multiple independently operating intelligent agents. Each agent is connected to a corresponding microgrid and is used to acquire the state parameters of the corresponding microgrid under the current scenario. Each agent is equipped with a pre-built and offline trained policy network, which outputs control actions based on the state parameters of the corresponding microgrid. These control actions are used to adjust the state of the corresponding microgrid under the current scenario. The microgrid control system provided in this application adopts a decentralized distributed architecture, utilizing multiple distributed agents to perform private observation and individual decision-making on their respective microgrids. This effectively handles fault conditions. Furthermore, the multiple agents do not communicate with each other, and the microgrid's decision-making process does not rely on a host computer or require global observation information to participate in the decision-making process. This not only greatly reduces communication pressure but also effectively improves the stability and operational safety of the power system. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0039] Figure 1 A schematic diagram of a power system provided in an embodiment of this application;

[0040] Figure 2 This is a schematic diagram of the overall process of a microgrid control method according to an embodiment of this application;

[0041] Figure 3 This is a schematic diagram of the software structure of the intelligent agent in an embodiment of this application;

[0042] Figure 4 This is a schematic diagram of the hardware structure of the intelligent agent in an embodiment of this application.

[0043] Reference numerals: 10, Microgrid; 11, New energy power generation unit; 12, Energy storage unit; 13, Load unit; 20, Intelligent agent; 21, Control device of microgrid; 211, Parameter acquisition module; 212, Action output module; 213, State adjustment module; 22, Control equipment of microgrid; 221, Memory; 222, Processor; 30, Control system of microgrid. Detailed Implementation

[0044] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0045] In the description of this application, it should be understood that the terms "upper," "lower," "front," "rear," "left," "right," "top," "bottom," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation on this application. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, and "at least one" can mean one, two, or more, unless otherwise explicitly specified.

[0046] With the increasing penetration rate of distributed energy resources in the power grid and the continuous development of energy storage technology, a new type of power system integrating "source, grid, load, and storage" is also constantly evolving. Setting up microgrids that integrate "source, grid, load, and storage" is an effective way to promote the transformation of regional energy structures. A microgrid is a small-scale power generation and distribution system composed of distributed power sources, energy storage devices, electrical loads, monitoring and protection devices, etc. Microgrids can be connected to the main power grid, and the coordinated control of microgrids plays a crucial role in promoting power balance in the power system and ensuring a reliable power supply.

[0047] Currently, power systems characterized by "source-grid-load-storage" are exhibiting new features. These are primarily reflected in the increasing proportion of traditional load lines possessing both source and load characteristics, and the growing stability issues arising from the lack of conventional power source inertia in scenarios where microgrids integrating "source-grid-load-storage" have a high proportion of renewable energy integration. To address these stability concerns, power systems can be configured with three lines of defense. The first line of defense is relay protection equipment, protecting individual components. The second line of defense is low-frequency, low-voltage load shedding devices based on local information, which can operate based on fault events and strategy tables. The third line of defense is a centralized low-frequency, low-voltage load shedding device based on overall grid information, which can indiscriminately disconnect predetermined load lines in stages when the frequency and voltage are detected to be below a threshold value. However, the new power system still faces the following problems: stability issues are extremely complex and difficult to enumerate; low-frequency and low-voltage load shedding devices based on local information cannot coordinate the overall behavior of the system; centralized low-frequency and low-voltage load shedding devices based on overall grid information rely on the decision-making of the host, and a failure of the host will seriously affect the stability and operational safety of the power system; the system also needs to be equipped with "second line of defense" and "third line of defense" devices, which are costly.

[0048] In view of this, embodiments of this application provide a control system, method, device, and power system for a microgrid. By setting up a decentralized distributed architecture, multiple intelligent agents that operate independently can perform private observations and make individual decisions on their respective microgrids. This can effectively handle fault situations. Moreover, the multiple intelligent agents do not communicate with each other, and the decision-making process of the microgrid does not rely on a host computer or require global observation information to participate in the decision-making. This can not only greatly reduce communication pressure but also effectively improve the stability and operational safety of the power system, thereby solving at least some of the above-mentioned technical problems.

[0049] Please see Figure 1 , Figure 1 This is a schematic diagram of a power system provided in an embodiment of this application. The power system may include at least two microgrids 10 and at least two intelligent agents 20, with the at least two intelligent agents 20 operating independently of each other. An intelligent agent 20 can represent a device deployed at a control node. The microgrids 10 are used for grid connection with the main power grid. Each microgrid 10 includes a new energy generation unit 11, an energy storage unit 12, and a load unit 13. Each intelligent agent 20 corresponds one-to-one with a microgrid 10, and is connected to the new energy generation unit 11, the energy storage unit 12, and the load unit 13, respectively. The intelligent agents 20 are used to employ the microgrid control method of the embodiments of this application, or include the control equipment of the microgrid of the embodiments of this application.

[0050] In some examples, the new energy generation unit 11 may include at least one of wind turbines and photovoltaics. The energy storage unit 12 may include energy storage devices. The load unit 13 may include at least one of charging piles, pumped storage power stations, general loads, air conditioning loads, and auxiliary equipment from coal-fired power plants. Taking a power system comprising three microgrids 10 as an example, the first microgrid 10 may include energy storage devices, wind turbines, photovoltaics, charging piles, pumped storage power stations, and general loads; the second microgrid 10 may include energy storage devices, wind turbines, charging piles, air conditioning loads, and general loads; and the third microgrid 10 may include energy storage devices, wind turbines, photovoltaics, auxiliary equipment from coal-fired power plants, air conditioning loads, and general loads. Air conditioning loads and charging piles are two examples of adjustable loads, while general loads are used to characterize conventional loads in the power system whose power output is not adjustable.

[0051] In this embodiment, at least two intelligent agents 20 perform status monitoring and fault handling on their respective microgrids 10, thus constituting the control system 30 of the microgrid in this application embodiment. The control system 30 of the microgrid in this application embodiment will be described below.

[0052] The microgrid control system 30 of this embodiment includes at least two independently operating agents 20. Each agent 20 is connected to a corresponding microgrid 10 and is used to acquire the state parameters of the corresponding microgrid 10 in the current scenario. Each agent 20 is equipped with a pre-built and offline trained policy network. The policy network is used to output control actions based on the state parameters of the corresponding microgrid 10, and the control actions are used to adjust the state of the corresponding microgrid 10 in the current scenario.

[0053] In some examples, state parameters may include grid connection point electrical quantity information, switch position information within the microgrid 10, and power information of controllable resources connected within the microgrid 10. In this way, an intelligent agent 20 can acquire multiple types of information from the microgrid 10, thereby reducing hardware complexity and enabling more intelligent decision-making and control.

[0054] For example, the grid connection point electrical quantity information includes the voltage, current, frequency, power, and phase angle of the grid connection point. Switch position information includes the position information of each switch within the microgrid topology. Controllable resources include adjustable energy sources, adjustable loads, and controllable loads. Adjustable resources include wind power, photovoltaics, and energy storage devices; adjustable loads characterize loads with adjustable power, such as charging piles, pumped storage power stations, auxiliary equipment from coal-fired power plants, and air conditioning loads; controllable loads characterize loads that can be disconnected or connected. Control actions include at least one of switching on / off, disconnecting controllable loads, adjusting the power of adjustable loads, and adjusting the power of adjustable energy sources. In this way, the intelligent agent 20 can monitor and control various types of information, greatly improving the flexibility and comprehensiveness of control, and thus enabling it to handle more complex fault conditions.

[0055] Through the above technical solution, the microgrid control system 30 of this application embodiment adopts a decentralized distributed architecture, which utilizes multiple distributed intelligent agents 20 to perform private observation and individual decision-making. In actual use scenarios, there is no need to rely on the host computer. The input of the policy network is only the local observation of the corresponding microgrid 10, and global observation information is not required to participate in the decision-making. There is no need for communication between the intelligent agents 20, and they are independent of each other, which can greatly reduce the communication pressure and greatly improve the operation stability and reliability of the entire power system.

[0056] In some embodiments, the control actions output by each policy network are configured as the optimal joint action for the current scenario. The optimal joint action characterizes the joint action with the best control effect, and the joint action characterizes the sum of the control actions corresponding to all microgrids 10. In this way, the overall behavior of the power system can still be coordinated through a decentralized distributed architecture. This not only achieves the optimal control effect of the entire power system in the current scenario and effectively handles fault situations, but also integrates the traditional "second line of defense" and "third line of defense," taking into account the functions of the "second line of defense" and "third line of defense," thereby eliminating the need for redundant configuration and greatly saving configuration costs.

[0057] In some embodiments, the intelligent agent 20 may be located at the grid connection point P of the corresponding microgrid 10. This allows for more efficient monitoring of electrical quantities at the grid connection point, which helps to save on cable laying costs.

[0058] It is understood that the microgrid control system 30 of this application embodiment adopts a decentralized distributed architecture and makes decisions using a multi-agent approach 20. It does not rely on a host and can coordinate the overall action behavior of the power system without communication participation in the actual use stage, so as to achieve the optimal control effect of the entire power system in the current scenario. This not only greatly reduces the communication pressure and improves the stability and reliability of the entire system, but also takes into account the functions of the "second line of defense" and the "third line of defense" without repeated configuration, thereby saving configuration costs.

[0059] The control method of the microgrid of this application embodiment executed by the intelligent agent 20 will be described below.

[0060] Please see Figure 2 , Figure 2 This is a schematic flowchart illustrating a microgrid control method according to an embodiment of this application. The microgrid control method includes the following steps:

[0061] Step 201: Obtain the state parameters of the microgrid connected in the current scenario.

[0062] In some examples, state parameters can be represented by a multi-dimensional data vector composed of the state parameters of each metric node in the microgrid under the current scenario. That is, state parameters can include information from each metric node, such as grid-connected electrical quantity information, switch position information within the microgrid, and controllable resource power information connected to the microgrid. Grid-connected electrical quantity information includes voltage, current, frequency, power, and phase angle, which can be directly acquired. Switch position information includes the position information of each switch within the microgrid topology, which can be obtained through switch position contacts, circuit breaker status, and open / closed contact signals around the access node. Controllable resources include adjustable energy sources, adjustable loads, and controllable loads. Controllable resource power information includes at least one of current power, adjustable power, and adjustable power, which can be obtained through GOOSE (Generic Object Oriented Substation Event). This improves data acquisition efficiency and ensures timely control.

[0063] Step 202: Input the state parameters of the microgrid in the current scenario into the policy network and obtain the control actions output by the policy network.

[0064] Specifically, each agent is equipped with a policy network, which takes as input parameters from its private observations and outputs control actions.

[0065] In some embodiments, the step of obtaining the policy network includes:

[0066] Step 1: Build the policy network and initialize its parameters.

[0067] Specifically, the policy network can be represented as a i =μ(o i ,θ i ), where a i o represents the control action output by policy network i. i θ represents the parameters of the private observations of the agent corresponding to policy network i. i The parameters represent the current deterministic gradient policy, which is used to directly map the state s to a specific action a (rather than the probability distribution of the action).

[0068] Step 2: Determine the state parameters of each microgrid in each simulation scenario.

[0069] In some examples, simulation software, such as Matlab, can be used to simulate various possible faults in each microgrid. The fault in each simulation scenario is characterized by private observations of all microgrids in the power system. That is, the fault in each simulation scenario includes a set of states of each node in the power system, such as changes in the position of a switch or abrupt changes in the current information of a node, which are reflected in the private observations of at least some agents.

[0070] Step 3: Based on the state parameters of the microgrid in each simulation scenario, train the policy network corresponding to the microgrid using a deterministic gradient policy to determine the control actions of the microgrid in each simulation scenario.

[0071] Specifically, Deterministic Policy Gradient (DPG) is a policy gradient method for reinforcement learning, specifically designed for problems involving continuous action spaces. DPG directly learns a deterministic policy, meaning that given a state s, it outputs a specific action a. A DPG can include an Actor (action) network and a value network (used to determine the Q-value of a given state-action pair). During training, the parameters θ of the DPG are updated using gradient ascent, with the goal of maximizing the expected Q-value, i.e., adjusting θ along the gradient direction of the Q-value so that the policy chooses the better action given the state.

[0072] Each policy network is trained offline based on the state parameters of the corresponding microgrid in each simulation scenario, and finally outputs the control action of the corresponding microgrid. The state parameters of the corresponding microgrid can be extracted from the set of a certain state of all metric nodes of the power system corresponding to the fault in each simulation scenario. For example, for simulation scenario X, policy network A1 is trained based on the private observation data of microgrid 1 in the private observation set corresponding to simulation scenario X, and outputs the control action of microgrid 1 in simulation scenario X. Policy network A2 is trained based on the private observation data of microgrid 2 in the private observation set corresponding to simulation scenario X, and outputs the control action of microgrid 2 in simulation scenario X.

[0073] Step four: Construct the objective function. The objective function is used to evaluate the control effect of executing joint actions in each simulation scenario. The joint actions are used to characterize the sum of control actions corresponding to all microgrids.

[0074] Specifically, the objective function is used to evaluate the quality of performing a certain joint policy action in the current state.

[0075] In some embodiments, the method for constructing the objective function includes:

[0076] Based on the impact of each microgrid on the power system, the timeliness coefficient of the joint action, the control actions corresponding to each microgrid, and the state parameters of each microgrid in the simulated scenario, an objective function is constructed. The timeliness coefficient of the joint action is used to characterize the impact of the joint action on the power system. In this way, the evaluation of whether the joint action is optimal can fully take into account the state parameters, control actions, the impact of the microgrid on the power system, and the timeliness of the joint action, thus enabling a more comprehensive and accurate evaluation of the control effect.

[0077] In some examples, the steps for obtaining the joint action timeliness coefficient include:

[0078] Determine the total time t of the joint actions in the simulated scenario. c .

[0079] Based on total time t c Based on the preset functional relationship with the joint action timeliness coefficient, the joint action timeliness coefficient a(t) in the simulated scenario is determined. c ).

[0080] Among them, the timeliness coefficient of joint action a(t) c This is used to measure the timeliness of action of all devices; the joint timeliness coefficient a(t) is used to measure the timeliness of action of all devices. c This is used to characterize the impact of action timeliness on the entire system. A first-order proportional function can be selected, or it can be adjusted according to the actual optimization effect. This application does not specifically limit this.

[0081] For example, the objective function y can be expressed by the following formula:

[0082] y = a(t) c )*r;

[0083]

[0084] Where r is the initial objective function, a(t) c ) represents the timeliness coefficient of the joint action, t c The total time for the coordinated action is represented by λ, where i is the measurement node, n is the number of frequency detection nodes in the entire power system, and λ represents the total time for the coordinated action. fi To measure the weight of the impact of the frequency deviation at node i on the entire power system, f i To measure the frequency of node i, f N To measure the rated power of node i, λ Ui To measure the weight of the impact of the voltage deviation at node i on the entire power system, U i To measure the voltage at node i, U N Let λ represent the rated voltage of node i, j be the number of the controllable load, m be the number of controllable loads, and λ be the value of the load. LiTo comprehensively consider the weights of the controllable load's importance, recovery difficulty, and the impact on industrial and residential production and daily life after its removal, dj is a coefficient characterizing whether the controllable load is removed; a value of 1 indicates removal, and 0 indicates otherwise. j For controllable load, k is the number of the adjustable resource, l is the total number of adjustable resources, and λ Ck To comprehensively consider the importance and adjustment difficulty of various adjustable loads, photovoltaic, wind power, energy storage, and other adjustable resources, C k The current power of the adjustable resource, h is the tie-line number, p is the number of tie-lines connecting each secondary grid to the outside world, and λ is the current power of the adjustable resource. Ph To measure the weight of the impact of changes in power interaction between each secondary power grid and the external environment on the overall system, P h P represents the current power value of the tie line. h_before This represents the power value of the tie line before the microgrid fault.

[0085] Step 5: Based on the state parameters of all microgrids, the corresponding control actions of all microgrids, and the optimal solution of the objective function in each simulation scenario, optimize the parameters of each policy network to obtain each trained policy network.

[0086] Specifically, for a certain joint action of all microgrids, the value of the objective function y can be calculated. For a fault in a certain simulated scenario, the optimal joint action of all microgrids can be found by continuously approximating the optimal solution of the objective function through the deterministic gradient method. That is, for the joint observation corresponding to the fault in the simulated scenario, the optimal joint action is obtained to obtain the policy network corresponding to each microgrid.

[0087] Through the above technical solutions, different fault scenarios can be quantified into specific private observation sets, which facilitates the output of control actions by the policy network. At the same time, it will continuously optimize based on the optimal solution in the simulated scenario, so that the sum of the control effects output by each policy network reaches the optimal level. This allows distributed decision-making that does not rely on the host to better coordinate the overall behavior of the power system in the actual use stage, achieve the optimal control effect of the entire power system in the current scenario, improve the operational stability and reliability of the entire system, and integrate the functions of the "second line of defense" and the "third line of defense" without repeated configuration, thereby saving configuration costs.

[0088] In some embodiments, the implementation method of step 202, based on the offline-trained policy network, includes the following steps:

[0089] In any simulated scenario, the control action corresponding to the microgrid in the simulated scenario is determined as the control action output by the strategy network.

[0090] If the current scene does not belong to any of the simulated scenes, determine the target scene that is closest to the current scene from all the simulated scenes.

[0091] Based on the Euclidean distance between the state parameters of the microgrid in the current scenario and the state parameters of the microgrid in the target scenario, and the corresponding control actions of the microgrid in the target scenario, the control actions output by the policy network in the current scenario are determined.

[0092] Specifically, the state parameters of the microgrid in the current scenario can be compared with those in the simulated scenario. If they match, the current scenario is determined to be a simulated scenario, and the corresponding control action can be directly inferred based on the trained policy network. If the state parameters of the microgrid in the current scenario do not match those in the simulated scenario, the current scenario does not belong to any simulated scenario. Based on the state parameters of each simulated scenario and the current scenario, the target scenario closest to the current scenario can be determined from among the simulated scenarios. This involves comparing the multidimensional data vector corresponding to the microgrid in each simulated scenario with the multidimensional data vector corresponding to the microgrid in the current scenario, and obtaining the simulated scenario with the closest distance as the target scenario. Then, based on the Euclidean distance between the state parameters of the microgrid in the current scenario and the state parameters of the microgrid in the target scenario, combined with the control action corresponding to the microgrid in the target scenario, the control action output by the policy network in the current scenario is determined.

[0093] Through the above technical solution, different fault scenarios will be quantified to form a specific private observation set, without having to exhaustively list all fault scenarios. When the input private observation is not within the fault samples simulated during training, the algorithm will output the most suitable policy action through the policy network via Euclidean distance, thereby being able to cope with the complex power system fault scenarios that actually occur.

[0094] Step 203: Adjust the state of the connected microgrid in the current scenario based on the control action.

[0095] Specifically, the state of each metric node in the microgrid is adjusted by the control actions output by the strategy network, such as the on / off state of each switch and the power value of adjustable loads.

[0096] It is understood that, under real fault scenarios, the control method of this application embodiment can make individual decisions and controls for each microgrid based on the local observations of each intelligent agent, without requiring global observation information to participate in the decision-making process. This can greatly reduce communication pressure and improve the operational stability and reliability of the power system. By training each policy network offline separately and optimizing them together, the overall action and behavior of the power system can still be coordinated even under individual decision-making, achieving the optimal control effect. In the simulation stage, it is not necessary to exhaustively enumerate all stability problems of the power system, and it can cope with complex power system fault scenarios that actually occur. At the same time, this method can simultaneously take into account the functions of the "second line of defense" and the "third line of defense", so that the microgrid does not need to be reconfigured, thereby saving configuration costs.

[0097] Please see Figure 3 , Figure 3 This is a schematic diagram of the software structure of the intelligent agent according to an embodiment of this application. The intelligent agent 20 in this embodiment may include a microgrid control device 21. The microgrid control device 21 includes a parameter acquisition module 211, an action output module 212, and a state adjustment module 213.

[0098] The parameter acquisition module 211 is used to acquire the status parameters of the microgrid connected in the current scenario.

[0099] The action output module 212 is used to input the state parameters of the microgrid in the current scenario into the policy network and obtain the control actions output by the policy network.

[0100] The state adjustment module 213 is used to adjust the state of the connected microgrid in the current scenario based on control actions.

[0101] It should be noted that the specific functional implementation and technical effects of each module in the microgrid control device 21 can be found in the microgrid control method of the aforementioned embodiment, and will not be described in detail here.

[0102] Please see Figure 4 , Figure 4 This is a schematic diagram of the hardware structure of the intelligent agent according to an embodiment of this application. The intelligent agent 20 in this embodiment may include a microgrid control device 22. The microgrid control device 22 includes a memory 221 and a processor 222. The memory 221 stores a program that can run on the processor 222. When the processor 222 executes the program, it can implement the microgrid control method of the aforementioned embodiment.

[0103] It is understood that, in real fault scenarios, the intelligent agent 20 of this application embodiment can make independent decisions and controls the corresponding microgrid based on local observations, without needing to communicate with each other or require global observation information to participate in the decision-making process. This can greatly reduce communication pressure and improve the operational stability and reliability of the power system. By setting up a policy network that has been trained offline and optimized together, the overall action behavior of the power system can still be coordinated even when making independent decisions, so as to achieve the best control effect. In the fault simulation stage, it is not necessary to exhaustively enumerate all stability problems of the power system, so as to cope with the complex power system fault scenarios that actually occur. At the same time, the intelligent agent 20 can simultaneously take on the functions of the "second line of defense" and the "third line of defense", so that the microgrid does not need to be reconfigured, thereby saving configuration costs.

[0104] The control system, method, device, and power system of a microgrid provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the technical solutions and core ideas of this application. Those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A control system for a microgrid, characterized in that, include: At least two intelligent agents, which operate independently of each other, and each intelligent agent is connected to the corresponding microgrid to obtain the state parameters of the corresponding microgrid in the current scenario; Each of the intelligent agents is equipped with a pre-built and offline trained policy network. The policy network is used to output control actions based on the state parameters of the corresponding microgrid. The control actions are used to adjust the state of the corresponding microgrid in the current scenario.

2. The control system for a microgrid according to claim 1, characterized in that, The control actions output by each of the policy networks are configured as the optimal joint actions in the current scenario. The optimal joint actions are used to characterize the joint actions with the best control effect, and the joint actions are used to characterize the sum of the control actions corresponding to all the microgrids.

3. The control system for a microgrid according to claim 1, characterized in that, The intelligent agent is set at the grid connection point corresponding to the microgrid.

4. The control system for a microgrid according to claim 1, characterized in that, The status parameters include grid connection point electrical quantity information, switch position information within the microgrid, and controllable resource power information connected within the microgrid.

5. The control system for a microgrid according to claim 4, characterized in that, The controllable resources include adjustable energy, adjustable load, and controllable load. The adjustable load is used to characterize a load whose power is adjustable, and the controllable load is used to characterize a load that can be disconnected or connected. The control action includes at least one of switching on / off, disconnecting the controllable load, adjusting the power of the adjustable load, and adjusting the power of the adjustable energy.

6. A control method for a microgrid, characterized in that, An intelligent agent applied in a control system for a microgrid as described in any one of claims 1-5, the method comprising: Obtain the state parameters of the microgrid connected in the current scenario; Input the state parameters of the microgrid in the current scenario into the policy network, and obtain the control actions output by the policy network; The state of the microgrid connected in the current scenario is adjusted based on the control action.

7. The microgrid control method according to claim 6, characterized in that, The steps for obtaining the policy network include: Construct the policy network and initialize the parameters of the policy network; Determine the state parameters of each microgrid under each simulation scenario; Based on the state parameters of the microgrid in each of the simulated scenarios, the policy network corresponding to the microgrid is trained by a deterministic gradient policy to determine the control action corresponding to the microgrid in each of the simulated scenarios. A target function is constructed to evaluate the control effect of executing joint actions in each of the simulation scenarios, and the joint actions are used to characterize the sum of control actions corresponding to all the microgrids. Based on the state parameters of all microgrids in each simulation scenario, the control actions corresponding to all microgrids, and the optimal solution of the objective function in each simulation scenario, the parameters of each policy network are optimized to obtain each trained policy network.

8. The microgrid control method according to claim 7, characterized in that, The state parameters of the microgrid in the current scenario are input into the policy network, and the control actions output by the policy network are obtained, including: If the current scenario is any of the simulated scenarios, the control action corresponding to the microgrid in the simulated scenario is determined as the control action output by the strategy network; If the current scene does not belong to any of the simulated scenes, determine the target scene that is closest to the current scene from each of the simulated scenes; Based on the Euclidean distance between the state parameters of the microgrid in the current scenario and the state parameters of the microgrid in the target scenario, and the control action corresponding to the microgrid in the target scenario, the control action output by the policy network in the current scenario is determined.

9. The microgrid control method according to claim 7, characterized in that, The objective function to be constructed includes: The objective function is constructed based on the degree of impact of each microgrid on the power system, the joint action timeliness coefficient, the control action corresponding to each microgrid, and the state parameters of each microgrid in the simulation scenario. The joint action timeliness coefficient is used to characterize the degree of impact of the joint action on the power system.

10. The microgrid control method according to claim 9, characterized in that, The steps for obtaining the timeliness coefficient of the joint action include: Determine the total time consumption of the combined actions in the simulated scenario; Based on the preset functional relationship between the total time consumption and the timeliness coefficient of the joint action, the timeliness coefficient of the joint action in the simulated scenario is determined.

11. A control device for a microgrid, characterized in that, include: The system includes a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the microgrid control method as described in any one of claims 6-10.

12. An electric power system, characterized in that, include: At least two microgrids, each microgrid comprising a new energy generation unit, an energy storage unit, and a load unit; At least two intelligent agents operate independently of each other, each intelligent agent corresponds to one of the microgrids, and each intelligent agent is connected to the new energy power generation unit, the energy storage unit and the load unit respectively. The intelligent agents are used to employ the microgrid control method as described in any one of claims 6-10, or include the microgrid control equipment as described in claim 11.