Distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning
By constructing a multi-agent reinforcement learning system, the coordinated control of distributed new energy sources and flexible loads is realized, which solves the problem of grid regulation imbalance and improves the flexibility and stability of grid regulation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGZHOU JINGLI ENG DESIGN CONSULTING CO LTD
- Filing Date
- 2026-04-02
- Publication Date
- 2026-05-29
AI Technical Summary
The lack of a unified coordination mechanism between distributed renewable energy and flexible load units in power grid regulation makes it difficult to fully leverage their synergistic effects, leading to imbalances in power grid regulation.
Distributed new energy units and flexible load units are constructed as multi-agent units. Reinforcement learning is used for collaborative training. By sharing the deviation penalty term, collaborative constraint signals are provided to update the control strategies of each agent to achieve collaborative control.
It improves the coordinated control effect of new energy output and flexible load regulation, avoids the coordination imbalance caused by independent optimization, and enhances the learning stability and convergence efficiency of multi-agent coordinated control.
Smart Images

Figure CN122118961A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart grid technology, specifically to a distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning. Background Technology
[0002] With the large-scale integration of distributed renewable energy sources, the output of renewable energy sources such as wind and solar power in the power grid exhibits significant randomness and volatility, posing a considerable challenge to the safe and stable operation of the power grid. Traditional power grid dispatching methods mainly rely on conventional generating units for power regulation. However, with the increasing proportion of distributed energy sources, relying solely on centralized dispatching methods is insufficient to address the power imbalance caused by fluctuations in renewable energy output in a timely manner.
[0003] On the other hand, with the development of demand-side response technology, flexible loads such as air conditioners, electric heating equipment, and electric vehicles have a certain power regulation capability and can participate in grid regulation within a certain range by adjusting their electricity consumption behavior, thereby improving the grid's regulation flexibility. However, in existing technologies, distributed renewable energy units and flexible load units usually adopt independent control or simple rule-based control methods, lacking a unified coordinated regulation mechanism, making it difficult to fully leverage their synergistic role in grid regulation. Summary of the Invention
[0004] (a) Purpose of the invention The purpose of this invention is to provide a distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning. By constructing distributed new energy units and flexible load units as multi-agents and using reinforcement learning for collaborative training, collaborative control of new energy output and flexible load regulation is achieved. In the multi-agent reinforcement learning training, a collaborative sharing deviation penalty term is used to provide collaborative constraint signals, so that the new energy agent and the flexible load agent simultaneously consider the overall regulation task sharing relationship when updating the strategy, thereby avoiding the collaborative imbalance problem caused by each agent optimizing independently based on its own reward.
[0005] (II) Technical Solution To address the above problems, this invention provides a distributed collaborative control method for new energy sources and flexible loads based on multi-agent reinforcement learning, comprising: The controllers and algorithms of the distributed new energy units and flexible load units are respectively used to construct new energy intelligent agents and flexible load intelligent agents; Based on the current control strategy, control each agent to output adjustment actions, and obtain the actual sharing factor of each agent. Based on the deviation between the actual sharing factor and the expected collaborative sharing factor, a collaborative sharing deviation penalty term is constructed; A reward function is constructed, which includes a reward function for a new energy intelligent agent, a reward function for a flexible load intelligent agent, and a collaborative sharing deviation penalty term. The collaborative sharing deviation penalty term applies to both the new energy intelligent agent and the flexible load intelligent agent. Based on the reward function, a preset multi-agent reinforcement learning algorithm is used for iterative training to update the control strategies of each agent. Based on the updated control strategy, each agent outputs adjustment actions to achieve coordinated adjustment.
[0006] In another aspect of the present invention, preferably, the new energy intelligent agent includes a first action space and a first state observation space, the first action space includes an active power output setpoint, and the first state observation space includes: adjustable output capacity, current active power output, grid demand information, and operating physical constraints; The flexible load agent includes a second action space and a second state observation space. The second action space includes the power adjustment range. The second state observation space includes: adjustable capacity, current load status, grid demand information, and comfort constraints.
[0007] In another aspect of the present invention, preferably, the desired collaborative sharing factor is obtained by the following method: Within the current time period, the adjustable output capacity of each new energy intelligent agent is obtained based on the first state observation space, and the adjustable capacity of each flexible load intelligent agent is obtained based on the second state observation space. Based on the adjustable output capacity and adjustable capacity, the capacity sharing coefficient of each intelligent agent is determined, and the capacity sharing coefficient is the expected collaborative sharing factor.
[0008] In another aspect of the present invention, preferably, the step of controlling the output adjustment actions of each agent based on the current control strategy to obtain the actual sharing factor of each agent includes: The adjustable output capacity, current active power output, operating physical constraints and grid demand information are used as inputs to the current control strategy of each new energy intelligent body, and the active power output set value is output. The active power output set value is the actual active power adjustment amount of each new energy intelligent body. The adjustable capacity, current load status, comfort constraints, and grid demand information are used as inputs to the current control strategy of each flexible load agent, and the output power adjustment range is the actual load power adjustment amount of the flexible load agent. Based on the actual active power adjustment and the actual load power adjustment, the actual sharing factor of each agent is obtained.
[0009] In another aspect of the present invention, preferably, the construction of a collaborative sharing deviation penalty term based on the deviation between the actual sharing factor and the expected collaborative sharing factor includes: Calculate the sharing deviation between the actual sharing factor and the corresponding expected collaborative sharing factor of each agent; The shared deviation values are squared and then weighted and summed to obtain the collaborative shared deviation amount. A collaborative deviation penalty term is constructed based on the system's collaborative deviation amount.
[0010] In another aspect of the present invention, preferably, the reward function of the new energy intelligent agent includes: a new energy output tracking reward item and an operational constraint penalty item; The reward function of the flexible load agent includes: a load adjustment response reward term and a comfort constraint penalty term.
[0011] In another aspect of the present invention, preferably, The new energy output tracking reward item is constructed based on the power deviation between the actual active power adjustment of the new energy intelligent agent and the grid demand information. The smaller the deviation, the larger the reward value. The operational constraint penalty item is constructed based on the constraint violation amount between the operational state of the new energy intelligent agent and the preset operational physical constraints. A penalty value is generated when the actual operational state of the new energy intelligent agent exceeds the range of the operational physical constraints.
[0012] In another aspect of the present invention, preferably, the load regulation response reward item is constructed based on the response deviation between the actual load power regulation amount of the flexible load agent and the grid demand information; the higher the response level, the greater the reward value. The comfort constraint penalty term is constructed based on the constraint violation amount between the power adjustment range of the flexible load agent and the preset comfort constraint range. A penalty is applied when the power adjustment range exceeds the preset comfort constraint range.
[0013] In another aspect of the present invention, preferably, the iterative training based on the reward function using a preset multi-agent reinforcement learning algorithm to update the control strategies of each agent includes: Within the current time period, each new energy intelligent agent and flexible load intelligent agent inputs the current policy network based on their respective state observation space and outputs the corresponding adjustment action; The reward and penalty values for each agent are calculated based on the adjustment actions and the power grid operating status. The state observations, adjustment actions, reward values, penalty values, and state observations for the next time period are used to form agent interaction samples. Update the policy network parameters of each agent based on the agent interaction samples; Repeat the above training process until the reward function converges or the preset training rounds are reached, to obtain the updated new energy intelligent agent control strategy and flexible load intelligent agent control strategy.
[0014] In another aspect of the present invention, preferably, the coordinated adjustment based on the updated control strategy by each agent outputting adjustment actions includes: During the real-time operation phase, each new energy intelligent agent, based on the updated control strategy, outputs the corresponding active power setpoint according to the current first state observation space, and performs active power adjustment according to the active power setpoint. Each flexible load agent, based on the updated control strategy, outputs the corresponding power adjustment amplitude according to the current second state observation space, and performs load power adjustment according to the power adjustment amplitude.
[0015] (III) Beneficial Effects The above-described technical solution of the present invention has the following beneficial technical effects: This invention constructs distributed new energy units and flexible load units as multi-agent entities and uses reinforcement learning for collaborative training to achieve coordinated control of new energy output and flexible load regulation. In the multi-agent reinforcement learning training, a collaborative sharing deviation penalty term is used to provide a collaborative constraint signal, enabling the new energy agent and the flexible load agent to simultaneously consider the overall regulation task sharing relationship when updating their strategies. This avoids the collaborative imbalance problem caused by each agent optimizing independently based solely on its own rewards. By penalizing the deviation between the actual sharing factor and the expected collaborative sharing factor, the invention helps guide the multi-agent strategy towards a more reasonable collaborative regulation direction, improving the learning stability and convergence efficiency of the multi-agent collaborative control strategy. Attached Figure Description
[0016] Figure 1 This is an overall flowchart of one embodiment of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.
[0018] Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0019] In the description of this invention, it should be noted that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0020] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0021] Example 1 A distributed collaborative control method for new energy sources and flexible loads based on multi-agent reinforcement learning. Figure 1 An overall flowchart of one embodiment of the present invention is shown, as follows: Figure 1 As shown, it includes: The controllers and algorithms of distributed renewable energy units and flexible load units are used to construct renewable energy agents and flexible load agents respectively. Distributed renewable energy units and flexible load units in the power grid are abstracted as independent agents. Distributed renewable energy units include photovoltaic power generation units, wind power generation units, or other renewable energy generation units with adjustable output capabilities; flexible load units include adjustable power electric vehicle charging loads, air conditioning loads, energy storage loads, or industrial interruptible loads. Each renewable energy agent includes at least one renewable energy unit, and each flexible load agent includes at least one flexible load unit. Each agent is an independent control unit capable of autonomously sensing environmental conditions, executing decision-making actions, and learning from the results, thus constructing a multi-agent collaborative control environment. Agents can be constructed through regional division, with all renewable energy units in each region belonging to one renewable energy agent and all flexible load units in each region belonging to one flexible load agent, or through a power system simulation platform. In this embodiment, the new energy intelligent agent includes a first action space and a first state observation space. The first action space includes an active power output setpoint, and the first state observation space includes: adjustable output capacity, current active power output, grid demand information, and operational physical constraints. The adjustable output capacity represents the maximum adjustable output capacity that the new energy power generation unit can provide at the current moment, which can be determined based on current meteorological conditions, equipment operating status, and historical output levels. The current active power output represents the actual output power of the new energy power generation unit within the current scheduling cycle. The grid demand information represents the current power balance demand on the grid side, which may include current load demand, system power deficit, or adjustment demand signals. The operational physical constraints represent the physical limitations that the new energy power generation equipment needs to meet during operation, such as temperature limits, maximum output limits, minimum output limits, and output ramp-up rate limits.
[0022] The flexible load agent includes a second action space and a second state observation space. The second action space includes the power adjustment amplitude, i.e., the power increase or decrease. The power adjustment amplitude can be positive or negative, where a positive value indicates an increase in load power and a negative value indicates a decrease in load power. Its range is determined by the adjustability of the flexible load and user-side constraints. The second state observation space includes: adjustable capacity, current load status, grid demand information, and comfort constraints. The adjustable capacity represents the maximum power adjustment amplitude that the flexible load can make in the current time period, and its size can be determined according to the equipment operating status, load type, and user habits. The current load status describes the actual power level or operating status of the flexible load at the current moment, such as the current power consumption, equipment operating level, or working mode. The grid demand information reflects the power adjustment demand of the grid side and is consistent with the grid demand information obtained by the new energy agent, used to guide the flexible load to participate in power balance adjustment. The comfort constraints describe the limits allowed by the user side for load adjustment behavior, such as the allowable range of indoor temperature changes, equipment operating time limits, or power adjustment limits.
[0023] Based on the current control strategy, the output adjustment actions of each agent are controlled to obtain the actual sharing factor of each agent, including: The adjustable power output capacity, current active power output, operational physical constraints, and grid demand information are used as inputs to the current control strategy of each new energy intelligent entity. The adjustable power output capacity, current active power output, operational physical constraints, and grid demand information can be obtained through the first state observation space. Based on the current control strategy corresponding to each new energy intelligent entity, the active power output setpoint is output. The active power output setpoint is the target active power output value of the new energy unit under each new energy intelligent entity in the next time period. The current control strategy of the new energy intelligent entity may include power allocation strategy, output tracking strategy, or fluctuation suppression strategy, etc. The active power output setpoint is the actual active power adjustment amount of each new energy intelligent entity. The adjustable capacity, current load status, and grid demand information are used as inputs to the current control strategy of each flexible load agent. The adjustable capacity, current load status, and grid demand information can be obtained through a second state observation space. The output is the power adjustment amplitude; the power adjustment amplitude represents the increase or decrease in the flexible load relative to the current load level. The current control strategy of the flexible load agent may include load response decision strategy, load adjustment and allocation strategy, demand response strategy, etc.; the power adjustment amplitude is the actual load power adjustment amount of the flexible load agent. Based on the actual active power regulation and the actual load power regulation, the actual sharing factor of each agent is obtained. According to the proportion of the power regulation actually undertaken by each agent to the total regulation of all agents, the actual sharing factor of each agent in the current control cycle is calculated. The actual sharing factor represents the actual degree of responsibility of each agent in the overall system regulation task.
[0024] Based on the deviation between the actual sharing factor and the expected collaborative sharing factor, a collaborative sharing deviation penalty term is constructed; wherein, the expected collaborative sharing factor is obtained through the following method: Within the current time period, the adjustable output capacity of each new energy intelligent body is obtained based on the first state observation space, and the adjustable capacity of each flexible load intelligent body is obtained based on the second state observation space. The adjustable output capacity represents the maximum adjustable active power output range that the new energy unit can use to participate in power regulation under the current operating state and equipment constraints. The adjustable capacity represents the maximum adjustable power range that the flexible load can participate in power regulation under the conditions of meeting user comfort constraints and load operation constraints. Based on the adjustable output capacity and adjustable capacity, the capacity sharing coefficient of each intelligent agent is determined, and the capacity sharing coefficient is the expected collaborative sharing factor. The adjustable output capacity of each new energy intelligent agent and the adjustable capacity of each flexible load intelligent agent are statistically analyzed to obtain the total adjustable capacity. Based on the proportion of the adjustable capacity of each intelligent agent in the total adjustable capacity, the capacity sharing coefficient of each intelligent agent is determined. The capacity sharing coefficient represents the theoretically required power adjustment ratio of each intelligent agent in the collaborative adjustment task under the current system adjustable capacity distribution conditions, and is used as the expected collaborative sharing factor. Specifically, it can be calculated using the following formula: in, This represents the adjustable power output capacity of the i-th new energy intelligent agent. This represents the adjustable capacity of the j-th flexible load agent. Indicates the total adjustable capacity. This represents the capacity sharing coefficient of new energy intelligent agents. This represents the capacity sharing coefficient of the flexible load agent.
[0025] Based on the deviation between the actual sharing factor and the expected collaborative sharing factor, a collaborative sharing deviation penalty term is constructed, including: Calculate the sharing deviation between the actual sharing factor and the corresponding expected collaborative sharing factor for each agent; square each sharing deviation and sum them by weights to obtain the collaborative sharing deviation, specifically expressed by the following formula: Where J represents the collaborative sharing of deviation, This represents the weight coefficient of the i-th new energy intelligent agent. This represents the weighting coefficient of the j-th flexible load agent. This represents the actual contribution factor of the i-th new energy intelligent agent. This represents the actual load-sharing factor of the j-th flexible load agent. This indicates that new energy intelligent agents expect to collaboratively share the burden of factors. This represents the expected collaborative sharing factor of the flexible load agent. This represents the total number of new energy intelligent entities. This represents the total number of flexible load agents.
[0026] To enhance the constraint effect on larger deviations, the shared deviation values of each agent are squared, so that agents with larger deviations are penalized more significantly. Simultaneously, corresponding weights are set according to the importance of different types of agents in system regulation, and the squared deviations of each agent are weighted and summed to obtain the overall collaborative shared deviation of the system. The weights here can be set according to the capacity, such as by calculating using the following formula: in, This represents the weight coefficient of the i-th new energy intelligent agent. This represents the weighting coefficient of the j-th flexible load agent. This represents the adjustable power output capacity of the i-th new energy intelligent agent. This represents the adjustable capacity of the j-th flexible load agent.
[0027] Based on the system's collaborative sharing deviation, a collaborative sharing deviation penalty term is constructed and introduced into the reward function of multi-agent reinforcement learning to form the collaborative sharing deviation penalty term.
[0028] A reward function is constructed, which includes a reward function for the new energy intelligent agent, a reward function for the flexible load intelligent agent, and a collaborative sharing deviation penalty term. The collaborative sharing deviation penalty term applies to both the new energy intelligent agent and the flexible load intelligent agent. The reward function for the new energy intelligent agent includes a new energy output tracking reward term and an operational constraint penalty term. The reward function for the flexible load intelligent agent includes a load adjustment response reward term and a comfort constraint penalty term.
[0029] The new energy output tracking reward is constructed based on the power deviation between the actual active power adjustment of the new energy intelligent agent and the grid demand information. The closer the actual active power adjustment of the new energy intelligent agent is to the target adjustment level indicated by the grid demand information, the smaller the deviation and the larger the reward value. This encourages new energy units to actively participate in grid power balance regulation, as expressed by the following formula: in, Let k1 represent the output tracking reward item for the i-th new energy intelligent agent, and k1 represent the weight coefficient of the output tracking reward item. This represents the actual regulation power of the i-th new energy intelligent agent. This represents the target regulation power for the grid demand allocation of the i-th new energy intelligent agent; the smaller the deviation, the greater the reward.
[0030] The operational constraint penalty term is constructed based on the constraint violation amount between the operating state of the new energy intelligent agent and the preset physical constraints, such as the output change rate. When the actual operating state of the new energy intelligent agent exceeds the range of the physical constraints, a penalty value is generated to avoid overload operation, as expressed by the following formula: in, Let k1 represent the operational constraint penalty term for the i-th new energy intelligent agent, and k2 represent the weight coefficient of the operational constraint penalty term. This represents the amount of violation of the operational constraints of the i-th new energy intelligent agent.
[0031] The load regulation response reward is constructed based on the response deviation between the actual load power regulation amount of the flexible load agent and the grid demand information. When the regulation direction and magnitude of the flexible load can effectively respond to the grid demand, the higher the degree of response, the greater the reward value, thus incentivizing the flexible load to actively participate in power regulation when the system needs it, as expressed by the following formula: in, Let k3 represent the load regulation response reward term for the j-th flexible load agent, and k3 represent the weight coefficient of the load regulation response reward term. This represents the actual power adjustment of the j-th flexible load agent. This represents the target power adjustment amount of the j-th flexible load agent. This represents the maximum allowable power adjustment range for the j-th flexible load agent.
[0032] The comfort constraint penalty term is constructed based on the constraint violation amount between the power adjustment range of the flexible load agent and the preset comfort constraint range. A penalty is applied when the power adjustment range exceeds the preset comfort constraint range. The comfort constraint penalty term is used to limit the adverse impact of flexible load regulation on user experience. By comparing the power adjustment range of the flexible load agent with the preset comfort constraint range, when the power adjustment range exceeds the user-acceptable comfort constraint range, the constraint violation amount is calculated according to the degree of exceedance, and a corresponding penalty value is applied to ensure that the flexible load still meets basic user comfort requirements when participating in grid regulation, as expressed by the following formula: in, This represents the comfort constraint penalty term, and k4 represents the weight coefficient of the comfort constraint penalty term. This indicates the maximum allowable adjustment range for comfort.
[0033] The collaborative sharing deviation penalty term is constructed based on the deviation between the actual sharing factor and the expected collaborative sharing factor of each agent, and applies to both the new energy agent and the flexible load agent. This ensures that each agent considers not only itself when updating its strategy, but also the overall collaborative adjustment effect, thereby promoting the formation of a reasonable power sharing relationship between the new energy unit and the flexible load.
[0034] The reward function is expressed by the following formula: in, Represents the reward function of new energy intelligent agents, Let λ represent the reward function of the flexible load agent, λ represent the weight coefficient of the collaborative sharing deviation penalty term, and J represent the collaborative sharing deviation amount.
[0035] Furthermore, the weighting coefficients for collaborative sharing of deviation penalties, comfort constraint penalties, load regulation response rewards, output tracking rewards, and operational constraint penalties can be set based on experience.
[0036] Based on the aforementioned reward function, a pre-defined multi-agent reinforcement learning algorithm is used for iterative training to update the control strategies of each agent. A centralized training-distributed execution framework is configured, such as using a multi-agent deep deterministic policy gradient algorithm, a proximal policy optimization algorithm, or a multi-agent soft actor critic algorithm. During the training phase, global state information is used to centrally optimize the policy networks of each agent to improve training stability and collaborative learning capabilities. During the execution phase, each new energy agent and flexible load agent independently outputs adjustment actions based solely on their respective state observation spaces, thereby achieving distributed decision-making and collaborative control. This approach balances the efficiency of multi-agent collaborative learning with the distributed control requirements of actual power grid operation.
[0037] Based on the reward function, a preset multi-agent reinforcement learning algorithm is used for iterative training to update the control policies of each agent, including: Within the current time period, each renewable energy agent and flexible load agent inputs its state observation space into the current strategy network and outputs corresponding adjustment actions. Each renewable energy agent first obtains current state information through a first state observation space, including adjustable output capacity, current active power output, grid demand information, and operational physical constraints. Each flexible load agent obtains corresponding state information through a second state observation space, including adjustable capacity, current load status, grid demand information, and comfort constraints. Each agent inputs its corresponding state observation information into its current strategy network, which calculates the adjustment actions for the current time period. Specifically, the renewable energy agent outputs the active power output setpoint, and the flexible load agent outputs the power adjustment range. These adjustment actions reflect the response strategies of each agent to grid regulation tasks under the current operating state.
[0038] The reward and penalty values for each agent are calculated based on the adjustment actions and the grid operating status. The state observations, adjustment actions, reward values, penalty values, and state observations for the next time period constitute an agent interaction sample. After obtaining the adjustment actions of each agent, these actions are applied to the grid operating environment to update the operating status of new energy units and flexible loads, thereby obtaining the corresponding actual active power adjustment and actual load power adjustment. Simultaneously, based on the grid operating status, the adjustment results of each agent, and the pre-constructed reward function, the new energy output tracking reward and operating constraint penalty for each new energy agent are calculated, as well as the load adjustment response reward and comfort constraint penalty for each flexible load agent. Combined with the collaborative sharing deviation penalty, the reward and penalty values for each agent in the current time period are obtained. This method allows for a comprehensive evaluation of the impact of agent adjustment behavior on grid power balance, equipment operating constraints, and user comfort. The state observation information, executed adjustment actions, obtained reward and penalty values, and state observation information for the next time period of each agent are combined to form a set of agent interaction samples between the agent and the grid environment. Interaction samples reflect the state transition relationships of the agent under the current policy and their corresponding rewards, and are stored in the experience sample set.
[0039] The policy network parameters of each agent are updated based on the agent interaction samples; the policy network parameters of each agent are updated through a loss function based on the agent interaction samples and a preset multi-agent reinforcement learning algorithm.
[0040] Repeat the above training process until the reward function converges or the preset training rounds are reached, to obtain the updated new energy intelligent agent control strategy and flexible load intelligent agent control strategy.
[0041] Based on the updated control policy, each agent outputs adjustment actions to perform coordinated adjustment, including: During the real-time operation phase, each new energy intelligent agent, based on the updated control strategy, outputs the corresponding active power setpoint according to the current first state observation space, and performs active power adjustment according to the active power setpoint. Each flexible load agent, based on the updated control strategy, outputs the corresponding power adjustment amplitude according to the current second state observation space, and performs load power adjustment according to the power adjustment amplitude.
[0042] Furthermore, training can be conducted periodically or conditionally.
[0043] This embodiment constructs distributed new energy units and flexible load units as multi-agents and uses reinforcement learning for collaborative training to achieve coordinated control of new energy output and flexible load regulation. In the multi-agent reinforcement learning training, a collaborative sharing deviation penalty term is used to provide a collaborative constraint signal, so that the new energy agent and the flexible load agent simultaneously consider the overall regulation task sharing relationship when updating the strategy. This avoids the collaborative imbalance problem caused by each agent optimizing independently based on its own reward. By penalizing the deviation between the actual sharing factor and the expected collaborative sharing factor, it helps guide the multi-agent strategy to converge towards a more reasonable collaborative regulation direction, improving the learning stability and convergence efficiency of the multi-agent collaborative control strategy.
[0044] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.
[0045] The present invention has been described above with reference to embodiments thereof. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. The scope of the invention is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.
[0046] Although embodiments of the present invention have been described in detail, it should be understood that various changes, substitutions, and modifications can be made to the embodiments of the present invention without departing from the spirit and scope of the invention.
[0047] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A distributed collaborative control method for new energy sources and flexible loads based on multi-agent reinforcement learning, characterized in that, include: The controllers and algorithms of the distributed new energy units and flexible load units are respectively used to construct new energy intelligent agents and flexible load intelligent agents; Based on the current control strategy, control each agent to output adjustment actions, and obtain the actual sharing factor of each agent. Based on the deviation between the actual sharing factor and the expected collaborative sharing factor, a collaborative sharing deviation penalty term is constructed; A reward function is constructed, which includes a reward function for a new energy intelligent agent, a reward function for a flexible load intelligent agent, and a collaborative sharing deviation penalty term. The collaborative sharing deviation penalty term applies to both the new energy intelligent agent and the flexible load intelligent agent. Based on the reward function, a preset multi-agent reinforcement learning algorithm is used for iterative training to update the control strategies of each agent. Based on the updated control strategy, each agent outputs adjustment actions to achieve coordinated adjustment.
2. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 1, characterized in that, The new energy intelligent agent includes a first action space and a first state observation space. The first action space includes an active power output setpoint, and the first state observation space includes: adjustable output capacity, current active power output, grid demand information, and operating physical constraints. The flexible load agent includes a second action space and a second state observation space. The second action space includes the power adjustment range. The second state observation space includes: adjustable capacity, current load status, grid demand information, and comfort constraints.
3. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 2, characterized in that, The expected collaborative sharing factor is obtained through the following method: Within the current time period, the adjustable output capacity of each new energy intelligent agent is obtained based on the first state observation space, and the adjustable capacity of each flexible load intelligent agent is obtained based on the second state observation space. Based on the adjustable output capacity and adjustable capacity, the capacity sharing coefficient of each intelligent agent is determined, and the capacity sharing coefficient is the expected collaborative sharing factor.
4. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 3, characterized in that, The process of controlling the output adjustment actions of each agent based on the current control strategy to obtain the actual sharing factor of each agent includes: The adjustable output capacity, current active power output, operating physical constraints and grid demand information are used as inputs to the current control strategy of each new energy intelligent body, and the active power output set value is output. The active power output set value is the actual active power adjustment amount of each new energy intelligent body. The adjustable capacity, current load status, comfort constraints, and grid demand information are used as inputs to the current control strategy of each flexible load agent, and the output power adjustment range is the actual load power adjustment amount of the flexible load agent. Based on the actual active power adjustment and the actual load power adjustment, the actual sharing factor of each agent is obtained.
5. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 4, characterized in that, The collaborative sharing deviation penalty term, constructed based on the deviation between the actual sharing factor and the expected collaborative sharing factor, includes: Calculate the sharing deviation between the actual sharing factor and the corresponding expected collaborative sharing factor of each agent; The shared deviation values are squared and then weighted and summed to obtain the collaborative shared deviation amount. A collaborative deviation penalty term is constructed based on the system's collaborative deviation amount.
6. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 5, characterized in that, The reward function for the new energy intelligent agent includes: a new energy output tracking reward item and an operational constraint penalty item; The reward function of the flexible load agent includes: a load adjustment response reward term and a comfort constraint penalty term.
7. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 6, characterized in that, The new energy output tracking reward item is constructed based on the power deviation between the actual active power adjustment of the new energy intelligent agent and the grid demand information. The smaller the deviation, the larger the reward value. The operational constraint penalty item is constructed based on the constraint violation amount between the operational state of the new energy intelligent agent and the preset operational physical constraints. A penalty value is generated when the actual operational state of the new energy intelligent agent exceeds the range of the operational physical constraints.
8. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 7, characterized in that, The load regulation response reward is constructed based on the response deviation between the actual load power regulation of the flexible load agent and the grid demand information. The higher the response level, the greater the reward value. The comfort constraint penalty term is constructed based on the constraint violation amount between the power adjustment range of the flexible load agent and the preset comfort constraint range. A penalty is applied when the power adjustment range exceeds the preset comfort constraint range.
9. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 8, characterized in that, The iterative training based on the reward function using a preset multi-agent reinforcement learning algorithm to update the control policies of each agent includes: Within the current time period, each new energy intelligent agent and flexible load intelligent agent inputs the current policy network based on their respective state observation space and outputs the corresponding adjustment action; The reward and penalty values for each agent are calculated based on the adjustment actions and the power grid operating status. The state observations, adjustment actions, reward values, penalty values, and state observations for the next time period are used to form agent interaction samples. Update the policy network parameters of each agent based on the agent interaction samples; Repeat the above training process until the reward function converges or the preset training rounds are reached, to obtain the updated new energy intelligent agent control strategy and flexible load intelligent agent control strategy.
10. The distributed new energy and flexible load collaborative control method based on multi-agent reinforcement learning according to claim 8, characterized in that, The coordinated adjustment based on the updated control strategy, whereby each agent outputs adjustment actions, includes: During the real-time operation phase, each new energy intelligent agent, based on the updated control strategy, outputs the corresponding active power setpoint according to the current first state observation space, and performs active power adjustment according to the active power setpoint. Each flexible load agent, based on the updated control strategy, outputs the corresponding power adjustment amplitude according to the current second state observation space, and performs load power adjustment according to the power adjustment amplitude.