Distributed resource autonomous regulation and control method based on multi-agent learning

By modeling distributed resources into multiple agents and using graph neural network to realize collaboration among agents, the problem of insufficient complexity and robustness of traditional power scheduling methods is solved, and the optimization and rapid response capabilities of global resource scheduling are achieved.

CN120109790APending Publication Date: 2025-06-06ZHUHAI JINDAO ENERGY TECH CO LTD

Patent Information

Application Number
CN202510216470.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

When facing large-scale distributed energy, the traditional centralized power scheduling method has too much computing volume, increased complexity, and cannot respond quickly to dynamic changes. The single agent method cannot achieve global optimal scheduling, and its dynamicity and robustness are insufficient.

Method used

The distributed resource autonomous regulation method based on multi-agent reinforcement learning (MARL) is adopted. By modeling distributed resources into multiple agents, each agent makes independent decisions based on local information, and efficient collaboration between agents is achieved through graph neural network to optimize the global goals.

Benefits of technology

Reliance on central control is reduced, system complexity and communication pressure is reduced, global resource scheduling is optimized, and load fluctuations and environmental changes are quickly responded to load fluctuations and environmental changes, improving the adaptability and robustness of the system, and improving the scheduling efficiency, stability and economics of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure SMS_40
    Figure SMS_40
  • Figure SMS_77
    Figure SMS_77
Patent Text Reader

Abstract

The invention discloses a distributed resource autonomous regulation and control method based on multi-agent learning, and relates to the field of intelligent power grids and resource scheduling. According to the method, a multi-agent reinforcement learning (MARL) framework is adopted, distributed resources are modeled into a plurality of agents, each agent makes an independent decision according to local information, and autonomous regulation and control of the resources are achieved. The system optimizes a global target and improves the economy, power balance and robustness of the system through a cooperation and competition mechanism between intelligent agents. In order to improve the adaptability of the system, a dynamic learning mechanism based on communication is provided, and the response capability of the system to real-time requirements and environmental changes is enhanced. Compared with a traditional centralized regulation and control method, the method has the advantages that dependence on central control can be reduced, regulation and control efficiency can be improved, resource allocation can be optimized, and the method has higher dynamic learning and adaptive capacity, shows higher regulation performance in a complex environment and has better economical efficiency, stability and robustness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart grid and resource scheduling, and in particular to a distributed resource autonomous control method based on multi-agent reinforcement learning (MARL). The method is applied to resource management and optimization in power systems, especially in the control of multiple distributed energy systems, energy storage equipment, loads and other resources, aiming to achieve intelligent autonomous scheduling of resources and optimize the economy, stability and robustness of the system. Background Art

[0002] With the rapid development of renewable energy, especially the large-scale access of new energy such as wind power and solar energy, traditional power systems are facing huge challenges. The traditional centralized power dispatching method relies on a central control system, which usually requires that the data of all equipment and resources be aggregated to the center for processing. However, this approach has significant defects when facing large-scale distributed energy. First, the centralized method has too much computational effort in large-scale systems, and the complexity of the system has increased significantly, resulting in an inability to respond quickly to dynamic changes. Secondly, due to the delay in information transmission, the system's real-time response capability is restricted.

[0003] In addition, although the single-agent approach can achieve good results on local problems through techniques such as reinforcement learning, it still has some obvious limitations. When dealing with distributed resource scheduling, a single agent ignores the collaboration and competition between multiple resources. Each agent only optimizes its own local goals and fails to consider global optimization, making it difficult to achieve the optimal scheduling of the overall system. Therefore, the single-agent approach cannot fully exert its advantages in practical applications, especially in large-scale distributed systems.

[0004] Existing technologies also have certain shortcomings when dealing with the dynamics and uncertainties of power systems. Traditional methods have weak response capabilities to real-time demands and are unable to adjust the system status in a timely manner when emergencies such as load fluctuations and equipment failures occur, thus affecting the stability and economy of the power system. These problems have prompted the demand for new dispatching methods, especially distributed dispatching methods that can adapt to environmental changes and have strong robustness and dynamic response capabilities.

[0005] In this context, multi-agent reinforcement learning (MARL) technology has gradually become a promising solution. By modeling each resource in the system as multiple agents, enabling them to make independent decisions based on local information and optimizing global goals through collaboration and competition among agents, the MARL method provides new ideas and solutions to solve the problems faced by traditional methods. Summary of the invention

[0006] In response to the problems existing in the prior art, the present invention provides a distributed resource autonomous control method based on multi-agent reinforcement learning, which solves the problems of complexity and communication delay of centralized scheduling methods, the problem that single-agent methods cannot achieve global optimality, and the problem of insufficient dynamics and robustness of existing scheduling methods.

[0007] The distributed resource autonomous control method based on multi-agent reinforcement learning proposed in the present invention comprises the following steps:

[0008] 1) Distributed resource modeling: Distributed resources are modeled as multiple agents, each of which represents a resource unit (such as distributed power supply, energy storage equipment or load, etc.) and makes scheduling decisions independently based on local information. The state representation of each agent includes active power, reactive power, energy storage status and load demand, and the action representation includes the adjustment of active and reactive power.

[0009] 2) Definition of global optimization objectives: Define global optimization objectives to optimize power balance, economy, and system reliability. Optimize resource allocation by maximizing system economy and minimizing control costs.

[0010] 3) Design of multi-agent reinforcement learning framework. We select the MARL framework of centralized training and decentralized execution (CTDE). In the centralized training phase, global information is used to optimize the agent strategy. In the decentralized execution phase, each agent makes independent decisions based on local information.

[0011] 4) Design of communication mechanism between agents, introducing graph neural network (GNN) to realize information interaction between agents. Through the communication mechanism, agents can share important information with neighboring agents, enhance global collaboration, and improve the overall optimization capability of the system.

[0012] 5) Reward function design. The reward function includes local rewards and global rewards. Local rewards reflect the contribution of each agent to the local goal, and global rewards reflect the optimization of the overall performance of the system.

[0013] 6) Online learning mechanism, which updates model parameters through real-time data, enhances the system's adaptability to environmental changes, and improves the system's robustness.

[0014] Preferably, in the method, the state of the agent is represented as:

[0015]

[0016] In the formula, and Respectively The active and reactive power of each resource, is the state of energy storage, is the current load demand. Preferably, in the method, the action of the agent is expressed as:

[0017]

[0018] In the formula, and are the adjusted active and reactive powers respectively.

[0019] Preferably, in the method, the communication mechanism between the agents is carried out through a graph neural network (GNN):

[0020]

[0021] In the formula, For the The neighbor set of an agent, For the Update information of an agent.

[0022] Preferably, in the method, the reward function includes a local reward and global rewards :

[0023]

[0024]

[0025] In the formula, is the cost weight, To control costs.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] Compared with the traditional centralized and single-agent dispatching methods, the method of the present invention based on multi-agent reinforcement learning reduces the dependence on central control through distributed autonomous dispatching, and alleviates the complexity and communication pressure of the system. Each agent makes decisions independently based on local information, and cooperates with other agents through efficient communication mechanisms to ensure global resource dispatch optimization. In addition, the method of the present invention can quickly respond to load fluctuations and environmental changes, improve the adaptability and robustness of the system, thereby greatly improving the dispatching efficiency, stability and economy of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for the specific embodiments or the description of the prior art. In all the drawings, similar elements or parts are generally identified by similar reference numerals. In the drawings, the elements or parts are not necessarily drawn according to the actual scale.

[0029] Figure 1 It is a flow chart of the technical solution of the present invention. DETAILED DESCRIPTION

[0030] Based on multi-agent reinforcement learning (MARL) technology, this paper proposes a distributed resource autonomous control method for optimizing the scheduling of distributed energy, energy storage equipment, loads and other resources. This method can improve the economy, stability and robustness of the power system, and has strong adaptability, especially in dealing with environmental changes, load fluctuations and emergencies. The specific implementation process of this method will be introduced in detail below.

[0031] S1. Distributed Resource Modeling and Problem Definition

[0032] In the present invention, it is first necessary to model the resources in the power system and abstract each resource unit into an agent. Each agent represents a resource unit, such as distributed power source, energy storage device, load, etc. These agents make independent decisions based on local information during the scheduling process, and achieve global optimization of the system through communication and cooperation between agents.

[0033] Specifically, the state of the agent includes the current active power , reactive power , energy storage state and load demand These status information describe the current working status of the resource unit. The action of each agent is expressed as the adjustment amount of active power and reactive power ( and ), that is, the agent adjusts its own power output to achieve coordination with other resources and ultimately achieve the global optimal goal.

[0034] When modeling the overall system, it is assumed that resource units, and the state information set of each resource unit is , where each Indicates The state of the resource. The action set of the agent is , each action Indicates Power adjustment of resources.

[0035] S2. Objective function and global optimization goal

[0036] In order to achieve the global optimal scheduling of the system, a global optimization goal needs to be defined. This goal not only considers power balance, but also involves multiple factors such as economy and reliability. The global optimization goal can be expressed by the following function:

[0037]

[0038] In the formula, is the regulation cost, which indicates the cost required to adjust the power under the current state; It is the economy of the system, which indicates the economic benefits of the system in its current state; It is a weight coefficient used to adjust the trade-off between regulation cost and economy.

[0039] Control costs Generally includes power generation cost, transmission loss and equipment maintenance cost, etc. It is related to factors such as the degree of load demand satisfaction and electricity price fluctuations.

[0040] S3. Multi-agent reinforcement learning framework design

[0041] The present invention adopts the multi-agent reinforcement learning (MARL) framework for resource scheduling. The MARL framework includes a centralized training and decentralized execution (CTDE) approach, that is, in the training phase, the strategy of the agent is updated through global information, and in the execution phase, the agent makes independent decisions based on local information.

[0042] In the specific implementation, the cooperation and competition between agents are optimized through the following local value function:

[0043]

[0044] In the formula, For the The reward for an agent, is the discount factor, which indicates the importance of the agent to future rewards. This formula reflects the Q-learning algorithm in reinforcement learning, which updates the agent's strategy by maximizing the expected return.

[0045] The global value function of the system is the sum of the local value functions of all agents:

[0046]

[0047] In the formula, is the global value function, indicating that the system is in state and actions Total return under For the The local value function of an agent represents the agent in state and take action Expected return when is the total number of agents in the system; For the Agents at time The state indicates the current power state, load demand and other information of the agent; For the Agents at time step The action taken represents the power adjustment amount for the agent.

[0048] This approach optimizes the strategy of each agent so that the entire system can make optimal decisions at different stages, thereby achieving the global optimization goal.

[0049] S4. Design of communication mechanism between agents

[0050] In order to enhance the collaborative ability of the system, agents need to communicate with each other. Traditional reinforcement learning algorithms usually ignore the interaction between agents, while the present invention achieves efficient communication between agents by introducing graph neural networks (GNNs). GNNs enable each agent to update its own state based on the state information of its neighboring agents, thereby achieving information sharing and collaboration between resources.

[0051] Specifically, the state update of agent 𝑖 is represented by the following graph neural network function:

[0052]

[0053] In the formula, For the The neighbor set of an agent, For the In this way, the agent can receive the status information of neighboring agents in a timely manner and adjust its own decision according to the global cooperation goal.

[0054] S5. Reward Function Design

[0055] In order to guide the intelligent agent to make decisions that meet the global goal, the design of the reward function is crucial. The present invention designs a composite reward mechanism including local rewards and global rewards.

[0056] The local reward of each agent is the difference between its scheduled action and the target, which is defined as:

[0057]

[0058] In the formula, is the desired power output, is the current agent’s power output. This local reward reflects the agent’s gap in reaching the predetermined power output target.

[0059] The global reward combines the local rewards of all agents and the control cost of the system, and is expressed as:

[0060]

[0061] In the formula, is the cost weight, To regulate costs, the global reward ensures that the agent not only pays attention to its own goals, but also takes into account the global optimization goals.

[0062] S6. Online learning mechanism

[0063] In practical applications, the state of the power system and the load demand are constantly changing. Therefore, the present invention also designs an online learning mechanism so that the intelligent agent can dynamically update the strategy according to real-time data. This mechanism enables the system to adjust resource scheduling according to real-time environmental changes to adapt to different load demands and emergencies.

[0064] Specifically, online learning collects the operating data of the power system in real time and feeds this data back to the reinforcement learning algorithm to update the policy network and value network of the agent. For example, the policy network is updated in the following ways:

[0065]

[0066] In the formula, is the learning rate, which controls the step size of parameter update; is the output of the policy network, indicating a given state Next, the agent chooses an action The probability distribution of is the reward value of the current time step; are the parameters of the policy network.

[0067] The value network is updated as follows:

[0068]

[0069] In the formula, are the parameters of the value network, representing the parameters of the model; It is the output of the target value function, which is used to train the value network and represents the desired state-action value; is the output of the current value network, representing the evaluation of the value function given a state-action pair; is the learning rate, which controls the step size of parameter update; is the value function with respect to the parameter The gradient of indicates the update direction of the parameters during the optimization process. This online learning mechanism ensures that the system can adjust the scheduling strategy in time to cope with dynamic power demand and environmental changes.

[0070] This online learning mechanism ensures that the system can adjust the dispatching strategy in a timely manner to cope with dynamic power demand and environmental changes.

[0071] The present invention proposes a solution for autonomous control of distributed resources through a multi-agent reinforcement learning method, which effectively solves the shortcomings of traditional power dispatching methods in complex and dynamic environments. By modeling distributed resources as multiple agents, each agent makes independent decisions based on local information, and realizes efficient collaboration between agents through graph neural networks, the present invention not only optimizes power balance and economy, but also improves the robustness of the system under uncertainty and emergencies. At the same time, through the online learning mechanism, the agent can adapt to real-time environmental changes and dynamically adjust strategies. The present invention has significant advantages in improving the efficiency, stability and adaptability of power systems, especially in large-scale distributed resource management, demonstrating its unique application value.

Claims

1. A distributed resource autonomous control method based on multi-agent learning, characterized in that: The following steps are involved: Step 1: Distributed resource modeling: Distributed resources are modeled as a multi-agent system. Each agent represents a resource unit. The state representation of the agent includes active power Reactive power Energy storage status and load demand Action indication includes active and reactive power adjustment ΔP i and ΔQ i ; Step 2: Define the global optimization goal and establish the objective function to optimize power balance, economy and system reliability; Step 3: Design a multi-agent reinforcement learning algorithm, select a centralized training and decentralized execution framework, and optimize the agent strategy through local value functions and global value functions; Step 4: Design a communication mechanism between agents and realize information interaction between agents through graph neural network (GNN); Step 5: The agent makes autonomous decisions based on local information while ensuring power balance and system stability through global coordination.

2. The distributed resource autonomous control method according to claim 1 is characterized in that: The global optimization goal is to optimize resource allocation by maximizing system economy and minimizing regulation cost.

3. The distributed resource autonomous control method according to claim 1 is characterized in that: The centralized training and decentralized execution framework updates the agent strategy through global information and enables the agent to make independent decisions based on local information.

4. The distributed resource autonomous control method according to claim 1 is characterized in that: The graph neural network improves global collaboration and optimizes resource scheduling efficiency through information interaction between intelligent agents.

5. The distributed resource autonomous control method according to claim 1 is characterized in that: The global coordination in step 5 synchronizes key information of the agents through a communication mechanism and ensures the stability and power balance of the system.

6. The distributed resource autonomous control method according to claim 1, characterized in that: The method includes an online learning mechanism to update model parameters through real-time data to enhance the adaptability and robustness of the system.

7. The distributed resource autonomous control method according to claim 1 is characterized in that: The reward function includes local rewards and global rewards to comprehensively consider the local objectives and global optimization objectives of the agent: In the formula, β is the cost weight.

8. The distributed resource autonomous control method according to claim 1, characterized in that: The method introduces adversarial training to enhance the system's adaptability to uncertainty and environmental changes.

Citation Information

Patent Citations

  • Distributed energy system cluster collaborative optimization method based on multi-agent reinforcement learning

    CN117350423A

  • Distributed energy coordination control method based on multi-agent reinforcement learning

    CN119323301A

Cited By

  • Micro-grid cooperative scheduling method and device

    CN121052610A

  • Cluster power grid dispatching method and system

    CN121189779A

  • A method and system for dispatching a power grid by clusters

    CN121189779B