Network construction type converter secondary frequency modulation method based on multi-agent deep reinforcement learning
Through the secondary frequency regulation method of grid-type converter based on multi-agent deep reinforcement learning, the traditional GFC frequency control method solves the problems of low frequency regulation accuracy, slow response speed and poor adaptability when the renewable energy volatility and grid scale expansion, and achieves more efficient grid frequency regulation.
Patent Information
- Application Number
- CN202510303149.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
When traditional GFC frequency control methods face renewable energy volatility, intermittentity and expansion of power grid scale, there are problems such as low frequency modulation accuracy, slow response speed and poor adaptability.
The secondary frequency modulation method of network-type converter based on multi-agent deep reinforcement learning is adopted, and the independent learning and collaborative decision-making of the agent is realized by constructing a GFC dynamic model, Markov game process and a multi-agent deep reinforcement learning framework.
It significantly improves the performance of GFC secondary frequency modulation, improves frequency modulation accuracy, response speed and adaptability, and enhances the fault tolerance and scalability of the system.
Smart Images

Figure CN120150179A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent frequency modulation of converters, and particularly relates to a secondary frequency modulation method for a grid-forming converter based on multi-agent deep reinforcement learning. Background Technique
[0002] With the rapid development of renewable energy, the power system is undergoing a transformation process from traditional centralized power generation to distributed power generation and smart grid. Renewable energy such as wind energy and solar energy brings great challenges to the frequency regulation and stability of the power grid due to its intermittency and volatility. In this context, grid-forming converters, as an important power conversion device, are widely used in fields such as wind power generation and photovoltaic power generation. Their frequency modulation function has grid friendliness and gradually becomes an important means to ensure the stability of the power grid. For example, the prior art performs frequency modulation control on the power grid system by adjusting control parameters and an optimized control strategy for pre-constructed adaptive inertia and damping parameters.
[0003] Secondary frequency modulation is a key link in the frequency modulation process of the power system. Its goal is to restore the grid frequency to the normal value by controlling the output power of devices such as generators. Usually, the secondary frequency modulation strategies of traditional generators rely on methods such as PID control and model predictive control, and these methods have achieved certain results in terms of stability and accuracy. However, with the expansion of the power grid scale and the increase in the proportion of renewable energy, the traditional GFC frequency control faces problems such as low frequency modulation accuracy, slow response speed, and poor adaptability. Therefore, the present invention proposes a secondary frequency modulation method for a grid-forming converter based on multi-agent deep reinforcement learning. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes a secondary frequency modulation method for a grid-forming converter based on multi-agent deep reinforcement learning to solve the problems existing in the above prior art.
[0005] To achieve the above object, the present invention provides a secondary frequency modulation method for a grid-forming converter based on multi-agent deep reinforcement learning, including:
[0006] Construct a GFC dynamic model based on the control logic and physical model of GFC;
[0007] Construct a Markov game process based on the GFC dynamic model;
[0008] Construct a multi-agent deep reinforcement learning framework based on the Markov game process;
[0009] Train the agents based on the multi-agent deep reinforcement learning framework, and update the policy parameters of the agents to obtain the GFC secondary frequency modulation strategy.
[0010] Optionally, the process of constructing a GFC dynamic model based on the GFC control logic and physical model includes:
[0011] Construct the active power-frequency control link of the GFC based on the GFC control logic;
[0012] Convert the physical model of the GFC into the power angle of the GFC by calculating the access power;
[0013] Calculate the GFC dynamic model based on the active power-frequency control link of the GFC and the power angle of the GFC.
[0014] Optionally, the expression of the power angle of the GFC is:
[0015]
[0016] In the formula, δ(t) is the power angle of the GFC at time t, ω and ω ref represent the output frequency of the GFC and the grid reference frequency respectively.
[0017] Optionally, the process of calculating the GFC dynamic model based on the active power-frequency control link of the GFC and the power angle of the GFC includes:
[0018] Convert the physical model of the GFC into a small-signal form to obtain a small-signal of the active power;
[0019] Obtain the GFC dynamic model based on the small-signal of the active power, the active power-frequency control link of the GFC and the power angle of the GFC.
[0020] Optionally, the expression of the GFC dynamic model is:
[0021]
[0022] In the formula, Δω represents the small-signal of the frequency, ΔP represents the small-signal of the active power, J represents the inertia coefficient; D represents the droop coefficient, E p represents the maximum voltage value at the PCC point, E represents the output electromotive force of the GFC, s represents the independent variable of the signal in the time domain after Laplace transform to the complex frequency domain signal, X d represents the equivalent reactance from the GFC to the connection point.
[0023] Optionally, the process of constructing a Markov game process based on the GFC dynamic model includes:
[0024] Partition the distribution network accessed by the GFC to obtain several sub-networks;
[0025] Create corresponding agents based on several of the sub-networks to obtain several agents;
[0026] Define the environment of the agent and define the state and action for each agent to obtain the agent state and agent action. The environment of the agent includes: global state and local state;
[0027] Define the reward function. Based on the reward function, the agent obtains the immediate reward after making corresponding actions based on the environment.
[0028] Optionally, the expression for calculating the immediate reward is:
[0029]
[0030] In the formula, is the instantaneous frequency of the i-th GFC in the sub-network u; m represents the number of sub-networks; ||·|| represents the 1-norm of the vector, R t represents the immediate reward, ω ref represents the grid reference frequency, and n represents the n-th GFC.
[0031] Optionally, the process of constructing a multi-agent deep reinforcement learning framework based on the Markov game process includes:
[0032] Determine the number of agents based on the number of sub-networks in the distribution network partition;
[0033] Configure an action output interface and an environment input interface for each agent;
[0034] The agent transmits the agent action to the GFC controller based on the action output interface;
[0035] The environment input interface receives the frequency information fed back by the GFC controller;
[0036] Construct a communication mechanism between agents based on the action output and environment input of the agents. The agents make collaborative decisions in the Markov game process based on the communication mechanism.
[0037] Optionally, the policy gradient method is used to update the policy function of each agent during the collaborative decision-making process of the agents;
[0038] Among them, the expression for updating the policy function of each agent is:
[0039]
[0040] In the formula, π i (a t |s t ) is the probability that the agent takes action a t in state s t ; θ i is the policy parameter of agent i; J i (θi ) is a function where the agent $i$ follows the policy parameter $\theta$ i ; represents the expectation calculation over all possible trajectories; represents the gradient operator.
[0041] Compared with the prior art, the present invention has the following advantages and technical effects:
[0042] The present invention proposes a secondary frequency regulation strategy for grid-forming converters (GFCs) based on multi-agent deep reinforcement learning (MADRL), aiming to solve the problems of low frequency regulation accuracy, slow response speed, and poor adaptability exposed by traditional frequency regulation methods in the face of the volatility and intermittency of renewable energy and the expansion of the power grid scale. By constructing a dynamic model of GFC and combining the Markov game process and the multi-agent deep reinforcement learning framework, the present invention realizes the autonomous learning and collaborative decision-making of agents, significantly improving the performance of GFC secondary frequency regulation.
[0043] In terms of technical effects, the present invention has the following advantages: First, the construction of the dynamic model provides accurate environmental feedback for the agents, enabling them to make optimal decisions based on real-time states; Second, the introduction of the Markov game process realizes the collaborative optimization among multiple agents, improving the overall frequency regulation effect of the system; Third, the deep reinforcement learning framework endows the agents with adaptive learning ability, enabling them to quickly respond and optimize the frequency regulation strategy in a complex dynamic environment; Finally, the distributed agent architecture reduces the dependence on the central controller, enhancing the fault tolerance and scalability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0045] Figure 1 is the flowchart of the GFC secondary frequency regulation in the embodiment of the present invention;
[0046] Figure 2 is the schematic diagram of the action output and environmental input of the agent in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0048] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] Embodiment 1
[0050] The implementation of the grid-forming converter (GFC) secondary frequency regulation strategy based on multi-agent deep reinforcement learning (MADRL) focuses on constructing a multi-agent deep reinforcement learning architecture and a Markov game process. The flowchart for realizing the GFC secondary frequency regulation is as Figure 1 shown. The present invention realizes the GFC secondary frequency regulation strategy by introducing MADRL, overcomes the limitations of traditional methods, improves the adaptability, response speed and system efficiency of the GFC frequency regulation strategy, and can perform efficient control and optimization in a large-scale distributed power grid.
[0051] As Figure 1 shown, in this embodiment, a method for secondary frequency regulation of a grid-forming converter based on multi-agent deep reinforcement learning is provided, including the following steps:
[0052] Step 1: Construct a dynamic model of the GFC. Based on the control logic and physical model of the GFC, construct a dynamic model of the GFC. The process of constructing a dynamic model of the GFC based on the control logic and physical model of the GFC includes: constructing an active power-frequency control link of the GFC based on the control logic of the GFC; converting the physical model of the GFC into an access power to calculate the power angle of the GFC; calculating the GFC dynamic model based on the active power-frequency control link of the GFC and the power angle of the GFC.
[0053] As a specific implementation manner of this embodiment, it includes:
[0054] Step 1-1: Construct an active power-frequency control link of the GFC.
[0055] According to the differences in the power control links, the grid-forming control can be divided into droop control, virtual synchronous generator control, power synchronization control, etc. The basic idea of the grid-forming control is to achieve frequency regulation by simulating the external characteristics of a synchronous generator. The forward channel of the grid-forming control can be uniformly expressed as:
[0056]
[0057] In the formula: ω and ω ref represent the output frequency of the GFC and the grid reference frequency respectively; E and E ref represent the output electromotive force of the GFC and the grid reference voltage respectively; P and P ref represent the output active power of the GFC and the active power reference value respectively; Q and Qref respectively represent the output reactive power and reactive power reference value of the GFC; J represents the inertia coefficient; D represents the droop coefficient.
[0058] Step 1-2: Derive the dynamic model of the GFC in combination with the physical model of the GFC.
[0059] The process of obtaining the GFC dynamic model based on the active power-frequency control link of the GFC and the power angle of the GFC includes: converting the physical model of the GFC into a small-signal form to obtain the small-signal of the active power; obtaining the GFC dynamic model based on the small-signal of the active power, the active power-frequency control link of the GFC and the power angle of the GFC.
[0060] The physical model of the GFC can be equivalent to an AC voltage source during operation, and the dynamic characteristics of the GFC are also included therein. The physical model formula of the GFC is the access power expression:
[0061]
[0062] In the formula, E p represents the maximum voltage value at the PCC point; is a complex number, representing the impedance from the GFC to the connection point, the real part represents the resistance, and the imaginary part represents the reactance; arg[·], Re[·] and Im[·] respectively represent the argument, real part and imaginary part of the complex number; |·| represents the absolute value of a real number or the modulus of a complex number. δ(t) is the power angle of the GFC at time t, and the calculation method is as follows:
[0063]
[0064] To derive the dynamic model of the GFC, it is necessary to convert the physical model of the GFC into a small-signal form, and the small-signal expression of its active power is:
[0065]
[0066] In the formula, ΔP and Δδ respectively represent the small-signal of the active power and the small-signal of the power angle. Therefore, in combination with the active power-frequency control link of the GFC and the power angle calculation formula, the dynamic model of the GFC is derived as:
[0067]
[0068] In the formula, Δω represents the small-signal of the frequency.
[0069] Step 2: Construct a Markov game process, and construct a Markov game process based on the GFC dynamic model.
[0070] Constructing a Markov game process based on the GFC dynamic model includes: partitioning the distribution network accessed by the GFC to obtain several sub-networks; creating corresponding agents based on the several sub-networks to obtain several agents; defining the environment of the agents and defining states and actions for each agent to obtain agent states and agent actions, where the environment of the agents includes: global state and local state; defining a reward function, and the agent obtains an immediate reward after making corresponding actions based on the environment based on the reward function.
[0071] As a specific implementation manner of this embodiment, it includes:
[0072] Step 2-1: Create agents.
[0073] In the Markov game, each sub-network represents an agent. Therefore, it is necessary to partition the distribution network accessed by the GFC to form sub-networks and create agents based on the number of sub-networks. Define Agent-u as any agent.
[0074] Step 2-2: Define the environment.
[0075] The GFC frequencies in all sub-networks are used as the global observation results; the GFC frequency in a single sub-network is used as the local observation result.
[0076] Step 2-3: Define the agent state.
[0077] The global state S set at time t t includes the states of all agents, that is, the global observation results. For Agent-u, the state s u includes the local observation of the u-th sub-network. There may be n GFCs in the u-th sub-network, so the state of Agent-u is defined as s u =(ω 1 , ω 2 , …, ω n ) T .
[0078] Step 2-4: Define the agent action.
[0079] The global action A set at time t t includes the actions of all agents. For Agent-u, the action a u contains the control variables within the u-th sub-network. When the GFC performs power regulation, the action of Agent-u is defined as a u =(a 1 , a 2 , …, a n ) T . Then, the action of Agent-u will become a control enable and be transmitted to the GFC within the sub-network.
[0080] Step 2-5: Define the agent reward.
[0081] The reward R set at time t t represents the immediate reward obtained by the agent when performing action A t in state S t . All agents share the same reward, that is:
[0082]
[0083] where is the instantaneous frequency of the i-th GFC in sub-network u; m represents the number of sub-networks; ||·|| represents the 1-norm of the vector.
[0084] Therefore, at each time, each agent obtains its corresponding sub-network global state S t , the local observation at time t, and makes action A t based on this local observation. Then, the control calculated according to the surrogate model enables the secondary frequency regulation of the GFC, and at the same time all agents obtain the immediate reward R t .
[0085] Step 3: Construct a multi-agent deep reinforcement learning architecture, and construct a multi-agent deep reinforcement learning framework based on the Markov game process.
[0086] The process of constructing a multi-agent deep reinforcement learning framework based on the Markov game process includes: determining the number of agents according to the number of sub-networks in the distribution network partition; configuring an action output interface and an environment input interface for each agent; the agent transmits the agent action to the GFC controller based on the action output interface; the environment input interface receives the frequency information fed back by the GFC controller; constructing a communication mechanism between agents based on the action output and environment input of the agent, and the agent makes collaborative decisions in the Markov game process based on the communication mechanism.
[0087] Step 3-1: Determine the number of agents.
[0088] The number of agents is determined according to the number of sub-networks in the distribution network partition. The basis for the distribution network partition needs to be determined according to the secondary frequency regulation requirements. Taking the GFC access node as the center, the reciprocal of the node frequency deviation is used as the partition basis, and the tie lines between nodes are used as constraints to achieve sub-network division. If n groups of GFCs are connected to the same node, they are regarded as one sub-network.
[0089] Step 3-2: Action output and environment input of the agent.
[0090] The schematic diagram of the action output and environment input of the agent is as shown in Figure 2as shown
[0091] All agents are carried in the dispatching terminal, and each agent corresponds to a sub-network. After encoding the GFC frequency control output, the GFC controller uses it as the environmental value encoding. After decoding at the dispatching center, it is used as the agent state and the reward value is calculated. The agent encodes according to the action, and then the GFC controller decodes it and inputs it into the tracking differentiator to form an analog quantity, and finally compensates it to the frequency output end of the GFC primary frequency modulation control.
[0092] Step 3-3: Update the agent.
[0093] Based on the reward obtained by the agent after executing the action and Q-learning in MADRL, the update rule of the Q value is:
[0094]
[0095] In the formula, Q t (S t , A t ) is the Q value of the agent taking action A t under state S t ; α is the learning rate; γ is the discount factor; the subscript t represents the time; a' i represents the maximum action of the i-th agent.
[0096] In multi-agent deep reinforcement learning, the policy gradient method can be used to update the policy function π i (a t |s t ) of each agent, so that the policy maximizes the expected reward.
[0097]
[0098] In the formula, π i (a t |s t ) is the probability of the agent taking action a t under state s t ; θ i is the policy parameter of agent i; J i (θ i ) is the function that agent i obeys the policy parameter θ i ; represents the expected calculation for all possible trajectories; represents the gradient operator.
[0099] Step 4: Model training. Based on the multi-agent deep reinforcement learning framework, train the agent and update the policy parameters of the agent to obtain the GFC secondary frequency modulation policy.
[0100] The goal of model training is to enable the agent to obtain the maximum reward in the environment. The gradient descent strategy is adopted to update the policy parameters of each agent. Usually, a small learning rate α is used to update the parameters (or parameter groups) at the t-th iteration, that is:
[0101]
[0102] where θ t represents the parameters (or parameter groups) to be updated at the t-th iteration. After the model training reaches a certain number of training steps, the reward value will be evaluated to prevent the divergence of the reward value.
[0103] The GFC secondary frequency modulation strategy based on MADRL has the following remarkable advantages:
[0104] Strong dynamic adaptability: Existing GFC frequency modulation strategies mostly rely on the primary frequency modulation control methods that simulate the external characteristics of synchronous machines. Although these methods perform well under static conditions, they are prone to response lags or over-adjustments when dealing with system dynamic changes, load fluctuations, and system disturbances. By adopting MADRL to achieve autonomous learning and adapt to different working conditions and environmental changes, the flexibility and real-time performance of the frequency modulation strategy are improved.
[0105] Integrated multi-agent system optimization: In the method of the present invention, the MADRL architecture is adopted, which enables multiple GFCs in the system to perform distributed decision-making during collaborative work and optimize the overall frequency modulation effect of the system. Each agent takes actions according to the environmental state and its own goals, and through Markov games with each other, realizes the globally optimal frequency modulation control.
[0106] Reward mechanism and self-adjustment ability: Traditional GFC control methods usually rely on externally set fixed rules or models for regulation and lack an adaptive adjustment mechanism. By designing a reward mechanism, during the deep reinforcement learning process, the system can self-optimize according to the actual operating conditions. Through trial and error and the feedback of reward signals, the agent continuously optimizes the control strategy, improving the frequency modulation accuracy and response speed.
[0107] Distributed computing and decision-making: Under the architecture of the present invention, each agent can independently execute decision-making tasks, and the computing and decision-making processes of the system do not depend on a central centralized computing unit. This distributed structure can reduce the system's dependence on the central controller, improve the system's fault tolerance and scalability, and realize the plug-and-play function.
[0108] Handling complex non-linear dynamic systems: Traditional frequency modulation techniques usually require designing control strategies based on complex system models, and complex non-linear dynamic systems pose high requirements on traditional methods. The present invention combines physical and data-based approaches to obtain an optimal control strategy through interaction with the environment, thus demonstrating strong adaptability and robustness when facing complex and non-linear problems.
[0109] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A secondary frequency modulation method for a grid-type converter based on multi-agent deep reinforcement learning, characterized in that: The following steps are involved: Construct GFC dynamic model based on GFC control logic and physical model; Constructing a Markov game process based on the GFC dynamic model; Construct a multi-agent deep reinforcement learning framework based on the Markov game process; Based on the multi-agent deep reinforcement learning framework, the agent is trained, and the strategy parameters of the agent are updated to obtain the GFC secondary frequency modulation strategy.
2. The secondary frequency modulation method of a networked converter based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The process of building a GFC dynamic model based on the GFC control logic and physical model includes: Construct the active power-frequency control link of GFC based on the control logic of GFC; Convert the physical model of GFC into access power to calculate the power angle of GFC; The GFC dynamic model is obtained based on the active power-frequency control link of the GFC and the power angle calculation of the GFC.
3. The secondary frequency modulation method of a networked converter based on multi-agent deep reinforcement learning according to claim 2 is characterized in that: The expression of the power angle of the GFC is: Where δ(t) is the power angle of GFC at time t, ω and ω ref Represent the GFC output frequency and grid reference frequency respectively.
4. The secondary frequency modulation method of a networked converter based on multi-agent deep reinforcement learning according to claim 3 is characterized in that: The process of obtaining the GFC dynamic model based on the active power-frequency control link of the GFC and the power angle calculation of the GFC includes: Converting the physical model of the GFC into a small signal form to obtain a small signal of active power; A GFC dynamic model is obtained based on the small signal of the active power, an active power-frequency control link of the GFC and a power angle of the GFC.
5. The secondary frequency modulation method of a networked converter based on multi-agent deep reinforcement learning according to claim 4 is characterized in that: The expression of the GFC dynamic model is: Where Δω represents the frequency signal, ΔP represents the active power signal, J represents the inertia coefficient, D represents the droop coefficient, and E represents the inertia coefficient. p represents the maximum voltage of the PCC point, E represents the output electromotive force of the GFC, s represents the independent variable of the complex frequency domain signal after Laplace transformation of the time domain signal, and X d It represents the equivalent reactance from GFC to the connection point.
6. The secondary frequency modulation method of a networked converter based on multi-agent deep reinforcement learning according to claim 1 is characterized in that: The process of constructing a Markov game based on the GFC dynamic model includes: The distribution network connected to the GFC is partitioned to obtain several sub-networks; Create corresponding intelligent agents based on the plurality of sub-networks to obtain a plurality of intelligent agents; Defining the environment of the agent and defining the state and action for each agent to obtain the agent state and agent action, wherein the environment of the agent includes: a global state and a local state; A reward function is defined, and the agent obtains an immediate reward after taking corresponding actions based on the environment based on the reward function.
7. The secondary frequency modulation method of a networked converter based on multi-agent deep reinforcement learning according to claim 6 is characterized in that: The expression for calculating the instant reward is: In the formula, is the instantaneous frequency of the i-th GFC in subnetwork u; m represents the number of subnetworks; ||·|| represents the 1-norm of the vector, R t represents the immediate reward, ω ref Represents the grid reference frequency and n represents the nth GFC.
8. The method for secondary frequency modulation of a grid-type converter based on multi-agent deep reinforcement learning according to claim 7 is characterized in that: The process of constructing a multi-agent deep reinforcement learning framework based on the Markov game process includes: Determine the number of intelligent agents based on the number of sub-networks of the distribution network partition; Configure action output interface and environment input interface for each agent; The agent transmits the agent action to the GFC controller based on the action output interface; The environmental input interface receives frequency information fed back by the GFC controller; A communication mechanism between intelligent agents is constructed based on the action output of the intelligent agents and the environment input, and the intelligent agents make collaborative decisions in the Markov game process based on the communication mechanism.
9. The method for secondary frequency modulation of a grid-type converter based on multi-agent deep reinforcement learning according to claim 8, characterized in that: The policy gradient method is used to update the policy function of each agent in the collaborative decision-making process of the agents; Among them, the expression for updating the strategy function of each agent is: In the formula, π i (a t |s t ) is the agent in state s t Take action a t The probability of i is the strategy parameter of agent i; J i (θ i ) is the policy parameter θ that agent i obeys i Function of Represents the expected calculation of all possible trajectories; Represents the gradient operator.