A multi-agent-based dual-mode converter multi-machine weak grid voltage frequency support control method

CN122533147APending Publication Date: 2026-08-07SOUTHEAST UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-04-28
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本发明所要解决的技术问题是:针对高比例新能源接入的低压配电弱电网,多台双模变流器并联运行时模式协同困难、参数调节不自适应、电压频率支撑能力不足

Benefits of technology

[0089]本发明达到的有益效果:采用分布式控制架构,无中心节点,避免单点故障,系统扩展性强,适配多机并联场景;基于MADDPG无模型控制,自适应捕捉多机非线性耦合特性,无需精确数学模型,适配高比例新能源台区;可根据本地短路比自主平滑切换GFL/VSG控制模式,兼顾强电网功率精度与弱电网电压频率支撑;自适应调节虚拟惯量、阻尼、电流环PI参数,抑制频率电压波动与次同步振荡,提升薄弱节点电压水平;实现多机按容量比例功率分配,协同性好,显著提升低压配电台区弱电网运行稳定性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122533147A_ABST
    Figure CN122533147A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-agent's dual-mode converter multi-machine weak power grid voltage frequency support control method, method includes: constructing the dual-mode converter multi-machine distributed control architecture of physical layer and communication layer cooperation, realize according to capacity proportion power distribution;Based on multi-agent deep reinforcement learning MADDPG, converter mode and parameter optimization are modeled as Markov decision process, adaptive switching and real-time adjustment of parameter are realized by offline training, online execution;Under the scene of dynamic change of power grid intensity and load disturbance, independently complete GFL / VSG control mode smooth switching and virtual inertia, damping, current loop PI parameter adaptive optimization, improve the stability of substation voltage frequency and weak power grid adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of new energy grid connection control technology in low-voltage distribution substations, specifically to a multi-agent dual-mode converter multi-machine weak grid voltage and frequency support control method. Background Technology

[0002] Low-voltage distribution substations, as key links directly facing users at the end of the distribution network, generally have typical characteristics such as high line impedance, complex topology, and dispersed and drastic load distribution. On the one hand, the long power supply radius and small conductor cross-section of the substation result in line impedance much higher than that of medium and high voltage distribution networks, leading to prominent voltage drop and fluctuation problems. On the other hand, the large spatial and temporal differences in residential and industrial loads and the strong randomness of electricity consumption further exacerbate the dynamic fluctuations of voltage and frequency in the substation, posing a severe challenge to the stable operation of the distribution network.

[0003] The large-scale integration of distributed renewable energy sources such as photovoltaics and wind power into low-voltage distribution substations has created a dual disturbance due to the intermittent, random, and volatile nature of renewable energy output and the strong uncertainty of substation loads. This has exacerbated voltage exceedance and frequency deviation issues in substations, making traditional voltage regulation methods inadequate for dynamic control requirements. Simultaneously, the increasing penetration rate of power electronic converters, the core interface for grid connection of distributed renewable energy and energy storage systems, has fundamentally altered the traditional power supply model centered on synchronous generators in distribution networks. The low inertia and weak damping characteristics of power electronic converters lead to a significant decrease in the equivalent rotating inertia of the substation grid, a substantial reduction in short-circuit capacity, and a continuous weakening of grid strength. Under weak grid conditions, the grid connection stability of converters and the system's ability to withstand disturbances decline sharply, easily triggering voltage oscillations and subsynchronous resonance, seriously threatening the power supply security and power quality of substations.

[0004] Currently, transformer control in power distribution areas is mainly divided into two categories: grid-connected GFLs and grid-connected VSGs. Existing solutions cannot fully utilize the complementary advantages of both and cannot adapt to sudden changes in grid strength. When multiple dual-mode transformers are connected in parallel, parameters such as mode, virtual inertia, and damping are difficult to coordinate, which can easily lead to subsynchronous oscillations and threaten the safe and stable operation of the distribution area. Traditional centralized control suffers from single-point failures and poor scalability. Control methods based on physical models are difficult to adapt to the nonlinear coupling characteristics of multiple machines. Therefore, a model-free, adaptive, and distributed multi-machine cooperative control method is needed. Summary of the Invention

[0005] The technical problem to be solved by this invention is that, for low-voltage power distribution networks with a high proportion of new energy access, multiple dual-mode converters operating in parallel face difficulties in mode coordination, non-adaptive parameter adjustment, and insufficient voltage and frequency support capabilities.

[0006] To address the aforementioned technical problems, this invention provides a multi-agent dual-mode converter multi-machine weak grid voltage and frequency support control method, characterized by comprising:

[0007] Step 1: Construct a distributed control architecture, including multiple ordinary nodes that communicate in series, and multiple dual-mode converters that are respectively connected to weak nodes in the ordinary nodes of the low-voltage distribution area line. The weak nodes are line nodes where the far-end power supply radius of the low-voltage distribution area line is greater than a set distance, the line impedance is greater than a set value, and the transient voltage fluctuation amplitude exceeds the limit; any adjacent nodes can send and receive operating status data bidirectionally.

[0008] The dual-mode converter node is a controllable node. An intelligent agent control unit is deployed at the dual-mode converter. The dual-mode converter node collects voltage and load power flow operation data of adjacent ordinary nodes for global transformer area status perception and power allocation decision-making.

[0009] The ordinary nodes are uncontrollable nodes, and they adopt a distributed adjacent node communication mode to perform natural power flow calculations and exchange node information with adjacent dual-mode converter nodes and ordinary nodes.

[0010] Step 2: Model the operation mode and control parameter adjustment problem of the dual-mode converter agent as a MADDPG model. With the goal of minimizing the frequency deviation and voltage deviation of the distribution area, set a reward function that includes frequency and voltage penalties, parameter stability penalties, and mode switching penalties.

[0011] Step 3: In offline mode, the transformer area jointly trains the MADDPG model in the dual-mode converter agent through global information. After training, each dual-mode converter is used for independent decision-making based on local information.

[0012] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0013] In step 1, the two dual-mode converter nodes adopt full-topology adjacent peer-to-peer distributed communication, bidirectionally exchanging power, voltage, and mode status information. Through a consensus protocol, the output power of the two dual-mode converters is allocated according to the controllable capacity ratio. The formulas for calculating power transmission and power exchange are as follows:

[0014] (1)

[0015] in, , They are nodes Local load power, , They are nodes , voltage amplitude, , They are nodes , voltage phase angle, For nodes Total active power injection For nodes Total reactive power injection For nodes With nodes Mutual inductance between lines, For nodes Self-inductance reactance;

[0016] Undirected connected graph The formula is:

[0017] (2)

[0018] Edge set The formula is:

[0019] (3)

[0020] in, Represents a connected node and nodes Undirected edges, adjacency matrix When, explain the node and There is a communication connection between them;

[0021] The formula for the node set V is:

[0022] (4)

[0023] Each node corresponds to a transformer area node, and k is the total number of nodes;

[0024] Communication Neighbor Node Set The formula is:

[0025] (5)

[0026] | indicates a condition separator;

[0027] Adjacency matrix The formula is:

[0028] (6)

[0029] Among them, the adjacency matrix =0 indicates a node and No direct communication or is itself a node;

[0030] The formula for the degree matrix D is:

[0031] (7)

[0032] in, It is a dual-mode converter node Number of neighbors; set of communicating neighbor nodes This is a matrix of numbers for all adjacent converters that can communicate directly.

[0033] Laplace matrix The formula is:

[0034] (8)

[0035] Definition of the first Local interaction information vector of dual-mode converter The formula is:

[0036] (9)

[0037] in, , , , The figures represent the voltage deviation and rate of change of the dual-mode converter in VSG or GFL modes, respectively. For virtual inertia, For virtual damping, This is the proportional gain of the current loop. This is the integral gain of the current loop; For dual-mode converter nodes Local short-circuit ratio, For operation mode, This indicates that the converter is operating in VSG control mode. This indicates that the system is operating in GFL control mode. For the rated capacity of the converter, For capacity weight, , These are the current active power and reactive power, respectively. This is the power deviation. The rate of change of power deviation, It is a continuous-time differential variable.

[0038] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0039] In step 1, a first-order consensus protocol with capacity weight is executed on active power and reactive power, and the power correction amount is output; the GFL control mode uses PQ control to track power commands, and the VSG control mode uses secondary frequency to recover received power commands, so as to realize dual-mode parallel power distribution according to capacity.

[0040] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0041] No. Capacity weighting of dual-mode converters The calculation formula is:

[0042] (10)

[0043] For the first The rated apparent capacity of the dual-mode converter To traverse the first [unit / area] within the [area / area] Rated apparent capacity of the dual-mode converter;

[0044] The formula for calculating the capacity-weighted first-order consensus protocol is:

[0045] (11)

[0046] in, , This is the power correction amount output by the distributed consensus protocol. for Time, Number The actual active power output of the dual-mode converter is currently... for Time, Node The communication neighbor number The actual active power output of the converter is currently... for Time, Number The actual reactive power output of the dual-mode converter is currently... for Time, Node Communication Neighbors The actual reactive power output of the converter is currently... It is a continuous-time differential variable;

[0047] The formula for calculating the power distribution based on capacity when dual-mode converters are connected in parallel is as follows:

[0048] (12)

[0049] The capacity-weighted first-order consensus protocol only requires each dual-mode converter to exchange local state information with its neighboring nodes. Without the participation of a central node, it can ensure that all dual-mode converters can achieve the effect of allocating active and reactive power according to the rated capacity ratio, provided that the communication topology is connected.

[0050] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0051] In step 2, the state dimension of the MADDPG model includes the short-circuit ratio (SCR). When the SCR is greater than the upper threshold, the system switches to GFL control mode. When the SCR is less than the lower threshold, the system switches to VSG control mode. When the SCR is in the middle range between the upper and lower thresholds, the agent in the dual-mode converter node makes autonomous optimization decisions based on the current operating parameters through the reinforcement learning neural network in the pre-trained MADDPG model, and dynamically selects the optimal operating control mode and the optimal parameter ratio.

[0052] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0053] In step 2, in the MADDPG model, the state dimension is set to 10, including voltage deviation, voltage change rate, virtual inertia, damping parameters, PI control parameters, system frequency deviation, frequency change rate, local short-circuit ratio (SCR), and operating mode identifier for VSG / GFL dual modes; the action dimension is set to 5, including virtual inertia adjustment, damping coefficient adjustment, PI proportional parameter adjustment, PI integral parameter adjustment, and grid-connected mode binary switching action.

[0054] The operating parameters include system frequency fluctuations. Voltage deviation in VSG mode Voltage deviation in GFL mode Virtual inertia Virtual damping Current loop proportional gain Current loop integral gain Dual-mode converter local short-circuit ratio Dual-mode converter operation mode .

[0055] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0056] In the MADDPG model, the objective function is calculated as follows:

[0057] (13)

[0058] The core objective is to minimize the time integral of the frequency deviation and voltage deviation of the transformer area. and For the first Virtual inertia and damping stability penalty term of dual-mode converter , These are the weighting coefficients. and For the first Virtual inertia of dual-mode converter and damping The current value, and For the first Virtual inertia of dual-mode converter and damping The rated value; and For the first The proportional gain of the dual-mode converter With integral coefficient Stability penalty item, , These are weighting coefficient one and weighting coefficient two, respectively. and The first Taiwan dual-mode converter and The current value of the coefficient. and The first Taiwan dual-mode converter and The rated value of the coefficient; The mode switching penalty coefficient, For dual-mode converter nodes Local short-circuit ratio; when operating mode At that time, the virtual inertia of the i-th dual-mode converter and damping The penalty item takes effect, and the dual-mode converter operates in VSG control mode; when the operating mode... hour, and With the penalty term applied, the dual-mode converter operates in GFL control mode; The total simulation duration is... For continuous-time differential variables, To enhance the learning of discrete decision step size, To optimize the objective function for multiple objectives;

[0059] The formula for calculating the reward function is:

[0060] (14)

[0061] in, The output value of the reward function represents the state. Execute action Then, transition to the next state. Instant rewards obtained at that time; The penalty weight representing the frequency deviation of the converter. , These represent the penalty weights for voltage deviation in the dual-mode converter, with each converter contributing a reward value based on its own mode. , , , These represent the penalty weights for parameter mutations in the dual-mode converter, respectively. This indicates the penalty weight for converter mode switching. This represents the reward weight when the voltage deviation in dual-mode converter is less than a threshold. , These represent the reward weights when the voltage deviation in dual-mode converter is less than the threshold. For the first Single adjustment step size of virtual inertia in VSG mode of Taiwan converter For the first The single adjustment step size of the damping coefficient in VSG mode of the converter. For the first The single adjustment step size of the PI proportional parameter in GFL mode of the converter. For the first The single adjustment step size of the PI integral parameter in GFL mode of the converter. For absolute value operators, System frequency deviation acceptable threshold The acceptable threshold for bus voltage deviation in VSG mode. The acceptable threshold for bus voltage deviation in GFL mode. This is an indicator function; the condition within the parentheses is 1 if satisfied, and 0 if not satisfied.

[0062] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0063] In step 3, the MADDPG model adopts a centralized training and decentralized execution framework. The policy network uses ReLU and Tanh activation functions to output continuous deterministic actions, and the value network input is a concatenated vector of state and action, which uses the ReLU activation function.

[0064] The aforementioned method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents...

[0065] The policy network is represented as follows:

[0066] (15)

[0067] in, The output of the policy network represents the state. Below, by parameters The optimal consecutive actions determined; The set of learnable parameters for the policy network, including weights 1 Weight 2 Bias 1 Bias 2 Apply the ReLU activation function to the output of the first layer and the tanh activation function to the output of the second layer, then multiply each element by the action scaling factor. and the offset Add;

[0068] Value networks are represented as:

[0069] (16)

[0070] in, The output of the value network represents the state. Execute action The expected long-term cumulative rewards that can be obtained afterward; This is the set of learnable parameters for the value network, including the weight matrix. Weight matrix four Bias three , offset four ;

[0071] The policy network loss function is expressed as:

[0072] (17)

[0073] in, For value network state The following are the actions output by the policy network. Value assessment The number of training rounds is given, and the optimization objective is to minimize the policy network loss. ;

[0074] The value network loss function is expressed as:

[0075] (18)

[0076] in, For the goal value, For prediction Value is the current value network. State Next action Value assessment; The mean squared error loss function of the value network is used to measure the target within a batch of samples. Value and current network prediction Deviation of values;

[0077] Target Q value The expression is derived from the Bellman equation, and the calculation formula is:

[0078] (19)

[0079] in, As a discount factor, and These are the output value network parameters and policy network parameters of the target network, respectively. For the first Instant reward for each sample For the first The next state of the power grid system after the converter corresponding to each sample executes the control action;

[0080] The state space calculation formula for the control strategy is as follows:

[0081] (20)

[0082] The state space S includes the dual-mode frequency and voltage states of each dual-mode converter, VSG and GFL parameters, the local SCR size identified by the dual-mode converter node, and historical operating mode signals. During operation, each dual-mode converter calculates the frequency and voltage states in parallel under both VSG and GFL modes, using only the mode signals... Select one of the output modes to take effect;

[0083] The formula for calculating the action space of the control strategy is:

[0084] (twenty one)

[0085] in, To adjust the step size for inertia, To adjust the damping step size, and Adjust the upper and lower limits for virtual inertia. and Adjust the upper and lower limits for damping; For the first The single adjustment step size of the PI proportional parameter in GFL mode of the converter. For the first The single adjustment step size of the integral parameters of the converter in GFL mode. and Adjust the upper and lower limits of the PI ratio parameter for GFL mode. and Adjust the upper and lower limits of the PI integral parameters for GFL mode respectively.

[0086] The aforementioned method for multi-agent dual-mode converter voltage and frequency support control of weak grid networks also includes:

[0087] Step 4: Online verification of the adaptive switching function of the dual-mode converter. When the grid short-circuit ratio fluctuates, the converter switches control modes. When the short-circuit ratio SCR is greater than the upper limit threshold, it switches to GFL control mode. When the short-circuit ratio SCR is less than the lower limit threshold, it switches to VSG control mode. The intermediate range between the upper and lower limit thresholds is determined autonomously by the trained agent. After the grid short-circuit ratio recovers to above the upper limit threshold, the converter switches back to GFL control mode.

[0088] A computer system includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described above.

[0089] The beneficial effects achieved by this invention are as follows: It adopts a distributed control architecture with no central node, avoiding single points of failure; the system has strong scalability and is adaptable to multi-machine parallel operation scenarios; based on MADDPG model-free control, it adaptively captures the nonlinear coupling characteristics of multiple machines, requiring no precise mathematical model, and is suitable for high-proportion renewable energy distribution areas; it can autonomously and smoothly switch between GFL / VSG control modes according to the local short-circuit ratio, balancing the power accuracy of strong power grids with the voltage and frequency support of weak power grids; it adaptively adjusts virtual inertia, damping, and current loop PI parameters to suppress frequency and voltage fluctuations and subsynchronous oscillations, improving the voltage level of weak nodes; it achieves power allocation among multiple machines according to capacity ratio, with good coordination, significantly improving the operational stability of low-voltage distribution areas in weak power grids. Attached Figure Description

[0090] Figure 1 This is a flowchart of the multi-agent dual-mode converter multi-machine weak grid voltage support control method in Embodiment 1 of the present invention;

[0091] Figure 2 This is a schematic diagram of the distributed control structure of the dual-mode converter in the transformer substation according to Embodiment 1 of the present invention;

[0092] Figure 3 This is a flowchart illustrating the offline training and online execution of the reinforcement learning network in Embodiment 1 of the present invention;

[0093] Figure 4 This is a reward curve diagram of the offline training process in Embodiment 1 of the present invention;

[0094] Figure 5 This is a graph showing the loss curves of the value network and the policy network in Embodiment 1 of the present invention;

[0095] Figures 6(a), 6(b), and 6(c) are schematic diagrams of the power output, current response, and frequency response of the dual-mode converter in Embodiment 1 of the present invention, respectively.

[0096] Figures 7(a) and 7(b) are curves showing the changes in inertia damping and proportional-integral control parameters of the dual-mode converter in Embodiment 1 of the present invention, respectively.

[0097] Figure 8 This is a graph showing the operating modes of the node SCR and the dual-mode converter in Embodiment 1 of the present invention.

[0098] Figure 9 This is a voltage comparison curve of the weak node in the substation area in Embodiment 1 of the present invention. Detailed Implementation

[0099] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings.

[0100] Example 1

[0101] like Figure 1 As shown, this embodiment provides a multi-agent dual-mode converter multi-machine weak grid voltage and frequency support control method, including the following steps:

[0102] Step 1: Construct a distributed control architecture, including a physical layer and a communication layer. The physical layer adopts a multi-node low-voltage distribution area model, including multiple ordinary nodes connected in series. Each dual-mode converter is connected to a weak node in the low-voltage distribution area line, and different dual-mode converters are deployed at weak nodes on different lines within the distribution area. The weak node is a line node in the low-voltage distribution area where the far-end power supply radius is greater than a set distance, the line impedance is greater than a set value, and the transient voltage fluctuation amplitude exceeds the limit.

[0103] The communication layer is a distributed undirected topology with no central master control node. The communication links between nodes are bidirectional and symmetrical, and information exchange is not divided into master and slave directions. Any adjacent node can send and receive operational status data bidirectionally, and there is no constraint of unidirectional command transmission.

[0104] The dual-mode converter nodes are controllable nodes with active power regulation capabilities, capable of participating in the optimization of transformer area voltage and frequency. Each dual-mode converter node can exchange information with the others. An intelligent agent control unit is deployed at each dual-mode converter, enabling bidirectional communication between dual-mode converter nodes and ordinary nodes. The dual-mode converter nodes collect voltage and load flow data from adjacent ordinary nodes for global transformer area status awareness and power allocation decision-making.

[0105] The ordinary nodes are uncontrollable nodes. Uncontrollable nodes have no autonomous power regulation capability. They adopt a distributed neighbor node communication mode, only perform natural power flow calculations and exchange node information with neighboring dual-mode converter nodes and ordinary nodes.

[0106] The two dual-mode converter nodes adopt full-topology adjacent peer-to-peer distributed communication, bidirectionally exchanging power, voltage, and mode status information. Through a consensus protocol, the output power of the two dual-mode converters is allocated according to a controllable capacity ratio. The formulas for calculating power transmission and exchange are as follows:

[0107] (1)

[0108] in, , They are nodes Local load power, , They are nodes , voltage amplitude, , They are nodes , voltage phase angle, For nodes Total active power injection For nodes Total reactive power injection For nodes With nodes Mutual inductance between lines, For nodes Self-inductance reactance.

[0109] Undirected connected graph The formula is:

[0110] (2)

[0111] Edge set The formula is:

[0112] (3)

[0113] in, Represents a connected node and nodes Undirected edges, adjacency matrix When, explain the node and There is a communication connection between them;

[0114] The formula for the node set V is:

[0115] (4)

[0116] Each node corresponds to a transformer area node, and k is the total number of nodes;

[0117] Communication Neighbor Node Set The formula is:

[0118] (5)

[0119] | indicates a condition separator;

[0120] Adjacency matrix The formula is:

[0121] (6)

[0122] Among them, the adjacency matrix =0 indicates a node and No direct communication or is itself a node;

[0123] The formula for the degree matrix D is:

[0124] (7)

[0125] in, It is a dual-mode converter node Number of neighbors; set of communicating neighbor nodes This is a matrix of numbers for all adjacent converters that can communicate directly.

[0126] Laplace matrix The formula is:

[0127] (8)

[0128] Definition of the first Local interaction information vector of dual-mode converter The formula is:

[0129] (9)

[0130] in, , , , The figures represent the voltage deviation and rate of change of the dual-mode converter in VSG or GFL modes, respectively. For virtual inertia, For virtual damping, This is the proportional gain of the current loop. This is the integral gain of the current loop; For dual-mode converter nodes Local short-circuit ratio, For operation mode, This indicates that the converter is operating in VSG control mode. This indicates that the system is operating in GFL control mode. For the rated capacity of the converter, For capacity weight, , These are the current active power and reactive power, respectively. This is the power deviation. The rate of change of power deviation, It is a continuous-time differential variable.

[0131] In step 1, a first-order consensus protocol with capacity weight is executed on active power and reactive power to output power correction. The GFL control mode uses PQ control to accurately track power commands, and the VSG control mode introduces secondary frequency recovery to receive power commands, so as to realize dual-mode parallel power distribution according to capacity.

[0132] No. Capacity weighting of dual-mode converters The calculation formula is:

[0133] (10)

[0134] For the first The rated apparent capacity of the dual-mode converter To traverse the first [unit / area] within the [area / area] Rated apparent capacity of the dual-mode converter;

[0135] The formula for calculating the capacity-weighted first-order consensus protocol is:

[0136] (11)

[0137] in, , This is the power correction amount output by the distributed consensus protocol. for Time, Number The actual active power output of the dual-mode converter is currently... for Time, Node The communication neighbor number The actual active power output of the converter is currently... for Time, Number The actual reactive power output of the dual-mode converter is currently... for Time, Node Communication Neighbors The actual reactive power output of the converter is currently... It is a continuous-time differential variable.

[0138] The formula for calculating the power distribution based on capacity when dual-mode converters are connected in parallel is as follows:

[0139] (12)

[0140] The capacity-weighted first-order consensus protocol only requires each dual-mode converter to exchange local state information with its neighboring nodes. Without the participation of a central node, it can ensure that all dual-mode converters can achieve the effect of allocating active and reactive power according to the rated capacity ratio, provided that the communication topology is connected.

[0141] Step 2, Multi-agent deep deterministic policy gradient model,

[0142] A MADDPG model is constructed, in which the state dimension is set to 10, including voltage deviation, voltage change rate, virtual inertia, damping parameters, PI control parameters, system frequency deviation, frequency change rate, local short-circuit ratio (SCR), and operating mode identifier for VSG / GFL dual modes; and the action dimension is set to 5, including virtual inertia adjustment, damping coefficient adjustment, PI proportional parameter adjustment, PI integral parameter adjustment, and grid-connected mode binary switching action.

[0143] When the short-circuit ratio (SCR) exceeds the upper threshold, the system switches to GFL control mode; when the SCR is less than the lower threshold, it switches to VSG control mode. When the SCR falls within the range between the upper and lower thresholds, the agents in the dual-mode converter node adjust their control based on the current system frequency fluctuations. Voltage deviation in VSG mode Voltage deviation in GFL mode Virtual inertia Virtual damping Current loop proportional gain Current loop integral gain Dual-mode converter local short-circuit ratio Dual-mode converter operation mode Based on operating parameters, the system autonomously selects the optimal operating control mode and parameter ratio through a reinforcement learning neural network in a pre-trained MADDPG model. This approach balances grid stability margin and dynamic response performance, smoothly connects the operating conditions of strong and weak grids, and avoids frequent mode oscillations and switching at threshold critical points.

[0144] The upper threshold is set to 3, and the lower threshold is set to 2.

[0145] When the agent outputs a decision action to switch between GFL / VSG operating modes, the Markov state transition probability is determined by the short-circuit ratio (SCR) value, and the algorithm discount factor is set to 0.95, which can balance speed and stability.

[0146] The reward function includes:

[0147] The frequency penalty weighting coefficient is set to 2.0. The greater the frequency deviation, the lower the reward, ensuring that the system frequency remains stable within the rated range.

[0148] The voltage penalty weighting coefficient is set to 2.0. The greater the voltage deviation from the rated value, the heavier the penalty, in order to maintain the bus voltage quality in compliance with standards.

[0149] The mode switching penalty weight coefficient is set to 1.0. Each time a mode switch is performed, 1 point is deducted to suppress meaningless frequent mode jumps.

[0150] The stability reward weighting coefficient is set to 10.0. Positive rewards are given for maintaining stable operation, which encourages long-term stable operation and avoids fluctuations.

[0151] The formula for calculating the objective function is:

[0152] (13)

[0153] The core objective is to minimize the time integral of the system frequency deviation and voltage deviation. and For the first Virtual inertia and damping stability penalty term of dual-mode converter , These are the weighting coefficients. and For the first Virtual inertia of dual-mode converter and damping The current value, and For the first Virtual inertia of dual-mode converter and damping The rated value. and For the first The proportional gain of the dual-mode converter With integral coefficient Stability penalty item, , These are weighting coefficient one and weighting coefficient two, respectively. and The first Taiwan dual-mode converter and The current value of the coefficient. and For the first Taiwan dual-mode converter and The rated value of the coefficient. This is a mode-switching penalty coefficient to suppress frequent switching of the dual-mode converter's operating modes. For dual-mode converter nodes Local short-circuit ratio. (When operating mode) At that time, the virtual inertia of the i-th dual-mode converter and damping The penalty item takes effect, and the dual-mode converter operates in VSG control mode; when the operating mode... hour, and With the penalty term applied, the dual-mode converter operates in GFL control mode. The total simulation duration is... For continuous-time differential variables, To enhance the learning of discrete decision step size, The objective function is used for multi-objective comprehensive optimization.

[0154] The formula for calculating the reward function is:

[0155] (14)

[0156] in, The output value of the reward function represents the state. Execute action Then, transition to the next state. Instant rewards received at that time. The penalty weight representing the frequency deviation of the converter. , These represent the penalty weights for voltage deviation in the dual-mode converter, with each converter contributing a reward value based on its own mode. , , , These represent the penalty weights for parameter mutations in the dual-mode converter, respectively. This indicates the penalty weight for converter mode switching, which avoids frequent switching of the dual-mode converter's operating mode and parameter oscillation during mode switching. This represents the reward weight when the voltage deviation in dual-mode converter is less than a threshold. , 1 represents the reward weight when the voltage deviation is less than the threshold in the dual mode of the converter, and Ⅱ represents the indicator function. The indicator function is 1 when the frequency or voltage deviation is less than the trigger threshold, and 0 otherwise. For the first Single adjustment step size of virtual inertia in VSG mode of Taiwan converter For the first The single adjustment step size of the damping coefficient in VSG mode of the converter. For the first The single adjustment step size of the PI proportional parameter in GFL mode of the converter. For the first The single adjustment step size of the PI integral parameter in GFL mode of the converter. For absolute value operators, System frequency deviation acceptable threshold The acceptable threshold for bus voltage deviation in VSG mode. The acceptable threshold for bus voltage deviation in GFL mode. This is an indicator function; the condition within the parentheses is 1 if satisfied, and 0 if not satisfied.

[0157] Step 3: The dual-mode converter is trained offline using the MADDPG model. During offline training, the agent of the dual-mode converter is jointly trained using global information. After training, each dual-mode converter makes independent decisions using only local information, which can adapt to the requirements of distributed control. The MADDPG learning rate parameter is set to 3e-4, the discount factor parameter is set to 0.99, the soft update coefficient parameter is set to 0.005, and the training epoch parameter is set to 1000.

[0158] The MADDPG model adopts a centralized training and decentralized execution framework. The policy network uses ReLU and Tanh activation functions to output continuous deterministic actions, and the value network input is a concatenated vector of state and action, using the ReLU activation function. The empirical replay pool breaks the temporal correlation of samples, and soft updates are used on the target network to ensure training stability.

[0159] The policy network is represented as follows:

[0160] (15)

[0161] in, The output of the policy network represents the state. Below, by parameters The optimal consecutive actions are determined. The set of learnable parameters for the policy network, including weights 1 Weight 2 Bias 1 Bias 2 Apply the ReLU activation function to the output of the first layer and the tanh activation function to the output of the second layer, then multiply each element by the action scaling factor. and the offset The values ​​are added together to adjust the baseline values ​​of the actions, ensuring that the final output falls within the parameter range allowed by the engineering.

[0162] Value networks are represented as:

[0163] (16)

[0164] in, The output of the value network represents the state. Execute action Then, the expected long-term cumulative rewards. This is the set of learnable parameters for the value network, including the weight matrix. Weight matrix four Bias three , offset four .

[0165] The policy network loss function is expressed as:

[0166] (17)

[0167] in, For value network state The following are the actions output by the policy network. Value assessment The number of training rounds is given, and the optimization objective is to minimize the policy network loss. .

[0168] The value network loss function is expressed as:

[0169] (18)

[0170] in, For the goal value, For prediction Value is the current value network. State Next action Value assessment. The mean squared error loss function of the value network is used to measure the target within a batch of samples. Value and current network prediction Deviation of values.

[0171] Target Q value The expression is derived from the Bellman equation, and the calculation formula is:

[0172] (19)

[0173] in, This is a discount factor, representing the degree of importance placed on rewards at future moments. and These are the output value network parameters and policy network parameters of the target network, respectively. For the first Instant reward for each sample For the first The next state of the power grid system after the converter corresponding to each sample executes the control action.

[0174] The state space calculation formula for the control strategy is as follows:

[0175] (20)

[0176] The state space S contains the dual-mode frequency and voltage states of each dual-mode converter, VSG and GFL parameters, the local SCR size identified by the dual-mode converter node, and historical operating mode signals. During operation, each dual-mode converter pre-calculates the frequency and voltage states in both VSG and GFL modes in parallel, but only through the mode signals... Select one of the modes to apply the output.

[0177] The formula for calculating the action space of the control strategy is:

[0178] (twenty one)

[0179] in, To adjust the step size for inertia, To adjust the damping step size, and Adjust the upper and lower limits for virtual inertia. and Adjust the upper and lower limits for damping. For the first The single adjustment step size of the PI proportional parameter in GFL mode of the converter. For the first The single adjustment step size of the integral parameters of the converter in GFL mode. and Adjust the upper and lower limits of the PI ratio parameter for GFL mode. and Adjust the upper and lower limits of the PI integral parameters for GFL mode respectively.

[0180] Step 4: Online verification of the adaptive switching function of the dual-mode converter. When the grid short-circuit ratio fluctuates, the converter automatically switches the control mode to adapt to stability requirements. When the short-circuit ratio SCR is greater than the upper limit threshold, it switches to GFL mode; when the short-circuit ratio SCR is less than the lower limit threshold, it switches to VSG mode. The intermediate range between the upper and lower limit thresholds is determined autonomously by the trained agent. After the grid short-circuit ratio recovers to above the upper limit threshold, the grid strength is high, and the converter switches back to GFL control mode. Under load disturbance, the frequency quickly recovers to 50Hz, the voltage fluctuation is small, and the parameter adjustment is smooth.

[0181] The reward curve during offline training is as follows: Figure 4As shown, the loss curves of the value network and policy network during the offline training process are as follows: Figure 5 As shown, the offline training process of the agent exhibits high convergence, with the reward curve rising steadily, indicating that the agent gradually masters the strategy for optimizing the operating mode and control parameters of the dual-mode converter. During training, the value network's evaluation of state value becomes increasingly accurate, and the parameter updates of the policy network gradually stabilize.

[0182] The power output, current response, and frequency response of the dual-mode converters are shown in Figures 6(a), 6(b), and 6(c). The output power of both dual-mode converters can quickly respond to load changes and achieve the power distribution control target, with no significant overshoot during the transition process. Under load disturbances, the frequency of the distribution area fluctuates slightly before quickly recovering to the rated value of 50Hz, demonstrating good frequency support capability.

[0183] The control parameter variation curves of the dual-mode converter are shown in Figures 7(a) and 7(b). During power distribution, the agent automatically increases the virtual inertia and damping parameters, and restores the preset inertia and damping parameters after the power disturbance ends, effectively suppressing frequency fluctuations and ensuring the long-term stability of the system. For converter 2 at node 8 operating in GFL control mode, the agent autonomously selects a new current loop bandwidth during power disturbances, achieving accurate power point tracking and stable grid-connected operation. In VSG control mode, the agent quickly adjusts the virtual inertia and virtual damping parameters, enabling the converter to rapidly improve the system's anti-disturbance capability, suppress frequency oscillations caused by power surges, and restore steady-state parameters after adjustment. In GFL control mode, the agent adjusts the proportional and integral gains of the current loop to new values. Before adjustment, the current loop bandwidth was approximately 7.96Hz, and after adjustment, the current loop bandwidth was approximately 3.41Hz. This reduces the power point tracking speed to some extent, but improves the stability of GFL under high-power disturbances and weak grid-connected conditions.

[0184] The operating mode curves of the node SCR and the dual-mode converter are as follows: Figure 8 As shown, before t=2s, both dual-mode converters operate in GFL control mode. When t=2s, the overall grid strength decreases, and the agent's decision causes dual-mode converter 1 to switch from GFL control mode to VSG control mode. When t=4s, the overall grid strength decreases again, and the agent's decision causes dual-mode converter 2 to also switch from GFL control mode to VSG control mode. When t=6s, the overall grid strength recovers, and the agent's decision causes both dual-mode converters 1 and 2 to switch from VSG control mode to GFL control mode.

[0185] Voltage comparison curves of weak nodes in the transformer area are as follows: Figure 9As shown, when an external distribution network fault occurs, the impedance of the transformer substation increases. The lowest point of the node voltage in the substation using only the GFL strategy is 0.968 pu. However, after adopting the mode adaptive switching control strategy based on multi-agent deep reinforcement learning dual-mode converter multi-machine parallel control proposed in this chapter, since dual-mode converters capable of mode switching are set at the weak nodes of the substation, when an external power grid fault occurs, the dual-mode converters switch to VSG control mode to actively support voltage and frequency. Therefore, the node voltage level of the substation is significantly improved, with the lowest point being 0.979 pu, an improvement of about 1.14% in voltage level. This is beneficial to the stable operation of the substation during periods of reduced external power grid strength.

[0186] Example 2

[0187] A computer system includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the steps of the method as described in Embodiment 1.

[0188] The above description is only a preferred embodiment of the present invention. The scope of protection of the present invention is not limited to the above embodiments. Any equivalent modifications or changes made by those skilled in the art based on the content disclosed in the present invention should be included within the scope of protection set forth in the claims.

Claims

1. A method for voltage and frequency support control of a multi-machine weak grid using a dual-mode converter based on multiple agents, characterized in that, include: Step 1: Construct a distributed control architecture, including multiple ordinary nodes that communicate in series, and multiple dual-mode converters that are respectively connected to weak nodes in the ordinary nodes of the low-voltage distribution area line. The weak nodes are line nodes where the far-end power supply radius of the low-voltage distribution area line is greater than a set distance, the line impedance is greater than a set value, and the transient voltage fluctuation amplitude exceeds the limit; any adjacent nodes can send and receive operating status data bidirectionally. The dual-mode converter node is a controllable node. An intelligent agent control unit is deployed at the dual-mode converter. The dual-mode converter node collects voltage and load power flow operation data of adjacent ordinary nodes for global transformer area status perception and power allocation decision-making. The ordinary nodes are uncontrollable nodes, and they adopt a distributed adjacent node communication mode to perform natural power flow calculations and exchange node information with adjacent dual-mode converter nodes and ordinary nodes. Step 2: Model the operation mode and control parameter adjustment problem of the dual-mode converter agent as a MADDPG model. With the goal of minimizing the frequency deviation and voltage deviation of the distribution area, set a reward function that includes frequency and voltage penalties, parameter stability penalties, and mode switching penalties. Step 3: In offline mode, the transformer area jointly trains the MADDPG model in the dual-mode converter agent through global information. After training, each dual-mode converter is used for independent decision-making based on local information.

2. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system as described in claim 1, characterized in that, In step 1, the two dual-mode converter nodes adopt full-topology adjacent peer-to-peer distributed communication, bidirectionally exchanging power, voltage, and mode status information. Through a consensus protocol, the output power of the two dual-mode converters is allocated according to the controllable capacity ratio. The formulas for calculating power transmission and power exchange are as follows: (1) in, , They are nodes Local load power, , They are nodes , voltage amplitude, , They are nodes , voltage phase angle, For nodes Total active power injection For nodes Total reactive power injection For nodes With nodes Mutual inductance between lines, For nodes Self-inductance reactance; Undirected connected graph The formula is: (2) Edge set The formula is: (3) in, Represents a connected node and nodes Undirected edges, adjacency matrix When, explain the node and There is a communication connection between them; The formula for the node set V is: (4) Each node corresponds to a transformer area node, and k is the total number of nodes; Communication Neighbor Node Set The formula is: (5) | indicates a condition separator; Adjacency matrix The formula is: (6) Among them, the adjacency matrix =0 indicates a node and No direct communication or is itself a node; The formula for the degree matrix D is: (7) in, It is a dual-mode converter node Number of neighbors; set of communicating neighbor nodes This is a matrix of numbers for all adjacent converters that can communicate directly. Laplace matrix The formula is: (8) Definition of the first Local interaction information vector of dual-mode converter The formula is: (9) in, , , , The figures represent the voltage deviation and rate of change of the dual-mode converter in VSG or GFL modes, respectively. For virtual inertia, For virtual damping, This is the proportional gain of the current loop. This is the integral gain of the current loop; For dual-mode converter nodes Local short-circuit ratio, For operation mode, This indicates that the converter is operating in VSG control mode. This indicates that the system is operating in GFL control mode. For the rated capacity of the converter, For capacity weight, , These are the current active power and reactive power, respectively. This is the power deviation. The rate of change of power deviation, It is a continuous-time differential variable.

3. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system as described in claim 1, characterized in that, In step 1, a first-order consensus protocol with capacity weight is executed on active power and reactive power, and the power correction amount is output; the GFL control mode uses PQ control to track power commands, and the VSG control mode uses secondary frequency to recover received power commands, so as to realize dual-mode parallel power distribution according to capacity.

4. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system according to claim 3, characterized in that, No. Capacity weighting of dual-mode converters The calculation formula is: (10) For the first The rated apparent capacity of the dual-mode converter To traverse the first [unit / area] within the [area / area] Rated apparent capacity of the dual-mode converter; The formula for calculating the capacity-weighted first-order consensus protocol is: (11) in, , This is the power correction amount output by the distributed consensus protocol. for Time, Number The actual active power output of the dual-mode converter is currently... for Time, Node The communication neighbor number The actual active power output of the converter is currently... for Time, Number The actual reactive power output of the dual-mode converter is currently... for Time, Node Communication Neighbors The actual reactive power output of the converter is currently... It is a continuous-time differential variable; The formula for calculating the power distribution based on capacity when dual-mode converters are connected in parallel is as follows: (12) The capacity-weighted first-order consensus protocol only requires each dual-mode converter to exchange local state information with its neighboring nodes. Without the participation of a central node, it can ensure that all dual-mode converters can achieve the effect of allocating active and reactive power according to the rated capacity ratio, provided that the communication topology is connected.

5. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system according to claim 1, characterized in that, In step 2, the short-circuit ratio (SCR) is included in the state dimension of the MADDPG model. When the SCR is greater than the upper threshold, the control mode is switched to GFL. When the SCR is less than the lower threshold, the control mode is switched to VSG. When the SCR is in the middle range between the upper and lower thresholds, the agent in the dual-mode converter node makes autonomous optimization decisions based on the current operating parameters through the reinforcement learning neural network in the pre-trained MADDPG model, and dynamically selects the optimal operating control mode and the optimal parameter ratio.

6. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system according to claim 5, characterized in that, In step 2, in the MADDPG model, the state dimension is set to 10, including voltage deviation, voltage change rate, virtual inertia, damping parameters, PI control parameters, system frequency deviation, frequency change rate, local short-circuit ratio (SCR), and operating mode identifier for VSG / GFL dual modes; the action dimension is set to 5, including virtual inertia adjustment, damping coefficient adjustment, PI proportional parameter adjustment, PI integral parameter adjustment, and grid-connected mode binary switching action. The operating parameters include system frequency fluctuations. Voltage deviation in VSG mode Voltage deviation in GFL mode Virtual inertia Virtual damping Current loop proportional gain Current loop integral gain Dual-mode converter local short-circuit ratio Dual-mode converter operation mode .

7. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system according to claim 5, characterized in that, In the MADDPG model, the objective function is calculated as follows: (13) The objective is to minimize the time integral of the frequency deviation and voltage deviation of the distribution area. and For the first Virtual inertia and damping stability penalty term of dual-mode converter , These are the weighting coefficients. and For the first Virtual inertia of dual-mode converter and damping The current value, and For the first Virtual inertia of dual-mode converter and damping The rated value; and For the first The proportional gain of the dual-mode converter With integral coefficient Stability penalty item, , These are weighting coefficient one and weighting coefficient two, respectively. and The first Taiwan dual-mode converter and The current value of the coefficient. and The first Taiwan dual-mode converter and The rated value of the coefficient; The mode switching penalty coefficient, For dual-mode converter nodes Local short-circuit ratio; when operating mode At that time, the virtual inertia of the i-th dual-mode converter and damping The penalty item takes effect, and the dual-mode converter operates in VSG control mode; when the operating mode... hour, and With the penalty term applied, the dual-mode converter operates in GFL control mode; The total simulation duration is... For continuous-time differential variables, To enhance the learning of discrete decision step size, To optimize the objective function for multiple objectives; The formula for calculating the reward function is: (14) in, The output value of the reward function represents the state. Execute action Then, transition to the next state. Instant rewards obtained at that time; The penalty weight representing the frequency deviation of the converter. , These represent the penalty weights for voltage deviation in the dual-mode converter, with each converter contributing a reward value based on its own mode. , , , These represent the penalty weights for parameter mutations in the dual-mode converter, respectively. This indicates the penalty weight for converter mode switching. This represents the reward weight when the voltage deviation in dual-mode converter is less than a threshold. , These represent the reward weights when the voltage deviation in dual-mode converter is less than the threshold. For the first Single adjustment step size of virtual inertia in VSG mode of Taiwan converter For the first The single adjustment step size of the damping coefficient in VSG mode of the converter. For the first The single adjustment step size of the PI proportional parameter in GFL mode of the converter. For the first The single adjustment step size of the PI integral parameter in GFL mode of the converter. For absolute value operators, System frequency deviation acceptable threshold The acceptable threshold for bus voltage deviation in VSG mode. The acceptable threshold for bus voltage deviation in GFL mode. This is an indicator function; the condition within the parentheses is 1 if satisfied, and 0 if not satisfied.

8. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system according to claim 1, characterized in that, In step 3, the MADDPG model adopts a centralized training and decentralized execution framework. The policy network uses ReLU and Tanh activation functions to output continuous deterministic actions. The input of the value network is a concatenated vector of state and action, which uses the ReLU activation function.

9. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system as described in claim 8, characterized in that, The policy network is represented as follows: (15) in, The output of the policy network represents the state. Below, by parameters The optimal consecutive actions determined; The set of learnable parameters for the policy network, including weights 1 Weight 2 Bias 1 Bias 2 Apply the ReLU activation function to the output of the first layer and the tanh activation function to the output of the second layer, then multiply each element by the action scaling factor. and the offset Add; Value networks are represented as: (16) in, The output of the value network represents the state. Execute action The expected long-term cumulative rewards that can be obtained afterward; This is the set of learnable parameters for the value network, including the weight matrix. Weight matrix four Bias three , offset four ; The policy network loss function is expressed as: (17) in, For value network state The following are the actions output by the policy network. Value assessment The number of training rounds is given, and the optimization objective is to minimize the policy network loss. ; The value network loss function is expressed as: (18) in, For the goal value, For prediction Value is the current value network. State Next action Value assessment; The mean squared error loss function of the value network is used to measure the target within a batch of samples. Value and current network prediction Deviation of values; Target Q value The expression is derived from the Bellman equation, and the calculation formula is: (19) in, As a discount factor, and These are the output value network parameters and policy network parameters of the target network, respectively. For the first Instant reward for each sample For the first The next state of the power grid system after the converter corresponding to each sample executes the control action; The state space calculation formula for the control strategy is as follows: (20) The state space S includes the dual-mode frequency and voltage states of each dual-mode converter, VSG and GFL parameters, the local SCR size identified by the dual-mode converter node, and historical operating mode signals. During operation, each dual-mode converter calculates the frequency and voltage states in parallel under both VSG and GFL modes, using only the mode signals... Select one of the output modes to take effect; The formula for calculating the action space of the control strategy is: (21) in, To adjust the step size for inertia, To adjust the damping step size, and Adjust the upper and lower limits for virtual inertia. and Adjust the upper and lower limits for damping; For the first The single adjustment step size of the PI proportional parameter in GFL mode of the converter. For the first The single adjustment step size of the integral parameters of the converter in GFL mode. and Adjust the upper and lower limits of the PI ratio parameter for GFL mode. and Adjust the upper and lower limits of the PI integral parameters for GFL mode respectively.

10. The method for voltage and frequency support control of a multi-machine weak grid based dual-mode converter using a multi-agent system according to claim 1, characterized in that, Also includes: Step 4: Online verification of the adaptive switching function of the dual-mode converter. When the grid short-circuit ratio fluctuates, the converter switches control modes. When the short-circuit ratio SCR is greater than the upper limit threshold, it switches to GFL control mode. When the short-circuit ratio SCR is less than the lower limit threshold, it switches to VSG control mode. The intermediate range between the upper and lower limit thresholds is determined autonomously by the trained agent. After the grid short-circuit ratio recovers to above the upper limit threshold, the converter switches back to GFL control mode.

11. A computer system comprising a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the steps of the method as claimed in any one of claims 1-9.