AGC and AVC cooperative control method based on multi-agent deep reinforcement learning algorithm

By using a multi-agent deep reinforcement learning algorithm, a resource coordination model for wind, solar, thermal, and energy storage was established. The agents were trained using the TD3 algorithm, which solved the coupling effect under independent AGC and AVC control and improved the stability and control quality of the power system.

CN120914903APending Publication Date: 2025-11-07CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510967965.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Traditional independent control methods of AGC and AVC result in significant coupling effects between active and reactive power in new power systems, affecting the normal and stable operation of the power system. Existing collaborative control methods fail to effectively combine coupling characteristics, making it difficult to achieve a coordinated and integrated control effect.

Method used

A collaborative control method for AGC and AVC based on multi-agent deep reinforcement learning algorithm is adopted. By establishing a control model for coordinating multiple types of resources such as wind, solar, thermal and storage, and combining the physical characteristics of the controlled objects of AGC and AVC, the TD3 algorithm is used to train the agents to establish a multi-agent collaborative control system and realize the collaborative control of AGC and AVC.

Benefits of technology

It effectively reduces control costs, eliminates the mutual coupling effect between AGC and AVC, improves the stability and control quality of the power system, and realizes efficient solution of multi-objective and multi-variable problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120914903A_ABST
    Figure CN120914903A_ABST
Patent Text Reader

Abstract

The AGC and AVC cooperative control method based on the multi-agent deep reinforcement learning algorithm comprises the steps that an AGC and AVC control model of multi-type resource coordination is established based on wind-light-fire-storage multi-type power generation resource output characteristics; establishing a multi-agent cooperative control system according to physical characteristics of AGC and AVC control objects; based on a TD3 algorithm of a deep reinforcement learning algorithm, combining AGC and AVC control models, and adopting a Markov decision chain to establish an AGC and AVC cooperative control model of the power system; tD3 algorithm network parameters are set, discrete training and centralized learning of multiple agents are carried out, and model training and verification are carried out in combination with actual power grid data. According to the method, the adjustment capability of multiple types of resources participating in AGC and AVC control is fully utilized, and the control cost is reduced; and the cooperative control of AGC and AVC is also realized, and the mutual coupling influence is effectively eliminated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of power system control, and particularly relates to an AGC and AVC collaborative control method based on a multi-agent deep reinforcement learning algorithm. BACKGROUND

[0002] With the construction of new power systems in China, large-scale wind, light and other new energy is connected to the grid, and the installed capacity and power generation of new energy in the power grid accounts for an increasing proportion year by year. A large number of power electronic devices change the structural characteristics of the power grid, and the coupling relationship between active and reactive power of the power system is increasingly strong. The traditional network-source coordination method, automatic generation control AGC and automatic voltage control AVC, uses an independent control method, which will have a greater mutual influence, thereby threatening the normal and stable operation of the power system to a certain extent.

[0003] Under this background, some researches on AGC and AVC collaborative control have emerged, such as a hierarchical control system based on refined decomposition of AGC and AVC coupling parameters, and a cross-iteration solution based on the idea of reducing mutual influence to an acceptable range. These two methods, to some extent, reduce the coupling influence of active and reactive power, but they have not really combined the coupling characteristics to achieve a collaborative control effect. With the rapid development of the field of artificial intelligence, emerging intelligent algorithms are constantly improving, and intelligent algorithms based on deep reinforcement learning are increasingly widely used in power systems. In particular, the multi-agent deep reinforcement learning algorithm can simplify the model of intelligent agents and coordinate the cooperation between intelligent agents, and has strong applicability to the AGC and AVC collaborative control problem involving multiple targets and multiple variables.

[0004] Therefore, how to establish a deep reinforcement learning model based on the control relationship between automatic generation control AGC and automatic voltage control AVC, combined with the characteristics of the new power system regulation resources, and the mutual cooperation and coordination of multiple intelligent agents, is of great significance for solving the active and reactive power coupling influence of AGC and AVC in the power system, and realizing the collaborative control of AGC and AVC, and also provides a new solution. SUMMARY

[0005] To solve the problem of mutual influence of AGC and AVC control under the background of strong coupling between active and reactive power in new power systems, the application provides an AGC and AVC collaborative control method based on a multi-agent deep reinforcement learning algorithm, which aims to solve the network-source coordination control problem under the background of high proportion of new energy grid connection. This method not only makes full use of the regulation ability of multiple types of resources participating in AGC and AVC control, and reduces the control cost, but also realizes the collaborative control of AGC and AVC, and effectively eliminates the mutual coupling influence.

[0006] The technical scheme adopted by the application is as follows:

[0007] The AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm comprises the following steps:

[0008] Step 1: Based on the output characteristics of wind, light, fire and storage multi-type power generation resources, an AGC and AVC control model coordinated by multiple types of resources is established.

[0009] Step 2: According to the physical characteristics of AGC and AVC control objects, a multi-agent collaborative control system is established.

[0010] Step 3: Based on the TD3 algorithm of deep reinforcement learning algorithm, combined with the AGC and AVC control model, the Markov decision chain is used to establish the AGC and AVC collaborative control model of the power system.

[0011] Step 4: Set the network parameters of TD3 algorithm, carry out discrete training and centralized learning of multi-agent, and combine with the actual power grid data to carry out model training and verification on IEEE39 node.

[0012] In step 1, in a large power grid system, the power output characteristics of wind, light, fire and storage multi-type power generation resources connected to the grid are considered, and active and reactive power output models are established.

[0013] 1.1: Wind and light unit output characteristics:

[0014] The active power output of a doubly-fed asynchronous wind turbine depends on the current wind speed, and within the installed capacity range, the power generation is proportional to the wind speed:

[0015]

[0016] In formula (1), P w is the actual active power output of the wind turbine, v in is the minimum cut-in wind speed, v max is the maximum cut-out wind speed, v w is the current wind speed, is the rated output power of the wind turbine, and η is the conversion coefficient.

[0017] The power and reactive power output limit of the grid-side converter is considered as:

[0018]

[0019] In formula (2), Q c,max and Q c,min are the upper and lower limits of the reactive power output of the grid-side converter; S c,max is the capacity of the grid-side converter; and P c is the active power passing through the grid-side converter.

[0020] The model of the stator-side reactive power output is:

[0021]

[0022] In formula (3), P s and Q s are stator-side active and reactive power, u s is stator voltage amplitude; i sd and i sq are d-axis and q-axis components of stator-side current.

[0023] The reactive output level of the stator side of the doubly-fed asynchronous wind turbine is also affected by the rotor winding side converter, therefore, the stator-side reactive output range under the maximum current constraint of the rotor-side converter is:

[0024]

[0025] In formula (4), Q q,max and Q q,min are upper and lower limits of stator-side reactive output, x s is stator impedance, x m is rotor-side impedance, is maximum current of rotor-side converter;

[0026] The above formula (1) to formula (4) are integrated to establish upper and lower limit constraints of active and reactive output of the wind turbine:

[0027] P w,min ≤ P w ≤ P w,max (5);

[0028] Q w,min ≤ Q w ≤ Q w,max (6);

[0029] In the above formula, P w,max and P w,min are upper and lower limits of active output of the wind turbine, Q w is reactive output power of the wind turbine, Q w,max and Q w,min are upper and lower limits of reactive output of the wind turbine.

[0030] The active output of the photovoltaic turbine is affected by current temperature and irradiance intensity, and the power output model is shown in formula (7):

[0031]

[0032] In formula (7), P pv is current active output of the photovoltaic turbine, is the rated power generation of the photovoltaic power station; a is the temperature conversion power coefficient of the photovoltaic; T is the air temperature at the current moment; T ref is the air temperature reference value; ST pv is the current moment of light intensity.

[0033] The relationship model between the active power, the reactive power and the photovoltaic inverter capacity of the photovoltaic unit is established as shown in formula (8):

[0034]

[0035] In formula (8), Q pv,max and Q pv,min are the upper and lower limits of the reactive output of the photovoltaic unit; S pv is the photovoltaic inverter capacity. The upper and lower limits of the active and reactive outputs of the photovoltaic unit are established by comprehensively considering the above formula (7) to formula (8):

[0036] P pv,min ≤ P pv ≤ P pv,max (9);

[0037] Q pv,min ≤ Q pv ≤ Q pv,max (10);

[0038] In the above formula, P pv,max and P pv,min are the upper and lower limits of the active output of the wind power unit, Q pv is the reactive output power of the wind power unit, Q pv,max and Q pv,min are the upper and lower limits of the reactive output of the wind power unit.

[0039] 1.2: Output characteristics of the energy storage unit:

[0040] The state of charge calculation formula of the energy storage:

[0041]

[0042] In formula (11), SOC is the current state of charge, SOC0 is the initial state of charge, P ess is the power change of the energy storage in the time period△t, η ess is the charging and discharging efficiency, E ess is the capacity of the energy storage system;△t is the unit time period.

[0043] The active and reactive output ranges of the energy storage in the state of charge range are as follows:

[0044] P ess,min ≤ P ess ≤ P ess,max(12);

[0045] Q ess,min ≤Q ess ≤Q ess,max (13);

[0046] In the formula, P ess,max P ess,min These are the upper and lower limits of active power, respectively; Q ess,max Q ess,min These represent the upper and lower limits of reactive power, respectively. Furthermore, since the energy storage is connected to the grid via an inverter, the reactive power output of the energy storage at the current moment is constrained by both the inverter capacity and the active power output, as shown in the following equation:

[0047]

[0048] In equation (14), S ess This refers to the capacity of the energy storage converter.

[0049] 1.3: Output characteristics of thermal power units:

[0050] Consider the upper and lower limits of the ramping power of thermal power units participating in AGC and AVC regulation, as well as the continuous ramping time constraint;

[0051]

[0052]

[0053] In the above formula, P G For the active power output of thermal power units. These are the upper and lower limits of the output of unit i, respectively; R G This represents the unit's current ramp rate. These are the upper and lower limits of the ramp power of unit i, respectively; This represents the current ramp-climbing duration of the unit. Q represents the upper and lower limits of the ramp-up time for unit i; G This represents the current reactive power output of the thermal power unit. These represent the upper and lower limits of reactive power output.

[0054] In step 1,

[0055] 1) AGC control model with coordinated and complementary regulation resources:

[0056] Considering the active power output of wind turbines and photovoltaic units participating in regulation, the active power of the energy storage system is used to suppress its volatility. Simultaneously, taking into account the economics of AGC regulation, an AGC control model is established with the objective of optimizing tie-line power stability. The objective function is as follows:

[0057] f AGC= min{f cost + η w P drop,w + η pv P drop,pv + η tie △P tie} (19);

[0058] In formula (19), f AGC represents an AGC control target function, f cost is a cost of AGC units participating in regulation; η w , η pv are penalty coefficients of abandoned wind and light; P drop,w , P drop,pv are abandoned wind and light powers; η tie is a tie-line deviation penalty coefficient; and △P tie is a tie-line deviation.

[0059] The constraint conditions include active power output upper and lower limit constraints of wind, light, fire and energy storage units, which are formula (5), formula (9), formula (12) and formula (15) respectively; a fire unit climbing constraint: formula (16); and a climbing time constraint: formula (17).

[0060] 2) An AVC control model of coordinated and complementary adjustment resources is established.

[0061] An electrical distance index is used to establish an adjustment cost estimation of reactive power participating in AVC adjustment, and an AVC control model is established with the minimum voltage deviation of a hub node as a control target.

[0062] f AVC = min{D cost + η p △V p} (20);

[0063] In formula (20), f AVC represents an AVC control target function; D cost is an electrical distance cost of reactive power adjustment participating in AVC adjustment; η p is a hub node voltage deviation penalty coefficient; and △V p is a hub node voltage deviation.

[0064] The constraint conditions include reactive power output constraints of wind, light, fire and energy storage units, which are formula (6), formula (10), formula (13) and formula (18) respectively.

[0065] In step 2, the multi-agent collaborative control system specifically includes: AGC and AVC agents are respectively established, and a layered progressive training method is used for collaborative training of the agents.

[0066] The specific cooperative control system includes two agents: the AVC agent and the AGC agent, and the interaction between them and the environment (env). The following is a detailed description of the framework:

[0067] The environment (env) represents the external environment or operating space in which the system is located. The agents (Agents) are responsible for different control tasks. Each agent takes action according to the state of the environment and its own strategy. The state (State) represents the state of the agent AVC and the agent AGC at time t. These states include the current parameters of the system, performance indicators, and other information. The action (Action) represents the action taken by the agent AVC and the agent AGC at time t. These actions are generated according to the current state and the strategy of the agent, with the goal of optimizing the performance of the system.

[0068] Interaction process: the agent takes corresponding action according to the current state s. These actions are sent to the environment (env), which updates its state according to these actions. The updated environment state is fed back to the agent, which continues to take action according to the new environment state, forming a cycle.

[0069] The interaction process in a specific cycle includes the following steps:

[0070] (1) Divide T0 time into T0.1 time and T0.2 time. At T0.1 time, the AVC action agent first observes the environment to obtain the observation value o t , and makes the adjustment action a P,t of AVC control;

[0071] (2) At T0.2 time, the AGC action agent observes the environment to obtain the observation value o t , and makes the adjustment action a Q,t of AGC control;

[0072] (3) Finally, the system environment gives the corresponding action reward r P,t , r Q,t according to the final state change caused by the actions of the two agents, and enters the next adjustment time T1, repeating the above process T times to complete a training cycle of optimization.

[0073] In step 3, the TD3 algorithm is as follows:

[0074] 1) The TD3 algorithm uses the current actor network to select the optimal action and uses the target critic network to evaluate the strategy:

[0075] y t = r(s t , a t ) + γQ θ′ (st+1 ,π φ (s t+1 )) (21);

[0076] In formula (21): y t is a target value function; γ is a discount rate; Q θ′ is a target value function under state s t and action π φ (s t ); a t is an action of the agent at the current time, r(s t ,a t ) represents a reward obtained by action a t under state s t , s t+1 is a state at the next time, π φ (s t+1 ) represents an action of the agent under policy π φ and state s t+1 , and π φ represents an action policy of the agent, with the input value being a state variable and the output being an action variable.

[0077] Formula (22) is a critic network target value function of the TD3 algorithm using the clipping double Q learning method:

[0078]

[0079] In formula (22): represents an action value under current state s t+1 and policy ; represents selecting the minimum value in and .

[0080] To save training costs, the TD3 algorithm uses an independent actor network and two critic networks: the actor network is updated according to the critic network , and the target value functions of the two critic networks are equal;

[0081] 2) Policy delay update: the TD3 algorithm updates the actor network once every d times of updating of the critic network;

[0082] 3) Target policy smoothing regularization: the TD3 algorithm introduces a regularization method to reduce the variance of the target value, and performs Q value estimation smoothing by bootstrapping similar state-action pair estimates;

[0083] y t = r(st a t )+E θ′ [Q θ′ (s t+1 ,π φ′ (s t+1 )+ε)] (23);

[0084] In equation (23), ε is the added noise; r(s t ,a t ) is the reward function at state s t and action a t ;

[0085] E θ′ [Q θ′ (s t+1 ,π φ′ (s t+1 )+ε)] is the expected return; Q θ′ (s t+1 ,π φ′ (s t+1 )+ε) represents the value function under policy π φ′ with added noise ε.

[0086] Similarly, the smooth regularization is achieved by adding a random noise to the target policy, and taking the average on the mini-batch:

[0087]

[0088] In equation (24), y t ′ represents the target value function under the current policy, represents the minimum value of and .

[0089] ε~clip(N(0,σ),-c,c) (25);

[0090] In equation (25), c is the clipping length of the noise value smooth regularization; N(0,σ) is a normal distribution; clip(·) represents the clipping function.

[0091] In step 3, the AGC agent modeling is included:

[0092] a. State space: including the active power output of the AGC unit of the power system, the output state of the new energy unit such as wind, light and storage, and the system external tie-line power;

[0093] S AGC ={P Gi ,…P W ,P PV ,Pess ,△P tie} (26);

[0094] In formula (26), S AGC represents the state space of the AGC agent, P G , P W , P PV , P ess are respectively the current active power values of the conventional unit, the wind turbine unit, the photovoltaic unit and the energy storage system; P Gi represents the current active power of the conventional unit, △P tie is the deviation of the current tie-line power from the rated value.

[0095] b. Action space: including the adjustment action (unit output increase value) of the AGC unit, the AGC adjustment output of the wind, light, storage unit.

[0096] A AGC ={△P Gi ,…△P W ,△P PV ,△P ess} (27);

[0097] In formula (27), A AGC represents the action space of the AGC agent, △P G , △P W , △P PV , △P ess are respectively the active power adjustment values of the conventional unit, the wind turbine unit, the photovoltaic unit and the energy storage system; △P Gi represents the active power adjustment value of the conventional unit.

[0098] c. Reward function: including the cost of the wind, light, fire and storage unit participating in the AGC adjustment and the penalty of abandoned wind and light.

[0099]

[0100] In formula (28), R AGC represents the reward function of the AGC agent, R tie is the reward corresponding to the tie-line power control; R cost is the reward corresponding to the cost; is the rated value of the tie-line power; P tie represents the current tie-line power, ω1 and ω2 are respectively the weight coefficients of the adjustment cost and the penalty of abandoned wind and light;

[0101] In step 3, the AVC agent modeling is included:

[0102] a. State space: the state of reactive power regulation resources within the AVC partition, the voltage amplitude of key pivotal nodes within the AVC control area.

[0103] S AVC = {Q G ,…Q w ,Q pv ,Q ess ,T CB ,Q SVG △V p} (29);

[0104] In formula (29), S AVC represents the state space of the AVC agent, including the current reactive power output of wind and light units, energy storage and thermal power units; Q SVG is the current reactive power value of SVG, T CB is the current gear position of series capacitor, and △V p is the deviation of the current pivotal node voltage amplitude from the rated value.

[0105] b. Action space: including the terminal voltage control of conventional units, the reactive power output increase value of AVC units, and the action of reactive power compensation devices;

[0106] A AVC = {△Q Gi ,…△Q W ,△Q PV ,△Q ess ,△Q SVG ,△T CB} (30);

[0107] In formula (30), A AVC represents the action space of the AVC agent, △Q Gi , △Q W , △Q PV , △Q ess are the reactive power adjustment values of conventional units, wind power units, energy storage units, respectively; △Q SVG is the SVG reactive power adjustment value, and △T CB is the adjustment state of series capacitor.

[0108] c. Reward function: including the amplitude of voltage deviation of regional pivotal nodes, and the electrical distance of reactive power resources participating in AVC control;

[0109]

[0110] In formula (31), R AVC represents the reward function of the AVC agent, ω3 and ω4 are weight coefficients, V p represents the pivotal node voltage, and λ AVCPenalty for voltage out-of-limit of the rest nodes; χ represents the total value of voltage out-of-limit penalty, χ i represents the voltage out-of-limit penalty value of bus i, V i represents the meaning of the voltage of bus i; sum(·) represents a summation function, R V represents the reward corresponding to the voltage of the hub node; R D represents the electrical distance reward of reactive power adjustment. represents the rated value of the voltage of the hub node.

[0111] The technical effects of the AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm are as follows:

[0112] 1) In step 1 of the present application, compared with the traditional AGC and AVC control, the control effect of wind, light, and storage new energy is fully considered, the active and reactive power output models of wind, light, and storage are established based on the active and reactive power output characteristics of wind, light, and storage units, and based on the AGC and AVC control principles, the AGC and AVC control models that take into account the economy and optimal operation quality of the power system are established.

[0113] 2) In step 2 of the present application, in order to solve the problem that a single agent often has poor convergence effect and slow convergence speed when facing the multi-objective and multi-variable solution of AGC and AVC collaboration, a hierarchical multi-agent collaborative system is proposed, and an agent learning and training framework is established based on the idea of "decentralized training and centralized learning". The AGC control model and the AVC control model are established into corresponding agents, the multi-variable and multi-objective problems are decomposed and collaboratively processed, which can effectively improve the model solving efficiency.

[0114] 3) In step 3 of the present application, the TD3 algorithm improved based on the DDPG algorithm is used for training the agent model, and the clipping double Q learning method and target policy smoothing regularization means can improve the agent training efficiency and save training time.

[0115] 4) Compared with the traditional AGC and AVC coordinated control method, the method fully considers the participation of various regulation resources, maximally utilizes the regulation ability of wind and light new energy, and establishes an AGC and AVC control model with economic efficiency and optimal control quality as the target. Based on the AGC and AVC control characteristics, a multi-agent decentralized training and centralized learning collaborative control system is proposed. Finally, the Markov decision process is used to establish the interaction model of the AGC agent and the AVC agent with the power system environment, and the TD3 algorithm of deep reinforcement learning is used for learning and training to improve the training efficiency of the agent. The above design provides a new idea and solution for the integration of artificial intelligence algorithm and power system control technology, and the method not only fully utilizes the regulation ability of multiple types of resources participating in AGC and AVC control, but also reduces the control cost; and realizes AGC and AVC collaborative control, effectively eliminating the mutual coupling effect. BRIEF DESCRIPTION OF DRAWINGS

[0116] The application will be further described below in combination with the drawings and examples:

[0117] Figure 1 The method flowchart of the application.

[0118] Figure 2 The power output characteristic diagram of the wind turbine in the application.

[0119] Figure 3 The power output characteristic diagram of the photovoltaic unit in the application.

[0120] Figure 4 The power output characteristic diagram of the energy storage unit in the application.

[0121] Figure 5 The multi-agent coordinated control framework diagram of the application.

[0122] Figure 6 The IEEE39 node system structure diagram used in the simulation example of the application.

[0123] Figure 7 The training reward convergence diagram of the simulation example of the application. DETAILED DESCRIPTION

[0124] This embodiment mainly focuses on how to eliminate the active and reactive power mutual coupling effect in the new power system construction background, explores the AGC and AVC collaborative control means and control method, analyzes and models the output characteristics of the main several power generation resources of the power grid, and establishes the AGC and AVC independent control model; secondly, a multi-agent collaborative system is proposed, and the TD3 algorithm of deep reinforcement learning is used for agent training; finally, the actual power grid operation data is used for simulation verification in the improved IEEE39 node.

[0125] In the method, the active output characteristics and the reactive output characteristics of the doubly-fed asynchronous wind power generator, the grid-connected characteristics of the photovoltaic unit inverter, and the relationship model of the active, the reactive and the inverter capacity are analyzed in detail; the AGC control model is established with the lowest cost of the regional tie-line power stability and control as the target, and the AVC control model is established with the minimum electrical distance of the voltage deviation value and the reactive adjustment of the key hub node as the target.

[0126] The AGC and AVC agents are respectively established, the fusion model is established based on the Markov decision process, the TD3 algorithm is adopted in combination with the multi-agent collaborative system proposed in the design to train the agent, and in the multi-agent training, the agent respectively makes independent decision in the respective Actor network, and finally learns in the Critic network to obtain the global reward.

[0127] Finally, based on the installed and generated data of the actual power grid, after scaling and desensitizing processing, the example analysis is carried out on the improved IEEE39 node.

[0128] Therefore, the method effectively solves the active and reactive collaborative output coordination problem when multiple types of resources participate in AGC and AVC control, effectively reduces the mutual influence of independent AGC and AVC control of the power system, improves the AGC and AVC control quality, and realizes the safe and stable operation of the power system. Secondly, the multi-agent collaborative control system is established, the artificial intelligence algorithm is adopted, the solution scheme of the multi-objective optimization control of the power system is expanded, and good technical support is provided for the multi-objective collaborative control of the power system.

[0129] As shown in Figure 1 the AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm comprises the following steps:

[0130] Step S101: design includes: based on the output characteristics of wind-light-fire-storage multi-type power generation resources, an AGC and AVC control model of multi-type resource coordination is established;

[0131] Step S102: according to the physical characteristics of the AGC and AVC control object, a multi-agent collaborative control system is established;

[0132] Step S103: based on the TD3 algorithm of the deep reinforcement learning algorithm, in combination with the AGC and AVC control model, a Markov decision chain is adopted to establish an AGC and AVC collaborative control model of the power system;

[0133] Step S104: set the network parameters of the TD3 algorithm, perform discrete training and centralized learning of the multi-agent, and perform online testing on the IEEE39 node system.

[0134] In step S101, the research object includes the power output characteristics of various types of power generation resources such as wind, solar, thermal and energy storage connected to the grid in the large power grid system, and establishes its active and reactive power output models.

[0135] 1.1 Analysis of the output characteristics of wind and solar turbines:

[0136] Taking a doubly-fed asynchronous wind turbine as an example, its active and reactive power output characteristics are as follows: Figure 2 As shown, both the stator and rotor of the wind turbine are electrically connected to the power grid. The stator is directly connected to the grid, while the rotor is connected to the grid through a set of reverse-linked converters. Specifically: the rotor-side converter (RSC) is connected to the generator windings, providing AC excitation to the rotor, controlling the generator speed, and achieving active and reactive power control on the generator stator side. The grid-side converter (GSC) is connected to the grid, enabling energy exchange with the grid side and simultaneously controlling the stability of the DC bus voltage. Therefore, the power output of the grid-connected wind turbine is constrained by the stator static stability limit, the RSC characteristics, and the GSC characteristics.

[0137] The active power output of a doubly fed asynchronous wind turbine mainly depends on the current wind speed. Within the installed capacity range, its power generation is directly proportional to the wind speed.

[0138]

[0139] In the formula, v in For the minimum cut-in wind speed, v max For the maximum cut-out wind speed, v w The current wind speed, Here, η represents the rated output power of the wind turbine, and η is the conversion factor. The reactive power output of the wind turbine is limited by the characteristics of its internal power electronic equipment, requiring analysis in conjunction with the stator and rotor grid connection structure and the current active power output.

[0140] Considering the power reactive power output limit of the grid-side converter is:

[0141]

[0142] In the formula, Q c,max Q c,min These represent the upper and lower limits of reactive power output of the grid-side converter, S c,max For the grid-side converter capacity, P c Let represent the active power passing through the current grid-side converter. A simplified model for the reactive power output of the rotor-side converter is as follows:

[0143]

[0144] In the formula, P s, Q s are rotor-side active and reactive power, u s , i sd are rotor-side voltage and current, respectively. Meanwhile, the stator-side reactive power output of DFIG is mainly affected by the rotor-side converter, so the stator-side reactive power output range under the maximum rotor-side converter current constraint is:

[0145]

[0146] where Q q,max , Q q,min are the upper and lower limits of stator-side reactive power output, x s is the stator impedance, x m is the rotor-side impedance, is the maximum rotor-side converter current.

[0147] The active power output of photovoltaic units is mainly affected by the current temperature and irradiance intensity. The simplified power output model is shown in equation (5). The photovoltaic unit grid-connected structure is much simpler than the DFIG, which uses an inverter as the grid-connected interface and only outputs power to the grid, equivalent to the converter only working in inverting state. The relationship model between the active power, reactive power and apparent power of photovoltaic units can be established as shown in equation (6), and the power output characteristics diagram is shown in Figure 3 .

[0148]

[0149] where P pv is the current active power output of photovoltaic units, is the rated power of photovoltaic power station; a is the temperature conversion power coefficient of photovoltaic; T is the current air temperature; T ref is the air temperature reference value; ST pv is the current irradiance intensity. Q pv,max , Q pv,min are the upper and lower limits of photovoltaic unit reactive power output, and C is the capacity of photovoltaic inverter.

[0150] 2.2, Power output characteristics of energy storage units:

[0151] The power output of energy storage is affected by the real-time state of charge (SOC), and it also uses a converter as the grid-connected interface, so when the SOC of the energy storage system is within the allowed range, its power output characteristics are consistent with the converter, which has four-quadrant operation capability, as shown in Figure 4 .

[0152] The state of charge calculation formula of energy storage is:

[0153]

[0154] where SOC is the current state of charge, SOC0 is the initial state of charge, P ess is the power change of the energy storage in the time period△t, η ess is the charge-discharge efficiency, E ess is the capacity of the energy storage system. The active and reactive power output ranges of the energy storage in the state of charge are as follows:

[0155] P ess,min ≤P ess ≤P ess,max (12);

[0156] Q ess,min ≤Q ess ≤Q ess,max (13);

[0157] where P ess,max , P ess,min , Q ess,max , Q ess,min are the upper and lower limits of the active and reactive power respectively. In addition, the energy storage is connected to the grid through a converter, so the reactive power output of the energy storage at the current time is jointly constrained by the converter capacity and the active power output, as follows:

[0158]

[0159] where S ess is the converter capacity of the energy storage

[0160] 2.3, thermal power unit output characteristics:

[0161] Traditional AGC and AVC control is mainly adjusted by thermal power units, but compared with new energy units controlled by power electronic devices, the adjustment speed of thermal power units is slow, and it is often difficult to effectively respond to short-term AGC and AVC control instructions, therefore, in this design, the upper and lower limits of the unit climbing power of the thermal power unit participating in AGC and AVC adjustment and the continuous climbing time constraint need to be considered.

[0162]

[0163] where P and P are the upper and lower limits of the output of unit i respectively, are the upper and lower limits of the climbing power of unit i respectively; are the upper and lower limits of the climbing time of unit i respectively;

[0164] (1) : Establish AGC control model of coordinated complementary of each regulation resource:

[0165] The control variable of AGC control is the active power output of AGC units including wind, light and storage, in order to ensure the full use of wind and light power generation and reduce the rate of abandoned wind and light, the active power of wind turbine and photovoltaic unit is given priority to participate in the regulation, and the active power of storage system is used to suppress its volatility. At the same time, considering the economy of AGC regulation, the AGC control model is established to optimize the stability of tie-line power.

[0166] The following objective function is established:

[0167] f AGC = min{f cost +η w P drop,w +η pv P drop,pv +η tie △P tie} (19);

[0168] In the formula, f cost is the cost of AGC unit participating in regulation, η w and η pv are wind and light abandoned punishment coefficients, P drop,w and P drop,pv are wind and light power, η tie is tie-line deviation punishment coefficient, and △P tie is tie-line deviation.

[0169] Constraint condition: considering the upper and lower limit constraints of active power output of wind, light and storage units, the climbing constraints and climbing time constraints of thermal power units.

[0170] (2) : Establish AVC control model of coordinated complementary of each regulation resource:

[0171] The control variable of AVC control is mainly the reactive power resources that can participate in the regulation in the power system, including the reactive power output of units and some discrete or continuous reactive power compensation devices such as SVG, SVC and switching capacitor. But usually the regulation cost of reactive power is difficult to measure quantitatively, this design uses the index of electrical distance to establish the regulation cost estimation of reactive power participating in AVC regulation, and establishes the AVC control model with the minimum voltage deviation of the center node as the control target.

[0172] f AVC = min{D cost +η p △V p} (20);

[0173] In the formula, D costη is the electrical distance cost for participating in the AVC regulation reactive power adjustment p ΔV is the central node voltage deviation penalty coefficient p ΔV is the central node voltage deviation penalty coefficient

[0174] The constraint condition includes the reactive power output constraint mentioned in step 3.

[0175] In step S102, based on the AGC and AVC control target and control means, a multi-agent collaborative control system is established. As shown in Figure 5 .

[0176] The specific steps are as follows:

[0177] (1) T0 time is divided into T0.1 time and T0.2 time, at T0.1 time, the AVC action agent first observes the environment and obtains the observation value o t , and makes the adjustment action a P,t of AVC control;

[0178] (2) At T0.2 time, the AGC action agent observes the environment and obtains the observation value o t , and makes the AGC control action adjustment a Q,t ;

[0179] (3) Finally, the system environment gives the corresponding action reward r P,t , r Q,t according to the final state change caused by the action of the two agents, and enters the next adjustment time T1, and the above process is repeated T times to complete an optimization period of training.

[0180] In step S103, a Markov decision chain is used to establish an AGC and AVC collaborative control model based on TD3 algorithm of deep reinforcement learning.

[0181] (1). TD3 algorithm principle:

[0182] Twin Delayed Deep Deterministic Policy Gradient (TD3) is a deep reinforcement learning algorithm in the actor-critic framework, which is extended based on DDPG. In order to solve the problem of overestimation of Q value in actor-critic framework algorithm, TD3 uses three key technologies to improve the stability and performance of the algorithm.

[0183] 1) Clipped Double Q-learning under actor-critic framework. Inspired by deep double Q-learning, TD3 uses the current actor network to select the optimal action and uses the target critic network to evaluate the policy:

[0184] y t = r(s t , a t ) + γQ θ′ (s t+1 , π φ (s t+1 )) (21);

[0185] where y t is the target value function; γ is the discount rate; Q θ′ is the target value function under state s t and action π φ (s t ). In the DDPG algorithm, the target actor network and the target critic network use "soft update" to make the real network and the target network too similar, making it difficult to effectively separate the action selection and policy evaluation. Therefore, the TD3 algorithm uses the clipping double Q learning method to calculate the target value:

[0186]

[0187] In order to reduce the training cost, the TD3 algorithm uses an independent actor network and two critic networks. The actor network is updated according to the critic network , and the target value of the critic network is equal to .

[0188] 2) Policy delay update. In deep reinforcement learning algorithms, the target network is used to provide a stable learning target. Through multi-step update, the critic network can gradually reduce the error between the target Q value; however, in the case of large critic network error, the update of the actor network makes the policy appear discrete behavior. Therefore, the update frequency of the actor network should be lower than that of the critic network, so as to ensure that the actor network can be updated under the condition of low Q value error and improve the update efficiency of the actor network. TD3 algorithm updates the actor network once every d times of critic network update.

[0189] 3) Target policy smoothing regularization. Similar to the DDPG algorithm, since the TD3 uses a deterministic policy, the target value is easily affected by the function approximation error during critic update, resulting in inaccurate target value. Therefore, TD3 introduces a regularization method to reduce the variance of the target value, and estimates the Q value by smoothing the estimated value of the bootstrap similar state action pair:

[0190] yt = r(s t , a t ) + E θ′ [Q θ′ (s t+1 , π φ′ (s t+1 ) + ε)] (23);

[0191] where ε is the added noise; r(s t , a t ) is the reward function at state s t and action a t ; E θ′ [Q θ′ (s t+1 , π φ′ (s t+1 ) + ε)] is the expected return. Meanwhile, the smooth regularization is achieved by adding a random noise to the target policy and taking the average over mini-batches:

[0192]

[0193] ε ~ clip(N(0, σ), -c, c) (25);

[0194] where c is the clipping length of the noise value smooth regularization, and N(0, σ) is a normal distribution.

[0195] (2). AGC agent modeling:

[0196] State space, including the active power output of the AGC unit of the power system, the output state of the wind-solar-storage new energy unit, and the system external tie-line power.

[0197] S AGC = {P G1 , P G2 , …, P W , P PV , P ess , ΔP tie} (32);

[0198] where P G , P W , P PV , P ess are the current active power values of the conventional unit, wind turbine, photovoltaic unit, and energy storage system, respectively, and ΔP tie is the deviation of the current tie-line power from the rated value.

[0199] Action space, including the adjustment action (unit output increase value) of the AGC unit and the AGC adjustment output of the wind-solar-storage unit.

[0200] A AGC = {△P G1 ,△P G2 ,…△P W ,△P PV ,△P ess} (33);

[0201] wherein,△P G ,△P W ,△P PV ,△P ess are the active power adjustment values of conventional units, wind turbine generators, photovoltaic generators, and energy storage systems, respectively.

[0202] The reward function includes the cost of wind, light, fire and storage units participating in AGC regulation and the penalty of abandoned wind and light.

[0203]

[0204] wherein, R cost is the reward corresponding to the cost, R tie is the reward corresponding to the tie-line power control, is the tie-line power rating.

[0205] (3). AVC agent modeling:

[0206] State space, the state of the reactive power regulation resource in the AVC partition, and the voltage amplitude of the key hub node in the AVC control area.

[0207] S AVC = {Q G ,…Q W ,Q PV ,Q ess ,T CB ,Q SVG △V p} (29);

[0208] wherein, Q SVG is the current reactive power value of SVG, T CB is the current gear position of series capacitor, and△V p is the deviation of the current hub node voltage amplitude from the rated value.

[0209] Action space, including the terminal voltage control of conventional units, the reactive power output increase value of AVC units, and the action of reactive power compensation devices.

[0210] A AVC = {△Q G1 ,△Q G2 ,…△Q W ,△Q PV△Q ess △Q SVG △T CB} (30);

[0211] wherein, △Q G , △Q W , △Q PV , △Q ess are reactive power adjustment values of conventional units, wind power units and energy storage units respectively, △Q SVG is the SVG reactive power adjustment value, △T CB is the adjustment state of series capacitors.

[0212] The reward function includes the amplitude of voltage deviation of the regional hub node and the electrical distance of the reactive power resource participating in AVC control.

[0213]

[0214] wherein, R V is the reward of the hub node voltage, R D is the electrical distance reward of reactive power adjustment, λ AVC is the penalty of the voltage out-of-limit of the remaining nodes, and V is the rated value of the hub node voltage.

[0215] In the step S104, the network hyperparameters of the TD3 algorithm of the multi-agent are set, and model training and verification are performed on the IEEE39 node in combination with actual power grid data.

[0216] Hyperparameter setting: the Actor and Critic network structures of the TD3 algorithm in the application each contain two hidden layers, the number of neurons in each layer is 512, the activation function of the last layer of the Actor network is tanh, so that the operation output of each layer is in the range of [-1, 1], as shown in Table 1.

[0217] Table 1 Hyperparameter setting

[0218] Symbol Parameter Value - Critic learning rate 0.0001 - Actor learning rate 0.001 τ Discount factor 0.99 batch-size Number of samples drawn 100 noise Target action noise variance 0.05 Experience Replay Buffer Experience pool capacity 10000

[0219] The actual wind and light installed capacity ratio and power generation ratio of a certain province in Central China are adopted to establish a 39-node system as shown in Figure 6 and the method is verified by simulation on the system. Based on the parameters in Table 1 and the neural network model of AGC and AVC coordinated control established in steps 3 and 5, the number of training rounds is set to 100, 200 steps are explored in each round, a total of 10 training is performed, and the target average reward function curve obtained is as shown in Figure 7It can be seen that the agent training reward function tends to converge at about 20 rounds, and basically stabilizes at about 80 rounds. The tie-line deviation of the IEEE 39-node system is within 0.1 MW, the node voltage deviation is within 0.001 p.u., and the control accuracy reaches a relatively ideal AGC and AVC control effect.

[0220] Optimal control effect: The simulation results of the control of the present design are compared with the control effect of the traditional grid-source coordinated control method, i.e., AGC and AVC independent control, and the results are shown in Table 2.

[0221] Table 2 Comparison of control effects

[0222]

[0223] Observing the above results, when active and reactive load disturbances occur, the AGC and AVC independent control method, although the target accuracy of control is consistent with the method of the present application, produces interactive coupling effects in the control process, so that the control effect is not very good. The AGC and AVC coordinated control method proposed in the present application greatly improves the control accuracy of tie-line power and node voltage on the basis of independent control, and effectively eliminates the influence of AGC and AVC independent control. Specific data analysis shows that, compared with independent control, the coordinated control scheme of the present application improves the control accuracy of tie-line power deviation and node voltage deviation by 0.9083 MW and 0.00128 p.u., respectively, reduces the grid loss by 3.29% compared with the independent control method, reduces the reactive power balance degree by 20.52%, and reduces the control cost by 17.74%, indicating that the proposed method significantly improves the safety and economic operation level of the power grid.

Claims

1. AGC and AVC collaborative control method based on multi-agent deep reinforcement learning algorithm, characterized in that The method comprises the following steps: Step 1: based on the output characteristics of wind, light, fire and storage multi-type power generation resources, an AGC and AVC control model of multi-type resource coordination is established; Step 2: according to the physical characteristics of AGC and AVC control objects, a multi-agent collaborative control system is established; Step 3: based on the TD3 algorithm of deep reinforcement learning algorithm, combined with the AGC and AVC control model, a Markov decision chain is used to establish an AGC and AVC collaborative control model of the power system; Step 4: the network parameters of the TD3 algorithm are set, the discrete training and centralized learning of the multi-agent are carried out, and the model training and verification are carried out combined with the actual power grid data.

2. The AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm according to claim 1, characterized in that: In step 1, in a large power grid system, the power output characteristics of wind, light, fire and storage multi-type power generation resources connected to the grid are considered, and active and reactive power output models are established; 1.1: wind, light unit output characteristics: The active power output of a doubly-fed asynchronous wind turbine depends on the current wind speed, and within the installed capacity range, the power generation is proportional to the wind speed: In formula (1), P w is the actual active power output of the wind turbine, v in is the minimum cut-in wind speed, v max is the maximum cut-out wind speed, v w is the current wind speed, is the rated output power of the wind turbine, and η is the conversion coefficient. The power and reactive power output limit of the grid-side converter is considered as: In formula (2), Q c,max , Q c,min are upper and lower limits of reactive power output of the grid-side converter, respectively; S c,max is the capacity of the grid-side converter; and P c is the active power currently passing through the grid-side converter. The model of the stator-side reactive power output is: In formula (3), P s and Q s are stator-side active and reactive power, u s is a stator voltage amplitude; i sd and i sq are d-axis and q-axis components of stator-side current; The stator-side reactive power output of the doubly-fed asynchronous wind turbine is also affected by the rotor winding side converter, so the stator-side reactive power output range under the maximum current constraint of the rotor side converter is: In formula (4), Q q,max , Q q,min are upper and lower limits of stator-side reactive power output, respectively, x s is stator impedance, x m is rotor-side impedance, and I rmax is maximum current of the rotor-side converter. Based on the above formulas (1)-(4), the upper and lower limit constraints of the active and reactive power output of the wind turbine are established: P w,min ≤P w ≤P w,max (5); Q w,min ≤Q w ≤Q w,max (6) In the above formula, P w,max , P w,min are upper and lower limits of active output of the wind turbine generator, Q w is reactive output power of the wind turbine generator, Q w,max , Q w,min are upper and lower limits of reactive output of the wind turbine generator. The active power output of the photovoltaic unit is affected by the current temperature and irradiance, and the power output model is shown in formula (7): In formula (7), P pv is the current active output of the photovoltaic unit, is the rated power generation of the photovoltaic power station; a is the temperature conversion power coefficient of the photovoltaic; T is the air temperature at the current time; T ref is the air temperature reference value; ST pv is the current light intensity at the current time; The relationship model between the active power, reactive power and photovoltaic inverter capacity of the photovoltaic unit is established as shown in formula (8): In formula (8), Q pv,max , Q pv,min are the upper and lower limits of the reactive power output of the photovoltaic unit, respectively; S pv is the photovoltaic inverter capacity. Based on the above formulas (7)-(8), the upper and lower limit constraints of the active and reactive power output of the photovoltaic unit are established: P pv,min ≤P pv ≤P pv,max (9) Q pv,min ≤Q pv ≤Q pv,max (10) In the above formula, P pv,max , P pv,min are upper and lower limits of active output of the wind turbine generator, Q pv is reactive output power of the wind turbine generator, Q pv,max , Q pv,min are upper and lower limits of reactive output of the wind turbine generator. 1.2: storage unit output characteristics: The state of charge calculation formula of the storage is: In formula (11), SOC is the current state of charge, SOC0 is the initial state of charge, P ess is the power change of the energy storage in the time period Δt, η ess is the charge and discharge efficiency, E ess is the capacity of the energy storage system; and Δt is the unit time period. The active and reactive power output range of the storage within the state of charge range is as follows: P ess,min ≤P ess ≤P ess,max (12) Q ess,min ≤Q ess ≤Q ess,max (13) In the formula, P ess,max , P ess,min are upper and lower limits of active power respectively; Q ess,max , Q ess,min are upper and lower limits of reactive power respectively; in addition, the energy storage is connected to the grid through a converter, so the reactive output of the energy storage at the current moment is jointly constrained by the converter capacity and the active output, as follows: In formula (14), S ess is the energy storage inverter capacity; 1.3: fire unit output characteristics: The upper and lower limit constraints of the unit climbing power of the fire unit participating in AGC and AVC adjustment, as well as the continuous climbing time constraint, are considered; In the above formula, P G is the active output of the thermal power unit, are the upper and lower limits of the output of the unit i, respectively; R G is the current ramp rate of the unit, are the upper and lower limits of the ramp power of the unit i, respectively; is the current ramp duration of the unit, are the upper and lower limits of the ramp time of the unit i, respectively; Q G is the current reactive output of the thermal power unit, are the upper and lower limits of the reactive output, respectively. 3.The AGC and AVC collaborative control method based on multi-agent deep reinforcement learning algorithm according to claim 2, characterized in that: The AGC control model of the step 1 includes the coordinated complementary AGC control model of each adjusting resource, which is as follows: The active power output of the wind turbine and the photovoltaic unit is considered to participate in the adjustment, and the active power of the energy storage system is used to suppress its fluctuation; at the same time, the economy of AGC adjustment is considered, and the AGC control model with the optimal stability of the tie line power as the target is established, and the objective function is as follows: f AGC = min{f cost + η w P drop,w + η pv P drop,pv + η tie △P tie} (19); In formula (19), f AGC represents the AGC control target function, f cost is the cost of the AGC unit participating in regulation; η w , η pv are wind and light abandoned penalty coefficients respectively; P drop,w , P drop,pv are wind and light abandoned power respectively; η tie is a tie-line deviation penalty coefficient; ΔP tie is a tie-line deviation; The constraint condition: considering the upper and lower limit constraints of the active power output of the wind, light, fire and storage units, which are formula (5), formula (9), formula (12) and formula (15) respectively; the climbing constraint of the fire unit: formula (16) and the climbing time constraint: formula (17).

4. The AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm according to claim 3, characterized in that: The step 1 includes: establishing the AVC control model of the coordinated complementary AVC control model of each adjusting resource, which is as follows: The adjustment cost estimation of the reactive power participating in the AVC adjustment is established by using the electrical distance index, and the AVC control model is established with the minimum voltage deviation of the central node as the control target; f AVC = min{D cost + η p △V p} (20); In formula (20), f AVC represents the AVC control target function; D cost is the electrical distance cost of participating in the AVC regulation of reactive power adjustment; η p is the central node voltage deviation penalty coefficient; △V p is the central node voltage deviation; Constraints: reactive power output constraints of wind, solar, thermal and storage units, which are formula (6), formula (10), formula (13) and formula (18) respectively.

5. The AGC and AVC collaborative control method based on multi-agent deep reinforcement learning algorithm according to claim 4, characterized in that: In step 2, the multi-agent collaborative control system includes two agents: AVC agent and AGC agent, and their interactions with the environment, as follows: The environment represents the external environment or operating space in which the system is located; the agents are responsible for different control tasks; each agent takes action according to the state of the environment and its own strategy; the state represents the state of agent AVC and agent AGC at time t; these states include the current parameters, performance indicators and other information of the system; the action represents the action taken by agent AVC and agent AGC at time t, which is generated according to the current state and the strategy of the agent; The interaction process: the agent takes corresponding action according to the current state s; these actions are sent to the environment, which updates its state according to these actions; the updated environment state is fed back to the agent, which continues to take action according to the new environment state, forming a cycle.

6. The AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm according to claim 5, characterized in that: The interaction process in a cycle includes the following steps: (1) T0 moment is divided into T0.1 moment and T0.2 moment, at T0.1 moment, the AVC action intelligent agent first observes the environment, obtains the observation value o t , and makes the adjustment action a of AVC control P,t ; (2) At time T0.2, the AGC action agent observes the environment to obtain observation o t , and makes an AGC control action adjustment a Q,t ; (3) Finally, the system environment gives the corresponding action reward r according to the final state change brought by the action of the two agents P,t Q,t , and enters the next adjustment time T1, and the above process is repeated T times to complete a training cycle of optimization.​ 7. The AGC and AVC collaborative control method based on multi-agent deep reinforcement learning algorithm according to claim 6, characterized in that: In step 3, the TD3 algorithm is as follows: 1) TD3 algorithm uses the current actor network to select the optimal action and uses the target critic network to evaluate the strategy: y t = r(s t , a t ) + γQ θ′ (s t+1 , π φ (s t+1 )) (21); In equation (21): y t γ is the objective function; Q is the discount rate; θ′ For state s t and action π φ (s t The objective value function under (a) t For the agent's current action, r(s) t ,a t ) represents action a t In state s t The reward obtained below, s t+1 For the state at the next moment, π φ (s t+1 ) represents the agent's policy π φ and state s t+1 The following action; Formula (22) is the critic network target value function of the TD3 algorithm using the clipped double Q learning method: In formula (22): π φ1 (s t+1 ) represents the action value under the current state s t+1 and policy π φ1 ; represents selecting the minimum value in and To save the cost of training, the TD3 algorithm uses an independent actor network and two critic networks: the actor network is updated according to the critic network , and the target value functions of the two critic networks are equal; 2) Policy delay update: TD3 algorithm updates the actor network once every d times of critic network update; 3) Target policy smoothing regularization: TD3 algorithm introduces a regularization method to reduce the variance of the target value, which smooths the Q value estimation by bootstrapping similar state-action pairs: y t = r(s t , a t ) + E θ′ [Q θ′ (s t+1 , π φ′ (s t+1 ) + ε)] (23) In equation (23): ε is the added noise; r(s t , a t ) is the reward function at state s t and action a t . E θ′ [Q θ′ (s t+1 ,π φ′ (s t+1 [Q) + ε] represents the expected return; θ′ (s t+1 ,π φ′ (s t+1 )+ε) represents π φ′ The value function after adding noise ε; Similarly, the smoothing regularization is achieved by adding a random noise to the target policy and taking the average on the mini-batch: In formula (24): y' t represents the target value function under the current policy, represents selecting the minimum value in and ε ~ clip(N(0, σ), -c, c) (25); In formula (25), c is the noise value smoothing regularization clipping length; N(0, σ) is a normal distribution; clip(·) represents the clipping function.

8. The AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm according to claim 7, characterized in that: In step 3, AGC agent modeling is included: a. State space: including the active power output of AGC units in the power system, the output state of wind, solar and storage new energy units, and the system external tie-line power; S AGC = {P Gi ,…P W ,P PV ,P ess ,△P tie} (26); In formula (26), S AGC represents the state space of the AGC agent, P G , P W , P PV , P ess are current active power values of the conventional unit, the wind turbine generator, the photovoltaic unit and the energy storage system respectively; P Gi represents the current active power of the conventional unit, △P tie is the deviation of the current tie-line power from the rated value; b. Action space: including the adjustment action of AGC units, and the output of wind, solar and storage units participating in AGC adjustment; A AGC = {△P Gi ,…△P W ,△P PV ,△P ess} (27); In formula (27), A AGC represents the action space of the AGC agent, and △P G represents the action space of the AGC agent, and △P W represents the action space of the AGC agent, and △P PV represents the action space of the AGC agent, and △P ess respectively represent the active power adjustment values of the conventional unit, the wind turbine generator, the photovoltaic unit, and the energy storage system; and △P Gi represents the active power adjustment value of the conventional unit. c. Reward function: including the cost of wind, solar, thermal and storage units participating in AGC adjustment, and the penalty for wind and solar curtailment; In formula (28), R AGC represents a reward function of the AGC agent, R tie is a reward corresponding to tie-line power control; R cost is a reward corresponding to cost; is a tie-line power rating; P tie represents current tie-line power, and ω1 and ω2 are weight coefficients of adjustment cost and wind and light curtailment penalty, respectively. 9.The AGC and AVC collaborative control method based on multi-agent deep reinforcement learning algorithm according to claim 8, characterized in that: In step 3, AVC agent modeling is included: a. State space: the state of reactive power regulation resources in the AVC partition, and the voltage amplitude of key hub nodes in the AVC control area; S AVC = {Q G ,…Q w ,Q pv ,Q ess ,T CB ,Q SVG △V p} (29); In formula (29): S AVC represents the state space of the AVC agent, including the current reactive power output of the wind-solar generator set, energy storage and thermal power unit; Q SVG is the current reactive value of SVG, T CB is the current position of series capacitor, △V p is the deviation of the current central node voltage amplitude from the rated value; b. Action space: including the terminal voltage control of conventional units, the reactive power output increase value of AVC units, and the action of reactive power compensation devices; A AVC = {△Q Gi ,…△Q W ,△Q PV ,△Q ess ,△Q SVG ,△T CB} (30); In formula (30): A AVC represents the action space of the AVC agent, △Q Gi , △Q W , △Q PV , △Q ess are the reactive power adjustment values of the conventional unit, the wind turbine generator unit and the energy storage unit respectively; △Q SVG is the SVG reactive power adjustment value, △T CB is the adjustment state of the series capacitor; c. Reward function: including the voltage deviation amplitude of regional hub nodes and the electrical distance of reactive power resources participating in AVC control; In formula (31): R AVC represents the reward function of the AVC agent, ω3, ω4 are weight coefficients, V p represents the central node voltage, λ AVC is the over-limit punishment of the voltage of the remaining nodes; χ represents the total value of the voltage over-limit punishment, χ i represents the over-limit punishment value of the bus i voltage, V i represents the meaning of the bus i voltage; sum(·) represents a summation function, R V is the reward corresponding to the central node voltage; R D is the electrical distance reward of reactive power adjustment; is the central node voltage rating.

10. The AGC and AVC collaborative control method based on the multi-agent deep reinforcement learning algorithm according to claim 9, characterized in that: In step 4, the hyperparameter settings: the Actor and Critic network structures of the TD3 algorithm each contain two hidden layers, with 512 neurons in each layer, and the activation function of the last layer of the Actor network is tanh, so that the output of each layer is in the range of [-1, 1].