Adaptive voltage control method for distributed power generation based on multi-agent reinforcement learning

By adopting the distributed power adaptive voltage control method of multi-agent reinforcement learning in the distribution network, the voltage fluctuation problem of distribution network is solved, the coordinated regulation of distributed power supplies is realized, and the voltage quality and system economy are improved.

CN115149542BActive Publication Date: 2025-06-06TIANJIN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210408704.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-06-06
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively deal with the problem of voltage fluctuations in the distribution network, especially in the complex operating environment after the distributed power supply is connected. The traditional centralized control method has high communication capabilities and is difficult to ensure privacy, and physical model optimization is difficult to ensure accuracy.

Method used

Adoptional voltage control method of distributed power supply based on multi-agent reinforcement learning is adopted, and by constructing a distribution network voltage control Markov decision-making process and a multi-agent deep deterministic strategy gradient network, the cooperative regulation of agents in each region is realized and the dependence on the network parameters of the power system is reduced.

Benefits of technology

Distributed power control of new multi-region distribution networks has been realized, which improves the voltage operation level of the distribution network, improves the voltage quality, improves the economical operation of the power system, and can respond to the system status in real time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115149542B_ABST
    Figure CN115149542B_ABST
Patent Text Reader

Abstract

A distributed power adaptive voltage control method based on multi-agent reinforcement learning: according to the selected distribution network containing distributed power, input the basic parameter information of the distribution network; form a Markov decision process for voltage control of the distribution network containing distributed power, and construct each distribution network regional agent and a distributed power reactive output dynamic boundary mask method based on a multi-agent deep deterministic policy gradient network; conduct offline training on each distribution network regional agent to obtain each trained distribution network regional agent; regulate the distributed power in the region, and each distribution network regional agent gives a control strategy for the distributed power in the region based on the real-time input distribution network status in the region, and is processed by the distributed power reactive output dynamic boundary mask method and sent to the distributed power converter in the region for execution. The present invention can improve the voltage operation level of the distribution network, improve the voltage quality, and enhance the economic efficiency of power system operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a distributed power supply adaptive voltage control method, and more particularly to a distributed power supply adaptive voltage control method based on multi-agent reinforcement learning. Background Art

[0002] The large number of distributed generators (DG) has further complicated the operation of the distribution network. Among them, overvoltage and voltage fluctuation issues have received widespread attention. In order to cope with the voltage fluctuations in the distribution network, a variety of voltage regulation devices have been integrated into the distribution network. Inverter-based DG can use the remaining capacity to provide reactive voltage support for the system. Since inverter-based DG has good economy and highly flexible and controllable reactive support capabilities, it has become an effective means of voltage control in the distribution network.

[0003] How to deal with the increasingly complex operating environment of the distribution network and alleviate the voltage fluctuation problem of the distribution network by adjusting the reactive power output of distributed power sources on-site has become a key issue that needs to be solved urgently. The traditional centralized control method requires the collection of data from the entire network, global optimization, and issuing instructions through the management system. It has high requirements on the communication capability of the system, and privacy issues are difficult to be guaranteed. In actual operation, due to the difficulty in obtaining accurate parameters of the distribution network, the distributed power reactive power optimization method based on the physical model cannot guarantee the accuracy of the model. Deep reinforcement learning, as an adaptive, model-free data-driven method, can be trained through historical data and adjust the reactive output of distributed power converters according to real-time status, so as to adaptively deal with voltage fluctuation problems. Multi-agent reinforcement learning enables agents to exchange historical information by sharing historical interaction experience, so that agents in each region can update their strategies according to the collaborative goals during training. After the strategy is updated, it can respond in real time based on local information to reduce communication requirements.

[0004] At present, the variables in the action space of the research on reinforcement learning control of distribution networks are mostly independent of each other, while distributed power sources have the problem of active / reactive output coupling. Secondly, the state of the distribution system changes rapidly, and the requirements for local control and real-time control are high. In order to solve the above problems, a distributed power adaptive voltage control method based on multi-agent reinforcement learning is proposed. Summary of the invention

[0005] The technical problem to be solved by the present invention is to overcome the shortcomings of the prior art and provide a distributed power supply adaptive voltage control method based on multi-agent reinforcement learning that can realize distributed power supply regulation of new distribution networks in multiple regions and realize flexible voltage control of new distribution networks.

[0006] The technical solution adopted by the present invention is: a distributed power supply adaptive voltage control method based on multi-agent reinforcement learning, comprising the following steps:

[0007] 1) According to the selected distribution network containing distributed power sources, input the basic parameter information of the distribution network, including the distribution network topology and parameter information, distribution network partition information, access location, capacity and observation node of each partition distributed power source, load, distribution network reference voltage and reference power, and input photovoltaic and load annual historical operation data;

[0008] 2) Based on the basic parameter information of the distribution network provided in step 1), a voltage control Markov decision process of the distribution network containing distributed power sources is formed, and the regional agents of each distribution network and the dynamic boundary mask method of the reactive power output of distributed power sources are respectively constructed based on the multi-agent deep deterministic policy gradient network;

[0009] 3) Based on the distribution network regional intelligent agents in step 2) and the photovoltaic and load annual historical operation data in step 1), each distribution network regional intelligent agent is offline trained to obtain each distribution network regional intelligent agent that has been trained;

[0010] 4) Based on the distribution network regional intelligent agents trained in step 3), the distributed power sources in the region are regulated. Each distribution network regional intelligent agent gives a control strategy for the distributed power sources in the region based on the real-time input distribution network status of the region, and is processed by the dynamic boundary mask method of the reactive output of the distributed power sources in step 2) and sent to the distributed power converters in the region for execution.

[0011] The Markov decision process for forming the voltage control of the distribution network containing distributed power sources in step 2) is expressed as:

[0012]

[0013]

[0014]

[0015] in, V represents the state space set of the intelligent agents in the nth distribution network area; i , P i and Q i They represent the voltage amplitude, injected active power, and injected reactive power of the observation node i respectively; represents the set of observation nodes of the nth distribution network area agent; represents the action space set of intelligent agents in the nth distribution network area; represents the reactive power injected into the system by the converter of the jth distributed generation; represents the set of distributed power sources in the area to which the nth distribution network regional agent belongs; r n V represents the reward of the intelligent agent in the nth distribution network area; 0 Indicates the system reference voltage amplitude.

[0016] Step 2) constructs each distribution network regional agent based on a multi-agent deep deterministic policy gradient network. Each distribution network regional agent is composed of four deep neural networks, namely, a current action network, a target action network, a current value network and a target value, and a corresponding optimizer. For the nth distribution network regional agent:

[0017]

[0018]

[0019]

[0020]

[0021] in, They represent the input dimension and output dimension of the current action network and target action network of the nth distribution network area agent respectively; represents the state space set of intelligent agents in the nth distribution network area; represents the action space set of intelligent agents in the nth distribution network area; N represents the input dimension and output dimension of the current value network and target value network of the nth distribution network area agent respectively; area Indicates the number of distribution network areas.

[0022] The method for constructing a dynamic boundary mask of reactive power output of distributed power sources based on a multi-agent deep deterministic policy gradient network in step 2) is expressed as:

[0023]

[0024]

[0025] In the formula, Q bound,j represents the absolute value of the dynamic boundary of reactive capacity of the jth distributed generation; P j represents the active output of the jth distributed generation; represents the apparent capacity of the jth distributed generation; It represents the reactive power injected into the system by the converter of the jth distributed generation.

[0026] The offline training of each distribution network area agent described in step 3) specifically includes:

[0027] 3.1) Input the information of the distribution network area;

[0028] 3.2) Set the training hyperparameters, randomly initialize the network parameters of the distribution network area agents in each area, and initialize the shared experience replay pool;

[0029] 3.3) Set the maximum number of training times M and the time point H for single-step training, and set the current training times m = 0;

[0030] 3.4) Initialize the current training time point h = 0;

[0031] 3.5) Each distribution network area agent respectively obtains the current state of the distribution network in its own area;

[0032] 3.6) Each distribution network area agent gives the actions of the distributed power converter in its own area according to the distribution network state in step 3.5), and the actions are processed and executed by the dynamic boundary masking method of the reactive power output of the distributed power source;

[0033] 3.7) Each area of the distribution network enters the next state and gives rewards, and each distribution network area agent stores its local experience in the shared experience replay pool;

[0034] 3.8) Each distribution network area agent samples from the shared experience replay pool and updates its respective network parameters using the backpropagation algorithm;

[0035] 3.9) If h < H, then h = h + 1, and return to step 3.5), otherwise enter the next step;

[0036] 3.10) If m < M, then m = m + 1, and return to step 3.4), otherwise enter the next step;

[0037] 3.11) Calculate the convergence index σ of each distribution network area agent n :

[0038]

[0039]

[0040] In the formula, μ n is the average value of the training rewards from the th to the Mth time of the nth distribution network area agent; M is the number of training times; R e is the reward for the e-th training; σ n is the convergence index of the nth distribution network area agent;

[0041] Let the convergence accuracy be ε. When the convergence indexes of all distribution network area agents satisfy σ nWhen <ε, all distribution network regional agents are considered to have converged and the offline training is stopped. Otherwise, return to step 3.2) to reset the training hyperparameters and train again.

[0042] The distributed power adaptive voltage control method based on multi-agent reinforcement learning of the present invention is based on solving the problem of voltage fluctuation of the distribution network under the access of distributed power sources, fully considering the uncertainty and real-time requirements of the distribution system, considering the distribution network partition and the power limit of the distributed power converter, and by establishing a voltage control Markov decision process based on the distributed power converter, each region of the new distribution network constructs a regional agent based on deep deterministic policy gradient to achieve distributed power control of multi-region new distribution networks. The present invention is trained through historical data, which reduces the dependence on the network parameters of the power system; during the training process, collaborative control between multiple regional agents can be achieved by sharing experience; in the control stage, the agent can intelligently handle the problem of active / reactive output coupling of distributed power sources according to the real-time status of the regional distribution network, and quickly give the control strategy of the distributed power source, thereby improving the voltage operation level of the distribution network, improving the voltage quality, and improving the economic efficiency of the power system operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a flow chart of a distributed power adaptive voltage control method based on multi-agent reinforcement learning of the present invention;

[0044] Figure 2 This is a framework diagram of distributed power supply voltage control technology based on deep reinforcement learning;

[0045] Figure 3 This is the topology diagram of the improved IEEE 33-node example with distributed power generation;

[0046] Figure 4 It is the record of changes in the agent training reward;

[0047] Figure 5 It is the load and photovoltaic operation curve of the test data;

[0048] Figure 6 It is a comparison chart of voltage extremes in two scenarios;

[0049] Figure 7 It is a comparison diagram of the voltage distribution of node 15 in two scenarios. DETAILED DESCRIPTION

[0050] The following is a detailed description of the distributed power supply adaptive voltage control method based on multi-agent reinforcement learning of the present invention in conjunction with the embodiments and drawings.

[0051] like Figure 1 , Figure 2As shown, the distributed power supply adaptive voltage control method based on multi-agent reinforcement learning of the present invention comprises the following steps:

[0052] 1) According to the selected distribution network containing distributed generation, input the basic parameter information of the distribution network, including the distribution network topology and parameter information, distribution network partition information, access location, capacity and observation node of each partition distributed generation, load, distribution network reference voltage and reference power, and input photovoltaic and load annual historical operation data.

[0053] For the embodiment of the present invention, the IEEE 33 node calculation including distributed power access is as follows: Figure 3 Detailed parameters are shown in Table 1 and Table 2. The reference voltage of the IEEE 33-bus example is 12.66 kV, and the total active power demand and total reactive power demand of the load are 3.715 MW and 2.300 Mvar respectively.

[0054] The access situation of distributed generation is shown in Table 3. The system base power is set to 1MVA. The safe operating range of the active distribution network voltage is 0.90pu~1.10pu.

[0055] Table 1 Load access locations and power for the IEEE 33-node example

[0056]

[0057]

[0058] Table 2 IEEE 33-bus line parameters

[0059]

[0060] Table 3 PV access location and capacity

[0061]

[0062]

[0063] 2) Based on the basic parameter information of the distribution network provided in step 1), a voltage control Markov decision process of the distribution network containing distributed power sources is formed, and the regional agents of each distribution network and the dynamic boundary mask method of the reactive power output of distributed power sources are respectively constructed based on the multi-agent deep deterministic policy gradient network; wherein,

[0064] (1) The Markov decision process for forming the voltage control of the distribution network containing distributed generation is expressed as:

[0065]

[0066]

[0067]

[0068] in, V represents the state space set of the intelligent agents in the nth distribution network area; i , P i and Q i They represent the voltage amplitude, injected active power, and injected reactive power of the observation node i respectively; represents the set of observation nodes of the nth distribution network area agent; represents the action space set of intelligent agents in the nth distribution network area; represents the reactive power injected into the system by the converter of the jth distributed generation; represents the set of distributed power sources in the area to which the nth distribution network regional agent belongs; r n V represents the reward of the intelligent agent in the nth distribution network area; 0 Indicates the system reference voltage amplitude.

[0069] (2) The above-mentioned multi-agent deep deterministic policy gradient network is used to construct each distribution network regional agent. Each distribution network regional agent is composed of four deep neural networks, namely, the current action network, the target action network, the current value network and the target value network, and the corresponding optimizer. For the nth distribution network regional agent:

[0070]

[0071]

[0072]

[0073]

[0074] in, They represent the input dimension and output dimension of the current action network and target action network of the nth distribution network area agent respectively; represents the state space set of intelligent agents in the nth distribution network area; represents the action space set of intelligent agents in the nth distribution network area; N represents the input dimension and output dimension of the current value network and target value network of the nth distribution network area agent respectively; area Indicates the number of distribution network areas.

[0075] The current action network and the target action network have the same structure, and the current value network and the target value network have the same structure. The activation function of the hidden layer of the above four deep neural networks is a linear rectifier activation function; the activation function of the output layer of the current action network and the target action network is a hyperbolic tangent activation function; the output layer activation function of the current value network and the target value network is a linear activation function.

[0076] The network uses the Adam optimizer to update parameters, and the loss function of the network is:

[0077]

[0078]

[0079] Where J(θ) and J(ω) represent the loss functions of the current action network and the current value network respectively; θ and ω represent the parameters of the current action network and the current value network respectively; m represents the number of samples; Q represents the value function; φ represents the strategy function; y j 、s j and a j They represent the target value, state and action of sample j respectively.

[0080] (3) The method for constructing a dynamic boundary mask of reactive power output of distributed power sources based on a multi-agent deep deterministic policy gradient network is expressed as:

[0081]

[0082]

[0083] In the formula, Q bound,j represents the absolute value of the dynamic boundary of reactive capacity of the jth distributed generation; P j represents the active output of the jth distributed generation; represents the apparent capacity of the jth distributed generation; It represents the reactive power injected into the system by the converter of the jth distributed generation.

[0084] 3) Based on each distribution network regional intelligent agent in step 2) and the photovoltaic and load annual historical operation data in step 1), each distribution network regional intelligent agent is offline trained to obtain each distribution network regional intelligent agent that has been trained; the offline training of each distribution network regional intelligent agent specifically includes:

[0085] 3.1) Input distribution network area information;

[0086] 3.2) Set training hyperparameters, randomly initialize the network parameters of the distribution network regional agent in each region, and initialize the shared experience replay pool;

[0087] 3.3) Set the maximum number of training times M and the time point H for single-step training, and set the current number of training times m = 0;

[0088] 3.4) Initialize the current training time point h = 0;

[0089] 3.5) Each distribution network area agent respectively obtains the current state of the distribution network within its own area;

[0090] 3.6) Each distribution network area agent gives the actions of the distributed power converters within its own area according to the distribution network state in step 3.5), and the actions are processed and executed by the dynamic boundary masking method of the reactive power output of the distributed power source;

[0091] 3.7) Each area of the distribution network enters the next state and gives a reward, and each distribution network area agent stores its local experience in the shared experience replay pool;

[0092] 3.8) Each distribution network area agent samples from the shared experience replay pool and updates its respective network parameters using the backpropagation algorithm;

[0093] 3.9) If h < H, then h = h + 1, and return to step 3.5), otherwise enter the next step;

[0094] 3.10) If m < M, then m = m + 1, and return to step 3.4), otherwise enter the next step;

[0095] 3.11) Calculate the convergence index σ of each distribution network area agent n :

[0096]

[0097]

[0098] In the formula, μ n is the average value of the training rewards from the th to the Mth time of the nth distribution network area agent; M is the number of training times; R e is the reward for the e-th training; σ n is the convergence index of the nth distribution network area agent;

[0099] Let the convergence accuracy be ε. When the convergence indexes of all distribution network area agents satisfy σ n < ε, it is considered that all distribution network area agents converge, stop the offline training, otherwise return to step 3.2) to reset the training hyperparameters and train again.

[0100] Considering the computing power limitation of the edge controller, the interaction data of each region will be uploaded to the same experience replay pool. By sharing the historical interaction experience, the regional agents of each distribution network can obtain the historical status and corresponding action information of the other regional agents of the distribution network. During the training process, the regional agents of the distribution network in this area update their strategies according to the cooperation goals, thus realizing multi-region and multi-agent coordinated control.

[0101] 4) Based on the distribution network regional intelligent agents trained in step 3), the distributed power sources in the region are regulated. Each distribution network regional intelligent agent gives a control strategy for the distributed power sources in the region based on the real-time input distribution network status of the region, and is processed by the dynamic boundary mask method of the reactive output of the distributed power sources in step 2) and sent to the distributed power converters in the region for execution.

[0102] In order to verify the feasibility and effectiveness of the distributed power adaptive voltage control method based on multi-agent reinforcement learning of the present invention, the following three scenarios are adopted for verification and analysis in the embodiment of the present invention:

[0103] Scenario 1: The distributed power converter is not controlled, that is, the distributed power only generates active power to obtain the voltage level of the distribution network in the initial state.

[0104] Scenario 2: The distributed power supply adaptive voltage control method based on multi-agent reinforcement learning of the present invention is used to control the distributed power supply converters in each region.

[0105] Scenario 3: Use centralized optimization methods to control the distributed power converters in each region.

[0106] First, train the agent, and the training record is as follows Figure 4 The parameters of the agent are shown in Table 4. After the training is completed, a test day is selected to test the agent. The load and photovoltaic level curves on the test day are shown in Figure 5 shown.

[0107] Table 4 Parameters of each agent

[0108] parameter size Learning Rate(Actor) 0.0001 Learning Rate (Critic) 0.001 Batch Size 32 Episodes 300 Memory Pool Size 10000 Discount Factor 0.1 ε 0.01

[0109] The computer hardware environment for performing training and test calculations is Intel(R) Xeon(R) W-2102CPU, with a main frequency of 2.90GHz and a memory of 64GB; the software environment is the Windows 10 operating system.

[0110] Select the maximum and minimum voltage amplitudes at each time of the day and draw the voltage extreme value curve as shown in Figure 6 Select the photovoltaic access point node 15 and draw the voltage distribution of the node during the day as shown in Figure 7Further numerical analysis of the voltage distribution during the day was performed, and the various indicators of voltage quality were obtained as shown in Table 5.

[0111] Table 5 Optimization results of each scenario

[0112] Scenario Maximum voltage (pu) Minimum voltage (pu) Average voltage deviation Scene 1 1.0283 0.9351 0.0188 Scene 2 1.0128 0.9720 0.0043 Scene 3 1.0089 0.9739 0.0037

[0113] Compared with scenario one in which the distributed power converter is not controlled, scenario two uses the deep reinforcement learning agent in the present invention to control the distributed power on-site, and the average voltage deviation is reduced by 77.13%; the minimum value of the uncontrolled voltage in scenario one is 0.9351, which is far below the lower limit of the safe operation constraint, while the voltages in scenario two are all within the safe operation constraint range, the system voltage level is significantly improved, and is very close to the optimal control level in scenario three.

[0114] From the comparison of scenes one, two and three, it can be seen that the distributed power adaptive voltage control method based on multi-agent reinforcement learning of the present invention can respond to the system status in real time, adjust the output of the distributed power converter, and is very close to the control effect of the centralized optimization method, thereby improving the system voltage distribution and the operating state.

Claims

1. A distributed power adaptive voltage control method based on multi-agent reinforcement learning, It is characterized in that The steps include: 1) According to the selected distribution network containing distributed power sources, input the basic parameter information of the distribution network, including the distribution network topology and parameter information, distribution network partition information, access location, capacity and observation node of each partition distributed power source, load, distribution network reference voltage and reference power, and input photovoltaic and load annual historical operation data; 2) Based on the basic parameter information of the distribution network provided in step 1), a voltage control Markov decision process of the distribution network containing distributed power sources is formed, and the regional agents of each distribution network and the dynamic boundary mask method of the reactive power output of distributed power sources are respectively constructed based on the multi-agent deep deterministic policy gradient network; wherein: The above-mentioned multi-agent deep deterministic policy gradient network is used to construct each distribution network regional agent. Each distribution network regional agent is composed of four deep neural networks, namely, current action network, target action network, current value network and target value, and corresponding optimizers. For the nth distribution network regional agent: in, Respectively represent the input dimension and output dimension of the current action network and target action network of the nth distribution network area agent; Σ n represents the state space set of the n-th distribution network area intelligent agent; n represents the action space set of intelligent agents in the nth distribution network area; N represents the input dimension and output dimension of the current value network and target value network of the nth distribution network area agent respectively; area Indicates the number of distribution network areas; The method for constructing a dynamic boundary mask of reactive power output of distributed power sources based on a multi-agent deep deterministic policy gradient network is expressed as: In the formula, Q bound,j represents the absolute value of the dynamic boundary of reactive capacity of the jth distributed generation; P j represents the active output of the jth distributed generation; represents the apparent capacity of the jth distributed generation; represents the reactive power injected into the system by the converter of the jth distributed generation; 3) Based on the distribution network regional intelligent agents in step 2) and the photovoltaic and load annual historical operation data in step 1), each distribution network regional intelligent agent is offline trained to obtain each distribution network regional intelligent agent that has been trained; 4) Based on the distribution network regional intelligent agents trained in step 3), the distributed power sources in the region are regulated. Each distribution network regional intelligent agent gives a control strategy for the distributed power sources in the region based on the real-time input distribution network status of the region, and is processed by the dynamic boundary mask method of the reactive output of the distributed power sources in step 2) and sent to the distributed power converters in the region for execution.

2. According to claim 1, the distributed power supply adaptive voltage control method based on multi-agent reinforcement learning, It is characterized in that The Markov decision process for forming the voltage control of the distribution network containing distributed power sources in step 2) is expressed as: Among them, Σ n V represents the state space set of the intelligent agents in the nth distribution network area; i , P i and Q i They represent the voltage amplitude, injected active power, and injected reactive power of the observation node i respectively; A represents the set of observation nodes of the nth distribution network area agent; n represents the action space set of intelligent agents in the nth distribution network area; represents the reactive power injected into the system by the converter of the jth distributed generation; represents the set of distributed power sources in the area to which the nth distribution network regional agent belongs; r n V represents the reward of the intelligent agent in the nth distribution network area; 0 Indicates the system reference voltage amplitude.

3. According to claim 1, the distributed power adaptive voltage control method based on multi-agent reinforcement learning, It is characterized in that The offline training of each distribution network area agent described in step 3) specifically includes: 3.1) Input distribution network area information; 3.2) Set training hyperparameters, randomly initialize the network parameters of the distribution network regional agent in each region, and initialize the shared experience replay pool; 3.3) Set the maximum number of training times M and the single-step training time point H, and set the current number of training times m = 0; 3.4) Initialize the current training time point h = 0; 3.5) Each distribution network regional agent obtains the current status of the distribution network in its area; 3.6) Each distribution network regional agent gives the actions of the distributed power converters in the region according to the distribution network status in step 3.5), and the distributed power reactive output dynamic boundary mask method is used to process and execute the actions; 3.7) Each area of ​​the distribution network enters the next state and gives rewards. Each distribution network area agent stores local experience in the shared experience playback pool; 3.8) Each distribution network area agent samples from the shared experience replay pool and updates its respective network parameters using the backpropagation algorithm with gradients; 3.9) If h < H, then h = h + 1, and return to step 3.5), otherwise proceed to the next step; 3.10) If m < M, then m = m + 1, and return to step 3.4), otherwise proceed to the next step; 3.11) Calculate the convergence index σ of each distribution network area intelligent agent n : In the formula, μ n is the first intelligent agent in the nth distribution network area. The average value of the training rewards from the th to the Mth time; M is the number of training times; R e is the reward for the e-th training; σ n is the convergence index of the intelligent agent in the nth distribution network area; Assume that the convergence accuracy is ε, when the convergence index of all distribution network regional agents meets σ n <ε, all distribution network regional agents are considered to have converged and the offline training is stopped. Otherwise, return to step 3.2) to reset the training hyperparameters and train again.