A power system power flow calculation convergence improvement method and device based on deep reinforcement learning multi-agent collaboration

By employing a deep reinforcement learning-based multi-agent collaborative method, the problem of non-convergence in power flow calculation was solved, enabling efficient power flow regulation and optimization of the power grid system and improving the stability and efficiency of power grid operation.

CN119695914BActive Publication Date: 2025-12-19STATE GRID JIANGSU ECONOMIC RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411483209.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-12-19
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The problem of non-convergence in power grid flow calculation is difficult to solve in existing technologies, especially in power systems with a high proportion of new energy and electric vehicle charging loads that are rapidly connected. This leads to high uncertainty and randomness in the characteristics of power grid operation. Existing methods are time-consuming and inefficient, and cannot cope with complex and ever-changing scenarios.

Method used

A multi-agent collaborative approach based on deep reinforcement learning is adopted. By constructing multiple agent models, each agent corresponds to a region or device in the power system. Deep reinforcement learning algorithms are used to realize collaborative work and information sharing among the agents, dynamically adjust active and reactive power, and optimize global power flow convergence.

Benefits of technology

It significantly improves the power flow convergence and global optimization capabilities of the power grid system, effectively copes with changes in complex power systems, provides efficient dispatching solutions, and improves the stability and efficiency of power grid operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119695914B_ABST
    Figure CN119695914B_ABST
Patent Text Reader

Abstract

The application provides a power system power flow calculation convergence improvement method based on deep reinforcement learning multi-agent collaboration, by constructing multiple reinforcement learning agent models, each agent corresponds to a region or a device in the power system, and the deep reinforcement learning algorithm is used to realize the collaboration and information sharing between the agents. Through real-time communication and cooperation, each agent quickly evaluates the operating state of the power system and dynamically adjusts the active and reactive power distribution to optimize the convergence of the global power flow. Specifically, under different load and power generation conditions, the agent continuously improves the decision-making process by learning feedback and reward mechanisms to meet the adjustment needs of active and reactive power, thereby significantly improving the convergence speed and accuracy of power flow calculation. The method has strong robustness and can effectively cope with the complexity changes of the power system, providing an efficient and intelligent scheduling solution for extreme scenarios of power flow divergence.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of power system simulation analysis, and relates to a power system load flow calculation convergence improvement method based on multi-depth reinforcement learning intelligent agent cooperation. BACKGROUND

[0002] The power grid load flow calculation is the basis for long-term planning and real-time scheduling analysis and decision-making in the power system. Due to the high proportion of new energy and the rapid access of electric vehicle charging loads in the construction process of the new-type power system, the operation characteristics of the power grid present high uncertainty and randomness. The load flow calculation for the long-term planning operation mode of the power grid has long existed the problem of non-convergence. In the actual planning work of the power grid, the convergence adjustment of the load flow highly depends on the artificial experience and mainly adopts the trial-and-error method, which consumes a large amount of time cost and cannot guarantee the convergence after adjustment, and the efficiency is very low. The problem of non-convergence of the load flow becomes one of the important problems that disturb the power operation and scheduling, and it is difficult to cope with complex and variable scenes. The load flow calculation is essentially the solution of the nonlinear equation set of the power grid model, and therefore the initial value setting and the initial mode arrangement greatly affect the convergence of the load flow calculation. Finally, it may fall into a pathological state, making it difficult to converge for the solution of the load flow equation with a feasible solution or directly leading to no solution of the load flow. Both cases present the feedback of the calculation result of the non-convergence of the load flow. The problem of non-convergence of the load flow calculation has long existed after the large-scale distributed energy is connected to the grid, and the variable dimension of the load flow equation of the large power grid is high, and the number of adjustable parameters is large, so the convergence adjustment is more difficult and the efficiency is lower.

[0003] For the problem of non-convergence of the load flow, the problem of non-convergence of the load flow can be classified into two kinds of adjustment of no solution of the load flow and adjustment of pathological load flow. The problem of pathological load flow is often in the improvement of the method for the load flow calculation, or focuses on the improvement of the Newton-Raphson method and the setting of the initial value. In the actual power system load flow adjustment process, the load flow algorithm and the initial value are often solved by using the packaged solver, so the improvement of this part has poor practical applicability. For the adjustment of no solution of the load flow, the earliest research was based on the characteristics of the solution boundary of the system, an index reflecting the degree of no solution of the load flow was studied, and a method for adjusting no solution of the load flow was proposed.

[0004] With the rise of artificial intelligence technology, scholars have made preliminary attempts to apply artificial intelligence technology to load flow analysis. Cluster analysis, expert system and other methods have been proposed to solve the problem of convergence of the load flow, but the first few methods are lacking in flexibility and self-exploration mechanism. SUMMARY

[0005] The purpose of the present application is to solve the defects in the prior art, and to provide a power system load flow calculation convergence improvement method based on deep reinforcement learning multi-agent cooperation. The application of the multi-agent system solves the problem of non-convergence of the load flow in the prior art.

[0006] The application adopts the technical scheme of a power system power flow calculation convergence improvement method based on multi-depth reinforcement learning agent cooperation, comprising the following steps:

[0007] S10: Obtain power grid data of the power grid system, wherein the power grid data comprises active power, reactive power balance, generator output, node capacitor and reactor switching state, and balancing machine output;

[0008] S20: According to the power grid data, determine the active power adjustment range of the power plant unit in the region and the switching range of the node capacitor and reactor through a deep reinforcement learning algorithm;

[0009] S30: Perform power flow calculation on the power grid system, and perform active power balance judgment and reactive power balance judgment;

[0010] S40: If the power grid system calculation does not converge, the deep reinforcement learning multi-agent adjusts by cooperating to adjust the active power of the power plant unit in the region and the switching of the reactive power compensation device of the reactive power compensation node in the region;

[0011] Among them, the deep reinforcement learning algorithm constructs a single agent state space according to the power grid system power flow state space, the active power flow adjustment state space, the capacitor and reactor switching state, the balancing machine state space, the important variables of the active power and reactive power balance of the power grid, and the output state of the generator., and the joint state space is composed of the states of multiple agents; the comprehensive reward is composed of the rewards of multiple agents; the action space is constructed according to the output change or effectiveness of the generator, and the effectiveness of the specific capacitor and reactor. The specific capacitor and reactor refers to the capacitor and reactor that is conducive to adjusting the reactive imbalance of the system.

[0012] The application constructs multiple reinforcement learning agent models, each agent corresponds to a region or a type of device in the power system, and uses a deep reinforcement learning algorithm to realize the cooperative work and information sharing between agents. The application uses the autonomous learning ability of the deep reinforcement learning agent to accumulate experience in the continuously accumulated optimization strategy, effectively improves the power flow convergence of the power grid system. At the same time, the real-time communication and cooperation of multiple agents are used to quickly evaluate the operating state of the power system, and the active and reactive power distribution is dynamically adjusted to further optimize the convergence of the global power flow.

[0013] Further, the above step S20 comprises:

[0014] S210: Take the power grid data as the input data of the deep reinforcement learning multi-agent, and set the initial state set;

[0015] S220: Set the joint state space and comprehensive reward value for multiple agents in the deep reinforcement learning algorithm;

[0016] The state space of a single intelligent agent includes the power flow state space of the power grid system, the active power flow adjustment state space, the reactive power flow adjustment state space, and the balancing machine state space; the comprehensive reward value includes the reward value of the calculation result and the reward value of the constraint condition.

[0017] S230: Based on the initial state set, determine the range of power plant units for regional active power adjustment, reactive power compensation nodes, and the number of reactive power compensation devices to be switched on and off.

[0018] More specifically, the joint state space is set up as follows:

[0019] (1) Power flow state space of power grid system:

[0020] P L ={P L1 ,P L2 ,P L3 ,...,P Li}

[0021] In the formula: P L Let P be the set of active power of all power grid lines. Li The actual active power output of the line is represented by i, where i is the line number.

[0022] (2) Active power flow adjustment state space:

[0023] P G ={P G1 ,P G2 ,P G3 ,...,P Gj}

[0024] In the formula: P G For the collection of active power output from all power plant units, P Gj This represents the active power output of a single generator, where j is the generator number.

[0025] (3) Reactive power flow adjustment state space:

[0026]

[0027] Where: δ CZ This is the set of the number of reactive capacitors and reactances switched at nodes. Z represents the number of capacitors and reactances switched at the k-th node;

[0028] (4) The state space of the balancing machine:

[0029] P s ={P s1P s2 P s3 P sh}

[0030] P s is the active power output of all balancing units, P sh is the active power output of a single balancing unit, and h is the balancing unit number;

[0031] The state space of a single agent is:

[0032] S n = {P L , P G , δ CZ , P s};

[0033] The joint state space of multiple agents is:

[0034] S = {S 1 , S 2 , S 3 ,..., S n ,..., S N};

[0035] In the formula: S is the set of all agent states, n is the agent number, and N is the total number of multiple agents.

[0036] The above comprehensive reward value is set as follows:

[0037] (1) Result reward value: The reward value is set according to the power flow calculation result, and the specific parameter setting is as follows:

[0038]

[0039] (2) Constraint condition reward value: The reward value is set through the constraint condition of rated power limit, and the specific parameter setting is as follows:

[0040]

[0041] In the formula: P sh is the active power output of a single balancing unit, P Gj is the active power output of a single generator, P Li is the actual active power output of a line; P Gjmax , P Gjmin are the upper and lower limits of the generator output; P shmax , P shmin are the upper and lower limits of the balancing unit output; P Limin , P Limax are the upper and lower limits of the active power output of the line.

[0042] The comprehensive reward value R of the multi-agent of the deep reinforcement learning algorithm is:

[0043] R=r1+r2+r3+r4+r5.

[0044] Further, before step S30, the following operation is performed: the network structure data of the power grid system is simplified, and the network structure data of the power grid system of the region to be analyzed is retained.

[0045] More specifically, the coordinated adjustment in step S40 is the coordinated adjustment of active power and reactive compensation, and the active power of the power plant in the balance output adjustment region and the number of capacitors and inductors (reactance) are adjusted according to the balance output.

[0046] The action space of the multi-agent of the deep reinforcement learning algorithm is as follows:

[0047]

[0048]

[0049] In the formula: ΔP G is the active power change of the generator; β is the effectiveness of the capacitor and the reactor, β=1 is the capacitor switching; β=-1 is the reactor switching; N G is the total number of generators; is the reactive power change, C is the capacitor model, and Z is the reactor model;

[0050] The generator output is discretized, and the following can be obtained:

[0051]

[0052] In the formula: P i GM is the upper limit of the output of the generator i, ΔP G is the unit size of the generator adjustment, and k is the adjustment multiple of each step of the power plant.

[0053] The application also provides a power system power flow calculation convergence improvement device, which comprises:

[0054] A data acquisition module acquires power grid data of a power grid system, including active power balance, reactive power balance, generator output, switching state of node capacitors and reactors, and balance machine output.

[0055] An environment construction module sets a calculation environment of a deep reinforcement learning algorithm according to the power grid data, including a state space construction unit, an action space construction unit, and a reward value construction unit.

[0056] A power flow calculation module performs power flow calculation on the power grid system, judges active power balance and reactive power balance; if the power grid system calculation does not converge, the deep reinforcement learning multi-agent adjusts the active power of the power plant unit in the region and the reactive power compensation device switching of the reactive power compensation node in the region through cooperation respectively.

[0057] The state space construction unit constructs a single agent state space according to the power grid system power flow state space, active power flow power flow adjustment state space, capacitor and reactor switching state, balance machine state space, important variables of power grid active power and reactive power balance, and generator output state, and the joint state space is composed of the states of multiple agents.

[0058] The action space construction unit constructs an action space according to the output change or effectiveness of the generator, and the effectiveness of the specific capacitor and reactor; the specific capacitor and reactor refer to the capacitor and reactor that are beneficial to adjusting the reactive imbalance of the system.

[0059] The reward value construction unit is used to construct the comprehensive reward value of the deep reinforcement learning algorithm multi-agent, including the calculation result reward value and the constraint condition reward value.

[0060] Compared with the prior art, the present application has the following advantages:

[0061] Through the deep reinforcement learning technology, each agent can intelligently allocate tasks and optimize adjustment strategies, handle complex nonlinear relationships and dynamic changes, and thus realize accurate power flow adjustment. In addition, the present application obtains the power flow data of the power grid system, so that the active power flow adjustment and the reactive power flow adjustment have directionality.

[0062] The method has strong robustness and can effectively cope with the complexity of the power system, providing an efficient and intelligent scheduling solution for extreme scenarios of power flow divergence, and has important theoretical and practical application value. BRIEF DESCRIPTION OF DRAWINGS

[0063] Figure 1 The figure is a step schematic diagram of the power system power flow calculation convergence improvement method of the present application;

[0064] Figure 2 The figure is a whole flow chart of the power system power flow calculation convergence improvement method of the present application;

[0065] Figure 3 The figure is the influence of load parameter change on power flow calculation convergence rate in embodiment 2 of the present application;

[0066] Figure 4 Convergence curve of reward function in embodiment 2 of the present application with the number of training rounds. DETAILED DESCRIPTION

[0067] The technical solutions in the embodiments of the present application will be clearly and completely described in combination with the drawings in the embodiments of the present application.

[0068] Embodiment 1

[0069] In the examples of the present application, Figure 1 and Figure 2 The specific steps and detailed flowchart of the power system load flow calculation convergence improvement method based on deep reinforcement learning multi-agent collaboration according to the present application are shown in Figs. Figure 1 and Figure 2 The present application comprises:

[0070] S10: Obtain power grid data of the power grid system, wherein the power grid data includes active power and reactive power balance of the power grid, and involves some regional characteristics such as output state of the generator, switching state of the capacitor / reactor, and inter-regional exchange power, etc.

[0071] The power grid data is obtained in real time from the power grid operation system by the DSP software, and then exported according to the byte method.

[0072] S20: According to the power grid data, determine the active power adjustment range and the reactive power compensation node by the deep reinforcement learning algorithm.

[0073] Specifically as follows:

[0074] S210: Take the power grid data as the input data of the deep reinforcement learning multi-agent, and set the initial state set.

[0075] S220: Set the state space and reward value of the deep reinforcement learning algorithm.

[0076] The state space includes important variables of the active power and reactive power balance of the power grid, and involves some regional characteristics such as output state of the generator, switching state of the capacitor / reactor, and inter-regional exchange power, etc. The load flow states of the multiple agents together constitute the state space, which is specifically as follows:

[0077] (1) Power grid system load flow state space:

[0078] P L ={P L1 ,P L2 ,P L3 ,...,P Li},

[0079] Wherein, P LP = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system Li P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system

[0080] (2) Active power flow adjustment state space:

[0081] P = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system G G1 G2 G3 Gj

[0082] P = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system G Gj

[0083] (3) Reactive power flow adjustment state space:

[0084]

[0085] P = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system CZ

[0086] (4) Balancer state space is:

[0087] P = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system s s1 s2 s3 sh

[0088] P = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system s sh

[0089] Single-agent space is:

[0090] P = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system n L G CZ s

[0091] Multi-agent state space is:

[0092] P = {P1, P2, P3,..., Pn} is the set of active power of all transmission lines in the power grid system 1 2 3 n N ​​​​​​​​​​​​​​​​​​​​​​​​​​

[0093] where S is the set of all agent states, n is the agent number, and N is the total number of multi-agents.

[0094] The reward value includes:

[0095] (1) The reward value of the calculation result: the calculation result of the power flow has only two kinds, convergence and non-convergence, and the convergence state is the adjustment target. In the case of non-adjustment convergence, the calculation result of the power flow is mostly non-convergence. Therefore, when the convergence is achieved, the reward value is set to a large positive number; when the convergence is not achieved, the reward value is set to a small negative number. The specific parameter settings are as follows:

[0096]

[0097] (2) The reward value of the constraint condition: the constraint condition is mainly the rated power limit, including the following two points (the specific parameter settings are shown in the following formula): ① The rated power constraint of the generator: if the output power is out of limit, a certain negative reward value will be obtained. ② The output constraint of the balancing machine: if the exchange power is out of limit, a certain negative reward value will be obtained.

[0098]

[0099] where P sh is the active power output of a single balancing machine, P Gj is the active power output of a single generator, P Li is the actual active power of the line; P Gjmax , P Gjmin are the upper and lower limits of the generator output; P shmax , P shmin are the upper and lower limits of the balancing machine output; P Limin , P Limax are the upper and lower limits of the line active power.

[0100] Therefore, the final reward value R of the power flow convergence improvement agent is:

[0101] R = r1 + r2 + r3 + r4 + r5.

[0102] After the multi-agents obtain the initial state space, the power flow calculation converges after multiple adjustment scenes, indicating that the algorithm has reached the final adjustment target state of the scene.

[0103] S230: The multi-agents determine the range of power plants for regional active power adjustment and the nodes for reactive power compensation according to the initial state set.

[0104] The initial state set usually includes the following information:

[0105] System load; power generation capacity, operating status, and cost characteristics of individual power plants; reactive power demand and status of compensation equipment; topology of the system, including the connection relationship of nodes and lines.

[0106] In determining the power plants within the region that can be adjusted for active power, the required active power is calculated according to the load demand. If the load demand exceeds the power generation capacity of a power plant, the power plant may be selected as the object of adjustment. Analyzing the power generation capacity of each power plant, including the minimum / maximum output, determines which power plants can participate in adjustment. According to the topology of the power system, ensure that the selected power plants can effectively meet the load demand in the region through the network.

[0107] Determine the node selection of reactive power compensation Analyze the reactive power demand of each node and identify the nodes that need compensation, such as nodes with lower voltage. Use the multi-agent model to select the best compensation node and device configuration to minimize voltage deviation or improve system stability.

[0108] Establish a dynamic adjustment mechanism to adapt to changes in load and system state, ensuring real-time compensation of reactive power.

[0109] In the multi-agent environment, each power plant and compensation node can be regarded as an independent agent. They need to be coordinated through the following ways:

[0110] Agents share information such as load demand, power generation capacity, and compensation status to make better decisions. Design coordination protocols to ensure that agents can coordinate with each other when adjusting active and reactive power, avoiding conflicts. Introduce a reward mechanism to encourage agents to maximize load demand and optimize operating costs while maintaining system stability. Once the range of power plants and reactive compensation nodes is determined, implement the control strategy, and monitor and feedback the operating state of the system to facilitate real-time adjustment and optimization.

[0111] Through the above steps, the range of active power adjustment of power plants in the region and the reactive compensation nodes in the multi-agent system can be effectively determined.

[0112] S30: Perform power flow calculation on the power grid system, and make judgments on active power balance and reactive power balance. Specifically as follows:

[0113] Active power balance is global balance, which needs to ensure the active power balance of all generators and loads, as well as the non-exceeding of transmission channels, which is relatively easy to adjust.

[0114] S310: Statistically analyze active power in each region and check the active power exchange between regions. If the exchange is too large, increase the start-up of the receiving end region and reduce the start-up of the sending end region, thereby reducing the transmission power of the tie line.

[0115] S320: Check whether the active power borne by the balancing machine is reasonable. If the active output of the balancing machine exceeds the bearing range, increase the generator output to make the active output of the balancing machine within the reasonable range.

[0116] S330: When adjusting, it is also necessary to check whether the generator output is out of range.

[0117] The reactive power balance is a hierarchical and zonal balance in place, and it is difficult to accurately locate the position of the reactive power imbalance. Only the area where the reactive power may be missing can be considered, so it is more difficult to adjust. The specific method is as follows:

[0118] S340: Estimate the reactive power shortage according to the voltage level or zone, and switch the capacitor or reactor nearby to adjust the reactive power balance.

[0119] S350: The positions where the reactive power may be imbalanced are the heavy-load substations with multiple incoming and outgoing lines and active power exchange, new energy access substations, direct current converter stations, and nearby regional transmission channels. Check whether the reactive power compensation configuration of these positions is sufficient. If the reactive power compensation configuration is insufficient, switch the nearby capacitor or reactor.

[0120] S40: If the grid system calculation does not converge, the multi-agent of reinforcement learning adjusts by cooperation, respectively adjusts the active power in the region, and switches the reactive power compensation device of the reactive power compensation node in the region. Specifically, it includes:

[0121] S410: Statistically analyze the active power in the zone, and check whether the active power exchange between regions is too large. If the exchange power is out of range, increase the start of the receiving end zone and reduce the start of the sending end zone, thereby reducing the transmission power of the tie line.

[0122] S420: Check whether the active power borne by the balancing machine is reasonable. If the active output of the balancing machine exceeds the bearing range, increase the generator output to make the active output of the balancing machine within the reasonable range.

[0123] S430: When adjusting, it is also necessary to check whether the generator output is out of range.

[0124] S440: Estimate the reactive power shortage according to the voltage level or zone, and switch the capacitor or reactor nearby to adjust the reactive power balance.

[0125] Further, the active power adjustment mode includes active output data of the regional generator set, and the switching mode of the reactive power compensation device includes the number of connected capacitors / reactors;

[0126] The action space of the deep reinforcement learning multi-agent is represented as: the action space, i.e., the adjustment variable of the reinforcement learning agent, includes the variables of the adjustable active power and reactive power in the system. In order to reduce the action space, the range of active and reactive adjustment is specified. The adjustable variables mainly include the output change amount or effectiveness of the generator, and the effectiveness of the specific capacitor / reactor. Here, the specific capacitor / reactor refers to the capacitor and reactor that is conducive to adjusting the reactive imbalance of the system, specifically the un-put-in capacitor and the put-in reactor near the heavy load and multi-outlet bus. Therefore, the action space can be constructed as shown in the formula:

[0127]

[0128] In the formula: ΔP G is the active power change amount of the generator; N G is the total number of generators; is the reactive power change amount;

[0129]

[0130] In the formula: β is the effectiveness of the capacitor and reactor, β=1 for capacitor switching, and β=-1 for reactor switching; C is the capacitor model, and Z is the reactor model;

[0131] The present application obtains the power flow data of the power grid system, and can make the active power flow adjustment and the reactive power flow adjustment directional.

[0132] Embodiment 2

[0133] Example analysis

[0134] (1) Description of the scene not converging

[0135] In this experiment, BPA is used for power flow calculation, and a regional power grid is used as the experimental scene to generate a large number of power flow calculation non-converging scenes as training data sets.

[0136] With the increase of the number of scenes, the success rate of the strategy adjustment power flow calculation presents a downward trend. After preliminary analysis, it is speculated that this phenomenon may be closely related to the complexity of the scene. In some complex scenes, the overall load level is too high, and the power output of the local power plant is unreasonable, which may lead to source-load imbalance. This imbalance increases the difficulty of the agent to achieve power flow calculation convergence through strategy adjustment. In some complex scenes, the strategy adjustment of the agent is unsuccessful, which may be because the constraint conditions and operating environment in these scenes put higher requirements on the decision-making ability of the agent.

[0137] During the test, 20 different load total size scenes with the same parameter settings were analyzed in depth to evaluate the influence of load parameter changes on the power flow calculation convergence rate, and the results are as followsFigure 3 As shown.

[0138] As shown in the figure, changes in the overall load factor have a significant impact on the system's convergence performance: when the load factor is close to 1, the system exhibits the highest convergence rate, indicating that the system's stability and efficiency are optimized within this range. However, when the load factor exceeds 1.1, the system's convergence rate decreases significantly. This trend suggests that excessively high or low load levels pose greater challenges to power flow calculations, increasing the difficulty of the system reaching steady state. Therefore, maintaining the load factor within an appropriate range can ensure the efficient and reliable operation of the power system.

[0139] (2) Reward function and number of rounds

[0140] In this study, the dataset was divided into training and test sets during training to facilitate effective model evaluation. Specifically, the training set contained 400 different scenarios, which underwent 80,000 rounds of iterative training to ensure the model could learn from diverse data. During agent training, two agents respectively adjusted active and reactive power, each adjusting for 10 steps. Adjustment stopped when the active or reactive power imbalance reached zero, proceeding to the next step until the power flow calculation converged. If the agent converged the scenario adjustment within the specified number of steps, the round ended early; otherwise, the agent received a penalty for unsuccessful policy adjustment. Figure 4 The results shown indicate that the two reward function curves representing the multi-agent system converge at approximately 27,000 rounds for the active power regulation agent and approximately 36,000 rounds for the reactive power regulation agent. This phenomenon suggests that the agents have found an effective strategy at this stage, enabling the power flow calculation to converge stably.

[0141] (2) Test results of part of the test set

[0142] Table 1 shows the test results for part of the test set.

[0143]

[0144]

[0145] In the performance evaluation phase of this study, 100 typical scenarios from the test set that initially did not converge in power flow calculation were selected for testing. Table 1 above shows some of the test results. These scenarios represent complex situations that may be encountered in practical applications, especially under conditions where the load fluctuation range is between 0.95 and 1.15 parameters. The test verifies the adaptability and adjustment efficiency of the agent under different overall load size conditions. The experimental results show that in each selected scenario, the agent system can achieve power flow calculation convergence within one round through no more than 20 steps of decision adjustment. In addition, through the optimization adjustment of the agent, the active power imbalance and reactive power imbalance in all scenarios are successfully adjusted to zero. This result verifies the effectiveness and reliability of the multi-agent system in complex power system environment, and the ability of the multi-agent system in responding to and solving the source-load imbalance problem. This ability is of great significance for improving the stability and efficiency of the power system.

Claims

1. A method for improving the convergence of power system load flow calculation based on deep reinforcement learning multi-agent collaboration, characterized in that, The method comprises the following steps: S10: obtaining power grid data of a power grid system, the power grid data comprising active power, reactive power balance, generator output, switching state of node capacitors and reactors, and balancing machine output; S20: determining active power adjustment range of power plants in the region and switching range of node capacitors and reactors according to the power grid data through a deep reinforcement learning algorithm; S30: performing power flow calculation on the power grid system, and performing active power balance judgment and reactive power balance judgment; S40: if the power grid system calculation does not converge, then the deep reinforcement learning multi-agent adjusts the active power of the power plants in the region and the switching of the reactive power compensation devices of the reactive power compensation nodes in the region through collaborative adjustment; The deep reinforcement learning algorithm constructs a single-agent state space according to the power grid system power flow state space, active power flow adjustment state space, capacitor and reactor switching state, balancing machine state space, important variables of active power and reactive power balance, and generator output state, and jointly constitutes a joint state space with the states of multiple agents; the rewards of multiple agents constitute a comprehensive reward function; the action space is constructed according to the output change or effectiveness of the generator and the effectiveness of specific capacitors and reactors; the specific capacitors and reactors refer to capacitors and reactors that are conducive to adjusting the reactive imbalance of the system; The step S20 comprises: S210: taking the power grid data as input data of the deep reinforcement learning multi-agent, and setting an initial state set; S220: setting the joint state space and comprehensive reward value of the deep reinforcement learning algorithm multi-agent; The single-agent state space comprises the power grid system power flow state space, active power flow adjustment state space, reactive power flow adjustment state space, and balancing machine state space; the comprehensive reward value comprises a calculation result reward value and a constraint condition reward value; S230: determining the active power adjustment power plant range, reactive power compensation node, and reactive power compensation device switching quantity in the region according to the initial state set; The setting of the comprehensive reward value is as follows: (1) Calculation result reward value: reward value setting according to the power flow calculation result, and the specific parameter setting is as follows: (2) Constraint condition reward value: reward value setting through the constraint condition of rated power limit, and the specific parameter setting is as follows: where: P sh is the single balancing machine active power output, P Gj is the single generator active power output, P Li is the actual line active power; P Gjmax , P Gjmin is the generator power upper and lower limits; P shmax , P shmin is the balancing machine power upper and lower limits; P Limin , P Limax line power upper and lower limits; The comprehensive reward value R of the deep reinforcement learning algorithm multi-agent is: R=r1+r2+r3+r4+r5; The collaborative adjustment in the step S40 is the collaborative adjustment of active power and reactive compensation, and the active power of the power plants in the adjustment region and the number of capacitor and inductor (reactor) connections are adjusted according to the balancing output; The action space of the deep reinforcement learning multi-agent is as follows: where: ΔP G is the active power change of the generator; β is the effectiveness of the capacitor and reactor, β = 1 for capacitor switching, β = -1 for reactor switching; N G is the total number of generators; is the reactive power change, C is the capacitor model, and Z is the reactor model. The generator output is discretized, and the following can be obtained: In the formula: Pmax is the generator output upper limit, ΔP G Pmax is the generator output upper limit, ΔP 2. The power system power flow calculation convergence enhancement method according to claim 1, wherein, The setting of the joint state space is as follows: (1) Power grid system power flow state space: P L = {P L1 , P L2 , P L3 ,..., P Li} where: P L P is the set of all grid system line active powers, P Li Pi is the actual active power output of line i, i is the line number; (2) Active power flow adjustment state space: P G = {P G1 , P G2 , P G3 ,..., P Gj} where: P G P is the set of all power plant unit active power outputs, P Gj Pj is the individual generator active power, j is the generator number; (3) Reactive power flow adjustment state space: wherein: δ CZ is the set of all node reactive capacitor and reactor switching quantities, is the switching quantity of the capacitor and reactor of the kth node; (4) The balancing machine state space: P s = {P s1 , P s2 , P s3 ,..., P sh} where: P s P is the set of all balanced unit outputs, P sh Ph is the active output of a single balanced unit, h is the balanced unit number; The state space of the single agent is: S n = {P L , P G , δ CZ , P s}; The joint state space of the multi-agent is: S = {S 1 ,S 2 ,S 3 ,...,S n ,..,S N}; Wherein: S is the set of all agent states, n is the agent number, and N is the total number of multi-agents.

3. The method of claim 1, wherein the method further comprises: Before step S30, the following operation is performed: the network structure data of the power grid system is simplified, and the network structure data of the power grid system in the region to be analyzed is retained.

4. A power system power flow calculation convergence improvement device using the power system power flow calculation convergence improvement method according to any one of claims 1 to 3, characterized by Comprise: A data acquisition module acquires power grid data of the power grid system, including active power, reactive power balance, generator output, switching state of node capacitors and reactors, and balancing machine output; An environment construction module sets a calculation environment of the deep reinforcement learning algorithm according to the power grid data, including a state space construction unit, an action space construction unit, and a reward value construction unit; A power flow calculation module performs power flow calculation on the power grid system, judges active power balance and reactive power balance; if the power grid system calculation does not converge, the deep reinforcement learning multi-agent adjusts the active power of the generator unit in the region and the switching of the reactive power compensation device of the reactive power compensation node in the region through collaborative adjustment; The state space construction unit constructs a single agent state space according to the power grid system power flow state space, active power flow adjustment state space, capacitor and reactor switching state, balancing machine state space, important variables of active power and reactive power balance, and generator output state, and a joint state space is formed by the states of multiple agents; The action space construction unit constructs an action space according to the output change amount or effectiveness of the generator, and the effectiveness of specific capacitors and reactors; the specific capacitors and reactors refer to capacitors and reactors that are conducive to adjusting system reactive imbalance; The reward value construction unit is used to construct the comprehensive reward value of the deep reinforcement learning algorithm multi-agent, including a calculation result reward value and a constraint condition reward value.

5. The power system power flow calculation convergence improvement apparatus according to claim 4, characterized by, The construction of the joint state space satisfies the following formula: (1) Power grid system power flow state space: P L = {P L1 , P L2 , P L3 ,..., P Li} where: P L is the set of all grid system line active powers, P Li is the actual active power output of line i, i is the line number; (2) Active power flow adjustment state space: P G = {P G1 , P G2 , P G3 ,..., P Gj} where: P G P is the set of all power plant unit active power outputs, P Gj Pj is the individual generator active power, j is the generator number; (3) Reactive power flow adjustment state space: wherein: δ CZ is the set of all node reactive capacitor and reactor switching quantities, is the switching quantity of the capacitor and reactor of the kth node; (4) The balancing machine state space: P s = {P s1 , P s2 , P s3 ,..., P sh} where: P s is the set of all balanced unit outputs, P sh is the active output of a single balanced unit, h is the balanced unit number; The state space of the single agent is: S n = {P L , P G , δ CZ , P s}; The joint state space of the multi-agent is: S = {S 1 ,S 2 ,S 3 ,...,S n ,..,S N}; Wherein: S is the set of all agent states, n is the agent number, and N is the total number of multi-agents.

6. The power system power flow calculation convergence improvement apparatus according to claim 5, wherein, The construction of the comprehensive reward value satisfies the following formula: (1) Calculation result reward value: The reward value is set according to the power flow calculation result, and the specific parameter setting is as follows: (2) Constraint condition reward value: The reward value is set through the constraint condition of rated power limit, and the specific parameter setting is as follows: where: P sh is the single balancer active power output, P Gj is the single generator active power output, P Li is the actual line active power; P Gjmax , P Gjmin are the generator active power upper and lower limits; P shmax , P shmin are the balancer active power upper and lower limits; P Limin , P Limax are the line active power upper and lower limits; The comprehensive reward value R of the deep reinforcement learning algorithm multi-agent is: R=r1+r2+r3+r4+r5。 7. The power system power flow calculation convergence improvement apparatus according to claim 6, characterized by, The construction of the action space satisfies the following formula: where ΔP G is the active power change of the generator; β is the effectiveness of the capacitor and reactor, β = 1 for capacitor switching, β = -1 for reactor switching; N G is the total number of generators; is the reactive power change, C is the capacitor model, and Z is the reactor model. Discretize the generator output to obtain: In the formula: Pmax is the generator output upper limit, ΔP G Pmax is the generator output upper limit, ΔP

Citation Information

Patent Citations

  • Automatic adjustment method and device for load flow calculation convergence

    CN111209710A

  • Power system operation mode intelligent generation method based on deep reinforcement learning

    CN115912367A