Load flow calculation convergence adjustment method and device and computer equipment

By constructing a Markov decision-making model based on active adjustment of trend calculation convergence and using PPO deep reinforcement learning algorithm for solving, the problem of slow convergence adjustment of trend calculation is solved, and faster and more efficient convergence adjustment is achieved to ensure the stable operation of the power system.

CN119994915APending Publication Date: 2025-05-13SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411982605.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When the prior art adjusts the convergence of trend calculations, the adjustment speed is slow and it is difficult to meet the real-time needs of the power system.

Method used

By obtaining the degree of non-convergence of the power system due to unreasonable active distribution, a Markov decision-making model for calculating convergence based on active adjustment is constructed, and a near-end strategy optimization PPO deep reinforcement learning algorithm is used to solve the model to obtain a convergence adjustment strategy.

Benefits of technology

The convergence adjustment speed of trend calculation is improved, and the active convergence can be adjusted in time according to changes in the power system state, ensuring the stable operation of the power system, and reducing the training difficulty and sample data requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119994915A_ABST
    Figure CN119994915A_ABST
Patent Text Reader

Abstract

The invention relates to a load flow calculation convergence adjustment method and device and computer equipment. The method comprises the steps of obtaining the non-convergence degree of the power system due to unreasonable active power distribution under the condition that the convergence of the power flow calculation of the power system does not reach the standard, and constructing a Markov decision model based on the convergence of the active power adjustment power flow calculation according to the non-convergence degree, and solving the Markov decision model by adopting a near-end strategy optimization PPO deep reinforcement learning algorithm to obtain a convergence adjustment strategy of load flow calculation, and adjusting the convergence of the load flow calculation according to the convergence adjustment strategy. By adopting the method, the active convergence adjustment strategy can be timely adjusted according to the state change of the power system, so that the convergence of load flow calculation is improved, and the stable operation of the power system is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of power grid technology, and in particular to a method, device and computer equipment for adjusting convergence of power flow calculation. Background Art

[0002] With the development of power electronics technology, medium- and long-term planning of power systems is crucial to ensure the reliability, economy and sustainability of power supply. Among them, power flow calculation can provide basic data for medium- and long-term planning of power systems. Therefore, improving the convergence of power flow calculation is particularly important for power systems.

[0003] In the related art, the convergence of the power flow calculation is mainly adjusted based on empirical knowledge.

[0004] However, in the process of adjusting the convergence of power flow calculation in the related art, there is a problem of slow adjustment speed. Summary of the invention

[0005] Based on this, it is necessary to provide a method, device and computer equipment for adjusting the convergence of power flow calculation in order to solve the above technical problems, so as to improve the adjustment speed of the convergence in the process of adjusting the convergence of power flow calculation.

[0006] In a first aspect, an embodiment of the present application provides a method for adjusting convergence of power flow calculation, the method comprising:

[0007] When the convergence of power flow calculation of the power system does not meet the standard, obtain the non-convergence degree of the power system due to unreasonable active power distribution;

[0008] According to the degree of non-convergence, a Markov decision model based on the convergence of active power flow calculation is constructed;

[0009] The proximal strategy optimization PPO deep reinforcement learning algorithm is used to solve the Markov decision model and obtain the convergence adjustment strategy of power flow calculation;

[0010] According to the convergence adjustment strategy, the convergence of the power flow calculation is adjusted.

[0011] In one embodiment, the method further comprises:

[0012] Obtaining the convergence of each target node in the power system; the target node is any node in the power system except the balance node;

[0013] When the convergence of at least one target node among the target nodes is not converged, it is determined that the convergence of the power flow calculation of the power system does not meet the standard.

[0014] In one embodiment, obtaining the convergence of each target node in the power system includes:

[0015] Obtaining the active power mismatch and reactive power mismatch of each target node;

[0016] For any target node, if the absolute value of the active mismatch amount and the absolute value of the reactive mismatch amount of the target node are both less than or equal to the preset convergence domain, the convergence of the power flow calculation of the target node is determined to be converged;

[0017] Otherwise, the convergence of the power flow calculation of the target node is determined to be non-convergent.

[0018] In one embodiment, obtaining the non-convergence degree of the power system due to unreasonable active power distribution includes:

[0019] According to the active power mismatch of each target node in the power system, a maximum active power mismatch is obtained; according to the convergence of the power flow calculation of each target node, a total number of non-convergent nodes in each target node is obtained; and a total active power mismatch of each target node is obtained;

[0020] The degree of non-convergence is determined based on the maximum active power mismatch, the total number of non-convergence nodes, and the sum of active power mismatch.

[0021] In one embodiment, according to the non-convergence degree, a Markov decision model based on the convergence of active power flow adjustment calculation is constructed, including:

[0022] Constructing a state space set of the power system and constructing an action space set for the intelligent agent to adjust the real-time output of the target generator unit in the power system; the target generator unit is the generator unit selected by the intelligent agent to perform the action;

[0023] Construct a reward and penalty function to adjust the agent's search direction based on the degree of non-convergence;

[0024] The state space set, action space set and reward and penalty function are determined as a Markov decision model.

[0025] In one embodiment, constructing a state space set of a power system includes:

[0026] Obtain a complete set of observed variables related to generators and balancing machines in the power system; obtain a set of mismatches in the power system; and obtain the total active regulation power allocated to each generator set in the power system in a single step during the intelligent agent's decision-making process;

[0027] The complete set, the mismatched quantity set and the total active regulated power are determined as a state space set.

[0028] In one embodiment, constructing an action space set for an intelligent agent to adjust the real-time output of a target generator set in a power system includes:

[0029] Determine the total amount of active power that needs to be adjusted by the target generator set according to the change in output power allocated to each generator in the target generator set;

[0030] According to the total active power, the action space set is determined.

[0031] In one embodiment, a reward and penalty function for adjusting the search direction of the agent is constructed according to the degree of non-convergence, including:

[0032] Constructing a convergence judgment reward function; constructing a non-convergence reduction reward function according to the degree of non-convergence; and constructing a safety constraint penalty function according to the difference between the real-time output of each generator in the power system before the action and the corresponding power lower limit and the difference between the real-time output of each generator after the action and the corresponding power lower limit;

[0033] The convergence judgment reward function, the non-convergence reduction reward function and the safety constraint penalty function are determined as the reward and penalty functions.

[0034] In a second aspect, an embodiment of the present application provides a device for adjusting convergence of power flow calculation, the device comprising:

[0035] The first acquisition module is used to obtain the degree of non-convergence of the power system due to unreasonable active power distribution when the convergence of the power flow calculation of the power system does not meet the standard;

[0036] A model building module is used to build a Markov decision model for convergence of active power flow adjustment calculation according to the degree of non-convergence;

[0037] The model solving module is used to solve the Markov decision model using the proximal strategy optimization PPO deep reinforcement learning algorithm to obtain the convergence adjustment strategy for power flow calculation;

[0038] The adjustment module is used to adjust the convergence of the power flow calculation according to the convergence adjustment strategy.

[0039] In a third aspect, an embodiment of the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method of any one of the embodiments of the first aspect when executing the computer program.

[0040] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method of any embodiment in the above-mentioned first aspect are implemented.

[0041] In a fifth aspect, an embodiment of the present application further provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method of any embodiment of the first aspect above.

[0042] The present application provides a method, device and computer equipment for adjusting the convergence of power flow calculation, including: when the convergence of power flow calculation of the power system does not meet the standard, obtaining the degree of non-convergence of the power system due to unreasonable active power distribution, and constructing a Markov decision model based on the convergence of active power flow calculation according to the degree of non-convergence, using the proximal strategy optimization PPO deep reinforcement learning algorithm to solve the Markov decision model, obtain the convergence adjustment strategy of power flow calculation, and adjust the convergence of power flow calculation according to the convergence adjustment strategy; the above method constructs a Markov decision model based on the convergence of active power flow calculation according to the degree of non-convergence of the power system due to unreasonable active power distribution, so that the constructed Markov decision model can well capture the dynamic characteristics of the power system and accurately It can reflect the actual operation of the power system. On this basis, it can timely adjust the active convergence adjustment strategy according to the changes in the power system state to improve the convergence of the power flow calculation and ensure the stable operation of the power system. At the same time, the above method solves the Markov decision model by optimizing the PPO deep reinforcement learning algorithm through the proximal strategy adopted, which can reduce the amount of sample data required in the solution process, speed up the acquisition of the convergence adjustment strategy, and reduce the difficulty of training, thereby improving the adjustment speed of the convergence of the power flow calculation and alleviating the subsequent reactive power regulation pressure. On this basis, it is further conducive to building a safer, more reliable and economical power system. In addition, the above method is suitable for the long-term planning scenarios of the power system, thereby improving the wide applicability of the convergence adjustment method of the power flow calculation. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in one embodiment;

[0044] Figure 2 A comparison diagram of the effect of adjusting the convergence of power flow calculation in an embodiment;

[0045] Figure 3 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in another embodiment;

[0046] Figure 4 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in another embodiment;

[0047] Figure 5 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in another embodiment;

[0048] Figure 6 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in another embodiment;

[0049] Figure 7 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in another embodiment;

[0050] Figure 8 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in another embodiment;

[0051] Fig. 9 A schematic diagram of a flow chart of a method for adjusting convergence of power flow calculation in another embodiment;

[0052] Fig.10 It is a structural block diagram of a power flow calculation convergence adjustment device in one embodiment;

[0053] Fig.11 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0054] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] In the field of power grids, medium- and long-term planning of power systems is crucial to ensuring the reliability, economy and sustainability of power supply. Among them, power flow calculation can provide basic data for the medium- and long-term planning of power systems. Therefore, improving the convergence of power flow calculation is particularly important for power systems. In the related art, the convergence of power flow calculation is mainly adjusted based on empirical knowledge. However, in the process of adjusting the convergence of power flow calculation in the related art, there is a problem of slow adjustment speed. Based on this, an embodiment of the present application provides a method for adjusting the convergence of power flow calculation, which can improve the adjustment speed of convergence in the process of adjusting the convergence of power flow calculation.

[0056] The power flow calculation convergence adjustment method provided in the embodiment of the present application can be applied to a computer device. The computer device can be, but is not limited to, a personal computer, a desktop computer, a laptop computer, a tablet computer, a smart phone, and other electronic devices with data processing functions. The following embodiment uses a computer device as the execution subject of the power flow calculation convergence adjustment method to introduce the specific process of the power flow calculation convergence adjustment method.

[0057] like Figure 1 FIG. 1 is a flow chart of a method for adjusting convergence of power flow calculation provided in an embodiment of the present application. The method can be implemented by the following steps:

[0058] S100. When the convergence of the power flow calculation of the power system does not meet the standard, obtain the degree of non-convergence of the power system due to unreasonable active power distribution.

[0059] In practical applications, computer equipment can realize power flow calculation, and then apply the power flow calculation results to the power system, call the convergence detection tool to detect the power system, and determine the convergence of the power flow calculation of the power system based on the detection results.

[0060] Specifically, when it is determined that the convergence of the power flow calculation of the power system does not meet the standard, the computer device can call a non-convergence degree acquisition instruction to obtain the non-convergence degree of the power system due to unreasonable active power distribution.

[0061] S200: According to the degree of non-convergence, a Markov decision model based on convergence of active power adjustment flow calculation is constructed.

[0062] The computer device may pre-train an algorithm model, and then input the non-convergence degree into the algorithm model, and the algorithm model constructs a Markov decision model based on the convergence of active power flow adjustment calculation according to the non-convergence degree.

[0063] S300, the proximal strategy optimization PPO deep reinforcement learning algorithm is used to solve the Markov decision model and obtain the convergence adjustment strategy of the power flow calculation.

[0064] Optionally, the computer device can solve the Markov decision model using dynamic programming, strategy iteration, value iteration, linear programming, Monte Carlo or time difference learning to obtain a convergence adjustment strategy for power flow calculation. The convergence adjustment strategy may include information such as the adjustment amount of parameters such as the output, terminal voltage, capacity and power factor of the generator in the power system, the transformation ratio of the transformer and the short-circuit impedance.

[0065] In an embodiment of the present application, a computer device may use a proximal strategy optimization PPO deep reinforcement learning algorithm to solve the Markov decision model and obtain a convergence adjustment strategy for power flow calculation.

[0066] S400. Adjust the convergence of the power flow calculation according to the convergence adjustment strategy.

[0067] Furthermore, the convergence of the power flow calculation can be adjusted based on the convergence adjustment strategy obtained in the previous steps. Figure 2 The figure shows the effect of the changing process of adjusting the convergence of the power flow calculation.

[0068] The technical solution in the embodiment of the present application, when the convergence of the power flow calculation of the power system does not meet the standard, obtains the degree of non-convergence of the power system due to unreasonable active power distribution, and constructs a Markov decision model based on the convergence of the active power flow calculation according to the degree of non-convergence, and adopts the proximal strategy optimization PPO deep reinforcement learning algorithm to solve the Markov decision model, obtains the convergence adjustment strategy of the power flow calculation, and adjusts the convergence of the power flow calculation according to the convergence adjustment strategy; the above method constructs a Markov decision model based on the convergence of the active power flow calculation according to the degree of non-convergence of the power system due to unreasonable active power distribution, so that the constructed Markov decision model can well capture the dynamic characteristics of the power system and accurately reflect the actual operation status, on this basis, the active power convergence adjustment strategy can be adjusted in time according to the changes in the power system state to improve the convergence of the power flow calculation and ensure the stable operation of the power system; at the same time, the above method solves the Markov decision model by optimizing the PPO deep reinforcement learning algorithm through the proximal strategy adopted, which can reduce the amount of sample data required in the solution process, speed up the acquisition of the convergence adjustment strategy, and reduce the difficulty of training, thereby improving the adjustment speed of the convergence of the power flow calculation and alleviating the subsequent reactive power regulation pressure. On this basis, it is further conducive to building a safer, more reliable and economical power system; in addition, the above method is suitable for the medium- and long-term planning scenarios of the power system, thereby improving the wide applicability of the convergence adjustment method of the power flow calculation.

[0069] The following is a description of the process of determining whether the convergence of the power flow calculation of the power system meets the standard. In one embodiment, before executing the steps in S100 above, Figure 3 As shown, the above method may further include the following steps:

[0070] S500: Obtain convergence of each target node in the power system, wherein the target node is any node in the power system except the balancing node.

[0071] In an embodiment of the present application, the target node may be any node in the power system except a balancing node, that is, a PQ node or a PV node.

[0072] Specifically, the computer device may call a convergence detection tool to detect the convergence of each target node in the power system.

[0073] S600: When at least one target node among the target nodes has non-convergence, determine that the convergence of the power flow calculation of the power system does not meet the standard.

[0074] In practical applications, based on the convergence of each target node obtained in the previous steps, it can be determined whether there is a target node among the target nodes whose convergence is not convergent. When it is determined that there is at least one target node among the target nodes whose convergence is not convergent, it is determined that the convergence of the power system flow calculation does not meet the standard.

[0075] In one embodiment, if Figure 4 As shown, the step of obtaining the convergence of each target node in the power system in the above S500 can be implemented in the following manner:

[0076] S510: Obtain active power mismatch and reactive power mismatch of each target node.

[0077] Among them, for any target node, the active power mismatch of the target node is p It can be determined by the difference between the calculated active power value injected by the target node and the actual value of the active power initially given; the reactive power mismatch of the target node mis q It can be determined by the difference between the calculated value of reactive power injected by the target node and the actual value of reactive power given initially.

[0078] Among them, in the process of adjusting the convergence of the power flow calculation, the modified equation of the Newton-Raphson method can be analyzed first. According to the analysis results, in order to obtain the convergence of each target node, it is necessary to determine the active mismatch and reactive mismatch of each target node.

[0079] S520. For any target node, if the absolute value of the active mismatch and the absolute value of the reactive mismatch of the target node are both less than or equal to the preset convergence domain, the convergence of the power flow calculation of the target node is determined to be convergent; otherwise, the convergence of the power flow calculation of the target node is determined to be non-convergent.

[0080] In practical applications, the above preset convergence domain may be user-defined or determined based on historical experience values.

[0081] Specifically, for any target node, the absolute value of the active power mismatch of the target node can be determined: Is it less than or equal to the preset convergence domain, and judge the absolute value of the reactive power mismatch of the target node Whether it is less than or equal to the convergence domain, if the absolute value of the active mismatch amount and the absolute value of the reactive mismatch amount of the target node are both less than or equal to the preset convergence domain and are both less than or equal to the convergence domain, the convergence of the power flow calculation of the target node is determined to be converged; otherwise, the convergence of the power flow calculation of the target node is determined to be non-convergent.

[0082] The technical solution in the embodiment of the present application obtains the active mismatch and reactive mismatch of each target node, and for any target node, if the absolute value of the active mismatch and the absolute value of the reactive mismatch of the target node are both less than or equal to the preset convergence domain, then the convergence of the power flow calculation of the target node is determined to be converged, otherwise, the convergence of the power flow calculation of the target node is determined to be non-convergent; the above method does not require the participation of complex algorithms in the process of determining the convergence of the power flow calculation of the target node, so that the processing process is relatively simple, and can speed up the determination speed of the convergence of the power flow calculation of the target node, improve the determination efficiency of the convergence of the power flow calculation of the target node, and the processing process does not involve a large amount of processed data, thereby also improving the accuracy of the convergence of the power flow calculation of the target node determined by the processing process.

[0083] The following is an explanation of the process of obtaining the non-convergence degree of the power system due to unreasonable active power distribution. Figure 5 As shown, the step of obtaining the degree of non-convergence of the power system due to unreasonable active power distribution in the above S100 may include:

[0084] S110, obtaining a maximum active power mismatch according to the active power mismatch of each target node in the power system; obtaining a total number of non-convergent nodes in each target node according to the convergence of the power flow calculation of each target node; and obtaining a sum of active power mismatches of each target node.

[0085] Specifically, the computer device can compare or find the maximum value of the active mismatch amount of each target node in the power system to obtain the maximum active mismatch amount among the active mismatch amounts of each target node. At the same time, the number of non-convergent nodes in each target node, that is, the total number of non-convergent nodes, can be obtained according to the convergence of the power flow calculation of each target node. In addition, the active mismatch amount of each target node can be summed to obtain the total active mismatch amount.

[0086] S120: Determine the degree of non-convergence according to the maximum active power mismatch amount, the total number of non-convergence nodes and the sum of active power mismatch amounts.

[0087] In the embodiment of the present application, the computer device may determine the maximum active power mismatch amount, the total number of non-convergence nodes and the sum of active power mismatch amounts as the degree of non-convergence of the power system due to unreasonable active power distribution.

[0088] The technical solution in the embodiment of the present application obtains the maximum active mismatch amount according to the active mismatch amount of each target node in the power system; obtains the total number of non-convergence nodes in each target node according to the convergence of the power flow calculation of each target node, and obtains the sum of the active mismatch amounts of each target node, and determines the degree of non-convergence according to the maximum active mismatch amount, the total number of non-convergence nodes and the sum of the active mismatch amounts; the above method can obtain the degree of non-convergence of the power system due to unreasonable active distribution, so that the convergence of the power flow calculation of the power system can be accurately adjusted based on the degree of non-convergence, which can greatly improve the convergence of the power flow calculation.

[0089] The following is an explanation of the process of constructing the Markov decision model based on the convergence of active power flow calculation according to the non-convergence degree. Figure 6 As shown, the step of constructing a Markov decision model based on the convergence of active power flow adjustment calculation according to the non-convergence degree in the above S200 may include:

[0090] S210, constructing a state space set of the power system and constructing an action space set for the intelligent agent to adjust the real-time output of the target generator set in the power system, wherein the target generator set is the generator set selected by the intelligent agent to perform the action.

[0091] Optionally, the power system may include multiple generator sets, and each generator set may include multiple generators.

[0092] In practical applications, the computer device can obtain the instantaneous value or integral value of the voltage and current of the power system, and then construct the state space set of the power system based on the instantaneous value or integral value of the voltage and current of the power system. At the same time, the computer device can construct the action space set for the intelligent agent to adjust the real-time output of the target generator unit in the power system based on the operating variables and value ranges of the range and manner in which each generator in the target generator unit can change its output power.

[0093] S220. Construct a reward and penalty function for adjusting the search direction of the intelligent agent according to the degree of non-convergence.

[0094] In an embodiment of the present application, a reward and penalty function for adjusting the agent's search direction can be constructed according to the maximum active mismatch amount, the total number of non-convergence nodes, and the total active mismatch amount in the non-convergence degree.

[0095] In one embodiment, a method for constructing a reward and penalty function for adjusting the search direction of an intelligent agent according to the maximum active mismatch, the total number of non-convergence nodes, and the total active mismatch in the degree of non-convergence can be to pre-train a reward and penalty function construction model, and then input the maximum active mismatch, the total number of non-convergence nodes, and the total active mismatch into the reward and penalty function construction model, and the reward and penalty function construction model outputs the constructed reward and penalty function for adjusting the search direction of the intelligent agent.

[0096] Optionally, the above-mentioned reward and penalty function construction model can be at least one of a convolutional neural network model, a fully connected neural network model, a long short-term memory neural network model, a recurrent neural network model, a residual neural network model, etc.

[0097] S230, determining the state space set, the action space set and the reward and penalty function as a Markov decision model.

[0098] In an embodiment of the present application, the computer device may determine the state space set, the action space set, and the reward-penalty function as a Markov decision model.

[0099] In one embodiment, if Figure 7 As shown, the step of constructing the state space set of the power system in the above S210 may include:

[0100] S211. Obtain a complete set of observed variables related to generators and balancing machines in the power system; obtain a set of mismatch quantities in the power system; and obtain the total active regulation power allocated to each generator set in the power system in a single step during the intelligent agent decision-making process.

[0101] In the embodiment of the present application, the complete set of observed variables related to the generator in the above power system is , the complete set of observed variables related to the balancing machine in the power system , mismatched quantity set , the total active power regulation power allocated to each generator set in the power system in a single step during the intelligent agent decision process They can be expressed by the following formulas:

[0102] (1)

[0103] in, Represents the total number of light-emitting machines in the power system, Indicates light machine The relevant observed variables;

[0104] (2)

[0105] in, represents the total number of balancing machines in the power system, Indicates balancing machine The relevant observed variables;

[0106] (3)

[0107] in, Represents the active power mismatch of all PQ nodes in the power system, Indicates the total number of PQ nodes in the power system whose active mismatch is greater than the preset convergence threshold. It represents the sum of the active mismatch of all PQ nodes in the power system. Represents the active mismatch of all PV nodes in the power system, Indicates the total number of PV nodes in the power system whose active mismatch is greater than the preset convergence threshold. It represents the sum of active mismatch of all PV nodes in the power system;

[0108] (4)

[0109] in, represents the total number of generating units in the power system, Indicates the power system Generator sets;

[0110] (5)

[0111] (6)

[0112] (7)

[0113] S212: Determine the complete set, the mismatched quantity set and the total active regulated power as a state space set.

[0114] In the embodiment of the present application, the complete set, the mismatch amount set and the total active regulation power may be determined as a state space set.

[0115] In one embodiment, if Figure 8 As shown, the step of constructing an action space set for the intelligent agent to adjust the real-time output of the target generator set in the power system in the above S210 may include:

[0116] S213: Determine the total amount of active power that needs to be adjusted for the target generator set according to the change in output power allocated to each generator in the target generator set.

[0117] In the embodiment of the present application, the generator The change in output power It can be expressed by formula (8):

[0118] (8)

[0119] + +......+ (9)

[0120] in, Indicates the total active power that needs to be adjusted by the target generator set. Indicates the total number of generators in the target generator set.

[0121] S214: Determine an action space set according to the total active power.

[0122] In an embodiment of the present application, the total active power that needs to be adjusted by the target generator set is determined as a set of action spaces for the intelligent agent to adjust the real-time output of the target generator set in the power system.

[0123] In one embodiment, if Fig. 9 As shown, the step of constructing a reward and penalty function for adjusting the agent's search direction according to the degree of non-convergence in the above S220 may include:

[0124] S221. Construct a convergence judgment reward function; construct a non-convergence reduction reward function according to the degree of non-convergence; and construct a safety constraint penalty function according to the difference between the real-time output of each generator in the power system before the action and the corresponding power lower limit, and the difference between the real-time output of each generator after the action and the corresponding power lower limit.

[0125] In the embodiment of the present application, the convergence judgment reward function It can be expressed by formula (10):

[0126] (10)

[0127] At the same time, the computer device can be used to calculate the maximum active mismatch amount in the non-convergence degree according to the non-convergence degree. , Total number of non-convergent nodes And the total active mismatch Perform arithmetic operations to construct a non-convergence reduction reward function ; Optionally, the above arithmetic operation can be implemented by at least one of addition, subtraction, multiplication, logarithm, exponential, division, etc.

[0128] In the embodiment of the present application, the maximum active mismatch amount in the non-convergence degree is , Total number of non-convergent nodes And the total active mismatch It can be processed according to formula (11) to construct a non-convergence reduction reward function :

[0129]

[0130] (11)

[0131] In addition, the computer equipment can calculate the difference between the real-time output of each generator in the power system before it operates and the corresponding power lower limit. And the difference between the real-time output of each generator after operation and the corresponding power lower limit Arithmetic operations are performed to construct a safety constraint penalty function.

[0132] In the embodiment of the present application, the difference between the real-time output of each generator in the power system before operation and the corresponding power lower limit is calculated. And the difference between the real-time output of each generator after operation and the corresponding power lower limit The security constraint penalty function can be constructed by processing according to the following formula, as shown in formulas (12)-(16):

[0133] (12)

[0134] (13)

[0135] (14)

[0136] (15)

[0137] (16)

[0138] S222. Determine the convergence judgment reward function, the non-convergence reduction reward function, and the safety constraint penalty function as a reward and penalty function.

[0139] In the embodiment of the present application, the convergence judgment reward function, the non-convergence reduction reward function and the safety constraint penalty function can be determined as the reward and penalty function.

[0140] The technical solution in the embodiment of the present application constructs a state space set of the power system, and constructs an action space set for an intelligent agent to adjust the real-time output of a target generator unit in the power system, constructs a reward and penalty function for adjusting the search direction of the intelligent agent according to the degree of non-convergence, and determines the state space set, the action space set and the reward and penalty function as a Markov decision model; the above method can convert the power flow calculation convergence problem of the power system into a Markov decision model, making it simple and efficient to improve the convergence of the power flow calculation.

[0141] In one embodiment, the present application also provides a method for adjusting convergence of power flow calculation, which is applied to a computer device. The method includes the following process:

[0142] (1) Obtaining the active power mismatch and reactive power mismatch of each target node; the target node is any node in the power system except the balancing node;

[0143] For any target node, if the absolute value of the active mismatch amount and the absolute value of the reactive mismatch amount of the target node are both less than or equal to the preset convergence domain, the convergence of the power flow calculation of the target node is determined to be converged; otherwise, the convergence of the power flow calculation of the target node is determined to be non-convergent;

[0144] When at least one of the target nodes has a non-convergence, it is determined that the convergence of the power flow calculation of the power system does not meet the standard;

[0145] In the case where the convergence of the power flow calculation of the power system does not meet the standard, the maximum active power mismatch amount is obtained according to the active power mismatch amount of each target node in the power system; the total number of non-convergent nodes in each target node is obtained according to the convergence of the power flow calculation of each target node; and the sum of the active power mismatch amount of each target node is obtained;

[0146] Determine the degree of non-convergence according to the maximum active power mismatch, the total number of non-convergence nodes and the sum of active power mismatch;

[0147] Obtain a complete set of observed variables related to generators and balancing machines in the power system; obtain a set of mismatches in the power system; and obtain the total active regulation power allocated to each generator set in the power system in a single step during the intelligent agent's decision-making process;

[0148] Determine the complete set, the mismatch set and the total active regulation power as a state space set; and

[0149] Determine the total amount of active power that needs to be adjusted by the target generator set according to the change in output power allocated to each generator in the target generator set;

[0150] According to the total active power, determine the action space set of the intelligent agent to adjust the real-time output of the target generator unit in the power system; the target generator unit is the generator unit selected by the intelligent agent to perform the action;

[0151] Constructing a convergence judgment reward function; constructing a non-convergence reduction reward function according to the degree of non-convergence; and constructing a safety constraint penalty function according to the difference between the real-time output of each generator in the power system before the action and the corresponding power lower limit and the difference between the real-time output of each generator after the action and the corresponding power lower limit;

[0152] The convergence judgment reward function, the non-convergence reduction reward function and the safety constraint penalty function are determined as the reward and penalty function;

[0153] The state space set, action space set and reward and penalty function are determined as a Markov decision model;

[0154] The proximal strategy optimization PPO deep reinforcement learning algorithm is used to solve the Markov decision model and obtain the convergence adjustment strategy of power flow calculation;

[0155] According to the convergence adjustment strategy, the convergence of the power flow calculation is adjusted.

[0156] The execution process of the above (1) to (5) can be specifically referred to the description of the above embodiment. The implementation principle and technical effect are similar and will not be repeated here.

[0157] It should be understood that, although the steps in the flowcharts involved in the above embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps is not strictly limited in order, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0158] Based on the same inventive concept, the embodiment of the present application also provides a power flow calculation convergence adjustment device for implementing the power flow calculation convergence adjustment method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more power flow calculation convergence adjustment device embodiments provided below can refer to the limitations of the power flow calculation convergence adjustment method above, and will not be repeated here.

[0159] In one embodiment, Fig.10 This is a schematic diagram of the structure of a power flow calculation convergence adjustment device in one embodiment of the present application. The power flow calculation convergence adjustment device provided in the embodiment of the present application can be applied to computer equipment. Fig.10 As shown, the power flow calculation convergence adjustment device of the embodiment of the present application may include: a first acquisition module 11, a model construction module 12, a model solution module 13 and an adjustment module 14, wherein:

[0160] The first acquisition module 11 is used to acquire the degree of non-convergence of the power system due to unreasonable active power distribution when the convergence of the power flow calculation of the power system does not meet the standard;

[0161] A model building module 12 is used to build a Markov decision model based on the convergence of active power adjustment flow calculation according to the degree of non-convergence;

[0162] The model solving module 13 is used to solve the Markov decision model using the proximal strategy optimization PPO deep reinforcement learning algorithm to obtain the convergence adjustment strategy of the power flow calculation;

[0163] The adjustment module 14 is used to adjust the convergence of the power flow calculation according to the convergence adjustment strategy.

[0164] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0165] In one embodiment, the power flow calculation convergence adjustment device further includes: a second acquisition module and a determination module, wherein:

[0166] The second acquisition module is used to acquire the convergence of each target node in the power system; the target node is any node in the power system except the balancing node;

[0167] The determination module is used to determine that the convergence of the power flow calculation of the power system does not meet the standard when at least one target node among the target nodes has a convergence that is not convergent.

[0168] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0169] In one embodiment, the second acquisition module is specifically used for:

[0170] Obtaining the active power mismatch and reactive power mismatch of each target node;

[0171] For any target node, if the absolute value of the active mismatch amount and the absolute value of the reactive mismatch amount of the target node are both less than or equal to the preset convergence domain, the convergence of the power flow calculation of the target node is determined to be converged;

[0172] Otherwise, the convergence of the power flow calculation of the target node is determined to be non-convergent.

[0173] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0174] In one embodiment, the first acquisition module 11 is specifically used for:

[0175] According to the active power mismatch of each target node in the power system, a maximum active power mismatch is obtained; according to the convergence of the power flow calculation of each target node, a total number of non-convergent nodes in each target node is obtained; and a total active power mismatch of each target node is obtained;

[0176] The degree of non-convergence is determined based on the maximum active power mismatch, the total number of non-convergence nodes, and the sum of active power mismatch.

[0177] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0178] In one embodiment, the model building module 12 includes: a first building unit, a second building unit and a determining unit, wherein:

[0179] The first construction unit is used to construct a state space set of the power system and to construct an action space set for the intelligent agent to adjust the real-time output of the target generator set in the power system; the target generator set is the generator set selected by the intelligent agent to perform the action;

[0180] The second construction unit is used to construct a reward and penalty function for adjusting the search direction of the agent according to the degree of non-convergence;

[0181] The determination unit is used to determine the state space set, the action space set and the reward and penalty function as a Markov decision model.

[0182] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0183] In one embodiment, the first construction unit includes: an acquisition subunit and a first determination subunit, wherein:

[0184] An acquisition subunit is used to acquire a complete set of observed variables related to generators and balancing machines in the power system; to acquire a set of mismatches in the power system; and to acquire the total active regulation power allocated to each generator set in the power system in a single step during the decision-making process of the intelligent agent;

[0185] The first determination subunit is used to determine the complete set, the mismatch amount set and the total active regulated power as a state space set.

[0186] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0187] In one embodiment, the first construction unit includes: a second determination subunit and a third determination subunit, wherein:

[0188] A second determining subunit is used to determine the total amount of active power that needs to be adjusted by the target generator set according to the change in output power allocated to each generator in the target generator set;

[0189] According to the total active power, the action space set is determined.

[0190] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0191] In one embodiment, the second building block is specifically used for:

[0192] Constructing a convergence judgment reward function; constructing a non-convergence reduction reward function according to the degree of non-convergence; and constructing a safety constraint penalty function according to the difference between the real-time output of each generator in the power system before the action and the corresponding power lower limit and the difference between the real-time output of each generator after the action and the corresponding power lower limit;

[0193] The convergence judgment reward function, the non-convergence reduction reward function and the safety constraint penalty function are determined as the reward and penalty functions.

[0194] The power flow calculation convergence adjustment device provided in the embodiment of the present application can be used to execute the technical solution in the power flow calculation convergence adjustment method embodiment described above in the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0195] For the specific definition of the power flow calculation convergence adjustment device, please refer to the definition of the power flow calculation convergence adjustment method mentioned above, which will not be repeated here. Each module in the above power flow calculation convergence adjustment device can be implemented in whole or in part by software, hardware and a combination thereof. The above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0196] In one embodiment, a computer device is also provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Fig.11 As shown. The computer device includes a processor, a memory and a network interface connected through a system bus. Among them, the processor of the computer device is used to provide processing power. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store convergence adjustment strategies. The network interface of the computer device is used to communicate with an external endpoint through a network connection. When the computer program is executed by the processor, a method for adjusting the convergence of power flow calculation is implemented.

[0197] Those skilled in the art will understand that Fig.11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0198] In one embodiment, a computer device is also provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the technical solution in the above-mentioned power flow calculation convergence adjustment method embodiment of the present application is implemented. The implementation principle and technical effect are similar and will not be repeated here.

[0199] In one embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, the technical solution of the above-mentioned power flow calculation convergence adjustment method of the present application is implemented. Its implementation principle and technical effect are similar and will not be repeated here.

[0200] In one embodiment, a computer program product is also provided, including a computer program, which, when executed by a processor, implements the technical solution of the above-mentioned power flow calculation convergence adjustment method of the present application. Its implementation principle and technical effect are similar and will not be repeated here.

[0201] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0202] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0203] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for adjusting convergence of power flow calculation, characterized in that: The method comprises: When the convergence of the power flow calculation of the power system does not meet the standard, obtaining the degree of non-convergence of the power system due to unreasonable active power distribution; According to the non-convergence degree, constructing a Markov decision model for adjusting the convergence of the power flow calculation based on active power; The proximal strategy optimization PPO deep reinforcement learning algorithm is adopted to solve the Markov decision model to obtain the convergence adjustment strategy of the power flow calculation; According to the convergence adjustment strategy, the convergence of the power flow calculation is adjusted.

2. The method according to claim 1, characterized in that The method further comprises: Obtaining convergence of each target node in the power system; the target node is any node in the power system except the balancing node; When the convergence of at least one target node among the target nodes is not converged, it is determined that the convergence of the power flow calculation of the power system does not meet the standard.

3. The method according to claim 2, characterized in that The obtaining of the convergence of each target node in the power system includes: Obtaining the active power mismatch amount and reactive power mismatch amount of each target node; For any target node, if the absolute value of the active mismatch amount and the absolute value of the reactive mismatch amount of the target node are both less than or equal to the preset convergence domain, then the convergence of the power flow calculation of the target node is determined to be converged; Otherwise, it is determined that the convergence of the power flow calculation of the target node is not converged.

4. The method according to any one of claims 1 to 3, characterized in that The obtaining of the non-convergence degree of the power system due to unreasonable active power distribution includes: Obtaining a maximum active mismatch amount according to the active mismatch amount of each target node in the power system; obtaining a total number of non-convergent nodes in each target node according to the convergence of the power flow calculation of each target node; and obtaining a sum of active mismatch amounts of each target node; The non-convergence degree is determined according to the maximum active power mismatch amount, the total number of non-convergence nodes and the sum of the active power mismatch amounts.

5. The method according to any one of claims 1 to 3, characterized in that: The step of constructing a Markov decision model for adjusting the convergence of the power flow calculation based on active power according to the non-convergence degree includes: Constructing a state space set of the power system, and constructing an action space set for an intelligent agent to adjust the real-time output of a target generator set in the power system; the target generator set is a generator set selected by the intelligent agent to perform an action; Constructing a reward and penalty function for adjusting the search direction of the agent according to the non-convergence degree; The state space set, the action space set and the reward-penalty function are determined as the Markov decision model.

6. The method according to claim 5, characterized in that The constructing the state space set of the power system comprises: Obtaining a complete set of observed variables related to generators and balancing machines in the power system; obtaining a set of mismatches in the power system; and obtaining a total active regulation power allocated to each generator set in the power system in a single step during the intelligent agent decision-making process; The complete set, the mismatch amount set and the total active regulated power are determined as the state space set.

7. The method according to claim 5, characterized in that The action space set of constructing the intelligent agent to adjust the real-time output of the target generator set in the power system includes: Determining the total amount of active power that needs to be adjusted for the target generator set according to the change in output power allocated to each generator in the target generator set; The action space set is determined according to the total active power.

8. The method according to claim 5, characterized in that The step of constructing a reward and penalty function for adjusting the agent's search direction according to the non-convergence degree includes: Constructing a convergence judgment reward function; constructing a non-convergence reduction reward function according to the non-convergence degree; and constructing a safety constraint penalty function according to the difference between the real-time output of each generator in the power system before the action and the corresponding power lower limit and the difference between the real-time output of each generator after the action and the corresponding power lower limit; The convergence judgment reward function, the non-convergence reduction reward function and the safety constraint penalty function are determined as the reward and penalty function.

9. A device for adjusting convergence of power flow calculation, characterized in that: The device comprises: A first acquisition module is used to acquire the degree of non-convergence of the power system due to unreasonable active power distribution when the convergence of the power flow calculation of the power system does not meet the standard; A model building module, used for building a Markov decision model for adjusting the convergence of the power flow calculation based on active power according to the non-convergence degree; A model solving module, used to solve the Markov decision model using the proximal strategy optimization PPO deep reinforcement learning algorithm to obtain a convergence adjustment strategy for the power flow calculation; The adjustment module is used to adjust the convergence of the power flow calculation according to the convergence adjustment strategy.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.