Voltage stability margin monitoring method and device, storage medium and computer equipment

By using a policy network and an accelerated convergence determination mechanism in the power system, the P-V curve of the voltage stability margin is quickly tracked, and the problem of complex and time-consuming calculation of the traditional method is solved, real-time monitoring and efficient calculation of the voltage stability margin is achieved.

CN120185205APending Publication Date: 2025-06-20ELECTRIC POWER RES INST CHINA SOUTHERN POWER GRID CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510349933.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The calculation process of the traditional power system voltage stability analysis method is complex and time-consuming, and it is difficult to meet the timeliness requirements of real-time monitoring, resulting in the inability to support the scheduling and operation decisions of the power system in a timely and effective manner.

Method used

By adopting a voltage stability margin monitoring method, by obtaining the current status information of the power system and inputting it into a preset policy network, the optimal action strategy corresponding to the status information is determined, the target P-V curve is tracked, the static stability limit power is determined, and the voltage stability margin is calculated based on this.

Benefits of technology

It realizes the rapid tracking of the target P-V curve while maintaining monitoring accuracy, improves the real-time monitoring capability of voltage stability margin, reduces the number of calculations, and improves the calculation efficiency of monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120185205A_ABST
    Figure CN120185205A_ABST
Patent Text Reader

Abstract

The invention provides a voltage stability margin monitoring method and device, a storage medium and computer equipment, and the method comprises the steps: obtaining the current state information of a power system, inputting the state information into a preset strategy network, and tracking a target P-V curve after an optimal action strategy is executed when the optimal action strategy corresponding to the state information is determined; as the strategy network is obtained by taking the state information of the power system as the state space and taking the increment of the active power as the action space to carry out optimization learning, the tracking speed of the P-V curve can be controlled by selecting the action strategy in the action space, and when the tracking speed is too high, the precision of the tracked P-V curve is reduced. Therefore, the optimal action strategy is sought based on the current state information of the power system, and the tracking speed and the positioning precision are balanced, so that the target P-V curve can be tracked as quickly as possible while the monitoring accuracy is maintained, and the real-time monitoring of the voltage stability margin is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and particularly to a method, device, storage medium, and computer device for monitoring voltage stability margin. Background Art

[0002] With the continuous development and complexity of power systems, voltage stability issues have increasingly become a major challenge in power system operation. Modern power systems include a large number of power generation, transmission, and distribution devices, as well as complex loads and renewable energy access, making the dynamic behavior and stability analysis of the system extremely complex. Voltage stability margin, as a key indicator for evaluating whether a power system can maintain voltage stability under disturbances, is crucial for ensuring the safe and stable operation of the power system.

[0003] However, traditional power system voltage stability analysis methods, such as the continuation power flow method (CPF), although able to provide relatively accurate voltage stability margin assessment, have a complex calculation process and long time consumption. Especially when performing online voltage stability analysis, this method is difficult to meet the timeliness requirements of real-time monitoring. As a result, it is unable to provide timely and effective support for power system dispatching and operation decisions, leading to difficulties in improving the monitoring efficiency while ensuring the accuracy of voltage stability analysis. Summary of the Invention

[0004] The purpose of this application aims to at least solve one of the above technical defects, especially the technical defect that the calculation process of the power system voltage stability analysis method in the prior art is complex and time-consuming. Especially when performing online voltage stability analysis, this method is difficult to meet the timeliness requirements of real-time monitoring. As a result, it is unable to provide timely and effective support for power system dispatching and operation decisions, leading to difficulties in improving the monitoring efficiency while ensuring the accuracy of voltage stability analysis.

[0005] In a first aspect, this application provides a method for monitoring voltage stability margin, the method including:

[0006] Obtain the current state information of the power system;

[0007] Input the state information into a preset policy network, execute the optimal action policy when determining the optimal action policy corresponding to the state information, and trace to obtain the target P-V curve; the policy network is optimized and learned with the state information of the power system as the state space and the growth of the active power of the new energy output nodes in the power system as the action space, and triggers an accelerated convergence determination mechanism to determine whether the P-V curve tracking task is completed when meeting the preset conditions;

[0008] Determine the static stability limit power according to the target P-V curve, and calculate the current voltage stability margin of the power system based on the static stability limit power.

[0009] In one embodiment, the optimization learning process of the policy network includes:

[0010] Determine a preset target network, and iteratively learn the action strategy of the target network during the P-V curve tracking process;

[0011] In each learning process, obtain a target sample set, determine the action strategy of the target network during the P-V curve tracking process according to the target sample set, and execute the action strategy. After determining the reward function based on the actual operating conditions of the power system after executing the action strategy, optimize the target network based on the determined reward function;

[0012] Judge whether the difference between the minimum unconverged power and the maximum converged power in the P-V curve obtained by tracking after executing the action strategy is less than a preset critical value. If it is less, trigger the accelerated convergence determination mechanism. When it is determined that the P-V curve tracking task is completed, determine the latest target network as the policy network, and determine the target P-V curve. When it is determined that the P-V curve tracking task is not completed, enter the next learning.

[0013] In one embodiment, the determination process of the accelerated convergence determination mechanism includes:

[0014] Calculate the first power flow state of the currently executed action strategy. If the first power flow state converges, add a preset convergence threshold to the currently executed action strategy as the first target strategy, and calculate the second power flow state of the first target strategy. If the second power flow state does not converge, determine that the P-V curve tracking task is completed. If the second power flow state converges, determine that the P-V curve tracking task is not completed;

[0015] If the first power flow state does not converge, subtract a preset convergence threshold from the currently executed action strategy as the second target strategy, and calculate the third power flow state of the second target strategy. If the third power flow state converges, determine that the P-V curve tracking task is completed. If the third power flow state does not converge, determine that the P-V curve tracking task is not completed.

[0016] In one embodiment, determining the reward function according to the actual operating conditions of the power system after executing the action strategy includes:

[0017] Determine reward terms regarding the average voltage, power output, power flow convergence state, and tracking task completion degree of the power system according to the actual operating conditions of the power system after executing the action strategy;

[0018] Determine the weights of each reward item according to the optimization objective of the policy network;

[0019] Generate a reward function based on each reward item and its corresponding weight; the reward function takes the increase in output power and the decrease in average voltage as the exploration direction.

[0020] In one embodiment, the obtaining of the target sample set includes:

[0021] Select target samples from a preset experience replay pool according to the sample priorities, and generate a target sample set based on the selected target samples; each sample in the experience replay pool has been marked with a sample priority, and the sample priority is used to quantify the error of the samples in the experience replay pool.

[0022] In one embodiment, the determining of the static stability limit power according to the target P-V curve and calculating the current voltage stability margin of the power system based on the static stability limit power includes:

[0023] Determine the maximum active power point in the target P-V curve, and determine the power corresponding to the maximum active power point as the static stability limit power;

[0024] Obtain the current operating power of the power system, and calculate the current voltage stability margin of the power system according to the operating power and the static stability limit power.

[0025] In one embodiment, the state space is represented as:

[0026]

[0027] Wherein, represents the active power of node n, represents the reactive power of node n, represents the active power load of node n, represents the reactive power load of node n, represents the voltage of node n, represents the voltage phase angle of node n, represents the power flow convergence information.

[0028] In a second aspect, the present application provides a voltage stability margin monitoring device, and the device includes:

[0029] An information acquisition module, configured to acquire the current state information of the power system;

[0030] A curve tracking module, configured to input the state information into a preset policy network, execute the optimal action policy when determining the optimal action policy corresponding to the state information, and track to obtain a target P-V curve; the policy network is optimized and learned with the state information of the power system as the state space and the growth rate of the active power of the new energy output nodes in the power system as the action space, and triggers an accelerated convergence determination mechanism to determine whether the P-V curve tracking task is completed when meeting the preset conditions;

[0031] A margin monitoring module, configured to determine the static stability limit power according to the target P-V curve and calculate the current voltage stability margin of the power system based on the static stability limit power.

[0032] In a third aspect, the present application provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, the one or more processors are caused to execute the steps of the voltage stability margin monitoring method according to any one of the above embodiments.

[0033] In a fourth aspect, the present application provides a computer device, including: one or more processors, and a memory;

[0034] The memory stores computer-readable instructions. When the one or more processors execute the computer-readable instructions, the steps of the voltage stability margin monitoring method according to any one of the above embodiments are executed.

[0035] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0036] The voltage stability margin monitoring method, device, storage medium, and computer device provided by this application. The method includes: obtaining the current state information of the power system, and inputting the state information into a preset policy network. After determining the optimal action policy corresponding to the state information and executing the optimal action policy, the target P-V curve is traced. Since the policy network is optimized and learned with the state information of the power system as the state space and the increase in the active power of the new energy output nodes in the power system as the action space, selecting an action policy in the action space can control the tracking speed of the P-V curve. When the tracking speed is too fast, the accuracy of the traced P-V curve will decrease. Therefore, an optimal action policy is sought based on the current state information of the power system to balance the tracking speed and positioning accuracy, so as to ensure that the target P-V curve can be traced as quickly as possible while maintaining the monitoring accuracy, thereby realizing the real-time monitoring of the voltage stability margin. And during the network convergence process, an accelerated convergence determination mechanism is set to determine whether the P-V curve tracking task is completed. The accelerated convergence determination mechanism will be triggered only when the preset conditions are met for additional verification, avoiding additional verification calculations at the early stage of the tracking task. In this way, convergence can be completed as soon as possible on the premise of reducing the number of calculations, further improving the calculation efficiency of voltage stability margin monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0038] Figure 1 It is a schematic flowchart of a voltage stability margin monitoring method provided by an embodiment of this application;

[0039] Figure 2 It is a schematic flowchart of the optimization learning process of the policy network provided by an embodiment of this application;

[0040] Figure 3 It is a schematic flowchart of the determination process of the accelerated convergence determination mechanism provided by an embodiment of this application;

[0041] Figure 4 It is a comparison example diagram of the accelerated convergence determination mechanism provided by an embodiment of this application;

[0042] Figure 5 It is a schematic flowchart of determining the reward function according to the actual operation situation of the power system after executing the action policy provided by an embodiment of this application;

[0043] Figure 6Schematic structural diagram of a voltage stability margin monitoring device provided by an embodiment of the present application;

[0044] Figure 7 Internal structure diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0046] In one embodiment, the present application provides a method for monitoring voltage stability margin. The following embodiments will be described by taking the application of this method to a server as an example. It can be understood that the server executing the voltage stability margin monitoring method can be a single server or a server cluster composed of multiple servers. The present application does not make specific limitations on this.

[0047] As Figure 1 shown, the present application provides a method for monitoring voltage stability margin, and the method includes:

[0048] S101: Obtain the current state information of the power system.

[0049] In this step, when the user needs to perform real-time detection on the voltage stability margin of the power system, multiple index parameters corresponding to the definition in the preset state space can be obtained, and then the current state information of the power system is generated according to the obtained index parameters.

[0050] Among them, the power system refers to a network composed of power transmission, power generation, power distribution, and power consumption. The index parameters refer to various operation indexes and parameters of the power system, such as active power, reactive power, voltage, current, etc. The state information is used to describe the operation state of the power system, and the state space refers to the set range composed of the state information of the power system.

[0051] S102: Input the state information into a preset policy network to execute the optimal action policy when determining the optimal action policy corresponding to the state information, and trace to obtain the target P-V curve.

[0052] Among them, the policy network is obtained through optimization learning with the state information of the power system as the state space and the increment of the active power of the new energy output nodes in the power system as the action space, and triggers an accelerated convergence determination mechanism to determine whether the P-V curve tracking task is completed when preset conditions are met. It can be understood that the new energy output nodes refer to the nodes in the power system that are connected to renewable energy power generation equipment such as wind energy and solar energy. The state information of the power system mainly includes the power information, load information, voltage information, and power flow state convergence information of each node.

[0053] In this step, the state information is input into a preset policy network. This policy network is obtained through optimization learning with the state information of the power system as the state space and the increment of the active power of the new energy output nodes in the power system as the action space. It can be understood that since the increment of the active power of the new energy output nodes is inversely proportional to the number of iterations required to reach the static stability limit power, that is, when the increment is larger, the number of iterations required to reach the static stability limit power is less. Therefore, by selecting the action strategy in the action space, the speed of tracking the P-V curve can be adjusted. However, when the tracking speed is fast, it is easy to affect the positioning accuracy of the P-V curve. Based on this, the policy network can use the optimization goal of tracking the target P-V curve as quickly as possible while maintaining the tracking accuracy. In the application stage of the policy network, by inputting the current state information of the power system, the policy network can calculate the corresponding optimal action strategy according to this state information, and then execute this optimal action strategy to adjust the active power of each new energy output node, and record the P-V curve, that is, the target P-V curve, when the static stability limit power is tracked.

[0054] In the optimization learning process of the policy network, an accelerated convergence determination mechanism can be triggered to determine whether the P-V curve tracking task is completed when preset conditions are met. It can be understood that the preset conditions can be set as the difference between the minimum unconverged power and the maximum converged power in the tracked P-V curve being less than a preset critical value. When performing the P-V curve tracking task, the power flow state is calculated after each execution of the action strategy to determine whether the static stability limit power is tracked. When the preset conditions are met, it means that it has reached the later stage of this tracking task. In order to quickly reach the static stability limit power, the accelerated convergence determination mechanism can be started, thereby accelerating the convergence speed of the policy network in the optimization learning process. This can also improve the decision-making efficiency of the policy network in application.

[0055] Specifically, the optimal action strategy refers to setting the increment of active power that fits the current state of the power system and takes into account both the tracking speed and accuracy. The static stability limit power refers to the maximum power that the power generation equipment can output when the power system is in a static stable state. The target P-V curve refers to the P-V curve when tracking the static stability limit power, and the P-V curve is a graphical curve used to describe the relationship between voltage (V) and active power (P) in the power system.

[0056] In one example, the action space can be expressed as:

[0057]

[0058] In the formula, represents the action space, is the action strategy, which is a scalar value representing the increment of the active power output of the new energy output node, represents the growth ratio of the active power of each new energy node, and m represents the node.

[0059] S103: Determine the static stability limit power according to the target P-V curve, and calculate the current voltage stability margin of the power system based on the static stability limit power.

[0060] Among them, the voltage stability margin is used to measure the ability of the power system to maintain voltage stability when suffering from disturbances or bearing voltage changes.

[0061] In this step, when tracking the target P-V curve, the static stability limit power can be determined according to the target P-V curve, and then the current voltage stability margin of the power system can be calculated based on the static stability limit power. It can be understood that since the calculation efficiency of the voltage stability margin is shortened each time, it has high real-time performance, so it can also support the real-time monitoring of the voltage stability margin.

[0062] The voltage stability margin monitoring method, device, storage medium and computer device provided by this application. The method includes: obtaining the current state information of the power system, and inputting the state information into a preset policy network. After determining the optimal action policy corresponding to the state information and executing the optimal action policy, the target P-V curve is traced. Since the policy network is optimized and learned with the state information of the power system as the state space and the growth rate of the active power of the new energy output nodes in the power system as the action space, selecting an action policy in the action space can control the tracking speed of the P-V curve. When the tracking speed is too fast, the accuracy of the traced P-V curve will decrease. Therefore, based on the current state information of the power system, the optimal action policy is sought to balance the tracking speed and positioning accuracy, so as to ensure that the target P-V curve can be traced as quickly as possible while maintaining the monitoring accuracy, thereby realizing the real-time monitoring of the voltage stability margin. And during the network convergence process, an accelerated convergence determination mechanism is set to determine whether the P-V curve tracking task is completed. The accelerated convergence determination mechanism will be triggered only when the preset conditions are met for additional verification, avoiding additional verification calculations at the early stage of the tracking task. In this way, convergence can be completed as soon as possible on the premise of reducing the number of calculations, further improving the calculation efficiency of voltage stability margin monitoring.

[0063] As Figure 2 shown, in one of the embodiments, the optimization and learning process of the policy network includes:

[0064] S201: Determine a preset target network, and perform iterative learning on the action policy of the target network during the P-V curve tracking process.

[0065] Among them, the target network can adopt the network architecture of DQN (Deep Q-Network).

[0066] S202: In each learning process, obtain a target sample set, determine the action policy of the target network during the P-V curve tracking process according to the target sample set, and execute the action policy.

[0067] Among them, the target sample set includes multiple target samples, and each target sample includes state information, action policy, reward value, and state information after executing the action policy.

[0068] S203: After determining the reward function according to the actual operation situation of the power system after executing the action policy, optimize the target network based on the determined reward function.

[0069] Among them, the reward function includes reward terms for the average voltage, power output, power flow convergence state, and tracking task completion degree of the power system.

[0070] S204: Determine whether the difference between the minimum unconverged power and the maximum converged power in the P-V curve obtained by tracking after executing the action strategy is less than a preset critical value.

[0071] S205: If it is less than, trigger the accelerated convergence determination mechanism.

[0072] S206: Determine whether the P-V curve tracking task has been completed.

[0073] S207: When it is determined that the P-V curve tracking task has been completed, determine the latest target network as the policy network and determine the target P-V curve.

[0074] S208: If it is not less than or when it is determined that the P-V curve tracking task has not been completed, enter the next learning.

[0075] In this embodiment, the target network is iteratively learned in the above manner to obtain the policy network. On the one hand, since the speed of tracking the P-V curve can be controlled by the action strategy, and on this basis, the optimal action strategy is searched for to improve the monitoring efficiency while maintaining the accuracy of the policy network. On the other hand, the accelerated convergence determination mechanism can accelerate the convergence process, enabling the policy network to have real-time response capabilities.

[0076] In one embodiment, a slave network can be established during the iterative learning of the target network. Every certain number of iterations, update the parameters of the slave network to the parameters of the target network to stabilize and optimize the iterative process.

[0077] In another embodiment, the e-greedy strategy can be adopted to balance exploration and exploitation. Specifically, at the initial stage of iterative learning, the action strategy can be determined with a relatively large exploration rate. As the iteration progresses, the exploration rate is gradually reduced, enabling the target network to make more decisions using the learned strategies. It can be understood that at the initial stage of iterative learning, in order to fully explore the environment and find potential reward sources, the target network will maintain a relatively high exploration rate. This means that the target network will frequently try new action strategies to understand which actions can bring higher rewards. However, as the iteration deepens, the target network gradually learns which actions are effective. Therefore, the target network will gradually reduce the exploration of new actions and instead execute more known behaviors that can bring high rewards. At this time, the exploration rate decreases and tends to be stable.

[0078] As Figure 3 shown, in one of the embodiments, the determination process of the accelerated convergence determination mechanism includes:

[0079] S301: Calculate the first power flow state of the currently executed action strategy.

[0080] S302: Determine whether the first power flow state has converged.

[0081] S303: If the first power flow state converges, add a preset convergence threshold to the currently executed action strategy as the first target strategy, and calculate the second power flow state of the first target strategy.

[0082] S304: Determine whether the second power flow state converges.

[0083] S305: If the second power flow state does not converge, it is determined that the P-V curve tracking task is completed.

[0084] S306: If the second power flow state converges, it is determined that the P-V curve tracking task is not completed.

[0085] S307: If the first power flow state does not converge, subtract a preset convergence threshold from the currently executed action strategy as the second target strategy, and calculate the third power flow state of the second target strategy.

[0086] S308: Determine whether the third power flow state converges.

[0087] S305: If the third power flow state converges, it is determined that the P-V curve tracking task is completed.

[0088] S309: If the third power flow state does not converge, it is determined that the P-V curve tracking task is not completed.

[0089] In an example, assume that the preset critical value is 20 and the convergence threshold is 0.5. When the difference between the minimum unconverged power and the maximum converged power in the P-V curve obtained by tracking after executing the action strategy is less than 20, the accelerated convergence determination mechanism is entered. In this stage, every time the target network executes a new action P, the corresponding power flow state will be calculated. If the current state power flow converges, calculate the power flow state at P + 0.5. If this state does not converge, it is determined that the P-V curve tracking task is completed, and it is confirmed that the static stability limit power has been tracked. If the current state power flow does not converge, calculate the power flow state at P - 0.5. If this state power flow converges, it is also determined that the P-V curve tracking task is completed, indicating that the static stability limit power has been tracked. Through this verification mechanism, the accuracy of the completion of the P-V curve tracking task can be effectively ensured. At the same time, since additional verification is only started when the difference between the minimum unconverged power and the maximum converged power is less than the preset critical value, this avoids performing additional verification calculations early in the task, thereby saving a large amount of calculation time and avoiding an unnecessary increase in the number of calculations.

[0090] As Figure 4 shown, Figure 4 is a comparison example diagram of the accelerated convergence determination mechanism provided by the embodiment of the present application. In Figure 4Among them, the P value represents power, the V value represents voltage, and the DQN method represents the P-V curve tracking method based on the acceleration convergence determination mechanism used in this solution. Compared with traditional P-V curve tracking methods such as the dichotomy method and the equal-distance halving method, the P-V curve tracking method based on the acceleration convergence determination mechanism proposed in this invention can track the static stable power limit only through 5 iterations, while the traditional equal-distance halving method and dichotomy tracking method need 19 and 13 iterations respectively. The P-V curve tracking method based on the acceleration convergence determination mechanism takes less time and improves the tracking efficiency of the P-V curve.

[0091] As Figure 5 shown, in one embodiment, the reward function is determined according to the actual operating conditions of the power system after executing the action strategy, including:

[0092] S401: Determine the reward terms regarding the average voltage, power output, power flow convergence state, and tracking task completion degree of the power system according to the actual operating conditions of the power system after executing the action strategy.

[0093] In this step, the required data can be obtained from the actual operating conditions of the power system after executing the action strategy, and then scores can be assigned to each reward term. In one example, the scores can be assigned to each reward term according to the following expression:

[0094]

[0095] In the formula, represents the reward term for the average voltage, represents the reward term for the power output, represents the reward term for the power flow convergence, represents the reward term for the tracking task completion degree, represents the average voltage drop value, represents that the power flow state is stable, represents that the power flow state is unstable, represents the increased value of the power output. When the tracking task completion degree is determined to be completed, a relatively large positive reward can be given at this time.

[0096] S402: Determine the weights of each reward term according to the optimization objective of the policy network.

[0097] In this step, the weights of each reward term can be determined by means of experimental exploration or theoretical analysis in combination with the optimization objective of the policy network.

[0098] Specifically, the weights of each reward item can be set and adjusted according to actual needs, and the present application does not make specific limitations thereon. By flexibly setting the weights of different reward items, the policy network can achieve more efficient learning and decision-making in a complex system environment, thereby improving the generalization ability of the policy network.

[0099] In one example, the weights of the reward item of average voltage, the reward item of power output, the reward item of power flow convergence, and the reward item of tracking task completion degree can be set to 0.1, 0.4, 1.0, and 1.0 respectively.

[0100] S403: Generate a reward function based on each reward item and its corresponding weight.

[0101] Among them, the reward function takes the increase of power output and the decrease of average voltage as the exploration direction.

[0102] In one example, the process of S403 can be determined according to the following expression:

[0103]

[0104]

[0105] In the formula, R represents the reward function, w represents the weight vector, represents the weight of the reward item of average voltage, represents the weight of the reward item of power output, represents the weight of the reward item of power flow convergence, represents the weight of the reward item of tracking task completion degree, V represents the reward vector, represents the reward item of average voltage, represents the reward item of power output, represents the reward item of power flow convergence, represents the reward item of tracking task completion degree.

[0106] In this embodiment, the reward function is determined through aspects such as average voltage, power output, power flow convergence state, and tracking task completion degree, which can comprehensively consider the state data of multiple dimensions, thereby accurately guiding the policy network to explore in the direction of increasing power output and decreasing average voltage, enabling the power system to be closer to the static stability limit power, and thus realizing the real-time monitoring of the voltage stability margin.

[0107] In one of the embodiments, obtain a target sample set, including:

[0108] Select target samples from a preset experience replay pool according to the sample priorities, and generate a target sample set based on the selected target samples; each sample in the experience replay pool has been marked with a sample priority, and the sample priority is used to quantify the error of the samples in the experience replay pool.

[0109] Specifically, when constructing the experience replay pool, the historical state information, the corresponding action strategy, the reward value, and the state information after executing the action strategy can be used as a sample and stored in the experience replay pool, and when storing, the sample priority of each stored sample is calculated to determine the sample priority of each sample in the experience replay pool. Also, the state information, action strategy, reward value, and state information after executing the action strategy generated during the iterative optimization learning process of the policy network will be stored in the experience replay pool to update the experience replay pool in a timely manner.

[0110] Furthermore, during the process of calculating the sample priority of each stored sample, the calculation can be performed according to the following expression:

[0111]

[0112] In the formula, represents the sample priority of sample t, represents the reward value of sample t, represents the discount factor, represents the state information of the value function, represents the state information after executing the action strategy of the value function.

[0113] In this embodiment, during the iterative learning process of the policy network, each iteration will re-obtain target samples from the experience replay pool to form the basic data for this iteration. This can improve the generalization ability and learning speed of the policy network. And, by selecting target samples according to the sample priorities, a target sample set is finally formed, and using target samples with smaller errors for optimization iteration can further accelerate the convergence speed and improve the computational efficiency of the policy network.

[0114] In one of the embodiments, determining the static stability limit power according to the target P-V curve and calculating the current voltage stability margin of the power system based on the static stability limit power includes:

[0115] S1: Determine the maximum active power point in the target P-V curve, and determine the power corresponding to the maximum active power point as the static stability limit power.

[0116] S2: Obtain the current operating power of the power system, and calculate the current voltage stability margin of the power system according to the operating power and the static stability limit power.

[0117] In one example, the current voltage stability margin of the power system can be calculated according to the following expression:

[0118]

[0119] Wherein, K is the voltage stability margin, represents the static stability limit power, represents the current operating power of the power system.

[0120] In this embodiment, the static stability limit power is determined by tracking the obtained target P-V curve, and then the voltage stability margin is calculated based on the static stability limit power and the current operating power, so as to quickly obtain the voltage stability margin to achieve on-line monitoring of the voltage stability margin of the power system.

[0121] In one of the embodiments, the state space is expressed as:

[0122]

[0123] Wherein, represents the active power of node n, represents the reactive power of node n, represents the active power load of node n, represents the reactive power load of node n, represents the voltage of node n, represents the voltage phase angle of node n, represents the power flow convergence information.

[0124] In this embodiment, from the above information, it can be seen that the state space includes the active power, reactive power, active power load, reactive power load, voltage, voltage phase angle of each node in the power system, and the power flow convergence information of the power system.

[0125] Specifically, the power flow convergence information is a boolean value. When the power flow state converges, the power flow convergence information is 1, and when the power flow state does not converge, the power flow convergence information is 0.

[0126] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the indications of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0127] The voltage stability margin monitoring device provided by the embodiments of the present application will be described below. The voltage stability margin monitoring device described below can be correspondingly referred to the voltage stability margin monitoring method described above.

[0128] As Figure 6 shown, the present application provides a voltage stability margin monitoring device 500, and the device includes:

[0129] An information acquisition module 501, configured to acquire the current state information of the power system;

[0130] A curve tracking module 502, configured to input the state information into a preset policy network, execute the optimal action policy when determining the optimal action policy corresponding to the state information, and track and obtain the target P-V curve; the policy network is optimized and learned with the state information of the power system as the state space and the increment of the active power of the new energy output node in the power system as the action space, and triggers an accelerated convergence determination mechanism to determine whether the P-V curve tracking task is completed when meeting the preset conditions;

[0131] A margin monitoring module 503, configured to determine the static stability limit power according to the target P-V curve, and calculate the current voltage stability margin of the power system based on the static stability limit power.

[0132] In one embodiment, the voltage stability margin monitoring device includes:

[0133] An iterative learning module, configured to determine a preset target network, and perform iterative learning on the action policy of the target network during the P-V curve tracking process;

[0134] A network optimization module, which is used to obtain a target sample set during each learning process, determine an action strategy of a target network during the P-V curve tracking process according to the target sample set, and execute the action strategy. After determining a reward function based on the actual operating conditions of the power system after executing the action strategy, the target network is optimized based on the determined reward function;

[0135] An acceleration convergence module, which is used to determine whether the difference between the minimum unconverged power and the maximum converged power in the P-V curve obtained after tracking by executing the action strategy is less than a preset critical value. If it is less than, an acceleration convergence determination mechanism is triggered. When it is determined that the P-V curve tracking task is completed, the latest target network is determined as the policy network, and a target P-V curve is determined. When it is determined that the P-V curve tracking task is not completed, the next learning is entered.

[0136] In one embodiment, the curve tracking module or the acceleration convergence module includes:

[0137] A first calculation sub-module, which is used to calculate the first power flow state of the currently executed action strategy. If the first power flow state converges, a preset convergence threshold is added to the currently executed action strategy as the first target strategy, and the second power flow state of the first target strategy is calculated. If the second power flow state does not converge, it is determined that the P-V curve tracking task is completed. If the second power flow state converges, it is determined that the P-V curve tracking task is not completed;

[0138] A second calculation sub-module, which is used to subtract a preset convergence threshold from the currently executed action strategy as the second target strategy if the first power flow state does not converge, and calculate the third power flow state of the second target strategy. If the third power flow state converges, it is determined that the P-V curve tracking task is completed. If the third power flow state does not converge, it is determined that the P-V curve tracking task is not completed.

[0139] In one embodiment, the network optimization module includes:

[0140] A reward item determination sub-module, which is used to determine reward items regarding the average voltage, transmitted power, power flow convergence state, and tracking task completion degree of the power system according to the actual operating conditions of the power system after executing the action strategy;

[0141] A weight determination sub-module, which is used to determine the weights of each reward item according to the optimization objective of the policy network;

[0142] A function determination sub-module, which is used to generate a reward function based on each reward item and its corresponding weight; the reward function takes the increase in transmitted power and the decrease in average voltage as the exploration direction.

[0143] In one embodiment, the information acquisition module includes:

[0144] An information acquisition sub-module, configured to select target samples from a preset experience replay pool according to sample priorities, and generate a target sample set based on the selected target samples; each sample in the experience replay pool has been marked with a sample priority, and the sample priority is used to quantify the error of the samples in the experience replay pool.

[0145] In one embodiment, the margin monitoring module includes:

[0146] A power determination sub-module, configured to determine a maximum active power point in a target P-V curve, and determine the power corresponding to the maximum active power point as the static stability limit power;

[0147] A margin calculation sub-module, configured to obtain the current operating power of the power system, and calculate the current voltage stability margin of the power system according to the operating power and the static stability limit power.

[0148] The division of each module in the above voltage stability margin monitoring device is only for illustrative purposes. In other embodiments, the voltage stability margin monitoring device can be divided into different modules as needed to complete all or part of the functions of the above voltage stability margin monitoring device. Each module in the above voltage stability margin monitoring device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in a computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to each of the above modules.

[0149] In one embodiment, the present application further provides a storage medium, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the voltage stability margin monitoring method as described in any one of the above embodiments.

[0150] In one embodiment, the present application further provides a computer device, in which computer-readable instructions are stored. When the computer-readable instructions are executed by one or more processors, one or more processors are caused to execute the steps of the voltage stability margin monitoring method as described in any one of the above embodiments.

[0151] Schematically, as Figure 7 shown, Figure 7 is an internal structural diagram of a computer device provided by an embodiment of the present application. The computer device 600 can be provided as a server. Referring to Figure 7, the computer device 600 includes a processing component 602, which further includes one or more processors, and memory resources represented by a memory 601 for storing instructions executable by the processing component 602, such as application programs. The application programs stored in the memory 601 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 602 is configured to execute instructions to perform the voltage stability margin monitoring method of any of the above embodiments.

[0152] The computer device 600 may further include a power supply component 603 configured to perform power management of the computer device 600, a wired or wireless network interface 604 configured to connect the computer device 600 to a network, and an input / output (I / O) interface 605. The computer device 600 may operate based on an operating system stored in the memory 601, such as Windows Server TM, Mac OS XTM, Unix TM, Linux TM, Free BSDTM or the like.

[0153] Those skilled in the art can understand that Figure 7 the structure shown in

[0154] Finally, it should also be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element. In this text, the singular forms "a", "an" and "the" may also include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms "comprising / including" or "having" etc. specify the existence of the stated features, wholes, steps, operations, components, parts or combinations thereof, but do not exclude the possibility of the existence or addition of one or more other features, wholes, steps, operations, components, parts or combinations thereof. At the same time, the term "and / or" used in this specification includes any and all combinations of the related listed items.

[0155] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0156] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A voltage stability margin monitoring method, characterized in that: The method comprises: Obtain current status information of the power system; The state information is input into a preset strategy network, so that when the optimal action strategy corresponding to the state information is determined, the optimal action strategy is executed, and the target PV curve is tracked; the strategy network is obtained by optimizing and learning by taking the state information of the power system as the state space and the growth amount of the active power of the new energy output node in the power system as the action space, and triggering the accelerated convergence judgment mechanism to judge whether the PV curve tracking task is completed when the preset conditions are met; A static stability limit power is determined according to the target PV curve, and a current voltage stability margin of the power system is calculated based on the static stability limit power.

2. The voltage stability margin monitoring method according to claim 1, characterized in that: The optimization learning process of the policy network includes: Determine a preset target network, and iteratively learn the action strategy of the target network during PV curve tracking; In each learning process, a target sample set is obtained, an action strategy of the target network in the PV curve tracking process is determined according to the target sample set, and the action strategy is executed, so as to determine a reward function according to the actual operation of the power system after the action strategy is executed, and then the target network is optimized based on the determined reward function; It is determined whether the difference between the minimum unconverged power and the maximum converged power in the PV curve tracked after executing the action strategy is less than a preset critical value. If so, the accelerated convergence determination mechanism is triggered. When it is determined that the PV curve tracking task is completed, the latest target network is determined as the strategy network, and the target PV curve is determined. When it is determined that the PV curve tracking task is not completed, the next learning is entered.

3. The voltage stability margin monitoring method according to claim 1 or 2, characterized in that: The determination process of the accelerated convergence determination mechanism includes: Calculate the first power flow state of the currently executed action strategy. If the first power flow state converges, add a preset convergence threshold to the currently executed action strategy as the first target strategy, and calculate the second power flow state of the first target strategy. If the second power flow state does not converge, it is determined that the PV curve tracking task has been completed. If the second power flow state converges, it is determined that the PV curve tracking task has not been completed. If the first power flow state has not converged, a preset convergence threshold is subtracted from the currently executed action strategy as the second target strategy, and the third power flow state of the second target strategy is calculated. If the third power flow state converges, it is determined that the PV curve tracking task has been completed. If the third power flow state has not converged, it is determined that the PV curve tracking task has not been completed.

4. The voltage stability margin monitoring method according to claim 2, characterized in that: The step of determining the reward function according to the actual operation status of the power system after executing the action strategy includes: Determine reward items related to the average voltage, output power, power flow convergence state and tracking task completion of the power system according to the actual operation status of the power system after executing the action strategy; Determining the weight of each reward item according to the optimization goal of the strategy network; A reward function is generated based on each reward item and its corresponding weight; the reward function explores the direction of increasing the output power and decreasing the average voltage.

5. The voltage stability margin monitoring method according to claim 2, characterized in that: The step of obtaining a target sample set includes: A target sample is selected from a preset experience replay pool according to a sample priority, and a target sample set is generated based on the selected target sample; each sample in the experience replay pool is marked with a sample priority, and the sample priority is used to quantify the error of the sample in the experience replay pool.

6. The voltage stability margin monitoring method according to claim 1, characterized in that: The step of determining the static stability limit power according to the target PV curve, and calculating the current voltage stability margin of the power system based on the static stability limit power, comprises: Determining a maximum active power point in the target PV curve, and determining the power corresponding to the maximum active power point as a static stability limit power; The current operating power of the power system is obtained, and the current voltage stability margin of the power system is calculated according to the operating power and the static stability limit power.

7. The voltage stability margin monitoring method according to claim 1, characterized in that: The state space is represented as: In the formula, represents the active power of node n, represents the reactive power of node n, represents the active power load of node n, represents the reactive power load of node n, represents the voltage at node n, represents the voltage phase angle at node n, Indicates power flow convergence information.

8. A voltage stability margin monitoring device, characterized in that: The device comprises: An information acquisition module is used to obtain the current status information of the power system; A curve tracking module is used to input the state information into a preset strategy network, so as to execute the optimal action strategy when determining the optimal action strategy corresponding to the state information, and track the target PV curve; the strategy network is obtained by optimizing and learning with the state information of the power system as the state space and the increase in the active power of the new energy output node in the power system as the action space, and triggers the accelerated convergence judgment mechanism to judge whether the PV curve tracking task is completed when the preset conditions are met; The margin monitoring module is used to determine the static stability limit power according to the target PV curve, and calculate the current voltage stability margin of the power system based on the static stability limit power.

9. A storage medium, characterized in that: The storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the voltage stability margin monitoring method as described in any one of claims 1 to 7.

10. A computer device, characterized in that: include: one or more processors, and memory; The memory stores computer-readable instructions, and when the computer-readable instructions are executed by the one or more processors, the steps of the voltage stability margin monitoring method according to any one of claims 1 to 7 are performed.