Power system fault control method and device facing stability boundary
By acquiring the composite stability margin index in real time and using reinforcement learning, key points for power system fault control are identified, solving the problems of accuracy and efficiency in fault control after the integration of new energy sources, and realizing the safe and stable operation of the power system.
Patent Information
- Application Number
- CN202511541038.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-27
- Publication Date
- 2026-01-30
AI Technical Summary
In the face of power systems with large-scale integration of new energy sources, existing fault control methods struggle to guarantee the efficiency and accuracy of fault identification. Traditional methods are unable to adapt to changes in system stability boundaries, leading to the omission of critical faults or waste of resources.
By acquiring the composite stability margin index in real time, combining fast time-domain simulation and reinforcement learning, the power flow section and the set of influencing nodes are identified, and a multi-dimensional reward function is constructed to output a fault control scheme, ensuring that the control scheme takes into account safety, stability, economy and operational feasibility.
This has improved the real-time performance and accuracy of fault control for new energy sources connected to the power system, avoiding the omission of critical faults and waste of resources, and ensuring the safe and stable operation of the power system.
Smart Images

Figure CN121440586A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power systems, and more particularly to a power system fault control method and apparatus oriented towards stability boundaries. Background Technology
[0002] In modern power systems, fault control is the lifeline for ensuring the safe and stable operation of the power grid. The task of fault control is to quickly and accurately identify the faults most likely to cause system collapse from thousands of potential faults, and to implement effective control measures to prevent local faults from evolving into large-scale blackouts. With the large-scale integration of fluctuating new energy sources such as wind power and photovoltaics, power electronic equipment has profoundly changed the dynamic characteristics of the system, making the system's stability boundary more complex and variable, which poses a challenge to traditional fault control methods.
[0003] In existing technologies, mainstream critical fault identification methods mainly rely on two types of techniques: one is a purely data-driven method based on clustering and machine learning; the other is an iterative screening method based on linearized physical models and heuristic rules. Because the first type of method heavily depends on the completeness of offline historical samples, it suffers from the problem of missing critical faults due to matching failures when the system's operation mode shifts to an uncovered new area due to fluctuations in renewable energy sources, lacking the ability to adapt to dynamic changes in stability boundaries.
[0004] Meanwhile, since the second type of method uses a fixed linearization model and filtering threshold to approximate the behavior of a strongly nonlinear system, when the system stability boundary shrinks, an overly lenient fixed threshold may cause dangerous faults to be missed; while when the system is stable and sufficient, an overly strict threshold will cause unnecessary waste of computing resources and cannot intelligently adjust the filtering strategy according to the real-time stress level of the system. Summary of the Invention
[0005] This invention provides a power system fault control method and apparatus oriented towards stability boundaries, which can solve the problem of improving the accuracy of fault control while ensuring fault identification efficiency in the prior art.
[0006] In a first aspect, embodiments of the present invention provide a power system fault control method oriented towards stability boundaries, comprising:
[0007] The current composite stability margin index of the power system is obtained in real time; wherein, the current composite stability margin index is calculated based on the real-time operating data of the power system using the composite stability margin index calculation formula;
[0008] If the current composite stability margin index is greater than or equal to the first preset threshold, the current operating scenario sample set of the power system is collected, and the current fault operating scenario and the expected disturbance set corresponding to the current fault operating scenario are obtained based on the preset offline stability scenario knowledge base and the current operating scenario sample set.
[0009] Based on the current fault operation scenario and the expected disturbance set corresponding to the current fault operation scenario, a fast time domain simulation is used to simulate and obtain the current dynamic response trajectory. Based on the current dynamic response trajectory, power flow identification is performed to obtain the current power flow section and the current set of affected nodes.
[0010] A state space is established based on the current composite stability margin index, the current power flow section, and the current set of affected nodes. The system is then trained in a preset reinforcement learning environment based on the state space and the reward function to output the current fault control scheme of the power system. The reward function is obtained by weighting the over-limit elimination reward, stability margin reward, adjustment cost reward, and constraint reward.
[0011] The power system is subjected to fault control according to the current fault control scheme.
[0012] This application embodiment obtains the current composite stability margin index calculated from real-time power system operation data in real time, enabling dynamic and accurate understanding of the system's current stability state. This avoids the inability to adapt to the complex changes in system stability boundaries after the integration of new energy sources due to reliance on offline data or static assessments. When the index reaches a first preset threshold, the current fault operation scenario and corresponding anticipated disturbance set are determined by collecting a sample set of current operation scenarios and combining it with a preset offline stability scenario knowledge base. This allows for targeted focus on potential fault risks, avoiding ineffective analysis of irrelevant scenarios to save computational resources and effectively preventing the omission of critical faults. Based on the above fault scenarios and anticipated disturbance sets, rapid time-domain simulation is used to obtain dynamic response trajectories and further identify power flow sections and sets of influencing nodes. This approach can accurately capture the dynamic changes of the system under fault conditions, identify the key links and core influencing nodes of fault propagation, and provide a basis for the precise application of subsequent control measures. Furthermore, a state space is constructed using a composite stability margin index, power flow sections, and sets of influencing nodes. A reward function that integrates multiple objectives—including limit elimination, stability margin, adjustment costs, and constraints—is then trained and output in a reinforcement learning environment. This ensures that the solution considers system safety, stability, economy, and operational feasibility, avoiding the shortcomings of control effects under a single objective. Finally, implementing fault control based on this solution can effectively address complex fault situations in the power system under large-scale integration of new energy sources, significantly improving the real-time performance, accuracy, and adaptability of fault control, and ensuring the safe and stable operation of the power system.
[0013] As a preferred example of the first aspect, the current composite stability margin index is calculated based on the real-time operating data of the power system using the composite stability margin index calculation formula, specifically:
[0014] Based on the real-time operating data of the power system, the voltage stability sub-item calculation formula, the thermal stability sub-item calculation formula, and the power angle stability sub-item calculation formula are used sequentially to calculate the voltage stability sub-item, the thermal stability sub-item, and the power angle stability sub-item.
[0015] The composite stability margin index is obtained by weighted geometric mean fusion of the voltage stability sub-item, the thermal stability sub-item, and the power angle stability sub-item.
[0016] In this preferred example, a clearly defined composite stability margin index calculation method overcomes the limitations of traditional single stability indices in assessing system state. By calculating voltage stability, thermal stability, and power angle stability sub-items separately, the core stability dimensions of power system operation are comprehensively covered, accurately reflecting the actual state of the system at different stability levels and avoiding misjudgments of system stability due to the one-sidedness of a single index. The use of a weighted geometric average to fuse each sub-item assigns appropriate weights based on the importance of different stability dimensions, making the final composite stability margin index more closely reflect the actual operating characteristics of the system. This provides a scientific and reliable quantitative basis for subsequent judgments on whether to initiate fault control procedures. This calculation method effectively adapts to the complex and variable stability characteristics of systems after the integration of new energy sources, improving the comprehensiveness and accuracy of system stability state assessment and laying a solid data foundation for fault control.
[0017] As a preferred example of the first aspect, the step of obtaining the current fault operation scenario and the expected disturbance set corresponding to the current fault operation scenario based on a preset offline stable scenario knowledge base and the current operating scenario sample set specifically involves:
[0018] The current running scenario sample set is matched with the scenario data in the preset offline stable scenario knowledge base to obtain the Mahalanobis distance factor, stability margin consistency factor and data completeness factor corresponding to each scenario.
[0019] The matching confidence level of each scenario data is obtained by weighting the Mahalanobis distance factor, stability margin consistency factor and data completeness factor.
[0020] Based on a preset confidence threshold, the matching confidence levels corresponding to each scenario data are sequentially filtered and matched to obtain the expected disturbance set corresponding to the current fault operation scenario.
[0021] In this preferred example, the proposed scenario matching and anticipated disturbance set acquisition method solves the matching failure problem caused by incomplete offline sample coverage in existing data-driven methods. By introducing Mahalanobis distance factor, stability margin consistency factor, and data completeness factor, the current operating scenario and the offline knowledge base scenario are matched from three dimensions: data similarity, stability characteristic consistency, and data quality. Compared with single-dimensional matching, this method is more comprehensive and reliable, and can effectively avoid scenario mismatch caused by a single matching standard. The matching confidence score is obtained by weighting each factor, and combined with preset threshold filtering, the fault operating scenario that best matches the current scenario and the corresponding anticipated disturbance set can be accurately located. This ensures that even if the system operation mode changes due to new energy fluctuations, suitable fault analysis basis can be quickly found, avoiding the omission of key faults. At the same time, it reduces the computational resource consumption caused by invalid scenario analysis and improves the pertinence and efficiency of fault scenario analysis.
[0022] As a preferred example of the first aspect, the step of identifying the power flow based on the current dynamic response trajectory to obtain the current power flow section and the current set of influencing nodes specifically involves:
[0023] Based on the current dynamic response trajectory, the current power flow section is obtained by using the stability margin slope decrease judgment method.
[0024] Based on the current power flow section, the electrical coupling group algorithm and the power flow entropy weight method are used for localization to obtain the current set of affected nodes.
[0025] In this preferred example, based on the current dynamic response trajectory, the stability margin slope descent judgment method can accurately capture the key power flow section in the dynamic process of system faults. This section can intuitively reflect the critical point of power flow change under the influence of faults, providing a clear direction for subsequent analysis of fault propagation paths. Combining the current power flow section, the electrical coupling group algorithm and power flow entropy weight method are used to locate the set of influencing nodes. From the perspectives of electrical correlation characteristics and the importance of power flow distribution, nodes that significantly affect fault development and system stability can be accurately identified, avoiding blind analysis of all nodes and significantly reducing computational load while ensuring that subsequent fault control measures can accurately target key nodes. This method improves the accuracy of power flow analysis and the targeting of influencing node location, providing crucial support for developing efficient fault control schemes and ensuring that control measures can quickly suppress fault propagation.
[0026] As a preferred example of the first aspect, the reward function is obtained by weighting the out-of-limit elimination reward, stability margin reward, adjustment cost reward, and constraint reward, specifically:
[0027] The original reward function is obtained by weighting the over-limit elimination reward, the stability margin reward, and the adjustment cost reward according to a preset weight set.
[0028] The reward function is obtained based on the original reward function and the constrained reward.
[0029] In this preferred example, a reward function is constructed to address the performance limitations of control schemes caused by the single reward factor in traditional reinforcement learning. By weighting the limit-avoidance elimination reward, stability margin reward, and adjustment cost reward using a pre-set weight set, the reward function can simultaneously consider the core objectives of fault control: the limit-avoidance elimination reward ensures that the control scheme can quickly eliminate system operational limits, guaranteeing system safety; the stability margin reward promotes the control scheme to improve system stability reserves and enhance disturbance rejection capabilities; and the adjustment cost reward controls the economic input during fault control, avoiding resource waste. The introduction of constraint rewards further ensures that the control scheme conforms to the physical constraints and operational specifications of the system, preventing infeasible control strategies. This reward function design comprehensively covers the safety, stability, economy, and feasibility objectives of fault control, effectively guiding reinforcement learning training, resulting in superior overall performance of the output fault control scheme. It adapts to the multi-objective balance requirements of fault control in systems with new energy integration, improving the practicality and reliability of the control scheme.
[0030] In a second aspect, the present invention provides a power system fault control device oriented towards stability boundaries, comprising: a data acquisition module, a judgment module, a first calculation module, a second calculation module, and a fault control module;
[0031] The data acquisition module is used to acquire the current composite stability margin index of the power system in real time; wherein, the current composite stability margin index is calculated based on the real-time operating data of the power system using the composite stability margin index calculation formula;
[0032] The judgment module is used to collect the current operating scenario sample set of the power system when the current composite stability margin index is greater than or equal to the first preset threshold, and obtain the current fault operating scenario and the expected disturbance set corresponding to the current fault operating scenario based on the preset offline stability scenario knowledge base and the current operating scenario sample set.
[0033] The first calculation module is used to perform simulation using fast time-domain simulation based on the current fault operation scenario and the expected disturbance set corresponding to the current fault operation scenario, to obtain the current dynamic response trajectory, and to perform power flow identification based on the current dynamic response trajectory, to obtain the current power flow section and the current set of affected nodes.
[0034] The second calculation module is used to establish a state space based on the current composite stability margin index, the current power flow section, and the current set of affected nodes, and to train the system in a preset reinforcement learning environment based on the state space and the reward function, and output the current fault control scheme of the power system; wherein, the reward function is obtained by weighting the over-limit elimination reward, stability margin reward, adjustment cost reward, and constraint reward.
[0035] The fault control module is used to perform fault control on the power system according to the current fault control scheme.
[0036] As a preferred example of the second aspect, the data acquisition module includes a first acquisition unit and a second acquisition unit;
[0037] The first acquisition unit is used to calculate the voltage stability sub-item, thermal stability sub-item, and power angle stability sub-item according to the real-time operation data of the power system in sequence using the voltage stability sub-item calculation formula, the thermal stability sub-item calculation formula, and the power angle stability sub-item calculation formula;
[0038] The second acquisition unit is used to perform a weighted geometric average fusion process on the voltage stability sub-item, the thermal stability sub-item, and the power angle stability sub-item to obtain the composite stability margin index.
[0039] As a preferred example of the second aspect, the judgment module includes a first judgment unit, a second judgment unit, and a third judgment unit;
[0040] As a preferred example of the second aspect, the first computing module includes a first computing unit and a second computing unit;
[0041] The first calculation unit is used to make a judgment based on the current dynamic response trajectory using the stability margin slope decrease judgment method to obtain the current power flow section;
[0042] The second calculation unit is used to locate the current influencing nodes based on the current power flow section using the electrical coupling group algorithm and the power flow entropy weight method.
[0043] As a preferred example of the second aspect, the second computing module includes a third computing unit and a fourth computing unit;
[0044] The third calculation unit is used to weight the over-limit elimination reward, the stability margin reward, and the adjustment cost reward according to a preset weight set to obtain the original reward function;
[0045] The fourth calculation unit is used to obtain the reward function based on the original reward function and the constraint reward.
[0046] In summary, this application embodiment, by acquiring the current composite stability margin index calculated from real-time power system operation data, can dynamically and accurately grasp the current stable state of the system, avoiding the inability to adapt to the complex changes in the system stability boundary after the access of new energy sources due to reliance on offline data or static evaluation. When the index reaches a first preset threshold, the current fault operation scenario and corresponding expected disturbance set are determined by collecting a sample set of current operation scenarios and combining it with a preset offline stability scenario knowledge base. This allows for targeted focus on potential fault risks, avoiding ineffective analysis of irrelevant scenarios to save computing resources and effectively avoiding the problem of missing key faults. Based on the above fault scenarios and expected disturbance sets, rapid time-domain simulation is used to obtain dynamic response trajectories and further identify power flow sections and influencing nodes. The set of data points can accurately capture the dynamic changes of the system under the influence of faults, identify the key links and core influencing nodes of fault propagation, and provide a basis for the precise application of subsequent control measures. Then, a state space is constructed using a composite stability margin index, power flow sections, and the set of influencing nodes. Combined with a reward function that integrates multiple dimensions of objectives such as limit elimination, stability margin, adjustment cost, and constraints, the system is trained in a reinforcement learning environment to output a control scheme. This ensures that the scheme takes into account system safety, stability, economy, and operational feasibility, avoiding the shortcomings of control effects under a single objective. Finally, the implementation of fault control based on this scheme can effectively cope with the complex fault situations of the power system under the large-scale integration of new energy sources, significantly improve the real-time performance, accuracy, and adaptability of fault control, and ensure the safe and stable operation of the power system.
[0047] Another embodiment of the present invention provides a terminal device, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements the steps of the power system fault control method oriented towards stability boundaries of the present invention.
[0048] Another embodiment of the present invention provides a computer-readable storage medium item, including: a stored computer program, which, when the computer program is running, controls the device where the computer-readable storage medium is located to perform the steps of the power system fault control method oriented towards stability boundaries of the present invention. Attached Figure Description
[0049] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0050] Figure 1 A flowchart illustrating an embodiment of a power system fault control method oriented towards stability boundaries provided by the present invention;
[0051] Figure 2 A simulation analysis flowchart of an embodiment of a power system fault control method oriented towards stability boundaries provided by the present invention;
[0052] Figure 3 A learning environment architecture diagram for one embodiment of a power system fault control method oriented towards stability boundaries provided by the present invention;
[0053] Figure 4 This is a module structure diagram of one embodiment of a power system fault control device oriented towards stability boundaries provided by the present invention. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0056] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0057] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0058] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0059] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0060] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0061] Example 1
[0062] See Figure 1 To address the challenge of improving fault control accuracy while maintaining efficient fault identification in existing technologies, an embodiment of the present invention provides a power system fault control method oriented towards stability boundaries, comprising:
[0063] S1. Obtain the current composite stability margin index of the power system in real time; wherein, the current composite stability margin index is calculated based on the real-time operating data of the power system using the composite stability margin index calculation formula.
[0064] In some embodiments of this application, the current composite stability margin index is calculated based on the real-time operating data of the power system using the composite stability margin index calculation formula, specifically as follows:
[0065] Based on the real-time operating data of the power system, the voltage stability sub-item calculation formula, the thermal stability sub-item calculation formula, and the power angle stability sub-item calculation formula are used sequentially to calculate the voltage stability sub-item, the thermal stability sub-item, and the power angle stability sub-item.
[0066] The composite stability margin index is obtained by weighted geometric mean fusion of the voltage stability sub-item, the thermal stability sub-item, and the power angle stability sub-item.
[0067] Specifically, real-time operational data can be collected through synchronous phasor measurement units, remote terminal units, and smart meters deployed in the power grid. This includes node data, generator data, load data, line / transformer data, and system-level data. Node data includes the voltage amplitude and voltage phase angle of each bus. Generator data includes the active and reactive power injected by conventional units and renewable energy plants. Load data includes the active and reactive power of important load nodes. Line / transformer data includes the active power flow and apparent power limits of key transmission lines and transformers. System-level data includes system frequency and inter-regional power exchange.
[0068] Specifically, the calculation formulas for the voltage stability sub-term, the thermal stability sub-term, and the power angle stability sub-term are as follows:
[0069] The voltage stability sub-item reflects the voltage stability margin in response to load increases, similar to the voltage sensitivity of various countries' computing systems. Its calculation formula is shown below:
[0070]
[0071] Where, N L ΔP represents the number of critical load nodes. L,ref For a given reference load increase, ΔV j The voltage change caused at the j-th critical load node is obtained by solving the corrected Jacobian matrix; I V The larger this value, the more sensitive the system voltage is to disturbances, and the smaller the voltage stability margin.
[0072] The thermal stability sub-item reflects the thermal stability margin by finding the line that is closest to its transmission limit, and its calculation formula is shown below:
[0073]
[0074] Among them, S current,ij S represents the current apparent power of lines i and j. max,ij For its limit, I T This sub-item is essentially the load rate of the most stressed line in the system. The closer the value is to 1, the more likely that at least one line is close to full load, and the higher the risk of thermal instability.
[0075] The power angle stability sub-item approximates the power angle stability level by using the maximum power angle difference in the system. Its calculation formula is shown below:
[0076]
[0077] Where, θ iLet be the voltage phase angle of node i. This calculation takes into account the voltage phase angles of all nodes within the generator. The larger the power angle difference, the weaker the system's ability to maintain synchronous operation and the lower the stability margin.
[0078] Specifically, the weighted geometric mean fusion of the voltage stability sub-item, the thermal stability sub-item, and the power angle stability sub-item to obtain the composite stability margin index can be processed by the following formula:
[0079]
[0080] Wherein, CSI is the composite stability margin index. and These are the normalized voltage stability sub-item, thermal stability sub-item, and power angle stability sub-item, respectively. ωV, ωT, and ωA are the weighting coefficients for each sub-item, which can be set by domain experts according to the actual operational concerns of the power grid. For example, ωV can be set to a larger value for the receiving-end power grid.
[0081] S2. If the current composite stability margin index is greater than or equal to the first preset threshold, the current operating scenario sample set of the power system is collected, and the current fault operating scenario and the expected disturbance set corresponding to the current fault operating scenario are obtained according to the preset offline stability scenario knowledge base and the current operating scenario sample set.
[0082] In some embodiments of this application, the step of obtaining the current faulty operating scenario and the expected disturbance set corresponding to the current faulty operating scenario based on a preset offline stable scenario knowledge base and a current operating scenario sample set specifically includes:
[0083] The current running scenario sample set is matched with the scenario data in the preset offline stable scenario knowledge base to obtain the Mahalanobis distance factor, stability margin consistency factor and data completeness factor corresponding to each scenario.
[0084] The matching confidence level of each scenario data is obtained by weighting the Mahalanobis distance factor, stability margin consistency factor and data completeness factor.
[0085] Based on a preset confidence threshold, the matching confidence levels corresponding to each scenario data are sequentially filtered and matched to obtain the expected disturbance set corresponding to the current fault operation scenario.
[0086] Specifically, to fully explain the process of obtaining the expected perturbation set, the following scheme will be used as an example:
[0087] ① Construct a pre-generated offline stable scene knowledge base (a preset offline stable scene knowledge base)
[0088] In the offline phase, a knowledge base is built. Each scenario in the knowledge base not only contains the feature vectors of traditional cluster centers, but also extends to include the following scenario fingerprints: the typical composite stability margin index of the scenario, the expected perturbation set confirmed after detailed time-domain simulation analysis of the lower diameter of the scenario, and the covariance matrix of the feature vectors of the scenario, which are used to describe the distribution characteristics of the running points within the cluster.
[0089] ② Calculate the matching confidence score
[0090] The confidence level in this embodiment is a comprehensive index determined by the following Mahalanobis distance factor, stability margin consistency factor, and data completeness factor, as shown in the following formula:
[0091]
[0092]
[0093] in, The current running feature vector, Let F be the feature vector of the cluster center in scene i. M F is the Mahalanobis distance factor. S As a stability margin consistency factor, CSI current To assess the stability level of the current operating point for this factor, CSI i For the typical stability level of offline scenario i, even if two running points are geometrically close, they are essentially different scenarios if one has sufficient margin and the other is on the verge of instability. This factor effectively avoids such mismatches. F D This is the data completeness factor, which assesses the current real-time data. The completeness of F, for example, if some PMU data is missing, resulting in incomplete feature vector dimensions, then F D <1, if the data is complete, then F D =1.
[0094] Finally, the matching confidence score (Conf) is calculated as follows:
[0095] Conf=w M ·F M +w S ·F S +w D ·F D
[0096] Among them, w M w S and w D These are the weights of each factor, and w M +w S +w D =1.
[0097] ③ Identify high-risk operating scenarios and screen anticipated disturbance sets
[0098] The anticipated disturbance set here refers to a set of faults or disturbances, pre-analyzed for a specific operating scenario, that are most likely to cause system instability if they occur. This includes not only traditional component switching faults such as N-1 and N-2, but also disturbance scenarios applicable to new power systems, such as large-scale grid disconnection of new energy sources and impulsive load switching. The identification and screening process is as follows:
[0099] Calculate the confidence score between the current running state and all scenarios in the knowledge base, and take the highest value. max and its scenario S match If Conf max ≥α high (α high If the value is 0.8, then the current scenario is determined to be high-risk, and the system immediately recalls scenario S. match The associated set of anticipated perturbations D match This was then taken as the core objective of subsequent analysis. At this point, the system essentially issued a warning: the current state is extremely similar to the historical high-risk state S. match Its vulnerable point D needs to be given special attention. match If α low ≤Conf max <α high (α low If the value is 0.5, then the system is considered to be in a scenario of moderate distraction or insufficient knowledge base coverage, and can recall Conf. max The corresponding D match This serves as a reference, but will also trigger a more in-depth evaluation later. If Conf max <α low If the current scenario is considered to be a new scenario not covered by the knowledge base, step S3 will be triggered directly for analysis.
[0100] S3. Based on the current fault operation scenario and the expected disturbance set corresponding to the current fault operation scenario, a fast time domain simulation is used to simulate and obtain the current dynamic response trajectory. Based on the current dynamic response trajectory, power flow identification is performed to obtain the current power flow section and the current set of affected nodes.
[0101] In some embodiments of this application, the step of identifying the power flow based on the current dynamic response trajectory to obtain the current power flow section and the current set of influencing nodes specifically includes:
[0102] Based on the current dynamic response trajectory, the current power flow section is obtained by using the stability margin slope decrease judgment method.
[0103] Based on the current power flow section, the electrical coupling group algorithm and the power flow entropy weight method are used for localization to obtain the current set of affected nodes.
[0104] Specifically, such as Figure 2 As shown, to fully explain the process of obtaining the current set of influencing nodes, the following scheme will be used as an example:
[0105] ① Rapid simulation analysis of high-risk disturbance scenarios
[0106] Input: High-risk operating scenarios (i.e., current operating state) from step S2 and their associated set of anticipated disturbances, such as a specific N-2 fault list;
[0107] Process: For each fault in the anticipated disturbance set, a fast time-domain simulation is initiated under the initial conditions of the current operating state. This simulation does not pursue millimeter-level accuracy of electromechanical transient processes, but rather adopts a quasi-steady-state model suitable for medium- and long-term stability analysis to balance calculation speed and evaluation accuracy, and meet the timeliness requirements of online applications.
[0108] Output: The dynamic response trajectory of the system over a period of time after the disturbance is obtained, including the voltage of each node, line power, generator power angle, etc.
[0109] ②Location of key tidal current sections
[0110] Input: The dynamic response trajectory obtained from the above simulation.
[0111] Procedure: Calculate the composite stability margin index (CSI) at each time point on the simulation trajectory according to the formula defined in step S1. t This forms a dynamic curve; it identifies points where the stability margin drops sharply, including the location of the CSI. t The point of descent with the steepest slope on the curve, t critical ;when K drop When the threshold value is set to 0, it is determined that the system's stability margin has deteriorated sharply at that moment; at the critical moment t... critical Analyze the power flow distribution of all transmission lines and transformers;
[0112] Here, the critical power flow section is defined as: a set of electrically closely related sections whose overall power flow direction is consistent, and which, at t critical The set of lines whose total power transmission level is extremely high at any given time (e.g., the total power accounts for more than 90% of the total transmission capacity limit of the section) or whose power growth is the main reason for the sharp decline in CSI; Location method: combine the electrical coupling group algorithm (to identify the group of lines with close electrical connections by clustering analysis of the network admittance matrix) and the power flow entropy weight method (to assign weights according to the importance of each line's power in the section) to automatically identify this set of critical lines.
[0113] ③ Extract the set of key influencing nodes
[0114] Input: The located critical power flow section C, and its power flow value F at the steady-state operating point (before or after the fault). C ;
[0115] Process: Using the perturbation method or an analytical method based on the transposed Jacobian matrix, calculate the active power P of all controllable generator nodes and interruptible load nodes in the system. m For the main tidal current F at the key section C Sensitivity S m : Among them, S m The physical meaning is how much the power flow at the critical section m will increase by a unit increase in active power at node m. Then, the generator nodes are categorized according to sensitivity S. m Sort the load nodes from positive to negative according to their sensitivity S. m Ordered from negative to positive. Generators with high positive sensitivity are the source of exacerbating cross-sectional blockage and are also key nodes requiring reduced output to alleviate blockage. Loads with high negative sensitivity (i.e., S...) m A large negative value is a sink point for alleviating section congestion because increasing the load at that point (or reducing its generation, if it's a distributed power source) can actually absorb the section's power and alleviate congestion. This is especially important when renewable energy power plants act as negative loads. Ideally, the top N values with the highest positive sensitivity should be selected. g Of the N generator nodes, the top N with the highest negative sensitivity are selected. l There are two sets of load (or negative power generation) nodes; these two sets of nodes together constitute the set of key influencing nodes. The nodes in this set are the most effective and economical control points for adjusting the cross-sectional power flow.
[0116] S4. Establish a state space based on the current composite stability margin index, the current power flow section, and the current set of affected nodes, and train the system in a preset reinforcement learning environment based on the state space and the reward function to output the current fault control scheme of the power system; wherein, the reward function is obtained by weighting the over-limit elimination reward, stability margin reward, adjustment cost reward, and constraint reward.
[0117] In some embodiments of this application, the reward function is obtained by weighting the excess elimination reward, stability margin reward, adjustment cost reward, and constraint reward, specifically:
[0118] The original reward function is obtained by weighting the over-limit elimination reward, the stability margin reward, and the adjustment cost reward according to a preset weight set.
[0119] The reward function is obtained based on the original reward function and the constrained reward.
[0120] Specifically, such as Figure 3As shown, to fully explain the training process of reinforcement learning described above, the following scheme will be used as an example:
[0121] ① Constructing the state space
[0122] state s t It is a vector that comprehensively describes the safe operating status of the system at time t, specifically including:
[0123] CSI t The composite stability margin index at the current moment is a quantitative benchmark for the overall stability level of the system.
[0124] F c,t The total active power at the critical power flow section C. This is a core safety indicator that needs to be directly controlled;
[0125] P G,t : The current active power output vector of all generator nodes in the key influence node set;
[0126] P L,t : The current active power vector of all controllable load nodes in the key impact node set;
[0127] Auxiliary status: Optional, total system load total and critical node voltage V key This will provide richer environmental information.
[0128] ② Constructing the action space
[0129] Action a t It is a vector representing the adjustment instructions sent to each node in the set of key affected nodes; for generator nodes, the action value is the adjustment amount ΔP of active power output. G Its range is determined by the generator's ramp rate and upper and lower output limits. For example, a Gi ∈[-20,20]MW, indicating that the generator can increase / decrease output by a maximum of 20MW within a single control cycle. For controllable load nodes, the action value is the adjustment amount ΔP of the load power. L For interruptible loads, this value is negative to indicate reduction; for demand-response loads, it can be positive or negative. The action space is continuous, which allows the agent to generate more refined and smoother control commands, superior to traditional discrete switching control.
[0130] ③ Construction of the reward function
[0131] Reward function R(s) t ,a t ,s t+1 The reward function (or reward function) is the guiding principle for the agent's learning. Here, a multi-objective weighted reward function is constructed, and its formula is shown below:
[0132] R = w1·R violation +w2·R CSI +w3·R cost +R constraint
[0133] Where R is the total reward obtained by the agent after taking the action, and w1, w2, and w3 are weighting coefficients used to adjust the importance of the over-limit elimination reward, stability margin reward, and adjustment cost reward in the total reward, respectively. violation To eliminate rewards for exceeding limits, R CSI To ensure stable margin rewards, R cost To adjust cost incentives, R constraint To constrain rewards.
[0134] For the reward R for exceeding the limit elimination violation Its composition formula is as follows:
[0135] R violation =-(λ) line ·∑max(0,|F line |-F max ) 2 +λ voltage ·∑max(0,|V i -V nom |
[0136] -ΔV max ) 2 )
[0137] Where, λ line and λ voltage F represents the weighting coefficients used to adjust the penalty for exceeding line power flow limits and voltage limits, respectively. line For the actual active power flow of the line, F max V is the maximum active power flow that the line is allowed to carry. i V is the actual voltage at node i. nom The rated voltage of the node, ΔV max This represents the maximum allowable voltage deviation.
[0138] Note: This penalty applies to overruns in power flow and voltage limits. When any limit is exceeded, the reward is negative (penalty), and the more severe the limit exceedance, the greater the penalty. To obtain higher rewards, the agent must learn how to eliminate all limits.
[0139] For the stable margin reward R CSI Its composition formula is as follows:
[0140] R CSI =η·(CSI) t+1 -CSI t )
[0141] Where η is a positive coefficient used to amplify or reduce the reward value resulting from changes in the stability margin, CSI t The CSI is the composite stability margin index at the current moment (before the action). t+1 It is the composite stability margin index for the next moment (after the action).
[0142] Note: This directly rewards actions that improve the composite stability margin index (CSI). The coefficient η is positive. If the system stability margin improves (CSI) after the action is executed... t+1 <CSI t (Note that a smaller CSI indicates greater security) results in a positive reward; conversely, a larger CSI results in a penalty, which encourages agents to make forward-looking and defensive adjustments.
[0143] Regarding the adjustment of cost incentive R cost Its composition formula is as follows:
[0144] R cost =-κ·(∑|ΔP G,i ·c G,i |+∑|ΔP L,j ·c L,j |)
[0145] Where κ is a positive coefficient used to adjust the weight of economic costs in the total reward, ΔP G,i c is the adjustment amount of the active power output of generator i. G,i Let ΔP be the output adjustment cost coefficient for generator i. L,j c is the active power adjustment of load j. L,j Let be the adjustment cost coefficient for load j.
[0146] Note: The economic cost of this punitive control action. G,i and c L,j These are the cost coefficients for generator output and load adjustment, respectively. This guides the agent to adopt adjustment schemes with the lowest possible economic cost, while ensuring safety.
[0147] S5. Perform fault control on the power system according to the current fault control scheme.
[0148] Specifically, to fully explain the above steps, the following scheme will be used as an example:
[0149] During online operation, the system initiates a closed-loop control cycle: First, steps S1 to S4 are executed in real time. This involves calculating the composite stability margin index by collecting grid measurement data, then matching it with an offline stability scenario knowledge base to identify high-risk operating scenarios and their anticipated disturbance sets. These disturbances are then rapidly simulated to locate key power flow sections and extract the set of key influencing nodes. Subsequently, the system inputs the current operating state (including the aforementioned index, section power, and node output) into the reinforcement learning optimal strategy model generated offline by step S4. The model immediately outputs precise adjustment commands for generators and controllable loads within the set of key influencing nodes. After safety verification, these commands are sent to the grid actuators, completing one automatic control intervention. After the control action is executed, the system monitors the new grid state. If the success criteria of "all power flow voltage exceedances eliminated, composite stability margin index improved, and adjustment costs controllable" are met, the complete experience data of this intervention (including pre-control state, control commands, reward obtained, and post-control state) is stored as a successful control sample in the online experience pool.
[0150] To achieve continuous adaptive optimization, the system operates two feedback mechanisms in parallel: Feedback Mechanism One periodically scans the online experience pool. If it finds that the matching confidence of the initial state corresponding to a certain type of successful control with all existing scenarios in the knowledge base is consistently low, it automatically extracts its features as a new scenario and adds it to the offline knowledge base, thereby expanding the system's scenario recognition capabilities. Feedback Mechanism Two utilizes new samples from the online experience pool to fine-tune the deployed reinforcement learning model online during off-peak periods, enabling the model parameters to adapt to the latest operating characteristics of the power grid. Through this complete closed loop of real-time perception, decision-making, execution, evaluation, and feedback, the system achieves a leap from static automation to dynamic self-learning. Its stability boundary perception, fault location, and control strategies can continuously evolve with operating time, ultimately forming a smart grid proactive defense system with lifelong learning capabilities.
[0151] In summary, this application embodiment, by acquiring the current composite stability margin index calculated from real-time power system operation data, can dynamically and accurately grasp the current stable state of the system, avoiding the inability to adapt to the complex changes in the system stability boundary after the access of new energy sources due to reliance on offline data or static evaluation. When the index reaches a first preset threshold, the current fault operation scenario and corresponding expected disturbance set are determined by collecting a sample set of current operation scenarios and combining it with a preset offline stability scenario knowledge base. This allows for targeted focus on potential fault risks, avoiding ineffective analysis of irrelevant scenarios to save computing resources and effectively avoiding the problem of missing key faults. Based on the above fault scenarios and expected disturbance sets, rapid time-domain simulation is used to obtain dynamic response trajectories and further identify power flow sections and influencing nodes. The set of data points can accurately capture the dynamic changes of the system under the influence of faults, identify the key links and core influencing nodes of fault propagation, and provide a basis for the precise application of subsequent control measures. Then, a state space is constructed using a composite stability margin index, power flow sections, and the set of influencing nodes. Combined with a reward function that integrates multiple dimensions of objectives such as limit elimination, stability margin, adjustment cost, and constraints, the system is trained in a reinforcement learning environment to output a control scheme. This ensures that the scheme takes into account system safety, stability, economy, and operational feasibility, avoiding the shortcomings of control effects under a single objective. Finally, the implementation of fault control based on this scheme can effectively cope with the complex fault situations of the power system under the large-scale integration of new energy sources, significantly improve the real-time performance, accuracy, and adaptability of fault control, and ensure the safe and stable operation of the power system.
[0152] Example 2
[0153] like Figure 4 As shown, based on the above method embodiments, corresponding device embodiments are provided;
[0154] An embodiment of the present invention provides a power system fault control device oriented towards stability boundaries, comprising: a data acquisition module 41, a judgment module 42, a first calculation module 43, a second calculation module 44, and a fault control module 45;
[0155] The data acquisition module 41 is used to acquire the current composite stability margin index of the power system in real time; wherein, the current composite stability margin index is calculated based on the real-time operating data of the power system using the composite stability margin index calculation formula;
[0156] The judgment module 42 is used to collect the current operating scenario sample set of the power system when the current composite stability margin index is greater than or equal to the first preset threshold, and obtain the current fault operating scenario and the expected disturbance set corresponding to the current fault operating scenario based on the preset offline stability scenario knowledge base and the current operating scenario sample set.
[0157] The first calculation module 43 is used to perform simulation using fast time-domain simulation based on the current fault operation scenario and the expected disturbance set corresponding to the current fault operation scenario, to obtain the current dynamic response trajectory, and to perform power flow identification based on the current dynamic response trajectory, to obtain the current power flow section and the current set of affected nodes.
[0158] The second calculation module 44 is used to establish a state space based on the current composite stability margin index, the current power flow section, and the current set of affected nodes, and to train the system in a preset reinforcement learning environment based on the state space and the reward function, and output the current fault control scheme of the power system; wherein, the reward function is obtained by weighting the over-limit elimination reward, stability margin reward, adjustment cost reward, and constraint reward.
[0159] The fault control module 45 is used to perform fault control on the power system according to the current fault control scheme.
[0160] In some embodiments of this application, the data acquisition module 41 includes a first acquisition unit and a second acquisition unit;
[0161] The first acquisition unit is used to calculate the voltage stability sub-item, thermal stability sub-item, and power angle stability sub-item according to the real-time operation data of the power system in sequence using the voltage stability sub-item calculation formula, the thermal stability sub-item calculation formula, and the power angle stability sub-item calculation formula;
[0162] The second acquisition unit is used to perform a weighted geometric average fusion process on the voltage stability sub-item, the thermal stability sub-item, and the power angle stability sub-item to obtain the composite stability margin index.
[0163] In some embodiments of this application, the determination module 42 includes a first determination unit, a second determination unit, and a third determination unit;
[0164] The first judgment unit is used to match the current running scenario sample set with the scenario data in the preset offline stable scenario knowledge base to obtain the Mahalanobis distance factor, stability margin consistency factor and data completeness factor corresponding to each scenario.
[0165] The second judgment unit is used to perform weighted calculation of the Mahalanobis distance factor, stability margin consistency factor and data completeness factor corresponding to each of the scenario data to obtain the matching confidence level corresponding to each of the scenario data.
[0166] The third judgment unit is used to sequentially filter and match the matching confidence levels corresponding to each scenario data according to a preset confidence threshold, so as to obtain the expected disturbance set corresponding to the current fault operation scenario.
[0167] In some embodiments of this application, the first computing module 43 includes a first computing unit and a second computing unit;
[0168] The first calculation unit is used to make a judgment based on the current dynamic response trajectory using the stability margin slope decrease judgment method to obtain the current power flow section;
[0169] The second calculation unit is used to locate the current influencing nodes based on the current power flow section using the electrical coupling group algorithm and the power flow entropy weight method.
[0170] In some embodiments of this application, the second computing module 44 includes a third computing unit and a fourth computing unit;
[0171] The third calculation unit is used to weight the over-limit elimination reward, the stability margin reward, and the adjustment cost reward according to a preset weight set to obtain the original reward function;
[0172] The fourth calculation unit is used to obtain the reward function based on the original reward function and the constraint reward.
[0173] For more detailed steps and working principles of this embodiment, please refer to the relevant description in Embodiment 1, but not limited to these descriptions.
[0174] In summary, this application embodiment, by acquiring the current composite stability margin index calculated from real-time power system operation data, can dynamically and accurately grasp the current stable state of the system, avoiding the inability to adapt to the complex changes in the system stability boundary after the access of new energy sources due to reliance on offline data or static evaluation. When the index reaches a first preset threshold, the current fault operation scenario and corresponding expected disturbance set are determined by collecting a sample set of current operation scenarios and combining it with a preset offline stability scenario knowledge base. This allows for targeted focus on potential fault risks, avoiding ineffective analysis of irrelevant scenarios to save computing resources and effectively avoiding the problem of missing key faults. Based on the above fault scenarios and expected disturbance sets, rapid time-domain simulation is used to obtain dynamic response trajectories and further identify power flow sections and influencing nodes. The set of data points can accurately capture the dynamic changes of the system under the influence of faults, identify the key links and core influencing nodes of fault propagation, and provide a basis for the precise application of subsequent control measures. Then, a state space is constructed using a composite stability margin index, power flow sections, and the set of influencing nodes. Combined with a reward function that integrates multiple dimensions of objectives such as limit elimination, stability margin, adjustment cost, and constraints, the system is trained in a reinforcement learning environment to output a control scheme. This ensures that the scheme takes into account system safety, stability, economy, and operational feasibility, avoiding the shortcomings of control effects under a single objective. Finally, the implementation of fault control based on this scheme can effectively cope with the complex fault situations of the power system under the large-scale integration of new energy sources, significantly improve the real-time performance, accuracy, and adaptability of fault control, and ensure the safe and stable operation of the power system.
[0175] It is understood that the above-described device embodiments correspond to the method embodiments of the present invention, and can implement the power system fault control method oriented towards stability boundaries provided by any of the above-described method embodiments of the present invention.
[0176] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, in the accompanying drawings of the device embodiments provided by this invention, the connection relationships between modules indicate that they have communication connections, which can specifically be implemented as one or more communication buses or signal lines. Those skilled in the art can understand and implement this without any creative effort.
[0177] Example 3
[0178] Based on the above embodiments of the power system fault control method oriented towards stability boundaries, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, it implements the power system fault control method oriented towards stability boundaries according to any embodiment of the present invention.
[0179] For example, in this embodiment, the computer program can be divided into one or more modules, which are stored in the memory and executed by the processor to complete the present invention. The one or more modules may be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of the computer program in the terminal device.
[0180] The terminal device may be a desktop computer, laptop, handheld computer, or cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0181] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the terminal device, connecting all parts of the terminal device via various interfaces and lines.
[0182] Example 4
[0183] Based on the above-described method embodiments, another embodiment of the present invention provides a computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to execute the power system fault control method oriented towards stability boundaries as described in any of the above-described method embodiments of the present invention.
[0184] The modules / units integrated in the device / terminal equipment, if implemented as software functional units and sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0185] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for power system fault control oriented to stable boundaries, characterized in that, The method comprises the following steps: obtaining a current composite stability margin index of the power system in real time; wherein the current composite stability margin index is obtained by calculating a composite stability margin index calculation formula according to real-time operation data of the power system; if the current composite stability margin index is greater than or equal to a first preset threshold, collecting a current operating scene sample set of the power system, and obtaining a current fault operating scene and a set of expected disturbances corresponding to the current fault operating scene according to a preset offline stability scene knowledge base and the current operating scene sample set; performing simulation by using fast time-domain simulation according to the current fault operating scene and the set of expected disturbances corresponding to the current fault operating scene, obtaining a current dynamic response trajectory, and performing power flow identification according to the current dynamic response trajectory to obtain a current power flow section and a current set of influence nodes; establishing a state space according to the current composite stability margin index, the current power flow section and the current set of influence nodes, and training in a preset reinforcement learning environment according to the state space and a reward function to output a current fault control scheme of the power system; wherein the reward function is obtained by weighting an out-of-limit elimination reward, a stability margin reward, an adjustment cost reward and a constraint reward; controlling the power system according to the current fault control scheme.
2. A power system fault control method oriented to stable boundaries as claimed in claim 1, characterized in that, The current composite stability margin index is obtained by calculating a composite stability margin index calculation formula according to real-time operation data of the power system, and specifically comprises the following steps: calculating a voltage stability sub-item, a thermal stability sub-item and an angle stability sub-item by sequentially using a voltage stability sub-item calculation formula, a thermal stability sub-item calculation formula and an angle stability sub-item calculation formula according to the real-time operation data of the power system; performing weighted geometric mean fusion processing on the voltage stability sub-item, the thermal stability sub-item and the angle stability sub-item to obtain the composite stability margin index.
3. A power system fault control method oriented to stable boundaries as claimed in claim 1, characterized in that, The current fault operating scene and the set of expected disturbances corresponding to the current fault operating scene are obtained according to the preset offline stability scene knowledge base and the current operating scene sample set, and specifically comprise the following steps: matching the current operating scene sample set with each scene data in the preset offline stability scene knowledge base to obtain a Mahalanobis distance factor, a stability margin consistency factor and a data completeness factor corresponding to each scene; performing weighted calculation on the Mahalanobis distance factor, the stability margin consistency factor and the data completeness factor corresponding to each scene data to obtain a matching confidence of each scene data; screening and matching the matching confidence of each scene data according to a preset confidence threshold to obtain a set of expected disturbances corresponding to the current fault operating scene.
4. A stable boundary oriented power system fault control method according to claim 1, characterized by, The current power flow section and the current set of influence nodes are obtained by performing power flow identification according to the current dynamic response trajectory, and specifically comprise the following steps: judging the current power flow section by using a stability margin slope drop judgment method according to the current dynamic response trajectory; positioning the current set of influence nodes by using an electrical coupling group algorithm and a power flow entropy weight method according to the current power flow section.
5. A stable boundary oriented power system fault control method according to claim 1, characterized by, The reward function is obtained by weighting an out-of-limit elimination reward, a stability margin reward, an adjustment cost reward and a constraint reward, and specifically: The out-of-limit elimination reward, the stability margin reward and the adjustment cost reward are weighted according to a preset weight set to obtain an original reward function; The reward function is obtained according to the original reward function and the constraint reward.
6. A stable boundary oriented power system fault control device, characterized by, Comprise: Data acquisition module, judgment module, first calculation module, second calculation module and fault control module; The data acquisition module is used for acquiring the current composite stability margin index of the power system in real time; wherein the current composite stability margin index is obtained by calculating the real-time operation data of the power system using a composite stability margin index calculation formula; The judgment module is used for collecting the current operating scene sample set of the power system if the current composite stability margin index is greater than or equal to the first preset threshold, and obtaining the current fault operating scene and the expected disturbance set corresponding to the current fault operating scene according to the preset offline stable scene knowledge base and the current operating scene sample set; The first calculation module is used for simulating the current dynamic response trajectory by using fast time domain simulation according to the current fault operating scene and the expected disturbance set corresponding to the current fault operating scene, and obtaining the current flow section and the current influence node set by performing power flow identification according to the current dynamic response trajectory; The second calculation module is used for establishing a state space according to the current composite stability margin index, the current flow section and the current influence node set, and outputting the current fault control scheme of the power system by training in a preset reinforcement learning environment according to the state space and the reward function; wherein the reward function is obtained by weighting the out-of-limit elimination reward, the stability margin reward, the adjustment cost reward and the constraint reward; The fault control module is used for controlling the power system according to the current fault control scheme.
7. A stable boundary oriented power system fault control device as claimed in claim 6, characterized in that, The data acquisition module comprises a first acquisition unit and a second acquisition unit; The first acquisition unit is used for calculating the voltage stability sub-item, the thermal stability sub-item and the power angle stability sub-item by using the voltage stability sub-item calculation formula, the thermal stability sub-item calculation formula and the power angle stability sub-item calculation formula in sequence according to the real-time operation data of the power system; The second acquisition unit is used for performing weighted geometric mean fusion processing on the voltage stability sub-item, the thermal stability sub-item and the power angle stability sub-item to obtain the composite stability margin index.
8. A stable boundary oriented power system fault control device as recited in claim 6, wherein, The judgment module comprises a first judgment unit, a second judgment unit and a third judgment unit; The first judgment unit is used for matching the current operating scene sample set with each scene data in the preset offline stable scene knowledge base to obtain the Mahalanobis distance factor, the stability margin consistency factor and the data completeness factor corresponding to each scene; The second judgment unit is used for performing weighted calculation on the Mahalanobis distance factor, the stability margin consistency factor and the data completeness factor corresponding to each scene data to obtain the matching confidence of each scene data; The third judging unit is configured to filter and match the matching confidence corresponding to each scene data in sequence according to a preset confidence threshold, and obtain a set of expected disturbances corresponding to the current fault operation scene.
9. A stable boundary oriented power system fault control device as recited in claim 6, wherein, The first calculation module comprises a first calculation unit and a second calculation unit. The first calculation unit is configured to judge the current dynamic response trajectory by using a stable margin slope drop judgment method, and obtain a current power flow section. The second calculation unit is configured to locate the current power flow section by using an electrical coupling group algorithm and a power flow entropy weight method, and obtain a current influence node set.
10. A stable boundary oriented power system fault control device as recited in claim 6, wherein, The second calculation module comprises a third calculation unit and a fourth calculation unit. The third calculation unit is configured to weight the out-of-limit elimination reward, the stable margin reward and the adjustment cost reward according to a preset weight set, and obtain an original reward function. The fourth calculation unit is configured to obtain the reward function according to the original reward function and the constraint reward.