New energy power system cascading failure identification method and system under extreme weather
By combining large language models and deep reinforcement learning fault chain search model, the chain failure of new energy power system under extreme weather is solved in real time, and the problem of not being able to capture dynamic changes in the power grid in real time is achieved, achieving more efficient fault path prediction and power system safety management.
Patent Information
- Application Number
- CN202510480942.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-04
AI Technical Summary
The existing technology cannot capture dynamic changes in the power grid in real time under extreme meteorological conditions, resulting in the prediction of chain fault paths deviating from actual evolution, and failing to effectively integrate historical fault cases, and the efficiency of chain fault path generation is inefficient.
The large language model (LLM) is used to process natural language information, and the fault chain search model is built in combination with deep reinforcement learning (DRL). The chain faults of the new energy power system under extreme weather are identified through the ∈-greed strategy and weighted current method, and the fault chain model is dynamically built to capture the power grid changes in real time.
It improves the accuracy and efficiency of chain fault identification, enhances the model's adaptability in the face of uncertainty and complexity, and ensures the reliability and safety of power supply.
Smart Images

Figure CN120257834A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power system fault detection, and more specifically, relates to a method and system for identifying cascading faults in a new energy power system under extreme weather conditions. Background Art
[0002] In modern power systems, with the rapid development of renewable energy and the continuous growth of power demand, the complexity and uncertainty of the system have increased significantly. This complexity makes the power system face various risks during operation, especially the risk of cascading faults. A cascading fault refers to a series of subsequent faults triggered by the occurrence of one fault, ultimately leading to the widespread failure of the system. This not only has a serious impact on the reliability of power supply but may also cause economic losses and safety hazards. Therefore, timely and accurately identifying the risk of cascading faults is crucial for ensuring the safe and stable operation of the power system.
[0003] The prior art document 1 (CN117374915A) discloses a method and device for generating cascading fault paths of a Markov Process starting from a meteorological disaster. Its disadvantages are that it cannot capture the dynamic changes of the power grid in real time, resulting in the deviation of the predicted path from the actual evolution, and it does not integrate historical fault cases, resulting in the lack of prior knowledge in the model and low efficiency in generating cascading fault paths. Summary of the Invention
[0004] To solve the deficiencies in the prior art, the present invention provides a method and system for identifying cascading faults in a new energy power system under extreme weather conditions, which uses an LLM (Large Language Model) to process and understand complex natural language information, extracts historical fault cases in the operation of the power system, and helps DRL (Deep Reinforcement Learning) reduce sample complexity. This combination not only improves the learning efficiency of the model but also enhances its adaptability in the face of uncertainty and complexity, ensuring the reliability and safety of power supply.
[0005] The present invention adopts the following technical solutions.
[0006] The first aspect of the present invention provides a method for identifying cascading faults in a new energy power system under extreme weather conditions, specifically including:
[0007] Establish a fault chain model for a new energy power system under extreme weather conditions;
[0008] According to the fault model of the new energy power system under extreme weather conditions, establish a search model for the maximum load loss fault chain;
[0009] Convert the maximum load loss fault chain search model into a POMDP, and construct the optimal action selection strategy of the POMDP;
[0010] Fuse the deep reinforcement learning algorithm of the large language model to construct the LLM-DRL fault chain search model;
[0011] Solve the optimal action selection strategy of the POMDP through the ∈-greedy strategy, generate a set of fault chains, and realize the identification of cascading faults in the new energy power system under extreme weather.
[0012] Among them, with a probability of ∈, the power system data under actual extreme weather is solved for fault chains using the weighted power flow method. With a probability of 1 - ∈, the power system data under actual extreme weather is input into the LLM-DRL fault chain search model for fault chain solution.
[0013] Preferably, the maximum load loss fault chain search model is represented by the following formula:
[0014]
[0015] In the formula,
[0016] represents the set of S fault chains of the maximum load loss,
[0017] represents the s-th fault chain of the maximum load loss,
[0018] S represents the number of fault chains of the maximum load loss,
[0019] represents the set of all fault chains,
[0020] represents the s-th fault chain causing the total load loss.
[0021] Preferably, the represents the total load loss caused by the s-th fault chain and is represented by the following formula:
[0022]
[0023] In the formula,
[0024] represents the total load loss caused by
[0025] represents the fault chain of the sequence of power equipment that fails in P stages,
[0026] Indicates the load loss caused by faults in the i-th stage,
[0027] Indicates the set of power equipment that has failed in , i ∈ [P], where P represents the total number of stages,
[0028] the topological structure of the power system in the i-th stage,
[0029] W xeternal represents the influence parameter of extreme meteorological conditions.
[0030] Preferably, the optimal action selection strategy for constructing the POMDP is expressed by the following formula:
[0031]
[0032] In the formula,
[0033] π * represents the optimal action selection strategy,
[0034] represents the expected value operator,
[0035] V π (S0) represents the total feedback obtained by the agent starting from the initial POMDP state S0 for any action selection strategy π.
[0036] Preferably, the deep reinforcement learning algorithm that integrates the large language model constructs an LLM-DRL fault chain search model, specifically including:
[0037] Construct a primary fault chain searcher based on the large language model;
[0038] Based on the primary fault chain searcher of the large language model and the deep reinforcement learning model, construct an LLM-DRL fault chain searcher.
[0039] Preferably, the construction of the primary fault chain searcher based on the large language model specifically includes:
[0040] Set the task description;
[0041] The large language model generates a description of the task based on the set task description;
[0042] The large language model generates the code of the primary fault chain searcher based on the set programming guide and in combination with the generated task description;
[0043] Perform the actual fault chain search operation according to the fault chain primary searcher code, and optimize the rules of the fault chain primary searcher code using the manual feedback mechanism to obtain the fault chain primary searcher based on the large language model.
[0044] Preferably, under the probability of 1 - ∈, input the power system data under the actual extreme weather into the LLM-DRL fault chain search model to solve the fault chain, specifically including:
[0045] The fault chain primary searcher based on the large language model generates initial fault chain action data according to the natural language description data in the power system data under the actual extreme weather environment, where the initial fault chain action data is stored in the large language model buffer;
[0046] After passing the initial fault chain action data through imitation learning, input it into the deep reinforcement learning algorithm to solve the current state, the action a taken by the agent with a probability of 1 - ∈ i,1-∈ , the reward discount factor, and the next state to obtain the deep reinforcement learning interaction data, where the deep reinforcement learning interaction data is stored in the deep reinforcement learning buffer;
[0047] Mix and sample data from the large language model buffer and the deep reinforcement learning buffer, update the evaluation network parameters and action network parameters of the evaluation network and action network of the deep reinforcement learning respectively, input the objective function of the LLM-DRL fault chain searcher for policy gradient calculation, and obtain the current optimal action according to the optimal action selection strategy;
[0048] Repeat the iteration until the training converges, and generate the fault chain according to all the optimal actions.
[0049] Preferably, the action a taken by the agent with a probability of 1 - ∈ i,1-∈ , is represented by the following formula:
[0050]
[0051] In the formula,
[0052] represents the Q value of the agent based on the current state and action,
[0053] Y i represents the input feature in stage i,
[0054] θ represents the parameter for learning the Q value function,
[0055] represents the number of times the power equipment i is selected when the agent is in the POMDP state S .
[0056] Preferably, the objective function of the LLM-DRL fault chain searcher is expressed by the following formula:
[0057]
[0058] where
[0059] represents the model of the interaction between the agent and the environment,
[0060] represents the expected value,
[0061] s0~p0 means that the initial state s0 is sampled according to the initial state distribution p0,
[0062] s t+1 ~p(·∣s t ,a t ) represents the state s t at time step t after the agent takes the action a t at time step t, and transfers to the next state s t ,a t ) according to the state transition probability p(·∣s t+1 ;
[0063] a t ~π(·∣s t ) represents that the agent selects the action a t at time step t under the state s t ) of time step t according to the policy π(·∣s t ,
[0064] γ t ∈[0,1] represents the reward discount factor;
[0065] represents the sparse reward function
[0066] The second aspect of the present invention provides a new energy power system chain fault identification system under extreme weather, which operates the new energy power system chain fault identification method described in the first aspect of the present invention, and specifically includes:
[0067] A fault chain construction module, which is used to establish a fault chain model of the new energy power system under extreme weather;
[0068] A loss model construction module, which is used to establish a maximum load loss fault chain search model according to the fault model of the new energy power system under extreme weather;
[0069] A policy construction module, which is used to convert the maximum load loss fault chain search model into a POMDP and construct an optimal action selection policy for the POMDP;
[0070] The fault chain search model construction module integrates the deep reinforcement learning algorithm of the large language model to construct the LLM-DRL fault chain search model;
[0071] The fault chain solving module is used to solve the optimal action selection strategy of POMDP through the ∈-greedy policy, generate a set of fault chains, and realize the identification of cascading faults in the new energy power system under extreme weather.
[0072] Among them, under the probability ∈, the power system data under actual extreme weather is solved for fault chains using the weighted power flow method, and under the probability 1 - ∈, the power system data under actual extreme weather is input into the LLM-DRL fault chain search model for fault chain solving.
[0073] Compared with the prior art, the beneficial effects of the present invention at least include: The present invention proposes a method and system for identifying cascading faults in a new energy power system under extreme weather based on the combination of LLM and DRL. Through the policy network iteratively selects the optimal action, captures the changes in the power grid in real time, dynamically constructs a fault chain model, can timely identify the changes in the power system, and improves the accuracy of the prediction path. The present invention uses LLM to process and understand complex natural language information, extracts historical fault cases in the operation of the power system, such as, but not limited to, valuable prior knowledge of historical data, fault modes, and environmental factors, and can better understand the dynamic changes and potential risks of the system. Deep reinforcement learning usually requires a large number of interaction samples to achieve high performance, while the knowledge extraction ability of LLM can provide valuable guidance in the learning process, thereby accelerating the learning process and reducing the dependence on environmental samples, helping DRL reduce sample complexity. This combination not only improves the learning efficiency of the model but also enhances its adaptability in the face of uncertainty and complexity. Through the combination of LLM and DRL, the present invention can quickly identify potential cascading faults, provide a new solution for the safety management of the power system, realize a more intelligent and efficient power network, and ensure the reliability and security of power supply. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] Figure 1 is a schematic diagram of the identification process of cascading faults in a new energy power system under extreme weather provided according to an embodiment of the present invention;
[0075] Figure 2 is a schematic diagram of the deep reinforcement learning algorithm architecture integrating a large language model provided according to an embodiment of the present invention;
[0076] Figure 3 is a schematic diagram of the architecture of the primary fault chain searcher based on LLM provided according to an embodiment of the present invention;
[0077] Figure 4 is a schematic diagram of the load loss accuracy provided according to an embodiment of the present invention;
[0078] Figure 5 is a schematic diagram of the discovery rate of risk fault chains provided according to an embodiment of the present invention. Detailed implementation manners
[0079] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0080] As Figure 1 shown, Embodiment 1 of the present invention provides a method for identifying cascading failures in a new energy power system under extreme weather conditions, including the following steps:
[0081] Step 1, establish a search model for the maximum load loss fault chain of the power system.
[0082] Step 1.1, establish a fault chain model for the new energy power system.
[0083] This model is used to describe the cascading fault chains caused by internal factors (such as system instability) and external factors (such as extreme weather conditions). The interaction of power equipment may trigger larger-scale power grid failures. The characteristics of the model are to dynamically identify the system topology changes caused by power equipment failures and quantify the impacts of each fault chain through TLL (Total Load Loss). Due to a series of internal factors such as but not limited to system instability) and external factors such as but not limited to extreme weather conditions, power equipment in the power system may fail. When the number of power equipment failures reaches a certain level, more failures may be triggered, thus forming a fault chain. The objective of the present invention is to dynamically identify potential fault chains in the system and evaluate their risks.
[0084] In a preferred but non-limiting implementation manner of the present invention, to describe the connection relationship of each bus in the system, the construction process of the topological structure and system state of a power system composed of N buses specifically includes:
[0085] Use an undirected graph to describe, where represents a bus, represents a transmission line. B ∈ {0, 1} N×N represents a binary adjacency matrix, Indicates that there is a transmission line E between bus u ∈ [N] and bus v ∈ [N]. Define as the power system state matrix, which represents the F types of system state parameters (e.g., voltage angle, injected active power) of all N buses, and use to represent any system state parameter f ∈ {1, …, F} of bus u ∈ [N].
[0086] Further preferably, represent the fault - free system topology before the occurrence of the fault chain as The corresponding system state is represented by X0. The possible fault chains that the system may face are specified as a series of consecutive power equipment failures. Consider a fault chain model consisting of at most P stages, where P can be selected according to the scope of concern of the risk assessment, and use the set of all power equipment in the system as to represent. It should be noted that the number of stages in different fault chains can be different. If the power equipment corresponding to the stage number has a fault chain that terminates before the stage, this means that the set will be an empty set.
[0087] More preferably, at each stage i ∈ [P], some additional power equipment fails. Define the set of power equipment that fails in stage i ∈ [P] as At any stage i, may consist of more than one power equipment failure. At the same time, when a transmission line fails and it has parallel channels, only the faulty line will be included in , while its parallel lines will be retained as normally operating power equipment. Obviously, and the fault chain of the sequence of power equipment that fails in P stages represents:
[0088]
[0089] where
[0090] represents the fault chain of the sequence of power equipment that fails in P stages,
[0091] represents the set of power equipment that fails in the i - th stage , i ∈ [P].[[]END]]
[0092] Use to represent the normal power equipment in the i - th stage of the fault chain. When the fault process is relatively slow (e.g., in the early stage of a cascading failure), the set There are fewer normal power devices in [it]; while when the fault process is rapid (e.g., in the final stage of a cascading fault), there are more normal power devices in these sets.
[0093] Step 1.2, according to the new energy power system fault chain model, construct the total load loss model caused by the fault chain.
[0094] In a preferred but non-limiting embodiment of the present invention, due to the occurrence of faults in power devices in each stage i, the power system topology will become The corresponding system adjacency matrix is B i and the system state X i . Continuous faults will continuously accumulate pressure on the remaining power devices, and these power devices need to ensure the minimum load reduction in the network. Represent the LL (Load Loss) caused by the fault of the power devices in [it] in stage i with and express it with the following formula:
[0095]
[0096] Further preferably, the total load loss model caused by the fault chain is calculated by the following formula:
[0097]
[0098] In the formula,
[0099] represents the total load loss caused by
[0100] represents the fault chain of the sequence of power devices that fail in P stages,
[0101] represents the load loss caused by the fault in the i-th stage,
[0102] represents the i-th stage the set of power devices that fail in [it], i ∈ [P],
[0103] represents the topology of the power system in the i-th stage,
[0104] the set of device faults represents the total load under the system topology
[0105] W xeternal An impact parameter representing extreme meteorological conditions, used to reflect external factors, such as but not limited to the impact of extreme meteorological conditions, can be dynamically assigned according to meteorological parameters such as but not limited to wind speed and temperature.
[0106] Step 1.3, establish a search model for the maximum load loss fault chain.
[0107] This model constructs multiple stages of fault chains by defining a set of all possible fault chains to dynamically capture the impact of system state changes on the load. The model emphasizes the combined effect of power equipment failures at each stage and quantifies the corresponding load loss.
[0108] After each stage of the fault chain, the power is redistributed in the transmission lines. Since the system state changes continuously over time, different loads and topological conditions face different risks. For any given initial system state X0, it is extremely important to efficiently and real-time search for the set of fault chains with the maximum total line loss.
[0109] The objective of the present invention is to determine S fault chains, each consisting of P stages, which will result in the maximum total line loss. Define as the set of all possible fault chains within the target time range P, and the objective is to determine S sequences with the maximum load loss in
[0110] In a preferred but non-limiting embodiment of the present invention, the search model for the maximum load loss fault chain is represented by the following formula:
[0111]
[0112] In the formula,
[0113] represents the set of S fault chains with the maximum load loss,
[0114] represents the s-th fault chain with the maximum load loss,
[0115] S represents the number of fault chains with the maximum load loss,
[0116] represents the set of all fault chains,
[0117] represents the s-th fault chain causing the total load loss.
[0118] Search model for the maximum load loss fault chain Aims to maximize the load loss accumulated due to S fault chains. Without loss of generality, assume the sequence set The total load losses are arranged in descending order, i.e., Solving this problem faces a huge computational burden because all the sets of fault chains The cardinality of which grows exponentially with the number of buses N, the number of power equipment in the system, and the number of stages P.
[0119] Step 2: Transform the search process of the fault chain with the maximum load loss into a POMDP, and construct the optimal action selection strategy of the POMDP.
[0120] In a preferred but non-limiting embodiment of the present invention, the search process of the fault chain with the maximum load loss is transformed into a partially observable Markov decision process. By defining the state, observation, and action of the agent, the TLL fault chain search process is modeled as a POMDP problem. In this model, the agent selects the action to be executed based on the currently observed system state and the past observation history to maximize the future cumulative reward. In particular, the impact of immediate rewards such as load loss is considered, and the time dependence between multiple stages is captured by updating the system state. This modeling method can effectively evaluate and select the most risky fault chain at each stage.
[0121] During the search process, in order to reduce the computational complexity, the present invention uses the system state at stage i to determine the system state at stage (i + 1) ( , X i+1 ). However, the load loss at stage (i + 1) depends on all the past i stages and the set of power equipment removed in those system states.
[0122] In a preferred but non-limiting embodiment of the present invention, it is defined that represents the system state of the agent at stage i, the POMDP state at stage i, and is expressed by the following formula:
[0123]
[0124] In the formula,
[0125] S i represents the POMDP state at the i-th stage,
[0126] represents the system state of the agent at the i-th stage,
[0127] represents the topological structure of the power system at the i-th stage,
[0128] X i represents the power system state at the i-th stage,
[0129] Let Si The POMDP state at stage i, which describes the entire past observation sequence O generated in the process i+1 .
[0130] Further preferably, at stage i, the agent has only the O i sequence information; at stage i + 1, when the agent receives O i , the agent will select a power device from the set of available power devices to remove in the next stage. To represent this process, the action of the agent is defined as the power device it selects, denoted by a i to represent the action at stage i, and the action space is defined as the set of all remaining components Once the agent takes an action at stage i the potential POMDP state for the next stage is randomly drawn from a transition probability distribution :
[0131]
[0132] where
[0133] S′ i+1 represents the potential POMDP state for the next stage i + 1 of S i ,
[0134] a i represents the action at stage i,
[0135] represents the set of all remaining components,
[0136] represents the transition probability distribution of the power system dynamics.
[0137] Further preferably, the probability distribution captures the randomness generated due to the dynamics of the power system, and it is determined by the rescheduling strategy at each stage of the fault chain. To quantify the immediate reward r(S′ i when transitioning from the POMDP state S i+1 to S′ i by taking the action a i+1 ∣S i ,a i ):
[0138]
[0139] where
[0140] r(S′ i+1 ∣S i ,ai ) represents the immediate reward for taking action a when transitioning from the POMDP state S i to S′ i+1 at time i ,
[0141] represents the total load under, and its result directly affects the evaluation of the agent's action strategy.
[0142] Further preferably, for any action selection strategy π, the total feedback obtained by the agent starting from the initial POMDP state is represented by the following formula:
[0143]
[0144] wherein,
[0145] V π (S0) represents the total feedback obtained by the agent starting from the initial POMDP state S0 for any action selection strategy π,
[0146] r(S i+1 ∣S i ,π(O i )) represents the immediate reward for taking action π(O i ) when transitioning from the POMDP state S i+1 to S′ i at time
[0147] represents the discount factor, which determines the preference degree of future feedback relative to immediate feedback,
[0148] π(O i ) represents the action selected by the agent at stage i ∈ [P] given the observation O i at time
[0149] More preferably, to find an optimal action selection strategy π * for the agent, the problem can be represented as follows:
[0150]
[0151] wherein,
[0152] π * represents the optimal action selection strategy,
[0153] represents the expected value operator, which is used to calculate the cumulative expected reward that the agent can obtain starting from the initial state S0 under the given policy π.
[0154] Step 3, use the ε-greedy strategy to solve the optimal action selection strategy of the POMDP.
[0155] In a preferred but non-limiting embodiment of the present invention, an ∈-greedy search strategy is used to solve the partially observable Markov decision process (POMDP). Among them, with a probability of ∈, the weighted power flow method is used for action exploration; with a probability of 1 - ∈, a deep reinforcement learning (DRL) algorithm integrating a large language model (LLM) is used for action exploration, and this algorithm is abbreviated as LLM-DRL.
[0156] In a preferred but non-limiting embodiment of the present invention, since the failure of power equipment carrying a high power will cause the remaining components to be easily overloaded, the detection strategy of the agent is to remove the power equipment carrying the maximum power flow. At any stage i of the fault chain, the agent has a probability of ∈ to select the jth power equipment according to the weighted power flow method, that is
[0157]
[0158] where
[0159] represents the absolute value of the power flowing through the power equipment ,
[0160] represents the number of times the power equipment i is selected when the agent is in the POMDP state S ,
[0161] a i,∈ represents the action at stage i with a probability of ∈,
[0162] represents the set of all remaining devices,
[0163] j represents the index used to traverse or reference the devices in the set .
[0164] k represents the index used to traverse the devices in the set .
[0165] In the remaining probability of 1 - ∈, the agent selects actions according to the Q-values learned through LLM-DRL.
[0166] In a preferred but non-limiting embodiment of the present invention, with a probability of 1 - ∈, a deep reinforcement learning algorithm integrating a large language model is used for solving. The selection of actions is proportional to the Q-value of each POMDP state-action visit count to avoid repeating previously discovered fault chains. The deep reinforcement learning algorithm integrating the large language model (LLM-DRL) is used as a fault chain searcher, specifically including: a primary fault chain searcher based on LLM and a fault chain searcher integrating LLM and DRL.
[0167] In a preferred but non-limiting embodiment of the present invention, a primary fault chain searcher based on LLM is constructed. The primary fault chain searcher based on LLM extracts the prior knowledge of the large language model through prompt engineering to generate a primary searcher for the fault chain search task. The generated primary searcher is based on the set rules and is written in a programming language (such as Python) to generate a primary fault chain searcher based on the set rules, and the primary fault chain searcher is optimized by combining observation results and human feedback, specifically including three steps: task description, programming guide, and human feedback. Among them, the task description and the programming guide are executed in an open-loop manner in sequence, while the human feedback is optimized according to the actual performance of the searcher.
[0168] Further preferably, the task description of the primary fault chain searcher based on LLM provides a standard framework. After receiving the task description, LLM generates a stage description of the task. This information will be used together with the code guide to guide LLM to generate the searcher code. The settings of the task description of the primary fault chain searcher based on LLM specifically include:
[0169] Searcher information, which describes the functions, logic, and related operation constraints of the primary fault chain searcher. This covers how the searcher conducts fault chain search based on preset rules and observed patterns;
[0170] Task processing process description, which details the core tasks that the primary fault chain searcher needs to complete, including the identification, classification of different fault situations, and their manifestations in the system. This part also needs to involve the processing of fault data and the feedback mechanism;
[0171] Problem description, which clarifies the key problems that LLM needs to solve, such as how to effectively identify actual faults from a large number of potential fault signals and predict possible fault propagation paths;
[0172] Answer template, which formulates a unified format and standard for standardizing the fault reports and warning information output by the searcher to ensure the consistency and accuracy of the output;
[0173] Rules, which set clear rules to help the primary fault chain searcher better understand and execute the task requirements, including the logical rules for fault determination and the processing flow for optimization feedback.
[0174] After receiving the task description, the LLM generates a description of the task, and this information will be used together with the code guidelines to guide the LLM to generate the searcher code.
[0175] Further preferably, the programming guidelines are used to provide detailed steps and technical requirements to ensure the generation of an efficient primary fault chain searcher from the LLM. Using the set programming guidelines and combining with the description of the generated task, the primary fault chain searcher code is generated, specifically including:
[0176] Control inputs. To ensure that the generated searcher can effectively identify faults in the system, all possible control inputs need to be predefined in advance, which includes the receiving format of fault data, the identification parameters of fault flags, etc. From these, the LLM will select a suitable subset for processing;
[0177] Rules. Provide specific programming guidelines to guide the use of programming (such as Python and its libraries) to create a rule-based searcher. For example, conditional statements and loops can be used to process fault data, and at the same time, exception handling is combined to optimize the fault response ability of the searcher;
[0178] The programming guidelines will ensure that developers can effectively construct and optimize the primary fault chain searcher according to the predefined standards and rules, thereby improving its accuracy and efficiency in actual operations.
[0179] Further preferably, during the actual operation of the primary searcher generated based on the programming guidelines, it may be found that the searcher performs poorly in specific situations. For example, when dealing with complex fault chains, the searcher may not be able to accurately identify the fault source or the fault propagation path. In this case, an artificial feedback mechanism is adopted for modification and optimization, specifically including:
[0180] After the initial application of the searcher, according to the actual performance, developers and maintainers can provide feedback on which rules or parameters need to be adjusted. For example, increasing the sensitivity to certain fault characteristics or adjusting the judgment logic of the fault chain. This feedback is not only based on technical analysis but also includes the experience and intuition of actual operators.
[0181] Compared with the traditional feedback process, this method encourages the use of simple natural language and intuitive operation feedback, enabling non-professionals to easily participate in the optimization process of the searcher, thus achieving performance improvement faster. Such a feedback mechanism enhances the adaptability and accuracy of the searcher, ensuring efficient fault handling performance in various operating environments.
[0182] The architecture of the primary fault chain searcher based on the LLM is as Figure 3 shown.
[0183] In a preferred but non-limiting embodiment of the present invention, the fault chain searcher integrating LLM and DRL constructs an LLM-DRL fault chain searcher, and the DRL algorithm architecture integrating LLM is as Figure 2 shown, specifically including:
[0184] The primary fault chain searcher based on the large language model generates initial fault chain action data according to the natural language description data of the current environment. Among them, the initial fault chain action data is stored in the large language model buffer. The primary fault chain searcher based on LLM is used to generate action samples during the interaction with the environment, thereby reducing the sample size required by DRL and improving the sample efficiency of DRL. The fault chain searcher integrating LLM and DRL combines the initial fault chain action data collected by LLM and the interaction data collected by the DRL online training strategy. The collected initial fault chain action data is stored in the LLM buffer (RLLM) and directly used for imitation learning to quickly improve the search ability of the searcher for common fault chains.
[0185] Further preferably, after the initial fault chain action data is processed by imitation learning, it is input into the deep reinforcement learning algorithm to solve the current state, the action taken by the agent with a probability of 1 - ∈, the reward discount factor, and the next state, and obtain the deep reinforcement learning interaction data. Among them, the deep reinforcement learning interaction data is stored in the deep reinforcement learning buffer, specifically including:
[0186] The data processed by imitation learning indirectly affects the DRL process, and the samples collected in the DRL buffer (RDRL) are optimized by adjusting the strategy in DRL training. This method of combining the efficient pattern recognition of LLM and the dynamic learning ability of DRL can significantly improve the adaptability and efficiency of the fault chain searcher in variable and complex fault environments. Through this method, the fault chain searcher can not only quickly learn from historical fault data, but also adjust its response strategy through continuous online interaction to ensure that it can effectively handle unknown or mutated fault situations in practical applications. Such a system design enables the fault chain searcher to have high flexibility and scalability while maintaining high precision.
[0187] LLM-DRL action loss function is constructed as follows:
[0188]
[0189] Among them,
[0190] λ IM represents a hyperparameter;
[0191] θ π represents the action network parameter;
[0192] represents the LLM action loss function;
[0193] represents the DRL action loss function.
[0194] To solve for θ π , an LLM-DRL objective function is established.
[0195] The agent has a probability of 1 - ∈ of taking action a i,1-∈ , that is:
[0196]
[0197] where
[0198] represents the Q value of the agent based on the current state and action,
[0199] Y i represents the input feature of stage i,
[0200] θ represents the parameter for learning the Q value function,
[0201] θ represents the parameter for learning the Q value function. During training, the probability ∈ is dynamically adjusted, that is:
[0202]
[0203] where
[0204] ∈0 represents the lowest level of exploration,
[0205] represents the set of candidate devices in the initial stage, that is, at the beginning of the algorithm, the set of all selectable devices for the agent in state S0.
[0206] Further preferably, data is sampled from the large language model buffer and the deep reinforcement learning buffer in a mixed manner to construct the objective function of the LLM-DRL fault chain searcher, specifically including:
[0207] The goal of LLM-DRL is to find an optimal action selection policy represented by the parameter θ π to maximize the objective function
[0208]
[0209] where
[0210] s t represents the state;
[0211] at Denote an action;
[0212] Denote the distribution of the initial state;
[0213] Denote the state transition probability, where Δ(·) is the simplex symbol;
[0214] Denote the sparse reward function for indicating whether the task is completed;
[0215] γ ∈ [0, 1] denotes the reward discount factor,
[0216] Denote the expectation value for all possible state trajectories s1, s2, s3, ..., s under the state distribution p(·∣s t , a t ), a t ~π(·∣s t ). t
[0217] s0 ~ p0 means that the initial state s0 is sampled according to the initial state distribution p0(s).
[0218] s t+1 ~ p(·∣s t , a t ) means that after the agent takes the action a t , the state s t transfers to the next state s t , a t ) according to the state transition probability p(·∣s t+1 .
[0219] a t ~ π(·∣s t ) means that the agent selects the action a t under the state s t ) according to the policy π(·∣s t .
[0220] π(·∣s t ) represents the policy function, which describes the probability distribution p(·∣s t ) of the agent selecting each action at any state
[0221] Denote the objective function, which describes the performance of the agent under the given model M and policy π.
[0222] π represents the policy.
[0223] Denote the model of the interaction between the agent and the environment
[0224] t represents the index of the time step and is used to track the decision-making sequence of the agent.
[0225] Further preferably, the evaluation network parameters and action network parameters of the deep reinforcement learning are updated respectively, and the policy gradient is calculated by inputting the objective function of the LLM-DRL fault chain searcher. After training converges, the optimal action selection policy is obtained, specifically including:
[0226] Use the LLM to improve the sample efficiency of the DRL algorithm. The DRL maximizes the objective function by finding a deterministic policy The approximation of the target evaluation network through the evaluation network loss function l is expressed by the following formula: critic as follows:
[0227]
[0228] where
[0229] N′ is the total number of time steps sampled from the replay buffer,
[0230] s n represents the state at time step n,
[0231] a n represents the action selected at time step n,
[0232] y n represents the best estimate of the evaluation network.
[0233] The y n represents the best estimate of the evaluation network and is expressed by the following formula:
[0234]
[0235] where
[0236] r n represents the immediate reward at the current time step,
[0237] γ represents the discount factor,
[0238] represents the target Q network,
[0239] represents the target action selection policy,
[0240] s ′ n represents the next state at time step n,
[0241] represents the evaluation network Parameter θ Q The mean value of
[0242] ∈ n represents the regularization value function.
[0243] The strategy network gradient update is adopted in the action network combined with the target evaluation network, which is represented by the following formula:
[0244]
[0245] In the formula,
[0246] represents the gradient of the target function J(θ π ),
[0247] N′ is the total number of time steps sampled from the replay buffer,
[0248] represents the target Q network The gradient of the action a,
[0249] represents the target action selection policy network The gradient of θ,
[0250] represents that in the state s n Under the condition, the target action selection policy network The output action a.
[0251] The policy network After being trained via formula 17, it provides the optimal action for each step, and then screens out high-risk actions through the Q-value evaluation of formula 15, and forms a complete fault chain through multiple-step iteration.
[0252] It should be noted that the prior art cannot capture the dynamic changes of the power system in real time, resulting in the deviation of the prediction path from the actual evolution. The lack of integration of historical fault cases leads to the lack of prior knowledge in the model, and the efficiency of generating the cascading fault path is low. Through the policy network Iteratively select the optimal action, capture the changes of the power grid in real time, dynamically construct a fault chain model, can timely identify the changes of the power system, improve the accuracy of the prediction path, process and understand complex natural language information through the LLM, so as to extract valuable prior knowledge of historical fault cases in the operation of the power system, such as but not limited to historical data, fault patterns and environmental factors, which can provide system dynamic changes and potential risk information for DRL, thus accelerating the learning process of DRL and reducing the dependence on environmental samples, improving the efficiency of the model to generate fault paths, and the use of LLM also enhances the adaptability of DRL in the face of uncertainty and complexity.
[0253] Step 4: Judge the accuracy of the set of fault chains solved in Step 3 to achieve the identification of cascading faults in the new energy power system under extreme weather.
[0254] Judge the accuracy of the set of fault chains, specifically including: the accuracy of load loss, the difference between the cumulative load loss searched by the fault chain searcher and the maximum cumulative load loss (given by the set For s ∈ [S], the accuracy of load loss is expressed by the following formula:
[0255]
[0256] Among them,
[0257] Regret(s) represents the accuracy of load loss,
[0258] A lower value of Regret(s) indicates higher accuracy.
[0259] Judge the discovery rate of risk fault chains, which is the proportion of fault chains regarded as risky in the set of fault chains. For s ∈ [S], it is expressed by the following formula:
[0260]
[0261] Among them,
[0262] Precision(s) represents the discovery rate of risk fault chains,
[0263] Indicates the indicator function. A higher value of Precision(s) indicates higher accuracy,
[0264] M represents the threshold for risk assessment to determine whether a fault chain belongs to the "high - risk" category.
[0265] Compared with the prior art, the beneficial effects of the present invention at least include: The present invention proposes a method and system for identifying cascading faults in a new energy power system under extreme weather based on the combination of LLM and DRL. Through the policy network Iteratively select the optimal actions, capture the changes in the power grid in real time, dynamically construct a fault chain model, and can promptly identify the changes in the power system, improving the accuracy of the predicted path. This invention uses LLM to process and understand complex natural language information, extracting valuable prior knowledge such as historical fault cases, for example but not limited to, historical data, fault patterns, and environmental factors during the operation of the power system, enabling a better understanding of the system's dynamic changes and potential risks. Deep reinforcement learning usually requires a large number of interaction samples to achieve high performance, while the knowledge extraction ability of LLM can provide valuable guidance during the learning process, thus accelerating the learning process and reducing the dependence on environmental samples, helping DRL reduce sample complexity. This combination not only improves the learning efficiency of the model but also enhances its adaptability in the face of uncertainty and complexity. Through the combination of LLM and DRL, this invention can quickly identify potential cascading faults, providing new solutions for the safety management of the power system, realizing a more intelligent and efficient power network, and ensuring the reliability and security of power supply.
[0266] Embodiment 2 of the present invention provides a system for identifying cascading faults in a new energy power system under extreme weather conditions, which operates the method for identifying cascading faults in a new energy power system under extreme weather conditions described in Embodiment 1, including:
[0267] A fault chain construction module for establishing a fault chain model of the new energy power system under extreme weather conditions;
[0268] A loss model construction module for establishing a search model for the fault chain with the maximum load loss according to the fault model of the new energy power system under extreme weather conditions;
[0269] A policy construction module for converting the search model for the fault chain with the maximum load loss into a POMDP and constructing an optimal action selection policy for the POMDP;
[0270] A fault chain search model construction module that integrates the deep reinforcement learning algorithm of the large language model to construct an LLM-DRL fault chain search model;
[0271] A fault chain solving module for solving the optimal action selection policy of the POMDP through the ∈-greedy policy, generating a set of fault chains, and realizing the identification of cascading faults in the new energy power system under extreme weather conditions,
[0272] Among them, with a probability of ∈, the power system data under the actual extreme weather conditions are solved for the fault chain using the weighted power flow method, and with a probability of 1 - ∈, the power system data under the actual extreme weather conditions are input into the LLM-DRL fault chain search model for solving the fault chain.
[0273] Embodiment 3 of the present invention provides a power system for testing under extreme weather conditions. The test system consists of 39 buses and 46 power equipment, including 12 transformers and 34 lines. Considering the load condition of 0.55 × base load, where the "base load" here represents the standard load data of the New England test case after power generation - load balancing in PYPOWER, to quantify the performance of this method.
[0274] Given that M is 5% of the total load (where the total load is 0.55 × base load). To quantify Regret(s), it is necessary to calculate the total load loss associated with the S most critical fault chains (true situation). Using the pre - calculated set Under the load condition of 0.55 × base load, using the fault chain searcher of the present invention, it is observed that there are a total of 3738 risky fault chains.
[0275] Test the performance of the fault chain searcher for different numbers of agents Num in DRL. The test results are shown in Table 1. It can be concluded that a larger Num will generate fault chains with a greater total load loss and will find more risky fault chains. This indicates that as Num increases, the accuracy metric is improved. This is because the more agents there are, the more accurately the fault chain searcher can predict the Q - value associated with each partially observable Markov decision process (POMDP) state S i For example, when Num = 3, Regret(s)=765.33×10 3 MW, which is 3.04% lower than Regret(s) when Num = 1. Similarly, the Precision(s) value when Num = 3 is 0.169, which is 26% higher than Precision(s) when Num = 1. To further evaluate how the Precision(s) metric varies with s ∈ [S], Figure 4 and Figure 5 shows the relationship between Precision(s) and s ∈ [S]. The results are consistent with Table 1, where it can be concluded that a higher Num will result in a lower Regret(s) and a higher Precision(s). It should be noted that the advantage of setting Num = 3 comes at the cost of a higher computational cost.
[0276] Table 1 Comparison of the performance of the fault chain searcher
[0277]
[0278]
[0279] To further illustrate the advantages of the proposed fault chain searcher, the present invention compares it with two other comparative methods. Comparative method 1 combines a power flow weighting strategy and a Q-learning algorithm. Specifically, the power flow weighting strategy guides the exploration process of the agent by considering the power flows of various components in the power system, enabling it to preferentially select those components carrying higher power flows for the search of fault chains. The agent uses the Q-learning algorithm to learn and optimize its decision-making strategy during this process to identify potential risk fault chains. The advantage of this method is that it can effectively utilize historical data to improve the accuracy of fault prediction. Comparative method 2 further introduces a transfer expansion (TE) mechanism on the basis of comparative method 1. This mechanism enables the agent to adapt to new system states more quickly in real-time implementation by leveraging the knowledge obtained previously under different load conditions (e.g., the Q-table obtained through offline training). The advantage of this method is that it can quickly adjust the strategy under new load conditions, thereby improving the efficiency and accuracy of fault chain prediction.
[0280] Table 1 also shows the fault chain search performance of comparative methods 1 and 2. It can be seen that the proposed method is consistently significantly better than comparative methods 1 and 2. For example, when Num = 3, the average Regret(s) generated by the proposed method is 6.2% less than that of comparative method 2. In addition, when Num = 3, the load loss of the fault chains searched by the proposed method is almost twice that of comparative methods 1 and 2.
[0281] In Figure 5 comparison, comparative method 2 is superior to the proposed method in the first fifty search iterations. However, as S increases, the proposed method can learn the fault dynamics more accurately, resulting in a better Precision(s) than the comparative methods. For example, when Num = 3, the Precision(s) of the proposed method is approximately 86% higher than that of comparative method 1 and 99% higher than that of comparative method 2.
[0282] The computational complexity for fault case search must be managed within each scheduling period. Therefore, it is necessary to strictly limit the evaluation of all methods within five minutes. Table 2 shows the relative performance of various methods under a limited budget, and three main results are observed. First, within the 5-minute computational time budget, for the case of Num = 3, the number of fault sequences found by the proposed method is significantly less than that of other methods, as shown in the second column of Table 2. This result is expected because a larger Num value requires more computational time for gradient updates in each fault search iteration, resulting in a reduced number of iterations available for search. Second, when comparing the evaluation metrics, the proposed method finds the largest average total load loss and the largest number of high-risk fault cases under the setting of Num = 2. Although Comparative Method 1 and Comparative Method 2 find the largest number of fault sequences respectively, they are inferior to the proposed method in terms of the quality of fault cases under the setting of Num = 2 because the cumulative total load loss and the number of high-risk fault cases of these two methods are less. Third, in terms of the comparison of accuracy metrics, the average Regret(s) and Precision(s) of the proposed method under the setting of Num = 3 are better than those of other algorithms. Although this method only finds 575 fault cases on average, it generates the highest-quality fault sequences.
[0283] Table 2 Comparison of Computational Complexity for Fault Case Search
[0284]
[0285] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.
Claims
1. A method for identifying cascading faults in a new energy power system under extreme weather, characterized in that: Establish a fault chain model for the new energy power system under extreme weather; According to the fault model of the new energy power system under extreme weather, establish a search model for the fault chain of the maximum load loss; Convert the search model for the fault chain of the maximum load loss into a POMDP, and construct an optimal action selection strategy for the POMDP; Integrate the deep reinforcement learning algorithm of the large language model to construct an LLM-DRL fault chain search model; Solve the optimal action selection strategy of the POMDP through the ∈-greedy strategy, generate a set of fault chains, and realize the identification of cascading faults in the new energy power system under extreme weather. Among them, with a probability of ∈, the power system data under actual extreme weather is solved for the fault chain using the weighted power flow method, and with a probability of 1 - ∈, the power system data under actual extreme weather is input into the LLM-DRL fault chain search model for fault chain solution.
2. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 1, characterized in that: The search model for the fault chain of the maximum load loss is expressed by the following formula: In the formula, Set of S fault chains representing the maximum load loss The fault chain representing the s-th largest load loss S represents the number of fault chains of the maximum load loss, Denote the set of all fault chains, Denote the s-th fault chain The total load loss caused 3. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 2, characterized in that: The total load loss caused by the sth fault chain is expressed by the following formula: In the formula, Indicates Total load loss caused A fault chain representing a sequence of power equipment that fails during P phases Indicates the load loss caused by the fault in the i-th stage, Indicates the set of power equipment that fails during the process, where i ∈ [P] and P represents the total number of stages represents the topological structure of the power system in the i-th stage, W xeternal An influence parameter representing extreme meteorological conditions.
4. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 3, characterized in that: The construction of the optimal action selection strategy for the POMDP is expressed by the following formula: In the formula, π * represents the optimal action selection strategy represents the expected value operator, V π (S0) represents the total feedback obtained by the agent starting from the initial POMDP state S0 for any action selection policy π.
5. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 1, characterized in that: The integration of the deep reinforcement learning algorithm of the large language model to construct the LLM-DRL fault chain search model specifically includes: Construct a primary fault chain searcher based on the large language model; Based on the primary fault chain searcher of the large language model and the deep reinforcement learning model, construct an LLM-DRL fault chain searcher.
6. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 5, characterized in that: The construction of the primary fault chain searcher based on the large language model specifically includes: Set the task description; The large language model generates a description of the task according to the set task description; The large language model generates the code of the primary fault chain searcher according to the set programming guide and in combination with the generated task description; Perform actual fault chain search operations according to the code of the primary fault chain searcher, and optimize the rules of the code of the primary fault chain searcher using the artificial feedback mechanism to obtain the primary fault chain searcher based on the large language model.
7. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 1, characterized in that: The input of the power system data under actual extreme weather into the LLM-DRL fault chain search model for fault chain solution with a probability of 1 - ∈ specifically includes: The primary fault chain searcher based on the large language model generates initial fault chain action data according to the natural language description data in the power system data under the actual extreme meteorological environment, where the initial fault chain action data is stored in the large language model buffer; After the initial fault chain action data is subjected to imitation learning, it is input into the deep reinforcement learning algorithm to solve the current state, the action a taken by the agent with a probability of 1 - ∈, i,1-∈ the reward discount factor, and the next state, and deep reinforcement learning interaction data is obtained, where the deep reinforcement learning interaction data is stored in the deep reinforcement learning buffer; Sample data from the large language model buffer and the deep reinforcement learning buffer, update the evaluation network parameters and action network parameters of the deep reinforcement learning for the evaluation network and the action network respectively, input the objective function of the LLM-DRL fault chain searcher for policy gradient calculation, and obtain the current optimal action according to the optimal action selection policy; Repeat the iteration until the training converges, and generate the fault chain according to all the optimal actions.
8. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 7, characterized in that: The action \(a\) taken by the agent with probability \(1-\epsilon\) i,1-∈ , which is expressed by the following formula: In the formula, Represents the Q-value of the agent based on the current state and action, Y i represents the input feature at stage i, θ represents the parameter for learning the Q-value function, Indicates the number of times an electrical device is selected when the agent is in the POMDP state S i at that time is selected.
9. The method for identifying cascading faults in a new energy power system under extreme weather according to claim 8, characterized in that: The objective function of the LLM-DRL fault chain searcher is expressed by the following formula: Among them, A model representing the interaction between an agent and the environment Indicates the expected value, s0~p0 means that the initial state s0 is sampled according to the initial state distribution p0, s t+1 ~p(·|s t ,a t ) represents the state s at time step t t After the agent takes the action a at time step t t , according to the state transition probability p(·|s t ,a t ), it transfers to the next state s t+1 ; a t ~π(·∣s t ) represents the state s of the agent at time step t t under which, according to the policy π(·∣s t ) selects the action a at time step t t , γ t ∈ [0, 1] represents the reward discount factor; Represents a sparse reward function.
10. A system for identifying cascading faults in a new energy power system under extreme weather, which runs the method for identifying cascading faults in a new energy power system under extreme weather according to any one of claims 1-9, characterized in that: A fault chain construction module for establishing a fault chain model of a new energy power system under extreme weather; A loss model construction module for establishing a maximum load loss fault chain search model according to the fault model of a new energy power system under extreme weather; A policy construction module for converting the maximum load loss fault chain search model into a POMDP and constructing an optimal action selection policy for the POMDP; A fault chain search model construction module for constructing an LLM-DRL fault chain search model by integrating the deep reinforcement learning algorithm of the large language model; A fault chain solving module for solving the optimal action selection policy of the POMDP through the ∈-greedy policy, generating a fault chain set, and realizing the identification of cascading faults in a new energy power system under extreme weather, Among them, with a probability of ∈, the power system data under the actual extreme meteorological environment is solved for the fault chain by the weighted power flow method, and with a probability of 1-∈, the power system data under the actual extreme meteorological environment is input into the LLM-DRL fault chain search model for the fault chain solution.
Citation Information
Patent Citations
Markov Process cascading failure path generation method and device starting from meteorological disasters
CN117374915A