Overload line emergency control strategy continuous learning method suitable for changing scene

By designing the large strategy model Π and the policy network model π, combining reinforcement learning and graph deep learning methods, the problem that emergency control strategies in the existing technology are difficult to retain the value of the old strategy in changing scenarios, and the continuous learning and efficient update of the overload line emergency control strategies is achieved, and emergency control strategies are suitable for changing scenarios.

CN120545987AActive Publication Date: 2025-08-26SOUTH CHINA UNIV OF TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510693487.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-26
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The existing emergency control strategy learning methods are difficult to retain the value of the old strategy after the operation mode changes, resulting in wasted training costs and difficult to apply to overload line emergency control in changing scenarios.

Method used

Design a large strategy model Π and a policy network model π, learn and update emergency control strategies in each time period through a continuous learning process, use the Transformer graph neural network to build a model, combine reinforcement learning and graph deep learning methods, filter and update samples of overload operation methods that cannot be processed, form a continuous learning sample set, and realize the update and accumulation of the strategy model.

Benefits of technology

It realizes continuous learning of overload line emergency control strategies in changing scenarios, and can be applied to emergency control strategies in various time periods, reducing training costs, retaining the value of the old strategy, and improving learning efficiency and applicability of the strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120545987A_ABST
    Figure CN120545987A_ABST
Patent Text Reader

Abstract

The invention provides an overload line emergency control strategy continuous learning method suitable for a changing scene. Firstly, a large strategy model and a strategy network model are designed. Then, aiming at an overload operation mode in a continuous time period, a large strategy model and an initial strategy network in a current time period are learned based on a large strategy model and an initial strategy network in a previous time period, and the specific process comprises the step of searching the overload operation mode in the time period; a large strategy model is adopted to implement a decision process, and overload operation modes which cannot be processed by the large strategy model are screened out; an overload operation mode which cannot be processed by the large strategy model is learned through a reinforcement learning method, and a strategy network is obtained; the overload operation mode and the emergency control scheme processed by the strategy network are used as newly added samples of the large strategy model, and the large strategy model is further learned. According to the invention, continuous learning of the overload line emergency control strategy in a changing scene can be realized, and the emergency control strategy suitable for each time period can be learned.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of power system emergency control, and in particular relates to a continuous learning method for an overload line emergency control strategy applicable to changing scenarios. Background Art

[0002] In the field of power system emergency control, an emergency control strategy must be generated in advance for each time period's operating mode and its potential risk scenarios to handle overloaded operating modes within that time period. With the development of artificial intelligence (AI) technology, reinforcement learning has been gradually applied to the generation of emergency control strategies. It can learn emergency control strategies based on the potential risky operating modes pre-searched for that time period.

[0003] However, the emergency control strategies generated by reinforcement learning methods for a specific time period have limited applicability. When the operating mode changes, the emergency control strategy becomes inapplicable, and a new emergency control strategy for that time period needs to be learned. Most current research follows this principle. However, each time a new control strategy is learned for the changed operating mode and its potential risk scenarios, the value of the old strategy is discarded, and its training costs are wasted. Clearly, existing emergency control strategy learning methods struggle to retain the value of old strategies, resulting in wasted training costs. There is an urgent need for a continuous learning method for overload line emergency control strategies that can continuously accumulate overload line emergency control experience. Summary of the Invention

[0004] In order to solve at least one of the problems existing in the prior art, the present invention provides a continuous learning method for an overload line emergency control strategy applicable to changing scenarios.

[0005] To achieve the purpose of the present invention, the present invention provides a method for continuous learning of an overload line emergency control strategy applicable to changing scenarios, the method comprising the following steps:

[0006] S1: Design a grand strategy model Π and a strategy network model π, and conduct a continuous learning process of emergency control strategies from S2 to S5 based on the grand strategy model Π and the strategy network model π;

[0007] S2: Define the symbol representing the continuous learning period as T, the number of time periods as n, and the sample set of the grand strategy model Π learning as D s ;

[0008] S3: Initialize the continuous learning time period index i=1, and initialize the initial large strategy model Π0 and the initial strategy network π0;

[0009] S4: To T i The overload operation mode under the time period is used to learn the emergency control strategy, based on the large strategy model Π i-1and the initial policy network π i-1 Learn the grand strategy model π i and the initial policy network π i , and output the latest sample set D' of the grand strategy model learning s ;

[0010] S5: Determine whether i is greater than n? If so, then end and output the grand strategy model Π after continuous learning i With the policy network π i ; If not, then i=i+1, D s =D' s , and jump to S4 to continue strategy learning for the next time period.

[0011] The large strategy model π and the strategy network model π are both constructed based on neural networks and are used to characterize high-dimensional emergency control schemes.

[0012] Furthermore, the structure of the neural network used to represent the grand strategy model π and the strategy network model π is a Transformer-based graph neural network. Both the grand strategy model π and the strategy network model π include a feature extraction module and a fitting module. The feature extraction module includes alternating Transformer convolution layers and batch normalization layers. The Transformer convolution layer is used to extract graph features, and the fitting module includes a fully connected network.

[0013] Furthermore, the inputs of the grand strategy model Π and the strategy network model π are both the operating state s = {X, A}, where X is the grid node feature, V is the node voltage, is the node phase, P is the node injected active power, Q is the node injected reactive power, and A represents the adjacency matrix of the power grid topology.

[0014] Furthermore, the outputs of the large strategy model π and the strategy network model π are both emergency control measures vector a for regulating line power, including the generator power regulation measure a G , load shedding measures a L and cutting machine measures a GC , a G 、a L and a GC All of them are between [0,1], representing the percentage of each measure’s regulation amount. Assuming that there are N generator nodes M There are N load nodes. D , then:

[0015]

[0016] Where, For N MThe adjustment amount of the power regulation measures of each generator node, For N D The load shedding measures for each load node are as follows: For N M The amount of cutting off of each generator node cutting measure.

[0017] Furthermore, the steps of the emergency control strategy learning method include:

[0018] S4-1: Search for T i The overload operation mode under the time period, the set of overload operation modes is recorded as D i ;

[0019] S4-2: Determine whether i is equal to 1? If so, then D' i =D i , and jump to S4-4; if not, jump directly to S4-3;

[0020] S4-3: Adopting a Grand Strategy Model i-1 For the set of overload operation modes D i The overload operation mode in the implementation of the decision process, obtain the decision effect, and screen out the large strategy model Π i-1 The scenes that cannot be handled are recorded as set D' i ;

[0021] S4-4: Using reinforcement learning method to analyze the set D' i The overload operation mode in the learning is used to learn the strategy network model π i ;

[0022] S4-5: Adopting a strategic network model π i For set D' i The decision process is implemented in the overload operation mode to form a mapping pair {s k ,a k}, where s k In the overload operation mode, a k is the emergency control scheme, k is the set D' i Run mode index; put all mappings to {s k ,a k}Store in the grand strategy model sample set D s In the example, we form the sample set D' s ;

[0023] S4-6: For sample set D' s In the samples, the graph deep learning method is used to train the large strategy model Π i-1 Conduct deep learning to obtain a large strategy model π i .

[0024] Furthermore, in step S4-1, the search T i The overload operation mode under the time period is as follows: i The basic operation mode under the time period generates a large number of operation scenarios after considering various uncertain factors.

[0025] Furthermore, the specific method of implementing the decision process is as follows: first, extract the state of the overload operation mode, including the node voltage V, node phase The node injects active power P and reactive power Q, and extracts the adjacency matrix of the current power grid to form the state of the operation mode s = {X, A}; then, the large strategy model Π i-1 The state s is input to obtain the emergency control measure vector a for regulating line power. Finally, the grid's operating mode, including generator node power and load power, is modified based on this emergency control measure vector a. The power flow calculation and time-domain simulation are then re-performed for this operating mode to determine the decision effect, which indicates whether the grid still experiences line overload or node voltage exceeding the limit after the decision is made.

[0026] Furthermore, the convergence condition of the reinforcement learning method is set to the policy network model π i Can process set D' i All overload operating modes in the operation mode are controlled to within the limit.

[0027] Furthermore, the graph deep learning method is to convert the previous time period T i-1 Grand Strategy Model Π i-1 As the basic model, the mean square error function is used as the loss function, and the new grand strategy model Π is continuously learned through multiple iterations. i When the number of iterations is reached, the model with the highest success rate in each iteration is selected as the grand strategy model Π i .

[0028] Compared with the prior art, the present invention can at least achieve the following beneficial effects:

[0029] The present invention provides a method for continuous learning of emergency control strategies for overloaded lines applicable to changing scenarios. A new continuous learning process is designed to solve the problem that the current emergency control strategy learning method is difficult to retain the value of old strategies and the training cost is wasted. For the overloaded operating mode under a continuous period, the large strategy model and strategy network model for the current time period are learned based on the large strategy model and strategy network in the previous time period. The specific process includes searching for the overloaded operating mode in the time period; implementing the decision-making process using the large strategy model to screen out the overloaded operating modes that the large strategy model cannot handle; learning the overloaded operating modes that the large strategy model cannot handle through the reinforcement learning method to obtain the strategy network; using the overloaded operating modes and emergency control schemes handled by the strategy network as new samples of the large strategy model to further learn the large strategy model. The present invention can realize continuous learning of emergency control strategies for overloaded lines under changing scenarios and learn emergency control strategies applicable to each time period. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a flowchart of the steps of a continuous learning method for an overload line emergency control strategy applicable to changing scenarios provided by an embodiment of the present invention.

[0031] Figure 2 Schematic diagram of the structure of the large strategy model Π or the strategy network model π in an embodiment of the present invention (the two have the same structure).

[0032] Figure 3 It is a flowchart of emergency control strategy learning in an embodiment of the present invention.

[0033] Figure 4 This is a topology diagram of the IEEE-39 node system in an embodiment of the present invention. DETAILED DESCRIPTION

[0034] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, the results obtained by ordinary technicians in this field are without creative work.

[0035] See also Figure 1 , an embodiment of the present invention provides a continuous learning method for an overload line emergency control strategy applicable to a changing scenario, comprising the following steps:

[0036] S1: Build a grand strategy model Π and a strategy network model π, and then perform the emergency control strategy continuous learning process S2 to S5 based on the grand strategy model Π and the strategy network model π.

[0037] To ensure that the emergency control strategy has a certain degree of generalization performance and can represent high-dimensional spaces, a neural network is used to represent the grand strategy model π and the strategy network model π. Since the grand strategy model π and the strategy network model π have similar functions, both calculating emergency control solutions based on the operating status of the power grid, the two models are defined as having the same structure.

[0038] For the large strategy model Π and the strategy network model π, the input of both models is defined as the state of the operation mode s = {X, A}, where X is the grid node characteristics, considering the node voltage V, node phase The node injects active power P and the node injects reactive power Q, that is, The voltage is expressed in per-unit values, and the phase is expressed in radians. A represents the adjacency matrix of the power grid topology. The adjacency matrix A changes as the lines or main transformers are connected. The elements at the corresponding positions in A are set to 1 only when the lines and main transformers are connected, and are set to 0 otherwise.

[0039] For the large strategy model Π and the strategy network model π, the outputs of both models are the emergency control measures vector for regulating line power, represented by symbol a, and the expression is shown in formula (1). The emergency control measures vector a for regulating line power includes the generator power measures a G , load shedding measures a L and cutting machine measures a GC , where a G 、a L and a GC are between [0,1], representing the percentage of each measure’s regulation. Assuming that there are N generator nodes M There are N load nodes. D , then For N M The adjustment amount of the power regulation measures of each generator node, For N D The load shedding measures for each load node are as follows: For N M The amount of cutting off of each generator node cutting measure.

[0040]

[0041] In one embodiment of the present invention, for the network structure of the large strategy model Π and the strategy network model π, considering that the power grid operation state is strongly related to the topology, the network structures of both use a Transformer-based graph neural network with strong power grid topology feature extraction capabilities. The Transformer-based graph neural network includes a feature extraction module and a fitting module, and the activation function adopts the Relu function. The feature extraction module includes a Transformer Convolution (TC) layer and a batch normalization (BN) layer. The fitting module includes a fully connected network (FCN), wherein the TC layer is used to extract graph features, the BN layer is used to adjust the distribution of data to make it more suitable for training, and the FCN layer is used to fit data. There are two fully connected networks, such as Figure 2 As shown, it should be noted that the number of each layer needs to be determined specifically according to the business situation. The current mainstream method for determining the number of each layer in the power field is mainly the trial method, that is, trying different configurations and selecting the number of layers with better results. In one embodiment of the present invention, the TC layer is set to three layers, the BN layer is set to four layers, and the BN layer and the TC layer are set alternately. In other embodiments, other numbers can be set. The input of the network model is the state s = {X, A}. After passing through the feature extraction module and the fitting module, the emergency control measure vector a for regulating the line power is output. Among them, the graph data after feature extraction is expanded into a one-dimensional vector as the input of FCN1 and FCN2 to fit the emergency control measure vector a for regulating the line power. Among them, FCN1 fits the continuous variable a G and a L , FCN2 fits discrete variable a GC , and finally synthesized into a, where a G For the purpose of regulating the generator power measure a G The vector composed of a L Load shedding measures a L The vector composed of a GC Cutting machine measures a GC The vector composed of .

[0042] S2: Define the symbol representing the continuous learning period as T, the number of time periods as n, and the sample set of the grand strategy model Π learning as D s .

[0043] S3: Initialize the continuous learning time period index i=1, and initialize the initial large strategy model Π0 and the initial strategy network model π0.

[0044] S4: To T iThe overload operation mode under the time period is used to learn the emergency control strategy, based on the large strategy model Π i-1 and policy network model π i-1 Learn the grand strategy model π i and policy network model π i , and output the latest sample set D' of the grand strategy model learning s .

[0045] S5: Determine whether i is greater than n? If so, then end and output the grand strategy model Π after continuous learning i With the policy network model π i ; If not, let i=i+1, D s =D' s , and jump to S4 to continue strategy learning for the next time period.

[0046] The specific process of emergency control strategy learning in step S4 is shown in Figure 3 The process can be summarized as follows:

[0047] S4-1: Search for T i The overload operation mode under the time period, the set of overload operation modes is recorded as D i .

[0048] S4-2: Determine whether i is equal to 1? If so, then D' i =D i , and jump to S4-4; if not, jump directly to S4-3.

[0049] S4-3: Adopting a Grand Strategy Model i-1 For the set of overload operation modes D i The overload operation mode in the implementation of the decision-making process to obtain the decision effect. Screen out the large strategy model Π i-1 The scenes that cannot be handled are recorded as set D' i .

[0050] S4-4: Using reinforcement learning method to analyze the set D' i The overload operation mode in the learning is used to learn the strategy network model π i .

[0051] S4-5: Adopting a strategic network model π i For set D' i The decision process is implemented in the overload operation mode to form a mapping pair {s k ,a k}, where s k In the overload operation mode, a k is the emergency control scheme, k is the set D' i The operation mode index in . Put all mapping pairs {sk ,a k}Store the sample set D for learning the grand strategy model s In the example, we form the sample set D' s .

[0052] S4-6: For sample set D' s In the samples, the graph deep learning method is used to train the large strategy model Π i-1 Learning, the grand strategy model after learning is recorded as Π i .

[0053] The following describes some of the above steps in detail.

[0054] The function of step S4-1 is to find T i Overload operation mode under time period. First, for T i The basic operation mode under the time period is generated by taking into account the uncertainty factors such as line disconnection, generator node disconnection and source load fluctuation. Then, the power flow calculation and time domain simulation are performed on the generated operation modes to calculate the line current of the entire network under each operation mode, and the current I j The current limit of the line I j,limit Compare and judge whether it is overloaded. If it satisfies formula (2), it is considered overloaded, and j is the jth line. Finally, all overload scenarios are screened out to form a set of overload operation modes D i .

[0055] I j >I j,limit (2)

[0056] The function of step S4-3 is to use the big strategy model to transform the set of overload operation modes D i The process of operating modes that can be processed is filtered out, and reinforcement learning is performed on other scenarios that cannot be processed, reducing the number of operating modes to be learned and improving learning efficiency.

[0057] In one embodiment of the present invention, a grand strategy model Π is used i-1 For the set of overload operation modes D i The specific decision-making method for the overload operation mode in the process is as follows: First, the state of the overload operation mode is extracted, including the node voltage V, node phase The node injects active power P and reactive power Q, and extracts the adjacency matrix of the current power grid to form the state of the operation mode s = {X, A}. Then, the large strategy model Π i-1The state s is input to obtain the emergency control measure vector a for regulating line power. Finally, the grid's operating mode, including generator node power and load power, is modified based on this emergency control measure vector a. The power flow calculation and time-domain simulation are then re-performed for this operating mode to determine the decision effect. This result indicates whether the grid still experiences line overload or node voltage overlimit after the decision is made.

[0058] In the above steps, the operation mode of the power grid is modified according to the emergency control measure vector a for adjusting the line power, which varies depending on the control object. The emergency control measure vector a for adjusting the line power includes three measures, and the generator power adjustment measure corresponds to a G The power regulation expression of the mth generator node under this measure is shown in formula (3), that is, according to the maximum adjustable power P of the generator node m m,max with a G,m The product of is used to adjust; the load shedding measure corresponds to a L , the power regulation expression of the dth load node under this measure is shown in equations (4) and (5); the corresponding machine-cutting measure is a GC , the power regulation expression of the mth generator node under this measure is shown in formula (6), that is, when a GC When it is greater than or equal to 0.5, P G,m Set to 0, otherwise leave unchanged.

[0059] P G,m =a G,m ·P m,max (3)

[0060] P′ L,d =a L,d ·(P L,d -P L,d,min ) (4)

[0061] Q′ L,d =a L,d ·(Q L,d -Q L,d,min ) (5)

[0062]

[0063] Among them, P G,m represents the power of the mth generator node; a G,m represents the regulation amount of the power regulation measure of the mth generator node; a L,d represents the load shedding amount of the d-th load node; P L,d represents the active power of the dth load node; P L,d,min represents the minimum active power of the dth load node; P′ L,drepresents the active power of the dth load node after load shedding; Q L,d The reactive power of the dth load node; Q L,d,min Represents the minimum reactive power of the dth load node, Q′ L,d Represents the reactive power after the load of the dth load node is removed.

[0064] In one embodiment of the present invention, in S4-4, the reinforcement learning method used is a common reinforcement learning method. i The overload operation mode in the strategy network model π is learned through reinforcement learning method i , it should be noted that the convergence condition of the reinforcement learning method is set to the policy network model π i Processable set D' i All overload operating modes in the operation mode are controlled to within the limit.

[0065] In one embodiment of the present invention, in S4-5, the policy network model π is used. i For set D' i The specific decision-making method for the overload operation mode in the process is as follows: First, the state of the overload operation mode is extracted, including the node voltage V, node phase The node injects active power P and reactive power Q, and extracts the adjacency matrix of the current power grid to form the state of the operation mode s = {X, A}. Then, the large strategy model Π i-1 Input state s and obtain the emergency control measure vector a for regulating line power. Finally, the operation mode of the power grid is modified according to the emergency control measure vector a, including the generator node power and load power, and the power flow calculation and time domain simulation of the operation mode are re-performed to obtain the decision effect. The decision effect is whether the power grid still has line overload or node voltage exceeding the limit after the decision is made. After the decision is completed, each operation mode state s is formed. k The corresponding control scheme a k , that is, the mapping pair {s k ,a k}, store all mapping pairs into the large strategy model sample set D s , forming D' s .

[0066] In one embodiment of the present invention, in S4-6, the graph deep learning method used is a common deep learning method. i-1 Grand Strategy Model Π i-1 As the basic model, the mean square error function is used as the loss function, and the new grand strategy model Π is continuously learned through multiple iterations. iIt should be noted that when the number of iterations is reached, the model with the highest success rate in each iteration is selected as the grand strategy model Π i .

[0067] In one embodiment of the present invention, the above continuous learning method is verified in an IEEE-39 node system.

[0068] like Figure 4 The following is a topology diagram of a 39-node system. Given four time periods, T1, T2, T3, and T4, after step S4-1, search for T i After analyzing the overload operation modes under different time periods, we obtain the overload operation mode sets D1, D2, D3 and D4 under the four time periods T1, T2, T3 and T4. The number of overload operation modes in these four sets is shown in Table 1. It can be seen that there are 20 overload scenarios in set D1, of which 10 are line overload operation modes after line disconnection, 6 are line overload operation modes after generator node disconnection, and 4 are line overload operation modes after source load fluctuation; there are 21 overload scenarios in D2, of which 10 are line overload operation modes after line disconnection, 6 are line overload operation modes after generator node disconnection, and 4 are line overload operation modes after source load fluctuation. There are 5 overload scenarios in D3; there are 22 overload scenarios in D3, of which 9 are line overloads after line disconnection, 6 are line overloads after generator node disconnection, and 7 are line overloads after source load fluctuations; there are 18 overload scenarios in D4, of which 9 are line overloads after line disconnection, 6 are line overloads after generator node disconnection, and 3 are line overloads after source load fluctuations.

[0069] Table 1. Overload scenario information in the 39-node system continuous learning case

[0070]

[0071] The continuous learning method of overload line emergency control strategy is used to continuously learn the sets D1, D2, D3 and D4 in turn, and the large strategy models learned at T1, T2, T3 and T4 are defined as Π1, Π2, Π3, Π4 and the strategy network models π1, π2, π3, π4 respectively.

[0072] Test the grand strategy model Π separately i and policy network model π i In T i Time period for D i The learning effect of , where i is 1, 2, 3 and 4. The test steps are as follows:

[0073] 1. First, through step S4-3, use T i-1 The grand strategy model Π trained before the time period i-1 To deal with T i Overload operation mode set D for the time period i , find out the operation mode of handling failure, recorded as D' i , and set the indicator "Use Π i-1 Processing D i The success rate of the operation mode in the macro strategy model is used to evaluate the T i Time period D i The handling capacity of medium overload operation mode.

[0074] 2. Then, for the set D' i Perform reinforcement learning in step S4-4 in the middle scenario to learn the processable set D' i Strategy network model for medium overload operation i , test the grand strategy model Π i-1 and policy network model π i Can we jointly process set D? i All overload operation modes are set in the indicator "Use strategy network model π i Processing D i The number of successful operations in the middle mode and the number of successful operations in the middle mode using the grand strategy model Π i-1 and policy network model π i Processing D i The effectiveness of the combined treatment was evaluated by the number of successful runs.

[0075] 3. Next, add the processing set D' to the large strategy model sample i New sample D for medium overload operation s , forming D' s , and through steps S4-6 the grand strategy model Π i-1 Perform deep learning to obtain a large strategy model π i , test the grand strategy model Π i In set D i performance, and set the indicator "Use the Grand Strategy Model Π i Can handle D i The effect of grand strategy model learning is evaluated by the number of successful operation modes, which verifies the effectiveness of grand strategy model learning.

[0076] 4. Finally, use the grand strategy model π4 to process the set D i The operation mode in the operation mode is used to count the number of successful processing of the operation mode, and the indicator "Using the large strategy model π4 to process D i The effectiveness of grand strategy model learning is verified by the number of successful runs.

[0077] The indicator results are shown in Table 2. It can be seen that:

[0078] 1) For the 20 overload operating modes in set D1, the learned policy network model π1 can handle all operating modes in set D1, and the large policy model Π1 can also handle 19 scenarios.

[0079] 2) For the 21 overload operation modes in set D2, the large strategy model π1 can handle 19 overload operation modes, so D'2 contains 2 overload operation modes. For the set D2 containing 2 overload operation modes, the learned strategy network model π2 can handle O 2,L In the two scenarios, the grand policy model π1 and the policy network model π2 can jointly handle all overload operating modes in set D2. By adding the decision samples of the two overload operating modes in D'2 to the training samples of the grand policy model π, the grand policy model π1 is further learned to obtain π2, which can handle all overload operating modes in set D2.

[0080] 3) For the 22 overload scenarios in set D3, the grand policy model π2 can handle 20 of them, resulting in two overload scenarios in D'3. After reinforcement learning on the two overload scenarios that failed to be handled in D'3, the grand policy model π2 and the learned policy network model π3 are jointly capable of handling all overload scenarios in set D3. The decision samples for the two scenarios in D'3 are added to the training samples of the grand policy model π. Grand policy model π2 is further trained to obtain the grand policy model π3, which can handle all scenarios in set D3.

[0081] 4) For the 18 overload scenarios in set D4, the grand policy model π3 can handle 17 of them, so D'4 contains one overload scenario. After reinforcement learning on the one scenario in D'4 that cannot be handled, the grand policy model π3 and the learned policy network model π4 can jointly handle all overload scenarios in set D4. The decision sample of the one scenario in D'4 is added to the training samples of the grand policy model π. The grand policy model π3 is further trained to obtain the grand policy model π4, which can handle all scenarios in set D4.

[0082] 5) The final learned grand strategy model π4 is used to make decisions and control a total of 81 overload operation modes in sets D1, D2, D3, and D4, and all scenarios can be successfully controlled.

[0083] Table 2 Continuous learning effect of 39-node system

[0084]

[0085] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A continuous learning method for overload line emergency control strategies suitable for changing scenarios, characterized by: The method comprises the following steps: S1: Design a grand strategy model Π and a strategy network model π, and conduct a continuous learning process of emergency control strategies from S2 to S5 based on the grand strategy model Π and the strategy network model π; S2: Define the symbol representing the continuous learning period as T, the number of time periods as n, and the sample set of the grand strategy model Π learning as D s ; S3: Initialize the continuous learning time period index i=1, and initialize the initial large strategy model Π0 and the initial strategy network π0; S4: To T i The overload operation mode under the time period is used to learn the emergency control strategy, based on the large strategy model Π i-1 and policy network model π i-1 Learn the grand strategy model π i and the initial policy network model π i , and output the latest sample set D' of the grand strategy model learning s ; S5: Determine whether i is greater than n? If so, then end and output the grand strategy model Π after continuous learning i With the policy network π i ; If not, then i=i+1, D s =D' s , and jump to S4 to continue strategy learning for the next time period.

2. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to claim 1 is characterized in that: The large strategy model π and the strategy network model π are both constructed based on neural networks and are used to characterize high-dimensional emergency control schemes.

3. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to claim 2 is characterized in that: The structure of the neural network used to represent the grand strategy model π and the strategy network model π is a Transformer-based graph neural network. Both the grand strategy model π and the strategy network model π include a feature extraction module and a fitting module. The feature extraction module includes alternating Transformer convolution layers and batch normalization layers. The Transformer convolution layer is used to extract graph features, and the fitting module includes a fully connected network.

4. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to claim 2 is characterized in that: The inputs of the grand strategy model Π and the strategy network model π are both the state of the operation mode s = {X, A}, where X is the grid node feature, V is the node voltage, is the node phase, P is the node injected active power, Q is the node injected reactive power, and A represents the adjacency matrix of the power grid topology.

5. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to any one of claims 2 to 4, characterized in that: The outputs of the large strategy model π and the strategy network model π are both emergency control measures vector a for regulating line power, including the generator power regulation measure a G , load shedding measures a L and cutting machine measures a GC , a G 、a L and a GC All of them are between [0,1], representing the percentage of each measure’s regulation amount. Assuming that there are N generator nodes M There are N load nodes. D There are: Where, For N M The adjustment amount of the power regulation measures of each generator node, For N D The load shedding measures for each load node are as follows: For N M The amount of cutting off of each generator node cutting measure.

6. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to claim 1 is characterized in that: The steps of the emergency control strategy learning method include: S4-1: Search for T i The overload operation mode under the time period, the set of overload operation modes is recorded as D i ; S4-2: Determine whether i is equal to 1? If so, then D' i =D i , and jump to S4-4; if not, jump directly to S4-3; S4-3: Adopting a Grand Strategy Model i-1 For the set of overload operation modes D i The overload operation mode in the implementation of the decision process, obtain the decision effect, and screen out the large strategy model Π i-1 The scenes that cannot be handled are recorded as set D' i ; S4-4: Using reinforcement learning method to analyze the set D' i The overload operation mode in the learning is used to learn the strategy network model π i ; S4-5: Adopting a strategic network model π i For set D' i The decision process is implemented in the overload operation mode to form a mapping pair {s k ,a k }, where s k In the overload operation mode, a k is the emergency control scheme, k is the set D' i Run mode index; put all mappings to {s k ,a k }Store in the grand strategy model sample set D s In the example, we form the sample set D' s ; S4-6: For sample set D' s In the samples, the graph deep learning method is used to train the large strategy model Π i-1 Conduct deep learning to obtain a large strategy model π i .

7. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to claim 6 is characterized in that: In step S4-1, the search T i The overload operation mode under the time period is as follows: i The basic operation mode under the time period generates a large number of operation scenarios after considering various uncertain factors.

8. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to claim 6 is characterized in that: The specific method of implementing the decision process is: first, extract the state of the overload operation mode, including the node voltage V, node phase The node injects active power P and reactive power Q, and extracts the adjacency matrix of the current power grid to form the state of the operation mode s = {X, A}; then, the large strategy model Π i-1 The state s is input to obtain the emergency control measure vector a for regulating line power. Finally, the grid's operating mode, including generator node power and load power, is modified based on this emergency control measure vector a. The power flow calculation and time-domain simulation are then re-performed for this operating mode to determine the decision effect, which indicates whether the grid still experiences line overload or node voltage exceeding the limit after the decision is made.

9. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to claim 6 is characterized in that: The convergence condition of the reinforcement learning method is set to the policy network model π i Can process set D' i All overload operating modes in the operation mode are controlled to within the limit.

10. The continuous learning method for overload line emergency control strategy applicable to changing scenarios according to any one of claims 6-9, characterized in that: The graph deep learning method is to convert the previous time period T i-1 Grand Strategy Model Π i-1 As the basic model, the mean square error function is used as the loss function, and the new grand strategy model Π is continuously learned through multiple iterations. i When the number of iterations is reached, the model with the highest success rate in each iteration is selected as the grand strategy model Π i .

Citation Information

Patent Citations

  • Deep reinforcement learning emergency control strategy extraction method for power system

    CN114004282A

  • Power grid low-voltage load shedding emergency control method based on graph deep reinforcement learning

    CN114865638A

  • Power grid rolling scheduling deep reinforcement learning decision-making method in extreme weather

    CN119627846A

  • Power grid risk disposal plan generation method and system based on reinforcement learning

    CN119990793A