Power grid safe operation boundary extraction method and device, storage medium, program product and computer equipment
By determining the unit combination model in the power grid and using reinforcement learning agents to generate data sets, the optimal strategy is extracted, and the problem of waste in grid operation complexity and limit setting under the conditions of new energy high penetration is solved, and the identification of optimal safe operation boundaries and the improvement of grid stability is achieved.
Patent Information
- Application Number
- CN202510149784.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-06-10
AI Technical Summary
Under the conditions of high penetration of new energy, the power grid operation mode is complex and changeable, and extreme scenarios are difficult to accurately identify, resulting in the conservative limit setting strategy causing the waste of transmission channel capabilities, and the determined safe operation boundary will lead to a large amount of stable operation space being wasted.
By determining the unit combination model, using pre-trained reinforcement learning agents, generating data sets based on the power system control model, performing policy extraction, and obtaining the optimal policy to indicate the optimal safe operation boundary.
On the premise of meeting the section limit constraints, the optimal safe operation boundary is efficiently and accurately identified, reducing the waste of operating space, and improving the stability and safety of the power grid.
Smart Images

Figure CN120127623A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of power systems, and in particular, to a method, device, storage medium, program product, and computer device for extracting the safe operation boundary of a power grid. Background Art
[0002] At present, the energy structure in the power grid continues to be optimized, and the installed capacity of clean energy has increased rapidly, driving the evolution of the power system towards a new type of power system. However, in the face of the challenges brought by load growth and large-scale access of new energy, the operating characteristics of the power grid have changed significantly. To ensure the safe and stable operation of the power grid, the dispatching and operation department usually selects the minimum value from multiple sets of transmission limit values calculated based on multiple typical scenarios as the transmission limit.
[0003] However, under the condition of high penetration of new energy, the system operation mode becomes more complex and changeable, and extreme scenarios are difficult to accurately identify. Although the conservative limit setting strategy can ensure system safety, it also causes serious waste of the transmission channel capacity, that is, the determined safe operation boundary will result in a large amount of stable operation space being wasted. Summary of the Invention
[0004] To solve the above technical problems, the embodiments of the present application propose a method, device, storage medium, program product, and computer device for extracting the safe operation boundary of a power grid, which can efficiently and accurately identify the optimal safe operation boundary and reduce the waste of operation space under the premise of meeting the section limit constraint.
[0005] In a first aspect, the embodiments of the present application provide a method for extracting the safe operation boundary of a power grid, including:
[0006] Determine a unit commitment model, where the objective function of the unit commitment model is suitable for characterizing the minimization of the target cost, the target cost includes the traditional unit power generation cost and the new energy abandonment cost, and the constraint conditions of the unit commitment model at least include a section limit constraint, and the section limit constraint is suitable for restricting the power flow value of the section of the power system between the negative direction limit value of the section transmission power and the positive direction limit value of the section transmission power;
[0007] Use a pre-trained reinforcement learning agent to generate data of the power system under several working conditions based on the power system control model to construct a data set, where the power system control model is restricted by the unit commitment model;
[0008] Perform policy extraction based on the data set, and obtain an optimal policy according to the policy extraction result, where the optimal policy is suitable for indicating the optimal safe operation boundary.
[0009] Optionally, the constraint conditions of the unit commitment model further include at least one of the following:
[0010] Output constraints of traditional units;
[0011] Output constraints of new energy units;
[0012] Power balance constraints between power sources and loads;
[0013] Unit start-up and shut-down time constraints;
[0014] System reserve constraints;
[0015] Unit ramp rate constraints.
[0016] Optionally, the power flow value of the section is expressed by the following formula:
[0017]
[0018] where P s (t) represents the power flow value of section s at time t, is the set of sections, and represents the set of tie lines l of section s, represents the set of traditional units g, represents the set of new energy units w, is the set of loads d, is the set of DC transmission lines dc, G g-l represents the power flow transfer factor of traditional unit g for tie line l, G w-l represents the power flow transfer factor of new energy unit w for tie line l, G d-l represents the power flow transfer factor of load d for tie line l, G dc-l represents the power flow transfer factor of DC transmission line dc for tie line l, P g (t) is the output of traditional unit g at time t, P w (t) is the output of new energy unit w at time t, P d (t) is the active power of load d at time t, P dc (t) is the DC transmission power plan value corresponding to DC transmission line dc at time t.
[0019] Optionally, the pre-trained reinforcement learning agent is obtained through iterative training, where each round of training in the iterative training includes:
[0020] Observing the current state data of the power system control model through the reinforcement learning agent to be trained;
[0021] Using the self-decision mechanism of the reinforcement learning agent to be trained to determine the output action according to the current state data;
[0022] Using a preset agent reward function, based on the output action and the current state data, perform current-round training on the reinforcement learning agent to be trained, where the agent reward function is determined by a control target, an action cost, and a system state;
[0023] Wherein, the output action is suitable for being executed by the power system control model, and the state of the power system control model changes as the output action is executed by the power system control model.
[0024] Optionally, the extracting a policy based on the data set and obtaining an optimal policy according to the policy extraction result includes:
[0025] Adopt a weighted oblique decision tree algorithm based on information gain ratio to extract a policy according to the data set, and obtain a policy extraction result for indicating a decision model;
[0026] Use a preset number of evaluation metrics to evaluate the decision model, so as to determine whether the decision model is an optimal decision model for indicating the optimal policy according to the evaluation result.
[0027] Optionally, the several evaluation metrics include at least one of the following:
[0028] A policy fidelity metric, which is used to measure whether the outputs obtained by the pre-trained reinforcement learning agent and the decision model are consistent under the same sample input;
[0029] A policy actual control performance metric, which is used to evaluate the performance difference between the pre-trained reinforcement learning agent and the decision model;
[0030] A model complexity metric, which is used to characterize the number of model parameters and / or the depth of the decision model.
[0031] In a second aspect, an embodiment of the present application provides a device for extracting the safe operation boundary of a power grid, including:
[0032] A unit commitment model determination module, which is used to determine a unit commitment model, wherein the objective function of the unit commitment model is suitable for characterizing the minimization of the target cost, and the target cost includes the power generation cost of traditional units and the cost of abandoning new energy. Wherein, the constraint conditions of the unit commitment model at least include a section limit constraint, and the section limit constraint is suitable for constraining the power flow value of a section of the power system between the negative direction limit value of the section transmission power and the positive direction limit value of the section transmission power;
[0033] A dataset construction module, configured to use a pre-trained reinforcement learning agent to generate data of a power system under several working conditions based on a power system control model, so as to construct a dataset, wherein the power system control model is limited by the unit commitment model;
[0034] A policy acquisition module, configured to extract a policy based on the dataset and obtain an optimal policy according to the policy extraction result, wherein the optimal policy is suitable for indicating an optimal safe operation boundary.
[0035] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0036] In a fourth aspect, an embodiment of the present application provides a computer program product, including computer instructions, which implement the steps of the method described in any one of the above when executed by a processor.
[0037] In a fifth aspect, an embodiment of the present application provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0038] In summary, the embodiments of the present application at least have the following beneficial effects:
[0039] By adopting the embodiments of the present application, by determining a unit commitment model, wherein the objective function of the unit commitment model is suitable for characterizing the minimization of the target cost, the target cost includes the power generation cost of traditional units and the cost of abandoning new energy, and the constraint conditions of the unit commitment model at least include section limit constraints, and the section limit constraints are suitable for restricting the power flow value of the section of the power system between the negative direction limit value of the section transmission power and the positive direction limit value of the section transmission power; using a pre-trained reinforcement learning agent to generate data of the power system under several working conditions based on a power system control model, so as to construct a dataset, wherein the power system control model is limited by the unit commitment model; extracting a policy based on the dataset and obtaining an optimal policy according to the policy extraction result, wherein the optimal policy is suitable for indicating an optimal safe operation boundary, so that the optimal safe operation boundary can be efficiently and accurately identified under the premise of meeting the section limit constraints, and the waste of the operation space can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a schematic flowchart of a method for extracting the safe operation boundary of a power grid provided by an embodiment of the present application;
[0041] Figure 2It is a schematic diagram of agent training provided by an embodiment of the present application;
[0042] Figure 3 It is a schematic diagram of the IGR-WODT algorithm provided by an embodiment of the present application;
[0043] Figure 4 It is a schematic structural diagram of a power grid security operation boundary extraction device provided by an embodiment of the present application;
[0044] Figure 5 It is a schematic diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0045] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0046] In the description of the present application, the terms "first", "second", "third", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more. In the description of the present application, the term "including" and its variants are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "according to" is "at least partially according to". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments".
[0047] In the description of the present application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific situations.
[0048] In the description of the present application, it should be noted that unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as those commonly understood by those skilled in the technical field to which the present application belongs. The terms used in the description of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.
[0049] The following explains some term concepts related to the embodiments of the present application:
[0050] Section. In the power system, a section usually refers to a specific part of the power grid, which consists of a group of transmission lines and / or transformers that connect two or more electrical regions. A section is usually used to describe the power flow situation of a specific part of the power system, that is, the power flow condition of this part.
[0051] In the first aspect, referring to Figure 1 , a flowchart of a method for extracting the power grid safe operation boundary provided by the embodiments of the present application is shown. This method includes steps S101 - S103, specifically as follows:
[0052] S101. Determine the unit commitment model. Among them, the objective function of the unit commitment model is suitable for characterizing the minimization of the target cost. The target cost includes the traditional unit power generation cost and the new energy abandonment cost. Among them, the constraint conditions of the unit commitment model at least include the section limit constraint, and the section limit constraint is suitable for restricting the power flow value of the section of the power system between the negative direction limit value and the positive direction limit value of the section transmission power;
[0053] In an example, the decision variables in this unit commitment model can include the generator output and the unit start - stop state, and the objective function can be expressed by the following formula:
[0054]
[0055] Among them, are respectively the set of traditional units and the set of new energy units in sequence, C g represents the quadratic fuel cost function of the traditional unit, a g , b g , c g are respectively the cost coefficients related to the traditional unit g, P g (t) represents the output power of the traditional unit g at time t, C w (t) represents the cost function of the new energy unit, c w represents the cost coefficient related to the new energy unit w, ΔP w(t) represents the deviation between the actual output and the planned output of the new energy unit w at time t.
[0056] S102, using a pre-trained reinforcement learning agent, based on the power system control model, generate data of the power system under several working conditions to construct a data set, wherein the power system control model is restricted by the unit commitment model;
[0057] S103, perform policy extraction based on the data set, and obtain an optimal policy according to the policy extraction result, wherein the optimal policy is suitable for indicating the optimal safe operation boundary.
[0058] In an alternative embodiment, the constraint conditions of the unit commitment model further include at least one of the following:
[0059] Conventional unit output constraint;
[0060] In one example, the conventional unit output constraint is represented by the following formula:
[0061]
[0062] Wherein, is the set of scheduling periods; P g (t) is the active power output of the gth conventional unit in the tth period u g (t) is the start-stop state of the gth conventional unit in the tth period; are the upper and lower limits of the output of the synchronous generator respectively.
[0063] New energy unit output constraint;
[0064] In one example, the new energy unit output constraint is represented by the following formula:
[0065]
[0066] Wherein, P w (t) is the output of the wth new energy unit in the tth period; is the predicted value of the available output of the new energy unit in the tth period.
[0067] Source-load power balance constraint;
[0068] In one example, the source-load power balance constraint is represented by the following formula:
[0069]
[0070] Wherein, is the load set; P d (t) is the active power of the dth load in the tth period is the set of DC transmission lines for external power transmission; P dc (t) is the planned value of the DC transmission power of the dc-th DC transmission line at time t
[0071] Unit start-up and shutdown time constraints;
[0072] In an example, the unit start-up and shutdown time constraints are expressed by the following formula:
[0073]
[0074]
[0075] Wherein, are the unit start-up and shutdown times respectively, T i on 、T i off are the minimum start-up and minimum shutdown times respectively.
[0076] System reserve constraints;
[0077] In an example, the system reserve constraints are expressed by the following formula:
[0078]
[0079] Wherein, R U (t), R D (t) respectively represent the lower limits of the system's positive and negative reserves, r U (t), r D (t) are intermediate variables for optimization, representing the system's positive and negative reserve levels during the optimization process respectively.
[0080] Unit ramp rate constraints.
[0081] In an example, the unit ramp rate constraints are expressed by the following formula:
[0082]
[0083] Wherein, are the upper limits of the unit's sliding and ramping powers respectively.
[0084] In some cases, in the face of the challenges brought about by load growth and large-scale access of new energy, the operating characteristics of the power grid have changed significantly. To ensure the safe and stable operation of the power grid, the dispatching and operation department usually selects the minimum value from multiple sets of section transfer limit values (Total Transfer Capacity, TTC) calculated based on multiple typical scenarios as the section transmission limit. However, under the condition of high penetration of new energy, the system operation mode becomes more complex and changeable, and extreme scenarios are difficult to accurately identify. Although the conservative limit setting strategy can ensure system safety, it also causes serious waste of the transmission channel capacity.
[0085] Through in-depth analysis, it is found that although TTC and section limit seem similar on the surface, there are significant differences in their essential connotations. TTC reflects the maximum transmission capacity under a specific operating state, and its calculation process is a process of gradually searching from a stable state to a critically stable state. The section limit, on the other hand, represents the minimum section power in the set of unstable operation modes and is used to divide the boundary between the stable region and the unstable region. Therefore, the TTC value based on a single operating state cannot be equated with the section limit. In addition, from the perspective of calculation characteristics, TTC is a dynamic parameter. Although it can accurately reflect the real-time transmission capacity, it has problems such as high calculation complexity and difficulty in adapting to the rapid changes of the power grid. Existing traditional calculation methods such as continuous power flow are difficult to solve the TTC problem containing multiple types of constraints in real time. Even if artificial intelligence methods are used, their engineering applicability and reliability still need to be verified. In contrast, the section limit, as a static parameter, has good engineering practicability, but its calculation also faces difficulties such as high-dimensional non-convex non-linear optimization. From the perspective of engineering application, as the scale of the power grid expands and the number of sections increases, the operation mode becomes more complex, and each section corresponds to different power limits under different operation modes. This poses higher requirements for the monitoring and control capabilities of operators. If the section limit is not updated in time, it may lead to the rapid expansion of local faults.
[0086] In some existing technologies, the determination of the section limit still mainly relies on traditional manual experience methods, and the section limit is set by checking under limited extreme working conditions. This method is not only inefficient but also often overly conservative. Especially in the scenario of high new energy penetration, the combination of operation modes shows an exponential growth, and the safety operation boundary determined by the traditional "one-size-fits-all" limit strategy is likely to result in a large amount of stable operation space being wasted.
[0087] In view of the above existing technologies, this application proposes an optional implementation method. In this implementation method, the power flow value of the section is expressed by the following formula:
[0088]
[0089] where, P s (t) represents the power flow value of section s at time t, is a set of cross-sections, and represents the set of tie lines l for cross-section s, represents the set of conventional units g, represents the set of new energy units w, is the set of loads d, is the set of DC external transmission lines dc, G g-l represents the power flow transfer factor of conventional unit g for tie line l, G w-l represents the power flow transfer factor of new energy unit w for tie line l, G d-l represents the power flow transfer factor of load d for tie line l, G dc-l represents the power flow transfer factor of DC external transmission line dc for tie line l, P g (t) is the output of conventional unit g at time t, P w (t) is the output of new energy unit w at time t, P d (t) is the active power of load d at time t, P dc (t) is the DC transmission power planned value corresponding to DC external transmission line dc at time t.
[0090] In an example, the cross-section limit constraint is suitable for constraining the power flow value of the cross-section of the power system between the negative limit value of the cross-section transmission power and the positive limit value of the cross-section transmission power, that is, this cross-section limit constraint can be expressed by the following formula:
[0091]
[0092] Among them, represents the negative limit value of the cross-section transmission power, represents the positive limit value of the cross-section transmission power.
[0093] It should be noted that the formula of this cross-section limit constraint can describe a single cross-section limit constraint without considering cross-section coupling.
[0094] In this embodiment, the operation stability of the power grid can be improved: by introducing a unit commitment model considering cross-section limit constraints, various constraint conditions in the power grid can be fully considered, including the output constraints of conventional units and new energy units, as well as constraints such as reserve and ramping, which helps to improve the stability and security of the power grid operation within a day under the background of new energy randomness.
[0095] In an alternative implementation, the pre-trained reinforcement learning agent is obtained through iterative training, where each round of training in the iterative training includes:
[0096] Observing the current state data of the power system control model through the reinforcement learning agent to be trained;
[0097] Using the self - decision mechanism of the reinforcement learning agent to be trained, determine the output action according to the current state data;
[0098] Using a preset agent reward function, based on the output action and the current state data, perform the current round of training on the reinforcement learning agent to be trained, where the agent reward function is determined by a control objective, an action cost, and a system state;
[0099] Wherein, the output action is suitable for being executed by the power system control model, and the state of the power system control model changes as the output action is executed by the power system control model.
[0100] In some cases, the power system control problem is a high - dimensional, non - linear optimal decision - making problem, so its formulation description usually can include:
[0101] Objective function:
[0102] minf(x t ,y t )
[0103] Power system dynamic model:
[0104]
[0105] Power system power flow equation:
[0106] Φ(x t ,y t )=0
[0107] Other constraint conditions (such as various safety - boundary constraints, etc., which can be set in specific control tasks):
[0108] ψ(x t ,y t )≤0
[0109] Wherein, x t is a vector composed of power system state variables, such as node power angle values; y t is a vector composed of power system control variables, such as generator output.
[0110] The problems established in the above situations can usually be modeled as an MDP (Markov decision process) and solved by DRL (Deep Reinforcement Learning). The DRL agent is usually designed as a deep neural network (DNN), and this agent is suitable for interacting with a real system or a simulation system. Specifically, the agent observes the real-time state of the system, denoted as x, and then the agent outputs an action according to its own decision-making mechanism, denoted as y. After the system executes the action y, the system state changes and is then observed by the agent again. Further, referring to Figure 2 , according to the control objective, considering the system state and action cost, design the agent reward function to guide the optimization of the agent's training.
[0111] After the reinforcement learning agent is trained, the agent interacts with the system under a large number of different working conditions, and then records the input x and output y of the agent, denoted as (x, y), which can be regarded as independently and identically distributed. Finally, construct a data set, denoted as S:
[0112] S = {(x 1 , y 1 ), (x 2 , y 2 ), (x i , y i ),..., (x N , y N )}
[0113] where x ∈ R n , y ∈ N, i represents the i-th sample of the data set S, and N is the total number of samples in the data set S.
[0114] In an alternative embodiment, extracting a policy based on the data set and obtaining an optimal policy according to the policy extraction result includes:
[0115] Adopt a weighted oblique decision tree algorithm based on the information gain ratio to extract a policy according to the data set, and obtain a policy extraction result for indicating a decision model;
[0116] Use a plurality of preset evaluation metrics to evaluate the decision model, so as to determine whether the decision model is an optimal decision model for indicating the optimal policy according to the evaluation result.
[0117] In an example, referring to Figure 3, continuing to use the concept of the above dataset S, the dataset S contains the decision-making knowledge learned by DRL during the training phase. In this dataset, each state-action pair (x, y) represents that when DRL observes the power grid state x, DRL will make its optimal decision y. According to imitation learning, regarding the DRL model as a "teacher", the strategy of DRL can be expressed as:
[0118] f(x) → y, x ∈ R n , y ∈ n
[0119] When the action space is discrete, y ∈ N; when the action space is continuous, y ∈ R. This application takes the discrete action space as an example for explanation.
[0120] In this way, the above formula can be naturally modeled as a classification task in supervised learning later.
[0121] For the WODT (Weighted Oblique Decision Tree) algorithm, under each internal node, first define the training sample set S train =(x i , y i ) ∈ S, i = 1, 2,..., M, M ≤ N. Then, WODT divides the sample set S train into left and right subsets based on the logistic regression model. Specifically, under the current node, WODT uses the logistic regression model to calculate the probability of each sample belonging to the left and right subsets, as shown in the following formula:
[0122]
[0123] where θ is the model parameter that needs to be continuously updated during the model training process; when , the i-th sample belongs to the left subset; when , the i-th sample belongs to the right subset. Then, define the weights of each sample as:
[0124]
[0125] Therefore, IGR-WODT (Information Gain Ratio Weighted Oblique Decision Tree, based on the information gain ratio of the weighted oblique decision tree algorithm, abbreviated as IGR-WODT) defines 2 sample sets associated with the sample weights:
[0126]
[0127] Therefore, WODT defines its objective function based on the weighted information entropy as:
[0128] E(θ) = W L H L + W R H R
[0129]
[0130] Where K is the total number of all sample categories in the sample set, and k represents the k-th sample category.
[0131] Furthermore, WODT can use the L-BFGS algorithm to solve the above "objective function defined by WODT based on weighted information entropy" to obtain the optimal parameter θ of the current node best , thereby obtaining the policy extraction result for indicating the decision model. It should be understood that on many open-source data sets, WODT has been proven to have better performance than other oblique decision trees.
[0132] In this embodiment, the accuracy and reliability of the policy can be enhanced: by training the reinforcement learning agent, the system can interactively generate data under various working conditions, thereby obtaining a more accurate and reliable operation policy. Combining the weighted oblique decision tree algorithm with information gain ratio (IGR-WODT), the optimal policy boundary can be effectively extracted to ensure that the policy has high fidelity and actual control performance.
[0133] In an alternative embodiment, the several evaluation metrics include at least one of the following:
[0134] The policy fidelity metric, which is used to measure whether the outputs obtained by the pre-trained reinforcement learning agent and the decision model are consistent under the same sample input;
[0135] In an example, the number of the same sample inputs can be N. Thus, the policy fidelity metric F p can be calculated by the following formula:
[0136]
[0137] where y and respectively represent the outputs of DRL and WODT under the same sample input conditions in sequence; I(·) is the indicator function.
[0138] The policy actual control performance metric, which is used to evaluate the performance difference between the pre-trained reinforcement learning agent and the decision model;
[0139] In an example, the policy actual control performance metric can be used to evaluate the effect of applying the IGR-WODT policy to the actual control scenario. Specifically, it represents the average return in each episode. Assume that the average return r of the DRL policy in the actual control scenario per episodee Correspondingly, the average return obtained by WODT in the corresponding scenario is denoted as r e ′ . Therefore, R e = r e ′ - r e > 0 indicates that the control performance of WODT is better than that of DRL, and vice versa.
[0140] The model complexity index is used to characterize the number of model parameters and / or the model depth of the decision model.
[0141] It should be noted that in addition to policy fidelity and actual control performance, model complexity is also an important consideration index. The model complexity can be measured by the number of model parameters or the depth. In other words, the model complexity index of the decision model can be determined by its number of model parameters and / or model depth.
[0142] In this embodiment, the model complexity can be simplified and the interpretability can be improved: IGR-WODT is used for policy extraction. While ensuring the policy effect, the model complexity can be controlled, thereby simplifying the structure of the policy model, improving its interpretability and operability, and being more suitable for application in the security domain search and decision support of the actual power grid.
[0143] In the second aspect, correspondingly, the embodiment of the present application further provides a device for extracting the power grid security operation boundary, which can implement all the processes of the method for extracting the power grid security operation boundary provided in the above embodiment.
[0144] See Figure 4 , which shows the structural schematic diagram of the device for extracting the power grid security operation boundary provided by the embodiment of the present application. The device includes:
[0145] The unit commitment model determination module 401 is used to determine the unit commitment model. Among them, the objective function of the unit commitment model is suitable for characterizing the minimization of the target cost. The target cost includes the power generation cost of traditional units and the cost of abandoning new energy. Among them, the constraint conditions of the unit commitment model at least include the section limit constraint, and the section limit constraint is suitable for restricting the power flow value of the section of the power system between the negative direction limit value of the section transmission power and the positive direction limit value of the section transmission power;
[0146] The data set construction module 402 is used to generate data of the power system under several working conditions based on the power system control model by using a pre-trained reinforcement learning agent, so as to construct a data set, where the power system control model is restricted by the unit commitment model;
[0147] A policy acquisition module 403, configured to extract a policy based on the data set and obtain an optimal policy according to the policy extraction result, where the optimal policy is suitable for indicating an optimal safe operation boundary.
[0148] In an alternative embodiment, the constraint conditions of the unit commitment model further include at least one of the following:
[0149] Output constraint of traditional units;
[0150] Output constraint of new energy units;
[0151] Power balance constraint between power sources and loads;
[0152] Unit start-up and shut-down time constraint;
[0153] System reserve constraint;
[0154] Unit ramp rate constraint.
[0155] In an alternative embodiment, the power flow value of the section is represented by the following formula:
[0156]
[0157] Where P s (t) represents the power flow value of section s at time t, is the set of sections, and represents the set of tie lines l of section s, represents the set of traditional units g, represents the set of new energy units w, is the set of loads d, is the set of DC external transmission lines dc, G g-l represents the power flow transfer factor of traditional unit g for tie line l, G w-l represents the power flow transfer factor of new energy unit w for tie line l, G d-l represents the power flow transfer factor of load d for tie line l, G dc-l represents the power flow transfer factor of DC external transmission line dc for tie line l, P g (t) is the output of traditional unit g at time t, P w (t) is the output of new energy unit w at time t, P d (t) is the active power of load d at time t, P dc (t) is the DC transmission power planned value corresponding to DC external transmission line dc at time t.
[0158] In an alternative embodiment, the pre-trained reinforcement learning agent is obtained through iterative training, where each round of training in the iterative training includes:
[0159] The current state data of the power system control model is observed by a reinforcement learning agent to be trained;
[0160] An output action is determined according to the current state data by using the self - decision - making mechanism of the reinforcement learning agent to be trained;
[0161] The reinforcement learning agent to be trained is trained in the current round by using a preset agent reward function based on the output action and the current state data, wherein the agent reward function is determined by a control target, an action cost, and a system state;
[0162] Wherein, the output action is suitable for being executed by the power system control model, and the state of the power system control model changes as the output action is executed by the power system control model.
[0163] In an alternative embodiment, the extracting a policy based on the data set and obtaining an optimal policy according to the policy extraction result includes:
[0164] A weighted oblique decision tree algorithm based on information gain ratio is adopted to extract a policy according to the data set, and a policy extraction result for indicating a decision model is obtained;
[0165] A preset number of evaluation metrics are used to evaluate the decision model, so as to determine whether the decision model is an optimal decision model for indicating the optimal policy according to the evaluation result.
[0166] In an alternative embodiment, the several evaluation metrics include at least one of the following:
[0167] A policy fidelity metric, which is used to measure whether the outputs obtained by the pre - trained reinforcement learning agent and the decision model are consistent under the same sample input;
[0168] A policy actual control performance metric, which is used to evaluate the performance difference between the pre - trained reinforcement learning agent and the decision model;
[0169] A model complexity metric, which is used to characterize the number of model parameters and / or the depth of the decision model.
[0170] In a third aspect, an embodiment of the present application provides a computer - readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method described in any one of the above are implemented.
[0171] Fourthly, an embodiment of the present application provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the method described in any one of the above.
[0172] Fifthly, an embodiment of the present application provides a computer device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the method described in any one of the above are implemented.
[0173] See Figure 5 , the computer device of this embodiment includes: a processor 501, a memory 502, and a computer program stored in the memory 502 and operable on the processor 501, such as a power grid security operation boundary extraction program. When the processor 501 executes the computer program, the steps in each of the above embodiments of the power grid security operation boundary extraction method are implemented, such as Figure 1 the steps S101 - S103 shown.
[0174] Exemplarily, the computer program can be divided into one or more modules / units. The one or more modules / units are stored in the memory 502 and executed by the processor 501 to complete the present application. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the computer device.
[0175] The computer device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include, but is not limited to, a processor 501 and a memory 502. Those skilled in the art can understand that the schematic diagram is only an example of the computer device, and does not constitute a limitation on the computer device. It may include more or fewer components than shown, or combine certain components, or different components. For example, the computer device may further include input / output devices, network access devices, a bus, etc.
[0176] The processor 501 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor 501 may also be any conventional processor, etc. The processor 501 is the control center of the computer device, and connects various parts of the entire computer device through various interfaces and lines.
[0177] The memory 502 can be used to store the computer programs and / or modules. The processor 501 realizes various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 502, and by calling the data stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0178] Among them, if the modules / units integrated in the computer device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 501, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate forms, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.
[0179] In summary, the embodiments of the present application have at least the following beneficial effects:
[0180] By adopting the embodiments of the present application, by determining a unit commitment model, wherein the objective function of the unit commitment model is suitable for characterizing the minimization of the target cost, the target cost includes the power generation cost of traditional units and the cost of abandoning new energy, and the constraint conditions of the unit commitment model at least include section limit constraints, and the section limit constraints are suitable for constraining the power flow value of the section of the power system between the negative direction limit value of the section transmission power and the positive direction limit value of the section transmission power; using a pre-trained reinforcement learning agent, based on the power system control model, generating data of the power system under several working conditions to construct a data set, wherein the power system control model is restricted by the unit commitment model; extracting a strategy based on the data set, and obtaining an optimal strategy according to the strategy extraction result, wherein the optimal strategy is suitable for indicating the optimal safe operation boundary, so that it is possible to efficiently and accurately identify the optimal safe operation boundary on the premise of meeting the section limit constraints and reduce the waste of the operation space.
[0181] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus a necessary hardware platform, and of course, it can also be implemented entirely through hardware. Based on such an understanding, all or part of the technical solution of this application that contributes to the background technology can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.
[0182] The above is the preferred embodiment of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of this application.
Claims
1. A method for extracting safe operation boundaries of a power grid, characterized in that: include: Determine a unit combination model, wherein the objective function of the unit combination model is suitable for characterizing the minimization of target costs, the target costs include the power generation costs of traditional units and the costs of abandoned new energy sources, wherein the constraints of the unit combination model at least include section limit constraints, and the section limit constraints are suitable for constraining the power flow value of the section of the power system to be between the negative direction limit value of the section transmission power and the positive direction limit value of the section transmission power; Using a pre-trained reinforcement learning agent, based on a power system control model, generating data of the power system under a plurality of operating conditions to construct a data set, wherein the power system control model is constrained by the unit commitment model; A strategy extraction is performed based on the data set, and an optimal strategy is obtained according to the strategy extraction result, wherein the optimal strategy is suitable for indicating an optimal safe operation boundary.
2. The method according to claim 1, characterized in that: The constraint conditions of the unit commitment model also include at least one of the following: Output constraints of traditional units; Output constraints of new energy units; Source-load power balance constraints; Unit start and stop time constraints; System standby constraints; Unit climbing constraints.
3. The method according to claim 1, characterized in that The tidal flow value of the section is expressed by the following formula: Among them, P s (t) represents the tidal flow value of section s at time t, is a set of cross sections, and represents the set of tie lines l of section s, represents the set of traditional units g, represents the set of new energy units w, is the set of loads d, is the set of DC transmission lines dc, G g-l G represents the power transfer factor of the traditional unit g to the tie line l, w-l G represents the power transfer factor of the new energy unit w to the tie line l, d-l G represents the power transfer factor of load d to tie line l, dc-l P represents the power transfer factor of the DC transmission line dc to the tie line l, g (t) is the output of the traditional unit g at time t, P w (t) is the output of the new energy unit w at time t, P d (t) is the active power of load d at time t, P dc (t) is the planned DC transmission power value corresponding to the DC transmission line dc at time t.
4. The method according to claim 1, characterized in that The pre-trained reinforcement learning agent is obtained through iterative training, wherein each round of training in the iterative training includes: Observing current state data of the power system control model through a reinforcement learning agent to be trained; Determine an output action according to the current state data by using the decision-making mechanism of the reinforcement learning agent to be trained; Using a preset agent reward function, based on the output action and the current state data, the reinforcement learning agent to be trained is trained in a current round, wherein the agent reward function is determined by a control target, an action cost, and a system state; The output action is suitable for execution by the power system control model, and the state of the power system control model changes as the output action is executed by the power system control model.
5. The method according to claim 1, characterized in that: The extracting strategy based on the data set and obtaining the optimal strategy according to the strategy extraction result includes: Using a weighted tilted decision tree algorithm based on information gain ratio, performing strategy extraction according to the data set, and obtaining a strategy extraction result for indicating a decision model; The decision model is evaluated using a plurality of preset evaluation indicators to determine whether the decision model is an optimal decision model for indicating the optimal strategy based on the evaluation result.
6. The method according to claim 5, characterized in that The evaluation indicators include at least one of the following: The policy fidelity index is used to measure whether the outputs obtained by the pre-trained reinforcement learning agent and the decision model are consistent under the same sample input; The actual control performance index of the strategy is used to evaluate the performance difference between the pre-trained reinforcement learning agent and the decision model; The model complexity index is used to characterize the model parameter quantity and / or model depth of the decision model.
7. A device for extracting safe operation boundaries of a power grid, characterized in that: include: A unit combination model determination module is used to determine a unit combination model, wherein the objective function of the unit combination model is suitable for minimizing the target cost, the target cost includes the power generation cost of traditional units and the cost of abandoned new energy, wherein the constraint condition of the unit combination model at least includes a section limit constraint, and the section limit constraint is suitable for constraining the power flow value of the section of the power system to be between the negative direction limit value of the section transmission power and the positive direction limit value of the section transmission power; A data set construction module, for generating data of the power system under a plurality of working conditions based on a power system control model using a pre-trained reinforcement learning agent, so as to construct a data set, wherein the power system control model is restricted by the unit commitment model; A strategy acquisition module is used to extract strategies based on the data set and obtain an optimal strategy according to the strategy extraction result, wherein the optimal strategy is suitable for indicating an optimal safe operation boundary.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor implements the method according to any one of claims 1 to 6 when executing the computer program.
Citation Information
Patent Citations
Power grid security boundary feature extraction method and system based on data driving
CN119005502A