Generator set control method and apparatus, and device

By training agents on generator nodes and combining them with the surrogate subgradient method, the sub-hour generator combination problem can be solved quickly, which solves the problem of slow computation speed in the existing technology and realizes efficient real-time scheduling of power systems.

WO2026011476A1PCT designated stage Publication Date: 2026-01-15TSINGHUA UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/105960
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2024-07-17
Publication Date
2026-01-15

Smart Images

  • Figure CN2024105960_15012026_PF_FP_ABST
    Figure CN2024105960_15012026_PF_FP_ABST
Patent Text Reader

Abstract

A generator set control method and apparatus, and a device. Comprising: constructing a state transition model for sub-problems of a single unit, and adding as a state in the model a penalty price corresponding to a Lagrange multiplier for each time period Using a reinforcement learning algorithm to train a startup / shutdown strategy and a power increase / decrease strategy for each unit; using a surrogate sub-gradient method to relax constraints coupled to different units in a UC problem, using the surrogate sub-gradient method to perform iteration and Lagrange multiplier updating, solving sub-problems in the iteration process using a trained reinforcement learning agent to perform sequential decision-making, and iterating repeatedly until convergence, so as to obtain an optimal solution to a dual problem; and performing a feasibility operation on a resulting unit commitment state, and controlling generator set nodes.
Need to check novelty before this filing date? Find Prior Art

Description

A control method, device and equipment for a generator set

[0001] Related applications

[0002] This application claims priority to Chinese invention patent filed on July 8, 2024, application number 202410908936.8, entitled "A control method, device and equipment for a generator set", all contents of which are incorporated herein by reference. Technical Field

[0003] The embodiments in this specification relate to the field of power system dispatching, and in particular to a control method, device and equipment for generator sets. Background Technology

[0004] The control of generator units is an important and complex control problem in power systems. Its research and application are of great significance for improving the economy, reliability, environmental protection and market operation efficiency of power systems. It is an important part of the macro-dispatch of power systems, determining whether all generator units in the power grid should be turned on or off at different times, and satisfying complex supply and demand balance constraints and security constraints, while improving the economy and environmental protection of power systems.

[0005] Unit Commitment (UC) problems typically use hours as the decision interval. However, with the widespread integration of renewable energy into the power system, the system needs to address the variability and instability brought about by renewable energy, requiring real-time dispatch capabilities. Therefore, rapidly resolving UC problems with sub-hour intervals (typically 15 minutes) in large-scale power systems is of great practical significance. However, these problems are far more complex than UC problems with hourly intervals, and obtaining feasible near-optimal solutions to large-scale sub-hourly interval UC problems within a short timeframe is both important and challenging.

[0006] The Common Unified Problem (UC) is typically modeled as a Mixed Integer Programming (MIP) problem. This model involves a large number of integer and continuous decision variables, and is constrained by numerous equality and inequality constraints. These constraints include system-level and individual-level constraints. System-level constraints cover things like system power balance constraints, reserve capacity constraints, and DC power flow constraints. Individual-level constraints describe generator operation, including ramping constraints and minimum start-stop time constraints. The UC problem is an NP-hard problem, and many methods exist for solving it, including Mixed Integer Linear Programming (MILP), Lagrange Relaxation (LR), heuristic methods, and dynamic programming. These methods aim to obtain a good feasible solution in a short time; however, their computational time remains insufficient when dealing with large-scale sub-hour interval UC problems. The SLR algorithm based on surrogate subgradient iteration and its improved methods are currently the best methods for calculating whether all generators in the power grid should be on or off at different times. However, their computational speed is still unsatisfactory, mainly because their iteration process requires repeatedly calculating the subproblems corresponding to each generator under different Lagrange multipliers.

[0007] Industrial-scale UC problems are typically solved using commercial solvers such as CPLEX or GUROBI, which employ advanced branch and cut (B&C) or branch and bound (B&B) algorithms. However, these solvers perform poorly when solving large-scale sub-hour UC problems. Lagrange relaxation methods are widely used, particularly for UC problems in large-scale power systems, because they can relax system-level constraints in UC problems and decouple the UC problem into independent subproblems for each cell. Lagrange relaxation requires iteratively updating the Lagrange multipliers using subgradients, which are obtained by solving all subproblems. SLR improves upon traditional LR by overcoming the "tug-of-war" phenomenon in obtaining subgradients and Lagrange multipliers. Despite these improvements, SLR still requires increased computational speed. During the iterative computation of the UC problem, SLR needs to repeatedly compute subproblems under different Lagrange multipliers, which often exhibit significant similarities. SLR typically uses B&C or other methods to compute these subproblems. This does not take into account the similarities and characteristics of the problems, but rather models and computes each subproblem under different Lagrange multipliers as a completely new problem. Improving the computation speed of these subproblems, and thus improving the computation speed of large-scale UC problems at sub-hour intervals, is crucial.

[0008] Summary of the Invention

[0009] The specific technical solutions of the embodiments in this specification are as follows:

[0010] On the one hand, embodiments of this specification provide a control method for a generator set, including:

[0011] Obtain the load of multiple power consumption nodes in the target power grid under multiple forecast periods and the current operating status of multiple generator nodes in the target power grid;

[0012] Receive the set first dual variable, which includes the Lagrange multipliers under the power balance constraint and DC power flow constraint of the target power grid;

[0013] Based on the first dual variable, the DC power flow factor of each generator node and its current operating state, a state vector corresponding to each generator node is constructed. The state vector is then input into the trained reinforcement learning model corresponding to each generator node for calculation to obtain the predicted operating state of each generator node under the first dual variable for the multiple prediction periods.

[0014] The first dual variable is updated based on the predicted operating status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period to obtain the second dual variable. The step of constructing the state vector corresponding to each generator node based on the second dual variable, the DC power flow factor of each generator node and its current operating status is performed to obtain the predicted operating status of each generator node under the second dual variable for the multiple predicted periods.

[0015] Determine whether the second dual variable converges;

[0016] If convergence is not achieved, the second dual variable is used as the first dual variable, and the step of updating the first dual variable based on the predicted working status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is repeated.

[0017] If convergence is achieved, the predicted working state of each generator node in each prediction period under the second dual variable will be taken as the first target working state of each generator node in the multiple prediction periods.

[0018] The operating state of each generator node that violates the DC power flow constraint in the first target operating state under each prediction period is determined as the operating state to be optimized, and the second target operating state is obtained by calculating the operating state to be optimized using a commercial solver.

[0019] The operating state of each generator node in each prediction period that does not violate the DC power flow constraint in the first target operating state and the corresponding second target operating state are taken as the final operating state of each generator node in each prediction period.

[0020] The corresponding generator nodes in the target power grid are controlled by utilizing the final operating state of each generator node in each predicted period.

[0021] On the other hand, embodiments of this specification also provide a control device for a generator set, including:

[0022] The data acquisition unit is used to acquire the load of multiple power consumption nodes in the target power grid under multiple forecast periods and the current working status of multiple generator set nodes in the target power grid.

[0023] The first dual variable receiving unit is used to receive the set first dual variable, which includes the Lagrange multipliers under the power balance constraint and DC power flow constraint of the target power grid.

[0024] The working state prediction unit is used to construct a state vector corresponding to each generator node based on the first dual variable, the DC power flow factor of each generator node and its current working state, and input the state vector into the trained reinforcement learning model corresponding to each generator node for calculation to obtain the predicted working state of each generator node under the first dual variable in the multiple prediction periods.

[0025] The dual variable update unit is used to update the first dual variable according to the predicted operating status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period, to obtain the second dual variable, and to perform the step of constructing the state vector corresponding to each generator node according to the second dual variable, the DC power flow factor of each generator node and its current operating status, to obtain the predicted operating status of each generator node under the second dual variable for the multiple predicted periods.

[0026] The dual variable convergence judgment unit is used to determine whether the second dual variable has converged. If it has not converged, the second dual variable is used as the first dual variable, and the step of updating the first dual variable based on the predicted working state of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is repeated. If it has converged, the predicted working state of each generator node under the second dual variable in each predicted period is used as the first target working state of each generator node in the multiple predicted periods.

[0027] The working state optimization unit is used to determine the working state of each generator node in the first target working state under each prediction period that violates the DC power flow constraint, as the working state to be optimized, and to calculate the working state to be optimized using a commercial solver to obtain the second target working state.

[0028] The final operating state determination unit is used to take the operating state of each generator node in the first target operating state under each prediction period, which does not violate the DC power flow constraint, and the corresponding second target operating state as the final operating state of each generator node under each prediction period.

[0029] A control unit is used to control the corresponding generator nodes in the target power grid using the final operating state of each generator node in each predicted period.

[0030] On the other hand, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the above-described method.

[0031] Using the embodiments in this specification, the load of multiple power-consuming nodes in the target power grid under multiple prediction time periods and the current operating state of multiple generator nodes are first obtained. Then, the coupling constraints between each generator node are relaxed, which can decouple them into solving sub-problems corresponding to each generator node. Solving the sub-problems is to obtain the optimal solution for each generator node. Specifically, the user-defined first dual variable is received, and then a state vector is constructed based on the first dual variable, the DC power flow factor of each generator node, and the current operating state of the machine. The state vector is input into a pre-trained reinforcement learning sub-model for multiple iterations to obtain the predicted operating state for multiple prediction time periods corresponding to the first dual variable.

[0032] Then, the embodiments of this specification use the surrogate subgradient method to solve the optimal solution of the dual problem. That is, the first dual variable is updated according to the predicted operating state of the generator node under the first dual variable for multiple prediction periods and the load of multiple power consumption nodes under the corresponding prediction periods to obtain the second dual variable. The reinforcement learning model is then recalculated using the second dual variable to obtain the predicted operating state corresponding to multiple prediction periods under the second dual variable. After the second dual variable converges, the optimal solution of the dual problem (the converged second dual variable) and the optimal solution of the corresponding relaxation problem (the predicted operating state corresponding to multiple prediction periods) are obtained.

[0033] However, the optimal solution to the relaxation problem may not satisfy all relaxed constraints. Therefore, the embodiments in this specification use a heuristic algorithm to find a feasible solution in the neighborhood, that is, to determine the working state of the generator node that violates the constraints in the optimal solution of the relaxation problem in each prediction period. Then, a commercial solver is used to solve these working states that violate the constraints. The solution results and the working states that do not violate the constraints in the optimal solution of the relaxation problem are used as the final working states. The generator node is controlled using the final working states.

[0034] The method described in this specification, utilizing a surrogate subgradient framework combined with reinforcement learning to solve subproblems, can quickly solve subproblems under different Lagrange multipliers during the iteration process of the surrogate subgradient method, thus greatly accelerating the surrogate subgradient iteration process. Finally, the operating states that violate constraints are identified and solved using commercial solvers, while operating states that do not violate constraints allow for direct control of the generator unit nodes. Compared to directly using commercial solvers to solve all operating states, the method described in this specification can find a near-optimal feasible solution to the unit combination problem in a faster time, improving the computational speed for calculating whether generator units should be turned on or off at different times in unit combination problems. Attached Figure Description

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0036] Figure 1 is a flowchart illustrating a control method for a generator set according to an embodiment of this specification;

[0037] Figure 2 shows a schematic diagram of the process by which the reinforcement learning model calculates the prediction working state of each generator node under the first dual variable for the multiple prediction periods in the embodiments of this specification.

[0038] Figure 3 shows a flowchart illustrating the process of updating the dual variable in an embodiment of this specification.

[0039] Figure 4 is a schematic diagram of the structure of a generator set control device in an embodiment of this specification.

[0040] Figure 5 shows a schematic diagram of the structure of the computer device in the embodiment of this specification.

[0041] [Explanation of reference numerals in the attached diagram]: 401, Data acquisition unit; 402, First dual variable receiving unit; 403, Working state prediction unit; 404, Dual variable update unit; 405, Dual variable convergence judgment unit; 406, Working state optimization unit; 407, Final working state determination unit; 408, Control unit; 502, Computer equipment; 504, Processing equipment; 506, Storage resources; 508, Drive system; 510, Input / output module; 512, Input device; 514, Output device; 516, Presentation device; 518, Graphical user interface; 520, Network interface; 522, Communication link; 524, Communication bus. Detailed Implementation

[0042] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0043] To address the problems existing in the prior art, this specification provides a generator set control method. The invention aims to improve the computational speed of surrogate subgradient methods for large-scale sub-hour interval UC problems by utilizing reinforcement learning (RL). Reinforcement learning can learn approximately optimal strategies for subproblems under different Lagrange multipliers, enabling rapid computation of subproblems and significantly improving the computational speed of surrogate subgradient methods. Ultimately, this improves the computational speed of generator set combination problems involving generator set states at different times. Figure 1 shows a schematic flowchart of the generator set control method in this specification. The figure illustrates the process of calculating the generator set's operating state and controlling the generator set using the calculated operating state. The order of steps listed in the embodiment is merely one possible execution order among many and does not represent the only possible execution order. In actual system or device products, the method can be executed sequentially or in parallel according to the embodiment or the accompanying drawings.

[0044] As shown in Figure 1, the method may include:

[0045] Step 101: Obtain the load of multiple power consumption nodes in the target power grid under multiple forecast periods and the current operating status of multiple generator set nodes in the target power grid;

[0046] In this step, the intervals between multiple forecast periods can be sub-hourly intervals (typically 15 minutes). The loads of multiple power consumption nodes in the target power grid during these multiple forecast periods are predictable or can be determined by users based on their electricity demand. The current operating status of multiple generator nodes in the target power grid includes their start / stop status, output power, etc.

[0047] Step 102: Receive the set first dual variable, which includes the power balance constraint and the Lagrange multiplier under the DC power flow constraint of the target power grid;

[0048] Step 103: Construct a state vector for each generator node based on the first dual variable, the DC power flow factor of each generator node and its current operating state, and input the state vector into the trained reinforcement learning model corresponding to each generator node for calculation to obtain the predicted operating state of each generator node under the first dual variable for the multiple prediction periods.

[0049] Step 104: Update the first dual variable according to the predicted operating status of the multiple generator nodes under the first dual variable for the multiple prediction periods and the load of the multiple power consumption nodes under the corresponding prediction periods to obtain the second dual variable. Then, perform the step of constructing the state vector corresponding to each generator node according to the second dual variable, the DC power flow factor of each generator node and its current operating status to obtain the predicted operating status of each generator node under the second dual variable for the multiple prediction periods.

[0050] Step 105: Determine whether the second dual variable has converged;

[0051] Step 106: If convergence is not achieved, the second dual variable is used as the first dual variable, and the step of updating the first dual variable based on the predicted working status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is repeated.

[0052] Step 107: If convergence is achieved, the predicted working state of each generator node under the second dual variable for each prediction period shall be taken as the first target working state of each generator node in the multiple prediction periods.

[0053] Step 108: Determine the operating state of each generator node that violates the DC power flow constraint in the first target operating state under each prediction period, and use it as the operating state to be optimized. Then, use a commercial solver to calculate the operating state to be optimized to obtain the second target operating state.

[0054] Step 109: Take the operating state of each generator node in the first target operating state under each prediction period that does not violate the DC power flow constraint and the corresponding second target operating state as the final operating state of each generator node under each prediction period;

[0055] Step 110: Control the corresponding generator nodes in the target power grid using the final operating state of each generator node in each predicted period.

[0056] Using the embodiments in this specification, the load of multiple power-consuming nodes in the target power grid under multiple prediction time periods and the current operating state of multiple generator nodes are first obtained. Then, the coupling constraints between each generator node are relaxed, which can decouple them into solving sub-problems corresponding to each generator node. Solving the sub-problems is to obtain the optimal solution for each generator node. Specifically, the user-defined first dual variable is received, and then a state vector is constructed based on the first dual variable, the DC power flow factor of each generator node, and the current operating state of the machine. The state vector is input into a pre-trained reinforcement learning sub-model for multiple iterations to obtain the predicted operating state for multiple prediction time periods corresponding to the first dual variable.

[0057] Then, the embodiments of this specification use the surrogate subgradient method to solve the optimal solution of the dual problem. That is, the first dual variable is updated according to the predicted operating state of the generator node under the first dual variable for multiple prediction periods and the load of multiple power consumption nodes under the corresponding prediction periods to obtain the second dual variable. The reinforcement learning model is then recalculated using the second dual variable to obtain the predicted operating state corresponding to multiple prediction periods under the second dual variable. After the second dual variable converges, the optimal solution of the dual problem (the converged second dual variable) and the optimal solution of the corresponding relaxation problem (the predicted operating state corresponding to multiple prediction periods) are obtained.

[0058] However, the optimal solution to the relaxation problem may not satisfy all relaxed constraints. Therefore, the embodiments in this specification use a heuristic algorithm to find a feasible solution in the neighborhood, that is, to determine the working state of the generator node that violates the constraints in the optimal solution of the relaxation problem in each prediction period. Then, a commercial solver is used to solve these working states that violate the constraints. The solution results and the working states that do not violate the constraints in the optimal solution of the relaxation problem are used as the final working states. The generator node is controlled using the final working states.

[0059] The method described in this specification, utilizing a surrogate subgradient framework combined with reinforcement learning to solve subproblems, can quickly solve subproblems under different Lagrange multipliers during the iteration process of the surrogate subgradient method, thus greatly accelerating the surrogate subgradient iteration process. Finally, the operating states that violate constraints are identified and solved using commercial solvers, while operating states that do not violate constraints allow for direct control of generator nodes. Compared to directly using commercial solvers to solve all operating states, the method described in this specification can find a near-optimal feasible solution to the generator combination problem in a faster time, improving the computational speed for calculating whether generator nodes should be turned on or off at different times in the generator combination problem.

[0060] In the embodiments of this specification, it is first necessary to use reinforcement learning to train the agent of each generator node, that is, the reinforcement learning model of each generator node. The purpose of the reinforcement learning model is to determine the predicted working state of the generator node in multiple prediction periods based on the current working state of the generator node and a given set of arbitrary dual variables.

[0061] The optimization objective of the UC problem is shown in Equation (1):

[0062] Where, N G p represents the total number of generator nodes. i,t C represents the power of the i-th generator node in the t-th prediction period, where T represents the total number of prediction periods. i (·) represents the fuel cost of the i-th generator node, which indicates how much carbon is emitted when the i-th generator node operates at a specified power, i.e., the carbon emissions, C. SU,i (·) represents the startup cost of the i-th generator node, which indicates the amount of carbon emitted during the process of the i-th generator node transitioning from a shutdown state to a normal operating state; it is also the carbon emission amount. i,t This indicates that the i-th generator node performs a start-up operation or other operation during the t-th prediction period. When performing the start-up operation, u i,t =1, when performing other operations, u i,t =0.

[0063] The constraints are as follows:

[0064] (1) System-level constraints: including power balance constraints and DC power flow constraints.

[0065] The formula for the power balance constraint is (2):

[0066] Where, N G p represents the total number of generator nodes.i,t N represents the power of the i-th generator node in the t-th prediction period. D d represents the total number of power consumption nodes. j,t This represents the load of the j-th power consumption node in the t-th prediction period, where T represents the total number of prediction periods;

[0067] The formula for the DC power flow constraint is (3):

[0068] Among them, F l This indicates the transmission capacity of line l. This represents the DC power flow factor of the i-th generator node with respect to line l. Let L represent the DC power flow factor of the j-th power consumption node with respect to line l, and L represent the total number of lines in the target power grid.

[0069] (2) Individual-level constraints, including power output range constraints, ramp-up constraints, and start-stop time constraints:

[0070] The formula for the power output range constraint is as follows (4):

[0071] Where, x i,t P represents the switching state of the i-th generator node in the t-th prediction period. i This represents the lower limit of the power of the i-th generator node. This represents the upper limit of the power of the i-th generator node.

[0072] The formula for the climbing constraint is as follows (5):

[0073] Among them, -RD i RU represents the maximum power reduction of the i-th generator node. i This represents the maximum increase in power for the i-th generator node.

[0074] The formulas for start-stop time constraints are (6)-(7):

[0075] Among them, v i,t This indicates that the i-th generator node performs a shutdown operation or other operation during the t-th predicted time period. When performing a shutdown operation, v i,t =1, when performing other operations, v i,t =0,TU i TD represents the shortest start-up hold time for the i-th generator node. i This represents the shortest shutdown hold time for the i-th generator node.

[0076] The UC problem is a mixed integer programming problem with (1) as the optimization objective and (2)-(7) as constraints. After Lagrangian relaxation of the system-level constraints (2) and (3), the Lagrangian function (8) can be obtained:

[0077] Where L(Ω, Ф) represents the Lagrangian function obtained after Lagrangian relaxation of the power balance constraint and DC power flow constraint, and Ω represents the decision variables, including x i,t p i,t u i,t v i,t Φ represents the dual variable, including μ. t , λ l,t , κ l,t μ t Let λ represent the Lagrange multiplier of the power balance constraint in the t-th prediction period. l,t κ represents the Lagrange multiplier for the upper limit constraint of DC power flow on line l in the target power grid during the t-th prediction period. l,t It represents the Lagrange multiplier for the lower limit constraint of DC power flow of line l in the target power grid during the t-th prediction period.

[0078] The dual function is (9):

[0079] st(4)-(7)

[0080] The dual problem is (10):

[0081] stλ l,t ≥0,κ l,t ≥0

[0082] The optimal solution to the dual problem is a lower bound of the primal problem. Furthermore, the dual problem is a concave function, which can be found using subgradient methods. In SLR, the surrogate subgradient is obtained by solving for the dual variable Φ after the given k-th iteration update. k Below The result is that, in L(Ω, Ф), the coupling constraints between the generator nodes are relaxed, and the problem can be decoupled into solving subproblems corresponding to each generator node:

[0083] The subproblem for generator node i is:

[0084] st(4),(5),(6),(7)

[0085] This problem is a sequential decision problem, trained using the reinforcement learning method DQN, so that the agent can make decisions on any Lagrange multiplier Ф. k Find the optimal strategy.

[0086] According to one embodiment of this specification, as shown in Figure 2, inputting the state vector into the trained reinforcement learning model corresponding to each generator node for calculation, and obtaining the predicted working state of each generator node under the first dual variable for the multiple prediction periods further includes:

[0087] For each generator node:

[0088] Step 201: Input the state vector of the predicted time period corresponding to the current working state into the trained reinforcement learning model corresponding to the generator node for calculation to obtain the action corresponding to the current predicted time period;

[0089] Step 202: Determine the forecasting status for the next forecasting period based on the actions corresponding to the current forecasting period;

[0090] Step 203: Take the predicted working state corresponding to the next prediction period as the current working state, and repeat the step of constructing the state vector corresponding to the generator node based on the first dual variable, the DC power flow factor of the generator node and its current working state, so as to finally obtain the predicted working state of the generator node under the first dual variable for the multiple prediction periods.

[0091] Specifically, formula (12) is a finite-time Markov decision problem with a duration of T. The state vector in reinforcement learning contains the switching state x of the generator node at that moment. i,t Power size p i,t Cumulative time for powering on or off (y) i,t And the "penalty price" sequence ws brought by Lagrange multipliers i,t The penalty price at time t is w i,t The additional cost at this generator node due to the Lagrange multipliers is defined as follows:

[0092] Among them, the "penalty price" sequence ws i,t The definition is as follows:

[0093] Padding with zeros is to ensure the vector ws i,t The length is the same at different times t, and the number of 0s can also reflect the current time.

[0094] Therefore, the formula for the state vector is: s i,t=(x i,t p i,t y i,t ws i,t ) T (15)

[0095] Among them, s i,t Let x represent the state vector of the i-th generator node in the t-th prediction period. i,t y represents the switching state of the i-th generator node in the t-th prediction period. i,t ws represents the cumulative start-up or shutdown time of the i-th generator node in the t-th prediction period. i,t w represents the penalty factor sequence for the i-th generator node in the t-th prediction period. i,t This represents the penalty factor for the i-th generator node in the t-th prediction period. This represents the Lagrange multiplier of the power balance constraint during the t-th prediction period when the dual variable is updated in the k-th round. This represents the Lagrange multiplier for the upper limit constraint of DC power flow of line l in the target power grid during the t-th prediction period when the dual variable is updated in the k-th round. Let represent the Lagrange multiplier for the DC power flow lower bound constraint of line l in the target power grid during the t-th prediction period when the dual variable is updated in the k-th round. express The transfer to, (x i,t p i,t y i,t ws i,t ) T It means (x) i,t p i,t y i,t ws i,t ) to.

[0096] In reinforcement learning, the action space is defined as the options at the current moment: power on, power off, power increase, power decrease, or power remain unchanged. It can be denoted as a. i,t =(Δx) t , Δp t In practical constraints, the generator node only needs to ensure that the changing power is within upper and lower limits. Here, to simplify the solution of reinforcement learning, the power change is discretized from a continuous interval. The action space also becomes discretized, and the DQN method can be used to train the agent to learn the optimal solution.

[0097] DQN is trained using the temporal difference method, and the loss for each action is: c i,t =C i (p i,t )-CSU,i min{Δx t ,0}+w i,t p i,t (16)

[0098] The optimization objective of this sequential decision problem is to minimize... Where γ is the discount factor, and when γ = 1, it aligns with the optimization objective of the subproblem. Using the temporal difference method to train DQN, the Bellman equation exists:

[0099] Where A represents the set of all actions. Q i Let Q represent the Q value of the i-th generator node. It should be noted that the Q value calculation method in the embodiments of this specification is the same as the mature Q value calculation method in reinforcement learning in this field, and will not be repeated here.

[0100] In DQN, the Q function is approximated using a deep neural network, with parameters θ. i The residual of the Q-function with respect to the i-th generator node in DQN is:

[0101] DQN uses methods such as gradient descent to train the parameters θ. i To minimize the residual, such that It approximates the true Q-function. When solving subproblems, a temporal difference method is used. The state at each time step is input sequentially, and the action with the smallest corresponding Q-value that satisfies the individual-level constraints is selected to obtain the state at the next time step. Repeating this process T times yields the sequential decision for the subproblem. When the Q-function is real, the solution obtained is denoted as the solution to the corresponding optimal subproblem in the discretized action space.

[0102] After training all generator nodes, reinforcement learning methods can be used to quickly solve the subproblems of each generator node given Lagrange multipliers, thereby obtaining... The optimal solution.

[0103] It should be noted that the calculation process of the reinforcement learning model in the embodiments of this specification involves calculating the Q-values ​​corresponding to all possible actions based on the state vector corresponding to the working state of a prediction time period, and then determining the optimal action based on the Q-values, which is the action corresponding to this prediction time period. After obtaining the action corresponding to this prediction time period, the working state of the next prediction time period can be obtained. Then, the action corresponding to the next prediction time period is calculated using the state vector corresponding to the working state of the next prediction time period. After multiple calculations, the predicted working state corresponding to each prediction time period is obtained.

[0104] The reinforcement learning models in the embodiments of this specification can be SAC, PPO, TD3, etc., and this specification does not limit the embodiments. The reinforcement learning models in the embodiments of this specification can be trained using training methods known in the art, which will not be described in detail here.

[0105] This can be understood as follows: the state vector of a reinforcement learning model is constructed based on the dual variable, and then iterative calculations are performed continuously to obtain the prediction working state for all prediction periods corresponding to that dual variable. However, this dual variable may not be optimal, so it is necessary to update and iterate the dual variable. Each time the dual variable is updated, another round of reinforcement learning model calculation is performed to obtain the prediction working state for all prediction periods corresponding to the updated dual variable.

[0106] After selecting the optimal dual variable, the optimal operating state for the prediction period is obtained. However, this calculation is performed after relaxing the constraints, and the optimal operating state of each generator node may not fully satisfy all the constraints. Therefore, the embodiments in this specification include subsequent steps 208-210.

[0107] According to one embodiment of this specification, the updating of the dual variable is first described.

[0108] As shown in Figure 3, the first dual variable is updated based on the predicted operating status of the multiple generator nodes under the first dual variable for the corresponding multiple prediction periods and the load of the multiple power consumption nodes under the corresponding prediction periods, resulting in the second dual variable, which further includes:

[0109] Step 301: Calculate the surrogate subgradient corresponding to the first dual variable based on the predicted working status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding prediction period.

[0110] In this step, the formula for calculating the surrogate subgradient corresponding to the first dual variable based on the predicted operating status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is as follows:

[0111] in, Let represent the surrogate subgradient of the Lagrange multipliers of the power balance constraint during the t-th prediction period when the dual variable is updated in the k-th round. This represents the power of the i-th generator node in the t-th prediction period when the dual variable is updated in the k-th round. Let represent the surrogate subgradient of the Lagrange multiplier with respect to the upper limit constraint of DC power flow of line l during the t-th prediction period when the dual variable is updated in the k-th round; Let represent the surrogate subgradient of the Lagrange multiplier of the DC power flow lower bound constraint of line l during the k-th prediction period when the dual variable is updated.

[0112] In the above definition The "agent optimality condition" should be satisfied, namely: L(Ω) k Φ k )<L(Ω k-1 , Φ k ) (twenty two)

[0113] Step 302: Update the first dual variable according to the surrogate subgradient corresponding to the first dual variable to obtain the second dual variable.

[0114] The formula for updating the first dual variable based on the surrogate subgradient corresponding to the first dual variable is as follows:

[0115] in, This represents the Lagrange multiplier of the power balance constraint during the t-th prediction period when the dual variable is updated in the (k+1)-th round. During the (k+1)th round of dual variable updates, the Lagrange multiplier for the upper limit constraint of DC power flow of line l in the target power grid during the t-th prediction period is... α represents the Lagrange multiplier for the DC power flow lower bound constraint of line l in the target power grid during the (k+1)th round of dual variable updates, in the t-th prediction period. k This represents the step size corresponding to the update of the dual variable in the k-th round;

[0116] in, M≥1, 0<r<1, α k-1 Let g(Ω) represent the step size corresponding to the update of the dual variable in the (k-1)th round. k ) represents the update of the dual variable in the k-th round with respect to L(Ω). k Φ k The surrogate subgradient of ) g(Ω) k-1 ) represents the update of the dual variable in the (k-1)th round with respect to L(Ω). k-1 , Φ k-1 The proxy subgradient;

[0117] Following the method described above, Φ kIt will converge to the dual optimal solution. Among them, Ω satisfies the "surrogate optimality condition" (22). k Reinforcement learning can be used to solve this problem. The "surrogate optimality condition" only requires that the new solution obtained is better than the previous solution, rather than finding Φ. k The optimal solution for L is obtained below. Since the function L is decoupled with respect to each generator node, all generator nodes are pre-divided into U groups, and Ω is solved each time. k Only one group of generator nodes needs to be solved using reinforcement learning, while the states of the other groups remain unchanged. Since the states of the other generator nodes remain unchanged, and all the generator nodes in this group have been solved to approximately optimal solutions, Ω... k The corresponding performance is better than Ω k-1 It has improved and meets the "optimal agent conditions".

[0118] In summary, the processing of agent subgradients in reinforcement learning is as follows:

[0119] Step 1: Initialize the iteration step number k = 0, and divide all generator nodes into U groups. Set the initial Lagrange multiplier Φ. 0 The trained reinforcement learning model is used to solve for Ω0, where:

[0120] Initialize step size in This is an estimate of the optimal solution to the dual problem.

[0121] Step 2: Determine α based on (26) k Update the Lagrange multipliers Φ according to (23)-(25) k .

[0122] Step 3: Update the group for this optimization to the next group after the group updated last time (if the last group is reached, start again from the first group). In the new Φ k The following uses reinforcement learning to calculate the decision variables for all generator nodes within the group. and Update Ω k .

[0123] Step 4: Determine if convergence has occurred; otherwise, return to Step 2.

[0124] The formula for determining whether the second dual variable converges is: ||Φ k -Φ k-1 ||<∈ (28)

[0125] Where, Φ k Let Φ represent the dual variable after the k-th round of updates. k-1 Let represent the dual variable after the (k-1)th round update, and ∈ represent the threshold.

[0126] It should be noted that the proxy subgradient framework in the embodiments of this specification can also be a proxy subgradient framework containing absolute values, and the embodiments of this specification do not impose any limitations.

[0127] According to one embodiment of this specification, after obtaining the first target operating state of each generator node in step 207 during the multiple prediction periods, the first target operating state may not necessarily satisfy all the already relaxed system-level constraints (Equations (2) and (3)). Therefore, this embodiment of the specification uses a heuristic algorithm to find a feasible solution in the neighborhood. This is also the mainstream method for finding feasible solutions after obtaining the dual optimal solution of the UC problem using subgradient methods. This embodiment of the specification fixes most of the decision variables, making their values ​​similar to Ω. k Since the values ​​of corresponding variables are the same, we use the remaining small subset of variables as decision variables to reconstruct and solve the mixed-integer programming model of the UC problem. This significantly reduces the number of variables compared to the original problem, allowing it to be solved in a very short time.

[0128] The method for selecting the working state to be optimized is as follows: First, find all the violated power flow constraints. For example, the constraint of line l at time t is violated, and the constraint is as follows:

[0129] If the upper limit (29) is violated, then choose For generator unit nodes that are positive and have large values, the starting point is... The shutdown generator node with a negative value and a small value is selected as the working state to be optimized. Similarly, if the lower limit (30) is violated, then the node is selected. The generator node that is negative and has a small value is the starting point. The generator set node group with a positive value and a large value is considered as the working state to be optimized.

[0130] Then, a commercial solver (such as Gurobi or IBM Cplex) is used to calculate the working state to be optimized, and the second target working state is obtained.

[0131] Then, the operating state of each generator node that does not violate the DC power flow constraint in the first target operating state under each prediction period and the corresponding second target operating state are taken as the final operating state of each generator node under each prediction period.

[0132] Finally, the final operating state of each generator node in each predicted period is used to control the corresponding generator node in the target power grid.

[0133] The resulting mixed-integer programming problem is much smaller than the original UC problem, and a feasible solution can be obtained quickly. Since it is near the dual optimal solution, it is a near-optimal feasible solution. However, compared to directly solving the original UC problem, this method can improve the solution speed by more than two orders of magnitude, which has significant practical implications.

[0134] The method described in this specification has the following beneficial effects:

[0135] (1) Quickly obtain an approximate optimal feasible solution to the large-scale sub-time interval unit combination problem. The large-scale integration of renewable energy into the grid introduces significant uncertainties to grid dispatch, necessitating a rapid solution to the large-scale sub-time interval unit combination problem; however, solving such problems is extremely slow. The method described in this specification uses a surrogate subgradient method as a framework, leveraging reinforcement learning to solve subproblems in its iterative process. This significantly accelerates the surrogate subgradient method, compressing the time required to solve the large-scale sub-time interval unit combination problem to a practically acceptable range.

[0136] (2) The time complexity of the solution is not affected by the form of the cost function, and the reinforcement learning can be used for life after successful training. When using reinforcement learning to solve the subproblems of the unit, it is only necessary to calculate the Q value of the discrete time intervals, which is not affected by the cost function. In contrast, the difficulty of solving the problem by the traditional branch and bound method is strongly correlated with the form of the cost function. Complex cost functions, such as quadratic functions, will greatly increase the difficulty of solving the problem.

[0137] Therefore, the reinforcement learning method combining surrogate subgradients in the embodiments of this specification can be applied to other large-scale control problems with similar characteristics (e.g., decomposing the problem into decoupled subproblems through Lagrange relaxation, and these subproblems are sequential decision problems). Since this method does not require repeatedly solving similar subproblems under different Lagrange multipliers, it greatly improves the solution efficiency of this type of problem.

[0138] Based on the same inventive concept, embodiments of this specification also provide a control device for a generator set, as shown in Figure 4, including:

[0139] The data acquisition unit 401 is used to acquire the load of multiple power consumption nodes in the target power grid under multiple predicted time periods and the current working status of multiple generator set nodes in the target power grid.

[0140] The first dual variable receiving unit 402 is used to receive the set first dual variable, which includes the power balance constraint and the Lagrange multiplier under the DC power flow constraint of the target power grid.

[0141] The working state prediction unit 403 is used to construct a state vector corresponding to each generator node based on the first dual variable, the DC power flow factor of each generator node and its current working state, and input the state vector into the trained reinforcement learning model corresponding to each generator node for calculation to obtain the predicted working state of each generator node under the first dual variable in the multiple prediction periods.

[0142] The dual variable update unit 404 is used to update the first dual variable according to the predicted working state of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period, to obtain the second dual variable, and to perform the step of constructing the state vector corresponding to each generator node according to the second dual variable, the DC power flow factor of each generator node and its current working state, to obtain the predicted working state of each generator node under the second dual variable for the multiple predicted periods.

[0143] The dual variable convergence judgment unit 405 is used to determine whether the second dual variable has converged; if it has not converged, the second dual variable is used as the first dual variable, and the step of updating the first dual variable according to the predicted working state of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is repeated; if it has converged, the predicted working state of each generator node under the second dual variable in each predicted period is used as the first target working state of each generator node in the multiple predicted periods.

[0144] The working state optimization unit 406 is used to determine the working state of each generator node in the first target working state under each prediction period that violates the DC power flow constraint, as the working state to be optimized, and to calculate the working state to be optimized using a commercial solver to obtain the second target working state.

[0145] The final operating state determination unit 407 is used to take the operating state of each generator node in the first target operating state under each prediction period, which does not violate the DC power flow constraint, and the corresponding second target operating state as the final operating state of each generator node under each prediction period.

[0146] Control unit 408 is used to control the corresponding generator node in the target power grid using the final operating state of each generator node in each predicted period.

[0147] Since the principle of the above-mentioned device in solving the problem is similar to that of the above-mentioned method, the implementation of the above-mentioned system can refer to the implementation of the above-mentioned method, and the repeated parts will not be described again.

[0148] Figure 5 shows a schematic diagram of the structure of a computer device according to an embodiment of this specification. The device in the embodiment of this specification can be the computer device in this embodiment, which executes the method described in the embodiment of this specification.

[0149] Computer device 502 may include one or more processing devices 504, such as one or more central processing units (CPUs), each of which may implement one or more hardware threads. Computer device 502 may also include any storage resource 506 for storing information of any kind, such as code, settings, data, etc. Non-limitingly, for example, storage resource 506 may include any type of RAM, any type of ROM, flash memory devices, hard disks, optical disks, etc. More generally, any storage resource can use any technology to store information.

[0150] Furthermore, any storage resource can provide volatile or non-volatile retention of information.

[0151] Furthermore, any storage resource can represent a fixed or removable component of the computer device 502. In one case, when the processing device 504 executes associated instructions stored in any storage resource or combination of storage resources, the computer device 502 can perform any operation of the associated instructions. The computer device 502 also includes one or more drive systems 508 for interacting with any storage resource, such as a hard disk drive system, an optical disk drive system, etc.

[0152] Computer device 502 may also include an input / output module 510 (I / O) for receiving various inputs (via input device 512) and providing various outputs (via output device 514). A specific output mechanism may include a presentation device 516 and an associated graphical user interface (GUI) 518. In other embodiments, the input / output module 510 (I / O), input device 512, and output device 514 may be omitted, and the device may function solely as a computer device within a network. Computer device 502 may also include one or more network interfaces 520 for exchanging data with other devices via one or more communication links 522. One or more communication buses 524 couple the components described above together.

[0153] Communication link 522 can be implemented in any way, such as via a local area network, a wide area network (e.g., the Internet), a point-to-point connection, or any combination thereof. Communication link 522 may include any combination of hardwired links, wireless links, routers, gateway functions, name servers, etc., governed by any protocol or combination of protocols.

[0154] It should be noted that when the Kubernetes platform in this embodiment implements the methods described in this embodiment for the computer device 502, it may not include the presentation device 516 and the associated graphical user interface (GUI) 518. For example, it may only include a minimal computer system consisting of a processing device 504, storage resources 506, and a network interface 520.

[0155] This specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0156] This specification also provides computer-readable instructions, wherein when a processor executes the instructions, the program therein causes the processor to perform the above-described method.

[0157] It should be understood that, in various embodiments of the present invention, the order of the above-mentioned process numbers does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

Claims

1. A control method for a generator set, characterized in that, The method includes: Obtain the load of multiple power consumption nodes in the target power grid under multiple forecast periods and the current operating status of multiple generator nodes in the target power grid; Receive the set first dual variable, which includes the Lagrange multipliers under the power balance constraint and DC power flow constraint of the target power grid; Based on the first dual variable, the DC power flow factor of each generator node and its current operating state, a state vector corresponding to each generator node is constructed. The state vector is then input into the trained reinforcement learning model corresponding to each generator node for calculation to obtain the predicted operating state of each generator node under the first dual variable for the multiple prediction periods. The first dual variable is updated based on the predicted operating status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period to obtain the second dual variable. The step of constructing the state vector corresponding to each generator node based on the second dual variable, the DC power flow factor of each generator node and its current operating status is performed to obtain the predicted operating status of each generator node under the second dual variable for the multiple predicted periods. Determine whether the second dual variable converges; If convergence is not achieved, the second dual variable is used as the first dual variable, and the step of updating the first dual variable based on the predicted working status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is repeated. If convergence is achieved, the predicted working state of each generator node in each prediction period under the second dual variable will be taken as the first target working state of each generator node in the multiple prediction periods. The operating state of each generator node that violates the DC power flow constraint in the first target operating state under each prediction period is determined as the operating state to be optimized, and the second target operating state is obtained by calculating the operating state to be optimized using a commercial solver. The operating state of each generator node in each prediction period that does not violate the DC power flow constraint in the first target operating state and the corresponding second target operating state are taken as the final operating state of each generator node in each prediction period. The corresponding generator nodes in the target power grid are controlled by utilizing the final operating state of each generator node in each predicted period.

2. The method according to claim 1, characterized in that, The state vector is input into the trained reinforcement learning model corresponding to each generator node for calculation, and the predicted working state of each generator node under the first dual variable for the multiple prediction periods further includes: For each generator node: The state vector of the predicted time period corresponding to the current working state is input into the trained reinforcement learning model corresponding to the generator node for calculation to obtain the action corresponding to the current predicted time period. Determine the forecasting status for the next forecasting period based on the actions corresponding to the current forecasting period. The predicted operating state corresponding to the next prediction period is taken as the current operating state. The step of constructing the state vector corresponding to the generator node based on the first dual variable, the DC power flow factor of the generator node and its current operating state is repeated to finally obtain the predicted operating state of the generator node under the first dual variable for the multiple prediction periods.

3. The method according to claim 1, characterized in that, The formula for the power balance constraint is: Where, N G p represents the total number of generator nodes. i,t N represents the power of the i-th generator node in the t-th prediction period. D d represents the total number of power consumption nodes. j,t This represents the load of the j-th power consumption node in the t-th prediction period, where T represents the total number of prediction periods; The formula for the DC power flow constraint is: Among them, F l This indicates the transmission capacity of line l. This represents the DC power flow factor of the i-th generator node with respect to line l. Let L represent the DC power flow factor of the j-th power consumption node with respect to line l, and L represent the total number of lines in the target power grid.

4. The method according to claim 3, characterized in that, The formula for constructing the state vector corresponding to each generator node based on the first dual variable, the DC power flow factor of each generator node, and its current operating state is as follows: s i,t =(x i,t p i,t y i,t ws i,t ); Among them, s i,t Let x represent the state vector of the i-th generator node in the t-th prediction period. i,t y represents the switching state of the i-th generator node in the t-th prediction period. i,t ws represents the cumulative start-up or shutdown time of the i-th generator node in the t-th prediction period. i,t w represents the penalty factor sequence for the i-th generator node in the t-th prediction period. i,t This represents the penalty factor for the i-th generator node in the t-th prediction period. This represents the Lagrange multiplier of the power balance constraint during the t-th prediction period when the dual variable is updated in the k-th round. This represents the Lagrange multiplier for the upper limit constraint of DC power flow of line l in the target power grid during the t-th prediction period when the dual variable is updated in the k-th round. Let represent the Lagrange multiplier for the DC power flow lower bound constraint of line l in the target power grid during the t-th prediction period when the dual variable is updated in the k-th round.

5. The method according to claim 4, characterized in that, The first dual variable is updated based on the predicted operating status of the multiple generator nodes under the first dual variable for the corresponding multiple prediction periods and the load of the multiple power consumption nodes under the corresponding prediction periods, to obtain the second dual variable, which further includes: The surrogate subgradient corresponding to the first dual variable is calculated based on the predicted working status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding prediction period. The first dual variable is updated based on the surrogate subgradient corresponding to the first dual variable to obtain the second dual variable.

6. The method according to claim 5, characterized in that, The formula for calculating the surrogate subgradient corresponding to the first dual variable based on the predicted operating status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is as follows: in, Let represent the surrogate subgradient of the Lagrange multipliers of the power balance constraint during the t-th prediction period when the dual variable is updated in the k-th round. This represents the power of the i-th generator node in the t-th prediction period when the dual variable is updated in the k-th round. Let represent the surrogate subgradient of the Lagrange multiplier with respect to the upper limit constraint of DC power flow of line l during the t-th prediction period when the dual variable is updated in the k-th round; Indicates the k-th round When updating the dual variables, the surrogate subgradient of the Lagrange multipliers of the DC power flow lower bound constraint of line l is used in the t-th prediction period.

7. The method according to claim 6, characterized in that, The formula for updating the first dual variable based on the surrogate subgradient corresponding to the first dual variable is as follows: in, This represents the Lagrange multiplier of the power balance constraint during the t-th prediction period when the dual variable is updated in the (k+1)-th round. During the (k+1)th round of dual variable updates, the Lagrange multiplier for the upper limit constraint of DC power flow of line l in the target power grid during the t-th prediction period is... α represents the Lagrange multiplier for the DC power flow lower bound constraint of line l in the target power grid during the (k+1)th round of dual variable updates, in the t-th prediction period. k This represents the step size corresponding to the update of the dual variable in the k-th round; in, α k-1 Let g(Ω) represent the step size corresponding to the update of the dual variable in the (k-1)th round. k ) represents the update of the dual variable in the k-th round with respect to L(Ω). k Ф k The surrogate subgradient of ) g(Ω) k-1 ) represents the update of the dual variable in the (k-1)th round with respect to L(Ω). k-1 Ф k-1 The proxy subgradient; L(Ω, Ф) represents the Lagrangian function obtained after Lagrangian relaxation of the power balance constraint and DC power flow constraint, and Ω represents the decision variables, including x. i,t p i,t u i,t v i,t , where u i,t This indicates that the i-th generator node performs a start-up operation or other operation during the t-th prediction period. When performing the start-up operation, u i,t =1, when performing other operations, u i,t =0, v i,t This indicates that the i-th generator node performs a shutdown operation or other operation during the t-th predicted time period. When performing a shutdown operation, v i,t =1, when performing other operations, v i,t =0; Φ represents the dual variable, including μ t , λ l,t , κ l,t .

8. The method according to claim 7, characterized in that, The formula for determining whether the second dual variable converges is: ||Φ k -Φ k-1 ||<∈; Where, Φ k Let Φ represent the dual variable after the k-th round of updates. k-1 Let represent the dual variable after the (k-1)th round update, and ∈ represent the threshold.

9. A control device for a generator set, characterized in that, include: The data acquisition unit is used to acquire the load of multiple power consumption nodes in the target power grid under multiple forecast periods and the current working status of multiple generator set nodes in the target power grid. The first dual variable receiving unit is used to receive the set first dual variable, which includes the Lagrange multipliers under the power balance constraint and DC power flow constraint of the target power grid. The working state prediction unit is used to construct a state vector corresponding to each generator node based on the first dual variable, the DC power flow factor of each generator node and its current working state, and input the state vector into the trained reinforcement learning model corresponding to each generator node for calculation to obtain the predicted working state of each generator node under the first dual variable in the multiple prediction periods. The dual variable update unit is used to update the first dual variable according to the predicted operating status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period, to obtain the second dual variable, and to perform the step of constructing the state vector corresponding to each generator node according to the second dual variable, the DC power flow factor of each generator node and its current operating status, to obtain the predicted operating status of each generator node under the second dual variable for the multiple predicted periods. The dual variable convergence judgment unit is used to determine whether the second dual variable has converged; if it has not converged, the second dual variable is used as the first dual variable, and the step of updating the first dual variable according to the predicted working status of the multiple generator nodes under the first dual variable and the load of the multiple power consumption nodes under the corresponding predicted period is repeated. If convergence is achieved, the predicted working state of each generator node in each prediction period under the second dual variable will be taken as the first target working state of each generator node in the multiple prediction periods. The working state optimization unit is used to determine the working state of each generator node in the first target working state under each prediction period that violates the DC power flow constraint, as the working state to be optimized, and to calculate the working state to be optimized using a commercial solver to obtain the second target working state. The final operating state determination unit is used to take the operating state of each generator node in the first target operating state under each prediction period, which does not violate the DC power flow constraint, and the corresponding second target operating state as the final operating state of each generator node under each prediction period. A control unit is used to control the corresponding generator nodes in the target power grid using the final operating state of each generator node in each predicted period.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Frequency control method and device based on Lagrange relaxation reinforcement learning

    CN115842380A

  • Optimizer and optimization method for real-time economic dispatch problem in micro-grid

    KR102582036B1

  • Distributed renewable energy grid controller

    US10326280B1

  • Method and device for identifying feasibility of transmission interface constraint in online rolling dispatching

    US20150088470A1

  • Security region based security-constrained economic dispatching method

    US20150310366A1