A security-reinforced learning load frequency control method, system, device and medium
By constructing a constrained Markov decision process and introducing a Lagrange multiplier optimization actor network, the problem of high computational resource consumption in existing methods is solved, achieving efficient load frequency control and improving the operating efficiency and frequency regulation performance of the power system.
Patent Information
- Application Number
- CN202411624478.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-14
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2044-11-14
AI Technical Summary
Existing load frequency control methods based on safety reinforcement learning require a large amount of computational resources to handle system constraints, resulting in the final control strategy failing to demonstrate good performance in practical applications.
Combining the frequency regulation characteristics and power constraints of flywheel energy storage systems, a frequency control model constrained by Markov decision process is constructed. Using long short-term memory networks and deep reinforcement learning controllers, the actor network is optimized through inequality constraints that transform the objective function using Lagrange multipliers to achieve real-time load frequency control.
It improves the accuracy of cost forecasting and the overall efficiency of the power system, ensures the auxiliary role of flywheel energy storage system in load frequency control, and takes into account both power system operating costs and frequency regulation performance.
Smart Images

Figure CN119602301B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of energy storage system frequency control, in particular to a safe reinforcement learning load frequency control method, system, device and medium. BACKGROUND
[0002] The frequency of a power system is a key indicator of power quality, and the imbalance of active power will cause frequency fluctuations. Load frequency control (LFC) is an effective means to cope with frequency deviation of a power system, which maintains the balance of active power between power generation and load by adjusting the power output at the power generation end, thereby maintaining the stability of the system frequency. Energy storage systems have become an important technical means for frequency control in modern power systems due to their fast response, precise control and ability to realize bidirectional regulation. Flywheel energy storage systems, as a power and physical energy storage device, have advantages such as environmental protection, fast response, and unlimited number of charge and discharge times, and have received widespread attention in load frequency control.
[0003] Traditional model-based load frequency control relies on accurate system models, which is often difficult to meet the requirements in complex power system scenarios, limiting its application range. In recent years, with the help of the fast computing power of neural networks, reinforcement learning-based load frequency control methods have gradually emerged, which can generate frequency regulation strategies for power systems by training neural networks. However, existing methods often ignore the frequency modulation characteristics and power constraints of flywheel energy storage systems, and fail to fully utilize the auxiliary frequency modulation capabilities of energy storage systems. In addition, when considering both power system operation cost and frequency modulation performance, existing methods often design reward functions by multiplying each target function value with a weight, and the determination of the weight often depends on simulation tests, which may affect the accuracy of the optimization process.
[0004] In addition, some load frequency control methods based on safe reinforcement learning require a large amount of computing resources when dealing with system constraints, resulting in control strategies that are difficult to exhibit good results in actual applications. Therefore, it is of great significance to further study safe reinforcement learning load frequency control methods that take into account the characteristics of flywheel energy storage, power constraints, and the economic efficiency and frequency modulation performance of power systems. SUMMARY
[0005] The embodiments of the present application provide a safe reinforcement learning load frequency control method, system, device and medium to solve the problem that some load frequency control methods based on safe reinforcement learning in the prior art require a large amount of computing resources when dealing with system constraints, resulting in control strategies that are difficult to exhibit good results in actual applications.
[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or to describe the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.
[0007] According to a first aspect of the present invention, a method for controlling the frequency of a secure reinforcement learning load is provided.
[0008] In one embodiment, the secure reinforcement learning load frequency control method includes:
[0009] Combining the frequency regulation characteristics and power constraints of flywheel energy storage systems, and constructing a frequency control model based on the core elements of constrained Markov decision processes;
[0010] Historical power system data is acquired, and a labeled dataset is constructed based on the cost function of the core elements. A long short-term memory network is trained based on the labeled dataset, and a cost prediction network is constructed based on the trained long short-term memory network. The state and actions of the frequency control model are used as inputs to the cost prediction network to predict the control cost at the current moment.
[0011] An upper limit value is set for the cost function, a deep reinforcement learning controller is constructed, and the inequality constraints of the objective function in the deep reinforcement learning controller are transformed based on Lagrange multipliers. According to the transformation result, the actor network of the deep reinforcement learning controller is optimized by predicting the control cost at the current moment, and the optimized actor network is applied to real-time load frequency control to achieve control of the load frequency.
[0012] In one embodiment, the construction of a frequency control model based on the frequency regulation characteristics and power constraints of the flywheel energy storage system and the core elements of a constrained Markov decision process includes:
[0013] Determine the generator generation rate constraints and power increment constraints of the diesel generator;
[0014] Based on the generator rate constraints and power increment constraints of diesel generators, and combined with the adjustment dead zone and power increment constraints of flywheel energy storage systems, the state space, action space, reward function, cost function and objective function of constrained Markov decision process are defined.
[0015] A frequency control model is constructed based on the state space, action space, reward function, cost function, and objective function of a constrained Markov decision process.
[0016] In one embodiment, the expression for the frequency control model is:
[0017]
[0018] In the formula, Δf is the power system frequency deviation; H is the power system inertial time constant; ΔP G The deviation in output power of the diesel generator; ΔP E The output power deviation of the flywheel energy storage system; ΔP W For wind power deviation; ΔP L ΔX represents the load power deviation; D represents the power system load damping coefficient; ΔX G For the speed controller output signal deviation; T t Δu is the governor constant; G For diesel generator control signals; T f R is the time constant of the diesel generator; f Δu is the droop coefficient of the diesel generator. E For flywheel energy storage system control signals; T E is the time constant of the flywheel energy storage system.
[0019] In one embodiment, the acquisition of historical power system data involves constructing a labeled dataset based on the cost function of the core elements; training a Long Short-Term Memory (LSTM) network using the labeled dataset; constructing a cost prediction network based on the trained LSM network; and using the state and actions of the frequency control model as input to the cost prediction network to predict the control cost at the current moment, including:
[0020] A labeled dataset is constructed using historical data, load frequency control is implemented through a PID controller, and the state and actions of the frequency control model are used as training data based on the cost function of a constrained Markov decision process.
[0021] Using the constructed labeled dataset, a pre-configured long short-term memory network model is trained;
[0022] Data pairs for constructing and training a long short-term memory network model were used to utilize the dynamic response data of a flywheel energy storage system.
[0023] A cost prediction network is constructed based on data pairs from the trained Long Short-Term Memory network.
[0024] The cost prediction network receives the labels of the training data at each time step and updates the current state through hidden states and memory units.
[0025] After obtaining the updated hidden state, the output of the cost prediction network is mapped to the predicted cost through a fully connected layer;
[0026] Supervised learning is used to establish a mapping relationship from state and action to cost, thereby completing the training of the cost prediction network.
[0027] In one embodiment, mapping the output of the cost prediction network to the predicted cost through a fully connected layer includes:
[0028] Q C (s t ,a t ) = h (n) [h (n-1) [...h (1) (h t )]];
[0029] In the formula, Q C (s t ,a t ) represents the output of the cost prediction network; h (n) h is the activation function for the nth layer; (n-1) h is the activation function for the (n-1)th layer. (1) The activation function for layer 1; (h t ) represents the hidden state output by the Long Short-Term Memory network.
[0030] In one embodiment, setting an upper limit for the cost function, constructing a deep reinforcement learning controller, and transforming the inequality constraints of the objective function in the deep reinforcement learning controller based on Lagrange multipliers; according to the transformation result, optimizing the actor network of the deep reinforcement learning controller using the predicted control cost at the current moment, and applying the optimized actor network to real-time load frequency control to achieve load frequency control includes:
[0031] Set an upper limit for the cost function;
[0032] Construct a deep reinforcement learning controller comprising an actor network, a reward critic network, an actor-goal network, and a reward critic-goal network;
[0033] Initialize the network parameters of the actor network and the reward critic network, and then assign the initialized network parameters of the actor network and the reward critic network to the actor target network and the reward critic network, respectively;
[0034] Initialize the interaction experience pool and dual variables of the Lagrange multipliers, obtain several interaction experiences between the actor network and the power system by predicting the control cost at the current moment, and construct the interaction experience pool based on these interaction experiences.
[0035] Optimize the network parameters of the actor network, rewarded commentator network, target actor network, and target rewarded commentator network based on the interactive experience pool;
[0036] Repeat the steps of building an interactive experience pool and optimizing using the interactive experience pool up to a preset number of times, and output the optimized actor network;
[0037] The optimized actor network is applied to real-time load frequency control to achieve control over the load frequency.
[0038] In one embodiment, optimizing the network parameters of the actor network, the rewarded critic network, the target actor network, and the target rewarded critic network based on the interactive experience pool includes:
[0039] Sample a preset number of interactive experiences from the interactive experience pool;
[0040] Calculate the target value of rewarding commentators based on the network parameters of the target actor network and the target reward commentator network.
[0041] Based on the reward critic objective value, the network parameters of the reward critic network are updated by minimizing the loss function;
[0042] Calculate the action gradient and update the network parameters of the actor network based on the action gradient using gradient descent.
[0043] The reward system for commentators is optimized by combining the chain rule with predictions of control costs at the current moment.
[0044] Calculate the gradient of the dual variable, and update the dual variable using gradient descent based on the gradient of the dual variable;
[0045] Optimize the target actor network and the target reward commentator network.
[0046] In one embodiment, the expression for calculating the gradient of the dual variable is:
[0047]
[0048] In the formula, d λ The gradient of the dual variable; K is the number of experiences sampled from the interactive experience pool; k is the experience number; Q C For cost prediction networks; s k For the state in experience k; μ(s) k |θ μ (This refers to an actor network;) d represents the parameters of the cost prediction network; d represents the constraint limit value.
[0049] According to a second aspect of the present invention, a secure reinforcement learning load frequency control system is provided.
[0050] In one embodiment, the secure reinforcement learning load frequency control system includes:
[0051] The frequency control model construction module is used to combine the frequency regulation characteristics and power constraints of the flywheel energy storage system, and to construct a frequency control model based on the core elements of the constrained Markov decision process.
[0052] The cost prediction network construction module is used to acquire historical data of the power system and construct a labeled dataset based on the cost function of the core elements; a long short-term memory network is trained according to the labeled dataset, and a cost prediction network is constructed based on the trained long short-term memory network. The state and actions of the frequency control model are used as the input of the cost prediction network to predict the control cost at the current moment.
[0053] The deep reinforcement learning controller optimization module is used to set the upper limit of the cost function, construct the deep reinforcement learning controller, and transform the inequality constraints of the objective function in the deep reinforcement learning controller based on Lagrange multipliers; according to the transformation result, the actor network of the deep reinforcement learning controller is optimized by predicting the control cost at the current moment, and the optimized actor network is applied to real-time load frequency control to achieve control of the load frequency.
[0054] In one embodiment, the frequency control model construction module includes: a diesel generator constraint determination module, a constraint Markov decision process element definition module, and a control model construction module;
[0055] The constraint determination module for diesel generators is used to determine the generator generation rate constraints and power increment constraints of diesel generators.
[0056] The Constrained Markov Decision Process (CDM) element definition module is used to define the state space, action space, reward function, cost function, and objective function of the constrained Markov decision process based on the generator rate constraint and power increment constraint of the diesel generator, combined with the adjustment dead zone and power increment constraint of the flywheel energy storage system.
[0057] The control model construction module is used to construct a frequency control model based on the state space, action space, reward function, cost function, and objective function of a constrained Markov decision process.
[0058] In one embodiment, the cost prediction network construction module includes: a labeled dataset module, a long short-term memory network model configuration module, a flywheel energy storage system data pair construction module, a prediction network construction module, an output mapping module, a cost prediction module, and a supervised learning training module;
[0059] The labeled dataset module is used to construct a labeled dataset using historical data, implement load frequency control through a PID controller, and use the state and actions of the frequency control model as training data based on the cost function of a constrained Markov decision process.
[0060] The Long Short-Term Memory Network Model Configuration Module is used to train a pre-configured Long Short-Term Memory Network Model using a constructed labeled dataset.
[0061] The flywheel energy storage system data pair construction module is used to construct and train a long short-term memory network model using the dynamic response data of the flywheel energy storage system.
[0062] The prediction network building module is used to construct a cost prediction network based on data pairs trained on the Long Short-Term Memory network.
[0063] The output mapping module is used to receive the labels of the training data at each time step through the cost prediction network and update the current state through the hidden state and memory unit.
[0064] The cost prediction module is used to map the output of the cost prediction network to the predicted cost through a fully connected layer after obtaining the updated hidden state.
[0065] The supervised learning training module is used to establish a mapping relationship from state and action to cost through supervised learning, thereby completing the training of the cost prediction network.
[0066] In one embodiment, the deep reinforcement learning controller optimization module includes: a cost function upper limit setting module, a deep reinforcement learning controller construction module, a network parameter initialization module, an interactive experience pool construction module, a network parameter optimization module, an optimization process iteration module, and a real-time load frequency control module.
[0067] The upper limit setting module for the cost function is used to set the upper limit value of the cost function;
[0068] The Deep Reinforcement Learning Controller Building Module is used to build a deep reinforcement learning controller that includes an actor network, a reward critic network, an actor target network, and a reward critic target network.
[0069] The network parameter initialization module is used to initialize the network parameters of the actor network and the reward critic network, and then assign the initialized network parameters of the actor network and the reward critic network to the actor target network and the reward critic network, respectively.
[0070] The interaction experience pool construction module is used to initialize the interaction experience pool and dual variables of the Lagrange multiplier. It obtains several interaction experiences between the actor network and the power system by predicting the control cost at the current moment, and constructs the interaction experience pool based on these interaction experiences.
[0071] The network parameter optimization module is used to optimize the network parameters of the actor network, the rewarded commentator network, the target actor network, and the target rewarded commentator network based on the interactive experience pool.
[0072] The optimization process iteration module is used to repeatedly execute the steps of building the interaction experience pool and optimizing using the interaction experience pool up to a preset number of times, and output the optimized actor network.
[0073] The real-time load frequency control module is used to apply the optimized actor network to real-time load frequency control, thereby controlling the load frequency.
[0074] In one embodiment, the network parameter optimization module includes: an interaction experience acquisition module, a reward commentator target value calculation module, a reward commentator network network parameter update module, an action gradient and update actor network network parameter calculation module, a reward commentator network optimization module, a dual variable update module, and a target actor network and target reward commentator network optimization module;
[0075] The interactive experience collection module is used to sample a preset number of interactive experiences from the interactive experience pool;
[0076] The reward commentator target value calculation module is used to calculate the reward commentator target value based on the target actor network and the network parameters of the target reward commentators.
[0077] The network parameter update module for the rewarded commentator network is used to update the network parameters of the rewarded commentator network by minimizing the loss function based on the rewarded commentator target value.
[0078] The action gradient and network parameter update module is used to calculate the action gradient and update the network parameters of the actor network based on the action gradient using the gradient descent method.
[0079] The rewarded commentator network optimization module is used to optimize the rewarded commentator network based on the chain rule and the prediction of the control cost at the current moment.
[0080] The dual variable update module is used to calculate the gradient of the dual variable and update the dual variable using the gradient descent method based on the gradient of the dual variable.
[0081] The Target Actor Network and Target Reward Commentator Network Optimization Module is used to optimize the Target Actor Network and Target Reward Commentator Network.
[0082] In one embodiment, the expression for calculating the gradient of the dual variable is:
[0083]
[0084] In the formula, d λ The gradient of the dual variable; K is the number of experiences sampled from the interactive experience pool; k is the experience number; Q C For cost prediction networks; s k For the state in experience k; μ(s) k |θ μ (This refers to an actor network;) d represents the parameters of the cost prediction network; d represents the constraint limit value.
[0085] According to a third aspect of the present invention, a computer device is provided.
[0086] In one embodiment, the computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method described above.
[0087] According to a fourth aspect of the present invention, a computer-readable storage medium is provided.
[0088] In one embodiment, a computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the above method.
[0089] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0090] This invention presents a safety-enhancing learning-based load frequency control method for flywheel energy storage systems. It fully considers the frequency regulation characteristics and power constraints of flywheel energy storage systems, constructing a constrained Markov decision process (CMP) with state space, action space, reward function, cost function, and objective function. To achieve a nonlinear mapping from state-action pairs to costs, a long short-term memory (LSTM) network is used, with the states and actions of the frequency control model as inputs to a cost prediction network to predict the control cost at the current moment. This method captures the complex temporal dependencies between states and actions during system operation, effectively improving the accuracy of cost prediction. Finally, Lagrange multipliers are introduced to transform the inequality constraint problem of the objective function into an unconstrained problem. A cost prediction neural network is then used to improve the efficiency of solving the Primal Dual-DDPG constrained Markov decision process, updating the actor network, and using the optimized actor network as the real-time load frequency control strategy model. This invention fully considers the auxiliary role and frequency regulation characteristics of flywheel energy storage systems in the load frequency control process. Compared to general Markov processes, this invention explicitly considers constraints while simultaneously taking into account power system operating costs and frequency regulation performance, ensuring the comprehensive benefits of the power system.
[0091] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0092] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0093] Figure 1 This is a flowchart illustrating a secure reinforcement learning load frequency control method according to an exemplary embodiment;
[0094] Figure 2 This is a schematic diagram of a security reinforcement learning load frequency control system according to an exemplary embodiment.
[0095] Figure 3 This is a schematic diagram of the structure of a computer device according to an exemplary embodiment;
[0096] Figure 4 This is a block diagram illustrating the principle of the frequency control model in a secure reinforcement learning load frequency control method according to an exemplary embodiment;
[0097] Figure 5 This is a schematic diagram illustrating the update of the reward commentator network and actor network in a secure reinforcement learning load frequency control method according to an exemplary embodiment. Detailed Implementation
[0098] The following description and accompanying drawings fully illustrate specific embodiments described herein to enable those skilled in the art to practice them. Some embodiments may include or substitute parts and features of other embodiments. The scope of the embodiments herein encompasses the entire scope of the claims and all available equivalents thereof. Throughout this document, the terms “first,” “second,” etc., are used only to distinguish one element from another without requiring or implying any actual relationship or order between the elements. Indeed, a first element can also be referred to as a second element, and vice versa. Furthermore, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a structure, apparatus, or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a structure, apparatus, or device. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the structure, apparatus, or device that includes said element. The various embodiments described herein are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.
[0099] The terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer" used in this document to indicate orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings. They are used solely for the convenience of describing the document and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description herein, unless otherwise specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two elements; they can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0100] In this document, unless otherwise stated, the term "multiple" means two or more.
[0101] In this article, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0102] In this article, the term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0103] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0104] The modules in the apparatus or system of this application can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0105] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0106] Figure 1An embodiment of a secure reinforcement learning load frequency control method of the present invention is shown.
[0107] In this optional embodiment, the secure reinforcement learning load frequency control method includes:
[0108] Step S101: Combine the frequency regulation characteristics and power constraints of the flywheel energy storage system, and construct a frequency control model based on the core elements of the constrained Markov decision process.
[0109] Step S103: Obtain historical data of the power system, construct a labeled dataset based on the cost function of the core elements; train a long short-term memory network according to the labeled dataset, construct a cost prediction network based on the trained long short-term memory network, and use the state and action of the frequency control model as the input of the cost prediction network to predict the control cost at the current moment.
[0110] Step S105: Set the upper limit of the cost function, construct a deep reinforcement learning controller, and transform the inequality constraints of the objective function in the deep reinforcement learning controller based on Lagrange multipliers; according to the transformation result, optimize the actor network of the deep reinforcement learning controller by predicting the control cost at the current moment, and apply the optimized actor network to real-time load frequency control to achieve load frequency control.
[0111] In this optional embodiment, the step of combining the frequency regulation characteristics and power constraints of the flywheel energy storage system, and constructing a frequency control model based on the core elements of a constrained Markov decision process, includes:
[0112] Determine the generator generation rate constraints and power increment constraints of the diesel generator;
[0113] Based on the generator rate constraints and power increment constraints of diesel generators, and combined with the adjustment dead zone and power increment constraints of flywheel energy storage systems, the state space, action space, reward function, cost function and objective function of constrained Markov decision process are defined.
[0114] Furthermore, it should be noted that the state space of a constrained Markov decision process is defined as follows: Define the action space of a constrained Markov decision process as follows: The reward function of a constrained Markov decision process is defined as follows:
[0115]
[0116] Define the cost function c of a constrained Markov decision process. t for:
[0117]
[0118] Define the objective function π* of the constrained Markov decision process as follows:
[0119]
[0120] In the formula: s t Let Δf be the agent's state at time t; t Let Δf be the frequency deviation of the power system at time t. t 2 c is the square of the power system frequency deviation at time t; t It is the cost function; To account for control signal differences and reduce control losses; For the function that enables the diesel generator's power generation rate to exceed the limit; For the flywheel energy storage system, the power generation rate exceeds the limit function; This is the function for exceeding the limit of incremental change in power generation; The control signal for the diesel generator at time t; This is the control signal for the diesel generator at time t-1; The control signal for the flywheel energy storage system at time t; This is the control signal for the flywheel energy storage system at time t-1; Let ΔP be the output power of the diesel generator at time t; Gmax ΔP is the upper limit of the diesel generator power increment. Gmin This represents the lower limit of the diesel generator power increment; ΔP Emin This represents the lower limit of the power increment of the flywheel energy storage system. Let ΔP be the output power of the flywheel energy storage system at time t. Emax This represents the upper limit of the power increment of the flywheel energy storage system; δ represents the incremental change in power generation; δ is the threshold for the diesel generator power ramp-up rate; π* is the optimal control trajectory; R(π) is the future discount reward. Let r be the expectation function; γ is the discount factor, γ≤1; i The reward obtained by the agent; a t Let C(π) be the agent's action at time t; C(π) be the future discount cost, and c i d is the cost function; α is the constraint limit value; α is the positive reward value. and These represent the penalty coefficients corresponding to different frequency deviation ranges; β is the penalty value, where the agent will be severely penalized when the frequency deviation exceeds 0.2Hz. To account for control signal differences and reduce control losses; and These are the over-limit functions for the incremental changes in the generator power generation rate and the generator output power of the diesel generator, respectively. This is the over-limit function for the power generation rate of the flywheel energy storage system.
[0121] A frequency control model is constructed based on the state space, action space, reward function, cost function, and objective function of a constrained Markov decision process.
[0122] In this alternative embodiment, such as Figure 4 As shown, the expression for the frequency control model is:
[0123]
[0124] In the formula, Δf is the power system frequency deviation; H is the power system inertial time constant; ΔP G The deviation in output power of the diesel generator; ΔP E The output power deviation of the flywheel energy storage system; ΔP W For wind power deviation; ΔP L ΔX represents the load power deviation; D represents the power system load damping coefficient; ΔX G For the speed controller output signal deviation; T t Δu is the governor constant; G For diesel generator control signals; T f R is the time constant of the diesel generator; f Δu is the droop coefficient of the diesel generator. E For flywheel energy storage system control signals; T E is the time constant of the flywheel energy storage system.
[0125] In this optional embodiment, the acquisition of historical power system data involves constructing a labeled dataset based on the cost function of the core elements; training a Long Short-Term Memory (LSTM) network using the labeled dataset; constructing a cost prediction network based on the trained LSM network; and using the state and actions of the frequency control model as input to the cost prediction network to predict the control cost at the current moment, including:
[0126] A labeled dataset is constructed using historical data, load frequency control is implemented through a PID controller, and the state and actions of the frequency control model are used as training data based on the cost function of a constrained Markov decision process.
[0127] Using the constructed labeled dataset, a pre-configured long short-term memory network model is trained;
[0128] Data pairs for constructing and training a long short-term memory network model were used to utilize the dynamic response data of a flywheel energy storage system.
[0129] A cost prediction network is constructed based on data pairs from the trained Long Short-Term Memory network.
[0130] The cost prediction network receives the labels of the training data at each time step and updates the current state through hidden states and memory units.
[0131] After obtaining the updated hidden state, the output of the cost prediction network is mapped to the predicted cost through a fully connected layer;
[0132] Supervised learning is used to establish a mapping relationship from state and action to cost, thereby completing the training of the cost prediction network.
[0133] Furthermore, it should be noted that a labeled dataset (x) is constructed by collecting real-time information from the power system. t ,y t ), x in the dataset t Includes the input and output of the PID controller, i.e. Meanwhile, in the actual control process y t The cost calculated using the cost function of the constrained Markov decision process is also used as a label for the training data.
[0134] Based on a Long Short-Term Memory (LSTM) network architecture, the cost prediction network is constructed as shown in the state update formula. The state update formula for the cost prediction network is as follows:
[0135]
[0136] After obtaining the updated hidden state h t Then, the output of the cost prediction network is mapped to the predicted cost through a fully connected layer:
[0137] Q C (s t ,a t ) = h (n) [h (n-1) [...h (1) (h t )]];
[0138] The cost prediction network is trained using supervised learning, as shown in the following expression:
[0139]
[0140] In the formula, f t The output of the forget gate is σ; the activation function is W. f h is the weight matrix of the forget gate; t-1 The hidden state at time t-1; x t b is the input to the Long Short-Term Memory network; f For the bias term of the forget gate; i t The output of the input gate; W i b is the weight matrix of the input gate; i This is the bias term for the input gate; For transitional memory units; Wc b is the weight matrix of the memory cells; c For the bias term of the memory unit; c t c is the memory unit at time t. t-1 For memory units at time t-1; o t The output of the output gate; W o b is the weight matrix of the output gate; o For the bias term of the output gate; h t Q is the hidden state at time t; C (s t ,a t ) represents the cost prediction function at all points; h (n) h is the activation function for the nth layer; (n-1) h is the activation function for the (n-1)th layer. (1) This is the activation function for layer 1; To find the minimum value; y is the sample label during supervised learning; ||yQ C (s t ,a t )|| 2 To monitor the loss value during the learning process.
[0141] In this optional embodiment, mapping the output of the cost prediction network to the predicted cost through a fully connected layer includes:
[0142] Q C (s t ,a t ) = h (n) [h (n-1) [...h (1) (h t )]];
[0143] In the formula, Q C (s t ,a t ) represents the output of the cost prediction network; h (n) h is the activation function for the nth layer; (n-1) h is the activation function for the (n-1)th layer. (1) The activation function for layer 1; (h t ) represents the hidden state output by the Long Short-Term Memory network.
[0144] In this alternative embodiment, such as Figure 5As shown, the process involves setting an upper limit for the cost function, constructing a deep reinforcement learning controller, and transforming the inequality constraints of the objective function in the deep reinforcement learning controller based on Lagrange multipliers. Based on the transformation result, the actor network of the deep reinforcement learning controller is optimized using the predicted control cost at the current moment. The optimized actor network is then applied to real-time load frequency control to achieve load frequency control.
[0145] Set an upper limit for the cost function;
[0146] Construct a deep reinforcement learning controller comprising an actor network, a reward critic network, an actor-goal network, and a reward critic-goal network;
[0147] Initialize the network parameters of the actor network and the reward critic network, and then assign the initialized network parameters of the actor network and the reward critic network to the actor target network and the reward critic network, respectively;
[0148] Initialize the interaction experience pool and dual variables of the Lagrange multipliers, obtain several interaction experiences between the actor network and the power system by predicting the control cost at the current moment, and construct the interaction experience pool based on these interaction experiences.
[0149] Optimize the network parameters of the actor network, rewarded commentator network, target actor network, and target rewarded commentator network based on the interactive experience pool;
[0150] Repeat the steps of building an interactive experience pool and optimizing using the interactive experience pool up to a preset number of times, and output the optimized actor network;
[0151] The optimized actor network is applied to real-time load frequency control to achieve control over the load frequency.
[0152] Furthermore, it should be noted that the introduction of the Lagrange multiplier λ transforms the inequality-constrained problem of the objective function into an unconstrained problem. Based on the primal-dual optimization method, the network parameters and dual variables are updated sequentially during iteration. The specific steps are as follows:
[0153] Construct a deep reinforcement learning controller, including an actor network μ(a|θ) μ ), Rewards for Critics Network Actor-target network μ'(a|θ) μ' ) and reward critic target network
[0154] Initialize the interaction experience pool and dual variable λ, and obtain several interaction experiences (s) between the actor network and the power system through simulation experiments. t ,a t ,r t ,ct ,s t+1 ), thus obtaining an interactive experience pool;
[0155] Optimize the network parameters of the actor network and the rewarded critic network, as well as the target actor network and the target rewarded critic network, based on the interactive experience pool;
[0156] Repeat the steps of obtaining the interactive experience pool, optimizing the actor network and the rewarded critic network, as well as the target actor network and the target rewarded critic network, until a preset number of times, and output the optimized actor network.
[0157] In this optional embodiment, optimizing the network parameters of the actor network, the reward commentator network, the target actor network, and the target reward commentator network based on the interaction experience pool includes:
[0158] Sample a preset number of interactive experiences from the interactive experience pool;
[0159] Calculate the target value of rewarding commentators based on the network parameters of the target actor network and the target reward commentator network.
[0160] Based on the reward critic objective value, the network parameters of the reward critic network are updated by minimizing the loss function;
[0161] Calculate the action gradient and update the network parameters of the actor network based on the action gradient using gradient descent.
[0162] The reward system for commentators is optimized by combining the chain rule with predictions of control costs at the current moment.
[0163] Calculate the gradient of the dual variable, and update the dual variable using gradient descent based on the gradient of the dual variable;
[0164] Optimize the target actor network and the target reward commentator network.
[0165] In this optional embodiment, the expression for calculating the gradient of the dual variable is:
[0166]
[0167] In the formula, d λ The gradient of the dual variable; K is the number of experiences sampled from the interactive experience pool; k is the experience number; Q C For cost prediction networks; s k For the state in experience k; μ(s) k |θ μ (This refers to an actor network;) denoted by , and d by , representing the parameters of the cost prediction network; d represents the constraint value. Additionally, it should be noted that K interaction experiences are sampled from the interaction experience pool.
[0168] The target value y for rewarding critics is calculated using the following formula. k The expression is:
[0169]
[0170] Based on target value y k By minimizing the loss function L R The expression for updating the network parameters of the rewarded commentator network is:
[0171]
[0172] The motion gradient is calculated using the following formula, and based on the motion gradient d... a The expression for updating the network parameters of the actor network using gradient descent is:
[0173]
[0174] The expression obtained using the chain rule is:
[0175]
[0176] The gradient of the dual variable is calculated using the following formula, and based on the gradient d of the dual variable... λ The expression for updating the dual variable using gradient descent is:
[0177]
[0178] The expressions for optimizing the goal actor network and the goal reward critic network are as follows:
[0179]
[0180] In the formula: τ is the preset soft update parameter; K is the number of experiences sampled from the interactive experience pool; k is the experience number; s k For the state in experience k; a k For the action in experience k; r k For the reward in experience k; c k Cost in experience k; s k+1 For the state in experience k+1; y k The target value is used to reward critics; γ is the discount factor; Q' R The target is a network of reward commentators; μ' is the target of the network of actor strategies; θ μ' For the parameters of the target actor network; The parameters of the target reward critic network; L R The network loss function is used to reward commenters; Q RThe network is used to reward critics; μ is the agent network strategy; θ μ For the actor network parameters; To reward critics for network parameters; d a For the action gradient; Q C For cost prediction network; μ(s|θ) μ ) represents the output of the actor network; d represents the constraint limit value; where, and The reward commentator network, cost prediction network, and actor network were obtained respectively using the backpropagation algorithm.
[0181] This invention presents a safety reinforcement learning-based load frequency control method for flywheel energy storage systems. It fully considers the frequency regulation characteristics and power constraints of flywheel energy storage systems, constructing a constrained Markov decision process with state space, action space, reward function, cost function, and objective function. To achieve a nonlinear mapping from state-action pairs to costs, a long short-term memory (LSTM) network is used, with the states and actions of the frequency control model as inputs to a cost prediction network to predict the control cost at the current moment. This method can capture the complex temporal dependencies between states and actions during system operation, effectively improving the accuracy of cost prediction. Finally, Lagrange multipliers are introduced to transform the inequality constraint problem of the objective function into an unconstrained problem. A cost prediction neural network is then used to improve the gradient of the primal dual deep deterministic policy (Primal)... This invention improves the efficiency of solving constrained Markov decision processes (Dual-DDPG) to update the actor network and uses the optimized actor network as a real-time load frequency control strategy model. The invention fully considers the auxiliary role of flywheel energy storage system in the load frequency control process and its frequency regulation characteristics. Compared with general Markov processes, this invention explicitly considers constraints, while taking into account the operating cost and frequency regulation performance of the power system, thus ensuring the comprehensive benefits of the power system.
[0182] Figure 2 An embodiment of the secure reinforcement learning load frequency control system of the present invention is shown.
[0183] In this optional embodiment, the secure reinforcement learning load frequency control system includes:
[0184] The frequency control model construction module 201 is used to combine the frequency regulation characteristics and power constraints of the flywheel energy storage system and construct a frequency control model based on the core elements of the constrained Markov decision process.
[0185] The cost prediction network construction module 203 is used to acquire historical data of the power system and construct a labeled dataset based on the cost function of the core elements; train a long short-term memory network according to the labeled dataset, construct a cost prediction network based on the trained long short-term memory network, and use the state and action of the frequency control model as the input of the cost prediction network to predict the control cost at the current moment.
[0186] The deep reinforcement learning controller optimization module 205 is used to set the upper limit of the cost function, construct the deep reinforcement learning controller, and transform the inequality constraints of the objective function in the deep reinforcement learning controller based on Lagrange multipliers; according to the transformation result, the actor network of the deep reinforcement learning controller is optimized by predicting the control cost at the current moment, and the optimized actor network is applied to real-time load frequency control to achieve control of the load frequency.
[0187] In this optional embodiment, the frequency control model construction module includes: a diesel generator constraint determination module, a constraint Markov decision process element definition module, and a control model construction module;
[0188] The constraint determination module for diesel generators is used to determine the generator generation rate constraints and power increment constraints of diesel generators.
[0189] The Constrained Markov Decision Process (CDM) element definition module is used to define the state space, action space, reward function, cost function, and objective function of the constrained Markov decision process based on the generator rate constraint and power increment constraint of the diesel generator, combined with the adjustment dead zone and power increment constraint of the flywheel energy storage system.
[0190] The control model construction module is used to construct a frequency control model based on the state space, action space, reward function, cost function, and objective function of a constrained Markov decision process.
[0191] In this optional embodiment, the cost prediction network construction module includes: a labeled dataset module, a long short-term memory network model configuration module, a flywheel energy storage system data pair construction module, a prediction network construction module, an output mapping module, a cost prediction module, and a supervised learning training module;
[0192] The labeled dataset module is used to construct a labeled dataset using historical data, implement load frequency control through a PID controller, and use the state and actions of the frequency control model as training data based on the cost function of a constrained Markov decision process.
[0193] The Long Short-Term Memory Network Model Configuration Module is used to train a pre-configured Long Short-Term Memory Network Model using a constructed labeled dataset.
[0194] The flywheel energy storage system data pair construction module is used to construct and train a long short-term memory network model using the dynamic response data of the flywheel energy storage system.
[0195] The prediction network building module is used to construct a cost prediction network based on data pairs trained on the Long Short-Term Memory network.
[0196] The output mapping module is used to receive the labels of the training data at each time step through the cost prediction network and update the current state through the hidden state and memory unit.
[0197] The cost prediction module is used to map the output of the cost prediction network to the predicted cost through a fully connected layer after obtaining the updated hidden state.
[0198] The supervised learning training module is used to establish a mapping relationship from state and action to cost through supervised learning, thereby completing the training of the cost prediction network.
[0199] In this optional embodiment, the deep reinforcement learning controller optimization module includes: a cost function upper limit setting module, a deep reinforcement learning controller construction module, a network parameter initialization module, an interactive experience pool construction module, a network parameter optimization module, an optimization process iteration module, and a real-time load frequency control module.
[0200] The upper limit setting module for the cost function is used to set the upper limit value of the cost function;
[0201] The Deep Reinforcement Learning Controller Building Module is used to build a deep reinforcement learning controller that includes an actor network, a reward critic network, an actor target network, and a reward critic target network.
[0202] The network parameter initialization module is used to initialize the network parameters of the actor network and the reward critic network, and then assign the initialized network parameters of the actor network and the reward critic network to the actor target network and the reward critic network, respectively.
[0203] The interaction experience pool construction module is used to initialize the interaction experience pool and dual variables of the Lagrange multiplier. It obtains several interaction experiences between the actor network and the power system by predicting the control cost at the current moment, and constructs the interaction experience pool based on these interaction experiences.
[0204] The network parameter optimization module is used to optimize the network parameters of the actor network, the rewarded commentator network, the target actor network, and the target rewarded commentator network based on the interactive experience pool.
[0205] The optimization process iteration module is used to repeatedly execute the steps of building the interaction experience pool and optimizing using the interaction experience pool up to a preset number of times, and output the optimized actor network.
[0206] The real-time load frequency control module is used to apply the optimized actor network to real-time load frequency control, thereby controlling the load frequency.
[0207] In this optional embodiment, the network parameter optimization module includes: an interaction experience acquisition module, a reward commentator target value calculation module, a reward commentator network network parameter update module, an action gradient and update actor network network parameter calculation module, a reward commentator network optimization module, a dual variable update module, and a target actor network and target reward commentator network optimization module;
[0208] The interactive experience collection module is used to sample a preset number of interactive experiences from the interactive experience pool;
[0209] The reward commentator target value calculation module is used to calculate the reward commentator target value based on the target actor network and the network parameters of the target reward commentators.
[0210] The network parameter update module for the rewarded commentator network is used to update the network parameters of the rewarded commentator network by minimizing the loss function based on the rewarded commentator target value.
[0211] The action gradient and network parameter update module is used to calculate the action gradient and update the network parameters of the actor network based on the action gradient using the gradient descent method.
[0212] The rewarded commentator network optimization module is used to optimize the rewarded commentator network based on the chain rule and the prediction of the control cost at the current moment.
[0213] The dual variable update module is used to calculate the gradient of the dual variable and update the dual variable using the gradient descent method based on the gradient of the dual variable.
[0214] The Target Actor Network and Target Reward Commentator Network Optimization Module is used to optimize the Target Actor Network and Target Reward Commentator Network.
[0215] In this optional embodiment, the expression for calculating the gradient of the dual variable is:
[0216]
[0217] In the formula, d λ The gradient of the dual variable; K is the number of experiences sampled from the interactive experience pool; k is the experience number; Q C For cost prediction networks; s k For the state in experience k; μ(s) k |θ μ (This refers to an actor network;) d represents the parameters of the cost prediction network; d represents the constraint limit value.
[0218] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 3 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores static and dynamic information data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above method embodiments.
[0219] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0220] In addition, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0221] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0222] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0223] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.
Claims
1. A method for controlling the frequency of a secure reinforcement learning load, characterized in that, The secure reinforcement learning load frequency control method includes: Combining the frequency regulation characteristics and power constraints of flywheel energy storage systems, and constructing a frequency control model based on the core elements of constrained Markov decision processes; Historical power system data is acquired, and a labeled dataset is constructed based on the cost function of the core elements. A long short-term memory network is trained based on the labeled dataset, and a cost prediction network is constructed based on the trained long short-term memory network. The state and actions of the frequency control model are used as inputs to the cost prediction network to predict the control cost at the current moment. An upper limit value is set for the cost function, a deep reinforcement learning controller is constructed, and the inequality constraints of the objective function in the deep reinforcement learning controller are transformed based on Lagrange multipliers. According to the transformation result, the actor network of the deep reinforcement learning controller is optimized by predicting the control cost at the current moment, and the optimized actor network is applied to real-time load frequency control to achieve control of the load frequency.
2. The method for controlling the frequency of a secure reinforcement learning load according to claim 1, characterized in that, The frequency control model, which combines the frequency regulation characteristics and power constraints of the flywheel energy storage system and is constructed based on the core elements of the constrained Markov decision process, includes: Determine the generator generation rate constraints and power increment constraints of the diesel generator; Based on the generator rate constraints and power increment constraints of diesel generators, and combined with the adjustment dead zone and power increment constraints of flywheel energy storage systems, the state space, action space, reward function, cost function and objective function of constrained Markov decision process are defined. A frequency control model is constructed based on the state space, action space, reward function, cost function, and objective function of a constrained Markov decision process.
3. The method for controlling the frequency of a secure reinforcement learning load according to claim 2, characterized in that, The expression for the frequency control model is: In the formula, Δf is the power system frequency deviation; H is the power system inertial time constant; ΔP G The deviation of the diesel generator output power; ΔP E The output power deviation of the flywheel energy storage system; ΔP W For wind power deviation; ΔP L ΔX represents the load power deviation; D represents the power system load damping coefficient; ΔX G For the speed controller output signal deviation; T t Δu is the governor constant; G For diesel generator control signals; T f R is the time constant of the diesel generator; f Δu is the droop coefficient of the diesel generator. E For flywheel energy storage system control signals; T E is the time constant of the flywheel energy storage system.
4. The method for controlling the frequency of a secure reinforcement learning load according to claim 1, characterized in that, The process of acquiring historical power system data involves constructing a labeled dataset based on the cost function of the core elements; training a Long Short-Term Memory (LSTM) network using the labeled dataset; constructing a cost prediction network based on the trained LSM network; and using the state and actions of the frequency control model as input to the cost prediction network to predict the control cost at the current moment, including: A labeled dataset is constructed using historical data, load frequency control is implemented through a PID controller, and the state and actions of the frequency control model are used as training data based on the cost function of a constrained Markov decision process. Using the constructed labeled dataset, a pre-configured long short-term memory network model is trained; Data pairs for constructing and training a long short-term memory network model were used to utilize the dynamic response data of a flywheel energy storage system. A cost prediction network is constructed based on data pairs from the trained Long Short-Term Memory network. The cost prediction network receives the labels of the training data at each time step and updates the current state through hidden states and memory units. After obtaining the updated hidden state, the output of the cost prediction network is mapped to the predicted cost through a fully connected layer; Supervised learning is used to establish a mapping relationship from state and action to cost, thereby completing the training of the cost prediction network.
5. The method for controlling the frequency of a secure reinforcement learning load according to claim 4, characterized in that, The step of mapping the output of the cost prediction network to the predicted cost through a fully connected layer includes: Q C (s t ,a t )=h (n) [h (n-1) [...h (1) (h t )]]; In the formula, Q C (s t ,a t () is the output of the cost prediction network; h (n) Let n be the activation function of the nth layer; h (n-1) Let n be the activation function of the (n-1)th layer. h (1) This is the activation function for layer 1; (h t ) represents the hidden state output by the Long Short-Term Memory network.
6. The method for controlling the frequency of a secure reinforcement learning load according to claim 1, characterized in that, The upper limit of the cost function is set, a deep reinforcement learning controller is constructed, and the inequality constraints of the objective function in the deep reinforcement learning controller are transformed based on Lagrange multipliers. Based on the transformation results, the actor network of the deep reinforcement learning controller is optimized using the predicted control cost at the current moment. The optimized actor network is then applied to real-time load frequency control to achieve load frequency control, including: Set an upper limit for the cost function; Construct a deep reinforcement learning controller comprising an actor network, a reward critic network, an actor-goal network, and a reward critic-goal network; Initialize the network parameters of the actor network and the reward critic network, and then assign the initialized network parameters of the actor network and the reward critic network to the actor target network and the reward critic network, respectively; Initialize the interaction experience pool and dual variables of the Lagrange multipliers, obtain several interaction experiences between the actor network and the power system by predicting the control cost at the current moment, and construct the interaction experience pool based on these interaction experiences. Optimize the network parameters of the actor network, rewarded commentator network, target actor network, and target rewarded commentator network based on the interactive experience pool; Repeat the steps of building an interactive experience pool and optimizing using the interactive experience pool up to a preset number of times, and output the optimized actor network; The optimized actor network is applied to real-time load frequency control to achieve control over the load frequency.
7. The method for controlling the frequency of a secure reinforcement learning load according to claim 6, characterized in that, The optimization of network parameters for the actor network, rewarded critic network, target actor network, and target rewarded critic network based on the interactive experience pool includes: Sample a preset number of interactive experiences from the interactive experience pool; Calculate the target value of rewarding commentators based on the network parameters of the target actor network and the target reward commentator network. Based on the reward critic objective value, the network parameters of the reward critic network are updated by minimizing the loss function; Calculate the action gradient and update the network parameters of the actor network based on the action gradient using gradient descent. The reward system for commentators is optimized by combining the chain rule with predictions of control costs at the current moment. Calculate the gradient of the dual variable, and update the dual variable using gradient descent based on the gradient of the dual variable; Optimize the target actor network and the target reward commentator network.
8. The method for controlling the frequency of a secure reinforcement learning load according to claim 7, characterized in that, The expression for calculating the gradient of the dual variable is as follows: In the formula, d λ The gradient of the dual variable; K is the number of experiences sampled from the interactive experience pool; k is the experience number; Q C For cost prediction networks; s k For the state in experience k; μ(s k |θ μ (This refers to an actor network;) These are the parameters for the cost prediction network; d represents the constraint limit value.
9. A safe reinforcement learning load frequency control system, characterized in that, include: The frequency control model construction module is used to combine the frequency regulation characteristics and power constraints of the flywheel energy storage system, and to construct a frequency control model based on the core elements of the constrained Markov decision process. The cost prediction network construction module is used to acquire historical data of the power system and construct a labeled dataset based on the cost function of the core elements; a long short-term memory network is trained according to the labeled dataset, and a cost prediction network is constructed based on the trained long short-term memory network. The state and actions of the frequency control model are used as the input of the cost prediction network to predict the control cost at the current moment. The deep reinforcement learning controller optimization module is used to set the upper limit of the cost function, construct the deep reinforcement learning controller, and transform the inequality constraints of the objective function in the deep reinforcement learning controller based on Lagrange multipliers. Based on the transformation results, the actor network of the deep reinforcement learning controller is optimized using the predicted control cost at the current moment. The optimized actor network is then applied to real-time load frequency control to achieve control of the load frequency.
10. A secure reinforcement learning load frequency control system according to claim 9, characterized in that, The frequency control model construction module includes: a diesel generator constraint determination module, a constraint Markov decision process element definition module, and a control model construction module; The constraint determination module for diesel generators is used to determine the generator generation rate constraints and power increment constraints of diesel generators. The Constrained Markov Decision Process (CDM) element definition module is used to define the state space, action space, reward function, cost function, and objective function of the constrained Markov decision process based on the generator rate constraint and power increment constraint of the diesel generator, combined with the adjustment dead zone and power increment constraint of the flywheel energy storage system. The control model construction module is used to construct a frequency control model based on the state space, action space, reward function, cost function, and objective function of a constrained Markov decision process.
11. A secure reinforcement learning load frequency control system according to claim 10, characterized in that, The cost prediction network construction module includes: a labeled dataset module, a long short-term memory network model configuration module, a flywheel energy storage system data pair construction module, a prediction network construction module, an output mapping module, a cost prediction module, and a supervised learning training module; The labeled dataset module is used to construct a labeled dataset using historical data, implement load frequency control through a PID controller, and use the state and actions of the frequency control model as training data based on the cost function of a constrained Markov decision process. The Long Short-Term Memory Network Model Configuration Module is used to train a pre-configured Long Short-Term Memory Network Model using a constructed labeled dataset. The flywheel energy storage system data pair construction module is used to construct and train a long short-term memory network model using the dynamic response data of the flywheel energy storage system. The prediction network building module is used to construct a cost prediction network based on data pairs trained on the Long Short-Term Memory network. The output mapping module is used to receive the labels of the training data at each time step through the cost prediction network and update the current state through the hidden state and memory unit. The cost prediction module is used to map the output of the cost prediction network to the predicted cost through a fully connected layer after obtaining the updated hidden state. The supervised learning training module is used to establish a mapping relationship from state and action to cost through supervised learning, thereby completing the training of the cost prediction network.
12. A secure reinforcement learning load frequency control system according to claim 11, characterized in that, The deep reinforcement learning controller optimization module includes: a cost function upper limit setting module, a deep reinforcement learning controller construction module, a network parameter initialization module, an interactive experience pool construction module, a network parameter optimization module, an optimization process iteration module, and a real-time load frequency control module. The upper limit setting module for the cost function is used to set the upper limit value of the cost function; The Deep Reinforcement Learning Controller Building Module is used to build a deep reinforcement learning controller that includes an actor network, a reward critic network, an actor target network, and a reward critic target network. The network parameter initialization module is used to initialize the network parameters of the actor network and the reward critic network, and then assign the initialized network parameters of the actor network and the reward critic network to the actor target network and the reward critic network, respectively. The interaction experience pool construction module is used to initialize the interaction experience pool and dual variables of the Lagrange multiplier. It obtains several interaction experiences between the actor network and the power system by predicting the control cost at the current moment, and constructs the interaction experience pool based on these interaction experiences. The network parameter optimization module is used to optimize the network parameters of the actor network, the rewarded commentator network, the target actor network, and the target rewarded commentator network based on the interactive experience pool. The optimization process iteration module is used to repeatedly execute the steps of building the interaction experience pool and optimizing using the interaction experience pool up to a preset number of times, and output the optimized actor network. The real-time load frequency control module is used to apply the optimized actor network to real-time load frequency control, thereby controlling the load frequency.
13. A secure reinforcement learning load frequency control system according to claim 12, characterized in that, The network parameter optimization module includes: an interactive experience acquisition module, a reward commentator target value calculation module, a reward commentator network network parameter update module, an action gradient and update actor network network parameter calculation module, a reward commentator network optimization module, a dual variable update module, and a target actor network and target reward commentator network optimization module; The interactive experience collection module is used to sample a preset number of interactive experiences from the interactive experience pool; The reward commentator target value calculation module is used to calculate the reward commentator target value based on the target actor network and the network parameters of the target reward commentators. The network parameter update module for the rewarded commentator network is used to update the network parameters of the rewarded commentator network by minimizing the loss function based on the rewarded commentator target value. The action gradient and network parameter update module is used to calculate the action gradient and update the network parameters of the actor network based on the action gradient using the gradient descent method. The rewarded commentator network optimization module is used to optimize the rewarded commentator network based on the chain rule and the prediction of the control cost at the current moment. The dual variable update module is used to calculate the gradient of the dual variable and update the dual variable using the gradient descent method based on the gradient of the dual variable. The Target Actor Network and Target Reward Commentator Network Optimization Module is used to optimize the Target Actor Network and Target Reward Commentator Network.
14. A secure reinforcement learning load frequency control system according to claim 13, characterized in that, The expression for calculating the gradient of the dual variable is as follows: In the formula, d λ The gradient of the dual variable; K is the number of experiences sampled from the interactive experience pool; k is the experience number; Q C For cost prediction networks; s k For the state in experience k; μ(s k |θ μ (This refers to an actor network;) These are the parameters for the cost prediction network; d represents the constraint limit value.
15. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Frequency modulation control strategy optimization method and system for power grid unit
CN118300128A
Model prediction-based frequency self-adaptive control method for virtual synchronizer inverter
WO2023088124A1