Load Frequency Control Method, Device and Electronic Equipment Based on Deep Reinforcement Learning
By applying a load frequency control method based on deep reinforcement learning in the power system, the training agent can independently adjust in a dynamic environment, solving the problem that traditional PID controllers cannot ensure grid frequency stability and load balancing in complex environments, and improving the robustness and performance of the system.
Patent Information
- Application Number
- CN202411001461.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-07-25
AI Technical Summary
Traditional PID controllers cannot ensure the stability of the grid frequency and the balance between load and power generation under random load changes and various operating environments.
The load frequency control method based on deep reinforcement learning is adopted, and the model (agent) is determined by training the initial controller parameters, empirical data is obtained in a dynamic power system environment, and behavior is automatically adjusted to adapt to changing conditions and load needs.
It improves the robustness and performance of the system under various uncertain conditions, effectively deals with nonlinear problems caused by load-power generation changes, and achieves frequency stability and energy management.
Smart Images

Figure CN119109079B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of power systems, and in particular, to a load frequency control method, device, and electronic device based on deep reinforcement learning. Background Art
[0002] Modern power systems face challenges of growing renewable energy demands and complex operating environments. Automatic Load Frequency Control (ALFC) is an important control link in power systems, aiming to maintain the balance between load and generation by regulating tie-line power flows and frequency oscillations between interconnected regions. Therefore, effective load frequency control strategies are crucial for achieving a balance between system reliability and efficiency under uncertain conditions.
[0003] Currently, Load Frequency Control (LFC) in power systems still uses classical Proportional-Integral-Derivative (PID) controllers, which have a simple structure, high reliability, and a good performance-cost ratio. However, for decades, the gain values of PID controllers have mainly relied on experience and have been adjusted through trial-and-error procedures and traditional tuning methods (such as the Ziegler-Nicholas method). In the face of random load changes and various operating environments, the control strategies of traditional PID controllers show limitations and cannot ensure the stability of the grid frequency and the balance between load and generation. Therefore, it is urgent to solve this technical problem. Summary of the Invention
[0004] In view of the above situation, embodiments of the present disclosure provide a load frequency control method, device, and electronic device based on deep reinforcement learning, aiming to solve the above problems or at least partially solve the above problems.
[0005] In a first aspect, embodiments of the present disclosure provide a load frequency control method based on deep reinforcement learning, the method including:
[0006] For any regional subsystem in a target interconnected power system, obtain corresponding current state data and an initial controller parameter determination model, where the current state data is area control error data, and the initial controller parameter determination model is generated based on a deep reinforcement learning algorithm;
[0007] For any target area subsystem, at each time step, a model is determined using the corresponding initial controller parameters to process the corresponding current state data, and target action data is obtained. The action data is a set of controller parameter adjustment amounts. Each of the target action data is executed in the target interconnected power system as the environment to obtain feedback data of the environment. The feedback data includes a reward value and the current state data of each target area subsystem. The reward value, along with the target action data, current state data, and next current state data corresponding to the target area subsystem, constitute an experience data, which is stored in the replay buffer corresponding to the target area subsystem.
[0008] If it is determined that each of the initial controller parameter determination models satisfies a preset termination condition, then each of the initial controller parameter determination models is used to generate a set of controller parameter adjustment amounts based on the current state data of the target interconnected power system to control the load frequency of the target interconnected power system.
[0009] If it is determined that each of the initial controller parameter determination models does not satisfy the preset termination condition, then for any target initial controller parameter determination model, a batch of sample experiences is extracted from the corresponding replay buffer, and the target initial controller parameter determination model is trained using the batch of sample experiences. The trained model obtained is used as a new initial controller parameter determination model. The target interconnected power system is reset, and the process returns to the step of obtaining the corresponding current state data for any area subsystem in the target interconnected power system.
[0010] In a second aspect, an embodiment of the present disclosure further provides a load frequency control device based on deep reinforcement learning. The device includes:
[0011] An acquisition module, configured to obtain the corresponding current state data and an initial controller parameter determination model for any area subsystem in the target interconnected power system. The current state data is area control error data, and the initial controller parameter determination model is generated based on a deep reinforcement learning algorithm.
[0012] An experience collection module, configured to, for any target area subsystem, at each time step, use the corresponding initial controller parameter determination model to process the corresponding current state data to obtain target action data. The action data is a set of controller parameter adjustment amounts. Each of the target action data is executed in the target interconnected power system as the environment to obtain feedback data of the environment. The feedback data includes a reward value and the current state data of each target area subsystem. The reward value, along with the target action data, current state data, and next current state data corresponding to the target area subsystem, constitute an experience data, which is stored in the replay buffer corresponding to the target area subsystem.
[0013] A control module, configured to, if it is determined that each of the initial controller parameter determination models satisfies a preset termination condition, use each of the initial controller parameter determination models to generate a set of controller parameter adjustment amounts according to the current state data of the target interconnected power system, so as to control the load frequency of the target interconnected power system;
[0014] A training module, configured to, if it is determined that each of the initial controller parameter determination models does not satisfy the preset termination condition, for any target initial controller parameter determination model, extract a batch of sample experiences from the corresponding replay buffer, and use the batch of sample experiences to train the target initial controller parameter determination model, and use the obtained trained model as a new initial controller parameter determination model; reset the target interconnected power system, and return to the step of obtaining the corresponding current state data for any regional subsystem in the target interconnected power system.
[0015] In a third aspect, an embodiment of the present disclosure further provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, the executable instructions, when executed, cause the processor to execute the steps of the above load frequency control method based on deep reinforcement learning.
[0016] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium, the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute the steps of the above load frequency control method based on deep reinforcement learning.
[0017] By means of the above technical solution, the load frequency control method, device and electronic device provided by the embodiment of the present disclosure, compared with the existing method of setting PID controller parameters according to experience to control the grid load frequency, the embodiment of the present disclosure proposes an automatic control method for grid load frequency based on deep reinforcement learning. Specifically, for multiple regional subsystems in the target interconnected power system, the corresponding initial controller parameter determination models (agents) are respectively trained, and experience data is obtained from the dynamic target interconnected power system environment through an interactive trial-and-error method, and the existing experience is effectively utilized for autonomous learning and training. Therefore, each agent can autonomously adjust its behavior to adapt to the changing power system conditions and load demands, and improve the robustness and performance of the system under various uncertain conditions. According to the current state of the target interconnected power system, a PID controller parameter adjustment amount is generated to adjust the parameters of the PID controller for controlling the target interconnected power system. The solution provided by the embodiment of the present disclosure can effectively cope with the nonlinear problems caused by load-generation changes, has a wide application prospect in the power system, and can effectively solve the frequency stability and energy management problems in complex environments.
[0018] The above description is only an overview of the technical solution of the present disclosure. In order to better understand the technical means of the present disclosure, it can be implemented according to the content of the specification. In addition, in order to make the above and other purposes, features, and advantages of the present disclosure more obvious and understandable, the specific embodiments of the present disclosure are specifically exemplified below. Description of the Drawings
[0019] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0020] Figure 1 A schematic flow chart of the load frequency control method based on deep reinforcement learning provided by an embodiment of the present disclosure is shown;
[0021] Figure 2 A schematic structural diagram of the controller parameter determination model provided by an embodiment of the present disclosure is shown;
[0022] Figure 3 A schematic structural diagram of the load frequency control device based on deep reinforcement learning provided by an embodiment of the present disclosure is shown;
[0023] Figure 4 A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown. Detailed Embodiments
[0024] In order to make the purpose, technical solution, and advantages of the present disclosure clearer, the technical solution of the present disclosure will be clearly and completely described below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0025] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0026] It should be noted that the terms "first", "second", etc. in the specification, claims, and drawings of the present disclosure are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such use can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the term "including" and its variants should be interpreted as an open term meaning "including but not limited to".
[0027] As introduced above, currently, in the power system load frequency control (LFC), the classical proportional integral derivative (PID) controller is still used. It has a simple structure, high reliability, and a good performance-cost ratio. However, for decades, the gain value of the PID controller has mainly relied on experience and has been adjusted through trial-and-error procedures and traditional tuning methods (such as the Ziegler-Nicholas method). In the face of random load changes and various operating environments, the control strategy of the traditional PID controller reveals limitations and cannot ensure the stability of the grid frequency and the balance between load and generation. Based on this, the present invention proposes a load frequency control method, device, and electronic device based on deep reinforcement learning, which will be described in detail below through specific embodiments.
[0028] Before introducing the specific embodiments, the professional terms involved in the embodiments of the present disclosure will be explained first:
[0029] 1) Interconnected Power Systems refers to a system in which several independent power systems are connected by tie lines or other connection devices.
[0030] 2) Area Control Error (ACE for short): An index used to measure the power generation and load conditions in a certain area.
[0031] 3) Environment: In deep reinforcement learning, the environment is the external world where the agent (the controller parameter determination model in the present disclosure) is located. It can be the real physical world or a virtual simulation environment. The agent receives the state information of the environment at each time step, selects an appropriate action, and then the environment gives feedback on this action, including rewards and new states. In the interconnected power system of the present disclosure, all things except the agent are called the environment. The interconnected power system adopts an integrated hybrid power system architecture, which can include wind energy, photovoltaic, electric vehicles, hydropower, and thermal power plants. At the same time, non-linear factors such as generation dead band (GDB) and generation rate constraint (GRC) should be considered.
[0032] 4) Observation: Observation is the information obtained by the agent from the environment, usually including the description of the state and other relevant information.
[0033] 5) Action: An action is one of the decisions that the agent can take in a given state, used to maximize the reward in a certain state. In the present disclosure, the action is in the form of a control signal output by the agent and is used to control the power plant.
[0034] 6) Reward: The reward is the feedback given by the environment after the agent executes an action. It is a scalar value and is usually used to measure the quality of that action.
[0035] 7) Time step: A basic unit for the agent to interact with the environment. After each time step, the agent takes an action based on the current state, and the environment gives a new state and a reward.
[0036] For ease of understanding of this embodiment, first, a load frequency control method based on deep reinforcement learning disclosed in the embodiments of the present disclosure will be introduced in detail. The execution subject of the load frequency control method based on deep reinforcement learning provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. Such a computer device includes, for example: a terminal device, a server, or other processing devices. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a personal digital assistant (PDA), a handheld device, a computing device, etc. In some possible implementation manners, the load frequency control method based on deep reinforcement learning may be implemented by a processor calling computer-readable instructions stored in a memory.
[0037] Figure 1 shows a schematic flowchart of the load frequency control method based on deep reinforcement learning provided in the embodiments of the present disclosure. From Figure 1 it can be seen that the embodiments of the present disclosure at least include steps S101 - S104:
[0038] S101: For any regional subsystem in the target interconnected power system, obtain the corresponding current state data and an initial controller parameter determination model. The current state data is area control error data, and the initial controller parameter determination model is generated based on a deep reinforcement learning algorithm.
[0039] S102: For any target regional subsystem, at each time step, use the corresponding initial controller parameter determination model to process the corresponding current state data to obtain target action data. The action data is a set of controller parameter adjustment amounts. Execute each of the target action data in the target interconnected power system as the environment to obtain feedback data from the environment. The feedback data includes a reward value and the current state data of each target regional subsystem. The reward value, and the target action data, current state data, and next current state data corresponding to the target regional subsystem form an experience data, which is stored in the replay buffer corresponding to the target regional subsystem.
[0040] S103: If it is determined that each of the initial controller parameter determination models meets the preset termination condition, then use each of the initial controller parameter determination models to generate a set of controller parameter adjustment amounts according to the current state data of the target interconnected power system, so as to control the load frequency of the target interconnected power system.
[0041] S104: If it is determined that each of the initial controller parameter determination models does not meet the preset termination condition, then for any target initial controller parameter determination model, extract a batch of sample experiences from the corresponding replay buffer, and use the batch of sample experiences to train the target initial controller parameter determination model, and use the obtained trained model as a new initial controller parameter determination model; reset the target interconnected power system, and return to the step of obtaining the corresponding current state data for any regional subsystem in the target interconnected power system.
[0042] From Figure 1 As can be seen from the method shown above, compared with the existing method of setting PID controller parameters according to experience to control the grid load frequency, the embodiments of the present disclosure propose an automatic control method for grid load frequency based on deep reinforcement learning. Specifically, for multiple regional subsystems in the target interconnected power system, the corresponding initial controller parameter determination models (agents) are respectively trained, and experience data is obtained from the dynamic target interconnected power system environment through an interactive trial-and-error method, and the existing experience is effectively utilized for autonomous learning and training. Therefore, each agent can autonomously adjust its behavior to adapt to the changing power system conditions and load demands, and improve the robustness and performance of the system under various uncertain conditions. According to the current state of the target interconnected power system, a PID controller parameter adjustment amount is generated to adjust the parameters of the PID controller for controlling the target interconnected power system. The solution provided by the embodiments of the present disclosure can effectively cope with the nonlinear problems caused by load-generation changes, has a wide application prospect in the power system, and can effectively solve the frequency stability and energy management problems in complex environments.
[0043] The above S101 - S104 will be described in detail below.
[0044] Regarding the above S101:
[0045] Exemplarily, the target interconnected power system includes two regional subsystems: regional subsystem A and regional subsystem B. Then, the regional control error data 1 of the system can be obtained from regional subsystem A as the current state data of system A, and the initial controller parameter determination model 1 can be obtained; the regional control error data 2 of the system can be obtained from regional subsystem B as the current state data of system B, and the initial controller parameter determination model 2 can be obtained.
[0046] In a possible implementation, the area control error data includes at least one of the following: area control error, integral of the area control error, and rate of change of the area control error. Exemplarily, the area control error data includes: area control error, integral of the area control error, and rate of change of the area control error. It should be noted that the examples here are only illustrative and do not limit the embodiments of the present disclosure.
[0047] In a possible implementation, the deep reinforcement learning algorithm is: Twin Delayed Deep Deterministic Policy Gradient algorithm, Deep Q Network, Double Deep Q Network, or Competitive Deep Q Network.
[0048] Exemplarily, if the deep reinforcement learning algorithm is the Twin Delayed Deep Deterministic Policy Gradient algorithm, for any area subsystem in the target interconnected power system, a corresponding initial controller parameter determination model can be created by initializing the actor network and the critic network, and initializing the target actor network and the target critic network using the parameters of the main actor-critic network.
[0049] The output of the fully connected layer in the traditional Twin Delayed Deep Deterministic Policy Gradient algorithm model is usually any real value, including negative values. It has been found through research that this network structure may cause the adjustment amount of the PID controller parameters output by the network to be negative during the gradient optimization process, thereby affecting the stability and effectiveness of the model. Based on this, in a possible implementation of the present disclosure, if the deep reinforcement learning algorithm is the Twin Delayed Deep Deterministic Policy Gradient algorithm; the main actor network and the target actor network in the initial controller parameter determination model include an input layer, a non-negative fully connected layer, and an output layer connected in sequence.
[0050] In this embodiment, a new fully connected layer is introduced to improve the main actor network and the target actor network in the initial controller parameter determination model. Specifically, the function form of the non-negative fully connected layer is y = abs(weights)*x, where x is the input of the network and abs(weights) represents the absolute value of the network weights. In this embodiment, the main actor network and the target actor network in the initial controller parameter determination model include an input layer, a non-negative fully connected layer, and an output layer connected in sequence. This design ensures that the network weights are always non-negative, thereby avoiding the problem of negative controller parameter adjustment amounts that may occur during the gradient optimization process, which helps to improve the stability and convergence speed of training, ensure the effectiveness of the model, and significantly reduce the computational complexity and save computational resources.
[0051] Regarding the above S102:
[0052] The target interconnected power system can be a real power system or a simulated power grid environment integrated with renewable energy and adding random disturbances. Exemplarily, at time step t = 1, for the aforementioned regional subsystem A, the model 1 can process the regional control error data 1 using the initial controller parameters to obtain the target action data 1, that is, the set 1 of controller parameter adjustment amounts; for the aforementioned regional subsystem B, the model 2 can process the regional control error data 2 using the initial controller parameters to obtain the target action data 2, that is, the set 2 of controller parameter adjustment amounts. Then, the target action data 1 and the target action data 2 are executed in the target interconnected power system, and the state of the target interconnected power system changes, and the current state data (the regional control error data 1' corresponding to the regional subsystem A, the regional control error data 2' corresponding to the regional subsystem B) and the reward value are fed back. The tuple (regional control error data 1, action data 1, reward value, regional control error data 1') can be stored in the replay buffer 1 corresponding to the initial controller parameter determination model 1; the tuple (regional control error data 2, action data 2, reward value, regional control error data 2') can be stored in the replay buffer 2 corresponding to the initial controller parameter determination model 2. And so on, at time steps t = 2,..., T, the corresponding tuples can be calculated and stored in each replay buffer.
[0053] To generate the optimal PID controller parameters, the reward function plays an important role in guiding the controller parameter determination model to take actions to solve the load frequency control problem. A reasonable reward function helps to converge quickly, reduce the calculation amount, and improve the performance. In a possible implementation manner of the present disclosure, the reward value is calculated according to the following reward function:
[0054]
[0055] where T is the total number of regional subsystems in the target interconnected power system, B i is the frequency deviation coefficient corresponding to the i-th regional subsystem, Δf i is the frequency deviation of the i-th regional subsystem, ΔP tie is the tie-line power.
[0056] Exemplarily, after the target action data 1 and the target action data 2 are executed in the target interconnected power system, the state of the target interconnected power system changes, and the tie-line power and the frequency deviation of each regional subsystem can be obtained from the target interconnected power system to calculate the reward value according to the above formula. Among them, the frequency deviation coefficient corresponding to each regional subsystem can be set according to actual needs, and the embodiments of the present disclosure do not limit this.
[0057] To achieve the goal of minimizing tie-line power and frequency fluctuations in each area, in this embodiment, the reward function is defined as the sum of the absolute values of the frequency deviation and the tie-line power. Through the design of the reward function, it can be ensured that the subsequent model for determining controller parameters can quickly converge to the optimal PID controller parameter settings, greatly improving the efficiency and stability of load frequency control.
[0058] Regarding the above S103:
[0059] In this step, it can be determined whether each initial controller parameter determination model meets the preset termination condition. If it is met, it indicates that the performance of each initial controller parameter determination model reaches the requirement at this time. Inputting the current state data of the target interconnected power system into each initial controller parameter determination model, multiple sets of controller parameter adjustment amounts can be obtained. Through the multiple sets of controller parameter adjustment amounts, the proportional-integral-derivative gains of the PID controllers for controlling each regional subsystem are adjusted, thereby realizing the control of the load frequency of the target interconnected power system.
[0060] In a possible implementation manner, the preset termination condition includes: the reward function converges, or the number of training rounds of each initial controller parameter determination model reaches the preset number of training rounds.
[0061] Among them, the preset number of training rounds can be set according to actual needs, and the embodiments of the present disclosure do not limit this.
[0062] Regarding the above S104:
[0063] If it is determined that each initial controller parameter determination model does not meet the preset termination condition, then for any target initial controller parameter determination model, a batch of sample experiences are extracted from the corresponding replay buffer. Specifically, during implementation, random sampling can be performed in the corresponding replay buffer to obtain multiple sample experiences. The parameters of the target initial controller parameter determination model can be updated using the batch of sample experiences to obtain a trained model, and the trained model is used as the new initial controller parameter determination model; reset the initial state of the target interconnected power system and return to step S101, and iterate continuously until each initial controller parameter determination model meets the preset termination condition, obtaining each initial controller parameter determination model that can accurately generate a set of PID parameter adjustment amounts according to the current state.
[0064] In a possible implementation manner, the initial controller parameter determination model is generated based on the twin delayed deep deterministic policy gradient algorithm; the training of the target initial controller parameter determination model by using the batch sample experience includes: for each experience in the batch sample experience, calculating a target Q value; based on a preset loss function and each of the target Q values, updating the parameters of the two main critic networks in the target initial controller parameter determination model; updating the parameters of the main actor network in the target initial controller parameter determination model through a policy gradient algorithm; and updating the parameters of the target actor network and the two target critic networks in the target initial controller parameter determination model through a soft update rule.
[0065] Figure 2 FIG. shows a schematic structural diagram of a controller parameter determination model provided by an embodiment of the present disclosure. It can be Figure 2 seen that the controller parameter determination model includes an actor network and a critic network; the actor network includes 1 main actor network and 1 target actor network; the critic network includes critic 1, critic 2 (for estimating long-term rewards), and target critic 1 and target critic 2; for the actor network: the network structure includes an input layer, a non-negative fully connected layer, and an output layer that are connected to each other. For the critic network: the network structure includes 2 input layers (respectively for processing the state s and the action a), 2 fully connected layers, a concatenation layer (for connecting the 2 inputs), a RELU layer, a fully connected layer, a RELU layer, and a Q value layer (output layer). The following will be combined with Figure 2 this embodiment for an exemplary description.
[0066] In this embodiment, the target Q value can be calculated for each piece of experience data first. Specifically, the target Q value can be calculated according to the following formula:
[0067]
[0068] where y i is the target Q value of the i-th sample experience, r i is the immediate reward value obtained at time step t, γ is the discount rate, Q′ i (s′, a′) represents the Q value estimation of the i-th critic network for the action a′ in the state s′, and s t+1 represents the state data at time step t + 1. That is, take and the minimum value of the two.
[0069] Then, based on a preset loss function and each of the target Q values, update the parameters φ1 and φ2 of the two main critic networks in the target initial controller parameter determination model. During implementation, the preset loss function is determined according to the following formula:
[0070]
[0071] Among them, M is the total number of samples in the batch sample experience Q(s i , a i ) is the Q value predicted by the critic network for the state-action pair of the i-th sample experience.
[0072] The parameters θ of the main actor network in the model can also be updated by the policy gradient algorithm to determine the initial target controller parameters. Specifically, the following update rule can be used to update the parameters θ of the main actor network:
[0073]
[0074] Among them, is the gradient of the objective function with respect to the actor network parameters, is the gradient of the Q value output by the critic network with respect to the action a, is the gradient of the action a output by the actor network with respect to the parameter θ.
[0075] The parameters φ'1 and φ'2 of the target actor network and the two target critic networks in the model can be updated by the soft update rule to finally obtain the trained model. Specifically, the parameters of the target actor network and the target critic network can be updated according to the following formula:
[0076] θ μ′ = τθ μ +(1 - τ)θ μ′
[0077] θ Q′ = τθ Q +(1 - τ)θ Q′
[0078] Among them, θ μ′ is the network parameter of the target actor network, τ is the smoothing factor, θ μ is the network parameter of the main actor network, θ Q′ is the network parameter of the target critic network, θ Q is the network parameter of the main critic network.
[0079] During implementation, the Adam optimizer can be used to update the parameters of the actor and critic networks, and the Glorot initializer is used for weight initialization of the fully connected layer.
[0080] The embodiments of the present disclosure also provide a load frequency control method based on deep reinforcement learning, including the following steps:
[0081] Step S1: For any regional subsystem in the target interconnected power system, obtain the corresponding current state data and the initial controller parameter determination model. The current state data includes the area control error, the integral of the area control error, and the rate of change of the area control error. The initial controller parameter determination model is generated based on the twin-delayed deep deterministic policy gradient algorithm. During implementation, the main actor network and the target actor network in the initial controller parameter determination model include an input layer, a non-negative fully connected layer, and an output layer connected in sequence.
[0082] Step S2: For any target regional subsystem, at each time step, use the corresponding initial controller parameter determination model to process the corresponding current state data to obtain target action data. The action data is a set of controller parameter adjustment amounts. Execute each target action data in the target interconnected power system as the environment to obtain the feedback data of the environment. The feedback data includes the reward value and the current state data of each target regional subsystem. The reward value, along with the target action data, the current state data, and the next current state data corresponding to the target regional subsystem, form an experience data, which is stored in the replay buffer corresponding to the target regional subsystem.
[0083] During implementation, the reward value is calculated according to the following reward function:
[0084]
[0085] where T is the total number of regional subsystems in the target interconnected power system, B i is the frequency deviation coefficient corresponding to the i-th regional subsystem, Δf i is the frequency deviation of the i-th regional subsystem, and ΔP tie is the tie-line power.
[0086] Step S3: If it is determined that each initial controller parameter determination model meets the preset termination condition, then use each initial controller parameter determination model to generate a set of controller parameter adjustment amounts according to the current state data of the target interconnected power system to control the load frequency of the target interconnected power system. During implementation, the preset termination conditions include: the reward function converges, or the number of training rounds of each initial controller parameter determination model reaches the preset number of training rounds.
[0087] Step S4: If it is determined that each initial controller parameter determination model does not meet the preset termination condition, then for any target initial controller parameter determination model, extract a batch of sample experiences from the corresponding replay buffer, use the batch of sample experiences to train the target initial controller parameter determination model, and use the obtained trained model as the new initial controller parameter determination model; reset the target interconnected power system, and return to the step of obtaining the corresponding current state data for any regional subsystem in the target interconnected power system.
[0088] Specifically, using the batch sample experience to train the target initial controller parameter determination model includes: for each experience in the batch sample experience, calculating the target Q value;
[0089] Based on a preset loss function and each target Q value, updating the parameters of the two main critic networks in the target initial controller parameter determination model; updating the parameters of the main actor network in the target initial controller parameter determination model through the policy gradient algorithm; updating the parameters of the target actor network and the two target critic networks in the target initial controller parameter determination model through the soft update rule.
[0090] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order, and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0091] It should be noted that in practical applications, all the above possible implementation manners can be combined arbitrarily to form possible embodiments of the present disclosure, which will not be elaborated here one by one.
[0092] Based on the same concept, the embodiments of the present disclosure also provide a key point detection device. Figure 3 FIG. shows a schematic structural diagram of a load frequency control device based on deep reinforcement learning provided by an embodiment of the present disclosure. Refer to Figure 3 As shown, the load frequency control device 300 based on deep reinforcement learning provided by the embodiment of the present disclosure includes:
[0093] An acquisition module 301, configured to acquire corresponding current state data and an initial controller parameter determination model for any regional subsystem in the target interconnected power system, where the current state data is area control error data, and the initial controller parameter determination model is generated based on a deep reinforcement learning algorithm;
[0094] An experience collection module 302, configured to, for any target regional subsystem, at each time step, use the corresponding initial controller parameter determination model to process the corresponding current state data to obtain target action data, where the action data is a set of controller parameter adjustment amounts; execute each of the target action data in the target interconnected power system as an environment to obtain feedback data of the environment, where the feedback data includes a reward value and the current state data of each target regional subsystem; the reward value, and the target action data, current state data, and next current state data corresponding to the target regional subsystem constitute an experience data, which is stored in the replay buffer corresponding to the target regional subsystem;
[0095] The control module 303 is configured to, if it is determined that each of the initial controller parameter determination models meets a preset termination condition, use each of the initial controller parameter determination models to generate a set of controller parameter adjustment amounts according to the current state data of the target interconnected power system, so as to control the load frequency of the target interconnected power system.
[0096] The training module 304 is configured to, if it is determined that each of the initial controller parameter determination models does not meet the preset termination condition, for any target initial controller parameter determination model, extract a batch of sample experiences from the corresponding replay buffer, use the batch of sample experiences to train the target initial controller parameter determination model, and use the obtained trained model as a new initial controller parameter determination model; reset the target interconnected power system, and return to the step of obtaining the corresponding current state data for any regional subsystem in the target interconnected power system.
[0097] In a possible implementation manner, in the above device, the area control error data includes at least one of the following: area control error, integral of the area control error, and change rate of the area control error.
[0098] In a possible implementation manner, in the above device, the deep reinforcement learning algorithm is: Twin Delayed Deep Deterministic Policy Gradient algorithm, Deep Q-Network, Double Deep Q-Network, or Competitive Deep Q-Network.
[0099] In a possible implementation manner, in the above device, if the deep reinforcement learning algorithm is the Twin Delayed Deep Deterministic Policy Gradient algorithm; the main actor network and the target actor network in the initial controller parameter determination model include an input layer, a non-negative fully connected layer, and an output layer connected in sequence.
[0100] In a possible implementation manner, in the above device, the reward value is calculated according to the following reward function:
[0101]
[0102] where T is the total number of regional subsystems in the target interconnected power system, B i is the frequency deviation coefficient corresponding to the i-th regional subsystem, Δf i is the frequency deviation of the i-th regional subsystem, ΔP tie is the tie-line power.
[0103] In a possible implementation manner, in the above device, the preset termination condition includes:
[0104] The reward function converges, or the number of training rounds of each of the initial controller parameter determination models reaches a preset number of training rounds.
[0105] In a possible implementation, in the above device, the initial controller parameter determination model is generated based on the twin delayed deep deterministic policy gradient algorithm; the training module 304 is configured to: for each piece of experience in the batch sample experience, calculate the target Q value; based on a preset loss function and each of the target Q values, update the parameters of the two main critic networks in the target initial controller parameter determination model; update the parameters of the main actor network in the target initial controller parameter determination model through a policy gradient algorithm; and update the parameters of the target actor network and the two target critic networks in the target initial controller parameter determination model through a soft update rule.
[0106] It should be noted that any of the above load frequency control devices based on deep reinforcement learning can correspondingly implement the foregoing load frequency control method based on deep reinforcement learning, which will not be elaborated here.
[0107] Figure 4 The structural schematic diagram of an electronic device provided by an embodiment of the present disclosure is shown. As Figure 4 shown, at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.
[0108] The processor, network interface, and memory can be interconnected through an internal bus, and the internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0109] The memory is used to store a program. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0110] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a load frequency control device based on deep reinforcement learning at the logical level. The processor executes the program stored in the memory and is specifically used to execute the foregoing method.
[0111] The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the foregoing method can be completed by the integrated logic circuit in the hardware of the processor or instructions in software form. The foregoing processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the foregoing method.
[0112] The electronic device can execute the load frequency control method based on deep reinforcement learning provided in multiple embodiments of the present disclosure and be implemented as a load frequency control device based on deep reinforcement learning in Figure 3 the functions of the illustrated embodiments, which are not described in detail herein for the embodiments of the present disclosure.
[0113] The embodiments of the present disclosure also propose a computer-readable storage medium that stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including multiple application programs, can enable the electronic device to execute the load frequency control method based on deep reinforcement learning provided in multiple embodiments of the present disclosure.
[0114] Those skilled in the art will understand that the embodiments of the present disclosure may be provided as a method, a system, or a computer program product. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0115] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0116] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0118] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0119] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0120] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0121] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity or device including the element.
[0122] It will be appreciated by those skilled in the art that the embodiments of the present disclosure may be provided as methods, systems or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0123] The above are only embodiments of the present disclosure and are not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various modifications and variations. Any modification, equivalent substitution, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the scope of the claims of the present disclosure.
Claims
1. A load frequency control method based on deep reinforcement learning, characterized in that: The method comprises: For any regional subsystem in the target interconnected power system, the corresponding current state data and the initial controller parameter determination model are obtained, wherein the current state data is the regional control error data, and the initial controller parameter determination model is generated based on the deep reinforcement learning algorithm; the regional control error data includes: the regional control error, the integral of the regional control error and the rate of change of the regional control error; For any target area subsystem, at each time step, the corresponding initial controller parameter determination model is used to process the corresponding current state data to obtain target action data, which is a set of controller parameter adjustment quantities; each target action data is executed in the target interconnected power system as the environment to obtain feedback data of the environment, which includes a reward value and the current state data of each target area subsystem; the reward value, and the target action data, current state data and next current state data corresponding to the target area subsystem constitute an experience data, which is stored in a playback buffer corresponding to the target area subsystem; If it is determined that each of the initial controller parameter determination models satisfies a preset termination condition, each of the initial controller parameter determination models is used to generate a controller parameter adjustment amount set according to current state data of the target interconnected power system to control the load frequency of the target interconnected power system; If it is determined that each of the initial controller parameter determination models does not meet the preset termination condition, then for any target initial controller parameter determination model, batch sample experience is extracted from the corresponding playback buffer, and the target initial controller parameter determination model is trained using the batch sample experience, and the trained model is used as a new initial controller parameter determination model; the target interconnected power system is reset, and the step of returning to any regional subsystem in the target interconnected power system to obtain the corresponding current state data; If the deep reinforcement learning algorithm is a twin delayed deep deterministic policy gradient algorithm; the main actor network and the target actor network in the initial controller parameter determination model include an input layer, a non-negative fully connected layer and an output layer connected in sequence; The reward value is calculated according to the following reward function: Where T is the total number of regional subsystems in the target interconnected power system, B i is the frequency deviation coefficient corresponding to the ith regional subsystem, Δf i is the frequency deviation of the ith regional subsystem, ΔP tie is the interconnection line power.
2. The method according to claim 1, characterized in that The deep reinforcement learning algorithm is: a twin delayed deep deterministic policy gradient algorithm, a deep Q network, a dual deep Q network or a competitive deep Q network.
3. The method according to claim 1, characterized in that The preset termination conditions include: The reward function converges, or the number of training rounds of each initial controller parameter determination model reaches a preset number of training rounds.
4. The method according to claim 2 or 1, characterized in that: The initial controller parameter determination model is generated based on the twin delay deep deterministic policy gradient algorithm; and the use of the batch sample experience to train the target initial controller parameter determination model includes: For each experience in the batch of sample experiences, calculate a target Q value; Based on a preset loss function and each of the target Q values, updating the parameters of the two main critic networks in the target initial controller parameter determination model; Updating the target initial controller parameters to determine the parameters of the main actor network in the model through a policy gradient algorithm; The parameters of the target actor network and two target critic networks in the target initial controller parameter determination model are updated through a soft update rule.
5. A load frequency control device based on deep reinforcement learning, characterized in that: The device comprises: An acquisition module is used to acquire corresponding current state data and an initial controller parameter determination model for any regional subsystem in the target interconnected power system, wherein the current state data is regional control error data, and the initial controller parameter determination model is generated based on a deep reinforcement learning algorithm; the regional control error data includes: regional control error, integral of the regional control error, and rate of change of the regional control error; An experience collection module is used for, for any target area subsystem, at each time step, using the corresponding initial controller parameter determination model, processing the corresponding current state data to obtain target action data, wherein the action data is a set of controller parameter adjustment amounts; executing each of the target action data in the target interconnected power system as the environment to obtain feedback data of the environment, wherein the feedback data includes a reward value and the current state data of each target area subsystem; the reward value, and the target action data, current state data and next current state data corresponding to the target area subsystem constitute an experience data, which is stored in a playback buffer corresponding to the target area subsystem; A control module, configured to generate a controller parameter adjustment amount set according to current state data of the target interconnected power system by using each of the initial controller parameter determination models if it is determined that each of the initial controller parameter determination models satisfies a preset termination condition, so as to control the load frequency of the target interconnected power system; A training module, for extracting batch sample experience from the corresponding playback buffer for any target initial controller parameter determination model if it is determined that each of the initial controller parameter determination models does not meet the preset termination condition, using the batch sample experience to train the target initial controller parameter determination model, and using the obtained trained model as a new initial controller parameter determination model; resetting the target interconnected power system, and returning to the step of obtaining the corresponding current state data for any regional subsystem in the target interconnected power system; Wherein, if the deep reinforcement learning algorithm is a twin delayed deep deterministic policy gradient algorithm; the main actor network and the target actor network in the initial controller parameter determination model include an input layer, a non-negative fully connected layer and an output layer connected in sequence; The reward value is calculated according to the following reward function: Where T is the total number of regional subsystems in the target interconnected power system, B i is the frequency deviation coefficient corresponding to the ith regional subsystem, Δf i is the frequency deviation of the ith regional subsystem, ΔP tie is the interconnection line power.
6. An electronic device comprising: processor; as well as A memory arranged to store computer executable instructions, wherein when the executable instructions are executed, the processor is caused to perform the steps of the method according to any one of claims 1 to 4.
7. A computer-readable storage medium storing one or more programs, characterized in that: When the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device executes the steps of the method as claimed in any one of claims 1 to 4.
Citation Information
Patent Citations
New energy station frequency control method based on deep reinforcement learning
CN116436029A
KR20220013884A