Automatic power generation control dynamic optimization method, system and equipment based on imitation reinforcement learning, and medium
By modeling and optimizing automatic power generation control (AGC) systems based on imitation reinforcement learning methods, the problem of difficult optimization and constraint violations of traditional AGC methods is solved, and safer and more efficient power system control is achieved.
Patent Information
- Application Number
- CN202411326592.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-09-23
AI Technical Summary
Traditional automatic power generation control (AGC) methods are difficult to optimize AGC units within 15 minutes, and cannot effectively deal with the frequency fluctuations and high uncertainty in the dynamic environment caused by high penetration new energy. Traditional reinforcement learning methods may violate key constraints in system operation during training.
Using a method based on imitation reinforcement learning, the operation constraints and control goals of AGC dynamic optimization problems are defined, and it is modeled as a Markov decision-making process model, and the offline expert empirical data is used for imitation training, and the strategy is initialized and the strategy is further optimized using the SAC algorithm to ensure the security of system constraints.
It effectively solves the problem that traditional AGC methods cannot meet key operational constraints, improves the performance of the initial strategy, avoids serious constraint violations during training, reduces the risk of system frequency default, and improves the adaptive control capabilities of the power system.
Smart Images

Figure CN120029049A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of automatic power generation control, and in particular relates to a dynamic optimization method, system, equipment and medium for automatic power generation control based on imitation reinforcement learning. Background Art
[0002] Automatic Generation Control (AGC) is an important part of the power system. The regional power grid dispatching center needs to implement closed-loop correction control based on the monitored real-time regional control error (ACE) deviation, which is of great significance for achieving system frequency stability and smoothing the power of the interconnected power grids. At present, the research on traditional AGC strategies has achieved fruitful results, such as proportional-integral-differential (PID) control, model predictive control, and learning-based intelligent control. However, traditional AGC is a typical control with feedback delay, which may lead to over-regulation or under-regulation when coordinating different AGC units (such as hydropower and thermal power units). In addition, with the increase in wind power penetration, centralized grid-connected wind power brings a large number of minute-level power fluctuations. This further complicates the AGC regulation task and puts higher requirements on the real-time coordinated control of AGC.
[0003] In order to cope with the increasing power fluctuation and hysteresis problems, the concept of AGC dynamic optimization has been proposed in recent years. The core idea is to optimize the dispatch of AGC units in advance based on the prediction of ultra-short-term load and renewable energy (such as wind power). The main advantage of introducing AGC dynamic optimization is that it can effectively handle short-term (within 15 minutes) fluctuations caused by renewable energy because it takes into account the information of future load and renewable energy. Therefore, AGC dynamic optimization is of great significance in power systems with stochastic renewable energy.
[0004] At present, the most common method to solve the AGC dynamic control problem is to use optimization programming based on the probability model of wind power. However, the optimization-based method relies heavily on the accurate probability model of renewable energy, which is difficult to obtain in practice. In addition, due to the uncertainty involved, the characterization model of the stochastic process is usually non-convex, difficult to solve the analytical solution and the computational burden is huge. Therefore, it is difficult to consider the fluctuation of renewable energy in the future in advance in the AGC scheduling process. In recent years, deep reinforcement learning (RL) has become increasingly popular in dealing with AGC dynamic optimization problems by adopting neural networks for uncertainty prediction. Many scholars use proximal policy optimization RL algorithms, double-delayed deep deterministic policy gradient methods with multi-experience pool replay, or soft actor-critic (SAC) algorithms to solve AGC dynamic optimization problems. However, these traditional RL algorithms must be trained through a large amount of "trial and error" interaction with real systems before they can converge and become intelligent. This means that the controller may make some "wrong" decisions during the training process, resulting in serious system frequency violations. This is very unsafe and unacceptable for real systems.
[0005] In summary, the traditional automatic generation control (AGC) method only considers economic dispatch every 15 minutes and inertial response at the second level, and lacks optimization of the AGC unit within 15 minutes. Therefore, it is difficult to cope with the fluctuations in the frequency of the power system caused by high penetration of new energy and the high uncertainty brought by the real-time changing dynamic environment.
[0006] Traditional reinforcement learning either only considers training on the simulation system, but this method will lead to poor application effect of the controller trained on the simulation system in the actual system due to the difference in data distribution between the simulation system and the real system; or it directly trains on the real system, but direct training cannot guarantee the safety of system operation during training, such as the key operating constraints of automatic power generation control cannot be met. This is because the initial strategy of the traditional reinforcement learning method is usually random, and the strategy needs to be iteratively updated in the process of continuous "trial and error", so the early strategy training process may bring catastrophic control results to the system, so it is impossible to theoretically guarantee the safety of key constraints such as system frequency and equipment limitations. Summary of the invention
[0007] The purpose of the present invention is to provide a method, system, device and medium for dynamic optimization of automatic power generation control based on imitation reinforcement learning, which solves the problem that the key operating constraints of automatic power generation control under the traditional artificial intelligence training framework cannot be met.
[0008] The present invention is achieved through the following technical solutions:
[0009] A dynamic optimization method for automatic power generation control based on imitation reinforcement learning, comprising:
[0010] Define the operation constraints and control objectives of AGC dynamic optimization based on the AGC dynamic optimization problem model;
[0011] Based on the operation constraints and control objectives of AGC dynamic optimization, the AGC dynamic optimization problem is modeled as a Markov decision process model, and the core elements of the Markov decision process model include state space, action space and reward function;
[0012] Based on offline expert experience data, the AGC controller is trained by imitation, the strategy is initialized, and the initialization AGC strategy is obtained;
[0013] Based on the initialization AGC strategy, the AGC controller is interactively trained online to solve the strategy of the Markov decision process model, realize the iterative update of the AGC controller, and obtain the optimal AGC strategy.
[0014] Further, the operation constraints of the AGC dynamic optimization include system power balance, AGC unit regulation characteristics, frequency deviation constraints and tie line power deviation constraints;
[0015] The expression of system power balance is:
[0016] in, represents the operating power of adjustable AGC unit i at time t; represents the operating power of the non-adjustable AGC unit i at time t; represents the grid-connected power of the wind farm at time t; represents the injected power on the demand side at time t; represents the power of the tie line at time t; represents the line loss power at time t;
[0017] The frequency deviation constraint expression is: Where Δf t Represents frequency deviation; Δ f Represents the lower limit of frequency deviation; Represents the upper limit of frequency deviation;
[0018] The tie line power deviation constraint expression is:
[0019] in, Represents the tie line power deviation; Δ P tie Represents the lower limit of the tie line power deviation; Represents the upper limit of the tie line power deviation.
[0020] Further, the control objective is to dispatch the AGC unit to meet the minimum economic cost of the ancillary service and the evaluation index of the control performance index;
[0021] The minimum economic cost of the ancillary services is expressed as:
[0022]
[0023] Among them, min f 1 represents the minimum economic cost of ancillary services, c i is the ancillary service cost coefficient of AGC unit i; is the ramp power of the adjustable AGC unit i at time t; u i,t ∈{-1,0,1} indicates the direction of change of power output; represents the operating power of adjustable AGC unit i at time t; Represents the adjustable AGC unit power generation at the initial moment;
[0024] The evaluation indicators of the control performance indicators include CPS1 and CPS2, where CPS1 is used to evaluate the correlation between the frequency deviation of the power system and the regional control deviation; CPS2 is defined as the average regional control deviation within 15 minutes, indicating that the regional control deviation remains within the tolerance range, which is used to ensure that the power exchange between regions does not exceed the specified limit.
[0025] Furthermore, the design of the state space takes into account the current power system operation state and the uncertain new energy generation information that needs to be predicted;
[0026] The current power system operating state includes the current operating power of the adjustable AGC unit i at time t Frequency deviation Δf t , Interconnection line power deviation and regional control deviation
[0027] For the uncertain renewable energy generation information that needs to be predicted, l historical wind power information is introduced into the state space to achieve better prediction. The state definition is as follows:
[0028]
[0029] in, represents the operating power of the current adjustable AGC unit i at time t, Δf t represents the frequency deviation, Represents the tie line power deviation, represents regional control deviation; Represents the power generated by wind power at a historical moment;
[0030] When designing the action space, considering that the control variables are the adjustment direction and adjustment power of each AGC unit, the action is defined as the power adjustment amount of the AGC unit:
[0031] Where, I represents the set of adjustable AGC units; Represents the power adjustment amount of the adjustable AGC unit; in S2, the expression of the reward function is:
[0032] r t =w 1 f 1 +w 2 K cps +w 3 r penalty ;
[0033] Among them, f 1 Represents the power system operation cost index; K cps Represents the performance evaluation index of AGC frequency modulation; r penalty represents the penalty index for violating the power system operation constraints; w 1 、w 2 、w 3 are weight factors, which are used to balance the trade-offs among the three sub-reward items.
[0034] Furthermore, the offline expert experience data is a set of state-action pairs, denoted as (s t ,a t );
[0035] The AGC controller is trained by imitation based on offline expert experience data to complete strategy initialization. The specific process is as follows:
[0036] First, collect offline expert experience data, build a fully connected neural network as the policy network, use the offline expert experience data to train the policy network, and update the policy network parameters φ. The goal is to find an imitation strategy π that best matches the provided state-action pair set. φ ;
[0037] The update of the policy network parameter φ uses maximum likelihood estimation, expressed as:
[0038]
[0039] Among them, φ * represents the optimal strategy network parameters; a i |s i Representative in s i The action probability under the state; t represents the time, t = 0, 1, 2...T; T represents the update step size.
[0040] Further, the update strategy network parameter φ is specifically:
[0041] The Adam stochastic gradient descent method is used to iteratively update the policy network parameters φ.
[0042] Furthermore, the SAC algorithm is used to solve the strategy of the Markov decision process model, and the solution goal is to maximize the expected reward;
[0043] The expression of reward expectation is:
[0044]
[0045] Among them, r t represents the reward function, H(π(·|s t )) is the equilibrium entropy, α is the hyperparameter of the equilibrium entropy and reward feedback parameter; J π represents reward expectation; s t Represents the state at time t;
[0046] The dual Q network is introduced into the SAC algorithm. The policy network parameter φ is continuously updated through the Adam stochastic gradient descent method. When it converges to the maximum reward expectation, the optimal AGC strategy is obtained. The optimal AGC strategy is used to control the AGC unit, output the control instructions of each AGC unit, and complete the automatic power generation control.
[0047] The present invention also discloses an automatic power generation control dynamic optimization system based on imitating expert experience combined with reinforcement learning, comprising:
[0048] AGC dynamic optimization problem definition module, used to define the operation constraints and control objectives of AGC dynamic optimization based on the AGC dynamic optimization problem model;
[0049] A modeling module is used to model the AGC dynamic optimization problem as a Markov decision process model based on the operation constraints and control objectives of the AGC dynamic optimization. The core elements of the Markov decision process model include a state space, an action space, and a reward function.
[0050] The imitation training module is used to perform imitation training on the controller based on offline expert experience data, complete strategy initialization, and obtain the initialized AGC strategy;
[0051] The online interactive training module is used to perform online interactive training on the AGC controller based on the initialization AGC strategy, solve the strategy of the Markov decision process model, realize iterative update of the AGC controller, and obtain the optimal AGC strategy.
[0052] The present invention also discloses a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the automatic power generation control dynamic optimization method based on imitation reinforcement learning are implemented.
[0053] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the automatic power generation control dynamic optimization method based on imitation reinforcement learning are implemented.
[0054] Compared with the prior art, the present invention has the following beneficial technical effects:
[0055] The present invention discloses a dynamic optimization method for automatic power generation control based on imitation reinforcement learning to ensure that the key operating constraints of the power system can be met during the training process. Considering that "trial and error" usually occurs in the initial stage of AGC controller training, because the initialization strategy is usually randomly generated. In order to avoid random exploration in the early stage of reinforcement learning, the present invention proposes to adopt imitation learning, first obtain an initialization strategy similar to expert experience based on offline training, then initialize the controller based on the imitation strategy, and finally use the SAC algorithm to further train online to obtain the optimal AGC strategy. The proposed imitation learning framework based on expert experience can improve the performance of the initial strategy, so that the early exploration strategy can imitate the historical experience of experts, thereby avoiding random exploration caused by the initial strategy being too poor, and avoiding serious constraint violations (such as frequency not meeting upper and lower limits) during the training process. Based on a mode of combining online training and offline training in stages, the present invention not only uses offline data to ensure the safety of system constraints, but also considers online training, avoiding the phenomenon of controller invalidity caused by errors between the simulation system and the actual system; introducing imitation learning given expert experience in the SAC algorithm framework of traditional reinforcement learning, effectively improving the performance of the initial strategy, avoiding random unsafe exploration caused by poor strategies in the early stage, and reducing the risk of system frequency default.
[0056] Furthermore, when constructing the Markov decision model, the AGC dynamic optimization process with a step size of 1 minute within 15 minutes is considered, and a reinforcement learning algorithm is introduced to achieve real-time dynamic prediction and early response to fluctuations in a high proportion of renewable energy, without relying on precise system models, making the power system more adaptively controllable.
[0057] Furthermore, compared to traditional RL algorithms, SAC uses a stochastic strategy that encourages exploration by adding entropy to the reward. Therefore, SAC is less likely to fall into local optimality and can better explore the action space.
[0058] In order to improve the stability and accuracy of value estimation, the SAC algorithm introduces a dual Q network, which uses the minimum value of the two Q networks to evaluate the Q function, which helps to alleviate the overestimation bias that may occur in the Q learning process and avoids the overestimation problem that may exist in the evaluation of a single Q function. Using two Q networks, the policy network parameter φ is continuously updated through the gradient descent method. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a specific implementation flow chart of an automatic power generation control dynamic optimization method based on imitation reinforcement learning of the present invention;
[0060] Figure 2 It is the AGC dynamic optimization problem model diagram of the present invention;
[0061] Figure 3 This is a framework diagram of a secure reinforcement learning algorithm combined with imitation learning of the present invention;
[0062] Figure 4 It is a core algorithm flow chart of an automatic power generation control dynamic optimization method based on imitation reinforcement learning of the present invention;
[0063] Figure 5 This is a principle block diagram of an automatic power generation control dynamic optimization system based on imitation reinforcement learning of the present invention. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical solution and advantages of the present invention more clear, the following is further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention, that is, the embodiments described are only part of the embodiments of the present invention, not all embodiments.
[0065] The detailed description of the embodiments of the present invention provided in the following drawings is not intended to limit the scope of the invention claimed for protection, but merely represents a selected embodiment of the present invention. Based on the drawings and embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0066] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0067] Example 1
[0068] like Figure 1 As shown, the present invention discloses a dynamic optimization method for automatic power generation control based on imitation reinforcement learning, comprising the following steps:
[0069] S1. Define the operation constraints and control objectives of AGC dynamic optimization based on the AGC dynamic optimization problem model;
[0070] S2. Based on the operation constraints and control objectives of AGC dynamic optimization, the AGC dynamic optimization problem is modeled as a Markov decision process model, wherein the core elements of the Markov decision process model include a state space, an action space, and a reward function;
[0071] S3, collecting offline expert experience data as training data, imitating and training the AGC controller based on the offline expert experience data, completing strategy initialization, and obtaining an initialized AGC strategy;
[0072] S4. Based on the initialization AGC strategy, the AGC controller is interactively trained online to solve the strategy of the Markov decision process model, implement iterative updates of the AGC controller, and obtain the optimal AGC strategy.
[0073] Example 2
[0074] Based on Example 1, the process of each step is described in detail.
[0075] like Figure 4 As shown, the present invention discloses a dynamic optimization method for automatic power generation control based on imitation reinforcement learning, comprising the following steps:
[0076] Step 1: Define the goals and constraints of the AGC dynamic optimization model
[0077] First, build Figure 2 The AGC dynamic optimization problem model diagram includes an AGC controller, a saturation module, a rate limiter, and a wind power grid-connected unit. The AGC controller, the saturation module, and the rate limiter are connected in sequence. The decision variables of the AGC controller are After passing through the limiting module and current limiter, the final power adjustment is output and then compared with the wind power change of the wind power grid-connected unit. Added together, they affect the system frequency through the power grid system.
[0078] The fluctuation of system frequency Δf will affect the power of inter-regional interconnection lines ΔP tie and conventional regional control error |e ACE |Change; AGC controller is based on the control error of the conventional area |e ACE |Monitoring and giving control instructions for the adjustable AGC unit After being limited by the limiting module and the climbing function, it is finally updated to the power system to stabilize the frequency of the power system.
[0079] The limiter module imposes upper and lower limits on the output value. When the output value exceeds the upper limit, the output is limited to the upper limit; when the output value is lower than the lower limit, the output is limited to the lower limit; when it is between the upper and lower limits, the original output is maintained. The limiter module contains one input port and one output port by default.
[0080] According to the dynamic process of power system operation, and then according to how the key variables are dynamically changed, the system dynamic differential equations of the key variables are derived. Among them, the key variables include system frequency f, tie line power P tie and regional control error e ACE .
[0081] In the present invention, the control objective is to schedule the AGC unit to meet the minimum economic cost of the auxiliary service and the evaluation index of the control performance index.
[0082] Therefore, the minimum economic cost of ancillary services can be expressed as:
[0083]
[0084] Among them, min f 1 represents the minimum economic cost of ancillary services, c i is the ancillary service cost coefficient of AGC unit i; is the ramp power of the adjustable AGC unit i at time t; u i,t ∈{-1,0,1} indicates the direction of change of power output; represents the operating power of adjustable AGC unit i at time t; Represents the operating power of the adjustable AGC unit at the initial moment.
[0085] The evaluation indicators of the Control Performance Standard (CPS) include CPS1 and CPS2, where CPS1 is used to evaluate the correlation between the frequency deviation of the power system and the area control error (ACE), while CPS2 is defined as the average ACE within 15 minutes, indicating that ACE is maintained within the tolerance range to ensure that the power exchange between regions does not exceed the specified limit.
[0086] In addition, when constructing the AGC dynamic optimization problem model, it is necessary to characterize the power system operation constraints as equality or inequality.
[0087] The operational constraints of AGC dynamic optimization include system power balance, AGC unit regulation characteristics (such as the control signal will be subject to saturation function and ramp rate constraints before execution), frequency deviation constraints, and tie line power deviation constraints.
[0088] The expression of system power balance is:
[0089] The frequency deviation constraint expression is: Where Δf t Represents frequency deviation; Δ f Represents the lower limit of frequency deviation; Represents the upper limit of frequency deviation;
[0090] The tie line power deviation constraint expression is: Where, represents the tie line power deviation; Δ P tie Represents the lower limit of the tie line power deviation; Represents the upper limit of the tie line power deviation.
[0091] in, represents the operating power of adjustable AGC unit i at time t; represents the operating power of the non-adjustable AGC unit i at time t; represents the grid-connected power of the wind farm at time t; represents the injected power on the demand side at time t; represents the power of the tie line at time t; represents the line loss power at time t; Δf t represents frequency deviation; Represents the tie line power deviation.
[0092] Step 2. Modeling as a Markov decision process
[0093] According to the AGC problem definition, the core elements of the Markov decision process (MDP) model corresponding to the control problem, such as state space, action space, reward function, etc., are obtained to obtain a sequential decision mathematical model equivalent to the original AGC dynamic optimization problem.
[0094] Specifically, 1) define the state space: The design of the state space needs to capture the necessary information from two aspects: the current operating status of the power system and the uncertain renewable energy generation information that needs to be predicted.
[0095] For the current power system operating state, the present invention considers four factors, including the operating power of the current adjustable AGC unit i at time t Frequency deviation Δf t , Interconnection line power deviation and regional control deviation
[0096] For the uncertain new energy generation information that needs to be predicted, the present invention only takes wind power prediction as an example, and introduces l historical wind power information into the state space to achieve better prediction. The state definition is as follows:
[0097]
[0098] in, Represents the power generated by wind power at a historical moment.
[0099] 2) Define the action space: In the AGC dynamic optimization problem, the control variables are the adjustment direction and adjustment power of each AGC unit. In order to simplify the action space with a smaller action dimension, the action is defined as the power adjustment amount of the AGC unit: Where, I represents the set of adjustable AGC units; Represents the power adjustment amount of the adjustable AGC unit.
[0100] 3) Define the reward function: The reward design of MDP should take into account the goals and constraints. Here, the reward function r is designed from three aspects: t =w 1 f 1 +w 2 K cps +w 3 r penalty ,f 1 Represents the power system operation cost index; K cps Represents the performance evaluation index of AGC frequency modulation; r penalty represents the penalty index for violating the power system operation constraints; here w 1 ,w 2 ,w 3 They are weight factors, which are used to balance the trade-offs between the three sub-reward items. penalty Used to penalize total violations beyond upper and lower limits, such as output power, slope power, tie-line power and frequency deviation limits in step 1.
[0101] Step 3. Initialize the AGC strategy based on imitation learning offline training
[0102] like Figure 3 As shown in Figure 1, the imitation learning method can directly imitate the demonstrator (i.e., expert experience) without interacting with the real environment, and then use a classifier or regressor to copy the expert's strategy based on the pre-collected state-action training data. Therefore, the learning goal of the AGC controller in the offline training phase is to obtain an imitation strategy as the initial strategy π 0 Given a pre-collected set of state-action pairs (s t ,a t ), the agent’s goal is to find an imitation policy π that best matches the provided set of state-action pairs φ , the policy network parameters φ are updated using maximum likelihood estimation:
[0103]
[0104] Considering that the action space designed by the present invention is continuous, it is assumed that the strategy of the present invention follows Gaussian distribution in each action dimension. The present invention uses a fully connected neural network to approximate the strategy π, and then uses the Adam stochastic gradient descent method to update the neural network parameters and solve the optimal strategy.
[0105] Step 4. Combine reinforcement learning with SAC algorithm online training to obtain the optimal AGC strategy
[0106] In order to find the optimal strategy π * ,The present invention adopts the SAC algorithm, which can maximize the ,accumulated reward of the agent while satisfying the safety constraints.
[0107] The SAC algorithm is based on the Adam stochastic gradient descent method and learns strategies by maximizing expected returns. It uses two neural networks: the Actor network and the Critic network. The Actor network is used to generate actions and select the optimal action based on the current state and policy parameters; the Critic network is used to estimate the state value function and the state-action value function.
[0108] In SAC, the policy is a probability distribution, the agent collects data through sampling, and uses this data to update the policy network parameters. SAC adopts the Softmax strategy to parameterize the policy network as a weighted sum of a set of basis functions, allowing the agent to explore the environment and obtain more information. Compared with traditional RL algorithms, SAC uses a random strategy to encourage exploration by adding entropy to rewards. Therefore, SAC is not prone to falling into local optimality and can better explore the action space.
[0109] The SAC algorithm is used to solve the strategy of the Markov decision process model. The solution goal is to maximize the expected reward, and then the optimal AGC strategy is obtained; the expression of the expected reward is:
[0110]
[0111] Among them, r t represents the reward function, H(π(·|s t )) is the equilibrium entropy, α is the hyperparameter of the equilibrium entropy and reward feedback parameter; J π represents reward expectation; s t Represents the state at time t.
[0112] In addition, in order to improve the stability and accuracy of value estimation, the SAC algorithm introduces a dual Q network, which uses the minimum value of the two Q networks to evaluate the Q function, which helps to alleviate the overestimation bias that may occur in the Q learning process and avoids the overestimation problem that may exist in the evaluation of a single Q function. Using two Q networks, the policy network parameters φ are continuously updated through the gradient descent method.
[0113] According to the real-time online interactive training results of the previous step, the AGC controller is continuously corrected through feedback, and the AGC controller is used to automatically adjust parameters to continuously balance between random exploration training and experience utilization optimization until the AGC controller converges to maximize the reward expectation, thus obtaining the optimal AGC strategy.
[0114] The trained reinforcement learning AGC controller is deployed to the regional power grid system to monitor in real time and output control instructions for each AGC unit based on the dynamic operating frequency of the power system to complete automatic power generation control.
[0115] Example 3
[0116] like Figure 5 As shown, the present invention also discloses an automatic power generation control dynamic optimization system based on imitating expert experience combined with reinforcement learning, comprising:
[0117] AGC dynamic optimization problem definition module, used to define the operation constraints and control objectives of AGC dynamic optimization based on the AGC dynamic optimization problem model;
[0118] A modeling module is used to model the AGC dynamic optimization problem as a Markov decision process model based on the operation constraints and control objectives of the AGC dynamic optimization. The core elements of the Markov decision process model include a state space, an action space, and a reward function.
[0119] The imitation training module is used to perform imitation training on the controller based on offline expert experience data, complete strategy initialization, and obtain the initialized AGC strategy;
[0120] The online interactive training module is used to perform online interactive training on the AGC controller based on the initialization AGC strategy, solve the strategy of the Markov decision process model, realize iterative update of the AGC controller, and obtain the optimal AGC strategy.
[0121] Example 4
[0122] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the automatic power generation control dynamic optimization method based on imitation reinforcement learning are implemented. The memory may include a memory, such as a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk memory, etc. The processor, the network interface, and the memory are interconnected through an internal bus, and the internal bus may be an industrial standard architecture bus, a peripheral component interconnection standard bus, an extended industrial standard structure bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include a program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0123] Example 5
[0124] The present invention also discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the automatic power generation control dynamic optimization method based on imitation reinforcement learning are implemented. Specifically, the computer-readable storage medium includes, but is not limited to, for example, a volatile memory and / or a non-volatile memory. The volatile memory may include a random access memory and / or a cache memory, etc. The non-volatile memory may include a read-only memory, a hard disk, a flash memory, an optical disk, a magnetic disk, etc.
[0125] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0126] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.
[0127] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0128] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.
Claims
1. A dynamic optimization method for automatic power generation control based on imitation reinforcement learning, characterized in that: include: Define the operation constraints and control objectives of AGC dynamic optimization based on the AGC dynamic optimization problem model; Based on the operation constraints and control objectives of AGC dynamic optimization, the AGC dynamic optimization problem is modeled as a Markov decision process model, and the core elements of the Markov decision process model include state space, action space and reward function; Based on offline expert experience data, the AGC controller is trained by imitation, the strategy is initialized, and the initialization AGC strategy is obtained; Based on the initialization AGC strategy, the AGC controller is interactively trained online to solve the strategy of the Markov decision process model, realize the iterative update of the AGC controller, and obtain the optimal AGC strategy.
2. According to claim 1, a dynamic optimization method for automatic power generation control based on imitation reinforcement learning is characterized in that: The operation constraints of the AGC dynamic optimization include system power balance, AGC unit adjustment characteristics, frequency deviation constraints and tie line power deviation constraints; The expression of system power balance is: in, represents the operating power of adjustable AGC unit i at time t; represents the operating power of the non-adjustable AGC unit i at time t; represents the grid-connected power of the wind farm at time t; represents the injected power on the demand side at time t; represents the power of the tie line at time t; represents the line loss power at time t; The frequency deviation constraint expression is: Where Δf t Represents frequency deviation; Δ f Represents the lower limit of frequency deviation; Represents the upper limit of frequency deviation; The tie line power deviation constraint expression is: in, Represents the tie line power deviation; Δ P tie Represents the lower limit of the tie line power deviation; Represents the upper limit of the tie line power deviation.
3. The method for dynamic optimization of automatic power generation control based on imitation reinforcement learning according to claim 1, characterized in that: The control objective is to dispatch the AGC unit to meet the minimum economic cost of the ancillary services and the evaluation index of the control performance index; The minimum economic cost of the ancillary services is expressed as: Where min f1 represents the minimum economic cost of ancillary services, c i is the ancillary service cost coefficient of AGC unit i; is the ramp power of the adjustable AGC unit i at time t; u i,t ∈{-1,0,1} indicates the direction of change of power output; represents the operating power of adjustable AGC unit i at time t; Represents the adjustable AGC unit power generation at the initial moment; The evaluation indicators of the control performance indicators include CPS1 and CPS2, where CPS1 is used to evaluate the correlation between the frequency deviation of the power system and the regional control deviation; CPS2 is defined as the average regional control deviation within 15 minutes, indicating that the regional control deviation remains within the tolerance range, which is used to ensure that the power exchange between regions does not exceed the specified limit.
4. The method for dynamic optimization of automatic power generation control based on imitation reinforcement learning according to claim 1, characterized in that: The design of the state space takes into account the current power system operation status and the uncertain new energy generation information that needs to be predicted; The current power system operating state includes the current operating power of the adjustable AGC unit i at time t Frequency deviation Δf t , Interconnection line power deviation and regional control deviation For the uncertain renewable energy generation information that needs to be predicted, l historical wind power information is introduced into the state space to achieve better prediction. The state definition is as follows: in, represents the operating power of the current adjustable AGC unit i at time t, Δf t represents the frequency deviation, Represents the tie line power deviation, represents regional control deviation; Represents the power generated by wind power at a historical moment; When designing the action space, considering that the control variables are the adjustment direction and adjustment power of each AGC unit, the action is defined as the power adjustment amount of the AGC unit: Where I represents the set of adjustable AGC units; Represents the power adjustment amount of the adjustable AGC unit; in S2, the expression of the reward function is: r t =w1f1+w2K Cps +w3r penalty ; Among them, f1 represents the power system operation cost index; K cps Represents the performance evaluation index of AGC frequency modulation; r penalty represents the penalty indicator for violating the power system operation constraints; w1, w2, and w3 are weight factors used to balance the trade-offs among the three sub-reward items.
5. The method for dynamic optimization of automatic power generation control based on imitation reinforcement learning according to claim 4 is characterized in that: The offline expert experience data is a set of state-action pairs, denoted as (s t , a t ); The AGC controller is trained by imitation based on offline expert experience data to complete strategy initialization. The specific process is as follows: First, collect offline expert experience data, build a fully connected neural network as the policy network, use the offline expert experience data to train the policy network, and update the policy network parameters φ. The goal is to find an imitation strategy π that best matches the provided state-action pair set. φ ; The update of the policy network parameter φ uses maximum likelihood estimation, expressed as: Among them, φ* represents the optimal strategy network parameters; a i |s i Representative in s i The action probability under the state; t represents the time, t = 0, 1, 2…T; T represents the update step size.
6. The method for dynamic optimization of automatic power generation control based on imitation reinforcement learning according to claim 5, characterized in that: The update strategy network parameter φ is specifically: The Adam stochastic gradient descent method is used to iteratively update the policy network parameters φ.
7. The method for dynamic optimization of automatic power generation control based on imitation reinforcement learning according to claim 1, characterized in that: The SAC algorithm is used to solve the strategy of the Markov decision process model, and the solution goal is to maximize the expected reward; The expression of reward expectation is: Among them, r t represents the reward function, H(π(·|s t )) is the equilibrium entropy, α is the hyperparameter of the equilibrium entropy and reward feedback parameter; J π represents reward expectation; s t Represents the state at time t; The dual Q network is introduced into the SAC algorithm. The policy network parameter φ is continuously updated through the Adam stochastic gradient descent method. When it converges to the maximum reward expectation, the optimal AGC strategy is obtained. The optimal AGC strategy is used to control the AGC unit, output the control instructions of each AGC unit, and complete the automatic power generation control.
8. An automatic power generation control dynamic optimization system based on imitating expert experience combined with reinforcement learning, characterized in that: include: AGC dynamic optimization problem definition module, used to define the operation constraints and control objectives of AGC dynamic optimization based on the AGC dynamic optimization problem model; A modeling module is used to model the AGC dynamic optimization problem as a Markov decision process model based on the operation constraints and control objectives of the AGC dynamic optimization. The core elements of the Markov decision process model include a state space, an action space, and a reward function. The imitation training module is used to perform imitation training on the controller based on offline expert experience data, complete strategy initialization, and obtain the initialized AGC strategy; The online interactive training module is used to perform online interactive training on the AGC controller based on the initialization AGC strategy, solve the strategy of the Markov decision process model, realize iterative update of the AGC controller, and obtain the optimal AGC strategy.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the automatic power generation control dynamic optimization method based on imitation reinforcement learning as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the automatic power generation control dynamic optimization method based on imitation reinforcement learning as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Intelligent power generation control method of virtual wolf cluster control strategy under island intelligent power distribution network
CN107589672A
Power system active power flow online optimization control method, storage medium and device
CN115293052A
Cited By
Data driving power distribution method and system for H-bridge module health degree balance
CN121030204A