PPO-based Wind Turbine Generator Power Prediction Control System
By adopting a control system based on PPO algorithm in the wind turbine group, the effective interaction between the wind turbine power prediction system model and the real-time operating state of the unit is solved, and the problem of mismatch between the model and the system in the existing technology is solved, and the accuracy of the wind turbine power prediction is improved.
Patent Information
- Application Number
- CN202111403350.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The existing wind turbine power prediction system model does not interact effectively with the real-time operating state of the unit, resulting in the mismatch between the model and the system, reducing the power prediction accuracy.
The wind turbine power prediction control system based on the PPO (near-end strategy optimization) algorithm is adopted to obtain historical and real-time power data and meteorological data through the operation data acquisition unit, establish a data-driven model, and use the PPO algorithm module for real-time learning and feedback to update the power prediction system model.
Through the PPO algorithm, the learning feedback on the real-time operation information of the wind turbine unit is completed, and the power prediction system model is improved, which improves the accuracy of unit power prediction.
Smart Images

Figure CN114060234B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wind turbine power prediction, and more specifically, to a PPO-based wind turbine power prediction control system. Background Art
[0002] The advantages of wind energy in terms of economy and environment make its utilization play a crucial role in social development. To reduce the threat posed by the randomness of wind energy to the power system, improving the power prediction accuracy of wind turbines is crucial for expanding their installed capacity. An accurate system model is the basis for power prediction of the turbines. Most of the existing wind turbine power prediction system models are established offline based on data or mechanisms and have no connection with the real-time operating state of the turbines, resulting in the problem of model-system mismatch. The deep reinforcement learning algorithm has the characteristic of being able to continuously and effectively learn the environment and has been successfully applied in many aspects of artificial intelligence. Therefore, by combining PPO in the deep reinforcement learning algorithm family to complete the interaction between the real-time operation information of wind turbines and their power prediction system models, the power prediction accuracy of the turbines can be improved. Summary of the Invention
[0003] The present invention aims to solve the problem of information interaction between the wind turbine power prediction model and the actual operating state of the turbine and improve the power prediction accuracy of the turbine.
[0004] A PPO-based wind turbine power prediction control system, which includes a wind turbine, an operating data acquisition unit, a power prediction model, and a PPO algorithm module; the wind turbine is driven by wind to output electricity, and its working state changes under the influence of wind speed; it is characterized in that
[0005] The operating data acquisition unit is connected to the wind turbine to acquire the historical power data and real-time power data of the wind turbine, and also acquire meteorological data;
[0006] Based on the meteorological data and the historical power data of the wind turbine, data modeling is carried out to obtain a data-driven model of the power prediction system;
[0007] The PPO algorithm module receives the real-time power data of the wind turbine as the actual operating state information and feeds back the effective learning result to the data-driven model of the power prediction system;
[0008] After the data-driven model of the power prediction system of the turbine receives the feedback information from the PPO algorithm module, it further compares the power prediction result with the actual power of the turbine. If the prediction error exceeds the threshold, the data-driven model of the power prediction system is updated.
[0009] Advantages of the Present Invention
[0010] 1. The sampling PPO algorithm conducts real-time learning on the real-time operating status of wind turbines, promoting the interaction between the power prediction system and the actual operating information of the turbines.
[0011] 2. Based on the learning feedback of the PPO algorithm on the real-time operating information of wind turbines, the power prediction system model is updated to improve the power prediction accuracy of the turbines. Description of the Drawings
[0012] Figure 1 Wind turbine power prediction system combined with the PPO algorithm;
[0013] Figure 2 Overall framework of the PPO algorithm; Detailed Implementation Modes
[0014] The present invention will be further described below in conjunction with the drawings. It should be understood that the content described herein is only used to illustrate and explain the present invention and is not used to limit the present invention.
[0015] List of Abbreviations, English and Definitions of Key Terms
[0016] 1. PPO: Proximal Policy Optimization
[0017] 2. agent: Agent
[0018] The embodiments will be described based on the operating data of a 100 MW wind farm in the past 3 days.
[0019] The scheme for optimizing the wind turbine power prediction system model based on the PPO algorithm and ultra-short-term operating data is as Figure 1 shown.
[0020] As Figure 1 shown, the wind turbine power prediction control system based on PPO includes a wind turbine, an operating data acquisition unit, a power prediction model, a PPO algorithm module, and a unit for optimizing the operation and maintenance of the turbine.
[0021] The wind turbine is driven by wind to output electricity, and its operating state changes affected by the wind speed.
[0022] The operating data acquisition unit is connected to the wind turbine to obtain the historical power data and real-time power data of the wind turbine. It can also obtain meteorological data, such as wind speed, through sensors or network servers.
[0023] Based on the meteorological data and the historical power data of the wind turbine, data modeling is carried out. Based on the historical actual operating data of the wind turbine in the past 3 days, a data-driven model of the turbine power prediction system is obtained through the subspace identification method.
[0024] Based on the identification of the 3-day ultra-short-term historical operation data of the wind turbine generator set, a data-driven model of its power prediction system can be obtained. This model is the model of the system at past moments and cannot fully reflect the real-time operation state of the unit. It is necessary to correct it by combining the PPO algorithm.
[0025] The PPO algorithm module receives the real-time power data of the wind turbine generator set, etc. as the actual operation state information, and the intelligent agent in the PPO algorithm conducts online effective learning on the actual operation state information of the wind turbine generator set, and feeds the effective learning result back to the data-driven model of the power prediction system.
[0026] The PPO algorithm is a kind of machine learning algorithm. Through the continuous active learning of the environmental information by the abstract intelligent agent in the algorithm, it outputs action information to the data-driven model of the power prediction system. By learning the real-time power data of the unit through the PPO algorithm intelligent agent, the real-time power data, wind speed and other environmental information that are beneficial to improving the power prediction accuracy after cleaning are fed back to the data-driven model of the power prediction system.
[0027] The power prediction system model of the unit receives the feedback information of the PPO algorithm module, that is, the action information of the intelligent agent, such as the real-time output power of the unit after cleaning, wind speed, pitch angle, etc. data that are beneficial to improving the accuracy of the data-driven model of the power prediction system. And further compares the power prediction result with the actual power of the unit. If the prediction error exceeds the threshold, the power prediction system model is updated.
[0028] The updated power prediction system model outputs power prediction information to the unit operation and maintenance optimization unit, and then adjusts the operation of the wind turbine generator set.
[0029] Repeat the above iterative update, and finally make the power prediction system model infinitely close to the real-time operation condition of the unit, so as to achieve the effect of improving the power prediction of the unit.
[0030] PPO algorithm module
[0031] Research shows that reinforcement learning (RL) algorithms are suitable for solving various complex stochastic problems, such as the frequency control of microgrids and the long-delay control loops of thermal power plants. However, RL algorithms are found to be ineffective in dealing with high-dimensional problems, that is, as the number of inputs increases, the computational complexity increases sharply, making it difficult to find good strategies in large state spaces. Therefore, it is necessary to find appropriate improvement schemes to expand their scope of application. Deep learning algorithms approximate arbitrary nonlinear functions by training deep neural networks to achieve the goal of learning the inherent laws and essential features of input data. Due to its powerful representation ability, deep learning has been widely and successfully applied in the field of artificial intelligence. Therefore, by combining reinforcement learning and deep learning, the deep reinforcement learning (DRL) algorithm with fast and efficient computing and strong fitting ability has emerged. Proximal Policy Optimization (PPO) is a new policy-based DRL algorithm that is less sensitive to hyperparameters and can avoid a large number of policy updates and poor action selections.
[0032] The overall framework of the PPO algorithm is as Figure 2 shown.
[0033] In DRL, an agent is a structure of a deep neural network (DNN). By setting the reward value and feeding it back to the agent to guide its effective actions, the learning task can be completed quickly and accurately. Correspondingly, in a wind turbine, a reward mechanism combined with the power prediction error is used to guide the agent to learn the real-time power generation, wind speed, and pitch angle of the turbine, and the learning results (cleaned real-time output power, wind speed, etc.) are fed back to the data-driven model of the wind turbine power prediction system in the form of agent actions, thereby improving the power prediction accuracy. The DNN consists of an input layer, an output layer, and several hidden layers. All layers are fully connected neural networks with parameters θ (weight matrix and bias vector). The state information of the wind turbine operation is input into the input layer of the DNN. Its output is the action or action value. PPO is a deep reinforcement learning algorithm with an actor-critic architecture. The actor and critic are two deep fully connected neural networks (DNNs) with weights θ μ and θ Q respectively. The actor network μ is used to estimate the policy function π(a|s,θ μ ). The critic network Q is used to estimate the value function V(s). The actor policy π(a|s,θ μ ) follows a normal distribution π(a|s) ∼ N(μ,σ 2 ). The output mean of the actor network is μ, and the standard deviation is σ. When the agent executes the policy, the specific action is sampled from the normal distribution.
[0034] The parameters between the actor and critic networks are not shared, so they are updated separately. During training, the actor network μ outputs continuous actions based on the real-time power generation data of the unit. Then, the critic network maps the state to a scalar V(s) to measure the quality of the state. The agent explores the environment periodically in the Markov decision process (MDP) and updates the DNN parameters for each epoch. At the end of each epoch (with length T), a series of unit operating condition information, the agent's actions, and reward values within the set constitute the trajectory τ = {s 1 ,a 1 ,r 1 ,s 2 ,a 2 ,r 2 ,…,s t ,a t ,r t ,…,s T ,a T ,r T}, which is considered to collect the unit power generation state information at past moments. Among them, s 1 represents the unit output power; a 1 is the wind speed; r 1 is the unit pitch angle.
[0035] The purpose of training the actor network with parameters θ μ is to maximize the cumulative reward value of the agent under the policy π, that is Based on the policy gradient method, the parameters θ μ of the actor network quality management are updated. The policy gradient algorithm relies on the sampled decision sequence when interacting with the real-time operating state of the unit, and uses the method of gradient ascent to update the policy parameters. The sampled policy loss function and its gradient are:
[0036]
[0037]
[0038] The gradient can improve the policy and develop in the direction of increasing the action probability and obtaining a larger reward value.
[0039] Adopting the actor-critic architecture has a significant impact on reducing the variance of gradient estimation by reconstructing the reward signal according to the advantage. Therefore, we define the advantage function to judge how much better the agent's action a is than the mean.
[0040] A(s,a) = Q(s,a) - V(a) (3)
[0041] Among them, the value function (Q-value) of an action is used to evaluate the quality of the policy π, and its definition is as follows:
[0042]
[0043] Correspondingly, the state value function for evaluating the quality of a certain policy state is defined as follows:
[0044]
[0045] It can be seen from the above definition that the value function V(s) is the long-term return under the state s and the policy π.
[0046] Actually, the actor network can be updated at each time step. And for the policy gradient method, after each step of optimization, it is necessary to resample from the new policy because the old samples are not suitable for the new policy. This will increase the training time. To solve this problem, another actor network μ' is added in the PPO algorithm, and this network is used to interact with the operating state of the wind turbine to obtain the trajectory. The actor network θ μ′ is called the old actor, and its parameter qm0 remains unchanged in a set. The actor network μ (also called the new actor) can use the trajectory to learn the policy multiple times. Therefore, the actor network μ of the PPO algorithm is updated using the piecewise agent loss function shown below:
[0047]
[0048] Piecewise
[0049]
[0050]
[0051] Among them, the core parameter ε ∈ (0, 1) is used to limit the probability difference; is the estimated value of the advantage function; α A is the learning rate of the actor network; p(θ) is the probability ratio between μ and μ', which is used to measure the difference between the two policies.
[0052] During the process of the agent learning the operating state information of the wind turbine, as the actor μ is updated, the old and new policies will diverge, increasing the variance of the estimate. Therefore, the old policy is updated periodically to match the new policy. If p(θ) falls outside the range [1 - ε, 1 + ε], the advantage function will be segmented, which helps to ensure the stability of the algorithm. The algorithm neither tends too greedily towards actions with positive advantages nor avoids too quickly actions with negative advantages from a small sample set. The minimum operating unit ensures that the agent objective function is a lower bound of the non-segmented objective.
[0053] A Generalized Advantage Estimation (GAE) as shown in the following formula is adopted:
[0054]
[0055] where, δ t = r t + γV(s t+1 ) - V(s t ).
[0056] Based on the temporal difference theory, the critic network is updated using the minimum loss function:
[0057] L(θ Q ) = E[(q t - V(s t )) 2 (11)
[0058] q t = r t + γV(s t+1 ) (12)
[0059]
[0060] where L(θ Q ) is the loss function; α Q represents the learning rate of the critic network. In fact, the training of the critic is a process of minimizing the difference between q t and V(s t ).
[0061] Finally, it should be noted that the above description is only an explanation of the present invention and is not used to limit the present invention. Although the present invention has been described in detail, those skilled in the art can still modify the aforementioned technical solutions or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A power prediction and control system for a wind turbine based on PPO, which includes a wind turbine, an operation data acquisition unit, a power prediction model, and a PPO algorithm module; the wind turbine is driven by wind to output electricity, and its working state changes affected by the wind speed. It is characterized in that the operation data acquisition unit is connected to the wind turbine to obtain the historical power data and real-time power data of the wind turbine, and also obtains meteorological data. Based on the meteorological data and the historical power data of the wind turbine, data modeling is carried out to obtain a data-driven model of the power prediction system. The PPO algorithm module receives the real-time power data of the wind turbine as the actual operation state information, and feeds back the effective learning result to the data-driven model of the power prediction system. After the data-driven model of the power prediction system of the unit receives the feedback information from the PPO algorithm module, it further compares the power prediction result with the actual power of the unit. If the prediction error exceeds the threshold, the data-driven model of the power prediction system is updated.
2. The power prediction and control system for a wind turbine based on PPO according to claim 1, It is characterized in that the meteorological data is obtained by sensor acquisition or network server method.
3. The power prediction and control system for a wind turbine based on PPO according to claim 1, It is characterized in that a data-driven model of its power prediction system is obtained by subspace identification method based on the ultra-short-term historical operation data of the wind turbine in the past several days.
4. The power prediction and control system for a wind turbine based on PPO according to claim 1, It is characterized in that the PPO algorithm module performs online effective learning on the actual operation state information of the wind turbine through the agent in the PPO algorithm, and feeds back the real-time power data and wind speed environment information that are beneficial to improving the power prediction accuracy after cleaning to the data-driven model of the power prediction system.
5. The power prediction and control system for a wind turbine based on PPO according to claim 1, It is characterized in that it further includes a unit operation and maintenance optimization unit, which uses the updated data-driven model of the power prediction system to output power prediction information to the unit operation and maintenance optimization unit, and then adjusts the operation of the wind turbine.
6. The power prediction and control system for a wind turbine based on PPO according to claim 5, It is characterized in that iterative update is repeatedly executed to make the data-driven model of the power prediction system close to the real-time operation status of the unit.
7. The power prediction and control system for a wind turbine based on PPO according to claim 1, It is characterized in that the agent in the PPO algorithm is a structure of a deep neural network. By setting the reward value and feeding it back to the agent to guide its effective actions, the learning task can be completed quickly and accurately. The deep neural network consists of an input layer, an output layer and several hidden layers; the state information of the environment is input into the input layer of the deep neural network, and the output is an action or an action value.
8. The power prediction and control system for a wind turbine based on PPO according to claim 7, It is characterized in that The PPO algorithm has an actor-critic architecture. The actor and critic are two deep neural networks with parameters θ μ and θ Q respectively. The actor network μ is used to estimate the policy function π(a|s,θ μ ). The critic network Q is used to estimate the value function V(s). The actor policy π(a|s,θ μ ) follows a normal distribution π(a|s) ~ N(μ,σ 2 ). The output mean of the actor network is μ and the standard deviation is σ.
Citation Information
Patent Citations
Generated power prediction apparatus, server, computer program, and generated power prediction method
JP2018007312A
Reducing curtailment of wind power generation
US10041475B1