Yaw pitch composite control parameter self-optimization method based on reinforcement learning
By constructing causal relationships and designing a causal intervention reward mechanism, the problems of insufficient causal correlation dynamic modeling and collaborative optimization of composite control parameters in the wind turbine control method are solved, and automatic optimization of composite control parameters of wind turbine yaw pitch is realized, improving operational efficiency and decision-making transparency.
Patent Information
- Application Number
- CN202510362964.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-26
AI Technical Summary
The existing wind turbine control method lacks dynamic modeling of causal correlations between multivariables in the reward function design, which leads to the easy entry of strategy updates into local optimality, and the coordinated optimization of composite control parameters faces dimensional disaster problems, the decision-making process lacks transparency, and it is impossible to reversely track the contribution of key variables to the control target.
By obtaining the historical operation data of the wind turbine, a causal relationship is constructed and the causal intensity is quantified to generate a causal graph with strength and weakness marks. Based on reinforcement learning methods, a causal intervention reward mechanism is designed and converted into interpretable decision-making rules. Combined with the cause and effect graph, reverse tracking of decision rules and updated causal maps, thereby realizing automatic optimization of the composite control parameters of the yaw pitch of the wind turbine.
It overcomes the limitations of fuzzy parameter relationships in traditional control methods, provides a scientific basis for control decisions, solves the "black box" problem, makes the generation process of reinforcement learning control strategies have clear physical significance and logical basis, and improves the operating efficiency of wind turbines in complex environments.
Smart Images

Figure CN120100629A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of intelligent control of wind turbine generator sets, and in particular to a self-optimization method for yaw-pitch composite control parameters based on reinforcement learning. Background Art
[0002] In recent years, the optimization of wind turbine control strategies has gradually evolved from traditional PID control and model predictive control to data-driven intelligent algorithms. With the breakthrough of machine learning technology, the control parameter adaptive method based on deep reinforcement learning (DRL) has shown potential in the optimization of dynamic response of wind turbines. Some current studies have shown that by introducing causal reasoning models, such as structural equation modeling (SEM) to analyze the influence path between variables, the physical interpretability of the control strategy can be improved. However, the existing technology still has the following shortcomings: First, traditional reinforcement learning often relies on empirical settings or single-target optimization in the design of reward functions, lacks dynamic modeling of causal relationships between multiple variables, and causes the strategy update to easily fall into local optimality. For example, in the process of yaw angle adjustment, the existing method fails to effectively distinguish the differentiated influence paths of wind speed mutation and mechanical inertia on the tip speed ratio, making it difficult to accurately evaluate the causal effect of action decision-making. Secondly, the collaborative optimization of composite control parameters faces the problem of dimensionality curse. Although the existing DRL algorithm based on the black box model can handle high-dimensional inputs, its decision-making process lacks transparency and cannot reversely track the contribution of key variables to the control target, resulting in parameter adjustment lagging behind operating condition changes. In addition, since the physical constraints implicit in the historical operating data (the nonlinear coupling of the torque equation and the power coefficient) are not fully encoded into the learning framework, the generalization ability of the algorithm under extreme conditions is limited. For example, when the pitch angle and yaw angle are adjusted at the same time, the existing methods find it difficult to analyze the competitive influence mechanism of the two on the output power, resulting in control command conflicts. These problems seriously affect the operating efficiency of wind turbines in complex environments. Summary of the invention
[0003] The purpose of this section is to summarize some aspects of embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the specification abstract and the invention title of this application to avoid blurring the purpose of this section, the specification abstract and the invention title, and such simplifications or omissions cannot be used to limit the scope of the present invention.
[0004] In view of the above existing problems, the present invention is proposed. Therefore, the present invention provides a self-optimization method for yaw-pitch composite control parameters based on reinforcement learning, which is used to solve the problems raised in the background technology.
[0005] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning, comprising:
[0006] Acquire historical operation data of the wind turbine generator set, construct a causal relationship based on the historical operation data, and quantify the causal strength in the causal relationship based on a constrained causal discovery algorithm to generate a causal graph with strong and weak marks;
[0007] Obtain real-time operation data of wind turbines, design a causal intervention reward mechanism based on reinforcement learning methods, and convert the causal intervention reward mechanism into an interpretable decision rule;
[0008] In combination with the cause-effect diagram, the decision rule is traced backwards, and the cause-effect diagram is updated with the traced result of the decision rule, thereby realizing automatic optimization of the yaw-pitch composite control parameters of the wind turbine generator set.
[0009] As a preferred solution of the method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning described in the present invention, the historical operation data of the wind turbine generator set is obtained, including:
[0010] Ambient wind speed, rotation speed of the wind turbine rotor, angle between the wind turbine nacelle and wind direction, angle between the blades and the rotation plane, and electrical power output by the wind turbine;
[0011] As well as, the torque equation and tip speed ratio.
[0012] As a preferred solution of the method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning described in the present invention, the historical operation data of the wind turbine generator set is obtained, and a causal relationship is established according to the historical operation data, including:
[0013] Based on the physical constraints of the wind turbine generator set itself, the causal relationships among wind speed → tip speed ratio, rotation speed → tip speed ratio, rotation speed → power, yaw angle → torque, pitch angle → power coefficient, power coefficient → torque, tip speed ratio → power coefficient, and torque → power are established;
[0014] Among them, the left side of the arrow is the influencing factor, and the right side of the arrow is the affected factor.
[0015] As a preferred solution of the method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning described in the present invention, the causal strength in the causal relationship is quantified based on a constrained causal discovery algorithm to generate a causal graph with strong and weak marks, including:
[0016] According to the causal relationship, a fully connected directed graph is constructed, and for each edge in the directed graph, its partial correlation coefficient is calculated, and the partial correlation coefficient is used as a measure of the causal strength in the causal relationship to generate a causal graph with strong and weak marks.
[0017] As a preferred solution of the method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning described in the present invention, real-time operation data of the wind turbine generator set is obtained, and a causal intervention reward mechanism is designed based on the reinforcement learning method, including:
[0018] The real-time operation data of wind turbines are used as the environmental state input of reinforcement learning to construct a dynamic causal reward function and delayed causal effect compensation.
[0019] As a preferred solution of the method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning described in the present invention, the construction of a dynamic causal reward function and delayed causal effect compensation includes:
[0020] Calculate the deviation between the real-time operation data and the historical operation data, extract the direction of change of the data action predicted by the causal graph with strong and weak marks, and calculate the reward value;
[0021] A sliding window is set. If the direction of change of the data action predicted by the causal graph shows an upward trend, the reward value is increased. If the direction of change of the data action predicted by the causal graph does not change or shows a downward trend, the reward value is not changed.
[0022] As a preferred solution of the method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning described in the present invention, the causal intervention reward mechanism is converted into an interpretable decision rule, including:
[0023] According to the causal diagram with strong and weak marks, the importance of each variable in the causal diagram is calculated, and each variable is sorted according to its importance. The decision tree is constructed based on the importance of the sorted variables.
[0024] As a preferred solution of the method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning described in the present invention, wherein: in combination with the causal graph, the decision rule is traced back, and the causal graph is updated with the tracking result of the decision rule, including:
[0025] A decision path is formed along the root node to the leaf node of the decision tree. If the order of the root node and the leaf node in the decision path is different from that in the originally constructed decision tree, the importance of each variable is reordered according to the decision path, and while reconstructing the decision tree, the causal diagram with strong and weak marks is updated in reverse.
[0026] Compared with the prior art, the invention has the following beneficial effects:
[0027] 1. This method quantitatively analyzes the cause-effect relationship between the operating parameters of the wind turbine generator set and adds the physical constraints of the wind turbine generator set itself to generate a cause-effect diagram with strong and weak marks, thus overcoming the limitation of the fuzzy parameter relationship in the traditional control method and providing a scientific basis for control decision-making;
[0028] 2. Convert the causal intervention reward mechanism into an interpretable decision rule, and construct a decision tree based on the importance ranking of variables in the causal graph, which solves the "black box" problem of traditional reinforcement learning, making the generation process of reinforcement learning control strategy have clear physical meaning and logical basis, and providing a transparent way for engineering implementation and fault diagnosis;
[0029] 3. Through the reverse tracking mechanism, a decision path is formed according to the constructed decision tree, and the dynamic update of the cause-effect diagram is realized, which effectively responds to the complex and changeable wind conditions and makes the control parameter optimization process in the wind turbine generator set more intelligent in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work. Among them:
[0031] Figure 1 This is an overall flow chart of a method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning according to an embodiment of the present invention;
[0032] Figure 2 The present invention is a flowchart of a decision tree for constructing a method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning according to an embodiment of the present invention. DETAILED DESCRIPTION
[0033] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.
[0034] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0035] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0036] The present invention is described in detail with reference to schematic diagrams. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.
[0037] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0038] In the present invention, unless otherwise clearly specified and limited, the terms "install, connect, connect" should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0039] Example 1
[0040] Reference Figure 1 and Figure 2 , which is the first embodiment of the present invention, and provides a method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning, comprising:
[0041] S1. Obtain historical operation data of the wind turbine generator set, construct a causal relationship based on the historical operation data, and quantify the causal strength in the causal relationship based on a constrained causal discovery algorithm to generate a causal graph with strong and weak marks;
[0042] Specifically, historical operation data and real-time operation data of the wind turbine are obtained from the wind turbine SCADA system;
[0043] Specifically, the historical operation data and real-time operation data include ambient wind speed, rotation speed of the wind turbine rotor, angle between the wind turbine nacelle and wind direction, angle between the blade and the rotation plane, electric power output by the wind turbine, torque equation and tip speed ratio;
[0044] It should be noted that both the historical operation data and the real-time operation data obtained need to be preprocessed, that is, to remove missing values and outliers in the data;
[0045] It should be explained that the torque equation refers to the torque driving the generator to rotate, and the tip speed ratio refers to the ratio of the rotor tip linear velocity to the wind speed;
[0046] Furthermore, the torque equation can be expressed as:
[0047]
[0048] Where ρ is the air density (kg / m 3 ), R represents the blade radius (m), C p (λ, β) is expressed as the power coefficient, which is composed of the tip speed ratio λ and the pitch angle β, C p It is represented by the efficiency of converting wind energy into mechanical energy; τ is represented by torque, and the torque in the scheme of the present invention refers to the torque of the wind turbine generator set; v is represented by wind speed (m / s);
[0049] Furthermore, the tip speed ratio is expressed as:
[0050]
[0051] Where ω is the rotation speed (rad / s);
[0052] Furthermore, the relationship between power, torque and speed is:
[0053] P=τω
[0054] Wherein, P represents power;
[0055] Specifically, taking the physical constraints of the wind turbine generator set as conditions, the following causal relationship is obtained:
[0056] Establish the cause-effect relationship between wind speed → tip speed ratio, rotation speed → tip speed ratio, rotation speed → power, yaw angle → torque, pitch angle → power coefficient, power coefficient → torque, tip speed ratio → power coefficient, and torque → power;
[0057] Among them, the left side of the arrow is the influencing factor, and the right side of the arrow is the affected factor;
[0058] It should be noted that for wind speed → tip speed ratio, changes in wind speed will lead to a reverse change in the tip speed ratio (i.e., when the wind speed increases, if the speed remains unchanged, the tip speed ratio decreases); for speed → tip speed ratio, the speed directly determines the size of the tip speed ratio and works together with the wind speed (i.e., an increase in speed will positively increase the tip speed ratio); for speed → power, the speed directly affects the power output, and even if the torque remains unchanged, changes in speed will linearly change the power; for yaw angle → torque, the yaw angle affects the effectiveness of the wind turbine in capturing wind energy (i.e., when the yaw angle increases, the effective wind speed decreases, and the torque decreases accordingly ); for pitch angle → power coefficient, the pitch angle directly affects the efficiency of converting wind energy into mechanical energy by changing the angle of attack of the wind turbine blades (i.e. increasing the pitch angle usually reduces the power coefficient); for power coefficient → torque, the power coefficient directly determines the size of the torque, and the change in the power coefficient will affect the torque in a proportional form; for tip speed ratio → power coefficient, the tip speed ratio reflects the matching degree between the rotor speed and the wind speed, and can directly affect the aerodynamic performance and wind energy capture efficiency of the wind turbine blades; for torque → power, the change in torque will affect the power output in a proportional state, and this relationship is the basic physical law of mechanical energy conversion;
[0059] Furthermore, based on the constrained causal discovery algorithm, it is assumed that there is a relationship between all variables in the causal relationship, that is, there is an edge between all variables in the initial causal relationship, and a fully connected directed graph is constructed according to the causal relationship;
[0060] It should be noted that the advantage of the constrained causal discovery (PC) algorithm is that it can automatically discover causal structures from data, but its disadvantage is that it only relies on data. In addition, in the actual application of wind turbines, relying solely on data may result in causal relationships that do not conform to physical laws. For example, if there is a relationship between all variables, then it will be shown that there is a direct correlation between wind speed and torque, but according to the physical constraints of the wind turbine, wind speed cannot directly affect torque, but indirectly affects it through tip speed ratio and power coefficient. Therefore, incorporating the physical constraints of the above wind turbines into the PC algorithm can solve this problem well. In addition, since the physical constraints of the wind turbine itself have been considered in the causal relationship before the directed graph is constructed by the PC algorithm, it is no longer necessary for the traditional PC algorithm to delete edges that do not meet the conditions through conditional independence tests, such as the influence between wind speed and torque, which further optimizes the PC algorithm.
[0061] Furthermore, for each edge in the directed graph, the partial correlation coefficient is calculated, and the partial correlation coefficient is used as a measure of the causal strength in the causal relationship to generate a causal graph with strong and weak marks;
[0062] Furthermore, the influencing factor is regarded as X, and the affected factor is regarded as Y. The other parent nodes of X are defined (i.e., the variables that directly affect X may indirectly affect Y through X), and the other parent nodes of Y are defined (i.e., the variables that directly affect Y may act on Y together with X), except X. For example, there are two edges of wind speed → blade tip speed ratio and rotation speed → blade tip speed ratio. Then, when calculating the partial correlation coefficient of wind speed → blade tip speed ratio, the influence of rotation speed needs to be controlled.
[0063] Specifically, the partial correlation coefficient ρ XY The calculation formula is expressed as:
[0064]
[0065] Among them, Cov and Var represent covariance and variance respectively, R X It is expressed as the residual value of X after controlling other variables, R Y It is expressed as the residual value of Y after controlling other variables;
[0066] Specifically, the value range of the partial correlation coefficient is [-1, 1]. The absolute value of the partial correlation coefficient is taken, and the absolute value is used to represent the strength of the causal strength. The larger the absolute value, the stronger the direct impact of X on Y.
[0067] Specifically, the causal strength of absolute value division is as follows:
[0068] 0.0~0.2: weak causal relationship;
[0069] 0.2-0.5: moderate causal relationship;
[0070] 0.5-1.0: strong causal relationship;
[0071] Specifically, if the value of the partial correlation coefficient is negative, it indicates a reverse causal relationship (X increases, Y decreases), and if the value of the partial correlation coefficient is positive, it indicates a positive causal relationship (X increases, Y increases);
[0072] Specifically, the absolute value of the partial correlation coefficient (causal strength) is used to mark each directed edge of the directed graph, and finally a causal graph with strong and weak marks is generated;
[0073] S2. Obtain real-time operation data of wind turbines, design a causal intervention reward mechanism based on reinforcement learning methods, and convert the causal intervention reward mechanism into an interpretable decision rule;
[0074] Furthermore, the real-time operation data of wind turbines are used as the environmental state input of reinforcement learning to construct a dynamic causal reward function and delayed causal effect compensation;
[0075] Furthermore, the dynamic causal reward function is constructed by calculating the deviation between the real-time operation data and the historical operation data, extracting the direction of change of the data action predicted by the causal graph with strong and weak marks, and calculating the reward value;
[0076] Specifically, the calculation of the deviation ΔD between the real-time operation data and the historical operation data can be expressed by the following formula:
[0077] ΔD=D rt_d -D h_d
[0078] Among them, D rt_d It is represented as real-time operation data, which can be any data contained in the real-time operation data. h_d Represented as historical operation data, similarly, it can be data contained in any historical operation data;
[0079] It should be noted that the deviation between the real-time operation data and the historical operation data needs to be calculated using the same data, for example, the deviation between the real-time wind speed and the historical wind speed;
[0080] Specifically, the direction of data action change refers to the control measures taken according to the deviation value. For example, if increasing the pitch angle → power reduction (strong causal relationship 0.6), then by reducing the pitch angle, the power reduction state is alleviated;
[0081] For example, assuming that the blade radius of a wind turbine is 50m, the rotation speed is 2 (rad / s), the historical wind speed is 10m / s, and the real-time wind speed is 12m / s, the effect of wind speed deviation on the tip speed ratio is as follows:
[0082] The historical tip speed ratio is calculated by the above formula:
[0083]
[0084] Calculate the real-time tip speed ratio:
[0085] (Keep two decimal places)
[0086] Then, through analysis, we can get that: when the wind speed increases from the historical value of 10m / s to the real-time value of 12m / s, the tip speed ratio decreases from 10 to 8.33; and the power coefficient of a wind turbine generator set is usually within a specific tip speed ratio range. Assuming that the historical tip speed ratio of 10 in the above example is the optimal value, then the real-time tip speed ratio of 8.33 deviates from this optimal value, which will lead to a decrease in the power coefficient of the wind turbine generator set, thereby reducing the actual power generation efficiency;
[0087] It should be noted that in order to make the competitive influence mechanism generated when the two variables are adjusted simultaneously, resulting in a conflict in control instructions, it is necessary to balance them through reward values;
[0088] Specifically, the reward value calculation includes the basic reward, which is the negative value of the specific deviation (such as -0.1 means a deviation of 10%); the causal compliance reward, which is a reward if the actual data change is consistent with the relationship in the causal graph, otherwise the reward is deducted; the total reward, which is the weighted sum of the basic reward and the causal compliance reward;
[0089] Furthermore, the delayed causal effect compensation adopts a sliding window method, and if the direction of the data action change predicted by the causal graph shows an upward trend, the reward value is increased; if the direction of the data action change predicted by the causal graph does not change or shows a downward trend, the reward value is not changed;
[0090] It should be noted that when the change direction of the data action predicted by the causal diagram shows an upward trend, it means that the control measures taken for the deviation value are positive (i.e., efforts are made to balance the parameters to an ideal state). In this case, increasing the reward value is equivalent to positive feedback in reinforcement learning, which helps wind turbines maintain stability and power generation efficiency under complex conditions. In addition, by using a sliding window (such as 5 minutes), the delayed effect after the predicted data action can be tracked and its continuous impact on the system state can be evaluated. For example, the wind speed may fluctuate violently within a few seconds, but the 5-minute sliding window can more accurately reflect the changing trend of the current predicted data action.
[0091] Furthermore, according to the causal diagram with strong and weak marks, the importance of each variable in the causal diagram is calculated, and each variable is sorted according to its importance, and a decision tree is constructed based on the importance of each variable after sorting;
[0092] Specifically, find the path from a certain variable to the target variable. For each path, calculate the product of all edge weights on the path, use the absolute value of the partial correlation coefficient (causal strength) as the edge weight, and add the products of all paths to get the total causal strength from a certain variable to the target variable.
[0093] For example, suppose there are two paths in the causal graph with strong and weak labels: X1→X2→Y and X1→Y;
[0094] Then for path 1 (X1→X2→Y), it contains two edges. Assuming that the edge weights are 0.8 for X1→X2 and 0.5 for X2→Y, the strength of path 1 is 0.8×0.5=0.4; for path 2 (X1→Y), it contains one edge. Assuming that the edge weight is 0.6, the strength of path 2 is 0.6; adding the products of all paths 0.4+0.6=1, the total causal strength of X1→Y is 1, and the causal strength of a certain variable X1 on the target variable Y is obtained as its importance;
[0095] Specifically, the importance of each variable is ranked, and the higher the strength of the variable, the greater its impact on the target variable;
[0096] It should be noted that in standard decision tree algorithms (such as CART), when selecting split variables at each step, the information gain (or Gini index reduction) of all candidate variables is calculated, and the variable with the largest gain is selected for splitting. In order to give priority to variables with high importance, causal strength can be incorporated into the decision tree. Figure 2 , the steps to construct a decision tree are as follows:
[0097] S201, determining the target variable;
[0098] S202, in each node of the causal graph with strong and weak labels, calculating the information gain of each variable from the candidate variables;
[0099] S203, multiplying the information gain by the causal strength of the variable to obtain a weighted information gain;
[0100] S204, selecting the variable with the highest weighted information gain for segmentation;
[0101] S205, recursively repeating this process for the child nodes under each node until the decision tree reaches the maximum depth, that is, all nodes in the causal graph are assigned;
[0102] For example, assume that the causal graph contains three variables X1, X2, X3 and the target variable Y, and the causal strengths are: X1 is 1; X2 is 0.5; X3 is 0.7; at the root node of the decision tree, assume that the information gains are: X1 is 0.3; X2 is 0.4; X3 is 0.35;
[0103] The calculated weighted information gain is:
[0104] X1: 1×0.3=0.3;
[0105] X2: 0.5×0.4=0.2;
[0106] X3: 0.7 × 0.35 = 0.245;
[0107] Then select X1 as the root node segmentation variable because it has the highest weighted information gain;
[0108] S3, combining the cause-effect diagram, reversely tracing the decision rule, and updating the cause-effect diagram with the tracing result of the decision rule, so as to realize automatic optimization of the yaw and pitch composite control parameters of the wind turbine generator set;
[0109] Furthermore, a decision path is formed along the root node to the leaf node of the decision tree. For example, in the yaw pitch control of a wind turbine, the decision path may be wind speed>10m / s→speed>15rad / s→power≈800kW.
[0110] It needs to be explained that a leaf node refers to a node without child nodes. In other words, a child node can be a leaf node;
[0111] Furthermore, if the order of the root node and the leaf node in the decision path is different from that in the originally constructed decision tree, the importance of each variable is reordered according to the decision path, and while reconstructing the decision tree, the causal graph with strong and weak marks is updated in reverse;
[0112] It should be noted that if the order of the root node and leaf node in the decision path is inconsistent with that in the originally constructed decision tree, it indicates that the operating environment of the wind turbine generator has changed and the original decision tree is no longer fully adapted to the current working conditions. Therefore, it is necessary to reconsider the importance of each variable and reconstruct the original decision tree.
[0113] Specifically, reverse updating means that if the decision path reveals a new causal relationship or negates the original relationship, then the edge structure of the original causal graph needs to be adjusted (deleting edges, adjusting edges, adding edges).
[0114] Those skilled in the art will appreciate that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the application can adopt the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program codes. The scheme in the embodiments of the present application can be implemented in various computer languages, for example, object-oriented programming language Java and literal scripting language JavaScript, etc.
[0115] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0116] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0117] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0118] Although the preferred embodiments of the present application have been described, those skilled in the art may make other changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0119] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A self-optimization method for yaw-pitch composite control parameters based on reinforcement learning, characterized in that: include: Acquire historical operation data of the wind turbine generator set, construct a causal relationship based on the historical operation data, and quantify the causal strength in the causal relationship based on a constrained causal discovery algorithm to generate a causal graph with strong and weak marks; Obtain real-time operation data of wind turbines, design a causal intervention reward mechanism based on reinforcement learning methods, and convert the causal intervention reward mechanism into an interpretable decision rule; In combination with the cause-effect diagram, the decision rule is traced backwards, and the cause-effect diagram is updated with the traced result of the decision rule, thereby realizing automatic optimization of the yaw-pitch composite control parameters of the wind turbine generator set.
2. The method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning as claimed in claim 1, characterized in that: Obtain historical operating data of wind turbines, including: Ambient wind speed, rotation speed of the wind turbine rotor, angle between the wind turbine nacelle and wind direction, angle between the blades and the rotation plane, and electrical power output by the wind turbine; As well as, the torque equation and tip speed ratio.
3. The method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning as claimed in claim 2, characterized in that: Acquiring historical operation data of the wind turbine generator set and building a causal relationship based on the historical operation data, including: Based on the physical constraints of the wind turbine generator set itself, the causal relationships among wind speed → tip speed ratio, rotation speed → tip speed ratio, rotation speed → power, yaw angle → torque, pitch angle → power coefficient, power coefficient → torque, tip speed ratio → power coefficient, and torque → power are established; Among them, the left side of the arrow is the influencing factor, and the right side of the arrow is the affected factor.
4. The method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning as claimed in claim 3, characterized in that: Based on the constrained causal discovery algorithm, the causal strength in the causal relationship is quantified to generate a causal graph with strong and weak marks, including: According to the causal relationship, a fully connected directed graph is constructed, and for each edge in the directed graph, its partial correlation coefficient is calculated, and the partial correlation coefficient is used as a measure of the causal strength in the causal relationship to generate a causal graph with strong and weak marks.
5. The method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning as claimed in claim 4, characterized in that: Obtain real-time operating data of wind turbines and design a causal intervention reward mechanism based on reinforcement learning methods, including: The real-time operation data of wind turbines are used as the environmental state input of reinforcement learning to construct a dynamic causal reward function and delayed causal effect compensation.
6. The method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning as claimed in claim 5, characterized in that: The construction of a dynamic causal reward function and delayed causal effect compensation includes: Calculate the deviation between the real-time operation data and the historical operation data, extract the direction of change of the data action predicted by the causal graph with strong and weak marks, and calculate the reward value; A sliding window is set. If the direction of change of the data action predicted by the causal graph shows an upward trend, the reward value is increased. If the direction of change of the data action predicted by the causal graph does not change or shows a downward trend, the reward value is not changed.
7. The method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning as claimed in claim 5, characterized in that: Convert the causal intervention reward mechanism into an interpretable decision rule, including: According to the causal diagram with strong and weak marks, the importance of each variable in the causal diagram is calculated, and each variable is sorted according to its importance. The decision tree is constructed based on the importance of the sorted variables.
8. The method for self-optimization of yaw-pitch composite control parameters based on reinforcement learning as claimed in claim 7, characterized in that: In combination with the cause-effect diagram, the decision rule is traced back, and the cause-effect diagram is updated with the traced result of the decision rule, including: A decision path is formed along the root node to the leaf node of the decision tree. If the order of the root node and the leaf node in the decision path is different from that in the originally constructed decision tree, the importance of each variable is reordered according to the decision path, and while reconstructing the decision tree, the causal diagram with strong and weak marks is updated in reverse.
Citation Information
Patent Citations
Method and equipment for automatically optimizing pitch angle under blade stall condition
CN114076060A
Draught fan yaw control method and system based on deep reinforcement learning
CN117052596A
Power consumption prediction method based on causal discovery and deep learning
CN118228853A
Wind power plant power generation control method and device based on wind power plant yaw model, electronic equipment and computer readable storage medium
CN118468710A
Wind power plant power generation control method and device based on wake flow estimation model, electronic equipment and computer readable storage medium
CN118468711A