A method and system for dynamic optimization and adjustment of wind power curves based on reinforcement learning
By adopting a reinforcement learning-based dynamic optimization and adjustment method for wind power curves, the problems of power capture deviation and load overload of wind turbines in complex environments are solved, achieving an optimal balance between power generation efficiency and mechanical life, and improving the adaptive capability and grid response capability of wind turbines.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINJIANG XINFENG XINNENG ENVIRONMENTAL PROTECTION TECH CO LTD
- Filing Date
- 2026-04-14
- Publication Date
- 2026-07-10
AI Technical Summary
In complex and ever-changing actual operating environments, traditional wind turbines cannot dynamically adjust their static power curves, resulting in large power capture deviations, overloads, and insufficient control flexibility, making it difficult to achieve an optimal balance between power output and mechanical life.
A reinforcement learning-based dynamic optimization and adjustment method for wind power curves is adopted. Through data acquisition, quality diagnosis, feature extraction, real-time clustering of operating conditions, and multi-scenario reinforcement learning strategies, the final control command is generated to achieve real-time optimization and adjustment of wind turbine units.
It significantly reduced the deviation between actual output and theoretical optimal value, improved power generation efficiency, suppressed fatigue load on the transmission chain and tower structure, and enhanced the unit's adaptability and grid dispatch response quality throughout its entire life cycle.
Smart Images

Figure CN122371334A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of wind power generation technology, specifically a method and system for dynamic optimization and adjustment of wind power curves based on reinforcement learning. Background Technology
[0002] With the profound transformation of the global energy structure, wind power, as a crucial pillar of the renewable energy sector, has seen its energy conversion efficiency and turbine operation reliability become a core focus of academic and industrial attention. In the operation and control system of wind turbines, the power curve is not only a key benchmark for measuring the aerodynamic performance and power generation efficiency of wind turbines, but also a core basis for wind farm energy dispatch, power prediction, and the formulation of optimal turbine control strategies. According to current international standards such as IEC61400-12-1, the traditional wind power curve is typically defined as a static mapping function between the instantaneous wind speed at hub height and the generator output power. In practical engineering applications, this static mapping is often achieved through a pre-defined lookup table method or low-order polynomial fitting. Its basic logic is based on assuming the turbine is operating under steady-state conditions with standard air density, no yaw error, and ideal aerodynamic characteristics on the blade surface, and determining the target torque or power command at a given wind speed through linear interpolation. However, a deeper exploration of the physical operating mechanism of wind turbines reveals that this control paradigm based on static mapping faces profound and irreconcilable technical contradictions when dealing with complex and ever-changing actual operating environments. First, wind turbines operate for extended periods in a highly unstable atmospheric boundary layer, where environmental parameters such as air density, turbulence intensity, and inflow angle exhibit significant time-varying characteristics. Static power curves neglect the dynamic adjustment of air density with altitude and temperature variations and fail to capture the nonlinear interference of turbulent fluctuations on energy capture efficiency, resulting in a significant deviation of over 15% between the actual output and design value under non-standard operating conditions. Further complicating matters, the aerodynamic performance of the blades, as the core component of energy conversion, is not constant. Over time, fouling, leading-edge wear, and even winter icing on the blade surface can lead to a decrease in the lift coefficient and an increase in the drag coefficient, causing an overall downward shift and reverse drift in the power characteristic curve.
[0003] Traditional control schemes, lacking the ability to identify these implicit aerodynamic parameter degradations online, often lead to the control system still performing ineffective "optimal" optimization based on factory-specified parameters. This not only results in significant power generation losses but also introduces additional drivetrain fatigue loads due to speed and torque mismatches. Further analysis from the perspective of system dynamics and control theory reveals a fundamental trade-off between maximizing power generation efficiency and ensuring generator load safety. To achieve maximum power point tracking, the control system needs to frequently adjust generator torque and blade pitch angle to maintain the optimal tip speed ratio. However, under conditions of strong turbulence or sudden gusts, such high-frequency, high-amplitude movements directly excite tower vibration modes, generating enormous tower base bending moments and instantaneous loads. Existing technologies typically design power optimization and load suppression as decoupled control loops or attempt rolling optimization within finite time bounds using simple model predictive control. However, the performance of model predictive control is highly dependent on the accuracy of the physical model, and in strongly nonlinear, multivariate coupled wind power systems, model mismatch often leads to the control sequence failing to converge. While recent studies have attempted to incorporate traditional reinforcement learning algorithms to address such nonlinear problems, most have been limited to discrete state space partitioning. When faced with dynamic conditions characterized by extremely high wind speed variations and continuous yet constrained action spaces, these algorithms often suffer from the curse of dimensionality and slow convergence, failing to meet the millisecond-level real-time control requirements of wind turbines. Crucially, as grid demands for wind farm participation in frequency regulation and dispatch response become increasingly stringent, turbines not only need to achieve efficiency across the entire wind speed range but also exhibit extremely high control flexibility under special conditions such as power-limited operation and grid frequency regulation. Existing static curve and single-objective optimization strategies fall short in handling smooth switching between multiple operating conditions, dynamic responses under load constraints, and the extraction of implicit environmental features.
[0004] Specifically, when environmental turbulence increases or blades show signs of icing, how to dynamically reconstruct the optimal energy capture function and maximize the expected power generation while ensuring the structural safety of the unit has become a core technical challenge in the current technological evolution.
[0005] In summary, how to break through the limitations of traditional static mapping and construct a dynamic optimization mechanism that can deeply integrate real-time environmental perception, aerodynamic parameter identification, and multi-objective trade-offs, so as to achieve the optimal balance between power output and mechanical life under varying operating conditions, has become a key challenge and an urgent technical problem for those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for dynamic optimization and adjustment of wind power curves based on reinforcement learning, so as to solve the technical problems in the prior art mentioned in the background, such as large power capture deviation, load over-limit, and insufficient control flexibility caused by the static power curve under complex time-varying environment and blade performance degradation conditions.
[0007] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A method for dynamic optimization and adjustment of wind power curves based on reinforcement learning includes the following steps: Step S1: Obtain real-time operating parameters through the data acquisition layer deployed on the wind turbine, and perform data quality diagnosis on the real-time operating parameters. When a sensor fault is detected, perform virtual wind speed reconstruction. When a data delay is detected to exceed a preset threshold, perform time-series interpolation compensation to obtain the operating data after quality repair. Step S2: Extract features from the operational data after quality repair, perform real-time clustering of operating conditions based on the extracted feature vectors, and output the current operating condition label based on the clustering results; Step S3: Based on the current operating condition label, activate the corresponding sub-agent from the pre-set multi-scenario reinforcement learning strategy library, and the sub-agent outputs preliminary control actions according to the current operating state; wherein, the multi-scenario reinforcement learning strategy library includes at least the efficiency-first reinforcement learning strategy corresponding to steady-state low-turbulence operating conditions, the load suppression reinforcement learning strategy corresponding to high-turbulence gust operating conditions, the aerodynamic parameter adaptive strategy corresponding to aerodynamic performance degradation operating conditions, and the power limiting dynamic curve strategy corresponding to power grid power limiting command operating conditions; Step S4: Perform fusion arbitration on the preliminary control actions generated by the switching of different operating conditions and strategies, generate the final control command using a gradual fusion method, and send the final control command to the wind turbine actuator.
[0008] According to the above technical solution, the virtual wind speed reconstruction in step S1 specifically includes: The generator speed and the effective wind speed at hub height are used as the state vector, the generator speed measurement and the generator power measurement are used as the observation vector, the drive chain rotational inertia balance equation is used as the basis for state transition, and the current effective wind speed is estimated recursively through extended Kalman filtering. When reference wind speed data from lidar or anemometers is available, the extended Kalman filter estimate is weighted and fused with the reference wind speed. The fusion weight is determined by the variance of the extended Kalman filter posterior error and the variance of the reference wind speed measurement noise according to the minimum variance criterion.
[0009] According to the above technical solution, step S2 specifically includes: Extract multi-dimensional feature vectors at a preset period. The multi-dimensional feature vectors include at least the average wind speed, turbulence intensity, high-frequency energy proportion of wind speed power spectral density, generator power variance, and mean tower base bending moment. After performing principal component analysis to reduce the dimensionality of the multidimensional feature vectors, the density-based DBSCAN clustering algorithm is used for clustering, and candidate working condition labels are determined according to the cluster centers and the preset mapping relationship. Hysteresis acknowledgment based on time continuity is performed on candidate operating condition labels. The candidate label is officially output as the current operating condition label only if the same candidate label remains unchanged for multiple consecutive periods.
[0010] According to the above technical solution, the efficiency-first reinforcement learning strategy in step S3 uses the following reward function to guide the agent training: in, Real-time generator power, Rated power, The pitch angle fine-tuning amount output by the sub-agent. For rotational acceleration, and These are the weighting coefficients calibrated based on the cost-per-kilowatt-hour model.
[0011] According to the above technical solution, the load suppression reinforcement learning strategy in step S3 uses the following reward function to guide agent training: in, The bending moment at the front and rear of the tower base, The design limit values of the bending moments before and after the tower base. The power loss is due to load shedding control. Generator rated power, For rotational acceleration, , , These are the weighting coefficients calibrated based on the wind turbine structural dynamics model.
[0012] According to the above technical solution, the aerodynamic parameter adaptive strategy in step S3 specifically includes: When aerodynamic performance degradation is detected, a periodic disturbance signal is superimposed on the original pitch angle command. The amplitude of the disturbance signal is dynamically adjusted according to the current turbulence intensity, so that the disturbance amplitude increases in low turbulence to ensure excitation intensity and decreases in high turbulence to suppress additional load. Using data collected during disturbance injection, the current aerodynamic characteristic surface is identified online through recursive least squares method with forgetting factor. Based on the identification results, the optimal tip speed ratio and speed reference value are updated, so that the action output of the sub-agent is transformed into tracking control of the speed reference value.
[0013] According to the above technical solution, in the recursive least squares method with forgetting factor, the forgetting factor is adaptively adjusted according to the current turbulence intensity. In steady-state low turbulence, the forgetting factor is increased to make full use of historical data, and in high turbulence, the forgetting factor is decreased to quickly track changes in aerodynamic characteristics.
[0014] According to the above technical solution, the gradient blending method in step S4 is as follows: When the operating condition label changes, the preliminary control action output by the strategy before the change and the preliminary control action output by the strategy after the change are weighted and fused. The fusion weight smoothly increases from zero to one within a preset transition period, and the rate of change of the fusion weight is zero at the beginning and end of the transition period.
[0015] A wind power curve dynamic optimization and adjustment system based on reinforcement learning, comprising: The data acquisition and quality diagnosis module is used to acquire real-time operating parameters of the wind turbine and perform data quality diagnosis on the real-time operating parameters. When a sensor fault is detected, virtual wind speed reconstruction is performed, and when the data delay exceeds a preset threshold, timing interpolation compensation is performed, and the operating data after quality repair is output. The working condition identification and diagnosis module is used to extract features and perform real-time clustering of working conditions on the operational data after quality repair, and output the current working condition label. The hierarchical reinforcement learning computing engine integrates multiple parallel computing units, which correspond to the efficiency-first reinforcement learning strategy for steady-state low-turbulence conditions, the load suppression reinforcement learning strategy for high-turbulence gust conditions, the aerodynamic parameter adaptive strategy for aerodynamic performance degradation conditions, and the power-limiting dynamic curve strategy for power grid power-limiting command conditions. The hierarchical reinforcement learning computing engine activates the corresponding computing unit based on the current operating condition label and outputs the initial control action. The action execution arbitration module is used to fuse and arbitrate the preliminary control actions generated by the switching of strategies under different operating conditions, and to generate the final control command using a gradual fusion method. And the underlying controller communication interface, used to send the final control commands to the wind turbine actuators.
[0016] According to the above technical solution, the action execution arbitration module is specifically used to perform weighted fusion of the preliminary control action output by the strategy before the switch and the preliminary control action output by the strategy after the switch when the working condition label is switched. The fusion weight smoothly increases from zero to one within a preset transition period, and the rate of change of the fusion weight is zero at the beginning and end of the transition period.
[0017] Compared with the prior art, the present invention has the following beneficial effects: This invention constructs a hierarchical parallel reinforcement learning architecture and designs specialized reward functions and state spaces for different operating conditions. This enables the control strategy of wind turbines to adaptively adjust under complex time-varying environments, effectively overcoming the performance degradation problem of traditional static power curves under non-standard operating conditions and significantly reducing the deviation between actual output and theoretical optimal value. By introducing a feedback penalty mechanism for tower base bending moment and rotational acceleration into the load suppression reinforcement learning strategy, and combining it with dynamic safety thresholds and hardware-level protection logic under extreme gust conditions, this invention effectively suppresses the accumulation of fatigue loads on the drive train and tower structure while improving power generation efficiency, achieving an optimal balance between power generation benefits and unit lifespan.
[0018] This invention also utilizes periodic small-disturbance excitation and a recursive online identification algorithm with a forgetting factor to enable the control system to track the degradation trend of blade aerodynamic performance in real time and automatically reconstruct the optimal operating reference trajectory. This avoids control mismatch and power generation loss caused by implicit performance degradation such as icing and fouling, enhancing the unit's adaptive capability throughout its entire life cycle. Through a dynamic translation mechanism of the reference curve under power-limited conditions and a power tracking deviation penalty design, the tracking accuracy and dynamic response quality of the wind turbine in response to grid dispatch commands are significantly improved. Simultaneously, the mechanical shock caused by control switching during power-limited operation is reduced, enhancing the unit's overall performance in grid-friendly operating modes.
[0019] The present invention also employs a gradual fusion arbitration mechanism during operating condition switching, using a smooth transition curve to weight and fuse the outputs of multiple strategies, thereby eliminating torque jumps and actuator shocks caused by abrupt changes in control logic, and ensuring control continuity and operational stability during the entire operating condition switching process. Attached Figure Description
[0020] Figure 1 This is a flowchart illustrating the reinforcement learning-based dynamic optimization and adjustment method for wind power curves in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of the wind power curve dynamic optimization and adjustment system based on reinforcement learning in an embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1 like Figure 1 and Figure 2As shown, this embodiment provides a method for dynamic optimization and adjustment of wind power curves based on reinforcement learning. This method is implemented based on an event-driven hierarchical parallel architecture. The architecture includes a data acquisition layer, a data quality diagnosis module, a real-time operating condition clustering module, a multi-scenario reinforcement learning strategy module, an action fusion arbitration module, and a low-level controller communication interface. Data interaction between modules is synchronized at the microsecond level through a real-time industrial Ethernet protocol bus.
[0023] The data acquisition layer uses a sampling frequency of 10Hz to acquire sensor data deployed at key parts of the wind turbine, specifically including wind speed sequence at hub height, generator speed, blade pitch angle, generator power, tower foundation bending moment, and the first-order vibration frequency of the tower. The 10Hz sampling frequency is chosen because: the first-order torsional vibration mode frequency of the wind turbine drivetrain is typically in the 1-3Hz range; according to Shannon's sampling theorem, a 10Hz sampling frequency can effectively capture torsional vibration dynamics. Simultaneously, the fundamental frequency of the blade's passing frequency is approximately three times the rotor speed; for large turbines near the rated speed, this frequency is approximately 0.8-1.5Hz, and 10Hz sampling can fully resolve the 3P frequency components. Furthermore, 10Hz is also a widely used standard sampling frequency in industrial wind power control systems, achieving a balance between data volume and processor load. The raw data, after hardware filtering to remove high-frequency noise, enters the data quality diagnostic module.
[0024] The data quality diagnostic module performs integrity checks and timeliness assessments on each data packet. Integrity checks use cyclic redundancy check (CRC) codes to verify the correctness of data transmission and check whether the values of each sensor are within the preset valid physical range. Timeliness assessment calculates the transmission delay by comparing the timestamp within the data packet with the system clock of the central control unit. When a wind speed sensor malfunction is detected, a virtual wind speed reconstruction algorithm based on extended Kalman filtering is triggered. When a data delay exceeding 2 seconds is detected, a timing interpolation compensation algorithm based on a long short-term memory network is triggered. The 2-second delay threshold is set to consider the upper limit of the wind power control system's tolerance for data delay: near the rated wind speed, the power change rate caused by sudden wind speed changes can reach more than 10% of the rated power per second. If the control command lags behind by more than 2 seconds, it will lead to significant power tracking deviation and load accumulation; therefore, this is used as the critical condition for triggering interpolation compensation.
[0025] Furthermore, in this embodiment, the dynamic model of the wind turbine drive chain is used as the basis for state transition, and the available non-wind speed sensor measurements are used as observation information. The current effective wind speed is estimated recursively through extended Kalman filtering.
[0026] The selection of the state vector is based on the following considerations: generator speed is a precisely measurable state variable, while effective wind speed is not directly measurable; together, they constitute the minimum state set describing the energy conversion of the drivetrain. The reason for selecting generator speed and power measurements as the observation vector is that the speed measurement directly corresponds to the state variable, while the power measurement indirectly reflects aerodynamic torque through the product of electromagnetic torque and speed, thus providing information about wind speed. The state transition relationship is established based on the drivetrain moment of inertia balance equation, which describes the torque balance relationship between aerodynamic torque, electromagnetic torque, damping torque, and speed acceleration. The aerodynamic torque function is calculated based on the aerodynamic characteristic surface calibrated at the wind turbine's factory, which comprehensively describes the energy conversion efficiency of the wind turbine under different operating conditions through two dimensionless parameters: tip speed ratio and pitch angle. The wind speed state follows a random walk model to reflect the natural fluctuations caused by turbulence. The rationale for this model is that, on a second-level time scale, the change in wind speed is much smaller than the wind speed itself, and the changes between adjacent moments have no directional trend, conforming to the statistical characteristics of a random walk.
[0027] The calibration of the process noise covariance matrix is based on typical turbulence intensities: the standard deviation of the process noise for rotational speed is set to 0.05 rad / s, corresponding to the fluctuation level of the transmission chain speed measurement; the standard deviation of the process noise for wind speed is set to 0.15 m / s, corresponding to the second-level fluctuation of Class A turbulence intensity at an average wind speed of 10 m / s in the IEC standard. The observation noise covariance matrix is set according to the nominal value of the sensor accuracy. The filter sequentially performs a prediction step and an update step within each sampling period: the prediction step uses the posterior estimate and state equation from the previous moment to calculate the prior estimate and prior error covariance for the current moment, where the state transition Jacobian matrix is composed of the partial derivatives of the aerodynamic torque function with respect to the state variables, and the calculation of the partial derivatives is achieved through local linearization of the aerodynamic characteristic surface; the update step calculates the Kalman gain based on the actual measured values, and then corrects the state estimate and updates the error covariance.
[0028] When a lidar is installed on the top of the nacelle or there is wind measurement tower data in the field, the posterior estimate of wind speed is fused with the reference wind speed using a weighted average. The fusion weights are determined by the adaptive fusion criterion proposed in this invention. The derivation of this criterion is based on the principle of minimum variance unbiased estimation: Let the error variance of the extended Kalman filter wind speed estimate be and the measurement noise variance of the reference wind speed be . The two estimates are independent of each other. Then, the minimum variance linear combination weights of the fused estimates should satisfy: in, This represents the fusion weight, which is the weight coefficient of the reference wind speed in the fusion result; This represents the variance of the error in Kalman filter wind speed estimation; The measurement noise variance serves as a reference for wind speed; the physical meaning of this weight is that when the extended Kalman filter estimation accuracy is high, When the value is relatively small, the fusion weights approach 0, meaning the filtered estimate is more trusted; when the reference sensor accuracy is high, The fusion weight is relatively small, approaching 1, indicating greater trust in the reference measurement. This adaptive fusion mechanism enables the system to maintain optimal wind speed estimation accuracy under varying sensor availability conditions. Simulation and field verification show that even with complete anemometer failure, the reconstructed wind speed deviation can be controlled within 0.2 m / s.
[0029] For data delay interpolation compensation, this embodiment employs a sequence-to-sequence long short-term memory network pre-trained and deployed on edge computing units. The network uses an encoder-decoder architecture. The input is a multivariate state sequence within a continuous time window before the delay occurs, including four normalized variables: wind speed, generator speed, pitch angle, and generator power. The output is a predicted sequence of these variables within the delay period. These four variables are chosen because they constitute a complete minimal set describing the wind turbine's operating state: wind speed determines the input energy, speed and pitch angle determine the operating point, and power is the output quantity. These four variables are coupled together through aerodynamic characteristics and drivetrain dynamics.
[0030] The training loss function, in addition to the conventional mean squared error, introduces a physical constraint residual term based on the drive chain power balance equation. The construction logic of the physical constraint residual term is as follows: for any prediction time, the four predicted variables should satisfy the drive chain rotational inertia balance equation, and the residual is defined as the difference between the left and right sides of the equation. The mean squared value of the residual is added as a penalty term to the loss function, with a constraint weight of 0.1, allowing the network to maintain the physical consistency of the prediction results while minimizing the prediction error. The mechanism of this physical constraint is that a purely data-driven model may learn spurious correlations in the data, producing predictions that are statistically reasonable but violate physical laws; the physical constraint term forces the network to implicitly learn the drive chain dynamics, improving the extrapolation reliability of the prediction model under extreme conditions with insufficient training data coverage.
[0031] The initial network model was trained offline using historical SCADA data. The training data was constructed by extracting continuous time periods from the historical data, randomly deleting intermediate portions to simulate data delays, and using the preceding sequence as input and the deleted portion as the output target for supervised learning. During online operation, the model was fine-tuned weekly using newly accumulated normal operation data to adapt to unit state drift. During fine-tuning, only the decoder layer parameters were updated while the encoder layer was frozen, to retain learned general feature representations while adapting to new data. When data delays occur, the latest model is directly called for forward inference, with a single inference time of less than 5ms, meeting real-time requirements.
[0032] After data cleaning and compensation, the system extracts and clusters feature vectors at 10-second intervals. The feature vectors are 7-dimensional, including the 10-second average wind speed, turbulence intensity, high-frequency energy proportion of wind speed power spectral density, generator power variance, mean tower base bending moment, ambient temperature, and mean blade pitch angle. The 10-second cycle length was determined by considering the following factors: too short a cycle would significantly affect statistical features due to measurement noise, reducing the reliability of operating condition identification; too long a cycle would fail to respond promptly to rapid changes in operating conditions, resulting in delayed control strategy switching. 10 seconds is approximately a suitable window for wind turbines within the atmospheric boundary layer turbulence integral timescale.
[0033] The introduction of the high-frequency energy proportion feature of wind speed power spectral density is based on the mechanism of turbulence's influence on unit load: the distribution of turbulent energy in the high-frequency band (above 0.1Hz) directly affects the dynamic angle of attack of the blades and the torque pulsation of the transmission chain, and is a key indicator for distinguishing between steady-state low-turbulence and high-turbulence gust conditions. The extraction method for this feature is as follows: A fast Fourier transform is performed on the wind speed sequence within 10 seconds to obtain the power spectral density, and the ratio of energy in the range from 0.1Hz to half the sampling frequency to the total energy across the entire frequency band is calculated.
[0034] To reduce the correlation between features and accelerate cluster convergence, principal component analysis (PCA) is first used to reduce the dimensionality of the feature vectors. Principal components with a cumulative variance contribution rate of 95% are selected as clustering inputs. The 95% contribution rate threshold is set to retain most of the information in the original features while achieving effective dimensionality reduction. Based on the analysis of historical data, typically retaining four principal components is sufficient to achieve this contribution rate.
[0035] The clustering algorithm employs the density-based DBSCAN algorithm, which has the advantage of automatically identifying outliers and adapting to non-spherical cluster distributions. The DBSCAN algorithm has two core parameters: neighborhood radius and minimum neighbor number. These are determined based on the statistical distribution of historical data. Specifically, the distance between each point in the historical dataset and its k-th nearest neighbor is calculated and sorted, a k-distance curve is plotted, and the distance value at the inflection point of the curve is selected as the neighborhood radius. The minimum neighbor number is twice the feature dimension. The DBSCAN algorithm outputs the cluster label of the current data point in real time. The mapping relationship between the cluster label and the preset operating conditions is established through prior knowledge of historical data.
[0036] To avoid frequent jumps in operating condition labels near boundaries, this invention designs a hysteresis confirmation mechanism based on temporal continuity. The mathematical expression of this mechanism is: in, This is the final output label for the current operating condition. These are the candidate labels output by the DBSCAN algorithm. This represents the label value from the past period. =3 represents the number of consecutive periods required for hysteresis confirmation. This is an indicator function; it takes the value 1 if the condition is met, and 0 otherwise. The mechanism works by confirming a change in operating condition only when a candidate tag remains unchanged for three consecutive cycles. The physical basis for this is that changes in wind conditions and unit status have inertia on a second-level timescale, while genuine changes in operating condition should be continuous; isolated tag changes caused by measurement noise or brief disturbances should not trigger a control strategy switch. The 30-second confirmation time effectively filters out false fluctuations in operating condition, and for extreme gust conditions that truly require rapid response, protection actions can be directly triggered through independent safety protection logic without being subject to this hysteresis limitation.
[0037] When the operating condition is labeled as steady-state low-turbulence, the system prioritizes efficiency through a reinforcement learning strategy, employing a soft actor-critic algorithm. This algorithm is chosen based on its excellent performance in continuous control tasks: the maximum entropy framework encourages the policy to maintain sufficient randomness while optimizing rewards, enhancing exploration capabilities and robustness to model errors, making it particularly suitable for control objects with aerodynamic uncertainties, such as wind turbines.
[0038] The state space is defined as a normalized vector, containing the ratio of real-time wind speed to rated wind speed, the ratio of real-time air density to standard air density, the ratio of generator speed to rated speed, the ratio of pitch angle to maximum pitch angle, and the ratio of the standard deviation of the power sequence over the past 10 seconds to the rated power. Normalization is necessary to ensure that the numerical ranges of each state component are similar, avoiding the gradient being affected by dimensional differences during training. The introduction of the power standard deviation as a state component allows the agent to perceive the stability of the output power, thus achieving a balance between pursuing high power and power quality.
[0039] The action space consists of two continuous variables: the torque gain coefficient and the pitch angle fine-tuning. The torque gain coefficient acts on the standard optimal mode gain, changing the turbine's operating point in a partial load region by adjusting the slope of the torque-speed curve. The pitch angle fine-tuning is superimposed on the basic pitch angle command and is used for fine-tuning aerodynamic efficiency in low wind speed ranges. The reason for limiting the action space to fine-tuning rather than absolute values is that the basic control logic of the wind turbine is already implemented in the underlying controller, and the reinforcement learning strategy only needs to perform incremental optimization. This design ensures control safety and narrows the action range that the strategy needs to explore, thereby accelerating convergence.
[0040] The core of this strategy lies in constructing a reward function that balances power generation efficiency and mechanical lifespan, expressed as: in, This is the single-step reward value, which is dimensionless. Real-time generator power, in kW or MW; Generator rated power, in kW or MW; This is the pitch angle fine-tuning amount, which is the pitch angle adjustment value output by the reinforcement learning strategy; This is the weighting coefficient for the pitch angle adjustment penalty term; This is the weighting coefficient for the rotational speed and acceleration penalty term.
[0041] The reward function is designed according to the following derivation: The first term is directly proportional to power output, aligning the agent's optimization direction with the fundamental goal of maximizing power generation. The second term imposes a squared penalty on the amplitude of pitch angle adjustment, based on the physical principle that the wear life of the pitch bearing is approximately proportional to the cumulative number of adjustments and the square of the adjustment amplitude: a larger pitch amplitude implies a significant increase in raceway contact stress. The third term imposes a squared penalty on rotational acceleration, based on the mechanical principle that fatigue damage in the transmission chain is related to the amplitude of alternating torque, and rotational acceleration is directly proportional to the imbalance of transmission chain torque; the squared form reflects the nonlinear relationship between fatigue damage and stress amplitude. The reason for using squared forms for each penalty term instead of absolute values is that the squared function has a smaller gradient near zero, allowing for small adjustments; while the penalty gradient increases rapidly during large adjustments, effectively constraining drastic actions. This characteristic allows the agent to find the optimal balance between fine adjustment and load protection.
[0042] The values of the weighting coefficients are determined based on the levelized cost of electricity (LCOE) model. The calibration process is as follows: First, an LCOE model is established for the 20-year lifespan of the wind turbine. LCOE is represented as the ratio of total cost to total power generation. Total cost includes initial investment, operation and maintenance costs, and lifespan depreciation costs due to fatigue damage. A Pareto front that minimizes LCOE is found through multi-objective optimization. A set of weights is then selected on this front that aligns the gradient of the reward function with the gradient of the LCOE. Based on simulation calibration of a typical 4MW unit, the recommended values are... =0.01, =0.005, and can be , The weighting coefficient is adjusted according to the wind farm's operational preferences within the specified range. The relationship between the weighting coefficient and the levelized cost of electricity (LCOE) can be understood as follows: a larger weighting coefficient means a greater emphasis on pitch wear and drivetrain fatigue, which is suitable for scenarios with lower electricity prices or higher operation and maintenance costs; a smaller weighting coefficient prioritizes maximizing power generation, which is suitable for scenarios with higher electricity prices or ample remaining lifespan.
[0043] Both the policy network and the value network employ a fully connected structure. The number of hidden layers and neurons is determined based on the state-action space complexity. The entropy regularization coefficient is automatically adjusted during training, aiming to maintain the policy's entropy at a preset target level, thereby encouraging exploration in the early stages of training and gradually converging in the later stages. A discount factor of 0.99 is used to encourage the agent to focus on the long-term rewards over approximately 100 control steps, and a small soft update coefficient is used to ensure training stability. The experience replay buffer has a capacity of millions of transition samples, and a batch of samples is randomly sampled each control cycle to update network parameters.
[0044] To address the power oscillations in blade passing frequency caused by wind shearing and tower shadow effects, this strategy introduces a wind speed fluctuation pre-compensation mechanism. The physical principle of this mechanism is as follows: during rotor rotation, the relative wind speeds experienced by the blades at different heights and azimuth angles exhibit periodic differences, resulting in aerodynamic torque containing a periodic component dominated by the third harmonic of the rotor speed. This component, when transmitted to the generator side, manifests as power oscillations. The core idea of the compensation is to introduce a feedforward compensation term in the torque command that has an opposite phase to the aforementioned periodic component, thus canceling out the oscillations at their source.
[0045] The calculation of the compensation torque is based on the frequency domain compensation relationship derived in this invention: in, The compensation torque value is expressed in N·m. This is the compensation gain coefficient; For the aerodynamic torque at the current operating point Regarding wind speed The partial derivative value, in units of N·m / (m / s), characterizes the change in aerodynamic torque caused by a unit change in wind speed; For time The time-domain signal of the changing wind speed sequence; It is a Fourier transform operator used to convert time-domain signals to frequency-domain representation; Let be the transfer function of the bandpass filter, where The imaginary unit, Angular frequency, in rad / s; It is an inverse Fourier transform operator used to convert the filtered frequency domain signal back to the time domain to obtain the 3P fluctuation component of wind speed.
[0046] The derivation logic of this relationship is as follows: First, the wind speed time-domain signal is converted to the frequency domain using Fourier transform; second, a bandpass filter centered at three times the rotor speed frequency is used to extract the fluctuation component near the 3P frequency. The passband width of the bandpass filter is designed according to the number of turbine blades and the rotor speed range to ensure coverage of the main periodic disturbance frequencies; then, the filtered frequency-domain signal is inversely transformed back to the time domain to obtain the 3P fluctuation component of the wind speed; finally, the wind speed fluctuation is converted into a compensating torque based on the sensitivity of aerodynamic torque to wind speed. (Sensitivity term) This represents the change in aerodynamic torque caused by a unit change in wind speed at the current operating point, calculated using the local derivative of the aerodynamic characteristic surface. Compensation gain. Based on the unit's structural dynamic characteristics, and considering the mechanical impedance of the transmission chain and the response delay of the converter, the value is typically between 0.5 and 0.8. This compensation mechanism, starting from the basic equations of aero-elastic coupling, achieves targeted suppression of 3P frequency power pulsations. Unlike conventional time-domain filtering methods, it has clear physical meaning and adjustable parameters.
[0047] When the operating condition label is switched to high turbulence gusts, the system runs a load suppression reinforcement learning strategy and employs a near-end policy optimization algorithm. This algorithm is chosen based on its training stability in non-stationary environments: by shearing the objective function to limit the differences between adjacent policies, it avoids the performance collapse that may occur in the policy gradient algorithm when the reward distribution changes drastically, making it particularly suitable for scenarios with large fluctuations in reward signals under gust conditions.
[0048] The state space is supplemented with two load-related dimensions on top of the basic state: the ratio of the bending moment at the tower base to the design limit value, and the ratio of the rotational acceleration to the upper limit of the allowable instantaneous acceleration of the transmission chain. The bending moment at the tower base is a key indicator reflecting the load on the tower structure, and its ratio to the limit value directly characterizes the safety margin of the structure; the rotational acceleration reflects the degree of torque imbalance in the transmission chain and is a driving factor for fatigue damage to the gearbox and main shaft. Incorporating these two physical quantities into the state space enables the agent to perceive the structural response and the dynamic load of the transmission chain.
[0049] The reward function was reconstructed as follows: in, This represents the single-step reward value under the load suppression strategy. This refers to the real-time measured bending moment values before and after the tower base; The design limit values of the bending moments before and after the tower base are expressed in kN·m. This refers to the power loss caused by load shedding control, which is the difference between the theoretical optimal power and the actual power. The generator's rated power is expressed in kW or MW; 5 represents the upper limit reference value for instantaneous acceleration allowed in the transmission chain design specifications. This is the weighting coefficient for the bending moment penalty term; This represents the weighting coefficient for the power loss penalty term; This is the weighting coefficient for the rotational speed acceleration survival reward item.
[0050] The reward function is constructed as follows: The first term constructs a penalty using the square of the moment ratio. The physical basis for this is that fatigue damage at tower welds and bolted connections is proportional to the square of the stress amplitude. According to Basquin's formula for material fatigue, the logarithm of the stress amplitude is linearly related to the logarithm of the fatigue life. Therefore, the squared penalty directly corresponds to the rate of fatigue damage accumulation. When the moment approaches the limit, the squared term increases rapidly, forming a "soft constraint" protection mechanism—unlike rigid constraints, soft constraints allow the moment to be slightly higher for a very short time in exchange for greater power generation benefits, but at the cost of accumulated damage penalties. The second term penalizes the power loss caused by load shedding control, driving the agent to find a control trajectory that achieves load suppression at the lowest power cost, avoiding overly conservative control strategies. The third term is a survival reward for rotational acceleration. The 5 rad / s² in the denominator corresponds to the upper limit of instantaneous acceleration allowed in the transmission chain design specifications; this term provides a positive reward when the acceleration is much lower than the upper limit, and decays to zero when it approaches or exceeds the upper limit. The introduction of this survival reward provides the agent with clear safety boundary guidelines, enabling it to maintain dynamic stability of the transmission chain throughout the load suppression process.
[0051] coefficient =0.3、 =0.2、 The value of 0.1 was determined through simulation analysis of the wind turbine's structural dynamics model. Specifically, fatigue damage calculation models for the tower, main shaft, and gearbox were established. Full-condition simulations were performed under different coefficient combinations to calculate the equivalent fatigue load and power generation of each key component. A comprehensive evaluation index was constructed, and the coefficient combination that optimizes the comprehensive index was selected. The relative importance of the three coefficients reflects the relative importance of each optimization objective: the tower base bending moment has the largest weight because it directly relates to the lifespan of the high-value, difficult-to-replace tower component; power loss has the second largest weight; and rotational speed acceleration has a relatively smaller weight, playing a supporting adjustment role.
[0052] The policy network outputs the action mean and variance to form a Gaussian policy, while the value network outputs the state value. The shearing parameters, generalized advantage estimation parameters, and number of updates per round are selected based on a trade-off between algorithm stability and sample efficiency. During training, multi-environment parallel sampling is employed, meaning multiple simulation environments are run simultaneously to collect experience, thereby improving data throughput and sample diversity, and accelerating training convergence.
[0053] Another important innovation of this invention lies in the adaptive threshold setting for the extreme gust safety protection trigger condition. Traditional fixed thresholds cannot adapt to the differences in the load-bearing capacity of the turbine under different average wind speeds: in low wind speed ranges, the turbine itself operates under partial load, with relatively small aerodynamic torque, and the drivetrain and tower have a large load margin; near the rated wind speed, the turbine is already operating close to full load, and the same rate of wind speed change will cause a larger load increment. Based on this physical understanding, this invention dynamically adjusts the trigger threshold according to the real-time average wind speed, and its calculation formula is: in, =6m / s 2 The base threshold is calculated by back-calculating the relationship between amplitude and rise time in the extreme operating gust model of the IEC61400-1 standard; The window length is set to 30 seconds to filter out instantaneous fluctuations, which is the moving average wind speed. Rated wind speed; =0.15 is the sensitivity coefficient, determined through simulation. This quadratic function takes its minimum value at the rated wind speed, and the threshold increases when deviating from the rated wind speed. The physical basis for this is that the sensitivity of the aerodynamic torque of the wind turbine to wind speed is the greatest near the rated wind speed, at which time the load change caused by gusts is the most drastic, so a relatively conservative triggering condition is required; in the low and high wind speed ranges, the sensitivity decreases, and the triggering condition can be appropriately relaxed to avoid unnecessary protection actions.
[0054] When the instantaneous wind speed change rate exceeds the dynamic threshold and lasts for 0.5 seconds, hardware-level safety protection is triggered. The 0.5-second duration requirement is to prevent false triggering caused by measurement noise spikes. The protection action sequence includes immediately pausing the action output of the reinforcement learning strategy, the pitch system feathering to a pitch angle of 5° at a rate of not less than 3° / s, and the converter linearly reducing the torque command to 0.5 times the rated torque within 0.2 seconds and maintaining this state for 2 seconds. The 2-second protection duration is based on the damping decay time calculation of the first-order torsional vibration mode of the drivetrain: the damping ratio of the torsional vibration mode is typically between 0.02 and 0.05, and the corresponding decay time constant (the time required for the amplitude to decay to 1 / e of the initial value) is approximately 1 to 3 seconds. The 2-second protection duration ensures that the torsional vibration energy is fully dissipated.
[0055] After the protection period ends, a soft recovery phase begins, during which the agent needs to smoothly transition from the safety protection state back to normal control. This invention designs an experience playback filtering and updating mechanism based on operational condition similarity. The formula for measuring similarity is: in, Let this be the current state vector. Let be the state vector of the i-th sample in the experience pool. This is the standard deviation vector for each dimension of the state, used to eliminate the influence of differences in the units of measurement across different dimensions. This represents the weighted Euclidean distance. The weight matrix is used to represent this. As a diagonal matrix, the diagonal elements are assigned according to the sensitivity of state variables to the control strategy: the weight of load-related dimensions (bending moment ratio, acceleration) is 1.5, the weight of operating state dimensions (wind speed, rotational speed) is 1.0, and the weight of environmental dimensions (temperature) is 0.5. This weighted design makes the similarity measurement focus more on the state components directly related to the current control decision.
[0056] The agent selects the 1000 most similar historical experiences and uses these experiences to perform five offline gradient updates, temporarily reducing the learning rate to half of its original value to increase update stability. After the update, the action output smoothly transitions from the protected state value to the new policy output value within 5 seconds. This soft recovery mechanism avoids the potential for abrupt changes in control commands that might occur when the policy network re-explores a random state after protection exits, enabling the unit to quickly and smoothly recover to its optimal operating state after extreme events.
[0057] For aerodynamic performance degradation conditions, when the independent monitoring logic detects that the power coefficient is below 85% of the factory-calibrated theoretical value for 30 consecutive seconds, it is determined that the blades have iced or fouled. The theoretical value is determined based on a lookup table value on the factory aerodynamic characteristic surface for the same tip speed ratio and pitch angle. The 30-second judgment duration ensures the reliability of the detection results and eliminates misjudgments caused by brief changes in wind direction or measurement noise.
[0058] The adaptive mechanism of this invention first superimposes a sinusoidal perturbation onto the original pitch angle command. The perturbation amplitude is dynamically adjusted according to the current turbulence intensity, and the adjustment relationship is given by the following formula: in, The amplitude of the disturbance signal is expressed in degrees (°). =0.6° is the maximum disturbance amplitude. =0.2° is the minimum disturbance amplitude. The current turbulence intensity is dimensionless and is defined as the ratio of the standard deviation of wind speed to the average wind speed. =0.05 is the turbulence attenuation constant. The derivation of this exponential attenuation relationship is based on the signal-to-noise ratio optimization principle: system identification theory shows that the accuracy of parameter estimation depends on the ratio of excitation signal energy to noise energy. In low-turbulence environments... <0.05 indicates relatively small environmental disturbances, requiring a large external excitation to ensure a sufficient signal-to-noise ratio; in highly turbulent environments... With a value >0.15, turbulence itself provides abundant natural excitation, and external excitation can be reduced accordingly to avoid introducing unnecessary loads. The exponential function form ensures that the amplitude decays smoothly with increasing turbulence intensity, avoiding abrupt changes.
[0059] The disturbance frequency was selected as 0.2Hz. This frequency was chosen to avoid several structural resonant frequencies: the first-order front and rear modal frequencies of the tower are typically 0.25-0.35Hz, the first-order flapping frequency of the blades is typically 0.5-0.7Hz, and the torsional vibration frequency of the drive train is typically 1.5-2.5Hz. 0.2Hz is outside these resonant frequency bands and meets the bandwidth requirements of the pitch system. Specifically, the maximum instantaneous pitch rate generated by the disturbance is... ≈0.75° / s, which is far below the typical pitch rate limit of 10° / s, and will not trigger pitch rate protection.
[0060] One hundred sets of data points were continuously collected after the disturbance injection, with the sampling time of each set of data points evenly distributed within the disturbance period to ensure complete capture of the system response. A recursive least squares method with a forgetting factor was used to identify the parameterized model of the aerodynamic characteristic surface online. The identification model used a second-order polynomial to parameterize the aerodynamic characteristic surface. The regression vector consisted of a constant term, a first-order term for the tip speed ratio, a first-order term for the blade pitch angle, a cross term, and the squared terms of the two variables, for a total of six parameters to be identified. The second-order polynomial can fully capture the nonlinear characteristics of the aerodynamic characteristic surface within the normal operating range, while avoiding overfitting and computational burden caused by excessively high orders.
[0061] The value of the forgetting factor has a significant impact on the convergence speed and noise resistance of the identification algorithm. This invention further proposes an adaptive adjustment strategy for the forgetting factor: in, The forgetting factor is dimensionless. =0.99 is the basic forgetting factor. =0.10 represents the reference turbulence intensity, which is a hyperbolic tangent function. The physical basis of this adjustment logic is that the forgetting factor determines the rate at which historical data decays in the current estimate, and its equivalent memory length is approximately... Under steady-state, low-turbulence conditions, aerodynamic characteristics change slowly, and a larger forgetting factor (close to 0.995) should be used to fully utilize historical data and reduce estimation variance. Under high-turbulence or drastic changes in conditions, the forgetting factor should be reduced (close to 0.99) to quickly track the actual changes in aerodynamic characteristics. The hyperbolic tangent function provides a smooth transition from the base value to 1. A reference turbulence intensity of 0.10 corresponds to a moderate turbulence level, at which point the forgetting factor is approximately 0.993, and the memory length is approximately 150 points.
[0062] After obtaining the updated aerodynamic characteristic surface, the optimal tip speed ratio under the current blade pitch angle reference value is solved. The solution method involves a one-dimensional search of the tip speed ratio on the surface model to find the tip speed ratio that maximizes the power coefficient. Then, the rotational speed reference value is calculated, with the conversion relationship being: the rotational speed reference value equals the product of the optimal tip speed ratio and the wind speed divided by the rotor radius. In this mode, the action output of the reinforcement learning strategy is transformed into tracking control of the rotational speed reference value, i.e., adjusting the electromagnetic torque through a proportional-integral controller to converge the actual rotational speed to the reference value. This design achieves decoupling between the control strategy and aerodynamic characteristics. Its key advantage is that when aerodynamic characteristics change due to icing or fouling, only the calculation basis of the rotational speed reference value needs to be updated, without retraining the reinforcement learning network, greatly improving the system's adaptability and deployment efficiency.
[0063] Upon receiving a power-limiting command from the power grid's AGC (Automatic Guided Vehicle) system, the system enters a dynamic power-limiting curve strategy. A normalized power margin dimension is added to the state space, representing the ratio of the difference between the target power and the current power to the rated power. This dimension allows the agent to perceive the magnitude and direction of the power tracking error. The reward function is reconstructed as follows: in, This is the single-step bonus value under power-limited operating conditions; This represents the actual output power of the generator at the current moment. The target power required by the power grid AGC directive; This refers to the generator's rated power. This is the amount of fine adjustment for the pitch angle.
[0064] The reward function is based on the square of the power tracking deviation. The squared form ensures that the reward is symmetrically penalized for both positive and negative deviations, guaranteeing timely bidirectional adjustment. The weighting coefficients are normalized so that when the power deviation is 10% of the rated power, the first penalty is approximately 0.005, which is comparable in magnitude to the second and third penalties, thus achieving multi-objective equilibrium.
[0065] To achieve a smooth transition from maximum power point tracking (MPPT) to limited power point tracking (LPPT), this invention designs a dynamic translation algorithm for the reference power curve. The translated reference power curve is represented as follows: in, This is the reference power value after translation; Real-time wind speed; This is the original optimal power curve, which is the power value that maximizes power generation at each wind speed without considering power rationing orders. The power reduction function is fitted online by a small neural network. The inputs are wind speed, turbulence intensity, and generator speed, and the output is the optimal power reduction under the current conditions. The three input variables are chosen based on the following: wind speed determines the upper limit of available power; turbulence intensity affects the amplitude of power fluctuations, thus impacting the tracking difficulty; and generator speed reflects the unit's current kinetic energy reserves, affecting response speed.
[0066] The network structure for the power reduction function is a single-hidden-layer fully connected network with 16 neurons in the hidden layer, using ReLU as the activation function. The network parameters are continuously updated during power-limited operation using a policy gradient method, with the update objective being to minimize the expected value of the first term in the reward function. The advantages of this method are: the network can learn the complex mapping relationship between the optimal reduction amount and the engine speed and pitch angle under different operating conditions without requiring manual definition of analytical expressions; simultaneously, since the network only outputs the reduction amount rather than directly outputting control commands, the reinforcement learning agent is still responsible for fine-tuning the pitch angle and torque, achieving a hierarchical control architecture.
[0067] In actual operation, the system first estimates the overall reduction ratio based on the target power and available power; then, it multiplies the original optimal torque curve by a correction coefficient related to the reduction ratio and turbulence intensity; the reinforcement learning agent further fine-tunes the pitch angle to ensure that the unit's operating point meets the power limit while minimizing drivetrain impact. This hierarchical and progressive control structure balances response speed and control accuracy.
[0068] When the operating condition label changes, the motion fusion arbitration module intervenes and executes the gradual fusion algorithm. The mathematical expression of this algorithm is: in, The actual output motion vector after fusion includes torque command and pitch angle command; The action value of the strategy before the switch. This represents the action value of the strategy after the switch. Fusion factor. The changing pattern directly affects the smoothness and response speed of the switching process. While linear transitions are simple, sudden acceleration changes at the start and end points can potentially trigger high-frequency responses in the drivetrain. This invention employs a time-based S-shaped transition curve, specifically a cosine transition function: in, The time is calculated from the start of the switch. The transition period is defined as follows. This cosine transition curve possesses the following excellent characteristics: at the start and end times ( =0 and A zero derivative means that the rate of change of the fusion factor is zero at the beginning and end of the transition, thus eliminating abrupt acceleration changes; the rate of change is greatest in the middle of the transition, ensuring timely switching response. The 60-second transition period takes into account the following considerations: the response time constant of the pitch system is approximately 0.5-1 seconds, the converter torque response time constant is approximately 0.1-0.2 seconds, the drivetrain torque build-up time is approximately 2-3 seconds, and the unit aerodynamic response delay is approximately 5-10 seconds. The 60-second transition duration is approximately 6-10 times the aerodynamic response delay, ensuring that the entire system fully responds to changes before accepting new changes, avoiding oscillations caused by the dynamic superposition of multiple time scales.
[0069] If the operating condition label changes again during the transition period, the current transition process will be terminated immediately, and the current actual output will be used as the new value. The latest strategy output is The system restarts a new 60-second transition timer. This processing logic ensures the stability of the system during continuous switching between multiple operating conditions: by resetting the timer and updating the baseline, the accumulation and interference of multiple transition processes are avoided. This mechanism ensures the C1 continuity of the actuator's actions, that is, both position and speed change continuously, and the torque command change rate is always less than the safety threshold, effectively eliminating torque jumps caused by control logic switching and protecting the generator main shaft and gearbox from impact loads.
[0070] The method in this embodiment operates in the central control unit of the wind turbine. This unit contains a multi-core embedded processor, with each reinforcement learning sub-agent running as an independent thread on a different core. Threads exchange data via shared memory. The data quality diagnosis module and the operating condition identification diagnosis module also run as independent threads, forming a pipelined data processing architecture. The outputs of each module are transmitted through a message queue, which employs a priority mechanism, with security-related messages having the highest priority to ensure timely response.
[0071] The storage array uses industrial-grade solid-state drives (SSDs) for persistent storage of experience pool data and model parameters. The storage strategy for experience pool data is as follows: new experiences are written to files hourly, while smaller files are periodically merged to reduce the number of files. Model parameters are stored using version management, retaining the three most recent versions after each update to facilitate rollback in case of policy performance degradation.
[0072] The redundancy protection module operates independently of the main control chip, monitoring the main control unit's heartbeat signal via a hardware watchdog. The heartbeat signal is a pulse sequence emitted by the main control software every 100ms. If no valid heartbeat is received for three consecutive control cycles, the redundancy protection module determines that the main control unit has failed and forcibly takes over the control of the pitch and converter. The control logic after takeover follows a preset safe shutdown curve: the pitch system feathers at a fixed rate to the feathering position, and the converter reduces torque to zero at a fixed slope. This multi-level defense system architecture, from the reward function constraints of the reinforcement learning agent to the smooth switching of the arbitration module, and then to the independent protection of hardware redundancy, forms a defense-in-depth system, providing robust security for the large-scale application of reinforcement learning algorithms with stochastic exploration characteristics in industrial-grade wind power control.
[0073] Example 2 This embodiment, based on Embodiment 1, further describes the system architecture details and provides simulation experimental data to verify the technical effects of the present invention.
[0074] The system comprises a data acquisition and preprocessing module, a data quality diagnosis module, a working condition identification and diagnosis module, a hierarchical reinforcement learning computing engine, an action execution arbitration module, a low-level controller communication interface, and a redundancy protection module. The functional division and interface relationships of each module are as follows: The data acquisition and preprocessing module is responsible for acquiring raw signals from physical sensors and performing hardware filtering, analog-to-digital conversion, and data packaging, outputting a unified format data frame to the data quality diagnosis module; the data quality diagnosis module verifies, compensates, and reconstructs the data frame, outputting cleaned data to the working condition identification and diagnosis module and the hierarchical reinforcement learning computing engine; the working condition identification and diagnosis module outputs the current working condition label based on feature extraction and clustering to the hierarchical reinforcement learning computing engine; the hierarchical reinforcement learning computing engine activates the corresponding sub-agent according to the working condition label and outputs control actions to the action execution arbitration module; the action execution arbitration module performs smooth working condition switching and outputs the final action command to the low-level controller communication interface; the low-level controller communication interface converts the command into a communication protocol recognizable by the pitch driver and converter and sends it out; the redundancy protection module independently monitors the status of each module and directly outputs safety commands in case of anomalies.
[0075] A comparative experiment was conducted on a high-fidelity simulation platform for a 4MW offshore wind turbine, using traditional lookup table-based static power curve control as a scale. The simulation platform integrates an aerodynamic-elastic-servo coupled model, which can accurately simulate the dynamic behavior of the turbine under the interaction of turbulent wind fields, structural vibrations, and control systems.
[0076] In the variable turbulence test, the average wind speed was set to 8.5 m / s, and the turbulence intensity jumped from 0.06 to 0.18 during the operation to simulate sudden gusts. The performance comparison after 2000 seconds of operation is shown in Table 1.
[0077] As shown in Table 1, the improvement in power coefficient stems from the combined optimization of torque gain and fine-tuning pitch angle by the efficiency-first strategy. During the low-turbulence phase, the reinforcement learning strategy uses small pitch adjustments to bring the unit's operating point closer to the theoretical optimal curve; during gusts, the strategy automatically adjusts the amplitude of the action to maintain stable power. A 34% reduction in the power standard deviation indicates that this invention effectively suppresses power fluctuations caused by turbulence, an improvement significant for meeting grid-connected power quality requirements. A 12.71% reduction in tower base fatigue load verifies the proactive load reduction effect of the load suppression strategy. According to the Miner criterion for fatigue damage, a 12.71% load reduction can extend the tower's fatigue life by approximately 30%. A moderate increase in pitch operation frequency reflects the proactive adjustment behavior of the agent to achieve performance optimization. Calculations show that even at a +25% operation frequency, the design life of the pitch bearing still has sufficient margin.
[0078] In the aerodynamic degradation experiment, the aerodynamic performance degradation caused by increased leading-edge roughness was simulated by modifying the blade aerodynamic data file. The comparative example still used the factory-installed aerodynamic surface lookup table, causing the operating point to deviate from the optimal region, resulting in a power generation drop to 82.5% of the normal value. This invention, upon detecting degradation, reconstructs the aerodynamic surface and updates the speed reference value within approximately 240 seconds through excitation signal injection and online identification, restoring power generation to 94.2% of the normal value, a relative improvement of 14.18 percentage points. The root mean square error between the identified converged aerodynamic surface and the actual degraded surface is less than 0.008, verifying the identification accuracy of the recursive least squares algorithm in noisy environments.
[0079] In the power limit command tracking experiment, the AGC command requires the generator to reduce its output power from full capacity to 80% of its rated power. The comparative method using a fixed pitch angle offset exhibits significant overshoot and steady-state error. This invention, by dynamically adjusting the reference power curve and supplementing it with reinforcement learning fine-tuning, reduces the root mean square error of power tracking from 45.2 kW to 12.8 kW, a reduction of 71.68%, with smooth pitch control and no overshoot. The transition time is approximately 45 seconds, far less than the grid's power response time requirements.
[0080] Sensitivity analysis of the reward function weighting coefficients and adaptive thresholds shows that, within the recommended parameter range, all performance indicators of this invention remain stable and superior to those of the comparative example, indicating that the scheme has good parameter robustness. Furthermore, in the test simulating wind speed sensor failure, the virtual wind speed reconstruction algorithm can control the wind speed estimation deviation within 0.2 m / s, and the power control performance degradation does not exceed 3%, verifying the effectiveness of the data quality diagnostic module.
[0081] In summary, this invention organically combines data quality diagnosis, adaptive clustering based on operating conditions, multi-scenario reinforcement learning, and a smooth arbitration mechanism to construct a dynamic optimization method for wind power curves applicable to all operating conditions and the entire life cycle. This method not only significantly improves the energy capture efficiency and grid-connected power quality of wind turbines but also extends the lifespan of key components through load-sensing control and endows the turbines with adaptive capabilities to cope with performance degradation through online aerodynamic parameter identification. More importantly, this invention solves the safety issues of applying reinforcement learning algorithms in industrial control through the design of a hierarchical parallel architecture and independent redundant protection modules, providing a practical and feasible technical path for the intelligent control of large-scale wind turbines.
[0082] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0083] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for dynamic optimization and adjustment of wind power curves based on reinforcement learning, characterized in that: Includes the following steps: Step S1: Obtain real-time operating parameters through the data acquisition layer deployed on the wind turbine, and perform data quality diagnosis on the real-time operating parameters. When a sensor fault is detected, perform virtual wind speed reconstruction. When a data delay is detected to exceed a preset threshold, perform time-series interpolation compensation to obtain the operating data after quality repair. Step S2: Extract features from the operational data after quality repair, perform real-time clustering of operating conditions based on the extracted feature vectors, and output the current operating condition label based on the clustering results; Step S3: Based on the current operating condition label, activate the corresponding sub-agent from the pre-set multi-scenario reinforcement learning strategy library, and the sub-agent outputs preliminary control actions according to the current operating state; wherein, the multi-scenario reinforcement learning strategy library includes at least the efficiency-first reinforcement learning strategy corresponding to steady-state low-turbulence operating conditions, the load suppression reinforcement learning strategy corresponding to high-turbulence gust operating conditions, the aerodynamic parameter adaptive strategy corresponding to aerodynamic performance degradation operating conditions, and the power limiting dynamic curve strategy corresponding to power grid power limiting command operating conditions; Step S4: Perform fusion arbitration on the preliminary control actions generated by the switching of different operating conditions and strategies, generate the final control command using a gradual fusion method, and send the final control command to the wind turbine actuator.
2. The method for dynamic optimization and adjustment of wind power curves based on reinforcement learning according to claim 1, characterized in that: The virtual wind speed reconstruction in step S1 specifically includes: The generator speed and the effective wind speed at hub height are used as the state vector, the generator speed measurement and the generator power measurement are used as the observation vector, the drive chain rotational inertia balance equation is used as the basis for state transition, and the current effective wind speed is estimated recursively through extended Kalman filtering. When reference wind speed data from lidar or anemometers is available, the extended Kalman filter estimate is weighted and fused with the reference wind speed. The fusion weight is determined by the variance of the extended Kalman filter posterior error and the variance of the reference wind speed measurement noise according to the minimum variance criterion.
3. The method for dynamic optimization and adjustment of wind power curves based on reinforcement learning according to claim 1, characterized in that: Step S2 specifically includes: Extract multi-dimensional feature vectors at a preset period. The multi-dimensional feature vectors include at least the average wind speed, turbulence intensity, high-frequency energy proportion of wind speed power spectral density, generator power variance, and mean tower base bending moment. After performing principal component analysis to reduce the dimensionality of the multidimensional feature vectors, the density-based DBSCAN clustering algorithm is used for clustering, and candidate working condition labels are determined according to the cluster centers and the preset mapping relationship. Hysteresis acknowledgment based on time continuity is performed on candidate operating condition labels. The candidate label is officially output as the current operating condition label only if the same candidate label remains unchanged for multiple consecutive periods.
4. The method for dynamic optimization and adjustment of wind power curves based on reinforcement learning according to claim 1, characterized in that: The efficiency-first reinforcement learning strategy in step S3 uses the following reward function to guide agent training: in, Real-time generator power, Rated power, The pitch angle fine-tuning amount output by the sub-agent. For rotational acceleration, and These are the weighting coefficients calibrated based on the cost-per-kilowatt-hour model.
5. The method for dynamic optimization and adjustment of wind power curves based on reinforcement learning according to claim 1, characterized in that: The load suppression reinforcement learning strategy in step S3 uses the following reward function to guide agent training: in, The bending moment at the front and rear of the tower base, The design limit values of the bending moments before and after the tower base. The power loss is due to load shedding control. Generator rated power, For rotational acceleration, , , These are the weighting coefficients calibrated based on the wind turbine structural dynamics model.
6. The method for dynamic optimization and adjustment of wind power curves based on reinforcement learning according to claim 1, characterized in that: The aerodynamic parameter adaptive strategy in step S3 specifically includes: When aerodynamic performance degradation is detected, a periodic disturbance signal is superimposed on the original pitch angle command. The amplitude of the disturbance signal is dynamically adjusted according to the current turbulence intensity, so that the disturbance amplitude increases in low turbulence to ensure excitation intensity and decreases in high turbulence to suppress additional load. Using data collected during disturbance injection, the current aerodynamic characteristic surface is identified online through recursive least squares method with forgetting factor. Based on the identification results, the optimal tip speed ratio and speed reference value are updated, so that the action output of the sub-agent is transformed into tracking control of the speed reference value.
7. The method for dynamic optimization and adjustment of wind power curves based on reinforcement learning according to claim 6, characterized in that: In the recursive least squares method with a forgetting factor, the forgetting factor is adaptively adjusted according to the current turbulence intensity. The forgetting factor is increased in steady-state low turbulence to make full use of historical data, and decreased in high turbulence to quickly track changes in aerodynamic characteristics.
8. The method for dynamic optimization and adjustment of wind power curves based on reinforcement learning according to claim 1, characterized in that: The gradient blending method in step S4 is as follows: When the operating condition label changes, the preliminary control action output by the strategy before the change and the preliminary control action output by the strategy after the change are weighted and fused. The fusion weight smoothly increases from zero to one within a preset transition period, and the rate of change of the fusion weight is zero at the beginning and end of the transition period.
9. A wind power curve dynamic optimization and adjustment system based on reinforcement learning, used to implement the method described in any one of claims 1 to 8, characterized in that: include: The data acquisition and quality diagnosis module is used to acquire real-time operating parameters of the wind turbine and perform data quality diagnosis on the real-time operating parameters. When a sensor fault is detected, virtual wind speed reconstruction is performed, and when the data delay exceeds a preset threshold, timing interpolation compensation is performed, and the operating data after quality repair is output. The working condition identification and diagnosis module is used to extract features and perform real-time clustering of working conditions on the operational data after quality repair, and output the current working condition label. The hierarchical reinforcement learning computing engine integrates multiple parallel computing units, which correspond to the efficiency-first reinforcement learning strategy for steady-state low-turbulence conditions, the load suppression reinforcement learning strategy for high-turbulence gust conditions, the aerodynamic parameter adaptive strategy for aerodynamic performance degradation conditions, and the power-limiting dynamic curve strategy for grid power-limiting command conditions. The hierarchical reinforcement learning computing engine activates the corresponding computing unit based on the current working condition label and outputs the initial control action; The action execution arbitration module is used to fuse and arbitrate the preliminary control actions generated by the switching of strategies under different operating conditions, and to generate the final control command using a gradual fusion method. And the underlying controller communication interface, used to send the final control commands to the wind turbine actuators.
10. A wind power curve dynamic optimization and adjustment system based on reinforcement learning according to claim 9, characterized in that: The action execution arbitration module is specifically used to perform weighted fusion of the preliminary control actions output by the strategy before the switch and the preliminary control actions output by the strategy after the switch when the working condition label changes. The fusion weight smoothly increases from zero to one within a preset transition period, and the rate of change of the fusion weight is zero at the beginning and end of the transition period.