A method, system, device and storage medium for optimizing active power shedding control in a wind farm

By constructing and training the SAC agent and constraint controller, the non-convex characteristics and convergence of the model in the fatigue load control of the wind farm are solved, and the optimal allocation of active power of the wind farm and the reduction of fatigue load is achieved, which improves the stability and reliability of the wind farm and reduces maintenance costs.

CN119628117BActive Publication Date: 2025-07-08SHANDONG UNIV
View PDF 2 Cites -1 Cited by

Patent Information

Application Number
CN202411963115.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-07-08
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

In the existing data-driven fatigue load control strategy of wind farm, some models present non-convex characteristics, which lead to difficulty in solving and easy to fall into local optimality. In the reinforcement learning method, the convergence of the agent is greatly affected by hyperparameters and weak generalization, making it difficult to achieve reasonable allocation of active power of wind farms and effective reduction of fatigue load.

Method used

Using SAC agent combined with constraint controller, by constructing and training the policy network, state value network, action-state value network and target action-state value network, deep learning is used to process the nonlinear factors of the wind farm, and constructing an ICNN-based sliding window equivalent fatigue load replacement model to realize the optimized allocation and real-time adjustment of the active power of the wind farm.

Benefits of technology

The reasonable allocation of active power of the wind farm is achieved, the equivalent fatigue load of the entire field is reduced, the fatigue damage of the wind turbine is reduced, the maintenance cost is reduced, the generalization and learning efficiency of the model are improved, and the stability and reliability of the wind farm are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119628117B_ABST
    Figure CN119628117B_ABST
Patent Text Reader

Abstract

The present invention provides a method, system, device, and storage medium for optimizing the active power shedding control of a wind farm, belonging to the technical field of active power control of wind farms, including: constructing and training a SAC agent through the SAC algorithm; obtaining the state quantities of the wind farm; inputting the state quantities of the wind farm into the SAC agent to obtain a vector composed of the initial reference values of the active power of each wind turbine generator; performing secondary calculation by inputting into a constraint controller to obtain a vector composed of the optimized reference values of the active power of each wind turbine generator; the wind farm executes; repeating the loop to adjust the active power of each wind turbine generator in real time. The model of the present invention has strong generalization ability, high stability, reliability, and learning efficiency, can realize the real-time online calculation of equivalent fatigue loads, can reasonably distribute the active power of the wind farm while satisfying the active power constraints of the wind farm, can reduce the overall equivalent fatigue load of the wind farm, and reduce the maintenance cost.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Wind power generation is an environmentally friendly renewable energy generation method. With the large-scale installation and use of wind turbines and the increase in the operating capacity and scale of wind farms, the safety and reliability of wind farm operation have received wide attention. Wind turbines are affected by time-varying wind fields. While generating energy, they are accompanied by structural fatigue. After long-term operation, it will lead to fatigue failure of wind turbines, thereby threatening the reliability of wind farms. In order to ensure the reliability of wind farms, it is necessary to frequently carry out maintenance operations on the wind turbines in the wind farm. Therefore, reducing the fatigue damage of wind turbines and the fatigue load of wind farms is crucial for reducing the operation and maintenance costs of wind farms.

[0003] In the research on reducing the fatigue load of wind farms, the active power control of wind farms is the key research direction. The current control strategies considering the fatigue load of wind farms are fatigue control strategies for wind farms based on physical mechanisms. By constructing the fatigue models of each wind turbine in the wind farm through physical mechanisms, with the goal of minimizing the total fatigue load of the wind farm or balancing the fatigue load distribution, combined with commercial solvers to solve the optimization strategy to achieve fatigue suppression. This method has good effects in specific application scenarios, but the mechanism models constructed by it are difficult to cover all the nonlinear dynamics of wind turbines, which will affect the optimization effect of model predictive control.

[0004] In order to improve the optimization effect of the fatigue load of wind farms, the prior art has also proposed an active power optimization control strategy based on data-driven methods, which has relatively low requirements for wind turbine modeling, and even some can use model-free control algorithms. Among them, some studies construct a wind turbine fatigue damage prediction model based on the RBF (Radial Basis Function) neural network or the residual neural network, and combine genetic algorithms or particle swarm algorithms to solve the control strategy. Its model has good prediction performance, but limited by the non-convex characteristics of the model, the solution is relatively difficult and it is often easy to fall into local optimum; there are also some studies that directly use the method of reinforcement learning for fatigue load suppression control, such as DDPG (Deep Deterministic Policy Gradient) and its derivative improved algorithms, which to a certain extent solve the problem of solving the control strategy for the strongly nonlinear model of wind farms, but the convergence of the agent is greatly affected by hyperparameters and the generalization ability is also relatively weak. Summary of the Invention

[0005] In view of the technical problems in the existing data-driven wind farm fatigue load control strategies, where some models exhibit non-convex characteristics, leading to difficult solutions and being prone to falling into local optima, and in the reinforcement learning method, the convergence of the agent is highly affected by hyperparameters and has weak generalization ability, the invention provides an active power reduction optimization control method, system, device, and storage medium for a wind farm. The model of the invention has strong generalization ability, high stability, reliability, and learning efficiency, can realize real-time online calculation of equivalent fatigue loads, can achieve reasonable distribution of active power in the wind farm under the condition of meeting the active power constraints of the wind farm, can reduce the equivalent fatigue load of the entire wind farm, and reduce the maintenance cost.

[0006] In the first aspect, the invention provides an active power reduction optimization control method for a wind farm. The steps include:

[0007] S1. Construct and train a SAC agent, and construct and train a SAC agent through the SAC algorithm;

[0008] S2. Obtain the state variables of the wind farm;

[0009] S3. Input the state variables of the wind farm into the SAC agent to obtain a vector composed of the initial reference values of the active power of each wind turbine in the wind farm, denoted as ;

[0010] S4. Input into the constraint controller, and the constraint controller performs secondary calculation on to obtain a vector composed of the optimized reference values of the active power of each wind turbine in the wind farm, denoted as ;

[0011] S5. The wind farm executes and adjusts the active power of each wind turbine in the wind farm to be the same as ;

[0012] S6. Repeat steps S2 - S5 to adjust the active power of each wind turbine in the wind farm in real time.

[0013] It should be further noted that the state variables of the wind farm include grid power command, current input wind speed of the wind turbine, current output power of the wind turbine, instantaneous change in the main shaft torque of the wind turbine, and instantaneous change in the tower thrust.

[0014] It should be further noted that in step S1, the steps of constructing and training the SAC agent include:

[0015] S101: Initialize the neural network, and construct a policy network, a state value network, a target state value network, an action-state value network, and a target action-state value network;

[0016] The function representing the policy network is: ;

[0017] In the formula, is the state variable of the wind farm and is the input value of the policy network;

[0018] is the vector composed of the active power values of each wind turbine in the wind farm and is the output value of the policy network;

[0019] The function representing the state value network is: ;

[0020] In the formula, is the state variable of the wind farm and is the input value of the state value network;

[0021] The output value of the state value network is the V value;

[0022] The function representing the target state value network is: ;

[0023] In the formula, is the state variable of the wind farm and is the input value of the target state value network;

[0024] The output value of the target state value network is the target V value;

[0025] The function representing the action-state value network is: ;

[0026] In the formula, s is the state variable of the wind farm, is the vector composed of the active power values of each wind turbine in the wind farm, and are both the input values of the action-state value network;

[0027] The output value of the action-state value network is the Q value;

[0028] The function representing the target action-state value network is: ;

[0029] In the formula, is the state variable of the wind farm, is the vector composed of the active power values of each wind turbine in the wind farm, and are both the input values of the target action-state value network;

[0030] The output value of the target action-state value network is the target Q value;

[0031] The state value network of the initial state is the same as the target state value network, and the action-state value network is the same as the target action-state value network;

[0032] S102: Experience replay buffer setting and hyperparameter setting;

[0033] S103: Collect experience, obtain the wind farm state variables at time t0, denoted as , and input into the policy network to obtain a vector composed of the initial training reference values of the active power of each wind turbine in the wind farm, denoted as ;

[0034] Input into the constraint controller. After secondary calculation by the constraint controller, obtain the optimized training reference values of the active power of each wind turbine in the wind farm, denoted as ;

[0035] The wind farm executes , and adjusts the active power of each wind turbine in the wind farm to be the same as . The execution end time is time t1;

[0036] Obtain the wind farm state variables, load data at time t1, and load data within a specified time window before time t1. The wind farm state variables at time t1 are denoted as , and then input the load data at time t1 and the load data within a specified time window before time t1 into the reward function to obtain the reward value for ;

[0037] Input , , the load data at time t1, and the reward value to form an experience tuple, and store the experience tuple in the experience replay buffer;

[0038] S104: Train the network, randomly extract a batch of experience samples from the experience replay buffer;

[0039] Use the extracted experience samples to calculate the target Q value using the target action-state value network and calculate the target V value using the target state value network;

[0040] Update the state value network according to the calculated target Q value, action entropy, and the output of the current state value network. The loss function formula of the state value network is:

[0041]

[0042] In the formula, are the weights of the state value network;

[0043] are the weights of the target action-value network;

[0044] is the weight of the policy network,

[0045] is the state quantity of the wind farm at time t0;

[0046] represents the experience replay buffer;

[0047] is the function representing the state value network;

[0048] is the function representing the target action-state value network;

[0049] is the function representing the policy network;

[0050] Update the action-state value network according to the calculated target V value, action entropy, and the output of the current action-state value network. The loss function formula of the action-state value network is:

[0051]

[0052] In the formula, is at state, take the reward value obtained by the action;

[0053] is the weight of the target state value network;

[0054] is the weight of the action-value network;

[0055] is the discount factor, and its value range is between 0 and 1;

[0056] is the function representing the target state value network;

[0057] is the function representing the action-state value network;

[0058] Calculate the update gradient of the policy network according to the extracted experience samples and the reinforcement learning algorithm, and update the policy network. The loss function formula of the policy network is:

[0059]

[0060] In the formula, is the function of adding noise;

[0061] is the noise vector, which conforms to the Gaussian spherical distribution;

[0062] Update the target action-state value network and the target state value network using the soft update method;

[0063] S105: Iterative training. Repeat S103 - 104 to perform iterative training of the SAC agent. After the iterative training is completed, the required SAC agent is obtained.

[0064] It should be further noted that the payload data includes the main shaft torque and the tower thrust.

[0065] It should be further noted that in step S103, the reward function is constructed by the sliding window equivalent fatigue load substitution model based on ICNN. The functional expression in the sliding window equivalent fatigue load substitution model based on ICNN is:

[0066]

[0067] In the formula, is the input vector, which is the payload data of the wind farm at time t1 and the payload data within the specified time window before time t1;

[0068] is the output vector of the layer;

[0069] is the output vector of the layer;

[0070] is the activation function, and the convex non-decreasing function Smoothed ReLU is used as the activation function ;

[0071] and are the weight vectors related to the layer;

[0072] is the bias term of the layer;

[0073] is a convex function with respect to , is a constant after training;

[0074] is the total number of layers;

[0075] The reward value calculated using the reward function is the equivalent fatigue load of the wind farm at time t1 and within the specified time window before time t1.

[0076] It should be further noted that in step S105, the total number of iterations of the iterative training ≥ 80000 times.

[0077] Further, it should be noted that the parameters of the constraint controller include the active power output constraint of the wind turbine, the dispatching instruction constraint of the transmission system operator, and the active power ramp rate constraint of the wind turbine. The solution condition of the constraint controller is as follows:

[0078]

[0079] In the formula, is the optimization objective, indicating the minimization of ;

[0080] is the lower limit of the theoretical output of the j-th wind turbine, is the upper limit of the theoretical output of the j-th wind turbine;

[0081] m is the total number of wind turbines;

[0082] is the active power dispatching instruction of the j-th wind turbine;

[0083] is the initial value of the active power dispatching instruction of the j-th wind turbine;

[0084] is the active power dispatching instruction of the j-th wind turbine at the previous moment;

[0085] is the dispatching instruction of the transmission system operator;

[0086] is the absolute value of the power change amount per unit time of the j-th wind turbine;

[0087] is the maximum change value of the power of the j-th wind turbine per unit time.

[0088] In the second aspect, the present invention provides a wind farm active power reduction optimization control system for implementing the above-mentioned wind farm active power reduction optimization control method, including:

[0089] A data acquisition module for acquiring the status quantities and load data of the wind farm;

[0090] An SAC agent construction and training module for constructing and training an SAC agent;

[0091] A constraint controller module for performing secondary calculation on the initial actions generated by the SAC agent;

[0092] An execution module for executing the actions optimized by the constraint controller module to adjust the active power of each wind turbine in the wind farm.

[0093] In a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is configured to implement the steps of the above-mentioned active power reduction optimization control method for a wind farm when executing the computer program.

[0094] In a fourth aspect, the present invention provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-mentioned active power reduction optimization control method for a wind farm.

[0095] The beneficial effects of the present invention are as follows:

[0096] 1. The active power reduction optimization control method, system, device, and storage medium for a wind farm provided by the present invention construct and train a SAC agent through the SAC algorithm, utilize deep learning to process the non-linear factors in the wind farm, realize the reasonable distribution of the active power of the wind farm, can reduce the overall equivalent fatigue load of the wind farm, reduce the fatigue damage of wind turbines, and thus reduce the maintenance cost; the SAC algorithm has low sensitivity to hyperparameters, strong model generalization, fast calculation speed, can adapt to different environments, high stability and reliability, and high learning efficiency.

[0097] 2. The present invention constructs the reward function of the SAC agent based on the ICNN-based sliding window equivalent fatigue load substitution model, can transform the complex fatigue load calculation in the wind farm into a convex function form, can improve the calculation speed while maintaining high accuracy, realize the real-time online calculation of the equivalent fatigue load, can improve the construction and training efficiency of the SAC agent, and further improve the convergence of reinforcement learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0099] Figure 1 is a flowchart of the active power reduction optimization control method for a wind farm in an embodiment of the present invention.

[0100] Figure 2 is a flowchart of the steps of constructing and training a SAC agent in an embodiment of the present invention.

[0101] Figure 3 is a schematic block diagram of the active power reduction optimization control system for a wind farm in an embodiment of the present invention.

[0102] Figure 4 is a schematic hardware structure diagram of an electronic device in an embodiment of the present invention. Detailed implementation manners

[0103] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0104] The active power reduction optimization control method for a wind farm involved in this application mainly focuses on the active power control of the wind farm. During the operation of the wind farm, by constructing and training a SAC agent, using deep learning to process the non-linear factors in the wind farm, and using the SAC agent and the constraint controller to process the state quantities of the wind farm in sequence, a vector composed of the optimized reference values of the active power of each wind turbine in the wind farm is obtained, and then the active power of each wind turbine in the wind farm is adjusted in real time to achieve the reasonable distribution of the active power of the wind farm, reduce the overall equivalent fatigue load of the wind farm, reduce the fatigue damage of the wind turbines, and thus reduce the maintenance cost. At the same time, this application constructs the reward function of the SAC agent based on the ICNN sliding window equivalent fatigue load substitution model, which can transform the complex fatigue load calculation in the wind farm into a convex function form, can improve the calculation speed while maintaining high accuracy, realize the real-time online calculation of the equivalent fatigue load, can improve the construction and training efficiency of the SAC agent, and further improve the convergence of reinforcement learning.

[0105] The active power reduction optimization control method for a wind farm involved in this application mainly aims at the technical problems in the existing data-driven fatigue load control strategy for wind farms, such as the non-convex characteristics of some models leading to difficult solution and easy to fall into local optimum; in the reinforcement learning method, the convergence of the agent is greatly affected by hyperparameters and the generalization ability is weak.

[0106] The following will describe in detail the active power reduction optimization control method for a wind farm involved in this application. For the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of this application. However, those skilled in the art should clearly understand that this application can also be implemented in other embodiments without these specific details.

[0107] In the active power load shedding optimization control method of the wind farm involved in this application, the term "including" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations. The terms "including", "comprising", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0108] For the convenience of clearly describing the technical solutions of this application, words such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit being different.

[0109] The statements such as "an embodiment" or "some embodiments" described in this application mean that the specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of this application. Thus, statements such as "in an embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" that appear in different places in this application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0110] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0111] The active power load shedding optimization control method of the wind farm provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the active power load shedding optimization control system of the wind farm runs in the computer device.

[0112] Figure 1 It is a flowchart of the active power load shedding optimization control method of the wind farm in an embodiment of the present invention. Among them, Figure 1 The execution subject can be an active power load shedding optimization control system of a wind farm. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.

[0113] As Figure 1 shown, the active power load shedding optimization control method of the wind farm includes:

[0114] S1. Construct and train a SAC agent through the SAC algorithm.

[0115] Construct an SAC agent through the SAC algorithm, aiming to enable the SAC agent to automatically learn how to adjust the active power of each wind turbine under different wind farm states by using the method of reinforcement learning, so as to achieve the optimal control of active power reduction in the wind farm. This intelligent optimization control can replace the traditional control method based on fixed rules or experience, and improve the accuracy and adaptability of control.

[0116] In some embodiments, as Figure 2 shown, the steps of constructing and training the SAC agent in step S1 include:

[0117] S101: Construct a policy network, a state value network, a target state value network, an action-state value network, and a target action-state value network;

[0118] The function representing the policy network is: ;

[0119] In the formula, is the wind farm state quantity, which is the input value of the policy network;

[0120] is a vector composed of the active power values of each wind turbine in the wind farm, which is the output value of the policy network;

[0121] The function representing the state value network is: ;

[0122] In the formula, is the wind farm state quantity, which is the input value of the state value network;

[0123] The output value of the state value network is the V value;

[0124] The function representing the target state value network is: ;

[0125] In the formula, is the wind farm state quantity, which is the input value of the target state value network;

[0126] The output value of the target state value network is the target V value;

[0127] The function representing the action-state value network is: ;

[0128] In the formula, s is the wind farm state quantity, is a vector composed of the active power values of each wind turbine in the wind farm, and are both input values of the action-state value network;

[0129] The output value of the action-state value network is the Q value;

[0130] The function representing the target action-state value network is: ;

[0131] In the formula, is the state quantity of the wind farm, is the vector composed of the active power values of each wind turbine in the wind farm, and are both input values of the target action-state value network;

[0132] The output value of the target action-state value network is the target Q value;

[0133] The state value network in the initial state is the same as the target state value network, and the action-state value network is the same as the target action-state value network;

[0134] Among them, the policy network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons, and the second hidden layer has 128 neurons;

[0135] The state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons, and the second hidden layer has 128 neurons;

[0136] The target state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons, and the second hidden layer has 128 neurons;

[0137] The action-state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons, and the second hidden layer has 128 neurons;

[0138] The target action-state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons, and the second hidden layer has 128 neurons.

[0139] Constructing the policy network, action-state value network, target action-state value network, state value network, and target state value network provides a framework for learning and decision-making for the SAC agent. These networks can output reasonable reference values for the active power of wind turbines according to different input parameters, thereby realizing intelligent adjustment of the active power of the wind farm, laying a foundation for subsequent training and optimization processes, and enabling the SAC agent to gradually learn the optimal decision-making strategy.

[0140] S102: Set up the experience replay buffer and hyperparameters. Set the control step size of the SAC algorithm to 1 s, the number of training steps to 162,000 steps, BatchSize to 256, Gamma to 0.9, and the learning rate of all networks to 0.001.

[0141] The experience replay buffer can store the experiences of the SAC agent during the interaction with the environment, avoiding starting from a completely new state every time training is performed, and improving the training efficiency. By randomly sampling experience samples for training, the SAC agent can better explore different state and action spaces, increasing the stability and generalization ability of the training; the hyperparameter settings can be adjusted according to the specific wind farm situation and optimization objectives to obtain better training results.

[0142] S103: Obtain the state variables of the wind farm at time t0, denoted as , and input into the policy network to obtain a vector composed of the initial training reference values of the active power of each wind turbine in the wind farm, denoted as ;

[0143] Input into the constraint controller. After the secondary calculation of the constraint controller, obtain the optimized training reference values of the active power of each wind turbine in the wind farm, denoted as ;

[0144] The wind farm executes , adjust the active power of each wind turbine in the wind farm to be the same as , and the end time of the execution is time t1;

[0145] Obtain the state variables of the wind farm, the load data at time t1, and the load data within the specified time window before time t1. The state variables of the wind farm at time t1 are denoted as , then input the load data at time t1 and the load data within the specified time window before time t1 into the reward function to obtain the reward value for ;

[0146] Input , , the load data at time t1, and the reward value to form an experience tuple, and store the experience tuple in the experience replay buffer.

[0147] Among them, the wind farm status variables reflect the real-time operating status and demands of the wind farm, and the load data is crucial for evaluating the operating status and fatigue life of wind turbines. Obtaining these data provides comprehensive decision-making basis for the SAC agent; the wind farm status variables are input into the policy network to obtain the initial reference value of active power training, and then the optimized reference value is obtained through the constraint controller. This process enables the SAC agent to learn how to adjust the active power of wind turbines to meet various constraint conditions and maximize the reward value under different wind farm states; the wind farm executes the optimized active power reference value, and inputs the executed status variables and load data into the reward function to obtain the reward value, and then forms an experience tuple for storage. This step provides a rich data source for subsequent training, enabling the SAC agent to continuously learn and improve from its own experiences.

[0148] S104: Randomly extract a batch of experience samples from the experience replay buffer;

[0149] Using the extracted experience samples, calculate the target Q value using the target action-state value network, and calculate the target V value using the target state value network;

[0150] According to the calculated target Q value, action entropy, and the output of the current state value network, update the state value network. The loss function formula of the state value network is:

[0151]

[0152] In the formula, is the weight of the state value network;

[0153] is the weight of the target action-value network;

[0154] is the weight of the policy network,

[0155] is the wind farm status variable at time t0;

[0156] represents the experience replay buffer;

[0157] is the function representing the state value network;

[0158] is the function representing the target action-state value network;

[0159] is the function representing the policy network;

[0160] Update the action-state value network according to the calculated target V value, the action entropy, and the output of the current action-state value network. The loss function formula of the action-state value network is as follows:

[0161]

[0162] In the formula, is the reward value obtained by taking the action in the state;

[0163] are the weights of the target state value network;

[0164] are the weights of the action-value network;

[0165] is the discount factor, and its value range is between 0 and 1;

[0166] is the function representing the target state value network;

[0167] is the function representing the action-state value network;

[0168] Calculate the update gradient of the policy network according to the extracted experience samples and the reinforcement learning algorithm, and update the policy network. The loss function formula of the policy network is as follows:

[0169]

[0170] In the formula, is the function for adding noise;

[0171] is the noise vector, which conforms to the Gaussian spherical distribution;

[0172] Update the target action-state value network and the target state value network using soft update.

[0173] The state value network is used to evaluate the value of different states, and the action-state value network is used to evaluate the value of different states and actions. By continuously updating the state value network and the action-state value network, the SAC agent can more accurately judge which actions are beneficial and which are not. The policy network is used to determine the actions taken by the SAC agent in different states. By continuously updating the policy network, the SAC agent can learn better decision-making strategies to maximize long-term rewards. The target state value network and the target action-state value network are updated using a soft update method, which avoids drastic changes in the target network and improves the stability of training. The soft update method can make the target network gradually approach the current network, thereby reducing fluctuations during training.

[0174] S105: Repeat S103 - 104 to perform iterative training of the SAC agent. After the iterative training ends, the required SAC agent is obtained.

[0175] Through iterative training, the SAC agent can continuously optimize its decision-making strategy. As the number of iterations increases, the SAC agent can gradually learn a better active power adjustment scheme for the wind farm, improving the operation efficiency and reliability of the wind farm.

[0176] In some embodiments, the load data includes main shaft torque and tower thrust.

[0177] In some embodiments, in step S103, a reward function is constructed based on the ICNN-based sliding window equivalent fatigue load surrogate model. The functional expression in the ICNN-based sliding window equivalent fatigue load surrogate model is:

[0178]

[0179] In the formula, is the input vector, which is the load data of the wind farm at time t1 and the load data within a specified time window before time t1;

[0180] is the output vector of the th layer;

[0181] is the output vector of the th layer;

[0182] is the activation function, and the convex non-decreasing function Smoothed ReLU is used as the activation function ;

[0183] and are the weight vectors related to the layer;

[0184] is the bias term of the layer;

[0185] is a convex function with respect to and is a constant after training; is a constant after training;

[0186] is the total number of layers;

[0187] The reward value calculated using the reward function is the equivalent fatigue load of the wind farm within a specified time window before and at time t1.

[0188] By constructing a reward function with a sliding window equivalent fatigue load substitution model based on ICNN, key fatigue load factors such as main shaft torque and tower thrust during the operation of the wind turbine are taken into consideration. The load conditions over a period of time are dynamically evaluated using a sliding window, and the ICNN is used to accurately process data features, thereby guiding the SAC agent to learn a strategy that can effectively adjust the active power and reduce the fatigue damage of the wind turbine.

[0189] In some embodiments, in step S105, the total number of iterations of iterative training ≥ 80000 times.

[0190] Limiting the total number of iterations of iterative training ≥ 80000 times can ensure that the SAC agent is fully trained, enabling the SAC agent to fully explore the state and action spaces and converge to a better decision-making strategy.

[0191] Step S2, obtain the wind farm state variables.

[0192] In some embodiments, the wind farm state variables include grid power command, current input wind speed of the wind turbine, current output power of the wind turbine, instantaneous change in main shaft torque of the wind turbine, and instantaneous change in tower thrust.

[0193] Obtaining the above wind farm state variable information can provide comprehensive wind farm operation state data for subsequent decision-making, enabling the SAC agent to make reasonable adjustments to the active power according to the actual situation.

[0194] Step S3, input the wind farm state variables into the SAC agent to obtain a vector composed of the initial reference values of the active power of each wind turbine in the wind farm, denoted as .

[0195] This step generates the initial actions, which can initially generate the initial reference values of the active power of each wind turbine and provide a basis for the constraint adjustment in the subsequent steps.

[0196] Step S4, input The input constraint controller, and the constraint controller performs a secondary calculation to obtain a vector composed of the reference values for optimizing the active power of each wind turbine in the wind farm, denoted as .

[0197] By performing a secondary calculation on the initial action to obtain a vector composed of the reference values for optimizing the active power, the stability and reliability of the operation of the wind farm are improved.

[0198] In some embodiments, the parameters of the constraint controller include the active power output constraint of the wind turbine, the dispatching instruction constraint of the transmission system operator, and the active power ramp rate constraint of the wind turbine. The solution conditions of the constraint controller are:

[0199]

[0200] In the formula, is the optimization objective, indicating the minimization of ;

[0201] is the lower limit of the theoretical output of the j-th wind turbine, is the upper limit of the theoretical output of the j-th wind turbine;

[0202] m is the total number of wind turbines;

[0203] is the active power dispatching instruction of the j-th wind turbine;

[0204] is the initial value of the active power dispatching instruction of the j-th wind turbine;

[0205] is the active power dispatching instruction of the j-th wind turbine at the previous moment;

[0206] is the dispatching instruction of the transmission system operator;

[0207] is the absolute value of the power change amount per unit time of the j-th wind turbine;

[0208] is the maximum change value of the power of the j-th wind turbine per unit time.

[0209] This constraint controller considers parameters such as the active power output constraint of the wind turbine, the dispatching instruction constraint of the transmission system operator, and the active power ramp rate constraint of the wind turbine, ensuring that the adjustment of the active power of the wind turbine is within a reasonable range and meets the requirements of actual operation.

[0210] Step S5, the wind farm executes , adjust the active power of each wind turbine in the wind farm to be the same as the same.

[0211] The wind farm executes the optimized active power reference value, enabling the wind turbines to adjust according to the decisions of the SAC agent, achieving optimized active power reduction control and improving the operation efficiency and reliability of the wind farm.

[0212] Step S6, repeat steps S2 - S5 to adjust the active power of each wind turbine in the wind farm in real time.

[0213] By repeating steps S2 - S5 in real time, the active power can be continuously adjusted according to the actual state of the wind farm, adapting to the dynamic changes of the wind farm operation and ensuring that the wind farm is always in the optimal operation state.

[0214] In a specific embodiment, the optimized active power reduction control method for the wind farm includes:

[0215] Step S1, construct and train the SAC agent. Construct and train the SAC agent through the SAC algorithm. Build a 50MW small - scale wind farm with 10 wind turbines based on the NERL5MW wind turbine model as the simulation environment to verify the algorithm performance. The hardware environment used is: 12th Gen Intel Core i7 - 12700H, NVIDIA GeForce RTX 3060 Laptop, 16Gb memory. The steps include:

[0216] S101: Initialize the neural network, construct the policy network, state value network, target state value network, action - state value network, and target action - state value network;

[0217] Among them, the policy network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons and the second hidden layer has 128 neurons;

[0218] The state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons and the second hidden layer has 128 neurons;

[0219] The action - state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons and the second hidden layer has 128 neurons;

[0220] The action - state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons and the second hidden layer has 128 neurons;

[0221] The target action-state value network includes an input layer, a first hidden layer, a second hidden layer, and an output layer. The first hidden layer has 256 neurons, and the second hidden layer has 128 neurons;

[0222] The function representing the policy network is: ;

[0223] In the formula, is the state quantity of the wind farm and is the input value of the policy network;

[0224] is the vector composed of the active power values of each wind turbine in the wind farm and is the output value of the policy network;

[0225] The function representing the state value network is: ;

[0226] In the formula, is the state quantity of the wind farm and is the input value of the state value network;

[0227] The output value of the state value network is the V value;

[0228] The function representing the target state value network is: ;

[0229] In the formula, is the state quantity of the wind farm and is the input value of the target state value network;

[0230] The output value of the target state value network is the target V value;

[0231] The function representing the action-state value network is: ;

[0232] In the formula, s is the state quantity of the wind farm, is the vector composed of the active power values of each wind turbine in the wind farm, and are both input values of the action-state value network;

[0233] The output value of the action-state value network is the Q value;

[0234] The function representing the target action-state value network is: ;

[0235] In the formula, is the state quantity of the wind farm, is the vector composed of the active power values of each wind turbine in the wind farm, and are both input values of the target action-state value network;

[0236] The output value of the target action-state value network is the target Q value;

[0237] The state value network of the initial state is the same as the target state value network, and the action-state value network is the same as the target action-state value network;

[0238] S102: Set the experience replay buffer and hyperparameters. Set the control step size of the SAC algorithm to 1 s, the number of training steps to 162,000 steps, the BatchSize to 256, the Gamma to 0.9, and the learning rate of all networks to 0.001;

[0239] S103: Obtain the wind farm state variables at time t0, denoted as , and input into the policy network to obtain a vector composed of the initial reference values of the active power training of each wind turbine in the wind farm, denoted as ;

[0240] Input into the constraint controller. After the secondary calculation of the constraint controller, obtain the optimized reference values of the active power training of each wind turbine in the wind farm, denoted as ;

[0241] The wind farm executes , and adjusts the active power of each wind turbine in the wind farm to be the same as , The execution end time is time t1;

[0242] Obtain the wind farm state variables, load data at time t1, and load data within a specified time window before time t1 of the wind farm. The wind farm state variables at time t1 are denoted as , and then input the load data at time t1 and the load data within the specified time window before time t1 into the reward function to obtain the reward value for ;

[0243] Construct a reward function based on the ICNN sliding window equivalent fatigue load substitution model. The ICNN has 4 hidden layers with dimensions of 128 each. The time window size is 10 s, the sliding step size is 1 s, the number of training times is 500,000 times, the BatchSize is 128, and the learning rate is 0.0001. The function expression in the ICNN sliding window equivalent fatigue load substitution model is:

[0244]

[0245] In the formula, is the input vector, which is the load data of the wind farm at time t1 and the load data within the specified time window before time t1;

[0246] is the output vector of the th layer;

[0247] is the output vector of the th layer;

[0248] is the activation function, and the convex non-decreasing function Smoothed ReLU is used as the activation function ;

[0249] and are weight vectors related to the layer;

[0250] is the bias term of the th layer;

[0251] is a convex function with respect to and is a constant after training; After training, it is a constant;

[0252] is the total number of layers;

[0253] The reward value calculated using the reward function is the equivalent fatigue load of the wind farm within the specified time window at and before time t1;

[0254] Combine , , the load data at time t1 and the reward value to form an experience tuple, and store the experience tuple in the experience replay buffer;

[0255] S104: Randomly sample a batch of experience samples from the experience replay buffer;

[0256] Using the sampled experience samples, calculate the target Q value using the target action-state value network and calculate the target V value using the target state value network;

[0257] Update the state value network according to the calculated target Q value, action entropy, and the output of the current state value network. The loss function formula of the state value network is:

[0258]

[0259] In the formula, are the weights of the state value network;

[0260] are the weights of the target action-value network;

[0261] are the weights of the policy network,

[0262] is the state quantity of the wind farm at time t0;

[0263] represents the experience replay buffer;

[0264] is a function representing the state value network;

[0265] is a function representing the target action-state value network;

[0266] is a function representing the policy network;

[0267] Update the action-state value network according to the calculated target V value, action entropy and the output of the current action-state value network. The loss function formula of the action-state value network is:

[0268]

[0269] In the formula, is at state, taking the reward value obtained by the action;

[0270] is the weight of the target state value network;

[0271] is the weight of the action-value network;

[0272] is the discount factor, and its value range is between 0 and 1;

[0273] is a function representing the target state value network;

[0274] is a function representing the action-state value network;

[0275] Calculate the update gradient of the policy network according to the extracted experience samples and the reinforcement learning algorithm, and update the policy network. The loss function formula of the policy network is:

[0276]

[0277] In the formula, is the function of adding noise;

[0278] is the noise vector, which conforms to the Gaussian spherical distribution;

[0279] Update the target action-state value network and the target state value network in a soft update manner;

[0280] S105: Repeat steps S103 - 104 to perform iterative training of the SAC agent for 162,000 steps. After the iterative training is completed, the required SAC agent is obtained;

[0281] Step S2: Obtain the wind farm state variables, which include the grid power command, the current input wind speed of the wind turbine, the current output power of the wind turbine, the instantaneous change in the main shaft torque of the wind turbine, and the instantaneous change in the tower thrust;

[0282] Step S3: Input the wind farm state variables into the SAC agent to obtain a vector composed of the initial reference values of the active power of each wind turbine in the wind farm, denoted as ;

[0283] Step S4: Input into the constraint controller. The parameters of the constraint controller include the active power output constraint of the wind turbine, the transmission system operator's dispatch instruction constraint, and the active power ramp rate constraint of the wind turbine. The solution condition of the constraint controller is:

[0284]

[0285] In the formula, is the optimization objective, indicating the minimization of ;

[0286] is the lower limit of the theoretical output of the j-th wind turbine, is the upper limit of the theoretical output of the j-th wind turbine;

[0287] m is the total number of wind turbines;

[0288] is the active power dispatch instruction of the j-th wind turbine;

[0289] is the initial value of the active power dispatch instruction of the j-th wind turbine;

[0290] is the active power dispatch instruction of the j-th wind turbine at the previous moment;

[0291] is the transmission system operator's dispatch instruction;

[0292] is the absolute value of the power change per unit time of the j-th wind turbine;

[0293] is the maximum value of the power change per unit time of the j-th wind turbine;

[0294] The constraint controller performs a secondary calculation to obtain a vector composed of the reference values for optimizing the active power of each wind turbine in the wind farm, denoted as ;

[0295] Step S5, the wind farm executes and adjusts the active power of each wind turbine in the wind farm to be the same as ;

[0296] Step S6, repeat steps S2 - S5 to adjust the active power of each wind turbine in the wind farm in real time.

[0297] The following are embodiments of the active power shedding optimization control system for a wind farm provided by the present disclosure. This active power shedding optimization system and the active power shedding optimization control methods for the wind farms in the above - mentioned various embodiments belong to the same inventive concept and are used to implement the active power shedding optimization control methods for the wind farms in the above - mentioned various embodiments. Details not described in detail in the embodiments of the active power shedding optimization control system for a wind farm can refer to the embodiments of the active power shedding optimization control method for a wind farm.

[0298] Now, mobile terminals implementing various embodiments of the present invention will be described with reference to the accompanying drawings. In the following description, suffixes such as "module", "component", or "unit" used to represent elements are only for the convenience of explaining embodiments of the present invention and have no specific meaning in themselves. Therefore, "module" and "component" can be used interchangeably.

[0299] As Figure 3 shown, the active power shedding optimization control system includes:

[0300] A data acquisition module for acquiring the state variables and load data of the wind farm;

[0301] An agent construction and training module for constructing and training a SAC agent;

[0302] A constraint controller module for performing a secondary calculation on the initial actions generated by the agent;

[0303] An execution module for executing the optimized actions of the constraint controller module to adjust the active power of each wind turbine in the wind farm.

[0304] This application also provides an electronic device implementing various embodiments of the present invention. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor.

[0305] Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not limit the electronic device. The electronic device may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.

[0306] Figure 4 Schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.

[0307] The electronic device 500 includes, but is not limited to: components such as a processor 501, a network module 502, an audio output unit 503, an input unit 504, a display unit 506, a user input unit 507, an interface unit 508, and a memory 509. Those skilled in the art can understand that the structure of the electronic device involved in the embodiments of the present invention does not limit the electronic device. The electronic device may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements.

[0308] In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments described and / or claimed in this application.

[0309] In the embodiments of the present application, the processor 501 may be implemented by using at least one of an application specific integrated circuit (ASIC), a digital signal processor (DSP), a digital signal processing device (DSPD), a processor, a controller, a microcontroller, a microprocessor, and an electronic unit designed to execute the functions described herein. In some cases, such an implementation may be implemented in the controller. For a software implementation, an implementation of a process or function may be implemented with a separate software module that allows execution of at least one function or operation. The software code may be implemented by a software application (or program) written in any suitable programming language. The software code may be stored in the memory and executed by the controller.

[0310] The display unit 506 is used to display the information input by the user or the information provided to the user. The display unit 506 may include a display panel, and the display panel may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.

[0311] The user input unit 507 may include, but is not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, a joystick, which will not be elaborated here.

[0312] The interface unit 508 is an interface for connecting an external device to the electronic device 500. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headset port, and so on.

[0313] In addition, the electronic device 500 includes some functional modules not shown here, which will not be elaborated here.

[0314] Those skilled in the art can understand that various aspects of the electronic device provided in this application can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0315] This application also provides a storage medium, in which a program product capable of implementing the active power reduction optimization control method of a wind farm is stored. In some possible implementation manners, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0316] The storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0317] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Thus, the invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An active power shedding optimization control method for a wind farm, characterized in that the steps Including: S1. Construct and train a SAC agent through the SAC algorithm; S2. Obtain the status quantities of the wind farm; S3. Input the status variables of the wind farm into the SAC agent to obtain a vector composed of the initial reference values of the active power of each wind turbine in the wind farm, denoted as ; S4. Apply to the input constraint controller, and the constraint controller performs a secondary calculation on to obtain a vector composed of the optimized reference values of the active power of each wind turbine in the wind farm, denoted as ; Among them, the parameters of the constraint controller include the active power output constraint of the wind turbine, the dispatching instruction constraint of the transmission system operator, and the active power ramp rate constraint of the wind turbine. The solution conditions of the constraint controller are: In the formula, is the optimization objective, indicating the minimization of ; is the lower limit of the theoretical output of the j-th wind turbine, is the upper limit of the theoretical output of the j-th wind turbine; m is the total number of wind turbines; is the active power scheduling instruction for the j-th wind turbine generator; is the initial value of the active power dispatch instruction for the j-th wind turbine generator unit; is the active power scheduling instruction for the j-th fan at the previous moment; It is a dispatching instruction for the transmission system operator; is the absolute value of the power change amount of the j-th wind turbine per unit time; is the maximum change value of the power of the j-th wind turbine per unit time; S5. The wind farm executes , and adjusts the active power of each wind turbine in the wind farm to be the same as ; S6. Repeat steps S2 - S5 to adjust the active power of each wind turbine in the wind farm in real time.

2. The active power shedding optimization control method according to claim 1, characterized in that The status quantities of the wind farm include the grid power instruction, the current input wind speed of the fan, the current output power of the fan, the instantaneous change in the main shaft torque of the fan, and the instantaneous change in the tower thrust.

3. The active power shedding optimization control method according to claim 2, wherein In step S1, the steps of constructing and training the SAC agent include: S101: Construct a policy network, a state value network, a target state value network, an action - state value network, and a target action - state value network; The function representing the policy network is: ; In the formula, is the state variable of the wind farm and is the input value of the policy network; is a vector composed of the active power values of each wind turbine in the wind farm and is the output value of the policy network; The function representing the state-value network is: ; wherein is the state quantity of the wind farm and is the input value of the state value network; The output value of the state value network is the V value; The function representing the target state value network is: ; Wherein, is the state quantity of the wind farm and is the input value of the target state value network; The output value of the target state value network is the target V value; The function representing the action-state value network is as follows: ; where \(s\) is the state variable of the wind farm, is a vector composed of the active power values of each wind turbine in the wind farm, and are both input values of the action-state value network; The output value of the action - state value network is the Q value; The function representing the target action-state value network is as follows: ; In the formula, is the state quantity of the wind farm, is a vector composed of the active power values of each wind turbine in the wind farm, and are both input values of the target action-state value network; The output value of the target action - state value network is the target Q value; The state value network and the target state value network are the same in the initial state, and the action - state value network and the target action - state value network are the same; S102: Set the experience replay buffer and set the hyperparameters; S103: Obtain the wind farm state variables at time t0, denoted as , and input into the policy network to obtain a vector composed of the initial training reference values of the active power of each wind turbine in the wind farm, denoted as ; The input constraint controller performs secondary calculations through the constraint controller to obtain the active power training optimization reference values for each wind turbine in the wind farm, denoted as ; The wind farm executes , and adjusts the active power of each wind turbine in the wind farm to be the same as the same . The end time of the execution is time t1; Obtain the wind farm status quantity, load data at time t1, and load data within a specified time window before time t1. The wind farm status quantity at time t1 is denoted as , and then input the load data at time t1 and the load data within the specified time window before time t1 into the reward function to obtain the reward value for ; Combine , , the load data and reward value at time t1 to form an experience tuple, and store the experience tuple in the experience replay buffer; S104: Randomly extract a batch of experience samples from the experience replay buffer; Using the extracted experience samples, calculate the target Q value using the target action - state value network, and calculate the target V value using the target state value network; Update the state value network according to the calculated target Q value, action entropy, and the output of the current state value network. The loss function formula of the state value network is: wherein, is the weight of the state value network; are the weights of the target action-value network; is the weight of the policy network, is the state quantity of the wind farm at time t0; Denotes an experience replay buffer; A function representing the state value network; is a function representing the target action-state value network; a function representing a policy network; Update the action - state value network according to the calculated target V value, action entropy, and the output of the current action - state value network. The loss function formula of the action - state value network is: In the formula, is the reward value obtained by taking actions in the state; is the weight of the target state value network; are the weights of the actor-critic network; is the discount factor, with a value range between 0 and 1; is a function representing the target state value network; A function representing the action-state value network; Calculate the update gradient of the policy network according to the extracted experience samples and the reinforcement learning algorithm, and update the policy network. The loss function formula of the policy network is: In the formula, is the function with added noise; is a noise vector that conforms to a Gaussian spherical distribution; Update the target action - state value network and the target state value network using the soft update method; S105: Repeat steps S103 - 104 for iterative training of the SAC agent. After the iterative training ends, the required SAC agent is obtained.

4. The active power shedding optimization control method according to claim 3, characterized in that The load data includes the main shaft torque and the tower thrust.

5. The active power shedding optimization control method according to claim 4, characterized in that In step S103, the reward function is constructed by the ICNN - based sliding window equivalent fatigue load surrogate model. The function expression in the ICNN - based sliding window equivalent fatigue load surrogate model is: In the formula, is the input vector, which is the load data of the wind farm at time t1 and the load data within a specified time window before time t1; is the output vector of the layer; is the output vector of the layer; As the activation function, the convex non-decreasing function Smoothed ReLU is used as the activation function ; and is a weight vector related to the layer; is the bias term for the th layer; is a convex function with respect to and is constant after training; ​ is the total number of layers; The reward value calculated using the reward function is the equivalent fatigue load of the wind farm at time t1 and within the specified time window before time t1.

6. The active power shedding optimization control method according to claim 3, characterized in that In step S105, the total number of iterations of the iterative training ≥ 80000 times.

7. A wind farm active power reduction optimization control system for implementing the active power reduction optimization control method of the wind farm according to any one of claims 1-6, characterized in that, Including: A data acquisition module for obtaining the status quantities of the wind farm and the load data; A SAC agent construction and training module for constructing and training a SAC agent; A constraint controller module for performing secondary calculations on the initial actions generated by the SAC agent; An execution module, configured to execute the actions optimized by the constraint controller module and adjust the active power of each wind turbine in the wind farm.

8. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor is configured to implement the steps of the active power reduction optimization control method for the wind farm as described in any one of claims 1-6 when executing the computer program.

9. A storage medium, characterized in that, A computer program is stored on a storage medium. When the computer program is executed by a processor, the steps of the active power reduction optimization control method for the wind farm as described in any one of claims 1-6 are implemented.

Citation Information

Patent Citations

  • SVG parameter optimization identification method based on SAC deep reinforcement learning

    CN115718478A

  • Wind power plant parameter identification method based on multi-agent SAC deep reinforcement learning

    CN116796644A