Wind storage primary frequency modulation optimization method and device based on improved depth deterministic strategy gradient
By using an improved deep deterministic strategy gradient method, the problems of low control accuracy and slow response speed in wind-storage primary frequency regulation technology are solved, and more efficient power system frequency regulation optimization is achieved.
Patent Information
- Application Number
- CN202511751330.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-27
AI Technical Summary
Existing wind-storage primary frequency regulation technology suffers from low regulation accuracy and slow response speed, making it difficult to meet the real-time regulation needs of power systems.
An improved deep deterministic policy gradient method is adopted. By establishing a wind-storage primary frequency regulation strategy, designing a Markov decision process, and using an early stopping mechanism, the deep deterministic policy gradient model is improved and trained to enhance frequency regulation performance.
While ensuring the accuracy of the solution, the overall frequency regulation performance of the system was improved, the frequency regulation accuracy and efficiency of the wind-storage system were enhanced, and the dynamic behavior of the wind-storage system was optimized.
Smart Images

Figure CN121584595A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power system primary frequency regulation optimization control technology, and in particular to a wind-storage primary frequency regulation optimization method and equipment based on an improved deep deterministic strategy gradient. Background Technology
[0002] In recent years, with the rapid development of wind power generation, its proportion in the power system has gradually increased. However, due to the randomness and volatility of wind energy, the output power of wind turbines also exhibits unstable characteristics, posing a challenge to the stable operation of the power system. To solve this problem, primary frequency regulation technology is needed to control the power output of wind turbine generators to balance the supply and demand relationship of the power system.
[0003] Existing wind-storage primary frequency regulation technologies typically rely on pre-set models and thresholds for decision-making. However, these methods suffer from low flexibility, low control accuracy, and slow response speed in practical applications. They also lack robustness and struggle to meet the real-time control requirements of power systems. Summary of the Invention
[0004] To address the problems of low control accuracy and slow response speed in existing technologies, the present invention aims to provide a wind-storage primary frequency regulation optimization method based on an improved deep deterministic strategy gradient, which improves the overall frequency regulation performance of the system while ensuring solution accuracy, thereby enhancing the accuracy and efficiency of the solution.
[0005] To achieve the above objectives, the present invention adopts the following technical solution: a wind-storage primary frequency regulation optimization method based on an improved deep deterministic strategy gradient, the method comprising the following sequential steps:
[0006] (1) Establish a wind-storage primary frequency regulation strategy, which includes the transfer function of wind power participating in primary frequency regulation, the transfer function of energy storage participating in primary frequency regulation, and the primary frequency regulation control strategy;
[0007] (2) Design a wind-storage primary frequency regulation optimization decision problem based on Markov decision process, wherein the wind-storage primary frequency regulation optimization decision problem includes state space Action space and reward function ;
[0008] (3) The early stopping mechanism is used to improve the deep deterministic policy gradient model, resulting in an improved deep deterministic policy gradient model;
[0009] (4) Train the improved deep deterministic policy gradient model to obtain the trained model;
[0010] (5) Solve the trained model using the wind-storage primary frequency regulation strategy and the wind-storage primary frequency regulation optimization decision problem, obtain the optimization results, and analyze the optimization results.
[0011] Step (1) specifically includes the following steps:
[0012] (1a) Establish the transfer function for wind power participation in primary frequency regulation:
[0013]
[0014]
[0015] in, This indicates that wind power participates in rotor inertia control during primary frequency regulation; This represents the wind speed-mechanical power involved in primary frequency regulation. Represents the Laplace operator; and All are the response time constants of wind turbines; Let be the inertial constant of the wind turbine;
[0016] (1b) Establish the transfer function for energy storage participating in primary frequency regulation:
[0017]
[0018] in, This indicates the inertial response of energy storage participating in primary frequency regulation. This represents the response time constant of energy storage;
[0019] (1c) Establish a primary frequency modulation control strategy:
[0020] A three-branch parallel architecture is adopted, consisting of a state path, an action path, and a common fusion path. The state path extracts state features through a signal input layer, a 50-neuron fully connected layer, an activation layer, and a 25-neuron fully connected layer. The action path extracts action features through a signal input layer and a 25-neuron fully connected layer. The state path and action path are fused at the element level through an additive layer via a common fusion path, and then the value of a single action is output through an activation layer and a fully connected layer.
[0021] Step (2) specifically includes the following steps:
[0022] (2a) The state space for:
[0023]
[0024] In the formula, Indicates the current moment; Indicates the frequency deviation at the current moment; The integral representing the frequency deviation at the current moment; This indicates the frequency regulation control signal that the controller is sending to the wind power at the current moment; This indicates the frequency modulation control signal sent by the controller to the energy storage at the current moment;
[0025] (2b) The action space for:
[0026]
[0027] (2c) The reward function for:
[0028]
[0029] In the formula, This represents the early stop identifier, and its value is 0 or 1; As a penalty for early termination; The current frequency is at The penalty coefficient corresponding to the interval; The current frequency is at The penalty coefficient corresponding to the interval; The current frequency is at The penalty coefficient corresponding to the interval.
[0030] Step (3) specifically refers to: selecting validation set indicators that are strongly correlated with the frequency modulation target, with the core indicator being the average frequency deviation at the current moment. At the same time, set indicator thresholds. After each training round, the current action network is run on the validation set and the metrics are calculated. The model parameters corresponding to the best metrics are saved. The system is continuously monitored for 20 to 50 rounds. If the core metrics do not improve, the auxiliary metrics are below the threshold, or the metrics rebound during this period, an early stop is triggered immediately to obtain an improved deep deterministic policy gradient model.
[0031] Step (4) specifically includes the following steps:
[0032] (4a) Initialization phase: Build an improved deep deterministic policy gradient model framework, including action network, evaluation network, initialize target network and experience replay buffer; at the same time, divide the training set and validation set, set early stopping judgment index and threshold, and define early stopping monitoring window;
[0033] (4b) Training loop initialization: Set the upper limit of the total number of training rounds and initialize the storage unit for the optimal model parameters;
[0034] (4c) Single round training process: Under the training set conditions, the state, action, reward and next state data are stored in the experience replay buffer. Data is randomly sampled from the buffer to update the evaluation network and action network, and the target network parameters are updated synchronously.
[0035] (4d) Validation set performance evaluation: After each training round, load the current action or evaluation network, run it under the validation set conditions, calculate the early stopping judgment index, and record the current validation set performance;
[0036] (4e) Early stopping judgment and parameter saving: Compare the current validation set performance with the historical best performance. If the current performance is better, update the optimal model parameters and check whether the early stopping count condition has reached the early stopping threshold. If the condition is met, terminate the training loop immediately; if not, return to step (4c) to continue the next round of training.
[0037] (4f) Final model determination: After training is terminated, load the stored optimal model parameters.
[0038] Another object of the present invention is to provide an electronic device comprising:
[0039] Processor; and
[0040] The memory stores computer program instructions that, when executed by the processor, cause the processor to perform the wind-storage primary frequency regulation optimization method based on the improved deep deterministic policy gradient as described above.
[0041] The present invention also provides a computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the wind-storage primary frequency regulation optimization method based on the improved deep deterministic strategy gradient as described above.
[0042] As can be seen from the above technical solution, the beneficial effects of the present invention are as follows: First, the present invention effectively coordinates the frequency regulation performance of wind-storage primary frequency regulation, improving the overall frequency regulation performance of the system while ensuring the accuracy of the solution; it helps the wind-storage system achieve its frequency regulation goals while better controlling costs and improving economic efficiency. Second, the present invention provides a detailed model of the wind-storage primary frequency regulation optimization problem, which helps to more accurately understand and solve the frequency regulation problem in the wind-storage system. Third, the present invention transforms the wind-storage primary frequency regulation optimization problem into a Markov decision process. By using the Markov decision process, the dynamic behavior of the wind-storage system can be better understood and optimized. Fourth, the present invention uses an improved deep deterministic policy gradient model to solve the Markov decision process. The improved deep deterministic policy gradient model is significantly better than traditional methods in handling complexity and uncertainty, thereby improving the accuracy and efficiency of the solution. Attached Figure Description
[0043] Figure 1 This is a flowchart of the method of the present invention;
[0044] Figure 2 This is a graph showing the frequency deviation change of the wind storage test system after adopting this invention.
[0045] Figure 3 The curve showing the change in wind power frequency regulation in the wind-storage testing system after adopting this invention;
[0046] Figure 4 The curve showing the change in energy storage frequency regulation power in the wind-storage test system after adopting this invention;
[0047] Figure 5 This is a schematic diagram comparing the reward results of the present invention with those of the original deep deterministic policy gradient model. Detailed Implementation
[0048] like Figure 1 As shown, a wind-storage primary frequency regulation optimization method based on an improved deep deterministic policy gradient is proposed, which includes the following sequential steps:
[0049] (1) Establish a wind-storage primary frequency regulation strategy, which includes the transfer function of wind power participating in primary frequency regulation, the transfer function of energy storage participating in primary frequency regulation, and the primary frequency regulation control strategy;
[0050] (2) Design a wind-storage primary frequency regulation optimization decision problem based on Markov decision process, wherein the wind-storage primary frequency regulation optimization decision problem includes state space Action space and reward function ;
[0051] (3) The early stopping mechanism is used to improve the deep deterministic policy gradient model, resulting in an improved deep deterministic policy gradient model;
[0052] (4) Train the improved deep deterministic policy gradient model to obtain the trained model;
[0053] (5) Solve the trained model using the wind-storage primary frequency regulation strategy and the wind-storage primary frequency regulation optimization decision problem, obtain the optimization results, and analyze the optimization results.
[0054] Step (1) specifically includes the following steps:
[0055] (1a) Establish the transfer function for wind power participation in primary frequency regulation:
[0056]
[0057]
[0058] in, This indicates that wind power participates in rotor inertia control during primary frequency regulation; This represents the wind speed-mechanical power involved in primary frequency regulation. Represents the Laplace operator; and All are the response time constants of wind turbines; Let be the inertial constant of the wind turbine;
[0059] (1b) Establish the transfer function for energy storage participating in primary frequency regulation:
[0060]
[0061] in, This indicates the inertial response of energy storage participating in primary frequency regulation. This represents the response time constant of energy storage;
[0062] (1c) Establish a primary frequency modulation control strategy:
[0063] A three-branch parallel architecture is adopted, consisting of a state path, an action path, and a common fusion path. The state path extracts state features through a signal input layer, a 50-neuron fully connected layer, an activation layer, and a 25-neuron fully connected layer. The action path extracts action features through a signal input layer and a 25-neuron fully connected layer. The state path and action path are fused at the element level through an additive layer via a common fusion path, and then the value of a single action is output through an activation layer and a fully connected layer.
[0064] Step (2) specifically includes the following steps:
[0065] (2a) The state space for:
[0066]
[0067] In the formula, Indicates the current moment; Indicates the frequency deviation at the current moment; The integral representing the frequency deviation at the current moment; This indicates the frequency regulation control signal that the controller is sending to the wind power at the current moment; This indicates the frequency modulation control signal sent by the controller to the energy storage at the current moment;
[0068] (2b) The action space for:
[0069]
[0070] (2c) The reward function for:
[0071]
[0072] In the formula, This represents the early stop identifier, and its value is 0 or 1; As a penalty for early termination; The current frequency is at The penalty coefficient corresponding to the interval; The current frequency is at The penalty coefficient corresponding to the interval; The current frequency is at The penalty coefficient corresponding to the interval.
[0073] Step (3) specifically refers to: selecting validation set indicators that are strongly correlated with the frequency modulation target, with the core indicator being the average frequency deviation at the current moment. At the same time, set indicator thresholds. After each training round, the current action network is run on the validation set and the metrics are calculated. The model parameters corresponding to the best metrics are saved. The system is continuously monitored for 20 to 50 rounds. If the core metrics do not improve, the auxiliary metrics are below the threshold, or the metrics rebound during this period, an early stop is triggered immediately to obtain an improved deep deterministic policy gradient model.
[0074] Step (4) specifically includes the following steps:
[0075] (4a) Initialization phase: Build an improved deep deterministic policy gradient model framework, including action network, evaluation network, initialize target network and experience replay buffer; at the same time, divide the training set and validation set, set early stopping judgment index and threshold, and define early stopping monitoring window;
[0076] (4b) Training loop initialization: Set the upper limit of the total number of training rounds and initialize the storage unit for the optimal model parameters;
[0077] (4c) Single round training process: Under the training set conditions, the state, action, reward and next state data are stored in the experience replay buffer. Data is randomly sampled from the buffer to update the evaluation network and action network, and the target network parameters are updated synchronously.
[0078] (4d) Validation set performance evaluation: After each training round, load the current action or evaluation network, run it under the validation set conditions, calculate the early stopping judgment index, and record the current validation set performance;
[0079] (4e) Early stopping judgment and parameter saving: Compare the current validation set performance with the historical best performance. If the current performance is better, update the optimal model parameters and check whether the early stopping count condition has reached the early stopping threshold. If the condition is met, terminate the training loop immediately; if not, return to step (4c) to continue the next round of training.
[0080] (4f) Final model determination: After training is terminated, load the stored optimal model parameters.
[0081] Figure 2 The frequency deviation of the wind storage test system after adopting the present invention is shown; the system was subjected to a frequency disturbance of -0.1Hz at 0 seconds.
[0082] Figure 3 The curves showing the change in wind power frequency regulation in the wind storage test system after adopting the present invention are presented.
[0083] Figure 4 The curves showing the change in energy storage frequency regulation power in the wind storage test system after adopting the present invention are presented.
[0084] Figure 5 To compare the reward results of this invention with those of the original deep deterministic policy gradient model, it can be seen that the initial average reward of this invention is higher and the convergence speed is faster than that of the original deep deterministic policy gradient model.
[0085] In summary, this invention effectively coordinates the frequency regulation performance of wind-storage primary frequency regulation, improving the overall frequency regulation performance of the system while ensuring solution accuracy. It helps the wind-storage system achieve its frequency regulation goals while better controlling costs and improving economic efficiency. This invention provides a detailed model of the wind-storage primary frequency regulation optimization problem, which helps to more accurately understand and solve the frequency regulation problem in the wind-storage system. This invention transforms the wind-storage primary frequency regulation optimization problem into a Markov decision process, which allows for a better understanding and optimization of the dynamic behavior of the wind-storage system. This invention uses an improved deep deterministic policy gradient model to solve the Markov decision process. The improved deep deterministic policy gradient model significantly outperforms traditional methods in handling complexity and uncertainty, thereby improving the accuracy and efficiency of the solution.
[0086] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention. The scope of protection claimed by the appended claims and their equivalents is defined.
Claims
1. A wind-storage primary frequency regulation optimization method based on an improved deep deterministic strategy gradient, characterized in that: The method includes the following steps in sequence: (1) Establish a wind-storage primary frequency regulation strategy, which includes the transfer function of wind power participating in primary frequency regulation, the transfer function of energy storage participating in primary frequency regulation, and the primary frequency regulation control strategy; (2) Design a wind-storage primary frequency regulation optimization decision problem based on Markov decision process, wherein the wind-storage primary frequency regulation optimization decision problem includes state space Action space and reward function ; (3) The early stopping mechanism is used to improve the deep deterministic policy gradient model, resulting in an improved deep deterministic policy gradient model; (4) Train the improved deep deterministic policy gradient model to obtain the trained model; (5) Solve the trained model using the wind-storage primary frequency regulation strategy and the wind-storage primary frequency regulation optimization decision problem, obtain the optimization results, and analyze the optimization results.
2. The wind-storage primary frequency regulation optimization method based on improved deep deterministic strategy gradient as described in claim 1, characterized in that: Step (1) specifically includes the following steps: (1a) Establish the transfer function for wind power participation in primary frequency regulation: ; ; in, This indicates that wind power participates in rotor inertia control during primary frequency regulation; This represents the wind speed-mechanical power involved in primary frequency regulation. Represents the Laplace operator; and All are the response time constants of wind turbines; Let be the inertial constant of the wind turbine; (1b) Establish the transfer function for energy storage participating in primary frequency regulation: ; in, This indicates the inertial response of energy storage participating in primary frequency regulation. This represents the response time constant of energy storage; (1c) Establish a primary frequency modulation control strategy: A three-branch parallel architecture is adopted, consisting of a state path, an action path, and a common fusion path. The state path extracts state features through a signal input layer, a 50-neuron fully connected layer, an activation layer, and a 25-neuron fully connected layer. The action path extracts action features through a signal input layer and a 25-neuron fully connected layer. The state path and action path are fused at the element level through an additive layer via a common fusion path, and then the value of a single action is output through an activation layer and a fully connected layer.
3. The wind-storage primary frequency regulation optimization method based on improved deep deterministic strategy gradient as described in claim 1, characterized in that: Step (2) specifically includes the following steps: (2a) The state space for: ; In the formula, Indicates the current moment; Indicates the frequency deviation at the current moment; The integral representing the frequency deviation at the current moment; This indicates the frequency regulation control signal that the controller is sending to the wind power at the current moment; This indicates the frequency modulation control signal sent by the controller to the energy storage at the current moment; (2b) The action space for: ; (2c) The reward function for: ; In the formula, This represents the early stop identifier, and its value is 0 or 1; As a penalty for early termination; The current frequency is at The penalty coefficient corresponding to the interval; The current frequency is at The penalty coefficient corresponding to the interval; The current frequency is at The penalty coefficient corresponding to the interval.
4. The wind-storage primary frequency regulation optimization method based on improved deep deterministic strategy gradient as described in claim 1, characterized in that: Step (3) specifically refers to: selecting validation set indicators that are strongly correlated with the frequency modulation target, with the core indicator being the average frequency deviation at the current moment. At the same time, set indicator thresholds. After each training round, the current action network is run on the validation set and the metrics are calculated. The model parameters corresponding to the best metrics are saved. The system is continuously monitored for 20 to 50 rounds. If the core metrics do not improve, the auxiliary metrics are below the threshold, or the metrics rebound during this period, an early stop is triggered immediately to obtain an improved deep deterministic policy gradient model.
5. The wind-storage primary frequency regulation optimization method based on improved deep deterministic strategy gradient as described in claim 1, characterized in that: Step (4) specifically includes the following steps: (4a) Initialization phase: Build an improved deep deterministic policy gradient model framework, including action network, evaluation network, initialize target network and experience replay buffer; at the same time, divide the training set and validation set, set early stopping judgment index and threshold, and define early stopping monitoring window; (4b) Training loop initialization: Set the upper limit of the total number of training rounds and initialize the storage unit for the optimal model parameters; (4c) Single round training process: Under the training set conditions, the state, action, reward and next state data are stored in the experience replay buffer. Data is randomly sampled from the buffer to update the evaluation network and action network, and the target network parameters are updated synchronously. (4d) Validation set performance evaluation: After each training round, load the current action or evaluation network, run it under the validation set conditions, calculate the early stopping judgment index, and record the current validation set performance; (4e) Early stopping judgment and parameter saving: Compare the current validation set performance with the historical best performance. If the current performance is better, update the optimal model parameters and check whether the early stopping count condition has reached the early stopping threshold. If the condition is met, terminate the training loop immediately; if not, return to step (4c) to continue the next round of training. (4f) Final model determination: After training is terminated, load the stored optimal model parameters.
6. An electronic device, comprising: processor; as well as A memory storing computer program instructions, which, when executed by the processor, cause the processor to perform the wind-storage primary frequency regulation optimization method based on the improved deep deterministic strategy gradient as described in any one of claims 1-5.
7. A computer-readable storage medium having stored thereon computer program instructions, which, when executed by a processor, cause the processor to perform the wind-storage primary frequency regulation optimization method based on an improved deep deterministic strategy gradient as described in any one of claims 1-5.