Active power and frequency coordination control method for wind storage combined system

By employing a three-layer architecture control method for wind-storage integrated systems, combined with Monte Carlo simulation and deep reinforcement learning, the problem of active power and frequency coordination in wind-storage integrated systems was solved, achieving a balance between frequency stability and energy storage lifetime, and improving the robustness and economy of the system.

CN121566491APending Publication Date: 2026-02-24GUONENG LIAONING NEW ENERGY DEVELOPMENT CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511526769.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively coordinate active power and frequency in wind-storage integrated systems. In particular, when faced with uncertainties in power generation and consumption, they cannot achieve a balance between frequency stability and the lifespan of energy storage devices. Furthermore, control strategies are prone to disconnect, making the system unable to adapt to rapid changes.

Method used

The control method adopts a three-layer architecture, including a predictive optimization layer, a real-time coordinated control layer, and an execution layer. It combines Monte Carlo simulation, deep reinforcement learning, and a proportional-integral controller, and uses scenario reduction technology and energy storage lifetime quantification indicators to achieve the handling of wind power and load uncertainties and the real-time management of energy storage status.

Benefits of technology

It enables the wind-storage integrated system to operate collaboratively at different time scales, improves the frequency stability of the system and the lifespan of energy storage equipment, reduces computational complexity and operation and maintenance costs, and enhances the robustness and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121566491A_ABST
    Figure CN121566491A_ABST
Patent Text Reader

Abstract

The invention discloses an active power and frequency coordination control method for a wind storage combined system, and belongs to the technical field of electric power system operation and control. The problems that in the prior art, due to wind power and load uncertainty, the frequency control effect is poor, energy storage life management is insufficient, and long-term and short-term control instructions are disjointed are solved. The problem is solved by constructing a three-layer collaborative architecture of a predictive optimization layer, a real-time coordination control layer and an execution layer: the predictive optimization layer randomly optimizes and generates a reference state track and a frequency modulation capacity instruction of an energy storage unit based on a scene method; the real-time coordination control layer utilizes a deep reinforcement learning coordination controller, synthesizes frequency deviation, SOC compensation signals and frequency modulation capacity reservation instructions, and dynamically distributes active power of a fan and energy storage; and the execution layer realizes rapid issuing of the instruction. The method is mainly used for active power coordination and frequency support of a wind storage combined system in a high-proportion new energy power system, and can effectively delay energy storage life attenuation while stabilizing frequency fluctuation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system operation and control technology, specifically relating to a method for coordinated control of active power and frequency in a wind-storage combined system. Background Technology

[0002] In power systems where wind energy penetration is constantly increasing, wind-storage integrated systems have become an important means of participating in system frequency regulation and improving grid stability. However, existing technologies still face several problems that urgently need to be solved in achieving coordinated control of active power and frequency in wind-storage integrated systems.

[0003] First, while optimization methods based on deterministic forecast data are widely used, their control effectiveness heavily relies on forecast accuracy. Wind power output and load demand inherently exhibit significant randomness and volatility; optimization using a single forecast value cannot adequately represent the uncertainties of future operating scenarios. This leads to discrepancies between the formulated pre-decision plans and actual conditions, resulting in insufficient or excessive reserved integrated frequency regulation capacity when the system faces power disturbances, making it difficult to effectively suppress frequency exceedances. Although stochastic optimization theory provides direction for addressing this problem, efficiently generating and reducing a large number of uncertain scenarios and transforming them into solvable optimization models remains a complex and computationally burdensome challenge.

[0004] Secondly, in the real-time control phase, traditional control strategies such as fixed-coefficient proportional-integral control or simple power distribution logic often struggle to achieve a dynamic optimal balance between rapidly smoothing frequency fluctuations and maintaining the safe and stable operation of the energy storage unit itself. If the control strategy focuses solely on reducing instantaneous frequency deviations, it may neglect the proper recovery of the energy storage unit's state of charge (SOC), causing its SOC to continuously deviate from the normal range, thus rendering it unable to regulate during subsequent critical periods. Simultaneously, frequent and deep charging and discharging operations accelerate the lifespan degradation of energy storage devices, and existing control methods typically lack online quantification and proactive management of this degradation, which is detrimental to the long-term economic operation of the system.

[0005] Furthermore, effectively coordinating long-term commands generated by the system-level optimization layer with the rapid actions of the real-time control layer presents challenges. Relying solely on the reference trajectory output by the optimization layer, without a coordination mechanism capable of dynamic compensation and command redistribution based on real-time operating conditions (such as frequency deviation, deviation between actual state of charge (SOC) and reference values), can easily lead to a disconnect at the control level. The system may exhibit rigid optimization commands, short-sighted real-time actions, and an inability to adapt to rapid changes in source load.

[0006] Finally, the coordinated control process involves multiple conflicting objectives, such as frequency stability and energy storage lifetime, and their dynamic relationships exhibit nonlinear and high-dimensional characteristics. Designing a controller that can simultaneously consider multiple objectives and adapt to the complex dynamic characteristics of the system is quite challenging. Traditional control design methods rely on precise mathematical models and empirical parameter tuning, which is not only cumbersome but also often fails to guarantee control performance and robustness when faced with unmodeled dynamics and changes in operating conditions. Therefore, seeking a coordinated control strategy that can adaptively learn system characteristics and achieve dynamic trade-offs among multiple objectives is one of the main challenges currently faced. Summary of the Invention

[0007] This invention provides a method for coordinated control of active power and frequency in a wind-storage integrated system. It can effectively coordinate the active power output of wind turbines and energy storage units, while addressing the uncertainties in power generation and consumption, and taking into account system frequency stability and the lifespan of energy storage equipment, thereby achieving comprehensive optimization of the operational safety and economy of the wind-storage integrated system.

[0008] To achieve these objectives and other advantages of the present invention, a method for coordinated control of active power and frequency in a wind-storage combined system is provided, comprising the following steps: S1. Predictive Optimization Layer: Executed in a rolling manner during the first cycle; Based on wind power and load forecast data, multiple scenarios representing future uncertainties are generated through Monte Carlo simulation, and a representative set of scenarios with probability weights is obtained by using scenario reduction technology; A stochastic optimization model is established with the goal of minimizing the comprehensive expected cost of system frequency limit exceedance severity and energy storage lifetime loss, and the reference state trajectory of the energy storage unit and the system comprehensive frequency regulation capacity reservation instruction are obtained by solving the model. S2. Real-time Coordination Control Layer: Executed in a second cycle shorter than the first cycle; real-time acquisition of grid frequency deviation, actual active power of wind turbines, and SOC status of energy storage units; conversion of the deviation between the energy storage SOC and the reference state trajectory into an SOC compensation signal; input of the frequency deviation, SOC compensation signal, and comprehensive frequency regulation capacity reservation command to a coordination controller based on deep reinforcement learning training; based on the output of the coordination controller, real-time dynamic allocation of wind turbine active power adjustment and energy storage unit active power command, wherein the reward function of the coordination controller is constructed as a weighted comprehensive function of the absolute value of frequency deviation, frequency change rate, and SOC tracking error; S3, Operation Execution Layer: The active power adjustment of the wind turbine and the active power command of the energy storage unit are respectively sent to the execution mechanism of the wind power generation unit and the energy storage unit.

[0009] Preferably, in step S1, the scene reduction technique employs a synchronous back-substitution reduction algorithm based on Kantorovich distance, specifically including: S110. Initialization: Use the complete set of scenes generated by the Monte Carlo simulation as the current scene set; S120, Scene Distance Calculation: Calculate the Kantorovich distance between every two scenes in the current scene set, where a scene consists of a time series of wind power and load forecast data; S130, Scene Reduction Iteration: In each iteration, calculate the probability density of each scene in the current scene set and the weighted sum of the Kantorovich distance of that scene to all other scenes. Select the scene with the smallest weighted sum as the scene to be reduced, add its probability weight to the nearest retained scene, and delete the scene to be reduced from the current scene set. S140. Termination judgment: Repeat step S130 until the number of scenes in the current scene set reaches the preset representative scene number threshold, forming a representative scene set with probability weights.

[0010] Preferably, in step S1, the objective function of the stochastic optimization model is specifically expressed as: Minimize: E [ α (max (0, |Δf |- Δf lim )) 2 + β (df / dt) 2 ] + λ L ESS , Where E[·] represents the mathematical expectation of all representative scenarios and their probability weights; Δf is the system frequency deviation, Δf lim df / dt is the frequency deviation limit; α and β are the weighting coefficients of the frequency-related terms; L ESS λ is a quantitative indicator of energy storage lifespan loss; λ is the weighting coefficient of the energy storage lifespan loss term. The energy storage life loss quantification index L ESS Calculate using the following steps: S150. Based on the planned charge and discharge power curves of the energy storage unit within the optimization cycle, the Rainflow counting algorithm is used to extract the half-cycles of the charge and discharge cycle and their corresponding cycle depths (DODs). k , where k = 1, 2, ..., K; S160. Based on the experimental lifetime characteristic curve of the energy storage unit, query the DOD for each cycle depth. k The corresponding maximum number of loops N max ; S170. Based on Miner's linear cumulative damage theory, calculate the total lifetime loss L. ESS =Σ(1 / N max,k This summation iterates through all extracted cyclic half-cycles within the optimization cycle.

[0011] Preferably, in step S2, the coordinating controller trained based on deep reinforcement learning is implemented through the following parallel interactive learning architecture, specifically: S210. Construct a central global neural network and multiple parallel learners: Create a central global neural network that includes a commentator module for evaluating state values ​​and an actor module for outputting control policies; simultaneously, instantiate copies of the central global neural network on multiple computing units as parallel interactive learners. S220, Distributed Environment Interaction and Experience Collection: Each of the parallel interactive learners runs in an independent wind-storage joint system simulation environment. Based on its local network parameters, it generates control actions according to the current system state and collects the resulting state transition sequences and corresponding reward signals to form local experience data. S230, Asynchronous gradient calculation and aggregation: Each parallel interactive learner periodically calculates the update gradient of the parameters of its local actor module and critic network module independently based on its local experience data; then, the gradient is asynchronously pushed to the central global neural network; S240, Global Parameter Update and Synchronization: The central global neural network continuously receives gradients from each parallel learner and uses the gradients to asynchronously update its central global neural network parameters; after the update, each parallel interactive learner pulls the latest parameters from the central global neural network to update its own local network. S250, Strategy Optimization Objective: The strategy optimization objective function of the actor module is composed of the state value function output by the critic module and the generalized advantage estimation function calculated based on multi-step temporal difference error, so as to accurately evaluate the long-term benefits of the action while reducing variance.

[0012] Preferably, in step S2, the deviation between the energy storage SOC and the reference state trajectory is converted into a SOC compensation signal, which is achieved through a proportional-integral controller with output limiting. The calculation process of the SOC compensation signal is as follows: S260. Calculate the SOC tracking error: e soc (t) = SOC ref (t) - SOC real (t), where SOC ref (t) represents the value of the energy storage unit reference state trajectory obtained in step S1 at the current time t, SOC real (t) represents the actual SOC state of the energy storage unit collected in real time; S270, Proportional-Integral Control Calculation: The SOC tracking error e is calculated... soc(t) is input to the proportional-integral controller, which outputs a preliminary value u of the compensation signal. comp '(t) is calculated by the following formula: u comp '(t) = Kp×e soc (t) + Ki×∫ e soc (τ) dτ, Wherein, the integral interval is from the start of the current control cycle to the current time t, Kp is the proportional gain coefficient, and Ki is the integral gain coefficient; S280, Output limiting processing: For the initial value u... comp '(t) is subjected to saturation limiting processing to obtain the final SOC compensation signal u. comp (t), which satisfies: u comp (t) = min( U max ,max( U min , u comp '(t) ) ), where, U max with U min These are the preset upper and lower limits for the SOC compensation signal, respectively.

[0013] Preferably, in step S2, the reward function of the coordination controller further includes a penalty term for the rate of change of wind turbine power, so as to limit the frequent adjustment of the active power of the wind turbine and reduce mechanical stress.

[0014] Preferably, between steps S1 and S2, an intermediate correction layer is further included, which is executed in a third cycle; the duration of the third cycle is between the first cycle and the second cycle, and the duration of the third cycle is longer than that of the second cycle but shorter than that of the first cycle; the operation process of the intermediate correction layer specifically includes: S180. Data Acquisition and Error Calculation: Collect historical actual operating data from the previous first cycle starting from the current moment, including actual wind power Pw. act (t) and actual load value Pl act (t), and compared with the wind power prediction value Pw for the corresponding time period in step S1. pre (t) and load forecast value P pre By comparing (t), the wind power prediction error sequence Ew(t) and the load prediction error sequence El(t) are calculated respectively; S190, Future Prediction Curve Correction: Based on the historical prediction error sequence, the wind power prediction curve and load prediction curve for the future period after the current moment are rolled corrected using the exponential smoothing method or Kalman filter algorithm, and an updated prediction curve is generated. S1100, Reference Trajectory and Capacity Command Re-optimization: Using the updated prediction curve as deterministic input, a rolling optimization model is established with the objective of minimizing frequency deviation and energy storage operation costs in the near future. The optimization period of this rolling optimization model covers the next third cycle, and the corrected reference state trajectory (SOC) of the energy storage unit is obtained by solving the model. ref,corrected Integrated frequency modulation capacity reservation instruction P reserve,corrected ; S1110, Instruction Update: Use the revised SOC. ref,corrected With P reserve,corrected Replace the corresponding instruction output in the original step S1 and send it to the real-time coordination and control layer in step S2.

[0015] Preferably, in step S2, the coordinating controller trained based on deep reinforcement learning is trained using a proximal policy optimization algorithm, wherein the policy update of the actor module is limited by pruning the objective function.

[0016] Preferably, in step S160, when querying the experimental lifetime characteristic curve, the operating environment temperature of the energy storage unit is further introduced as a correction factor; the maximum number of cycles N max Based on the actual operating temperature T, the following corrections are made: N max,corrected =N max × exp[-E a / R×(1 / T -1 / T ref ) ], where E a The activation energy characterizes the chemical aging rate of the energy storage unit, where R is the gas constant and T is the activation energy. ref The reference temperature used to obtain the experimental lifetime characteristic curve.

[0017] The present invention has at least the following beneficial effects: First, the method of this invention achieves coordinated operation of the wind-storage integrated system across different time scales by constructing a three-layer architecture of predictive optimization, real-time coordinated control, and execution. The predictive optimization layer handles uncertainty through scenario-based methods, providing the system with a forward-looking and robust reference plan; the real-time coordinated control layer utilizes the powerful decision-making capabilities of deep reinforcement learning to balance frequency regulation and energy storage state management in real time in complex dynamic environments. This layered and progressive structure effectively combines long-term economic optimization with short-term stability control, improving the system's frequency stability in response to source-load fluctuations and enhancing the system's overall lifecycle economics by incorporating energy storage lifetime considerations. The execution layer ensures the accurate implementation of control commands, forming a complete closed loop from decision-making to execution.

[0018] Secondly, this invention employs a synchronous back-substitution reduction algorithm based on Kantorovich distance for scene reduction, which can reduce the number of scenes participating in stochastic optimization while minimizing distortion of the initial scene probability distribution. This directly reduces the size and computational complexity of the optimization model, enabling scene-based stochastic optimization to meet the computational efficiency requirements of online rolling execution. This method ensures that the retained scene set can represent the characteristics of the original uncertainty space to the greatest extent, thereby improving computational efficiency while maintaining the accuracy and robustness of the optimization results as much as possible, providing key technical support for the practical application of predictive optimization layers.

[0019] Third, this invention quantifies the energy storage lifespan loss index L. ESS By explicitly introducing the objective function of the stochastic optimization model, the optimization decision can proactively balance the system's frequency regulation performance with the long-term cost of energy storage equipment. The Rainflow counting algorithm is used to accurately identify charge-discharge cycles, and lifetime loss is calculated by combining experimental lifetime characteristic curves and Miner's cumulative damage theory. This method achieves a scientific and quantitative assessment of energy storage lifetime. Its effect is that it guides the optimization algorithm to automatically avoid harsh charge-discharge strategies that significantly damage energy storage lifetime, favoring more gradual and sustainable power plans. This significantly extends the lifespan of the energy storage system while ensuring system functionality, reducing the overall lifespan maintenance and replacement costs.

[0020] Fourth, this invention utilizes a parallel interactive learning architecture to train a deep reinforcement learning controller, significantly improving training efficiency and learning stability. Multiple parallel learners independently interact with the environment to collect experience, achieving parallel data acquisition and accelerating experience accumulation. Asynchronous gradient calculation and global parameter update mechanisms prevent the local experience of a single learner from dominating policy updates, facilitating the exploration of better policy spaces and suppressing oscillations during training. This architecture effectively utilizes distributed computing resources, accelerating convergence and helping to escape local optima. The resulting coordinated controller policy is of higher quality, more robust, and better suited for handling complex real-time control tasks in wind-storage integrated systems.

[0021] Fifth, this invention employs a proportional-integral (PI) controller with output limiting to convert SOC tracking deviation into a compensation signal, providing a simple, reliable, and computationally efficient method that tightly couples long-term optimization objectives with the real-time control process. PI control ensures the smoothness of the SOC compensation signal, avoiding command jumps, while its output limiting ensures the compensation signal is always confined within a reasonable range, preventing excessive SOC compensation from interfering with the main frequency regulation function. This design enables the real-time coordination controller to dynamically sense and respond to the long-term planned execution status of the energy storage SOC, ensuring that while participating in frequent frequency regulation, the energy state of the energy storage can stably return to the ideal trajectory set by the optimization layer, maintaining its continuous regulation capability.

[0022] Sixth, introducing a penalty term for the rate of change of wind turbine power into the reward function of the coordinated controller can suppress drastic fluctuations in the active power of the wind turbine at the source of control. This design allows the deep reinforcement learning agent to spontaneously choose a power adjustment scheme that is more favorable to the mechanical transmission system of the wind turbine when learning control strategies. Its beneficial effects include significantly reducing the mechanical stress of the wind turbine and reducing fatigue wear on key components such as gearboxes and bearings, thereby improving the operational reliability and service life of the wind turbine. This is an important measure for protecting the wind turbine itself while realizing grid ancillary services, reflecting the comprehensiveness and refinement of system-level control.

[0023] Seventh, this invention introduces an intermediate correction layer, establishing a dynamic feedback loop between the predictive optimization layer and the real-time coordinated control layer. It utilizes the latest actual operating data to perform rolling re-optimization of the prediction curve and optimization instructions, effectively compensating for the performance degradation of optimization layer instructions caused by prediction errors and model mismatch. This design enhances the system's adaptability to actual operating conditions, combining feedforward optimization with feedback correction.

[0024] Eighth, this invention employs a proximal policy optimization algorithm and limits the magnitude of policy updates by pruning the objective function, providing crucial stability assurance for the training of deep reinforcement learning controllers. The pruning mechanism effectively prevents drastic performance degradation or training collapse caused by excessively large single policy update steps, allowing the learning process to evolve smoothly and monotonically towards performance improvement. This robust update characteristic reduces the difficulty of training parameter tuning and improves the reproducibility of training results.

[0025] Ninth, this invention introduces ambient temperature as a correction factor into the energy storage life assessment model, improving the accuracy of energy storage life prediction. The chemical aging rate of a battery is closely related to temperature; ignoring the temperature effect will lead to significant deviations in the life model under different operating environments. By applying temperature correction to the maximum cycle life using the Arrhenius equation, the quantification index L of life loss is improved. ESSIt can more accurately reflect the aging of energy storage under actual operating conditions.

[0026] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description

[0027] Figure 1 This is a flowchart illustrating the active power and frequency coordination control method for a wind-storage combined system according to the present invention. Detailed Implementation

[0028] The present invention will now be described in further detail so that those skilled in the art can implement it based on the description.

[0029] It should be understood that terms such as “having,” “comprising,” and “including” as used herein do not exclude the presence or addition of one or more other elements or combinations thereof.

[0030] like Figure 1As shown, this embodiment of the invention provides a method for coordinated control of active power and frequency in a wind-storage integrated system. The predictive optimization layer is a high-level decision-making module in the control architecture of this embodiment. It executes in a rolling first cycle, which is typically set to a relatively long time interval, such as tens of minutes to several hours. The specific value can be adjusted according to the actual system requirements, such as 30 minutes or 1 hour. The main function of this predictive optimization layer is to generate a large number of scenarios representing future uncertainties based on wind power and load forecast data through Monte Carlo simulation. Monte Carlo simulation is a statistical method based on random sampling. It covers various potential operating states by simulating multiple possible change paths of wind power and load to cope with the randomness and volatility of renewable energy output. Since the number of generated scenarios is large, directly using them for optimization calculations would lead to excessive computational burden. Therefore, scenario reduction techniques are needed to obtain a representative set of scenarios with probability weights. Scenario reduction techniques can use algorithms based on probabilistic distance, such as calculating the Kantorovich distance between scenarios to evaluate similarity and iteratively reducing redundant scenarios, ultimately retaining a small number of representative scenarios while maintaining the characteristics of the original uncertainty distribution. Based on this, a stochastic optimization model is established, aiming to minimize the combined expected cost of system frequency limit exceedance severity and energy storage lifetime loss. Frequency limit exceedance severity is typically quantified by the square of the frequency deviation exceeding the limit, while energy storage lifetime loss is represented by quantitative indicators, such as a cumulative damage model based on charge-discharge cycles. After solving the stochastic optimization model, the model outputs the reference state trajectory of the energy storage unit (e.g., an ideal curve of state of charge (SOC) over time) and the system's comprehensive frequency regulation capacity reservation instruction (i.e., the power capacity reserved by the system to cope with frequency fluctuations). During implementation, the predictive optimization layer periodically acquires the latest prediction data and continuously executes the above-mentioned scenario generation, reduction, and optimization solution steps, thereby dynamically adjusting the long-term operation plan to ensure the system remains robust in the face of uncertainty. By incorporating uncertainty into the optimization framework, this predictive optimization layer ensures that decisions are not based on a single prediction value but consider the expected effects of multiple possible scenarios, thus providing more reliable reference instructions for lower-level control.

[0031] The real-time coordination control layer is the mid-level control module in this embodiment of the invention. It executes in a second cycle shorter than the first cycle, typically set to the second to minute level, for example, 1 to 10 seconds, with the specific value selected according to the system response speed requirements. This real-time coordination control layer collects operational data in real time, including grid frequency deviation, actual active power of the wind turbine, and the SOC status of the energy storage unit. The grid frequency deviation reflects the instantaneous state of the system's power balance, while the energy storage SOC status characterizes its energy level. To combine long-term optimization goals with real-time control, this real-time coordination control layer converts the deviation between the energy storage SOC and the reference state trajectory output by the predictive optimization layer into an SOC compensation signal. This conversion process can be implemented using a proportional-integral controller with output limiting. The proportional-integral operation generates a smooth compensation signal, while the output limiting ensures the signal remains within a preset range, avoiding over-compensation. Subsequently, the frequency deviation, SOC compensation signal, and integrated frequency regulation capacity reservation command are input together to a coordination controller trained using deep reinforcement learning. This coordination controller is implemented using a deep neural network, and its training process involves interaction with the wind-storage integrated system simulation environment to learn the optimal control strategy. The reward function of the coordinating controller is constructed as a weighted composite function of the absolute value of frequency deviation, the rate of frequency change, and the SOC tracking error. This allows the coordinating controller to simultaneously consider frequency stability, frequency change smoothness, and energy storage state tracking accuracy when making decisions. Based on the output of the coordinating controller, the active power adjustment of the wind turbine and the active power command of the energy storage unit are dynamically allocated in real time. During implementation, the real-time coordinating control layer continuously collects data, calculates compensation signals, calls the coordinating controller to generate control commands, and dynamically adjusts the allocation strategy according to the system state at intervals of two cycles. This layer, through the adaptive capabilities of deep reinforcement learning, achieves multi-objective trade-offs in complex operating environments, ensuring rapid frequency recovery while maintaining the sustainable operation of the energy storage unit.

[0032] The execution layer is the bottom-level implementation module of the control architecture in this embodiment of the invention. It is responsible for distributing the active power adjustment of the wind turbine and the active power commands of the energy storage unit generated by the real-time coordination control layer to the execution mechanisms of the wind power generation unit and the energy storage unit, respectively. The execution mechanisms of the wind power generation unit include converters, pitch control systems, etc., while the execution mechanisms of the energy storage unit involve charge and discharge controllers. The execution layer typically operates at a high frequency, such as milliseconds or synchronized with the second cycle, to ensure timely execution of control commands. During implementation, the execution layer receives control commands through a communication network and converts them into operational signals that can be executed by specific devices, such as adjusting the output power of the wind turbine or controlling the charging and discharging power of the energy storage. The key to this layer is to ensure the accurate and rapid distribution of commands and to ensure that the execution mechanisms respond correctly, thereby forming a closed-loop control from decision-making to execution. Through the operation of the execution layer, the wind-storage combined system can actually play a role in active power regulation and frequency support, transforming the upper-level optimization and coordination control into actual system behavior.

[0033] This embodiment achieves coordinated operation of the wind-storage integrated system across different time scales by constructing a three-layer control architecture of predictive optimization, real-time coordinated control, and execution. Compared with existing technologies, this invention effectively overcomes the dependence of optimization methods based on deterministic prediction data on prediction accuracy, and the shortcomings of traditional real-time control strategies in balancing frequency fluctuations and energy storage state management. By introducing uncertainty scenario optimization and deep reinforcement learning coordinated control, this invention improves the frequency stability of the system when facing random fluctuations in source load, and enhances the system's full lifecycle economy by actively managing energy storage lifetime loss. Furthermore, the layered and progressive structure ensures close integration of long-term planning and short-term actions, avoiding control disconnect and improving the overall system's adaptability and robustness. Compared with traditional methods, this invention shows significant improvements in frequency regulation accuracy, energy storage lifetime protection, and system response speed.

[0034] In one specific implementation, during the initialization phase of the scenario reduction technique, a large set of scenarios generated by Monte Carlo simulation is first processed as the current scenario set. Monte Carlo simulation is a statistical method based on random sampling that generates hundreds to thousands of future operating scenarios by simulating various possible paths of change in wind power and load demand. Each scenario consists of a series of time-series data representing predicted values ​​of wind power and load over a specific time period. This complete set of scenarios aims to cover the extensive uncertainties in renewable energy output and load demand, ensuring that all potential operating states are taken into account. The purpose of initialization is to provide a complete starting point for the subsequent reduction process. The number of scenarios can be adjusted according to the actual system scale and computing resources; for example, the initial number of scenarios can be set to 500 or 1000, with the specific value determined based on prediction accuracy and computational efficiency requirements. Through initialization, the scenario reduction algorithm starts with a comprehensive set of scenarios, laying the foundation for subsequent distance calculation and iterative reduction, ensuring that the algorithm can efficiently handle uncertain data.

[0035] In the scene distance calculation phase, it is necessary to evaluate the similarity between every two scenes in the current scene set, which is achieved by calculating the Kantorovich distance. The Kantorovich distance is a probabilistic metric used to quantify the difference between two probability distributions. Each scene is considered a multi-dimensional time series composed of wind power and load forecast data points; therefore, distance calculation involves comparing the overall shape and fluctuation characteristics of these time series. For example, some scenes may exhibit high wind power fluctuations, while others are relatively stable. The core principle of distance calculation is based on the distribution comparison of scene data points. By considering the cumulative differences over the entire time range and weighting them to reflect the influence of probability weights, it captures the structural differences between scenes. In practice, distance calculation is typically solved using numerical optimization methods, requiring the construction of a distance matrix, which has high computational complexity. Therefore, in practical applications, parallel computing or approximation algorithms can be used to accelerate processing, especially when the number of scenes is large. This step provides crucial data for subsequent iterative reduction, ensuring that the algorithm can identify and retain the most representative scenes.

[0036] In the scene reduction iteration and termination judgment phase, the algorithm gradually reduces the number of scenes through multiple iterations until a preset representative scene number threshold is reached. In each iteration, the algorithm first calculates the weighted sum of the probability density of each scene in the current scene set and its Kantorovich distance to all other scenes. This weighted sum reflects the importance of each scene in the overall distribution; scenes with smaller weighted sums are more similar to other scenes and can therefore be reduced without significantly affecting the distribution characteristics. Specifically, the algorithm selects the scene with the smallest weighted sum as the scene to be reduced, adds its probability weight to the nearest retained scene to maintain the overall probability quality, and then deletes the reduced scene from the current scene set. This process is repeated, reducing one scene at a time until the number of scenes is reduced to the target threshold. The representative scene number threshold is a key parameter that balances computational efficiency and model accuracy. For example, the representative scene number threshold can be set to 20 to 50 scenes; the actual value needs to be adjusted according to the optimization model's solution time constraints and the system's uncertainty level. Termination is determined based on the number of scenarios. Once the number of scenarios in the current scenario set reaches or falls below a threshold, iteration stops, ultimately forming a representative scenario set with probability weights. This set retains the core characteristics of the original uncertainty space while reducing the scale of the optimization problem. It should be noted that during scenario reduction, the Kantorovich distance between two scenarios composed of wind power-load time series is calculated by solving their optimal transmission problem. The basic cost function is defined as the sum of the Euclidean distance of power values ​​and the time misalignment penalty term, and can be efficiently approximated using the Sinkhorn iterative algorithm, thus ensuring computational feasibility while preserving the differences in sequence morphology. Specifically, when calculating the Kantorovich distance between two scenarios composed of wind power-load time series, the basic cost function c(i,j) is defined as the sum of the Euclidean distance of power values ​​and the time misalignment penalty term, with the following specific form: For scenario S... a The i-th data point (t) i P a,i ) and Scene S b The j-th data point (t) j P b,j The point-to-point basic cost c(i,j) is calculated as follows: , where P a,i and P b,j Representing scene S respectively a and S b At the corresponding time point t i and t j The power value (including the joint state vector of wind power and load power). It is the square of the Euclidean distance between power values, used to quantify the differences in power state at the same moment; It is a time misalignment penalty term used to quantify two time points t. i and t j The degree of offset between them. This ensures that in optimal transmission matching, points with similar times are matched first, which conforms to the causal and morphological characteristics of time series; This is a preset weighting coefficient used to balance the relative importance of power differences and time misalignment in the total cost. Its value is usually determined through pre-experiments or sensitivity analysis, and can be selected within the range [0.01, 0.1] to ensure that the influence of the time dimension is moderate and does not dominate the matching process. Based on this basic cost function, the two scenarios S... a and S b The Kantorovich distance between them is defined as the distance between S at this cost. a The probability of optimal quality transmission to S b The minimum total cost required. In practical solutions, the entropy-regularized Sinkhorn iterative algorithm can be used for efficient approximate calculations.

[0037] The technical effect achieved by this implementation is that, through a synchronous back-substitution reduction algorithm based on Kantorovich distance, a large number of initial scenarios can be efficiently reduced to a compact set of representative scenarios, while minimizing the distortion of the original probability distribution. This reduces the computational complexity of the stochastic optimization model, enabling scenario-based optimization methods to be executed online in real-world systems without sacrificing the robustness and accuracy of decision-making. The scenario reduction process ensures that the representative scenario set fully captures the key patterns of wind power and load uncertainties, thereby providing reliable input for the predictive optimization layer and improving the frequency stability and economy of the system when facing fluctuations. Furthermore, through the reasonable allocation of probability weights, this method maintains the statistical consistency of the scenario set, helping the optimization model generate reference instructions that better conform to actual operating conditions, ultimately enhancing the overall coordinated control capability of the wind-storage integrated system.

[0038] In one specific implementation, the core objective in constructing the objective function of the stochastic optimization model is to minimize the combined expected cost of system frequency limit violation severity and energy storage lifetime loss. This objective function uses a mathematical expectation operator to perform a weighted average of all representative scenarios and their probability weights, ensuring the robustness of the optimization result under different uncertainty scenarios. Frequency limit violation severity is quantified by the square term of the frequency deviation exceeding the limit, which amplifies the penalty for larger deviations, thus more effectively suppressing frequency limit violation events. Simultaneously, the square term of the frequency change rate is used to smooth frequency fluctuations and improve system stability. Weighting coefficients α and β are used to balance the relative importance of frequency deviation and change rate in the objective function. These coefficients can be adjusted according to system operating requirements; for example, α can be set in the range of 0.5 to 1.5, and β in the range of 0.1 to 0.5. Actual values ​​need to be determined through system simulation. The energy storage lifetime loss term is weighed against the frequency-related term using a weighting coefficient λ. The value of λ may be small, such as 0.05 to 0.2, to reflect the balance between long-term loss and short-term stability. Δf lim The upper limit for system frequency deviation is typically set to 0.2 Hz or determined according to grid standards. Weighting coefficients α, β, and λ need to be determined through sensitivity analysis using a wind-storage integrated system simulation model to balance frequency stability and energy storage lifetime. During implementation, this objective function is embedded in a rolling optimization framework. Each optimization cycle is re-solved based on an updated set of scenarios, outputting the reference state trajectory of the energy storage unit and the system's overall frequency regulation capacity reservation command, providing forward-looking guidance for lower-level control.

[0039] In the process of quantifying energy storage lifetime loss, the Rainflow counting algorithm is first used to extract charge-discharge cycle characteristics based on the planned charge-discharge power curve of the energy storage unit within the optimization cycle. The Rainflow algorithm is a fatigue analysis technique capable of identifying complete charge-discharge cycle half-cycles and their corresponding cycle depths (DODs) from complex power fluctuations. k In practice, the Rainflow algorithm analyzes the time series of the power curve, identifying complete fluctuation segments from trough to peak or peak to trough, and quantifies the depth of each fluctuation segment as a cycle depth value. For example, it can identify multiple cycles of different sizes, such as 0.3 and 0.5. These cycle depth values ​​reflect the stress level experienced by the energy storage unit in actual operation; the greater the depth, the more significant the impact on lifetime. Through this method, seemingly chaotic power commands can be transformed into a series of standardized cyclic events, providing structured input for subsequent lifetime assessment.

[0040] In the lifetime attrition calculation phase, the DOD (Demand of Detail) for each cycle depth is first queried based on the experimental lifetime characteristic curve of the energy storage unit. k The corresponding maximum number of loops N max,kExperimental lifetime characteristic curves are typically provided by battery manufacturers and obtained through accelerated aging experiments. These curves describe the maximum number of cycles a battery can withstand at different cycle depths; for example, a cycle depth of 0.8 might correspond to a maximum of 3000 cycles, while a cycle depth of 0.3 might correspond to 15000 cycles. Then, based on Miner's linear cumulative damage theory, the damage to lifetime from each extracted cycle is quantified as 1 / N. max,k The total lifetime loss L is obtained by summing the damage values ​​over all half-cycles. ESS During implementation, this calculation is tightly coupled with the optimization model. Each candidate power scheme is evaluated for its lifetime loss simultaneously, enabling the optimization algorithm to automatically avoid operational strategies with high damage rates. For example, when generating charge-discharge plans, the optimization solver tends to select power curves with shallower cycle depths and more uniform distributions, even if these schemes are slightly inferior in frequency regulation, they are more economical from a total lifecycle cost perspective.

[0041] The technical effect achieved by this implementation is that, by incorporating precisely quantified energy storage lifespan loss into the optimization objective, an effective balance is achieved between system frequency regulation performance and long-term equipment economics. This method ensures that optimization decisions not only focus on immediate frequency stability but also proactively consider the durability of energy storage equipment, guiding the system to choose an operating strategy more favorable to battery aging. The combination of the Rainflow algorithm and Miner's theory provides a scientifically reliable means of assessing lifespan loss, avoiding the coarse estimation or complete neglect of energy storage lifespan found in traditional methods. This integrated optimization framework can significantly extend the lifespan of the energy storage system while ensuring grid frequency quality, reducing the overall life-cycle operation and maintenance costs, and improving the comprehensive economic benefits of wind-storage combined systems. Simultaneously, the flexible configuration of weighting coefficients allows system operators to adjust the priority of frequency stability and equipment lifespan according to actual needs, enhancing the adaptability of the control strategy.

[0042] In one specific implementation, the coordinating controller based on deep reinforcement learning training is implemented through the following parallel interactive learning architecture. In the implementation of the parallel interactive learning architecture, a central global neural network is first constructed. This central global neural network serves as the core knowledge base and contains two functional modules: a commentator module and an actor module. The commentator module is responsible for evaluating the value of the system state, i.e., predicting the long-term cumulative reward that can be obtained by following the current policy in the current state. Its principle is based on the value function approximation, learning the mapping relationship between the state and the expected reward through a deep learning network. The actor module is responsible for outputting the control policy, i.e., generating specific control actions based on the current system state, such as frequency deviation, SOC compensation signal, and frequency modulation capacity reservation instructions, such as the active power adjustment of wind turbines and the active power instructions of energy storage units. Its principle is based on the policy gradient method, directly optimizing policy parameters to maximize long-term rewards. The central global neural network adopts a fully connected network, containing two hidden layers, each with 128 neurons, using the ReLU activation function; the optimizer uses Adam. This invention provides clear tuning guidelines for key hyperparameters (including learning rate and batch size) in deep reinforcement learning controller training: the learning rate is typically selected in the range of 0.0001-0.01, and a decay strategy can be used to balance convergence speed and stability; the batch size is recommended to be set proportionally (e.g., 1 / 50 to 1 / 100) based on the empirical replay buffer capacity to coordinate gradient estimation accuracy and training efficiency. Through a systematic process (including baseline configuration, sensitivity analysis, and grid search), these parameters are collaboratively optimized to ensure the controller converges efficiently and stably to a high-performance strategy, thereby guaranteeing the effectiveness and reproducibility of the overall control method. In this embodiment, the central global neural network acts as a shared model, and its parameters are continuously optimized during training to coordinate the active power allocation of the wind-storage combined system. Simultaneously, copies of the central global neural network are instantiated on multiple computing units, such as using 4 to 16 CPU cores or GPU instances (the specific number can be flexibly adjusted according to computing resources), and these copies act as parallel interactive learners. Each parallel interactive learner runs in an independent wind-storage joint system simulation environment, which simulates real grid operating conditions, including wind power fluctuations, load changes, and frequency dynamics. Based on local network parameters, each learner generates and executes control actions in real time according to the current system state, and then collects state transition sequences, including new states, actions, and reward signals, forming local experience data. The reward signal is calculated based on a weighted synthesis function of the absolute value of frequency deviation, the rate of frequency change, and the SOC tracking error, and is used to quantify the immediate effect of the actions. Through distributed environment interaction, multiple learners explore different state spaces and action strategies in parallel, accelerating experience accumulation.During implementation, the simulation environment can run at second-level cycles, such as 1 to 10 seconds. Learners store experience data after completing a certain number of interaction steps (e.g., 1000 steps), ensuring data diversity and coverage. This architecture effectively utilizes computing resources through parallel processing, making the learning process more efficient. Furthermore, the independent environment avoids experience correlation, enhancing exploration capabilities. It should be noted that this invention ensures the real-time decision-making capability of the deep reinforcement learning controller by adopting an "offline training, online deployment" mode and designs a degradation control mechanism with three levels of response: when controller performance degrades or outputs abnormally, the system first dynamically adjusts the reward weight for compensation; if this does not improve the situation, it switches to a backup controller based on fixed rules or PID control; finally, it can trigger a device-level protection strategy, forcing energy storage and wind turbines into a safe operating mode, thereby achieving industrial-grade reliability while ensuring intelligent control performance. The degradation control mechanism designed in this invention has clear quantitative triggering conditions: when the average reward value of the coordinating controller is continuously lower than 70% of the benchmark, the root mean square value of the frequency deviation exceeds 0.25Hz, or the key signal is abnormal, the system automatically switches to the backup controller based on fixed rules; if the communication interruption exceeds 3 cycles or the frequency continues to deteriorate to the dangerous limit of 0.5Hz, the equipment-level protection is immediately triggered, forcing the energy storage and wind turbine to switch to the preset safety mode, thereby forming a complete fault response chain from performance compensation to emergency shutdown.

[0043] Following the experience collection phase, each parallel interactive learner periodically and independently calculates the update gradients of the parameters of its local actor and critic network modules based on local experience data. This calculation cycle can be triggered based on the size of the experience buffer or a fixed time interval, such as performing gradient calculations after every 500 to 2000 experience samples collected. However, the specific value needs to be adjusted according to training stability and hardware performance; this is only one possible setting. Gradient calculation uses the backpropagation algorithm, based on loss functions (such as the mean squared error loss for the critic module and the policy gradient loss for the actor module). Its principle is to minimize prediction error or maximize expected reward by optimizing network parameters. The calculated gradients reflect the direction and magnitude of adjustments needed to the local network parameters to adapt to the current environment dynamics. Subsequently, these gradients are asynchronously pushed to the central global neural network. As a preferred implementation, to improve the utilization efficiency of training data and the stability of the algorithm, each of the parallel interactive learners also maintains a local experience replay buffer. The experience replay buffer is a first-in, first-out data storage area used to cache a certain capacity of historical experience data. Each piece of experience data typically includes a state, action, reward, next state, and termination flag. When each parallel interactive learner computes an update gradient, it doesn't just use the most recently collected experience; instead, it randomly samples a mini-batch of experience data from its local experience replay buffer. This random sampling breaks the temporal correlation between experience data, effectively reducing variance during training and improving learning efficiency by reusing historical experience. This means that each learner sends its gradient immediately after computation, without waiting for other learners to synchronize, thus reducing idle time and communication latency. The central global neural network continuously receives gradients from each parallel learner and uses these gradients to asynchronously update its parameters. The update method can employ stochastic gradient descent or its optimized variants (such as the Adam optimizer), and the learning rate can be set in the range of 0.001 to 0.01, with the specific value determined experimentally to balance convergence speed and stability. After the update, each parallel interactive learner pulls the latest parameters from the central global neural network to update its own local network, ensuring that all learners interact based on a unified strategy. During operation, the frequency of gradient pushing and parameter fetching can be coordinated. For example, synchronization can be performed after each gradient calculation, but the interval can also be dynamically adjusted based on network bandwidth and computational load (e.g., every few seconds). This asynchronous mechanism avoids the local experience of a single learner dominating global updates, promotes policy diversity and stability, and accelerates the convergence process through distributed aggregation.

[0044] The policy optimization objective function of the actor module is composed of the state value function output by the critic module and the generalized advantage estimation function calculated based on multi-step temporal difference error. The state value function provides an estimate of the long-term reward of the current state, and its principle is to approximate the true value function through a neural network to evaluate the potential reward of the state. The generalized advantage estimation function is calculated using multi-step temporal difference error, which combines reward information from multiple time steps, such as using rewards from 5 to 10 steps. The choice of the number of steps requires a trade-off between bias and variance, and the actual value may vary dynamically depending on the environment. It measures the superiority or inferiority of a specific action relative to the average action, and its principle is to reduce the variance of the value estimation and improve accuracy. The policy optimization objective of the actor module is to adjust the policy parameters by maximizing the advantage estimation, making it more likely to select actions that bring positive long-term rewards, while the critic module provides reliable state evaluation by minimizing the prediction error of the value function. In practice, the optimization process can use pruning or regularization techniques to limit the policy update magnitude. For example, a proximal policy optimization algorithm can be used to ensure that each update does not deviate significantly from the current policy, thereby maintaining training stability. During operation, the objective function is calculated in batches based on empirical data, for example, processing 64 to 256 samples per batch, with the batch size adjustable according to memory and computing power. This combination enables policy optimization to accurately capture the long-term effects of actions while reducing estimation variance, thereby guiding the coordinating controller to learn intelligent strategies that can both quickly respond to frequency fluctuations and ensure the sustainability of energy storage.

[0045] This implementation improves the efficiency and stability of deep reinforcement learning training through a parallel interactive learning architecture. Multiple parallel learners interact with the environment simultaneously, accelerating the collection of empirical data and enabling the training process to converge to a high-performance policy more quickly. The asynchronous gradient update mechanism reduces oscillations and local optima risks during training, promoting the discovery of the globally optimal policy. The policy optimization objective, combining the state value function and generalized advantage estimation, ensures the accuracy of action evaluation and full consideration of long-term benefits, enabling the coordinated controller to make more intelligent decisions in complex dynamic environments and effectively balance frequency regulation and energy storage state management. Overall, this architecture enhances the adaptability and robustness of the controller, improves the overall coordinated control performance of the wind-storage integrated system, and reduces training time and dependence on a single resource through distributed learning.

[0046] In one specific implementation, during the generation of the SOC compensation signal, it is first necessary to calculate the tracking error of the energy storage SOC. This step is based on the real-time acquisition of the actual SOC state of the energy storage unit (SOC). real (t) and the reference state trajectory (SOC) output by the predictive optimization layer. refThe difference between (t) and (t). The actual SOC is acquired in real time through sensors in the battery management system and is usually expressed as a percentage, reflecting the current energy level of the energy storage unit; the reference state trajectory is the ideal SOC change curve obtained by the upper-level optimization model, which aims to guide the energy storage unit to maintain the optimal energy state during long-term operation. Its principle is to proactively optimize the balance between frequency regulation and energy storage lifetime. SOC tracking error e soc The formula for calculating (t) is SOC. ref (t) minus SOC real (t), where a positive error value indicates that the actual SOC is lower than the reference value, requiring charging compensation; conversely, a negative error value requires discharging compensation. During implementation, this calculation is performed in a rolling second cycle (e.g., a short cycle of 1 to 10 seconds, the specific value can be selected according to system response requirements, such as a 5-second cycle). Real-time data is transmitted to the coordination control layer via a communication network (such as a SCADA system). The error calculation module continuously compares the reference value with the actual value, generating a time-series error signal to provide input for subsequent control calculations. Through this step, the system can dynamically sense the deviation between the energy storage state and the long-term plan, ensuring consistency between real-time control and optimization objectives.

[0047] The calculated SOC tracking error is then input to a proportional-integral (PI) controller with output limiting. This controller is a classic feedback control element, based on a combination of proportional and integral actions: the proportional part (Kp) directly responds to the instantaneous value of the error, providing rapid compensation; the integral part (Ki) accumulates historical errors, eliminating steady-state deviations and ensuring long-term accuracy. The PI controller converts the error e... soc (t) is converted into a preliminary compensation signal u comp The calculation process involves a linear combination of the proportional gain coefficient Kp and the integral gain coefficient Ki, where the integral interval extends from the start of the current control cycle (i.e., the second cycle) to the current time t, and is implemented in discrete form (e.g., using the trapezoidal integration method). The values ​​of parameters Kp and Ki need to be tuned according to the system's dynamic characteristics; possible ranges include Kp between 0.5 and 2.0, and Ki between 0.05 and 0.2, but the actual values ​​need to be determined through simulation or on-site debugging. This is only one possible setting. During implementation, the controller operates at a high frequency (e.g., synchronized with the second cycle), reading the error value in real time and updating the output: the proportional term is directly multiplied by the current error, while the integral term is obtained by accumulating past error values, resulting in the calculated preliminary signal u. comp '(t) reflects the power adjustment required to correct the SOC deviation. This step transforms the error into a continuous control command through smooth calculation, avoiding abrupt changes. At the same time, the integral action ensures that the SOC can gradually return to the reference trajectory, enhancing the system's stability and tracking performance.

[0048] Preliminary compensation signal ucomp '(t) is processed by output limiting to generate the final SOC compensation signal u. comp (t), the limiting function is achieved through saturation processing. Its principle is to limit the signal within a preset upper and lower limit range to prevent overcharging or over-discharging of the energy storage unit due to overcompensation, thereby protecting equipment safety and maintaining system stability. Limiting parameter U max and U min These represent the upper and lower limits of the compensation signal, respectively. These values ​​are typically set based on the power capacity and operating constraints of the energy storage unit, such as U. max It can be set to a positive limit (e.g., 20% to 50% of the rated power corresponding to the charging power), U min The limit value is a negative value (e.g., a similar proportion to the corresponding discharge power), but the specific value needs to be adjusted according to the actual system design. During implementation, the limiting module adjusts the initial value u. comp '(t) performs a min-max operation: first, it is compared with the lower limit U min Compare the values, taking the larger one to avoid the signal being too low, and then compare it with the upper limit U. max The smaller value is selected during the comparison to avoid excessively high signal levels, and the final output u is obtained. comp (t) Ensure that in [U min U max Within the specified range. This processing is performed in real time, updated once per cycle, and tightly coupled with the preceding PI controller, ensuring that the compensation signal remains smooth while being physically constrained. Through this step, the system dynamically adjusts the SOC while avoiding command overshoot, ensuring that the energy storage unit operates within a safe range, and coordinates with other control signals (such as frequency deviation) to achieve multi-objective balance.

[0049] This implementation converts SOC tracking deviation into a compensation signal using a PI controller with output limiting, effectively linking long-term optimization goals with real-time control. This approach improves the accuracy and smoothness of SOC tracking, preventing energy storage state control failure due to error accumulation, while the limiting function effectively prevents equipment risks caused by overcompensation. Overall, this method enhances the coordinated control capability of the wind-storage integrated system, enabling energy storage to maintain a sustainable energy state when participating in frequency regulation, thereby improving the system's operational reliability and economy.

[0050] In one specific implementation, a penalty term for the wind turbine power change rate is added to the reward function design of the coordinating controller. This technical feature aims to limit frequent adjustments in the active power of the wind turbine. The wind turbine power change rate refers to the magnitude of change in the wind turbine's output power per unit time. It is calculated based on the real-time collected active power sequence of the wind turbine, and the rate of power change is obtained through differential operations. In principle, frequent power fluctuations will generate cyclic stress on the mechanical transmission system of the wind turbine generator, especially causing cumulative fatigue damage to key components such as gearboxes and bearings. This penalty term, as a component of the reward function, will generate a corresponding negative incentive in the reward value when the controller's output action causes a rapid change in wind turbine power. In a specific implementation, the calculation of the power change rate can be based on the power difference between adjacent control cycles (e.g., an interval of 1-5 seconds, the specific value can be adjusted according to the wind turbine characteristics; this is only one possible reference choice), and a weighted coefficient is used to incorporate it into the overall reward calculation. This design allows the deep reinforcement learning agent to spontaneously tend to choose control actions with relatively gradual power changes when exploring the optimal strategy.

[0051] During the operation of the real-time coordinated control layer, the implementation of the power change rate penalty term involves multiple stages. First, the system continuously collects the actual active power data of the wind turbine at a frequency of two cycles (usually on the order of seconds) and calculates the power change within each control cycle. This change can be characterized by the absolute difference between the current power value and the power value of the previous cycle, or a more complex sliding window calculation method can be used to capture short-term fluctuation characteristics. Subsequently, the calculated power change rate is multiplied by a preset penalty coefficient, which serves as the negative component of the reward function. The selection of the penalty coefficient needs to strike a balance between suppressing power fluctuations and ensuring frequency regulation effectiveness; the possible value range is between 0.1 and 0.5, but the specific value needs to be determined through debugging based on the actual system operation requirements and wind turbine mechanical characteristics. During the training phase of the coordinated controller, this penalty term, along with other reward components such as frequency deviation and SOC tracking error, influences the optimization direction of the strategy, guiding the agent to learn a control strategy that satisfies both grid frequency regulation requirements and equipment protection.

[0052] The power change rate penalty term does not function independently, but rather works closely with other technical features of the coordinating controller to form a complete multi-objective optimization framework. In the reward function, the absolute value of the frequency deviation term ensures the system's rapid response to grid frequency fluctuations, the frequency change rate term focuses on the smoothness of frequency stability, the SOC tracking error term guarantees the sustainable operation of the energy storage unit, and the newly added power change rate penalty term specifically targets the mechanical protection of the wind turbine equipment. These reward components are weighted and fused using different weighting coefficients to form a comprehensive evaluation index that guides the decision-making process of the deep reinforcement learning agent. The specific form of the reward function is a weighted negative sum of the absolute value of the frequency deviation, the frequency change rate, the SOC tracking error, and the wind turbine power change rate. Each component needs to be normalized before weighting, and the weighting coefficients are determined through sensitivity analysis using system simulation to achieve an optimal balance between frequency regulation performance and equipment lifespan. During implementation, the weighting coefficients can be dynamically adjusted according to operational needs; for example, in conditions where wind turbine fatigue accumulation is severe, the weight of the power change rate penalty term can be appropriately increased. This multi-objective coordination mechanism enables the control strategy to achieve an intelligent trade-off between frequency regulation performance and equipment lifespan, ensuring that grid stability is not sacrificed for excessive protection of the wind turbine, nor that mechanical stress management of the wind turbine is neglected in pursuit of optimal frequency control.

[0053] This implementation effectively limits frequent adjustments to the active power of wind turbines by introducing a penalty term for the rate of change of wind turbine power into the reward function of the coordinating controller. This design reduces the cyclic stress on the mechanical transmission system of the wind turbine generator, reduces fatigue wear on key components, thereby extending the service life of the wind turbine and improving operational reliability. Simultaneously, the synergistic effect of this penalty term with other reward components enables the control system to intelligently balance different operational objectives while ensuring grid frequency stability, avoiding excessive equipment wear caused by solely pursuing frequency performance. Overall, this integrated reward function design enhances the comprehensiveness and precision of the wind-storage integrated system control, improving the quality of grid ancillary services while also protecting the generator equipment itself, providing crucial assurance for the long-term economic and stable operation of the system.

[0054] In one specific implementation, an intermediate correction layer is also included, executed in a third cycle; the duration of the third cycle is between the first and second cycles. During the operation of the intermediate correction layer, data acquisition and error calculation are first performed. This step involves collecting historical actual operating data from the previous first cycle up to the current moment, including actual wind power and actual load values, and comparing them with the predicted values ​​for the corresponding time period in the predictive optimization layer. Actual wind power is acquired in real-time through the wind farm's SCADA system, while actual load values ​​come from measurement data from the power grid dispatch center. These actual values ​​reflect the true operating status of the system. The prediction error sequence is obtained by calculating the difference between the actual and predicted values. The principle is to use historical deviations to reveal the systematic errors and random fluctuations of the prediction model. Error calculation is performed in time series form, for example, one data point every 5 minutes, continuously collecting data for several hours to form an error sequence, but the specific sampling frequency and duration can be adjusted according to system requirements. During implementation, the system establishes an error database to store error information from the most recent first cycles (e.g., 24 hours), providing a data foundation for subsequent corrections. Next, future forecast curve correction is performed. Based on historical forecast error sequences, either exponential smoothing or Kalman filtering algorithms are used to perform rolling corrections on the future forecast curves output by the predictive optimization layer. Exponential smoothing corrects the forecast trend by giving higher weight to recent errors, based on the inertial characteristics of time series. Kalman filtering, on the other hand, dynamically adjusts the forecast values ​​through a state-space model, effectively handling noise and uncertainty. The correction process is executed at intervals of a third period (e.g., 10-30 minutes, with specific values ​​between tens of minutes in the first period and seconds in the second period). Each correction uses the most recent error sequence (e.g., error data from the past 2-4 hours) as input, and the algorithm generates updated wind power and load forecast curves. This dynamic correction mechanism allows the forecast values ​​to adjust according to changes in actual operating conditions, significantly improving forecast accuracy and adaptability to actual operating conditions. To clarify the collaboration between the intermediate correction layer and the prediction layer, their time node connection logic is as follows: the predictive optimization layer is triggered to execute at the top of the hour (T0, T0+30, ...) with a fixed first cycle (e.g., 30 minutes); the intermediate correction layer is initiated at specific times (e.g., T0+10, T0+20, ...) during the execution interval of the prediction layer with a shorter third cycle (e.g., 10 minutes). During each correction, it collects and utilizes historical actual operating data from the current moment back to the beginning of a complete first cycle (i.e., the most recent 30 minutes) to perform rolling correction on the prediction curve and re-optimizes to generate correction instructions covering the next third cycle (i.e., the next 10 minutes). This design ensures that long-term optimization instructions can be dynamically refreshed with high frequency and short cycles based on the latest actual system state, thereby achieving closed-loop feedback and fine-tuning within the large-cycle framework of the prediction layer.

[0055] After obtaining the updated prediction curve, the system enters the reference trajectory and capacity command re-optimization phase. This step uses the updated prediction curve as deterministic input to establish a rolling optimization model for the short-term future (covering the next third cycle). The goal of this rolling optimization model is to minimize the frequency deviation and energy storage operation cost in the short term. The frequency deviation is quantified by the square of the deviation, while the energy storage operation cost considers factors such as power variation and cycle count. The principle of the rolling optimization model is to perform refined recalculation within a shorter time domain, fully utilizing the latest prediction information to correct potential deviations in long-term optimization results. The objective function of the rolling optimization model is min... , among which, T c This is the duration of the third cycle. and P is the weighting coefficient. ESS The energy storage capacity is defined as follows. In practical implementation, the optimization period can be set from 30 minutes to 2 hours. This duration needs to be sufficient to cover the system's inertial response time, but not too long to avoid losing the ability to respond to the latest state. The actual value needs to be determined based on the system characteristics. The solution process uses optimization algorithms such as linear programming or quadratic programming to complete the calculation within a finite time (e.g., within a few minutes), outputting the corrected reference state trajectory of the energy storage unit and the system's comprehensive frequency regulation capacity reservation command. These corrected commands are closer to actual operating requirements, considering both the short-term requirements for frequency stability and the operating constraints of the energy storage equipment. The entire re-optimization process is closely integrated with the prediction and correction stage, forming a complete closed loop from error analysis to command correction.

[0056] Command updates are the final step in the intermediate correction layer and a crucial bridge connecting the predictive optimization layer and the real-time coordinated control layer. In this step, the system replaces the corresponding commands output by the original predictive optimization layer with the corrected reference state trajectory and frequency regulation capacity reserved commands, and immediately sends them to the real-time coordinated control layer. The principle of command updates is to establish a dynamic overlay mechanism, replacing potentially unrealistic earlier plans with optimization results calculated based on the latest data. During implementation, update operations are performed synchronously in a three-cycle rhythm to ensure the timeliness and consistency of commands. When new commands arrive at the real-time coordinated control layer, the coordinated controller immediately incorporates them into its decision-making, adjusting the SOC tracking target and frequency regulation capacity benchmark. Simultaneously, the system retains command version information for tracing and analyzing optimization effects. This update mechanism complements the long-term planning of the predictive optimization layer and the rapid response of the real-time coordinated control layer: the predictive optimization layer provides strategic direction, the intermediate correction layer makes tactical adjustments, and the real-time coordinated control layer is responsible for tactical execution. Command transmission and data exchange between the three layers are conducted through a standard interface, ensuring that the entire control system maintains both long-term optimization foresight and flexibility to respond to unforeseen circumstances. In abnormal situations, such as communication interruption or computation timeout, the system also has instruction caching and degradation strategies to ensure the continuity and reliability of control.

[0057] This implementation introduces an intermediate correction layer, establishing a dynamic feedback correction mechanism between predictive optimization and real-time control, significantly improving the system's adaptability to actual operating conditions. This scheme effectively compensates for the performance degradation of optimization commands caused by prediction errors and model mismatch, enabling the system to maintain superior control performance even when facing rapid changes in source load. Through rolling re-optimization and command updates, the synergy between feedforward optimization and feedback correction is enhanced, improving the accuracy and reliability of the entire wind-storage integrated system. Simultaneously, this hierarchical and progressive control architecture allows for seamless integration of long-term planning and short-term actions, greatly improving the overall robustness and economy of the system, and providing more comprehensive technical support for the wind-storage integrated system to participate in grid frequency regulation services.

[0058] In one specific implementation, the proximal policy optimization algorithm is used as the core training method in the coordinating controller trained by deep reinforcement learning. This is an advanced policy gradient algorithm specifically designed to handle continuous or high-dimensional action space problems in complex environments. The basic principle of the proximal policy optimization algorithm is to gradually optimize control performance by iteratively updating policy parameters, while introducing a special mechanism to ensure the stability of the training process. This algorithm is applied to the actor module and the critic module of the coordinating controller. The actor module is responsible for generating active power commands for wind turbines and energy storage units, while the critic module evaluates the long-term value of these actions. In practice, the training process is carried out in a simulation environment. The system state includes grid frequency deviation, SOC compensation signal, and frequency regulation capacity reservation commands, while the action space corresponds to the active power adjustment of wind turbines and the power commands of energy storage. The training cycle can be set to tens of thousands to hundreds of thousands of iterations, with each iteration including multiple experience collection and parameter update steps. However, the specific number of iterations needs to be determined based on the complexity of the environment and computational resources; this is only one possible reference choice. Through the framework of proximal policy optimization, the coordinating controller can adaptively learn an optimized strategy that balances frequency regulation and energy storage state management under varying operating conditions.

[0059] The policy update of the actor module limits the magnitude of policy changes by pruning the objective function, which is a core feature of proximal policy optimization algorithms. The basic principle of the pruning mechanism is to constrain the probability ratio between the old and new policies during gradient calculation, limiting it to a preset neighborhood range to prevent drastic policy changes in a single update. The pruning operation is controlled by a pruning parameter, which defines the upper and lower bounds of the allowable fluctuation of the ratio, typically set between 0.8 and 1.2, but the specific value needs to be adjusted according to task characteristics and training stability. During implementation, each time the policy is updated, the system calculates the action probability ratio of the old and new policies in the same state and compares it with the pruning boundary: if the ratio exceeds the upper bound, the upper bound value is used for gradient calculation; if it is below the lower bound, the lower bound value is used; otherwise, the actual ratio is used. This pruning mechanism is combined with advantage function estimation, which reflects the superiority or inferiority of a specific action relative to the average level and is calculated through multi-step temporal difference error. The entire update process is performed in batches, with batch size potentially set between 64 and 512 empirical samples, depending on memory capacity and training efficiency requirements. By designing this pruning objective function, the policy update, while pursuing performance improvement, strictly limits the magnitude of change in each iteration.

[0060] Training a coordinated controller based on proximal policy optimization is a systematic iterative process involving cycles of data collection, policy evaluation, and parameter updates. In the initial training phase, the controller primarily explores, randomly trying different power allocation strategies. As training progresses, the proportion of existing knowledge utilized gradually increases. The training process is typically divided into multiple stages, each containing a certain number of environmental interaction steps and parameter update rounds. For example, each stage might collect 4000 steps of experience and then perform 10-15 updates, but these values ​​can be adjusted according to actual conditions. During implementation, the system maintains an experience replay buffer, storing tuples of states, actions, rewards, and transitions to the next state for subsequent batch training. The training frequency can be set according to actual needs, such as large-scale offline training in a simulation environment or continuous fine-tuning online in a real system. After training is complete, the obtained neural network parameters are deployed to the real-time coordinated control layer, executing online decisions at a high frequency in the second cycle. The entire scheme ensures training stability through a pruning mechanism, enabling the controller to gradually learn robust and adaptable control strategies, effectively coordinating the active power allocation of the wind-storage combined system, while simultaneously considering multiple objectives such as frequency stability and energy storage device protection.

[0061] This implementation significantly improves the stability and reliability of deep reinforcement learning controller training by employing a near-end policy optimization algorithm and introducing a pruning mechanism. This scheme effectively prevents drastic fluctuations and crashes in policy performance during training, ensuring smooth convergence to a high-performance policy. The pruning mechanism reduces sensitivity to hyperparameter adjustments by limiting the policy update magnitude, greatly alleviating the burden of training and parameter tuning. The resulting coordinated controller exhibits superior robustness and adaptability, enabling intelligent decision-making in the complex and ever-changing operating environment of wind-storage integrated systems. It effectively balances multiple objectives such as frequency regulation, energy storage state management, and equipment protection, thereby comprehensively improving the control performance and operational economy of the wind-storage integrated system.

[0062] In one specific implementation, the introduction of ambient temperature as a correction factor during energy storage lifetime assessment is based on the scientific principle that battery chemical aging rate is highly correlated with temperature. The actual operating temperature of the energy storage unit is monitored in real time by temperature sensors installed inside the battery compartment, typically in degrees Celsius. Monitoring points may be distributed at different locations within the battery module to obtain representative temperature data. Experimental lifetime characteristic curves are usually obtained by battery manufacturers under standard laboratory conditions. However, in actual operation, energy storage systems may face temperature variations ranging from -10°C to 50°C or even wider. These temperature differences significantly affect the electrochemical reaction rate inside the battery, leading to deviations between the actual aging rate and standard experimental data. The temperature correction mechanism establishes a quantitative relationship between temperature and aging rate, enabling the lifetime assessment model to adapt to different operating environmental conditions. During implementation, the system periodically collects temperature data from the energy storage unit, with a sampling frequency potentially set from once per minute to once every five minutes, depending on the rate of temperature change. This temperature data is stored synchronously with cycle depth data, providing necessary input parameters for subsequent lifetime correction calculations.

[0063] The temperature correction for the maximum cycle life is achieved through the Arrhenius equation, which describes the exponential relationship between chemical reaction rates and temperature and is the fundamental model for battery aging kinetics analysis. In practical applications, the activation energy parameter Ea characterizes the energy barrier of the chemical aging process of the energy storage unit. Its value depends on the battery's chemical system and may be in the range of 50 kJ / mol to 70 kJ / mol for lithium-ion batteries, but the specific value needs to be determined through accelerated aging experiments. The gas constant R is a physical constant, approximately 8.314 J / (mol·K), and is used as a proportionality constant in the calculations. T and T0 ref The absolute temperature used to obtain the experimental lifetime characteristic curve is expressed in Kelvin (K) and the absolute temperature is expressed in T. ref Corresponding to the standard test conditions for obtaining experimental lifetime characteristic curves, the range is typically chosen between 293K and 298K. During implementation, when the system queries the standard maximum number of cycles N corresponding to a certain cycle depth... max,k Subsequently, the current average operating temperature T is acquired synchronously and then substituted into the Arrhenius equation to calculate the correction coefficient. The temperature data may be processed using a moving average, for example, taking the average operating temperature of the most recent 24 hours to eliminate the impact of instantaneous fluctuations, but the specific time window setting needs to balance response speed and stability requirements. This correction calculation is performed during each life assessment to ensure that the output maximum cycle life accurately reflects the battery durability under the current temperature conditions.

[0064] The temperature correction mechanism does not operate independently, but is tightly integrated with the entire energy storage lifetime assessment system to jointly form a more accurate lifetime prediction model. During implementation, the system first extracts the half-cycle and its corresponding cycle depth (DOD) from the planned charge / discharge power curve using the Rainflow counting algorithm. k This step is the same as the uncorrected process. Subsequently, when querying the experimental lifetime characteristic curve, the system will simultaneously acquire the current operating temperature data of the energy storage unit and calculate the maximum number of cycles N after temperature correction based on the Arrhenius equation. max,corrected This correction was subsequently applied to calculations based on Miner's linear cumulative damage theory, where the damage value for each half-cycle became 1 / N. max,corrected Finally, summing these values ​​yields L, a quantitative index of energy storage lifespan loss that better reflects actual operating conditions. ESS The entire calculation process is executed in a rolling fashion, with optimization cycles as the basis, for example, re-evaluating every 30 minutes or 1 hour. The specific frequency can be determined based on computing resources and accuracy requirements. In the stochastic optimization model, this temperature-corrected lifetime degradation index, together with frequency-related terms, constitutes the objective function, guiding the optimization algorithm to generate an energy storage operation strategy that considers both system frequency regulation requirements and actual aging conditions. This integrated design allows the lifetime assessment to consider not only the mechanical stress of charge-discharge cycles but also the impact of ambient temperature on chemical aging, achieving a more comprehensive assessment of the health status of energy storage equipment.

[0065] This implementation significantly improves the accuracy and practicality of the lifespan prediction model by introducing an operating environment temperature correction factor into the energy storage lifespan assessment. This approach more realistically reflects the actual aging of energy storage under different operating environments, avoiding lifespan assessment biases caused by temperature differences. Through temperature adaptive correction, optimization decisions can better balance short-term operational performance and long-term equipment durability, enhancing the economic efficiency of the wind-storage integrated system throughout its entire lifecycle.

[0066] The control method of this invention has broad equipment adaptability. Through parameterized and modular design, it can be flexibly applied to various energy storage technologies such as lithium-ion batteries and flow batteries. In specific implementation, only the experimental lifetime characteristic curve (cycle depth-lifetime relationship) and activation energy parameters of a specific energy storage unit need to be used as preset inputs. The lifetime assessment and optimization module in the system can automatically adapt to the characteristics of this type of energy storage, thereby accurately balancing the frequency support effect and equipment lifespan loss in coordinated control, and giving full play to the advantages of different energy storage technologies.

[0067] The number of devices and processing scale described herein are for the purpose of simplifying the description of the invention. Applications, modifications, and variations of the invention will be readily apparent to those skilled in the art.

[0068] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details.

Claims

1. A method for coordinated control of active power and frequency in a wind-storage combined system, characterized in that, Includes the following steps: S1. Predictive Optimization Layer: Executed in a rolling manner during the first cycle; Based on wind power and load forecast data, multiple scenarios representing future uncertainties are generated through Monte Carlo simulation, and a representative set of scenarios with probability weights is obtained by using scenario reduction technology; A stochastic optimization model is established with the goal of minimizing the comprehensive expected cost of system frequency limit exceedance severity and energy storage lifetime loss, and the reference state trajectory of the energy storage unit and the system comprehensive frequency regulation capacity reservation instruction are obtained by solving the model. S2, Run the real-time coordination control layer: Execute in a second cycle shorter than the first cycle; The system collects real-time data on grid frequency deviation, actual active power of wind turbines, and SOC status of energy storage units; converts the deviation between the energy storage SOC and the reference state trajectory into an SOC compensation signal; inputs the frequency deviation, SOC compensation signal, and comprehensive frequency regulation capacity reservation command to a coordination controller trained using deep reinforcement learning; based on the output of the coordination controller, it dynamically allocates the adjustment amount of active power of wind turbines and the active power command of energy storage units in real time, wherein the reward function of the coordination controller is constructed as a weighted comprehensive function of the absolute value of frequency deviation, frequency change rate, and SOC tracking error; S3, Operation Execution Layer: The active power adjustment of the wind turbine and the active power command of the energy storage unit are respectively sent to the execution mechanism of the wind power generation unit and the energy storage unit.

2. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 1, characterized in that, In step S1, the scene reduction technique employs a synchronous back-substitution reduction algorithm based on Kantorovich distance, specifically including: S110. Initialization: Use the complete set of scenes generated by the Monte Carlo simulation as the current scene set; S120, Scene Distance Calculation: Calculate the Kantorovich distance between every two scenes in the current scene set, where a scene consists of a time series of wind power and load forecast data; S130, Scene Reduction Iteration: In each iteration, calculate the probability density of each scene in the current scene set and the weighted sum of the Kantorovich distance of that scene to all other scenes. Select the scene with the smallest weighted sum as the scene to be reduced, add its probability weight to the nearest retained scene, and delete the scene to be reduced from the current scene set. S140. Termination judgment: Repeat step S130 until the number of scenes in the current scene set reaches the preset representative scene number threshold, forming a representative scene set with probability weights.

3. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 1, characterized in that, In step S1, the objective function of the stochastic optimization model is specifically expressed as: Minimize: E [ α (max(0,|Δf |- Δf lim )) 2 + β (df / dt) 2 ] + λ L ESS , Where E[·] represents the mathematical expectation of all representative scenarios and their probability weights; Δf is the system frequency deviation, Δf lim df / dt is the frequency deviation limit; α and β are the weighting coefficients of the frequency-related terms; L ESS λ is a quantitative indicator of energy storage lifespan loss; λ is the weighting coefficient of the energy storage lifespan loss term. The energy storage life loss quantification index L ESS Calculate using the following steps: S150. Based on the planned charge and discharge power curves of the energy storage unit within the optimization cycle, the Rainflow counting algorithm is used to extract the half-cycles of the charge and discharge cycle and their corresponding cycle depths (DODs). k , where k = 1, 2, ..., K; S160. Based on the experimental lifetime characteristic curve of the energy storage unit, query the DOD for each cycle depth. k The corresponding maximum number of loops N max ; S170. Based on Miner's linear cumulative damage theory, calculate the total lifetime loss L. ESS =Σ(1 / N max,k This summation iterates through all extracted cyclic half-cycles within the optimization cycle.

4. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 1, characterized in that, In step S2, the coordinating controller trained based on deep reinforcement learning is implemented through the following parallel interactive learning architecture, specifically: S210. Construct a central global neural network and multiple parallel learners: Create a central global neural network that includes a commentator module for evaluating state values ​​and an actor module for outputting control policies; simultaneously, instantiate copies of the central global neural network on multiple computing units as parallel interactive learners. S220, Distributed Environment Interaction and Experience Collection: Each of the parallel interactive learners runs in an independent wind-storage joint system simulation environment. Based on its local network parameters, it generates control actions according to the current system state and collects the resulting state transition sequences and corresponding reward signals to form local experience data. S230, Asynchronous gradient calculation and aggregation: Each parallel interactive learner periodically calculates the update gradient of the parameters of its local actor module and critic network module independently based on its local experience data; then, the gradient is asynchronously pushed to the central global neural network; S240, Global Parameter Update and Synchronization: The central global neural network continuously receives gradients from each parallel learner and uses the gradients to asynchronously update its central global neural network parameters; after the update, each parallel interactive learner pulls the latest parameters from the central global neural network to update its own local network. S250, Strategy Optimization Objective: The strategy optimization objective function of the actor module is composed of the state value function output by the critic module and the generalized advantage estimation function calculated based on multi-step temporal difference error, so as to accurately evaluate the long-term benefits of the action while reducing variance.

5. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 1, characterized in that, In step S2, the deviation between the energy storage SOC and the reference state trajectory is converted into a SOC compensation signal, which is achieved through a proportional-integral controller with output limiting. The calculation process of the SOC compensation signal is as follows: S260. Calculate the SOC tracking error: e soc (t) = SOC ref (t) - SOC real (t), where SOC ref (t) represents the value of the energy storage unit reference state trajectory obtained in step S1 at the current time t, SOC real (t) represents the actual SOC state of the energy storage unit collected in real time; S270, Proportional-Integral Control Calculation: The SOC tracking error e is calculated... soc (t) is input to the proportional-integral controller, which outputs a preliminary value u of the compensation signal. comp '(t) is calculated by the following formula: u comp '(t) = Kp×e soc (t) + Ki×∫ e soc (τ)dτ, Wherein, the integral interval is from the start of the current control cycle to the current time t, Kp is the proportional gain coefficient, and Ki is the integral gain coefficient; S280, Output limiting processing: For the initial value u... comp '(t) is subjected to saturation limiting processing to obtain the final SOC compensation signal u. comp (t), which satisfies: u comp (t) = min(U max ,max(U min u comp '(t))), where U max with U min These are the preset upper and lower limits for the SOC compensation signal, respectively.

6. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 1, characterized in that, In step S2, the reward function of the coordination controller also includes a penalty term for the rate of change of wind turbine power, in order to limit the frequent adjustment of the active power of the wind turbine and reduce mechanical stress.

7. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 1, characterized in that, Between steps S1 and S2, an intermediate correction layer is included, which is executed in a third cycle. The duration of the third cycle is between the first and second cycles, and the duration of the third cycle is longer than that of the second cycle but shorter than that of the first cycle. The operation of the intermediate correction layer specifically includes: S180. Data Acquisition and Error Calculation: Collect historical actual operating data from the previous first cycle starting from the current moment, including actual wind power Pw. act (t) and actual load value Pl act (t), and compared with the wind power prediction value Pw for the corresponding time period in step S1. pre (t) and load forecast value P pre By comparing (t), the wind power prediction error sequence Ew(t) and the load prediction error sequence El(t) are calculated respectively; S190, Future Prediction Curve Correction: Based on the historical prediction error sequence, the wind power prediction curve and load prediction curve for the future period after the current moment are rolled corrected using the exponential smoothing method or Kalman filter algorithm, and an updated prediction curve is generated. S1100, Reference Trajectory and Capacity Command Re-optimization: Using the updated prediction curve as deterministic input, a rolling optimization model is established with the objective of minimizing frequency deviation and energy storage operation costs in the near future. The optimization period of this rolling optimization model covers the next third cycle, and the corrected reference state trajectory (SOC) of the energy storage unit is obtained by solving the model. ref,corrected Integrated frequency modulation capacity reservation instruction P reserve,corrected ; S1110, Instruction Update: Use the revised SOC. ref,corrected With P reserve,corrected Replace the corresponding instruction output in the original step S1 and send it to the real-time coordination and control layer in step S2.

8. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 1, characterized in that, In step S2, the coordinating controller based on deep reinforcement learning is trained using a proximal policy optimization algorithm, wherein the policy update of the actor module is limited by pruning the objective function to restrict the policy change range.

9. The active power and frequency coordinated control method for a wind-storage combined system as described in claim 3, characterized in that, In step S160, when querying the experimental lifetime characteristic curve, the operating environment temperature of the energy storage unit is further introduced as a correction factor; the maximum number of cycles N max Based on the actual operating temperature T, the following corrections are made: N max,corrected =N max × exp[-E a / R×(1 / T -1 / T ref ) ], where E a The activation energy is used to characterize the chemical aging rate of the energy storage unit, where R is the gas constant, and T and T0 are given. ref The absolute temperature used to obtain the experimental lifetime characteristic curve is expressed in K.

Citation Information

Cited By

  • Power coordination control method of wind-storage combined power generation system

    CN121749280A

  • Distributed control technology-based wind storage integrated network-related performance PHM detection system and method

    CN122159501A