A power grid frequency control method, device, equipment and storage medium

CN122532996APending Publication Date: 2026-08-07YUNNAN POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUNNAN POWER GRID CO LTD
Filing Date
2026-05-18
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0004]本申请的主要目的在于提供一种电网频率控制方法、装置、设备及存储介质,旨在解决高比例新能源并网带来的系统惯量低、抗扰动能力弱、传统调频手段不足以及极端故障下黑启动恢复困难的问题,实现电网从常态运行、扰动抗扰到黑启动恢复的全场景频率协调控制

Benefits of technology

[0014]此外,为实现上述目的,本申请还提出一种存储介质,所述存储介质为计算机可读存储介质,所述存储介质上存储有计算机程序,所述计算机程序被处理器执行时实现如上文所述的电网频率控制方法的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122532996A_ABST
    Figure CN122532996A_ABST
Patent Text Reader

Abstract

The application discloses a power grid frequency control method and device, equipment and storage medium, relates to the power grid frequency control technical field, and includes: obtaining an optimal energy storage configuration scheme, wherein the optimal energy storage configuration scheme is obtained by global optimization and solution based on a preset energy storage optimization configuration model; collecting power grid real-time operation data according to the optimal energy storage configuration scheme, extracting power grid operation characteristics, and generating a feedforward power compensation reference value based on the power grid operation characteristics; generating an adaptive optimal control instruction using a preset decision model based on the power grid operation characteristics and the feedforward power compensation reference value; checking and correcting the adaptive optimal control instruction based on a preset full-scenario safety constraint rule, and issuing the corrected executable control instruction to an execution unit to realize power grid frequency coordinated control. The application can realize full-scenario frequency stability support in normal state, disturbance and black start, and significantly improve the inertia level, disturbance resistance and fault recovery speed of high-proportion new energy power grids.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power grid frequency control technology, and in particular to a power grid frequency control method, apparatus, equipment and storage medium. Background Technology

[0002] Driven by the "dual carbon" goals, the installed capacity of new energy sources such as wind power and photovoltaics in the power system continues to increase. The power grid is gradually exhibiting the "dual high" characteristics of high proportion of new energy grid connection and high power electronic equipment. New energy units are connected to the grid through power electronic interfaces, and they almost do not have the inherent rotational inertia and damping characteristics of traditional synchronous generator units, resulting in a significant decrease in the overall inertia level of the system and a significant weakening of its anti-disturbance capability. However, current power grid frequency stability control still mainly relies on the frequency regulation methods of traditional thermal power and hydropower synchronous units. Related energy storage frequency regulation technologies mostly focus on optimizing control strategies for single daily frequency regulation scenarios, while black start technologies are mostly researched for single-site grid construction and voltage regulation control, which can only meet the frequency support needs under local or specific operating conditions.

[0003] Existing technologies cannot cover the frequency coordination control requirements of all scenarios, including normal operation, disturbance rejection, black start of extreme faults and grid recovery. They have problems such as poor adaptability to operating conditions, complex strategy switching, and disconnect between energy storage configuration and operation control. They are unable to cope with the frequency instability risk caused by strong random fluctuations of new energy sources, and cannot provide reliable support for rapid recovery after grid accidents. They are prone to frequency overruns, secondary collapses and even large-scale power outages. Summary of the Invention

[0004] The main purpose of this application is to provide a power grid frequency control method, device, equipment and storage medium, which aims to solve the problems of low system inertia, weak anti-disturbance capability, insufficient traditional frequency regulation methods and difficulty in black start recovery under extreme faults caused by high proportion of new energy grid connection, and realize the full-scenario frequency coordination control of the power grid from normal operation, disturbance immunity to black start recovery.

[0005] To achieve the above objectives, this application proposes a power grid frequency control method, the method comprising: Obtain the optimal energy storage configuration scheme, wherein the optimal energy storage configuration scheme is obtained by global optimization solution based on a preset energy storage optimization configuration model; Real-time grid operation data is collected according to the optimal energy storage configuration scheme, grid operation characteristics are extracted, and feedforward power compensation reference values ​​are generated based on the grid operation characteristics. Based on the power grid operation characteristics and the feedforward power compensation reference value, an adaptive optimal control command is generated using a preset decision model. The adaptive optimal control command is verified and corrected based on the preset full-scenario safety constraint rules, and the corrected executable control command is sent to the execution unit to realize the coordinated control of power grid frequency.

[0006] In one possible implementation, before obtaining the optimal energy storage configuration, the method further includes: Acquire power grid structure data, equipment operating parameters, and historical operating data to determine energy storage performance requirements under normal frequency regulation, disturbance immunity, and black start recovery scenarios, respectively. Obtain relevant power grid operation data and full-scenario operation constraints. Based on the relevant power grid operation data, the full-scenario operation constraints, and the energy storage performance requirements, construct a model framework with the optimization objectives of optimal cost and optimal frequency stability. Embed the physical constraints of energy storage devices, power grid frequency security constraints, and black-start emergency backup capacity constraints to form an energy storage optimization configuration model.

[0007] In one possible implementation, obtaining the optimal energy storage configuration scheme includes: The optimization variables are determined with the goal of achieving the lowest total lifecycle cost of energy storage and the best frequency stability across all scenarios. Based on the optimization variables, an improved particle swarm optimization algorithm is used to perform global optimization on the energy storage optimization configuration model. Iterative optimization is completed under the constraints of the entire scenario operation, and the values ​​of the optimization variables are continuously updated during the iteration process. Based on the iterative update results, select the value combination that satisfies the constraints and has the best optimization effect. Based on the value combination, determine the power configuration, capacity configuration and site layout of energy storage to form the optimal energy storage configuration scheme.

[0008] In one possible implementation, the step of collecting real-time grid operation data according to the optimal energy storage configuration scheme, extracting grid operation characteristics, and generating feedforward power compensation reference values ​​based on the grid operation characteristics includes: Real-time grid operation data is collected according to the energy storage optimal configuration scheme, and grid operation characteristics are extracted from the real-time grid operation data through multi-scale sliding window processing. Based on the aforementioned power grid operation characteristics, a multi-source disturbance joint prediction model is constructed, and the power grid power imbalance is predicted through the multi-source disturbance joint prediction model. Based on the power imbalance in the power grid, the reference value for feedforward power compensation is calculated.

[0009] In one possible implementation, the step of generating adaptive optimal control commands based on the power grid operating characteristics and the feedforward power compensation reference value using a preset decision model includes: The power grid frequency coordination control process is modeled as a discrete-time Markov decision process, and a decision model is constructed that includes a continuous state space, a continuous action space, a multi-objective reward function, a state transition probability matrix, and a discount factor. Based on the long short-term memory network, the long-term time-series dependency features of the power grid operation characteristics are extracted, and the improved dual-delay deep deterministic strategy gradient algorithm is used to complete the training of the decision model. The power grid operating characteristics and the feedforward power compensation reference value are input into the trained decision model to generate adaptive optimal control commands that are suitable for different operating conditions.

[0010] In one possible implementation, the long-range temporal dependency features of the power grid operation characteristics extracted based on the Long Short-Term Memory network are used to train the decision model, and an improved dual-delay deep deterministic policy gradient algorithm is employed, including: The time series of the power grid operation characteristics are processed by a long short-term memory network to extract long-term time-series dependency features that characterize the frequency change and the trend of new energy fluctuations. The long-range temporal dependency features are input into the actor network, and the actor network outputs the control action vector in the continuous action space. The long-range temporal dependency features and the control action vector are input into a dual critic network to obtain two sets of value evaluation results, and the smaller value is selected as the target value evaluation result. Based on the target value assessment results, the network parameters of the actor network and the dual critic network are updated at intervals according to the delayed update strategy, and an experience replay buffer is constructed to store samples of the agent's interaction with the power grid environment. The sample priority is determined based on the time-series differential error, and samples are extracted according to the priority to complete the iterative optimization of the decision model.

[0011] In one possible implementation, the adaptive optimal control command is verified and corrected based on preset full-scenario safety constraint rules, and the corrected executable control command is sent to the execution unit to achieve grid frequency coordinated control, including: The adaptive optimal control command is weighted and fused with the feedforward power compensation reference value to obtain the control command to be verified. Based on the full-scenario security constraint rules, the control command to be verified is subjected to amplitude limiting verification and command correction to generate an executable control command. The executable control commands are sent to the execution unit and the power grid operating status is fed back to achieve closed-loop coordinated control of the power grid frequency.

[0012] Furthermore, to achieve the above objectives, this application also proposes a power grid frequency control device, which includes: The acquisition module is used to acquire the optimal energy storage configuration scheme, wherein the optimal energy storage configuration scheme is obtained by global optimization based on a preset energy storage optimization configuration model; The extraction module is used to collect real-time power grid operation data according to the optimal energy storage configuration scheme, extract power grid operation characteristics, and generate feedforward power compensation reference values ​​based on the power grid operation characteristics. The generation module is used to generate adaptive optimal control commands based on the power grid operation characteristics and the feedforward power compensation reference value using a preset decision model. The control module is used to verify and correct the adaptive optimal control command based on preset full-scenario safety constraint rules, and send the corrected executable control command to the execution unit to realize grid frequency coordinated control.

[0013] In addition, to achieve the above objectives, this application also proposes a power grid frequency control device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the power grid frequency control method as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the power grid frequency control method described above.

[0015] In addition, to achieve the above objectives, this application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the power grid frequency control method described above.

[0016] This application provides a power grid frequency control method, apparatus, device, and storage medium. The power grid frequency control method obtains an optimal energy storage configuration scheme, which is obtained by global optimization based on a preset energy storage optimization configuration model. Then, real-time power grid operation data is collected according to the optimal energy storage configuration scheme, power grid operation characteristics are extracted, and a feedforward power compensation reference value is generated based on the power grid operation characteristics. Based on the power grid operation characteristics and the feedforward power compensation reference value, an adaptive optimal control command is generated using a preset decision model. The adaptive optimal control command is then verified and corrected based on preset full-scenario safety constraint rules, and the corrected executable control command is sent to the execution unit to achieve coordinated control of power grid frequency. Thus, through global optimization configuration of energy storage, multi-source disturbance feedforward compensation, and adaptive decision-making collaborative control, stable frequency support is achieved in all scenarios including normal, disturbance, and black start, significantly improving the inertia level, anti-disturbance capability, and fault recovery speed of high-proportion renewable energy power grids. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an embodiment of the power grid frequency control method of this application. Figure 2 A comparison diagram showing the implementation effects of the power grid frequency control method of this application in a normal frequency regulation scenario and a traditional PID control scenario; Figure 3 A comparison of the implementation effects of the power grid frequency control method of this application under disturbance immunity scenarios with those of traditional PID control. Figure 4 A comparison of the implementation effects of the power grid frequency control method of this application in black start and power grid recovery scenarios with traditional PID control. Figure 5 This is a schematic diagram of the equipment structure of the hardware operating environment involved in the power grid frequency control method in the embodiments of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0023] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, or mobile phone, or an electronic device, big data service platform, or power grid frequency control system capable of realizing the above functions. The following description uses a power grid frequency control system as an example to illustrate this embodiment and the subsequent embodiments.

[0024] Based on this, the embodiments of this application provide a power grid frequency control method, referring to... Figure 1 ,Figure 1 This is a flowchart illustrating an embodiment of the power grid frequency control method of this application.

[0025] In this embodiment, the power grid frequency control method includes steps S11 to S14: Step S11: Obtain the optimal energy storage configuration scheme, wherein the optimal energy storage configuration scheme is obtained by global optimization based on a preset energy storage optimization configuration model; It should be noted that the optimal energy storage configuration scheme refers to a comprehensive scheme that simultaneously meets the requirements of normal frequency regulation, disturbance immunity, and black start recovery, while taking into account the full life cycle cost of energy storage and frequency stability. The pre-built energy storage optimization configuration model refers to a mathematical model that is pre-constructed and coupled with full-scenario operational constraints and a dual-objective optimization orientation. Global optimization solution refers to the computational process of traversing and searching the entire feasible domain using intelligent optimization algorithms to determine the optimal solution. The system uses a pre-built model rather than online real-time model construction, which significantly reduces online computational overhead, ensures the efficiency and reliability of scheme generation, provides hardware configuration basis and safe operating boundaries for subsequent real-time frequency control, and achieves synergy between energy storage planning and operation control.

[0026] Furthermore, the global optimization process uses energy storage power and capacity as optimization variables, embedding equipment constraints, grid security constraints, and black-start backup constraints to ensure the solution is feasible and executable. In one possible implementation, the system employs an improved particle swarm optimization algorithm to perform the global optimization, dynamically updating the values ​​of optimization variables during the iteration process until the objective function converges.

[0027] Specifically, the system directly calls the preset energy storage optimization configuration model that has been built and verified offline. With the lowest cost of energy storage throughout its entire life cycle and the best frequency stability effect in all scenarios as the optimization direction, it completes iterative optimization under all scenario constraints, selects the value combination that meets the constraints and has the best optimization effect, and then determines the power configuration, capacity configuration and site layout of energy storage to form the optimal energy storage configuration scheme.

[0028] For example, for regional power grids with a high proportion of new energy access, the system uses an improved particle swarm optimization algorithm to perform global optimization based on a pre-built full-scenario coupled energy storage optimization configuration model. Finally, it determines a combination scheme of configuring centralized energy storage at key hub nodes and distributed energy storage at each new energy power station, forming an optimal energy storage configuration scheme that simultaneously meets the needs of normal frequency regulation, disturbance support and black start emergency backup.

[0029] Step S12: Collect real-time grid operation data according to the optimal energy storage configuration scheme, extract grid operation characteristics, and generate feedforward power compensation reference values ​​based on the grid operation characteristics; It should be noted that real-time grid operation data refers to the real-time frequency and bus voltage at the grid's point of common coupling, real-time output data of each new energy power plant, real-time state of charge and charging / discharging output data of each energy storage system, real-time load data of the entire grid, grid connectivity status, and line power flow data acquired through data acquisition equipment. Grid operation characteristics refer to the frequency dynamics, disturbance change trends, system inertia level, and support capacity characteristics extracted from the raw data. Feedforward power compensation reference values ​​refer to reference instructions calculated based on disturbance prediction results and used to compensate for grid power imbalances in advance. The system determines the acquisition range and feature dimensions based on the optimal energy storage configuration scheme, ensuring a high degree of matching between data and control requirements. Through multi-scale processing and disturbance prediction, it achieves advance adjustment, effectively improving the response speed and anti-disturbance capability of frequency control.

[0030] Furthermore, the power grid operation characteristics are extracted using a multi-scale sliding window, which can simultaneously capture instantaneous frequency fluctuations and continuous changing trends, improving the comprehensiveness of feature representation. In one possible implementation, the system sets a short-time window to capture instantaneous frequency fluctuation characteristics and a long-time window to identify frequency changing trend characteristics, simultaneously extracting features such as new energy output fluctuations, load changes, system inertia, and adjustable reserve capacity.

[0031] Specifically, the system determines the data acquisition range and dimensions according to the optimal energy storage configuration scheme, collects grid operation data in real time, extracts grid operation characteristics through multi-scale sliding window processing, constructs a multi-source disturbance joint prediction model based on grid operation characteristics, predicts grid power imbalance by the multi-source disturbance joint prediction model, and calculates and generates feedforward power compensation reference values ​​based on grid power imbalance.

[0032] For example, the system defines the acquisition range according to the optimal energy storage configuration scheme, collects real-time operating data through the SCADA system and synchronous phasor measurement device, and completes multi-scale processing using a 200ms short window and a 2s long window to extract grid operation characteristics such as frequency deviation, frequency change rate, new energy output fluctuation, load change, and system inertia level. The characteristics are input into a multi-source disturbance joint prediction model based on a long short-term memory network to obtain the grid power imbalance at future moments, and calculate and generate feedforward power compensation reference values ​​accordingly.

[0033] Step S13: Based on the power grid operation characteristics and the feedforward power compensation reference value, generate adaptive optimal control commands using a preset decision model; It should be noted that the preset decision model refers to an intelligent decision model that has been pre-modeled and trained, and constructed based on an improved dual-delay deep deterministic strategy gradient algorithm that integrates long short-term memory networks. The adaptive optimal control command refers to the energy storage charging and discharging power command, new energy virtual inertia parameters, and primary frequency regulation droop parameters that can adapt to all operating conditions including normal, disturbance, and black start, and achieve optimal frequency regulation. The system inputs the grid operating characteristics and feedforward power compensation reference values ​​into the preset decision model, and performs inference based on time-series trend characteristics. It can output a control strategy with strong foresight and excellent adaptability, solving the problems of lag and poor robustness of traditional control methods.

[0034] Furthermore, the preset decision model employs an offline training and online invocation mode to ensure the real-time performance and stability of command output. In one possible implementation, the system models the power grid frequency coordination control process as a discrete-time Markov decision process, constructing a continuous state space, a continuous action space, a multi-objective reward function, a state transition probability matrix, and a discount factor to complete the framework of the decision model.

[0035] Specifically, the system models the power grid frequency coordination control process as a discrete-time Markov decision-making process, constructs a decision model that includes a continuous state space, a continuous action space, a multi-objective reward function, a state transition probability matrix, and a discount factor, extracts long-range time-series dependency features of power grid operation characteristics based on a long short-term memory network, and completes the training of the decision model using an improved dual-delay deep deterministic strategy gradient algorithm. The power grid operation characteristics and feedforward power compensation reference values ​​are input into the trained decision model to generate adaptive optimal control commands that are adapted to different operating conditions.

[0036] For example, the system inputs the extracted grid operation characteristics and feedforward power compensation reference values ​​into a pre-trained improved TD3 decision model that integrates a long short-term memory network, so that the improved TD3 decision model can complete forward inference calculation based on long-term time-series dependency characteristics, and output energy storage charging and discharging power commands, new energy power station virtual inertia control parameters and primary frequency regulation droop control parameters adapted to the current operating conditions, forming an adaptive optimal control command.

[0037] Step S14: Based on the preset full-scenario safety constraint rules, the adaptive optimal control command is verified and corrected, and the corrected executable control command is sent to the execution unit to realize grid frequency coordinated control.

[0038] It should be noted that the preset full-scenario safety constraint rules refer to a pre-defined set of constraints covering the physical constraints of energy storage devices, the safety operation constraints of the power grid, and the operating limits of new energy devices; executable control commands refer to control commands that have been safety-limited and corrected, meet the safety boundaries of the equipment and the power grid, and can be directly issued and executed; execution units refer to field execution devices with frequency regulation capabilities, such as energy storage systems, new energy power plants, and controllable loads. The system eliminates out-of-limit commands through safety verification, which can avoid equipment damage and power grid operation risks. Through command issuance and status feedback, a closed-loop control is formed, continuously improving the accuracy and stability of frequency control.

[0039] Furthermore, the verification and correction process involves a weighted fusion of control commands and feedforward compensation values ​​to ensure smooth and shock-free control actions. In one possible implementation, the system weightedly fuses the adaptive optimal control command with the feedforward power compensation reference value to obtain the control command to be verified, and then performs amplitude limiting verification and correction component by component based on full-scenario safety constraint rules.

[0040] Specifically, the system weightedly fuses the adaptive optimal control command with the feedforward power compensation reference value to obtain the control command to be verified. Based on the full-scenario safety constraint rules, the system performs amplitude limiting verification and command correction on the control command to be verified, generates an executable control command that meets the safety requirements, sends the executable control command to the execution unit, and receives the grid operation status feedback from the execution unit to realize grid frequency closed-loop coordinated control.

[0041] For example, the system weights and fuses the adaptive optimal control command and the feedforward power compensation reference value according to preset weights. It then performs verification and amplitude correction based on full-scenario safety constraints such as energy storage charging and discharging power limits, state of charge safety range, grid frequency deviation limits, and voltage deviation limits. This generates executable control commands, which are then sent to the execution units of each energy storage power station and new energy power station. At the same time, the real-time grid operation status after execution is transmitted back to the status perception module, thus completing closed-loop frequency coordination control.

[0042] This embodiment quickly obtains the optimal energy storage configuration scheme through a preset energy storage optimization configuration model, thereby providing hardware support and safety boundaries. It then generates feedforward compensation to achieve advanced adjustment and realizes full-condition adaptive optimal control through a preset decision model. The control process is ensured to be safe and stable through full-scenario safety verification and closed-loop execution. Finally, it realizes frequency coordination control of a high proportion of new energy power grid in all scenarios of normal operation, disturbance rejection and black start recovery, significantly improving the system's inertia support capability, disturbance rejection capability and fault rapid recovery capability.

[0043] In one feasible implementation, prior to obtaining the optimal energy storage configuration, the method further includes: Step S21: Obtain grid structure data, equipment operating parameters and historical operating data, and determine the energy storage performance requirements under normal frequency regulation scenario, disturbance immunity scenario and black start recovery scenario respectively; It should be noted that: power grid structure data refers to the basic data characterizing the connection relationship of power grid lines, node distribution, and topology; equipment operating parameters refer to the rated parameters, operating limits, and regulation characteristics of energy storage systems, new energy power plants, and synchronous generator units; historical operating data refers to the recorded data of frequency fluctuations, power disturbances, fault recovery, and load and new energy output changes of the power grid over a period of time; normal frequency regulation scenario refers to the daily frequency regulation scenario in which the power grid is in a stable operating state and needs to smooth out random fluctuations in new energy and load; disturbance immunity scenario refers to the emergency support scenario in which the power grid experiences power disturbances such as sudden changes in new energy output, large load switching, and local generator disconnection; black start recovery scenario refers to the extreme emergency scenario in which power supply is gradually restored from a completely dark state after the power grid is disconnected from the grid and a large-scale power outage occurs; energy storage performance requirements refer to the index requirements for energy storage power regulation rate, response time, peak power, support capacity, and emergency backup capacity under different scenarios.

[0044] Furthermore, the system comprehensively analyzes the grid's operating characteristics and boundary conditions through multi-source data, enabling precise matching of energy storage capacity requirements for different scenarios and avoiding insufficient or excessive energy storage configuration. In addition, the determination of energy storage performance requirements covers the entire lifecycle of operating conditions, ensuring that the subsequent configuration scheme simultaneously meets both economic and control performance requirements.

[0045] Specifically, the system collects and analyzes power grid structure data, operating parameters of various equipment, and historical operating data. Based on the power grid operation patterns and fault characteristics reflected in the data, it clarifies the performance requirements of energy storage for regulation rate, response time, peak power, support capacity, and emergency backup capacity in normal frequency regulation scenarios, disturbance immunity scenarios, and black start recovery scenarios.

[0046] For example, for normal frequency regulation scenarios, the power regulation rate, millisecond response time, and daily cycle number requirements of energy storage are determined, as well as the continuous regulation capacity requirements needed to smooth out random fluctuations in new energy sources and loads. For disturbance and anti-disturbance scenarios, the peak power and support capacity requirements of energy storage are determined when the grid experiences sudden changes in new energy output, large load surges and drops, or small-scale unit disconnection, to meet the maximum power gap of the system and the duration of disturbances. For black start scenarios, the power requirements for starting new energy power plants, the power impact compensation requirements for the entire process of grid restoration and load switching, and the minimum reserve capacity requirements for continuous emergency support are determined under extreme blackout or grid disconnection accidents, forming a complete set of energy storage performance requirements for all scenarios.

[0047] Step S22: Obtain relevant power grid operation data and full-scenario operation constraints. Based on the relevant power grid operation data, the full-scenario operation constraints, and the energy storage performance requirements, construct a model framework with the optimization objectives of optimal cost and optimal frequency stability. Embed the physical constraints of energy storage devices, power grid frequency security constraints, and black-start emergency backup capacity constraints to form an energy storage optimization configuration model.

[0048] It should be noted that the grid-related operational data refers to operational status data such as grid operation mode, load level, renewable energy output characteristics, and system inertia level; the full-scenario operational constraints refer to the operational restrictions that the grid and equipment must comply with under all operating conditions including normal, disturbance, and black start; cost optimization refers to minimizing the total life-cycle cost of energy storage, including initial investment, operation and maintenance, and equipment replacement costs; optimal frequency stability refers to minimizing grid frequency deviation, fluctuation, and maximizing stability across all scenarios; the model framework refers to the basic structure of energy storage optimization configuration that includes optimization objectives, decision variables, and constraints; physical constraints of energy storage devices refer to the power, capacity, state of charge, and charge / discharge rate limitations determined by the characteristics of energy storage itself; grid frequency security constraints refer to the constraints that ensure grid frequency operation within safe limits; black start emergency backup capacity constraints refer to the minimum energy storage capacity constraints that must be retained to ensure grid recovery after extreme accidents; and the energy storage optimization configuration model refers to the dual-objective optimization mathematical model used to solve for the optimal energy storage power, capacity, and layout scheme.

[0049] Furthermore, the system constructs an energy storage optimization configuration model guided by dual objectives and bounded by multi-dimensional constraints. This enables synergistic optimization of economic efficiency and control performance, ensuring that the configuration scheme output by the energy storage optimization configuration model is feasible, executable, and covers all scenarios. In addition, the model construction process of the energy storage optimization configuration model organically integrates requirements, objectives, and constraints, avoiding the limitations of single-objective optimization and improving the comprehensive applicability of the solution.

[0050] In one possible implementation, the system takes the lowest life-cycle cost of energy storage and the minimum cumulative frequency stability deviation across all scenarios as dual optimization objectives, builds a model framework, and embeds equipment constraints, grid security constraints, and black-start backup constraints in sequence to finally form a solvable energy storage optimization configuration model.

[0051] Specifically, the system acquires relevant power grid operation data and sorts out the operation constraints of the entire scenario. Based on the relevant power grid operation data, the operation constraints of the entire scenario, and the energy storage performance requirements of each scenario, a model framework is built with the optimization objectives of optimal cost and optimal frequency stability. The physical constraints of energy storage devices, power grid frequency security constraints, and black start emergency backup capacity constraints are embedded in the model framework, and finally a complete full-scenario coupled energy storage optimization configuration model is formed.

[0052] In one embodiment, a dual-objective optimization objective function is constructed, comprehensively considering the economic efficiency of energy storage and the frequency stability performance across all scenarios. The expression is:

[0053]

[0054]

[0055] In the formula: Cost of energy storage throughout its entire lifecycle; This represents the cumulative frequency stability deviation across all scenarios. , These are the weighting coefficients for the cost objective and the stability objective, respectively, satisfying... ; n represents the number of energy storage sites; n represents the total lifespan of the energy storage project. The timeline for replacing energy storage devices; , These are the investment cost per unit power of energy storage and the investment cost per unit capacity of energy storage, respectively. For the first The rated power of each energy storage site; For the first The rated capacity of each energy storage site; The annual operation and maintenance cost of energy storage is positively correlated with the installed capacity of energy storage. Number of typical scenarios; Let be the deviation between the real-time frequency and the rated frequency of the power grid in the k-th typical scenario. is the discount rate.

[0056] Furthermore, embedding full-scenario safety constraints ensures that the energy storage configuration solution meets the safety boundaries for operation under all conditions. The constraints are as follows:

[0057] In the formula: , The first Upper and lower limits of the rated power of an energy storage site; , The first The upper and lower limits of the rated capacity of each energy storage site; This refers to the safety limit for power grid frequency deviation; , The first The black start emergency backup capacity requirement and daily frequency regulation capacity requirement of each energy storage site; , These represent the upper and lower safety limits of the energy storage state of charge, thus forming a fully coupled energy storage optimization configuration model suitable for high-proportion new energy power grids.

[0058] This embodiment accurately determines the energy storage performance requirements for all scenarios by comprehensively collecting basic and historical data of the power grid. Guided by dual objectives and bounded by multiple constraints, it constructs a coupled energy storage optimization configuration model for all scenarios, thereby achieving coordinated optimization of economic efficiency and frequency stability in the energy storage planning stage.

[0059] In one feasible implementation, obtaining the optimal energy storage configuration scheme includes: Step S31: Determine the optimization variables with the goal of minimizing the total lifecycle cost of energy storage and achieving optimal frequency stability across all scenarios. It should be noted that the total lifecycle cost of energy storage refers to the total cost of energy storage from initial investment, operation and maintenance to equipment replacement; optimal frequency stability across all scenarios refers to the minimum frequency deviation, minimum fluctuation, and highest stability level of the power grid under all operating conditions, including normal frequency regulation, disturbance immunity, and black start recovery; optimization direction refers to the objective orientation used to guide the optimization process; and optimization variables refer to the core parameters that can be adjusted and optimized during the solution process of the energy storage optimization configuration model. The system uses both economic efficiency and frequency stability as optimization directions, which can avoid the problem of unreasonable configuration caused by a single objective. By clarifying the optimization variables, it provides a solvable parameter basis for subsequent global optimization, ensuring that the configuration scheme takes into account both cost and control performance.

[0060] Furthermore, the selection of optimization variables is consistent with the actual control objects of the power grid, ensuring that the solution results can be directly implemented in engineering. In one possible implementation, the system uses the rated power and rated capacity of each energy storage site as optimization variables, while also including the site layout location in the optimization scope, forming a complete set of optimization variables.

[0061] Specifically, the system takes the lowest life-cycle cost of energy storage and the best frequency stability effect across all scenarios as dual optimization directions. Combining the grid structure and control requirements, it selects energy storage power, energy storage capacity, and energy storage layout location as optimization variables to complete the determination of optimization variables.

[0062] For example, for power grids in areas with a high proportion of new energy, the system takes the lowest life-cycle cost of energy storage and the best frequency stability effect across all scenarios as the optimization direction, and determines the rated power and rated capacity of each energy storage site as optimization variables.

[0063] Step S32: Based on the optimization variables, the improved particle swarm optimization algorithm is used to perform global optimization on the energy storage optimization configuration model. Iterative optimization is completed under the constraints of the entire scenario operation, and the values ​​of the optimization variables are continuously updated during the iteration process. It should be noted that the improved particle swarm optimization algorithm refers to an intelligent optimization algorithm that optimizes parameters such as inertia weight and learning factor based on the standard particle swarm optimization algorithm to improve optimization accuracy and convergence speed; global optimization solution refers to the calculation process of searching for the optimal solution in the entire feasible domain; full-scenario operation constraints refer to the safety operation restrictions that the equipment and power grid must meet under normal, disturbance, and black start conditions; iterative optimization refers to the process of gradually approaching the optimal solution through multiple iterations; the values ​​of optimization variables refer to the specific values ​​of parameters such as energy storage power and capacity updated in each iteration. The system uses the improved particle swarm optimization algorithm for solution, which can quickly converge to the global optimum under complex constraints, avoid getting trapped in local optima, and ensure the legality, effectiveness, and stability of the optimization process by iteratively updating the variable values ​​within the constraint range.

[0064] Furthermore, the iterative process simultaneously calculates the objective function value, guiding the optimization direction and improving the rationality of the optimization search. In one possible implementation, the system sets an appropriate population size and a maximum number of iterations, linearly decreases the inertia weight during iteration, and dynamically adjusts the particle position and velocity to achieve continuous updating of the optimization variable values.

[0065] Specifically, based on the determined optimization variables, the system calls the improved particle swarm optimization algorithm to perform global optimization on the energy storage optimization configuration model. Under the premise of meeting the constraints of the entire scenario operation, it performs iterative optimization and updates the values ​​of the optimization variables according to the algorithm rules in each iteration, gradually approaching the optimal solution.

[0066] Step S33: Based on the iterative update results, select the value combination that satisfies the constraints and has the best optimization effect, and determine the power configuration, capacity configuration and site layout of energy storage based on the value combination to form the optimal energy storage configuration scheme.

[0067] It should be noted that the iterative update result refers to the multiple sets of optimized variable values ​​and objective function calculation results obtained after all iterations are completed; satisfying the constraints means that the optimized variable values ​​are all within the safe range of the full-scenario operation constraints; optimal optimization effect means that the objective function value is optimal, that is, the overall cost and frequency stability effect is the best; value combination refers to a complete set of parameters composed of all optimized variables; power configuration refers to the rated power of each energy storage site; capacity configuration refers to the rated capacity of each energy storage site; site layout refers to the installation location of energy storage in the power grid.

[0068] Specifically, based on the results of iterative optimization, the system selects the optimal variable value combination that simultaneously meets the constraints of operation in all scenarios and has the best optimization effect. Based on this value combination, the power configuration, capacity configuration and site layout of energy storage are determined, and finally a complete optimal energy storage configuration scheme is formed.

[0069] In one embodiment, the optimization objectives are to achieve the lowest total lifecycle cost of energy storage and the best frequency stability across all scenarios. The objective function in step S22 can be referenced, where: To comprehensively optimize the objectives, The total lifecycle cost of energy storage, including initial investment cost, operation and maintenance cost, and equipment replacement cost, is calculated over a 15-year lifecycle using the discount rate. Take 5% as the unit power investment cost Taking 1.2 million yuan / MW as the unit capacity investment cost Taking 1.8 million yuan / MWh as the benchmark, the annual operation and maintenance cost is 2% of the initial investment cost, and the equipment replacement cost over the entire life cycle is calculated as 60% of the initial investment cost. This represents the cumulative frequency stability deviation across all scenarios. , These are the weighting coefficients for the cost target and the stability target, respectively, with values ​​of 0.4 and 0.6.

[0070] Furthermore, an improved particle swarm optimization algorithm is used to globally optimize the dual-objective optimization model (i.e., the energy storage optimization configuration model). The algorithm population size is set to 100, the maximum number of iterations is set to 500, the inertia weight is linearly decreased from 0.9 to 0.4, and the learning factor is... After iterative optimization and convergence, the optimal energy storage configuration scheme for the target area power grid is output: One centralized energy storage power station with a rated power of 300MW and a rated capacity of 600MWh is configured at key hub nodes of the regional power grid, serving as the main frequency regulation power source and black start emergency power source for the power grid; eight new energy power stations are equipped with distributed energy storage, with a rated power of 25MW and a rated capacity of 50MWh per station, for a total installed capacity of 200MW and a total capacity of 400MWh, used for power station-level output stabilization and local frequency support; the total installed capacity of energy storage in the entire region is 500MW and the total capacity is 1000MWh, meeting the demand boundaries and safety constraints of all scenarios, while achieving the comprehensive optimization of full life cycle cost and frequency stability effect.

[0071] This embodiment clarifies the dual-objective optimization direction and optimization variables, uses an improved particle swarm optimization algorithm to complete global optimization under full-scenario constraints, and determines the energy storage power, capacity and layout based on the optimal value combination, ultimately forming an optimal energy storage configuration scheme that takes into account both economy and frequency stability, providing a reliable hardware foundation and safe operation boundary for subsequent real-time frequency coordination control of the power grid.

[0072] In one feasible implementation, the step of collecting real-time grid operation data according to the optimal energy storage configuration scheme, extracting grid operation characteristics, and generating feedforward power compensation reference values ​​based on the grid operation characteristics includes: Step S41: Collect real-time grid operation data according to the optimal energy storage configuration scheme, and extract grid operation features from the real-time grid operation data through multi-scale sliding window processing; It should be noted that multi-scale sliding window processing refers to a data processing method that sets time windows of different durations to capture instantaneous frequency fluctuations and long-term trends. Power grid operation characteristics refer to key feature quantities extracted from raw operation data that characterize the dynamic properties, disturbance trends, and support capabilities of the power grid, including frequency deviation, frequency change rate, new energy output fluctuation characteristics, load change characteristics, system inertia level, and adjustable reserve capacity characteristics. The system uses the optimal energy storage configuration scheme as the basis for data acquisition, ensuring a high degree of matching between data dimensions and control requirements. Multi-scale sliding window processing can simultaneously capture short-term fluctuations and long-term trends, improving the completeness and accuracy of feature representation. In one possible implementation, the system sets a short-term window to capture instantaneous frequency fluctuation characteristics and a long-term window to identify continuous frequency change trend characteristics, obtaining the corresponding feature quantities through differential calculation and linear fitting.

[0073] Specifically, the system collects real-time grid operation data in accordance with the data collection range and dimensions determined by the optimal energy storage configuration scheme. It then uses a multi-scale sliding window to filter, differentiate, and fit the real-time grid operation data to extract features such as frequency dynamic characteristics, disturbance source change trends, and system real-time inertia levels that support full-scenario control.

[0074] For example, the system defines the acquisition range according to the optimal energy storage configuration scheme, and collects grid frequency, voltage, new energy output, energy storage state of charge, load and grid status data at a sampling frequency of 100Hz through the SCADA system and synchronous phasor measurement device. Multi-scale processing is completed by using a short time window of 200ms and a long time window of 2s to extract grid operation characteristics such as frequency deviation, frequency change rate, new energy fluctuation rate, load change rate and system inertia level.

[0075] In one embodiment, the instantaneous frequency change rate is obtained through differential calculation. The slope 'a' of the frequency change trend is obtained by linear fitting using the least squares method, which characterizes the dynamic change characteristics of the frequency and represents the instantaneous rate of change of the frequency. for:

[0076] In the formula: Let be the power grid frequency deviation at time t. At the current data sampling time, The sampling time interval for power grid operation data. This represents the power grid frequency deviation at the previous sampling time.

[0077] The slope characteristic of the frequency change trend was obtained by linear fitting using the least squares method, and the fitting formula is as follows:

[0078] In the formula: The duration of the long-term window; The slope characteristic of the frequency change trend obtained by fitting is used to characterize the continuous change trend of the frequency. This is the intercept of the fitted line.

[0079] Furthermore, the power output fluctuation rate and load change rate of new energy sources are calculated to identify the sources and trends of power disturbances. At the same time, the real-time inertia level of the entire network, the adjustable capacity and adjustable power margin of energy storage, and the frequency regulation reserve capacity of conventional units are calculated to characterize the current anti-disturbance capability of the system. The above features are integrated into a core feature set and synchronously input into the multi-source disturbance joint prediction model and the subsequent intelligent decision-making core.

[0080] Step S42: Construct a multi-source disturbance joint prediction model based on the power grid operation characteristics, and predict the power grid power imbalance through the multi-source disturbance joint prediction model; It should be noted that the multi-source disturbance joint prediction model refers to a time-series prediction model built on a long short-term memory network, using the time series of power grid operating characteristics as input, capable of simultaneously predicting fluctuations in renewable energy output and sudden load changes. Power grid power imbalance refers to the difference between power output and load demand in the future, characterizing the system's active power deficit or surplus. The system, based on power grid operating characteristics, constructs a multi-source disturbance joint prediction model that can identify power fluctuations caused by various types of disturbances in advance, accurately predict power grid power imbalance, provide a reliable basis for feedforward compensation, achieve proactive adjustment before disturbances occur, and significantly improve the response speed and disturbance rejection capability of frequency control.

[0081] Furthermore, the multi-source disturbance joint prediction model is constructed using a long short-term memory network, which can effectively capture the long-range dependencies of time-series data and improve prediction accuracy. In one possible implementation, the system uses historical grid operation characteristics at multiple time points as the model input of the multi-source disturbance joint prediction model, and uses future renewable energy output and load power at multiple time points as the model output of the multi-source disturbance joint prediction model, thus completing the construction and training of the multi-source disturbance joint prediction model.

[0082] Specifically, the system takes time-series data of power grid operation characteristics as input, constructs a multi-source disturbance joint prediction model using a long short-term memory network, uses the multi-source disturbance joint prediction model to predict the fluctuation of new energy output and load changes in the future time domain, and calculates the power imbalance of the power grid based on the prediction results.

[0083] For example, a multi-source disturbance joint prediction model based on LSTM is constructed. The model input is the core feature time series sequence of the past 10 time points, and the output is the predicted values ​​of renewable energy output and load power for the next 5 time points. The multi-source disturbance joint prediction model structure is set with a 2-layer LSTM network, with 64 neurons in each layer. The activation function is ReLU. The model is trained by the ADAM optimizer with a learning rate of 0.001. It outputs the predicted value sequence of total output of renewable energy power plants such as wind power and photovoltaic power plants in the entire network and the predicted value sequence of total active power load in the entire network within the preset time domain. At the same time, it outputs the predicted value of the adjustable output limit of conventional synchronous generators in the grid, providing a complete basic data source for the calculation of power imbalance and completing the prediction results of multi-source disturbances.

[0084] Furthermore, for different operating conditions of the power grid, such as normal frequency regulation, disturbance immunity, and black start recovery, the real-time adjustable active power benchmark value of the system at the current moment is calculated to complete the system adjustable power benchmark calculation. Then, based on the basic logic of power system active power balance, the total predicted load value in the future time domain is used as the total active power demand of the system, and the sum of the predicted total output of new energy sources, the predicted adjustable output of conventional units, and the system adjustable power benchmark value under the corresponding operating conditions is used as the total active power supply of the system. By calculating the difference between active power demand and active power supply, the power grid active power imbalance at each moment in the future time domain is obtained, completing the final calculation of the power grid power imbalance.

[0085] Step S43: Calculate the feedforward power compensation reference value based on the power imbalance of the power grid.

[0086] It should be noted that the feedforward power compensation reference value refers to a power regulation reference command calculated based on the predicted power imbalance in the power grid, used to offset power fluctuations in advance. The system generates the feedforward power compensation reference value based on the power imbalance in the power grid, enabling it to output compensation commands before the actual disturbance occurs, thus smoothing out frequency fluctuations in advance. This avoids the frequency overshoot problem caused by the lag in traditional feedback control and improves the inertia support capability and operational stability of high-proportion renewable energy power grids. Furthermore, the feedforward power compensation reference value can be adaptively adjusted based on the system's real-time inertia level; the lower the inertia, the greater the compensation intensity, to adapt to different operating conditions.

[0087] Specifically, the system uses the predicted power imbalance of the power grid as the basis for calculation, and performs calculations in combination with the preset feedforward compensation coefficient to directly generate a feedforward power compensation reference value for advanced regulation.

[0088] For example, a multi-source disturbance joint prediction model can be constructed. Using time-series data of the core feature set as input, a time-series prediction model (i.e., a multi-source disturbance joint prediction model) is built through a Long Short-Term Memory (LSTM) network. This model can then predict future fluctuations in renewable energy output and load surges. Based on the prediction results, the power imbalance of the power grid at future moments can be calculated, generating a feedforward power compensation reference value. The expression is:

[0089] In the formula: Forward compensation coefficient; The predicted future active power imbalance of the power grid is represented by a positive value to indicate the power gap. In this embodiment, the feedforward compensation coefficient... The system adaptively adjusts based on its real-time inertia level; the lower the system inertia, the better. The larger the value, the range is 0.6 to 1.2.

[0090] This embodiment guides data acquisition through the optimal energy storage configuration scheme, extracts grid operation characteristics through a multi-scale sliding window, constructs a multi-source disturbance joint prediction model based on a long short-term memory network to obtain the grid power imbalance, and calculates the feedforward power compensation reference value accordingly. This achieves advanced perception and compensation of multi-source disturbances, effectively improving the foresight and stability of grid frequency control.

[0091] In one feasible implementation, the step of generating adaptive optimal control commands based on the power grid operating characteristics and the feedforward power compensation reference value using a preset decision model includes: Step S51: Model the power grid frequency coordination control process as a discrete-time Markov decision process and construct a decision model that includes a continuous state space, a continuous action space, a multi-objective reward function, a state transition probability matrix, and a discount factor. It should be noted that the power grid frequency coordination control process refers to the control flow for achieving frequency stability regulation of the power grid under all scenarios including normal operation, disturbance rejection, and black start recovery; the discrete-time Markov decision process refers to a sequential decision model that combines system state transitions, action selection, and reward feedback within a discrete time step; the continuous state space refers to a continuous set of states consisting of power grid frequency deviation, frequency change rate, energy storage state of charge, renewable energy output, load, and grid on / off state; the continuous action space refers to a continuous set of actions consisting of energy storage charging and discharging power commands, renewable energy virtual inertia parameters, and primary frequency regulation droop parameters; the multi-objective reward function refers to a multi-dimensional evaluation function that integrates frequency stability, energy storage safety, and control smoothness; the state transition probability matrix refers to the probability distribution of the system transitioning from the current state to the next state after executing a certain action; the discount factor refers to the coefficient used to balance the weights of immediate and long-term rewards; and the decision model refers to an intelligent decision model used to achieve adaptive frequency control under all operating conditions.

[0092] The continuous state space and continuous action space can cover the full operating conditions of the power grid, enabling the decision model to adapt to the complex operating scenarios of a high proportion of renewable energy power grids. In one possible implementation, the system divides the control process into multiple decision-making moments using discrete time steps, and sequentially defines the state, action, reward, state transition probability, and discount factor to complete the modeling of the Markov decision process.

[0093] Specifically, the system transforms the full-scenario power grid frequency coordination control process into a discrete-time Markov decision process, sequentially constructing a continuous state space, a continuous action space, a multi-objective reward function, and a state transition probability matrix, and setting discount factors to form a complete decision model.

[0094] For example, the full-scenario power grid frequency coordination control problem can be modeled as a discrete-time Markov decision process (MDP), defined as a quintuple. ,in: It is a continuous state space, representing the real-time state of the power grid operation in all dimensions; The continuous action space represents the full-dimensional control actions that can be executed by this invention; It is a multi-objective reward function used to evaluate the quality of control actions and guide the agent to learn the optimal control strategy; The state transition probability matrix represents the probability distribution of the system transitioning from the current state to the next state after a control action is executed. This is a discount factor used to weigh immediate rewards against future long-term rewards.

[0095] Specifically, the discrete time step is set to 0.1s, a continuous state space is defined, and the core features of full-dimensional perception are fused to construct a state vector covering the entire scene's operational information. The expression is as follows:

[0096] In the formula: For power grid frequency deviation, The rate of change of frequency, Here are the charge state vectors for each energy storage system. To provide real-time power output for wind farms, To provide real-time power output for photovoltaic power plants, For real-time load across the entire network, This indicates the on / off state of the power grid structure.

[0097] Further define the continuous action space at time t The expression is:

[0098] In the formula: The charging and discharging power commands for each energy storage system are given, with negative values ​​used for charging. For virtual inertia control parameters of new energy power plants. These are the primary frequency regulation droop control parameters for new energy power plants.

[0099] Design a multi-objective comprehensive reward function. The expression is:

[0100] In the formula: Basic rewards, This is a frequency deviation penalty term. Penalty term for deviation of energy storage state of charge. To control the penalty items for drastic changes in movement, This is a penalty item for frequency exceeding limits or grid recovery failure.

[0101] Define a state transition probability matrix P, simulate the system state transition after executing control actions in a power grid simulation environment, and define a discount factor. Set to 0.99 to balance the weight of immediate rewards and long-term rewards.

[0102] Step S52: Extract long-term time-series dependency features of the power grid operation characteristics based on the long short-term memory network, and complete the training of the decision model using an improved dual-delay deep deterministic strategy gradient algorithm; It should be noted that Long Short-Term Memory (LSTM) networks refer to recurrent neural networks that can effectively capture long-term dependencies in time-series data; long-range temporal dependency features refer to temporal features that characterize the historical trends and future evolution of the power grid state; the improved dual-delay deep deterministic policy gradient algorithm refers to an enhanced deep reinforcement learning algorithm that integrates LSM networks, employs dual commentators, delayed updates, target policy smoothing, and priority experience replay; and the training of the decision model refers to the process of learning the optimal control strategy through agent interaction with the power grid environment, updating network parameters, and learning the optimal control strategy. The system extracts temporal trend features through LSM networks, which improves the foresight of decisions; the use of the improved dual-delay deep deterministic policy gradient algorithm for training alleviates value overestimation, improves training stability and convergence speed, and enables the decision model to learn the optimal control strategy suitable for all scenarios. In addition, the training process is divided into offline training and online optimization. Offline training completes policy learning, while online training ensures real-time performance. During the offline training phase, a power grid simulation environment covering all scenarios, including normal operation, disturbance, and black start, is constructed. This allows the agent to continuously interact with the environment, collect interaction samples to update network parameters, and learn the optimal frequency coordination control strategy for all scenarios by maximizing the accumulated reward value. During the online operation phase, the real-time collected system state vector is input into the trained agent network, and the adaptive optimal control command under the current operating condition is directly output through forward propagation. This includes the charging and discharging power commands of each energy storage site and the virtual inertia and droop control parameters of the new energy power station.

[0103] Specifically, the system inputs the time-series data of power grid operation characteristics into a long short-term memory network, extracts long-term time-series dependency features to characterize trend changes, and uses an improved dual-delay deep deterministic policy gradient algorithm to train the decision model through interactive samples and iterative updates.

[0104] Step S53: Input the power grid operating characteristics and the feedforward power compensation reference value into the trained decision model to generate adaptive optimal control commands that are suitable for different operating conditions.

[0105] It should be noted that the trained decision model refers to an intelligent decision model that has undergone offline training, parameter convergence, and can be directly inferred online; the adaptive optimal control command refers to a control command that can automatically match different operating conditions such as normal operation, disturbance, and black start to achieve optimal frequency regulation; the different operating conditions refer to three typical scenarios: normal stable operation of the power grid, high-power disturbance immunity, and black start recovery from extreme faults. The system inputs the power grid operating characteristics and feedforward power compensation reference values ​​into the decision model, and can output the optimal command by combining real-time status and advanced prediction information to achieve adaptive regulation under all operating conditions, solving the problems of poor adaptability, regulation lag, and insufficient robustness of traditional control methods.

[0106] Specifically, the system inputs the grid operation characteristics and feedforward power compensation reference values ​​into the trained decision model. The model combines long-range time-series dependency characteristics to complete inference calculations and outputs adaptive optimal control commands that can adapt to different operating conditions such as normal, disturbance, and black start.

[0107] For example, the system inputs the real-time grid operation characteristics and feedforward power compensation reference values ​​into the trained fusion LSTM improved TD3 decision model, so that the improved TD3 decision model can automatically identify the current operating conditions and infer the energy storage charging and discharging power, new energy virtual inertia and primary frequency regulation parameters to form adaptive optimal control commands.

[0108] In one embodiment, the temporal feature extraction layer uses a single-layer LSTM network with 128 neurons to process the input temporal sequence of states and extract long-range temporal dependency features of the power grid operation state; the actor network uses three fully connected layers with 64, 32, and 11 neurons respectively, the input of which is the temporal feature extracted by the LSTM, and the output is the control action vector of the continuous action space; the dual critic network uses two fully connected networks with identical structures, each with 64, 32, and 1 neurons, the input of which is the concatenated vector of state and action, and the output is the Q-value of the state-action pair; the target actor network and the target critic network are completely identical in structure to the main network, and a soft update strategy is used to synchronize parameters, with a soft update coefficient of 0.005; the critic network is updated twice, and the actor network and the target network are updated once.

[0109] Furthermore, in a power grid simulation environment jointly built with MATLAB and Python, a training sample set covering all scenarios including normal, disturbance, and black start was constructed. The number of training rounds was set to 2000, and the maximum step size per round was 2000. The ADAM optimizer was used to update the network parameters. The learning rate of the actor network was set to 0.0001, and the learning rate of the critic network was set to 0.0002. After training, a converged optimal control strategy model for all scenarios was obtained.

[0110] This embodiment constructs a standard decision model by modeling frequency coordination control as a discrete-time Markov decision process, extracts time-series features using a long short-term memory network, and completes model training using an improved TD3 algorithm. Finally, it combines real-time features and feedforward compensation to generate adaptive optimal control commands for all operating conditions, thus realizing intelligent, forward-looking, and robust frequency coordination control for high-proportion renewable energy power grids.

[0111] In one feasible implementation, the step of extracting long-range temporal dependency features of the power grid operation characteristics based on a long short-term memory network, and training the decision model using an improved dual-delay deep deterministic policy gradient algorithm, includes: Step S61: Process the time series of the power grid operation characteristics through a long short-term memory network to extract long-term time-series dependency features that characterize the frequency change and new energy fluctuation trends. It should be noted that the time series of power grid operation characteristics refers to multiple sets of power grid operation characteristic data arranged in chronological order; frequency change trends refer to the rising, falling, or stable changes in power grid frequency over a period of time; new energy fluctuation trends refer to the fluctuating changes in wind power and photovoltaic output over a period of time; long-range time-series dependency characteristics refer to time-series characteristics that can reflect the historical patterns and future evolution directions of the power grid state. The system utilizes a Long Short-Term Memory (LSTM) network to process the time series, which can fully explore the long-term changing patterns of the power grid operation state and effectively improve the foresight and adaptability of the control strategy. Furthermore, long-range time-series dependency characteristics can correlate historical states with future trends, enabling the decision-making model to have trend prediction capabilities, rather than relying solely on instantaneous states for decision-making.

[0112] Specifically, the system inputs the power grid operation characteristics arranged in chronological order into a Long Short-Term Memory (LSTM) network, encodes and extracts features from the time-series data, and outputs long-range temporal dependency features that characterize the frequency change trend and the new energy output fluctuation trend. For example, by using a LSTM network to process the input state vector time-series sequence, the system extracts the long-range temporal dependency features of the power grid operation state, captures the trend characteristics of frequency changes and new energy output fluctuations, and improves the foresight of decision-making.

[0113] Step S62: Input the long-range temporal dependency features into the actor network, and output the control action vector in the continuous action space from the actor network; It should be noted that the actor network refers to the neural network responsible for outputting the deterministic control strategy in the improved dual-delay deep deterministic policy gradient algorithm; the continuous action space refers to the set of continuously adjustable actions composed of energy storage charging and discharging power, new energy virtual inertia parameters, and primary frequency modulation droop parameters; and the control action vector refers to a one-dimensional or multi-dimensional vector containing control instructions from multiple execution units. By inputting long-range temporal dependency features into the actor network, the system can fully integrate trend information into the control actions, outputting more forward-looking instructions. Simultaneously, the continuous action space ensures fine and smooth control actions, avoiding step-like shocks. Furthermore, the actor network is a fully connected neural network structure, with temporal features as input and continuous actions satisfying constraints as output. In one possible implementation, the system uses a multi-layer fully connected network to construct the actor network, directly mapping the long-range temporal dependency features to obtain the control action vector within the continuous action space.

[0114] Specifically, the system inputs the extracted long-range temporal dependency features into the actor network, which then performs forward inference calculations based on the input features and outputs control action vectors within the legal range of the continuous action space. In one specific implementation, the system inputs the long-range temporal dependency features into a multi-layer fully connected actor network, and the network inference outputs control action vectors that include various energy storage charging and discharging power commands, new energy virtual inertia parameters, and primary frequency modulation droop parameters.

[0115] Step S63: Input the long-range temporal dependency features and the control action vector into the dual critic network to obtain two sets of value evaluation results, and select the smaller value as the target value evaluation result; It should be noted that the dual-critic network refers to two independent and parallel value evaluation networks used in the improved dual-delay deep deterministic policy gradient algorithm; the value evaluation result refers to the Q-value output by the critic network, used to evaluate the quality of the current state-action pair; the target value evaluation result refers to the final value evaluation index used to update the network parameters. By employing a dual-critic network and selecting the smaller value, the system can effectively alleviate the value overestimation problem commonly found in traditional deep reinforcement learning, improving training stability and policy reliability. Furthermore, the two critic networks have identical structures and their parameters are updated independently, allowing for the evaluation of action value from different perspectives, thus improving evaluation robustness.

[0116] Specifically, the system inputs long-range temporal dependency features and control action vectors into two structurally independent dual-critic networks, obtaining two sets of value evaluation results. The set with the smaller value is selected as the final target value evaluation result. The dual-critic network (Critic) uses two independent Q-network structures to evaluate the value of the state-action pair, and the smaller value of the two network outputs is taken as the value evaluation result. That is, the system concatenates the long-range temporal dependency features and control action vectors and inputs them into the dual-critic network to obtain two sets of Q-value evaluation results. The smaller Q-value is selected as the target value evaluation result for subsequent network parameter updates, thereby alleviating the value overestimation problem commonly found in traditional deep reinforcement learning algorithms and improving training stability.

[0117] Step S64: Based on the target value assessment result, update the network parameters of the actor network and the dual critic network at intervals according to the delayed update strategy, and construct an experience replay buffer to store samples of the agent's interaction with the power grid environment. Determine the sample priority based on the time-series differential error, and extract samples according to the priority to complete the iterative optimization of the decision model.

[0118] It should be noted that the delayed update strategy refers to the parameter update rule that makes the actor network update frequency lower than the critic network update frequency; network parameters refer to the weights and biases of each layer in the neural network; the experience replay buffer refers to the dataset used to store interaction samples between the agent and the power grid environment; the agent refers to the reinforcement learning subject that performs decisions, obtains rewards, and updates the strategy; the power grid environment refers to the simulated or actual power grid operation state and feedback mechanism; the interaction sample refers to the training data unit containing state, action, reward, and next state; the temporal difference error refers to the error index used to measure the accuracy of value assessment; the sample priority refers to the weight of samples selected for training based on the error magnitude; and the iterative optimization of the decision model refers to the process of gradually converging the model strategy to the optimal state through multiple rounds of sample training and parameter updates. The system adopts delayed updates to avoid strategy oscillations, and adopts priority experience replay to improve the utilization rate of high-value samples, accelerate the convergence speed, and improve the performance of the control strategy.

[0119] Furthermore, the target actor network and the target critic network synchronize their parameters from the main network using a soft update method, further improving training stability. In one possible implementation, the system updates the dual critic networks multiple times before updating the parameters of the actor network and the target network in a single update, and prioritizes the samples according to their temporal difference error from largest to smallest.

[0120] Specifically, the system calculates the gradient based on the target value assessment result and updates the parameters of the dual critic network. It then updates the actor network parameters at intervals according to a delayed update strategy. Finally, it synchronizes the main network parameters to the target network via soft updates. The delayed update strategy sets the actor network's update frequency lower than the critic network's; every N updates to the critic network, the actor and target network parameters are updated once more, avoiding training oscillations caused by overly frequent strategy updates. Additionally, a small amount of random noise is added to the target action to smooth the target Q-values ​​of adjacent states, reducing the value function's sensitivity to action changes and improving the policy's robustness.

[0121] Simultaneously, an experience replay buffer is constructed to store samples of interactions between the agent and the environment. Based on the temporal difference error, the priority of samples is calculated, and samples that contribute more to strategy optimization are sampled first, improving sample utilization efficiency and algorithm convergence speed. For example, continuous collection of power grid operation data and control effect data in all scenarios is performed, and an operation sample set containing state sequences, action sequences, reward values, and frequency control effects is formed in 1-hour cycles. A priority experience replay buffer is constructed, and the sample priority is calculated based on the temporal difference error. High-value samples are stored in the buffer, with a maximum buffer capacity of 1,000,000. A 24-hour sliding time window is set, and within each time window, the network parameters of the improved TD3 algorithm are incrementally trained based on the incremental training sample set. The learning rate is set to 1 / 10 of the offline training rate. After the parameter update is completed, the optimized decision model is synchronized to the online decision-making process to achieve adaptive optimization of the control strategy.

[0122] For example, the system updates the dual critic network based on the target value assessment results, updates the actor network once after every two updates of the critic network, constructs an experience replay buffer to store state, action, reward, and next state samples, determines the priority and extracts samples based on the temporal difference error, and continuously iterates and optimizes until the decision model converges, thus completing the iterative optimization of the decision model.

[0123] This embodiment extracts long-term temporal dependency features through a long short-term memory network, outputs control actions through an actor network, completes value evaluation through a dual critic network, and uses the minimum Q value to mitigate overestimation. It combines delayed updates, soft updates, and priority experience replay mechanisms to complete the iterative optimization of the decision model, and finally obtains a stable, efficient, and robust full-scene frequency coordination control strategy.

[0124] In one feasible implementation, the adaptive optimal control command is verified and corrected based on preset full-scenario safety constraint rules, and the corrected executable control command is sent to the execution unit to achieve grid frequency coordinated control, including: Step S71: The adaptive optimal control command and the feedforward power compensation reference value are weighted and fused to obtain the control command to be verified. It should be noted that weighted fusion refers to a processing method in which two types of commands are weighted separately according to preset weight coefficients and then superimposed and combined; the control commands to be verified refer to control commands that have completed feedforward and feedback fusion but have not yet undergone safety constraint verification. The system weights and fuses the adaptive optimal control command with the feedforward power compensation reference value, which can simultaneously leverage the advantages of high accuracy of feedback control and fast response of feedforward control, achieving synergy between proactive and precise regulation, and significantly improving the dynamic performance and steady-state effect of frequency control. In addition, the weight coefficients used in weighted fusion can be adaptively adjusted according to the real-time operating conditions of the power grid, the system inertia level, and the disturbance intensity.

[0125] Specifically, the system weights and superimposes the adaptive optimal control command and the feedforward power compensation reference value according to the set weight ratio, and then performs numerical fusion to form a control command to be verified that has both feedforward prediction and feedback optimization effects.

[0126] In one embodiment, the adaptive optimal control command output by the intelligent agent is weighted and fused with the feedforward power compensation reference value generated by multi-source disturbance prediction to obtain the control command to be verified, wherein the fusion expression of the energy storage charging and discharging power command is:

[0127] In the formula: The merged energy storage power command to be verified; The energy storage power control command output by the deep reinforcement learning agent; , The weight coefficients for the feedback decision instruction and the feedforward compensation instruction are respectively, satisfying... In this embodiment, the feedback decision instruction weight Feedforward compensation instruction weight .

[0128] Step S72: Based on the full-scenario safety constraint rules, perform amplitude limiting verification and instruction correction on the control instruction to be verified, and generate an executable control instruction; It should be noted that the full-scenario safety constraint rules refer to a set of constraints covering all operating conditions, including normal frequency regulation, disturbance immunity, and black start recovery, encompassing physical constraints of energy storage devices, grid safety operation constraints, and operating limits of new energy devices. Limit verification refers to the process of comparing the control command to be verified with the upper and lower limits of the constraints to determine whether the command exceeds the limits. Command correction refers to the adjustment operation of limiting command components exceeding safety boundaries to within the allowable range and eliminating illegal components. Executable control commands refer to the final control commands that satisfy all safety constraints, comply with the physical limits of the equipment, and can be directly issued to field units for execution. By implementing limit verification and command correction through full-scenario safety constraint rules, the system can avoid equipment damage, frequency violations, or grid operation risks caused by control commands exceeding limits, ensuring the safety and reliability of the control process.

[0129] In addition, the safety constraint rules for all scenarios are derived from the safety operation boundaries determined by the optimal energy storage configuration scheme, and are highly consistent with the previous planning.

[0130] In one possible implementation, the system performs amplitude limiting verification on components such as energy storage charging and discharging power, state of charge, power change rate, grid frequency deviation, and voltage deviation one by one, and corrects the out-of-limit components to within the constraint limit.

[0131] Specifically, the system compares the control command to be verified with the full-scenario safety constraint rules item by item, performs limit violation judgment, limit pruning and command correction, and finally generates an executable control command that meets all safety constraints.

[0132] For example, a comprehensive set of safety constraints can be constructed, based on optimal energy storage configuration schemes, power grid safety operation procedures, and equipment physical limits, to build a multi-dimensional safety constraint system covering energy storage, power grid, and new energy equipment, including:

[0133] In the formula: For the first Energy storage sites The charging and discharging power command at any given time; , The first Upper and lower limits of charging and discharging power for each energy storage site; , The first Safety upper and lower limits of the state of charge of an energy storage site; For the first Power change rate limits for individual energy storage sites; , These are the upper and lower safety limits for grid frequency deviation; , These are the upper and lower safety limits for grid voltage deviation, respectively.

[0134] Furthermore, based on the full-scenario safety constraint set, component-by-component safety verification and amplitude limiting processing are performed on the control commands to be verified. Command components that exceed the safety boundary are limited to the allowable range. After eliminating non-compliant command components, the final executable control commands are generated to ensure that all commands meet the physical limits of the equipment and the safety requirements of the power grid.

[0135] Step S73: The executable control command is sent to the execution unit and the power grid operating status is fed back to realize closed-loop coordinated control of power grid frequency.

[0136] It should be noted that the execution unit refers to the field equipment in the power grid that has power regulation and frequency support capabilities, including energy storage systems, new energy power plants, and controllable loads; the power grid operating status refers to the real-time frequency, voltage, new energy output, energy storage state of charge, load, and grid status information of the power grid after the command is executed; closed-loop coordinated control refers to the operation mode that forms a closed control loop by issuing commands, collecting status data, providing effect feedback, and optimizing strategies. The system issues executable control commands to the execution unit and collects the power grid operating status in real time, forming a complete control closed loop. It continuously senses the control effect and provides the latest input for subsequent decisions, ensuring that the control strategy always matches the actual operating status of the power grid, thereby improving the stability and adaptability of frequency coordinated control across all scenarios.

[0137] In addition, the power grid's operating status is transmitted back to the status perception and feedforward prediction module in real time, forming a continuous closed-loop iterative mechanism. In one possible implementation, the system issues executable control commands to each execution unit through the power grid dispatch automation system. After the execution unit completes its actions, the status perception module collects the latest operating data and transmits it back to the decision-making stage.

[0138] Specifically, the system sends executable control commands to each execution unit, which then performs adjustment actions according to the commands. At the same time, it feeds back the real-time operating status of the power grid after execution to the status sensing module, forming a closed control loop and realizing closed-loop coordinated control of the power grid frequency.

[0139] For example, the final executable control commands are sent to various new energy power plants, energy storage power plants, controllable loads and other execution units through the power grid dispatch automation system. The execution units complete the corresponding charging and discharging, parameter adjustment and other actions according to the commands. At the same time, the real-time operation status data of the power grid after the command is executed is transmitted back to the all-dimensional power grid status perception module at a frequency of 100Hz to complete the closed-loop iteration of frequency coordination control and continuously optimize the control effect.

[0140] This embodiment improves control dynamic performance by weighted fusion of adaptive optimal control commands and feedforward power compensation reference values, ensures operational safety through full-scenario safety constraint verification, and forms closed-loop control through command issuance and status feedback. Ultimately, it achieves safe, stable, and adaptive frequency coordination control of a high-proportion renewable energy power grid under all scenarios, including normal, disturbance, and black start conditions.

[0141] Understandably, this method uses the state vector... The system automatically identifies grid operating conditions, enabling adaptation to all scenarios. When the system is operating under normal and stable conditions, the core control objective is to finely smooth out fluctuations in renewable energy output, control the grid frequency within ±0.1Hz, and simultaneously optimize the energy storage SOC to maintain around 0.5, reducing equipment losses. The relevant implementation effects can be seen in Figure 2.

[0142] When the system operates under disturbance and immunity scenarios, the core control objective is to rapidly smooth frequency fluctuations, with energy storage providing millisecond-level response to fill power gaps, and renewable energy power plants simultaneously enhancing their virtual inertia support capabilities. The goal is to control the maximum frequency overshoot within ±0.5Hz to prevent the escalation of accidents. The relevant implementation results can be found in [reference needed]. Figure 3 .

[0143] When the system operates under black start and grid restoration conditions, the core control objectives are to establish stable frequency and voltage references, mitigate power surges caused by load switching, reserve 30% emergency backup capacity for energy storage, and fully support the entire process of grid restoration from a blackout state. The relevant implementation effects can be found in [reference needed]. Figure 4 .

[0144] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0145] This application also provides a power grid frequency control device, the power grid frequency control device comprising: The acquisition module is used to acquire the optimal energy storage configuration scheme, wherein the optimal energy storage configuration scheme is obtained by global optimization based on a preset energy storage optimization configuration model; The extraction module is used to collect real-time power grid operation data according to the optimal energy storage configuration scheme, extract power grid operation characteristics, and generate feedforward power compensation reference values ​​based on the power grid operation characteristics. The generation module is used to generate adaptive optimal control commands based on the power grid operation characteristics and the feedforward power compensation reference value using a preset decision model. The control module is used to verify and correct the adaptive optimal control command based on preset full-scenario safety constraint rules, and send the corrected executable control command to the execution unit to realize grid frequency coordinated control.

[0146] The power grid frequency control device provided in this application, employing the power grid frequency control method in the above embodiments, can solve the technical problems in the background art. Compared with the prior art, the beneficial effects of the power grid frequency control device provided in this application are the same as the beneficial effects of the power grid frequency control method provided in the above embodiments, and other technical features in the power grid frequency control device are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.

[0147] This application provides a power grid frequency control device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the power grid frequency control method in Embodiment 1 above.

[0148] The following is for reference. Figure 5 The diagram illustrates a structural schematic of a power grid frequency control device suitable for implementing embodiments of this application. The power grid frequency control device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The power grid frequency control device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0149] like Figure 5 As shown, the power grid frequency control device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the power grid frequency control device. The processing unit 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. Communication device 1009 allows the power grid frequency control equipment to communicate wirelessly or wiredly with other equipment to exchange data. Although the figure shows power grid frequency control equipment with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0150] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0151] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0152] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the power grid frequency control methods provided by the above methods.

[0153] The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0154] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0155] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.

Claims

1. A power grid frequency control method, characterized in that, include: Obtain the optimal energy storage configuration scheme, wherein the optimal energy storage configuration scheme is obtained by global optimization solution based on a preset energy storage optimization configuration model; Real-time grid operation data is collected according to the optimal energy storage configuration scheme, grid operation characteristics are extracted, and feedforward power compensation reference values ​​are generated based on the grid operation characteristics. Based on the power grid operation characteristics and the feedforward power compensation reference value, an adaptive optimal control command is generated using a preset decision model. The adaptive optimal control command is verified and corrected based on the preset full-scenario safety constraint rules, and the corrected executable control command is sent to the execution unit to realize the coordinated control of power grid frequency.

2. The power grid frequency control method as described in claim 1, characterized in that, Before obtaining the optimal energy storage configuration scheme, the following steps are also included: Acquire power grid structure data, equipment operating parameters, and historical operating data to determine energy storage performance requirements under normal frequency regulation, disturbance immunity, and black start recovery scenarios, respectively. Obtain relevant power grid operation data and full-scenario operation constraints. Based on the relevant power grid operation data, the full-scenario operation constraints, and the energy storage performance requirements, construct a model framework with the optimization objectives of optimal cost and optimal frequency stability. Embed the physical constraints of energy storage devices, power grid frequency security constraints, and black-start emergency backup capacity constraints to form an energy storage optimization configuration model.

3. The power grid frequency control method as described in claim 2, characterized in that, The process of obtaining the optimal energy storage configuration includes: The optimization variables are determined with the goal of achieving the lowest total lifecycle cost of energy storage and the best frequency stability across all scenarios. Based on the optimization variables, an improved particle swarm optimization algorithm is used to perform global optimization on the energy storage optimization configuration model. Iterative optimization is completed under the constraints of the entire scenario operation, and the values ​​of the optimization variables are continuously updated during the iteration process. Based on the iterative update results, select the value combination that satisfies the constraints and has the best optimization effect. Based on the value combination, determine the power configuration, capacity configuration and site layout of energy storage to form the optimal energy storage configuration scheme.

4. The power grid frequency control method as described in claim 1, characterized in that, The step of collecting real-time grid operation data according to the optimal energy storage configuration scheme, extracting grid operation characteristics, and generating feedforward power compensation reference values ​​based on the grid operation characteristics includes: Real-time grid operation data is collected according to the energy storage optimal configuration scheme, and grid operation characteristics are extracted from the real-time grid operation data through multi-scale sliding window processing. Based on the aforementioned power grid operation characteristics, a multi-source disturbance joint prediction model is constructed, and the power grid power imbalance is predicted through the multi-source disturbance joint prediction model. Based on the power imbalance in the power grid, the reference value for feedforward power compensation is calculated.

5. The power grid frequency control method as described in claim 1, characterized in that, The step of generating adaptive optimal control commands based on the power grid operating characteristics and the feedforward power compensation reference value using a preset decision model includes: The power grid frequency coordination control process is modeled as a discrete-time Markov decision process, and a decision model is constructed that includes a continuous state space, a continuous action space, a multi-objective reward function, a state transition probability matrix, and a discount factor. Based on the long short-term memory network, the long-term time-series dependency features of the power grid operation characteristics are extracted, and the improved dual-delay deep deterministic strategy gradient algorithm is used to complete the training of the decision model. The power grid operating characteristics and the feedforward power compensation reference value are input into the trained decision model to generate adaptive optimal control commands that are suitable for different operating conditions.

6. The power grid frequency control method as described in claim 5, characterized in that, The process involves extracting long-range temporal dependency features of the power grid operation characteristics based on a long short-term memory network, and training the decision model using an improved dual-delay deep deterministic strategy gradient algorithm, including: The time series of the power grid operation characteristics are processed by a long short-term memory network to extract long-term time-series dependency features that characterize the frequency change and the trend of new energy fluctuations. The long-range temporal dependency features are input into the actor network, and the actor network outputs the control action vector in the continuous action space. The long-range temporal dependency features and the control action vector are input into a dual critic network to obtain two sets of value evaluation results, and the smaller value is selected as the target value evaluation result. Based on the target value assessment results, the network parameters of the actor network and the dual critic network are updated at intervals according to the delayed update strategy, and an experience replay buffer is constructed to store samples of the agent's interaction with the power grid environment. The sample priority is determined based on the time-series differential error, and samples are extracted according to the priority to complete the iterative optimization of the decision model.

7. The power grid frequency control method as described in claim 1, characterized in that, The adaptive optimal control command is verified and corrected based on preset full-scenario safety constraint rules, and the corrected executable control command is sent to the execution unit to achieve grid frequency coordinated control, including: The adaptive optimal control command is weighted and fused with the feedforward power compensation reference value to obtain the control command to be verified. Based on the full-scenario security constraint rules, the control command to be verified is subjected to amplitude limiting verification and command correction to generate an executable control command. The executable control commands are sent to the execution unit and the power grid operating status is fed back to achieve closed-loop coordinated control of the power grid frequency.

8. A power grid frequency control device, characterized in that, include: The acquisition module is used to acquire the optimal energy storage configuration scheme, wherein the optimal energy storage configuration scheme is obtained by global optimization based on a preset energy storage optimization configuration model; The extraction module is used to collect real-time power grid operation data according to the optimal energy storage configuration scheme, extract power grid operation characteristics, and generate feedforward power compensation reference values ​​based on the power grid operation characteristics. The generation module is used to generate adaptive optimal control commands based on the power grid operation characteristics and the feedforward power compensation reference value using a preset decision model. The control module is used to verify and correct the adaptive optimal control command based on preset full-scenario safety constraint rules, and send the corrected executable control command to the execution unit to realize grid frequency coordinated control.

9. A power grid frequency control device, characterized in that, The power grid frequency control device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the power grid frequency control method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the power grid frequency control method as described in any one of claims 1 to 7.