AI-driven power system dynamic frequency stability optimization scheduling method, device and equipment and storage medium

CN121097845BActive Publication Date: 2026-08-21GUANGXI UNIV FOR NATITIES
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511190532.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-25
Publication Date
2026-08-21
Estimated Expiration
2045-08-25

AI Technical Summary

Technical Problem

然而,可再生能源发电具有波动性、间歇性等特点,其大规模接入使得电网的频率稳定性受到严峻挑战

Benefits of technology

[0071]本申请通过构建集成潮流计算、时域仿真及奖励计算的强化学习环境,精确建模频率动态响应过程与调度约束;结合深度强化学习算法训练智能体生成优化调度策略,可在多类扰动工况下进行策略优化训练,获得适应系统实际运行状态的智能控制策略;该方法能够基于电力系统实时运行状态生成兼顾动态频率稳定性和经济性的优化调度指令,实现对频率响应过程的快速调节与稳定控制,有效提升高比例新能源接入背景下电力系统抵御扰动的能力,降低频率失稳及系统崩溃的风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121097845B_ABST
    Figure CN121097845B_ABST
Patent Text Reader

Abstract

The application provides an AI-driven power system dynamic frequency stability optimization scheduling method, device, equipment and storage medium. It relates to the technical field of power system optimization scheduling. The method comprises: constructing a reinforcement learning environment for simulating the dynamic response of the power system frequency; defining the state space, action space and reward function, wherein the state space includes voltage, power and other information, the action space includes generator output adjustment instructions, and the reward function integrates frequency stability and economic indicators; training the agent using a deep reinforcement learning algorithm to obtain the optimal scheduling strategy; deploying the trained optimal strategy to the actual power system scheduling process to achieve fast control and economic optimization of the system frequency response under large disturbance. The application can improve the frequency stability and operating efficiency of the power system under high proportion of new energy access.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of power system optimization and dispatching technology, and in particular to an AI-driven method, apparatus, equipment and storage medium for dynamic frequency stability optimization and dispatching of power systems. Background Technology

[0002] With the continuous growth of clean energy installed capacity, the traditional power structure dominated by thermal power generation is undergoing profound changes, and the proportion of thermal power units in the system is gradually decreasing. However, renewable energy generation is characterized by volatility and intermittency, and its large-scale integration poses a severe challenge to the frequency stability of the power grid. On the one hand, the decline in the proportion of traditional thermal power units leads to a decrease in system inertia, causing the frequency of the power system to drop faster and its inertial support capacity to weaken when subjected to disturbances. On the other hand, the number of controllable frequency regulation resources is relatively insufficient, making it difficult to respond effectively to frequency fluctuations in a timely manner, thereby increasing the risk of frequency instability or even frequency collapse in the power grid. Therefore, how to ensure the dynamic stability of the power system frequency under the background of high proportion of new energy integration into the grid has become one of the key issues that urgently need to be addressed in the field of power grid dispatching.

[0003] Meanwhile, with the rapid development of artificial intelligence technology and the continuous improvement of computing resources, data-driven modeling and intelligent decision-making methods have been widely applied in power system operation analysis and control optimization. Especially in frequency control scenarios, extracting key features of system operating status from large-scale, multi-source, heterogeneous wide-area measurement data (such as SCADA and PMU data) using machine learning techniques to construct disturbance identification models and frequency response prediction models has become an important research direction for next-generation frequency stability control. These methods can predict key indicators such as the lowest frequency after a disturbance and frequency recovery capability, providing data support and a modeling foundation for building more intelligent and adaptive frequency regulation strategies. Summary of the Invention

[0004] This application provides an AI-driven dynamic frequency stability optimization scheduling method, device, equipment, and storage medium for power systems. By constructing a dynamic frequency optimization scheduling model that reflects the system's disturbance response characteristics, and introducing a deep reinforcement learning algorithm to train and solve the model, a preventive scheduling scheme can be formulated based on the strategy learned by the agent. This enables the power system to achieve a primary frequency regulation response when subjected to disturbances such as load changes or unit tripping, effectively suppressing frequency drops and reducing the risk of frequency instability and system collapse.

[0005] In a first aspect, this application provides an AI-driven dynamic frequency stability optimization scheduling method for power systems, comprising:

[0006] A reinforcement learning environment is constructed; wherein the environment integrates power flow calculation, time-domain simulation and reward calculation; the time-domain simulation is based on the dynamic frequency stability model of the power system and is used to simulate the dynamic trajectory of the system frequency under disturbance;

[0007] The elements of deep reinforcement learning are determined; wherein, the elements of deep reinforcement learning include a state space, an action space, and a reward function, the state space includes measurement data reflecting the operating state and dynamic frequency stability of the power system, the action space includes dispatch control instructions for adjustable generation resources or controllable loads, and the reward function is a composite function that integrates dynamic frequency stability indicators and system economic indicators, and is calculated in real time by the reward based on the system state.

[0008] In the environment described above, based on the state space, action space, and reward function, a deep reinforcement learning algorithm is used to train the agent, enabling the agent to learn an optimized scheduling strategy; wherein, the optimized scheduling strategy is used to select actions to maximize the cumulative reward.

[0009] The trained agent is deployed to the real-time dispatching process of the power system to generate dispatching instructions based on real-time status input, thereby achieving joint optimization of the system's dynamic frequency stability and economy.

[0010] In one possible design, the time-domain simulation is constructed using the system frequency dynamic response equations, which include the generator swing equation, the speed regulation system dynamic response equation, and the node power balance equation; the specific methods for constructing the time-domain simulation include:

[0011] A swing equation for a synchronous generator is established, which is constrained by electromagnetic power, generator angular frequency, generator power angle, and generator inertial time constant. The expression is as follows:

[0012]

[0013] Where, δ i Let ω be the generator power angle of node i. i Let ω be the rotor angular velocity of the generator connected to node i. ref H is the rated rotor angular velocity of the generator connected to node i. i Let P be the inertial time constant of the generator at node i. m,i P represents the mechanical power of the generator prime mover connected to node i. e,i E represents the electromagnetic power generated by the generator at node i. i The potential of the generator. This represents the differential power angle variable of the generator connected to node i. Let V represent the differential angular velocity of the generator connected to node i, where m is the node number of the generator connected to the node.i Let θ be the voltage amplitude at node i connected to the generator. i Let x' be the voltage phase angle at node i where the generator is connected. d,i E is the transient reactance of the generator. i The generator's electromotive force;

[0014] Based on the speed governor's parameter data, the response equation of the generator speed control system is established, and the expression is:

[0015]

[0016] P m,i =(1-F HP,i )z g3,i +F HP,i z g2,i i = 1, ..., m;

[0017] Among them, T G,i T CH,i T RH,i z represents the time constant of each stage. g1,i z g2,i z g3,i K represents the output variable of each integral stage. G P represents the unit regulating power of the generator. m0,i To affect the mechanical power of the generator before the disturbance, F HP,i P represents the percentage of steam in the high-pressure cylinder relative to the total steam volume. m,i Let dt be the mechanical power of the generator connected to node i, t be a time variable, and dt be the derivative with respect to time t.

[0018] The nodal power balance equations are established as follows:

[0019]

[0020] Among them, Y aij α represents the (i,j) element in the extended network node admittance matrix that takes into account the generator potential. aij Let (i,j) be the (i,j) element of the phase angle matrix of the network node admittance. Let be the voltage amplitude at node i after the disturbance. Let i be the voltage phase angle at node i after the disturbance. To represent the j-th element of the extended network node voltage vector that takes into account the generator potential, To represent the j-element of the phase angle vector of the extended network nodes taking into account the generator electromotive force, P d,i Q represents the active load of node i before the disturbance. d,i For the reactive load of node i before the disturbance, ΔP d,i Let be the active power lost by node i due to the sudden disconnection of renewable energy from the grid, and n be the number of nodes in the system.

[0021] In one possible design, the state space includes at least one of the following:

[0022] Generator active power output, generator reactive power output, load node active power, load node reactive power, voltage amplitude, system frequency, critical node frequency, generator speed or power angle, power flow on critical lines, renewable energy output, and frequency change rate signal.

[0023] In one possible design, the action space includes at least one of the following:

[0024] Traditional generator sets include active power adjustment commands, generator terminal voltage regulation commands, and energy storage device charging and discharging control commands.

[0025] In one possible design, the reward function considers indicators of device security, dynamic frequency stability, and system economy, specifically including the following constraints and reward quantification mechanisms:

[0026] Under normal operating conditions prior to the disturbance, the node voltage amplitude, generator output, and line current should meet the following requirements:

[0027] V min,i ≤V i ≤V max,i i = 1, ..., n

[0028] P gmin,i ≤P g,i ≤P gmax,i i = 1, ..., m

[0029] Q gmin,i ≤Q g,i ≤Q gmax,i i = 1, ..., m

[0030] Among them, V min,i V max,i P represents the upper and lower limits of the voltage amplitude at node i, respectively. g,i P represents the active power of the generator connected to node i. gmin,i P gmax,i These are the upper and lower limits of the active power of the generator at node i, respectively, and Q. g,i Q represents the reactive power of the generator connected to node i. gmin,i Q gmax,i These are the upper and lower limits of the reactive power of the generator at access node i, respectively.

[0031] The increased power generation and primary frequency regulation reserve capacity of the generator after the disturbance are set to meet the following requirements:

[0032] P m,i-P g,i ≥0, i=1,...,m

[0033] P m,i -P g,i ≤P rmax,i i = 1, ..., m

[0034] Among them, P m,i P represents the mechanical power of the generator at node i. rmax,i The maximum value of the primary frequency regulation reserve for the generator at access node i;

[0035] Set the minimum frequency constraint and the rate of change of frequency constraint, with the following expressions:

[0036]

[0037] Among them, f lim ω is the minimum frequency allowed for frequency stability requirements. i This represents the angular frequency of the generator connected to node i. Δω is the maximum permissible rate of frequency change. i The change in angular frequency of the generator at access node i, where Δt is the change in time;

[0038] Reward function R t Represented as:

[0039]

[0040] Where max_penalty represents the maximum penalty value set by the power flow divergence system caused by the agent's scheduling instructions, λ dyn dyn_violation represents the penalty value for violating constraints during the system's dynamic frequency response, λ. sts sts_violation represents the penalty value for violations of runtime and technical constraints. This represents the reward value when no constraint is violated;

[0041] During the system's dynamic frequency response, the penalty value for violating constraints is determined by λ. dyn Composed of dyn_violation, λ dyn Here, is the set penalty coefficient, and dyn_violation is the degree of constraint violation, expressed as follows:

[0042] dyn_violation=v flim +v RoCoF +v mg +v pr

[0043] Among them, v flimThe minimum frequency f during the response process nadir The degree to which the frequency falls below the lower frequency limit constraint is expressed as:

[0044] v flim =min(0,f nadir -f lim )v RoCoF The expression representing the degree to which the rate of change of frequency violates the specified value during the response process is:

[0045]

[0046] v mg The expression representing the degree of violation of the power constraint during the response process is:

[0047] v mg =min(0,P) m,i -P g,i )v pr The expression representing the degree of violation of the maximum value of primary frequency regulation reserve during the response process is:

[0048] v pr =min(0,P) rmax,i -P m,i -P g,i ))

[0049] During the inspection of operational and technical constraints, the penalty value for violating the constraints is determined by λ. sts Composed of sts_violation, λ sts Here, `sts_violation` is the set penalty coefficient, and `sts_violation` is the degree of constraint violation, expressed as follows:

[0050]

[0051] in, Indicates the degree of violation of node voltage. This indicates the degree of violation of the generator's active power output. The degree of violation of generator reactive power output is expressed as follows:

[0052]

[0053] Where max is the maximum value function.

[0054] In one possible design, the reward function considers system economic indicators, specifically including:

[0055] Based on the optimal power flow model, the active power output of generators is optimized and rescheduled. The scheduling objective is set as minimizing the generator output adjustment cost, expressed as:

[0056]

[0057] in, The unit cost of increasing the active power output of the generator at node i. The active power added to the generator of access node i. To reduce the unit cost of active power output of the generator at node i. The active power reduction for the generator at node i is denoted by m, where m is the number of generators.

[0058] In one possible design, the constructed time-domain simulation module includes a set of initial condition equations used to set the initial operating states of various components in the power system; wherein, the initial condition equations include:

[0059]

[0060] In the formula, To change the rotor angular velocity of the synchronous generator before the disturbance, and E is the initial value of the integral variable before the disturbance. i P is the generator potential. g,i Q represents the active power generated by the generator at node i before the disturbance. g,i Let be the reactive power generated by the generator at node i before the disturbance. The power angle of the synchronous generator before the disturbance.

[0061] Secondly, this application provides an AI-driven dynamic frequency stability optimization scheduling device for power systems, the device comprising:

[0062] The reinforcement learning environment construction module is configured to build a simulation-based reinforcement learning environment, which integrates a power flow calculation unit, a time-domain simulation unit, and a reward calculation unit; wherein, the time-domain simulation unit is based on the dynamic frequency response model of the power system and is used to simulate the dynamic frequency trajectory of the system under disturbance.

[0063] The learning element determination module is configured to determine deep reinforcement learning elements; wherein, the deep reinforcement learning elements include a state space, an action space, and a reward function, the state space includes measurement data reflecting the operating state and dynamic frequency stability of the power system, the action space includes dispatch control commands for adjustable generation resources or controllable loads, and the reward function is a composite function that integrates dynamic frequency stability indicators and system economic indicators, and is calculated in real time by the reward calculation unit based on the simulation output state;

[0064] The agent training module is configured to train the agent using a deep reinforcement learning algorithm based on the state and reward data output by the reinforcement learning environment; enabling the agent to learn and optimize scheduling strategies; wherein the agent maximizes cumulative rewards through action selection;

[0065] The optimized scheduling module is configured to deploy the trained agent to the real-time scheduling process of the power system, generate scheduling instructions based on the real-time status input, and achieve joint optimization of the system's dynamic frequency stability and economy.

[0066] The output of the reinforcement learning environment construction module is connected to the input of the agent training module; the input of the optimization scheduling module is connected to the real-time status data stream of the power system.

[0067] Thirdly, embodiments of this application provide an electronic device, including: at least one processor and a memory; the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to execute the AI-driven dynamic frequency stability optimization scheduling method for power systems as described in the first aspect and various possible designs of the first aspect.

[0068] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions. When a processor executes the computer-executable instructions, it implements the AI-driven dynamic frequency stability optimization scheduling method for power systems as described in the first aspect and various possible designs of the first aspect.

[0069] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the AI-driven dynamic frequency stability optimization scheduling method for power systems as described in the first aspect and various possible designs of the first aspect.

[0070] The AI-driven dynamic frequency stability optimization scheduling method, apparatus, equipment, and storage medium for power systems provided in this application have at least the following beneficial effects:

[0071] This application constructs a reinforcement learning environment integrating power flow calculation, time-domain simulation, and reward calculation to accurately model the frequency dynamic response process and scheduling constraints. By combining deep reinforcement learning algorithms to train agents to generate optimized scheduling strategies, it can perform strategy optimization training under various disturbance conditions to obtain intelligent control strategies that adapt to the actual operating state of the system. This method can generate optimized scheduling instructions that take into account both dynamic frequency stability and economy based on the real-time operating state of the power system, realize rapid adjustment and stable control of the frequency response process, effectively improve the power system's ability to resist disturbances under the background of high proportion of new energy access, and reduce the risk of frequency instability and system collapse. Attached Figure Description

[0072] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0073] Figure 1 A flowchart illustrating an AI-driven dynamic frequency stability optimization scheduling method for power systems, provided as an embodiment of this application;

[0074] Figure 2 A flowchart illustrating the construction of a power system dynamic frequency stability reinforcement learning environment provided in this application embodiment;

[0075] Figure 3 A flowchart for determining deep reinforcement learning elements provided in the embodiments of this application;

[0076] Figure 4 A flowchart of agent training provided for embodiments of this application;

[0077] Figure 5 A flowchart illustrating the implementation of the optimization strategy provided in this application embodiment;

[0078] Figure 6 A flowchart illustrating the interaction between the DRL agent and the reinforcement learning environment provided in this embodiment of the application;

[0079] Figure 7 A structural diagram of the AI-driven dynamic frequency stability optimization scheduling device for power systems provided in this application embodiment.

[0080] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0081] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0082] The collection, storage, use, processing, transmission, provision, and disclosure of financial data or user data involved in the technical solution of this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0083] It should be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0084] The technical solution of this application and how it solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0085] To address the shortcomings of existing technologies in generating dynamic frequency control strategies, this application provides an AI-driven method for optimizing dynamic frequency stability scheduling of power systems. This optimization scheduling method includes:

[0086] 1. Constructing a reinforcement learning environment: Establishing a reinforcement learning environment that integrates power flow calculation, time-domain simulation, and reward calculation to accurately model the frequency dynamic response process and scheduling constraints;

[0087] 2. Define the key elements of deep reinforcement learning: including state space, action space, and reward function;

[0088] 3. Agent Training: Train the agent using a DRL algorithm (such as DDPG); maximize cumulative rewards through environmental interactions;

[0089] 4. Strategy Deployment and Application: Deploy the trained agents to the real-time scheduling system, and dynamically output generator output adjustment, load switching, or energy storage control commands according to the system status.

[0090] This method maps the operating state of the power system to optimized scheduling actions using a state-action-reward framework, with the training objective being to maximize long-term cumulative rewards. The reinforcement learning environment is built upon the system's dynamic frequency response characteristics, providing a unified simulation platform for the evaluation and optimization of scheduling strategies.

[0091] Specifically, such as Figure 1 The diagram shows a flowchart of an AI-driven dynamic frequency stability optimization scheduling method for power systems provided in an embodiment of this application. The AI-driven dynamic frequency stability optimization scheduling method for power systems includes the following steps S10-S40.

[0092] S10: Construct a reinforcement learning environment for dynamic frequency stability of power systems.

[0093] In this embodiment, the dynamic frequency stability reinforcement learning environment of the power system can simulate the dynamic frequency response of the system under disturbance, and is used to support the training and policy evaluation of reinforcement learning agents.

[0094] In some embodiments, such as Figure 2 The diagram shown is a flowchart of the construction of a power system dynamic frequency stability reinforcement learning environment provided in this application embodiment. Step S10 specifically includes the following steps S101-S104.

[0095] S101: Input of agent scheduling instructions: The reinforcement learning environment receives the scheduling control instructions output by the agent based on the current system state as the input of the environment.

[0096] S102: Power Flow Calculation: Input the environmental input and the current system operating status into the power flow calculation module, and obtain updated node voltage, line power and other power flow information through power flow calculation.

[0097] S103: Time-domain dynamic simulation: Using power flow calculation results as input, the power system dynamic model is driven to perform time-domain simulation to simulate the frequency dynamic response process of the system under disturbance conditions.

[0098] The system frequency dynamic response process is modeled based on a set of differential-algebraic equations, which includes the generator swing equation, the speed regulation system dynamic response equation, and the node power balance equation.

[0099] The swing equation for a synchronous generator is:

[0100]

[0101] Where, δ i Let ω be the generator power angle of node i. i Let ω be the rotor angular velocity of the generator connected to node i. ref H is the rated rotor angular velocity of the generator connected to node i.i Let P be the inertial time constant of the generator at node i. m,i P represents the mechanical power of the generator prime mover connected to node i. e,i E represents the electromagnetic power generated by the generator at node i. i The potential of the generator. This represents the differential power angle variable of the generator connected to node i. Let represent the differential variable of the angular velocity of the generator at node i.

[0102] The response equation of the generator speed control system is:

[0103]

[0104] P m,i =(1-F HP,i )z g3,i +F HP,i z g2,i ,i=1,...,m;

[0105] Among them, T G,i T CH,i T RH,i z represents the time constant of each stage. g1,i z g2,i z g3,i K represents the output variable of each integral stage. G P represents the unit regulating power of the generator. m0,i To affect the mechanical power of the generator before the disturbance, F HP,i P represents the percentage of steam in the high-pressure cylinder relative to the total steam volume. m,i This represents the mechanical power of the generator connected to node i.

[0106] The node power balance equation is:

[0107]

[0108] S104: Based on the evolution of key performance indicators during the simulation, including system frequency deviation, frequency change rate, generator output change, etc., calculate the reward value corresponding to the current moment and feed it back to update the reinforcement learning strategy.

[0109] S20: Define the elements of deep reinforcement learning, including state space, action space, and reward function.

[0110] In this embodiment, the state space includes measurement data reflecting the operating status and dynamic frequency stability of the power system, the action space includes dispatch control commands for adjustable generation resources or controllable loads, and the reward function includes a composite function that integrates dynamic frequency stability index and system economic index.

[0111] In some embodiments, such as Figure 3 The diagram shown is a flowchart of the deep reinforcement learning element determination process provided in this application embodiment. Step S20 specifically includes the following steps S201-S203.

[0112] S201: Define the state space of the power system.

[0113] In this embodiment, the power system state space is represented as follows:

[0114] s t ={P g,i (t), Q g,i (t), V(t), P d,i (t), Q d,i (t)}

[0115] Where P g,i Q(t) represents the active power output of the i-th generator at time step t. g,i P(t) represents the reactive power output of the i-th generator at time step t, V(t) represents the voltage amplitude at the node at time step t, and P(t) represents the reactive power output of the i-th generator at time step t. d,i (t) represents the active load of the i-th load node at time step t, Q d,i (t) represents the reactive load of the i-th load node at time step t, s t As input to the agent;

[0116] S202: Define the action space of the intelligent agent.

[0117] In this embodiment, the action space a of the intelligent agent t Represented as:

[0118] a t ={ΔP g,i V g}

[0119] Where ΔP g,i V represents the active power adjustment of the generator excluding the balancing node. g This represents the terminal voltage of all generators.

[0120] S203: Define the reward function.

[0121] In this embodiment, the reward function R t Represented as:

[0122]

[0123] Where max_penalty represents the maximum penalty value set by the power flow divergence system caused by the agent's scheduling instructions, λ dyndyn_violation represents the penalty value for violating constraints during the system's dynamic frequency response, λ. sts sts_violation represents the penalty value for violations of runtime and technical constraints. This represents the reward value when there is no constraint violation.

[0124] The constraints that should be followed during the system's dynamic frequency response process include the minimum frequency after disturbance, the rate of frequency change, the increased power output, the primary frequency regulation reserve capacity, and the system scheduling cost, specifically including:

[0125] The increased power after the disturbance is set to meet the following requirements:

[0126] P m,i -P g,i ≥0, i=1,...,m

[0127] Among them, P m,i This represents the mechanical power of the generator connected to node i.

[0128] The generator frequency regulation standby after a disturbance is set to meet the following requirements:

[0129] P m,i -P g,i ≤P rmax,i i = 1, ..., m

[0130] Among them, P rmax,i This is the maximum value of the primary frequency regulation reserve for the generator at access node i.

[0131] To ensure the frequency stability of the power system, a minimum frequency constraint is set, expressed as:

[0132]

[0133] Among them, f lim ω is the minimum frequency allowed for frequency stability requirements. i This represents the angular frequency of the generator connected to node i.

[0134] When the frequency decreases beyond a certain upper limit, the relevant protection will activate, causing system power loss. Therefore, it is essential to constrain the generator's frequency change rate, as expressed in the following expression:

[0135]

[0136] in, The maximum allowable rate of change of frequency is less than the operating threshold value of the relevant protection.

[0137] During the system's dynamic frequency response, the penalty value for violating constraints is determined by λ. dynComposed of dyn_violation, λ dyn Here, is the set penalty coefficient, and dyn_violation is the degree of constraint violation, expressed as follows:

[0138] dyn_violation=v flim +v RoCoF +v mg +v pr

[0139] Among them, v flim f represents the minimum frequency during the response process. nadir The degree to which the frequency falls below the lower frequency limit constraint is expressed as:

[0140] v flim =min(0,f nadir -f lim )v RoCoF The expression representing the degree to which the rate of change of frequency violates the specified value during the response process is:

[0141]

[0142] v mg The expression representing the degree of violation of the power constraint during the response process is:

[0143] v mg =min(0,P) m,i -P g,i )v pr The expression representing the degree of violation of the maximum value of primary frequency regulation reserve during the response process is:

[0144] v pr =min(0,P) rmax,i -P m,i -P g,i ))

[0145] Operational and technical constraints are used to evaluate whether the dispatch instructions of the intelligent agent conform to the actual operation of the power system. Specifically, under the normal operating conditions before the disturbance occurs, the node voltage amplitude, generator output, and line current should meet the following requirements:

[0146] V min,i ≤V i ≤V max,i i = 1, ..., n

[0147] P gmin,i ≤P g,i ≤P gmax,i i = 1, ..., m

[0148] Q gmin,i ≤Q g,i ≤Qgmax,i i = 1, ..., m

[0149] Among them, V i Let P be the voltage magnitude at node i. gmin,i P gmax,i These are the upper and lower limits of the active power of the generator at node i, respectively, and Q. gmin,i Q gmax,i These are the upper and lower limits of the reactive power of the generator at access node i, respectively.

[0150] During the inspection of operational and technical constraints, the penalty value for violating the constraints is determined by λ. sts Composed of sts_violation, λ sts Here, `sts_violation` is the set penalty coefficient, and `sts_violation` is the degree of constraint violation, expressed as follows:

[0151]

[0152] in, Indicates the degree of violation of node voltage. This indicates the degree of violation of the generator's active power output. The degree of violation of generator reactive power output is expressed as follows:

[0153]

[0154] Among them, V max,i and V min,i P represents the maximum and minimum values ​​of the node voltage, respectively. gmax,i and P gmin,i Q represents the maximum and minimum active power of the generator at the node, respectively. gmax,i and Q gmin,i These represent the maximum and minimum reactive power of the generators at the nodes, where max is the maximum value function, and V... i Let P be the voltage at node i. g,i Q represents the active power of the generator at the node. g,i Let be the reactive power of the generator at node i.

[0155] S30: In a reinforcement learning environment, based on the state space, action space and reward function, a deep reinforcement learning algorithm is used to train the agent, enabling the agent to learn and optimize the scheduling strategy.

[0156] In this embodiment, the agent learns and optimizes the scheduling strategy to select actions to maximize cumulative rewards.

[0157] In some embodiments, the deep reinforcement learning algorithm used in step S30 is applicable to a continuous action space, including but not limited to policy gradient algorithms, value function-based methods, policy optimization-based methods, or algorithms with an actor-critic structure. This algorithm is used to train the agent to learn an optimized scheduling policy.

[0158] In some embodiments, the DDPG algorithm is selected to train the agent. For example... Figure 4 The diagram shown is a flowchart of the intelligent agent training process provided in an embodiment of this application. Step S30 specifically includes the following steps S301-S303.

[0159] S301: Initialize network structure parameters.

[0160] In the actor network, the initialization involves neural network parameters including a state input layer, multiple fully connected hidden layers, and an output layer, where the input data is the power system state set s. t The output data is in the current state s t The following action a is generated t The number of nodes in the input layer is equal to the state space dimension of the power system, dim. state The number of nodes in the output layer is equal to the action space dimension of the power system, dim. action .

[0161] In the critic network, the initialization includes neural network parameters for a state input layer, an action input layer, multiple fully connected hidden layers, and an output layer. The input data consists of the power system state and the set of output actions {s}. t ,a t The output layer has 1 node, and the output data is the Q-value Q(s) of the combination of the corresponding state and output action. t ,a t ):

[0162]

[0163] Where Q(s) t ,a t ) represents state s t Take action a t Q value; R t Represents state s t Take action a t The immediate reward; γ is a discount factor used to balance current and future rewards; s t+1 Indicates taking action a t The state after which it transitions to the next state; a t+1 s t+1 The action selected in the current state; Indicates the next state s t+1Choose the Q value for the optimal action.

[0164] S302: Construction of the experience replay pool.

[0165] Set the storage capacity of the experience playback buffer, and store the collected data in chronological order using e t =(s t ,a t ,R t ,s t+1 The data is stored in a buffer in the format of () to construct an experience replay buffer. An experience replay buffer for DDPG is constructed for storing and retrieving samples during training. By constructing the experience replay buffer, the correlation of data can be reduced and the training efficiency of the model can be improved.

[0166] S303: Execute the training process.

[0167] Five independent replicas of the reinforcement learning environment were launched on the training server. Each replica loaded the same power grid topology parameters but used different initial running points and disturbance scenario sequences.

[0168] After receiving the action commands from the agent, each environment replica synchronously executes power flow calculations and dynamic frequency simulations, and processes each interaction data according to the transition tuple e. t =(s t ,a t ,R t ,s t+1 The format is stored in the buffer;

[0169] After accumulating 512 interaction data points, the central training module (Learner) randomly samples a mini-batch of 128 data points from the experience replay pool and performs the following network update operations: Every 512 sets of transition data accumulated, the central trainer (Learner) randomly samples a batch of samples from the experience pool and performs the following operations:

[0170] The objective of calculating the temporal difference (TD) of the Critic network is:

[0171] y = r + γ·Q target (s t+1 ,μ target (s t+1 ))

[0172] Where y is the TD target value; r is the value in state s t Perform action a t Then, the immediate reward returned by the environment; γ is a discount factor used to measure the current value of future rewards; Q targetFor the Target Critic network, the next state-action pair (s) is generated. t+1 ,a t+1 The Q-value estimate of μ is used to construct the TD target y; target For the Target Actor network in the next state s t+1 Deterministic action a of the output t+1 =μ target (s t+1 ).

[0173] Update Critic network parameters:

[0174]

[0175] Where L is the mean squared error loss function of the Critic network, and N is the number of mini-batch samples used in one update.

[0176] Update Actor network parameters:

[0177]

[0178] in, Let the policy objective function J be the parameter θ of the Actor network. μ The gradient is used to guide the parameter update direction of the Actor network; μ(s) is the gradient of the Actor network in state s. t The deterministic action output is the action given by the current policy. This is the gradient of the Critic network with respect to action 'a', used to measure the impact of the action on the Q-value in the current state; this gradient is multiplied by... The parameters of the Actor network are updated using a chain rule, so that the actions output by the network are conducive to improving the corresponding Q-value.

[0179] Soft update target network:

[0180] θ target ←τθ+(1-τ)θ target (τ=0.01)

[0181] Where, θ target Let τ be the target network parameters, τ be the soft update coefficient, and θ be the current network parameters.

[0182] S40: The trained agent is deployed to the real-time dispatch process of the power system to generate dispatch instructions based on the real-time status input, thereby achieving joint optimization of the system's dynamic frequency stability and economy.

[0183] In some embodiments, in step S40, the trained agent is deployed on the power system energy management system or a dedicated control server, and the state space information is obtained in real time through the SCADA system or PMU system.

[0184] In some embodiments, a security check is performed in step S40 before or after the output of the scheduling control instruction to verify the feasibility and security of the generated scheduling instruction under system constraints.

[0185] In some embodiments, such as Figure 5 The diagram shown is a flowchart illustrating the implementation of the optimization strategy provided in this application embodiment. Step S40 specifically includes the following steps S401-S405.

[0186] S401: Obtain the real-time status of the power system. t ;s t ={P g,i (t), Q g,i (t), V(t), P d,i (t), Q d,i (t)}.

[0187] S402: Input the current state into the actor network of the already trained DDPG model to generate scheduling action α. t ;

[0188] S403: Apply the generated scheduling actions to the power system to implement optimized control strategies.

[0189] S404: Real-time monitoring of the power system status response after implementation, including dynamic frequency stability indicators such as voltage and frequency, and using the response information to update the real-time status representation in the actor network to maintain the accuracy and generalization capability of the network.

[0190] S405: If the system response is as expected, continue with the current strategy; if the system response is not ideal, add the state information to the experience replay buffer for subsequent offline training and improvement of the control strategy.

[0191] In some embodiments, such as Figure 6 The diagram shows the interaction flowchart between the DRL agent and the reinforcement learning environment provided in the embodiment of the application. As shown in the figure, the proposed power system dynamic frequency stability optimization scheduling method based on deep reinforcement learning includes two main components: a reinforcement learning environment for simulating the dynamic frequency stability of the power system, and an agent for generating optimized scheduling strategies. These two components form a closed-loop interactive system, continuously training and optimizing the scheduling strategy.

[0192] After receiving state information (state s) and reward information (reward r) from the environment, the agent first improves the Critic network: based on the current policy and reward result, the Critic network evaluates the performance of the current policy and updates it; then, it improves the Actor network: the updated Critic network guides the updating of the Actor network; next, it generates actions: the improved Actor network generates actions (i.e., scheduling policies) based on the current state; this action is then sent to the environment module to trigger a new round of scheduling and simulation.

[0193] After receiving the action, the environment first uses the power flow calculation module to determine the system power distribution under the action, and then uses the time-domain simulation module to analyze the system's frequency dynamic response under disturbance conditions. After the time-domain simulation is completed, the reward calculation module calculates the reward value r based on the simulation results, and then the environment returns the new state and reward to the agent. Through repeated interactions with the environment, the agent continuously optimizes the policy network, enabling the generated actions to maximize the long-term cumulative reward, thereby achieving joint optimization of system frequency stability and economy.

[0194] in:

[0195] "Current state" and "current action" are the core inputs and outputs for the agent's perception and response;

[0196] "Power flow calculation" is used to handle power balance and constraint verification;

[0197] "Time-domain simulation" provides a characterization of frequency dynamics after a disturbance;

[0198] The "Reward Calculation" generates real-time feedback based on frequency deviation, RoCoF, economic indicators, and other factors.

[0199] "State update" then sends the system characteristics back to the agent after the simulation ends.

[0200] Through the above cycle, the agent continuously improves its scheduling decision-making capabilities, ultimately forming an optimized strategy that can be deployed in actual power system scheduling scenarios.

[0201] This application also provides an AI-driven dynamic frequency stability optimization scheduling device for power systems, such as... Figure 7 As shown, the AI-driven dynamic frequency stability optimization scheduling device for power systems includes:

[0202] The reinforcement learning environment construction module 701 is configured to build a simulation-based reinforcement learning environment, which integrates a power flow calculation unit, a time-domain simulation unit, and a reward calculation unit; the time-domain simulation unit is based on the dynamic frequency response model of the power system and is used to simulate the dynamic frequency trajectory of the system under disturbance.

[0203] The learning element determination module 702 is configured to determine deep reinforcement learning elements; wherein, the deep reinforcement learning elements include a state space, an action space, and a reward function, the state space includes measurement data reflecting the operating state and dynamic frequency stability of the power system, the action space includes dispatch control instructions for adjustable generation resources or controllable loads, and the reward function includes a composite function that integrates dynamic frequency stability indicators and system economic indicators.

[0204] The agent training module 703 is configured to train the agent using a deep reinforcement learning algorithm based on the state space, action space, and reward function in the reinforcement learning environment, so that the agent learns an optimized scheduling strategy; wherein the optimized scheduling strategy is used to select actions to maximize the cumulative reward.

[0205] The optimized scheduling module 704 is configured to deploy the trained agent to the real-time scheduling process of the power system, generate scheduling instructions based on the real-time status input, and achieve joint optimization of the system's dynamic frequency stability and economy.

[0206] This application provides an electronic device. The electronic device may include a processor and a memory, wherein the processor and the memory can communicate; exemplarily, the processor and the memory communicate via a communication bus.

[0207] The processor executes computer execution instructions stored in memory, causing the processor to perform the scheme in the above embodiments. The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0208] The communication bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The system bus can be divided into address bus, data bus, control bus, etc. Transceivers are used to enable communication between database access devices and other computers (e.g., clients, read-write libraries, and read-only libraries). Memory may include random access memory (RAM) and may also include non-volatile memory.

[0209] The electronic device provided in this application embodiment can be the terminal device described in the above embodiments.

[0210] This application also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on a computer, the computer performs the technical solution of the AI-driven power system dynamic frequency stability optimization scheduling method described in the above embodiments.

[0211] This application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, it can implement the technical solution of the AI-driven power system dynamic frequency stability optimization scheduling method in the above embodiments.

[0212] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.

[0213] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.

[0214] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.

[0215] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.

[0216] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0217] The memory may include high-speed RAM, and may also include non-volatile storage (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.

[0218] Buses can be Industry Standard Architecture (ISA) buses, Peripheral Component Interconnect (PCI) buses, or Extended Industry Standard Architecture (EISA) buses, etc. Buses can be categorized into address buses, data buses, control buses, etc.

[0219] The aforementioned storage medium can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium accessible to general-purpose or special-purpose computers.

[0220] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. The processor and storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic control unit or main control device.

[0221] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. An AI-driven dynamic frequency stability optimization scheduling method for power systems, characterized in that, The method includes: A reinforcement learning environment is constructed; wherein the environment integrates power flow calculation, time-domain simulation and reward calculation; the time-domain simulation is based on the dynamic frequency stability model of the power system and is used to simulate the dynamic trajectory of the system frequency under disturbance; The elements of deep reinforcement learning are determined; wherein, the elements of deep reinforcement learning include a state space, an action space, and a reward function, the state space includes measurement data reflecting the operating state and dynamic frequency stability of the power system, the action space includes dispatch control instructions for adjustable generation resources or controllable loads, and the reward function is a composite function that integrates dynamic frequency stability indicators and system economic indicators, and is calculated in real time by the reward based on the system state. In the environment described above, based on the state space, action space, and reward function, a deep reinforcement learning algorithm is used to train the agent, enabling the agent to learn an optimized scheduling strategy; wherein, the optimized scheduling strategy is used to select actions to maximize the cumulative reward. The trained agent is deployed to the real-time dispatching process of the power system to generate dispatching instructions based on real-time status input, thereby achieving joint optimization of the system's dynamic frequency stability and economy. The reward function Represented as: in, This indicates that the agent's scheduling instructions result in the maximum penalty value set by the power flow divergence system. This represents the penalty value for violating constraints during the system's dynamic frequency response. This represents the penalty value for violations of operational and technical constraints. This represents the reward value when there is no constraint violation; During the system's dynamic frequency response process, the penalty value for violating constraints is determined by... constitute, The set penalty coefficient, The degree of constraint violation is expressed as follows: in, Represents the minimum frequency during the response process. The degree to which the frequency falls below the lower frequency limit constraint is expressed as: in, The minimum frequency allowed for frequency stability requirements; The expression representing the degree to which the rate of change of frequency violates the specified value during the response process is: in, For the maximum allowable rate of frequency change, Access Node The change in the angular frequency of the generator, It is a quantity that changes over time; The expression representing the degree of violation of the power constraint during the response process is: in, For access nodes The mechanical power of the generator; For access nodes The active power of the generator; The expression representing the degree of violation of the maximum value of primary frequency regulation reserve during the response process is: in, For access nodes The maximum value of the primary frequency regulation reserve of the generator; During the inspection of operational and technical constraints, the penalty value for violating the constraints is determined by... constitute, The set penalty coefficient, The degree of constraint violation is expressed as follows: in, Indicates the degree of violation of node voltage. This indicates the degree of violation of the generator's active power output. The degree of violation of generator reactive power output is expressed as follows: in, It is a function for maximizing the value; For generator access node voltage amplitude, V min,i , V max,i They are nodes The upper and lower limits of the voltage amplitude, 、 Access nodes The upper and lower limits of the active power of the generator. Q g,i For access nodes The reactive power of the generator , Access nodes The upper and lower limits of the reactive power of the generator.

2. The AI-driven dynamic frequency stability optimization scheduling method for power systems according to claim 1, characterized in that, The time-domain simulation is constructed using the system frequency dynamic response equation, which includes the generator swing equation, the speed regulation system dynamic response equation, and the node power balance equation. The specific methods for constructing the time-domain simulation include: A swing equation for a synchronous generator is established, which is constrained by electromagnetic power, generator angular frequency, generator power angle, and generator inertial time constant. The expression is as follows: in, For access nodes The generator power angle, For access nodes The rotor angular velocity of the generator, For access nodes The rated rotor angular velocity of the generator, For access nodes The inertial time constant of the generator, The electromagnetic power generated by the generator, The potential of the generator. Indicates access node The differential variable of the generator's power angle, Indicates access node The differential variable of the generator's angular velocity, m This refers to the sequence number of the generator connection node. For generator access node voltage phase angle, For the generator transient reactance, E i The generator's electromotive force; Based on the speed governor's parameter data, the response equation of the generator speed control system is established, and the expression is: in, This represents the time constant for each stage. This represents the output variables of each integral step. For the unit regulating power of the generator, To reduce the mechanical power of the generator before the disturbance, The percentage of steam in the high-pressure cylinder relative to the total steam volume. t For time variables, dt Indicates time t The derivative; The nodal power balance equations are established as follows: in, This represents the (i,j) element in the extended network node admittance matrix that takes into account the generator potential. Let (i,j) be the (i,j) element of the network node admittance phase angle matrix. For the nodes after the disturbance voltage amplitude, For the nodes after the disturbance voltage phase angle, To represent the j-th element of the extended network node voltage vector that takes into account the generator potential, To represent the j-element of the phase angle vector of the extended network nodes taking into account the generator electromotive force, For the perturbation of the previous node Active load, For the perturbation of the previous node reactive load, For nodes Active power lost due to sudden disconnection of new energy sources from the grid. n This represents the number of system nodes.

3. The AI-driven dynamic frequency stability optimization scheduling method for power systems according to claim 1, characterized in that, The state space includes at least one of the following: Generator active power output, generator reactive power output, load node active power, load node reactive power, voltage amplitude, system frequency, critical node frequency, generator speed or power angle, critical line power flow, new energy output, and frequency change rate signal.

4. The AI-driven dynamic frequency stability optimization scheduling method for power systems according to claim 1, characterized in that, The action space includes at least one of the following: Traditional generator sets include active power adjustment commands, generator terminal voltage regulation commands, and energy storage device charging and discharging control commands.

5. The AI-driven dynamic frequency stability optimization scheduling method for power systems according to claim 1, characterized in that, The reward function considers indicators such as equipment safety, dynamic frequency stability, and system economy, and specifically includes the following constraints and reward quantification mechanisms: Under normal operating conditions prior to the disturbance, the node voltage amplitude, generator output, and line current should meet the following requirements: The increased power generation and primary frequency regulation reserve capacity of the generator after the disturbance are set to meet the following requirements: Set the minimum frequency constraint and the rate of change of frequency constraint, with the following expressions: in, Indicates access node The angular frequency of the generator.

6. The AI-driven dynamic frequency stability optimization scheduling method for power systems according to claim 1 or 5, characterized in that, The reward function considers system economic indicators, specifically including: Based on the optimal power flow model, the active power output of generators is optimized and rescheduled. The scheduling objective is to minimize the generator output adjustment cost, expressed as: min in, For access nodes The unit cost of increasing active power output of generators For access nodes The increased active power of the generator For access nodes Generators reduce the unit cost of active power output. For access nodes The reduced active power of the generator This represents the number of generators.

7. The AI-driven dynamic frequency stability optimization scheduling method for power systems according to claim 2, characterized in that, The constructed time-domain simulation module includes a set of initial condition equations used to set the initial operating states of various components in the power system; wherein, the initial condition equations include: In the formula, To change the rotor angular velocity of the synchronous generator before the disturbance, , and Let these be the initial values ​​of the integral variables before the perturbation. E i For the generator's electromotive force, For access nodes The active power output of the generator before the disturbance. Q g,i For access nodes The reactive power generated by the generator before the disturbance. The power angle of the synchronous generator before the disturbance.

8. An AI-driven dynamic frequency stability optimization scheduling device for power systems, used to implement the method as described in any one of claims 1-7, characterized in that, The device includes: The reinforcement learning environment construction module is configured to build a simulation-based reinforcement learning environment, which integrates a power flow calculation unit, a time-domain simulation unit, and a reward calculation unit; wherein, the time-domain simulation unit is based on the dynamic frequency response model of the power system and is used to simulate the dynamic frequency trajectory of the system under disturbance. The learning element determination module is configured to determine deep reinforcement learning elements; wherein, the deep reinforcement learning elements include a state space, an action space, and a reward function, the state space includes measurement data reflecting the operating state and dynamic frequency stability of the power system, the action space includes dispatch control commands for adjustable generation resources or controllable loads, and the reward function is a composite function that integrates dynamic frequency stability indicators and system economic indicators, and is calculated in real time by the reward calculation unit based on the simulation output state; The agent training module is configured to train the agent using a deep reinforcement learning algorithm based on the state and reward data output by the reinforcement learning environment; enabling the agent to learn and optimize scheduling strategies; wherein the agent maximizes cumulative rewards through action selection; The optimized scheduling module is configured to deploy the trained agent to the real-time scheduling process of the power system, generate scheduling instructions based on the real-time status input, and achieve joint optimization of the system's dynamic frequency stability and economy. The output of the reinforcement learning environment construction module is connected to the input of the agent training module; the input of the optimization scheduling module is connected to the real-time status data stream of the power system.

9. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes the computer execution instructions stored in the memory to implement the AI-driven dynamic frequency stability optimization scheduling method for power systems as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the AI-driven dynamic frequency stability optimization scheduling method for power systems as described in any one of claims 1-7.

Citation Information

Patent Citations

  • AGC unit dynamic optimization method based on deep reinforcement learning

    CN112186811A