Power grid stability risk intelligent boundary decision-making method based on deep reinforcement learning

The intelligent boundary decision-making method for power grid stability risk, which utilizes deep reinforcement learning and employs a hybrid reward function to guide model training, achieves rapid and adaptive decision-making for power grid stability boundaries. This solves the problems of computational conservatism and insufficient real-time performance of traditional methods, thereby improving the economic benefits and efficiency of power grid operation.

CN120933900APending Publication Date: 2025-11-11CHINA SOUTHERN POWER GRID NEW POWER SYSTEM (BEIJING) RESEARCH INSTITUTE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510866698.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional methods for determining the stability boundary of a power grid yield conservative results and lack real-time adaptability, failing to meet the needs of modern power grids for refined and intelligent operation. In particular, they impose an excessive computational burden on large-scale power grids and high-proportion renewable energy grid integration, making them difficult to apply online.

Method used

A smart boundary decision-making method for power grid stability risk based on deep reinforcement learning is constructed. The method uses a deep neural network model with an actor-commentator architecture, combined with a hybrid reward function that includes boundary proximity reward, instability risk penalty and adjustment economic cost. The method is trained offline, and complex simulations are performed offline. In the online stage, only the actor network is deployed for real-time decision-making.

Benefits of technology

It achieves millisecond-level rapid determination of power grid stability risk boundaries, possesses high real-time performance and adaptive capabilities, dynamically adjusts stability boundaries, and improves the utilization efficiency and operational economy of power grid assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120933900A_ABST
    Figure CN120933900A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid stability risk intelligent boundary decision-making method based on deep reinforcement learning, and belongs to the technical field of power system control. According to the method, an actor-reviewer model is constructed, training is carried out through interaction with a power grid simulation environment in an offline stage, and a mixed reward function fusing boundary proximity reward, instability risk penalty and economic cost is adopted in training to optimize a decision strategy. During online application, the optimal decision action vector can be quickly output only by deploying the training convergence lightweight actor network and receiving the real-time power grid state vector, so that the intelligent stability boundary considering safety and economy is determined. According to the method, decoupling of the decision process and complex simulation is realized, the method has the advantages of high decision speed and high adaptability, and the operation economic benefit of the power grid can be remarkably improved on the premise of guaranteeing the system safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system control, and specifically relates to a smart boundary decision method for power grid stability risk based on deep reinforcement learning. Background Technology

[0002] The safe and stable operation of the power grid is the cornerstone of the entire power system, and determining the stability boundary of the power grid is a crucial step in ensuring its safe operation. The stability boundary defines the maximum operating limit that the system can withstand without the risk of instability. Traditional methods for determining the power grid stability boundary mainly rely on offline simulation analysis based on mechanistic models, such as time-domain simulation. These methods typically require a large number of time-consuming simulation calculations for a pre-set, finite set of faults, resulting in a set of static, fixed operating rules or limit sections.

[0003] However, these traditional methods have significant drawbacks. First, their calculations are typically conservative because, to ensure absolute safety under all anticipated contingencies, the stability margin is set too large, resulting in the underutilization of the grid's actual transmission potential and sacrificing operational economics. Second, these methods lack real-time adaptability, failing to dynamically adjust stability boundaries according to the rapidly changing operating conditions of the grid, making it difficult to meet the online decision-making requirements of modern grids' refined and intelligent operation. Furthermore, with the ever-expanding scale of power grids and the integration of a high proportion of renewable energy, the dynamic characteristics of power systems have become unprecedentedly complex, causing the computational burden of traditional simulation methods to increase exponentially, facing the "curse of dimensionality," which further limits their feasibility for online application. Therefore, there is an urgent need for an intelligent decision-making method that can quickly, accurately, and adaptively assess grid stability risks under current operating conditions and provide optimal operating boundaries online. Summary of the Invention

[0004] To address the aforementioned problems in existing technologies, the present invention aims to provide a smart boundary decision-making method for power grid stability risk based on deep reinforcement learning, comprising the following steps:

[0005] Step S1: Construct a power grid state vector by acquiring real-time operation data and topology data of the power grid and processing them into a state vector that represents the current operating condition of the power grid.

[0006] Step S2: Construct a deep reinforcement learning decision model. Construct a deep reinforcement learning decision model based on an actor-commentator architecture. The model includes an actor network for generating decision action vectors based on the state vectors, and a commentator network for evaluating the values ​​of the state vectors and the decision action vectors.

[0007] Step S3: Calculate the hybrid reward function, execute the decision action vector in the power grid simulation environment, and perform transient stability simulation on the power grid state after execution. Based on the simulation results and the adjustment amount of the decision action vector, calculate a hybrid reward value that integrates boundary proximity reward, instability risk penalty and adjustment economic cost.

[0008] Step S4: Offline model training. Using the state vector, decision action vector, and the hybrid reward value, the deep reinforcement learning decision model is trained offline in the power grid simulation environment until the model converges.

[0009] Step S5, online boundary decision-making: The actor network that has been trained and converged is deployed in the power grid control system. The actor network receives the state vector generated in real time and outputs the optimal decision action vector. The running point defined by the decision action vector is the intelligent boundary of power grid stability risk under the current operating condition.

[0010] Furthermore, in step S1, the real-time operating data of the power grid acquired includes node voltage amplitude, node voltage phase angle, branch active power, branch reactive power, generator active output, generator reactive output, and system frequency; the construction of the state vector also includes normalizing the real-time operating data.

[0011] Furthermore, in step S2, the input of the actor network is the state vector, and the output is the decision action vector; the input of the commenter network is the state vector and the decision action vector output by the actor network, and the output is a scalar evaluation value that assesses the value of the decision action.

[0012] Furthermore, in step S3, the process of calculating the mixed reward value includes:

[0013] Step S301: Calculate the boundary proximity reward based on the adjustment amount of the decision action vector.

[0014] Step S302: Determine whether the power grid is unstable based on the transient stability simulation results. If it is unstable, apply a preset negative instability penalty.

[0015] Step S303: Calculate the adjustment economic cost based on the magnitude of the decision action vector.

[0016] Step S304: The boundary proximity reward, the instability penalty, and the adjustment economic cost are weighted and summed to obtain the final mixed reward value.

[0017] Furthermore, in step S302, the process of determining whether the power grid is unstable involves comparing the dynamic trajectory data output from the transient stability simulation with preset instability criteria to generate a clear binary penalty signal for deep reinforcement learning training. The instability criteria include at least one of the following: the relative power angle difference of at least one generator exceeds a preset power angle instability threshold during the simulation period; the voltage of at least one critical bus remains below a preset voltage safety lower limit threshold for a period exceeding a preset duration threshold after fault clearance; and the absolute value of the system frequency deviating from the nominal frequency remains above a preset frequency offset threshold after fault clearance. The power angle instability threshold, the voltage safety lower limit threshold, the duration threshold, and the frequency offset threshold are all configurable parameters set before the offline training of the model begins, based on the operating procedures and safety requirements of the target power grid.

[0018] Furthermore, the mixed reward value R(s) t ,a t R(s) is calculated using the following formula: t ,a t ) = w prox ×f(P k,t+1 )-w pen ×I(M t+1 <M min )-w econ ×C(a t ), where: s t The state vector at the current moment; a t f(P) represents the decision action vector at the current moment; k,t+1 () represents the target power P of the critical transmission section after the action is performed. k,t+1 The associated boundary proximity reward; I(M) t+1 <M min ) represents the instability risk penalty term, where M t+1 M is the stability margin index calculated through transient stability simulation after the action is performed. min This represents the preset minimum safety margin threshold, and I(×) represents the indicator function, when M t+1 <M min Its value is 1 when it is active, and 0 otherwise; C(a) t ) indicates the economic cost of the adjustment performed; w prox w pen w econ These are the weighting coefficients for the boundary proximity reward, the instability risk penalty, and the adjustment economic cost, respectively, and w pen The value is greater than w prox and w econ The value of .

[0019] Furthermore, the offline training of the model in step S4 includes:

[0020] Step S401: Store the experience data, including state vectors, decision action vectors, mixed reward values, and new state vectors generated by the interaction between the agent and the power grid simulation environment, into an experience playback buffer.

[0021] Step S402: Randomly extract data samples from the experience replay buffer and calculate the loss function based on the samples to update the network parameters of the actor network and the commentator network.

[0022] Furthermore, the online boundary decision in step S5 decouples the decision-making process from the complex simulation. In the online stage, only the trained and converged actor network is deployed, without deploying the commentator network and the power grid simulation environment. The actor network, as a policy function that implicitly incorporates the power grid stability mechanism, directly maps the real-time state vector to the optimal decision action vector through a single forward propagation calculation, thus completing the determination of the stability boundary.

[0023] Furthermore, the decision action vector is a multi-dimensional continuous vector, each dimension of which uniquely corresponds to the adjustment amount of a preset key control variable in the power grid, thereby realizing multi-variable coordinated control of the power grid; the key control variables include: the active power output adjustment amount of one or more generators, the target active power value of one or more key transmission sections, and the reactive power output setting value of one or more dynamic reactive power compensation devices; the decision action vector as a whole represents a coordinated and comprehensive control strategy aimed at pushing the system operating point toward the stable boundary.

[0024] Furthermore, the power grid simulation environment is an interactive platform that integrates a power grid differential algebraic equation solver and is specifically encapsulated for deep reinforcement learning training. The platform has an automated interactive interface and can perform the following operations in a loop: receive the decision action vector output by the actor network, modify the initial power flow value in the simulation model accordingly, then automatically apply a preset severe fault, perform time-domain simulation, and automatically calculate the stability margin index and the mixed reward value based on the simulation results. Finally, the new state and reward value are returned to the training algorithm, forming a closed-loop training iteration.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] This invention successfully decouples the complex power grid transient stability simulation process from the complex simulation by placing it in the offline training phase, while the online decision-making phase only requires the deployment of a lightweight actor network. In online applications, a single fast forward propagation calculation can directly map the real-time power grid state vector to the optimal decision action vector, achieving millisecond-level rapid determination of the power grid stability risk boundary. This approach completely changes the limitations of traditional methods that rely on extensive and time-consuming computations for online application. It endows the decision-making system with high real-time performance and adaptability to dynamic changes in the power grid, enabling dynamic adjustment of the stability boundary based on rapidly changing operating conditions, thus ensuring the safe operation of the power grid in complex and volatile environments.

[0027] This invention constructs a hybrid reward function that integrates boundary proximity, instability risk, and economic cost to guide the training of deep reinforcement learning models. This method not only uses ensuring grid instability as a hard constraint but also incentivizes the agent to learn a strategy that actively pushes the system's operating point towards the stability boundary through a reward mechanism, thereby fully exploring and utilizing the grid's potential transmission capacity. Simultaneously, by introducing multi-dimensional continuous decision-making action vectors, it achieves coordinated optimization control of multiple key variables such as generator output and cross-sectional power. The stability boundary found is no longer the conservative, static operating section obtained by traditional methods but a comprehensive optimal operating point that balances safety and operational economy, significantly improving the utilization efficiency of grid assets and the overall economic benefits of operation. Attached Figure Description

[0028] Figure 1 This is an exemplary flowchart of the boundary decision method of the present invention.

[0029] Figure 2 This is an exemplary flowchart of the steps involved in calculating the hybrid reward function according to the present invention. Detailed Implementation

[0030] The present invention will be further described below with reference to specific embodiments.

[0031] This application provides a method for intelligent boundary decision-making on power grid stability risk based on deep reinforcement learning. It aims to use artificial intelligence technology to determine the safety and stability boundary of the power system under the current operating conditions in real time and dynamically, thereby maximizing the operating efficiency of the power grid while ensuring system safety.

[0032] like Figure 1 The diagram shown is an exemplary flowchart of the boundary decision method in this embodiment, including the following steps:

[0033] Step S1: Construct a power grid state vector by acquiring real-time operation data and topology data of the power grid and processing them into a state vector that represents the current operating condition of the power grid.

[0034] In one embodiment, the system first acquires key operational data of the power grid in real time through an interface with the power grid's monitoring, control, and data acquisition system or a wide-area measurement system. Exemplarily, this real-time operational data includes, but is not limited to: voltage amplitude and phase angle of each node, active and reactive power on each transmission branch, active and reactive power output of each generator, and the system's global frequency. Simultaneously, the system also acquires current power grid topology data to clarify the connection relationships between various devices. Subsequently, this multi-source, heterogeneous data is integrated and arranged into a fixed-length one-dimensional vector, namely the state vector s. To eliminate differences between different physical dimensions and improve the training efficiency and stability of subsequent neural network models, the system normalizes this state vector, for example, by linearly scaling all values ​​to the interval [0,1] or [-1,1].

[0035] Step S2: Construct a deep reinforcement learning decision model. Construct a deep reinforcement learning decision model based on an actor-commentator architecture. The model includes an actor network for generating decision action vectors from state vectors, and a commentator network for evaluating the values ​​of state vectors and decision action vectors.

[0036] In step S2, the input to the actor network is the state vector, and the output is the decision action vector; the input to the commentator network is the state vector and the decision action vector output by the actor network, and the output is a scalar evaluation value that assesses the value of the decision action.

[0037] In one embodiment, the present invention employs an advanced actor-commentator architecture. This architecture comprises two deep neural networks: an actor network, which functions as a policy-making network. It receives the grid state vector *s* generated in step S1 as input and outputs a multidimensional, continuous decision action vector *a* through a series of nonlinear transformations. Each dimension of this vector corresponds to a specific, continuously adjustable control variable, such as the active power output adjustment of a generator. The commentator network functions as a value assessment network. It receives the state vector *s* and the action vector *a* output by the actor network as common input and outputs a single scalar evaluation value to estimate the long-term value of performing the action in the current state. During the training phase, this value assessment guides the actor network to optimize its policy.

[0038] Step S3: Calculate the hybrid reward function, execute the decision action vector in the power grid simulation environment, and perform transient stability simulation on the power grid state after execution. Based on the simulation results and the adjustment amount of the decision action vector, calculate a hybrid reward value that integrates boundary proximity reward, instability risk penalty and adjustment economic cost.

[0039] like Figure 2 The diagram shows an exemplary step flow for calculating the hybrid reward function in step S3 of this embodiment, including:

[0040] Step S301: Calculate the boundary proximity reward based on the adjustment amount of the decision action vector; calculate the target power of the key transmission section after the system executes the action based on the adjustment amount of the decision action vector a, and calculate the reward item accordingly.

[0041] Step S302: Determine whether the power grid is unstable based on the transient stability simulation results. If unstable, apply a preset negative instability penalty. Apply the action vector a to the power grid simulation model, simulate a preset severe fault, and perform transient stability simulation. Determine whether the power grid is unstable based on the simulation results. If unstable, apply a preset, large negative instability penalty.

[0042] Step S303: Calculate the economic cost of adjustment based on the magnitude of the decision action vector; calculate the economic cost of this control adjustment based on the magnitude of the decision action vector a.

[0043] Step S304: The boundary proximity reward, instability penalty, and adjustment economic cost are weighted and summed to obtain the final mixed reward value. The aforementioned boundary proximity reward, instability risk penalty, and adjustment economic cost are weighted and summed to obtain the final mixed reward value.

[0044] Mixed reward value R(s) t ,a t Calculated using the following formula:

[0045] R(s t ,a t ) = w prox ×f(P k,t+1 )-w pen ×I(M t+1 <M min )-w econ ×C(a t ), where: s t Represents the state vector at the current moment; a t f(P) represents the decision action vector at the current moment; k,t+1 () represents the target power P of the critical transmission section after the action is performed. k,t+1 Related boundary proximity reward; I(M t+1 <Mmin ) represents the penalty item for instability risk, where M t+1 M is the stability margin index calculated through transient stability simulation after the action is performed. min This represents the preset minimum safety margin threshold, and I(×) represents the indicator function, when M t+1 <M min Its value is 1 when it is active, and 0 otherwise; C(a) t ) indicates the adjustment cost of performing the action; w prox w pen w econ These are the weighting coefficients for boundary proximity reward, instability risk penalty, and adjustment economic cost, respectively, and w pen The value is greater than w prox and w econ The value of .

[0046] Step S4: Offline training of the model. Using the state vector, decision action vector, and mixed reward value, the deep reinforcement learning decision model is trained offline in the power grid simulation environment until the model converges.

[0047] The offline training of the model in step S4 includes:

[0048] Step S401: Store the empirical data generated by the interaction between the agent and the power grid simulation environment, which includes state vectors, decision action vectors, mixed reward values ​​and new state vectors, into an empirical replay buffer.

[0049] Step S402: Randomly sample data from the experience replay buffer and calculate the loss function based on the samples to update the network parameters of the actor network and the commentator network.

[0050] Step S5, online boundary decision-making: The trained and converged actor network is deployed in the power grid control system. The actor network receives the state vector generated in real time and outputs the optimal decision action vector. The operating point defined by the decision action vector is the intelligent boundary of power grid stability risk under the current operating condition.

[0051] In step S1, the real-time operating data of the power grid obtained includes node voltage amplitude, node voltage phase angle, branch active power, branch reactive power, generator active output, generator reactive output, and system frequency; the construction of the state vector also includes normalization processing of the real-time operating data.

[0052] In step S302, the process of determining whether the power grid is unstable involves comparing the dynamic trajectory data output from the transient stability simulation with preset instability criteria to generate a clear binary penalty signal for deep reinforcement learning training. The instability criteria include at least one of the following: the relative power angle difference of at least one generator exceeds a preset power angle instability threshold during the simulation period; the voltage of at least one critical bus remains below a preset voltage safety lower limit threshold for a duration exceeding a preset duration threshold after fault clearance; and the absolute value of the system frequency deviating from the nominal frequency exceeds a preset frequency offset threshold after fault clearance. The power angle instability threshold, voltage safety lower limit threshold, duration threshold, and frequency offset threshold are all configurable parameters set before the offline training of the model, based on the operating procedures and safety requirements of the target power grid.

[0053] The online boundary decision in step S5 decouples the decision-making process from the complex simulation. In the online stage, only the trained and converged actor network is deployed, without the need to deploy the commentator network and the power grid simulation environment. The actor network, as a policy function that implicitly contains the power grid stability mechanism, directly maps the real-time state vector to the optimal decision action vector through a single forward propagation calculation, thus completing the determination of the stability boundary.

[0054] In one embodiment, the key advantage of this invention lies in its decoupling of the decision-making process from complex simulation. After offline training, only the trained, lightweight actor network needs to be deployed to the actual power grid control system, while the commentator network and the large simulation environment do not require online deployment. During online operation, the control system collects the power grid state in real time and constructs a state vector s, which is then input into the deployed actor network. The actor network can directly output the optimal decision action vector a under the current operating condition through a single fast forward propagation calculation. The operating point defined by this vector is the "intelligent boundary of power grid stability risk" that balances safety and economy under the current operating condition. The control system can issue commands based on this result to achieve online, closed-loop, and intelligent stability control of the power grid.

[0055] The decision action vector is a multi-dimensional continuous vector, with each dimension uniquely corresponding to the adjustment amount of a preset key control variable in the power grid, thereby realizing multi-variable coordinated control of the power grid. Key control variables include: the active power output adjustment amount of one or more generators, the target active power value of one or more key transmission sections, and the reactive power output setpoint of one or more dynamic reactive power compensation devices. The decision action vector as a whole represents a coordinated and comprehensive control strategy aimed at pushing the system operating point toward the stability boundary.

[0056] The power grid simulation environment is an interactive platform that integrates a solver for power grid differential-algebraic equations and is specifically packaged for deep reinforcement learning training. The platform has an automated interactive interface that can perform the following operations in a loop: receive the decision action vectors output by the actor network, modify the initial power flow values ​​in the simulation model accordingly, automatically apply a preset severe fault, perform time-domain simulation, automatically calculate the stability margin index and mixed reward value based on the simulation results, and finally return the new state and reward value to the training algorithm, forming a closed-loop training iteration.

[0057] The above description is merely an example and illustration of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the scope defined by the invention, and all such modifications and additions should fall within the protection scope of the present invention.

Claims

1. A method for intelligent boundary decision-making on power grid stability risk based on deep reinforcement learning, characterized in that, Includes the following steps: Step S1: Construct a power grid state vector by acquiring real-time operation data and topology data of the power grid and processing them into a state vector that represents the current operating condition of the power grid. Step S2: Construct a deep reinforcement learning decision model. Construct a deep reinforcement learning decision model based on an actor-commentator architecture. The model includes an actor network for generating decision action vectors based on the state vectors, and a commentator network for evaluating the values ​​of the state vectors and the decision action vectors. Step S3: Calculate the hybrid reward function, execute the decision action vector in the power grid simulation environment, and perform transient stability simulation on the power grid state after execution. Based on the simulation results and the adjustment amount of the decision action vector, calculate a hybrid reward value that integrates boundary proximity reward, instability risk penalty and adjustment economic cost. Step S4: Offline training of the model. Using the state vector, decision action vector and the hybrid reward value, the deep reinforcement learning decision model is trained offline in the power grid simulation environment until the model converges. Step S5, online boundary decision-making: The actor network that has been trained and converged is deployed in the power grid control system. The actor network receives the state vector generated in real time and outputs the optimal decision action vector. The running point defined by the decision action vector is the intelligent boundary of power grid stability risk under the current operating condition.

2. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: In step S1, the real-time operating data of the power grid obtained includes node voltage amplitude, node voltage phase angle, branch active power, branch reactive power, generator active output, generator reactive output, and system frequency; the construction of the state vector also includes normalizing the real-time operating data.

3. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: In step S2: the input of the actor network is the state vector, and the output is the decision action vector; the input of the commenter network is the state vector and the decision action vector output by the actor network, and the output is a scalar evaluation value that assesses the value of the decision action.

4. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the process of calculating the mixed reward value includes: Step S301: Calculate the boundary proximity reward based on the adjustment amount of the decision action vector; Step S302: Determine whether the power grid is unstable based on the transient stability simulation results. If it is unstable, apply a preset negative instability penalty. Step S303: Calculate the adjustment economic cost based on the magnitude of the decision action vector; Step S304: The boundary proximity reward, the instability penalty, and the adjustment economic cost are weighted and summed to obtain the final mixed reward value.

5. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: In step S302, the process of determining whether the power grid is unstable involves comparing the dynamic trajectory data output from the transient stability simulation with preset instability criteria to generate a clear binary penalty signal for deep reinforcement learning training. The instability criteria include at least one of the following: the relative power angle difference of at least one generator exceeds a preset power angle instability threshold during the simulation period; the voltage of at least one critical bus remains below a preset voltage safety lower limit threshold for a duration exceeding a preset duration threshold after fault clearing; and the absolute value of the system frequency deviating from the nominal frequency exceeds a preset frequency offset threshold after fault clearing. The power angle instability threshold, the voltage safety lower limit threshold, the duration threshold, and the frequency offset threshold are all configurable parameters set before the offline training of the model begins, based on the operating procedures and safety requirements of the target power grid.

6. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: The mixed reward value R(s) t ,a t R(s) is calculated using the following formula: t ,a t ) = w prox ×f(P k,t+1 )-w pen ×I(M t+1 <M min )-w econ ×C(a t ), where: s t The state vector at the current moment; a t f(P) represents the decision action vector at the current moment; k,t+1 () represents the target power P of the critical transmission section after the action is performed. k,t+1 The associated boundary proximity reward; I(M) t+1 <M min ) represents the instability risk penalty term, where M t+1 M is the stability margin index calculated through transient stability simulation after the action is performed. min This represents the preset minimum safety margin threshold, and I(×) represents the indicator function, when M t+1 <M min Its value is 1 when it is active, and 0 otherwise; C(a) t ) indicates the economic cost of the adjustment performed; w prox w pen w econ These are the weighting coefficients for the boundary proximity reward, the instability risk penalty, and the adjustment economic cost, respectively, and w pen The value is greater than w prox and w econ The value of .

7. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: The offline training of the model in step S4 includes: Step S401: Store the empirical data generated by the interaction between the agent and the power grid simulation environment, which includes state vectors, decision action vectors, mixed reward values ​​and new state vectors, into an empirical replay buffer. Step S402: Randomly extract data samples from the experience replay buffer and calculate the loss function based on the samples to update the network parameters of the actor network and the commentator network.

8. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: The online boundary decision in step S5 decouples the decision-making process from the complex simulation. In the online stage, only the trained and converged actor network is deployed, without deploying the commentator network and the power grid simulation environment. The actor network, as a policy function that implicitly incorporates the power grid stability mechanism, directly maps the real-time state vector to the optimal decision action vector through a single forward propagation calculation, thus completing the determination of the stability boundary.

9. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: The decision action vector is a multi-dimensional continuous vector, each dimension of which uniquely corresponds to the adjustment amount of a preset key control variable in the power grid, thereby realizing multi-variable coordinated control of the power grid. The key control variables include: the active power output adjustment of one or more generators, the target active power value of one or more key transmission sections, and the reactive power output setting value of one or more dynamic reactive power compensation devices; the decision action vector as a whole represents a coordinated and comprehensive control strategy aimed at pushing the system operating point toward the stable boundary.

10. The intelligent boundary decision-making method for power grid stability risk based on deep reinforcement learning according to claim 1, characterized in that: The power grid simulation environment is an interactive platform that integrates a power grid differential algebra equation solver and is specifically packaged for deep reinforcement learning training. The platform has an automated interactive interface and can perform the following operations in a loop: receive the decision action vector output by the actor network, modify the initial power flow value in the simulation model accordingly, then automatically apply a preset severe fault, perform time-domain simulation, and automatically calculate the stability margin index and the mixed reward value based on the simulation results. Finally, the new state and reward value are returned to the training algorithm, forming a closed-loop training iteration.

Citation Information

Patent Citations

  • New energy consumption power dispatching method based on artificial intelligence

    CN115345380A

  • Multi-section power transmission limit calculation method based on multi-target migratable reinforcement learning

    CN119691331A

  • Intra-day ideal scheduling intelligent decision-making method based on safe deep reinforcement learning

    CN119886671A