Power grid section limit dynamic optimization control system and method based on reinforcement learning

Through the dynamic optimization control system for grid section limits based on reinforcement learning, key sections are identified and optimal control instructions are generated, the problem of lag in traditional grid control methods in dynamic changing scenarios is solved, and precise adjustment and adaptive optimization of grid section trends are achieved.

CN120357474AActive Publication Date: 2025-07-22HEFEI POWER SUPPLY COMPANY OF STATE GRID ANHUI ELECTRIC POWER +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510827886.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-22
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Traditional grid cross-section control methods are difficult to quickly update the control strategy in dynamic changing scenarios, resulting in lagging responses and insufficient control accuracy, and the inability to accurately adjust the current of key sections. There is also a lack of a predictive feedback mechanism for the system state after the FACTS control effect, resulting in the inability to dynamically optimize the control strategy.

Method used

The grid section limit dynamic optimization control system based on reinforcement learning identifies the controlled section set through sparse modeling and feature contribution analysis, constructs a dynamic evaluation model of section limit and response sensitivity matrix, and combines the multi-objective optimization function to generate the optimal control instructions of the FACTS controller to realize rolling optimization control.

Benefits of technology

It significantly improves the timeliness perception and dynamic accuracy of cross-sectional safety margin, realizes intelligent, dynamic and adaptive optimization of power grid cross-sectional limit control, ensures that the control effectiveness and security margin meet, and avoids invalid or redundant control resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357474A_ABST
    Figure CN120357474A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid section limit dynamic optimization control system and method based on reinforcement learning, and relates to the technical field of section limit dynamic control, and the method comprises the following steps: recognizing a controlled section set, and constructing a section limit dynamic evaluation model; a section response sensitivity matrix is constructed, and a multi-objective optimization function is constructed in combination with the section quota dynamic evaluation model; solving the multi-objective optimization function based on a reinforcement learning method to obtain an optimal control instruction set, and predicting a section power flow change result after control; a prediction result is fed back to the section quota dynamic evaluation model for judgment, a control instruction set is issued based on a judgment result, and dynamic optimization of a control process is achieved based on a rolling optimization control mechanism; according to the method, the optimal control instruction of the FACTS controller is generated through reinforcement learning, and the FACTS controller is accurately controlled through the control instruction to adjust the key section power flow distribution, so that the problem of insufficient section accommodating margin is effectively relieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-section limit dynamic control, and more specifically, to a power grid cross-section limit dynamic optimization control system and method based on reinforcement learning. Background Art

[0002] In the automated operation management of complex systems, especially in the scheduling control scenarios of large-scale wide-area distributed networks such as power systems, the unbalanced distribution of cross-section power flow often causes local channel overload, restricts the global operation margin, and further affects the achievement rate and robustness of the overall objective function of the control system. Especially in scenarios where the system has disturbances (such as load fluctuations, component removal, planned maintenance, etc.) or the operating constraints change dynamically, the traditional fixed threshold setting and rule-based control logic cannot accurately represent the real-time safety margin of the cross-section, resulting in the lack of fine dynamic adjustable ability of the control system.

[0003] Currently, the cross-section control function of power systems mostly relies on the sequential logic program control method driven by manually set parameters or the heuristic optimization control algorithm based on empirical rules.

[0004] For example, an adaptive control method for a multi-heat-source thermal management system based on reinforcement learning model predictive control disclosed in the invention patent announcement with the publication number CN119535989A includes: environmental state reception, reward function calculation, policy optimization, decision correction, optimization of the objective function and construction of the control law, temperature error calculation, action optimization, and device control. By integrating the policy optimization of the intelligent agent and the action optimization in the traditional model predictive control, the present invention realizes the real-time optimization of the control strategies for devices such as fans, water pumps, and valves in the cooling system, and has the adaptability of reinforcement learning and the stability of model predictive control; by setting the intelligent agent and the feedback correction module, the present invention realizes the adjustment of the prediction model parameters according to the actual state of the system, improving the robustness and adaptability of the control system; through the rolling optimization of the output control law combined with multiple real-time feedbacks and prediction corrections, the present invention realizes the efficient temperature control of the system in dynamic changes.

[0005] For example, a method for constructing a power grid dispatching control model and a power grid dispatching control method disclosed in the invention patent announcement with the announcement number of CN111864743B. The method for constructing the power grid dispatching control model includes: obtaining multiple historical sectional power flow data of the power grid; constructing a power grid dispatching control model based on the maximum entropy reinforcement learning algorithm according to the preset safe operation requirements and control objectives; extracting training samples from the multiple historical sectional power flow data, and inputting the training samples into the power grid dispatching control model for model training to obtain various power grid control actions; according to the power grid operation characteristics after executing each power grid control action based on the historical sectional power flow data, updating the model parameters of the power grid dispatching control model, and returning to the step of extracting the power grid operation characteristics corresponding to the current power grid operation index from the multiple historical sectional power flow data as training samples until all training samples are trained; determining the optimal power grid dispatching control model according to the training results. This lays a foundation for improving the efficiency and performance of power grid dispatching control.

[0006] In the above disclosed technical solution, there are at least the following technical problems: Although some control strategies have introduced FACTS controllers to improve flexibility, the current methods generally have deficiencies. For the control of FACTS controllers, traditional optimization methods often use static algorithms such as linear or nonlinear programming, genetic algorithms, etc., which are difficult to quickly update control strategies in dynamic change scenarios, resulting in response lags and insufficient control accuracy, and unable to achieve precise regulation of the critical sectional power flow to improve the sectional accommodation margin. And there is a lack of a prediction feedback mechanism for the system state after the FACTS control action, resulting in the control strategy being unable to be dynamically corrected and optimized according to the system response. In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0007] In order to overcome the above defects of the prior art, the embodiments of the present invention provide a power grid sectional limit dynamic optimization control system and method based on reinforcement learning, which generate the optimal control instructions of the FACTS controller through reinforcement learning, and precisely control the FACTS controller to adjust the critical sectional power flow distribution through the control instructions, thereby effectively alleviating the problem of insufficient sectional accommodation margin.

[0008] To achieve the above object, the present invention provides the following technical solutions: A power grid sectional limit dynamic optimization control system and method based on reinforcement learning, including the following steps: Identify the controlled section set based on sparse modeling and feature contribution analysis, and construct a dynamic assessment model for section limits; construct a section response sensitivity matrix based on the sensitivity of each FACTS controller to the section power flow change, and construct a multi-objective optimization function in combination with the dynamic assessment model for section limits; solve the multi-objective optimization function based on the reinforcement learning method to obtain the optimal control instruction set of the FACTS controller, and predict the section power flow change result after the control action; feedback the prediction result to the dynamic assessment model for section limits for judgment, and issue the control instruction set based on the judgment result and realize the dynamic optimization of the control process based on the rolling optimization control mechanism.

[0009] In a preferred embodiment, identifying the controlled section set based on sparse modeling and feature contribution analysis is specifically as follows: construct input features and output features, where the input feature is the control variable control matrix of the FACTS controller, and the output feature is the matrix of power flow change amounts of each section; based on the sparse modeling method, construct a section power flow change prediction model through the input and output features, traverse all power grid sections, and record the sections with non-zero outputs as high-response sections to form a candidate set of controlled sections; apply the SHAP explanation method to the output of the prediction model to obtain the corresponding contribution values, and obtain the contribution value index through statistical methods; for each section in the candidate set of controlled sections, perform a weighted score based on the number of its non-zero model outputs and the contribution value index, and screen the sections with scores greater than the preset threshold to obtain the controlled section set.

[0010] In a preferred embodiment, constructing a dynamic assessment model for section limits is specifically as follows: obtain the section limit impact data of the controlled section set, where the section limit impact data includes power offset rate characteristics, node voltage distribution gradient characteristics, and frequency disturbance response characteristics; input the section impact data as a feature vector, and construct a dynamic assessment model for section limits based on the XGBoost method, and output the accommodation margin of each section.

[0011] In a preferred embodiment, constructing a section response sensitivity matrix based on the sensitivity of each FACTS controller to the section power flow change is specifically as follows: obtain the FACTS controller data, and construct an initial section response sensitivity matrix through the power flow sensitivity analysis method; set a rolling time window and obtain the historical equipment data within this time window; input the historical equipment data set into the section power flow change prediction model to output a response sensitivity correction coefficient matrix; apply the response sensitivity correction coefficient matrix to the initial section response sensitivity matrix to dynamically correct the matrix, and use the corrected section response sensitivity matrix as the final section response sensitivity matrix.

[0012] In a preferred embodiment, a multi-objective optimization function is constructed in combination with the section limit dynamic assessment model, specifically: based on the section response sensitivity matrix and the section limit dynamic assessment model, a multi-objective optimization function with the goal of maximizing the section accommodation margin and minimizing the control cost is constructed, and constraint conditions are set to obtain a constrained optimization model. The constraint conditions include node voltage constraints, power balance constraints, and FACTS controller boundary constraints.

[0013] In a preferred embodiment, the optimal control instruction set of the FACTS controller is obtained by solving the multi-objective optimization function based on the reinforcement learning method, specifically: based on the agent, its state space and action space are defined, and the multi-objective optimization function is used as the reward function, where the positive reward represents the increase in the section accommodation margin, and the negative penalty term represents the control cost; the FACTS controller instruction is output through the policy network method, the quality of the current policy is evaluated through the value network method, and gradient update and policy iteration are performed; continuously collect the section power flow response, accommodation margin change, and control cost feedback indicators after the execution of the control instruction, and update the training data set; apply the trained reinforcement learning policy to the output FACTS controller to output the optimal control instruction set, and the optimal control instruction set includes the optimal control quantity of each FACTS controller.

[0014] In a preferred embodiment, the section power flow change result after the predictive control is predicted, specifically: update the parameter state of the corresponding FACTS controller in the power grid model according to the optimal control instruction set and construct an equipment parameter model; based on the AC power flow algorithm, use the equipment parameter model to calculate the predicted result of the whole network power flow distribution after the implementation of the FACTS control action, and extract the current power flow data of the controlled section.

[0015] In a preferred embodiment, the prediction result is fed back to the section limit dynamic assessment model for judgment, and based on the judgment result, the control instruction set is issued and the dynamic optimization of the control process is realized based on the rolling optimization control mechanism, specifically: use the current power flow data as the input parameter, feed it back to the section limit dynamic assessment model, and calculate the accommodation margin after control of the section after control; judge whether the accommodation margin after control meets the preset safety margin threshold. If it is greater than the preset section safety margin threshold, it is considered that the control instruction set is effective and the section safety margin meets the requirements; otherwise, return to regenerate the control instruction set to form a closed-loop control mechanism; if all sections meet the safety margin threshold, terminate the control process, issue the FACTS control instruction, and update the equipment state.

[0016] In a preferred embodiment, the dynamic optimization of the control process is achieved based on a rolling optimization control mechanism, specifically: constructing a cross-section limit optimization control process based on a rolling time window, setting a fixed-period control execution interval, and repeating the entire process of cross-section state evaluation, control optimization decision-making, control instruction issuance, and state feedback update within each period; at the start of each rolling period, updating the input data set; based on the data at the current moment, re-evaluating the cross-section limits, constructing a FACTS response sensitivity matrix, and updating the weights of the accommodation margins and control cost terms in the objective function to form an optimization model for the current period; using an intelligent optimization algorithm to solve the optimization model, forming a control instruction set applicable to the current period, outputting and executing it; and feeding back the system state after control execution to the evaluation model and optimizer for parameter update.

[0017] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By introducing sparse modeling and feature contribution analysis, accurately identify the set of controlled cross-sections that can be controlled and have a key impact on the safe operation of the system, avoid the allocation of ineffective or redundant control resources, and improve control efficiency; the dynamic evaluation model of cross-section limits constructed by integrating power offset, node voltage gradient, and frequency disturbance response characteristics breaks through the conservative bottleneck of traditional static limits, significantly enhancing the timeliness perception and dynamic accuracy of cross-section safety margins; combined with the cross-section response sensitivity matrix, clarify the marginal response capabilities of each FACTS controller to cross-section power flows, and achieve control decisions with strong responsiveness and clear directionality; finally, through a multi-objective optimization function aiming at maximizing the accommodation margin and minimizing the control cost, realize control adaptability and dynamic balance for the actual operation scenario.

[0018] 2. Construct a multi-objective optimization and rolling feedback control mechanism based on reinforcement learning with FACTS controllers as the core, realizing the intelligent, dynamic, and adaptive optimization of power grid cross-section limit control. This solution takes maximizing the cross-section accommodation margin and minimizing the control cost as dual objectives, and through state-action modeling and policy iteration learning, effectively identifies bottleneck cross-sections and generates an optimal control instruction set, with the ability of continuous learning and rolling update. At the same time, after control, closed-loop verification is carried out through power flow prediction and dynamic evaluation models to ensure that the control effectiveness and safety margin are met, realizing the real-time linkage of control-feedback-optimization. Description of the Drawings

[0019] Figure 1 It is a schematic flow chart of the dynamic optimization control method for power grid cross-section limits based on reinforcement learning provided by the embodiments of the present application.

[0020] Figure 2 It is a schematic structural diagram of the dynamic optimization control system for power grid cross-section limits based on reinforcement learning provided by the embodiments of the present application. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0022] Embodiment 1 Figure 1 The following is a schematic flowchart of a dynamic optimization control method for grid section limits based on reinforcement learning provided by an embodiment of this application, including the following steps: S1. Identify the controlled section set based on sparse modeling and feature contribution analysis and construct a dynamic assessment model for section limits.

[0023] In this embodiment, the controlled section set refers to a set of network sections that can be used as effective adjustment objects in the control system. The power flow distribution of this set can be adjusted by control means (such as FACTS devices, power output adjustment, load redistribution, etc.) deployed in the system. The identification process of this set is a key link in the identification and screening mechanism of controllable objects in the control system, aiming to improve the allocation efficiency of control resources and the system response performance. By identifying the controlled section set, it is possible to significantly avoid the indiscriminate allocation of limited control resources to all network sections, effectively avoiding the problems of the expansion of the control system model scale and the increase in operation complexity. Only optimize the control strategy for the key sections that can be accurately modeled and controlled and have a substantial impact on the system operation bottleneck, so as to improve the targeting and real-time performance of the control strategy, and provide target support for subsequent response modeling and optimal control.

[0024] Identifying the controlled section set based on sparse modeling and feature contribution analysis specifically includes: Obtain the historical operation data of the power grid and construct input features and output features. The input features are the control variable control matrices of FACTS controllers, and the output features are the power flow change amount matrices of each section. The historical operation data of the power grid includes the action amounts of FACTS controllers at multiple moments and the corresponding power flow change amounts of sections at the corresponding moments; Based on the sparse modeling method, construct a power flow change prediction model through input and output features, and traverse all sections. Mark the sections with at least one non-zero model output as high-response sections to obtain a candidate set of controlled sections; Obtain the corresponding contribution values based on the SHAP method according to the output of the power flow change prediction model, and obtain the contribution value index through statistical methods. The statistical method is to calculate the average value; For the cross-sections in the candidate set of controlled cross-sections, weighted scoring is performed based on the number of non-zero model outputs and the contribution value index thereof, and the cross-sections greater than the preset threshold are screened to obtain the controlled cross-section set.

[0025] The cross-section tidal current change prediction model has the following specific calculation formula:

[0026] In the formula, is the response sensitivity of the th cross-section to each feature, is the control variable control matrix, is the tidal current change amount matrix, is the regularization coefficient, is the sum of squared residuals, is the L1 norm.

[0027] Furthermore, a dynamic assessment model for cross-section limits is constructed, mainly to realize the dynamic assessment and prediction of the maximum tolerable tidal current capacity (i.e., real-time limit) of each key cross-section at the current moment. This model integrates the power offset rate feature, the node voltage distribution gradient feature, and the frequency disturbance response feature, and can output the "accommodation margin" of each cross-section, thereby providing a more fine-grained and dynamically perceivable basis for the cross-section operation boundary for the control system. And it dynamically replaces the static limit value, overcoming the problem of overly conservative safety margins in conventional control. In the subsequent steps, the response sensitivity ranking of the controlled cross-sections and the constraint optimization of the FACTS control amount both need to be based on the dynamic accommodation margin value of each cross-section. Therefore, this model constitutes the basic assessment module for the optimization of the entire control strategy.

[0028] The construction of the dynamic assessment model for cross-section limits is specifically as follows: Obtain the cross-section limit impact data of the controlled cross-section set, and the cross-section limit impact data includes the power offset rate feature, the node voltage distribution gradient feature, and the frequency disturbance response feature; Input the cross-section impact data as a feature vector, and construct a dynamic assessment model for cross-section limits based on XGBoost, and output the accommodation margin of each cross-section.

[0029] Furthermore, the power offset rate feature is used to measure the intensity of cross-section power fluctuations, which can reflect the impact intensity of factors such as load disturbances and intermittent power output fluctuations on the operating boundary of this cross-section. The larger the value, the more unstable the system is, and the more conservative the cross-section limit should be. By constructing a dynamic assessment model of cross-section limits using the power offset rate feature, the trend changes caused by cross-section load or source-end fluctuations can be captured in real time, enhancing the perception ability of the assessment model for sudden power flow disturbances. This helps to enhance the feedforward regulation ability of the control system's disturbance response, enabling the controller to adjust the control quantity in advance before perceiving the potential unstable trend of the key cross-section, and improving the system's ability to suppress sudden disturbances.

[0030] The specific calculation formula of the power offset rate feature is as follows:

[0031] In the formula, is the power offset rate feature, is the power flow value of a certain controlled cross-section at the current moment, is time.

[0032] The node voltage distribution gradient feature is used to reflect the difference in reactive power support capabilities on both sides of the cross-section, indirectly indicating the voltage stability risk. The larger the voltage difference, the more serious the voltage imbalance degree is, and the worse the operating stability of the cross-section is. Constructing a dynamic assessment model of cross-section limits using the node voltage distribution gradient feature helps to reflect the stability margin of the cross-section in terms of voltage support capabilities.

[0033] The specific calculation formula of the node voltage distribution gradient feature is as follows:

[0034] In the formula, is the node voltage distribution gradient feature, is the average voltage of the left region, is the average voltage of the right region, is the set of nodes in the left region of the cross-section, containing nodes, is the set of nodes in the right region of the cross-section, containing nodes, is the voltage amplitude of the th node in the left region of the cross-section, is the voltage amplitude of the th node in the right region of the cross-section.

[0035] The frequency disturbance response characteristic is used to measure the sensitivity of a section to system frequency disturbances. The larger the value, the more intense the response of the section to frequency disturbances, and there may be a large energy tide, which affects the transient stability of the power grid. By constructing a dynamic assessment model of section limits using the frequency disturbance response characteristic, the frequency fluctuation speed in the adjacent area of the section at the initial stage of the disturbance can be reflected, and the prediction ability of the dynamic safety margin of the section can be improved.

[0036] The specific calculation formula of the frequency disturbance response characteristic is as follows:

[0037] In the formula, is the frequency disturbance response characteristic, and are the power flow power values of the section at two moments before and after the disturbance occurs, and are the values of the network-wide frequency at two moments before and after the disturbance.

[0038] The specific calculation formula of the dynamic assessment model of section limits is as follows:

[0039] In the formula, is the section accommodation margin, is the power offset rate characteristic, is the node voltage distribution gradient characteristic, is the frequency disturbance response characteristic, and and are the weight coefficients after training.

[0040] S2. Based on the sensitivity of each FACTS controller to the change of section power flow, construct a section response sensitivity matrix, and combine it with the dynamic assessment model of section limits to construct a multi-objective optimization function.

[0041] In this embodiment, the section response sensitivity matrix refers to a two-dimensional matrix structure constructed through power flow sensitivity analysis, which represents the influence degree of each FACTS controller variable (such as injected reactive power voltage, series impedance change, etc.) on the power flow change of each controlled section in the power grid. Each element in the matrix represents the marginal influence of a certain FACTS controller on the power flow of a certain section, reflecting the linear or nearly linear response relationship between the control action and the target response. By constructing this sensitivity matrix, the response intensity of each FACTS controller to the power flow of each section can be systematically and quantitatively grasped, and the equipment with large control redundancy and weak response ability can be distinguished from the equipment with strong target response and large influence, providing a theoretical basis for subsequent control optimization.

[0042] Based on the sensitivity of each FACTS controller to the change of section power flow, a section response sensitivity matrix is constructed as follows: Obtain the FACTS controller data deployed in the target power grid, where the FACTS controller data includes device type, installation location, adjustable parameter range, control cost, and real-time operating status; Based on the FACTS controller data, through the power flow sensitivity analysis method, obtain the influence degree of the control variable of each FACTS controller on the change of section power flow, and construct an initial section response sensitivity matrix; Set a rolling time window, and obtain the historical device data within the time window, where the historical device data includes historical power flow data, FACTS controller command data, and corresponding controlled section power flow response data; Based on the section power flow change prediction model, use the historical device data as input to obtain a response sensitivity correction coefficient matrix; Apply the response sensitivity correction coefficient matrix to the initial section response sensitivity matrix to dynamically correct the matrix; Take the corrected initial section response sensitivity matrix as the final section response sensitivity matrix.

[0043] It should be noted that the response sensitivity correction coefficient matrix refers to a proportionality factor for quantitatively correcting the initially calculated response sensitivity value in order to more accurately evaluate the actual impact of FACTS devices on the change of controlled section power flow in power grid control modeling and response analysis.

[0044] By introducing a section limit dynamic assessment model, the current accommodation margin of each section can be reflected in real time. Taking this accommodation margin index as one of the optimization objectives can ensure that the control behavior is directly related to the system safety objective, no longer relying on static limits, and improving the responsiveness and pertinence of control. Using the section response sensitivity matrix can avoid trial-and-error-based control, improve the predictability and linear solvability of control actions, and make the multi-objective optimization process have a clear physical basis and mathematical solvability. The multi-objective optimization function simultaneously considers maximizing the accommodation margin and minimizing the control cost, avoiding control imbalance or redundant control caused by only focusing on a single dimension of section safety or control cost in the past; it supports flexible adjustment of the weight factor according to the operation scenario to achieve strategy adaptive control.

[0045] The construction of the multi-objective optimization function by combining the section limit dynamic assessment model is as follows: Based on the section response sensitivity matrix and the section limit dynamic assessment model, a multi-objective optimization function aiming to maximize the section accommodation margin and minimize the control cost is constructed, and a constrained optimization model is obtained by setting constraint conditions. The constraint conditions include node voltage constraints, power balance constraints, and FACTS controller boundary constraints. The node voltage constraint means that after the FACTS controller is applied, the voltage amplitude of any node should be maintained within its preset safe allowable range. The power balance constraint means that after the control is executed, the sum of the active power outputs of all generating nodes should be equal to the sum of the active powers of all load nodes to maintain the active power balance of the power grid. The FACTS controller boundary constraint means that the control value of each FACTS controller needs to satisfy the range defined by its minimum control capacity and maximum control capacity.

[0046] The specific calculation formula for the minimization of the control cost is as follows:

[0047] In the formula, is the minimization of the control cost, is the total number of FACTS controllers, is the FACTS controller index, is the control variable of the FACTS controller, is the quadratic cost sensitivity of the variable of this FACTS controller, is the linear cost sensitivity of the variable of this FACTS controller.

[0048] The specific calculation formula for the multi-objective optimization function is as follows:

[0049] In the formula, is the multi-objective optimization function, is the minimization of the control cost, is the maximization of the section accommodation margin, , are the weight coefficients.

[0050] It should be noted that the minimization of the control cost is based on the response sensitivity matrix as a constraint to measure the feasible control solution with the minimum cost.

[0051] S3. Solve the multi-objective optimization function based on the reinforcement learning method to obtain the optimal control instruction set of the FACTS controller, and predict the section power flow change result after the control action.

[0052] In this embodiment, due to the limited number of FACTS devices and their limited regulation capabilities, while there are numerous power flow sections to be regulated in the system, if a direct intervention method is adopted, not only will the control cost be high and the influence range be wide, but also operational feasibility constraints may be faced. Therefore, it is necessary to construct a multi-objective optimization model and integrate a reinforcement learning algorithm to dynamically and intelligently screen the optimal control strategy combination from a global perspective of the system, so as to minimize the control cost while improving the tolerance margin of key sections.

[0053] Among them, the section tolerance margin can be defined as the remaining active power disturbance margin that the section can withstand under the current operating state. By combining the real-time state estimation model and the control response sensitivity matrix, the intelligent solution mechanism can identify the bottleneck section of the current system and directionally guide the control action to the control path that can most improve its tolerance margin, avoiding the ineffective allocation of control resources on non-critical sections.

[0054] Different from the traditional control logic based on a rule base or manual strategy formulation, intelligent algorithms such as reinforcement learning have the capabilities of strategy adaptation and environmental feedback learning, and can effectively explore in the multi-strategy space and converge to the control combination that has the most improvement effect on the target section. This method not only improves the pertinence and effectiveness of control actions, but also can continuously evolve the control strategy in the rolling time domain, realize the self-optimization of control actions driven by feedback, and ultimately improve the accuracy and stability of the regulation of the section tolerance margin.

[0055] The optimal control instruction set of the FACTS controller is obtained by solving the multi-objective optimization function based on the reinforcement learning method, specifically as follows: Based on the intelligent agent, the state space is defined as the current section tolerance margin, the FACTS control variable state, and the network power flow state, the action space is defined as the controllable quantity of the FACTS controller, and the multi-objective optimization function is used as the reward function, where the positive reward represents the improvement amount of the section tolerance margin, and the negative penalty term represents the control cost; Output the FACTS controller instruction through the policy network method, evaluate the pros and cons of the current policy through the value network method, and perform gradient update and policy iteration; Through the experience replay mechanism and the target network to stabilize the training process, improve the convergence speed and policy stability; Continuously collect the section power flow response, the change of the tolerance margin, and the control cost feedback index after the execution of the control instruction, update the training data set, and improve the generalization ability and practicality of the reinforcement learning model; After the training is completed, apply the trained reinforcement learning strategy to output the optimal control instruction set of the FACTS controller under the current power grid state, and the optimal control instruction set includes the optimal control quantity of each FACTS controller.

[0056] Furthermore, the design and implementation of control measures should not solely rely on the verification and evaluation of ex-post results, but should possess predictive control capabilities, that is, perform feed-forward modeling and evaluation of the response changes in the system power flow state before the control action is executed, so as to achieve the pre-judgment of potential operation risks and the pre-analysis of control effects, and enhance the reliability and robustness of control strategies in a dynamic operating environment.

[0057] Based on the ability to predict the power flow evolution process after control, the system can actively identify and avoid secondary risks caused by control, such as line overload, voltage violation, frequency abnormality, etc. At the same time, coordinated control optimization among multiple FACTS devices can be achieved. With the support of the control coupling effect modeling, the operating safety of the entire system can be improved through the collaborative action of multiple devices, and the dynamic balance between control cost and system operating efficiency can be achieved under constraints, ultimately achieving the unified goal of safety and economy.

[0058] The cross-section power flow change results after the predictive control action are specifically as follows: Update the parameter status of the corresponding FACTS controller in the power grid model according to the optimal control instruction set, form the initial values of the adjusted topology structure and power flow parameters, and construct the device parameter model; Based on the AC power flow algorithm, use the device parameter model to calculate the predicted results of the whole network power flow distribution after the FACTS control action is implemented, and extract the current power flow data of the controlled cross-section.

[0059] S4. Feed the prediction results back to the cross-section limit dynamic assessment model for judgment, and issue the control instruction set based on the judgment result and realize the dynamic optimization of the control process based on the rolling optimization control mechanism.

[0060] In this embodiment, the power grid operating state has highly real-time and uncertain characteristics. The control instruction of any FACTS device may trigger new power flow distribution changes. If there is a lack of secondary effect evaluation of the system response after control execution, it may lead to ineffective control and even induce new cross-section overload risks. Therefore, it is necessary to construct a feedback evaluation mechanism based on predictive power flow to form a closed-loop control system of control-evaluation-correction to ensure the effectiveness and feasibility of control instructions under dynamic operating conditions. Although the control strategy initially stems from a multi-objective optimization solution model aimed at achieving a balance between safety and economy, due to factors such as certain structural approximations in the model, input data perturbations, and the non-linear dynamic response of the power system, the actual control execution results often have the risk of deviating from the expectations. By introducing a feedback link, online re-evaluation of whether the system reaches the preset cross-section accommodation margin threshold after control execution can be carried out, the control deviation can be identified, and the secondary optimization adjustment mechanism can be triggered, thereby enhancing the robustness and accuracy of the control process.

[0061] The prediction results are fed back to the dynamic evaluation model of the section limit, and the control instruction set is issued based on the judgment result output by the model to perform dynamic optimization of the section limit, specifically: The current flow data is used as input parameters and fed back to the dynamic assessment model of section limits, and the controlled accommodation margin of the controlled section is calculated; The controlled accommodation margin is judged to see whether it meets the preset safety margin threshold. If it is greater than the preset section safety margin threshold, it is considered that the control instruction set is valid and the section safety margin meets the requirements; otherwise, the control instruction set is returned to be regenerated to form a closed-loop control mechanism; If all sections meet the safety margin threshold, the control process is terminated, the FACTS control instructions are issued, and the equipment status is updated.

[0062] The specific calculation formula of the controlled accommodation margin is as follows:

[0063] In the formula, is the accommodation margin after control, is the current flow data, It is the section accommodation margin.

[0064] It should be noted that the controlled flow state is re-input into the dynamic evaluation model to obtain the estimated value of the controlled capacity margin based on the current state, which has higher timeliness and authenticity, and avoids the error accumulation caused by static models or one-time predictions. When the control result does not meet the capacity margin threshold, the feedback optimization module recalculates the control strategy and adjusts the model parameters according to the new state to achieve the coordinated evolution of adaptive control optimization and system state. Only when the control strategy is verified to be effective in the evaluation model (that is, the capacity margins of all sections meet the threshold), the control instructions are actually issued, which greatly improves the robustness and safety of the control decision. The closed-loop control mechanism can give priority to strategies with obvious control effects and lower costs to avoid wasting control resources. At the same time, through continuous feedback and optimization, the potential overload risk of the section is actively identified to achieve risk forward shift and control pre-positioning.

[0065] Furthermore, the rolling optimization control mechanism is designed. Its fundamental purpose is to build a real-time responsive control system for the evolution of complex power grid operation status, and to achieve dynamic and continuous optimization of section limit accommodation margin and risk pre-control. This step is not only the logical convergence of all the aforementioned module functions, but also the key link for the entire control system to achieve practical, intelligent and stable operation. It has the following functions and technical benefits: The power grid operation environment is highly dynamic: factors such as power grid load changes, renewable energy fluctuations, and equipment status changes all make the section power flow and limit margin have obvious time-evolution characteristics. Fixed control schemes are difficult to cope with real-time changes, and a rolling control mechanism is needed for dynamic adaptation.

[0066] Avoid the risks of control failure and lag: If the control strategy is formulated based on outdated data or a single control model, control lag or control deviation may occur in actual operation, leading to operation risks such as section over-limits. The rolling update mechanism can effectively reduce the timeliness problem of control strategies.

[0067] The control strategy needs to be dynamically corrected with the change of state: The response of the FACTS controller has certain nonlinearity and hysteresis. The rolling mechanism can achieve continuous feedback on the control execution results, so as to dynamically evolve and adaptively correct the control strategy.

[0068] Implement a self-closed-loop structure of control-feedback-re-optimization: This mechanism constructs a closed-loop chain of section evaluation, optimization decision-making, instruction issuance, state feedback, and re-evaluation, ensuring that each control operation has subsequent feedback evaluation and strategy correction, and improving the system robustness.

[0069] The dynamic optimization of the control process is realized based on the rolling optimization control mechanism, specifically as follows: Set a fixed control period, construct a rolling time window, and execute the optimization control process within each control period; At the start of each control period, obtain and update the current system operation data, where the system operation data includes section power flow, FACTS controller status, and available regulation resource capacity; Based on the updated system operation data, perform real-time evaluation of the section limit status and reconstruct the FACTS controller response sensitivity matrix; According to the real-time evaluation results and the sensitivity matrix, update the weight coefficients of the section accommodation margin and control cost in the optimization objective function, and construct an optimization model applicable to the current period; Use an intelligent optimization algorithm to solve the optimization model and generate a set of FACTS controller instructions for the current period; Send the control instructions to the corresponding FACTS controller and execute the corresponding control actions; Obtain the system state data after control execution and input it as feedback information to update the section limit evaluation model and optimization model parameters in the next control period. The system state data after control execution includes the section power flow change result after control action execution, the new state of the FACTS controller after control, and the section limit dynamic evaluation result.

[0070] Embodiment 2 Figure 2Schematic diagram of the dynamic optimization control system for grid section limits based on reinforcement learning provided by the embodiments of the present application, including a section state evaluation module, an objective optimization module, a control instruction generation and prediction module, and a feedback optimization module. There are connections between the modules: The section state evaluation module is used to identify the controlled section set based on sparse modeling and feature contribution analysis and construct a dynamic evaluation model for section limits; The objective optimization module is used to construct a section response sensitivity matrix based on the sensitivity of each FACTS controller to the change in section power flow, and construct a multi-objective optimization function in combination with the dynamic evaluation model of section limits; The control instruction generation and prediction module is used to solve the multi-objective optimization function based on the reinforcement learning method to obtain the optimal control instruction set of the FACTS controller, and predict the change result of the section power flow after the control action; The feedback optimization module is used to feedback the prediction result to the dynamic evaluation model of section limits for judgment, issue the control instruction set based on the judgment result, and realize the dynamic optimization of the control process based on the rolling optimization control mechanism.

[0071] The above formulas are all dimensionless and take their numerical calculations. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0072] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0073] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0074] In addition, the functional modules in each embodiment of the present application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0075] As described above, this is only the specific implementation manner of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

[0076] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A dynamic optimization control method for grid section limits based on reinforcement learning, characterized in that It includes the following steps: Based on sparse modeling and feature contribution analysis, identify the controlled section set and construct a dynamic evaluation model for section limits; Based on the sensitivity of each FACTS controller to the section power flow change, construct a section response sensitivity matrix, and combine it with the dynamic evaluation model of section limits to construct a multi-objective optimization function; Based on the reinforcement learning method, solve the multi-objective optimization function to obtain the optimal control instruction set of the FACTS controller, and predict the section power flow change result after the control action; Feed the prediction result back to the dynamic evaluation model of section limits for judgment, and based on the judgment result, issue the control instruction set and realize the dynamic optimization of the control process based on the rolling optimization control mechanism.

2. The dynamic optimization control method for grid section limit based on reinforcement learning according to claim 1, wherein The identification of the controlled section set based on sparse modeling and feature contribution analysis is specifically as follows: Construct input features and output features, where the input feature is the control variable control matrix of the FACTS controller, and the output feature is the matrix of power flow change amounts of each section; Based on the sparse modeling method, construct a section power flow change prediction model through the input and output features, traverse all power grid sections, and record the sections with non-zero outputs as high-response sections to form a candidate set of controlled sections; Apply the SHAP interpretation method to the output of the prediction model to obtain the corresponding contribution values, and obtain the contribution value index through statistical methods; For each section in the candidate set of controlled sections, perform a weighted score based on the number of non-zero model outputs and the contribution value index, and screen out the sections with scores greater than the preset threshold to obtain the controlled section set.

3. The dynamic optimization control method for grid section limit based on reinforcement learning according to claim 2, characterized in that The construction of the dynamic evaluation model for section limits is specifically as follows: Obtain the section limit impact data of the controlled section set, and the section limit impact data includes power offset rate characteristics, node voltage distribution gradient characteristics, and frequency disturbance response characteristics; Take the section impact data as the input of the feature vector, and based on the XGBoost method, construct a dynamic evaluation model for section limits, and output the accommodation margin of each section.

4. The method for dynamically optimizing and controlling the grid section limit based on reinforcement learning according to claim 3, wherein, The construction of the section response sensitivity matrix based on the sensitivity of each FACTS controller to the section power flow change is specifically as follows: Obtain the FACTS controller data, and construct an initial section response sensitivity matrix through the power flow sensitivity analysis method; Set a rolling time window and obtain the historical equipment data within this time window; Input the historical equipment data set into the section power flow change prediction model, and output the response sensitivity correction coefficient matrix; Apply the response sensitivity correction coefficient matrix to the initial section response sensitivity matrix, dynamically correct the matrix, and take the corrected section response sensitivity matrix as the final section response sensitivity matrix.

5. The dynamic optimization control method for grid section limit based on reinforcement learning according to claim 4, characterized in that The construction of the multi-objective optimization function by combining the dynamic evaluation model of section limits is specifically as follows: Based on the section response sensitivity matrix and the dynamic evaluation model of section limits, construct a multi-objective optimization function with the goal of maximizing the section accommodation margin and minimizing the control cost, and set the constraint conditions to obtain a constrained optimization model. The constraint conditions include node voltage constraints, power balance constraints, and FACTS controller boundary constraints.

6. The dynamic optimization control method for grid section limit based on reinforcement learning according to claim 1, characterized in that, The solution of the multi-objective optimization function based on the reinforcement learning method to obtain the optimal control instruction set of the FACTS controller is specifically as follows: Define its state space and action space based on the agent, and use the multi-objective optimization function as the reward function, where the positive reward represents the increase in the section accommodation margin, and the negative penalty term represents the control cost; Output the FACTS controller command through the policy network method, evaluate the pros and cons of the current policy through the value network method, and perform gradient update and policy iteration; Continuously collect the section power flow response, accommodation margin change and control cost feedback index after the control command is executed, and update the training data set; Apply the trained reinforcement learning policy to the output FACTS controller to output the optimal control instruction set, and the optimal control instruction set includes the optimal control amount of each FACTS controller.

7. The dynamic optimization control method for grid section limit based on reinforcement learning according to claim 1, characterized in that The result of the section power flow change after the predictive control action is specifically: Update the parameter state of the FACTS controller according to the optimal control instruction set and construct the equipment parameter model; Based on the AC power flow algorithm, use the equipment parameter model to calculate the predicted result of the whole network power flow distribution after the FACTS control action is implemented, and extract the current power flow data of the controlled section.

8. The method for dynamically optimizing and controlling the grid section limit based on reinforcement learning according to claim 7, characterized in that, The prediction result is fed back to the section limit dynamic evaluation model for judgment, and the control instruction set is issued based on the judgment result and the dynamic optimization of the control process is realized based on the rolling optimization control mechanism. Specifically: Use the current power flow data as the input parameter, feed it back to the section limit dynamic evaluation model, and calculate the accommodation margin after control; Judge whether the controlled accommodation margin meets the preset safety margin threshold. If it is greater than the preset section safety margin threshold, it is considered that the control instruction set is effective and the section safety margin meets the requirements; otherwise, return to regenerate the control instruction set; If all sections meet the safety margin threshold, terminate the control process, issue the FACTS control instruction and update the equipment state.

9. The method for dynamically optimizing and controlling the grid section limit based on reinforcement learning according to claim 8, characterized in that The dynamic optimization of the control process based on the rolling optimization control mechanism is specifically: Set a fixed control period, construct a rolling time window, and execute the optimization control process within each control period; At the beginning of each control period, obtain and update the current system operation data; Based on the updated system operation data, perform real-time evaluation of the section limit state and reconstruct the FACTS controller response sensitivity matrix; According to the real-time evaluation result and the sensitivity matrix, update the weight coefficients of the section accommodation margin and control cost in the objective function, and construct an optimization model applicable to the current period; Use the intelligent optimization algorithm to solve the optimization model and generate the FACTS controller instruction set within the current period; Send the control instruction to the corresponding FACTS controller and execute the corresponding control action; Obtain the system state data after the control execution and use it as the feedback information input to update the section limit evaluation model and optimization model parameters in the next control period.

10. A system using the dynamic optimization control method for grid section limit based on reinforcement learning according to any one of claims 1-9, characterized in that, It includes a section state evaluation module, an objective optimization module, a control instruction generation and prediction module, and a feedback optimization module, and there are connections between the modules: The section state evaluation module is used to identify the controlled section set based on sparse modeling and feature contribution analysis and construct a section limit dynamic evaluation model; The target optimization module is used to construct a section response sensitivity matrix based on the sensitivity of each FACTS controller to the section power flow change, and construct a multi-objective optimization function in combination with the section limit dynamic assessment model; The control instruction generation and prediction module is used to solve the multi-objective optimization function based on the reinforcement learning method to obtain the optimal control instruction set of the FACTS controller, and predict the section power flow change result after the control action; The feedback optimization module is used to feedback the prediction result to the section limit dynamic assessment model for judgment, issue the control instruction set based on the judgment result, and realize the dynamic optimization of the control process based on the rolling optimization control mechanism.

Citation Information

Patent Citations

  • Section power flow optimization control method based on N-1 static security constraints

    CN109301832A

  • Multi-index evaluation method and system for direct-current support strength degree of alternating-current power grid

    CN110070200A

  • Method and system for establishing parallel deep reinforcement learning model for power flow state adjustment

    CN113517684A

  • Method and system for improving transmission capacity of alternating-current and direct-current hybrid power transmission corridor

    CN119813402A

  • Linearized robust optimal power flow generation method and system

    WO2024259880A1