Power grid control method and device based on hybrid expert network, and medium

By constructing a power grid control method based on a hybrid expert network, the computational complexity and communication failure problems of traditional power grid control methods under high-proportion renewable energy access are solved. This enables second-level optimization and millisecond-level decision-making for large-scale power grids, improving the security and responsiveness of the power grid.

CN121507944APending Publication Date: 2026-02-10GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511495494.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-20
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Traditional power grid control methods face challenges such as high computational complexity, easy global failure due to communication failures, lack of topology adaptation capability in static partitioning strategies, economic losses caused by conservative constraints in robust optimization, and poor generalization of data-driven models when a high proportion of renewable energy is integrated and power electronic equipment is proliferating. These problems make it difficult to meet the real-time response requirements and safe and efficient operation of large-scale power grids.

Method used

A power grid control method based on a hybrid expert network is constructed. Through a multimodal data-driven intelligent scheduling model, a decoupling-coordination mechanism for electrical coupling quantification, and a two-layer reinforcement learning framework, dynamic expert decision-making and distributed autonomous optimization are achieved. Spatiotemporal graph convolutional networks are used to improve the accuracy of the renewable energy consumption model. Expert weights are dynamically allocated by a gating network for parallel optimization and cross-regional coordination. A lightweight communication protocol is used to achieve efficient synchronization.

Benefits of technology

It reduces the computational complexity of global optimization from O(N³) to O(N), supports second-level collaborative optimization of power grids with thousands of nodes, achieves millisecond-level local decision-making, improves the power grid's response to load changes and new energy fluctuations, and ensures the safe and efficient operation of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121507944A_ABST
    Figure CN121507944A_ABST
Patent Text Reader

Abstract

The invention discloses a power grid control method and device based on a hybrid expert network and a medium, and belongs to the technical field of power grid optimization, and the method comprises the steps: obtaining power grid data which comprises node voltage, load demands, generator output and energy storage charging and discharging data; based on the power grid data, an intelligent scheduling model driven by multi-modal data is constructed, and the intelligent scheduling model comprises a plurality of expert sub-networks composed of a steady-state regulation and control model, a deep reinforcement learning model, a model prediction control model, a new energy consumption model and dynamic safety evaluation, and dynamically allocating expert weights through the gating network to realize self-adaptive decision making. Through the combination of the dynamic expert network and the lightweight gating mechanism, the multi-modal data of the power grid can be analyzed in real time, the optimal expert combination is activated in a self-adaptive manner, and dynamic spatial-temporal characteristic capture is realized. Compared with a traditional deep reinforcement learning method, the method is remarkably improved, and excellent self-adaptability and real-time performance are shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power grid optimization technology, and more specifically to power grid control methods, equipment, and media based on hybrid expert networks. Background Technology

[0002] With the high proportion of renewable energy access and the surge in power electronic equipment, traditional power grid control methods face multiple challenges: centralized optimization is difficult to meet the real-time response requirements of large-scale power grids due to its high computational complexity (O(N³)), and communication failures can easily lead to global failures; the control success rate of a single model drops sharply when wind and solar fluctuations exceed 30%, while fixed-rule expert systems cannot adapt to the dynamic characteristics of new equipment; static partitioning strategies and master-slave collaborative mechanisms lack topology adaptation capabilities, leading to an increased risk of cascading overloads under "N-2" faults; traditional robust optimization suffers economic losses due to conservative constraints, data-driven models have poor generalization to extreme events, and the transient recovery efficiency under unmodeled disturbances is less than 20% improved.

[0003] Existing technologies such as Hybrid Expert Networks (MoE) and Distributed Autonomous Optimization Models (ADMM) also have shortcomings. Shallow gated networks struggle to capture meteorological-electrical coupling characteristics, ADMM iteration efficiency is low in strongly coupled power grids, and offline adversarial training is disconnected from online state management. Therefore, there is an urgent need to construct a novel control framework that integrates dynamic expert decision-making, distributed autonomous optimization, and online robust learning to overcome bottlenecks in multi-objective conflicts, strong uncertainty disturbances, and wide-area collaboration, thereby ensuring the safe and efficient operation of new power systems. Summary of the Invention

[0004] To address the aforementioned technical issues, a power grid control method based on a hybrid expert network is proposed, which includes acquiring power grid data, including node voltage, load demand, generator output, and energy storage charging and discharging data.

[0005] Based on the power grid data, a multimodal data-driven intelligent scheduling model is constructed. The intelligent scheduling model includes an expert subnetwork consisting of a steady-state control model, a deep reinforcement learning model, a model predictive control model, a new energy consumption model, and a dynamic security assessment. Expert weights are dynamically allocated through a gating network to achieve adaptive decision-making.

[0006] By quantifying electrical coupling, a decoupling-coordination mechanism for the wide-area power grid is constructed. The power grid is divided into electrically weakly coupled sub-regions. Each sub-region is optimized in parallel based on local expert strategies, and cross-regional power balance and safety boundary coordination are achieved through global consistency constraints.

[0007] Based on the aforementioned intelligent scheduling model and decoupling-coordination mechanism, through online evolution of data-knowledge fusion, a two-layer reinforcement learning framework is adopted to dynamically adjust the expert network parameters. The two-layer reinforcement learning framework includes a bottom-layer offline pre-training and an upper-layer online rolling optimization to cope with the uncertainty interference of load mutations and new energy fluctuations.

[0008] As a preferred embodiment of the power grid control method based on a hybrid expert network described in this invention, the construction of the multimodal data-driven intelligent scheduling model includes:

[0009] A steady-state control model is constructed, integrating the economic optimization of deep reinforcement learning with the real-time safety constraints of model predictive control. The objective function is a weighted sum of economic costs and safety constraints, expressed as follows:

[0010]

[0011] in, It is for variables This minimizes the summation expression. Indicates time The state vector; The scheduling period; For dynamic weights, For time step index; The economic cost objective function for deep reinforcement learning, The objective function for safety constraints in model predictive control; T-1 represents the start time step within the scheduling period; T-1 represents the last time step within the scheduling period.

[0012] The deep reinforcement learning model adopts the Q-learning framework and defines the action value function, expressed as follows:

[0013]

[0014] in, Indicates the expected cumulative reward. Represents the mathematical expectation; Includes real-time power grid status. For control commands, It is a discount factor; For a moment +k is the immediate reward; k represents the time offset from the current time step t. The whole is expressed as a long-term cumulative reward expression for mathematical expectation;

[0015] The predictive control model is solved through rolling optimization and is expressed as follows:

[0016]

[0017] in, To predict the control sequence in the time domain, H is the prediction time domain length. and These are the cost coefficients for power generation and energy storage, respectively. and This represents the generator output and energy storage charging / discharging power at the corresponding moment.

[0018] As a preferred embodiment of the power grid control method based on a hybrid expert network described in this invention, the renewable energy consumption model employs a spatiotemporal graph convolutional network to improve the accuracy of ultra-short-term power prediction, including:

[0019] Constructing the spatiotemporal diagram of new energy power plants G=( , ):node , Indicates the first A wind and solar (wind power and photovoltaic) power station, Represents the set of all nodes and edges. , Represents a node and The connection between them Represents the set of connections between nodes;

[0020] Spatial convolution aggregates neighborhood node information, represented as follows:

[0021]

[0022] in, For the first Layer features, For the first Layer features; Spatial convolution weights; ReLU is the activation function; Represented as an adjacency matrix;

[0023] Temporal dependencies are captured through temporal convolution, represented as follows:

[0024]

[0025] in, It is the first convolution after temporal convolution in a spatiotemporal graph convolutional network. Layer feature output; 1D is for capturing temporal features through convolution, where Conv1D represents a one-dimensional convolution operation. These are the temporal convolution weights.

[0026] As a preferred embodiment of the power grid control method based on hybrid expert networks described in this invention, the method includes: constructing a decoupling-coordination mechanism for a wide-area power grid, comprising:

[0027] Electrical Coupling , is represented as ,

[0028]

[0029] in, Gather for the neighbors, It is an index variable. For electrical coupling, This represents an element of the admittance matrix, and |.| represents taking the absolute value;

[0030] The power grid is partitioned by minimizing the coupling between regions, which is expressed as follows:

[0031]

[0032] Regional size constraints, expressed as,

[0033]

[0034] Power balance constraints are expressed as follows:

[0035]

[0036] in, This represents the node of the first sub-region; For the set of sub-region nodes, For the number of partitions, This indicates the introduction of constraints. This indicates the maximum number of nodes allowed in a single sub-region. Represents a node The generator output, This represents the total load demand of the entire power grid;

[0037] Kalman filtering eliminates noise, as shown below.

[0038]

[0039] in, It is a quantitative indicator of the state of a sub-region. It is the state transition matrix. For PMU measurement data, It is the observation matrix. This is the Kalman gain.

[0040] As a preferred embodiment of the power grid control method based on a hybrid expert network described in this invention, the method includes: dynamically allocating expert weights through a gating network, including:

[0041] Expert weights are generated using a multilayer perceptron, and are represented as follows:

[0042]

[0043] The normalized exponential function transforms any real vector into a probability distribution, ensuring that the output values ​​are non-negative and sum to 1. Attention function. By calculating the query matrix Key matrix Value matrix The correlation is used to extract key features from multimodal inputs. Representation matrix The transpose of the expression, in the attention mechanism formula. , , From coding features Obtained by linear transformation, The dimension of the key matrix;

[0044] As a preferred embodiment of the power grid control method based on a hybrid expert network described in this invention, the two-layer reinforcement learning framework includes:

[0045] The underlying layer performs offline pre-training through near-end policy optimization, defining a reward function as follows:

[0046]

[0047] in, For a moment Instant rewards; Weighting coefficients used to adjust for the impact of power generation costs. Weighting coefficients used to adjust the effects of voltage deviation. Weighting coefficients used to adjust the influence of changes in motion; Represents the cost of electricity generation; It is the real-time voltage value. Compared with reference value The 2-norm of the deviation; For the current control action Actions at the previous moment The 2-norm of the change;

[0048] The optimization objective is denoted as,

[0049]

[0050] in, Optimize the objective function for pruning; The parameter is The policy network outputs in the state Next action The probability of. This represents the probability output of the old policy network, used to calculate the ratio of the old and new policies to limit the update magnitude; Indicates time The control action vector; Indicates time The power grid state vector; This is the dominant value.

[0051] As a preferred embodiment of the power grid control method based on hybrid expert networks described in this invention, the online evolution includes adversarial training to improve generalization to unknown disturbances, including:

[0052] Adversarial perturbations are generated using the fast gradient sign method, denoted as follows:

[0053]

[0054] in, To counteract disturbances; It is a hyperparameter that controls the amplitude of the disturbance; This is the sign function, which outputs the sign of the gradient direction. Indicates the direction in which the perturbation is added; It is the gradient of the state with respect to the loss function. The action distribution output by the policy network, The target action.

[0055] Another objective of this invention is to provide a power grid control system based on a hybrid expert network. By combining a dynamic expert network (MoE) with a lightweight Transformer gating mechanism, the system can analyze multimodal power grid data in real time, adaptively activate the optimal expert combination, and achieve dynamic spatiotemporal feature capture.

[0056] As a preferred embodiment of the power grid control system based on a hybrid expert network described in this invention, the online rolling optimization includes dynamically adjusting the expert network parameters through real-time data feedback and coordinating the safety boundaries between sub-regions under global consistency constraints.

[0057] A computer device includes a memory and a processor, the memory storing a computer program, the processor executing the computer program to implement the steps of the power grid control method based on a hybrid expert network.

[0058] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the power grid control method based on a hybrid expert network.

[0059] The beneficial effects of this invention are as follows: Based on the dynamic partitioning algorithm of electrical coupling degree and the improved ADMM framework, the system reduces the global optimization computation complexity from O(N³) to O(N), and shortens the time for collaborative optimization of a thousand-node power grid from minutes to seconds.

[0060] The system adopts a parallel distributed architecture, supports millisecond-level local decision-making in sub-regions based on a 100Hz PMU data sampling rate, and achieves efficient cross-region data synchronization through a lightweight communication protocol (DDS-RTPS). Attached Figure Description

[0061] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a general flowchart of a power grid control method based on a hybrid expert network, provided as an embodiment of the present invention.

[0063] Figure 2 This is a schematic diagram of the multi-expert mechanism of a power grid control method based on a hybrid expert network, provided as an embodiment of the present invention.

[0064] Figure 3 A comparison diagram of the robustness of power grid operation of a power grid control method based on a hybrid expert network provided in an embodiment of the present invention.

[0065] Figure 4 A comparison diagram of the power grid operation stability effect of a power grid control method based on a hybrid expert network provided in an embodiment of the present invention during communication failure.

[0066] Figure 5 A comparison diagram of the power flow fluctuation stability of a power grid control method based on a hybrid expert network, provided in an embodiment of the present invention.

[0067] Figure 6 A comparison diagram of the grid stability effects of a grid control method based on a hybrid expert network under voltage sag and wind power fluctuations, provided in an embodiment of the present invention. Detailed Implementation

[0068] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0069] Example 1, referring to Figures 1-2 This is the first embodiment of the present invention, which provides a power grid control method based on a hybrid expert network, including:

[0070] Step S1: Obtain grid data, which includes node voltage, load demand, generator output, and energy storage charging and discharging data.

[0071] Step S2: Based on the power grid data, construct a multimodal data-driven intelligent scheduling model. The intelligent scheduling model includes an expert subnetwork consisting of a steady-state control model, a deep reinforcement learning model, a model predictive control model, a new energy consumption model, and a dynamic security assessment. Expert weights are dynamically allocated through a gating network to achieve adaptive decision-making.

[0072] Step S3: By quantifying electrical coupling, a decoupling-coordination mechanism for the wide-area power grid is constructed. The power grid is divided into electrically weakly coupled sub-regions. Each sub-region is optimized in parallel based on local expert strategies, and cross-regional power balance and safety boundary coordination are achieved through global consistency constraints.

[0073] Step S4: Based on the intelligent scheduling model and decoupling-coordination mechanism, through online evolution of data-knowledge fusion, the expert network parameters are dynamically adjusted using a two-layer reinforcement learning framework. The two-layer reinforcement learning framework includes a bottom-layer offline pre-training and an upper-layer online rolling optimization to cope with the uncertainty interference of load mutations and new energy fluctuations.

[0074] Specifically, it involves acquiring power grid data; obtaining real-time power grid data, such as voltage, current, and frequency, from the monitoring systems of companies like State Grid and China Southern Power Grid.

[0075] Construct a steady-state control model. Integrate the economic optimization of deep reinforcement learning (DRL) with the real-time safety constraints of model predictive control (MPC). The scheduling period is T, and the objective function is a weighted sum of economic costs and safety constraints:

[0076]

[0077] Where, min is for the variable Find the minimum value of the summation expression from t=0 to t=T−1, that is, by choosing an appropriate... This minimizes the summation expression. T is the scheduling period, referring to the duration of a complete optimal scheduling operation of the distribution network. Due to the rapid changes in renewable energy and load in the distribution network, a shorter period (e.g., 15 minutes) is typically set to ensure timely response to state changes and stable system operation. t is the time step index, used to mark each discrete time point within the scheduling period (from the start time to time T-1). t=0 represents the starting time step within the scheduling period, and T-1 represents the last time step within the scheduling period. This refers to the status of the distribution network system (including node voltage, load demand, etc.). It is the scheduling cycle. Dynamic weights (adaptively adjusted based on real-time safety margins, such as...) ), and These are the economic cost objective function of deep reinforcement learning and the safety constraint objective function of model predictive control, respectively. This formula integrates the economic optimization of DRL with the real-time safety constraints of MPC, using dynamic weights to balance their priorities. DRL learns the economic patterns of historical scheduling strategies, while MPC solves the rolling optimization problem under safety constraints in real time. When the risk of voltage exceeding limits increases... Approaching zero allows MPC to dominate decision-making to ensure safety, while DRL dominates to optimize economic costs.

[0078] The Deep Reinforcement Learning (DRL) model uses the Q-learning framework to define the action value function:

[0079]

[0080] in, Indicates the expected cumulative reward. Represents the mathematical expectation; Includes real-time power grid status. For control commands, It is the discount factor; k represents the time offset; The overall long-term cumulative reward expression, as a mathematical expectation, is subsequently derived through... (Mathematical expectation) Calculate the value of an action. This provides the core basis for "quantifying the long-term value of control actions" in the adaptive decision-making of the deep reinforcement learning model for power grids.

[0081] Instant reward function The cost of power generation and voltage deviation will be used as penalties. ; Weighting coefficients; This framework uses reinforcement learning to learn the optimal scheduling strategy, aiming to maximize long-term cumulative rewards, and guides the agent to choose economical and safe control actions. Indicates the current time step After that, the Instant rewards for every moment. This represents the time offset, used to calculate the cumulative discount of rewards at future times. At that time, the corresponding "current moment" (i.e., time step) The reward at this time is , is the current action In the current state The instant rewards generated below. This represents the time step index, used to mark discrete time points (from 0 to T-1) within a scheduling period.

[0082] Each step solves the rolling time-domain optimization problem:

[0083]

[0084] In the rolling optimization problem of the MPC model, To predict the control sequence in the time domain, It predicts the length of the time domain. and These are the cost coefficients for power generation and energy storage, respectively. and This formula represents the generator output and energy storage charging / discharging power at the corresponding moment; it is based on the principle of model predictive control, which predicts future... The system behavior at each time step is analyzed to find the control sequence that minimizes the cost of power generation and energy storage. Rolling optimization is performed at each step to adapt to real-time changes in the power grid and ensure that the decision meets the current security constraints.

[0085] Through a dynamic weighted fusion model, based on real-time security risks Adjust the weights:

[0086]

[0087] In the dynamic weight adjustment formula, For a moment Dynamic weights, It is the Sigmoid function. m is the sensitivity coefficient. It is a voltage over-limit risk indicator ( For voltage, (As a threshold); this formula uses the Sigmoid function to transform the risk of voltage exceeding the limit into a dynamic weight. When the risk of voltage exceeding the limit increases, Approaching 0 allows MPC to lead decision-making and ensure safety, while DRL leads optimization of economy, achieving an adaptive balance between "safety and economy" and avoiding the shortcomings of traditional fixed weights.

[0088] New energy consumption model: Spatiotemporal Graph Convolutional Network (ST-GCN), which improves the accuracy of ultra-short-term power prediction by utilizing spatiotemporal correlation. Constructing the spatiotemporal graph of new energy power plants, G = ( , ):

[0089] Constructing the spatiotemporal diagram of new energy power plants G=( , ):node , Indicates the first A scenic station, Represents the set of all nodes and edges. , Represents a node and The connection between them Represents the set of connections between nodes;

[0090] The weights are quantized by the adjacency matrix A, and the edge weights reflect the electrical coupling strength and geographical correlation between stations. The adjacency matrix A is constructed based on electrical distance and geographical correlation, with the following weights: The electrical distance between nodes is represented by the admittance matrix elements derived from the power grid topology calculation. The smaller the electrical distance, the closer the electrical connection between the two stations in the power grid, and the more direct the power interaction. This represents the bandwidth parameter of the Gaussian kernel, which controls the rate of weight decay.

[0091] Each layer contains spatial convolution and temporal convolution. Spatial convolution aggregates neighborhood node information and is expressed as:

[0092]

[0093] Spatial convolution formula utilizes adjacency matrix (weight) Aggregate information from neighboring stations. For the first Layer features, For the first Layer features, For spatial convolution weights, For activation functions;

[0094] Temporal convolution uses 1D convolution to capture temporal dependencies:

[0095]

[0096] It is the first convolution after temporal convolution in a spatiotemporal graph convolutional network. Layer feature output, the temporal convolution formula captures temporal features through 1D convolution, where This represents a one-dimensional convolution operation. These are the temporal convolution weights.

[0097] Multi-scale prediction fusion final output power prediction value:

[0098]

[0099] in, It is the output of the new energy consumption model through multi-scale prediction fusion. Predicted new energy power at any given time This indicates the number of feature fusion layers (i.e., the number of feature levels involved in multi-scale fusion). ( ) represents a multilayer perceptron, used to map features to power predictions. The multi-scale prediction formula utilizes attention weights. Integrating features from each layer The model predicts power through a multilayer perceptron; it utilizes the spatiotemporal correlation of the power grid and improves the accuracy of new energy power prediction through graph convolution and temporal convolution, adapting to the multi-frequency fluctuation characteristics of wind and solar power.

[0100] Gated networks utilize lightweight Transformer dynamic weight allocation to dynamically activate optimal expert combinations based on multimodal data. By inputting feature encodings, multi-source data is projected onto a unified dimension.

[0101]

[0102] In the attention calculation of gating networks, Through projection matrix , , Power grid status Meteorological data Other auxiliary data The `oth` parameter is mapped to a unified dimension and concatenated; the query is then calculated. ),key( ),value( )matrix:

[0103]

[0104] Attention score:

[0105]

[0106] The normalized exponential function transforms any real vector into a probability distribution, ensuring that the output values ​​are non-negative and sum to 1. Attention function. By calculating the query matrix Key matrix Value matrix The correlation is used to extract key features from multimodal inputs. Representation matrix The transpose of the expression, in the attention mechanism formula. , , From coding features Obtained by linear transformation, The dimension of the key matrix;

[0107] Only the top-3 attention points are retained to reduce computational load.

[0108] Expert weights are generated by mapping the attention output to a weight vector through a fully connected layer.

[0109]

[0110] The expert weight vector is generated by a multilayer perceptron that transforms the attention output into expert weights. This mechanism captures multimodal feature associations through self-attention and dynamically activates the optimal expert combination (such as activating energy storage frequency regulation experts when wind speed changes abruptly), thereby achieving adaptive weight allocation.

[0111] For example, when input features Includes voltage sag indicator and When a sudden change in wind speed is indicated, the MLP outputs:

[0112]

[0113] in, This indicates that the weight of the "voltage recovery expert subnetwork" is 0.6. This indicates that the weight of the "energy storage frequency regulation expert sub-network" is 0.3. The total weight of other expert subnetworks is 0.1. Furthermore, the system realizes multimodal data perception, dynamic expert weight allocation and safety-economic collaborative optimization, providing highly adaptable decision support for complex power grid scenarios.

[0114] A dynamic partitioning algorithm, based on electrical coupling degree, decouples a large-scale power grid into autonomous sub-regions with low electrical coupling degree. Electrical coupling degree is quantified. :

[0115]

[0116] in, Gather for the neighbors, It is an index variable. For electrical coupling, represents the element of the admittance matrix, and | represents taking the absolute value;

[0117] The dynamic zoning optimization model constructs an optimization problem that minimizes the coupling between regions. The power grid is then partitioned by minimizing this inter-regional coupling problem, denoted as follows:

[0118]

[0119] Regional size constraints are expressed as:

[0120]

[0121] Power balance constraints are expressed as:

[0122]

[0123] The optimization objective is to partition the power grid by minimizing inter-regional coupling. For the set of sub-region nodes, Let N be the number of partitions, and N be the number of individual sub-regions. It is composed of nodes in the first sub-region of the entire power grid partition, and... Together they form all the zones of the power grid. This indicates the introduction of constraints. This represents the maximum number of nodes allowed in a single sub-region, i.e. , This represents the generator output at node i. This represents the total load demand of the entire power grid; the algorithm quantifies the node coupling strength based on the admittance matrix and divides the power grid into electrically weakly coupled sub-regions through spectral clustering, reducing the computational complexity of distributed optimization (from...). Down to N is the problem size. Requires cubic-scale time for the problem. It only requires N problem-scale time to execute, reducing the computation time from cubic to constant, thus ensuring power balance across regions.

[0124] Based on PMU state estimation, each sub-region updates its state vector using PMU measurement data (sampling rate 100Hz). :

[0125]

[0126] Indicates the voltage phase angle; Subregion At any moment The active power of q nodes.

[0127] Noise removal using Kalman filtering:

[0128]

[0129] It is a sub-region state quantization index obtained by updating the measurement data (sampling rate 100Hz) through PMU (Synchronous Phasor Measurement Unit) and eliminating noise through Kalman filtering. The Kalman filter formula includes... sub-region At any moment The state vector, It is the state transition matrix. For PMU measurement data, It is the observation matrix. Kalman gain; sub-region At any moment The state vector; this formula eliminates noise in PMU data through recursive calculation of state prediction and measurement update, realizes high-precision estimation of sub-region state, and supports accurate acquisition of real-time power grid state.

[0130] Boundary information exchange protocol, adjacent sub-regions and Exchange boundary information via DDS-RTPS protocol:

[0131] Sending data indicates:

[0132]

[0133] Update rules: ,

[0134] In the update rules of the border information exchange protocol, and These are shared values ​​for the phase angle and power at the boundaries of adjacent sub-regions. These represent the voltage phase angles at the boundary nodes of the sub-regions, , Representing sub-regions and subregions At any moment The active power at the boundary, It is a weighting factor based on line capacity; this rule achieves cross-regional data synchronization by weighted averaging of boundary information of adjacent sub-regions, ensuring the consistency of boundary conditions in distributed optimization.

[0135] The improved ADMM algorithm, a distributed solution to cross-regional power flow balancing problems, introduces a virtual impedance model:

[0136] Insert virtual impedance at the region boundary , Let R represent the virtual impedance at the boundary of the insertion region, where R represents the resistive component (real part) of the virtual impedance, used to simulate the active power loss characteristics of the boundary line, and X represents the reactive component (imaginary part) of the virtual impedance, used to correct the power flow equations.

[0137]

[0138] In the power flow equations of the virtual impedance model, and The active and reactive power at the region boundary. , For voltage amplitude and phase angle, , For nodes Voltage amplitude and phase angle, It is the reactance component of the virtual impedance; Represents the sine function. Representing the cosine function, this model modifies the power flow equations by inserting virtual impedances at the region boundaries, providing a unified mathematical expression for improving the ADMM algorithm and supporting distributed solutions to cross-regional power flow balance problems.

[0139] The local optimization problem for region k is defined through the ADMM iterative process:

[0140]

[0141] During the iterative process of the ADMM algorithm, in the local optimization problem, For the local objective function of the sub-region, As a penalty factor, and For globally consistent variables and Lagrange multipliers, It is a sub-region The local security constraints are input as sub-region decision variables. The output is the constraint deviation (which must meet the following conditions). ), " This refers to the core principle of safe operation of the power system, namely the single fault safety principle.

[0142] Alternating updates:

[0143]

[0144] argmin represents finding the independent variable within a given range of variables that minimizes the objective function. In the update formula of the ADMM iterative process, Let f represent the augmented Lagrangian function, and let f represent the number of sub-regions into which the power grid is dynamically divided. For sub-regional decision variables, It is a globally consistent variable. For Lagrange multipliers, It is a penalty factor; this iterative mechanism decomposes the global optimization problem into parallel sub-problems through a loop of sub-region local optimization, global variable aggregation, and multiplier update, thereby achieving efficient collaborative optimization of large-scale power grids.

[0145]

[0146] Denotes the iterative convergence error of the globally consistent variable, where This is the threshold for algorithm convergence. Its value ranges from 0 to 1.5, and the threshold is selected based on the maximum and minimum average error of the optimized power grid data.

[0147] Define communication health indicators through flexible communication – autonomous decision-making, communication status monitoring:

[0148]

[0149] In the communication health indicators of the dual-mode switching mechanism, For communication health, It is the delay weighting coefficient. This is the maximum allowable latency; this metric monitors the communication status in real time through the quantitative calculation of packet loss rate and latency. The islanding mode is triggered in case of communication failure to ensure the stable operation of the power grid.

[0150] RMPC control in island mode, sub-region Autonomous solution of robust optimization problems:

[0151]

[0152] In the RMPC control objective function in islanded mode, To control the sequence, the formula " "" indicates the set of uncertainties All disturbances Take the maximum value. This represents a time index, with a value range of [value range missing]. arrive Used to mark discrete time points within the prediction time domain. This indicates the prediction time domain length, i.e., from the current time. Start predicting the future time steps. It is a set of uncertainties that includes events such as extreme load fluctuations. Indicates its transpose; Indicates time The state vector contains parameters such as node voltage, frequency, and distributed generation output. It is a moment The control action vector, Indicates time Uncertainty disturbances belong to sets , Represents the state transition matrix. Represents the control input matrix. This represents the perturbation input matrix. , These represent the lower and upper limits of the node voltage, used to constrain the voltage within a safe range.

[0153] This function minimizes control costs by minimax optimization, taking into account the worst-case scenario of all possible disturbances, and ensures that the sub-region autonomously maintains voltage and frequency stability when communication is interrupted.

[0154] Historical scene database matching, real-time status Similarity calculation with the j-th scene in the scene library:

[0155]

[0156] In the similarity calculation of historical scene matching, It is an exponential function. Real-time status With historical scenes similarity, The formula controls the rate of similarity decay by quantifying the Euclidean distance difference of state vectors through an exponential function, quickly matching the most similar historical scenarios, loading preset control strategies, and achieving rapid response in fault scenarios.

[0157] Multi-objective policy optimization is achieved through offline pre-training of underlying PPO:

[0158] By training an expert network with historical data and extreme scenarios, a scheduling strategy that balances economy, security, and response speed can be learned.

[0159] Markov Decision Process (MDP) modeling: The state space S contains the real-time state of the power grid.

[0160]

[0161] It is a moment The state vector; The node voltage amplitude directly reflects the grid voltage level. It represents the system frequency and reflects the active power balance. It is the output of distributed power sources; It is in a state of energy storage and charging.

[0162] Reward function design:

[0163]

[0164] in, For a moment Instant rewards; Weighting coefficients used to adjust for the impact of power generation costs. Weighting coefficients used to adjust the effects of voltage deviation. Weighting coefficients used to adjust the influence of changes in motion; Represents the cost of electricity generation; It is the real-time voltage value. Compared with reference value The 2-norm of the deviation; For the current control action Actions at the previous moment The 2-norm of the variable. This formula transforms the economic cost of power grid dispatch, voltage security, and regulation stability into penalty values, which the reinforcement learning agent minimizes during training. We learn optimization strategies that take into account "low power generation cost, small voltage deviation, and stable control action" to adapt to the multi-objective operation requirements of the power grid.

[0165] Define the policy network through PPO policy optimization. Value network The optimization objective is:

[0166]

[0167] It is the pruning optimization objective function of the PPO algorithm. It ensures training stability by limiting the update range of the old and new strategies and averaging the loss values ​​of multiple training samples. The parameter is The policy network outputs in the state Next action The probability of. This represents the probability output of the old policy network, used to calculate the ratio of the old to new policies to limit the update magnitude. Indicates time Control action vectors, such as generator output regulation and energy storage charging and discharging commands. Indicates time The power grid state vector.

[0168] The advantage function in PPO strategy optimization is: middle, As the dominant value, For instant rewards, It is the output of the value network; this function measures the merits of the current action relative to the average policy by accumulating the difference between the discounted reward and the value function, providing a basis for policy updates and improving the efficiency and stability of reinforcement learning. This is a clipping function that restricts the input values ​​to a specific range. Within the range. The trimming range parameter controls the allowable fluctuation range of the ratio between the old and new strategies. The parameter is The current policy network, The parameter is The old policy network (i.e., the policy network before the update).

[0169] Upper-layer online rolling optimization, robust adjustment of opportunity constraints, and a real-time prediction-based dynamic correction strategy ensure a high probability of meeting safety constraints. Ultra-short-term prediction integration is also included. Let the current time be... Predicting the time domain (15 minutes) New Energy Output Forecast Load forecasting Generated using LSTM.

[0170] Construct a rolling optimization problem using a chance-constrained robust model:

[0171]

[0172] Indicates the prediction time domain (from time...) arrive The sequence of control actions in a chance-constrained robust model The conditional value of risk (95% confidence level) representing voltage deviation, combined with the objective function. While minimizing the expected control cost, through Control the risk of voltage deviation in extreme scenarios and ensure that safety constraints are met with a 95% probability. Indicates voltage amplitude. Indicates time The system frequency, This represents the vector of control cost coefficients. express transpose, m represents the risk aversion coefficient, and the adjusted conditional value at risk (VaR) The weight of the item, It is a benchmark value for measuring the stability of grid node voltage and is used to assess actual voltage. Deviation from the target state. ( ) represents the power flow equations describing the relationship between the state and control actions of the power grid system. Indicates time The state vector, Indicates time The control action vector, Indicates the system's rated frequency. It is the unit of frequency, "Hertz". H represents the current time, and H represents the prediction time domain length. and These represent the maximum and minimum values ​​of the node voltage amplitude, respectively.

[0173] Attention weights are defined by adjusting the gating network parameters online. Corresponding to the One expert adjusted the rules:

[0174]

[0175] This indicates that after one parameter update, the first... e experts at the time step The latest weights, Indicates at time step The gating network is the first A gated network with decision weights assigned to each expert subnetwork. The online adjustment formula for the gated network parameters... For the first The attention weights of e experts, It's the learning rate. This indicates finding the partial derivative. To robustly optimize the loss function, this formula uses gradient descent to dynamically adjust expert weights based on real-time optimization results, thereby improving the adaptability of the gated network to the current power grid scenario.

[0176] The policy is forced to converge in a flat region of the parameter space, enhancing its robustness against disturbances. To combat disturbance generation, a disturbance δ is injected into the state observations:

[0177]

[0178] Indicates the disturbance amplitude constraint Below, "argmax" represents the perturbation vector that maximizes the loss function L, indicating the independent variable that maximizes the objective function under given constraints. In the perturbation generation formula for adversarial training, To constrain the amplitude of the disturbance, For loss function, This indicates that it includes real-time operating parameters of the power grid. For real-world actions, this formula generates perturbation-injected state observations using the fast gradient sign method, forcing the expert network to find a "flat" extreme value region in the parameter space, thus improving robustness to unknown perturbations. By modifying the PPO loss function and adding a KL divergence penalty term, the consistency of policy distribution before and after the perturbation is constrained, ensuring decision stability under perturbation.

[0179] Approximate solution using FGSM (Fast Gradient Sign Method):

[0180]

[0181] in, To counteract disturbances; It controls the amplitude of the disturbance; This is the sign function, which outputs the sign of the gradient direction. Indicates the direction in which the perturbation is added; It is the gradient of the state with respect to the loss function. The action distribution output by the policy network, The target action.

[0182] This formula calculates the gradient of the state with respect to the loss, generates a small perturbation along the direction of the steepest gradient sign, injects it into the original state, and then trains the policy. This allows the policy to make stable decisions even under perturbed conditions, improving the robustness of the policy in scenarios such as power grid dispatch. For example, it can resist state interference caused by fluctuations in new energy power and communication noise, and enhance the policy's ability to cope with uncertainties in the actual power grid.

[0183] By adversarial training objectives, the PPO loss function is modified by adding an adversarial perturbation penalty term:

[0184]

[0185] This represents the loss function for adversarial training, specifically the PPO loss modification during adversarial training. It is the original PPO loss. To counter the weight, The mathematical expectation operator, It is the KL divergence; This represents a policy network with parameter θ. This represents the real-time state vector of the power grid. The loss function, by introducing a KL divergence penalty term for the policy distribution before and after the disturbance, forces the policy to converge in the flat region of the parameter space, thereby improving the generalization ability to unknown disturbances.

[0186] Flattening minimization optimization, through the SAML (Sharpness-Aware Minimization) optimizer, jointly minimizes the loss value and the sharpness of the loss surface:

[0187]

[0188] In the optimization of flat minimization The amplitude does not exceed The disturbance This indicates a constraint on the magnitude of parameter perturbation. The parameters represent the policy network. Let the loss function be denoted as . This optimization finds a flat extremum that is insensitive to perturbations by maximizing the loss under perturbations and minimizing the maximum value, and then updates the direction of the parameters accordingly. This further enhances the robustness of the model.

[0189] To address the strong randomness and unknown disturbances in new energy sources, a two-layer reinforcement learning framework is proposed. The bottom layer employs the Proximal Policy Optimization (PPO) algorithm for offline pre-training of the expert network. The training dataset includes tens of thousands of historical operational data sets and simulated extreme scenarios (such as multi-machine cascading failures caused by typhoons, and a sudden 50% drop in photovoltaic power). A reward function design (balancing economic cost, safety margin, and adjustment speed) guides the agent to learn the optimal policy. The upper online optimization layer performs rolling updates every 5 minutes: based on ultra-short-term new energy forecast results (15-minute scale) and load fluctuation trends, it dynamically adjusts the attention weight allocation parameters of the gating network and generates scheduling instructions through a robust optimization model with opportunistic constraints, ensuring that hard constraints such as voltage deviation and frequency fluctuations are met with a 95% probability. For unmodeled disturbances, an adversarial training mechanism is introduced: Gaussian noise and random equipment failures are injected during training, forcing the expert network to find flat extreme value regions in the parameter space, thereby improving its generalization ability to unknown scenarios.

[0190] Example 2, refer to Figures 3-6 This is the second embodiment of the present invention, which provides a power grid control method based on a hybrid expert network. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through experiments.

[0191] Verification of the robustness of power grid operation, control algorithms, including "a method for robust operation of parallel power grid optimization control based on hybrid expert networks", "model predictive control", "distributed control algorithm" and "reinforcement learning algorithm".

[0192] The robustness index of "a method for robust operation of parallel power grid optimization control based on hybrid expert networks" is the highest.

[0193] Figure 3 This study demonstrates the performance of different control algorithms in terms of robustness to power grid operation. The parallel control method based on hybrid expert networks exhibits the highest robustness, better able to cope with uncertainties in power grid operation, and ensure stable grid operation. The robustness of model predictive control, distributed control algorithms, and reinforcement learning algorithms decreases in that order.

[0194] Verification of power grid operation stability during communication failures, control algorithms, including "a method for robust operation of parallel power grid optimization control based on hybrid expert networks", "model predictive control", "distributed control algorithm" and "reinforcement learning algorithm".

[0195] The method for robust operation of parallel power grid optimization control based on hybrid expert networks has the highest stability index.

[0196] Figure 4 In the event of communication failures, the parallel control method based on hybrid expert networks still maintains high grid operation stability, followed by distributed control algorithms and model predictive control, while reinforcement learning algorithms show relatively low stability. This indicates that the proposed method has stronger adaptability and stability assurance capabilities under abnormal conditions such as communication failures.

[0197] Verification of power flow fluctuations and grid operation stability, control algorithms, including "a robust operation method for parallel grid optimization control based on hybrid expert networks", "model predictive control", "distributed control algorithm" and "reinforcement learning algorithm".

[0198] The method for robust operation of parallel power grid optimization control based on hybrid expert networks has the highest stability index.

[0199] Figure 5 Under power flow fluctuations, the parallel control method based on hybrid expert networks exhibits the best grid operation stability and can effectively cope with power flow fluctuations. Reinforcement learning algorithms also demonstrate good stability, while model predictive control and distributed control algorithms show relatively lower stability. This indicates that this method can better maintain the stable operation of the power grid under power flow fluctuation scenarios.

[0200] Verification of grid stability under voltage sag and wind power fluctuations, control algorithms, including "a robust operation method for parallel grid optimization control based on hybrid expert networks", "model predictive control", "distributed control algorithm" and "reinforcement learning algorithm".

[0201] The method for robust operation of parallel power grid optimization control based on hybrid expert networks has the highest stability index.

[0202] Figure 6 Under voltage sag and wind power fluctuation conditions, the "method for robust operation of parallel grid optimization control based on hybrid expert network" has the highest stability index, indicating that the method can maintain high grid stability under voltage sag and wind power fluctuation conditions.

[0203] Figure 3 , Figure 4 , Figure 5 and Figure 6 This is a comparison chart of the effects of the present invention considering comprehensive evaluation indicators and other methods; wherein, Figure 3 The parallel power grid optimization control based on a hybrid expert network proposed in this invention exhibits significantly better robustness in power grid operation than other comparative optimization algorithms. Figure 4 , Figure 5 and Figure 6This invention demonstrates that the parallel power grid optimization control based on a hybrid expert network proposed in this invention can effectively improve the stability of power grid operation under communication failures, power flow fluctuations, voltage drops, and wind power fluctuations.

[0204] Example 3 is the third embodiment of the present invention, which differs from the previous two embodiments in that:

[0205] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0206] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0207] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0208] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0209] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A power grid control method based on a hybrid expert network, characterized in that: include, Acquire power grid data, including node voltage, load demand, generator output, and energy storage charging and discharging data; Based on the power grid data, a multimodal data-driven intelligent scheduling model is constructed. The intelligent scheduling model includes an expert subnetwork consisting of a steady-state control model, a deep reinforcement learning model, a model predictive control model, a new energy consumption model, and a dynamic security assessment. Expert weights are dynamically allocated through a gating network to achieve adaptive decision-making. By quantifying electrical coupling, a decoupling-coordination mechanism for the wide-area power grid is constructed. The power grid is divided into electrically weakly coupled sub-regions. Each sub-region is optimized in parallel based on local expert strategies, and cross-regional power balance and safety boundary coordination are achieved through global consistency constraints. Based on the aforementioned intelligent scheduling model and decoupling-coordination mechanism, through online evolution of data-knowledge fusion, a two-layer reinforcement learning framework is adopted to dynamically adjust the expert network parameters. The two-layer reinforcement learning framework includes a bottom-layer offline pre-training and an upper-layer online rolling optimization to cope with the uncertainty interference of load mutations and new energy fluctuations.

2. The power grid control method based on a hybrid expert network as described in claim 1, characterized in that: The construction of the multimodal data-driven intelligent scheduling model includes, A steady-state control model is constructed, integrating the economic optimization of deep reinforcement learning with the real-time safety constraints of model predictive control. The objective function is a weighted sum of economic costs and safety constraints, expressed as follows: in, It is for variables This minimizes the summation expression. Indicates time The state vector; The scheduling period; For dynamic weights, For time step index; The economic cost objective function for deep reinforcement learning, The objective function for safety constraints in model predictive control; T-1 represents the start time step within the scheduling period; T-1 represents the last time step within the scheduling period. The deep reinforcement learning model adopts the Q-learning framework and defines the action value function, expressed as follows: in, Indicates the expected cumulative reward. Represents the mathematical expectation; This represents the real-time status of the power grid. For control commands, It is a discount factor; For a moment +k instant reward; Indicates the time offset; The whole is expressed as a long-term cumulative reward expression for mathematical expectation; The predictive control model is solved through rolling optimization and is expressed as follows: in, To predict the control sequence in the time domain, H is the prediction time domain length. and These are the cost coefficients for power generation and energy storage, respectively. and This represents the generator output and energy storage charging / discharging power at the corresponding moment.

3. The power grid control method based on a hybrid expert network as described in claim 2, characterized in that: The new energy consumption model employs a spatiotemporal graph convolutional network to improve the accuracy of ultra-short-term power prediction. include, Constructing the spatiotemporal diagram of new energy power plants G=( , ):node , Indicates the first A scenic station, Represents the set of all nodes and edges. , Represents a node and The connection between them Represents the set of connections between nodes; Spatial convolution aggregates neighborhood node information, represented as follows: in, For the first Layer features, For the first Layer features; Spatial convolution weights; ReLU is the activation function; Represented as an adjacency matrix; Temporal dependencies are captured through temporal convolution, represented as follows: in, It is the first convolution after temporal convolution in a spatiotemporal graph convolutional network. Layer feature output; 1D is for capturing temporal features through convolution, where Conv1D represents a one-dimensional convolution operation. These are the temporal convolution weights.

4. The power grid control method based on a hybrid expert network as described in claim 3, characterized in that: Constructing a decoupling-coordination mechanism for wide-area power grids includes, Electrical Coupling , is represented as , in, Gather for the neighbors, It is an index variable. For electrical coupling, represents the element of the admittance matrix, and | represents taking the absolute value; The power grid is partitioned by minimizing the coupling between regions, which is expressed as follows: Regional size constraints, expressed as, Power balance constraints are expressed as follows: in, This represents the node of the first sub-region; For the set of sub-region nodes, For the number of partitions, This indicates the introduction of constraints. This indicates the maximum number of nodes allowed in a single sub-region. Represents a node The generator output, This represents the total load demand of the entire power grid; Kalman filtering eliminates noise, as shown below. in, It is a quantitative indicator of the state of a sub-region. It is the state transition matrix. For PMU measurement data, It is the observation matrix. For Kalman gain.

5. The power grid control method based on a hybrid expert network as described in claim 4, characterized in that: Expert weights are dynamically assigned through a gating network, including: Expert weights are generated using a multilayer perceptron, and are represented as follows: Normalized exponential function, attention function By calculating the query matrix Key matrix Value matrix The degree of correlation, Representation matrix transpose, is the dimension of the key matrix.

6. The power grid control method based on a hybrid expert network as described in claim 5, characterized in that: The two-layer reinforcement learning framework includes, The underlying layer performs offline pre-training through near-end policy optimization, defining a reward function as follows: in, For a moment Instant rewards; Weighting coefficients used to adjust for the impact of power generation costs. Weighting coefficients used to adjust the effects of voltage deviation. Weighting coefficients used to adjust the influence of changes in motion; Represents the cost of electricity generation; It is the real-time voltage value. Compared with reference value The 2-norm of the deviation; For the current control action Actions at the previous moment The 2-norm of the change; The optimization objective is denoted as, in, Optimize the objective function for pruning; ; Conditional control variables; The parameter is The policy network outputs in the state Next action The probability, This represents the probability output of the old policy network; Indicates time The control action vector; Indicates time The power grid state vector; This is the dominant value.

7. The power grid control method based on a hybrid expert network as described in claim 6, characterized in that: The online evolution includes adversarial training to improve generalization to unknown perturbations, including, Adversarial perturbations are generated using the fast gradient sign method, denoted as follows: in, To counteract disturbances; It is a hyperparameter that controls the amplitude of the disturbance; This is the sign function, which outputs the sign of the gradient direction. Indicates the direction in which the perturbation is added; It is the gradient of the state with respect to the loss function. The action distribution output by the policy network. The target action.

8. The power grid control method based on a hybrid expert network as described in claim 7, characterized in that: The online rolling optimization includes dynamically adjusting expert network parameters through real-time data feedback and coordinating the security boundaries between sub-regions under global consistency constraints.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the power grid control method based on a hybrid expert network as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the power grid control method based on a hybrid expert network as described in any one of claims 1 to 8.