Power system dispatching control method, device, equipment, medium and program product

By converting the heterogeneous diagram structure of the power system into isomorphic diagrams and splicing them, combined with the deep reinforcement learning model of graph attention, the problem that traditional methods are difficult to provide accurate scheduling control is solved, and precise scheduling control is achieved in the event of power system failure.

CN118970923BActive Publication Date: 2025-07-01CHINA THREE GORGES CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411043804.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-07-01
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

Traditional graph convolutional networks are difficult to train heterogeneous information networks and cannot provide accurate scheduling and control guidance when power system failures.

Method used

By obtaining the heterogeneous graph structure of the power system in the target scenario, using the pre-constructed transformation matrix of different metapaths to convert it into isomorphic graphs, and after splicing, input the graph attention depth reinforcement learning model to output the value of different scheduling control strategies, and use the most valuable strategies for scheduling control.

Benefits of technology

It provides accurate scheduling control guidance when a power system fails, can accurately capture the topological connection relationship of the power system, and improves the accuracy and effectiveness of scheduling control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118970923B_ABST
    Figure CN118970923B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of electronic technologies, and discloses a dispatching control method, device, equipment, medium and program product for a power system. The method provided by the present invention converts the heterogeneous graph structure of the power system in a target scenario into a homogeneous graph, and splices the homogeneous graphs to obtain a spliced graph structure, which is convenient for subsequent feature recognition; calculates the values corresponding to different regulation strategies in the target scenario by using a pre-constructed graph attention deep reinforcement learning model; regulates the power system by using the regulation strategy with the highest value. The graph attention deep reinforcement learning model can accurately capture the topological connection relationship of the power system, and outputs the values of different dispatching control strategies based on the spliced graph structure, solving the problem that accurate regulation guidance cannot be provided when a fault occurs in the power system in the related art. An interpretable method for the graph deep reinforcement learning model is proposed, solving the problem of lack of credibility in the application of the related art in the field of power systems.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic technologies, and particularly to a scheduling control method, device, equipment, medium and program product for a power system. Background Art

[0002] The power system is a multi-layer network structure system. Due to the multi-dimensionality and heterogeneous nature of grid data, the data in the power system is complex and intertwined. The transient data of the power system belongs to non-Euclidean data, which is a kind of graph data structure containing components (nodes) and topological link relationships (edges). In order to study the rich heterogeneous semantic information in the power system and achieve smooth data flow and information sharing, constructing a heterogeneous information network for the power system has become a popular means. However, most of the existing data-driven methods are mainly designed for Euclidean data analysis. For example, deep neural networks usually model the grid operating state as a long vector composed of feature vectors. These methods often fail to effectively capture the topological connection relationships of the power system, and it is difficult to train a heterogeneous information network with traditional graph convolutional networks, and it is impossible to provide accurate scheduling control guidance when a fault occurs in the power system. Summary of the Invention

[0003] In view of this, the present invention provides a scheduling control method, device, equipment, medium and program product for a power system to solve the problem that it is difficult to train a heterogeneous information network with traditional graph convolutional networks and it is impossible to provide accurate scheduling control guidance when a fault occurs in the power system.

[0004] In a first aspect, the present invention provides a scheduling control method for a power system, the method comprising: obtaining a heterogeneous graph structure of the power system in a target scenario, the heterogeneous graph structure being composed of a feature matrix and an adjacency matrix, the feature matrix being used to represent the feature vectors of multiple nodes, the adjacency matrix being used to represent the association relationships between different nodes, the heterogeneous graph structure containing multiple meta-paths, the meta-paths being used to represent the paths connecting different types of nodes; converting the heterogeneous graph structure by using the conversion matrices respectively corresponding to different pre-constructed meta-paths to obtain multiple homogeneous graphs; splicing the multiple homogeneous graphs to obtain a spliced graph structure; inputting the spliced graph structure into a pre-constructed graph attention deep reinforcement learning model so that the graph attention deep reinforcement learning model outputs the values corresponding to different scheduling control strategies of the power system in the target scenario; and performing scheduling control on the power system by using a target scheduling control strategy with the highest value among different scheduling control strategies.

[0005] The dispatching control method of the power system provided by the present invention uses the conversion matrices respectively corresponding to different meta-paths constructed in advance to convert the heterogeneous graph structure, obtaining multiple homogeneous graphs; splicing the multiple homogeneous graphs to obtain the spliced graph structure; inputting the spliced graph structure into the pre-constructed graph attention deep reinforcement learning model, so that the graph attention deep reinforcement learning model outputs the values respectively corresponding to different dispatching control strategies of the power system under the target scenario; using the target dispatching control strategy with the highest value among different dispatching control strategies to perform dispatching control on the power system. The method provided by the present invention, by converting the heterogeneous graph structure of the power system under the target scenario into a homogeneous graph and splicing the homogeneous graphs to obtain the spliced graph structure, facilitates subsequent feature recognition; uses the pre-constructed graph attention deep reinforcement learning model to calculate the values respectively corresponding to different regulation strategies under the target scenario; uses the regulation strategy with the highest value to regulate the power system. The graph attention deep reinforcement learning model can accurately capture the topological connection relationship of the power system, and based on the spliced graph structure, outputs the values of different dispatching control strategies, solving the problem in the related technology that accurate dispatching control guidance cannot be provided when the power system fails.

[0006] In an alternative embodiment, the graph attention deep reinforcement learning model includes a graph attention deep learning model sub-model and a reinforcement learning sub-model; the graph attention deep learning model sub-model is used to implement the self-attention mechanism for each node in the spliced graph structure to obtain the first weight matrix respectively corresponding to each node under different meta-paths; based on the first weight matrix respectively corresponding to each node under different meta-paths, add the multi-head attention mechanism to the corresponding nodes in the spliced graph structure to obtain the target feature vector respectively corresponding to the corresponding nodes under different meta-paths; determine the weights respectively corresponding to different meta-paths based on the target feature vector respectively corresponding to each node under different meta-paths; normalize the weights respectively corresponding to different meta-paths and the target feature vector respectively corresponding to each node under different meta-paths to obtain the processing result, and output the processing result; the reinforcement learning sub-model includes a state space, an action space and a reward function, and the reinforcement learning sub-model is used to use the processing result output by the graph attention deep learning model sub-model to determine the corresponding target state, determine the corresponding multiple actions based on the target state, calculate the value corresponding to each action to obtain the calculation result, and output the calculation result.

[0007] The method provided by this alternative embodiment, for the dynamically changing graph structure data, the features of the nodes may change. By introducing the attention mechanism, different weights can be adaptively assigned according to the importance between the nodes, so as to model the complex patterns and information in the graph structure, help the model focus on the relatively more important parts in the input data, and thus reduce the computational cost of processing irrelevant information.

[0008] In an alternative embodiment, the method includes: inputting the heterogeneous graph structure and the graph attention deep reinforcement learning model of the power system in a target scenario into a pre-constructed graph interpretation model, so that the model outputs a corresponding target subgraph.

[0009] For the method provided in this alternative embodiment, the target subgraph is used to characterize the features of the nodes and edges that have the greatest impact on the prediction. The target subgraph can provide a comprehensive and intuitive explanation for the model, provide a more efficient decision-making basis for power system dispatchers, avoid the "black box" characteristic of the artificial intelligence model, and enable dispatchers to understand and trust the decision-making results of the artificial intelligence model.

[0010] In an alternative embodiment, the graph interpretation model calculates the target subgraph through the following steps: transforming the heterogeneous graph structure into multiple homogeneous graphs; determining the target set of each node in each homogeneous graph, where the target set is used to characterize all node sets in the corresponding homogeneous graph whose shortest path starting from the corresponding node does not exceed a preset value; determining the first subgraph of the corresponding homogeneous graph based on the target set of each node in each homogeneous graph, and the first subgraph includes a feature matrix and an adjacency matrix; respectively splicing the feature matrices and adjacency matrices of the first subgraphs of the multiple homogeneous graphs to obtain a spliced feature matrix and an adjacency matrix; constructing multiple feature selectors based on the spliced feature matrix and multiple edge selectors based on the spliced adjacency matrix; using the multiple feature selectors to respectively perform feature screening on the spliced feature matrix to obtain a target feature matrix after feature screening by each feature selector; using the multiple edge selectors to respectively perform edge screening on the spliced adjacency matrix to obtain a target adjacency matrix after edge screening by each edge selector; determining multiple second subgraphs based on the target feature matrix after feature screening by each feature selector and the target adjacency matrix after edge screening by each edge selector; using the graph attention deep reinforcement learning model to determine the scheduling control strategy of each second subgraph; and determining the second subgraph with the smallest difference from the target scheduling control strategy among the multiple second subgraphs as the target subgraph.

[0011] In an alternative embodiment, using multiple edge selectors to respectively perform feature screening on the adjacency matrix to obtain a target adjacency matrix after feature screening by each edge selector includes: multiplying different feature selectors by relevant elements of the spliced feature matrix respectively to obtain a first matrix corresponding to each feature selector; determining first target elements in the first matrix corresponding to each feature selector whose element values are less than a preset threshold; and deleting the node feature vectors corresponding to the target elements in each first matrix from the spliced feature matrix to obtain a target feature matrix after feature screening by the corresponding feature selector.

[0012] In an alternative embodiment, multiple edge selectors are used to separately screen the edges of the spliced adjacency matrix to obtain a target adjacency matrix after edge screening by each edge selector, including: multiplying the relevant elements of the spliced feature matrix by different edge selectors respectively to obtain a second matrix corresponding to each edge selector; determining second target elements in the second matrix corresponding to each edge selector whose element values are less than a preset threshold; deleting the edge features corresponding to the second target elements in each second matrix from the spliced adjacency matrix to obtain a target adjacency matrix after edge screening by the corresponding edge selector.

[0013] In a second aspect, the present invention provides a dispatching control device for a power system, the device includes: an acquisition module, configured to acquire a heterogeneous graph structure of the power system in a target scenario, the heterogeneous graph structure is composed of a feature matrix and an adjacency matrix, the feature matrix is used to characterize the feature vectors of multiple nodes, the adjacency matrix is used to characterize the association relationship between different nodes, and the heterogeneous graph structure contains multiple meta-paths, and the meta-paths are used to characterize the paths connecting different types of nodes; a conversion module, configured to convert the heterogeneous graph structure by using conversion matrices respectively corresponding to different pre-constructed meta-paths to obtain multiple homogeneous graphs; a splicing module, configured to splice the multiple homogeneous graphs to obtain a spliced graph structure; a first determination module, configured to input the spliced graph structure into a pre-constructed graph attention deep reinforcement learning model, so that the graph attention deep reinforcement learning model outputs the values corresponding to different dispatching control strategies of the power system in the target scenario; a dispatching control module, configured to perform dispatching control on the power system by using a target dispatching control strategy with the highest value among different dispatching control strategies.

[0014] In a third aspect, the present invention provides a computer device, including: a memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the dispatching control method of the power system in the first aspect or any corresponding embodiment thereof.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the dispatching control method of the power system in the first aspect or any corresponding embodiment thereof.

[0016] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the dispatching control method of the power system in the first aspect or any corresponding embodiment thereof. Description of the Drawings

[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of a dispatching control method for a power system according to an embodiment of the present invention;

[0019] Figure 2 It is a schematic flowchart of a dispatching control method for another power system according to an embodiment of the present invention;

[0020] Figure 3 It is a structural block diagram of a dispatching control device for a power system according to an embodiment of the present invention;

[0021] Figure 4 It is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. Specific Embodiments

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0023] Currently, most existing data-driven methods are mainly designed for Euclidean data analysis. For example, deep neural networks usually model the power grid operation state as a long vector composed of feature vectors. These methods often fail to effectively capture the topological connection relationship of the power system, and traditional graph convolutional networks are difficult to train heterogeneous information networks and cannot provide accurate dispatching control guidance when faults occur in the power system.

[0024] In view of this, a scheduling control method for a power system provided by an embodiment of the present application can be applied to a server to implement the scheduling control of the power system. The method provided by the present invention converts the heterogeneous graph structure of the power system in the target scenario into a homogeneous graph, and splices the homogeneous graphs to obtain the spliced graph structure, which is convenient for subsequent feature recognition; calculates the values corresponding to different regulation strategies in the target scenario by using the pre-constructed graph attention deep reinforcement learning model; and regulates the power system by using the regulation strategy with the highest value. The graph attention deep reinforcement learning model can accurately capture the topological connection relationship of the power system and output the values of different scheduling control strategies based on the spliced graph structure, solving the problem in the related art that accurate scheduling control guidance cannot be provided when the power system fails.

[0025] According to an embodiment of the present invention, an embodiment of a scheduling control method for a power system is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0026] In this embodiment, a scheduling control method for a power system is provided, which can be used for the above-mentioned server. Figure 1 is a flowchart of a scheduling control method for a power system according to an embodiment of the present invention, as Figure 1 shown, the process includes the following steps:

[0027] Step S101, obtain the heterogeneous graph structure of the power system in the target scenario.

[0028] Exemplarily, the power system can be a power system that needs to be scheduled and controlled, and the target scenario can include, but is not limited to, the scenario when one or more components in the power system fail. The heterogeneous graph structure is composed of a feature matrix and an adjacency matrix. The feature matrix is used to represent the feature vectors of multiple nodes, and the adjacency matrix is used to represent the association relationship between different nodes. The heterogeneous graph structure contains multiple meta-paths, and the meta-paths are used to represent the paths connecting different types of nodes. In the active power correction control scenario of the present application embodiment, electrical components such as generators, loads, and transmission lines can be regarded as nodes, and the connection relationship between the electrical components is defined as an edge, so as to realize the graph regularization of the topological transformation grid state. The feature matrix in the proposed active power correction control scenario contains the active power of various nodes and the load rate on the transmission line. The adjacency matrix represents the connection relationship between nodes such as generators, loads, and transmission lines in the power grid. Taking the active power of each node and the load rate of the transmission line as the observation features, the feature matrix can be specifically expressed as:

[0029]

[0030] An adjacency matrix is established based on the connection relationships between various power devices, and the adjacency matrix can be expressed as:

[0031]

[0032] where P L , P TL and are the active power values on the load, transmission line, and generator respectively, and ρ is the transmission line load rate.

[0033] Step S102: Use the transformation matrices corresponding to different pre-constructed meta-paths to transform the heterogeneous graph structure to obtain multiple homogeneous graphs.

[0034] Exemplarily, a path connecting two types of objects in the heterogeneous graph structure is defined as a meta-path, and the transformation matrix W m (p) is used to transform the heterogeneous graph structure. When the number of meta-paths is P, P homogeneous graphs can be obtained through transformation.

[0035] Step S103: Stitch multiple homogeneous graphs to obtain a stitched graph structure.

[0036] Exemplarily, in the embodiments of the present application, the feature vector of the stitched node can be shown as the following formula:

[0037]

[0038] where, is the feature vector of node i in the l-th convolutional layer after stitching, || represents the stitching operation, represents the weight matrix, and H (l) (p) represents the input of the l-th convolutional layer in the homogeneous graph p.

[0039] Step S104: Input the stitched graph structure into a pre-constructed graph attention deep reinforcement learning model so that the graph attention deep reinforcement learning model outputs the values corresponding to different scheduling control strategies of the power system under the target scenario.

[0040] Exemplarily, in the embodiments of the present application, after the graph attention deep reinforcement learning model obtains the stitched graph structure, it performs feature recognition and calculation on the graph structure, and outputs the values corresponding to different scheduling control strategies of the power system under the target scenario.

[0041] Step S105: Use the target scheduling control strategy with the highest value among different scheduling control strategies to perform scheduling control on the power system.

[0042] The scheduling control method of the power system provided in this embodiment converts the heterogeneous graph structure of the power system in the target scenario into a homogeneous graph, and splices the homogeneous graphs to obtain the spliced graph structure, which is convenient for subsequent feature recognition; uses the pre-constructed graph attention deep reinforcement learning model to calculate the values corresponding to different regulation strategies in the target scenario; uses the regulation strategy with the highest value to regulate the power system. The graph attention deep reinforcement learning model can accurately capture the topological connection relationship of the power system and output the values of different scheduling control strategies based on the spliced graph structure, solving the problem that in the related art, accurate scheduling control guidance cannot be provided when the power system fails.

[0043] In this embodiment, a scheduling control method of a power system is provided, which can be used for the above-mentioned server. Figure 2 It is a flowchart of the scheduling control method of the power system according to an embodiment of the present invention, as Figure 2 shown, and the process includes the following steps:

[0044] Step S201, obtain the heterogeneous graph structure of the power system in the target scenario. The heterogeneous graph structure is composed of a feature matrix and an adjacency matrix. The feature matrix is used to represent the feature vectors of multiple nodes, and the adjacency matrix is used to represent the association relationship between different nodes. The heterogeneous graph structure contains multiple meta-paths, and the meta-paths are used to represent the paths connecting different types of nodes. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.

[0045] Step S202, use the pre-constructed transformation matrices corresponding to different meta-paths to transform the heterogeneous graph structure to obtain multiple homogeneous graphs. For details, please refer to Figure 1 Step S101 of the embodiment shown, which will not be elaborated here.

[0046] Step S203, splice the multiple homogeneous graphs to obtain the spliced graph structure. For details, please refer to Figure 1 Step S103 of the embodiment shown, which will not be elaborated here.

[0047] Step S204, input the spliced graph structure into the pre-constructed graph attention deep reinforcement learning model, so that the graph attention deep reinforcement learning model outputs the values corresponding to different scheduling control strategies of the power system in the target scenario. For details, please refer to Figure 1 Step S104 of the embodiment shown, which will not be elaborated here.

[0048] In some optional implementation manners, the graph attention deep reinforcement learning model includes a graph attention deep learning model sub-model and a reinforcement learning sub-model;

[0049] The sub-model of the graph attention deep learning model is used to implement the self-attention mechanism for each node in the spliced graph structure, obtaining the first weight matrix corresponding to each node under different meta-paths; adding the multi-head attention mechanism to the corresponding nodes in the spliced graph structure based on the first weight matrix corresponding to each node under different meta-paths, obtaining the target feature vector corresponding to the corresponding node under different meta-paths; determining the weights corresponding to different meta-paths based on the target feature vector corresponding to each node under different meta-paths; normalizing the weights corresponding to different meta-paths and the target feature vector corresponding to each node under different meta-paths, obtaining the processing result, and outputting the processing result.

[0050] Exemplarily, in the embodiments of the present application, the graph attention deep learning sub-model can be constructed through the following ideas: Construct a Graph Convolutional Network (GCN). GCN is a deep learning algorithm based on convolutional neural networks, used to process image data. Its core is to represent image data in the form of a graph and perform convolutional operations on the graph. GCN can process data in non-Euclidean spaces. Its input is a graph structure, and the graph G=(V, E) can be represented as a set of nodes and edges. Each node has its corresponding feature information. G=(V, E) can be described by the feature matrix and the adjacency matrix where S represents the number of input samples in one training step, N is the number of nodes V in the graph G=(V, E), and F is the number of node input features. The adjacency matrix represents the connection relationship between nodes in the graph G=(V, E). It represents the adjacency relationship between any two vertices. If they are adjacent, it is 1; if not, it is 0.

[0051] In GCN, the feature vector of each node in the graph G=(V, E) is convolved with the feature vectors of its adjacent nodes to obtain a new feature representation of the node. The information H (l+1) of each node in the next layer is obtained by combining the information H (l) of the previous layer itself and the information of adjacent nodes. Perform a linear transformation on the feature vector of each node, then perform a weighted sum of the feature vector of the node and the feature vectors of its adjacent nodes, and finally perform an operation of the activation function.

[0052]

[0053] where I is the identity matrix. To prevent ignoring the features of the node itself, add the identity matrix I to E. is the degree matrix of. The degree of each node refers to the number of nodes it is connected to. Η is the feature matrix of each layer, and l represents the number of convolutional layers. For the first layer, Η(0) = X. σ(·) is a non - linear activation function, and W (l) is the weight matrix of the l - th convolutional layer.

[0054] The GCN model assumes that the features of nodes in the model are invariant, that is, they have the same features at different positions and different time points in the graph. For dynamically changing graph - structured data, the features of nodes may change, and it is difficult for the GCN model to directly handle this situation. In this case, an attention mechanism can be introduced, which can adaptively assign different weights according to the importance between nodes. The Graph Attention Network (GAT) can more flexibly capture the relationships between nodes through the attention mechanism, improving the expressive ability of the model. It can learn the weights between different nodes, thus modeling the complex patterns and information in the graph structure, helping the model focus on the relatively more important parts of the input data, and thus reducing the computational cost of processing irrelevant information. At the same time, the embedding of the attention mechanism can enhance the transparency of the original model and the interpretability of the neural network, achieving intrinsic interpretability in model construction.

[0055] In the embodiments of this application, it is hoped to learn the importance between each node and its neighbor nodes in the power grid topology through the graph attention mechanism. The input Η of each layer in the GAT network (l) can be expressed as:

[0056]

[0057] where N is the number of nodes and F is the feature dimension of each node. To calculate the attention weight between node i and neighbor node j, an attention mechanism function a(·) is introduced, which is represented by a feed - forward neural network. The relevance between the current central node i and all other neighbor nodes j is obtained:

[0058]

[0059] where [·||·] represents the concatenation operation, and W l is the learnable weight matrix of the l - th attention layer. A shared linear transformation is parameterized through the weight matrix to generate a low - dimensional node feature representation, and this transformation is applied to each node to implement the attention mechanism of self - attention. e ij represents the original attention weight score between node i and neighbor node j. To better allocate weights, it is necessary to calculate the relevance e ij between the current central node and all other neighbor nodes, and softmax is used for unified normalization:

[0060]

[0061] Among them, Att ij is the weight matrix of the node, which reflects the regions or features that the model focuses on during the decision-making process, ensuring that the weighted sum of the weight coefficients of all neighbor nodes of the current central node is 1. The activation function used is Leaky ReLU. To further improve the expressive ability of the attention layer, a multi-head attention mechanism is introduced.

[0062]

[0063] Among them, N i is the set of neighbor nodes at a distance of 1 from node i, σ is a non-linear function, and K represents the number of heads in the multi-head attention mechanism. By introducing multiple groups of independent attention mechanisms, the multi-head attention mechanism calculates multiple different attention weight matrices simultaneously, enabling the model to learn multiple different feature representations. This further enhances the ability of the multi-head attention mechanism to distribute attention to multiple relevant features between the central node and neighbor nodes, making the learning ability of the system more powerful.

[0064] For the spliced graph structure, the new node feature vector is:

[0065]

[0066] Among them, || represents the splicing operation. Assuming the current key node is i, a shared linear transformation is parameterized through the weight matrix to increase the dimension of the features of the vertices, and this transformation is applied to each node to implement the attention mechanism of self-attention:

[0067]

[0068] Among them, a is a calculation function of the attention mechanism, which maps the spliced high-dimensional features to a real number. represents the importance of node j to node i under the meta-path p. To better allocate weights, the relevance calculated between the current central node and all other neighbor nodes is uniformly normalized using sotfmax:

[0069]

[0070] Among them, is the weight matrix of the node, ensuring that the weighted sum of the weight coefficients of all neighbor nodes of the current central node is 1. The activation function used is Leaky ReLU.

[0071] In the embodiments of the present application, to further improve the expressive ability of the attention layer, a multi-head attention mechanism is added. The feature vector of node i after adding the multi-head attention mechanism is shown as follows:

[0072]

[0073] Among them, N i is the set of neighbor nodes of node v i with a distance of 1. σ is a non-linear function, and K represents the number of heads of multi-head attention. Adding multiple sets of independent attention mechanisms enables the multi-head attention mechanism to allocate attention to multiple relevant features between the central node and neighbor nodes, making the learning ability of the system stronger.

[0074] Due to different choices of meta-paths, it is necessary to obtain the weights of different meta-paths. The weights of different meta-paths can be determined by the following formula:

[0075]

[0076] Among them, Att mp ∈R N×P is the weight matrix of meta-path P, Ω ∈ R F×M , μ ∈ R M . M is the attention size of the meta-path, Tanh is an activation function that maps the input between [-1, 1], and then the softmax function is used for normalization.

[0077] Furthermore, the output of the sub-model of the graph attention deep learning model can be obtained as shown in the following formula:

[0078]

[0079] Among them, O represents the output of the sub-model of the graph attention deep learning model, O ∈ R N×C , C is the number of categories. Att′ mp = R N ×P×1 is the quantity after Att mp is changed and an additional dimension is added. and Att′ mp are summed up in other dimensions except the second dimension.

[0080] The reinforcement learning sub-model includes a state space, an action space, and a reward function. The reinforcement learning sub-model is used to determine the corresponding target state by using the processing result output by the sub-model of the graph attention deep learning model, determine multiple corresponding actions based on the target state, calculate the value corresponding to each action to obtain the calculation result, and output the calculation result.

[0081] Exemplarily, the processing results output by the graph attention deep learning model sub-model may include the outputs of multiple convolutional layers. In the embodiments of the present application, the number of convolutional layers may be two, and the results of the two convolutional layers and are concatenated. The dimensions of the output features of the two convolutional layers are the same, where N is the number of nodes and C is the number of categories. To make the features of each node contain more information, the outputs of the two convolutional layers containing the attention mechanism are concatenated in the second dimension, and the concatenation result can be shown as follows:

[0082] H′ = (H (1) ⊕ H (2) ) | dim=1

[0083]

[0084] where ⊕ represents the element-wise addition at the corresponding positions. The concatenation of the convolutional layers can make the features of each node contain more information, so as to better describe the attributes and relationships of the nodes. At the same time, it can avoid losing some feature information in a single convolutional layer and improve the prediction performance of the model. After that, each node in the fully connected layer is connected to each node in the previous layer, and the node feature vectors are mapped to a low-dimensional space to obtain a new low-dimensional vector h′ to capture the global information between the nodes and learn the relationships between the nodes.

[0085] h′ = Linear(h) = [h1′, h2′,... h′ C′

[0086] where C′ is the output parameter set by the fully connected layer. After passing through the two convolutional networks containing the attention mechanism and the fully connected layer, a Deep Q-Learning (DQN) network structure is adopted. The specific DQN model architecture and algorithm process will not be introduced here. The values of the state and action are calculated through a competitive structure to obtain the output layer, that is, the Q value of the output action in the current state.

[0087] The DQN model constructs the corresponding state space, action space, and reward function for the active power correction control of the power system according to the Markov decision process to achieve the optimal control of the system.

[0088] Specifically, the state is defined as the observable information at the current time step, and the state space ​Describes the observed state of the agent in the power dispatch system, including the relevant characteristics of generators, loads, transmission lines, and buses. Specifically, it can include the active power of generators, loads, AC transmission lines, and buses, as well as the load rate on the transmission line. At each time step t, the environment generates an observation S t , Specifically, it can be expressed as:

[0089]

[0090] Where X represents the features in the set of states observed at time step t, and N e is the number of components of the observed feature. P G , P L , P TL , P B represent the active power of generators, loads, AC transmission lines, and buses respectively, and ρ represents the load rate of the transmission line.

[0091] Action space Contains all the action sets that the agent can execute. The common active power correction control actions include adjusting the active power output of generators, load shedding, and topological adjustment. The agent can perceive the state S of the environment t , and take action A t at time t to change the environment state.

[0092]

[0093] Among them, represent the h-th actions of adjusting the active power output of generators, load shedding, and topological adjustment actions at time step t respectively. h represents the category of the actions output by the model, and H is the number of actions, h ∈ H. The agent will set the reward value for the selected action and use the learned optimal policy to provide the optimal action for the environment according to the observed environment state to change the environment state. The goal is to apply the optimal action given the current state so that the agent can accumulate most of the reward values over time.

[0094] By designing the reward function to meet various objectives and constraints of security correction, which is used to describe the reward value obtained by the agent when executing a certain action in the current state. The reward function of active power correction control usually can include the penalty for transmission line power over-limit and the penalty for generator rescheduling adjustment amount. The penalty at time t of the system can be specifically expressed as:

[0095]

[0096] Among them, A represents the reward value for the normal operation of the system, and N L is the number of transmission lines, and N G is the number of adjustable generators. ρ L,i represents the load rate of the i-th transmission line. μ and ν are the penalty factors for transmission line overload and generator rescheduling adjustment amount respectively. The system is controlled in a time series manner, hoping that the power system can survive longer. Therefore, the immediate reward function r t can be defined as:

[0097]

[0098] Among them, λ is a negative number greater than the preset threshold.

[0099] Step S205, use the target scheduling control strategy with the highest value among different scheduling control strategies to perform scheduling control on the power system. For details, please refer to Figure 1 step S105 of the illustrated embodiment, which will not be elaborated here.

[0100] Step S206, input the heterogeneous graph structure and the graph attention deep reinforcement learning model of the power system under the target scenario into the pre-constructed graph interpretation model, so that the model outputs the corresponding target subgraph.

[0101] In some alternative embodiments, the graph interpretation model calculates the target subgraph through the following steps:

[0102] Step a1, transform the heterogeneous graph structure into multiple homogeneous graphs.

[0103] Exemplarily, in the embodiment of the present application, the initial heterogeneous graph is divided into P homogeneous graphs G m after the transformation matrix W p .

[0104] Step a2, determine the target set of each node in each homogeneous graph. The target set is used to represent all node sets in the corresponding homogeneous graph whose shortest path starting from the corresponding node does not exceed a preset value.

[0105] Step a3, determine the first subgraph of the corresponding homogeneous graph based on the target sets of each node in each homogeneous graph. The first subgraph includes a feature matrix and an adjacency matrix.

[0106] Exemplarily, the preset value can be determined according to requirements. In the embodiment of the present application, in order to obtain a more compact effect, for the graph G p for any node v i in it, determine the set of all nodes whose shortest path starting from v i does not exceed B (preset value), so as to obtain their B-hop neighborhood information to obtain the subgraph B is a positive integer. X S (p) and E S (p) are the characteristic matrix and the adjacency matrix of the first subgraph (B-hop subgraph), respectively.

[0107]

[0108] where v i , v j and v l are nodes in the subgraph . X S (p) is consistent with the input features of the original graph, specifically including the active power of generators, loads, AC transmission lines, and buses, as well as the load rate on the transmission lines. (P G , P L , P TL , P B , ρ) i is the eigenvector of node v j , and e jl is the edge connecting nodes v j and v l .

[0109] Step a4: Concatenate the characteristic matrices and adjacency matrices of the first subgraphs of multiple isomorphic graphs to obtain the concatenated characteristic matrix and adjacency matrix.

[0110] Exemplarily, by concatenating the characteristic matrices and adjacency matrices in P subgraphs, a new characteristic matrix X S (p) pre and a new adjacency matrix E S (p) pre are obtained.

[0111]

[0112] Step a5: Construct multiple feature selectors based on the concatenated characteristic matrix, and construct multiple edge selectors based on the concatenated adjacency matrix.

[0113] Exemplarily, in the embodiments of the present application, in order to identify which node features and edge information have important impacts on the model decision result, a feature selector F mask ∈ {0, 1} N×F and an edge selector E mask ∈ {0, 1} N×N are determined in the subgraph interpreter. N is the number of nodes in the graph , and F is the number of input features of each node.

[0114] ​Step a6: Use multiple feature selectors to separately perform feature screening on the concatenated feature matrix to obtain a target feature matrix after feature screening using each feature selector.

[0115] Specifically, step a6 includes:

[0116] Step a61: Multiply different feature selectors with relevant elements of the concatenated feature matrix respectively to obtain a first matrix corresponding to each feature selector.

[0117] Exemplarily, in the embodiments of the present application, multiple feature selectors F mask ∈{0,1} N×F are respectively multiplied with relevant elements in X S (p) to obtain a first matrix corresponding to each feature selector.

[0118] Step a62: Determine first target elements in the first matrix corresponding to each feature selector whose element values are less than a preset threshold.

[0119] Step a63: Delete the node feature vectors corresponding to the target elements in each first matrix from the concatenated feature matrix to obtain a target feature matrix after feature screening using the corresponding feature selector.

[0120] Exemplarily, the target elements in each first matrix are features less than the preset threshold. In the embodiments of the present application, the node feature vectors corresponding to the target elements in each first matrix are deleted from the concatenated feature matrix to obtain a target feature matrix after feature screening using the corresponding feature selector which contains features important for the model decision result.

[0121] Step a7: Use multiple edge selectors to separately perform edge screening on the concatenated adjacency matrix to obtain a target adjacency matrix after edge screening using each edge selector.

[0122] Specifically, step a7 includes:

[0123] Step a71: Multiply different edge selectors with relevant elements of the concatenated feature matrix respectively to obtain a second matrix corresponding to each edge selector.

[0124] Exemplarily, in the embodiments of the present application, different edge selectors are respectively multiplied with relevant elements of the concatenated feature matrix to obtain a second matrix corresponding to each edge selector

[0125] Step a72: Determine second target elements in the second matrix corresponding to each edge selector whose element values are less than a preset threshold.

[0126] Exemplarily, the preset threshold can be determined based on requirements.

[0127] Step a73, delete the edge features corresponding to the second target elements in each second matrix from the spliced adjacency matrix, and obtain the target adjacency matrix after edge screening using the corresponding edge selector.

[0128] Exemplarily, delete the edge features corresponding to the second target elements in each second matrix from the spliced adjacency matrix, and obtain the target adjacency matrix after edge screening using the corresponding edge selector

[0129] Step a8, determine multiple second subgraphs based on the target feature matrix after feature screening by each feature selector and the target adjacency matrix after feature screening by each edge selector.

[0130] Exemplarily, determine multiple second subgraphs based on the target feature matrix after feature screening by each feature selector and the target adjacency matrix after feature screening by each edge selector

[0131] Among them,

[0132] The second subgraph only contains the most important nodes and edges in the corresponding B-hop subgraph among them.

[0133] Step a9, use the graph attention deep reinforcement learning model to determine the scheduling control strategy for each second subgraph.

[0134] Exemplarily, use the graph attention deep reinforcement learning model to determine the scheduling control strategy feature matrix in each second subgraph and the adjacency matrix of the decision result

[0135]

[0136] The action space of the graph attention deep reinforcement learning model includes generator active power adjustment, load shedding, and topology adjustment.

[0137] Step a10, determine the second subgraph with the smallest difference from the target scheduling control strategy among the multiple second subgraphs as the target subgraph.

[0138] Exemplarily, assume is the structure most relevant to the decision result Y of the initial graph G=(V, E), then the overall goal of the subgraph interpreter is to minimize the information difference between the target subgraph and the original graph G=(V, E), as shown in the following formula:

[0139] Y(a G ,a L ,aT ) = Φ(X(P G , P L , P TL , P B , ρ), E)

[0140]

[0141] where Y(a G , a L , a T ) is the decision result (target scheduling control strategy) of the graph attention deep reinforcement learning model F with respect to the initial graph G = (V, E). IM(·) represents the total amount of information, and the concept of importance is formalized using IM(·). Since the amount of information in the original graph G = (V, E) is fixed, the final optimization goal is to maximize the information contained in the target subgraph , that is, to minimize the information difference between the target subgraph and the B-hop subgraph .

[0142]

[0143] where Y S is the decision result of the graph attention deep reinforcement learning model F with respect to the B-hop subgraph . By minimizing the difference between the target subgraph and the B-hop subgraph , the feature selector F mask ∈ {0, 1} N×F and the edge selector E mask ∈ {0, 1} N×N can be adjusted by feedback. Finally, the simplest but most important target subgraph structure

[0144] In this embodiment, a scheduling control device for a power system is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0145] This embodiment provides a scheduling control device for a power system, as Figure 3 shown, including:

[0146] An acquisition module 301, configured to acquire a heterogeneous graph structure of a power system in a target scenario. The heterogeneous graph structure is composed of a feature matrix and an adjacency matrix. The feature matrix is used to represent the feature vectors of multiple nodes, and the adjacency matrix is used to represent the association relationships between different nodes. The heterogeneous graph structure contains multiple meta-paths, and the meta-paths are used to represent the paths connecting different types of nodes;

[0147] A conversion module 302, configured to convert the heterogeneous graph structure by using conversion matrices respectively corresponding to different pre-constructed meta-paths to obtain multiple homogeneous graphs;

[0148] A splicing module 303, configured to splice the multiple homogeneous graphs to obtain a spliced graph structure;

[0149] A first determination module 304, configured to input the spliced graph structure into a pre-constructed graph attention deep reinforcement learning model, so that the graph attention deep reinforcement learning model outputs the values respectively corresponding to different scheduling control strategies of the power system in the target scenario;

[0150] A scheduling control module 305, configured to perform scheduling control on the power system by using a target scheduling control strategy with the highest value among different scheduling control strategies.

[0151] In some alternative embodiments, the graph attention deep reinforcement learning model includes a graph attention deep learning model sub-model and a reinforcement learning sub-model;

[0152] The graph attention deep learning model sub-model is configured to implement a self-attention mechanism for each node in the spliced graph structure to obtain first weight matrices respectively corresponding to each node under different meta-paths; add a multi-head attention mechanism to the corresponding nodes in the spliced graph structure based on the first weight matrices respectively corresponding to each node under different meta-paths to obtain target feature vectors respectively corresponding to the corresponding nodes under different meta-paths; determine the weights respectively corresponding to different meta-paths based on the target feature vectors respectively corresponding to each node under different meta-paths; perform normalization processing on the weights respectively corresponding to different meta-paths and the target feature vectors respectively corresponding to each node under different meta-paths to obtain a processing result, and output the processing result;

[0153] The reinforcement learning sub-model includes a state space, an action space, and a reward function. The reinforcement learning sub-model is configured to use the processing result output by the graph attention deep learning model sub-model to determine a corresponding target state, determine a corresponding multiple actions based on the target state, calculate the value corresponding to each action to obtain a calculation result, and output the calculation result.

[0154] In some alternative embodiments, the apparatus further includes:

[0155] A second determination module, configured to input the heterogeneous graph structure and the graph attention depth reinforcement learning model of the power system in a target scenario into a pre-constructed graph interpretation model, so that the model outputs a corresponding target subgraph.

[0156] In some alternative embodiments, the graph interpretation model calculates the target subgraph through the following steps: transforming the heterogeneous graph structure into a plurality of homogeneous graphs; determining the target set of each node in each homogeneous graph, where the target set is used to represent all node sets in the corresponding homogeneous graph whose shortest path starting from the corresponding node does not exceed a preset value; determining the first subgraph of the corresponding homogeneous graph based on the target set of each node in each homogeneous graph, where the first subgraph includes a feature matrix and an adjacency matrix; respectively splicing the feature matrices and adjacency matrices of the first subgraphs of the plurality of homogeneous graphs to obtain a spliced feature matrix and an adjacency matrix; constructing a plurality of feature selectors based on the spliced feature matrix and constructing a plurality of edge selectors based on the spliced adjacency matrix; using the plurality of feature selectors to respectively perform feature screening on the spliced feature matrix to obtain a target feature matrix after feature screening by each feature selector; using the plurality of edge selectors to respectively perform edge screening on the spliced adjacency matrix to obtain a target adjacency matrix after edge screening by each edge selector; determining a plurality of second subgraphs based on the target feature matrix after feature screening by each feature selector and the target adjacency matrix after edge screening by each edge selector; using the graph attention depth reinforcement learning model to determine the scheduling control strategy of each second subgraph; and determining the second subgraph with the smallest difference from the target scheduling control strategy among the plurality of second subgraphs as the target subgraph.

[0157] In some alternative embodiments, using the plurality of feature selectors to respectively perform feature screening on the spliced feature matrix to obtain a target feature matrix after feature screening by each feature selector includes:

[0158] Multiplying different feature selectors by relevant elements of the spliced feature matrix respectively to obtain a first matrix corresponding to each feature selector; determining first target elements in the first matrix corresponding to each feature selector whose element values are less than a preset threshold; and deleting the node feature vectors corresponding to the target elements in each first matrix from the spliced feature matrix to obtain a target feature matrix after feature screening by the corresponding feature selector.

[0159] In some alternative embodiments, using the plurality of edge selectors to respectively perform edge screening on the spliced adjacency matrix to obtain a target adjacency matrix after edge screening by each edge selector includes:

[0160] Multiply different edge selectors with relevant elements of the spliced feature matrix respectively to obtain a second matrix corresponding to each edge selector; determine second target elements in the second matrix corresponding to each edge selector whose element values are less than a preset threshold; delete the edge features corresponding to the second target elements in each second matrix from the spliced adjacency matrix to obtain a target adjacency matrix after edge screening using the corresponding edge selector.

[0161] The further functional descriptions of the above-mentioned various modules and units are the same as those in the corresponding above embodiments, and will not be elaborated here.

[0162] The dispatching control device of the power system in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0163] An embodiment of the present invention further provides a computer device having the above Figure 3 shown dispatching control device of the power system.

[0164] Please refer to Figure 4 , Figure 4 is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 4 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways according to needs. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 4 In

[0165] Processor 10 can be a central processor, a network processor, or a combination thereof. Among them, processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0166] Among them, the memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.

[0167] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the computer device, etc. In addition, the memory 20 may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided relative to the processor 10, and these remote memories can be connected to the computer device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0168] The memory 20 may include volatile memory, for example, random access memory; the memory may also include non-volatile memory, for example, flash memory, a hard disk, or a solid-state drive; the memory 20 may further include a combination of the above types of memory.

[0169] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.

[0170] The embodiment of the present invention further provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored as such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may further include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.

[0171] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0172] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A dispatching control method for a power system, characterized in that: The method comprises: Acquire a heterogeneous graph structure of a power system in a target scenario, wherein the heterogeneous graph structure is composed of a feature matrix and an adjacency matrix, wherein the feature matrix is ​​used to characterize feature vectors of a plurality of nodes, and the adjacency matrix is ​​used to characterize associations between different nodes, and the heterogeneous graph structure contains a plurality of meta-paths, and the meta-paths are used to characterize paths connecting nodes of different categories; The heterogeneous graph structure is transformed using the pre-constructed transformation matrices corresponding to different meta-paths to obtain multiple isomorphic graphs; Splicing the multiple isomorphic graphs to obtain a spliced ​​graph structure; Inputting the spliced ​​graph structure into a pre-built graph attention deep reinforcement learning model, so that the graph attention deep reinforcement learning model outputs the values ​​corresponding to different dispatching control strategies of the power system under the target scenario; The target dispatch control strategy with the highest value among different dispatch control strategies is used to dispatch and control the power system; The graph attention deep reinforcement learning model includes a graph attention deep learning model sub-model and a reinforcement learning sub-model; The graph attention deep learning model sub-model is used to implement a self-attention mechanism on each node in the spliced ​​graph structure to obtain the first weight matrices corresponding to each node under different meta-paths; based on the first weight matrices corresponding to each node under different meta-paths, a multi-head attention mechanism is added to the corresponding nodes in the spliced ​​graph structure to obtain the target feature vectors corresponding to the corresponding nodes under different meta-paths; based on the target feature vectors corresponding to each node under different meta-paths, the weights corresponding to different meta-paths are determined; the weights corresponding to different meta-paths and the target feature vectors corresponding to each node under different meta-paths are normalized to obtain a processing result, and the processing result is output; The reinforcement learning sub-model includes a state space, an action space and a reward function. The reinforcement learning sub-model is used to determine the corresponding target state using the processing results output by the graph attention deep learning model sub-model, determine the corresponding multiple actions based on the target state, calculate the value corresponding to each action, obtain the calculation results, and output the calculation results. The multiple actions are used to characterize different scheduling control strategies.

2. The method according to claim 1, characterized in that The method further comprises: The heterogeneous graph structure and graph attention deep reinforcement learning model of the power system in the target scenario are input into a pre-built graph interpretation model so that the model outputs the corresponding target subgraph.

3. The method according to claim 2, characterized in that The graph interpretation model calculates the target subgraph through the following steps: Transforming the heterogeneous graph structure into multiple isomorphic graphs; Determine a target set for each node in each isomorphic graph, where the target set is used to represent a set of all nodes in the corresponding isomorphic graph whose shortest path found starting from the corresponding node does not exceed a preset value; Determine a first subgraph of the corresponding isomorphic graph based on a target set of each node in each isomorphic graph, wherein the first subgraph includes a feature matrix and an adjacency matrix; Concatenating the feature matrices and adjacency matrices of the first subgraphs of the multiple isomorphic graphs respectively to obtain concatenated feature matrices and adjacency matrices; Construct multiple feature selectors based on the concatenated feature matrix, and construct multiple edge selectors based on the concatenated adjacency matrix; Using multiple feature selectors to perform feature screening on the concatenated feature matrix respectively, and obtaining a target feature matrix after feature screening by each feature selector; Using multiple edge selectors to perform edge screening on the spliced ​​adjacency matrix respectively, and obtaining a target adjacency matrix after edge screening by each edge selector; Determine a plurality of second subgraphs based on a target feature matrix after feature selection by each feature selector and a target adjacency matrix after feature selection by each edge selector; Determining a scheduling control strategy for each second subgraph using the graph attention deep reinforcement learning model; The second subgraph with the smallest difference from the target scheduling control strategy among the multiple second subgraphs is determined as the target subgraph.

4. The method according to claim 3, characterized in that Use multiple feature selectors to perform feature screening on the concatenated feature matrix respectively, and obtain the target feature matrix after feature screening by each feature selector, including: Multiplying different feature selectors with the related elements of the concatenated feature matrix respectively to obtain a first matrix corresponding to each feature selector; Determine a first target element whose element value in the first matrix corresponding to each feature selector is less than a preset threshold; The node feature vectors corresponding to the target elements in each first matrix are deleted from the concatenated feature matrix to obtain the target feature matrix after feature screening using the corresponding feature selector.

5. The method according to claim 3, characterized in that: Use multiple edge selectors to perform edge filtering on the concatenated adjacency matrix respectively, and obtain the target adjacency matrix after edge filtering by each edge selector, including: Multiplying different edge selectors by the related elements of the concatenated feature matrix respectively to obtain a second matrix corresponding to each edge selector; Determine a second target element in the second matrix corresponding to each edge selector whose element value is less than a preset threshold; The edge features corresponding to the second target elements in each second matrix are deleted from the concatenated adjacency matrix to obtain the target adjacency matrix after edge screening using the corresponding edge selector.

6. A dispatching control device for a power system, characterized in that: For executing the method according to claim 1, the device comprises: An acquisition module is used to acquire a heterogeneous graph structure of a power system in a target scenario, wherein the heterogeneous graph structure is composed of a feature matrix and an adjacency matrix, wherein the feature matrix is ​​used to characterize feature vectors of multiple nodes, and the adjacency matrix is ​​used to characterize associations between different nodes. The heterogeneous graph structure contains multiple meta-paths, and the meta-paths are used to characterize paths connecting nodes of different categories. A conversion module, used to convert the heterogeneous graph structure using pre-constructed conversion matrices corresponding to different meta-paths to obtain multiple isomorphic graphs; A splicing module, used for splicing the multiple isomorphic graphs to obtain a spliced ​​graph structure; A first determination module is used to input the spliced ​​graph structure into a pre-built graph attention deep reinforcement learning model, so that the graph attention deep reinforcement learning model outputs the values ​​corresponding to different dispatch control strategies of the power system under the target scenario; The dispatching control module is used to dispatch and control the power system using the target dispatching control strategy with the highest value among different dispatching control strategies.

7. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the dispatching control method of the power system according to any one of claims 1 to 5 by executing the computer instructions.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the dispatching control method for the power system according to any one of claims 1 to 5.

9. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the dispatching control method for a power system according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Meta-path learning method for high-order heterogeneous graph classification

    CN112148931A

  • Heterogeneous graph attention networks for scalable multi-robot scheduling

    US20220226994A1