A power grid regulation method based on human-computer cooperation combined with inverse reinforcement learning

CN115309908BActive Publication Date: 2026-09-25ELECTRIC POWER SCI & RES INST OF STATE GRID TIANJIN ELECTRIC POWER CO +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210721460.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2026-09-25
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

[0003]经检索,未发现与本发明相同或相近似的现有技术的公开文献

Benefits of technology

[0077]1、本发明在基于离线的历史电网调控决策信息结合知识图谱构建的关联决策图上对节点之间的边赋予权重,进而通过相邻节点之间的距离远近表达它们之间的相关性。与现有的方法不同,现存的算法中直接对离线数据信息结合知识图片通过迪杰斯特拉方法来计算从源节点到目标节点的最短距离,未考虑到节点之间的边其实是具有远近区别的而并非均为1。本发明中通过先计算节点加邻接边与邻居节点欧式距离差,将结果转换为节点与邻居节点的相关性即为权重,然后基于带有权重的节点关联决策图来计算从源节点到目标节点的最短距离,以这些具有最短距离的实例路径作为监督性质的状态转移路径,用于监督约束使用逆强化学习策略学习到的路径,进而更合理地更新策略,提高生成用于解释调控过程的可解释性路径的合理性以及提升调控的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115309908B_ABST
    Figure CN115309908B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of power grid regulation methods based on man-machine cooperation combined with inverse reinforcement learning, comprising the following steps: step 1, input the data set of power grid;Step 2, construct the knowledge graph of power grid equipment node state and regulation behavior;Step 3, obtain the Embedding of equipment node state and regulation action;Step 4, define multi-hop scoring function according to the situation from current state to target state;Step 5, construct regulation meta-path based on state using prior knowledge of artificial expert;Step 6, generate the first part of reward function of reinforcement learning;Step 7, generate total reward function;Step 8, define the Markov process of inverse reinforcement learning and the inverse reinforcement learning strategy update framework based on actor-critic;Step 9, training produces power grid regulation strategy based on man-machine cooperation combined with inverse reinforcement learning.The present application can improve the accuracy of online and offline decision-making of power grid regulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power grid control technology, and relates to a power grid control method, particularly a power grid control method based on human-machine collaborative inverse reinforcement learning. Background Technology

[0002] As the power grid continues to expand, it needs to address diverse demands, leading to increasingly complex power grid control operations and a heavier workload for control personnel. This places higher demands on the intelligence and precision of control operations. Existing power grid control applications based on technologies such as deep learning rely on offline sample training for control decisions, which cannot handle all the complex and variable operating conditions of the power grid. This results in low decision-making accuracy of the trained models on online data, and the existing models also suffer from poor interpretability and interactivity. Therefore, how to propose a power grid control method that achieves better performance in optimization, decision-making, and reasoning tasks, and improves the accuracy, interactivity, and interpretability of decisions, is a technical challenge that urgently needs to be solved by those skilled in the art.

[0003] A search revealed no publicly available literature of the same or similar prior art as this invention. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and propose a power grid control method based on human-machine collaboration and inverse reinforcement learning, which is rationally designed, highly interactive and interpretable, and has high decision-making accuracy.

[0005] The present invention solves its practical problem by adopting the following technical solution:

[0006] A power grid control method based on human-machine collaboration combined with inverse reinforcement learning includes the following steps:

[0007] Step 1: Input the power grid dataset;

[0008] Step 2: Using prior knowledge of power grid regulation and combining it with the status of power grid equipment entities and corresponding regulation actions in the offline power grid historical dataset, construct a knowledge graph that includes the status of power grid equipment nodes and regulation behaviors in the dataset.

[0009] Step 3: Using the knowledge graph constructed in Step 2 and the relationship between the state transitions of each device entity in the power grid dataset, perform graph representation learning on the device nodes and control actions in the power grid dataset in Step 1, and finally obtain the embedding of device node states and control actions.

[0010] Step 4: Select the knowledge graph constructed in Step 2, and define a multi-hop scoring function based on the current state to the target state to evaluate the correlation between two states. The score is calculated by using the embedding of the device node state as the input of the scoring function.

[0011] Step 5: Based on the multi-hop scoring function defined in Step 4, construct a state-based regulation meta-path using the prior knowledge of human experts;

[0012] Step 6: Use the meta-path of state-based regulation action obtained in Step 5 as the prior guidance in the reinforcement learning decision-making process, generate regulation action selection constraints, generate the path from the source state to the target state, use the scoring function to calculate the score evaluation of multi-hop nodes in the path, and generate the first part of the reward function for reinforcement learning.

[0013] Step 7: Based on Step 2 and Step 3, obtain the dual-supervised reward function under offline historical data constraints and online human-computer interaction constraints respectively, and combine it with the first part of the reward function obtained from Step 6 to generate the total reward function;

[0014] Step 8: Based on the reward function obtained in Step 7, define the Markov process for inverse reinforcement learning and the actor-critic-based inverse reinforcement learning policy update framework.

[0015] Step 9: Train and generate a power grid control strategy based on human-machine collaboration combined with inverse reinforcement learning.

[0016] Furthermore, the specific steps of step 2 include:

[0017] (1) Obtain the control action record of each power grid equipment node in its initial state;

[0018] (2) The state of each power grid equipment node is regarded as an entity node in the knowledge graph, and the control actions made for the state of each power grid equipment node are regarded as the association edges between entity nodes.

[0019] (3) The status of power grid equipment nodes in the entire power grid dataset is associated with the edges corresponding to the control actions, and finally a knowledge graph containing the status of power grid equipment nodes and control actions in the dataset is formed.

[0020] Furthermore, the specific steps of step 3, which utilizes the knowledge graph constructed in step 2 and the state transition relationships of various equipment entities in the power grid dataset to perform graph representation learning on the control actions of the equipment nodes included in the power grid dataset in step 1, include:

[0021] (1) Based on the state of the power grid equipment node, define the entity class corresponding to each state of the power grid equipment node, and define the number of entity classes as n; at the same time, define the dimension size of each state input in reinforcement learning as embed_size.

[0022] (2) The entity class is initialized for representation learning based on the number m of the corresponding power grid equipment node states contained in each entity class. The dimension of the initialization vector is m*embed_size.

[0023] (3) Initialize the device node information in the power grid dataset. The dimension of the initialization vector is embed_size.

[0024] (4) Define the dimension of the initialization vector for fault handling actions as 1*embed_size;

[0025] Based on the relevant state-based control dataset, corresponding records are extracted. Each record contains instance records corresponding to n entity classes, forming an n-tuple. From these n-tuples, triplets (state i, control action r, state j) with corresponding relationships are generated. The number of such triplets is denoted as k. These k triplets are used as input to the mature graph representation learning algorithm TransR for loss training, resulting in a representation model capable of embedding the current node state and control action. This model is then used to obtain the embedding representations of the node and control action.

[0026] Furthermore, in step 4, the knowledge graph constructed in step 2 is selected, and a multi-hop scoring function is defined based on the current state to the target state. The specific steps include:

[0027] (1) First, define the entities in the multi-hop path. The first entity in the path is defined as e0, and the last entity is defined as e. t Based on knowledge graphs, if e0 and e t There exists a series of entities in the middle, such as {e0, e1, ..., e...} t-1}, and the t relationships between them. That is, {r1,r2,...,r t Therefore, we can define a specific and effective multi-hop path based on the knowledge graph.

[0028] (2) After defining the multi-hop path, it is necessary to define the scoring function for the multi-hop path, targeting the two entities e0 and e in the multi-hop path. t The scoring function can be defined as:

[0029]

[0030] Where j represents the index of any entity node in the multi-hop path, b et This is the bias value set here. When t=0 and j=0, this scoring function represents the similarity between two entity vectors, that is:

[0031]

[0032] When t=1 and j=1, the scoring function represents the similarity between the head entity (after adding the relation) and the tail entity, that is:

[0033]

[0034] Based on the above, a multi-hop scoring function based on knowledge graphs is defined to evaluate the correlation between two states.

[0035] Furthermore, the specific steps of step 5 include:

[0036] (1) Generate a series of triples based on the power grid equipment node state types and control action types contained in the knowledge graph.

[0037] (2) Based on the prior knowledge of artificial experts, these triples with relationships are associated, and finally multiple meta-paths with prior guidance significance are abstracted, which can effectively guide the reinforcement learning agent to make control action selection in the corresponding state.

[0038] Furthermore, the specific steps of step 6 include:

[0039] (1) Define multiple meta-paths based on expert prior knowledge;

[0040] (2) In the path exploration and attempt process of the agent in reinforcement learning, the current power equipment state is guided to select the control action according to the defined meta-path, so that the equipment is transferred to the next state, and so on until the end of the cycle, and finally the state transition path from the source state of the power equipment to the target state is generated.

[0041] (3) The correlation between the source state and the target state is calculated by using the predefined multi-hop scoring function to obtain the first part of the reward function for reinforcement learning.

[0042] Furthermore, the specific steps of step 7 include:

[0043] (1) Based on the knowledge graph of power grid control instances extracted from offline historical data in step 2 and the representation of nodes and edges in step 3, a knowledge graph of power grid control instances with practical significance is obtained.

[0044] (2) Based on the nodes and edges in the knowledge graph of power grid control examples with practical significance, calculate the Euclidean distance of a node to its neighboring node through the relation edge, use this distance to express the weight of the correlation between nodes, and then replace the relationship between the instance nodes with the weight.

[0045] (3) Based on the knowledge graph of power grid regulation instances with weights, the Dijkstra algorithm is used to calculate the shortest state transition path from the source node to the target node, which is used as supervision information of offline historical data to counteract the state transition path generated by the inverse reinforcement learning strategy and generate the second part of the reward function.

[0046] (4) After obtaining the second part of the reward function, use the control actions in the online human interaction information to counteract the decision of a specified step to generate the third part of the reward function;

[0047] (5) Superimpose the three reward functions to generate a reward function that drives the update of the entire inverse reinforcement learning policy.

[0048] Furthermore, the specific steps of step 8 include:

[0049] (1) Select an inverse reinforcement learning network framework based on actor-critic;

[0050] (2) State definition. At time t, state s t Defined as a triple (u, e) t ,h t ), where u belongs to the entity set U of the power grid equipment node state type, which here refers to the starting point of the decision-making process, and e t This represents the entity that the agent reaches after t steps, and the final h. t This represents the historical records up to step t. These records constitute the current state.

[0051] Based on the above definition, the initial state is clearly represented as:

[0052]

[0053] The state at termination time T can be represented as:

[0054] s T =(u,e T ,h T )

[0055] (3) The definition of an action is the state s at a certain time t. t Under these conditions, each intelligent agent will have a corresponding action space, which includes the action space of entity e at time t. t The set of all out-degree edges, and the entity does not contain entities that exist in the history, i.e.:

[0056]

[0057] (4) Definition of soft reward in reinforcement learning: This soft reward mechanism is based on a multi-hop scoring function, and the reward R obtained by the corresponding state at the termination time T is based on this function. T Defined as:

[0058]

[0059] (5) State transition probability, that is, in the Markov decision process, assuming the state s at the current time t is known. t =(u,e t ,h t ), and in the current state, according to the path search strategy π θ Then execute action a t =(r t+1 ,e t+1 After an action is performed, the agent will reach the next state. There is a definition of state transition probability during this process, which is defined here as:

[0060]

[0061] The initial state is determined by the initial state of the power grid equipment nodes.

[0062] (6) In a given period of a deterministic Markov decision process, what is the total reward G in the state corresponding to a certain time t? t It can be defined as:

[0063] G t =R t+1 +γR t+2 +γ2R t+3 +…+γ T-t-1 R T

[0064] (7) Obtain the reward function under the dual-supervision mechanism of offline historical data combined with online human-computer interaction information at time t, which is R. offline,t and R online,t The definition is as follows:

[0065] R offline,t =log(D p (s t ,at))-log(1-D p (s t ,a t ))

[0066] R online,t =(Embedding(At )-Embedding(a t )) 2

[0067] Where s t a represents the state at time t. t D represents the action generated by the inverse reinforcement learning policy at time t. p It is used to obtain (s) at time t t ,a t The discriminator is derived from the probability of historical experience data, and the embedding is the encoder obtained in step 3. t The time t represents the action command given by the staff based on human-computer interaction.

[0068] (8) The reward function under the dual-supervision mechanism of offline historical data combined with online human-computer interaction information at time t, the strategy optimization, that is, in the Markov decision process, our goal is to learn an excellent search strategy. This search strategy can maximize the cumulative reward of any starting state of the power grid equipment node within the search period, that is, the formula is defined as:

[0069]

[0070] (9) Perform gradient updates for the inverse reinforcement learning policy. The definition is as follows:

[0071]

[0072] Where R all This represents the transition from state s to the final state s. T The discount to the reward plus the R corresponding to the time t when state s is in offline,t and R online,t The sum of .

[0073] (10) Finally, a trainable actor-critic-based inverse reinforcement learning model framework is obtained.

[0074] Furthermore, the specific method for step 9 is as follows:

[0075] The system first constructs a meaningful knowledge graph of power grid control instances based on the power grid control instance knowledge graph obtained in step 2 and the embeddings of power equipment node states and control actions obtained in step 3. This knowledge graph is then used as input to the edge weighting module to obtain a weighted power grid control instance knowledge graph. Next, the Dijkstra algorithm is used to calculate the shortest state transition path based on the offline historical data of power grid control. Then, control actions are extracted based on human-computer interaction information and encoded using the embedding module from step 3. Following this, the Markov process and inverse reinforcement learning policy update framework defined in step 8 are used to input the node states from the offline historical data of power grid control into the inverse reinforcement learning model. The inverse reinforcement learning policy guides the generation of actions and action paths. The shortest state transition path based on the offline historical data of power grid control and the control actions extracted based on human-computer interaction information are used as supervisory constraints to generate a reward function, driving policy updates. Finally, a power grid control policy based on human-machine collaborative inverse reinforcement learning is trained and generated.

[0076] Advantages and beneficial effects of the present invention:

[0077] 1. This invention assigns weights to edges between nodes on an association decision graph constructed based on offline historical power grid control decision information and a knowledge graph, thereby expressing the correlation between adjacent nodes through their distance. Unlike existing methods, which directly calculate the shortest distance from the source node to the target node using Dijkstra's method on offline data and a knowledge graph, neglecting to consider that edges between nodes have varying distances and are not all equal to 1, this invention first calculates the Euclidean distance difference between a node and its neighboring edges, converting the result into the correlation between the node and its neighbors, which is then used as the weight. The shortest distance from the source node to the target node is then calculated based on the weighted node association decision graph. These shortest distance instance paths serve as supervised state transition paths, used to supervise and constrain the paths learned using inverse reinforcement learning, thereby updating the strategy more rationally, improving the rationality of generating interpretable paths for explaining the control process, and enhancing the accuracy of control.

[0078] 2. This invention employs a dual-supervision model combining offline data and online human-computer interaction information to implement a power grid control method based on human collaboration and inverse reinforcement learning. Addressing the issue of low accuracy in processing online data using models trained on offline data, this invention first extracts effective decision paths for power grid control using offline historical information combined with a knowledge graph. These effective decision paths serve as target constraints for the paths inferred by the inverse reinforcement learning model. Then, human-computer interaction information is used as target constraints for each decision step. This dual-supervision constraint approach effectively optimizes inverse reinforcement learning decisions, thus resolving, to some extent, the issues of insufficient decision quality and poor interactivity in reinforcement learning. The difference between this invention and previous power grid control methods lies in the adoption of a dual-supervision constraint mechanism for the output of the reinforcement learning framework. By collecting decision information data under different states, such as offline historical data and online human-computer interaction data, the input data is enhanced, and the output strategy is constrained, thereby improving the training quality and online effectiveness of the inverse reinforcement learning model.

[0079] 3. The inverse reinforcement learning proposed in this invention incorporates effective paths extracted from historical offline data as self-supervised information into an unsupervised trial-and-error learning process, while also integrating control information from online human-computer interaction. This process eliminates the need for dataset labeling and effectively constrains the trial-and-error process through online and offline supervision information, thereby improving the accuracy of online and offline decision-making in power grid regulation. Attached Figure Description

[0080] Figure 1 This invention provides a flowchart of the process of assigning weights to the edges between nodes on an association decision graph constructed based on offline historical power grid control decision information and knowledge graph, and then expressing the correlation between adjacent nodes by the distance between them.

[0081] Figure 2 This is a schematic diagram illustrating the processing flow of the power grid control method based on human collaboration combined with inverse reinforcement learning, implemented by the present invention using a dual-supervision mode technology that combines offline data information and online human-computer interaction information.

[0082] Figure 3 The network framework diagram for updating the power grid control strategy based on inverse reinforcement learning of this invention is shown in the figure. Detailed Implementation

[0083] The present invention will be further described in detail below with reference to the accompanying drawings:

[0084] A power grid control method based on human-machine collaboration combined with inverse reinforcement learning, such as Figures 1 to 3 As shown, it includes the following steps:

[0085] Step 1: Input the power grid dataset;

[0086] The power grid dataset in step 1 includes information on the equipment nodes in the power grid and a set of actions for regulating the status of the power grid equipment.

[0087] Step 2: Using prior knowledge of power grid regulation and combining it with the status of power grid equipment entities and corresponding regulation actions in the offline power grid historical dataset, construct a knowledge graph that includes the status of power grid equipment nodes and regulation behaviors in the dataset.

[0088] Based on the power grid equipment entity states and corresponding control actions contained in the historical power grid dataset from step 1, a knowledge graph is constructed that includes the relationships between entities, encompassing the power grid equipment node states and control behaviors in the dataset.

[0089] The specific steps of step 2 include:

[0090] (1) Obtain the control action record of each power grid equipment node in its initial state;

[0091] (2) The state of each power grid equipment node is regarded as an entity node in the knowledge graph, and the control actions made for the state of each power grid equipment node are regarded as the association edges between entity nodes.

[0092] (3) The status of power grid equipment nodes in the entire power grid dataset is associated with the edges corresponding to the control actions, and finally a knowledge graph containing the status of power grid equipment nodes and control actions in the dataset is formed.

[0093] Step 3: Using the knowledge graph constructed in Step 2 and the relationship between the state transitions of each device entity in the power grid dataset, perform graph representation learning on the device nodes and control actions in the power grid dataset in Step 1, and finally obtain the embedding of device node states and control actions.

[0094] Step 3 utilizes the knowledge graph constructed in Step 2 and the relationships between the state transitions of various equipment entities in the power grid dataset to perform graph representation learning on the control actions of the equipment nodes in the power grid dataset from Step 1. Specific steps include:

[0095] (1) Based on the state of the power grid equipment node, define the entity class corresponding to each state of the power grid equipment node, and define the number of entity classes as n; at the same time, define the dimension size of each state input in reinforcement learning as embed_size.

[0096] (2) The entity class is initialized for representation learning based on the number m of the corresponding power grid equipment node states contained in each entity class. The dimension of the initialization vector is m*embed_size.

[0097] (3) Initialize the device node information in the power grid dataset. The dimension of the initialization vector is embed_size.

[0098] (4) Define the dimension of the initialization vector for fault handling actions as 1*embed_size;

[0099] Based on the relevant state control dataset, corresponding records are extracted. Each record contains instance records corresponding to n entity classes, forming an n-tuple. Based on the n-tuple, triplets (state i, control action r, state j) with corresponding relationships are generated. The number of such triplets is denoted as k. These k triplets are used as input to the mature graph representation learning algorithm TransR for loss training, obtaining a representation model capable of embedding the current node state and control action. This model is then used to obtain the embedding representations of the node and control action.

[0100] Step 4: Select the knowledge graph constructed in Step 2, and define a multi-hop scoring function based on the current state to the target state to evaluate the correlation between two states. The score is calculated by using the embedding of the device node state as the input of the scoring function.

[0101] In step 4, the knowledge graph constructed in step 2 is selected, and a multi-hop scoring function is defined based on the current state to the target state. Specific steps include:

[0102] (1) First, define the entities in the multi-hop path. The first entity in the path is defined as e0, and the last entity is defined as e. t Based on knowledge graphs, if e0 and e t There exists a series of entities in the middle, such as {e0, e1, ..., e...} t-1}, and the t relationships between them. That is, {r1,r2,...,r t Therefore, we can define a specific and effective multi-hop path based on the knowledge graph.

[0103] (2) After defining the multi-hop path, it is necessary to define the scoring function for the multi-hop path, targeting the two entities e0 and e in the multi-hop path. t The scoring function can be defined as:

[0104]

[0105] Where j represents the index of any entity node in the multi-hop path, b et This is the bias value set here. When t=0 and j=0, this scoring function represents the similarity between two entity vectors, that is:

[0106]

[0107] When t=1 and j=1, the scoring function represents the similarity between the head entity (after adding the relation) and the tail entity, that is:

[0108]

[0109] Based on the above, a multi-hop scoring function based on knowledge graphs is defined to evaluate the correlation between two states.

[0110] Step 5: Based on the multi-hop scoring function defined in Step 4, construct a state-based regulation meta-path using the prior knowledge of human experts;

[0111] The specific steps of step 5 include:

[0112] Multiple meta-paths can be defined using prior knowledge from human experts in the relevant field. A specific method could be:

[0113] (1) Generate a series of triples based on the power grid equipment node state types and control action types contained in the knowledge graph.

[0114] (2) Based on the prior knowledge of artificial experts, these triples with relationships are associated, and finally multiple meta-paths with prior guidance significance are abstracted, which can effectively guide the reinforcement learning agent to make control action selection in the corresponding state.

[0115] Step 6: Use the meta-path of state-based regulation action obtained in Step 5 as the prior guidance in the reinforcement learning decision-making process, generate regulation action selection constraints, generate the path from the source state to the target state, use the scoring function to calculate the score evaluation of multi-hop nodes in the path, and generate the first part of the reward function for reinforcement learning.

[0116] The specific steps of step 6 include:

[0117] In step 6, the meta-path obtained from step 5 is used to constrain the search path of the reinforcement learning agent. Specifically, the method can be as follows:

[0118] (1) Define multiple meta-paths based on expert prior knowledge;

[0119] (2) In the path exploration and attempt process of the agent in reinforcement learning, the current power equipment state is guided to select the control action according to the defined meta-path, so that the equipment is transferred to the next state, and so on until the end of the cycle, and finally the state transition path from the source state of the power equipment to the target state is generated.

[0120] (3) The correlation between the source state and the target state is calculated by using the predefined multi-hop scoring function to obtain the first part of the reward function for reinforcement learning.

[0121] Step 7: Based on Step 2 and Step 3, obtain the dual-supervised reward function under offline historical data constraints and online human-computer interaction constraints respectively, and combine it with the first part of the reward function obtained from Step 6 to generate the total reward function;

[0122] In this embodiment, the first part of the reward function obtained in step 6 is used to update the policy of the reinforcement learning network. However, its constraint-driven approach is insufficient, which leads to the policy not being well applied to online and offline regulation. Therefore, in this step, we first transform the edges in the knowledge graph with actual representational meaning from relations to weights based on the knowledge graph obtained in step 2 and the entity and relation embeddings obtained in step 3. Then, based on the weighted edges, we use the Dijkstra algorithm to find the shortest state transition path from the source node to the target node in the knowledge graph as a supervisory constraint for the state transition path generated by the reinforcement learning policy. The error loss is used as the second part of the reward function. At the same time, the regulation action in the human-computer interaction information is used as a supervisory constraint for the decision of a certain step, thus obtaining the third part of the reward function. Finally, the three parts of the reward function are combined as the reward function of reinforcement learning to drive the policy update.

[0123] The specific steps of step 7 include:

[0124] (1) Based on the knowledge graph of power grid control instances extracted from offline historical data in step 2 and the representation of nodes and edges in step 3, a knowledge graph of power grid control instances with practical significance is obtained.

[0125] (2) Based on the nodes and edges in the knowledge graph of power grid control examples with practical significance, calculate the Euclidean distance of a node to its neighboring node through the relation edge, use this distance to express the weight of the correlation between nodes, and then replace the relationship between the instance nodes with the weight.

[0126] (3) Based on the knowledge graph of power grid regulation instances with weights, the Dijkstra algorithm is used to calculate the shortest state transition path from the source node to the target node, which is used as supervision information of offline historical data to counteract the state transition path generated by the inverse reinforcement learning strategy and generate the second part of the reward function.

[0127] (4) After obtaining the second part of the reward function, use the control actions in the online human interaction information to counteract the decision of a specified step to generate the third part of the reward function;

[0128] (5) Superimpose the three reward functions to generate a reward function that drives the update of the entire inverse reinforcement learning policy.

[0129] Step 8: Based on the reward function obtained in Step 7, define the Markov process for inverse reinforcement learning and the actor-critic-based inverse reinforcement learning policy update framework.

[0130] The specific steps of step 8 include:

[0131] In step 8, the Markov process for inverse reinforcement learning and the actor-critic inverse reinforcement learning policy update network based on temporal difference are defined, specifically as follows:

[0132] (1) Select an inverse reinforcement learning network framework based on actor-critic;

[0133] (2) State definition. That is, at time t, the state s t Defined as a triple (u, e) t ,h t ), where u belongs to the entity set U of the power grid equipment node state type, which here refers to the starting point of the decision-making process, and e t This represents the entity that the agent reaches after t steps, and the final h. t This represents the historical records up to step t. These records constitute the current state.

[0134] Based on the above definition, the initial state is clearly represented as:

[0135]

[0136] The state at termination time T can be represented as:

[0137] s T =(u,e T ,h T )

[0138] (3) The definition of an action is the state s at a certain time t. t Under these conditions, each intelligent agent will have a corresponding action space, which includes the action space of entity e at time t. t The set of all out-degree edges, and the entity does not contain entities that exist in the history, i.e.:

[0139]

[0140] (4) Definition of soft reward in reinforcement learning: This soft reward mechanism is based on a multi-hop scoring function, and the reward R obtained by the corresponding state at the termination time T is based on this function. T Defined as:

[0141]

[0142] (5) State transition probability, that is, in the Markov decision process, assuming the state s at the current time t is known. t =(u,e t ,h t ), and in the current state, according to the path search strategy π θ Then execute action a t =(r t+1 ,e t+1 After an action is performed, the agent will reach the next state. There is a definition of state transition probability during this process, which is defined here as:

[0143]

[0144] The initial state is determined by the initial state of the power grid equipment nodes.

[0145] (6) The discount factor refers to the fact that in a Markov decision-making process, in order to obtain more rewards, the agent often considers not only the immediate reward obtained at present, but also the immediate reward obtained in future states. In a given period of a deterministic Markov decision-making process, the total reward G corresponding to a certain state at a certain time t is... t It can be defined as:

[0146] G t =R t+1 +γR t+2 +γ 2 R t+3 +…+γ T-t-1 R T

[0147] This is the sum of the current immediate reward and the discounted future reward, where T represents the final state. Because the environment is often random, performing a specific action doesn't necessarily lead to a specific state. Therefore, future rewards should be diminished compared to the reward in the current state. This is the purpose of using a discount factor γ, where γ belongs to [0,1], indicating that the further away the reward is from the current state, the greater the discount needs to be. If it equals 0, it means only the reward in the current state needs to be used; if it equals 1, it means the environment is deterministic, and the same action can obtain the same reward. Therefore, in practice, values ​​like 0.8 or 0.9 are often chosen. Thus, our ultimate task is to train a policy to maximize the final reward R.

[0148] (7) Obtain the reward function under the dual-supervision mechanism of offline historical data combined with online human-computer interaction information at time t, which is R. offline,t and R online,t The definition is as follows:

[0149] R offline,t =log(D p (s t ,at))-log(1-D p (s t ,a t ))

[0150] R online,t =(Embedding(A t )-Embedding(a t )) 2

[0151] Where s t a represents the state at time t. t D represents the action generated by the inverse reinforcement learning policy at time t. p It is used to obtain (s) at time t t ,a t The discriminator is derived from the probability of historical experience data, and the embedding is the encoder obtained in step 3. t The time t represents the action command given by the staff based on human-computer interaction.

[0152] (8) The reward function under the dual-supervision mechanism of offline historical data combined with online human-computer interaction information at time t, the strategy optimization, that is, in the Markov decision process, our goal is to learn an excellent search strategy. This search strategy can maximize the cumulative reward of any starting state of the power grid equipment node within the search period, that is, the formula is defined as:

[0153]

[0154] (9) Perform gradient updates for the inverse reinforcement learning policy. The definition is as follows:

[0155]

[0156] Where R all This represents the transition from state s to the final state s. T The discount to the reward plus the R corresponding to the time t when state s is in offline,t and R online,t The sum of .

[0157] (10) Finally, a trainable actor-critic-based inverse reinforcement learning model framework is obtained.

[0158] Step 9: Develop a power grid control strategy based on human-machine collaboration and inverse reinforcement learning.

[0159] The specific method for step 9 is as follows:

[0160] The system first constructs a meaningful knowledge graph of power grid control instances based on the power grid control instance knowledge graph obtained in step 2 and the embeddings of power equipment node states and control actions obtained in step 3. This knowledge graph is then used as input to the edge weighting module to obtain a weighted power grid control instance knowledge graph. Next, the Dijkstra algorithm is used to calculate the shortest state transition path based on the offline historical data of power grid control. Then, control actions are extracted based on human-computer interaction information and encoded using the embedding module from step 3. Following this, the Markov process and inverse reinforcement learning policy update framework defined in step 8 are used to input the node states from the offline historical data of power grid control into the inverse reinforcement learning model. The inverse reinforcement learning policy guides the generation of actions and action paths. The shortest state transition path based on the offline historical data of power grid control and the control actions extracted based on human-computer interaction information are used as supervisory constraints to generate a reward function, driving policy updates. Finally, a power grid control policy based on human-machine collaborative inverse reinforcement learning is trained and generated.

[0161] The specific steps of step 9 include:

[0162] (1) The inverse reinforcement learning used in this invention is based on the speaker-critic algorithm framework, where the reward-driven approach comes from three parts. First, the offline historical dataset of power grid regulation is input, and a meaningful knowledge graph of power grid regulation instances is constructed based on the knowledge graph of power grid regulation instances obtained in step 2 and the embedding of power equipment node states and the embedding set of regulation actions obtained in step 3. Then, this knowledge graph is used as the input of the association edge weighting module to obtain a weighted knowledge graph of power grid regulation instances. Next, the Dijkstra algorithm is used to calculate the shortest state transition path based on the offline historical data of power grid regulation. Then, the regulation actions are extracted based on human-computer interaction information, and the embedding module in step 3 is used for action encoding. Next, in step 8, we define a speaker network (also known as an actor network), which is mainly used to learn a path search strategy to calculate the probability distribution of each action being selected in the effective action space corresponding to the node in the current state. The actor network takes the action space of the current node and its current state as input, and outputs the probability distribution of each action in the action space. Then, a masking operation is used to remove invalid actions, and the result is fed into a softmax layer to generate the final action probability distribution. Its network architecture is as follows: Figure 3As shown in the Actor module. Next, in step 8, the critic network (also known as the critic network) is defined. The critic network is mainly used to evaluate the value of the current state. Its input is the current state of the node, and its output is the value evaluation of that state. Its network architecture is as follows... Figure 3 As shown in the Critic module.

[0163] (2) Set the number of training iterations epochs, and start training from epochs equal to 1.

[0164] (3) The representation learning, namely Embedding, is performed on the node data and control actions in the overall dataset in step 3. Then, the data is input into the actor network and critic network in batches to obtain the probability distribution (fault handling) of each action in the action space and the value evaluation (state quality) of the state.

[0165] (4) Calculate the predicted value of the state by the critic network and the sum of the three rewards obtained in the state to minimize the loss function, and calculate the product of the probability of the current action and the reward brought by the current action to maximize the operation. At the same time, define an entropy to ensure the balance between model exploration and development, and maximize the entropy.

[0166] (5) Within the range of values ​​defined by epochs, repeat steps (3) and (4) in step 9 to finally complete the training of inverse reinforcement learning and obtain the power grid control strategy based on human-machine collaborative inverse reinforcement learning.

[0167] In reinforcement learning application systems, the main focus is on obtaining more reward function drivers. This exploration process, lacking historical experience data or human interaction constraints and supervision, is ultimately too free-form. Such models have not been well applied online or offline, especially when data is insufficient. The innovation of this invention lies in two aspects: First, it transforms reinforcement learning into inverse reinforcement learning, adding supervised constraints to reinforcement learning. This patent adds two constraints: one from offline historical power grid control data, and the other from online human-computer interaction information on control actions. Second, based on a power grid control instance knowledge graph, it transforms the relationships between nodes into weights that express the strength of those relationships. This better serves the generation of experience-providing node state transition paths from offline historical power grid control data, supervising and constraining the update of inverse reinforcement learning strategies. The method proposed in this paper differs from previous reinforcement learning approaches, primarily by combining more self-supervision and supervisory information to enhance the learning ability of inverse reinforcement learning, thereby improving the quality of control action strategies and obtaining more reasonable control suggestions.

[0168] The method in this invention assigns weights to the edges between nodes on an association decision graph constructed from offline historical power grid control decision information and a knowledge graph. The correlation between adjacent nodes is then expressed by their distance. Dijkstra's algorithm is used to sample the shortest path, which serves as a supervisory constraint for the state transition paths generated later based on the reinforcement learning strategy, assisting in updating the model strategy. For the dual-supervision mode technology using offline data and online human-computer interaction information, effective decision paths for power grid control are first extracted using offline historical information combined with the knowledge graph. These effective decision paths serve as the target constraints for the paths inferred by the inverse reinforcement learning model. Then, human-computer interaction information is used as the target constraints for each decision step. This dual-supervision constraint-based decision optimization for inverse reinforcement learning effectively addresses, to some extent, the problems of insufficient decision quality and poor interactivity in reinforcement learning.

[0169] Based on the above improvements, the proposed human-machine collaborative inverse reinforcement learning-based power grid control method has been realized. This method can better improve the performance of the strategy in optimization, decision-making, and reasoning tasks, and enhance the accuracy, interactivity, and interpretability of the decision-making process.

[0170] Figure 1 This invention presents a flowchart illustrating the process of assigning weights to edges between nodes in an association decision graph constructed using offline historical power grid control decision information and a knowledge graph. The flowchart first takes a set of power grid equipment states and a set of actions to control these states as input. It then constructs an association decision graph containing nodes and relationships. Based on the nodes and their corresponding relationships, the invention calculates the Euclidean distance between a node and a node in the chain. This Euclidean distance is used to express the strength of the correlation between adjacent nodes. Finally, the relationship edges are transformed into weighted edges expressing the strength of the relationship, thus assigning weights to the edges of the entire association graph.

[0171] Figure 2 This diagram illustrates the processing flow of a power grid control method based on human collaboration combined with inverse reinforcement learning, implemented using a dual-supervised mode technology that combines offline data and online human-computer interaction. This module aims to use empirical data from historical control information and control information from online human-computer interaction to impose dual constraints on the unsupervised trial-and-error reinforcement learning control strategy, thereby improving the accuracy of online and offline decisions. Its inputs are historical power grid control datasets and control information from online human-computer interaction, and its output is a decision-making strategy that can automatically control the power grid based on its state.

[0172] Figure 3This invention presents a network framework diagram for updating power grid control strategies based on inverse reinforcement learning. The framework comprises three parts: the first part is a module for weighting the associated edges of the power grid control knowledge graph; the second part is a module for learning decision strategies using inverse reinforcement learning; and the third part is a module for updating the inverse reinforcement learning strategy using dual-supervised constraints of offline data information and online human-computer interaction information.

[0173] The working principle of this invention is:

[0174] This invention first formats the centralized equipment nodes and corresponding control actions in the power grid dataset. Based on the processed data, a knowledge graph is constructed using prior knowledge of power grid control. Then, a graph representation learning method is used to learn the graph representations of the power grid equipment node states and control actions. The graph representation learning model is used to embed node and edge information from historical data. Next, the relationships between nodes are transformed into distance-weighted edges by calculating the distance between each node and its corresponding neighboring edge and adjacent connected nodes. The Dijkstra algorithm is then used to find the shortest path edge from the source node to the destination node, and a portion of the path is sampled as a supervisory condition for subsequent inverse reinforcement learning. Subsequently, the graph representation learning model is also used to embed online human-computer interaction control commands as supervisory conditions for each time step of the subsequent inverse reinforcement learning. Finally, based on the knowledge graph, a... The process involves a multi-hop scoring function from the initial grid state to the target grid state; then, prior knowledge from human experts is used to construct a meta-path for grid regulation, providing reasonable regulatory action selection for the transition of current grid equipment node states; the grid equipment node state information is used as input for reinforcement learning, where the reinforcement learning framework consists of an actor-network and a critic-network. The actor-network outputs action selections, ultimately generating a temporally sequenced regulation command and state transition path. At each time step, the embedded representation of the human-computer interaction command and the action output by the actor-network are used for supervision and constraint. Periodic supervision and constraint are applied using the sampled partial paths and the state transition path generated based on the reinforcement learning strategy. This dual-supervision mechanism constrains the reinforcement learning strategy and updates its parameters. The inverse reinforcement learning solution process in this patent application uses a temporal difference method, enabling the reinforcement learning to ultimately obtain a strategy that better guides regulation actions during the model's policy update process.

[0175] The innovation of this invention lies in:

[0176] 1. This invention proposes a power grid control method based on human-machine collaboration combined with inverse reinforcement learning. The difference from previous power grid control methods lies in the adoption of a dual-supervision constraint mechanism, combining online and offline methods, for updating the reinforcement learning framework's strategy. On one hand, empirical data from offline historical datasets are used for supervision and constraint; on the other hand, control commands based on online human-machine interaction are used to locally supervise and constrain the model's strategy decision-making actions at each time step. These two constraints improve the model's online and offline performance accuracy, interactivity, and interpretability.

[0177] 2. Unlike previous inverse reinforcement learning methods that relied on sampling based on offline historical experience supervision data, this invention transforms the relational edges in the association graph constructed by combining the offline historical experience dataset with the knowledge graph into Euclidean distance weights representing the transition probabilities between two adjacent nodes, instead of defaulting the edge weights to 1. As a result, the state transition paths sampled by the Dijkstra algorithm are more reasonable, thus serving as better supervision constraints for the state transition paths generated later based on reinforcement learning strategies. This improves the interpretability of the generated paths used to explain the regulation process and enhances the accuracy and effectiveness of regulation.

[0178] It should be emphasized that the embodiments described in this invention are illustrative rather than limiting. Therefore, this invention includes, but is not limited to, the embodiments described in the specific implementation. Any other implementations derived by those skilled in the art based on the technical solutions of this invention are also within the scope of protection of this invention.

Claims

1. A power grid control method based on human-machine collaboration combined with inverse reinforcement learning, characterized in that: Includes the following steps: Step 1: Input the power grid dataset; Step 2: Using prior knowledge of power grid regulation and combining it with the status of power grid equipment entities and corresponding regulation actions in the offline power grid historical dataset, construct a knowledge graph that includes the status of power grid equipment nodes and regulation behaviors in the dataset. Step 3: Using the knowledge graph constructed in Step 2 and the relationship between the state transitions of each device entity in the power grid dataset, perform graph representation learning on the device nodes and control actions in the power grid dataset in Step 1, and finally obtain the embedding of device node states and control actions. Step 4: Select the knowledge graph constructed in Step 2, and define a multi-hop scoring function based on the current state to the target state to evaluate the correlation between two states. The score is calculated by using the embedding of the device node state as the input of the scoring function. Step 5: Based on the multi-hop scoring function defined in Step 4, construct a state-based regulation meta-path using the prior knowledge of human experts; Step 6: Use the meta-path of state-based regulation action obtained in Step 5 as the prior guidance in the reinforcement learning decision-making process, generate regulation action selection constraints, generate the path from the source state to the target state, use the scoring function to calculate the score evaluation of multi-hop nodes in the path, and generate the first part of the reward function for reinforcement learning. Step 7: Based on Step 2 and Step 3, obtain the dual-supervised reward function under offline historical data constraints and online human-computer interaction constraints respectively, and combine it with the first part of the reward function obtained from Step 6 to generate the total reward function; Step 8: Based on the reward function obtained in Step 7, define the Markov process for inverse reinforcement learning and the actor-critic-based inverse reinforcement learning policy update framework. Step 9: Train and generate a power grid control strategy based on human-machine collaboration combined with inverse reinforcement learning; The specific steps of step 2 include: (1) Obtain the control action record of each power grid equipment node in its initial state; (2) The state of each power grid equipment node is regarded as an entity node in the knowledge graph, and the control actions made for the state of each power grid equipment node are regarded as the association edges between entity nodes; (3) The status of power grid equipment nodes in the entire power grid dataset is associated with the edges corresponding to the control actions, and finally a knowledge graph containing the status of power grid equipment nodes and control actions in the dataset is formed. Step 3, which utilizes the knowledge graph constructed in Step 2 and the state transition relationships of various equipment entities in the power grid dataset, performs graph representation learning on the equipment nodes and control actions included in the power grid dataset from Step 1. The specific steps include: (1) Based on the state of the power grid equipment node, define the entity class corresponding to each state of the power grid equipment node, and define the number of entity classes as n; at the same time, define the dimension size of each state input in reinforcement learning as embed_size; (2) The entity class is initialized for representation learning based on the number m of the corresponding power grid equipment node states contained in each entity class. The dimension of the initialization vector is m*embed_size. (3) Initialize the device node information in the power grid dataset. The dimension of the initialization vector is embed_size. (4) Define the dimension of the initialization vector for fault handling actions as 1*embed_size; Based on the control dataset under the relevant states, the corresponding records are obtained. Each record contains instance records corresponding to n entity classes, forming an n-tuple. Based on the n-tuple, triplets (state i, control action r, state j) with corresponding relationships are generated. The number of such triplets is denoted as k. These k triplets are used as input to the mature graph representation learning algorithm TransR for loss training, obtaining a representation model that can embed the current node state and control action. The embedding representation of the node and control action is then obtained using this representation model. In step 4, the specific steps for selecting the knowledge graph constructed in step 2 and defining the multi-hop scoring function based on the current state to the target state include: (1) First, define the entities in the multi-hop path. The first entity of the path is defined as... The final entity is defined as Based on knowledge graphs, if and There exists a series of entities in the middle { , ,..., }, and the t relationships between them. Right now{ , ,..., }, Based on the knowledge graph, a definite effective multi-hop path is defined { }; (2) After defining the multi-hop path, it is necessary to define the scoring function of the multi-hop path, which is applied to the two entities in the multi-hop path. and The scoring function can be defined as: Where j represents the index of any entity node in the multi-hop path, This is the bias value set here; when t=0 and j=0, this scoring function represents the similarity between two entity vectors, that is: , When t=1 and j=1, the scoring function represents the similarity between the head entity (after adding the relation) and the tail entity, that is: , Based on the above, the definition of a knowledge graph-based multi-hop scoring function is completed, which is used to evaluate the correlation between two states. The specific steps of step 5 include: (1) Generate a series of triples based on the state types of power grid equipment nodes and the types of control actions contained in the knowledge graph; (2) Based on the prior knowledge of experts in the artificial domain, the association of triples with related relationships is sorted out, and multiple meta-paths with prior guidance are abstracted from them to guide the reinforcement learning agent to make decisions on the selection of control actions in different state scenarios. The specific steps of step 6 include: (1) Define multiple meta-paths based on expert prior knowledge; (2) In the process of path exploration and attempt of the agent in reinforcement learning, the current state of the power equipment is guided to select the control action according to the defined meta path, so that the equipment is transferred to the next state, and so on until the end of the cycle, and finally the state transition path from the source state of the power equipment to the target state is generated. (3) The correlation between the source state and the target state is calculated by using the predefined multi-hop scoring function to obtain the first part of the reward function for reinforcement learning; The specific steps of step 7 include: (1) Based on the knowledge graph of power grid control instances extracted from offline historical data in step 2 and the representation of nodes and edges in step 3, a knowledge graph of power grid control instances with practical significance is obtained. (2) Based on the nodes and edges in the knowledge graph of power grid control examples with practical significance, calculate the Euclidean distance of a node to its neighbor through the relation edge, use this distance to express the weight of the correlation between nodes, and then replace the relationship between the instance nodes with the weight. (3) Based on the knowledge graph of power grid regulation instances with weights, the Dijkstra algorithm is used to calculate the shortest state transition path from the source node to the target node, which is used as the supervision information of offline historical data to counteract the state transition path generated by the inverse reinforcement learning strategy and generate the second part of the reward function. (4) After obtaining the second part of the reward function, use the control actions in the online human interaction information to counteract the decision of a specified step to generate the third part of the reward function; (5) Superimpose the three reward functions to generate a reward function that drives the update of the entire inverse reinforcement learning policy.

2. The power grid control method based on human-machine collaborative inverse reinforcement learning according to claim 1, characterized in that: The specific steps of step 8 include: (1) Select an inverse reinforcement learning network framework based on actor-critic; (2) State definition: At time t, the state is Defined as a triple Where u belongs to the entity set U of the power grid equipment node state type, and here it refers to the starting point of the decision-making process, while This represents the entity the agent reaches after t steps, the last one. This represents the historical records up to step t; these records constitute the current state. Based on the above definition, the initial state is clearly represented as: The state at termination time T can be represented as: (3) The definition of an action is the state at a certain time t. Each intelligent agent will have a corresponding action space, which contains the entity at time t. The set of all out-degree edges, and the entity does not contain entities that exist in the history, i.e.: (4) Definition of soft reward in reinforcement learning: This soft reward mechanism is based on a multi-hop scoring function, and the reward obtained from the state corresponding to the termination time T. Defined as: (5) State transition probability is the probability that, in the Markov decision process, the state at the current time t is known. And in the current state, based on the path search strategy Then perform the action. The agent will reach the next state; there is a definition of state transition probability in the process of moving from one action to the next state. Here, the state transition probability is defined as: The initial state is determined by the initial state of the power grid equipment nodes; (6) In a given period of a deterministic Markov decision process, what is the total reward at a certain time t corresponding to the state? It can be defined as: (7) Obtain the reward function under the dual-supervision mechanism of offline historical data combined with online human-computer interaction information at time t, which are respectively and The definition is as follows: in This represents the state at time t. This represents the action produced by the inverse reinforcement learning policy at time t. It is used to obtain time t The discriminator is derived from probabilities in historical experience data, and the embedding is the encoder obtained in step 3. The time t represents the action command given by the staff based on human-computer interaction. (8) The reward function under the dual supervision mechanism of offline historical data combined with online human-computer interaction information at time t, the strategy optimization is, in the Markov decision process, its optimization objective is to learn the optimal search strategy so that the power grid equipment nodes can maximize the cumulative revenue within the search period starting from any initial state, that is, the formula is defined as: (9) Perform gradient updates for the inverse reinforcement learning policy. The definition is as follows: in This represents the transition from state s to the final state. The discount for receiving the reward plus the time t corresponding to state s The sum of all additions; (10) Finally, a trainable actor-critic-based inverse reinforcement learning model framework is obtained.

3. The power grid control method based on human-machine collaborative inverse reinforcement learning according to claim 1, characterized in that: The specific method for step 9 is as follows: Inputting the offline historical dataset of power grid control, firstly, based on the power grid control instance knowledge graph obtained in step 2 and the embeddings of power equipment node states and control actions obtained in step 3, a practically meaningful power grid control instance knowledge graph is constructed; then, this knowledge graph is used as input to the association edge weighting module to obtain a weighted power grid control instance knowledge graph; next, the Dijkstra algorithm is used to calculate the shortest state transition path based on the offline historical data of power grid control; then, control actions are extracted based on human-computer interaction information, and the action encoding is performed using the embedding module in step 3; Next, the Markov process and inverse reinforcement learning policy update framework defined in step 8 are used to input the node states in the offline historical dataset of power grid regulation into the inverse reinforcement learning model. The inverse reinforcement learning policy guides the generation of actions and action paths. Then, the shortest state transition path based on the offline historical data of power grid regulation and the regulation actions extracted based on human-machine interaction information are used as supervision constraints to generate a reward function, drive policy updates, and finally train and generate a power grid regulation policy based on human-machine collaboration and inverse reinforcement learning.

Citation Information

Patent Citations

  • Multi-mode reinforcement learning-based power grid regulation and control method

    CN113947320A