A Method and System for Optimizing Power Grid Dispatching Knowledge Graph Data

By using graph database-based models and deep learning technologies in the grid scheduling knowledge graph, entity recognition and relationship extraction are realized, and the problems of insufficient knowledge extraction capabilities and low update efficiency in the existing technology are solved, and high-precision and high-efficiency grid scheduling optimization decision support are achieved.

CN114077674BActive Publication Date: 2025-06-03NARI TECH CO LTD +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111279160.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-31
Publication Date
2025-06-03
Estimated Expiration
2041-10-31

AI Technical Summary

Technical Problem

In the case of the expansion of the power grid scale and frequent knowledge updates, the existing grid scheduling knowledge graphs have problems such as insufficient knowledge extraction capabilities, insensitive time and space, and the need to reconstruct the knowledge graphs for each update, resulting in wasting of computing power and time.

Method used

The grid scheduling knowledge graph model based on graph database is adopted, combined with deep learning technology, entity recognition and relationship extraction are realized, accurate entities and relationships are generated, and incremental updates are carried out through natural language learning knowledge fusion technology to achieve sustainable learning of dynamic knowledge graphs.

Benefits of technology

It improves the knowledge extraction ability and update efficiency of the power grid scheduling knowledge graph, reduces computing resources and time consumption, and achieves high-precision optimization decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114077674B_ABST
    Figure CN114077674B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for optimizing power grid dispatching knowledge graph data. The method of the present invention first uses deep learning methods to automatically mine high-quality domain phrases, complete the automatic recognition of dispatching entities and equivalent disambiguation; then complete the extraction of global relationships of dispatching entities according to deep learning techniques, so as to complete the recognition and verification of entity relationships, and achieve the purpose of establishing an initial power grid dispatching knowledge graph; on the basis of completing the above two steps, use natural language learning knowledge fusion technology to perform incremental training on newly added dispatching plan data based on timestamps; at the same time, introduce the life cycle management of the knowledge content of the knowledge graph during the completion of each step; finally, complete the dynamically sustainable learning knowledge graph under the joint cooperation of the above steps. The present invention ensures the high accuracy of the power grid dispatching optimization decision-making knowledge graph, ensures the dynamic update of incremental knowledge, and reduces the computational resources and time consumption during update training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of power grid dispatching, and particularly relates to a method for optimizing power grid dispatching knowledge graph data. Background Art

[0002] With the continuous expansion of the scale of the power system and the increasing proportion of new energy, the difficulty of active power dispatching is rising day by day. As one of the most important system control means of the power system, accurate and safe dispatching not only involves the development of the national economy, but also ensures the safe and efficient operation of the entire power system.

[0003] At the current stage, due to the rapid growth of the number of power grid devices and the total amount of power system knowledge, traditional knowledge organization and management means can no longer meet the requirements. Compared with basic databases, knowledge bases such as intelligent decision-making systems and transmission grid planning decision-making systems that contain rules and are based on expert systems have been put into operation and have been widely used in the power system due to their obvious advantages. However, the current knowledge bases rely on experts. Extracting, sorting out and finally storing relevant content in the form of charts in the database not only severely limits the storage structure, but also requires a large number of professionals and time for updating. Especially for professional fields such as existing power dispatching with frequent environmental changes, high rule requirements, fast update iterations and rich cases, the industry urgently needs a more automated and intelligent knowledge extraction, storage, management and reasoning method system.

[0004] For the above reasons, the power grid dispatching knowledge graph system came into being. A knowledge graph is a structured semantic knowledge base that represents and stores entities and their relationships in the form of a graph. In a knowledge graph, the basic unit of an entity and its relationship is a triple of "entity - relationship - entity", and the attributes of the entity itself are represented and stored using "attribute - value" pairs. The unique storage of facts, instances and relationships in the knowledge graph perfectly meets the requirements of power grid dispatching auxiliary tools, and comprehensively improves the original system in terms of professionalism, relevance, collaboration and constructability. On the other hand, the construction technology of the knowledge graph also includes knowledge update and learning capabilities based on artificial intelligence, which solves the weakness of traditional databases based on strings and links in functions such as fuzzy search and similar case query.

[0005] However, the existing power grid dispatching knowledge graph still has certain limitations, such as the need to improve the knowledge extraction ability under the condition of the continuous increase of the existing power grid scale, the insensitivity to operation and environment in time and space, and the waste of computing power and time caused by the need to reconstruct the knowledge graph every time it is updated. If a truly practical power grid dispatching knowledge graph is to be achieved, these problems still need to be solved urgently. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to realize the full life cycle management of the power grid dispatching knowledge graph: knowledge extraction - knowledge fusion - knowledge graph storage update - expired knowledge elimination, and solve the technical problem of insufficient intelligence level of the dispatching system in the prior art.

[0007] To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0008] A method for optimizing power grid dispatching knowledge graph data, comprising the following steps:

[0009] Step 1. Establish a power grid dispatching knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph. The records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment.

[0010] Step 2. Identify and extract power grid dispatching entities.

[0011] Step 3. For the original corpus of the extracted power grid dispatching entity data, use an open-source word segmentation tool to perform Chinese word segmentation on the sentences and evaluate the word segmentation results.

[0012] Step 4. Initialize the entity relationship triple.

[0013] Step 5. Construct and train an entity recognition model and a relationship recognition model based on a deep neural network to generate accurate entities and relationships.

[0014] A power grid dispatching knowledge graph data optimization system, characterized by including the following program modules:

[0015] Model establishment module: Establish a power grid dispatching knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph. The records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment.

[0016] Entity extraction module: Identify and extract power grid dispatching entities.

[0017] Word segmentation module: For the original corpus of the extracted power grid dispatching entity data, use an open-source word segmentation tool to perform Chinese word segmentation on the sentences and evaluate the word segmentation results.

[0018] Entity relationship triple module: Initialize the entity relationship triple.

[0019] Neural network training module: Construct and train an entity recognition model and a relationship recognition model based on a deep neural network to generate accurate entities and relationships.

[0020] Advantages of the present invention: The present invention provides a method for optimizing power grid dispatching knowledge graph data, which includes the following steps: First, use deep learning technology to automatically mine high-quality domain phrases, and on this basis, complete the automatic recognition and equivalent disambiguation of dispatching entities; then, according to deep learning technology, complete the extraction of global relationships between dispatching entities, so as to complete the recognition and verification of entity relationships, generate accurate entities and relationships, and achieve the purpose of establishing an initial power grid dispatching knowledge graph; on the basis of completing the above two steps, use natural language learning knowledge fusion technology to perform incremental training on newly added dispatching plan data based on timestamps; finally, complete a dynamically sustainable learning knowledge graph under the joint cooperation of the above steps. The present invention ensures the high accuracy of the power grid dispatching optimization decision-making knowledge graph, ensures the dynamic update of incremental knowledge while reducing the computational resources and time consumption during update training. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 Schematic diagram of the topology structure diagram model of the power grid dispatching knowledge graph provided by the embodiment of the present invention;

[0022] Figure 2 Schematic diagram of the method for automatically identifying and extracting dispatching entities provided by the embodiment of the present invention;

[0023] Figure 3 Schematic diagram of the method for obtaining relationships between dispatching entities provided by the embodiment of the present invention;

[0024] Figure 4 Schematic diagram of the incremental update method of the knowledge graph model based on timestamps provided by the embodiment of the present invention;

[0025] Figure 5 Schematic diagram of the storage scheme provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] The present invention will be further described below. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention, and cannot be used to limit the protection scope of the present invention.

[0027] Embodiment 1

[0028] This embodiment provides a method for optimizing power grid dispatching knowledge graph data, including the following steps:

[0029] Step 1. Establish a power grid dispatching knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph. As Figure 1 shown, the records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment to construct a power grid dispatching knowledge graph model.

[0030] In addition to the above-mentioned internal connection relationships and knowledge graph structures in the production station, such data often shows an associated relationship in terms of geographical location. Many data in the system contain geographical location information, such as company stations, substations, distribution stations, power generation stations, etc.

[0031] China's administrative system is generally divided into four levels. The first level is provincial level, including provinces, autonomous regions, municipalities directly under the Central Government and special administrative regions; the second level is prefecture level, including prefecture-level cities, autonomous prefectures, regions and leagues; the third level is county level, including counties, county-level cities, autonomous counties, special zones, forest areas and city districts; the fourth level is township level, including townships, ethnic townships, towns and sub-district offices, etc. Based on China's administrative system, State Grid Corporation of China often uses large administrative regions to divide and manage its subsidiaries. Considering the above factors, the grid dispatching knowledge graph model structure of the present invention adopts a cascaded geographical knowledge structure to achieve better structuring and accuracy. The first level is the large region, the second level is the provincial level, the third level is the municipal level, and the fourth level is the station or transfer station. Lower-level locations are connected to their corresponding upper-level locations using the superior-subordinate relationship. When the original corpus data only contains low-level geographical location information, the connection relationship can be used to automatically complete the high-level geographical location information and other corresponding relationship attributes for unconnected stations. In this way, when querying, the selection of associated knowledge can be effectively increased, and at the same time, the accuracy and query efficiency can be increased.

[0032] Step 2. Identify and extract grid dispatching entities, including the following steps:

[0033] Dispatch optimization decision-making involves multiple types of data, including structured data and text data;

[0034] The structured data includes load forecasting data, new energy forecasting data, and security constraint sections, which can be directly exported from the grid real-time database, and are extracted into triples using rules and stored in the grid dispatching knowledge graph;

[0035] The text data includes tie-line plans, maintenance plans, grid operation modes, and grid abnormal events. The document content is traversed using methods such as data cleaning or interface conversion, and the document is divided according to the discourse structure association model to unify the original data format;

[0036] Step 3. For the original corpus of the extracted grid dispatching entity data, use an open-source word segmentation tool to perform Chinese word segmentation on the sentences, and evaluate the word segmentation results. According to the evaluation results, complete the entity object association model and store it in the grid dispatching knowledge graph model. The specific steps are as follows:

[0037] 31) Word Segmentation: Use Chinese word segmentation methods to segment the original corpus data into phrases; according to power core dictionaries such as power terms and power grid topology structure models and related extended dictionaries, maximize the effective division of the data content involved in dispatching optimization decisions, and the data content includes power grid operation data, dispatching regulations, experience data, etc.;

[0038] 32) Evaluation:

[0039] Use the semi-supervised learning method based on deep neural network to complete the iterative calculation of statistical index features;

[0040] And use the semi-supervised learning method to evaluate the phrase quality;

[0041] Use the semi-supervised learning method based on deep neural network to iteratively mine the vocabulary, and mine high-quality vocabulary and new words such as power grid entities, attributes, operation terms, and qualification constraints;

[0042] Adopt the semi-supervised learning method to establish a named entity classification system;

[0043] 33) Use open-source natural language preprocessing models (such as Transformer, BERT, etc.) for entity recognition, classify and cluster the entities of dispatching optimization decision types, extract the proper nouns of basic elements such as dispatching entities and attributes, and extract entity information such as time;

[0044] 34) Use the iterative training method in deep learning to make inductive distinctions for entities with the same name but different meanings and entities with different names but the same meaning, so as to achieve the disambiguation effect.

[0045] Step Four. Initialize the entity relationship triple, which specifically includes the following steps:

[0046] 41) Use the distance limit between entities and the position limit of relationship indicators to automatically obtain the triple of "entity-relationship-entity" and verify and label the entities;

[0047] 42) Mark the credible and non-credible relationship triples, and use the Naive Bayes classifier to train and classify the relationship triples into credible and non-credible, so as to obtain the relationship representation model;

[0048] 43) Through the relationship representation model obtained by training, superimpose the feature data, and perform relationship recognition on the trained classifier (which can be the Naive Bayes classifier) to obtain candidate relationship triples; the feature data includes syntactic features, semantic features, part of speech, sequence, etc.;

[0049] 44) Merge all approximate relationship candidate triples, and calculate the credibility of each relationship triple by statistically analyzing the probability distribution.

[0050] Step 5. Build and train an entity recognition model and a relationship recognition model based on a deep neural network, and perform verification to automatically generate accurate entities and relationships, which specifically include the following steps:

[0051] 51) Build an entity recognition model based on a deep neural network. After completing the training of the entity recognition model, it can be re-verified. Within the scope of all data in the power grid dispatching knowledge graph, annotate the entity relationships in the dispatching plan text, mark the specific types corresponding to the entities, and then use the annotated corpus to train a relationship recognition model based on a convolutional neural network;

[0052] 52) Use the trained relationship recognition model to recognize the dispatching entity relationships in the unannotated dispatching plan text;

[0053] 53) Further verify the entity relationships based on the recognition results in step 52) to achieve the consistency of entity relationships, and check whether the relationship types exist in the relationship set according to the power grid dispatching knowledge graph. If not, prompt for review; conversely, search for the matching entities in the power grid dispatching knowledge graph through entity semantic features. If the entity-relationship-entity triples obtained by matching recognition are inconsistent with the entity-relationship-entity triples obtained by searching the power grid dispatching knowledge graph, then prompt for re-review. If they are consistent, continue the review without prompting.

[0054] A power grid dispatching knowledge graph data optimization system, characterized in that it includes the following program modules:

[0055] Model establishment module: Establish a power grid dispatching knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph. The records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment;

[0056] Entity extraction module: Identify and extract power grid dispatching entities;

[0057] Word segmentation module: Use an open-source word segmentation tool to perform Chinese word segmentation on the original corpus of the extracted power grid dispatching entity data, and evaluate the word segmentation results;

[0058] Entity relationship triple module: Initialize entity relationship triples;

[0059] Neural network training module: Build and train an entity recognition model and a relationship recognition model based on a deep neural network to generate accurate entities and relationships.

[0060] Embodiment 2

[0061] On the basis of steps 1 to 5 of Embodiment 1, it further includes:

[0062] Step 6. Perform incremental data update and knowledge fusion, which specifically includes the following steps:

[0063] 61) Based on the trained entity recognition model and relationship recognition model in Step 5, construct training sets for the entity recognition model and relationship recognition model using power grid equipment information and power grid dispatching knowledge. The training sets are respectively the core sets of the entity recognition model and relationship recognition model.

[0064] 62) For the newly added data of the new dispatching plan type that changes over time, use the existing entity recognition model and relationship recognition model to automatically complete entity extraction and relationship extraction, construct a newly added data training set, and on the already constructed core set, the entity recognition model and relationship recognition model incrementally learn the newly added data training set. The newly added data of the new dispatching plan type includes maintenance plans, power grid operation modes, etc.

[0065] 63) According to the corresponding knowledge graph instance layer update rules for different scenarios, use the key power information in the dispatching plan text after deep learning to obtain entities and entity relationships.

[0066] 64) For data such as incremental correction cases and regulation experiences, directly construct sub-graphs to expand the original knowledge graph.

[0067] 65) For time-sensitive basic data such as dispatching plans, maintenance plans, and grid connection plans, copy the power grid topology graph and overlay the time-sensitive data to construct a new topology graph.

[0068] Example: The equipment entities and relationships extracted from the maintenance plan act on the power grid topology graph to update the power grid topology entity status and connection relationships.

[0069] 66) After incrementally learning and updating the power grid dispatching knowledge graph model, perform entity alignment and attribute alignment; complete conflict detection and resolve conflicts; perform concept merging, concept super-subordinate relationship merging, and concept attribute definition merging, which specifically includes the following steps:

[0070] Align entities between different power knowledge graphs by calculating the semantic similarity between two power entities.

[0071] Use the tool methods of the power grid dispatching knowledge graph itself to detect and resolve conflicts. The tool methods include the method based on voting and the method based on quality assessment.

[0072] After completing the above steps, the fusion update of the newly added knowledge sub-graph and the original knowledge can be realized, and the knowledge graph entities and the relationships between entities are automatically updated.

[0073] A power grid dispatching knowledge graph data optimization system, characterized in that it includes the following program modules:

[0074] Model building module: Build a power grid dispatching knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph. The records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment;

[0075] Entity extraction module: Identify and extract power grid dispatching entities;

[0076] Word segmentation module: For the original corpus of the extracted power grid dispatching entity data, use an open-source word segmentation tool to perform Chinese word segmentation on the sentences and evaluate the word segmentation results;

[0077] Entity relationship triple module: Initialize entity relationship triples;

[0078] Neural network training module: Build and train an entity recognition model and a relationship recognition model based on a deep neural network to generate accurate entities and relationships;

[0079] Incremental module: Perform incremental data update and knowledge fusion.

[0080] Embodiment 3

[0081] On the basis of Steps 1 to 5 of Embodiment 1 or Embodiment 2, it further includes:

[0082] Step 7. Knowledge dynamic update, storage, and recovery based on time sections, specifically including the following steps:

[0083] Use an open-source graph database such as Neo4j for graph storage. First, mark the entities, relationships, and entity and relationship attributes in the graph database of the power grid dispatching knowledge graph with timestamps. The timestamp contains two timestamps, a start timestamp and an end timestamp. The start timestamp is the time when it is added to the knowledge graph, and the end timestamp is the time when it is deleted from the knowledge graph;

[0084] When an end timestamp is marked for a certain entity or relationship, it indicates the end of the full life cycle of the relevant knowledge;

[0085] Backups are made for the deletion of batch entities and relationships in the power grid dispatching knowledge graph for later graph recovery queries;

[0086] When inserting, deleting, or updating entities, relationships, and entity and relationship attributes, the corresponding timestamps are also updated;

[0087] When there is a need to restore the power grid dispatching knowledge graph at a specific historical moment, only the changes in the database after the specified moment need to be deleted to restore the database to the specified moment.

[0088] Power grid data has its business characteristics. Changes to the core equipment of the power grid are few, and there will be no significant increase or decrease in substations and lines within several years. Power grid fault texts have the characteristics of scattered data and rapid growth. The storage and processing of power grid fault text data are mainly considered.

[0089] Power grid fault text data grows over time. Sliced according to time information, the fault text data in the same time period is stored in adjacent memory spaces. Sliced by time and equipment, a sectional view of the knowledge graph of specific fault cases is formed, including the operating modes before and after the corresponding faults, disposal operations, risk warnings of associated equipment, equipment operations and maintenance, live working logs, limits, over-limit situations, etc.

[0090] Build a highly available cluster storage based on Neo4j. Neo4j is a high-performance network cluster graph database. Its characteristic is to store structured data on a graph rather than in a table. As an existing and feasible high-performance graph engine, it has all the characteristics of a mature database and provides the functions required for the implementation of the present invention, ensuring that even when network or hardware failures occur, the graph cluster can still continue to provide services. If a node in the cluster is damaged, or the network connection is interrupted, the graph cluster should be able to continue to provide services without completely losing the ability to serve. In a distributed environment of the knowledge graph, data can be loaded in parallel, real-time data can be updated, and multi-threaded queries can be performed, so that the query speed of the graph and the number of nodes storing the graph can grow in a weakly linear manner.

[0091] The highly available Neo4j cluster adopts a master-slave replication structure. The cluster can provide two key capabilities: one is the resilience and fault tolerance in case of hardware failures, and the other is the ability to expand the Neo4j read-intensive data scenario.

[0092] The storage solution is as Figure 5 shown. The power grid fault data such as newly added fault trip logs and dispatching logs is sliced by time, and the newly added data is distributed and stored on the cluster nodes through load balancing, realizing the scalability of the knowledge graph storage.

[0093] Load balancing performs balancing processing on the data. Each slice of data has two copies stored on different cluster nodes to ensure the balance of the data volume of each cluster node.

[0094] The specific method of load balancing is:

[0095] 1. Fault knowledge is stored in slices according to the time of fault occurrence;

[0096] 2. Cluster nodes store according to the memory weight ratio;

[0097] 3. Monitor the data volume stored in the cluster nodes. If the number of knowledge graph nodes exceeds the set number, add cluster nodes to maintain the linear growth of the graph.

[0098] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0099] A power grid dispatching knowledge graph data optimization system, characterized by including the following program modules:

[0100] Model establishment module: Establish a power grid dispatching knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph, and the records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment;

[0101] Entity extraction module: Identify and extract power grid dispatching entities;

[0102] Word segmentation module: Use an open-source word segmentation tool to perform Chinese word segmentation on the original corpus of the extracted power grid dispatching entity data, and evaluate the word segmentation results;

[0103] Entity relationship triple module: Initialize entity relationship triples;

[0104] Neural network training module: Construct and train an entity recognition model and a relationship recognition model based on a deep neural network to generate accurate entities and relationships;

[0105] Incremental module: Perform incremental data update and knowledge fusion;

[0106] Knowledge dynamic update module: Perform knowledge dynamic update, storage, and recovery based on time sections.

[0107] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the specified functions in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0108] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more of the processes and / or blocks Figure 1 of one or more of the processes and / or blocks Figure 1 specified in the flowchart(s) and / or block diagram(s).

[0109] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more of the processes and / or blocks Figure 1 of one or more of the processes and / or blocks Figure 1 specified in the flowchart(s) and / or block diagram(s).

[0110] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for optimizing power grid dispatching knowledge graph data, characterized in that, it includes the following steps: Step 1. Establish a power grid dispatching knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph. The records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment; Step 2. Identify and extract power grid dispatching entities; Step 3. For the original corpus of the extracted power grid dispatching entity data, use an open-source word segmentation tool to perform Chinese word segmentation on the sentences and evaluate the word segmentation results; Step 4. Initialize entity relationship triples; Step 5. Construct and train an entity recognition model and a relationship recognition model based on a deep neural network , Generate accurate entities and relationships; Step 6. Perform incremental data update and knowledge fusion, including: 61) Based on the trained entity recognition model and relationship recognition model in Step 5, construct training sets for the entity recognition model and relationship recognition model with power grid equipment information and power grid dispatching knowledge. The training sets are the core sets of the entity recognition model and relationship recognition model respectively; 62) For the newly added data of the new scheduling plan that is constantly changing over time, use the existing entity recognition model and relationship recognition model to automatically complete entity extraction and relationship extraction, construct a newly added data training set, and on the already constructed core set, the entity recognition model and relationship recognition model incrementally learn the newly added data training set; 63) Adopt corresponding knowledge graph instance layer update rules according to different scenarios, and use the key power information in the scheduled plan text after deep learning to obtain entities and entity relationships; 64) For incremental correction cases and regulation experience data, construct sub-graphs to expand the original knowledge graph; 65) For time-sensitive basic data, copy the power grid topology graph and overlay the time-sensitive data to construct a new topology graph; 66) After incrementally learning and updating the power grid dispatching knowledge graph model, perform entity alignment and attribute alignment; complete conflict detection and resolve conflicts; perform concept merging, concept upper and lower relationship merging, and concept attribute definition merging; Step 7. Perform knowledge dynamic update, storage, and recovery based on time slices, including: Use an open-source graph database for graph storage. First, mark the entities, relationships, and entity and relationship attributes in the graph database of the power grid dispatching knowledge graph with timestamps. The timestamp contains two timestamps, a start timestamp and an end timestamp. The start timestamp is the time when it is added to the knowledge graph, and the end timestamp is the time when it is deleted from the knowledge graph; When an entity or relationship is marked with an end timestamp, it indicates the end of the full life cycle of the relevant knowledge; Back up the deletion of batch entities and relationships in the power grid dispatching knowledge graph for later graph recovery query; When inserting, deleting, or updating entities, relationships, and entity and relationship attributes, the corresponding timestamps are also updated; When there is a need to restore the power grid dispatching knowledge graph at a specific historical moment, only delete the changes in the database after the specified moment, that is, restore to the database at the specified moment; Achieve the scalability of knowledge graph storage through load balancing. The specific method of load balancing is: Store the newly added power grid fault data in slices according to the time of fault occurrence; The newly added power grid fault data is distributed and stored on cluster nodes, and the cluster nodes store according to the memory weight storage ratio; Monitor the data volume stored by cluster nodes. If the data volume stored by a knowledge graph node exceeds the set data volume, then maintain the linear growth of the graph by adding cluster nodes.

2. The power grid dispatching knowledge graph data optimization method according to claim 1, characterized in that: In step two, identify and extract power grid dispatching entities, including the following steps: The dispatching optimization decision involves multiple types of data, including structured data and text data; The structured data is exported from the power grid real-time database, and is extracted into triples using rules and stored in the power grid dispatching knowledge graph; The text data traverses the document content using data cleaning or interface conversion methods, and divides the document according to the discourse structure association model to unify the original data format.

3. The power grid dispatching knowledge graph data optimization method according to claim 1, characterized in that: In step three, it specifically includes the following steps: 31) Use Chinese word segmentation method to segment the original corpus data into phrases; 32) Use the semi-supervised learning method based on deep neural network to complete the iterative calculation of statistical index features; And use the semi-supervised learning method to evaluate the phrase quality; Use the semi-supervised learning method based on deep neural network to iteratively mine the vocabulary, and mine high-quality vocabulary and new words; Use the semi-supervised learning method to establish a named entity classification system; 33) Use an open-source natural language preprocessing model for entity recognition, classify and cluster the dispatching optimization decision entities accordingly, extract the proper nouns of dispatching entities and attributes, and extract entity information; 34) Use the iterative training method in deep learning to make inductive distinctions between entities with the same name but different meanings and entities with different names but the same meaning.

4. The power grid dispatching knowledge graph data optimization method according to claim 1, characterized in that: In step four, it specifically includes the following steps: 41) Use the distance limit between entities and the position limit of relationship indicator words to automatically obtain entity-relationship-entity triples, and perform verification and annotation on the entities; 42) Mark the trustworthy and untrustworthy relationship triples, and use the Naive Bayes classifier to train and classify the relationship triples into trustworthy and untrustworthy, so as to obtain a relationship representation model; 43) Through the relationship representation model obtained by training, superimpose the feature data, and perform relationship recognition on the trained classifier to obtain candidate relationship triples; 44) Merge all approximate relationship candidate triples, and calculate the credibility of each relationship triple by statistically calculating the probability distribution.

5. The power grid dispatching knowledge graph data optimization method according to claim 1, characterized in that: In step five, it specifically includes the following steps: 51) Construct an entity recognition model based on deep neural network. After completing the training of the entity recognition model, within the scope of all data in the power grid dispatching knowledge graph, annotate the entity relationships in the dispatching plan text, mark the specific types corresponding to the entities, and then use the annotated corpus to train a relationship recognition model based on convolutional neural network; 52) Use the trained relationship recognition model to recognize the scheduling entity relationships in the unannotated scheduling plan text; 53) Further verify the entity relationships based on the recognition results in step 52) to achieve the consistency of entity relationships, and verify whether the relationship types exist in the relationship set according to the power grid scheduling knowledge graph. If not, prompt for review; otherwise, search for the matching entities in the power grid scheduling knowledge graph through entity semantic features. If the entity-relationship-entity triples obtained by matching recognition and the entity-relationship-entity triples obtained by searching the power grid scheduling knowledge graph are inconsistent, prompt for review again. If they are consistent, continue the review without prompting.

6. A power grid scheduling knowledge graph data optimization system Characterized in that It includes the following program modules: Model establishment module: Establish a power grid scheduling knowledge graph model based on a graph database, including a power grid historical text sub-graph and a power grid equipment sub-graph. The records of each power grid historical text sub-graph are connected to the power grid equipment sub-graph through relevant substations, lines, and equipment; Entity extraction module: Identify and extract power grid scheduling entities; Word segmentation module: Use an open-source word segmentation tool to perform Chinese word segmentation on the original corpus of the extracted power grid scheduling entity data, and evaluate the word segmentation results; Entity relationship triple module: Initialize entity relationship triples; Neural network training module: Construct and train an entity recognition model and a relationship recognition model based on a deep neural network to generate accurate entities and relationships; Incremental data update and knowledge fusion module, which performs the following steps: 61) Based on the trained entity recognition model and relationship recognition model, construct training sets for the entity recognition model and the relationship recognition model with power grid equipment information and power grid scheduling knowledge. The training sets are the core sets of the entity recognition model and the relationship recognition model respectively; 62) For the newly added scheduling plan type of new data that is constantly changing over time, use the existing entity recognition model and relationship recognition model to automatically complete entity extraction and relationship extraction, construct a new data training set, and on the already constructed core set, the entity recognition model and the relationship recognition model incrementally learn the new data training set; 63) Adopt corresponding knowledge graph instance layer update rules according to different scenarios, and use the key power information in the scheduled plan text after deep learning to obtain entities and entity relationships; 64) For incremental correction cases and regulation experience data, construct sub-graphs to expand the original knowledge graph; 65) For time-sensitive basic data, copy the power grid topology graph and overlay the time-sensitive data to construct a new topology graph; 66) After incrementally learning and updating the power grid scheduling knowledge graph model, perform entity alignment and attribute alignment; complete conflict detection and resolve conflicts; perform concept merging, concept upper and lower relationship merging, and concept attribute definition merging; Knowledge dynamic update, storage and recovery module, which performs the following steps: Using an open-source graph database for graph storage, first, timestamps are marked on the entities, relationships, and entity and relationship attributes in the graph database of the power grid dispatching knowledge graph. The timestamp contains two timestamps, a start timestamp and an end timestamp. The start timestamp is the time when the entity is added to the knowledge graph, and the end timestamp is the time when the entity is deleted from the knowledge graph; When an end timestamp is marked on a certain entity or relationship, it indicates the end of the full life cycle of the relevant knowledge; Backups are made for the deletion of a batch of entities and relationships in the power grid dispatching knowledge graph for later graph recovery queries; When inserting, deleting, or updating entities, relationships, and entity and relationship attributes, the corresponding timestamps are also updated; When there is a need to restore the power grid dispatching knowledge graph at a specific historical moment, only the changes to the database after the specified moment are deleted, that is, the database is restored to the specified moment; The scalability of knowledge graph storage is achieved through load balancing. The specific method of load balancing is as follows: The newly added power grid fault data is stored in slices according to the time of fault occurrence; The newly added power grid fault data is distributed and stored on cluster nodes, and the cluster nodes store according to the memory weight storage ratio; Monitor the data volume stored on the cluster nodes. If the data volume stored on the knowledge graph nodes exceeds the set data volume, add cluster nodes to maintain the linear growth of the graph.

Citation Information

Patent Citations

  • Knowledge graph construction method for power grid main equipment

    CN112612902A

  • Dynamic updating method and device for power grid dispatching knowledge graph

    CN112905804A