A power grid knowledge graph construction method based on multi-source data
By collecting, cleaning, and standardizing multi-source power grid data, a knowledge graph ontology framework is constructed and entities are fused, solving the data integration problem in the power grid system and realizing efficient management and fault response optimization of power grid equipment.
Patent Information
- Application Number
- CN202411809306.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-10
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-12-10
AI Technical Summary
In existing technologies, multi-source heterogeneous data in power grid systems cannot be effectively integrated and analyzed, resulting in information silos, which increases the difficulty of equipment management and fault diagnosis. There is a lack of efficient methods to convert multi-source heterogeneous data into visual knowledge graphs, making it difficult to achieve intelligent data analysis and dynamic querying.
The system collects structured, semi-structured, and unstructured data related to the power grid, cleans, standardizes, and semantically normalizes the data, constructs a knowledge graph ontology framework, defines the core entities of the power grid and their attributes and relationships, uses semantic matching and instance fusion technology to unify the same or similar entity information into a single node in the knowledge graph, and analyzes the potential relationships between devices through graph structure learning algorithms. Finally, the data is transformed into a graph form for visualization.
It enables efficient integration and unified management of multi-source power grid data, accurately expresses the logical and spatial relationships between devices, improves the intelligence level and query efficiency of data analysis, reduces data redundancy, and optimizes the management and fault response process of power grid equipment.
Smart Images

Figure CN119886298B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power grid knowledge graph construction technology, and in particular to a method for constructing a power grid knowledge graph based on multi-source data. Background Technology
[0002] With the expansion of power grid scale and the increasing demand for intelligent systems, power grid systems generate a large amount of multi-source heterogeneous data, including structured data from power production management systems, energy management systems, and monitoring sensors, as well as semi-structured and unstructured data (fault reports, environmental data) from external systems. This data comes from complex sources and is diverse in format, often making direct integration and analysis impossible, leading to information silos and increasing the difficulty of equipment management, operational status monitoring, and fault diagnosis. Current technologies commonly use data warehouses and databases to store this information, but these still have limitations in semantic understanding, unified management, and cross-system integration. Furthermore, there is a lack of an efficient method to convert multi-source heterogeneous data into a visual knowledge graph, hindering intelligent data analysis and dynamic querying. Summary of the Invention
[0003] In view of the problems existing in the prior art, the present invention is proposed.
[0004] Therefore, the technical problem to be solved by this invention is that while data warehouses and databases are commonly used to store this information, there are still limitations in terms of semantic understanding, unified management and cross-system integration of the data.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for constructing a power grid knowledge graph based on multi-source data, comprising,
[0006] Collect structured, semi-structured, and unstructured data related to the power grid to obtain multi-source data;
[0007] Cleaning, standardization, and semantic normalization of multi-source data;
[0008] Construct a knowledge graph ontology framework and define the core entities of the power grid, their attributes and relationships, and clarify the associations and logical relationships between devices;
[0009] Using semantic matching and instance fusion technology, the same or similar entity information from different data sources is unified into a single node in the knowledge graph;
[0010] The fused data is transformed into a graph format to obtain a visualized power grid knowledge graph.
[0011] As a further aspect of the present invention, the steps of cleaning, standardizing, and semantically normalizing multi-source data include:
[0012] Remove duplicate data from multi-source data to ensure that each record is unique;
[0013] Use imputation, padding, or deletion methods to fill in missing parts of multi-source data, and identify and remove outliers;
[0014] Check and adjust the consistency of field formats in multi-source data;
[0015] The categorical data is uniformly encoded, and the timestamps are aligned.
[0016] The requirements for semantic standardization of multi-source data should be clarified, including unified field naming and terminology, data mapping and fusion, and unified time format.
[0017] As a further aspect of the present invention, the steps of constructing a knowledge graph ontology framework and defining the core entities of the power grid, their attributes and relationships, and clarifying the associations and logical relationships between devices are as follows:
[0018] Based on the equipment type and function of the power grid system, entities are identified, and the acquired entities are classified according to equipment category and management level to form a set of basic nodes in the knowledge graph;
[0019] Assign specific attributes to each core entity, including equipment status, rated capacity, operating parameters, and geographical location information;
[0020] Based on the operation logic of the power grid, the relationship types between different entities are clearly defined to describe the connection methods between different devices, and entity nodes and attribute structures are linked together.
[0021] As a further aspect of the present invention: the association and logical relationship between devices are clarified by analyzing the potential relationship between devices in multi-source data using a graph structure learning algorithm.
[0022] As a further aspect of the present invention: the steps of using a graph structure learning algorithm to analyze the potential relationships between devices in multi-source data and adjusting the direction and hierarchy of relationship connections are as follows:
[0023] Multi-source data is converted into nodes and edges to construct a preliminary graph structure;
[0024] Calculate the strength of the relationship between each pair of nodes, calculate the relationship weight, and determine whether to add an edge to the graph structure.
[0025] After determining the connection strength of the edges, the direction and level of the edges are set based on the relationship strength;
[0026] Adjust the direction and hierarchy of the edges to form the final knowledge graph structure.
[0027] As a further aspect of the present invention: the step of using semantic matching and instance fusion technology to unify identical or similar entity information from different data sources into a single node in a knowledge graph is as follows:
[0028] Using the BERT word embedding model, entity names are converted into vector representations, and the original text data is mapped to the semantic space;
[0029] Cosine similarity is calculated by applying it to entity vectors from each data source to obtain similarity scores between entity pairs.
[0030] Set a matching threshold and apply the set matching threshold to the calculated entity pair similarity scores;
[0031] Analyze the attribute values of each entity pair in the candidate matching pairs, and merge identical or similar entities into one node based on semantic similarity;
[0032] The instance fusion algorithm is used to cluster entities, merging similar entities into a single node;
[0033] The merged entity nodes are inserted into the node set of the knowledge graph, and duplicate entity nodes are removed.
[0034] For merged nodes, their attributes and relationships are integrated to form a unified node representation in the knowledge graph;
[0035] Identical or similar entities from all data sources are unified into a single node.
[0036] As a further aspect of the present invention: the step of using an instance fusion algorithm to cluster entities and merging similar entities into a single node is as follows:
[0037] Extract key features from each entity from different data sources;
[0038] The features of each entity are combined into a feature vector, and the similarity matrix is generated using the feature vector.
[0039] The DBSCAN density clustering algorithm is used to cluster the similarity matrix, grouping entities with high similarity into the same cluster;
[0040] The boundaries of each cluster are determined by density threshold and minimum number of samples, and similar entities that meet the conditions are divided into a single group.
[0041] For each cluster in the clustering results, all entities within the cluster are merged into a single node.
[0042] As a further aspect of the present invention: the step of analyzing the attribute values of each entity pair in the candidate matching pairs is as follows:
[0043] Extract the key attribute values of each candidate entity pair, and use the feature similarity calculation method to compare the attribute values of each candidate entity pair item by item;
[0044] The similarity scores of each attribute are aggregated into a single overall similarity score;
[0045] Based on the overall similarity score and the set matching threshold, candidate pairs with similarity scores higher than the threshold are selected as the final matching pairs.
[0046] As a further aspect of the present invention: the step of converting the fused data into a graph format to obtain a visualized power grid knowledge graph is as follows:
[0047] In the constructed graph structure, add attribute values to each node and edge;
[0048] The nodes and edges in the graph structure are arranged, and the positions of the nodes are adjusted in two-dimensional space by simulating the attractive and repulsive forces between the nodes.
[0049] Import the redesigned graph structure into a visualization tool to generate a visual representation of the graph.
[0050] By setting different colors and sizes for each node in the visualization interface, a diagram of the structure and interrelationships of power grid equipment can be obtained.
[0051] As a further aspect of the present invention: the nodes and edges in the graph structure are arranged, and the attractive and repulsive forces between the nodes are calculated using the following formula:
[0052]
[0053] Among them, f repel (d) represents the repulsive force between nodes, f attract (d) represents the attraction between edges, where d is the distance between nodes and k is the layout constant.
[0054] Compared with existing technologies, the advantages of this invention are as follows: This method for constructing a power grid knowledge graph based on multi-source data converts multi-source power grid data into nodes and edges to build a preliminary graph structure, and introduces methods for calculating relationship strength and weight, making the addition of edges and directions in the graph structure more accurate. By quantifying edge weights based on factors such as feature similarity and geographical distance between nodes, and setting directionality and hierarchy, the logical and spatial relationships between power grid equipment can be more accurately reflected. Attached Figure Description
[0055] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0056] Figure 1 This is a schematic diagram of the overall process described in the embodiments provided by the present invention.
[0057] Figure 2 This is a flowchart illustrating step S2 of an embodiment provided by the present invention.
[0058] Figure 3 This is a flowchart illustrating step S3 of an embodiment provided by the present invention.
[0059] Figure 4 This is a flowchart illustrating step S5 of an embodiment provided by the present invention. Detailed Implementation
[0060] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0061] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0062] Secondly, the present invention will be described in detail with reference to the schematic diagrams. When describing the embodiments of the present invention, for ease of explanation, the cross-sectional views illustrating the device structure will be partially enlarged, not according to the usual scale. Furthermore, the schematic diagrams are merely examples and should not limit the scope of protection of the present invention. In addition, actual fabrication should include the three-dimensional spatial dimensions of length, width, and depth.
[0063] Furthermore, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in an embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0064] Example 1
[0065] like Figures 1-3As shown, this invention provides a technical solution: a method for constructing a power grid knowledge graph based on multi-source data, comprising: S1: collecting structured, semi-structured, and unstructured data related to the power grid to obtain multi-source data; S2: cleaning, standardizing, and semantically normalizing the multi-source data; S3: constructing a knowledge graph ontology framework and defining the core entities of the power grid and their attributes and relationships, clarifying the associations and logical relationships between devices; S4: using semantic matching and instance fusion technology to unify the same or similar entity information from different data sources into a single node in the knowledge graph; S5: converting the fused data into a graph form to obtain a visualized power grid knowledge graph.
[0066] The multi-source data includes structured, semi-structured and unstructured data related to the power grid; the structured data includes structured data such as equipment ledgers and operation records collected from power production management systems, energy management systems and monitoring sensors; semi-structured data extracted from equipment monitoring reports and alarm information; and unstructured text data collected from fault analysis reports, maintenance records and external public data (meteorological and geographical data).
[0067] It should be noted that: data in the data source is collected through API interfaces; structured data from the power production management system, energy management system, and monitoring sensors can be obtained directly from the equipment management system and monitoring platform through API interfaces; equipment monitoring reports and alarm information are also extracted from the monitoring system through API interfaces, and this data is usually semi-structured; fault analysis reports, maintenance records, and external public data (meteorological and geographic data) are obtained from external data sources such as fault management systems, meteorological platforms, and geographic information systems through API interfaces.
[0068] Specifically, the main difference between structured, semi-structured, and unstructured data lies in their organization and storage format. Structured data has a fixed format, like tables in a database, where data is stored in rows and columns, making it easy to query and analyze. Semi-structured data falls between the two; although it doesn't have a strict table format, the data has a certain structure, such as XML or JSON files, which contain tags or field names. Unstructured data has no fixed format at all and can be in the form of text, images, audio, etc., requiring additional processing to extract useful information.
[0069] Furthermore, equipment ledger data includes, but is not limited to, information such as the equipment's unique identifier, type, model, installation date, location, status, manufacturer, and maintenance records, as well as the equipment's rated capacity and lifespan; structured operation record data includes, but is not limited to, real-time monitoring of power data such as current, voltage, power, and frequency, equipment operating status, load information, fault records, alarm information, environmental parameters (temperature, humidity), energy consumption data, and power grid performance indicators.
[0070] More specifically, the semi-structured data extracted from equipment monitoring reports and alarm information typically includes the following: equipment fault type (short circuit, overload, equipment aging), alarm level (critical, warning, informational), fault description (including the specific circumstances of the fault occurrence, abnormal equipment operation, etc.), equipment ID (used to identify the specific equipment that has failed), fault occurrence time (time stamp of the alarm or fault occurrence), fault location (location of the equipment or area where the fault occurred), recovery measures (repair, replacement, configuration adjustment), and processing result (repaired, pending processing).
[0071] More specifically, unstructured data collected from fault analysis reports, maintenance records, and external public data (meteorological and geographical data) includes fault descriptions (cause and occurrence of equipment failure), time and location of the failure, fault handling measures (repair or replacement of equipment), and maintenance activity records (description of equipment inspections, operators, and maintenance results). In addition, it includes external meteorological data (the impact of weather on equipment), geographical data (the geographical location and environmental characteristics of the equipment), and environmental influencing factors (the impact of temperature changes on the equipment).
[0072] Furthermore, the multi-source data undergoes cleaning, standardization, and semantic normalization, including removing redundancy, standardizing formats, and performing word segmentation and part-of-speech tagging on unstructured data to ensure data consistency. The steps involved in cleaning, standardizing, and semantically normalizing the multi-source data include:
[0073] S21: Remove duplicate data from multi-source data to ensure that each record is unique;
[0074] S22: Fill in missing parts of multi-source data using imputation, padding, or deletion methods, and identify and remove outliers;
[0075] S23: Check and adjust the consistency of field formats for multi-source data; for example, unify dates to the “YYYY-MM-DD” format to ensure seamless data integration.
[0076] S24: Standardize the categorical data and align the timestamps; specifically, in data standardization, standardize the numerical units or ranges to ensure that data from different sensors are comparable; standardize the categorical data to avoid ambiguity caused by different expressions; at the same time, align the timestamps to ensure that data from different sources can be compared within the same time frame.
[0077] S25: Clearly define the semantic standardization requirements for multi-source data, unify field naming and terminology, and perform data mapping and fusion to unify time formats; specifically, data semantic standardization requires unifying field naming and terminology, for example, unifying "temperature" and "temp" into "temperature" to avoid different words representing the same meaning; then, perform data mapping and fusion to ensure that data from different sources can be correctly matched; finally, unify the time format to ensure that all time data conforms to the same standard (such as ISO 8601) to avoid processing problems caused by inconsistent formats.
[0078] Furthermore, the steps for constructing a knowledge graph ontology framework and defining the core entities of the power grid, their attributes and relationships, and clarifying the associations and logical relationships between devices are as follows:
[0079] S31: Based on the equipment type and function of the power grid system, identify entities and classify the acquired entities according to equipment category and management level to form a set of basic nodes in the knowledge graph;
[0080] S32: Assign specific attributes to each core entity, including equipment status, rated capacity, operating parameters, and geographical location information;
[0081] S33: Based on the operation logic of the power grid, clarify the relationship types between different entities to describe the connection methods between different devices, and connect entity nodes and attribute structures.
[0082] It should be noted that: structured data (such as equipment ledgers and operation records) provides basic information about power grid equipment (equipment number, model, rated capacity, etc.), providing a direct basis for identifying core entities in the power grid (such as transformers, circuit breakers, cables, etc.); semi-structured data (such as equipment monitoring reports and alarm information) supports the definition of attributes such as the operating status, fault type, and alarm information of power grid equipment; and unstructured data (such as fault analysis reports, maintenance records, and meteorological data) provides background information on the relationships between power grid equipment, the scope of fault impact, and the impact of the external environment.
[0083] The acquired structured data provides the foundation for identifying core entities in the power grid (such as transformers and circuit breakers) and assigning them attributes (such as equipment number, rated capacity, and operating status); semi-structured and unstructured data help define the types of relationships between devices and provide deeper levels of equipment status and fault background (such as fault causes and impact range in fault analysis reports); through this step, these different types of data are merged and structured to form a comprehensive knowledge graph that reflects the power grid equipment, attributes, and interrelationships.
[0084] It should be noted that clarifying the connections and logical relationships between devices involves analyzing the potential relationships between devices in multi-source data using graph structure learning algorithms. Specifically, after constructing the relationships between entities, graph structure learning algorithms are used to analyze the potential relationships between devices in multi-source data, and the direction and hierarchy of relationship connections are adjusted. Finally, the defined entity nodes, attributes, and relationships are integrated into the ontology framework of the knowledge graph, and a graph database is used to implement the structured storage of the knowledge graph. By constructing the ontology framework of the knowledge graph and defining the attributes and relationships of the core entities of the power grid, the hierarchical structure, operating status, and spatial distribution characteristics of power grid equipment can be systematically organized and represented. By analyzing and optimizing the potential relationships between entities through graph structure learning algorithms, the connection methods and hierarchical relationships between devices become more accurate. After integrating all entities, attributes, and relationships into the knowledge graph, the graph can be stored in a structured manner through a graph database, ensuring the efficiency of data querying, analysis, and visualization. It should be noted that multi-source data specifically refers to structured data collected from different sources, including equipment ledgers and operation records in power production management systems and energy management systems; semi-structured data includes equipment monitoring reports and alarm information, which are usually stored in XML or JSON format; unstructured data includes fault analysis reports, maintenance records, meteorological data, and geographic data.
[0085] The steps for using graph structure learning algorithms to analyze the potential relationships between devices in multi-source data and to adjust the direction and hierarchy of relationship connections are as follows: convert the multi-source data into the form of nodes and edges to construct a preliminary graph structure; calculate the strength of the relationship between each pair of nodes, calculate the relationship weight, and determine whether to add an edge to the graph structure; after determining the connection strength of the edge, set the direction and hierarchy of the edge based on the relationship strength; adjust the direction and hierarchy of the edge to form the final knowledge graph structure.
[0086] Specifically, in this embodiment, the steps of using graph structure learning algorithms to analyze the potential relationships between devices in multi-source data and adjusting the direction and hierarchy of relationship connections are as follows:
[0087] Multi-source power grid data is converted into the form of nodes and edges to construct a preliminary graph structure G = (V, E), where V represents the set of nodes in the graph, i.e. the equipment in the power grid (transformers, load points, circuit breakers, etc.), and E represents the set of edges in the graph, i.e. the potential relationships between different equipment (power supply relationships, spatial adjacency, etc.).
[0088] Subsequently, the strength of the relationship between each pair of nodes (u, v) ∈ V is calculated using the formula w. uv =α·S uv +β·D uv Calculate relation weight w uv Determine whether to add an edge to the graph structure G, where w uvThe weight representing the relationship between nodes u and v determines whether an edge is created in the graph structure. S uv D represents the similarity score between nodes u and v, quantified by feature similarity (voltage level, geographical location). uv This represents the geographic or physical distance between nodes u and v, quantifying the spatial relationship between devices. α and β are adjustable parameters used to balance the effects of similarity and distance.
[0089] After determining the connection strength w of the edge uv Then, based on the strength of the relationship, the direction and level of the edge are set, and the directionality of the edge is defined as follows:
[0090] Among them, R uv It is the strength of the relationship from node u to v, determined by the weight w. uv The calculation determines the direction of the relationship, when R uv >R vu When, the edge points from u to v; when R uv <R vu When, the edge points from v to u; when R uv =R vu When setting bidirectional edges, the direction and hierarchical relationship of the edges are adjusted through relation strength and directionality calculations to form the final knowledge graph structure. By converting multi-source power grid data into the form of nodes and edges, a preliminary graph structure is constructed, and the calculation methods of relation strength and weight are introduced to make the addition of edges and directions in the graph structure more accurate. The weight of the edges is quantified based on factors such as feature similarity and geographical distance between nodes, and the directionality and hierarchy are set to more accurately reflect the logical and spatial relationships between power grid equipment.
[0091] Furthermore, the steps for unifying identical or similar entity information from different data sources into a single node in the knowledge graph using semantic matching and instance fusion techniques are as follows:
[0092] S41: Using the BERT word embedding model, entity names are converted into vector representations, and the original text data is mapped to the semantic space;
[0093] S42: Apply cosine similarity calculation to the entity vectors in each data source to obtain the similarity score between entity pairs;
[0094] S43: Set a matching threshold and apply the set matching threshold to the calculated entity pair similarity scores;
[0095] S44: Analyze the attribute values of each entity pair in the candidate matching pairs, and merge identical or similar entities into one node based on semantic similarity;
[0096] S45: Use the instance fusion algorithm to cluster entities and merge similar entities into a single node;
[0097] S46: Insert the merged entity nodes into the node set of the knowledge graph and remove duplicate entity nodes;
[0098] S47: For merged nodes, integrate their attributes and relationships to form a unified node representation in the knowledge graph;
[0099] S48: Identical or similar entities from all data sources are unified into a single node.
[0100] Specifically, in this embodiment, the steps of using semantic matching and instance fusion technology to unify identical or similar entity information from different data sources into a single node in the knowledge graph are as follows:
[0101] The BERT word embedding model is used to convert entity names into vector representations, mapping the original text data into a semantic space. Then, the formula is used... Cosine similarity is calculated for entity vectors from each data source to obtain the similarity score between entity pairs. Here, u and v are the word vector representations of the two entities, u·v represents the inner product of the vectors, and |u| and |v| represent the magnitudes of the vectors, respectively. The similarity score sim(u,v) is used to measure the degree of similarity between the two entities. The higher the similarity, the more likely they are the same entity.
[0102] Set a matching threshold T, apply the set matching threshold T to the calculated entity pair similarity scores, retain only entity pairs with similarity greater than the matching threshold T as candidate matching pairs and filter out combinations with low similarity. Then, analyze the attribute values of each entity pair in the candidate matching pairs, and merge the same or similar entities into one node according to semantic similarity. Subsequently, apply the instance fusion algorithm to cluster the entities and merge similar entities into a single node.
[0103] The merged entity nodes are inserted into the knowledge graph's node set, and duplicate entity nodes are removed. Then, for the merged nodes, their attributes and relationships are integrated to form a unified node representation in the knowledge graph. Ultimately, identical or similar entities from all data sources are unified into a single node, providing a consistent data structure for knowledge graph analysis. Cosine similarity is used to calculate similarity scores, accurately measuring the similarity of entities from different data sources. By setting a matching threshold, only entity pairs with similarity higher than the threshold are retained as candidate matching pairs, ensuring matching accuracy. Attribute values are further analyzed in the candidate pairs. Through semantic similarity and instance fusion algorithms, identical or similar entities are aggregated into a single node, and the merged entity is inserted into the knowledge graph. Redundant nodes are removed, and attributes and relationships are integrated to form a unified node representation. This effectively reduces duplicate data in the graph, achieves the merging of identical or similar entities from different data sources, and provides a standardized and consistent structure for subsequent knowledge graph analysis.
[0104] Furthermore, the steps for clustering entities using an instance fusion algorithm, merging similar entities into a single node, are as follows:
[0105] Extract key features from each entity from different data sources;
[0106] The features of each entity are combined into a feature vector, and the similarity matrix is generated using the feature vector.
[0107] The DBSCAN density clustering algorithm is used to cluster the similarity matrix, grouping entities with high similarity into the same cluster;
[0108] The boundaries of each cluster are determined by density threshold and minimum number of samples, and similar entities that meet the conditions are divided into a single group.
[0109] For each cluster in the clustering results, all entities within the cluster are merged into a single node.
[0110] Specifically, in this embodiment, the steps of using the instance fusion algorithm to cluster entities and merge similar entities into a single node are as follows:
[0111] Extract key features from each entity in different data sources, including device type, status, and location. It is the similarity between entities i and j, x i and x j These are the feature vectors of two entities, |x i | and |x j | represents the magnitude of the vector;
[0112] The DBSCAN density clustering algorithm is used to cluster the similarity matrix, grouping entities with high similarity into the same cluster. The DBSCAN algorithm determines the boundary of each cluster by using the density threshold ∈ and the minimum number of samples, dividing similar entities that meet the conditions into a single group, which facilitates the subsequent merging of entities in the same group.
[0113] For each cluster in the clustering results, all entities within the cluster are merged into a single node. The merging operation integrates the attribute values and relationships of each entity, forming a new unified node. This ensures that duplicate entities in the graph are removed, and the clustered nodes are retained as standard nodes for the knowledge graph. Finally, the merged single node is inserted into the knowledge graph, removing duplicate nodes and updating the relationship connections between nodes. After the graph update, similar entities are represented by only a single node in the knowledge graph, thus simplifying the graph structure and reducing redundant information. By extracting key features and representing each entity as a feature vector, a similarity matrix is generated, and then the DBSCAN density clustering algorithm is applied to group entities with high similarity into the same cluster, achieving efficient clustering of similar entities. Cluster boundaries are set using density thresholds and minimum sample size, allowing entities with similar conditions to be merged into a single node, integrating their attributes and relationships, thereby removing duplicate data in the knowledge graph. Finally, the merged node is inserted into the graph, updating node relationships, making the graph structure more concise and clear, effectively reducing redundant information, and providing a standardized structural foundation for knowledge graph management and subsequent analysis.
[0114] Furthermore, the step of analyzing the attribute values of each entity pair in the candidate matching pairs is as follows:
[0115] Extract the key attribute values of each candidate entity pair, and use the feature similarity calculation method to compare the attribute values of each candidate entity pair item by item;
[0116] The similarity scores of each attribute are aggregated into a single overall similarity score;
[0117] Based on the overall similarity score and the set matching threshold, candidate pairs with similarity scores higher than the threshold are selected as the final matching pairs.
[0118] Specifically, in this embodiment, the step of further analyzing the attribute values of each entity pair in the candidate matching pairs is as follows:
[0119] Key attribute values are extracted from each candidate entity pair. Using a feature similarity calculation method, the attribute values of each candidate entity pair are compared item by item. For textual attributes, semantic similarity is calculated using cosine similarity; for numerical attributes, numerical similarity is calculated. The expression for numerical similarity is: Where a and b are two values of a numerical attribute, sim num (a, b) represents numerical similarity, with values closer to 1 indicating greater similarity;
[0120] Then, through the formula The similarity scores of each attribute are aggregated into a single overall similarity score S. uv To assess the overall similarity of candidate entity pairs, where S uv It is the combined similarity score between candidate entities u and v, where n represents the number of attributes, sim i (u, v) is the similarity score of the i-th attribute;
[0121] Based on the comprehensive similarity score S uv Based on a set matching threshold, candidate pairs with similarity scores higher than the threshold are selected as final matching pairs. Only entity pairs with similarity scores reaching the threshold are considered the same entity and proceed to the next instance fusion stage. By extracting key attribute values from candidate entity pairs and calculating similarity for each attribute, the similarity of textual and numerical attributes is accurately measured. Cosine similarity is used to evaluate the semantic similarity of textual attributes, and numerical similarity is used to quantify the closeness of numerical attributes. The similarity scores of each attribute are aggregated into an overall similarity score, which is then compared with the set matching threshold. Only entity pairs with an overall similarity score higher than the threshold are retained as final matching pairs, thereby improving matching accuracy and ensuring that candidate pairs have high similarity, thus providing a reliable foundation for subsequent instance fusion.
[0122] In this embodiment, various types of data are collected from power production management, energy management, monitoring sensors, and external data sources. Through data cleaning, standardization, and semantic normalization, consistency of multi-source data is achieved. By constructing an ontology framework of a knowledge graph, the core entities of the power grid and their attributes and relationships are defined, thereby accurately reflecting the associations and logical relationships between devices in the knowledge graph. Using semantic matching and instance fusion technology, the same or similar entity information from different data sources is merged into a single node, eliminating data redundancy and duplication. Finally, the processed data is transformed into a graph form and visualized, realizing the associated display and efficient management of power grid data.
[0123] Example 2
[0124] like Figure 4As shown, this embodiment differs from the previous embodiment in that it further describes the processing steps for converting the fused data into a graph form to obtain a visualized power grid knowledge graph. The steps are as follows: S51: In the constructed graph structure, add attribute values to each node and edge; S52: Layout the nodes and edges in the graph structure and adjust the positions of the nodes in two-dimensional space by simulating the attraction and repulsion forces between nodes; S53: Import the layout-adjusted graph structure into a visualization tool to generate a visualization display interface for the graph; S54: Set different colors and sizes for each node in the visualization interface to obtain a graph of the structure and interrelationships of power grid equipment.
[0125] Specifically, the steps to transform the fused data into a graph format and visualize the power grid knowledge graph are as follows:
[0126] In the constructed graph structure G, attribute values are added to each node and edge, including the node type, device status, operating parameters, and the relationship type and directionality of the edge. Node attribute values are automatically assigned through the node feature allocation model, while edge attributes are set through association rules. The attribute configuration of nodes and edges maps device features and relationship type information to the graph, providing a complete attribute representation for the graph.
[0127] Through formula The nodes and edges in the graph structure are arranged, and the attractive and repulsive forces between nodes are simulated to adjust the position of nodes in two-dimensional space to reduce node overlap and edge intersection. Here, f repel (d) represents the repulsive force between nodes, f attract (d) represents the attractive force between edges, where d is the distance between nodes and k is the layout constant used to balance the repulsive and attractive forces.
[0128] The redesigned graph structure is imported into a visualization tool to generate a visual representation of the graph. In this visualization, each node is assigned a different color and size based on its device type and status, and each edge uses a different style and line color according to its relationship type, allowing the graph to display the structure and interrelationships of the power grid equipment. By assigning detailed attribute values to the nodes and edges in the graph structure, a comprehensive representation of the characteristics and relationship types of the power grid equipment is ensured, and automatic configuration is achieved using a node feature allocation model and association rules. The placement of nodes and edges is optimized using attraction and repulsion models, effectively reducing node overlap and edge intersections, resulting in a clear distribution of the graph structure in two-dimensional space. Finally, the adjusted graph structure is imported into the visualization tool, where nodes and edges are assigned different colors and styles based on their device type and relationship type, thus intuitively presenting the structure and interrelationships of the power grid equipment.
[0129] Example 3
[0130] This embodiment differs from the previous embodiment in that it further describes a comparison between the present invention and rule-based data modeling methods in the prior art.
[0131] Specifically, rule-based data modeling uses predefined rules to model power grid equipment and the relationships between them. These rules are manually set, and logical reasoning is used to construct the connections between equipment, such as equipment fault analysis and state reasoning. The same power grid dataset is used in the comparative experiments, including: Equipment information: equipment number, type, model, rated capacity, and geographical location; Operational data: real-time operational data of equipment such as voltage, current, and power; Alarm and fault data: equipment fault records, alarm information, and processing status. During the comparative experiments, a dataset of the same size is selected, containing 10,000 equipment information entries and 50,000 fault records.
[0132] All experiments were run on the same hardware configuration, using a server with an 8-core CPU, 16GB of memory, and 500GB of storage. The operating system was Ubuntu 20.04, and the database platform was Neo4j. This invention uses the Neo4j database platform to construct and store the knowledge graph of power grid equipment; existing solutions use rule-based data modeling methods, with rule reasoning performed through expert systems or custom rule engines.
[0133] Key indicators of the comparative experiment include: data processing efficiency: the time required for different schemes to process the same amount of data; device relationship accuracy: the accuracy of different schemes in device connection and dependency identification; query response time: the response time for operations such as device information query and fault analysis; after comparison, the beneficial effects of the present invention are demonstrated through data or charts.
[0134] By comparing the processing time of different schemes with the same amount of data, the efficiency advantage of this invention in processing power grid equipment data can be demonstrated. Assume that 10,000 devices and 50,000 fault records need to be processed:
[0135] plan Data processing time (seconds) This invention (based on graph database) 120 Existing solutions (rule-based solutions) 220
[0136] Among them, the accuracy comparison of device relationships: by comparing the performance of the present invention with that of the prior art in the accuracy of device relationship identification (such as connection relationship, dependency relationship, control relationship, etc. between devices), the relationship expression capability of the present invention is demonstrated.
[0137] plan Equipment relationship accuracy (%) This invention (based on graph database) 95 Existing solutions (rule-based solutions) 85
[0138] Among them, the query response time comparison demonstrates the advantages of this invention in query efficiency by comparing the response times of different solutions when performing operations such as device status query and fault location.
[0139] plan Query response time (milliseconds) This invention (based on graph database) 50 Existing solutions (rule-based solutions) 200
[0140] Comparative experimental data shows that the present invention demonstrates significant advantages over existing technologies in data processing efficiency, accuracy of device relationships, query response time, and accuracy of fault detection and location. Through a power grid knowledge graph supported by a graph database, the present invention can more efficiently and accurately express the complex relationships and logical structures of power grid equipment, optimizing power grid management and fault response processes. These experimental results effectively demonstrate the outstanding performance of the present invention in expressing the association and logical relationships of power grid equipment.
[0141] It is important to note that the constructions and arrangements of this application shown in several different exemplary embodiments are merely illustrative. Although only a few embodiments are described in detail in this disclosure, those who consult this disclosure will readily understand that many modifications are possible without substantially departing from the novel teachings and advantages of the subject matter described in this application. For example, variations in the size, dimensions, structure, shape, and proportions of various elements, as well as parameter values such as temperature, pressure, etc., installation arrangements, use of materials, color, orientation, etc. For instance, an element shown as integrally formed may be composed of multiple parts or elements, the position of elements may be inverted or otherwise altered, and the nature or number or position of discrete elements may be changed or altered. Therefore, all such modifications are intended to be included within the scope of the invention. The order or sequence of any process or method steps may be changed or rearranged according to alternative embodiments. In the claims, any "device plus function" clause is intended to cover the structure performing the function described herein, and not only structural equivalents but also equivalent structures. Other substitutions, modifications, alterations, and omissions may be made in the design, operation, and arrangement of the exemplary embodiments without departing from the scope of the invention. Therefore, the present invention is not limited to the specific embodiments, but extends to various modifications that still fall within the scope of the appended claims.
[0142] Furthermore, in order to provide a concise description of exemplary embodiments, not all features of actual embodiments may be described, i.e., those features that are not relevant to the currently considered best mode for carrying out the invention, or those features that are not relevant to implementing the invention.
[0143] It should be understood that numerous specific implementation decisions can be made during the development of any practical implementation, such as in any engineering or design project. Such development efforts may be complex and time-consuming, but for those of ordinary skill in the art who benefit from this disclosure, the development effort will be a routine task in design, manufacturing, and production without requiring extensive experimentation.
[0144] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for constructing a power grid knowledge graph based on multi-source data, characterized in that: include, Collect structured, semi-structured, and unstructured data related to the power grid to obtain multi-source data; Cleaning, standardization, and semantic normalization of multi-source data; Construct a knowledge graph ontology framework and define the core entities of the power grid, their attributes and relationships, and clarify the associations and logical relationships between devices; Using semantic matching and instance fusion technology, the same or similar entity information from different data sources is unified into a single node in the knowledge graph; The fused data is transformed into a graph format to obtain a visualized power grid knowledge graph; The steps involved in constructing a knowledge graph ontology framework, defining core power grid entities and their attributes and relationships, and clarifying the associations and logical relationships between devices are as follows: Based on the equipment type and function of the power grid system, entities are identified, and the acquired entities are classified according to equipment category and management level to form a set of basic nodes in the knowledge graph; Assign specific attributes to each core entity, including equipment status, rated capacity, operating parameters, and geographical location information; Based on the operation logic of the power grid, the relationship types between different entities are clearly defined to describe the connection methods between different devices, and entity nodes and attribute structures are linked together. Among these methods, clarifying the connections and logical relationships between devices involves using graph structure learning algorithms to analyze the potential relationships between devices in multi-source data; The steps involved in using graph structure learning algorithms to analyze the potential relationships between devices in multi-source data and adjusting the direction and hierarchy of these relationships are as follows: Multi-source data is converted into nodes and edges to construct a preliminary graph structure; Calculate the strength of the relationship between each pair of nodes, calculate the relationship weight, and determine whether to add an edge to the graph structure. After determining the connection strength of the edges, the direction and level of the edges are set based on the relationship strength; Adjust the direction and hierarchy of the edges to form the final knowledge graph structure.
2. The method for constructing a power grid knowledge graph based on multi-source data as described in claim 1, characterized in that: The steps involved in cleaning, standardizing, and semantically normalizing multi-source data include: Remove duplicate data from multi-source data to ensure that each record is unique; Use imputation, padding, or deletion methods to fill in missing parts of multi-source data, and identify and remove outliers; Check and adjust the consistency of field formats in multi-source data; The categorical data is uniformly encoded, and the timestamps are aligned. The requirements for semantic standardization of multi-source data should be clarified, including unified field naming and terminology, data mapping and fusion, and unified time format.
3. The method for constructing a power grid knowledge graph based on multi-source data as described in claim 1 or 2, characterized in that: The steps for unifying identical or similar entity information from different data sources into a single node in a knowledge graph using semantic matching and instance fusion techniques are as follows: Using the BERT word embedding model, entity names are converted into vector representations, and the original text data is mapped to the semantic space; Cosine similarity is calculated by applying it to entity vectors from each data source to obtain similarity scores between entity pairs. Set a matching threshold and apply the set matching threshold to the calculated entity pair similarity scores; Analyze the attribute values of each entity pair in the candidate matching pairs, and merge identical or similar entities into one node based on semantic similarity; The instance fusion algorithm is used to cluster entities, merging similar entities into a single node; The merged entity nodes are inserted into the node set of the knowledge graph, and duplicate entity nodes are removed. For merged nodes, their attributes and relationships are integrated to form a unified node representation in the knowledge graph; Identical or similar entities from all data sources are unified into a single node.
4. The method for constructing a power grid knowledge graph based on multi-source data as described in claim 3, characterized in that: The steps for clustering entities using an instance fusion algorithm, merging similar entities into a single node, are as follows: Extract key features from each entity from different data sources; The features of each entity are combined into a feature vector, and the similarity matrix is generated using the feature vector. The DBSCAN density clustering algorithm is used to cluster the similarity matrix, grouping entities with high similarity into the same cluster; The boundaries of each cluster are determined by density threshold and minimum number of samples, and similar entities that meet the conditions are divided into a single group. For each cluster in the clustering results, all entities within the cluster are merged into a single node.
5. The method for constructing a power grid knowledge graph based on multi-source data as described in claim 4, characterized in that: The steps for analyzing the attribute values of each entity pair in the candidate matching pairs are as follows: Extract the key attribute values of each candidate entity pair, and use the feature similarity calculation method to compare the attribute values of each candidate entity pair item by item; The similarity scores of each attribute are aggregated into a single overall similarity score; Based on the overall similarity score and the set matching threshold, candidate pairs with similarity scores higher than the threshold are selected as the final matching pairs.
6. The method for constructing a power grid knowledge graph based on multi-source data as described in claim 4 or 5, characterized in that: The steps to transform the fused data into a graph format to obtain a visualized power grid knowledge graph are as follows: In the constructed graph structure, add attribute values to each node and edge; The nodes and edges in the graph structure are arranged, and the positions of the nodes are adjusted in two-dimensional space by simulating the attractive and repulsive forces between the nodes. Import the redesigned graph structure into a visualization tool to generate a visual representation of the graph. By setting different colors and sizes for each node in the visualization interface, a diagram of the structure and interrelationships of power grid equipment can be obtained.
7. The method for constructing a power grid knowledge graph based on multi-source data as described in claim 6, characterized in that: The nodes and edges in the graph structure are arranged, and the attractive and repulsive forces between the nodes are calculated using the following formula: in, This represents the repulsive force between nodes. Indicates the attraction between edges. It is the distance between nodes. It is a layout constant.
Citation Information
Patent Citations
Power grid monitoring field knowledge graph construction method based on multi-source data fusion
CN113094516A
Multi-source data difference traceability retrieval method based on knowledge graph
CN115809345A