Knowledge Graph Construction Method and Related System for Multi-Source Heterogeneous Data of Substation Main Equipment

Through the acquisition, processing and fusion of multi-source heterogeneous data, a knowledge graph of the substation main equipment is built, the problem of insufficient data integrity is solved, and a comprehensive understanding of substation equipment is achieved and efficient operation and maintenance is improved, and the accuracy of fault diagnosis and operation and maintenance strategies is improved.

CN119474148BActive Publication Date: 2025-07-11STATE GRID JIANGSU ELECTRIC POWER CO LTD SUZHOU BRANCH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411441208.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-16
Publication Date
2025-07-11
Estimated Expiration
2044-10-16

AI Technical Summary

Technical Problem

The existing knowledge graph construction method for multi-source heterogeneous data of substation main equipment. When the data integrity is poor, the knowledge graph constructed is not comprehensive enough to achieve a comprehensive understanding of substation equipment and efficient operation and maintenance.

Method used

Through multi-source heterogeneous data acquisition, primary processing, data augmentation algorithm, knowledge extraction, fusion and optimization processing, a knowledge graph of the substation main device is built, and technologies such as database connection, natural language processing, neural network model, graph database storage are used to identify and fuse the device entities, relationships and attributes, and give key nodes higher attention weights.

Benefits of technology

It improves the comprehensiveness of data utilization and knowledge graph, enhances the diversity and integrity of data, and improves the efficiency and accuracy of equipment fault diagnosis and operation and maintenance strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119474148B_ABST
    Figure CN119474148B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data processing, and specifically to a method for constructing a knowledge graph of multi-source heterogeneous data of main substation equipment and related systems, including multi-source heterogeneous data collection. For structured database data, a database connection tool is used to obtain the required data by writing SQL query statements; for semi-structured data, corresponding parsing libraries are used for parsing to extract key information and convert it into a structured data format; for unstructured text data and web data, natural language processing technology and web crawler technology are adopted. The method for constructing a knowledge graph of multi-source heterogeneous data of main substation equipment in the present invention improves data utilization rate, and in the secondary data processing, by introducing a data enhancement algorithm, data can be enhanced and supplemented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and specifically to a method for constructing a knowledge graph of multi-source heterogeneous data of main substation equipment and related systems. Background Technique

[0002] A knowledge graph is a graphical structure used to represent knowledge, where entities and relationships are represented as nodes and edges in the graph. The construction process of a knowledge graph includes extracting entities and relationships from structured and unstructured data and organizing them into a meaningful graph. The construction of a knowledge graph for multi-source heterogeneous data of main substation equipment refers to the process of constructing a knowledge graph for data related to main substation equipment from different sources and with different structures. By integrating multi-source heterogeneous data, a comprehensive understanding and unified management of substation equipment can be achieved, improving the efficiency and accuracy of equipment operation and maintenance. In existing methods for constructing a knowledge graph of multi-source heterogeneous data of main substation equipment, only simple data cleaning, denoising, and other normalization operations can be performed based on multi-source heterogeneous data. In cases where the data integrity is poor, the constructed knowledge graph is not comprehensive enough. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for constructing a knowledge graph of multi-source heterogeneous data of main substation equipment and related systems to solve the problems raised in the above background technique.

[0004] To achieve the above purpose, the present invention provides the following technical solution: A method for constructing a knowledge graph of multi-source heterogeneous data of main substation equipment, the method includes the following steps:

[0005] S1. Multi-source heterogeneous data collection. For structured database data, a database connection tool is used to obtain the required data by writing SQL query statements. For semi-structured data, corresponding parsing libraries are used to parse and extract key information and convert it into a structured data format. For unstructured text data and web data, natural language processing technology and web crawler technology are used. The web crawler uses the Scrapy framework to collect data according to the set target website and data extraction rules. After collecting unstructured text data, it is stored in a local file or database for subsequent processing;

[0006] S2. Primary data processing. The primary data processing is divided into structured data processing, semi-structured data processing, and unstructured data processing;

[0007] Structured data processing: Perform data cleaning, remove duplicate records. By writing database query statements or using data processing tools, compare key fields to identify and delete duplicate record rows; correct incorrect formats. For inconsistent date formats, convert them to a standard format; for outliers in numerical data, such as temperature values significantly outside the normal range, screen and correct them by setting a reasonable threshold range.

[0008] Data integration: Integrate relevant structured data obtained from different database tables or systems. Combine the basic device information in the device management system with the operation data in the monitoring system through key fields such as device numbers to form a dataset containing multi-faceted device information.

[0009] Semi-structured data processing: Perform format conversion. For XML-format maintenance reports, parse the XML node content, extract key information, and convert it into a structured table form; for JSON-format device configuration files, parse the key-value pair information, extract device parameter configuration information, and convert it into a structured format for easy processing.

[0010] Data normalization: Normalize the extracted semi-structured data content.

[0011] Unstructured data processing: Text cleaning, remove punctuation marks, special characters, and meaningless character noise. Use Chinese word segmentation tools to segment the text into a sequence of words, and remove common stop words with no practical meaning.

[0012] Text annotation: For inspection records and fault description text data, perform manual annotation or use supervised machine learning algorithms for annotation to identify key information such as device entities and status descriptions.

[0013] S3. Secondary data processing: Introduce a data augmentation algorithm, set up a generator and a discriminator. In multi-source heterogeneous data processing, the generator is used to generate new data similar to the real data, and the discriminator is used to distinguish between real data and generated data. Through the adversarial game between the two, the generator continuously improves the quality of the generated data, thereby achieving data augmentation and supplementing data sources with less data volume.

[0014] S4. Perform knowledge extraction on the data obtained in S3, which includes entity extraction, relationship extraction, and attribute extraction. Entity extraction identifies entities by formulating rule templates; relationship extraction constructs a neural network model, using a convolutional neural network (CNN) or a recurrent neural network (RNN) and its variants to build a relationship extraction model. The input of the model is the sentence text vector representation containing two entities. Through the network layer, the model learns the semantic features of the sentence and outputs the result of the relationship type between entities. Attribute extraction is divided into extraction from structured data and extraction from unstructured data. When extracting from structured data, the inherent attribute information of the device is directly obtained from the database table fields. For the attribute information clearly marked in semi-structured data, it is extracted according to the format. When extracting from unstructured data, an attribute extraction rule template is defined. When the text conforms to the template rules, the corresponding attribute information is extracted. At the same time, natural language processing technology is combined for semantic analysis to improve the accuracy of attribute extraction;

[0015] S5. Perform knowledge fusion, which includes ontology construction, entity alignment, and conflict detection and resolution;

[0016] S6. Knowledge graph formation. Store the extracted and fused knowledge through a graph database. Store the attribute information of entities as nodes and store the attribute and direction information of relationships as edges to form a knowledge graph. The graph database includes Neo4j and JanusGraph;

[0017] S7. Optimization and highlighting processing. Assign different attention weights to nodes according to the features of the knowledge graph nodes and the information of neighboring nodes through an optimization algorithm. In the substation equipment knowledge graph, some key equipment nodes and nodes with important connection relationships have important impacts on the operation and fault propagation of the entire system. Identify these important nodes through the algorithm and assign higher attention weights;

[0018] S8. Application implementation. Based on the knowledge graph further optimized and highlighted in S7, implement equipment fault diagnosis, analysis, and operation and maintenance strategy optimization.

[0019] Preferably, the data enhancement algorithm in S3 is specifically as follows:

[0020] The generator G receives a random noise z and partial substation equipment data features as inputs to generate new data samples , and the discriminator D receives real data or generated data , and outputs a probability value indicating the probability that the input data is real data. The goal of the discriminator is to maximize the probability of correctly distinguishing real data and generated data, while the goal of the generator is to make the discriminator unable to distinguish generated data and real data;

[0021] The loss function of the discriminator and the loss functions of the generator are defined as follows respectively:

[0022]

[0023] where is the distribution of real data is the distribution of random noise. By alternately training the discriminator and the generator and continuously optimizing their parameters, the generated data can better simulate the distribution of real data, enhancing the diversity and integrity of multi-source heterogeneous data.

[0024] Preferably, for the ontology construction in S5: First, define the concept hierarchy, determine the core concepts of substation equipment, including equipment, components, operating status, and operations, and construct a concept hierarchy structure. Define clear semantics and connotations for each concept, and ensure the accuracy and consistency of the concepts through a combination of natural language description and formal definition; Define the relationship types, determine the relationship types between entities, including the composition relationship of equipment, the relationship between equipment and the environment, and the operating status relationship of equipment. Define names, semantics, and constraint conditions for each relationship type; Create an ontology model, use an ontology modeling tool to construct an ontology model of substation equipment, and present the defined concepts and relationships in a visual and formal way. In the model, classes, attributes, and instance elements can be defined, and the logical relationships between them can be established. The ontology model provides a unified framework and standard for subsequent knowledge fusion.

[0025] Preferably, for the entity alignment in S5, calculate the entity similarity through the following method:

[0026] For entity pairs that may represent the same entity from different data sources, use the edit distance algorithm to calculate the name similarity, and combine the attribute similarity to calculate the comprehensive similarity; Let the names of two entities and be and respectively, the edit distance be the name similarity be ,

[0027] For the attribute similarity, if the attribute value is numerical, use the Euclidean distance to calculate; if it is text-based, use the cosine similarity to calculate. The comprehensive similarity is the weighted sum of the name similarity and the attribute similarity; When the comprehensive similarity of the entity pair exceeds the set threshold, merge them into one entity.

[0028] Preferably, the conflict detection and resolution in S5 includes attribute value conflict detection and relationship conflict detection. Attribute value conflict detection: When integrating knowledge from different sources, it detects the situation where the attribute values of the same entity are inconsistent, and conducts conflict detection by comparing the attribute values, the reliability of the data sources, and the data collection time. Relationship conflict detection: It detects the situation where the relationships between entities are inconsistent, and conducts relationship conflict detection by analyzing the text semantics and the knowledge of related devices. For the conflicting data in multiple data sources, a voting method is used to determine the final value or relationship. When the majority of data sources consider that a certain attribute value of a device is X, then X is used as the final attribute value, and expert knowledge and domain rules are combined for auxiliary judgment to improve the accuracy of conflict resolution.

[0029] Preferably, the optimization algorithm in S7 is specifically as follows.

[0030] For each node in the knowledge graph , its feature vector representation is , and the set of neighbor nodes is . First, calculate the attention coefficient between node and neighbor node :

[0031]

[0032] Among them, is a learnable parameter vector, represents vector concatenation, is an activation function. Then, normalize the attention coefficient:

[0033]

[0034] Finally, the new feature vector of node is obtained by weighted summation of the neighbor node feature vectors:

[0035]

[0036] Among them, is an activation function. After multiple rounds of calculation, the role of important nodes in the knowledge graph is highlighted, and higher attention is assigned to the corresponding nodes in subsequent analysis and applications.

[0037] A system includes a storage medium and a processor. A computer program is provided in the storage medium. When the computer program is executed by the processor, the processor executes the steps of the above method.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] The knowledge graph construction method for multi-source heterogeneous data of main power transformation equipment in the present invention improves data utilization rate. By introducing a data enhancement algorithm in the secondary data processing, it can enhance and supplement the data, enhance the diversity and integrity of multi-source heterogeneous data, and make the construction of the knowledge graph more comprehensive and accurate.

[0040] In the knowledge graph construction method of the present invention, through optimized highlighting processing, different attention weights can be assigned to nodes according to the characteristics of the knowledge graph nodes and the information of neighbor nodes. Higher attention weights are given to prominent important nodes, so that in the use of the knowledge graph, important nodes can be more easily noticed, the work processing efficiency is increased, and the probability of overlooking important nodes is reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0043] Please refer to Figure 1 , the present invention provides a technical solution: a knowledge graph construction method for multi-source heterogeneous data of main power transformation equipment, and this method includes the following steps:

[0044] S1. Multi-source heterogeneous data collection. For structured database data, a database connection tool such as JDBC is used to obtain the required data by writing SQL query statements;

[0045] For semi-structured data such as monitoring data in XML and JSON formats, corresponding parsing libraries are used to parse and extract key information. XML can use the xml.etree.ElementTree library in Python and be converted into a structured data format; for unstructured text data and web data such as operation and maintenance records and documents, natural language processing technology and web crawler technology are adopted. The web crawler uses the Scrapy framework to collect data according to the set target website and data extraction rules. After the unstructured text data is collected, it is stored in a local file or database waiting for subsequent processing;

[0046] S2. Primary data processing. Primary data processing is divided into structured data processing, semi-structured data processing, and unstructured data processing;

[0047] Structured data processing: Perform data cleaning, remove duplicate records. By writing database query statements or using data processing tools, compare key fields to identify and delete duplicate record rows; correct incorrect formats. For inconsistent date formats, uniformly convert them to the standard format; for outliers in numerical data, such as temperature values significantly outside the normal range, screen and correct them by setting a reasonable threshold range.

[0048] Data integration, integrate relevant structured data obtained from different database tables or systems. Combine the basic device information in the device management system with the operation data in the monitoring system through key fields such as device numbers to form a data set containing multi-faceted information about the devices.

[0049] Semi-structured data processing: Perform format conversion. For XML format maintenance reports, parse the XML node content, extract key information, and convert it into a structured table form; for JSON format device configuration files, parse the key-value pair information, extract the device parameter configuration information, and convert it into a structured format convenient for processing.

[0050] Data normalization, perform normalization processing on the extracted semi-structured data content. For example, unify the expression methods of device component names, such as unifying "switch contact" and "switch contact head" to "switch contact".

[0051] Unstructured data processing: Text cleaning, remove punctuation marks, special characters, and meaningless character noises. Use Chinese word segmentation tools to segment the text into a sequence of words, and remove common stop words with no actual meaning.

[0052] Text annotation, for inspection records and fault description text data, perform manual annotation or use supervised machine learning algorithms for annotation to identify key information such as device entities and status descriptions.

[0053] S3. Secondary data processing, introduce data augmentation algorithms, set up a generator and a discriminator. In multi-source heterogeneous data processing, the generator is used to generate new data similar to the real data, and the discriminator is used to distinguish between real data and generated data. Through the adversarial game between the two, the generator continuously improves the quality of the generated data, thereby achieving data augmentation and supplementing data sources with less data volume. The data augmentation algorithm is as follows:

[0054] The generator G receives random noise z and some power transformation equipment data features as inputs to generate new data samples , the discriminator D receives real data or generated data , output a probability value representing the probability that the input data is real data. The goal of the discriminator is to maximize the probability of correctly distinguishing real data from generated data, while the goal of the generator is to make the discriminator unable to distinguish generated data from real data;

[0055] The loss function of the discriminator and the loss function of the generator are defined as follows:

[0056]

[0057] where is the distribution of real data, is the distribution of random noise. By alternately training the discriminator and the generator and continuously optimizing their parameters, the generated data can better simulate the distribution of real data, enhancing the diversity and integrity of multi-source heterogeneous data.

[0058] S4. Perform knowledge extraction on the data obtained in S3, which includes entity extraction, relationship extraction, and attribute extraction. Entity extraction identifies entities by formulating rule templates. For example, define the rule: If the text contains a pattern like "Device Model: Specific Model", then extract "Specific Model" as the entity. A series of rules are formulated for common entity types such as device names and manufacturers for extraction;

[0059] Relationship extraction constructs a neural network model, using a convolutional neural network (CNN) or a recurrent neural network (RNN) and its variants to construct a relationship extraction model. The input of the model is the sentence text vector representation containing two entities. Through the network layer, the semantic features of the sentence are learned, and the result of the relationship type between the entities is output;

[0060] Attribute extraction is divided into extraction from structured data and extraction from unstructured data. When extracting from structured data, the inherent attribute information of the device is directly obtained from the database table fields. For the attribute information clearly marked in semi-structured data, it is parsed and extracted according to the format. When extracting from unstructured data, an attribute extraction rule template is defined. When the text conforms to the template rules, the corresponding attribute information is extracted, and at the same time, natural language processing technology is combined for semantic analysis to improve the accuracy of attribute extraction;

[0061] S5. Perform knowledge fusion, which includes ontology construction, entity alignment, and conflict detection and resolution;

[0062] Ontology Construction: First, define the concept hierarchy, determine the core concepts of substation equipment, including equipment, components, operating status, and operations, and construct a concept hierarchy structure. Define clear semantics and connotations for each concept. Ensure the accuracy and consistency of the concepts through a combination of natural language description and formal definition. Define the relationship types, determine the relationship types between entities, including the composition relationship of equipment, the relationship between equipment and the environment, and the operating status relationship of equipment. Define the name, semantics, and constraint conditions for each relationship type. Create an ontology model, use an ontology modeling tool to construct an ontology model of substation equipment, and present the defined concepts and relationships in a visual and formal way. In the model, classes, properties, and instance elements can be defined, and the logical relationships between them can be established. Provide a unified framework and standard for subsequent knowledge fusion through the ontology model.

[0063] Entity Alignment: Calculate the entity similarity through the following methods:

[0064] For entity pairs that may represent the same entity from different data sources, use the edit distance algorithm to calculate the name similarity, and combine the attribute similarity to calculate the comprehensive similarity. Suppose two entities and have names and respectively, the edit distance is , the name similarity is . For the attribute similarity, if the attribute value is numerical, use the Euclidean distance to calculate; if it is text-based, use the cosine similarity to calculate. The comprehensive similarity is the weighted sum of the name similarity and the attribute similarity. When the comprehensive similarity of the entity pair exceeds the set threshold, merge them into one entity.

[0065] Conflict Detection and Resolution, including attribute value conflict detection and relationship conflict detection. Attribute value conflict detection: When integrating knowledge from different sources, detect the situation where the attribute values of the same entity are inconsistent. Conduct conflict detection by comparing the attribute values, the reliability of the data sources, and the data collection time. Relationship conflict detection: Detect the situation where the relationships between entities are inconsistent. For example, in one data source, equipment A and equipment B are in a parallel relationship, while in another data source, they are in a series relationship. Conduct relationship conflict detection by analyzing the text semantics and relevant equipment knowledge. For the conflicting data in multiple data sources, use the voting method to determine the final value or relationship. When the majority of data sources consider that a certain attribute value of a device is X, then use X as the final attribute value, and combine expert knowledge and domain rules for auxiliary judgment to improve the accuracy of conflict resolution.

[0066] S6. Knowledge graph formation: Store the extracted and fused knowledge through a graph database. Store the entity's attribute information as nodes and the relationship's attribute and direction information as edges to form a knowledge graph. The graph database includes Neo4j and JanusGraph. Neo4j is characterized by being simple to use and having high query performance, suitable for storing and querying knowledge graphs of small and medium scales; JanusGraph has distributed storage capabilities and good scalability, suitable for large-scale knowledge graph applications. Select a suitable graph database according to factors such as substation equipment data volume and application requirements.

[0067] S7. Optimization and highlighting processing: Use an optimization algorithm to assign different attention weights to nodes based on the characteristics of the knowledge graph nodes and the information of neighboring nodes. In the substation equipment knowledge graph, some key equipment nodes and nodes with important connection relationships have important impacts on the operation and fault propagation of the entire system. Identify these important nodes through the algorithm and assign them higher attention weights. The optimization algorithm is as follows:

[0068] For each node in the knowledge graph , its feature vector is represented as , and the set of neighboring nodes is . First, calculate the attention coefficient between node and neighboring node :

[0069]

[0070] Among them, is a learnable parameter vector, represents vector concatenation, is an activation function. Then, normalize the attention coefficient:

[0071]

[0072] Finally, the new feature vector of node is obtained by weighted summation of the neighboring node feature vectors:

[0073]

[0074] Among them, is an activation function. After multiple loop calculations, highlight the role of important nodes in the knowledge graph and assign higher attention to the corresponding nodes in subsequent analysis and applications.

[0075] S8. Application implementation: Implement equipment fault diagnosis, analysis, and operation and maintenance strategy optimization based on the knowledge graph further optimized and highlighted in S7.

[0076] A system includes a storage medium and a processor. A computer program is provided in the storage medium. When the computer program is executed by the processor, the processor is caused to execute the steps of the above method.

[0077] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a knowledge graph of multi-source heterogeneous data of main substation equipment, characterized in that, The method includes the following steps: S1. Multi-source heterogeneous data collection: For structured database data, a database connection tool is used to obtain the required data by writing SQL query statements; for semi-structured data, a corresponding parsing library is used for parsing to extract key information and convert it into a structured data format; for unstructured text data and web data, natural language processing technology and web crawler technology are adopted. The web crawler uses the Scrapy framework to collect data according to the set target website and data extraction rules. After the unstructured text data is collected, it is stored in a local file or database waiting for subsequent processing; S2. Primary data processing, which is divided into structured data processing, semi-structured data processing, and unstructured data processing; Structured data processing: Data cleaning is carried out to remove duplicate records. By writing database query statements or using data processing tools, key fields are compared to identify and delete duplicate record rows; incorrect formats are corrected. For the case of inconsistent date formats, they are uniformly converted into a standard format; for outliers in numerical data, such as temperature values that are significantly outside the normal range, reasonable threshold ranges are set for screening and correction; Data integration: The relevant structured data obtained from different database tables or systems is integrated. The basic device information in the device management system and the operation data in the monitoring system are associated and merged through key fields such as device numbers to form a data set containing multi-faceted information about the device; Semi-structured data processing: Format conversion is carried out. For the XML format maintenance report, the XML node content is parsed to extract key information and convert it into a structured table form; for the JSON format device configuration file, the key-value pair information is parsed to extract the device parameter configuration information and convert it into a structured format convenient for processing; Data standardization: The extracted semi-structured data content is subjected to standardization processing; Unstructured data processing: Text cleaning is carried out to remove punctuation marks, special characters, and meaningless character noises. A Chinese word segmentation tool is used to segment the text, splitting the continuous text into a sequence of words, and removing common stop words without practical significance; Text annotation: For inspection records and fault description text data, manual annotation or supervised machine learning algorithms are used for annotation to identify key information such as device entities and status descriptions; S3. Secondary data processing: A data augmentation algorithm is introduced, and a generator and a discriminator are set. In multi-source heterogeneous data processing, the generator is used to generate new data similar to the real data, and the discriminator is used to distinguish between real data and generated data. Through the adversarial game between the two, the generator continuously improves the quality of the generated data, thereby achieving data augmentation and supplementing data sources with less data volume; S4. Perform knowledge extraction on the data obtained in S3, which includes entity extraction, relationship extraction, and attribute extraction. Entity extraction identifies entities by formulating rule templates; relationship extraction constructs a neural network model, and uses a convolutional neural network (CNN) or a recurrent neural network (RNN) and its variants to construct a relationship extraction model. The input of the model is the sentence text vector representation containing two entities. After learning the sentence semantic features through the network layer, the relationship type result between entities is output; attribute extraction is divided into extraction from structured data and extraction from unstructured data. When extracting from structured data, the inherent attribute information of the device is directly obtained from the database table fields. For the attribute information clearly marked in semi-structured data, it is parsed and extracted according to the format. When extracting from unstructured data, an attribute extraction rule template is defined. When the text conforms to the template rules, the corresponding attribute information is extracted, and natural language processing technology is combined for semantic analysis to improve the accuracy of attribute extraction; S5. Perform knowledge fusion, where the knowledge fusion includes ontology construction, entity alignment, and conflict detection and resolution; S6. Knowledge graph formation. The extracted and fused knowledge is stored in a graph database. The entities are stored as nodes with their attribute information, and the relationships are stored as edges with the attribute and direction information of the relationships, forming a knowledge graph. The graph databases include Neo4j and JanusGraph; S7. Optimization and highlighting processing. Different attention weights are assigned to nodes through an optimization algorithm according to the features of the knowledge graph nodes and the information of neighboring nodes. In the substation equipment knowledge graph, some key equipment nodes and nodes with important connection relationships have important impacts on the operation and fault propagation of the entire system. These important nodes are identified by the algorithm and given higher attention weights; S8. Application implementation. Based on the knowledge graph further optimized and highlighted in S7, device fault diagnosis, analysis, and operation and maintenance strategy optimization are realized; The data enhancement algorithm described in S3 is specifically as follows: The generator G receives random noise z and partial power transformation equipment data features as inputs and generates new data samples . The discriminator D receives real data or generated data and outputs a probability value indicating the probability that the input data is real data. The goal of the discriminator is to maximize the probability of correctly distinguishing real data from generated data, while the goal of the generator is to make the discriminator unable to distinguish between generated data and real data; The loss function of the discriminator and the loss function of the generator are defined as follows respectively: ; Among them, is the distribution of real data, is the distribution of random noise. By alternately training the discriminator and the generator and continuously optimizing their parameters, the generated data can better simulate the distribution of real data, enhancing the diversity and integrity of multi-source heterogeneous data.

2. The method for constructing a knowledge graph of multi-source heterogeneous data of main power transformation equipment according to claim 1, wherein: The ontology construction described in S5: First, define the concept hierarchy, determine the core concepts of substation equipment, including equipment, components, operating status, and operations, and construct a concept hierarchy structure. Clear semantics and connotations are defined for each concept. Through a combination of natural language description and formal definition, the accuracy and consistency of the concepts are ensured; define the relationship types, determine the relationship types between entities, including the composition relationship of equipment, the relationship between equipment and the environment, and the operating status relationship of equipment. Names, semantics, and constraint conditions are defined for each relationship type; Create an ontology model. Use an ontology modeling tool to construct a substation equipment ontology model, and present the defined concepts and relationships in a visual and formal way. Classes, attributes, and instance elements can be defined in the model, and logical relationships between them are established. The ontology model provides a unified framework and standard for subsequent knowledge fusion.

3. The method for constructing a knowledge graph of multi-source heterogeneous data of main power transformation equipment according to claim 1, characterized in that: The entity alignment described in S5 calculates the entity similarity through the following methods: For entity pairs that may represent the same entity from different data sources, the edit distance algorithm is used to calculate the name similarity, and the comprehensive similarity is calculated by combining the attribute similarity; assume two entities and whose names are and respectively, the edit distance is , the name similarity is . For the attribute similarity, if the attribute value is numerical, the Euclidean distance is used for calculation; if it is text-based, the cosine similarity is used for calculation. The comprehensive similarity is the weighted sum of the name similarity and the attribute similarity; when the comprehensive similarity of the entity pair exceeds the set threshold, they are merged into one entity.

4. The method for constructing a knowledge graph of multi-source heterogeneous data of main power transformation equipment according to claim 1, wherein: The conflict detection and resolution described in S5 includes attribute value conflict detection and relationship conflict detection. Attribute value conflict detection: When integrating knowledge from different sources, it detects the situation where the attribute values of the same entity are inconsistent, and conducts conflict detection by comparing attribute values, the reliability of data sources, and data collection times; Relationship conflict detection: It detects the situation where the relationships between entities are inconsistent, and conducts relationship conflict detection by analyzing text semantics and relevant device knowledge; For the conflicting data in multiple data sources, a voting method is used to determine the final value or relationship. When the majority of data sources consider that a certain attribute value of a device is X, then X is used as the final attribute value, and at the same time, expert knowledge and domain rules are combined for auxiliary judgment to improve the accuracy of conflict resolution.

5. The method for constructing a knowledge graph of multi-source heterogeneous data of main power transformation equipment according to claim 1, characterized in that: The optimization algorithm described in S7 is specifically as follows. For each node in the knowledge graph , Its eigenvector is expressed as , and the set of neighbor nodes is . First, calculate the attention coefficient between node and its neighbor node : ; Among them, is a learnable parameter vector, represents vector concatenation, is an activation function, and then the attention coefficients are normalized: ; Finally, the node 's new feature vector is obtained by weighted summation of the feature vectors of its neighbor nodes: ; Among them, is an activation function. After multiple loop calculations, it highlights the role of important nodes in the knowledge graph and assigns higher attention to the corresponding nodes in subsequent analysis and applications.

6. A knowledge graph construction system for multi-source heterogeneous data of main substation equipment, characterized in that, It includes a storage medium and a processor. A computer program is provided in the storage medium. When the computer program is executed by the processor, the processor executes the steps of the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Asset management knowledge base construction method based on automatic data generation

    CN112287117A

  • Automatic knowledge graph construction method based on multi-source heterogeneous power equipment data

    CN112860908A

  • Construction method for multi-source heterogeneous data knowledge graph of power distribution network

    CN117273133A