Entity similarity calculation method and system for electric power safety knowledge graph

By subdividing the numerical, list-type and text-type attribute calculation methods of power equipment entities, combined with multiple similarity calculation methods, the problem of inaccurate entity similarity calculation in power equipment status evaluation is solved, and more efficient and accurate equipment status evaluation is achieved.

CN120336866APending Publication Date: 2025-07-18SHENZHEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510270883.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-07-18

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_3
    Figure QLYQS_3
  • Figure QLYQS_8
    Figure QLYQS_8
Patent Text Reader

Abstract

The embodiment of the invention discloses an entity similarity calculation method and system for an electric power security knowledge graph. The method comprises the following steps: a numerical attribute calculation step: calculating the distance between numerical vectors of entities by adopting an Euclidean distance or a Manhattan distance; in the list type attribute calculation step, the similarity between the entities is obtained by calculating the number of attribute value intersections of the entities or through Jaccard similarity calculation, and the similarity is output; and a text type attribute calculation step: calculating the similarity of the text type attributes of the entities and outputting the similarity. According to the method, the similarity calculation of the entities is more accurate, so that the efficiency and the precision of state evaluation of the power equipment can be remarkably improved, and more intelligent operation and maintenance decisions are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of smart grids, and particularly to a method and system for calculating entity similarity of a power safety knowledge graph. Background Art

[0002] The state assessment of power equipment (such as generators, transformers, relay protection devices, etc.) (such as the remaining life of transformers, transmission line condition monitoring, etc.) has important theoretical significance for ensuring the safe and stable operation of the power system. With the continuous improvement of the smart grid system, the types, quantities, and operating condition complexities of power equipment have increased significantly. How to effectively manage and utilize this knowledge is one of the important contents of the refined operation and health management of power equipment. The knowledge graph formally describes core elements such as "entities", "attributes", and "relationships" through triples, and effectively organizes large-scale network information at the lowest cost, realizing the visual storage, organization, and management of hundreds of millions of entities, attributes, and association relationships in the Internet field.

[0003] The knowledge graph better strings and presents knowledge through a network topology structure. It can integrate and extract knowledge from multi-source heterogeneous data, contains richer semantic associations between entities, and can provide personalized services by combining implicit information obtained through reasoning. The knowledge graph can integrate multi-source heterogeneous data from sensors, historical records, and expert experience, etc., to form a unified knowledge base, which is convenient for comprehensively analyzing the equipment state, quickly identifying the types and causes of faults, improving the diagnosis efficiency, and reducing the maintenance cost.

[0004] The knowledge graph has improved the intelligent level of equipment management and enhanced the reliability and security of the power system by integrating multi-source data, providing intelligent analysis, and decision support in the state assessment of power equipment. For a knowledge base containing numerous entities, in addition to paying attention to the information of the entities themselves, it is also necessary to pay attention to the association information between entities. One of the problems faced is how to determine whether two entities are similar and calculate the degree of similarity.

[0005] The similarity between entities refers to the similarity at the deep semantic level, rather than the traditional similarity that only focuses on surface information. For a circuit breaker and a disconnector, although both names imply disconnection, the circuit breaker is used to cut off fault currents, while the disconnector is used to isolate the power supply and cannot cut off currents; a 500 kV hub substation and a 220 kV terminal substation, although they seemingly have little relationship, both belong to the substation entity; a high-voltage switchgear and a low-voltage switchgear, although their names are quite different, have similar functions and are both devices used to control and protect the power system. Therefore, to determine the similarity between entities, it is necessary to first understand the semantic information of the entities, and the traditional character similarity method is definitely not feasible.

[0006] Existing research on entity similarity is not comprehensive enough, without sufficient consideration of entity attributes and the characteristics of the entities themselves, resulting in low accuracy of similarity calculation, and further leading to poor efficiency and accuracy in power equipment status assessment. Summary of the Invention

[0007] The technical problem to be solved by the embodiments of the present invention is to provide a method and system for calculating entity similarity of a power safety knowledge graph to improve the efficiency and accuracy of power equipment status assessment.

[0008] To solve the above technical problem, the embodiments of the present invention propose a method for calculating entity similarity of a power safety knowledge graph, including: Numeric attribute calculation step: Obtain the values corresponding to the numeric attributes of each entity in the power safety knowledge entity data, calculate the distance between the numeric vectors of each entity using Euclidean distance or Manhattan distance, and output the corresponding similarity value according to the calculated distance; List attribute calculation step: Obtain the attribute values corresponding to the list attributes of each entity in the power safety knowledge entity data, calculate the similarity between each entity by calculating the number of intersections of the attribute values of each entity and output, or calculate the similarity between each entity using Jaccard similarity and output; Text attribute calculation step: Obtain the text data corresponding to the text attributes of each entity in the power safety knowledge entity data, and calculate and output the similarity of the text attributes of each entity.

[0009] Correspondingly, the embodiments of the present invention also provide a system for calculating entity similarity of a power safety knowledge graph, including: Numeric attribute calculation module: Obtain the values corresponding to the numeric attributes of each entity in the power safety knowledge entity data, calculate the distance between the numeric vectors of each entity using Euclidean distance or Manhattan distance, and output the corresponding similarity value according to the calculated distance; List attribute calculation module: Obtain the attribute values corresponding to the list attributes of each entity in the power safety knowledge entity data, calculate the similarity between each entity by calculating the number of intersections of the attribute values of each entity and output, or calculate the similarity between each entity using Jaccard similarity and output; Text attribute calculation module: Obtain the text data corresponding to the text attributes of each entity in the power safety knowledge entity data, and calculate and output the similarity of the text attributes of each entity.

[0010] The beneficial effects of the present invention are as follows: The present invention subdivides the similarity calculation of entities into numerical attributes, text attributes, and list attributes, and considers the structural properties, which can more comprehensively reflect the similarity between entities. This subdivision helps to improve the accuracy and pertinence of similarity calculation. For the similarity calculation of numerical attributes, the similarity is directly calculated through mathematical formulas such as Euclidean distance and Manhattan distance, so as to obtain accurate similarity values, and the processing is relatively simple without complex preprocessing steps. For the similarity calculation of list attributes, calculating the intersection, union, etc. of these values can better reflect the diversity and overlap between entities; through Jaccard similarity calculation, it is possible to flexibly process attribute value sets of different lengths. For the similarity calculation of text attributes, the TF-IDF method can effectively quantify the similarity between text entity attributes. By calculating the TF-IDF values of each word, the keywords in the text can be highlighted, and the similarity between two texts can be calculated through cosine similarity; calculating the similarity of text attribute entities using the TransE model has multiple advantages. The edge-based loss function is easy to implement and optimize. TransE maps entities and relationships to a low-dimensional vector space, which helps to reduce the demand for computing resources and can effectively capture the relationships between entities; at the same time, as new entities and relationships are added, the TransE model can update the vector representations of entities and relationships through incremental training without retraining the entire model; TransE can be used for entity linking tasks, and by calculating the similarity between unknown entities and known entities in the knowledge graph, the identity of the unknown entities can be inferred. Finally, the similarity calculation of structural properties considers the relationship structure between entities, such as the position of entities in the network, connection patterns, etc., which is very important for understanding the complex relationships between entities, more comprehensively considers the deep semantic information of entities, makes the similarity calculation of entities more accurate, and thus can significantly improve the efficiency and accuracy of power equipment status assessment and support more intelligent operation and maintenance decisions. Detailed implementation manners

[0011] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The present invention will be further described in detail below with reference to specific embodiments.

[0012] In a knowledge base where data is structured and stored, the attributes of entities can be used as the basis for similarity judgment, and at the same time, the entities themselves can calculate similarities through different methods, so as to calculate the similarity of entities by combining these two aspects (attributes and entities themselves). The entity similarity calculation method of the power safety knowledge graph in the embodiment of the present invention includes a numerical attribute calculation step, a list attribute calculation step, and a text attribute calculation step.

[0013] The power grid knowledge graph should consist of entity, attribute, and relationship triples, defined as a set of entity, attribute, and relationship triples: . In the formula, E refers to entities. Entities in the power safety knowledge graph refer to objectively existing and distinguishable power grid equipment, personnel, event records, etc.; P refers to attributes. Attributes in the power safety knowledge graph refer to different characteristic dimensions for describing entities from multiple perspectives, such as the name and voltage level of a substation; R refers to relationships. Relationships in the power safety knowledge graph refer to the association methods between entities and entities, between entities and entity types, and between entity types, including classification relationships, composition relationships, attribute relationships, etc. Common entity types in the power grid knowledge field include: equipment entities, personnel entities, environmental entities, operation entities, accident entities, etc. Their corresponding attributes are equipment model, rated capacity, operating status, length of service, qualification certificates, safety training records, illegal operation records, temperature, humidity, wind speed, geographical information, environmental pollution level, operation type, operation time, operation risk level, accident date, accident location, accident cause, casualty situation, economic loss, handling measures, accident level, etc. It can be found that entity attributes are diverse, including numerical, text, and list types. Different similarity calculation methods are used for different types of data to measure the similarity between them. The present invention calculates the similarity of entity attributes based on these three situations, and methods such as edit distance, set similarity (Jaccard coefficient, Dice), vector similarity (cosine similarity, Euclidean distance), etc. are adopted.

[0014] Calculation steps for numerical attributes: Obtain the numerical values corresponding to the numerical attributes of each entity in the power safety knowledge entity data, calculate the distance between the numerical vectors of each entity using Euclidean distance or Manhattan distance, and obtain and output the similarity values between the numerical attributes of each entity based on the calculated distance.

[0015] Numerical attributes such as the rated capacity of equipment, the working years of personnel, temperature and humidity attributes, etc. usually use Euclidean distance to calculate the distance between two numerical vectors. The smaller the distance, the more similar they are, as shown in formula (1).

[0016] (1) Among them, , represents the numerical value of the numerical attribute to be calculated, 、 represent the maximum and minimum values of all numerical attributes to be calculated, m is the total number of numerical attributes to be calculated, and k, j ∈ m; It can be seen from this formula that the value range of D is [0,1], and and The greater the difference between them, the greater the D value, indicating the smaller the similarity between them. In high-dimensional space, the Manhattan distance is usually used to calculate the distance between values. The cosine similarity is used to measure the similarity degree of two vectors in terms of direction. For the similarity calculation of numerical attributes, the similarity is directly calculated through the above mathematical formula to obtain an accurate similarity value, and the processing is relatively simple without complex preprocessing steps.

[0017] The present invention subdivides the similarity calculation of entities into numerical attributes, text attributes, and list attributes, and considers the structural properties, which can more comprehensively reflect the similarity between entities. This subdivision helps to improve the accuracy and pertinence of the similarity calculation. For the similarity calculation of numerical attributes, the present invention directly calculates the similarity through mathematical formulas such as the Euclidean distance and the Manhattan distance, so as to obtain an accurate similarity value, and the processing is relatively simple without complex preprocessing steps.

[0018] Calculation steps for list attributes: Obtain the attribute values corresponding to the list attributes of each entity in the power safety knowledge entity data, calculate the similarity between the list attributes of each entity by calculating the number of intersections of the attribute values of each entity, or calculate the similarity between the list attributes of each entity through the Jaccard similarity, and output. For the similarity calculation of list attributes in the present invention, by calculating the intersections, unions, etc. of these values, the diversity and overlap between entities can be better reflected; through the Jaccard similarity calculation, attribute value sets of different lengths can be flexibly processed.

[0019] List attributes indicate that their attribute values are one or more elements in a certain set. For example, there are 500 kV substations, 220 kV substations, 110 / 35 kV substations, etc. for substations, and the main technical parameters of distribution transformers include rated voltage, model, capacity, impedance parameters, etc. List data can be processed as a set. For the attribute values of list attributes, two metrics are used to measure their similarity. One is to calculate the number of intersections. The larger the number of intersections, the more similar they are; the other is the Jaccard similarity, and the calculation formula is shown in Equation (2).

[0020] (2) For two sets C and F, the value range of the Jaccard similarity is [0, 1]. The larger the value, the higher the similarity between them. For the similarity calculation of list attributes, by calculating the intersections, unions, etc. of these values, the diversity and overlap between entities can be better reflected; and through the Jaccard similarity calculation, attribute value sets of different lengths can be flexibly processed.

[0021] Textual attribute calculation steps: Obtain the textual data corresponding to the textual attributes of each entity in the power safety knowledge entity data, calculate the similarity between the textual attributes of each entity, and output the result.

[0022] The value of the textual attribute is a piece of text information, such as text information like fault descriptions and safety operation steps. The textual data contains a lot of potential corpus information, and comparing them will largely reflect the similarity between two entities. The present invention uses the TF-IDF method to effectively quantify the similarity between textual entity attributes. By calculating the TF-IDF value of each word, the keywords in the text can be highlighted, and the similarity between two texts can be calculated through cosine similarity; in addition, using the TransE model to calculate the similarity of textual attribute entities has many advantages. The margin-based loss function is easy to implement and optimize. TransE maps entities and relationships to a low-dimensional vector space, which helps reduce the demand for computing resources and can effectively capture the relationships between entities; at the same time, as new entities and relationships are added, the TransE model can update the vector representations of entities and relationships through incremental training without having to retrain the entire model; TransE can be used for entity linking tasks, by calculating the similarity between unknown entities and known entities in the knowledge graph to infer the identity of the unknown entities. The following is a detailed introduction to the two methods.

[0023] As an implementation manner, the present invention uses the TF-IDF method based on the vector space model to calculate the similarity of textual attributes. TF-IDF (Term Frequency-Inverse Document Frequency) is a commonly used text analysis method and a statistical method for evaluating the importance of a word in a document or a set of documents. The TF-IDF method has three important formulas.

[0024] (1) TF (Term Frequency) - Word frequency - represents the frequency of a word appearing in a document. It is usually defined as the number of times the word appears in the document divided by the total number of words in the document. The formula is shown in Equation (3): (3) (2) IDF (Inverse Document Frequency) - Inverse document frequency - represents the rarity of a word in the entire document collection. If a word appears in many documents, its IDF is lower; if it appears rarely, its IDF is higher. The calculation formula is shown in Equation (4): (4) (3) TF-IDF value - The TF-IDF value is the product of TF and IDF and is used to measure the importance of a word to a document, as shown in Equation (5): (5) The steps for calculating the similarity of text-based attributes by the TF-IDF method include preprocessing, constructing a bag-of-words model, calculating TF-IDF weights, and calculating similarity. 1) Preprocessing: Preprocess the text, including steps such as word segmentation, stop-word removal, stemming, or lemmatization, to reduce noise and improve the accuracy of similarity calculation. 2) Constructing a bag-of-words model: Represent the text attributes of each entity as a bag-of-words model, that is, regard each word in the text as a feature, and each document is represented as a vector, where each dimension of the vector corresponds to the occurrence frequency of a word. 3) Calculating TF-IDF weights: Calculate the TF-IDF value for each word to represent the importance of this word in the document. Usually, a TF-IDF matrix is used to represent a set of documents. 4) Calculating similarity: Use methods such as cosine similarity to calculate the similarity between the attribute vectors of two entities, and use to represent the value of cosine similarity. For two n-dimensional TF-IDF vectors A and B, their cosine similarity CosSim can be calculated by Equation (6).

[0025] (6) It can be seen that the value range of cosine similarity is [0, 1], and the larger the value, the higher the similarity.

[0026] As an implementation method, in the scenario of power safety knowledge, the similarity calculation of entities mainly includes the similarity calculation of equipment, faults, maintenance processes, safety measures, etc. in the power system. Different entities use different algorithms to calculate the similarity between them. For example, methods based on word vectors (Word2Vec, FastText) can calculate the similarity between entities such as equipment and fault types. Methods based on graph neural networks (GNNs) model the topological structure of the power system to infer the similarity between entities. Rule-based matching methods calculate similarity, and rules are formulated according to expert knowledge to determine the similarity between entities. For example, if two devices belong to the same type of transformer, then the similarity between them is relatively high. There are also aggregation methods (obtaining entity similarity by weighted averaging of attribute similarity vectors), clustering methods (clustering entities into clusters and calculating the similarity between cluster entities instead of calculating the similarity between pairwise entities), and methods of representation learning to calculate similarity. For example, TransX series models - TransE, TransH, TransR, etc. can use the TransE model to map entities and relationships in the knowledge graph to low-dimensional dense space vectors to calculate entity similarity, which is used to represent the relationship between devices and devices, and between devices and fault types in the power system.

[0027] The TransE model is a knowledge graph embedding method that maps entities and relations into a low-dimensional vector space and uses vector operations to model the relations between entities. In the TransE model, both entities and relations are represented as vectors, and the L1 or L2 norm (the L1 norm and L2 norm are the Manhattan norm and Euclidean norm respectively) is usually used to measure the distance between vectors. The basic idea is that if there is a relation connecting entity and entity , that is , then the vector of entity plus the vector of relation should be close to the vector of entity . Mathematically, it is expressed as: . The TransE model mainly consists of two aspects: (1) Initialization, randomly initialize the vector representations and for each entity e and relation r. (2) Loss function, define a loss function to optimize these vector representations, and the commonly used loss function is the margin-based maximum likelihood estimation. The steps to calculate entity similarity using the TransE model are as follows: (1) Organize the data and train the TransE model, collect the triple data of entities and relations from the knowledge graph, and use the maximum likelihood estimation loss function to train the TransE model to optimize the vector representations of entities and relations.

[0028] (2) Calculate entity vectors. For each entity e, obtain the corresponding vector representation using the trained TransE model.

[0029] (3) Calculate similarity: Use the L1 or L2 norm to calculate the distance between entity vectors. The smaller the distance, the higher the similarity.

[0030] The present invention can effectively quantify the similarity between text-based entity attributes using the TF-IDF method. By calculating the TF-IDF value of each word, the keywords in the text can be highlighted, and the similarity between two texts can be calculated through cosine similarity. Calculating the similarity of text-based attribute entities using the TransE model has multiple advantages. The margin-based loss function is easy to implement and optimize. TransE maps entities and relationships to a low-dimensional vector space, which helps reduce the demand for computing resources and can effectively capture the relationships between entities. At the same time, as new entities and relationships are added, the TransE model can update the vector representations of entities and relationships through incremental training without retraining the entire model. Especially when the data volume is continuously growing, efficient similarity calculation can be achieved through incremental learning. TransE can be used for entity linking tasks by calculating the similarity between unknown entities and known entities in the knowledge graph to infer the identity of the unknown entities.

[0031] As an implementation manner, the present invention uses a path similarity algorithm and a graph edit distance algorithm to calculate structural similarity. The entity attribute similarity in the power safety knowledge graph also includes structural similarity, such as relationship type similarity, whether the relationship types between entities are the same, both being "subordinate to" or "maintain", etc.; such as relationship path similarity, the similarity of other entities reached by entities through different relationship paths. Structural similarity not only considers the attributes of the entities themselves but also the relationships between entities and the layout of these relationships in the entire knowledge graph. It usually includes the following aspects: (1) Relationship type: Whether the relationship types between entities are the same, such as "belong to", "connect", "maintain", etc. For example, a substation includes a transformer, a busbar, a circuit breaker, etc.; (2) Relationship path: The similarity of other entities reached by entities through different relationship paths. For example, substations are connected by lines; (3) Neighbor nodes: Whether the direct neighbor nodes (i.e., directly connected entities) around the entity are similar. For example, a main transformer consists of a winding diagram, an iron core, a main transformer tank, transformer oil, a voltage regulating device, etc. The neighbor nodes of the iron core are the winding diagram, the main transformer tank, transformer oil, and the voltage regulating device; (3) Network position: The position or role of the entity in the network, whether it is on a critical path, etc. For example, transmission equipment transports the electric energy of a power plant to a substation through an overhead line, and the overhead line is on the path from the power plant to the substation; There are various methods for calculating structural similarity. This invention generally adopts the path similarity algorithm and the graph edit distance algorithm: (1) Path similarity: Calculate the length of the shortest path between two entities. The shorter the path, the higher the similarity; or count how many direct neighbor nodes two entities have in common. The more the number, the higher the similarity. (2) Graph edit distance: Measure the similarity between two entities by calculating the minimum number of edit operations (such as adding, deleting nodes or edges) required to transform one graph into another. The similarity calculation based on structural properties takes into account the relationship structure between entities, such as the position of entities in the network, connection patterns, etc. This is very important for understanding the complex relationships between entities, thus more comprehensively considering the deep semantic information of entities and making the similarity calculation of entities more accurate.

[0032] The similarity calculation based on the structural properties of this invention takes into account the relationship structure between entities, such as the position of entities in the network, connection patterns, etc. This is very important for understanding the complex relationships between entities, more comprehensively considering the deep semantic information of entities and making the similarity calculation of entities more accurate.

[0033] The entity similarity calculation system of the power safety knowledge graph in the embodiment of this invention includes a numerical attribute calculation module, a list attribute calculation module, a text attribute calculation module, and a structural similarity calculation module.

[0034] Numerical attribute calculation module: Obtain the numerical values corresponding to the numerical attributes of each entity in the power safety knowledge entity data, calculate the distance between the numerical vectors of two entities using the Euclidean distance or Manhattan distance, and obtain and output the similarity values between the numerical attributes of each entity according to the calculated distance.

[0035] List attribute calculation module: Obtain the attribute values corresponding to the list attributes of each entity in the power safety knowledge entity data, obtain the similarity between the list attributes of each entity by calculating the number of intersections of the attribute values of two entities, or calculate and output the similarity between the list attributes of each entity using the Jaccard similarity.

[0036] Text attribute calculation module: Obtain the text data corresponding to the text attributes of each entity in the power safety knowledge entity data, calculate and output the similarity between the text attributes of each entity.

[0037] As an implementation method, the text attribute calculation module calculates the similarity according to the following steps: 1) Preprocessing: Preprocess the text data; 2) Construct a bag-of-words model: Represent the text data of each entity as a bag-of-words model, that is, regard each word in the text data as a feature, and represent each document as a vector. Each dimension of the vector corresponds to the occurrence frequency of a word; 3) Calculate TF-IDF weights: Calculate the TF-IDF value for each word to represent the importance of this word in the document; 4) Calculate similarity: Use cosine similarity to calculate the similarity between the text attributes of two entities. The calculation formula is as follows: ; where A and B are respectively the n-dimensional TF-IDF vector sets of the two entities to be compared, i ∈ n, and A i , B i are respectively the TF-IDF vectors of the i-th dimension of the two entities to be compared.

[0038] As an implementation, the text attribute calculation module calculates similarity according to the following steps: (1) Construct a TransE model, collect triple data of entities and relationships, and use the maximum likelihood estimation loss function to train the TransE model to optimize the vector representations of entities and relationships; (2) Calculate entity vectors. For each entity e, obtain the corresponding vector representation using the trained TransE model ; (3) Calculate similarity: Use the L1 or L2 norm to calculate the distance between entity vectors.

[0039] As an implementation, the numerical attribute calculation module performs Euclidean distance calculation according to the following formula: ; where, , represents the value of the numerical attribute to be calculated, , represent the maximum and minimum values of all numerical attributes to be calculated, m is the total number of numerical attributes to be calculated, and k, j ∈ m; The list attribute calculation module calculates the similarity between two entities according to the following formula: ; where C and F are the sets of list attribute data of the two entities to be compared.

[0040] As an implementation, the structural similarity calculation module uses the path similarity algorithm and the graph edit distance algorithm to calculate structural similarity.

[0041] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalent scope.

Claims

1. A method for calculating entity similarity of an electric power safety knowledge graph, characterized in that Including: Numerical attribute calculation steps: Obtain the numerical values corresponding to the numerical attributes of each entity in the power safety knowledge entity data, calculate the distance between the numerical vectors of each entity using Euclidean distance or Manhattan distance, and obtain and output the corresponding similarity value according to the calculated distance; List attribute calculation steps: Obtain the attribute values corresponding to the list attributes of each entity in the power safety knowledge entity data, calculate and output the similarity between each entity by calculating the number of intersections of the attribute values of each entity, or calculate and output the similarity between each entity by using Jaccard similarity; Text attribute calculation steps: Obtain the text data corresponding to the text attributes of each entity in the power safety knowledge entity data, calculate and output the similarity of the text attributes of each entity.

2. The method for calculating the entity similarity of the power safety knowledge graph according to claim 1, wherein In the text attribute calculation steps, calculate the similarity according to the following steps: 1) Preprocessing: Preprocess the text data; 2) Construct a bag-of-words model: Represent the text data of each entity as a bag-of-words model, that is, regard each word in the text data as a feature, and represent each document as a vector, where each dimension of the vector corresponds to the occurrence frequency of a word; 3) Calculate TF-IDF weights: Calculate the TF-IDF value for each word to represent the importance of this word in the document; 4) Calculate similarity: Use cosine similarity to calculate the similarity between the text attributes of each entity. The calculation formula is as follows: ; Where A and B are respectively n-dimensional TF-IDF vector sets of two entities to be compared, i ∈ n, A i , B i are respectively the i-th dimensional TF-IDF vectors of the two entities to be compared.

3. The method for calculating entity similarity of the power safety knowledge graph according to claim 1, wherein, In the text attribute calculation steps, calculate the similarity according to the following steps: (1) Construct a TransE model, collect triple data of entities and relationships, train the TransE model using the maximum likelihood estimation loss function, and optimize the vector representations of entities and relationships; (2) Calculate the entity vector. For each entity e, use the trained TransE model to obtain the corresponding vector representation ; (3) Calculate similarity: Calculate the distance between entity vectors using the L1 or L2 norm.

4. The method for calculating entity similarity of the power safety knowledge graph according to claim 1, wherein In the numerical attribute calculation steps, perform Euclidean distance calculation according to the following formula: ; Among them, , represents the value of the numerical attribute to be calculated, and represent the maximum and minimum values of all numerical attributes to be calculated. m is the total number of numerical attributes to be calculated, and k, j ∈ m; In the list attribute calculation steps, calculate the similarity of each entity according to the following formula: ; where C and F are the sets of list attribute data of two entities to be compared.

5. The method for calculating entity similarity of the power safety knowledge graph according to claim 1, wherein Use path similarity algorithm and graph edit distance algorithm to calculate structural similarity.

6. An entity similarity calculation system for an electric power safety knowledge graph, characterized in that, Including: Numerical attribute calculation module: Obtain the numerical values corresponding to the numerical attributes of each entity in the power safety knowledge entity data, calculate the distance between the numerical vectors of each entity using Euclidean distance or Manhattan distance, and obtain and output the corresponding similarity value according to the calculated distance; List attribute calculation module: Obtain the attribute values corresponding to the list attributes of each entity in the power safety knowledge entity data, calculate and output the similarity between each entity by calculating the number of intersections of the attribute values of each entity, or calculate and output the similarity between each entity by using Jaccard similarity; Text attribute calculation module: Obtain the text data corresponding to the text attributes of each entity in the power safety knowledge entity data, calculate and output the similarity of the text attributes of each entity.

7. The entity similarity calculation system of the power safety knowledge graph according to claim 6, characterized in that The text attribute calculation module calculates the similarity according to the following steps: 1) Preprocessing: Preprocess the text data; 2) Construct the bag-of-words model: Represent the text data of each entity as a bag-of-words model, that is, regard each word in the text data as a feature, and represent each document as a vector, where each dimension of the vector corresponds to the occurrence frequency of a word; 3) Calculate the TF-IDF weights: Calculate the TF-IDF value for each word to represent the importance of this word in the document; 4) Calculate the similarity: Use the cosine similarity to calculate the similarity between the text attributes of each entity. The calculation formula is as follows: ; Among them, A and B are respectively the n-dimensional TF-IDF vector sets of two entities to be compared, i ∈ n, A i , B i are respectively the i-th dimensional TF-IDF vectors of the two entities to be compared.

8. The entity similarity calculation system of the power safety knowledge graph according to claim 6, characterized in that The text attribute calculation module calculates the similarity according to the following steps: (1) Construct a TransE model, collect the triple data of entities and relationships, and train the TransE model using the maximum likelihood estimation loss function to optimize the vector representations of entities and relationships; (2) Calculate the entity vectors. For each entity e, use the trained TransE model to obtain the corresponding vector representation ; (3) Calculate the similarity: Calculate the distance between entity vectors using the L1 or L2 norm.

9. The entity similarity calculation system of the power safety knowledge graph according to claim 6, characterized in that, The numerical attribute calculation module performs Euclidean distance calculation according to the following formula: ; Among them, , represents the value of the numerical attribute to be calculated, , represent the maximum and minimum values of all numerical attributes to be calculated, m is the total number of numerical attributes to be calculated, k, j ∈ m; In the calculation steps of list-type attributes, calculate the similarity of each entity according to the following formula: ; Among them, C and F are the sets of list-type attribute data of two entities to be compared.

10. The entity similarity calculation system of the power safety knowledge graph according to claim 6, characterized in that, It also includes a structural similarity calculation module, and the structural similarity calculation module uses the path similarity algorithm and the graph edit distance algorithm to calculate the structural similarity.

Citation Information

Patent Citations

  • Method and system for calculating entity similarity in knowledge graph

    CN108846080A

  • Knowledge graph construction method for power grid main equipment

    CN112612902A

  • Entity updating method and system of knowledge graph

    CN115809340A

  • Knowledge representation learning model construction method based on intelligent equipment vulnerability knowledge graph

    CN118316662A