A knowledge graph entity disambiguation fusion method applicable to fault records of power distribution network equipment
By establishing a knowledge graph in the fault records of power distribution network equipment, and using name and structural similarity calculation and weight adjustment, the problems of entity merging and differentiation were solved, the matching accuracy was improved, and the structure of the knowledge graph was perfected.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2026-04-03
AI Technical Summary
During the expansion of the knowledge graph of fault records of power distribution network equipment, the entity names or structural differences make it impossible to effectively merge or distinguish them, resulting in identification errors.
By establishing a basic knowledge graph of power distribution network equipment faults, and using name similarity and structural similarity calculation methods combined with weight adjustment, the entity matching degree is optimized, and identical entities are merged and new branches are opened to eliminate ambiguity.
It improves the matching accuracy of the knowledge graph of fault records of power distribution network equipment, eliminates ambiguity caused by entity duplication and similar factors, and improves the completeness of the knowledge graph.
Smart Images

Figure CN115238084B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power distribution network equipment fault record processing technology, and relates to a knowledge graph entity disambiguation fusion method, particularly a knowledge graph entity disambiguation fusion method applicable to power distribution network equipment fault records. Background Technology
[0002] During the operation, maintenance and repair of power distribution network equipment, a large number of fault troubleshooting records have been accumulated. However, due to their unstructured nature, most of these records remain idle in the operation and maintenance system. Therefore, it is of great significance to deeply explore these resources for the fault handling of power distribution networks.
[0003] Knowledge graphs, as advanced data retrieval and mining technologies, can extract entities from unstructured data and further mine and establish correspondences and attributes between entities, achieving structured representation and knowledge-based display of data. However, during the entity expansion process, knowledge graphs are prone to errors in recognition due to differences in entity names or structures, making effective merging or differentiation difficult.
[0004] A search revealed no publicly available literature of the same or similar prior art as this invention. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and propose a knowledge graph entity disambiguation fusion method applicable to fault records of distribution network equipment. This method can solve the technical problem that during the entity expansion process of the knowledge graph of faults in distribution network equipment, it is easy to fail to effectively merge or distinguish entities due to differences in entity names or structures, resulting in identification errors.
[0006] The present invention solves its practical problem by adopting the following technical solution:
[0007] A knowledge graph entity disambiguation fusion method for fault records of power distribution network equipment includes the following steps:
[0008] Step 1: Establish a basic knowledge graph of power distribution network equipment faults;
[0009] Step 2: Extract entities from the fault records of the power distribution network equipment.
[0010] Step 3: Compare the entity names extracted from each fault record with the power distribution equipment fault knowledge graph from Step 1, and calculate the name similarity and structural similarity respectively.
[0011] Step 4: Based on the name similarity and structure similarity results calculated in Step 3, calculate the comprehensive matching degree between the entities in each fault record and the knowledge graph, select appropriate weights, obtain the optimal matching result, merge the same entities, open new entity branches for different entities to eliminate ambiguity, and improve the knowledge graph of power distribution network equipment faults.
[0012] Furthermore, the specific method of step 1 is as follows:
[0013] By using specialized terminology related to equipment and faults, a basic knowledge graph of power distribution network equipment faults is established and stored in the neo4j graph database.
[0014] Furthermore, the specific method for step 2 is as follows:
[0015] The "character window" word segmentation method was used to extract entities from the text of the power distribution network equipment fault record. The text was truncated character by character in the order of 5, 4, 3, 2 characters, and compared with the power distribution network equipment fault knowledge graph in step 1 to extract entity nouns including equipment name, equipment type, equipment attributes and related faults.
[0016] Furthermore, the specific method for step 3 is as follows:
[0017] The entity names obtained in step 2 are compared with the knowledge graph. First, the name similarity is calculated:
[0018]
[0019] In the formula, m and n are the lengths of characters w1 and w2, D(m,n) is the minimum number of basic character editing operations (i.e., the total number of deletion, addition, reordering, and replacement operations) required to transform character w1 into w2, and max(m,n) is the maximum character length of characters w1 and w2. Further, entity structure similarity is calculated. The entity to be matched is taken as the root node, and the child entity nodes adjacent to the root node in the graph are taken as leaf nodes, forming a tree structure. Then, the similarity is calculated for the leaf nodes:
[0020]
[0021] In the formula, stru(w) represents the number of sub-entities, |stru(w1)∪stru(w2)| represents the total number of sub-entities, and N(w1,w2) represents the sub-entity name similarity. Step 4: Based on the name similarity and structure similarity results calculated in Step 3, calculate the comprehensive matching degree between the entities in each fault record and the knowledge graph, select appropriate weights, obtain the optimal matching result, merge identical entities, and open new entity branches for different entities to eliminate ambiguity and improve the knowledge graph of power distribution network equipment faults.
[0022] The specific steps of step 4 include:
[0023] Based on the entity name similarity and structural similarity calculated in step 3, the overall matching degree of the entities is further calculated:
[0024]
[0025] In the formula, Q represents the weighted sum of entity name similarity N(w1,w2) and structural similarity S(w1,w2) when the entity to be matched and the entity in the knowledge graph do not completely match. The specific value of the weight Q is determined by the matching accuracy P. r Match completeness rate P a The harmonic mean P c The calculations are as follows: Taking a small number of entity samples, we perform manual matching and automatic matching respectively. Y1 represents the number of entities that failed to match automatically, Y2 represents the number of entities that matched correctly automatically, and Y3 represents the number of entities that matched incorrectly automatically. The results are:
[0026]
[0027]
[0028]
[0029] By adjusting the weight Q, P is made c When Q is at its maximum, it serves as the optimal weight value for calculating the overall matching degree and performing automatic entity matching. After obtaining the optimal matching result, identical entities are merged, and new entity branches are created for different entities to eliminate ambiguity and improve the knowledge graph of power distribution network equipment faults.
[0030] Advantages and beneficial effects of the present invention:
[0031] This invention proposes a knowledge graph entity disambiguation fusion method applicable to fault records of distribution network equipment. After initially establishing the knowledge graph, based on name similarity and structural similarity, the comprehensive matching degree between entities in each fault record and the knowledge graph is calculated, and appropriate weights are selected to obtain the optimal matching result. This effectively improves the matching accuracy, eliminates ambiguity caused by entity duplication and similarity during the expansion of the knowledge graph of fault records of distribution network equipment, merges identical entities, and opens new branches for different entities, effectively improving the completeness of the knowledge graph of fault records of distribution network equipment. Attached Figure Description
[0032] Figure 1 This is a flowchart of the knowledge graph entity disambiguation and fusion method for fault records of power distribution network equipment according to the present invention;
[0033] Figure 2This is a schematic diagram of the entities to be matched extracted from the original fault elimination record information based on relevant knowledge graphs according to the present invention.
[0034] Figure 3 This is an example diagram of the knowledge graph obtained by adding entities based on the comprehensive matching degree according to the present invention. Detailed Implementation
[0035] The present invention will be further described in detail below with reference to the accompanying drawings:
[0036] A knowledge graph entity disambiguation fusion method applicable to fault records of power distribution network equipment, such as Figure 1 As shown, it includes the following steps:
[0037] Step 1: Establish a basic knowledge graph of power distribution network equipment faults;
[0038] The specific method for step 1 is as follows:
[0039] Based on the equipment and fault terminology specified in the "Technical Specification for Real-Type Test of Single-Phase Ground Fault in 10kV Distribution Network", "Standardized Operation Process and Application for Emergency Repair of Distribution Network Faults", "Power Supply Bureau Distribution Network Fault Management Standard" and "Technical Specification for Distribution Network Fault Monitoring (2016)", a basic knowledge graph of distribution network equipment faults is established and stored in the neo4j graph database.
[0040] In this embodiment, in step 1, based on natural language processing technology and neo4j graph database software, the equipment and fault professional terms specified in the "Technical Specification for Real-Type Test of Single-Phase Ground Fault in 10kV Distribution Network", "Standardized Operation Process and Application of Emergency Repair of Distribution Network Faults", "Power Supply Bureau Distribution Network Fault Management Standard" and "Technical Specification for Distribution Network Fault Monitoring (2016)" are processed to establish a basic knowledge graph of distribution network equipment faults.
[0041] Step 2: Extract entities from the text of fault records of power distribution network equipment;
[0042] The specific method for step 2 is as follows:
[0043] Entity extraction was performed on the fault records of distribution network equipment using the "character window" word segmentation method. Specific text examples include... Figure 2 As shown, the fault elimination record of a low-voltage distribution area detection terminal disconnection is used. The text is truncated character by character in the order of 5, 4, 3, and 2 characters respectively, and compared with the distribution network equipment fault knowledge graph in step 1 to extract entity nouns including equipment name, equipment type, equipment attributes, and related faults.
[0044] In this embodiment, in step 2, the text is truncated character by character using the "character window" segmentation method, with 5, 4, 3, and 2 characters respectively, and entities are extracted. The extracted entities are compared with the knowledge graph of power distribution equipment faults in step 1 to extract entity nouns including equipment name, equipment type, equipment attributes, and related faults, as shown in Figure 1.
[0045] Table 1. List of Fault Entities in Distribution Network Equipment
[0046]
[0047]
[0048] Step 3: Compare the entity names extracted from each fault record with the power distribution equipment fault knowledge graph from Step 1, and calculate the name similarity and structural similarity respectively.
[0049] The specific method for step 3 is as follows:
[0050] The entity names obtained in step 2 are compared with the knowledge graph. First, the name similarity is calculated:
[0051]
[0052] In the formula, m and n are the lengths of characters w1 and w2, D(m,n) is the minimum number of basic character editing operations (i.e., the total number of deletion, addition, reordering, and replacement operations) required to transform character w1 into w2, and max(m,n) is the maximum character length of characters w1 and w2. Further, entity structure similarity is calculated. The entity to be matched is taken as the root node, and the child entity nodes adjacent to the root node in the graph are taken as leaf nodes, forming a tree structure. Then, the similarity is calculated for the leaf nodes:
[0053]
[0054] In the formula, stru(w) represents the number of sub-entities, |stru(w1)∪stru(w2)| represents the total number of sub-entities, and N(w1,w2) represents the sub-entity name similarity. Step 4: Based on the name similarity and structure similarity results calculated in Step 3, calculate the comprehensive matching degree between the entities in each fault record and the knowledge graph, select appropriate weights, obtain the optimal matching result, merge identical entities, and open new entity branches for different entities to eliminate ambiguity and improve the knowledge graph of power distribution network equipment faults.
[0055] The specific steps of step 4 include:
[0056] Based on the entity name similarity and structural similarity calculated in step 3, the overall matching degree of the entities is further calculated:
[0057]
[0058] In the formula, Q represents the weighted sum of entity name similarity N(w1,w2) and structural similarity S(w1,w2) when the entity to be matched and the entity in the knowledge graph do not completely match. The specific value of the weight Q is determined by the matching accuracy P. r Match completeness rate P a The harmonic mean P c The calculations are as follows: Taking a small number of entity samples, we perform manual matching and automatic matching respectively. Y1 represents the number of entities that failed to match automatically, Y2 represents the number of entities that matched correctly automatically, and Y3 represents the number of entities that matched incorrectly automatically. The results are:
[0059]
[0060]
[0061]
[0062] By adjusting the weight Q, P is made c At its maximum, Q is used as the optimal weight value to calculate the overall matching degree and perform automatic entity matching. After obtaining the optimal matching result, identical entities are merged, and new entity branches are created for different entities to eliminate ambiguity and improve the knowledge graph of distribution network equipment faults. Some graphs are shown below. Figure 3 As shown, the relationships between entities can be seen more clearly through knowledge graphs.
[0063] This invention pre-establishes a basic knowledge graph of power distribution network equipment faults using specialized terminology related to equipment and faults. Entities to be matched are selected through word segmentation, and the overall matching degree between entities in each fault record and the knowledge graph is calculated based on name similarity and structural similarity. Appropriate weights are selected, and after obtaining the optimal matching result, identical entities are merged, while new entity branches are created for different entities to eliminate ambiguity and improve the knowledge graph of power distribution network equipment faults.
[0064] It should be emphasized that the embodiments described in this invention are illustrative rather than limiting. Therefore, this invention includes, but is not limited to, the embodiments described in the specific implementation. Any other implementations derived by those skilled in the art based on the technical solutions of this invention are also within the scope of protection of this invention.
Claims
1. A knowledge graph entity disambiguation and fusion method applicable to fault records of power distribution network equipment, characterized in that: Includes the following steps: Step 1: Establish a basic knowledge graph of power distribution network equipment faults; Step 2: Extract entities from the fault records of the power distribution network equipment. Step 3: Compare the entity names extracted from each fault record with the power distribution equipment fault knowledge graph from Step 1, and calculate the name similarity and structural similarity respectively. Step 4: Based on the name similarity and structure similarity results calculated in Step 3, calculate the comprehensive matching degree between the entities in each fault record and the knowledge graph, select appropriate weights, obtain the optimal matching result, merge the same entities, open new entity branches for different entities to eliminate ambiguity, and improve the knowledge graph of faults in the power distribution network equipment. The specific method for step 3 is as follows: The entity names obtained in step 2 are compared with the knowledge graph. First, the name similarity is calculated: In the formula, m and n are the lengths of characters w1 and w2, D(m,n) is the minimum number of basic character editing operations required to transform character w1 into w2, i.e., the total number of operations for deletion, addition, reordering, and replacement, and max(m,n) is the maximum character length of characters w1 and w2; further, the entity structure similarity is calculated; the entity to be matched is taken as the root node, and the child entity nodes adjacent to the root node in the graph are leaf nodes, forming a tree structure, and then the similarity is calculated for the leaf nodes: In the formula, stru(w) represents the number of child entities, |stru(w1)∪stru(w2)| represents the total number of child entities, and N(w1,w2) represents the similarity of child entity names; The specific steps of step 4 include: Based on the entity name similarity and structural similarity calculated in step 3, the overall matching degree of the entities is further calculated: In the formula, Q represents the weighted sum of entity name similarity N(w1,w2) and structural similarity S(w1,w2) when the entity to be matched and the entity in the knowledge graph do not completely match; the specific value of the weight Q is determined by the matching accuracy P. r Match completeness rate P a The harmonic mean P c The calculations show that, by taking a small number of entity samples and performing manual and automatic matching respectively, Y1 represents the number of entities that failed to match automatically, Y2 represents the number of entities that matched correctly automatically, and Y3 represents the number of entities that matched incorrectly automatically. The results are as follows: By adjusting the weight Q, P is made c At its maximum, Q is used as the optimal weight value to calculate the overall matching degree and perform automatic entity matching. After obtaining the optimal matching result, identical entities are merged, and new entity branches are opened for different entities to eliminate ambiguity and improve the knowledge graph of power distribution network equipment faults.
2. The knowledge graph entity disambiguation and fusion method for fault records of power distribution network equipment according to claim 1, characterized in that: The specific method for step 1 is as follows: By using specialized terminology related to equipment and faults, a basic knowledge graph of power distribution network equipment faults is established and stored in the neo4j graph database.
3. The knowledge graph entity disambiguation and fusion method for fault records of power distribution network equipment according to claim 1, characterized in that: The specific method for step 2 is as follows: The "character window" word segmentation method was used to extract entities from the text of the power distribution network equipment fault record. The text was truncated character by character in the order of 5, 4, 3, 2 characters, and compared with the power distribution network equipment fault knowledge graph in step 1 to extract the equipment name, equipment type, equipment attributes and related fault entity nouns.
Citation Information
Patent Citations
Knowledge graph disambiguation method, system and equipment based on standard text, and medium
CN113569060A
Information representation method and apparatus
WO2021018154A1