Knowledge graph construction method and construction system

By constructing a dynamic knowledge graph with multi-dimensional associations, the problem that existing power grid knowledge graphs cannot capture deep associations is solved, enabling efficient fault reasoning and decision support for large power grid models, and improving the safety and operation and maintenance response speed of the power system.

CN120911563APending Publication Date: 2025-11-07STATE GRID JIANGSU ELECTRIC POWER CO LTD NANTONG POWER SUPPLY BRANCH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510831220.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing methods for constructing power grid knowledge graphs cannot capture complex causal chains or deep connections between appearances and essences. This makes large power grid models susceptible to interference from local data during fault diagnosis, making it difficult to generate global and interpretable decision recommendations, thus affecting the safe and stable operation of the power system.

Method used

We employ a multi-dimensional, dynamically updatable knowledge graph construction method. Through function mapping analysis, we extract causal, substantive, and parallel relationships from structured data. Combined with event topic modeling, we construct event models for unstructured data, thereby building a three-dimensional knowledge graph that enables dynamic matching and updating of structured and unstructured data.

Benefits of technology

It improves the inference accuracy and efficiency of the large power grid model, solves the problems of high misjudgment rate and difficulty in tracing the source caused by data fragmentation and static updates, and enhances the safety and operation and maintenance response speed of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120911563A_ABST
    Figure CN120911563A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge graph construction method and system, and the method comprises the steps: obtaining the multi-source data of a power grid, and dividing the multi-source data into structured data and unstructured data; for the structured data, extracting causal, table and parallel relationships among data vectors; constructing an event model containing weight elements for the unstructured data; performing approximate matching on elements of the event model and the structured data vector, and defining causal, table and parallel relationships among events; constructing a three-dimensional knowledge graph based on the relationship between events; through three-dimensional event relation modeling and combination of dynamic matching verification of structured and unstructured data, intelligence of fault reasoning and decision support is realized, reasoning accuracy and efficiency of a large power grid model are improved, the problems of high misjudgment rate and difficult traceability caused by data splitting, flat structure and static updating in the prior art are solved, and the method is suitable for popularization and application. And the safety and the operation and maintenance response speed of the power system are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to artificial intelligence technology, and in particular to a knowledge graph construction method and a construction system. BACKGROUND

[0002] Under the background of the intelligent development of the power system, the knowledge graph technology has become a key tool to support the reasoning ability of the power grid large model. In the prior art, the construction of the power grid knowledge graph mainly relies on two methods: one is entity relationship extraction based on structured data, which establishes a simple association between entities through predefined rules or statistical models; the other is topic modeling or event extraction based on unstructured text, which extracts key words or event fragments using natural language processing technology. However, these existing technologies have significant limitations: first, the structured data processing method can ensure data accuracy, but can only express static two-dimensional relationships and cannot capture complex causal chains or deep relationships between appearance and essence; second, the unstructured text analysis technology can extract event topics, but lacks dynamic association with structured data, resulting in event elements that cannot be matched and verified with real-time monitoring data. And the existing knowledge graph is mostly flat or tree structure, which is difficult to expand horizontally for similar events, cannot trace back to the root cause of the fault vertically, and cannot be dynamically updated to reflect the changes in the power grid state. These problems seriously restrict the reasoning ability of the power grid large model, making it vulnerable to local data interference in fault diagnosis and difficult to generate global and interpretable decision recommendations, ultimately affecting the safe and stable operation of the power system. SUMMARY

[0003] The purpose of the present application is to provide a knowledge graph construction method with multi-dimensional association, dynamic update and multi-source data fusion, and the other purpose of the present application is to provide a construction system of the knowledge graph.

[0004] Technical solution: The knowledge graph construction method provided by the present application comprises the following contents:

[0005] Obtaining power grid multi-source data, and dividing the data into structured data and unstructured data according to the format;

[0006] For structured data, the causal, surface and parallel relationships between data vectors are extracted through function mapping analysis; for unstructured data, an event model containing weight elements is constructed through event topic modeling;

[0007] Approximate matching of the elements of the event model and the structured data vectors based on text similarity is performed to obtain the association pairs of the elements and the vectors, and the causal, surface and parallel relationships between the structured data vectors are mapped to the corresponding event model elements based on the association pairs, and the corresponding relationships between events are defined;

[0008] A three-dimensional knowledge graph is constructed based on the causality, surface-internal and parallel relationship between events, wherein: the breadth dimension is associated with parallel relationship events, the length dimension is associated with causality event chains, and the depth dimension is associated with surface-internal relationship events.

[0009] Preferably, the structured data includes databases, tables, CSV files, etc. This data type has the following characteristics: clear data attributes, and a data set that has been manually screened, has a clear meaning or direction, and can be directly used for verification and association of various relationships in the knowledge graph.

[0010] Unstructured data includes text files, web content, research papers, etc. This data type has the following characteristics: it implies potential event descriptions and logical relationships between events, and is the main information source for knowledge graph construction, but it lacks corresponding evidence and associated data for constructing logical relationships between events.

[0011] Preferably, the function mapping analysis includes:

[0012] Each row or column of structured data is defined as a data vector X_i;

[0013] The mapping relationship between data vectors X_n = f(X_i,…,X_j) is fitted by a neural network regression method;

[0014] Based on the fitted mapping relationship type, the causality, surface-internal and parallel relationship between data vectors are defined:

[0015] If there is a mapping relationship X_n = f(X_i,…,X_j), the data vector corresponding to the independent variable vector X_i,…,X_j has a causal relationship with the dependent variable vector X_n;

[0016] If there are two mapping relationships: X_1 = f1(X_i,…,X_j), X_2 = f2(X_m,…,X_n), and the vector X_1 and the vector X_2 have a one-to-one mapping relationship g, i.e. X_1 = g(X_2), then the vectors X_i,…,X_j and X_m,…,X_n have a surface-internal relationship;

[0017] If there are two mapping relationships X_1 = f1(X_i,…,X_j), X_1 = f2(X_m,…,X_n), i.e. f1(X_i,…,X_j) = f2(X_m,…,X_n), then the vectors X_i,…,X_j and X_m,…,X_n have a parallel relationship.

[0018] Preferably, the event model containing weight elements is generated by event topic modeling, specifically:

[0019] The unstructured text is segmented into a sentence set S = {s1, s2,…,sn}, where each si represents a sentence;

[0020] annotating the nouns in the set of sentences S by natural language processing techniques;

[0021] calculating the frequency of occurrence of each noun t in each sentence si, F(t,si), where F(t,si) = the number of occurrences of t in sentence si divided by the total number of nouns in sentence si;

[0022] generating a matrix of noun frequencies M for each unstructured text passage;

[0023] extracting the nouns in the sentences and calculating the matrix of frequencies M = [F(tm,sn)], n being the number of sentences and m the number of nouns;

[0024] performing an eigenvalue analysis of the matrix of noun frequencies M, selecting the sentence si corresponding to the eigenvector with the largest eigenvalue as the key sentence and determining the associated nouns tj as the element attributes; based on the key sentence si, extracting the event by syntactic structure matching and converting the element attributes tj into event elements, the weight of the event elements being determined by the corresponding F(tj,si) value, F(tj,si) being positively correlated with the weight value.

[0025] Preferably, the approximate matching of the text similarity uses the edit distance algorithm to associate the event elements with a higher similarity than the threshold value with the data vector.

[0026] Preferably, it further comprises:

[0027] When new power grid data is added, the structured data analysis, event modeling and association relationship construction are re-executed at a preset period to update the event association relationship of the three-dimensional knowledge graph.

[0028] A system for constructing the above knowledge graph, comprising:

[0029] a data storage module for storing structured data and unstructured data respectively; a structured processing module for performing function mapping analysis to extract the relationship between data vectors; an unstructured processing module for performing event theme modeling and generating an event model; a graph construction module for associating the event model elements with the structured data and constructing a three-dimensional knowledge graph.

[0030] Preferably, the structured processing module comprises: a vector definition unit for converting data rows or columns into vectors; a relationship analysis unit for fitting mapping relationships by machine learning and defining cause-and-effect, inside-and-outside, and parallel relationships.

[0031] Preferably, the unstructured processing module comprises: a text segmentation unit for segmenting unstructured text into a set of sentences; a matrix generation unit for constructing a matrix of noun frequencies; an eigenvalue processing unit for screening key sentences and element attributes.

[0032] Preferably, the graph construction module comprises: a fuzzy matching unit using an edit distance algorithm to match event elements with data vectors, and a three-dimensional linking unit for associating events in terms of breadth, length and depth dimensions.

[0033] Preferably, it further comprises a hardware acceleration module for performing machine learning and matrix operations using parallel processors.

[0034] A computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the knowledge graph construction method described above.

[0035] Advantages: Compared with the prior art, the present application has the following significant advantages: through three-dimensional event relationship modeling and combined with dynamic matching verification of structured and unstructured data, the intelligence of fault reasoning and decision support is realized, which not only improves the reasoning accuracy and efficiency of the power grid large model, but also solves the problems of high misjudgment rate and difficult traceability caused by data fragmentation, flat structure and static update of existing devices, significantly improving the safety and operation response speed of the power system. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 The knowledge graph structure of the present application is shown.

[0037] Figure 2 The system structure of the present application is shown. DETAILED DESCRIPTION

[0038] The technical solutions of the present application will be further described below with reference to the accompanying drawings Figures 1-2 The technical solutions of the present application will be further described below with reference to the accompanying drawings

[0039] As shown in the accompanying drawings, the present embodiment provides a knowledge graph construction method and construction system, which comprises the following contents: Figure 1

[0040] Obtain power grid multi-source data, and divide the data into structured data and unstructured data according to the format.

[0041] Structured data includes databases, tables, CSV files, etc. The data type has the following characteristics: it has clear data properties and is a data set selected by humans, and has clear meaning or direction, and can be directly used for verification and association of various relationships in the knowledge graph.

[0042] Unstructured data includes text files, web content, research papers, etc. The data type has the following characteristics: it implies potential event descriptions and logical relationships between events, and is the main information source for knowledge graph construction, but lacks corresponding evidence and associated data for constructing logical relationships between events.

[0043] ​For structured data, the causal, surface and parallel relationships between data vectors are extracted through function mapping analysis; for unstructured data, an event model containing weight elements is constructed through event topic modeling.

[0044] The function mapping analysis includes:

[0045] Each row or column of structured data is defined as a data vector X_i;

[0046] The mapping relationship between data vectors X_n = f(X_i,…,X_j) is fitted through a neural network regression method;

[0047] Based on the fitted mapping relationship type, the causal, surface and parallel relationships between data vectors are defined:

[0048] If there is a mapping relationship X_n = f(X_i,…,X_j), the data vectors corresponding to the independent variable vectors X_i,…,X_j have a causal relationship with the dependent variable vector X_n;

[0049] If there are two mapping relationships: X_1 = f1(X_i,…,X_j), X_2 = f2(X_m,…,X_n), and there is a one-to-one mapping relationship g between vector X_1 and vector X_2, i.e. X_1 = g(X_2), then vectors X_i,…,X_j and X_m,…,X_n have a surface relationship;

[0050] If there are two mapping relationships X_1 = f1(X_i,…,X_j), X_1 = f2(X_m,…,X_n), i.e. f1(X_i,…,X_j) = f2(X_m,…,X_n), then vectors X_i,…,X_j and X_m,…,X_n have a parallel relationship.

[0051] The event model containing weight elements generated through event topic modeling is specifically:

[0052] The unstructured text is segmented into a sentence set S = {s1,s2,…,sn}, where each si represents a sentence;

[0053] The nouns in the sentence set S are labeled through natural language processing technology;

[0054] The frequency F(t,si) of each noun t in each sentence si is calculated, where F(t,si) = the number of occurrences of t in sentence si divided by the total number of nouns in sentence si;

[0055] For each unstructured text paragraph, a noun frequency matrix M is generated;

[0056] Extract the nouns in the sentence and calculate the frequency matrix M = [F(tm,sn)], n is the number of sentences, m is the number of nouns;

[0057] Perform eigenvalue analysis on the noun frequency matrix M, select the sentence si corresponding to the largest eigenvalue as the key sentence, and determine the associated noun tj as the element attribute; based on the key sentence si, extract events through syntax structure matching, and convert the element attribute tj into event elements, wherein the weight of the event elements is determined by the corresponding F(tj,si) value, and F(tj,si) is positively correlated with the weight value.

[0058] The element of the event model is matched with the structured data vector using the edit distance algorithm, and the event elements with a similarity higher than a threshold value are associated with the data vector to obtain an association pair of elements and vectors, and based on the association pair, the cause-effect, surface-internal and parallel relationship between the structured data vectors is mapped to the corresponding event model elements to define the corresponding relationship between events.

[0059] Based on the cause-effect, surface-internal and parallel relationship between events, a three-dimensional knowledge graph is constructed, wherein: the breadth dimension is associated with parallel relationship events, the length dimension is associated with cause-effect relationship event chains, and the depth dimension is associated with surface-internal relationship events.

[0060] When new power grid data is added, the structured data analysis, event modeling and association relationship construction are re-executed at a preset period to update the event association relationship of the three-dimensional knowledge graph.

[0061] As shown in Figure 2 The embodiment also provides a knowledge graph construction system, which comprises: a data storage module for storing structured data and unstructured data respectively; the data storage module can use a relational database such as MySQL or OceanBase; a structured processing module for performing function mapping analysis to extract the relationship between data vectors; an unstructured processing module for performing event theme modeling and generating an event model; a graph construction module for associating event model elements with structured data and constructing a three-dimensional knowledge graph; and a hardware acceleration module for using a parallel processor to perform machine learning and matrix operation.

[0062] The structured processing module comprises: a vector definition unit for converting data rows or columns into vectors; and a relationship analysis unit for fitting a mapping relationship through machine learning and defining cause-effect, surface-internal and parallel relationships. The unstructured processing module comprises: a text segmentation unit for segmenting unstructured text into a sentence set, and converting unstructured data into an electronic document through an OCR tool; a matrix generation unit for constructing a noun frequency matrix; and an eigenvalue processing unit for screening key sentences and element attributes. The graph construction module comprises: a fuzzy matching unit for matching event elements with data vectors using an edit distance algorithm; and a three-dimensional linking unit for associating events according to the breadth, length and depth dimensions.

[0063] A computer readable storage medium stores computer program instructions, when the computer program instructions are executed by a processor, the knowledge graph construction method is realized.

Claims

1. A method for constructing a knowledge graph, characterized in that, The method comprises the following steps: Obtain power grid multi-source data, and divide the data into structured data and unstructured data according to formats; For structured data, analyze and extract the cause-and-effect, surface-and-depth and parallel relationship between data vectors through function mapping; For unstructured data, build an event model containing weight elements through event theme modeling; Approximate match the elements of the event model with the structured data vectors based on text similarity, obtain the association pairs of the elements and the vectors, and map the cause-and-effect, surface-and-depth and parallel relationship between the structured data vectors to the corresponding event model elements based on the association pairs, and define the corresponding relationship between events; Build a three-dimensional knowledge graph based on the cause-and-effect, surface-and-depth and parallel relationship between events, wherein: the breadth dimension is associated with parallel relationship events, the length dimension is associated with cause-and-effect relationship event chains, and the depth dimension is associated with surface-and-depth relationship events.

2. The method of claim 1, wherein, The function mapping analysis comprises: Define each row or column of structured data as a data vector X_i; Fit the mapping relationship between data vectors X_n=f(X_i,…,X_j) through a neural network regression method; Define the cause-and-effect, surface-and-depth and parallel relationship between data vectors based on the fitted mapping relationship types: If there is a mapping relationship X_n=f(X_i,…,X_j), the data vectors corresponding to the independent variable vectors X_i,…,X_j have a cause-and-effect relationship with the dependent variable vector X_n; If there are two mapping relationships: X_1=f1(X_i,…,X_j), X_2=f2(X_m,…,X_n), and there is a one-to-one mapping relationship g between vector X_1 and vector X_2, i.e. X_1=g(X_2), then vectors X_i,…,X_j and X_m,…,X_n have a surface-and-depth relationship; If there are two mapping relationships X_1=f1(X_i,…,X_j), X_1=f2(X_m,…,X_n), i.e. f1(X_i,…,X_j)=f2(X_m,…,X_n), then vectors X_i,…,X_j and X_m,…,X_n have a parallel relationship.

3. The method of claim 1, wherein, Generating an event model containing weight elements through event theme modeling comprises: Divide unstructured text into a sentence set S={s1,s2,…,sn}, wherein each si represents a sentence; Annotate the nouns in the sentence set S through natural language processing technology; Calculate the frequency F(t,si) of each noun t in each sentence si, wherein F(t,si)=the number of occurrences of t in sentence si divided by the total number of nouns in sentence si; Generate a noun frequency matrix M for each unstructured text paragraph; Extract the nouns in the sentence and calculate the frequency matrix M=[F(tm,sn)], n is the number of sentences, and m is the number of nouns; Eigenvalue analysis is performed on the noun frequency matrix M, a sentence si corresponding to a characteristic vector with the largest eigenvalue is selected as a key sentence, and an associated noun tj is determined as an element attribute; based on the key sentence si, an event is extracted through syntax structure matching, and the element attribute tj is converted into an event element, wherein the weight of the event element is determined by the corresponding F(tj,si) value, and F(tj,si) is positively correlated with the weight value.

4. The method of claim 1, wherein, The approximate matching of the text similarity adopts an edit distance algorithm, and event elements with a similarity higher than a threshold value are associated with data vectors.

5. The method of claim 1, wherein, Further comprising: When new power grid data is added, the structured data analysis, event modeling and association relationship construction are re-executed at a preset period, and the event association relationship of the three-dimensional knowledge graph is updated.

6. The system for constructing a knowledge graph of claim 1, wherein, Comprising: A data storage module that respectively stores structured data and unstructured data; A structured processing module that performs function mapping analysis to extract relationships between data vectors; An unstructured processing module that performs event theme modeling and generates an event model; a graph construction module that associates event model elements with structured data and constructs a three-dimensional knowledge graph.

7. The system of claim 6, wherein, The structured processing module comprises: a vector definition unit that converts data rows or columns into vectors; and a relationship analysis unit that fits mapping relationships through machine learning and defines cause-and-effect, inside-and-outside and parallel relationships.

8. The system of claim 6, wherein, The unstructured processing module comprises: a text segmentation unit that segments unstructured text into a sentence set; a matrix generation unit that constructs a noun frequency matrix; and an eigenvalue processing unit that screens key sentences and element attributes.

9. The system of claim 6, wherein, The graph construction module comprises: a fuzzy matching unit that matches event elements and data vectors using an edit distance algorithm; and a three-dimensional linking unit that associates events according to the breadth, length and depth dimensions.

10. A computer-readable storage medium, characterized in that, A computer program instruction is stored, and when the computer program instruction is executed by a processor, the knowledge graph construction method of claims 1-5 is realized.