A method, apparatus, storage medium, and server for creating a knowledge graph
Through an automated method, the graph knowledge layer schema is used to analyze material text, extract instance element data, and perform horizontal fusion, which solves the problem of low efficiency in knowledge graph creation in the existing technology, and realizes an efficient and automated knowledge graph construction process.
Patent Information
- Application Number
- CN202011015057.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-09-24
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2040-09-24
AI Technical Summary
In the prior art, the creation of knowledge graphs is relatively inefficient, which is mainly due to the need for a lot of manual operations, which leads to the complex and time-consuming process.
By obtaining the material text of the knowledge field to which the knowledge graph to be created belongs, the material text is parsed using the pre-constructed graph knowledge layer schema, the instance element data is extracted, and data fused with the graph knowledge layer schema, to generate a knowledge graph in the vertical field. Then, find the associated graph nodes from the knowledge graph library to generate the final knowledge graph horizontally.
It realizes automation of the knowledge graph creation process, reduces the need for manual operations, improves the efficiency and accuracy of knowledge graph creation, and reduces the communication costs between business personnel and model developers.
Smart Images

Figure CN112163098B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of computer technology, and particularly relates to a method, apparatus, storage medium, and server for creating a knowledge graph. Background Art
[0002] Currently, the process of creating a knowledge graph mainly includes: business personnel sort out the knowledge framework, nodes, relationships, and triples in the vertical domain in Excel, and output the knowledge layer schema in Excel format; hand over the Excel-format knowledge layer schema to the modeling personnel, who write code and store it in the graph database; process and clean the data offline according to the knowledge layer schema, and process the unstructured and semi-structured data into structured data corresponding to the knowledge layer schema; the modeling personnel write code to fuse the knowledge layer schema with the structured data to generate a complete knowledge graph. The above process involves a large amount of manual operations, and the efficiency of creating the knowledge graph is relatively low. Summary of the Invention
[0003] In view of this, this application proposes a method for creating a knowledge graph, which can improve the efficiency of creating the knowledge graph.
[0004] In a first aspect, an embodiment of this application provides a method for creating a knowledge graph, including:
[0005] Obtain the material text of the knowledge field to which the knowledge graph to be created belongs;
[0006] Parse the material text according to the pre-constructed graph knowledge layer schema, and extract instance element data;
[0007] Fuse the extracted instance element data with the graph knowledge layer schema to obtain the first knowledge graph in the vertical domain;
[0008] Search in a preset knowledge graph library for other knowledge graphs that have at least one associated graph node with the first knowledge graph;
[0009] Taking the associated graph nodes as connection points, horizontally fuse the first knowledge graph and the other knowledge graphs to obtain the created knowledge graph.
[0010] In the embodiments of the present application, the user only needs to prepare the schema of the graph knowledge layer of the knowledge graph to be created and the corresponding material text in advance. The system will automatically extract the instance element data in the material text and fuse it with the graph knowledge layer to obtain a knowledge graph in a vertical domain. Then, it will search for other knowledge graphs in the knowledge graph library that have at least one associated graph node with the knowledge graph in the vertical domain. Finally, the individual knowledge graphs will be horizontally fused to obtain the finally created knowledge graph. By setting it like this, the construction progress of the knowledge graph in the vertical domain can be accelerated, and knowledge editing visualization, element parsing automation, model training standardization, and knowledge fusion unification in the process of knowledge graph construction can be realized, reducing the communication cost between business personnel and model developers, and effectively improving the creation efficiency of the knowledge graph.
[0011] Further, parsing the material text according to the pre-constructed schema of the graph knowledge layer to extract instance element data may include:
[0012] Detect the structured data in the material text, and find out the structured data and unstructured data contained in the material text;
[0013] Parse the structured data using a preset rule model to obtain the first instance element data contained therein;
[0014] Parse the unstructured data using a pre-constructed NLP recognition model to obtain the second instance element data contained therein;
[0015] Fuse the first instance element data and the second instance element data to obtain the extracted instance element data.
[0016] For example, for structured data with a high degree of structuring such as name, gender, and age, the rule model can be directly used to extract the instance element data contained therein. For unstructured data, such as in legal field judgment documents, to extract a certain type of dispute focus from all judgment documents under a certain case type, it is necessary to manually label the data, train the NLP model, and then optimize and iterate. After reaching a certain index, it can be parsed to extract the instance element data contained in the unstructured data.
[0017] Further, after searching for other knowledge graphs in the preset knowledge graph library that have at least one associated graph node with the first knowledge graph, it may further include:
[0018] Determine the knowledge field to which the first knowledge graph belongs, and respectively determine the knowledge fields to which each of the found other knowledge graphs belongs;
[0019] From a pre-constructed knowledge domain comparison table, respectively search for the degree of association between the knowledge domain to which each of the other knowledge graphs belongs and the knowledge domain to which the first knowledge graph belongs;
[0020] The horizontal integration of the first knowledge graph and the other knowledge graphs may specifically be:
[0021] The first knowledge graph and the target knowledge graphs whose correlation degree is greater than a preset threshold in each of the other found knowledge graphs are horizontally integrated.
[0022] Through this setting, we can avoid fusing too many other knowledge graphs with low correlation, thereby further improving the accuracy and practicality of knowledge graph fusion.
[0023] Furthermore, horizontally fusing the first knowledge graph and the target knowledge graphs whose correlation is greater than a preset threshold in each of the other found knowledge graphs may include:
[0024] For each of the target knowledge graphs, the length of the horizontal connection line for horizontal fusion with the first knowledge graph is determined according to the respective correlation degree, and then the horizontal connection line of the corresponding length is drawn with the associated graph nodes of each as the connection point to complete the horizontal fusion with the first knowledge graph;
[0025] The greater the correlation degree is, the shorter the length of the corresponding horizontal connection line is.
[0026] When knowledge graphs are horizontally integrated, the knowledge graphs with greater correlation have shorter horizontal connection lines when connected, that is, the distance between the knowledge graphs is closer, which can further improve the practicality of the integrated knowledge graph.
[0027] Furthermore, horizontally fusing the first knowledge graph and the target knowledge graphs whose correlation is greater than a preset threshold in each of the other found knowledge graphs may include:
[0028] Counting the number of the target knowledge graphs;
[0029] If the number of the target knowledge graphs is less than or equal to the set upper limit, all the target knowledge graphs are horizontally merged with the first knowledge graph;
[0030] If the number of the target knowledge graphs is greater than the upper limit, the target knowledge graphs are sorted in descending order according to the degree of association, and the knowledge graphs within the upper limit that are ranked higher in the target knowledge graphs are horizontally merged with the first knowledge graph.
[0031] To control the number of knowledge graphs in the vertical domain for fusion, each of the target knowledge graphs can be sorted in descending order according to the degree of association, and then a set number of the knowledge graphs with the highest rankings among the target knowledge graphs are horizontally fused with the first knowledge graph.
[0032] Furthermore, the horizontal fusion of the first knowledge graph and the other knowledge graphs may include:
[0033] Count the number of associated graph nodes in each of the other knowledge graphs respectively;
[0034] For each of the other knowledge graphs, respectively determine the length of the horizontal connection line for horizontal fusion with the first knowledge graph according to the number of associated graph nodes each has, and then draw a horizontal connection line of the corresponding length with the associated graph nodes each has as connection points to complete the horizontal fusion with the first knowledge graph;
[0035] Among them, the more the number of associated graph nodes, the shorter the length of the corresponding horizontal connection line.
[0036] When the knowledge graphs in the vertical domain are horizontally fused, the knowledge graph with more associated graph nodes has a shorter horizontal connection line during connection, that is, the closer the distance to the first knowledge graph, which can further improve the rationality of knowledge graph fusion and the practicality of the fused knowledge graph.
[0037] Furthermore, the obtaining of the material text of the knowledge domain to which the knowledge graph to be created belongs may include:
[0038] Obtain the graph creation task issued by the user on the specified platform;
[0039] Determine the knowledge domain to which the knowledge graph to be created belongs according to the graph creation task;
[0040] Search for the material text from the storage path corresponding to the knowledge domain to which the knowledge graph to be created belongs.
[0041] The specified platform can be a knowledge graph management platform. Users can solidify professional knowledge of the created knowledge graphs, support operations such as adding, deleting, and modifying nodes of the knowledge graph, adding, deleting, and modifying triples, and adding, deleting, and modifying relationships. The results of the operations are directly synchronized to the underlying graph database to achieve the effect of online visual editing. In addition, if a new knowledge graph needs to be created, users can also issue a graph creation task through this platform. In this graph creation task, the knowledge domain to which the knowledge graph to be created belongs can be selected, and then the corresponding material text can be searched from the corresponding storage path.
[0042] Second aspect, an embodiment of the present application provides an apparatus for creating a knowledge graph, including:
[0043] A material text acquisition module, configured to acquire material texts in the knowledge domain to which the knowledge graph to be created belongs;
[0044] An instance element extraction module, configured to parse the material texts according to a pre-constructed schema of the graph knowledge layer, and extract instance element data;
[0045] A data fusion module, configured to perform data fusion on the extracted instance element data and the schema of the graph knowledge layer to obtain a first knowledge graph in a vertical domain;
[0046] A knowledge graph search module, configured to search for other knowledge graphs having at least one associated graph node with the first knowledge graph from a preset knowledge graph library;
[0047] A knowledge graph fusion module, configured to perform horizontal fusion on the first knowledge graph and the other knowledge graphs with the associated graph nodes as connection points to obtain the created knowledge graph.
[0048] Third aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for creating a knowledge graph proposed in the first aspect of the embodiments of the present application are implemented.
[0049] Fourth aspect, an embodiment of the present application provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for creating a knowledge graph proposed in the first aspect of the embodiments of the present application are implemented.
[0050] It can be understood that the beneficial effects of the above second aspect to the fourth aspect can refer to the relevant descriptions in the above first aspect, and will not be elaborated here. Description of the Drawings
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0052] Figure 1 It is a flowchart of the first embodiment of a method for creating a knowledge graph provided by an embodiment of the present application;
[0053] Figure 2 It is a flowchart of the second embodiment of a method for creating a knowledge graph provided by an embodiment of the present application;
[0054] Figure 3 It is a flowchart of the third embodiment of a method for creating a knowledge graph provided by an embodiment of the present application;
[0055] Figure 4 It is a structural diagram of an embodiment of an apparatus for creating a knowledge graph provided by an embodiment of the present application;
[0056] Figure 5 It is a schematic diagram of a server provided by an embodiment of the present application. Detailed implementation manners
[0057] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application. Additionally, in the description of the specification of the present application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.
[0058] The present application proposes a method for creating a knowledge graph, which can improve the efficiency of creating the knowledge graph.
[0059] It should be understood that the execution subject of the method for creating a knowledge graph proposed in each embodiment of the present application is a server.
[0060] Please refer to Figure 1 , the first embodiment of a method for creating a knowledge graph in an embodiment of the present application includes:
[0061] 101. Obtain the material text of the knowledge domain to which the knowledge graph to be created belongs;
[0062] The knowledge graph includes a knowledge layer schema and an instance layer. The knowledge layer is a data model created manually, containing meaningful concept types in the relevant field and the attributes of these types, that is, the types and attributes of each object in the graph. The instance layer is the instance data of each object in the knowledge layer. For example, the element data in the instance layer corresponding to the object "name" in the knowledge layer includes Zhang San, Li Si, etc. Since the knowledge layer schema involves relatively in-depth professional knowledge, it generally needs to be created manually by experts in the field, while the element data in the instance layer can be extracted by means of artificial intelligence. In the embodiment of the present application, the knowledge layer schema of the knowledge graph to be created is created manually in advance, and in order to extract the instance element data, it is necessary to first obtain some material texts in the relevant knowledge field. For example, for the legal field, the material text can be various types of judgment documents.
[0063] Further, step 101 may include:
[0064] (1) Obtain the graph creation task issued by the user on the specified platform;
[0065] (2) Determine the knowledge field to which the knowledge graph to be created belongs according to the graph creation task;
[0066] (3) Search for the material text from the storage path corresponding to the knowledge field to which the knowledge graph to be created belongs.
[0067] The specified platform can be a knowledge graph management platform. Users can solidify professional knowledge of the created knowledge graph, support operations such as adding, deleting, and modifying nodes of the knowledge graph, adding, deleting, and modifying triples, and adding, deleting, and modifying relationships. The results of the operations are directly synchronized to the underlying graph database to achieve the effect of online visual editing. In addition, if a new knowledge graph needs to be created, the user can also issue a graph creation task through this platform. In this graph creation task, the knowledge field to which the knowledge graph to be created belongs can be selected, and then the corresponding material text can be searched from the corresponding storage path. For example, a storage area for material texts can be constructed, divided into multiple areas according to different knowledge fields, with each knowledge field corresponding to one area, and the material texts of each different knowledge field are stored in the corresponding area in advance.
[0068] 102. Parse the material text according to the pre-constructed graph knowledge layer schema to extract instance element data;
[0069] Next, parse the material text according to the pre-constructed schema of the knowledge graph layer to extract instance element data. The schema of a knowledge graph is equivalent to a data model within a domain, which contains meaningful concept types and their attributes in this domain. The schema of any domain is mainly expressed by types and properties. The instance layer is the instance data of each object in the knowledge layer. By parsing and identifying the corresponding material text, the instance element data contained therein can be extracted.
[0070] Further, step 102 may include:
[0071] (1) Detect the structured data in the material text to find out the structured data and unstructured data contained in the material text;
[0072] (2) Parse the structured data using a preset rule model to obtain the first instance element data contained therein;
[0073] (3) Parse the unstructured data using a pre-constructed NLP recognition model to obtain the second instance element data contained therein;
[0074] (4) Integrate the first instance element data and the second instance element data to obtain the extracted instance element data.
[0075] For example, for structured data with a high degree of structuring such as name, gender, and age, the rule model can be directly used to extract the instance element data contained therein. For unstructured data, such as in judgment documents in the legal field, to extract a certain type of dispute focus from all judgment documents under a certain case type, it is necessary to manually label the data, train the NLP model, and then optimize and iterate. After reaching a certain index, perform parsing to extract the instance element data contained in the unstructured data.
[0076] In actual operation, the data parsing task can be divided into two categories: one is the instance element data that can be directly extracted by rules, which is directly parsed by the rule model of the parsing platform; the other is the instance element data that requires NLP model training for parsing. This part of the task is then dispatched to a preset annotation platform. After the task is dispatched, professional personnel perform data annotation to solidify professional knowledge, and then use the annotated data as the training data set of the model for model training, and continuously optimize and iterate. When the model reaches a certain accuracy, parse such instance elements, and then integrate the elements parsed by the rule class and the elements parsed by the NLP model to obtain the final extracted instance element data.
[0077] 103. Fuse the extracted instance element data with the schema of the knowledge graph layer of the atlas to obtain the first knowledge graph in the vertical domain;
[0078] Then, fuse the extracted instance element data with the schema of the knowledge graph layer of the atlas, that is, the fusion of vertical knowledge, and then complete the construction of the knowledge graph in the vertical domain. Specifically, this data fusion is to associate each instance element data with each graph node (object) in the schema of the knowledge graph layer of the atlas respectively. For example, associate the instance element data "request for repayment of the principal of the loan" with the graph node "claim".
[0079] 104. Search in the preset knowledge graph library for other knowledge graphs that have at least one associated graph node with the first knowledge graph;
[0080] After obtaining the first knowledge graph in the vertical domain, search in the knowledge graph library for other knowledge graphs that have at least one associated graph node with the first knowledge graph. This knowledge graph library stores the created knowledge graphs in each different vertical domain or non-vertical domain. The system will respectively obtain the graph nodes of each knowledge graph in this knowledge graph library, and judge whether there is at least one associated graph node between each knowledge graph and the first knowledge graph, and find out those parts of these knowledge graphs that have at least one associated graph node with the first knowledge graph. Specifically, the associated graph nodes can be graph nodes with the same or similar names, or the same or similar attributes.
[0081] 105. Using the associated graph nodes as connection points, horizontally fuse the first knowledge graph and the other knowledge graphs to obtain the created knowledge graph.
[0082] Finally, using the associated graph nodes as connection points, horizontally fuse the first knowledge graph and the other knowledge graphs to obtain the final knowledge graph. For example, both the legal knowledge graph and the risk control knowledge graph have a graph node of "subject", so the legal knowledge graph and the risk control knowledge graph can complete the horizontal knowledge fusion in different vertical domains through this graph node of "subject" to obtain the final knowledge graph.
[0083] In the embodiment of the present application, the user only needs to prepare the schema of the graph knowledge layer of the knowledge graph to be created and the corresponding material text in advance. The system will automatically extract the instance element data in the material text and fuse it with the graph knowledge layer to obtain a knowledge graph in a vertical domain. Then, it will search for other knowledge graphs in the knowledge graph library that have at least one associated graph node with the knowledge graph in the vertical domain. Finally, the knowledge graphs in each vertical domain will be horizontally fused to obtain the finally created knowledge graph. By setting it like this, the construction progress of the knowledge graph in the vertical domain can be accelerated, and the knowledge editing visualization, element analysis automation, model training standardization, and knowledge fusion unification in the process of knowledge graph construction can be realized, reducing the communication cost between business personnel and model developers, and effectively improving the creation efficiency of the knowledge graph.
[0084] Please refer to Figure 2 , the second embodiment of a method for creating a knowledge graph in the embodiment of the present application includes:
[0085] 201. Obtain the material text of the knowledge domain to which the knowledge graph to be created belongs;
[0086] 202. Analyze the material text according to the pre-constructed schema of the graph knowledge layer, and extract the instance element data;
[0087] 203. Perform data fusion on the extracted instance element data and the schema of the graph knowledge layer to obtain the first knowledge graph in the vertical domain;
[0088] 204. Search for other knowledge graphs in the preset knowledge graph library that have at least one associated graph node with the first knowledge graph;
[0089] Steps 201-204 are the same as steps 101-104, and the relevant descriptions of steps 101-104 can be specifically referred to.
[0090] 205. Determine the knowledge domain to which the first knowledge graph belongs, and respectively determine the knowledge domains to which each of the other found knowledge graphs belongs;
[0091] 206. Search for the association degrees between the knowledge domains to which each of the other knowledge graphs belongs and the knowledge domain to which the first knowledge graph belongs in the pre-constructed knowledge domain comparison table;
[0092] 207. Using the associated graph nodes as connection points, horizontally fuse the first knowledge graph and the target knowledge graphs among the other found knowledge graphs whose association degrees are greater than the preset threshold to obtain the created knowledge graph.
[0093] For steps 205-207, an example is as follows: Assume that the knowledge domain of the first knowledge graph is civil law, and the knowledge domains of the three other found knowledge graphs are criminal law, administrative law, and public education respectively. Then, obtain a pre-constructed knowledge domain comparison table, which records the correlation degrees between different knowledge domains. Assume that the correlation degrees between civil law and criminal law, administrative law found from the knowledge domain comparison table are both 0.9, while the correlation degree between civil law and public education is 0.3. Then, the knowledge graph with a correlation degree less than the set threshold of 0.6, that is, the knowledge graph in the public education knowledge domain, can be removed. Finally, the knowledge graphs for horizontal fusion are the first knowledge graph, the knowledge graph in the criminal law knowledge domain, and the knowledge graph in the administrative law knowledge domain. By setting it this way, it is possible to avoid fusing too many other knowledge graphs with low correlation degrees, thereby further improving the accuracy and practicality of knowledge graph fusion.
[0094] Further, the horizontal fusion of the first knowledge graph and the target knowledge graphs with a correlation degree greater than the preset threshold among the found other knowledge graphs may include:
[0095] (1) Count the number of the target knowledge graphs;
[0096] (2) If the number of the target knowledge graphs is less than or equal to the set upper limit of the number, then horizontally fuse all the target knowledge graphs with the first knowledge graph;
[0097] (3) If the number of the target knowledge graphs is greater than the upper limit of the number, then sort the target knowledge graphs in descending order of the correlation degree, and horizontally fuse the top upper limit number of knowledge graphs among the target knowledge graphs with the first knowledge graph.
[0098] To control the number of knowledge graphs for fusion, the target knowledge graphs can be sorted in descending order of the correlation degree, and then the top set upper limit number of knowledge graphs among the target knowledge graphs are horizontally fused with the first knowledge graph.
[0099] Further, the horizontal fusion of the first knowledge graph and the target knowledge graphs with a correlation degree greater than the preset threshold among the found other knowledge graphs may include:
[0100] For each of the target knowledge graphs, the length of the horizontal connection line for horizontal fusion with the first knowledge graph is determined respectively according to their respective correlation degrees, and then, with the associated graph nodes they each have as connection points, horizontal connection lines of corresponding lengths are drawn to complete the horizontal fusion with the first knowledge graph; wherein, the greater the correlation degree, the shorter the length of the corresponding horizontal connection line.
[0101] For example, the target knowledge graphs are A, B, and C. Among them, the correlation degree between A and the knowledge domain to which the first knowledge graph belongs is the highest, and the correlation degree between B and the knowledge domain to which the first knowledge graph belongs is the lowest. Then, when horizontally fusing the knowledge graphs, the length of the horizontal connection line between A and the first knowledge graph is the shortest, the length of the horizontal connection line between B and the first knowledge graph is the longest, and the length of the horizontal connection line between C and the first knowledge graph is in the middle of the two. When horizontally fusing the knowledge graphs, the greater the correlation degree of a knowledge graph, the shorter the horizontal connection line during connection, that is, the closer the distance between the knowledge graphs, which can further improve the practicability of the fused knowledge graph.
[0102] In the embodiments of the present application, the user only needs to prepare in advance the schema of the graph knowledge layer of the knowledge graph to be created and the corresponding material text. The system will automatically extract the instance element data in the material text and fuse it with the graph knowledge layer to obtain a knowledge graph in a vertical domain; then, other knowledge graphs having at least one associated graph node with the knowledge graph in the vertical domain will be searched from the knowledge graph library, and finally, each knowledge graph will be horizontally fused to obtain the finally created knowledge graph. Moreover, when horizontally fusing each knowledge graph in this embodiment, the correlation degree between the knowledge domains to which each knowledge graph belongs will be determined, and the knowledge graphs with a correlation degree greater than a preset threshold will be fused, which can avoid fusing too many other knowledge graphs with a low correlation degree, thereby further improving the accuracy and practicability of knowledge graph fusion.
[0103] Please refer to Figure 3 , the third embodiment of a method for creating a knowledge graph in the embodiments of the present application includes:
[0104] 301. Obtain the material text of the knowledge domain to which the knowledge graph to be created belongs;
[0105] 302. Analyze the material text according to the pre-constructed schema of the graph knowledge layer, and extract the instance element data;
[0106] 303. Perform data fusion on the extracted instance element data and the schema of the graph knowledge layer to obtain the first knowledge graph in the vertical domain;
[0107] 304. Search for other knowledge graphs in the preset knowledge graph library that have at least one associated graph node with the first knowledge graph;
[0108] Steps 301-304 are the same as steps 101-104, and the specific descriptions can be referred to the relevant descriptions of steps 101-104.
[0109] 305. Count the number of the associated graph nodes of each of the other knowledge graphs respectively;
[0110] 306. For each of the other knowledge graphs, determine the length of the horizontal connection line for horizontal fusion with the first knowledge graph respectively according to the number of the associated graph nodes each has, and then draw a horizontal connection line with the corresponding length with the associated graph nodes each has as connection points to complete the horizontal fusion with the first knowledge graph, and obtain the created knowledge graph.
[0111] For steps 305-306, the following is an example: Assume that the found other knowledge graphs are A, B, and C respectively. Among them, the number of associated graph nodes between A and the first knowledge graph is 1, the number of associated graph nodes between B and the first knowledge graph is 3, and the number of associated graph nodes between C and the first knowledge graph is 2. Then when horizontally fusing the knowledge graphs, the length of the horizontal connection line between A and the first knowledge graph is the longest, the length of the horizontal connection line between B and the first knowledge graph is the shortest, and the length of the horizontal connection line between C and the first knowledge graph is in the middle of the two. When horizontally fusing the knowledge graphs in the vertical domain, the more the number of associated graph nodes a knowledge graph has, the shorter the horizontal connection line during connection, that is, the closer the distance to this first knowledge graph, which can further improve the rationality of knowledge graph fusion and the practicality of the fused knowledge graph.
[0112] In the embodiment of the present application, the user only needs to prepare the schema of the graph knowledge layer of the knowledge graph to be created and the corresponding material text in advance. The system will automatically extract the instance element data in the material text and fuse it with the graph knowledge layer to obtain a knowledge graph in a vertical domain; then, it will search for other knowledge graphs in the knowledge graph library that have at least one associated graph node with the knowledge graph in this vertical domain, and finally horizontally fuse each knowledge graph to obtain the finally created knowledge graph. Moreover, in this embodiment, when horizontally fusing each knowledge graph, it will count the number of the associated graph nodes of each of the other knowledge graphs respectively. When horizontally fusing the knowledge graphs, the more the number of associated graph nodes a knowledge graph has, the shorter the horizontal connection line during connection, which can further improve the rationality of knowledge graph fusion and the practicality of the fused knowledge graph.
[0113] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0114] Corresponding to the method for creating a knowledge graph described in the above embodiments, Figure 4 The structural block diagram of a device for creating a knowledge graph provided by an embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0115] Referring to Figure 4 , the device includes:
[0116] A material text acquisition module 401, configured to acquire material texts in the knowledge field to which the knowledge graph to be created belongs;
[0117] An instance element extraction module 402, configured to parse the material text according to a pre-constructed schema of the graph knowledge layer, and extract instance element data;
[0118] A data fusion module 403, configured to perform data fusion on the extracted instance element data and the schema of the graph knowledge layer to obtain a first knowledge graph in the vertical domain;
[0119] A knowledge graph search module 404, configured to search for other knowledge graphs having at least one associated graph node with the first knowledge graph from a preset knowledge graph library;
[0120] A knowledge graph fusion module 405, configured to perform horizontal fusion of the first knowledge graph and the other knowledge graphs with the associated graph nodes as connection points to obtain the created knowledge graph.
[0121] Further, the instance element extraction module may include:
[0122] A data structure detection unit, configured to detect the structured data of the material text, and find out the structured data and unstructured data included in the material text;
[0123] A first data parsing unit, configured to parse the structured data by using a preset rule model to obtain first instance element data included therein;
[0124] A second data parsing unit, configured to parse the unstructured data by using a pre-constructed NLP recognition model to obtain second instance element data included therein;
[0125] A data fusion unit for fusing the first instance feature data and the second instance feature data to obtain the extracted instance feature data.
[0126] Further, the knowledge graph creation device may further include:
[0127] A knowledge domain determination module for determining the knowledge domain to which the first knowledge graph belongs, and respectively determining the knowledge domains to which each of the found other knowledge graphs belongs;
[0128] A domain correlation degree search module for respectively searching, from a pre-constructed knowledge domain comparison table, the correlation degree between the knowledge domain to which each of the other knowledge graphs belongs and the knowledge domain to which the first knowledge graph belongs;
[0129] The knowledge graph fusion module may be used for:
[0130] Horizontally fusing the first knowledge graph and the target knowledge graphs among the found other knowledge graphs whose correlation degree is greater than a preset threshold.
[0131] Even further, the knowledge graph fusion module may specifically be used for:
[0132] For each of the target knowledge graphs, respectively determine the length of the horizontal connection line for horizontal fusion with the first knowledge graph according to their respective correlation degrees, and then draw horizontal connection lines of corresponding lengths with their respective associated graph nodes as connection points to complete the horizontal fusion with the first knowledge graph; wherein, the greater the correlation degree, the shorter the length of the corresponding horizontal connection line.
[0133] Even further, the knowledge graph fusion module may include:
[0134] A graph quantity statistics unit for counting the quantity of the target knowledge graphs;
[0135] A first graph fusion unit for horizontally fusing all the target knowledge graphs with the first knowledge graph if the quantity of the target knowledge graphs is less than or equal to a set quantity upper limit;
[0136] A second graph fusion unit for sorting the target knowledge graphs in descending order of the correlation degree if the quantity of the target knowledge graphs is greater than the quantity upper limit, and horizontally fusing the top quantity upper limit of the knowledge graphs among the target knowledge graphs with the first knowledge graph.
[0137] Further, the knowledge graph fusion module may include:
[0138] An associated node quantity statistical unit, configured to respectively count the quantity of the associated graph nodes each of the other knowledge graphs has;
[0139] A third graph fusion unit, configured to, for each of the other knowledge graphs, respectively determine the length of the horizontal connection line for horizontal fusion with the first knowledge graph according to the quantity of the associated graph nodes each has, and then draw a horizontal connection line with the corresponding length taking the associated graph nodes each has as connection points, to complete the horizontal fusion with the first knowledge graph; wherein, the more the quantity of the associated graph nodes is, the shorter the length of the corresponding horizontal connection line is.
[0140] Further, the material text acquisition module may include:
[0141] A task acquisition unit, configured to acquire a graph creation task issued by a user on a specified platform;
[0142] A knowledge field determination unit, configured to determine the knowledge field to which the knowledge graph to be created belongs according to the graph creation task;
[0143] A material text search unit, configured to search for the material text from the storage path corresponding to the knowledge field to which the knowledge graph to be created belongs.
[0144] The embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of any one of the knowledge graph creation methods as Figures 1 to 3 shown are implemented.
[0145] The embodiment of the present application further provides a server, including a memory, a processor, and computer-readable instructions stored in the memory and executable on the processor. When the processor executes the computer-readable instructions, the steps of any one of the knowledge graph creation methods as Figures 1 to 3 shown are implemented.
[0146] Figure 5 is a schematic diagram of a server provided by an embodiment of the present application. As Figure 5 shown, the server 5 of this embodiment includes: a processor 50, a memory 51, and computer-readable instructions 52 stored in the memory 51 and executable on the processor 50. When the processor 50 executes the computer-readable instructions 52, the steps in the above-mentioned embodiments of the knowledge graph creation method are implemented, such as Figure 1 the steps 101 to 105 shown. Or, when the processor 50 executes the computer-readable instructions 52, the functions of each module / unit in the above-mentioned device embodiments are implemented, such as Figure 4 the functions of the modules 401 to 405 shown.
[0147] Exemplarily, the computer-readable instructions 52 may be divided into one or more modules / units, which are stored in the memory 51 and executed by the processor 50 to complete this application. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer-readable instructions 52 in the server 5.
[0148] The server 5 may be a computing device such as a smart phone, a notebook, a palm computer, or a cloud server. The server 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 merely examples of the server 5, which do not constitute a limitation on the server 5. It may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the server 5 may further include input / output devices, network access devices, a bus, etc.
[0149] The processor 50 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.
[0150] The memory 51 may be an internal storage unit of the server 5, such as the hard disk or memory of the server 5. The memory 51 may also be an external storage device of the server 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the server 5. Further, the memory 51 may also include both the internal storage unit and the external storage device of the server 5. The memory 51 is used to store the computer-readable instructions and other programs and data required by the server 5. The memory 51 may also be used to temporarily store data that has been output or is to be output.
[0151] It should be noted that for the content such as information interaction and execution process between the above-mentioned devices / units, since they are based on the same concept as the method embodiments of the present application, for their specific functions and the technical effects brought about, reference can be specifically made to the method embodiment part, and details are not described herein again.
[0152] Those skilled in the art can clearly understand that for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiment, and details are not described herein again.
[0153] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, a computer program can be used to instruct the relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form. The computer-readable medium can at least include: any entity or device that can carry the computer program code to the photographing device / terminal device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.
[0154] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0155] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included within the protection scope of the present application.
Claims
1. A method for creating a knowledge graph, characterized in that, Including: Obtain the material text of the knowledge domain to which the knowledge graph to be created belongs; Parse the material text according to the pre-constructed schema of the knowledge graph knowledge layer, and extract instance element data; Fuse the extracted instance element data with the schema of the knowledge graph knowledge layer to obtain the first knowledge graph of the vertical domain; Search in the preset knowledge graph library for other knowledge graphs that have at least one associated graph node with the first knowledge graph; Using the associated graph nodes as connection points, horizontally fuse the first knowledge graph and the other knowledge graphs to obtain the created knowledge graph; Wherein, after searching in the preset knowledge graph library for other knowledge graphs that have at least one associated graph node with the first knowledge graph, it further includes: Determine the knowledge domain to which the first knowledge graph belongs, and respectively determine the knowledge domains to which each of the found other knowledge graphs belongs; Search in the pre-constructed knowledge domain comparison table for the association degrees between the knowledge domains to which each of the other knowledge graphs belongs and the knowledge domain to which the first knowledge graph belongs; The horizontally fusing the first knowledge graph and the other knowledge graphs includes: Horizontally fuse the first knowledge graph and the target knowledge graphs among the found other knowledge graphs whose association degrees are greater than a preset threshold; Respectively count the number of the associated graph nodes of each of the other knowledge graphs; For each of the other knowledge graphs, respectively determine the length of the horizontal connection line for horizontally fusing with the first knowledge graph according to the number of the associated graph nodes each has, and then use the associated graph nodes each has as connection points to draw horizontal connection lines of corresponding lengths to complete the horizontal fusion with the first knowledge graph; Wherein, the more the number of the associated graph nodes, the shorter the length of the corresponding horizontal connection line.
2. The method for creating a knowledge graph according to claim 1, characterized in that, The parsing the material text according to the pre-constructed schema of the knowledge graph knowledge layer and extracting instance element data includes: Detect the structured data of the material text to find out the structured data and unstructured data contained in the material text; Parse the structured data using a preset rule model to obtain the first instance element data contained therein; Parse the unstructured data using a pre-constructed NLP recognition model to obtain the second instance element data contained therein; Fuse the first instance element data and the second instance element data to obtain the extracted instance element data.
3. The method for creating a knowledge graph according to claim 1, characterized in that, The horizontally fusing the first knowledge graph and the target knowledge graphs among the found other knowledge graphs whose association degrees are greater than a preset threshold includes: For each of the target knowledge graphs, respectively determine the length of the horizontal connection line for horizontally fusing with the first knowledge graph according to their respective association degrees, and then use the associated graph nodes each has as connection points to draw horizontal connection lines of corresponding lengths to complete the horizontal fusion with the first knowledge graph; The greater the correlation degree is, the shorter the length of the corresponding horizontal connection line is.
4. The method for creating a knowledge graph according to claim 1, characterized in that, The horizontal fusion of the first knowledge graph and the target knowledge graphs whose correlation degree is greater than a preset threshold in each of the other found knowledge graphs includes: Counting the number of the target knowledge graphs; If the number of the target knowledge graphs is less than or equal to the set upper limit, all the target knowledge graphs are horizontally merged with the first knowledge graph; If the number of the target knowledge graphs is greater than the upper limit, the target knowledge graphs are sorted in descending order according to the degree of association, and the knowledge graphs within the upper limit that are ranked higher in the target knowledge graphs are horizontally merged with the first knowledge graph.
5. The method for creating a knowledge graph according to any one of claims 1 to 4, characterized in that, The method of obtaining the material text of the knowledge field to which the knowledge graph to be created belongs includes: Get the graph creation tasks issued by users on the specified platform; Determine the knowledge domain to which the knowledge graph to be created belongs according to the graph creation task; The material text is searched from the storage path corresponding to the knowledge field to which the knowledge graph to be created belongs.
6. A device for creating a knowledge graph, characterized in that, include: The material text acquisition module is used to obtain the material text of the knowledge field to which the knowledge graph to be created belongs; An instance element extraction module is used to parse the material text according to a pre-built graph knowledge layer schema to extract instance element data; A data fusion module, used to fuse the extracted instance element data with the graph knowledge layer schema to obtain a first knowledge graph in a vertical field; A knowledge graph search module, used to search for other knowledge graphs having at least one associated graph node with the first knowledge graph from a preset knowledge graph library; A knowledge graph fusion module, used to horizontally fuse the first knowledge graph with the other knowledge graphs using associated graph nodes as connection points to obtain a created knowledge graph; The knowledge graph creation device also includes: A knowledge domain determination module, used to determine the knowledge domain to which the first knowledge graph belongs, and to respectively determine the knowledge domain to which each other knowledge graph found belongs; A domain relevance search module, used to search for the relevance between the knowledge domain to which each of the other knowledge graphs belongs and the knowledge domain to which the first knowledge graph belongs from a pre-constructed knowledge domain comparison table; The knowledge graph fusion module is used to: Horizontally fusing the first knowledge graph with the target knowledge graphs whose correlation degree is greater than a preset threshold in each of the other knowledge graphs found; The knowledge graph fusion module includes: An associated node number counting unit, used to count the number of associated graph nodes of each of the other knowledge graphs; The third graph fusion unit is configured to, for each of the other knowledge graphs, respectively determine the lengths of the horizontal connection lines for horizontal fusion with the first knowledge graph according to the number of the associated graph nodes each has, and then draw horizontal connection lines of corresponding lengths with the associated graph nodes each has as connection points to complete the horizontal fusion with the first knowledge graph; wherein, the more the number of the associated graph nodes, the shorter the length of the corresponding horizontal connection line.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method for creating a knowledge graph according to any one of claims 1 to 5 are implemented.
8. A server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method for creating a knowledge graph according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Knowledge graph construction and query method based on power enterprise
CN110929042A
Automatic knowledge graph construction method for legal consultation and retrieval system thereof
CN111241299A
Knowledge graph construction method and device in field of oil and gas exploration and development
CN111475653A