A knowledge graph generation method and device, a storage medium, and an electronic device
By encapsulating the data processing steps into configurable tool nodes through a visual graph generation system, the problem of fragmented steps in knowledge graph construction is solved, enabling a flexible and efficient graph construction process and improving operational efficiency and system integration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ZHEJIANG LAB
- Filing Date
- 2025-05-12
- Publication Date
- 2026-05-29
AI Technical Summary
The fragmented data processing stages in the existing knowledge graph construction process result in high operational complexity and poor connectivity, making it difficult to meet the needs for flexible and efficient construction.
This paper provides a knowledge graph generation method. Through a visual graph generation system, data extraction, fusion, reasoning and other processes are encapsulated into configurable tool nodes. Users can customize the execution order and parameter configuration to achieve flexible assembly and automated processing of the construction process.
It reduces operational complexity, improves the efficiency of process connections, achieves flexibility and efficiency in map construction, and avoids manual intervention in cross-tool data conversion and verification.
Smart Images

Figure CN120542537B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of data processing technology, and in particular to a method, apparatus, storage medium, and electronic device for generating knowledge graphs. Background Technology
[0002] With the rapid development of big data and artificial intelligence technologies, knowledge graphs, as a structured semantic network, have become a core technology supporting fields such as intelligent search, recommendation systems, medical diagnosis, and financial risk control. Through the formalized expression of concepts, relationships, and attributes, knowledge graphs can efficiently integrate multi-source heterogeneous data, enabling knowledge association analysis and deep reasoning.
[0003] However, the construction of knowledge graphs usually involves multiple independent data processing steps. The overall process of existing graph construction is relatively fixed, and each data processing step needs to be implemented through different software tools or program modules. This divide-and-conquer model leads to frequent manual conversion and verification between each step, poor connectivity, and significantly increases the complexity of operation, making it difficult to meet the needs of flexible and efficient graph construction. Summary of the Invention
[0004] This specification provides a method, apparatus, storage medium, and electronic device for generating knowledge graphs, in order to partially solve the aforementioned problems existing in the prior art.
[0005] The following technical solution is adopted in this specification:
[0006] This specification provides a method for generating knowledge graphs, including:
[0007] The system receives a creation request for a knowledge graph generation system and displays a creation page for the knowledge graph generation system based on the creation request. The creation page contains multiple tool nodes, each of which is used to perform different data processing steps in the knowledge graph generation process.
[0008] Based on the user's operations on each tool node, each target tool node and its corresponding execution order are determined. Furthermore, based on the user's configuration operations on the specified target tool nodes, the graph configuration information of the knowledge graph to be generated is determined. Each target tool node includes: data nodes and non-data nodes, where the non-data nodes are used to process data provided by data nodes with which they have connections.
[0009] Based on the graph configuration information and the execution order of each target tool node, the graph generation system is created;
[0010] After receiving the data to be processed, the graph generation system processes the data to generate a knowledge graph based on the data.
[0011] Optionally, the tool node includes: a main tool node and sub-tool nodes;
[0012] Based on the user's operations on each tool node, determine each target tool node and its corresponding execution order, specifically including:
[0013] Based on the user's expansion operation on any main tool node, at least one sub-tool node under that main tool node is displayed;
[0014] Based on the drag-and-drop operation performed by the user on the at least one sub-tool node, the execution order of each target tool node and the corresponding execution order of each target tool node are determined; wherein, the drag-and-drop operation is used to drag the sub-tool node to the node canvas of the creation page, and the node canvas displays multiple target tool nodes connected in sequence.
[0015] Optionally, the data processing stage includes at least: a data import stage, a data cleaning stage, and a graph generation stage, wherein the graph generation stage includes: a knowledge definition sub-stage, a knowledge extraction sub-stage, and a graph output sub-stage.
[0016] Optionally, the specified target tool node includes: the tool node corresponding to the knowledge definition stage, and / or other target tool nodes located in the same pipeline branch as the tool node corresponding to the knowledge definition stage, and whose execution order is before the tool node corresponding to the knowledge definition stage.
[0017] Optionally, based on the configuration operation performed by the user on the specified target tool node, the graph configuration information of the knowledge graph to be generated is determined, specifically including:
[0018] Based on the user's selection operation on the other target tool nodes, display a configuration option overlay;
[0019] Based on the configuration options selected by the user in the configuration options overlay, a quick configuration window corresponding to those options is displayed; wherein, the configuration options include: concept configuration options and relationship configuration options. A drop-down list is provided for quick configuration of each parent concept and its child concepts.
[0020] Optionally, based on the configuration operation performed by the user on the specified target tool node, the graph configuration information of the knowledge graph to be generated is determined, specifically including:
[0021] Based on the configuration operation performed by the user on the specified target tool node, a configuration page for graph configuration information is displayed to the user; wherein, the configuration page is used to configure at least one of the concepts and relationships between concepts in the knowledge graph to be generated.
[0022] Optionally, the concept includes: a parent concept and a child concept. If two child concepts belong to two parent concepts with a relationship, then the relationship configuration between the two child concepts is the same as the relationship configuration between the parent concepts of the two child concepts.
[0023] Optionally, the atlas generation system includes multiple pipeline branches, which include a main branch and new branches;
[0024] After receiving the data to be processed, the graph generation system processes the data to generate a knowledge graph based on it, specifically including:
[0025] After the tool node of the main branch receives the initial data, it processes the initial data through the main branch of the graph generation system to obtain the initial knowledge graph.
[0026] After the tool node of the newly added branch receives the new data, it processes the new data through the newly added branch of the graph generation system, and updates the initial knowledge graph based on the processed new data to obtain the updated knowledge graph.
[0027] Optionally, the method further includes:
[0028] Based on the verification operations performed by the user on the tool nodes corresponding to any data processing stage, the data processing results of that data processing stage are displayed.
[0029] This specification provides a knowledge graph generation apparatus, including:
[0030] The receiving module is used to receive a creation request for the knowledge graph generation system and display the creation page of the knowledge graph generation system according to the creation request; wherein, the creation page is provided with multiple tool nodes, each of which is used to perform different data processing steps in the knowledge graph generation process;
[0031] The determination module is used to determine each target tool node and the execution order corresponding to each target tool node based on the user's operations on each tool node, and to determine the graph configuration information of the knowledge graph to be generated based on the configuration operations performed by the user on the specified target tool nodes; wherein, each target tool node includes: data nodes and non-data nodes, and the non-data nodes are used to process the data provided by the data nodes with which they have a connection;
[0032] A creation module is used to create the graph generation system based on the graph configuration information and the execution order of each target tool node;
[0033] A generation module is used to process the data to be processed through the graph generation system after receiving the data to be processed, so as to generate a knowledge graph based on the data to be processed. This specification provides a knowledge graph generation apparatus, including:
[0034] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for generating a knowledge graph.
[0035] This specification provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described method for generating a knowledge graph.
[0036] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects:
[0037] The knowledge graph generation method provided in this specification receives a creation request for a knowledge graph generation system and displays a creation page for the system based on the request. The creation page contains multiple tool nodes, each used to execute different data processing steps in the knowledge graph generation process. Based on the user's specified operations on each tool node, the system determines each target tool node and its corresponding execution order. Furthermore, based on the user's configuration operations on the specified target tool nodes, the system determines the knowledge graph configuration information to be generated. The system is then created based on the configuration information and the execution order of the target tool nodes. Upon receiving data to be processed, the system processes the data to generate a knowledge graph based on it.
[0038] As can be seen from the above methods, this application effectively solves the technical problems of fragmented processes, scattered tools, and frequent manual intervention in the traditional knowledge graph construction process by constructing a visual graph generation system. Specifically, by encapsulating independent processing steps such as data extraction, fusion, and reasoning into configurable tool nodes, and supporting user-defined execution order and parameter configuration, it achieves flexible assembly of the construction process, breaking through the limitations of existing fixed processes. All tool nodes are integrated into a unified system, eliminating manual intervention in cross-tool data conversion and verification, significantly improving the efficiency of process connection. The system automatically executes end-to-end processing based on configuration information, avoiding manual interaction between multiple software modules and reducing operational complexity to the level of visual configuration. This solution, through modular design and process automation, achieves a dual improvement in flexibility and efficiency while ensuring the professionalism of graph construction. Attached Figure Description
[0039] The accompanying drawings, which are included to provide a further understanding of this specification and form part of this specification, illustrate exemplary embodiments and are used to explain this specification, but do not constitute an undue limitation thereof. In the drawings:
[0040] Figure 1 This document provides a flowchart illustrating a method for generating a knowledge graph.
[0041] Figure 2 This is a schematic diagram of the creation page of a map generation system provided in this specification;
[0042] Figure 3 This is a schematic diagram of a conceptual configuration page provided in this specification;
[0043] Figure 4 This is a schematic diagram of a relationship configuration page provided in this specification;
[0044] Figure 5 This is a schematic diagram of a node configuration page provided in this specification;
[0045] Figure 6 This is a schematic diagram of a quick configuration window provided in this manual;
[0046] Figure 7 A schematic diagram of a knowledge graph generation device provided in this specification;
[0047] Figure 8 This specification provides a corresponding Figure 1 A schematic diagram of an electronic device. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this specification will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0049] Knowledge graphs are a semantically based method of knowledge representation, forming a relational network that connects different types of information. It's a graph-based data structure composed of nodes (points) and edges (edges). Each node represents an "entity," and each edge represents a "relationship" between entities. By structurally encoding and linking knowledge, it reduces information redundancy and increases query accuracy.
[0050] Knowledge graphs are typically defined by concepts and relation configurations. Concepts refer to the classification names of entities of the same type in the knowledge graph. For example, under the concept of "building," there might be related entities such as "school," "hospital," and "office building." Relation configurations define and describe the relationships between nodes in the knowledge graph. The purpose of relation configurations is to specify the types of relationships between nodes (such as parent-child relationships, synonym relationships, etc.) and relationship attributes (such as parent node, synonym similarity, etc.), aiming to provide a clear and standardized descriptive framework that allows different entities and applications to share and use the relationships established in the knowledge graph.
[0051] Knowledge graph construction refers to the process of converting data into triples based on different structured formats and methods, then fusing knowledge from these triples to ultimately create a knowledge graph. The steps involved in building a knowledge graph can be broadly categorized as: knowledge definition, knowledge extraction, and finally, the completion of the knowledge graph.
[0052] Most existing knowledge graph construction tools adopt a modular construction process, which divides the knowledge graph construction process into several large modules, requiring each module to be built independently. With the accumulation and application of big data, big data research and exploration based on knowledge graphs has begun to be widely used in various fields.
[0053] In reality, knowledge graphs often require continuous integration of new data to form iterative graphs, and the data sources in the graphs are often multimodal. For the construction of multimodal graphs, each modality of data often requires the development of a separate module to complete specific knowledge extraction, such as: text data entering the text recognition function area, and image data entering the object extraction function area.
[0054] Furthermore, current knowledge graph construction tools often place data analysis visualization and knowledge graph in different modules, preventing data analysis visualization and knowledge graph from being displayed in parallel or simultaneously.
[0055] Based on this, this manual provides a method for generating knowledge graphs that supports the execution of data processing steps such as data import, data cleaning, concept definition, relationship definition, knowledge extraction, and graph viewing within the same analysis flow. All steps of graph construction can be completed through interactive operations such as clicking on configuration and node connection.
[0056] The technical solutions provided in the various embodiments of this specification are described in detail below with reference to the accompanying drawings.
[0057] Figure 1 The flowchart of a knowledge graph generation method provided in this specification includes the following steps:
[0058] S101: Receive a creation request for the knowledge graph generation system, and display the creation page of the knowledge graph generation system according to the creation request; wherein, the creation page is provided with multiple tool nodes, each tool node is used to perform different data processing steps in the knowledge graph generation process.
[0059] In this specification, the execution subject for performing a knowledge graph generation method can be an application client or a cloud service client set on a terminal device such as a mobile phone, tablet computer, laptop computer, or desktop computer. Of course, it can also be a server. For ease of description, the following will only use the client as the execution subject to illustrate the knowledge graph generation method provided in this specification.
[0060] The client can receive a creation request for the map generation system and display the creation page of the map generation system according to the creation request.
[0061] In practical applications, the above creation request can be generated after the user opens or logs into the client, or it can be generated after the user selects the creation function of the atlas generation system provided by the client. This manual does not specifically limit this. For ease of understanding, this manual provides a schematic diagram of the creation page of the atlas generation system, such as... Figure 2 As shown.
[0062] Figure 2 This is a schematic diagram of the creation page of a map generation system provided in this specification.
[0063] The creation page includes a node canvas and multiple tool nodes. Each tool node is located in area ① and is used to perform different data processing steps in the knowledge graph generation process.
[0064] In this specification, the data processing stage includes at least the following: data import, data cleaning, and map generation. The aforementioned tool nodes may include a main node and child nodes under each main node.
[0065] The data import stage is the starting point of the data processing flow, used to import the multimodal data to be processed uploaded by users into the map generation system.
[0066] The main node corresponding to the data import stage can include multiple child nodes, each of which is used to import different types of data, such as... Figure 1 The structure data (node 10, node 1) and image data (node 4) are included. Of course, in practical applications, other types of data, such as audio data, may also be included as child nodes, but this specification does not specifically limit this.
[0067] The data cleaning process aims to correct data quality issues and ensure data accuracy, consistency, and integrity. Similarly, the main node for data cleaning can include multiple sub-nodes, each corresponding to a different data cleaning method, such as handling missing values, handling duplicate values, outlier detection, format standardization, and error correction.
[0068] The knowledge graph generation process includes three sub-processes: knowledge definition, knowledge extraction, and instance mapping. The sub-nodes for each sub-process are as follows: Figure 1 As shown in region ②.
[0069] Among them, the knowledge definition sub-step is the cornerstone of building a knowledge graph, used to define the semantic framework of the knowledge graph, including entity types (concepts), relationships and attributes.
[0070] The knowledge extraction sub-step is used to extract key information such as subjects, relationships, and events from unstructured or semi-structured data such as text, images, and videos, and map them onto concepts, relationships, and attributes in the knowledge graph.
[0071] The graph output sub-stage is used to output the knowledge graph and supports naming and previewing the knowledge graph.
[0072] Of course, in practical applications, other data processing steps such as Structured Query Language (SQL) injection and algorithm processing can also be included. The main node corresponding to the algorithm processing step can also include multiple child nodes, each of which corresponds to a different processing algorithm (such as node 5).
[0073] S102: Based on the user's operations on each tool node, determine each target tool node and the execution order corresponding to each target tool node, and based on the configuration operations performed by the user on the specified target tool node, determine the graph configuration information of the knowledge graph to be generated; wherein, each target tool node includes: data nodes and non-data nodes, and the non-data nodes are used to process the data provided by the data nodes with which they have a connection.
[0074] Specifically, the client can display at least one sub-tool node under any main tool node based on the user's expand operation on any main tool node, and then determine each target tool node and the execution order corresponding to each target tool node based on the user's drag operation on the at least one sub-tool node.
[0075] In this specification, the node canvas is used to construct the topology of each tool node in the graph generation system. It displays multiple target tool nodes connected in sequence. Users can drag and drop child tool nodes as target tool nodes onto the node canvas and establish connections with other existing target tool nodes. This allows the client to determine the target tool nodes in the graph generation system to be constructed and their corresponding execution order.
[0076] The target tool nodes can include: data nodes and non-data nodes. Data nodes are the tool nodes corresponding to the data import stage, mainly used for importing data. Non-data nodes are the main tool nodes or sub-tool nodes corresponding to other data processing stages besides the data import stage, used for data processing of the data provided by the data nodes with which they are connected.
[0077] by Figure 2 For example, nodes 1, 4, and 10 are data nodes, and the remaining nodes are non-data nodes. Since non-data nodes 2 and 3 are connected to data node 1, they are used to process the structured data uploaded by the user to data node 1. Similarly, since non-data nodes 5 and 6 are connected to data node 4, they are used to process the image data uploaded by the user to data node 4. And since non-data node 8 is connected to data nodes 1, 4, and 10, it can be used to process the data provided by data nodes 1, 4, and 10.
[0078] It should be further noted that if there is a connection between two non-data nodes, the preceding non-data node will input its processed data into the following non-data node, so that the following non-data node can process the processed data output by the preceding non-data node.
[0079] Furthermore, the client can determine the graph configuration information of the knowledge graph to be generated based on the configuration operations performed by the user on the specified target tool node.
[0080] In this specification, the target tool nodes specified above may include: the tool node corresponding to the knowledge definition sub-step, and / or other target tool nodes located in the same pipeline branch as the tool node corresponding to the knowledge definition sub-step, and whose execution order is prior to the tool node corresponding to the knowledge definition sub-step (e.g., Figure 1 (Nodes 1 to 6 in the data).
[0081] In this specification, the client displays a configuration page for the knowledge graph configuration information to the user based on the configuration operations performed by the user on the specified target tool node; wherein, the configuration page is used to configure at least one of the concepts, relationships between concepts, and graph nodes in the knowledge graph to be generated.
[0082] Specifically, the client can display a configuration page for the knowledge graph based on the configuration operations performed by the user on the tool nodes corresponding to the knowledge definition sub-step. This configuration page is used to configure at least one of the following: concepts, concept attributes, relationships between concepts, and relationship attributes in the knowledge graph to be generated. The configuration page is shown below. Figure 3 and Figure 4 As shown.
[0083] Figure 3 This is a schematic diagram of a conceptual configuration page provided in this specification.
[0084] Figure 3 The diagram shows the concept configuration panel for defining knowledge nodes. Area ③ is used to switch between the concept configuration page and the relationship configuration page. Area ④ in the concept configuration page is used to edit, add, and delete concepts. Area ⑤ is used to display entities under a concept. Area ⑥ is used to edit the attributes of a concept.
[0085] Figure 4 This is a schematic diagram of a relationship configuration page provided in this specification.
[0086] Figure 4 The panel shows the relationship configuration for defining knowledge nodes. Area ⑦ is used to switch between the concept configuration page and the relationship configuration page. Area ⑧ is used to edit, add, and delete relationships in the relationship tree. Area ⑨ is used to display a preview of the relationships between concepts. Area ⑩ is used to edit the attributes of the relationships.
[0087] Additionally, the client can display the node configuration page of the graph generation system to the user based on the configuration operations performed by the tool nodes corresponding to the knowledge extraction sub-stages, such as... Figure 5 As shown.
[0088] Figure 5 This is a schematic diagram of a node configuration page provided in this specification.
[0089] Figure 5 The node configuration page is displayed. The task configuration area is extracted. Used to add tasks. Each added task creates a new configuration item. The preceding item is the tool node, and the following item is the created concept or relationship. You can select from the dropdown list of the preceding item in the designated area. The system displays configurable tool nodes, which contain all tool nodes connected to the data flow steps preceding the knowledge node definition. It supports dropdown selection as well as node name search. Clicking the dropdown list of the following item allows you to select concepts or relationships, and select multi-level concepts or relationships through a tree structure.
[0090] The client can display a configuration options overlay based on the user's selections made on the other target tool nodes mentioned above. Then, based on the configuration options selected by the user in the overlay, a quick configuration window corresponding to those options will be displayed. These configuration options include concept configuration options and relationship configuration options, allowing users to quickly configure concepts and relationships within the quick configuration window. The quick configuration window's page structure is as follows: Figure 6 As shown.
[0091] Figure 6 This is a schematic diagram of a quick configuration window provided in this manual.
[0092] in, Figure 6 The document illustrates another convenient, targeted extraction configuration operation: by right-clicking any data node or processed data node in the same branch pipeline connected to the defined knowledge node, the client can further display a configuration options overlay. User in configuration options overlay After selecting the option that needs to be configured, the client can further display the corresponding quick configuration window. The first option in the drop-down selection list defaults to the data item of the current node, while the subsequent options are configured as single or multiple selections from the concept tree list or relationship tree list (i.e., for the same data, multiple fields can be configured with different concepts or relationships).
[0093] It should be noted that the concept of a knowledge graph can include parent concepts and child concepts. If two child concepts belong to two parent concepts with a relationship, then the relationship configuration between these two child concepts is the same as the relationship configuration between the parent concepts described in these two child concepts.
[0094] S103: Create the map generation system according to the map configuration information and the execution order of each target tool node.
[0095] S104: After receiving the data to be processed, the data to be processed is processed by the graph generation system to generate a knowledge graph based on the data to be processed.
[0096] After determining the target tool nodes in the graph generation system, the execution order of each target tool node, and the graph configuration information, the client can build the graph generation system based on this information.
[0097] Once the client receives the data to be processed, it can process it through the graph generation system. This allows the graph generation system to execute all data processing steps in the graph generation process, generating a knowledge graph based on the data to be processed.
[0098] In this specification, the atlas generation system includes multiple pipeline branches, including main branches and new branches;
[0099] After the initial data is received at the data import node of the main branch (the tool node corresponding to the data import process), the initial data can be processed through the main branch of the graph generation system to obtain the initial knowledge graph. After the new data is received at the data import node of the new branch, the new data can be processed through the new branch of the graph generation system, and the initial knowledge graph can be updated based on the processed new data to obtain the updated knowledge graph.
[0100] by Figure 1 For example, nodes 1, 2, 3, 8, and 9 form the main branch of a pipeline. When a user uploads structured data at node 1, node 8 extracts knowledge from the data processed by its preceding tool nodes and outputs the initial knowledge graph (graph version V1.0) based on the extracted knowledge through node 9.
[0101] When new structural data is added, users can upload the new data at node 10 and add newly defined knowledge at node 7. Then, knowledge is extracted from the data processed by the preceding tool nodes through node 12, and a new knowledge graph (graph version V1.1) is generated based on the new knowledge extracted by node 12 and the knowledge already extracted in node 8.
[0102] In addition, during the knowledge graph construction process, the client can display the data processing results of any data processing step based on the verification operations performed by the user on the tool nodes corresponding to any data processing step.
[0103] For example, users can click on the data cleaning node to check it, and then the client can show them the data after it has been cleaned.
[0104] In practical applications, the data to be processed can include domain data or global data from various fields such as medicine, aerospace, and finance. The client can perform tasks such as intelligent search, information recommendation, medical diagnosis, and financial risk control based on the generated knowledge graph.
[0105] As can be seen from the above methods, this solution can achieve the following effects:
[0106] The system allows for the tracing of knowledge graph data sources, rapid data replacement, modification, and updates, followed by automatic updates of the complete knowledge graph through a data analysis workflow. This avoids the tedious process of users frequently importing new data and deleting old data, achieving a visualized solution where the data analysis and knowledge graph construction processes use the same workflow.
[0107] It supports interactive operations that connect multiple analysis flow branches, enabling the connection of existing knowledge graphs with knowledge extraction nodes based on new data sources to complete knowledge expansion capabilities.
[0108] It satisfies the need for different roles to collaborate in completing tasks. Data processing and graph construction often require different roles such as data analysts and graph experts to complete the final graph. Data analysts can focus on data cleaning, algorithm application, etc., and import the prepared data in the data preparation module of the knowledge extraction node; while the core work of graph experts is to define the knowledge tree.
[0109] By adopting a visualization scheme based on analytical flow, the branch data streams connected to knowledge graph nodes are regarded as the reserve data of the graph. Knowledge extraction tasks can be quickly configured by right-clicking, which greatly reduces the efficiency of the graph.
[0110] The above describes one or more implementations of knowledge graph generation methods in this specification. Based on the same approach, this specification also provides corresponding knowledge graph generation devices, such as... Figure 7 As shown.
[0111] Figure 7 A schematic diagram of a knowledge graph generation device provided in this specification includes:
[0112] The receiving module 701 is used to receive a creation request for the knowledge graph generation system and display the creation page of the knowledge graph generation system according to the creation request; wherein, the creation page is provided with multiple tool nodes, each tool node is used to perform different data processing steps in the knowledge graph generation process;
[0113] The determination module 702 is used to determine each target tool node and the execution order corresponding to each target tool node based on the user's operations on each tool node, and to determine the graph configuration information of the knowledge graph to be generated based on the configuration operations performed by the user on the specified target tool nodes; wherein, each target tool node includes: data nodes and non-data nodes, and the non-data nodes are used to process the data provided by the data nodes with which they have a connection;
[0114] A creation module 703 is used to create the map generation system based on the map configuration information and the execution order of each target tool node;
[0115] The generation module 704 is used to process the data to be processed through the graph generation system after receiving the data to be processed, so as to generate the knowledge graph based on the data to be processed.
[0116] Optionally, the tool node includes: a main tool node and sub-tool nodes;
[0117] The determining module 702 is specifically used to: display at least one sub-tool node under the main tool node based on the expand operation performed by the user on any main tool node; and determine each target tool node and the execution order corresponding to each target tool node based on the drag operation performed by the user on the at least one sub-tool node; wherein the drag operation is used to drag the sub-tool node to the node canvas of the creation page, and the node canvas displays multiple target tool nodes connected in sequence.
[0118] Optionally, the data processing stage includes at least: a data import stage, a data cleaning stage, and a graph generation stage, wherein the graph generation stage includes: a knowledge definition sub-stage, a knowledge extraction sub-stage, and a graph output sub-stage.
[0119] Optionally, the specified target tool node includes: the tool node corresponding to the knowledge definition sub-step, and / or other target tool nodes located in the same pipeline branch as the tool node corresponding to the knowledge definition sub-step, and whose execution order is before the tool node corresponding to the knowledge definition sub-step.
[0120] Optionally, the determining module 702 is specifically used to: display a configuration option overlay based on the user's selection operation on the other target tool node; and display a shortcut configuration window corresponding to the configuration option selected by the user in the configuration option overlay; wherein the configuration option includes: concept configuration option and relationship configuration option.
[0121] Optionally, the determining module 702 is specifically used to display a configuration page of graph configuration information to the user based on the configuration operation performed by the user on the specified target tool node; wherein, the configuration page is used to configure at least one of the concepts, relationships between concepts, and graph nodes in the knowledge graph to be generated.
[0122] Optionally, the concept includes: a parent concept and a child concept. If two child concepts belong to two parent concepts with a relationship, then the relationship configuration between the two child concepts is the same as the relationship configuration between the parent concepts of the two child concepts.
[0123] Optionally, the atlas generation system includes multiple pipeline branches, which include a main branch and new branches;
[0124] The generation module 704 is specifically used to: after the tool node of the main branch receives the initial data, process the initial data through the main branch of the graph generation system to obtain an initial knowledge graph; after the tool node of the new branch receives the new data, process the new data through the new branch of the graph generation system, and update the initial knowledge graph based on the processed new data to obtain an updated knowledge graph.
[0125] Optionally, the generation module 704 is further configured to display the data processing results of any data processing stage based on the verification operation performed by the user on the tool node corresponding to any data processing stage.
[0126] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 This provides a method for generating knowledge graphs.
[0127] This instruction manual also provides Figure 8 One of the corresponding Figure 1 A schematic diagram of the structure of an electronic device. (e.g.) Figure 8 At the hardware level, the electronic device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for the business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1The method for generating the knowledge graph is described above. Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution subject of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0128] Improvements in a technology can be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many improvements to the methodology can now be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that an improvement in methodology cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program and "integrate" a digital system onto a PLD themselves, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog are the most commonly used. Those skilled in the art understand that by simply performing some logic programming on the method flow using one of these hardware description languages and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.
[0129] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0130] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0131] For ease of description, the above devices are described in terms of function, divided into various units. Of course, in implementing this specification, the functions of each unit can be implemented in one or more software and / or hardware.
[0132] Those skilled in the art will understand that embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0133] This specification is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this specification. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0134] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0135] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0136] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0137] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0138] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0139] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0140] Those skilled in the art will understand that the embodiments of this specification can be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this specification may take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0141] This specification can be described in the general context of computer-executable instructions that are executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This specification can also be practiced in distributed computing environments, where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0142] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0143] The above description is merely an embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of the claims of this specification.
Claims
1. A method for generating a knowledge graph, characterized in that, include: The system receives a creation request for a knowledge graph generation system and displays a creation page for the knowledge graph generation system based on the creation request. The creation page contains multiple tool nodes, each of which is used to perform different data processing steps in the knowledge graph generation process. Based on the user's operations on each tool node, each target tool node and its corresponding execution order are determined. Furthermore, based on the user's configuration operations on the specified target tool nodes, the graph configuration information of the knowledge graph to be generated is determined. Each target tool node includes: data nodes and non-data nodes. The non-data nodes are used to process data provided by data nodes with which they have connections. The data processing stage includes a graph generation stage, which includes a knowledge definition sub-stage for defining entity types, relationships, and attributes of the knowledge graph. The specified target tool nodes include: the tool node corresponding to the knowledge definition sub-stage, and other target tool nodes located in the same pipeline branch as the tool node corresponding to the knowledge definition sub-stage and whose execution order precedes that of the tool node corresponding to the knowledge definition sub-stage. Based on the graph configuration information and the execution order of each target tool node, the graph generation system is created; After receiving the data to be processed, the knowledge graph is processed by the graph generation system to generate the knowledge graph based on the data to be processed.
2. The method as described in claim 1, characterized in that, The tool nodes include: main tool nodes and sub-tool nodes; Based on the user's operations on each tool node, determine each target tool node and the execution order corresponding to each target tool node, specifically including: Based on the user's expansion operation on any main tool node, at least one sub-tool node under that main tool node is displayed; Based on the drag-and-drop operation performed by the user on the at least one sub-tool node, the target tool node and the execution order corresponding to the target tool node are determined; wherein, the drag-and-drop operation is used to drag the sub-tool node to the node canvas of the creation page, and the node canvas displays multiple target tool nodes connected in sequence.
3. The method as described in claim 1, characterized in that, The data processing stage includes at least: a data import stage, a data cleaning stage, and a graph generation stage. The graph generation stage includes: a knowledge definition sub-stage, a knowledge extraction sub-stage, and a graph output sub-stage.
4. The method as described in claim 3, characterized in that, Based on the configuration operations performed by the user on the specified target tool node, the graph configuration information of the knowledge graph to be generated is determined, specifically including: Based on the user's selection operation on the other target tool nodes, display a configuration option overlay; Based on the configuration options selected by the user in the configuration options overlay, a quick configuration window corresponding to the configuration options is displayed; wherein, the configuration options include: concept configuration options and relationship configuration options.
5. The method as described in claim 1, characterized in that, Based on the configuration operations performed by the user on the specified target tool node, the graph configuration information of the knowledge graph to be generated is determined, specifically including: Based on the configuration operation performed by the user on the specified target tool node, a configuration page for graph configuration information is displayed to the user; wherein, the configuration page is used to configure at least one of the concepts and relationships between concepts in the knowledge graph to be generated.
6. The method as described in claim 5, characterized in that, The concepts include: parent concepts and child concepts. If two child concepts belong to two parent concepts with a relationship, then the relationship configuration between these two child concepts is the same as the relationship configuration between the parent concepts described above.
7. The method as described in claim 1, characterized in that, The map generation system includes multiple pipeline branches, which include main branches and newly added branches; After receiving the data to be processed, the graph generation system processes the data to generate a knowledge graph based on it, specifically including: After the tool node of the main branch receives the initial data, it processes the initial data through the main branch of the graph generation system to obtain the initial knowledge graph. After the tool node of the newly added branch receives the new data, it processes the new data through the newly added branch of the graph generation system, and updates the initial knowledge graph based on the processed new data to obtain the updated knowledge graph.
8. The method as described in claim 4, characterized in that, The method further includes: Based on the verification operations performed by the user on the tool nodes corresponding to any data processing stage, the data processing results of that data processing stage are displayed.
9. A knowledge graph generation device, characterized in that, include: The receiving module is used to receive a creation request for the knowledge graph generation system and display the creation page of the knowledge graph generation system according to the creation request; wherein, the creation page is provided with multiple tool nodes, each of which is used to perform different data processing steps in the knowledge graph generation process; The determination module is used to determine each target tool node and its corresponding execution order based on the user's operations on each tool node, and to determine the graph configuration information of the knowledge graph to be generated based on the configuration operations performed by the user on the specified target tool nodes. Each target tool node includes: data nodes and non-data nodes. The non-data nodes are used to process data provided by data nodes with which they have connections. The data processing stage includes a graph generation stage, which includes a knowledge definition sub-stage for defining entity types, relationships, and attributes of the knowledge graph. The specified target tool nodes include: the tool node corresponding to the knowledge definition sub-stage, and other target tool nodes located in the same pipeline branch as the tool node corresponding to the knowledge definition sub-stage and whose execution order precedes that of the tool node corresponding to the knowledge definition sub-stage. A creation module is used to create the graph generation system based on the graph configuration information and the execution order of each target tool node; The generation module is used to process the data to be processed through the graph generation system after receiving the data to be processed, so as to generate the knowledge graph based on the data to be processed.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the method described in any one of claims 1 to 7.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method described in any one of claims 1 to 7.