Knowledge graph automatic construction system and method based on langgraph workflow
By combining multi-source heterogeneous data and large language models with the LangGraph workflow, the problems of accuracy and redundant entities in existing knowledge graph construction are solved, and high-quality knowledge graph construction is achieved.
Patent Information
- Application Number
- CN202511301744.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-12
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-09-12
AI Technical Summary
Existing knowledge graph construction methods rely on manual editing or rule-based approaches, and machine learning-based methods have accuracy issues when understanding complex contexts and implicit relationships, which may result in inaccurate or non-compliant content.
A system and methodology based on LangGraph workflow are adopted. Data of target entities are obtained through a multi-source heterogeneous data acquisition module. Knowledge extraction and graph construction are performed using a large language model. Entity attributes are enriched and relationships are supplemented by structured attribute data from Wikidata. Finally, entity fusion and disambiguation are performed to form a high-quality knowledge graph.
It significantly improves the knowledge density and accuracy of knowledge graphs, solves the problem of redundant entities caused by the diversity of entity names, and improves the quality of the constructed data.
Smart Images

Figure CN120781950B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of knowledge graph, and in particular to a knowledge graph automatic construction system and method based on LangGraph workflow. BACKGROUND
[0002] As a structured knowledge base, knowledge graph is used to describe concepts, entities and their relationships in the real world, and has been widely applied in search engines, intelligent question answering, recommendation systems and other fields. Extracting information from semi-structured or unstructured text such as Wikipedia to construct knowledge graph is an important research direction in knowledge engineering field. Early knowledge graph construction mainly relies on manual editing or rule-based methods. With the development of natural language processing (NLP) technology, methods based on machine learning have emerged, such as named entity recognition (NER), relation extraction (RE), etc. These methods usually require a large amount of labeled data, and have limited understanding ability for complex context and implicit relationships.
[0003] In recent years, large language models (LLM), such as models based on Transformer architecture (GPT series, BERT, etc.), have made significant progress in natural language understanding and generation. They can directly extract entities and relationships from text, and even generate structured output. This provides a new way for automated knowledge graph construction. However, the output of LLM is not always accurate and reliable, and the generated content may have factual errors or not conform to the preset pattern. SUMMARY
[0004] To solve the technical problems in the background art, the present application proposes a knowledge graph automatic construction system and method based on LangGraph workflow.
[0005] The knowledge graph automatic construction system and method based on LangGraph workflow proposed by the present application comprises:
[0006] A multi-source heterogeneous data acquisition module is used to acquire multi-source heterogeneous data of target entities according to a name list of target entities; wherein the multi-source heterogeneous data includes unstructured text data of Wikipedia and structured attribute data of Wikidata;
[0007] A knowledge extraction and graph construction module is used to perform knowledge extraction and graph construction on the unstructured text data of Wikipedia of target entities based on LangGraph workflow, to obtain a preliminary graph structure, and store it in a graph database;
[0008] Wikidata structured attribute integration module, configured to enrich entity attributes and supplement relationships in the preliminary graph structure in the graph database according to the structured attribute data of Wikidata of the target entity, to obtain an intermediate graph structure;
[0009] Entity fusion and disambiguation processing module, configured to perform entity fusion and disambiguation processing on the intermediate graph structure in the graph database, to obtain a final knowledge graph structure.
[0010] Preferably, the knowledge extraction and graph construction module comprises:
[0011] A document reading unit is configured to read unstructured text data of a Wikipedia page of the target entity from a cache according to a name list of the target entity;
[0012] A text chunking unit is configured to perform intelligent chunking operation on the unstructured text data of the Wikipedia page of each target entity, to obtain a plurality of text blocks of each target entity, and obtain a text block list of each target entity according to the plurality of text blocks of each target entity, and cache the text block list of each target entity;
[0013] An information extraction unit is configured to perform information extraction on each text block in the text block list of each target entity by using a preset large language model, and aggregate entity data and relationship data obtained by performing information extraction on all text blocks of each target entity to form an extraction result list of each target entity, and cache the extraction result list of each target entity;
[0014] An extraction post-processing unit is configured to perform post-processing on the entity data and relationship data in the extraction result list corresponding to each target entity, to obtain a candidate entity dictionary and a candidate relationship dictionary corresponding to each target entity;
[0015] A relationship review unit is configured to perform relationship review on each relationship in the candidate relationship dictionary corresponding to each target entity based on a preset large language model, to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity;
[0016] A graph data preparation unit is configured to convert the entity data and relationship data in the final entity dictionary and the final relationship dictionary corresponding to each target entity into a standard format required by a batch write operation of a graph database;
[0017] A graph writing unit is configured to construct a preliminary graph structure according to the converted entity data and relationship data in the standard format, and write the preliminary graph structure into the graph database.
[0018] Preferably, the entity data comprises entity names, entity types and descriptive attributes; and the relationship data comprises head entity names, tail entity names, relationship types and descriptive information of the relationships.
[0019] Preferably, the intelligent chunking operation process of the text chunking unit comprises:
[0020] The original text is preliminarily divided into a plurality of paragraphs by identifying the markers of the chapter titles as natural paragraph separators based on the inherent chapter structure of the Wikipedia article;
[0021] For the paragraphs whose lengths still exceed the preset maximum length threshold after the division, a sentence division logic is called to refine the long paragraphs into a plurality of text blocks composed of one or more continuous sentences according to standard sentence ending punctuation marks; and a filtering operation is performed on the text blocks to remove text blocks with lengths less than a preset minimum length threshold.
[0022] Preferably, the information extraction process comprises: performing information extraction on each text block in the list of text blocks of each target entity by using a preset large language model according to an extraction prompt; wherein the extraction prompt comprises: a task instruction, a constraint condition, a structured output format requirement and a high-quality few-shot example.
[0023] Preferably, the post-processing comprises data cleaning, relationship subject verification, data deduplication, graph structure consistency maintenance and entity pruning operation.
[0024] Preferably, each relationship in the candidate relationship dictionary corresponding to each target entity is subjected to relationship review based on the preset large language model to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity, specifically comprising:
[0025] Querying whether each relationship in the candidate relationship dictionary corresponding to each target entity has been reviewed and recorded in the cache;
[0026] If a certain relationship has been reviewed and recorded in the cache, the review result in the cache is directly adopted;
[0027] If a certain relationship has been reviewed but not recorded in the cache or not reviewed, a review prompt is constructed;
[0028] Performing relationship review on the relationship that has been reviewed but not recorded in the cache or not reviewed by using the preset large language model according to the review prompt to obtain a review response of the large language model;
[0029] Extracting the review response of the large language model to obtain a relationship review result; wherein the relationship review result comprises yes or no;
[0030] Based on the relationship review result, the candidate entity dictionary is filtered to retain the entities in the relationship whose relationship review result is yes, to obtain the final entity dictionary and the final relationship dictionary corresponding to each target entity.
[0031] Preferably, the review prompt includes: the role of the large language model, the format of the relationship triple to be reviewed, and the basis for judgment.
[0032] Preferably, the knowledge extraction and graph construction module further comprises an error judgment unit and an error processing unit.
[0033] The error judgment unit is used to check whether the error information field in the current state object is set after each unit is executed; if yes, a preset first routing instruction is returned to guide the workflow to the error processing unit; if not, a preset second routing instruction is returned to make the workflow flow to the next processing unit in a predetermined order.
[0034] The error processing unit is used to record detailed error information when an abnormal error is captured, and terminate the current processing flow.
[0035] Preferably, according to the structured attribute data of the target entity in Wikidata, the entity attribute enrichment and relationship supplement of the preliminary graph structure are performed in the graph database to obtain an intermediate graph structure, specifically including:
[0036] According to the structured attribute data of the target entity in Wikidata, the descriptive attributes of the target entity in the graph database are enriched.
[0037] According to the structured attribute data of the target entity in Wikidata and the preset mapping table, new relationships are supplemented on the graph database after the descriptive attribute enrichment of the entity.
[0038] Based on the graph database after supplementing new relationships, the preliminary graph structure is updated to obtain an intermediate graph structure.
[0039] Preferably, according to the structured attribute data of the target entity in Wikidata, the descriptive attributes of the target entity in the graph database are enriched, specifically including:
[0040] All JSON format attribute files under the specified path in the object cache are traversed, the entity name of the target entity in the JSON format attribute file and the graph database is updated to the expression in the standard label in the corresponding JSON format attribute file, and all the remaining descriptive attributes contained in the JSON format attribute file are added as new attributes of the target entity or used to update the same-named attribute already on the target entity.
[0041] Preferably, the mapping table comprises mapping the relational properties in the structured property data of Wikidata to the recommended types of the relationship types expected to be created in the graph database and the recommended types of the associated entities of each relationship.
[0042] Preferably, the new relationships are supplemented on the graph database enriched with the descriptive properties of the entities according to the structured property data of Wikidata of the target entity and the preset mapping table, specifically comprising:
[0043] The associated entities are merged or newly created in the graph database enriched with the descriptive properties of the entities according to the relational properties in the structured property data of Wikidata of the target entity; wherein, if the associated entities are newly created, the entity type and name of the newly created associated entities are set according to the preset mapping table; a relationship of the relationship type specified by the mapping table is created between the target entity and the merged or newly created associated entities.
[0044] Preferably, the entity fusion and disambiguation processing is performed on the intermediate graph structure in the graph database to obtain the final knowledge graph structure, specifically comprising:
[0045] Querying all target entities with alias properties in the graph database.
[0046] For each target entity with alias information found, parsing the alias properties to obtain an alias list comprising all known aliases of the target entity.
[0047] Traversing each alias in the alias list corresponding to each target entity, for each alias, determining whether there is an alias entity of the same type in the graph database whose entity name is exactly the same as the alias; if there is, copying or migrating all the attribute information possessed by the alias entity to the target entity, and redirecting all the relationships pointing to or originating from the alias entity to the target entity, and completely removing the successfully merged alias entity and all its associated relationships from the graph database.
[0048] Updating the intermediate graph structure based on the graph database to obtain the final knowledge graph structure.
[0049] Preferably, it further comprises:
[0050] An iterative graph expansion support module for querying whether there is an entity in the final knowledge graph structure that already exists but whose key structured information has not been completely supplemented, if not, outputting the knowledge graph structure, if yes, finding the entity that already exists but whose key structured information has not been completely supplemented, and forming a new to-be-crawled entity text file according to the entity that already exists but whose key structured information has not been completely supplemented.
[0051] Preferably, the iterative graph expansion support module is further configured to feed back the new entity text file to be crawled as a name list of the target entity to the multi-source heterogeneous data acquisition module.
[0052] In a second aspect, the present application further provides a LangGraph workflow-based automatic knowledge graph construction method, comprising:
[0053] According to the name list of the target entity, multi-source heterogeneous data of the target entity is acquired; wherein the multi-source heterogeneous data comprises unstructured text data of Wikipedia and structured attribute data of Wikidata;
[0054] Based on the LangGraph workflow, knowledge extraction and graph construction are performed on the unstructured text data of Wikipedia of the target entity, to obtain a preliminary graph structure, which is stored in a graph database;
[0055] According to the structured attribute data of Wikidata of the target entity, entity attribute enrichment and relationship supplement are performed on the preliminary graph structure in the graph database, to obtain an intermediate graph structure;
[0056] Entity fusion and disambiguation processing are performed on the intermediate graph structure in the graph database, to obtain a final knowledge graph structure.
[0057] In the present application, the LangGraph workflow-based automatic knowledge graph construction system and method proposed in the present application acquire multi-source heterogeneous data of the target entity in the name list of the target entity through the multi-source heterogeneous data acquisition module, and perform knowledge extraction and graph construction on the unstructured text data of Wikipedia of the target entity based on the LangGraph workflow through the knowledge extraction and graph construction module, i.e., a stateful, controllable and robust knowledge extraction and preliminary graph construction workflow based on the LangGraph framework is constructed and driven to obtain a preliminary graph structure with higher data quality; the Wikidata structured attribute integration module is used to enrich entity attributes and supplement relationships of the preliminary graph structure, to obtain an intermediate graph structure with more comprehensive and accurate knowledge coverage, thereby significantly improving the knowledge density and accuracy of the knowledge graph; the entity fusion and disambiguation processing module solves the problem of multiple redundant entities in the graph that may be caused by the diversity of entity names, and further improves the data quality of the constructed knowledge graph structure. BRIEF DESCRIPTION OF DRAWINGS
[0058] Figure 1 The structure diagram of the LangGraph workflow-based automatic knowledge graph construction system in an embodiment of the present application is shown.
[0059] Figure 2A flowchart of an automatic construction method of a knowledge graph based on a LangGraph workflow in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0060] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict. The present application will be described in detail below with reference to the accompanying drawings and in combination with embodiments.
[0061] In a first aspect, as Figure 1 shown, the present application proposes an automatic construction system of a knowledge graph based on a LangGraph workflow, comprising:
[0062] A multi-source heterogeneous data acquisition module is configured to acquire multi-source heterogeneous data of a target entity according to a name list of the target entity, wherein the multi-source heterogeneous data comprises unstructured text data of Wikipedia and structured attribute data of Wikidata;
[0063] A knowledge extraction and graph construction module is configured to perform knowledge extraction and graph construction on the unstructured text data of Wikipedia of the target entity based on a LangGraph workflow, to obtain a preliminary graph structure, and store the preliminary graph structure in a graph database;
[0064] A Wikidata structured attribute integration module is configured to enrich entity attributes and supplement relationships of the preliminary graph structure in the graph database according to the structured attribute data of Wikidata of the target entity, to obtain an intermediate graph structure;
[0065] An entity fusion and disambiguation processing module is configured to perform entity fusion and disambiguation processing on the intermediate graph structure in the graph database, to obtain a final knowledge graph structure.
[0066] The multi-source heterogeneous data acquisition module acquires multi-source heterogeneous data of a target entity in a name list of the target entity, and the knowledge extraction and graph construction module performs knowledge extraction and graph construction on unstructured text data of Wikipedia of the target entity based on a LangGraph framework, forming a stateful, controllable and robust knowledge extraction and preliminary graph construction workflow, which improves the data quality of the obtained preliminary graph structure; the Wikidata structured attribute integration module enriches entity attributes and supplements relationships of the preliminary graph structure, to obtain an intermediate graph structure with more comprehensive and accurate knowledge coverage, thereby significantly improving the knowledge density and accuracy of the knowledge graph; the entity fusion and disambiguation processing module solves the problem of multiple redundant entities in the graph that may be caused by the diversity of entity names, further improving the data quality of the constructed knowledge graph structure.
[0067] Of course, the embodiment also includes a cache module for caching multi-source heterogeneous data, such as object storage, to facilitate subsequent multi-source heterogeneous data retrieval.
[0068] In the embodiment, the multi-source heterogeneous data acquisition module includes a first acquisition submodule and a second acquisition submodule. The first acquisition submodule is configured to acquire unstructured text data of the target entity from Wikipedia, and the second acquisition submodule is configured to acquire structured attribute data of the target entity from Wikidata.
[0069] The first acquisition submodule is configured to receive a list of names of target entities, and automatically query and download the pure text content of the corresponding Wikipedia page for each name of the target entity in the list using the public Wikipedia query interface provided by LangChain, such as WikipediaQueryRun and WikipediaAPIWrapper, and cache the pure text content of the Wikipedia page corresponding to each target entity as unstructured text data in the form of an independent text file. For example, the unstructured text data is cached in an object storage for subsequent knowledge extraction processes. The object storage is a MinIO object storage.
[0070] To ensure the accuracy of the data and the robustness of the processing, the first acquisition submodule is further configured to match and verify the title of the Wikipedia page with the name of the target entity when automatically querying the corresponding Wikipedia page for each name of the target entity, to avoid data misplacement.
[0071] The structured attribute information includes the ID (Wikidata ID) of the target entity in the Wikidata knowledge base, as well as standard labels (Label), brief descriptions (Description), commonly used aliases (Aliases), birth time, gender, industry, position, and a series of relational attributes such as educational background, work unit, and family member relationships.
[0072] The structured attribute data of each target entity from Wikidata is cached as an independent JSON format attribute file, and the WIKIDATA_ID of the target entity serves as the unique identifier for the entity. The file name of the JSON format attribute file of the target entity corresponds to the name of the target entity.
[0073] The second obtaining sub-module is configured to receive a name list of target entities, and automatically query and obtain structured attribute information of each target entity in the name list in a Wikidata knowledge base by using a public Wikidata query interface provided by LangChain, such as WikidataQueryRun and WikidataAPIWrapper; and perform data analysis on the Wikidata structured attribute information of each target entity to extract Wikidata structured attribute data therefrom and cache the Wikidata structured attribute data.
[0074] The workflow of the knowledge extraction and graph construction process of the knowledge extraction and graph construction module in the embodiment includes document reading, text chunking, information extraction, post-extraction processing, relationship review, graph data preparation, and graph writing.
[0075] The knowledge extraction and graph construction module in the embodiment includes:
[0076] The document reading unit is configured to read unstructured text data of a Wikipedia page of a target entity from the cache according to a name list of the target entity.
[0077] The text chunking unit is configured to perform intelligent chunking on the unstructured text data of the Wikipedia page of each target entity to obtain a plurality of text blocks of each target entity, obtain a text block list of each target entity according to the plurality of text blocks of each target entity, and cache the text block list of each target entity.
[0078] The information extraction unit is configured to perform information extraction on each text block in the text block list of each target entity by using a preset large language model, aggregate entity data and relationship data obtained by performing information extraction on all text blocks of each target entity to form an extraction result list of each target entity, and cache the extraction result list of each target entity.
[0079] The post-extraction processing unit is configured to perform post-processing on the entity data and relationship data in the extraction result list corresponding to each target entity to obtain a candidate entity dictionary and a candidate relationship dictionary corresponding to each target entity; the post-processing includes data cleaning, relationship subject verification, data deduplication, graph structure consistency maintenance, and entity pruning.
[0080] The relationship review unit is configured to perform relationship review on each relationship in the candidate relationship dictionary corresponding to each target entity based on a preset large language model to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity.
[0081] The graph data preparation unit is configured to convert the entity data and relationship data of the final entity dictionary and the final relationship dictionary corresponding to each target entity into a standard format required by a batch write operation of a graph database.
[0082] The graph writing unit is configured to construct a preliminary graph structure according to the entity data and relationship data converted into the standard format, and write the preliminary graph structure into the graph database.
[0083] In this embodiment, the document reading unit reads the plain text content of the Wikipedia page of the specified target entity from the object storage, and the text chunking unit performs intelligent chunking operation on the unstructured text data of the Wikipedia page of the specified target entity in the cache to divide the possibly very lengthy Wikipedia article into a series of text blocks with moderate length and relatively complete semantics, so as to adapt to the input length limit of the large language model, while retaining as much context information as possible to facilitate accurate extraction. A text block list of each target entity is obtained according to a plurality of text blocks of each target entity, and the text block list of each target entity is cached. The information extraction unit accurately and efficiently extracts information from each text block in the text block list of each target entity in the cache by using a preset large language model, and the extracted original entities and relationship data from all text blocks of each target entity are summarized to form an extraction result list of each target entity, and the extraction result list of each target entity is cached. The preliminary extraction results in the extraction result list are processed by the post-extraction processing unit to filter out data that may still contain noise or do not fully meet the requirements output by the preset large language model. The relationship review unit performs relationship review based on the preset large language model to further improve the accuracy and reliability of the extracted knowledge. The graph data preparation unit converts the entity data and relationship data of the final entity dictionary and the final relationship dictionary into a standard format required by a batch write operation of a graph database. The graph writing unit constructs a preliminary graph structure according to the entity data and relationship data converted into the standard format, and writes the preliminary graph structure into the graph database.
[0084] In one specific embodiment, the graph database is a Neo4j graph database.
[0085] To cut the potentially very long Wikipedia article into a series of moderate length and relatively complete semantic text blocks to adapt to the input length limit of large language models, while retaining as much context information as possible to facilitate accurate extraction, in further embodiments, the intelligent blocking operation process of the text blocking unit includes: first, using the inherent chapter structure of the Wikipedia article, by identifying the markers of chapter titles as natural paragraph separators, the text is preliminarily cut into paragraphs; then, for the paragraphs whose length still exceeds the preset maximum length threshold after cutting, call the sentence cutting logic, and according to the standard sentence ending punctuation (such as period, question mark, exclamation mark) to refine the long paragraphs into smaller text blocks composed of one or more continuous sentences; after the blocking is completed, a filtering operation is performed to remove those too short text blocks whose length is less than the preset minimum length threshold, to reduce the potential noise data interference to the subsequent extraction process. After this series of processing by the text blocking unit, the original long Wikipedia text is converted into a series of well-structured and appropriately lengthened text blocks, which form a text block list and can be used as the input unit for the subsequent large language model-driven information extraction unit.
[0086] To guide the preset large language model to accurately and efficiently complete the task, the information extraction unit in this embodiment adopts a precise prompt engineering strategy. In the prompt engineering strategy, the extraction prompt constructed for each text block in the text block list includes: 1) clear task instructions, informing the large language model that its current role is to assist in knowledge graph construction and is processing the Wikipedia page of a specific target entity; 2) strict constraint conditions, limiting the types of entities and relationships that can be extracted, and especially emphasizing that the extracted relationships must be directly related to the current target entity; 3) structured output format requirements, using the Pydantic data model to constrain the output of the large language model to conform to a specific JSON structure, and using the Langchain framework to provide a formatted output function to bind this structured requirement to the large language model call, thereby ensuring the normativity and ease of parsing of the output; 4) high-quality few-shot examples, by providing input text fragments and their corresponding expected extraction result samples, to help the large language model better understand the task requirements and output style. Among them, the entity data and relationship data obtained in the information extraction process of each text block using the preset large language model conform to the predetermined Pydantic model.
[0087] In order to filter out the data of the original output of the large language model that may still contain noise or not fully meet the requirements, in the post-processing process, through data cleaning, entries that do not meet the predefined entity type or relationship type or lack necessary fields such as entity name and relationship type are filtered out, and through relationship subject verification, it is strictly ensured that the source entity in all extracted relationships must be the core target currently being processed, and the entity list is de-duplicated based on entity names, and the relationship list is de-duplicated based on the triplets composed of the source entity, the target entity and the relationship type; and through graph structure consistency maintenance, only the relationships in which the source entity and the target entity involved are present in the current identified valid entity list are retained, avoiding hanging relationships; finally, through entity pruning, entities that are not referenced by any valid relationship after the above processing are removed. In this way, the initial entity dictionary and the initial relationship dictionary that are preliminarily sorted and purified can be obtained through the extraction post-processing unit.
[0088] In order to further improve the accuracy and reliability of the extracted knowledge, in the embodiment, the relationship review based on the preset large language model is performed on each relationship in the candidate relationship dictionary corresponding to each target entity, to obtain the final entity dictionary and the final relationship dictionary corresponding to each target entity, which specifically includes:
[0089] querying whether each relationship in the candidate relationship dictionary corresponding to each target entity has been reviewed and recorded in the cache; if a certain relationship has been reviewed and recorded in the cache, the review result in the cache is directly adopted;
[0090] if a certain relationship has been reviewed but not recorded in the cache or not reviewed, a review prompt is constructed;
[0091] using the preset large language model to perform relationship review on the relationship that has been reviewed but not recorded in the cache or not reviewed according to the review prompt, to obtain a review response of the large language model;
[0092] extracting the review response of the large language model to obtain a relationship review result; wherein the relationship review result includes "yes" or "no";
[0093] based on the relationship review result, filtering the candidate entity dictionary to retain the entities in the relationship whose relationship review result is "yes", to obtain the final entity dictionary and the final relationship dictionary corresponding to each target entity.
[0094] The relationship review unit introduced by the embodiment utilizes the relationship review mechanism based on the preset large language model, and this relationship review unit strictly reviews each relationship in the relationship dictionary output by the extraction post-processing node.
[0095] In the review process, first, the relationship review unit implements a relationship review cache mechanism. Before reviewing the relationship by the large language model, the cache will be queried first. If a relationship has been reviewed and recorded in the cache, the verification result in the cache will be directly adopted, thereby avoiding repeated large language model calls for the same relationship and significantly improving processing efficiency.
[0096] For relationships not recorded in the cache or not reviewed, the relationship review unit will construct a more rigorous and detailed review prompt. The large language model is used to review the relationship according to the review prompt, and the large language model review response is obtained. A special result analysis function is called to accurately extract the final "yes" or "no" judgment from the large language model review response that may contain the large language model thinking process, that is, the large language model relationship review result is obtained. Based on the large language model relationship review result, the final relationship dictionary is obtained. Based on the large language model relationship review result, the relationship review unit will filter the candidate entity dictionary again, and only keep the entities that appear in the valid relationship judged by the large language model as "yes". Finally, the relationship review unit outputs the final entity dictionary and the final relationship dictionary strictly reviewed by the large language model.
[0097] The review prompt includes: setting the role of the large language model as a knowledge graph relationship audit expert; clearly informing the large language model of the format and "judgment basis" of the input relationship triple to be reviewed, where the "judgment basis" is the descriptive text of the relationship given by the large language model during information extraction; providing a very detailed, phased review process and judgment standard, for example: the first stage checks the compliance of the head and tail entities in the relationship triple (such as whether the entity type meets the preset range, whether the entity name is specific and not general, etc.), the second stage deeply reviews the logical rationality of the relationship itself (such as whether the behavior relied on by the relationship is clear and matches the relationship type, whether the relationship reflects a certain degree of continuity rather than an occasional event, whether the subject and object of the relationship are accurately corresponding without generalization or attribution errors, etc.); providing several detailed analysis examples and their judgment logic on both positive and negative aspects to help the large language model accurately grasp the review scale; and strictly requiring the final output of the large language model to be only a single Chinese character "yes" (indicating that the verification is passed) or "no" (indicating that the verification is not passed).
[0098] In the graph data preparation process, the entity data includes the entity name and the entity type; the relationship data includes the head entity name, the tail entity name, the relationship type, and the description information of the relationship.
[0099] In one of the specific graph writing processes, the graph writing unit instantiates a Neo4j writer and calls its bulk write method to build a preliminary graph structure from the entity data and relationship data converted into a standard format and write it into the Neo4j graph database. During the writing process, the MERGE statement of the Cypher query language is usually used to create entities and relationships, which ensures that the graph writing unit is based on the uniqueness of its ID (if the entity already exists, it will not be created repeatedly, and only the attributes can be updated), and on this basis, the relationships between them are created.
[0100] In further embodiments, the knowledge extraction and graph construction module further comprises:
[0101] An error judgment unit is configured to call a condition judgment function to check whether an error information field in the current state object is set after each unit is executed; if yes, a preset first routing instruction is returned by the condition judgment function to guide the workflow to the error handling unit; if no, a preset second routing instruction is returned by the condition judgment function to make the workflow flow to the next normal processing unit in a predetermined order.
[0102] An error handling unit is configured to record detailed error information when an abnormal error is captured, and terminate the current processing flow.
[0103] It should be understood that the error handling unit is a unified error handling node. In any processing step of the LangGraph workflow, if an abnormal error is captured, the process will be directed to the error handling unit. The error handling unit is responsible for recording detailed error information (including the node where the error occurred, the error type and the specific content) to the log system, and will usually terminate the processing flow of the target entity to avoid error propagation.
[0104] The entire LangGraph workflow in this embodiment operates around a centralized state object GraphState. The state object encapsulates the data and control information of each stage in the process, including the storage object name currently being processed, the target entity name, the original text content, the text block list, the intermediate data extracted at each step, the entity data and relationship data ready to be written into the graph database, and the error information and error source node identifier for error tracking. This design makes the data transfer in the process clear and traceable, and facilitates the implementation of complex conditional routing.
[0105] In the LangGraph workflow, nodes and connections between nodes are defined, where a node is a processing unit in the workflow, and each node is a function or a Langchain executable object for performing a preset task.
[0106] In order to significantly improve the knowledge density and accuracy of the graph, in the embodiment, according to the structured attribute data of the target entity in Wikidata, the entity attribute enrichment and relationship supplement are performed on the preliminary graph structure in the graph database to obtain the intermediate graph structure, which specifically includes: enriching the descriptive attributes of the target entity in the graph database according to the structured attribute data of the target entity in Wikidata; supplementing new relationships on the graph database after the enrichment of the descriptive attributes of the entity according to the structured attribute data of the target entity in Wikidata and the preset mapping table; updating the preliminary graph structure based on the graph database after supplementing new relationships to obtain the intermediate graph structure.
[0107] Wherein, the descriptive attributes of the entity in the Neo4j graph database are enriched according to the structured attribute data of the target entity in Wikidata, which specifically includes: traversing all JSON format attribute files under the specified path in the cache, updating the entity name in the graph database to the expression in the standard label in the corresponding JSON format attribute file, and adding all the remaining descriptive attributes contained in the JSON format attribute file as new attributes of the entity or updating the existing same-named attributes on the entity.
[0108] Wherein, the mapping table includes mapping the relational attribute name in the structured attribute data of the target entity in Wikidata in the intermediate graph structure to the relationship type expected to be created in the graph database and the recommended type of the associated entity of each relationship.
[0109] Wherein, new relationships are supplemented on the graph database after the enrichment of the descriptive attributes of the entity according to the structured attribute data of the target entity in Wikidata and the preset mapping table, which specifically includes: merging or newly creating associated entities in the graph database after the enrichment of the descriptive attributes of the entity according to the relational attributes in the structured attribute data of the target entity in Wikidata; wherein, if the associated entity is newly created, the entity type and name of the newly created associated entity are set according to the preset mapping table; a relationship of the relationship type specified by the mapping table is created between the target entity and the merged or newly created associated entity.
[0110] In implementation, the input of the Wikidata structured attribute integration module is the JSON format attribute file of each target entity generated by the aforementioned Wikidata structured attribute data crawling submodule and cached in the object storage. The core processing logic of the Wikidata structured attribute integration module is to traverse all the JSON format attribute files under the specified path in the object storage, update the name of the target entity in the graph database to a more standardized expression in the standard label in the corresponding JSON format attribute file. At the same time, all other descriptive attributes contained in the JSON format attribute file will be added as new attributes of the entity, or used to update the same-named attributes already existing on the entity, so as to enrich the descriptive attributes of each entity; after the entity attributes are enriched, the Wikidata structured attribute integration module will also create new edge relationships in the graph based on the relationship attributes explicitly declared in the Wikidata structured attribute data in the JSON format attribute file. The Wikidata structured attribute integration module internally predefines a mapping table that maps the relationship attribute names (for example, educational experience, kinship, etc.) in the Wikidata structured attribute data to the relationship types expected to be created in the graph database and the recommended types of the associated entities of the relationship. For each target entity, the relationship supplement unit will attempt to merge or create the associated entity in the graph database, and if the associated entity is newly created, it will be assigned a corresponding entity type and name according to the mapping table. Finally, a relationship of the type specified by the mapping table is created between the target entity and the associated entity determined or newly created.
[0111] Through the processing of the Wikidata structured attribute integration module described above, the attributes of the entities in the graph database that were originally mainly based on text extraction are greatly enriched and standardized by the authoritative Wikidata structured attribute data, and the edge set of the Neo4j graph database is also effectively expanded based on the explicit relationship declarations in the Wikidata structured attribute data, thereby obtaining an intermediate graph structure with more comprehensive and accurate knowledge coverage.
[0112] In order to further improve the data quality of the graph, especially to solve the problem of multiple redundant entities pointing to the same object in the graph due to the diversity of entity names, in this embodiment, the intermediate graph structure is subjected to entity fusion and disambiguation processing to obtain a final knowledge graph structure, which specifically includes:
[0113] Query all entities with alias attributes in the graph database;
[0114] For each target entity with alias information found, parse its alias attribute to obtain an alias list including all known aliases of the target entity; traverse each alias in the alias list, and for each alias, determine whether there is an alias entity of the same type in the graph database whose entity name is exactly the same as the name of the target entity; if so, copy or migrate all attribute information owned by the alias entity to the target entity, and redirect all relationships pointing to or originating from the alias entity to the target entity, and completely remove the alias entity and all its associated relationships from the graph database after successful merging;
[0115] Update the intermediate graph structure based on the graph database to construct the final knowledge graph structure.
[0116] In specific implementation, the input of the entity fusion and disambiguation processing module is the intermediate graph structure stored in the Neo4j graph database and enhanced by Wikidata attributes, which particularly focuses on the alias attributes already existing on the entity and derived from Wikidata. The core processing logic of the entity fusion and disambiguation processing module is as follows: first, query all target type entities with alias attributes in the Neo4j graph database, and for each target entity with alias information found, parse its alias attribute string to obtain a list containing all known aliases of the target entity. Subsequently, the entity fusion and disambiguation processing module traverses each alias string in the alias list. For each alias, it attempts to accurately find whether there is an alias entity of the same type in the Neo4j graph database whose name is exactly the same as the name of the target entity. If a successful alias entity is found and it is confirmed that the alias entity is not the same entity object as the target entity being processed in the database, it is determined that the two entities may point to the same real-world entity, and an entity merging operation is triggered to copy or migrate all attribute information owned by the alias entity to the target entity, and redirect all relationships pointing to or originating from the alias entity to the target entity by using the APOC extension library. Finally, the successfully merged alias entity and all its associated relationships are completely removed from the graph database using the Cypher delete command.
[0117] By performing the above entity fusion and disambiguation processing on the entire graph database, redundant entities that may be introduced in the initial stage of graph construction due to entity reference diversity (such as full name, commonly used name, nickname, etc.) can be effectively reduced, thereby significantly improving the uniqueness and standardization of entity representation in the graph, and further improving the overall data quality, density, and usability of the knowledge graph.
[0118] To support the continuous dynamic construction of the knowledge graph, the gradual discovery of knowledge, and the continuous improvement and completion of the existing graph information, the embodiment further includes an iterative graph expansion support module configured to query whether there is an entity in the final knowledge graph structure that already exists but whose key structured information has not been completely supplemented; if not, output the knowledge graph structure; if yes, find the entity that already exists but whose key structured information has not been completely supplemented, and form a new entity text file to be crawled according to the entity that already exists but whose key structured information has not been completely supplemented.
[0119] It should be understood that the entity text file to be crawled contains the names of all the entities that have been judged and screened out and whose information is not complete, which can serve as a name list of new target entities, so as to facilitate the acquisition of multi-source heterogeneous data, knowledge extraction and graph construction, entity attribute enrichment and relationship supplement, and entity fusion and disambiguation processing, thereby facilitating the subsequent completion of the node information and the expansion of the graph knowledge.
[0120] In specific implementation, the input of the iterative graph expansion support module is the final knowledge graph structure stored in the Neo4j graph database. The core processing logic is to execute a specific Cypher query in the Neo4j database, and the goal of the query is to find and identify entities that already exist in the graph but whose key structured information has not been completely supplemented. A typical judgment criterion is to find entities whose WIKIDATA_ID attribute is empty or missing. The existence of these entities indicates that although they have been identified as part of the graph, their detailed background information has not been systematically crawled and integrated into the graph. The output of the iterative graph expansion support module is an entity text file to be crawled that records entities to be crawled, which contains the names of all the entities that have been filtered out by the query and whose information is not complete, and it can directly serve as the input of the aforementioned “multi-source heterogeneous data acquisition module” in the system for the next round of Wikipedia text data crawling and Wikidata structured attribute data crawling. By re-feeding these “to-be-supplemented information” entities into the data acquisition process, the system can specifically crawl the missing unstructured text data of Wikipedia and structured attribute data of Wikidata for these entities, and then the newly acquired unstructured text data of Wikipedia and structured attribute data of Wikidata will enter the LangGraph core knowledge extraction process and the Wikidata attribute integration process, thereby achieving the completion of the entity information and the iterative expansion of the graph knowledge.
[0121] In a further embodiment, the iterative graph expansion support module is further configured to feed the to-be-crawled entity text files as a list of names of target entities to the multi-source heterogeneous data acquisition module.
[0122] The embodiment can periodically scan the Neo4j graph database through the iterative graph expansion support module, find and identify entities that already exist in the knowledge graph structure but whose key structured information has not been fully supplemented, automatically generate a new list of to-be-processed entities, and feed the list to the multi-source heterogeneous data acquisition module, thereby supporting the rolling and iterative discovery and information perfection of the knowledge graph, forming a closed loop of graph construction and optimization, and enabling the entire knowledge graph construction system to have the ability of iterative growth and self-improvement. The graph is not built once and for all, but can be expanded through multiple iterations to discover new entities, supplement entity attributes, and extract new relationships, thereby gradually expanding the coverage of knowledge and improving the granularity and accuracy of knowledge.
[0123] In summary, the present application has the following advantages:
[0124] (1) Higher automation and end-to-end integration: The present application realizes a complete automated process from multi-source data acquisition, preprocessing, complex extraction and verification driven by large language models, structured data fusion, entity disambiguation, to graph database storage and iterative expansion through the LangGraph workflow engine, significantly reducing the complexity of manual intervention and integration between modules;
[0125] (2) Significantly improved knowledge extraction accuracy and reliability: The present application can more accurately and reliably extract entities and relationships from complex Wikipedia texts through two-stage application of large language models (initial extraction + relationship review), combined with structured output constraints and carefully designed prompt engineering, effectively addressing the problems caused by the "hallucination" phenomenon of large language models;
[0126] (3) More comprehensive multi-source information fusion and entity characterization: The present application not only extracts relationships from Wikipedia texts, but also deeply integrates structured attributes and alias information from Wikidata, achieving more comprehensive and accurate characterization of entities, and eliminating some data redundancy through entity fusion;
[0127] (4) Support for continuous evolution and self-improvement of the graph: The iterative expansion mechanism provided by the present application enables the knowledge graph to dynamically discover information gaps based on the current graph state and automatically trigger a new round of knowledge supplementation process, achieving continuous growth and self-improvement of the knowledge graph, which is particularly important for building large-scale and dynamically updated knowledge graphs.
[0128] In a second aspect, as Figure 2As shown, the present application also proposes a knowledge graph automatic construction method based on LangGraph workflow, comprising:
[0129] According to the name list of the target entity, obtain the multi-source heterogeneous data of the target entity; wherein the multi-source heterogeneous data includes unstructured text data of Wikipedia and structured attribute data of Wikidata;
[0130] Based on the LangGraph workflow, the unstructured text data of Wikipedia of the target entity is subjected to knowledge extraction and graph construction, and a preliminary graph structure is obtained and stored in a graph database;
[0131] According to the structured attribute data of Wikidata of the target entity, the preliminary graph structure is subjected to entity attribute enrichment and relationship supplement in the graph database, and an intermediate graph structure is obtained;
[0132] In the graph database, the intermediate graph structure is subjected to entity fusion and disambiguation processing, and a final knowledge graph structure is obtained.
[0133] In this embodiment, based on the LangGraph workflow, the unstructured text data of Wikipedia of the target entity is subjected to knowledge extraction and graph construction, and a preliminary graph structure is obtained, specifically comprising:
[0134] According to the name list of the target entity, read the unstructured text data of the Wikipedia page of the target entity from the cache;
[0135] Perform intelligent block operation on the unstructured text data of the Wikipedia page of each target entity, obtain multiple text blocks of each target entity, and obtain a text block list of each target entity according to the multiple text blocks of each target entity, and cache the text block list of each target entity;
[0136] Use a preset large language model to extract information from each text block in the text block list of each target entity, and summarize the entity data and relationship data obtained by information extraction from all text blocks of each target entity to form an extraction result list of each target entity, and cache the extraction result list of each target entity;
[0137] Post-process the entity data and relationship data in the extraction result list corresponding to each target entity to obtain a candidate entity dictionary and a candidate relationship dictionary corresponding to each target entity;
[0138] Each relationship in the candidate relationship dictionary corresponding to each target entity is subjected to relationship review based on the preset large language model, and a final entity dictionary and a final relationship dictionary corresponding to each target entity are obtained;
[0139] Convert the entity data and relationship data of the final entity dictionary and the final relationship dictionary corresponding to each target entity into a standard format required by the batch write operation of the Neo4j graph database;
[0140] Construct a preliminary graph structure according to the entity data and relationship data converted into the standard format, and write the preliminary graph structure into the graph database.
[0141] In further embodiments, each relationship in the candidate relationship dictionary corresponding to each target entity is subjected to relationship review based on a pre-set large language model, to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity, specifically including: querying whether each relationship in the candidate relationship dictionary corresponding to each target entity has been reviewed and recorded in the cache; if a certain relationship has been reviewed and recorded in the cache, the review result in the cache is directly adopted; if a certain relationship has been reviewed but not recorded in the cache or not reviewed, a review prompt is constructed;
[0142] Relationship review is performed on the relationship that has been reviewed but not recorded in the cache or not reviewed, according to the review prompt, using a pre-set large language model, to obtain a review response of the large language model; the review response of the large language model is extracted to obtain a relationship review result; wherein the relationship review result includes "yes" or "no"; based on the relationship review result, the entity list is filtered to retain the entities in the relationship with a "yes" relationship review result, to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity.
[0143] Wherein, in the process of knowledge extraction and graph construction of the unstructured text data of the target entity's Wikipedia based on the LangGraph workflow, after each unit is executed, it is checked whether the error information field in the current state object is set; if yes, a preset first routing instruction is returned to guide the workflow to the error handling unit; if no, a preset second routing instruction is returned to make the workflow flow to the next processing unit in a predetermined order; when an abnormal error is captured, the error handling unit records detailed error information and terminates the current processing flow.
[0144] Wherein, the multi-source heterogeneous data is obtained, specifically including: obtaining the unstructured text data of the target entity's Wikipedia and caching it in an independent text file; obtaining the structured attribute data of the target entity's Wikidata and caching it in an independent JSON format attribute file; wherein the file name of the target entity's JSON format attribute file corresponds to the name of the target entity.
[0145] Wherein, the structured attribute data of Wikidata includes the standard label, brief description, common alias, descriptive attribute and relational attribute of the target entity.
[0146] In the embodiment, the preliminary graph structure is enriched with entity attributes and supplemented with relationships in the graph database according to the structured attribute data of Wikidata of the target entity, to obtain an intermediate graph structure, which specifically includes: enriching the descriptive attributes of the target entity in the graph database according to the structured attribute data of Wikidata of the target entity; supplementing new relationships on the graph database enriched with entity attributes according to the structured attribute data of Wikidata of the target entity and a preset mapping table; updating the preliminary graph structure based on the graph database supplemented with new relationships to obtain the intermediate graph structure.
[0147] In a further embodiment, the descriptive attributes of the entity in the Neo4j graph database are enriched according to the structured attribute data of Wikidata of the target entity, which specifically includes: traversing all JSON format attribute files under the specified path in the cache, updating the entity name of the target entity in the JSON format attribute file and in the graph database to the expression in the standard label of the corresponding JSON format attribute file, and adding all the remaining descriptive attributes contained in the JSON format attribute file as new attributes of the target entity or updating the existing attributes with the same name on the target entity.
[0148] The mapping table includes mapping the relational attribute name in the structured attribute data of Wikidata to the relationship type expected to be created in the Neo4j graph database and the recommended type of the associated entity of each relationship.
[0149] According to the structured attribute data of Wikidata of the target entity and the preset mapping table, new relationships are supplemented on the graph database enriched with the descriptive attributes of the entity, which specifically includes: merging or newly creating the associated entity in the graph database enriched with the descriptive attributes of the entity according to the relational attribute in the structured attribute data of Wikidata of the target entity; wherein, if the associated entity is newly created, the entity type and name of the newly created associated entity are set according to the preset mapping table; a relationship of the relationship type specified by the mapping table is created between the target entity and the merged or newly created associated entity.
[0150] In the embodiment, the intermediate graph structure is processed for entity fusion and disambiguation in the graph database to obtain a final knowledge graph structure, which specifically includes: querying all target entities with alias attributes in the graph database; for each target entity with alias information queried, the alias attribute of the target entity is parsed to obtain an alias list including all known aliases of the target entity;
[0151] traverse each of the alias list of the target entity, for each alias, determine whether there is an entity name in the graph database with the same type of alias entity that is exactly the same; if there is, copy or migrate all attribute information owned by the alias entity to the target entity, and all relationships pointing to or from the alias entity are redirected to the target entity, and the alias entity and all associated relationships have been successfully merged from the graph database is completely removed;
[0152] Based on the graph database, the intermediate knowledge graph structure is updated to obtain a final knowledge graph structure.
[0153] In the embodiment, after the intermediate knowledge graph structure is subjected to entity fusion and disambiguation processing to obtain the final knowledge graph structure, the method further includes: querying whether there is an entity in the knowledge graph structure that already exists but whose key structured information has not been completely supplemented; if not, outputting the knowledge graph structure; if yes, finding out the entity that has been identified in the knowledge graph structure but whose key structured information has not been completely supplemented, and forming a new to-be-crawled entity text file according to the entity that has been identified in the knowledge graph structure but whose key structured information has not been completely supplemented.
[0154] In a further embodiment, after the new to-be-crawled entity text file is formed according to the entity that has been identified in the knowledge graph structure but whose key structured information has not been completely supplemented, the method further includes: taking the new to-be-crawled entity text file as a name list of a target entity for iteration, and re-performing the automatic construction and optimization of the knowledge graph.
[0155] The above description is only a preferred embodiment of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can make equivalent replacements or changes to the technical solution and the inventive concept of the present application within the technical range disclosed by the present application, which should be covered within the protection scope of the present application.
Claims
1. A LangGraph workflow-based knowledge graph automatic construction system, characterized in that, The method comprises the following steps: a multi-source heterogeneous data acquisition module for acquiring multi-source heterogeneous data of a target entity according to a name list of the target entity; wherein the multi-source heterogeneous data comprises unstructured text data of Wikipedia and structured attribute data of Wikidata; a knowledge extraction and graph construction module for performing knowledge extraction and graph construction on the unstructured text data of Wikipedia of the target entity based on a LangGraph workflow, obtaining a preliminary graph structure, and storing the preliminary graph structure in a graph database; a Wikidata structured attribute integration module for enriching the preliminary graph structure with entity attributes and supplementing the preliminary graph structure with relationships in the graph database based on the structured attribute data of Wikidata of the target entity, and obtaining an intermediate graph structure; an entity fusion and disambiguation processing module for performing entity fusion and disambiguation processing on the intermediate graph structure in the graph database, and obtaining a final knowledge graph structure; wherein the LangGraph workflow comprises: document reading: reading the unstructured text data of the Wikipedia page of the target entity from the cache according to the name list of the target entity; text blocking: performing intelligent blocking operation on the unstructured text data of the Wikipedia page of each target entity to obtain multiple text blocks of each target entity, obtaining a text block list of each target entity according to the multiple text blocks of each target entity, and caching the text block list of each target entity; information extraction: using a pre-set large language model to perform information extraction on each text block in the text block list of each target entity, and summarizing the entity data and relationship data obtained by information extraction from all text blocks of each target entity to form an extraction result list of each target entity, and caching the extraction result list of each target entity; post-extraction processing: performing post-processing on the entity data and relationship data in the extraction result list corresponding to each target entity to obtain a candidate entity dictionary and a candidate relationship dictionary corresponding to each target entity; relationship review: performing relationship review on each relationship in the candidate relationship dictionary corresponding to each target entity based on the pre-set large language model to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity; graph data preparation: converting the entity data and relationship data of the final entity dictionary and the final relationship dictionary corresponding to each target entity into a standard format required for batch writing operation of the graph database; graph writing: constructing a preliminary graph structure according to the entity data and relationship data converted into the standard format, and writing the preliminary graph structure into the graph database. 2.The LangGraph workflow-based knowledge graph automated construction system according to claim 1, wherein, performing relationship review on each relationship in the candidate relationship dictionary corresponding to each target entity based on the pre-set large language model to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity, which specifically comprises: querying whether each relationship in the candidate relationship dictionary corresponding to each target entity has been reviewed and recorded in the cache; if a certain relationship has been reviewed and recorded in the cache, the review result in the cache is directly adopted; if a certain relationship has been reviewed but not recorded in the cache or not reviewed, a review prompt is constructed; The relationship review is performed on the relationship that has been reviewed but not recorded in the cache or not reviewed by using a preset large language model according to the review prompt, to obtain a review response of the large language model; The review response of the large language model is extracted to obtain a relationship review result; wherein the relationship review result includes yes or no; Based on the relationship review result, the candidate entity dictionary is filtered to retain the entities in the relationship with the relationship review result being yes, to obtain a final entity dictionary and a final relationship dictionary corresponding to each target entity. 3.The LangGraph workflow based knowledge graph automated construction system of claim 1, wherein, The knowledge extraction and graph construction module further comprises an error judgment unit and an error processing unit. The error judgment unit is used to check whether the error information field in the current state object is set after each unit is executed; if yes, a preset first routing instruction is returned to guide the workflow to the error processing unit; if no, a preset second routing instruction is returned to make the workflow flow to the next processing unit in a predetermined order; The error processing unit is used to record detailed error information when an abnormal error is captured, and terminate the current processing flow. 4.The LangGraph workflow-based knowledge graph automated construction system of claim 1, wherein, The structured attribute data of Wikidata of the target entity includes a standard label, a brief description, a commonly used alias, a descriptive attribute, and a relational attribute of the target entity; The structured attribute data of Wikidata of the target entity is used to enrich the entity attribute and supplement the relationship of the preliminary graph structure in the graph database, to obtain an intermediate graph structure, specifically including: The structured attribute data of Wikidata of the target entity is used to enrich the descriptive attribute of the target entity in the graph database; The structured attribute data of Wikidata of the target entity and a preset mapping table are used to supplement new relationships on the graph database after the entity attribute enrichment; The preliminary graph structure is updated based on the graph database after the new relationships are supplemented, to obtain the intermediate graph structure. 5.The LangGraph workflow-based knowledge graph automated construction system of claim 4, wherein, The structured attribute data of Wikidata of each target entity is cached in an independent JSON format attribute file, and the file name of each JSON format attribute file corresponds to the name of the target entity; The structured attribute data of Wikidata of the target entity is used to enrich the descriptive attribute of the target entity in the graph database, specifically including: All JSON format attribute files under the specified path in the cache are traversed, the entity name of the target entity in the JSON format attribute file and the graph database is updated to the expression in the standard label of the corresponding JSON format attribute file, and all the remaining descriptive attributes contained in the JSON format attribute file are added as new attributes of the target entity or used to update the same-named attributes already existing on the target entity. 6.The LangGraph workflow-based knowledge graph automated construction system of claim 5, wherein, The mapping table includes mapping the relational attribute in the structured attribute data of Wikidata of the target entity in the intermediate graph structure to the relationship type expected to be created in the graph database and the recommended type of the associated entity of each relationship; The structured attribute data of Wikidata of the target entity and the preset mapping table are used to supplement new relationships on the graph database after the descriptive attribute of the entity is enriched, specifically including: According to the relational attribute in the structured attribute data of the Wikidata of the target entity, a related entity is merged or newly created in a descriptive attribute-rich graph database of the entity; if the related entity is newly created, the entity type and name of the newly created related entity are set according to a preset mapping table; A relationship of a relationship type specified by the mapping table is created between the target entity and the merged or newly created related entity. 7.The LangGraph workflow based knowledge graph automated construction system of claim 1, wherein, In the graph database, entity fusion and disambiguation processing are performed on the intermediate graph structure to obtain a final knowledge graph structure, specifically including: In the graph database, all target entities with alias attributes are queried; For each target entity with alias information, the alias attribute of the target entity is parsed to obtain an alias list including all known aliases of the target entity; each alias in the alias list is traversed, and for each alias, it is determined whether there is an alias entity of the same type in the graph database whose name is completely identical with the alias; if there is, all attribute information possessed by the alias entity is copied or migrated to the target entity, all relationships pointing to or originating from the alias entity are redirected to the target entity, and the alias entity and all associated relationships that have been successfully merged are completely removed from the graph database; Based on the graph database, the intermediate graph structure is updated to obtain the final knowledge graph structure. 8.The LangGraph workflow based knowledge graph automated construction system of claim 1, wherein, Further including: An iterative graph expansion support module is configured to query whether there is an entity in the final knowledge graph structure that already exists but whose key structured information has not been completely supplemented; if not, the knowledge graph structure is output; if yes, the entity that already exists but whose key structured information has not been completely supplemented is found, and a new to-be-crawled entity text file is formed according to the entity that already exists but whose key structured information has not been completely supplemented. 9.The LangGraph workflow-based knowledge graph automated construction system of claim 8, wherein, The iterative graph expansion support module is further configured to feed back the new to-be-crawled entity text file as a name list of a target entity to the multi-source heterogeneous data acquisition module.
10. A LangGraph workflow-based knowledge graph automatic construction method applied to the LangGraph workflow-based knowledge graph automatic construction system of any one of claims 1-9, characterized in that, Including: According to the name list of the target entity, multi-source heterogeneous data of the target entity is acquired; the multi-source heterogeneous data includes unstructured text data of Wikipedia and structured attribute data of Wikidata; Based on the LangGraph workflow, knowledge extraction and graph construction are performed on the unstructured text data of Wikipedia of the target entity to obtain a preliminary graph structure, and the preliminary graph structure is stored in the graph database; According to the structured attribute data of Wikidata of the target entity, entity attribute enrichment and relationship supplement are performed on the preliminary graph structure in the graph database to obtain an intermediate graph structure; In the graph database, entity fusion and disambiguation processing are performed on the intermediate graph structure to obtain a final knowledge graph structure.
Citation Information
Patent Citations
Knowledge graph construction method and device based on pre-trained large language model
CN117851610A
Multi-stage serial knowledge graph construction method and system based on cue words
CN120124729A