Construction and extension method and system based on dynamic semantic knowledge graph
By constructing and expanding the knowledge graph and using large language models for inference and verification, the problems of insufficient content depth and insufficient data volume of the knowledge graph in the existing technology are solved, and high-quality knowledge graph expansion and deepening are achieved, which is suitable for scenarios such as intelligent question-and-answer and information retrieval.
Patent Information
- Application Number
- CN202510736200.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
During the construction process, the existing knowledge graphs have problems such as insufficient content depth, weak semantic understanding ability and insufficient data size, which is difficult to meet the knowledge representation needs in complex scenarios.
By obtaining the head entity, tail entity and entity relationship of text blocks, attaching attribute key-value pairs, generating relationship instance triplets, building structured and semantic knowledge graphs, and using the Large Language Model (LLM) for inference and verification, generating trusted knowledge graphs, realizing the expansion and deepening of the knowledge graphs.
It realizes high-quality expansion of the knowledge graph, improves semantic accuracy and data coverage, is suitable for intelligent applications in complex scenarios, and ensures the accuracy and credibility of inference results.
Smart Images

Figure CN120258115A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and specifically relates to a method and system for constructing and expanding a dynamic semantic knowledge graph. Background Art
[0002] As a complex network structure for organizing and expressing data, the Knowledge Graph (KG) has become a key tool for the efficient utilization of data elements. Through the structured representation of entities and their relationships, the knowledge graph realizes the transformation from data to knowledge and is widely used in scenarios such as information retrieval, personalized recommendation, and intelligent question answering services. With the continuous growth of data scale and complexity, existing knowledge graph construction technologies have problems such as insufficient content depth, weak semantic understanding ability, and insufficient data volume.
[0003] Although existing knowledge graph construction technologies have made certain progress, the following key problems are still faced in the construction process: First, currently, most knowledge graph construction stays in the shallow learning stage, mainly reflecting the organizational structure of data and hierarchical entity relationships, and failing to deeply mine and express the content depth of knowledge. This results in insufficient capture of the internal logic and details of knowledge, limiting its application effect in solving complex problems. Second, the knowledge graph still has insufficient understanding at the semantic level. Existing technologies mostly focus on the processing of surface features of entities and direct entity relationships, and fail to deeply analyze semantic attributes and complex entity relationship networks. This shallow semantic expression method is difficult to meet users' requirements for knowledge accuracy in complex scenarios. Third, in the current knowledge graph construction process, the data volume is limited, and it is difficult to perform reasoning and expansion through automated means. This deficiency leads to limited knowledge coverage and affects the performance of the knowledge graph in large-scale applications. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide a method and system for constructing and expanding a dynamic semantic knowledge graph, which can solve the problems of insufficient content depth, weak semantic understanding ability, and insufficient data volume existing in the prior art during the construction of the knowledge graph.
[0005] To solve the above technical problems, this application is implemented as follows: First aspect, an embodiment of the present application provides a method for constructing and expanding a dynamic semantic knowledge graph. The method includes: obtaining the head entity, tail entity, and entity relationship of each text block, where the entity relationship represents the semantic entity relationship between the head entity and the tail entity; respectively attaching corresponding attribute key-value pair sets to the head entity, tail entity, and entity relationship to obtain an attached head entity, an attached entity relationship, and an attached tail entity. The attribute key-value pair set includes multiple attribute keys and attribute values corresponding to the head entity and the tail entity; generating a relationship instance triple corresponding to each text block based on the attached head entity, entity relationship, and attached tail entity; generating a structured knowledge graph according to the relationship instance triples; generating a semantic knowledge graph according to the attached head entity, attached tail entity, and attached entity relationship after respectively adding corresponding index information; respectively selecting any one of the attached head entity, attached tail entity, and entity relationship in the relationship instance triple as the masking object, reasoning about the masking object based on the other two in the relationship instance triple, generating multiple inferred relationship instance triples and then validating them to generate a credible knowledge graph; fusing the structured knowledge graph, semantic knowledge graph, and credible knowledge graph to generate a target knowledge graph.
[0006] As an alternative implementation manner of the first aspect of the present application, the specific process of obtaining the relationship instance triples includes: constructing an entity set and an entity relationship set according to the entities and entity relationships of multiple text blocks. The entity set includes multiple head entities and multiple tail entities, and the entity relationship set includes multiple entity relationships; fusing the entity set based on a preset fusion rule to generate an attribute key-value pair set; respectively obtaining the entity type and attribute key-value pair set corresponding to each entity in the entity set according to the head entity and the tail entity; generating relationship instance triples based on the entity set, entity relationship set, and attribute key-value pair set.
[0007] As an alternative implementation manner of the first aspect of the present application, the preset fusion rule includes: obtaining the entity type, entity name, and attribute value of each text block. When the entity types and entity names are the same but the attribute values are different, determine whether the attribute key corresponding to the attribute value exists. When the attribute key exists, replace the attribute value of the later batch with the attribute value of the previous batch. When the attribute key does not exist, add a new attribute key-value pair according to the updated attribute key and the attribute value of the later batch.
[0008] As an alternative implementation manner of the first aspect of the present application, the process of generating the semantic knowledge graph includes: ; where, represents the semantic knowledge graph, represents the head entity, represents the tail entity, represents the entity relationship, Represents the index of the head entity, Represents the index of the entity relationship, Represents the index of the tail entity, Represents the set of attribute key-value pairs of the head entity, Represents the set of attribute key-value pairs of the entity relationship, Represents the set of attribute key-value pairs of the tail entity.
[0009] As an alternative implementation of the first aspect of this application, the process of reasoning about the obscured object includes: Respectively obscure the head entity, tail entity, and entity relationship within the relational instance triple, and correspondingly obtain multiple relational instance triples after partial information concealment processing; use the LLM to drive the inference of missing information for the relational instance triples after partial information concealment processing to obtain the inferred head entity, the inferred entity relationship, and the inferred tail entity; use the LLM to drive the attribute inference for the inferred head entity and the inferred tail entity respectively to obtain the inferred attribute set; combine the inferred attribute set, the inferred head entity, and the inferred entity relationship to generate multiple relational instance triples after inference.
[0010] As an alternative implementation of the first aspect of this application, the process of generating a trustworthy knowledge graph includes: searching for web content blocks associated with the relational instance triples after inference through a Web search engine, where each relational instance triple after inference corresponds to a web content block; calculating the cosine similarity between each relational instance triple after inference and the corresponding web content block; based on the magnitude relationship between the cosine similarity and a preset threshold, verifying each relational instance triple after inference to generate multiple verified relational instance triples; combining multiple verified relational instance triples to generate a trustworthy knowledge graph.
[0011] As an alternative implementation of the first aspect of this application, the process of verifying each relational instance triple after inference based on the magnitude relationship between the cosine similarity and a preset threshold includes evidence sufficiency verification and multi-way verification. After the evidence sufficiency verification passes, enter the multi-way verification to generate multiple verified relational instance triples. The process of evidence sufficiency verification includes: The process of multi-way verification includes: Among them, Represents the evidence sufficiency verification result, Represents the success of the evidence sufficiency verification, Represents the failure of the evidence sufficiency verification, Represents the preset threshold, represents the intersection operation, and represents the vector generated after being encoded by the embedding model, represents the cosine similarity between the inferred relational instance triple and a single web page content block, 、 、 represent three different prompt words, represents the inferred relational instance triple, represents the multi-way verification result, represents the inference of the large language model.
[0012] In a second aspect, an embodiment of the present application provides a construction and extension system for a dynamic semantic knowledge graph, and the system includes: An extraction module, configured to obtain the head entity, tail entity, and entity relationship of each text block; An attribute attachment module, configured to respectively attach corresponding attribute key-value pairs to the head entity, tail entity, and entity relationship, and respectively obtain the attached head entity, attached entity relationship, and attached tail entity; A relational instance triple construction module, configured to generate a relational instance triple corresponding to each text block based on the attached head entity, entity relationship, and attached tail entity; A structured knowledge graph generation module, configured to generate a structured knowledge graph according to the relational instance triple; A semantic knowledge graph generation module, configured to generate a semantic knowledge graph according to the attached head entity, attached tail entity, and attached entity relationship after respectively adding corresponding index information; A trustworthy knowledge graph generation module, configured to respectively select any one of the attached head entity, attached tail entity, and entity relationship within the relational instance triple as the masking object, perform inference on the masking object according to the other two within the relational instance triple, generate multiple inferred relational instance triples and then perform verification to generate a trustworthy knowledge graph.
[0013] A target knowledge graph generation module, configured to fuse the structured knowledge graph, semantic knowledge graph, and trustworthy knowledge graph to generate a target knowledge graph.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, and the electronic device includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of the method in the first aspect are implemented.
[0015] In a fourth aspect, an embodiment of the present application provides a readable storage medium, and a program or instruction is stored on the readable storage medium. When the program or instruction is executed by the processor, the steps of the method in the first aspect are implemented.
[0016] Compared with the prior art, the beneficial effects of a method for constructing and expanding a dynamic semantic knowledge graph provided by the present invention are as follows: First, through the knowledge graph extraction method based on LLM, the present invention adopts the step-by-step block parsing and semantic-driven entity relationship extraction technology to achieve the unified fusion of multi-data sources and multi-modal information, forming a structured semantic knowledge graph. When processing text data, it can automatically maintain the context integrity, ensuring the semantic accuracy and structural consistency of the knowledge graph. By using the fusion algorithm, the redundancy and inconsistency problems between different data blocks are solved, making the generated knowledge graph have high quality and high availability, and meeting the knowledge representation requirements in complex scenarios.
[0017] Second, through the reasoning method based on LLM, the present invention can automatically infer a new set of entities, entity relationships, and attribute key-value pairs on the basis of the existing knowledge graph. During the reasoning process, through the strategies of concealment and filling, the model can not only verify the integrity of the existing data, but also expand the data volume of the knowledge graph. This method is particularly suitable for fields with insufficient data or incomplete annotation, and can realize the semantic expansion of the knowledge graph, significantly improving its coverage and depth, and providing richer data support for subsequent intelligent applications.
[0018] Third, the present invention combines boolean queries and verification strategies to ensure the correctness and credibility of the reasoning results. For each newly added reasoning relationship instance triple, this method can quickly collect the evidence sources and perform associated matching based on the semantic similarity of the RAG system, effectively improving the efficiency and accuracy of verification. In the case of insufficient evidence, the reliability of the results is further strengthened through a three-way verification mechanism. This verification method combining the evidence chain enhances the usability and user trust of the knowledge graph, and is suitable for scenarios with high requirements for data quality.
[0019] Fourth, the present invention innovatively realizes the closed-loop operation of knowledge graph reasoning and verification, generates new knowledge through reasoning, verifies to ensure data quality, and continuously optimizes the knowledge graph in an iterative manner. The entire system has strong adaptability and scalability, can not only process complex semantic scenarios, but also dynamically update and expand the content of the knowledge graph to meet the knowledge update requirements in practical applications. Description of the Drawings
[0020] Figure 1 is a flowchart of a method for constructing and expanding a dynamic semantic knowledge graph provided by the first embodiment of the present application; Figure 2 is a full flowchart of knowledge graph construction provided by the first embodiment of the present application; Figure 3 is a flowchart of the process of generating a structured knowledge graph provided by the first embodiment of the present application; Figure 4 It is the internal structure diagram of a construction and expansion system based on a dynamic semantic knowledge graph provided by the second embodiment of the present application. Detailed implementation manners
[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0022] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally represents an "or" entity relationship between the associated objects before and after.
[0023] Next, in conjunction with the accompanying drawings, a construction and expansion method and system based on a dynamic semantic knowledge graph provided by the embodiments of the present application will be described in detail through specific embodiments and their application scenarios.
[0024] Embodiment 1 Please refer to Figure 1 , which represents the flowchart of a construction and expansion method based on a dynamic semantic knowledge graph provided by the present invention. The method includes steps S1 to S7.
[0025] Step S1: Obtain the head entity, tail entity, and entity relationship of each text block. The entity relationship represents the semantic entity relationship between the head entity and the tail entity.
[0026] Specifically, the present invention first preprocesses the text data, gradually divides the text data into blocks, that is, evenly divides it into multiple text blocks of equal length, and ensures the integrity of the context where the text block is located during the cutting process.
[0027] Step S2: Attach the corresponding attribute key-value pair sets to the head entity, tail entity, and entity relationship respectively, and obtain the attached head entity, attached entity relationship, and attached tail entity respectively. The attribute key-value pair set includes multiple attribute keys and attribute values corresponding to the head entity and the tail entity.
[0028] Specifically, the present invention first, based on the defined entity type , uses the LLM to identify the entity set one by one , and then based on the defined entity relationship types , extract the set of entity relationships between entities from each text block , each entity relationship links two entities and , and finally for each entity within the entity set , obtain the set of attribute keys held by each entity type. Among them, represents the entity types within, represents the entity set within each entity, represents the relationship type within each relationship type, represents the relationship set within each entity relationship.
[0029] Among them, multiple attribute keys and attribute values corresponding to the head entity and the tail entity are obtained by extracting from the set of attribute keys. For example, for the operating system, this entity type has multiple instance entities such as Linux, HarmonyOS, and IOS. This type has attribute keys "architecture" and "compatibility", and the values corresponding to the same attribute key for each instance entity are different, finally forming a set of attribute key-value pairs , among which, represents each different attribute key-value pair, , , represents each different attribute key, , , represents each different attribute value.
[0030] Step S3: Generate a relationship instance triple corresponding to each text block based on the additional head entity, entity relationship, and additional tail entity.
[0031] Specifically, the specific process of obtaining the relationship instance triple includes: constructing an entity set and an entity relationship set according to the entities and entity relationships of multiple text blocks, where the entity set includes multiple head entities and multiple tail entities, and the entity relationship set includes multiple entity relationships; fusing the entity set based on a preset fusion rule to generate a set of attribute key-value pairs; respectively obtaining the entity type and the set of attribute key-value pairs corresponding to each entity in the entity set according to the head entity and the tail entity; generating a relationship instance triple based on the entity set, the entity relationship set, and the set of attribute key-value pairs.
[0032] Among them, the preset fusion rules include: obtaining the entity type, entity name, and attribute value of each text block. When the entity types and entity names are the same but the attribute values are different, it is judged whether the attribute key corresponding to the attribute value exists. When the attribute key exists, the attribute value of the later batch is replaced with the attribute value of the previous batch. When the attribute key does not exist, an attribute key-value pair is newly added according to the updated attribute key and the attribute value of the later batch. The specific algorithm implementation formula is: ; Among them, represents the finally fused attribute set, represents the number of text blocks, represents the attribute key, represents the attribute value of the previous batch, represents the attribute value of the later batch, represents the intersection operation, represents the union operation, represents the set of attribute key-value pairs.
[0033] Step S4: Generate a structured knowledge graph according to the relational instance triples.
[0034] Specifically, the specific process of generating a structured knowledge graph includes: generating a structured knowledge graph based on the extracted entities and entity relationships. The mathematical expression is: ; Among them, represents the structured knowledge graph, represents the attribute set of the head entity, represents the attribute set of the tail entity, represents the head entity, represents the tail entity, represents the entity relationship.
[0035] Furthermore, combining multiple to obtain a structured knowledge graph .
[0036] Step S5: Generate a semantic knowledge graph according to the additional head entity, additional tail entity, and additional entity relationship after respectively adding the corresponding index information.
[0037] Specifically, the present invention integrates the entity set , the entity relationship set , and the finally fused attribute set into the final semantic knowledge graph , where represents each entity in the entity set , and represents the relationship set For each relationship within it, obtain the final semantic knowledge graph The specific process is as follows: First, focus on the entity relationship set. For each entity relationship, the head and tail entities will be extracted accordingly. Similar to extracting a complete sentence from a document, the predicate will be extracted first, and then the subject and object will be extracted in sequence to form a complete subject-predicate-object structure. The relationship instance triple is the basic component (not the smallest) that constitutes the knowledge graph. Its integration process is similar. It also traverses each entity relationship in the entity relationship set. The entity relationship structure contains the head and tail entities, but does not contain information such as types and attributes. Therefore, it is necessary to integrate the information in the entity set and the attribute set. For example, when traversing an entity relationship: ibuprofen - relieves -> pain, at this time, only the basic information for forming the relationship instance triple is available, lacking more important details: entity types and entity attributes. Therefore, the head and tail entities are used as indexes respectively to find their respective type and attribute information in the entity set and the attribute combination. For example, in the entity set, there is ibuprofen: drug. Using the index "ibuprofen", its type information "drug" can be obtained. Similarly, its attribute information can also be obtained.
[0038] Among them, the mathematical expression of the semantic knowledge graph is: ; Among them, represents the semantic knowledge graph, represents the head entity, represents the tail entity, represents the entity relationship, represents the index of the head entity, represents the index of the entity relationship, represents the index of the tail entity, represents the set of attribute key-value pairs of the head entity, represents the set of attribute key-value pairs of the entity relationship, represents the set of attribute key-value pairs of the tail entity.
[0039] Step S6: Select any one of the additional head entity, additional tail entity, and entity relationship within the relationship instance triple as the masking object, and perform reasoning on the masking object based on the other two within the relationship instance triple. After generating multiple reasoned relationship instance triples, perform verification to generate a credible knowledge graph.
[0040] Specifically, the process of reasoning about the obscured object includes: respectively obscuring the head entity, the tail entity, and the entity relationship within the relational instance triple, and correspondingly obtaining multiple relational instance triples after partial information concealment processing; using the LLM to drive the inference of missing information for the relational instance triples after partial information concealment processing to obtain the inferred head entity, the inferred entity relationship, and the inferred tail entity; using the LLM to drive the attribute inference for the inferred head entity and the inferred tail entity respectively to obtain the inferred attribute set; combining the inferred attribute set, the inferred head entity, and the inferred entity relationship to generate multiple inferred relational instance triples.
[0041] Among them, the specific process of obscuring the head entity, the tail entity, and the entity relationship within the relational instance triple includes: # Note: Entity structure h -> { Type: "Th", name: "h", Attributes: { "attr1": "value1", ... } }; After the concealment process, the above entity will be processed into: -> { Type: "Th" # The type information is still retained to guide the LLM inference direction and prevent a large amount of noisy data from being introduced during the inference process name: "?", # Conceal the name Attributes: { "attr1": "?", # Retain the attribute key and conceal the attribute value } Among them, Type, name, Attributes, and "attr1" represent the attributes within the entity structure, "Th" represents the attribute value corresponding to the Type attribute before the concealment process, "h" represents the attribute value corresponding to the name attribute before the concealment process, "value1" represents the attribute value corresponding to the "attr1" attribute before the concealment process, h- represents the entity before the concealment process, - represents the entity after the concealment process, "Th" represents the attribute value corresponding to the Type attribute after the concealment process, "?" represents the attribute value corresponding to the name attribute and the "attr1" attribute after the concealment process, and Attributes represents multiple attributes.
[0042] Then, perform partial information concealment on the entity part and complete concealment of the entity relationship in the relationship instance triples in the knowledge graph to form three types of concealed relationship instance triples, namely: concealed head entity, hidden entity relationship, and concealed tail entity.
[0043] Furthermore, use the LLM to drive the inference of missing information for the relationship instance triples after partial information concealment to generate the missing information. The inference formula is: ; where is the output generated by the LLM, is the input of the concealed relationship instance triple, is the designed prompt word used to guide the LLM to generate accurate missing information.
[0044] Furthermore, for the entities generated by inference, use the LLM to generate their attribute key-value pair sets , and the inference formula is: ; where represents the set of attributes generated by inference, represents the set of target attribute keys, represents the attribute key, represents the attribute value.
[0045] The process of generating a trustworthy knowledge graph in the present invention includes: searching for web content blocks associated with the inferred relationship instance triples through a Web search engine, where each inferred relationship instance triple corresponds to a web content block; calculating the cosine similarity between each inferred relationship instance triple and the corresponding web content block; based on the magnitude relationship between the cosine similarity and a preset threshold, verifying each inferred relationship instance triple to generate multiple verified relationship instance triples; and combining the multiple verified relationship instance triples to generate a trustworthy knowledge graph.
[0046] Specifically, the process of verifying each inferred relationship instance triple based on the magnitude relationship between the cosine similarity and the preset threshold includes evidence sufficiency verification and multi-way verification. After the evidence sufficiency verification passes, enter the multi-way verification to generate multiple verified relationship instance triples. The process of sufficiency verification includes: The process of multi-way verification includes: where represents the sufficiency verification result, represents the success of sufficiency verification, Indicates that the sufficiency verification fails, Indicates the preset threshold, Indicates the intersection operation, and Indicates the vector generated after encoding by the embedding model, Indicates the cosine similarity between the inferred relationship instance triple and a single web page content block, , , Indicates three different prompt words, Indicates the inferred relationship instance triple, Indicates the multi-way verification result, Indicates the inference of the large language model.
[0047] Furthermore, when both the sufficiency verification and the multi-way verification pass, the inferred entity relationship implementation triple that passes the verification is regarded as the verified relationship instance triple, and multiple such verified relationship instance triples are combined to obtain a credible knowledge graph.
[0048] The specific steps of the present invention to obtain a credible knowledge graph include: First, traverse each inference triple one by one and organize it into a Boolean query statement, filter out the most relevant web pages through an advanced search engine, and extract the web page content. Then, through the RAG system, the web page content is embedded and encoded into vectors, and the inference triple is used as the query query statement to retrieve the most relevant web page content from the vector library and try to use it as evidence. If the evidence is sufficient and reasonable, it is considered that the inference triple is correct and reliable, with a high credibility, and is directly marked as a credible triple. Otherwise, it further enters the multi-way verification module, and through three designed verification logics, drives the LLM as an auditor to comprehensively judge the correctness or credibility of the triple from three perspectives. If all three judgment results are correct, it is also considered that the inference triple is correct and marked as a credible triple. Finally, all credible triples are integrated into a complete credible knowledge graph.
[0049] In summary, the present invention ensures the correctness and credibility of the inference result through search engine-based and LLM inference, combined with Boolean query and multi-way verification strategies. For each newly added inference triple, the present invention can quickly collect the evidence source and perform association matching based on the semantic similarity of the RAG system, effectively improving the efficiency and accuracy of verification. In the case of insufficient evidence, the reliability of the result is further strengthened through the three-way verification mechanism. This verification method combining the evidence chain enhances the usability and user trust of the knowledge graph, and is suitable for scenarios with high requirements for data quality.
[0050] Step S7: Integrate the structured knowledge graph, the semantic knowledge graph, and the credible knowledge graph to generate the target knowledge graph.
[0051] Specifically, the process of generating the target knowledge graph includes: ; Among them, represents the target knowledge graph, is a structured knowledge graph, represents the semantic knowledge graph, represents the trustworthy knowledge graph, and the target knowledge graph will be used as the input for subsequent intelligent question-answering systems or information retrieval. represents the union operation.
[0052] Please refer to Figure 2 , which shows the full flow chart of constructing the knowledge graph of the present invention, including: First, the input file is processed through structuring to generate a structured knowledge graph. Then, a knowledge graph with semantic information is constructed by parsing the text and its index relationships. After the semantic knowledge graph is constructed, reasoning is performed on it to generate new semantic association relationships and entities to complement the information that may be missing in the knowledge graph. Subsequently, the knowledge generated by reasoning is verified to ensure accuracy and reliability, and finally, the verified knowledge is fused with the original knowledge to generate the final semantic knowledge graph.
[0053] Please refer to Figure 3 , which shows the flow chart of the process of generating the structured knowledge graph of the present invention, Figure 3 showing how to extract structured information (such as a table of contents with structural meta-information like Chapter 1 -> Section 1 -> 1.1 -> 1.1.1, and other such as papers: abstracts, experiments, etc.) from files in formats such as PDF, DOCX, or TXT, and generate a structured knowledge graph. The process is as follows: Convert the structured files (PDF files, DOCX files, and TXT files) into Markdown text, then extract relationship instance structure triples according to the subordinate relationship and precedence relationship to generate subordinate relationship instance triples and precedence relationship instance triples, and then assign entity text attributes to the triples to generate a structural knowledge graph.
[0054] Specifically, first read the file content and convert it into Markdown format. By parsing the hierarchical and sequential relationships (such as Chapter 1 preceding Chapter 2, and Section 1 being in a relatively lower hierarchical relationship belonging to Chapter 1), extract instance triples. The extracted triples will be appended with text attribute information (such as in the node "Chapter 1", the value of the attribute "text" is the summary of the content of this chapter), for example, the chapter and title to which each entity belongs, thus generating a complete structured knowledge graph. The output form of the structured knowledge graph includes (Chapter1 [Text: Text1], has, First Title1 [Text: Text2]), etc., and at the same time, generate the correspondence between the index and the text, providing a basis for the subsequent construction of the semantic knowledge graph.
[0055] The present invention is based on the construction of an indexed structured knowledge graph and a semantic knowledge graph, and includes the following steps: 1) File conversion and structured information extraction; Convert the input PDF, DOCX, TXT and other files into Markdown format through automated tools or manual marking, and parse the file content to generate instance relationship instance triples based on the hierarchical and sequential entity relationships (such as the markdown content: # Chapter 1 Chapter Abstract 1 ## 1.1.1 Specific Content 1.1.1 # Chapter 2 Chapter Abstract 2 ## 2.1.1 Specific Content 2.1.1 According to the markdown syntax, the syntax marker of Chapter 1 is in a higher hierarchical relationship with the syntax marker of 1.1.1, so it is extracted as a structural relationship instance triple (Chapter 1, higher level, 1.1.1). Also, because the syntax marker of Chapter 2 is at the same level as that of Chapter 1 but appears after the latter, it is extracted as a structural relationship instance triple (Chapter 1, precedes, Chapter 2). At the same time, the corresponding content is used as the attribute of the entity. The process of structured information extraction is as follows: { Head entity: { "name": "Chapter 1", "Type": "Chapter 2", "Attributes": { "text": "Chapter Abstract 1" }, Entity relationship: { "name": "1.1.1" "Type": "Third-level heading" "Attributes": { "text": "Specific content 1.1.1" }} }} }; Each entity will be attached with attribute information in the following form: ; Among them, represents the attribute information corresponding to entity , represents the entity, represents the attribute name, represents the corresponding attribute value, represents the entity corresponding number of attributes.
[0056] 2) Construction of the structured knowledge graph; After structured information extraction, based on the extracted entities and entity relationships, a structured knowledge graph is generated, and the expression is: ; Among them, represents the structured knowledge graph, represents the set of attributes of the head entity, represents the set of attributes of the tail entity, represents the head entity, represents the tail entity, represents the entity relationship.
[0057] 3) Construction of the semantic knowledge graph; Based on the structured knowledge graph, an index attribute is added to each relationship instance triple, and the semantic knowledge graph is represented as follows: ; Among them, represents the semantic knowledge graph, is the index attribute, representing the text source of the relationship instance triple (e.g., if idx = \"C++ Primer -> Chapter 1 Introduction\", it can be proved that this node comes from Chapter 1 of this book), represents the set of attributes of entity relationship r. The finally generated semantic knowledge graph serves as the basis for subsequent reasoning and verification.
[0058] The construction process of the inference knowledge graph without index based on the present invention includes the following steps: 1) Semantic information extraction; By chunking the text, entities and entity relationships are extracted from it: ; ; Among them, respectively represent the set of entities and entity relationships that conform to the target type, represent the types of entities and entity relationships, are respectively the set of target entity types and the set of target entity relationship types, represents an entity, represents an entity relationship, represents a set of entities, represents a set of entity relationships.
[0059] 2) Preliminary generation of relationship instance triples; Combining the extracted entities and entity relationships to generate semantic relationship instance triples : ; 3) Semantic reasoning; Infer potential semantic entity relationships and entities through existing relationship instance triples. Through information obfuscation processing, obtain obfuscated relationship instance triples, and drive the LLM to perform inference and completion of hidden information according to the originally retained entity types, entity relationship types, and attribute key sets. The mathematical expression is: I ; ; Among them, I represents the inference result, that is, the inference relationship instance triple, represents the prompt word designed to drive the LLM for inference and completion, represents the input obfuscated relationship instance triple, represents the entity relationship generated after LLM inference, represents the head entity generated after LLM inference, represents the tail entity generated after LLM inference.
[0060] 4) Integration of inference results; The newly generated relationship instance triples inferred are integrated as: ; Combine multiple newly generated relationship instance triples inferred to generate an unindexed inference knowledge graph.
[0061] In summary, through the inference method based on the LLM, the present invention can automatically infer new entities, entity relationships, and attribute information on the basis of the existing knowledge graph. During the inference process, through the strategies of concealment and filling, the model can not only verify the integrity of the existing data, but also expand the data volume of the knowledge graph. This method is particularly suitable for fields with insufficient data or incomplete annotation, can achieve semantic expansion of the knowledge graph, significantly improve its coverage and depth, and provide richer data support for subsequent intelligent applications.
[0062] Embodiment 2 Please refer to Figure 4 , which shows the internal structure diagram of a construction and expansion system based on a dynamic semantic knowledge graph proposed in the second embodiment of the present application. The system includes: An extraction module 100, configured to obtain the head entity, tail entity, and entity relationship of each text block; An attribute attachment module 200, configured to respectively attach corresponding attribute key-value pairs to the head entity, tail entity, and entity relationship, and respectively obtain an attached head entity, an attached entity relationship, and an attached tail entity; A relationship instance triple construction module 300, configured to generate a relationship instance triple corresponding to each text block based on the attached head entity, entity relationship, and attached tail entity; A structured knowledge graph generation module 400, configured to generate a structured knowledge graph according to the relationship instance triples; A semantic knowledge graph generation module 500, configured to generate a semantic knowledge graph according to the attached head entity, attached tail entity, and attached entity relationship after respectively adding corresponding index information; A reliable knowledge graph generation module 600, configured to respectively select any one of the attached head entity, attached tail entity, and entity relationship in the relationship instance triple as an object to be obscured, infer the object to be obscured according to the other two in the relationship instance triple, generate multiple inferred relationship instance triples and then verify them to generate a reliable knowledge graph.
[0063] A target knowledge graph generation module 700, configured to fuse the structured knowledge graph, semantic knowledge graph, and reliable knowledge graph to generate a target knowledge graph.
[0064] The beneficial effects of a construction and expansion system based on a dynamic semantic knowledge graph provided by the present invention are as follows: First, through the extraction module 100, the unified fusion of multi-data sources and multi-modal information is achieved, forming a structured semantic knowledge graph. When processing text data, it can automatically maintain the context integrity, ensuring the semantic accuracy and structural consistency of the knowledge graph. The redundancy and inconsistency problems between different data blocks are solved through the fusion algorithm, making the generated knowledge graph have high quality and high availability, meeting the knowledge representation requirements in complex scenarios; Second, through the trusted knowledge graph generation module 600, new entity, entity relationship, and attribute information are automatically inferred based on the existing knowledge graph. During the reasoning process, through the strategies of concealment and filling, the model can not only verify the integrity of the existing data but also expand the data volume of the knowledge graph. This method is particularly suitable for fields with insufficient data or incomplete annotation, enabling the semantic expansion of the knowledge graph, significantly improving its coverage and depth, and providing richer data support for subsequent intelligent applications. At the same time, through the trusted knowledge graph generation module 600, combined with the Boolean query and multi-way verification strategy, the correctness and credibility of the reasoning results are ensured. For each newly added reasoning triple, this method can quickly collect the evidence source and perform associated matching based on the semantic similarity of the RAG system, effectively improving the efficiency and accuracy of verification. In the case of insufficient evidence, the reliability of the results is further strengthened through a three-way verification mechanism. This verification method combining the evidence chain enhances the usability and user trust of the knowledge graph and is suitable for scenarios with high requirements for data quality.
[0065] The construction and expansion system of a dynamic semantic knowledge graph in the embodiments of the present application can be a device, or a component, an integrated circuit, or a chip in a terminal. The device can be a mobile electronic device or a non-mobile electronic device. Exemplarily, the mobile electronic device can be a mobile phone, a tablet computer, a laptop computer, a handheld computer, a vehicle-mounted electronic device, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc., and the non-mobile electronic device can be a server, a Network Attached Storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiments of the present application do not make specific limitations.
[0066] A construction and extension system based on a dynamic semantic knowledge graph in an embodiment of the present application can be a device with an operating system. The operating system can be an Android operating system, an iOS operating system, or other possible operating systems, which are not specifically limited in the embodiments of the present application.
[0067] A construction and extension system based on a dynamic semantic knowledge graph provided in an embodiment of the present application can implement Figures 1 to 3 each process implemented by a construction and extension system based on a dynamic semantic knowledge graph in a method embodiment. To avoid repetition, it will not be described in detail here.
[0068] Optionally, an embodiment of the present application further provides an electronic device, including a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, it implements each process of the above method embodiment of a construction and extension method based on a dynamic semantic knowledge graph and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0069] An embodiment of the present application further provides a readable storage medium, on which a program or instruction is stored. When the program or instruction is executed by a processor, it implements each process of the above method embodiment of a construction and extension method based on a dynamic semantic knowledge graph and can achieve the same technical effect. To avoid repetition, it will not be described in detail here.
[0070] Among them, the processor is the processor in the electronic device described in the above embodiment. The readable storage medium includes a computer-readable storage medium, such as a computer read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.
[0071] It should be noted that, in this text, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising that element. In addition, it should be pointed out that the scope of the methods and apparatuses in the embodiments of the present application is not limited to performing functions in the order shown or discussed, but may also include performing functions in a substantially simultaneous manner or in the reverse order according to the functions involved. For example, the described methods may be performed in an order different from that described, and various steps may be added, omitted, or combined. Additionally, the features described with reference to certain examples may be combined in other examples.
[0072] From the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0073] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Those of ordinary skill in the art, under the inspiration of the present application and without departing from the spirit of the present application and the scope protected by the claims, can still make many forms, all of which fall within the protection scope of the present application.
Claims
1. A construction and extension method based on a dynamic semantic knowledge graph, characterized in that, Including: Obtain the head entity, tail entity, and entity relationship of each text block, where the entity relationship represents the semantic entity relationship between the head entity and the tail entity; Attach corresponding sets of attribute key-value pairs to the head entity, the tail entity, and the entity relationship respectively, and obtain an additional head entity, an additional entity relationship, and an additional tail entity respectively. The set of attribute key-value pairs includes multiple attribute keys and attribute values corresponding to the head entity and the tail entity; Generate a relationship instance triple corresponding to each text block based on the additional head entity, the entity relationship, and the additional tail entity; Generate a structured knowledge graph according to the relationship instance triples; Generate a semantic knowledge graph based on the additional head entity, the additional tail entity, and the additional entity relationship after adding corresponding index information respectively; Select any one of the additional head entity, additional tail entity, and entity relationship in the relationship instance triple as the masked object respectively, reason about the masked object based on the other two in the relationship instance triple, generate multiple reasoned relationship instance triples and then verify them to generate a trustworthy knowledge graph; Fuse the structured knowledge graph, the semantic knowledge graph, and the trustworthy knowledge graph to generate a target knowledge graph.
2. A construction and extension method based on a dynamic semantic knowledge graph according to claim 1, characterized in that The specific process of obtaining the relationship instance triples includes: Construct an entity set and an entity relationship set according to the entities and entity relationships of multiple text blocks. The entity set includes multiple head entities and multiple tail entities, and the entity relationship set includes multiple entity relationships; Fuse the entity set based on a preset fusion rule to generate a set of attribute key-value pairs; Obtain the entity type and the set of attribute key-value pairs corresponding to each entity in the entity set according to the head entity and the tail entity respectively; Generate relationship instance triples based on the entity set, the entity relationship set, and the set of attribute key-value pairs.
3. A construction and extension method of a dynamic semantic knowledge graph according to claim 2, characterized in that The preset fusion rule includes: obtaining the entity type, entity name, and attribute value of each text block. When the entity types and entity names are the same but the attribute values are different, determine whether the attribute key corresponding to the attribute value exists. When the attribute key exists, replace the attribute value of the later batch with the attribute value of the previous batch. When the attribute key does not exist, add a new attribute key-value pair according to the updated attribute key and the attribute value of the later batch.
4. A construction and extension method of a dynamic semantic knowledge graph according to claim 1, characterized in that The process of generating the semantic knowledge graph includes: ; Among them, represents the semantic knowledge graph, represents the head entity, represents the tail entity, represents the entity relationship, represents the index of the head entity, represents the index of the entity relationship, represents the index of the tail entity, represents the set of attribute key-value pairs of the head entity, represents the set of attribute key-value pairs of the entity relationship, represents the set of attribute key-value pairs of the tail entity.
5. A construction and extension method of a dynamic semantic knowledge graph according to claim 1, characterized in that The process of reasoning about the masked object includes: Mask the head entity, tail entity, and entity relationship in the relationship instance triple respectively, and obtain multiple relationship instance triples after partial information concealment processing; Use LLM-driven to reason about the missing information in the relationship instance triples after partial information concealment processing, and obtain the reasoned head entity, reasoned entity relationship, and reasoned tail entity; Use LLM-driven to perform attribute reasoning on the reasoned head entity and the reasoned tail entity respectively to obtain a set of reasoned attributes; Combine the set of reasoned attributes, the reasoned head entity, and the reasoned entity relationship to generate multiple reasoned relationship instance triples.
6. A construction and extension method of a dynamic semantic knowledge graph according to claim 1, characterized in that The process of generating a trusted knowledge graph includes: Searching, through a Web search engine, for web content chunks associated with the inferred relation instance triples, where each of the inferred relation instance triples corresponds to a web content chunk; Calculating the cosine similarity between each of the inferred relation instance triples and the corresponding web content chunk; Verifying each of the inferred relation instance triples based on the magnitude relationship between the cosine similarity and a preset threshold to generate multiple verified relation instance triples; Combining the multiple verified relation instance triples to generate a trusted knowledge graph.
7. A construction and extension method based on a dynamic semantic knowledge graph according to claim 6, characterized in that The process of verifying each of the inferred relation instance triples based on the magnitude relationship between the cosine similarity and a preset threshold includes evidence sufficiency verification and multi-way verification. After the evidence sufficiency verification passes, multi-way verification is entered to generate multiple verified relation instance triples. The process of the sufficiency verification includes: The process of the multi-way verification includes: Among them, represents the sufficiency verification result, represents successful sufficiency verification, represents failed sufficiency verification, represents the preset threshold, represents the intersection operation, and represents the vector generated after encoding by the embedding model, represents the cosine similarity between the inferred relational instance triple and a single web page content block, 、 、 represent three different prompt words, represents the inferred relational instance triple, represents the multi-way verification result, represents large language model inference.
8. A construction and extension system based on a dynamic semantic knowledge graph, characterized in that, Includes: An extraction module for obtaining the head entity, tail entity, and entity relation of each text chunk; An attribute attachment module for respectively attaching corresponding sets of attribute key-value pairs to the head entity, the tail entity, and the entity relation to obtain an attached head entity, an attached entity relation, and an attached tail entity; A relation instance triple construction module for generating a relation instance triple corresponding to each text chunk based on the attached head entity, the entity relation, and the attached tail entity; A structured knowledge graph generation module for generating a structured knowledge graph according to the relation instance triples; A semantic knowledge graph generation module for generating a semantic knowledge graph according to the attached head entity, the attached tail entity, and the attached entity relation after respectively adding corresponding index information; A trusted knowledge graph generation module for respectively selecting any one of the attached head entity, attached tail entity, and entity relation within the relation instance triple as an object to be masked, reasoning about the object to be masked based on the other two within the relation instance triple, generating multiple inferred relation instance triples and then verifying them to generate a trusted knowledge graph; A target knowledge graph generation module for fusing the structured knowledge graph, the semantic knowledge graph, and the trusted knowledge graph to generate a target knowledge graph.
9. An electronic device, characterized in that, Includes a processor, a memory, and a program or instruction stored on the memory and executable on the processor. When the program or instruction is executed by the processor, the steps of a method for constructing and expanding a dynamic semantic knowledge graph as described in any one of claims 1-7 are implemented.
10. A readable storage medium, characterized in that, A program or instruction is stored on the readable storage medium. When the program or instruction is executed by a processor, the steps of a method for constructing and expanding a dynamic semantic knowledge graph as described in any one of claims 1-7 are implemented.
Citation Information
Patent Citations
Knowledge graph multi-hop reasoning method based on Transform deep reinforcement learning
CN115455146A
Knowledge graph automatic construction system and working method thereof
CN115618006A
Fusion reasoning method, system and equipment based on large language model and medium
CN117709468A
Knowledge reasoning method based on multi-training global perception
CN118674033A
Composite symbolic and non-symbolic artificial intelligence system for advanced reasoning and semantic search
US20240386015A1
Cited By
Intelligent question and answer implementation method and system based on large model and semantic map
CN120723876A
Inference method, system and equipment of wide constraint large language model and medium
CN121146067A