Incremental knowledge graph generation system and method of large language model
Through the predefined architecture and iterative traversal method of the large language model, a consistent incremental knowledge graph is built, which solves the ambiguity problem in multilingual document processing, and generates a more accurate knowledge graph, which is applied to smart cities, forest fire emergency rescue and smart industrial systems.
Patent Information
- Application Number
- CN202510158486.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-12
- Publication Date
- 2025-07-11
Smart Images

Figure CN120296175A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to text analysis, and in particular to an incremental knowledge graph generation system and method for large language models, which can be used in various applications such as smart city information interaction, forest fire prevention emergency rescue system information networking, and smart industrial system information transmission. Background Art
[0002] With the rapid development of Internet technology, the amount of multimedia data has grown exponentially. However, most of the data is unstructured, and if not effectively utilized, it will lead to a large amount of information loss. This unstructured data lacks a pre-set format, posing a great challenge to traditional data processing methods. Therefore, advanced text understanding and information extraction technologies are needed to effectively analyze and extract meaningful parts from this data.
[0003] A knowledge graph (KG) is an important tool for realizing data structuring and efficient information access. It is a structured knowledge representation form that organizes interrelated information through a graph structure, where entities and relationships are represented as nodes and edges respectively, and it has wide applications in fields such as information retrieval, reasoning, and analysis.
[0004] However, most existing methods rely on specific topics, there is no fixed predefined pattern, the applicability is low, and it is difficult to process multi-language knowledge files. The processing paradigm of large language models is mainly strictly constrained, and the generated entities are ambiguous. Summary of the Invention
[0005] To solve the above problems existing in the prior art, the present disclosure proposes a text incremental knowledge graph generation system, which is a system that gradually constructs a consistent incremental knowledge graph from multi-language initial documents using a large language model.
[0006] Specifically, an incremental knowledge graph generation system for a large language model, the system includes a text semantic block extraction and unification module, a multi-layer entity and relationship extraction triple module, and a graph generation module; wherein: the text semantic block extraction and unification module is configured to be implemented based on a large language model, which takes a text segment as input, performs unified language processing on the text segment, and according to a preset architecture paradigm, forms a set of semantic blocks with the required information through a prompt (prompt); the multi-layer entity and relationship extraction triple module is configured to iteratively traverse the set of semantic blocks to extract entities and iteratively traverse the set of semantic blocks to extract the relationships between entities, thereby obtaining triples, and the triples are represented as (entity i, entity j, the relationship between entity i and entity j); the graph generation module is configured to construct a knowledge graph based on the triples.
[0007] In one implementation of the above technical solution, the unified language is preferably English.
[0008] In one implementation of the above technical solution, the large language model is preferably llama3.
[0009] In one implementation of the above technical solution, the step of iteratively traversing the semantic block set to extract entities includes: taking out a semantic block d0 from the semantic block set, and extracting entities with unique concepts based on the semantic block d0 to form a global entity set E; Step P1: taking out a semantic block from the remaining semantic blocks as the current semantic block, and extracting entities with unique concepts from the current semantic block to form a local entity set E d ; For each entity e d in E i , if e i belongs to E, then add ei to the entity matching set E {d,matched} ; Otherwise, obtain its similarity with each entity e j in E. If the similarity between e i and e j is greater than the set threshold Threshold, then put the entity e j into the set ; If the set is an empty set, then add the entity e i to the entity matching set E {d,matched} ; Otherwise, select an entity with the highest similarity from E and add it to the entity matching set E {d,matched} ; After processing the current semantic block, merge the entity matching set E {d,matched} and the global entity set E to form a new global entity set. If there are still semantic blocks in the remaining semantic blocks, return to step P1.
[0010] In one implementation of the above technical solution, the step of iteratively traversing the semantic block set to extract the relationships between entities includes: based on the global entity set E, obtaining the relationships between pairwise entities in E to form a global relationship set R; Step P2: taking out a semantic block from the remaining semantic blocks except the semantic block d0 as the current semantic block, and extracting a local relationship set R {d,matched} based on the current semantic block and the entity matching set E d ; For each relationship r d in R i , if it belongs to R, then add it to the relationship matching set R {d,matched} ; Otherwise, obtain its similarity with each relationship r j in R. If the similarity between r i and r j is greater than the set threshold Threshold, then put the relationship r j into the set ; If the set is an empty set, then add the relationship ri Add it to the relationship matching set R {d,matched} ; Otherwise, select a relationship with the highest similarity from R and add it to the relationship matching set R {d,matched} ; After processing the current semantic block, merge the relationship matching set R {d,matched} with the global relationship set R to form a new global relationship set. If there are still semantic blocks in the remaining semantic blocks, return to step P2.
[0011] In one implementation of the above technical solution, the knowledge graph uses a large language model to output a specified language representation and generates a graph file for visual display.
[0012] In one implementation of the above technical solution, the nodes and relationships in the knowledge graph are displayed in the form of search.
[0013] According to the above system technical solution, correspondingly, the present disclosure proposes an incremental knowledge graph generation method for a large language model. The steps include: taking a text segment as the input of the large language model, performing unified language processing on the text segment, and forming a set of semantic blocks for the required information according to a preset architecture paradigm through prompts (prompts); iteratively traversing the set of semantic blocks to extract entities, and iteratively traversing the set of semantic blocks to extract the relationships between entities, so as to obtain triples, where the triples are expressed as (entity i, entity j, the relationship between entity i and entity j); constructing a knowledge graph based on the triples.
[0014] The beneficial technical effects of the present disclosure are as follows: First, when extracting document semantic blocks, the text language is first unified and standardized to reduce the understanding deviation of the large model. Second, the preset architecture paradigm of the large language model is used to guide the large language model to extract text semantic blocks in related fields through prompts. Third, when the present disclosure extracts global text entities and text relationships, an iterative traversal method is used to generate multiple times to solve possible ambiguities. By generating multiple times, it is ensured that each entity is clearly defined and distinguished from other entities, thereby reducing irrelevant noise. This solution is applied in aspects such as smart city information interaction, forest fire prevention emergency rescue system information networking, and smart industrial system information transmission. Description of the Drawings
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 、 oneSchematic diagram of the principle of the knowledge graph generation system based on the large language model in a certain implementation manner.
[0017] Figure 2 , one Schematic diagram of extracting pseudo code from a document in a certain implementation manner.
[0018] Figure 3 , one Schematic diagram of extracting pseudo code of entity relationships in a document in a certain implementation manner.
[0019] Figure 4 , one Schematic diagram of a part of the visual result of the experiment in a certain implementation manner. Specific implementation manner
[0020] On the one hand, most of the existing technologies rely on specific knowledge topics and do not have a fixed predefined pattern, making it difficult to adapt to other knowledge topics. This leads to poor entity resolution and relationship extraction effects. On the other hand, most of the existing technologies are difficult to process multi-language knowledge files. The processing paradigm of the large language model requires strict constraints, and the generated entities may be ambiguous.
[0021] In view of the above two shortcomings, the present disclosure proposes an incremental knowledge graph generation system based on the large language model. Different from the existing technologies, the solution of the present disclosure utilizes the strong English processing ability of the large language model. By making the large language model follow a predefined architecture, the initial document is standardized into a unified English language and reorganized into multiple semantic blocks. This architecture is similar to a predefined JSON structure, guiding the large language model to extract text information related to specific keys from each document. The processing of the unified language reduces the ambiguity caused by text block segmentation. When processing the semantic blocks, unique semantic entities are identified and possible ambiguities are resolved. Multiple generations are performed to ensure that each entity is clearly defined and distinguished from other entities.
[0022] Next, the implementation of the technical solution of this case will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described implementation manners are only a part of the implementation manners of this case, rather than all of the implementation manners.
[0023] See Figure 1 , an incremental knowledge graph generation system of a large language model, which sequentially processes the input text fragments through a text semantic block extraction and unification module, a multi-layer entity and relationship extraction module, and a graph generation module, and outputs the constructed knowledge graph.
[0024] (1) Text semantic block extraction and unification module
[0025] The text semantic block extraction and unification module is configured to use a large language model to reorganize the initial document into multiple semantic blocks according to a predefined architecture. Exemplarily, the large language model is the llama3 version.
[0026] Those skilled in the art are familiar that a predefined JSON structure is usually a fixed data structure defined in JSON data for easy exchange and storage during application. The predefined architecture of the present disclosure is similar to the predefined JSON structure.
[0027] The predefined architecture can be used to guide the large language model to extract text information related to specific keys from each document, and at the same time uniformly process this text information into English. Through the predefined architecture, the large language model is guided to be biased towards specific categories while also ensuring a certain degree of flexibility in other categories, so as to achieve the unified standardization of the language of each document through the large language model, and the problem of different document languages can be solved.
[0028] The specific key is the text in a specific field proposed in the large language model prompt, and the extracted text information is used as the specific value of the specific key.
[0029] For each document, if the required information exists in the document, we will obtain a partially filled JSON. Then, we aggregate all these partially filled JSONs to form the semantic blocks of the document. We used the JSON Parser tool of Langchain to define the schema. The main objectives of the module are: improving the signal-to-noise ratio, reducing the redundant information noise that may affect the knowledge graph, and guiding the graph construction using the predefined architecture paradigm of the large language model.
[0030] (2) Multi-layer entity and relationship extraction triple module
[0031] In the multi-level entity and relationship extraction module, first iterate through all semantic blocks to extract global document entities. The pseudocode for extraction is shown in Figure 2 as shown.
[0032] Specifically, entities are extracted from the first semantic block d0 using the large language model to form a global entity set E. This solution assumes that entities are pairwise independent and its constraint condition (C1) is that we prompt the large language model to extract only entities representing unique concepts to avoid semantic mixing. The entity of the unique concept means that the name of the entity is unique.
[0033] For subsequent semantic blocks d ∈ D, where D is the document, use the algorithm to extract the local entity set E d . Then try to match these local entities with the global entity set E. If a certain local entity e i is found in E, then add it to the matching set E{d,matched} If not found, the algorithm will use cosine similarity measurement and combine with a predefined threshold to find entities with lower similarity in E. If there is still no match, the local entity will be directly added to E {d,matched} ; if none of the above cases occur, the global entity e′ with the highest similarity to it i will be added to the matching set.
[0034] Subsequently, the global entity set E is updated by merging E with E {d,matched} . This process is repeated for each document in D, and finally a comprehensive global entity set is formed.
[0035] Process the semantic chunks, identify the unique semantic entities therein, and resolve possible ambiguities by generating multiple times to ensure that each entity is clearly defined and distinguished from other entities.
[0036] The global entity set E is provided as context to the incremental relation extraction module and used together with each semantic chunk to extract the global relation set R, see the Figure 3 pseudo-code shown. By detecting semantically unique relations, entity connections are formed, and then triples are obtained, which are represented as (entity i, entity j, the relation between entity i and entity j). Entities i and j are pairwise different entities in the set of semantic chunks.
[0037] We observe that the behavior of relation extraction varies depending on whether global entities or local entities are used as the context for semantic chunks. When global entities are used as the context, the relations extracted by the large language model include not only the relations explicitly stated in the semantic chunks but also the implicit relations, especially for entities not explicitly present in the semantic chunks. This approach can enrich the potential information in the graph but also increases the possibility of irrelevant relations. On the contrary, when using locally matched entities as the context, the large language model only extracts the relations explicitly stated in the context. This approach reduces the richness of the graph but also decreases the probability of irrelevant relations.
[0038] (3) Graph Generation Module
[0039] Finally, based on the triples obtained from the global entities and relations, a knowledge graph is constructed, and the large language model is used to output the specified language representation, such as uniformly translating the entities and relations into Chinese and generating a graph file, which is saved in html format for visual display. The graph nodes and relations can be highlighted through search.
[0040] Verify the above technical solution on the text data set related to smart cities, and the expected results are obtained, see the Figure 4 partial visualization results shown. In Figure 4Among them, there are atlas nodes such as Smart City and Smart City Initiative. Entities with multiple relationships are regarded as nodes. Among the entities related to the entity Smart City Initiative, there are research, key success factors, efficiency improvement, overall approach, urban vision, stakeholder reference, Giffinger's model, etc.
[0041] In summary, in view of the problems existing in the prior art in Incremental KnowledgeGraphs Construction, namely the large language model paradigm limitation on the scope of knowledge topics and the ambiguity of generated entities, the present disclosure proposes a knowledge graph construction solution for iteratively traversing and generating triples based on the preset paradigm of the large language model, which is a solution for gradually constructing a consistent incremental knowledge graph from multilingual initial documents using the large language model. Specifically, given some texts collected by a collection device, this solution analyzes the content in the texts through a designed algorithm system, uses the large language model to identify the relationships between various entities in the texts, and establishes a knowledge graph network composed of multiple entity nodes and multiple relationship edges.
[0042] Through the description of the above embodiments, those skilled in the art can clearly understand that through the system of the present disclosure, a method for generating an incremental knowledge graph of a large language model can be correspondingly proposed. The steps include: taking a text segment as the input of the large language model, performing unified language processing on the text segment, and forming a set of semantic chunks for the required information through prompting according to the preset architecture paradigm; iteratively traversing the set of semantic chunks to extract entities, and iteratively traversing the set of semantic chunks to extract the relationships between entities, so as to obtain triples, where the triples are represented as (entity i, entity j, the relationship between entity i and entity j); constructing a knowledge graph based on the triples.
[0043] Through the description of the above embodiments, those skilled in the art can clearly understand that the system and method of the present disclosure can be implemented by means of software plus necessary general-purpose hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions accomplished by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, in more cases for the present disclosure, software program implementation is a better implementation manner.
[0044] Although the embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, and these all fall within the scope of protection of the present disclosure.
Claims
1. An incremental knowledge graph generation system for large language models, characterized in that, The system includes a text semantic block extraction and unification module, a multi-layer entity and relationship extraction triple module, and a graph generation module; where: The text semantic block extraction and unification module is configured to be implemented based on a large language model. It takes a text fragment as input, performs unified language processing on the text fragment, and according to a preset architecture paradigm, forms a semantic block set of required information through prompts. The multi-layer entity and relationship extraction triple module is configured to iteratively traverse the semantic block set to extract entities and iteratively traverse the semantic block set to extract the relationships between entities, thereby obtaining triples, where the triples are represented as (entity i, entity j, the relationship between entity i and entity j). The graph generation module is configured to construct a knowledge graph based on the triples.
2. The system according to claim 1, wherein, The unified language is preferably English.
3. The system according to claim 1, characterized in that, The large language model is preferably llama3.
4. The system according to claim 1, wherein The steps of iteratively traversing the semantic block set to extract entities include: Taking out a semantic block d0 from the semantic block set, and extracting entities with unique concepts based on the semantic block d0 to form a global entity set E. Step P1: Take a semantic chunk from the remaining semantic chunks as the current semantic chunk, and extract the entities of the unique concepts from the current semantic chunk to form the local entity set E d ; For E d For each entity e i in it, if e i belongs to E, then add e i to the entity matching set E {d,matched} ; otherwise, obtain its similarity with each entity e j in E. If the similarity between e i and e j is greater than the set threshold Threshold, then put the entity e j into the set ; If the set is an empty set, then add the entity e i to the entity matching set E {d,matched} ; otherwise, select an entity with the highest similarity from E and add it to the entity matching set E {d,matched} ; After processing the current semantic block, merge the entity matching set E {d,matched} with the global entity set E to form a new global entity set. If there are still semantic blocks in the remaining semantic blocks, return to step P1.
5. The system according to claim 4, characterized in that, The steps of iteratively traversing the semantic block set to extract the relationships between entities include: Based on the global entity set E, obtaining the relationships between pairwise entities in E to form a global relationship set R. Step P2: Take a semantic chunk from the remaining semantic chunks after removing the semantic chunk d0 as the current semantic chunk, and extract the local relationship set R based on the current semantic chunk and the entity matching set E {d,matched} Extract the local relationship set R d ; For R d For each relationship r i in it, if it belongs to R, add it to the relationship matching set R {d,matched} ; otherwise, obtain its similarity with each relationship r j in R. If the similarity between r i and r j is greater than the set threshold Threshold, then put the relationship r j into the set ; If the set is an empty set, then add the relationship r i to the relationship matching set R {d,matched} ; otherwise, select a relationship with the highest similarity from R and add it to the relationship matching set R {d,matched} ; After processing the current semantic block, merge the relation matching set R {d,matched} with the global relation set R to form a new global relation set. If there are still semantic blocks in the remaining semantic blocks, return to step P2.
6. The system according to claim 1, wherein The knowledge graph uses the large language model to output a specified language representation and generates a graph file for visual display.
7. The system according to claim 1, wherein The nodes and relationships in the knowledge graph are displayed in the form of search.
8. An incremental knowledge graph generation method for large language models, characterized in that, The method includes the following steps: Taking the text fragment as the input of the large language model, performing unified language processing on the text fragment, and according to a preset architecture paradigm, forming a semantic block set of required information through prompts. Iteratively traversing the semantic block set to extract entities and iteratively traversing the semantic block set to extract the relationships between entities, thereby obtaining triples, where the triples are represented as (entity i, entity j, the relationship between entity i and entity j). Constructing a knowledge graph based on the triples.
9. The method according to claim 8, wherein The steps of iteratively traversing the semantic block set to extract entities include: Taking out a semantic block d0 from the semantic block set, and extracting entities with unique concepts based on the semantic block d0 to form a global entity set E. Step P1: Take a semantic chunk from the remaining semantic chunks as the current semantic chunk, and extract the entities with unique concepts from the current semantic chunk to form a local entity set E d ; For E d For each entity e i , if e i belongs to E, then e i Add entity matching set E {d,matched} Otherwise, get its corresponding entity e in E. j The similarity of i With e j If the similarity is greater than the set threshold, entity e j Add to collection middle; If the set is an empty set, then add the entity e i to the entity matching set E {d,matched} ; otherwise, select the entity with the highest similarity from E and add it to the entity matching set E {d,matched} ; After processing the current semantic block, merge the entity matching set E {d,matched} with the global entity set E to form a new global entity set. If there are still semantic blocks in the remaining semantic blocks, return to step P1.
10. The method according to claim 9, wherein The steps of iteratively traversing the semantic block set to extract the relationships between entities include: Based on the global entity set E, obtaining the relationships between pairwise entities in E to form a global relationship set R. Step P2: Take a semantic chunk from the remaining semantic chunks after removing the semantic chunk d0 as the current semantic chunk, and extract the local relationship set R based on the current semantic chunk and the entity matching set E {d,matched} Extract the local relationship set R d ; For R d For each relationship r i in it, if it belongs to R, add it to the relationship matching set R {d,matched} ; otherwise, obtain its similarity with each relationship r j in R. If the similarity between r i and r j is greater than the set threshold Threshold, then put the relationship r j into the set ; If the set is an empty set, then add the relationship r i to the relationship matching set R {d,matched} ; otherwise, select a relationship with the highest similarity from R and add it to the relationship matching set R {d,matched} ; After processing the current semantic block, merge the relationship matching set R {d,matched} with the global relationship set R to form a new global relationship set. If there are still semantic blocks in the remaining semantic blocks, return to step P2.
Citation Information
Cited By
Knowledge graph generation method and device, storage medium and electronic equipment
CN121365722A
Knowledge graph generation methods, devices, storage media and electronic equipment
CN121365722B