Multi-stage serial knowledge graph construction method and system based on cue words
By adopting multi-stage serial extraction chains and dynamic text chunking strategies in the construction of knowledge graphs, the problem of dependence on single prompt words and model parameters in the existing technology is solved, and more efficient and accurate knowledge extraction is achieved, improving the quality and flexibility of knowledge graph construction.
Patent Information
- Application Number
- CN202510023083.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-06-10
AI Technical Summary
The existence of a single prompt word in the construction of knowledge graphs in the existing technology leads to poor model performance, large dependence on model parameter volume, and unreasonable text blocking strategies, resulting in a decrease in the accuracy and consistency of knowledge extraction.
The multi-stage serial knowledge graph construction method based on prompt words is adopted, and the knowledge extraction task is decomposed into entity recognition, attribute completion and relationship extraction stages through dynamic text chunking and multi-stage serial extraction chain to ensure the accuracy and consistency of the extraction results.
It significantly improves the accuracy and consistency of knowledge extraction, reduces dependence on model complexity, enhances the flexibility and controllability of the system, and ensures the high reliability and integrity of structured data.
Smart Images

Figure CN120124729A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of artificial intelligence and natural language processing. Specifically, it particularly relates to a knowledge graph construction method and system based on a multi-stage serial extraction chain and a dynamic text chunking strategy. Background Art
[0002] In the era of information explosion, efficiently extracting and organizing valuable knowledge from a vast amount of unstructured data has become an important research direction in the fields of natural language processing and artificial intelligence. As a powerful knowledge representation form, a knowledge graph expresses entities, concepts, and the relationships between them through a graph structure, and demonstrates significant application value in many fields such as information retrieval, intelligent question answering, recommendation systems, medical diagnosis, and financial analysis.
[0003] Traditional knowledge graph construction methods mainly rely on complex rule sets or a large amount of labeled data, and have the following limitations: 1. Strong rule dependence. Traditional methods require manual design of complex rules to extract entities and relationships. This process is not only time-consuming and laborious, but also difficult to cope with the diversity of natural language expressions. 2. High data annotation cost. Supervised learning-based methods require a large amount of high-quality labeled data, and the process of obtaining this data is complex and costly. Especially in vertical fields, it often requires the in-depth participation of domain experts. 3. Limited model generalization ability. Existing methods have insufficient generalization ability when dealing with cross-domain or complex texts, and it is difficult to ensure the accuracy and consistency of extraction results.
[0004] In recent years, with the development of large language models, a new knowledge graph construction method has gradually attracted attention. This method generally includes the following key steps: 1. Text preprocessing. First, clean and format the original text to remove noise data and ensure the quality of the input text. 2. Text chunking. Divide the preprocessed text into appropriately sized chunks so that the large language model can process it efficiently. 3. Prompt design. Design a prompt template for knowledge extraction to guide the large language model to complete tasks such as entity extraction, attribute extraction, and relationship extraction. 4. Knowledge extraction. The large language model extracts entities, attributes, and relationships from the text chunks according to the prompts to generate preliminary knowledge triples. 5. Knowledge fusion and verification. Fusion and verification of the extracted knowledge triples are performed to eliminate redundancy and conflicts and ensure the consistency and accuracy of the knowledge graph.
[0005] Although the method for constructing a knowledge graph based on large language models has significant advantages in terms of flexibility and efficiency, the existing technologies still face the following challenges: 1. Single prompt words lead to poor model performance. Using a single prompt word to complete multiple different tasks (such as entity extraction, relationship extraction, etc.) easily distracts the model's attention and affects the performance of each task. For example, when performing entity recognition and relationship extraction simultaneously, the model may not be able to fully focus on the specific requirements of each task, resulting in a decrease in the accuracy and consistency of the extraction results. 2. High dependence on model parameters. Constructing a high-quality knowledge graph usually requires a large number of model parameters, which poses challenges to resource-constrained scenarios, not only increasing the training time cost but also the deployment and maintenance costs. For example, in resource-limited environments such as edge computing or mobile devices, the operating efficiency of large-scale models may be significantly reduced, making it difficult to meet real-time requirements. 3. Unreasonable text chunking strategy. Existing text chunking strategies usually adopt hard-coded chunk sizes, which may result in too many or too few chunks, thus affecting the large language model's understanding and processing of context and reducing the quality of knowledge graph construction. For example, when the chunks are too small, the model may not be able to capture cross-chunk semantic relationships; while when the chunks are too large, the model's computational burden will increase, affecting the processing efficiency.
[0006] In summary, although the method for constructing a knowledge graph based on large language models has significant advantages in terms of flexibility and efficiency, the existing technologies still have many limitations, restricting the accuracy and application scope of knowledge extraction. Therefore, there is an urgent need for a more optimized method to address the above challenges to achieve more efficient and accurate knowledge graph construction. Summary of the Invention
[0007] The purpose of the present invention is to overcome the defects and deficiencies of the existing technologies and provide a multi-stage serial knowledge graph construction method and system based on prompt words. By dynamically calculating the size of text chunks through dynamic text chunking, the situation of too many or too few chunks is reduced, improving the efficiency and accuracy of knowledge extraction. The knowledge extraction tasks are decomposed into an entity recognition stage, an attribute completion stage, and a relationship extraction stage through a multi-stage serial extraction chain to ensure the accuracy and consistency of the extraction results.
[0008] To achieve the above purpose, the technical solutions adopted by the present invention are as follows: A multi-stage serial knowledge graph construction method based on prompt words, comprising the following steps: Pre-configuration: Initialize various parameters and resources required by the system through large language model configuration, prompt word set configuration, and example set configuration; Knowledge extraction: Achieve accurate knowledge extraction through dynamic text chunking and a multi-stage serial extraction chain, extract entities, attributes, and relationships from the text, and generate preliminary knowledge triples; Knowledge Graph Construction: Through database configuration, knowledge storage and construction, as well as interfaces and queries, organize knowledge triples into a structured knowledge graph and achieve querying and visual display.
[0009] Furthermore, the configuration of the large language model includes the following steps: Model Selection and Loading: Select a suitable large language model according to the task requirements and load the model parameters; Parameter Setting: Configure the inference parameters of the model, specifically including context length, temperature value, maximum output length, presence penalty value, and confidence call; Model Optimization: Adjust the computational resource allocation of the model according to the hardware resources.
[0010] Furthermore, the configuration of the prompt set includes the following steps: Prompt Template Design: Design dedicated prompt templates for each knowledge extraction stage; Prompt Storage: Store the optimized prompt templates in the system; Prompt Version Iteration Management: Conduct version iteration management of the prompt templates through experiments and evaluations, guide the large language model to complete tasks, and continuously optimize the prompts.
[0011] Furthermore, adopt few-shot learning to configure the example set, including the following steps; Example Design and Collection: Design example data for each knowledge extraction stage to help the large language model understand the task requirements; Example Storage: Store the optimized example data in the system; Example Set Version Iteration Management: Dynamically adjust the example content according to the task requirements and improve the model's performance through version iteration management.
[0012] Furthermore, dynamic text chunking includes the following steps; Text Preprocessing: Clean and format the original text, remove noise data, and ensure the quality of the input text; Chunk Size Calculation: Dynamically calculate the text chunk size according to the maximum context length supported by the large language model and the length of the prompt. The calculation formula is:
[0013] In the formula, is the dynamically calculated text chunk size; is the maximum context length supported by the large language model; is the length of the prompt; is the empirical coefficient, used to reserve a part of the context length for the large language model to perform inference; Text Chunk Optimization: Divide the preprocessed text into multiple text chunks of dynamic sizes, ensuring that the length of each text chunk does not exceed the dynamically calculated chunk size while avoiding semantic breaks; Chunk Input: Input the optimized text chunks into the large language model for processing.
[0014] Furthermore, the multi-stage serial extraction chain decomposes the knowledge extraction task into an entity recognition stage, an attribute completion stage, and a relationship extraction stage, including the following steps: In the entity recognition stage, use prompt templates to guide the large language model to extract entities from the text, including rough extraction, consistency evaluation, entity review, final quality inspection, and output links; In the attribute completion stage, based on the output of the entity recognition stage, use prompt templates to guide the large language model to complete the attribute information for entities, including rough extraction, consistency evaluation, attribute review, final quality inspection, and output links; In the relationship extraction stage, based on the output of the attribute completion stage, use prompt templates to guide the large language model to extract the relationships between entities, including rough extraction, consistency evaluation, relationship review, final quality inspection, and output links.
[0015] Furthermore, database configuration includes the following steps: Database Selection: Select a database system that supports efficient storage and query of graph-structured data; Database Initialization: Set up the operating environment of the database, including specifying the location of data storage and configuring the caching mechanism to ensure the high efficiency of the database operation; Data Model Definition: Define the data structure in the database according to the results of knowledge extraction, including determining different types of data elements and their attributes, as well as the association relationships between data elements.
[0016] Furthermore, knowledge storage and construction include the following steps: Data Import: Import the knowledge triple data extracted from the text into the database to establish the corresponding data elements and association relationships; Data Verification: Check the imported data to ensure the integrity and consistency of data elements and association relationships and avoid data errors; Index Optimization: Create indexes for data elements and association relationships in the database to improve the efficiency of data query.
[0017] Furthermore, interfaces and queries include the following steps: Interface Definition: Provide a set of standardized interfaces for external systems to query and operate on the knowledge graph; Query Optimization: Optimize the query process, support complex query requirements, and ensure that the query can respond in real time; Visualization display: Provide visualization tools to support users to view data elements and association relationships in the knowledge graph in a dynamic and interactive manner.
[0018] A multi-stage serial knowledge graph construction system based on prompting words, applying the multi-stage serial knowledge graph construction method described in any one of the above, includes: A pre-configuration module, a knowledge extraction module, and a graph construction module. The pre-configuration module is used to initialize various parameters and resources required by the system. The knowledge extraction module is used to extract entities, attributes, and relationships from the text to generate preliminary knowledge triples. The graph construction module is used to organize the knowledge triples into a structured knowledge graph and implement query and visualization display; The pre-configuration module includes a large language model configuration, a prompt set configuration, and an example set configuration sub-module. The large language model configuration sub-module is used for model selection and loading, parameter setting, and model optimization. The prompt set configuration sub-module is used for prompt design, prompt storage, and prompt version iteration management. The example set configuration sub-module is used for example design, example storage, and example version iteration management; The knowledge extraction module includes a dynamic text chunking and multi-stage serial extraction chain sub-module. The dynamic text chunking sub-module is used to dynamically calculate the size of text chunks. The multi-stage serial extraction chain sub-module is used to decompose the knowledge extraction task into an entity recognition stage, an attribute completion stage, and a relationship extraction stage; The graph construction module includes a database configuration, knowledge storage and construction, and an interface and query sub-module. The database configuration sub-module is used for database selection, database initialization, and data model definition. The knowledge storage and construction sub-module is used for data import, data verification, and index optimization. The interface and query sub-module is used for interface definition, query optimization, and visualization display.
[0019] Compared with the prior art, the present invention has the following advantages and beneficial effects: The present invention improves the quality of structured data construction. By adopting a multi-stage serial extraction chain and a dynamic text chunking strategy, the accuracy and consistency of entity, attribute, and their relationship extraction are significantly improved. The present invention can effectively reduce errors and ambiguities in the data extraction process, ensuring that the finally generated structured data has higher reliability and integrity.
[0020] The present invention reduces the dependence on model complexity. Through modular design and few-shot learning strategies, the dependence on large-scale model parameters is reduced. Even using a general model, high-quality structured data can be constructed, reducing the complexity of technical implementation and resource consumption.
[0021] The present invention enhances flexibility and controllability. The modular design enables the prompt words at each stage to be configured and optimized individually, thereby improving the flexibility and controllability of the entire data extraction and construction process. Users can adjust the parameters and rules at each stage according to specific requirements to adapt to different application scenarios and data characteristics. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a schematic flowchart of a multi-stage serial knowledge graph construction method based on prompt words.
[0023] Figure 2 It is a schematic structural diagram of a multi-stage serial extraction chain proposed by the present invention.
[0024] Figure 3 It is a schematic structural diagram of a multi-stage serial knowledge graph construction system based on prompt words. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following further describes the multi-stage serial knowledge graph construction method and system based on prompt words of the present invention with reference to the accompanying drawings and specific embodiments.
[0026] Please refer to Figure 1 , the present invention discloses a multi-stage serial knowledge graph construction method based on prompt words, including the following steps: Pre-configuration: Initialize various parameters and resources required by the system through large language model configuration, prompt word set configuration, and example set configuration; Knowledge extraction: Achieve accurate knowledge extraction through dynamic text chunking and multi-stage serial extraction chain, extract entities, attributes, and relationships from the text, and generate preliminary knowledge triples; Graph construction: Organize the knowledge triples into a structured knowledge graph through database configuration, knowledge storage and construction, and interfaces and queries, and implement query and visualization display.
[0027] In the pre-configuration step, initialize various parameters and resources required by the system through large language model configuration, prompt word set configuration, and example set configuration to ensure the smooth progress of the knowledge extraction process.
[0028] Specifically, the large language model configuration includes the following steps: Model selection and loading: Select a suitable large language model according to the task requirements and load the model parameters. For example, select a model that supports long text processing and high-precision inference, such as GPT-4, etc. The loading of model parameters can be achieved through application programming interface calls or local file reading, etc.
[0029] Parameter setting: Configure the inference parameters of the model, specifically including context length, temperature value, maximum output length, presence penalty value, and confidence call.
[0030] The context length is used to set the maximum context length for the model to process text, such as 1024 characters, to ensure that the model can fully understand the semantic information of long texts.
[0031] The temperature value adjusts the diversity of the generated content. A lower temperature value (such as 0.2) makes the output more deterministic and concentrated, while a higher temperature value (such as 0.8) increases the randomness and diversity of the output.
[0032] The maximum output length is used to limit the length of the content generated by the model, such as 256 characters, to avoid generating overly long or irrelevant text.
[0033] The presence penalty value reduces the occurrence of repeated words or phrases in the generated content, such as setting it to 1.5, to enhance the diversity and quality of the generated content.
[0034] The trust call technology improves the reliability and accuracy of the output results by automatically detecting and fixing the error responses generated by the model. It introduces an additional verification mechanism to ensure that every part of the content generated by the model meets the expected requirements.
[0035] Model optimization: Adjust the computational resource allocation of the model according to the hardware resources, such as using distributed computing or model pruning techniques, to improve the running efficiency.
[0036] Specifically, the configuration of the prompt set includes the following steps: Prompt template design: Design dedicated prompt templates for each knowledge extraction stage (entity recognition, attribute completion, relation extraction). For example, the prompt template for the entity recognition stage is: "Please extract all entities from the following text and give their types." In prompt design, prompt engineering techniques are adopted to guide the model to reason step by step through appropriate prompts to ensure high-quality output. At the same time, the chain of thought technique is introduced to enable the model to solve problems step by step, improving the accuracy and logic of reasoning.
[0037] Prompt storage: Store the optimized prompt templates in the system, which can be done in various ways such as using a database, file system, or in-memory cache.
[0038] Prompt version iteration management: Conduct version iteration management of the prompt templates through experiments and evaluations to ensure that they can accurately guide the large language model to complete tasks and continuously optimize the effectiveness and advancement of the prompts.
[0039] Specifically, few-shot learning is adopted for the configuration of the example set, including the following steps; Example Design and Collection: Design example data for each knowledge extraction stage to help the large language model better understand the task requirements. For example, an example for the entity recognition stage is: "Text: 'John studies at the Massachusetts Institute of Technology.', Entities: 'John' (person), 'Massachusetts Institute of Technology' (organization).". Example Storage: Store the optimized example data in the system, which can be done in various ways such as using a database, file system, or in-memory cache.
[0040] Iterative Management of Example Set Versions: Dynamically adjust the example content according to the task requirements, and improve the model's performance through iterative management of versions. For example, regularly update the example set and add new examples to cover more scenarios.
[0041] To further improve the quality of knowledge extraction, the method proposed in the present invention allows adding examples in the prompt words and adopts a few-shot learning strategy. By providing a small number of representative examples, it helps the large language model better understand the task requirements and improve the quality of knowledge extraction. This strategy can effectively improve the model's performance in low-resource scenarios and reduce the dependence on model parameters and pre-trained data.
[0042] Knowledge Extraction Steps: Achieve efficient and accurate knowledge extraction through dynamic text chunking and multi-stage serial extraction chains. To improve the context utilization rate of the large language model, the present invention proposes a dynamic text chunking strategy that dynamically calculates the size of text chunks based on the maximum context window length supported by the large language model and the length of the prompt words. This strategy avoids the problems caused by traditional hard-coded chunk sizes, can make full use of the performance of the large language model, reduce the situation of too many or too few chunks, and thus improve the efficiency and accuracy of knowledge extraction.
[0043] Specifically, dynamic text chunking includes the following steps; Text Preprocessing: Clean and format the original text, remove noisy data, and ensure the quality of the input text. The specific steps include removing HTML tags, converting to lowercase, removing special characters, etc.
[0044] Chunk Size Calculation: Dynamically calculate the size of text chunks based on the maximum context length supported by the large language model and the length of the prompt words. The calculation formula is:
[0045] In the formula, is the dynamically calculated text chunk size; is the maximum context length supported by the large language model; is the dynamically calculated length of the prompt words; is an empirical coefficient used to reserve a part of the context length for the large language model to perform reasoning.
[0046] Text Chunk Optimization: The preprocessed text is divided into multiple text chunks of dynamic sizes, ensuring that the length of each text chunk does not exceed the dynamically calculated chunk size while avoiding semantic breaks. Specific methods include the sliding window method, the fixed-size chunking method, etc.
[0047] Chunk Input: The optimized text chunks are input into the large language model for processing. Techniques such as batch processing or asynchronous processing can be used to improve processing efficiency.
[0048] Please refer to Figure 2 , the multi-stage serial extraction chain decomposes the knowledge extraction task into an entity recognition stage, an attribute completion stage, and a relationship extraction stage. Each stage includes multiple parallel rough extractions, self-consistency evaluations, and quality inspection links to ensure the accuracy and consistency of the extraction results.
[0049] In the entity recognition stage, prompt templates are used to guide the large language model to extract entities from the text, including rough extraction, consistency evaluation, entity review, final quality inspection, and output links.
[0050] Rough Extraction: The large language model initially extracts named entities and their types from the text according to preset prompt templates. To improve the coverage and diversity of extraction, a multiple parallel rough extraction strategy is adopted to generate multiple groups of candidate named entities and their types. This strategy can effectively avoid omissions or biases that may be caused by a single extraction path, and improve the comprehensiveness and accuracy of the extraction results. Specific methods include using multiple different prompt templates, calling the model multiple times, etc.
[0051] Consistency Evaluation: Based on multiple groups of candidate named entities obtained from rough extraction, a majority voting method is used for consistency evaluation. Specifically, the occurrence frequencies of each candidate named entity in multiple groups of extraction results are counted, and the named entity and its type with the highest occurrence frequency are selected as the result of consistency evaluation. For example, if an entity is recognized as "person" 8 times in 10 parallel extractions, then "person" is selected as the type of this entity.
[0052] This method can significantly reduce the impact of accidental errors and improve the robustness and reliability of the extraction results. For candidate named entities that do not pass the consistency evaluation, they enter a separate review process, and further decisions are made in combination with the content of the text materials to determine whether to enter the final quality inspection stage.
[0053] Entity Review: For candidate named entities that do not pass the consistency evaluation, they enter a separate review process. Find the complete sentence or sentence fragment corresponding to the named entity in the context to ensure that sufficient context information is provided. Analyze and judge whether the naming and classification of the named entity are correct, and give the modification conclusion for this named entity. Return the analysis process and modification conclusion for each named entity. If the conclusion is that no modification is required, the result should be the same as the original version.
[0054] Final quality inspection: Based on the results of consistency evaluation and review, duplicate removal and disambiguation of named entities are carried out. Specifically, it includes operations such as merging duplicate named entities, eliminating ambiguity in named entity types, and correcting named entity boundary errors. Through this step, it is ensured that the final output of named entities and their types have a high degree of accuracy and consistency. For example, merge "Beijing" and "Beijing City" into "Beijing", and unify "John" and "Jon" into "John".
[0055] Output: Store the named entities and their types that have passed the final quality inspection as structured data for use in subsequent stages, and they can be stored in formats such as JSON and XML.
[0056] The attribute completion stage is based on the output of the entity recognition stage, and uses prompt templates to guide the large language model to complete attribute information for entities, including rough extraction, consistency evaluation, attribute review, final quality inspection, and output links.
[0057] Rough extraction: Based on the output of the entity recognition stage, the large language model extracts relevant attributes for each named entity according to the preset prompt templates. The multiple parallel rough extraction strategy is also adopted to generate multiple groups of candidate attributes and their values to improve the coverage and diversity of attribute extraction. Specific methods include using multiple different prompt templates, calling the model multiple times, etc.
[0058] Consistency evaluation: Based on multiple groups of candidate attributes extracted roughly, the method of majority voting is used for consistency evaluation. Count the occurrence frequencies of each candidate attribute in multiple extraction results, and select the attribute and its value with the highest occurrence frequency as the result of consistency evaluation. This method can effectively reduce accidental errors in attribute extraction and improve the robustness of the results. For candidate attributes that fail the consistency evaluation, they enter a separate review process to further decide whether to enter the final quality inspection stage in combination with the content of the text materials. For example, if a certain attribute is recognized as "Age: 30 years old" 8 times out of 10 extractions, then select "30 years old" as the value of this attribute.
[0059] Attribute review: For candidate attributes that fail the consistency evaluation, they enter a separate review process. Find the complete sentence or sentence fragment corresponding to the attribute in the context to ensure that sufficient context information is provided. Analyze and judge whether the naming and classification of the attribute are correct, and give the modification conclusion of the attribute. Return the analysis process and modification conclusion of each attribute. If the conclusion is that no modification is required, the result should be the same as the original version.
[0060] Final Quality Inspection: Based on the results of consistency assessment and review, duplicate attributes are removed and ambiguities are eliminated. Specifically, it includes operations such as merging duplicate attributes, correcting attribute value errors, and eliminating attribute type ambiguities. Through this step, it is ensured that the final output attributes and their values are highly accurate and consistent. For example, unify "Age: 30 years old" and "Age: 30" into "Age: 30 years old".
[0061] Output: Store the named entities and their attributes that have passed the final quality inspection as structured data for use in subsequent stages. It can be stored in formats such as JSON and XML.
[0062] The relationship extraction stage is based on the output of the attribute completion stage. Using prompt templates to guide the large language model to extract the relationships between entities, including rough extraction, consistency assessment, relationship review, final quality inspection, and output links.
[0063] Rough Extraction: Based on the output of the attribute completion stage, the large language model extracts the relationships between named entities from the text according to the preset prompt templates. To improve the coverage and diversity of relationship extraction, a multiple parallel rough extraction strategy is adopted to generate multiple groups of candidate relationships and their types. The specific execution process is as follows: Traverse the named entity list: Traverse each named entity in the provided named entity list, assumed to be E.
[0064] Relationship Extraction: For the currently indexed named entity E, search for the relationships related to this named entity in the context, where E needs to be the head of the relationship.
[0065] Relationship Fields: The fields of the relationship include the relationship head (including the named entity name and type), the relationship type (relationship predicate), and the relationship tail (including the named entity name and type).
[0066] Relationship Saving: Save all the relevant relationships of the currently indexed named entity E in the form of a list.
[0067] If there is indeed no relationship with this named entity E as the head, it is allowed to set the relationship fields to empty, and it is absolutely not allowed to fabricate or create out of nothing. After completing the relationship extraction of one named entity, continue to process the relationship extraction of the next named entity in the named entity list.
[0068] Consistency Assessment: Based on multiple groups of candidate relationships obtained from rough extraction, a majority voting method is used for consistency assessment. Count the occurrence frequencies of each candidate relationship in multiple extraction results, and select the relationship and its type with the highest occurrence frequency as the result of consistency assessment. For example, if a certain relationship is recognized as "John lives in Beijing" 8 times out of 10 extractions, then select "John lives in Beijing" as the description of this relationship.
[0069] This method can significantly reduce accidental errors in relation extraction and improve the robustness and reliability of the results. For candidate relations that fail the consistency assessment, they enter a separate review process, and further decisions are made in combination with the content of the text materials to determine whether to enter the final quality inspection stage.
[0070] Relation review: Find the complete sentence or sentence fragment corresponding to the relation in the context to ensure sufficient context information is provided. Analyze and judge whether the naming and classification of the head, tail, and relation type of the relation are correct. Combine the context information to judge whether the relation is reasonable, whether it conforms to the semantics, and give the modification conclusion of the relation. Return the corrected relation list, including the corrected version of the relation. If no modification is required, retain the original version.
[0071] Final quality inspection: Based on the results of the consistency assessment and review, perform deduplication and disambiguation of relations. Specifically, it includes operations such as merging duplicate relations, correcting relation type errors, and eliminating relation ambiguities. For example, unify "John lives in Beijing City" and "John lives in Beijing" into "John lives in Beijing City". Through this step, ensure that the finally output relations and their types have high accuracy and consistency.
[0072] Output: Store the relations that have passed the final quality inspection as structured data to form the triples of the knowledge graph, which can be stored in formats such as JSON and XML.
[0073] Steps for knowledge graph construction: Through database configuration, knowledge storage and construction, and interfaces and queries, organize the knowledge triples extracted from the text into a structured knowledge graph and implement querying and visualization display.
[0074] Specifically, database configuration includes the following steps: Database selection: Select a database system or graph database engine that supports efficient storage and query of graph-structured data, such as Neo4j, ArangoDB, etc.
[0075] Database initialization: Set the running environment of the database, including specifying the location of data storage, configuring the caching mechanism, etc., to ensure the efficient operation of the database. For example, set the connection method and connection string of the database, configure the cache size, and whether to support pooling, etc.
[0076] Data model definition: According to the results of knowledge extraction, define the data structure in the database, including determining different types of data elements (such as entities) and their attributes, as well as the association relationships between data elements. For example, define the entity type as Person, with attributes including name, age, etc., and the relation type as LIVES_IN.
[0077] Specifically, knowledge storage and construction includes the following steps: Data Import: Import the knowledge triple data extracted from the text into the database, establish corresponding data elements and association relationships, and techniques such as batch import or streaming import can be used to improve the import efficiency.
[0078] Data Validation: Check the imported data to ensure the integrity and consistency of data elements and association relationships, and avoid data errors. Specific methods include data verification, data cleaning, etc.
[0079] Index Optimization: Create indexes for data elements and association relationships in the database to improve the efficiency of data query. For example, create a full-text index for entity names, create an index for relationship types, etc.
[0080] Specifically, the interface and query include the following steps: Interface Definition: Provide a set of standardized interfaces for external systems to query and operate on the knowledge graph. For example, define REST API interfaces that support HTTP methods such as GET and POST.
[0081] Query Optimization: Optimize the query process, support complex query requirements, and ensure that the query can respond in real time. Specific methods include using indexes, query caching, etc.
[0082] Visualization Display: Provide visualization tools to support users to view data elements and association relationships in the knowledge graph in a dynamic and interactive manner. For example, use third-party development tools to generate charts on the client side.
[0083] Please refer to Figure 3 This invention also discloses a multi-stage serial knowledge graph construction system based on prompt words, which applies the multi-stage serial knowledge graph construction method based on prompt words described in any one of the above, and includes: A pre-configuration module, a knowledge extraction module, and a graph construction module. The pre-configuration module is used to initialize various parameters and resources required by the system. The knowledge extraction module is used to extract entities, attributes, and relationships from the text to generate preliminary knowledge triples. The graph construction module is used to organize the knowledge triples into a structured knowledge graph and implement query and visualization display; The pre-configuration module includes a large language model configuration, a prompt word set configuration, and an example set configuration sub-module. The large language model configuration sub-module is used for model selection and loading, parameter setting, and model optimization. The prompt word set configuration sub-module is used for prompt word design, prompt word storage, and prompt word version iteration management. The example set configuration sub-module is used for example design, example storage, and example version iteration management; The knowledge extraction module includes a dynamic text chunking and a multi-stage serial extraction chain sub-module. The dynamic text chunking sub-module is used to dynamically calculate the size of text chunks, and the multi-stage serial extraction chain sub-module is used to decompose the knowledge extraction task into an entity recognition stage, an attribute completion stage, and a relationship extraction stage; The knowledge graph construction module includes a database configuration, a knowledge storage and construction, and an interface and query sub-module. The database configuration sub-module is used for database selection, database initialization, and data model definition. The knowledge storage and construction sub-module is used for data import, data verification, and index optimization. The interface and query sub-module is used for interface definition, query optimization, and visualization display.
[0084] The pre-configuration module initializes various parameters and resources required by the system through a large language model configuration sub-module, a prompt set configuration sub-module, and an example set configuration sub-module to ensure the smooth progress of the knowledge extraction process.
[0085] Specifically, the large language model configuration sub-module is used to set various parameters and configurations required for interacting with the large language model, and implement the following steps: Model selection and loading: Select a suitable large language model according to the task requirements and load the model parameters. For example, select a model that supports long text processing and high-precision inference, such as GPT-4, etc. The loading of model parameters can be achieved through application programming interface calls or local file reading, etc.
[0086] Parameter setting: Configure the inference parameters of the model, specifically including context length, temperature value, maximum output length, presence penalty value, and trust call.
[0087] The context length is used to set the maximum context length for the model to process text, such as 1024 characters, to ensure that the model can fully understand the semantic information of long texts.
[0088] The temperature value adjusts the diversity of the generated content. A lower temperature value (such as 0.2) makes the output more deterministic and concentrated, while a higher temperature value (such as 0.8) increases the randomness and diversity of the output.
[0089] The maximum output length is used to limit the length of the content generated by the model, such as 256 characters, to avoid generating overly long or irrelevant texts.
[0090] The presence penalty value reduces the occurrence of repeated words or phrases in the generated content, such as setting it to 1.5, to improve the diversity and quality of the generated content.
[0091] The trust call technology improves the reliability and accuracy of the output results by automatically detecting and fixing the error responses generated by the model. It introduces an additional verification mechanism to ensure that each part of the content generated by the model meets the expected requirements.
[0092] Model Optimization: Adjust the computational resource allocation of the model according to the hardware resources, such as using distributed computing or model pruning techniques to improve the running efficiency.
[0093] Specifically, the prompt set configuration sub-module is used to provide various types of prompt templates and implement the following steps: Prompt Template Design: Design dedicated prompt templates for each knowledge extraction stage (entity recognition, attribute completion, relation extraction). For example, the prompt template for the entity recognition stage is: "Please extract all entities from the following text and give their types." In prompt design, adopt prompt engineering techniques to guide the model to reason step by step through appropriate prompts to ensure high-quality output. At the same time, introduce the chain of thought technique to enable the model to solve problems step by step and improve the accuracy and logic of reasoning.
[0094] Prompt Storage: Store the optimized prompt templates in the system, which can be done in various ways such as using a database, file system, or in-memory cache.
[0095] Prompt Version Iteration Management: Manage the version iteration of the prompt templates through experiments and evaluations to ensure that they can accurately guide the large language model to complete tasks and continuously optimize the effectiveness and advancement of the prompts.
[0096] Specifically, the example set configuration sub-module is used to store and manage the example data for few-shot learning and implement the following steps: Example Design and Collection: Design example data for each knowledge extraction stage to help the large language model better understand the task requirements. For example, an example for the entity recognition stage is: "Text: 'John studies at the Massachusetts Institute of Technology.', Entities: 'John' (person), 'Massachusetts Institute of Technology' (organization)." Example Storage: Store the optimized example data in the system, which can be done in various ways such as using a database, file system, or in-memory cache.
[0097] Example Set Version Iteration Management: Dynamically adjust the example content according to the task requirements and improve the model's performance through version iteration management. For example, regularly update the example set and add new examples to cover more scenarios.
[0098] The present invention adds examples to the prompts and adopts a few-shot learning strategy. By providing a small number of representative examples, it helps the large language model better understand the task requirements and improve the quality of knowledge extraction.
[0099] The knowledge extraction module realizes efficient and accurate knowledge extraction through the dynamic text chunking sub-module and the multi-stage serial extraction chain sub-module.
[0100] The dynamic text chunking sub-module dynamically calculates the size of text chunks according to the maximum context window length supported by the large language model and the length of the prompt. This avoids the problems caused by traditional hard-coded chunk sizes, can make full use of the performance of the large language model, reduce the situation of excessive or insufficient chunking, and thus improve the efficiency and accuracy of knowledge extraction.
[0101] Specifically, the dynamic text chunking sub-module implements the following steps: Text preprocessing: Clean and format the original text, remove noisy data, and ensure the quality of the input text. The specific steps include removing HTML tags, converting to lowercase, removing special characters, etc.
[0102] Chunk size calculation: Dynamically calculate the size of text chunks according to the maximum context length supported by the large language model and the length of the prompt. The calculation formula is:
[0103] In the formula, is the dynamically calculated text chunk size; is the maximum context length supported by the large language model; is the dynamically calculated length of the prompt; is an empirical coefficient used to reserve a part of the context length for the large language model to perform reasoning.
[0104] Text chunk optimization: Divide the preprocessed text into multiple text chunks of dynamic sizes, ensure that the length of each text chunk does not exceed the dynamically calculated chunk size, and avoid semantic breaks. Specific methods include the sliding window method, the fixed-size chunking method, etc.
[0105] Chunk input: Input the optimized text chunks into the large language model for processing, and batch processing or asynchronous processing and other technologies can be used to improve the processing efficiency.
[0106] The multi-stage serial extraction chain sub-module decomposes the knowledge extraction task into an entity recognition stage, an attribute completion stage, and a relationship extraction stage. Each stage includes multiple parallel rough extractions, self-consistency evaluations, and quality inspection links to ensure the accuracy and consistency of the extraction results.
[0107] The entity recognition stage uses a prompt template to guide the large language model to extract entities from the text, including rough extraction, consistency evaluation, entity review, final quality inspection, and output links.
[0108] Rough extraction: The large language model initially extracts named entities and their types from the text according to a preset prompt template. To improve the extraction coverage and diversity, a multiple parallel rough extraction strategy is adopted to generate multiple sets of candidate named entities and their types. This strategy can effectively avoid omissions or biases that may be caused by a single extraction path, and enhance the comprehensiveness and accuracy of the extraction results. Specific methods include using multiple different prompt templates, calling the model multiple times, etc.
[0109] Consistency evaluation: Based on multiple sets of candidate named entities obtained from rough extraction, a majority voting method is used for consistency evaluation. Specifically, count the occurrence frequencies of each candidate named entity in multiple sets of extraction results, and select the named entity and its type with the highest occurrence frequency as the result of consistency evaluation. For example, if an entity is recognized as "person" 8 times in 10 parallel extractions, then select "person" as the type of this entity.
[0110] This method can significantly reduce the impact of accidental errors and enhance the robustness and reliability of the extraction results. For candidate named entities that do not pass the consistency evaluation, they enter a separate review process, and further decisions are made in combination with the content of the text materials to determine whether to enter the final quality inspection stage.
[0111] Entity review: For candidate named entities that do not pass the consistency evaluation, they enter a separate review process. Locate the complete sentence or sentence fragment corresponding to the named entity in the context to ensure that sufficient context information is provided. Analyze and judge whether the naming and classification of the named entity are correct, and give the modification conclusion for this named entity. Return the analysis process and modification conclusion for each named entity. If the conclusion is that no modification is required, the result should be the same as the original version.
[0112] Final quality inspection: Based on the results of consistency evaluation and review, perform deduplication and disambiguation of named entities. Specifically, it includes operations such as merging duplicate named entities, eliminating ambiguity in named entity types, and correcting named entity boundary errors. Through this step, ensure that the final output named entities and their types have high accuracy and consistency. For example, merge "Beijing" and "Beijing Municipality" into "Beijing", and unify "John" and "Jon" into "John".
[0113] Output: Store the named entities and their types that have passed the final quality inspection as structured data for use in subsequent stages, and JSON, XML, etc. formats can be used for storage.
[0114] Based on the output of the entity recognition stage, the attribute completion stage uses a prompt template to guide the large language model to complete the attribute information for entities, including rough extraction, consistency evaluation, attribute review, final quality inspection, and output links.
[0115] Rough extraction: Based on the output of the entity recognition stage, the large language model extracts relevant attributes for each named entity according to a preset prompt template. The multiple parallel rough extraction strategy is also adopted to generate multiple sets of candidate attributes and their values, so as to improve the coverage and diversity of attribute extraction. Specific methods include using multiple different prompt templates, calling the model multiple times, etc.
[0116] Consistency evaluation: Based on multiple sets of candidate attributes obtained from rough extraction, the method of majority voting is used for consistency evaluation. The occurrence frequency of each candidate attribute in multiple sets of extraction results is counted, and the attribute and its value with the highest occurrence frequency are selected as the result of consistency evaluation. This method can effectively reduce accidental errors in attribute extraction and improve the robustness of the results. For candidate attributes that fail the consistency evaluation, they enter a separate review process, and further decisions are made in combination with the content of the text materials to determine whether to enter the final quality inspection stage. For example, if a certain attribute is recognized as "Age: 30 years old" in 8 out of 10 extractions, then "30 years old" is selected as the value of this attribute.
[0117] Attribute review: For candidate attributes that fail the consistency evaluation, they enter a separate review process. Locate the complete sentence or sentence fragment corresponding to the attribute in the context to ensure that sufficient context information is provided. Analyze and judge whether the naming and classification of the attribute are correct, and give the modification conclusion of the attribute. Return the analysis process and modification conclusion of each attribute. If the conclusion is that no modification is required, the result should be the same as the original version.
[0118] Final quality inspection: Based on the results of consistency evaluation and review, duplicate removal and disambiguation of attributes are performed. Specifically, operations such as merging duplicate attributes, correcting incorrect attribute values, and eliminating ambiguity in attribute types are included. Through this step, it is ensured that the final output attributes and their values have high accuracy and consistency. For example, unify "Age: 30 years old" and "Age: 30" into "Age: 30 years old".
[0119] Output: Store the named entities and their attributes that have passed the final quality inspection as structured data for use in subsequent stages. JSON, XML, etc. formats can be used for storage.
[0120] Based on the output of the attribute completion stage, the relationship extraction stage uses prompt templates to guide the large language model to extract the relationships between entities, including rough extraction, consistency evaluation, relationship review, final quality inspection, and output links.
[0121] Rough extraction: Based on the output of the attribute completion stage, the large language model extracts the relationships between named entities from the text according to a preset prompt template. To improve the coverage and diversity of relationship extraction, the multiple parallel rough extraction strategy is adopted to generate multiple sets of candidate relationships and their types. The specific execution process is as follows: Traverse the list of named entities: Traverse each named entity in the provided list of named entities, assumed to be E.
[0122] Relation extraction: For the named entity E at the current index, search for the relationships related to this named entity in the context, where E needs to be the head of the relationship.
[0123] Relation fields: The fields of the relationship include the head of the relationship (including the named entity name and type), the relationship type (relationship predicate), and the tail of the relationship (including the named entity name and type).
[0124] Relation saving: Save all the relevant relationships of the named entity E at the current index in the form of a list.
[0125] If the named entity E really has no relationships with it as the head, it is allowed to leave the relationship fields empty, and it is absolutely not allowed to fabricate or create out of nothing. After completing the relation extraction of one named entity, continue to process the relation extraction of the next named entity in the list of named entities.
[0126] Consistency evaluation: Based on multiple groups of candidate relationships extracted roughly, use the method of majority voting for consistency evaluation. Count the occurrence frequencies of each candidate relationship in the multiple groups of extraction results, and select the relationship and its type with the highest occurrence frequency as the result of the consistency evaluation. For example, if a certain relationship is recognized as "John lives in Beijing City" 8 times out of 10 extractions, then select "John lives in Beijing City" as the description of this relationship.
[0127] This method can significantly reduce accidental errors in relation extraction and improve the robustness and reliability of the results. For candidate relationships that do not pass the consistency evaluation, enter a separate review process, and further decide whether to enter the final quality inspection stage in combination with the content of the text materials.
[0128] Relation review: Find the complete sentence or sentence fragment corresponding to the relationship in the context to ensure that sufficient context information is provided. Analyze and judge whether the naming and classification of the head, tail, and relationship type of the relationship are correct. Combine the context information to judge whether the relationship is reasonable and consistent with the semantics, and give the modification conclusion of this relationship. Return the corrected list of relationships, including the corrected version of the relationship. If no modification is required, keep the original version.
[0129] Final quality inspection: Based on the results of the consistency evaluation and review, perform deduplication and disambiguation of the relationships. Specifically, it includes operations such as merging duplicate relationships, correcting relationship type errors, and eliminating relationship ambiguities. For example, unify "John lives in Beijing City" and "John lives in Beijing" into "John lives in Beijing City". Through this step, ensure that the final output relationships and their types have high accuracy and consistency.
[0130] Output: Store the relations that have passed the final quality inspection as structured data to form the triples of the knowledge graph, which can be stored in formats such as JSON and XML.
[0131] Specifically, the database configuration sub-module implements the following steps: Database selection: Select a database system or graph database engine that supports efficient storage and query of graph-structured data, such as Neo4j, ArangoDB, etc.
[0132] Database initialization: Set the running environment of the database, including specifying the location of data storage, configuring the caching mechanism, etc., to ensure the efficient operation of the database. For example, set the connection method and connection string of the database, configure the cache size, and whether to support pooling, etc.
[0133] Data model definition: According to the results of knowledge extraction, define the data structure in the database, including determining different types of data elements (such as entities) and their attributes, as well as the association relationships between data elements. For example, define the entity type as Person, with attributes including name, age, etc., and the relationship type as LIVES_IN.
[0134] Specifically, the knowledge storage and construction sub-module implements the following steps: Data import: Import the knowledge triple data extracted from the text into the database, and establish the corresponding data elements and association relationships. Technologies such as batch import or streaming import can be used to improve the import efficiency.
[0135] Data verification: Check the imported data to ensure the integrity and consistency of data elements and association relationships, and avoid data errors. Specific methods include data verification and data cleaning.
[0136] Index optimization: Create indexes for data elements and association relationships in the database to improve the efficiency of data query. For example, create a full-text index for entity names and an index for relationship types, etc.
[0137] Specifically, the interface and query include the following steps: Interface definition: Provide a set of standardized interfaces for external systems to query and operate on the knowledge graph. For example, define REST API interfaces that support HTTP methods such as GET and POST.
[0138] Query optimization: Optimize the query process, support complex query requirements, and ensure that the query can respond in real time. Specific methods include using indexes and query caching.
[0139] Visualization display: Provide visualization tools to support users to view data elements and association relationships in the knowledge graph in a dynamic and interactive manner. For example, use third-party development tools to generate charts on the client side.
[0140] The prompt-based multi-stage serial knowledge graph construction system of the present invention can execute the prompt-based multi-stage serial knowledge graph construction method of the present invention, can execute any combination of the implementation steps of the method embodiments, and has the corresponding functions and beneficial effects of the method.
[0141] The present invention separates the prompts of each task stage from the large language model, realizes modular design, and facilitates the independent configuration and optimization of the prompts of each stage. This design allows for optimization for specific task stages. For example, the prompts for the entity recognition stage can be optimized to improve the accuracy of entity recognition without affecting the prompts and task performance of the attribute completion and relationship extraction stages. The modular design significantly improves the flexibility, controllability, and maintainability of the system and is applicable to different scenarios and requirements.
[0142] The above description is a detailed description of the preferred feasible embodiments of the present invention, but the embodiments are not intended to limit the scope of the patent application of the present invention. Any equivalent changes or modifications made under the technical spirit disclosed by the present invention shall fall within the scope of the patent covered by the present invention.
Claims
1. A method for constructing a multi-stage serial knowledge graph based on prompt words, characterized in that: The following steps are involved: Pre-configuration: Initialize various parameters and resources required by the system through large language model configuration, prompt word set configuration, and example set configuration; Knowledge extraction: Accurate knowledge extraction is achieved through dynamic text segmentation and multi-stage serial extraction chain, extracting entities, attributes and relations from text to generate preliminary knowledge triples; Graph construction: Through database configuration, knowledge storage and construction, and interface and query, knowledge triples are organized into a structured knowledge graph, and query and visualization are realized.
2. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: The configuration of a large language model includes the following steps: Model selection and loading: Select a suitable large language model according to task requirements and load model parameters; Parameter setting: configure the inference parameters of the model, including context length, temperature value, maximum output length, existence penalty value, and trust call; Model optimization: Adjust the computing resource allocation of the model according to the hardware resources.
3. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: The prompt word set configuration includes the following steps: Prompt word template design: Design a dedicated prompt word template for each knowledge extraction stage; Prompt word storage: store the optimized prompt word template in the system; Prompt word version iteration management: Through experiments and evaluations, the prompt word template is managed in version iteration to guide the large language model to complete the task and continuously optimize the prompt words.
4. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: Using few-shot learning, the example set configuration includes the following steps: Example design and collection: Design example data for each knowledge extraction stage to help the large language model understand task requirements; Sample storage: Store the optimized sample data in the system; Example set version iteration management: Dynamically adjust example content according to task requirements and improve model performance through version iteration management.
5. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: Dynamic text segmentation includes the following steps: Text preprocessing: clean and format the original text, remove noise data, and ensure the quality of the input text; Chunk size calculation: The size of the text chunk is dynamically calculated based on the maximum context length and prompt word length supported by the large language model. The calculation formula is: In the formula, The text chunk size is calculated dynamically; The maximum context length supported for large language models; is the length of the prompt word; It is an empirical coefficient used to reserve a portion of the context length for the large language model to perform reasoning; Text block optimization: Divide the preprocessed text into multiple text blocks of dynamic sizes, ensuring that the length of each text block does not exceed the dynamically calculated block size, while avoiding semantic discontinuity; Chunk input: Input the optimized text blocks into the large language model for processing.
6. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: The multi-stage serial extraction chain decomposes the knowledge extraction task into entity recognition stage, attribute completion stage and relation extraction stage, including the following steps: The entity recognition phase uses prompt word templates to guide the large language model to extract entities from the text, including rough extraction, consistency assessment, entity review, final quality inspection and output. The attribute completion phase is based on the output of the entity recognition phase. It uses the prompt word template to guide the large language model to complete the attribute information for the entity, including rough extraction, consistency assessment, attribute review, final quality inspection and output. The relationship extraction stage is based on the output of the attribute completion stage. It uses prompt word templates to guide the large language model to extract the relationship between entities, including rough extraction, consistency assessment, relationship review, final quality inspection and output.
7. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: Database configuration includes the following steps: Database selection: Choose a database system that supports efficient storage and query of graph structured data; Database initialization: Set up the database operating environment, including specifying the location of data storage and configuring the cache mechanism to ensure the efficient operation of the database; Data model definition: Based on the results of knowledge extraction, define the data structure in the database, including determining different types of data elements and their attributes, as well as the relationship between data elements.
8. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: Knowledge storage and construction includes the following steps: Data import: import the knowledge triples extracted from the text into the database and establish corresponding data elements and associations; Data validation: Check the imported data to ensure the integrity and consistency of data elements and relationships and avoid data errors; Index optimization: Create indexes for data elements and relationships in the database to improve the efficiency of data queries.
9. The method for constructing a multi-stage serial knowledge graph based on prompt words according to claim 1, characterized in that: Interface and query, including the following steps: Interface definition: Provide a set of standardized interfaces for external systems to query and operate knowledge graphs; Query optimization: optimize the query process, support complex query requirements, and ensure that queries can be responded to in real time; Visual display: Provides visualization tools to support users to view data elements and relationships in the knowledge graph in a dynamic and interactive way.
10. A system for constructing a multi-stage serial knowledge graph based on prompt words, applying the method for constructing a multi-stage serial knowledge graph based on prompt words according to any one of claims 1 to 9, characterized in that: include: Pre-configuration module, knowledge extraction module and graph construction module. The pre-configuration module is used to initialize the parameters and resources required by the system. The knowledge extraction module is used to extract entities, attributes and relationships from the text and generate preliminary knowledge triples. The graph construction module is used to organize knowledge triples into a structured knowledge graph and realize query and visualization. The front-end configuration module includes large language model configuration, prompt word set configuration and example set configuration submodules. The large language model configuration submodule is used for model selection and loading, parameter setting and model optimization; the prompt word set configuration submodule is used for prompt word design, prompt word storage and prompt word version iteration management; The example set configuration submodule is used for example design, example storage, and example version iteration management; The knowledge extraction module includes a dynamic text segmentation submodule and a multi-stage serial extraction chain module. The dynamic text segmentation submodule is used to dynamically calculate the size of the text segmentation. The multi-stage serial extraction chain module is used to decompose the knowledge extraction task into an entity recognition stage, an attribute completion stage, and a relationship extraction stage. The graph construction module includes database configuration, knowledge storage and construction, and interface and query sub-modules. The database configuration sub-module is used for database selection, database initialization and data model definition. The knowledge storage and construction sub-module is used for data import, data verification and index optimization. The interface and query sub-module is used for interface definition, query optimization and visualization display.
Citation Information
Cited By
Knowledge graph automatic construction system and method based on LangGraph workflow
CN120781950A
Knowledge graph automatic construction system and method based on langgraph workflow
CN120781950B
Self-adaptive software stack optimization method and system based on GPU (Graphics Processing Unit) server configuration
CN120803541A
Self-guiding knowledge graph construction method and device based on large model
CN121119063A
Large model-based self-guided knowledge graph construction method and device
CN121119063B