Multi-level retrieval generation method based on atlas enhancement
By introducing knowledge graph technology into the RAG method, a multi-level search generation method is constructed, and the problems of multi-level related information extraction and global information perspective support in complex unstructured texts are solved, efficient and accurate information retrieval and integration are achieved, and user experience and decision-making capabilities are improved.
Patent Information
- Application Number
- CN202510073665.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
When existing RAG methods deal with complex unstructured text, it is difficult to effectively extract multi-level related information, lack the support of the global information perspective, cannot effectively handle multi-sense entities, and lack the coordination mechanism between local and global queries.
A multi-level search generation method based on graph enhancement is adopted, through the knowledge graph construction stage and the running query engine stage, text segmentation, entity and relationship information extraction, event information extraction, community classification and multi-level information integration are realized, and a multi-dimensional information retrieval plan is provided in combination with local query and global query mechanisms.
It improves the accuracy and comprehensiveness of information retrieval, can effectively extract and integrate multi-level related information, support complex queries, and improves user experience and decision-making capabilities.
Abstract
Description
Technical Field
[0001] The present invention relates to the field of natural language processing and knowledge graph technology, and specifically to a multi-level retrieval generation method based on graph enhancement. Background Art
[0002] In the field of natural language processing, retrieval-augmented generation (RAG) technology has become an important method for solving open-domain question-answering tasks. RAG usually retrieves knowledge points related to the question by segmenting and vectorizing the knowledge base text to calculate the vector similarity between texts to measure the semantic relevance, and then inputs these knowledge points into the large language model (LLM) to generate answers. This method performs well in processing general question-answering scenarios, especially when the knowledge base structure is simple and the information is directly searchable, and can meet the needs of most users.
[0003] However, when the knowledge base content is more complex, such as containing long unstructured texts such as news reports, novels, and professional analysis reports, ordinary RAG methods often face limitations. For these complex contents, user needs may involve deep-level related information, comprehensive analysis of cross-chapter content, and even correlation mining between multiple knowledge points. For example, in the field of shipping analysis, users may want to query the evolution of an event or the influence of multiple related factors; in literary works, users may want to understand the relationship between characters in a certain plot and their evolution. Ordinary RAG is difficult to deal with effectively when dealing with such complex queries, because its single semantic similarity matching method cannot deeply explore the multi-level associations between texts.
[0004] Specifically include the following issues:
[0005] 1. Insufficient extraction of relevant information from unstructured text: Traditional RAG methods only rely on vector similarity for retrieval, which makes it difficult to effectively extract multi-level relevant information from the text, especially when the knowledge base content involves cross-paragraph, cross-chapter or even cross-article associations. The accuracy and comprehensiveness of information extraction are low.
[0006] 2. Lack of support for global information perspective: When answering complex questions, users may not only need directly matching knowledge points, but also need to analyze macro information across multiple knowledge units. Existing RAG technology is difficult to effectively integrate global information, resulting in incomplete answers to complex queries (such as summary questions).
[0007] 3. Disambiguation and unification of polysemous entities in the knowledge base: Complex texts often contain the same entity with different names or descriptions, such as multiple titles of the same character in literary works, or multiple mentions of the same event in shipping reports. Traditional RAG lacks an effective mechanism to uniformly handle polysemous entities, resulting in missing or repeated information in the answering process.
[0008] 4. Lack of collaborative mechanism for local and global queries: The current RAG method mainly relies on local similarity matching and fails to establish a collaborative retrieval mechanism for local and global queries. When users need to integrate local and global information to answer complex questions, the existing method cannot flexibly improve retrieval accuracy by filtering global information.
[0009] To this end, existing technologies have gradually introduced knowledge graphs (KGs) to improve the structuring of information. Knowledge graphs extract entities and their relationships to construct multi-level association networks for complex content. However, existing knowledge graph applications still have shortcomings, mainly reflected in the lack of segmentation processing of long and complex texts, maintenance of semantic consistency, and cross-knowledge point association analysis. Ordinary knowledge graphs usually only support direct matching of local entities and are difficult to provide global association information integration, especially when dealing with complex association problems across paragraphs, chapters, and even chapters, lacking support from a global perspective. Summary of the invention
[0010] In order to solve the above problems in the prior art, the present invention provides a multi-level retrieval generation method based on graph enhancement.
[0011] The technical solution of the present invention is as follows:
[0012] A multi-level retrieval generation method based on graph enhancement, characterized by comprising a knowledge graph construction stage and a query engine operation stage;
[0013] The query engine operation phase includes the following steps:
[0014] (1) Text segmentation: divide the document to be processed into multiple text units for subsequent processing;
[0015] (2) Entity information extraction: For each text unit, a large language model is used to extract entity information, including the entity’s name, type, and description.
[0016] (3) Relationship information extraction: While extracting entity information, a large language model is used to identify and extract relationship information between entities in text units, including source entities, target entities, and description information;
[0017] (4) Entity and relationship description summary: The extracted description information is summarized through a large language model, and multiple descriptions are summarized into a concise sentence;
[0018] (5) Event information extraction: Use a large language model to extract event information from text units, including the initiator, reporter, event type, status, start and end dates, description information, and event source;
[0019] (6) Community classification: classify the extracted entities to construct community information;
[0020] (7) Generate community report: Generate a community report that describes the composition of the community, including the entities, relationships, and events in the community;
[0021] (8) Vectorization: vectorize the description information in entities, relationships, and communities, and vectorize the original text units;
[0022] (9) Local query and global query mechanism:
[0023] Local query: Based on vector similarity, retrieve entity information, relationship information, event information and community information related to the question, select the knowledge points with the highest matching degree and input them into the large language model to generate the answer;
[0024] Global query: Analyze the importance of the community through a large language model, evaluate the matching degree between the community and the question, filter out the communities and their information related to the question, and then generate the answer through the large language model;
[0025] (10) Multi-level information integration: Integrate the results of local and global queries, combine them with specific problem requirements, and provide diversified information support from different levels to ensure the comprehensiveness and accuracy of the answers.
[0026] Preferably, in step (1), the maximum length of each text unit is set to 300 tokens by default.
[0027] Preferably, in step (2), the description information is stored in a list format.
[0028] Preferably, in step (3), the relationship information is stored in a list format.
[0029] Preferably, the knowledge graph construction stage specifically includes:
[0030] (1) Data preparation: Ensure that all documents to be processed are in txt format. For documents that are not in txt format, format conversion is required to meet system requirements.
[0031] (2) Data input: Create a folder named input in the project directory and put all the prepared data files into this folder;
[0032] (3) Initialize the workspace: Run the initialization command in the project directory to create two files: .env and settings.yaml;
[0033] (4) Run the indexing process: After configuring the settings.yaml file, run the indexing command to start the indexing process. After the process is completed, a file named output / <timestamp> / artifacts, which contains a series of parquet files, indicating that the knowledge graph was built successfully.
[0034] Preferably, in step (3) of the knowledge graph construction phase, the .env file contains environment variables for accessing corresponding API services; and the settings.yaml file contains process configuration settings.
[0035] The technical effects of the present invention are as follows:
[0036] 1. Improve the accuracy of information retrieval: By segmenting the text and extracting entities, relationships, events and other information, the system can more accurately understand the semantics of user queries, thereby improving the accuracy of knowledge point retrieval. This precise matching method enables the system to quickly locate relevant information and avoid ambiguity and ambiguity in the information retrieval process.
[0037] 2. Rich knowledge structure: The present invention transforms unstructured knowledge into structured knowledge graphs through community classification and multi-level information integration, forming a clear knowledge hierarchy. This structured information enables users to obtain more comprehensive and relevant knowledge points when querying, improving the user experience.
[0038] 3. Flexible query methods: Combining local query and global query mechanisms, the system can provide adaptive retrieval strategies for different types of query requirements. Users can choose the corresponding query method according to the specific nature of the question to obtain more accurate and relevant answers.
[0039] 4. Efficient information integration capability: The present invention can not only provide answers from the specific entity level, but also integrate knowledge from the macro community level, which is applicable to a variety of complex query scenarios. This information integration capability enables users to quickly obtain comprehensive information related to the question, meeting the needs of complex queries.
[0040] 5. Support multiple application scenarios: Since the present invention can process different types of text data (such as news reports, novels, academic documents, etc.), it is suitable for a variety of application scenarios, including education, scientific research, business analysis and other fields, providing users with more extensive information support.
[0041] 6. Improve user decision-making ability: By providing high-quality knowledge support, the present invention helps users make more informed decisions in complex situations. This decision-making support capability provides powerful assistance to users in various practical applications.
[0042] The present invention proposes a multi-level retrieval generation method based on graph enhancement, which combines graph technology with multi-level retrieval methods to realize a multi-dimensional information retrieval solution combining local query and global query. This method realizes efficient storage and management of complex related information through graph construction, and provides users with high-quality knowledge association and question-answering support through local matching and global screening in the query process, which is particularly suitable for processing complex and long unstructured text knowledge bases.
[0043] This method is suitable for building and retrieving knowledge from complex unstructured text data to provide high-quality question-answering support. By combining retrieval generation (RAG) technology with knowledge graph construction, the present invention can achieve more accurate knowledge point positioning and related information query in long text and cross-domain knowledge base.
[0044] In summary, the implementation of the present invention can significantly improve the efficiency and quality of knowledge retrieval and generation, provide users with richer and more accurate information services, and promote technological progress and application expansion in related fields. DETAILED DESCRIPTION
[0045] In order to better understand the present invention, the present invention is further explained below in conjunction with specific embodiments.
[0046] Example 1
[0047] In order to implement the multi-level retrieval generation method based on graph enhancement described in the present invention, the following is a detailed implementation step, including a knowledge graph construction stage and a query engine operation stage.
[0048] The knowledge graph construction phase includes:
[0049] 1. Data preparation:
[0050] - Ensure that all documents to be processed are in txt format. The text content should be in a standardized format to avoid extraction errors.
[0051] - For documents in non-txt format, such as PDF or Word documents, format conversion is required to adapt to the system requirements. Note that scanned PDF files are not supported.
[0052] 2. Data input:
[0053] -Create a folder called input in your project directory and place all your prepared data files into this folder.
[0054] 3. Initialize the workspace:
[0055] - Run the initialization command in your project directory, which will create two files: .env and settings.yaml.
[0056] -.env file contains environment variables such as API_KEY, which is used to access the corresponding API service. Replace this key with your own API key.
[0057] The -settings.yaml file contains the process configuration settings and can be modified according to specific needs.
[0058] 4. Run the indexing process:
[0059] -After configuring the settings.yaml file, run the index command to start the indexing process. The running time of this process depends on the size of the input data, the model used, and the text block size (these parameters can be configured in the .env file).
[0060] -After the process is completed, a file named output / <timestamp> / artifacts, which contains a series of parquet files. This indicates that the knowledge graph has been successfully built.
[0061] The query engine operation phase includes:
[0062] 1. Text segmentation: The document to be processed is divided into multiple text units (TextUnit), and the maximum length of each text unit is 300 tokens by default, which is convenient for subsequent processing. This step aims to reduce the length of the text and improve the efficiency of subsequent information extraction.
[0063] 2. Entity information extraction: For each text unit, a large language model (LLM) is used to extract entity information, including the entity's name, type, and description. The description information is stored in a list format to facilitate the management of multiple descriptions of the same entity from different text units.
[0064] 3. Relationship information extraction: While extracting entity information, LLM is used to identify and extract relationship information between entities in text units, including source entity (source), target entity (target) and description information (description). These relationship information are also saved in list form to handle multiple relationship descriptions between the same entities.
[0065] 4. Entity and relationship description summary: The extracted description information is summarized through LLM, and multiple descriptions are summarized into a concise sentence to reduce the complexity of the query and facilitate fast matching.
[0066] 5. Event information extraction: Use LLM to extract event information that occurs in text units, including the initiator, reporter, event type, status, start and end date, description information, and event source. The extraction of this information helps enrich the knowledge base content and provide support for complex queries.
[0067] 6. Community classification: Classify the extracted entities to construct community information. For example, classify the relevant entities into communities such as "the pilgrim team" or "the heaven", thus forming a hierarchical knowledge structure to facilitate subsequent queries.
[0068] 7. Generate community report: Generate a community report that describes the composition of the community, including the entities, relationships, and events in the community. The community description information will be used as the matching condition for subsequent queries.
[0069] 8. Vectorization: Vectorize the description information of entities, relationships, and communities so that semantically similar content can be quickly found when querying. At the same time, vectorize the original text units to provide basic support for subsequent knowledge base retrieval.
[0070] 9. Local query and global query mechanism:
[0071] -Local Search: Based on vector similarity, retrieve entity information, relationship information, event information and community information related to the question, select the knowledge point with the highest matching degree and input it into LLM to generate the answer. This method is suitable for questions that require specific entity answers.
[0072] - Local search example: Run a query command, such as calling the query module using a Python script, and specify the query method as local search, and the question to be queried, such as "Who is the protagonist of the story of Scrooge, and what are his main relationships?" This command will retrieve entity information, relationship information, event information, and community information related to the question based on vector similarity, and input the knowledge point with the highest matching degree into the large language model (LLM) to generate an answer.
[0073] -Global Search: LLM is used to analyze the importance of communities, evaluate the matching degree between communities and questions, screen out communities and their information related to the questions, and then generate answers through LLM. This method is more suitable for complex problems that require comprehensive information.
[0074] -Global search example: Run a query command, such as calling the query module using a Python script, and specify the query method as global search, and the question to be queried, such as "What is the theme of this story?" This command will start the global query process, and the system will analyze the importance of the community and select the most relevant community information to answer the given question.
[0075] 10. Multi-level information integration: Integrate the results of local and global queries, combine them with specific problem requirements, and provide diversified information support from different levels to ensure the comprehensiveness and accuracy of the answers.
[0076] Through the above steps, this embodiment constructs a retrieval generation framework that can effectively process complex unstructured knowledge bases, improves the ability to extract knowledge points, generate answers and disambiguate entities, thereby better meeting the needs of users in complex query scenarios.
[0077] Through the above steps, the multi-level retrieval generation method based on graph enhancement described in the present invention can be implemented to effectively process complex unstructured knowledge bases and provide high-quality question-answering support. This method is particularly suitable for scenarios where local and global information needs to be integrated to answer complex questions.< / timestamp> < / timestamp>
Claims
1. A multi-level retrieval generation method based on graph enhancement, characterized in that It includes the knowledge graph construction phase and the query engine operation phase; The query engine operation phase includes the following steps: (1) Text segmentation: divide the document to be processed into multiple text units for subsequent processing; (2) Entity information extraction: For each text unit, a large language model is used to extract entity information, including the entity’s name, type, and description. (3) Relationship information extraction: While extracting entity information, a large language model is used to identify and extract relationship information between entities in text units, including source entities, target entities, and description information; (4) Entity and relationship description summary: The extracted description information is summarized through a large language model, and multiple descriptions are summarized into a concise sentence; (5) Event information extraction: Use a large language model to extract event information from text units, including the initiator, reporter, event type, status, start and end dates, description information, and event source; (6) Community classification: classify the extracted entities to construct community information; (7) Generate community report: Generate a community report that describes the composition of the community, including the entities, relationships, and events in the community; (8) Vectorization: vectorize the description information in entities, relationships, and communities, and vectorize the original text units; (9) Local query and global query mechanism: Local query: Based on vector similarity, retrieve entity information, relationship information, event information and community information related to the question, select the knowledge points with the highest matching degree and input them into the large language model to generate the answer; Global query: Analyze the importance of the community through a large language model, evaluate the matching degree between the community and the question, filter out the communities and their information related to the question, and then generate the answer through the large language model; (10) Multi-level information integration: Integrate the results of local and global queries, combine them with specific problem requirements, and provide diversified information support from different levels to ensure the comprehensiveness and accuracy of the answers.
2. The method according to claim 1, characterized in that In step (1), the maximum length of each text unit is 300 tokens by default.
3. The method according to claim 1, characterized in that In step (2), the description information is stored in a list format.
4. The method according to claim 1, characterized in that In step (3), the relationship information is stored in a list format.
5. The method according to claim 1, characterized in that The knowledge graph construction stage specifically includes: (1) Data preparation: Ensure that all documents to be processed are in txt format. For documents that are not in txt format, format conversion is required to meet system requirements. (2) Data input: Create a folder named input in the project directory and put all the prepared data files into this folder; (3) Initialize the workspace: Run the initialization command in the project directory to create two files: .env and settings.yaml; (4) Run the indexing process: After configuring the settings.yaml file, run the indexing command to start the indexing process. After the process is completed, a file named output / <timestamp> / artifacts, which contains a series of parquet files, indicating that the knowledge graph was built successfully.< / timestamp> 6. The method according to claim 5, characterized in that In step (3), the .env file contains environment variables for accessing the corresponding API service; the settings.yaml file contains process configuration settings.
Citation Information
Cited By
Intelligent film and television media asset retrieval method and system based on image enhancement generation
CN121524374A