A progressive retrieval question and answer method and system
Patent Information
- Application Number
- CN202611017512.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-09
- Publication Date
- 2026-09-29
AI Technical Summary
1.本发明提出的一种渐进式检索问答方法,通过信息充足度门控机制,改变了喂给大语言模型冗长无序文本的现状,通过在信息充足时及时停止检索,避免了现有技术强行拼凑固定数量(如Top-10)片段造成的冗长上下文,降低了大语言模型推理的上下文Token消耗,大幅缩减了API调用成本;实现了“按需检索”。
Smart Images

Figure CN122838522A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of natural language processing and artificial intelligence information retrieval technology, and in particular to a progressive retrieval question-answering method and system. Background Technology
[0002] In intelligent question-answering systems based on large language models, Retrieval-augmented Generation (RAG) is widely used. Its standard workflow is as follows: receiving user questions, vectorizing the questions, performing a one-time similarity search in a vector database to recall the top-K text fragments, and finally concatenating these fragments before inputting them into the large language model to generate the answer. It employs a fixed-number recall strategy, meaning that regardless of the complexity of the user question, a one-time similarity search is performed, and a fixed number of text fragments are concatenated before being input into the large language model.
[0003] However, the existing technology has the following significant technical shortcomings: First, there is a serious waste of token resources. For simple questions, a large amount of redundant text is recalled, which leads to excessively long inference time for the large language model, significantly increasing computational costs and inference latency, resulting in wasted computing resources and response delays. For complex questions, a fixed number of recall strategies are used, lacking the ability to dynamically judge the quality of search results and failing to confirm whether the information obtained is sufficient to support an accurate answer.
[0004] Second, contextual interference leads to frequent illusions. The large number of text fragments recalled at once often contain distracting information that is literally similar to the question but semantically unrelated. This redundant information interferes with the attention mechanism of large language models, inducing the model to produce factually incorrect outputs.
[0005] Third, there is a lack of ability to judge the quality of search results. Existing technology is in a state of passively executing the search. It cannot dynamically adjust the search starting point according to the question's intent, nor can it terminate the search process in a timely manner when there is sufficient information. Furthermore, it lacks a precise control mechanism for the search scope and cannot determine whether the retrieved information is sufficient to support the answer to the user's question before generating an answer. This leads to a double dilemma of information deficiency when encountering complex questions and information overload when encountering simple questions.
[0006] Therefore, there is an urgent need to propose a progressive retrieval question answering method and system to solve the problems of serious waste of token resources, frequent illusions caused by contextual interference, and lack of ability to judge the quality of retrieval results. Summary of the Invention
[0007] The purpose of this invention is to provide a progressive retrieval question-answering method and system. It breaks away from the traditional one-time flat retrieval mode, constructs a tree-like hierarchical index structure, introduces an evaluation agent, and proposes a closed-loop retrieval algorithm of "intent-adaptive starting point - multi-level retrieval gating - dynamic termination and drill-down". It has the advantages of being able to dynamically adjust the retrieval starting point and depth according to the complexity of the question, avoiding token resource waste, context interference, information overload and information loss, and significantly improving the efficiency and accuracy of question answering.
[0008] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a progressive retrieval and question-answering method, the method comprising: S1. Establish a multi-level indexed database; S2. Vectorize the data in the multi-level index database to obtain numerical data that can be understood by computers, and store it in the vector database; S3. Configure the intent classifier, which includes intent routing. The intent classifier performs the following steps: S30. Receive text input by the user and identify the intent in the text; S31. Vectorize the intent and output the intent label; S32. Input the intent label into the intent route; the intent route determines the initial level of the retrieval hierarchy; S33. Intent routing distributes the intent to the vector database for similarity matching, and retrieves the text of the Top-N nodes and their associated identifiers; S4. Configure the evaluation agent, which performs the following steps: S40. Receive the text input by the user and the text of the Top-N nodes retrieved for intent routing, along with their associated identifiers; S41. Parse the core entities and interrogative words in the user's input text; S42. Check whether the text of the Top-N nodes contains answer entities that support the question word, and execute the judgment logic.
[0009] As one possible implementation, multi-level index databases include: S20. Configure the text data and parse the text data into a tree structure; S21. Construct multi-level indexes based on a tree structure; S22. A multi-level indexed database consists of text data and multi-level indexes.
[0010] As one possible implementation, a multi-level index has at least three levels.
[0011] As one possible implementation, a multi-level index includes a document summary layer: storing metadata and a global summary of the entire document, with each node carrying a unique document ID.
[0012] As one possible implementation, a multi-level index includes a document chapter layer: storing the title and summary of each chapter, with each node having a document ID and a chapter ID.
[0013] As one possible implementation, a multi-level index contains slice layers: storing specific text blocks, each text block with a document ID and a chapter ID as metadata tags.
[0014] As one possible implementation, the evaluation agent is an agent with autonomous perception, decision-making, execution, and interaction capabilities; it can autonomously identify whether the text block corresponding to the chapter ID sufficiently contains the text fragments required by the user.
[0015] As one possible implementation, the evaluation agent's decision logic involves checking whether the text of the Top-N nodes contains answer entities that support the interrogative word; When the check results do not contain the answer entity or contain only part of the answer entity, it is determined that the results are insufficient and have not reached the lowest level of the retrieval level. The document IDs of the currently recalled Top-N node texts are extracted, the document IDs are set as metadata filters, and the intent is further distributed to the vector database corresponding to the next level of the retrieval level for similarity matching. If the search results do not contain the answer entity or contain only part of the answer entity, it is determined to be insufficient and has reached the lowest level of the search hierarchy, and a "not found" message is generated to answer the user. If the check result contains the answer entity, it is deemed sufficient, and together with the text entered by the user, it is input into the core large language model to generate the final natural language answer to the user.
[0016] As one possible implementation, metadata filters are used to filter candidate documents in real time based on their structured attributes, either before or simultaneously with vector similarity searches.
[0017] Secondly, the present invention provides a progressive retrieval question answering system, including a multi-level index database, a vector database, an intent classifier, and an evaluation agent; The system receives text input by the user, sets the initial retrieval level through intent routing, performs a retrieval, recalls the retrieval results and inputs them into the evaluation agent for judgment, performs a new round of evaluation based on the judgment results, and finally determines that the concise text fragment is sufficiently complete. This fragment, along with the text input by the user, is then input into the retrieval enhancement and generation model to generate the final natural language answer to the user.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The progressive retrieval question answering method proposed in this invention changes the current situation of feeding long and disordered text to large language models through an information sufficiency gating mechanism. By stopping the retrieval in time when the information is sufficient, it avoids the long context caused by the forced piecing together of a fixed number (such as Top-10) fragments in the existing technology, reduces the context token consumption of large language model inference, and greatly reduces the API call cost; thus realizing "on-demand retrieval".
[0019] 2. The progressive retrieval question answering method proposed in this invention suppresses information interference and illusions at the source. It uses a progressive funnel drilling mechanism to narrow the retrieval scope layer by layer by utilizing metadata filters. The context input to the generative model achieves an extremely high signal-to-noise ratio, cuts off the interference of irrelevant or similar conflicting content on the large language model, significantly improves the factual accuracy of the answers, and effectively suppresses "illusions".
[0020] 3. The progressive retrieval question-answering method proposed in this invention is designed with intent adaptation as the starting point. It can dynamically adjust the retrieval starting point and depth according to the complexity of the question. When dealing with macro-level questions (such as summarizing the whole text) and micro-level questions (such as querying specific parameters), it can dynamically select the best path, avoiding information overload and information loss. It also avoids the loss of summarization ability caused by the fragmentation of underlying text. It has the advantages of significantly improving the efficiency and accuracy of question answering, while improving the overall experience of question answering in all scenarios and significantly improving the flexible and adaptive response capability. Attached Figure Description
[0021] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is an overall business process diagram proposed in an embodiment of the present invention. Detailed Implementation
[0022] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" are not necessarily different.
[0023] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0024] In this invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, "at least one of a, b, or c" can represent: a, b, c, a combination of a and b, a combination of a and c, a combination of b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0025] Existing technologies in Retrieval-augmented Generation (RAG) suffer from several technical problems, including significant waste of token resources, frequent illusions caused by contextual interference, and a lack of ability to judge the quality of search results. Specifically, a fixed-quantity recall strategy leads to redundant text, increasing computational costs and inference latency; a large number of text fragments recalled at once contain noise, interfering with the attention mechanism of large language models and inducing factual errors; and the system cannot determine whether the information is sufficient to support the answer before generating it, resulting in information loss or overload.
[0026] In a first aspect, the present invention provides a progressive retrieval and question-answering method, the method comprising: S1. Establish a multi-level index database. A multi-level index database is constructed as a hierarchical data structure designed to store and manage large amounts of text information. Its core function is to divide and index raw text data according to different granularities, such as organizing it by document, chapter, paragraph, etc., thereby supporting information retrieval at different depths and scopes.
[0027] S2. The data in the multi-level index database is vectorized to obtain a computer-understandable numerical form, which is then stored in a vector database. The vector database is specifically designed to store high-dimensional vector data and supports efficient similarity searching. In the progressive retrieval question-answering method, text data, after vectorization, is converted into a computer-understandable numerical form, i.e., vectors, and stored in this database. By calculating the similarity between the query vector and the vectors stored in the database, semantically relevant text fragments can be quickly retrieved.
[0028] S3. Configure the intent classifier, which includes intent routing. The intent classifier performs the following steps: S30. Receive text input by the user and identify the intent in the text; S31. Vectorize the intent and output the intent label; S32. Input the intent label into the intent route; the intent route determines the initial level of the retrieval hierarchy; S33. Intent routing distributes the intent to the vector database for similarity matching, and retrieves the text of the Top-N nodes and their associated identifiers.
[0029] The intent classifier is configured to analyze user-input text and identify the user's intent. Its operation typically involves natural language processing techniques, mapping user queries to predefined intent categories or tags. Internally, the classifier contains intent routing, responsible for guiding subsequent retrieval processes based on the identified intent.
[0030] Intent routing: As part of the intent classifier, this component dynamically determines the initial level of the retrieval process based on the intent tags output by the intent classifier. For example, for general questions, the retrieval starting point might be set at a shallower level; for specific fact queries, it might be set at a deeper level. This routing is also responsible for distributing user intents to a vector database for similarity matching.
[0031] S4. Configure the evaluation agent, which performs the following steps: S40. Receive the text input by the user and the text of the Top-N nodes retrieved for intent routing, along with their associated identifiers; S41. Parse the core entities and interrogative words in the user's input text; S42. Check whether the text of the Top-N nodes contains answer entities that support the question word, and execute the judgment logic.
[0032] The evaluation agent is configured to receive text input from the user and text fragments intended for routing recall, and to evaluate the information sufficiency of the recalled text. Its core function is to parse key information from the user query and check whether the recalled text contains answer entities that can support the response. Based on the check results, the agent executes pre-defined decision logic to determine whether further retrieval is needed or whether to directly generate an answer.
[0033] This invention provides a progressive retrieval and question-answering method that achieves refined management of the knowledge base by constructing a multi-level index database. The method first establishes a multi-level index database, where data can be organized in various ways. For example, all documents can be simply stored flat as a single-level index, with each document serving as an independent retrieval unit. Alternatively, a more structured approach can be adopted, dividing document content into several logical units, such as chapters or paragraphs, and creating an index for each unit, without establishing a clear hierarchical relationship between these units.
[0034] Data in the multi-level index database is vectorized to obtain computer-understandable numerical data, which is then stored in a vector database. Vectorization can be implemented in ways including, but not limited to: converting text content into sparse vectors using traditional methods such as the Bag-of-Words model or TF-IDF (term frequency–inverse document frequency); or mapping specific keywords or phrases in the text to predefined numerical codes using rule-based or dictionary-based methods. This numerical data is stored in the vector database for subsequent similarity matching.
[0035] See Figure 1 Configure an intent classifier, which includes intent routing. This intent classifier performs the steps of receiving text input from the user and identifying the intent within the text. For example, the intent classifier can identify the user's intent by comparing specific words in the user's input text with a pre-defined list of intent keywords based on keyword matching rules. The identified intent is vectorized, and an intent label is output. Intent vectorization simply maps the identified intent label to a unique ID.
[0036] Intent tags are input into the intent route, which determines the initial search level. For example, the intent route can directly associate different intent tags with a fixed search level by consulting a pre-defined mapping table. If the intent tag indicates that the user is querying a broad topic, the intent route might set the initial search level to a shallower level, such as the document level. If the intent tag indicates that the user is querying a specific fact, the intent route might set the initial search level to a deeper level, such as the paragraph level.
[0037] Intent routing distributes the intent to a vector database for similarity matching, retrieving the top-N node texts and their associated identifiers. For example, intent routing can directly calculate the similarity between the vector representation of the user's intent and data at all levels in the vector database, retrieving the top-N results with the highest similarity. The retrieved node texts and their associated identifiers, such as document IDs or chapter IDs, are used for subsequent information processing.
[0038] The evaluation agent performs the steps of receiving the user-input text and the text of the Top-N nodes retrieved for intent routing, along with their associated identifiers. For example, the evaluation agent can simply concatenate the user-input text and the retrieved node text to form a text block to be evaluated. The evaluation agent then parses the core entities and interrogative words in the user-input text. For instance, it can extract noun phrases as core entities from the user-input text using predefined grammar rules or regular expressions, and identify interrogative words such as "what" and "how much".
[0039] The evaluation agent checks whether the text of the Top-N nodes contains an answer entity that supports the interrogative word and executes the decision logic. For example, the evaluation agent can simply check whether the recalled Top-N node text contains a keyword that directly matches the parsed core entity and interrogative word. If a directly matching keyword exists, it is determined that the answer entity is contained; otherwise, it is determined that it is not contained.
[0040] This invention, through the introduction of a multi-level index database and an intent classifier, achieves adaptive recognition of user query intent and dynamically adjusts the retrieval starting point based on the intent. The configuration of the evaluation agent enables the system to assess the information sufficiency of the recalled text fragments, thus avoiding the problem of recalling a large amount of redundant information at once in traditional retrieval methods. This method helps reduce unnecessary computational resource consumption and improves the response efficiency and information accuracy of the question-answering system when handling problems of varying complexity.
[0041] In some embodiments of the present invention described above, a multi-level index database is proposed, and the data therein is vectorized for storage in a vector database. However, in practical applications, if the construction of the multi-level index database lacks a clear structured method, it may lead to chaotic data organization, making it difficult to effectively support subsequent vectorization processing and similarity matching, thereby affecting the accuracy and efficiency of retrieval.
[0042] As one possible implementation, multi-level index databases include: S20. Configure text data and parse it into a tree structure. This refers to processing raw, unstructured, or semi-structured text information, such as documents, articles, and reports, in an automated or semi-automated manner to identify its internal logical hierarchy and structural relationships. For example, Natural Language Processing (NLP) technology can be used to identify structural elements in a document, such as titles, subheadings, paragraphs, and lists, and organize them into a hierarchical tree data structure based on the inclusion, parallel, or subordinate relationships between these elements. Each node in the tree structure represents a logical unit of text, such as an entire document, a chapter, a paragraph, or a sentence. Nodes are connected through parent-child relationships, clearly showing the organizational structure of the text content.
[0043] S21. Build a multi-level index based on a tree structure. This involves creating an indexing mechanism on top of the parsed tree structure that enables efficient access and retrieval of information at different levels. This index assigns a unique identifier to each node (i.e., each logical unit of text) in the tree structure and records its position within the entire document structure, its level, and its relationships with other nodes (such as parent node ID, child node ID, etc.). For example, index entries can be created for information of different granularities, such as documents, chapters, and paragraphs. These index entries not only contain the text content itself or its summary but also metadata pointing to its position in the tree structure. This multi-level index design allows the system to quickly locate information of a specific granularity based on retrieval needs without traversing the entire text content.
[0044] S22. A multi-level indexed database consists of text data and multi-level indexes. This database is a comprehensive storage system that contains not only the structured text data itself (which may be stored in its raw form or a pre-processed tree structure) but also multi-level indexes for efficient retrieval and navigation of this text data. The two are closely integrated to form an organic whole, enabling the system to quickly locate target text fragments through the indexes and obtain their complete contextual information. This database can be implemented using various technologies, such as combining relational databases to store metadata and index information, while using document databases or file systems to store the actual text content and linking them through indexes.
[0045] This invention parses raw text data into a tree structure and builds a multi-level index based on this structure, effectively solving the problems of disorganized data and inefficient retrieval. This structured data organization allows information in the multi-level index database to be stored and managed in a clear and logical manner. In subsequent vectorization processing, more semantically representative vectors can be generated based on this structure, significantly improving the accuracy and efficiency of similarity matching in the vector database. Furthermore, the establishment of the multi-level index provides a solid foundation for progressive retrieval, enabling the system to flexibly switch between different granular information levels based on user intent and evaluation results, achieving precise information retrieval from macro to micro levels. This avoids blind retrieval of massive amounts of unstructured data and significantly optimizes the overall performance and user experience of the question-answering system.
[0046] In some embodiments of the present invention, a progressive retrieval question-answering method is proposed. This method involves establishing a multi-level index database and performing vectorization processing, combined with an intent classifier and an evaluation agent for information retrieval and judgment. However, when constructing a multi-level index, if the hierarchical structure of the index is not fine enough or lacks sufficient depth, the granularity of the recalled node text may be inappropriate during similarity matching. This can affect the efficiency and accuracy of the evaluation agent in checking the answer entity, and may even require multiple unnecessary iterative searches, thus reducing the overall performance of the question-answering system.
[0047] As a possible implementation, a multi-level index has at least three levels. This means that when building a multi-level index, its structural depth includes at least three different levels of abstraction or granularity. This ensures that the index can organize and represent text data from macro to micro levels, providing the necessary granularity selection for progressive retrieval. It allows the system to flexibly perform information matching and retrieval at different levels of abstraction based on user intent and the current retrieval stage. For example, coarse-grained matching can be performed at higher levels to quickly filter out relevant documents; if more detailed information is needed, fine-grained matching can be performed at intermediate or lowest levels, thereby gradually focusing on the precise answer required by the user. This multi-level index can be divided according to the logical structure or semantic relevance of the text content. For example, the first level can be set as a "topic overview layer" to store the overall topic or core concepts of the document; the second level can be set as a "content module layer" to store the various main content modules or sub-arguments related to the topic; and the third level can be set as a "detailed information layer" to store specific facts, data, or detailed descriptions supporting the content modules. This hierarchical division of at least three levels enables the system to perform efficient retrieval and information focusing at different levels of abstraction, based on the granularity of the user's query requirements.
[0048] By employing the aforementioned technical solution, the multi-level index is configured with at least three levels. This allows the intent routing in the intent classifier to select an appropriate level for similarity matching based on the user's intent and the initial retrieval level, choosing from at least three different granularity index levels. For example, for intents with strong generalization, the system can prioritize matching at higher levels to quickly narrow the search scope; for intents requiring specific details, matching can proceed directly or gradually to lower levels. When the evaluation agent determines that the currently recalled Top-N node texts are insufficient, this multi-level structure can be used to extract identifiers (such as document IDs) and use them as metadata filters to continue searching at the next finer-grained level, thus achieving progressive information retrieval from coarse to fine. This at least three-level hierarchical structure enables progressive retrieval question-answering methods to locate the answers users need more efficiently and accurately. It avoids overly detailed searches in the initial stage, reducing unnecessary computational resource consumption; simultaneously, it avoids the problem of insufficient detailed information due to too few levels. By providing multi-level retrieval granularity, the system can flexibly adjust retrieval strategies according to the complexity of user queries and the accuracy of required information, significantly improving the recall efficiency and accuracy of answers in the question-answering system, and optimizing the judgment process for evaluating the agent, reducing invalid judgments and iterations.
[0049] As one possible implementation, a multi-level index includes a document summary layer: storing metadata and a global summary of the entire document, with each node carrying a unique document ID.
[0050] The document summary layer is a specific level within a multi-level index, its primary function being to provide a highly summarized view of the entire document's content. This layer is designed as an entry point for users or systems to perform initial information filtering and understanding, avoiding direct processing of all the details of the original document. Through the document summary layer, the relevance of a document to a query can be quickly determined, thus guiding subsequent progressive retrieval. Each node in the document summary layer is configured to store the entire document's metadata and a global summary.
[0051] Metadata can include, but is not limited to, structured attributes such as document title, author, creation date, document type, keywords, and source information. These attributes help to classify, filter, and retrieve documents.
[0052] The global summary is a concise summary of the core content of the entire document. It is usually generated through automatic summarization technology (such as extractive summarization or generative summarization) or manual editing, and aims to convey the main points and information of the document in concise language.
[0053] To ensure that each document in the document summary layer can be uniquely identified and tracked, each node is assigned a unique document ID. This document ID serves as a global identifier for the document, not only for indexing and retrieval within the document summary layer, but also as a crucial link connecting the document summary layer with other lower levels in the multi-level index (such as the document chapter layer and slice layer), ensuring accurate location of the original document and its specific content fragments during progressive retrieval.
[0054] By introducing a document summary layer and storing the metadata and global summary of the entire document, with each node bearing a unique document ID, this invention significantly improves the efficiency and accuracy of progressive retrieval question answering methods. Once a user's query intent is identified, the system can first perform preliminary similarity matching at the document summary layer. Because this layer stores high-level summary information about the documents, the system can quickly filter out the set of documents most relevant to the user's intent from a massive amount of data without needing to deeply analyze the detailed content of each document. This high-level filtering mechanism effectively reduces the amount of data processed subsequently and lowers computational overhead. Simultaneously, through the unique document ID, the system can ensure that after the initial screening, it can accurately trace back to the original document and further conduct in-depth searches at finer-grained levels (such as chapter or slice levels), thereby achieving progressive information acquisition from macro to micro levels and ultimately providing users with more accurate and efficient answers.
[0055] As one possible implementation, a multi-level index includes a document chapter layer: storing the title and summary of each chapter, with each node having a document ID and a chapter ID.
[0056] The document chapter layer is a key component of the multi-level index, primarily responsible for organizing and indexing the structured information within a document. Located below the overall document level but above the finest-grained text block level, it aims to provide a medium-granularity information retrieval entry point. Within this layer, each chapter extracts and stores its corresponding title and summary information. The chapter title provides a general identifier of the chapter's content, while the chapter summary further condenses the core content. This information is crucial for initial semantic matching and content filtering, enabling the system to quickly determine whether a chapter is relevant to the user's intent during retrieval. To ensure accurate location and traceability of information within the document chapter layer, each chapter node is assigned a unique identifier. The document ID identifies the original document to which the chapter belongs, while the chapter ID uniquely identifies a specific chapter within that document. These IDs, serving as metadata tags, not only facilitate information association and navigation between different levels but also provide precise location capabilities for subsequent progressive retrieval, ensuring that the source and location within the document are clearly known when recalling relevant chapters.
[0057] By introducing a document chapter layer and storing the titles and summaries of each chapter, as well as assigning a document ID and a chapter ID to each node, this invention addresses the issues of low retrieval efficiency and inaccurate positioning that can occur when relying solely on the entire document or the finest-grained text blocks for searching. When a user's intent relates to a specific topic or subdomain of a document, the intent classifier and evaluation agent can first perform matching and evaluation at the document chapter layer. Through the semantic information of chapter titles and summaries, the system can quickly filter out chapters highly relevant to the user's query, avoiding blind searching of the entire document or large amounts of fragmented text blocks. Simultaneously, the introduction of document IDs and chapter IDs ensures that the recalled chapter information can be accurately traced back to its original document and specific location, providing a clear context for subsequent in-depth retrieval or direct answer generation. This significantly improves the efficiency and accuracy of progressive retrieval, enabling users to obtain the information they need faster and more accurately, thereby optimizing the overall question-and-answer experience.
[0058] As one possible implementation, a multi-level index contains slice layers: storing specific text blocks, each text block with a document ID and a chapter ID as metadata tags.
[0059] The slice layer is the lowest level of the multi-level index structure. Its main function is to store the smallest, semantically independent text units in the original text data. This layer is designed to provide the finest-grained access to text content, enabling the system to precisely locate the specific text fragments that a user query might involve, rather than just a macro-level overview of the document or chapter. A specific text block refers to a text fragment of a certain length and complete semantics, extracted from the original text, such as a sentence, a paragraph, or a logically complete set of phrases. These text blocks are the basic processing units of information retrieval and question-answering systems. By storing these specific text blocks, the system can avoid recalling the entire document or chapter during retrieval, thus significantly improving the accuracy and relevance of search results. The document ID is a globally unique identifier used to uniquely identify the original document, while the chapter ID is used to uniquely identify the chapter to which the text block belongs. Attaching these metadata tags to each text block ensures that even at the finest-grained slice layer, each text block can clearly trace its position and context within the original document structure. These metadata tags play a crucial role in subsequent retrieval, filtering, and evaluation processes. For example, they can serve as metadata filters for precise screening or help the evaluation agent understand the source and context of text blocks.
[0060] By introducing a slicing layer and storing specific text blocks, while attaching document IDs and chapter IDs as metadata tags to each text block, this embodiment of the invention enables deeper refinement and organization of text content. This allows intent routing to directly recall at the finest-grained text block level during similarity matching, significantly improving the accuracy of the recall results. When the evaluation agent receives the recalled Top-N node texts, since these node texts are already specific text blocks with clear context labels, the evaluation agent can more efficiently and accurately check whether they contain answer entities supporting the question words, reducing the processing of irrelevant information. Furthermore, through document IDs and chapter IDs, even if the recalled text blocks are small, their complete context in the original document can be quickly reconstructed, providing a solid foundation for the subsequent generation of high-quality natural language answers by the core large language model, thereby improving the efficiency and user experience of the entire question-answering system.
[0061] As one possible implementation, the evaluation agent is an agent with autonomous perception, decision-making, execution, and interaction capabilities; it can autonomously identify whether the text block corresponding to the chapter ID sufficiently contains the text fragments required by the user.
[0062] The evaluation agent's autonomous perception capability refers to its ability to proactively acquire and understand information without relying on external instructions. It can perceive the semantics and contextual information of the user-input text, as well as the content, structure, and association identifiers of the top-N nodes of text retrieved via intent routing. This perception capability can be achieved through Natural Language Processing (NLP) techniques, such as named entity recognition, sentiment analysis, and topic modeling, to gain a deeper understanding of the input information. Its decision-making capability refers to the evaluation agent's ability to autonomously make judgments and choices based on the information it perceives. For example, when determining whether the retrieved text is sufficient, it can decide, based on preset rules, models, or learned knowledge, whether to continue the next round of retrieval, generate prompts, or submit the current text to the core large language model. The decision-making process can be based on machine learning models or expert system rules. The evaluation agent's execution capability refers to its ability to translate decisions into concrete actions. For example, when the decision requires further retrieval, it can proactively trigger intent routing for the next level of retrieval; when the decision requires generating an answer, it can pass relevant information to the core large language model. Execution capability is typically achieved by calling corresponding interfaces or modules. Furthermore, evaluating the interactive capabilities of an intelligent agent refers to its ability to exchange information and collaborate with other system components or users. It can interact with internal components such as intent classifiers, vector databases, and core large language models to exchange data and instructions, and can also engage in clarifying or guiding interactions with users when necessary to obtain more specific needs. This interactive capability ensures the smoothness and efficiency of the entire question-and-answer process.
[0063] Simultaneously, the evaluation agent can autonomously identify whether the text block corresponding to the chapter ID sufficiently contains the text fragments required by the user. Autonomous identification here means that the evaluation agent can automatically perform deep analysis of the text content without human intervention. This goes beyond simple keyword matching or entity existence checks; it uses advanced NLP techniques such as semantic understanding and contextual analysis to determine the deep relationship between the text block and the user's query. For example, pre-trained language models can be used to calculate semantic similarity, or question-answering pair generation models can be used to evaluate whether the text block can directly answer the user's question. The text block corresponding to the chapter ID refers to the specific text content fragment located by the chapter ID in a multi-level index database. These text blocks are structured and have clear contextual boundaries. Whether they sufficiently contain the text fragments required by the user is a key judgment of the quality and completeness of the recalled text. It is not merely a matter of determining whether a certain entity exists, but rather evaluating the amount of information, level of detail, relevance, and completeness of the information provided by the text block. For example, if a user asks for the definition and application scenarios of a concept, the evaluation agent needs to determine whether the recalled text block simultaneously contains the definition and application scenarios, and whether the description is clear and comprehensive enough. This judgment can be made by constructing a specialized evaluation model that combines the complexity of the user query with the granularity of the expected answer.
[0064] Through the aforementioned technical solution, the evaluation agent is designed to possess autonomous perception, decision-making, execution, and interaction capabilities. It can autonomously identify whether the text block corresponding to the chapter ID sufficiently contains the text fragments required by the user, thereby significantly improving the intelligence level and question-answering quality of the progressive retrieval question-answering method. The evaluation agent no longer merely performs simple entity existence checks but can conduct deep semantic understanding and contextual analysis of the recalled text content, accurately judging its matching degree with the user query and the completeness of information. This capability allows the system to more accurately determine when to stop the retrieval, when to drill down to a finer granular level, or when to directly submit the recalled content to the core large language model to generate an answer. For example, when a user asks a question that requires the integration of multiple information points to answer, the evaluation agent, with its autonomous perception and decision-making capabilities, can identify that although the currently recalled text block contains some answer entities, the information is insufficient to form a complete answer, thus proactively triggering a deeper level of retrieval until a sufficiently complete text fragment is found. This avoids generating low-quality answers due to incomplete information, reduces unnecessary retrieval iterations, and improves retrieval efficiency. Meanwhile, its interactive capabilities enable the system to better communicate with users when necessary, obtain clearer needs, and further optimize question-and-answer performance. Ultimately, by ensuring that the core large language model receives high-quality, sufficient, and accurate text fragments, the accuracy, completeness, and user satisfaction of the final natural language response are greatly improved.
[0065] As one possible implementation, the evaluation agent's decision logic involves checking whether the text of the Top-N nodes contains answer entities that support the interrogative word.
[0066] The evaluation agent receives the user's input text and the top-N node texts retrieved from the intent routing, along with their associated identifiers. It then parses the core entities and question words within the user's input text. The evaluation agent performs in-depth analysis of the top-N node texts, using techniques such as natural language processing, named entity recognition, or semantic matching to determine if any entity information directly answers or supports the user's question words exists. This step aims to initially filter out the text fragments most relevant to the user's query, providing a foundation for subsequent decision-making.
[0067] When the check results do not contain the answer entity or contain only part of the answer entity, it is determined that the results are insufficient and have not reached the lowest level of the retrieval level. The document IDs of the currently recalled Top-N node texts are extracted, the document IDs are set as metadata filters, and the intent is further distributed to the vector database corresponding to the next level of the retrieval level for similarity matching. The role of metadata filters is to pre-screen the data in the vector database before subsequent vector similarity matching. This means that when retrieving the next level, the system will only perform similarity matching on the vector data associated with these specific document IDs, thus precisely limiting the search scope to the set of documents relevant to the initial results, effectively improving the efficiency and accuracy of the search. For example, if the current level recalls a document summary, but the summary does not contain the complete answer, the system can use that document ID to perform a more granular search on the detailed content of that document at the next level (such as the chapter or slice level). Based on this, the system will continue to distribute the intent to the vector database corresponding to the next level of the search level for similarity matching.
[0068] If the search results do not contain the answer entity or contain only part of the answer entity, it is determined to be insufficient and has reached the lowest level of the search hierarchy, and a "not found" message is generated to answer the user. This step defines the termination condition for progressive retrieval. If the evaluating agent, after checking all available retrieval levels, still fails to find a sufficient number of answer entities from the recalled node text—that is, after exhausting all possible retrieval paths—the system will generate a clear "not found" message and return it to the user as the answer. This avoids the system getting stuck in an infinite loop or providing irrelevant information when it cannot provide a valid answer, thus ensuring clarity of user experience and robustness of the system.
[0069] If the check result contains the answer entity, it is deemed sufficient, and together with the text entered by the user, it is input into the core large language model to generate the final natural language answer to the user.
[0070] This step describes the processing flow after a successful retrieval. The system will select highly relevant text fragments, deem them "sufficiently sufficient," and input these, along with the user's original text, into the core language model. The core language model, leveraging its powerful natural language understanding and generation capabilities, integrates, refines, and organizes this information, ultimately generating a fluent, accurate, and human-readable natural language response, presented directly to the user. This allows users to obtain a high-quality, easily understood answer, rather than the raw text fragment.
[0071] Through the above technical solution, this invention effectively solves the problem in progressive retrieval question-answering methods where there is a lack of clear judgment logic and further retrieval strategies when the initial retrieval results are insufficient. By introducing a refined evaluation agent judgment logic, intelligent judgment of retrieval results is achieved. When the answer entity is insufficient and the lowest level of retrieval has not been reached, the system can use the document ID as a metadata filter to intelligently focus the retrieval scope on deeper-level related documents, thereby performing more accurate and in-depth similarity matching, significantly improving the recall and accuracy of the answer. At the same time, a clear termination condition ensures that the system can provide clear feedback after exhausting all retrieval levels, avoiding invalid repeated retrieval. Finally, by inputting sufficient retrieval results along with user input into the core large language model, high-quality natural language answers can be generated, greatly improving the efficiency and experience of users obtaining information. This progressive and intelligent strategy combining retrieval and generation makes the question-answering system more robust and effective when handling complex queries.
[0072] As one possible implementation, metadata filters are used to filter candidate documents in real time based on their structured attributes, either before or simultaneously with vector similarity searches.
[0073] Metadata filters are mechanisms used to filter data based on its non-content attributes (i.e., metadata). In vector databases, in addition to the vector itself, each data point is typically associated with structured metadata, such as document ID, section ID, title, author, publication date, document type, topic tags, keywords, summary length, and hierarchical information. Metadata filters utilize these structured attributes to narrow down the search scope. Their role is to quickly eliminate candidate documents that do not meet specific criteria before or simultaneously with the time-consuming vector similarity search, thereby improving search efficiency and result accuracy. In terms of implementation, vector databases typically build inverted indexes, B-tree indexes, or hash indexes for metadata fields. When filter conditions are received, the system first uses these indexes to quickly locate the set of document IDs or vector IDs that meet the conditions.
[0074] Defining the timing of metadata filtering before or simultaneously with vector similarity search aims to maximize its impact on search efficiency. Filtering before vector similarity search means that before performing vector similarity calculations (e.g., using cosine similarity, Euclidean distance, etc.), the system first selects a small subset of candidate vectors from the entire vector database based on metadata filtering conditions, and then performs vector similarity calculations only on this subset. This significantly reduces the number of vector pairs that need to be calculated. Filtering simultaneously with vector similarity search typically requires the support of the underlying index structure of the vector database. For example, in some graph-based vector indexes (such as HNSW), during graph traversal, when a node is visited, its metadata can be checked synchronously to see if it meets the filtering conditions. If not, the node and its child nodes can be skipped, thus dynamically pruning the search path during the search process.
[0075] Based on the document's structured attributes, richer and more granular filtering dimensions are provided than simply using document IDs. A document's structured attributes refer to information stored in a structured form, beyond the main text content, that can be used for classification, retrieval, and filtering. These attributes are typically in key-value pair form, such as document ID, chapter ID, title, author, publication date, document type, topic tags, keywords, summary length, and hierarchical information. By utilizing these attributes, more precise search scope can be achieved, such as searching only documents within a specific chapter, author, or time range. When data is stored in a vector database, in addition to the vector itself, a series of structured metadata fields are defined and stored for each document or text block. Vector databases typically provide specialized query languages or APIs that allow users to specify metadata filtering conditions in vector search requests.
[0076] "Real-time filtering of candidate documents" emphasizes the immediacy and efficiency of the filtering process, ensuring that filtering is completed within the user's waiting time without introducing significant delays. This relies on the efficient indexing of metadata fields by the vector database, enabling metadata queries to return results quickly. Furthermore, frequently used metadata or indexes may be loaded into memory to reduce disk I / O latency. In large-scale systems, metadata filtering operations can be parallelized, executing simultaneously on multiple nodes to accelerate the filtering process.
[0077] By using a metadata filter to screen candidate documents in real time based on their structured attributes before or simultaneously with vector similarity search, this invention significantly improves the efficiency and accuracy of progressive retrieval question answering methods. When the evaluation agent determines that the current recall results are insufficient and a next-level retrieval is needed, the metadata filter no longer relies solely on document IDs but can utilize richer structured attributes (such as chapter IDs and hierarchical information) for pre-filtering. This refined filtering mechanism can significantly reduce the vector space to be searched before or simultaneously with the computationally intensive operation of vector similarity calculation, thereby accelerating retrieval speed and reducing system resource consumption. Simultaneously, due to more precise filtering conditions, the recalled Top-N node texts will be more focused on specific content fragments pointed to by user intent and question words, effectively improving the success rate of subsequent evaluation agent's entity checking of answers and the accuracy of the final answer. This allows the entire progressive retrieval process to converge to the answer needed by the user more efficiently and intelligently.
[0078] Secondly, the present invention provides a progressive retrieval question answering system that executes the aforementioned progressive retrieval question answering method. The system includes a multi-level index database, a vector database, an intent classifier, and an evaluation agent.
[0079] After receiving a user's query, the system first performs a query complexity and intent analysis to dynamically determine the starting level for the initial search. A multi-level knowledge base index structure is pre-built (e.g., first-level document summary layer, second-level document chapter layer, third-level paragraph slice layer). If the question is determined to be a macro-level summary question, the search starting point is set to a shallow level (first-level document summary layer); if the question is determined to be a specific fact query question, the search starting point is set to a deeper level (second-level document chapter layer or third-level paragraph slice layer).
[0080] A lightweight "evaluation agent" is instantiated within the system. After completing the initial retrieval at each level, the text is not sent directly to the answer generation module, but rather first sent to the evaluation agent. The evaluation agent receives the user's question and the current recall context, uses preset logical reasoning rules to calculate the coverage of the current context for answering the question, and outputs an information sufficiency judgment result (Boolean value: true / false, or quantified score).
[0081] Execute branch routing based on the output of the evaluated agent: If the information sufficiency is determined to be "true" (or the score reaches the preset threshold), a dynamic termination instruction is triggered, cutting off the subsequent retrieval process and directly generating the final natural language answer from the user using the core large language model of the current context input answer generation module. If the information sufficiency is determined to be "false", a drill-down command is triggered. The system automatically extracts the association identifier (document ID or chapter ID) of the node with the highest relevance in the current level of search results, sets it as a metadata filter, and carries this filter to the next finer-grained level to initiate a precise range search.
[0082] The system receives text input by the user, sets the initial retrieval level through intent routing, performs a retrieval, recalls the retrieval results and inputs them into the evaluation agent for judgment, performs a new round of evaluation based on the judgment results, and finally determines that the concise text fragment is sufficiently complete. This fragment, along with the text input by the user, is then input into the retrieval enhancement and generation model to generate the final natural language answer to the user.
[0083] The core innovation of this invention lies in combining intent routing with an evaluation agent in a closed-loop control manner. This dynamically assesses information sufficiency during the retrieval process and triggers drill-down or termination as needed. This solves the problems of token resource waste, contextual interference, and lack of quality judgment capabilities in existing technologies, achieving significant reductions in the inference context length of large language models, elimination of irrelevant information interference, and improved question-answering accuracy. Intent routing dynamically sets the initial retrieval level based on user intent, avoiding unnecessary deep searches for simple questions. The evaluation agent executes information sufficiency judgment logic on the recall results. When the information is deemed sufficient, it immediately truncates the subsequent retrieval process, effectively reducing redundant text input. When the information is deemed insufficient, the system extracts the associated identifiers of the current recall node to construct a metadata filter, precisely drilling down to a finer-grained level to initiate a retrieval, ensuring a high signal-to-noise ratio in the input context. Because the above technical solution realizes the closed-loop process of "intent-adaptive starting point - multi-level retrieval gating - dynamic termination and drill-down", the system terminates the retrieval in a timely manner when there is sufficient information, which greatly reduces the context token consumption of large language model inference. At the same time, the layer-by-layer application of metadata filters eliminates context interference caused by the introduction of irrelevant information, suppresses model illusion from the root, and significantly improves the factual accuracy and response efficiency of question answering results.
[0084] To facilitate understanding of the technical solution of this application, further explanation is provided below with reference to specific embodiments.
[0085] Example 1 As a preferred embodiment, the solution of the present invention is implemented as follows: A corporate knowledge management system contains a large number of structured and unstructured documents, such as annual financial reports, product manuals, and internal policy documents. To improve the efficiency and accuracy of users searching for these documents, the system employs a progressive question-and-answer retrieval method.
[0086] Application Example: Intelligent Question-Answering System for Enterprise Financial Reports A user asked: "What was the specific revenue figure for the new energy vehicle business in the European market in the 2023 annual report?" (1) Intent analysis: The system determines that it is a micro-fact query and sets the initial starting point as the document chapter layer.
[0087] (2) Initial search of document chapter level: The summary of the chapter "2023 Annual Report - Core Business Segment Analysis" was retrieved by searching the entire document chapter level.
[0088] (3) Evaluation Agent Judgment: The evaluation agent found that the document chapter layer summary only mentioned "new energy vehicles have achieved growth in the European market", but did not provide specific "revenue" figures. The output is insufficient (Sufficient: False).
[0089] (4) Dynamic drill-down: The system extracts the section identifier id = 'Auto_EU_2023' and constructs a metadata filter filter={"section_id":"Auto_EU_2023"}.
[0090] (5) Secondary search slice layer: The search is carried out in the slice layer with the metadata filter. Due to the limitation of the metadata filter, the progressive search question answering system no longer performs full database matching, but only searches in the specific paragraphs of the chapter, accurately recalling the text block containing "new energy vehicles achieved revenue of 5.02 billion yuan in the European market".
[0091] (6) Re-evaluate: The evaluation agent reads the text block and confirms that it contains the specific number "5.02 billion yuan", which meets the query requirements. The output is sufficient (Sufficient: True).
[0092] (7) Generating the answer: The progressive retrieval question-answering system truncates the retrieval and inputs only this precise text block into the large language model. The model outputs: "The specific revenue of the new energy vehicle business in the European market in 2023 was RMB 5.02 billion." Since only one highly relevant piece of content was input, the token consumption is extremely low, and there is absolutely no interference from other business data.
[0093] The above application examples describe the following steps: 1. Preliminary preparation and data construction First, the system establishes a multi-level index database. The system parses various documents within the enterprise and configures the text data, parsing it into a tree structure. Based on this tree structure, a multi-level index is built, containing at least three levels. For example, the first level is the document summary layer, storing the metadata and a global summary for each document, with each node bearing a unique document ID. The second level is the document chapter layer, storing the titles and summaries of each chapter, with each node bearing both a document ID and a chapter ID. The third level is the slice layer, storing specific text blocks, with each text block bearing both a document ID and a chapter ID as metadata tags. Ultimately, the text data and the multi-level index constitute the multi-level index database.
[0094] The system vectorizes the data in the multi-level index database to obtain computer-understandable numerical data, which is then stored in a vector database. This means that document summaries, chapter summaries, and specific text blocks are all converted into high-dimensional vectors for similarity matching.
[0095] 2. User Query and Intent Recognition When user A enters a query text, such as "What was the company's revenue from new energy vehicle business in the European market in 2023?", the system executes step S3, and the configured intent classifier begins to work.
[0096] The system receives a question from user A. It then calls the intent classifier to categorize the question.
[0097] The classification rules are as follows: if the question contains macro-level features such as "summary" or "what is the main content", the initial retrieval level is set to the document summary level; if the question contains micro-level features such as "how much" or "specific parameters", the initial retrieval level is set to the document chapter level or slice level.
[0098] The intent classifier receives text input from user A and identifies the intent within the text. For example, the query might be identified as a "micro-fact query" or a "specific numerical query." The intent classifier then vectorizes the identified intent and outputs the corresponding intent label.
[0099] Based on the initial level determined by the intent classifier after classifying the questions (assuming it is the document summary level), the question vector is used to perform similarity matching in the vector database corresponding to that level, and the top-N node texts and their associated identifiers are retrieved.
[0100] The intent classifier inputs intent labels into the intent route. The intent route determines the initial retrieval level based on the intent labels. For specific numerical queries like "What was the company's revenue from new energy vehicles in the European market in 2023?", the intent route determines that the initial retrieval level should be set to the document chapter level, rather than the more macro-level document summary level. This avoids the inefficiency of traditional systems that start searching from the shallowest level regardless of the complexity of the question, resulting in the initial retrieval of a large amount of irrelevant information.
[0101] 3. Initial Search and Information Recall Intent routing distributes the intent to the corresponding document chapter layer in the vector database for similarity matching. Within the document chapter layer, the system recalls the Top-N node texts most similar to the user's query, along with their associated identifiers. For example, it recalls the summary of the chapter "2023 Annual Report - Core Business Segment Analysis," which may mention "significant growth in the European market for new energy vehicles."
[0102] 4. Assessment and Dynamic Drilling The system-configured evaluation agent begins to work. The evaluation agent is an agent with autonomous perception, decision-making, execution, and interaction capabilities. It can autonomously identify whether the text block corresponding to the chapter ID sufficiently contains the text fragments required by the user.
[0103] Input the user A's question and the recalled node text into the "evaluation agent". The evaluation agent executes the following internal decision logic: Analyze the core entities and interrogative words in the question, and check whether the recalled text contains the exact answer entity that supports the interrogative word.
[0104] Output the judgment result: If the answer entity is included, output the sufficiency indicator "Sufficient: True"; if the answer entity is not included but relevant clues are found, or if the answer entity is not included at all, output the sufficiency indicator "Sufficient: False".
[0105] The evaluation agent first receives the text input by user A and the text of the Top-N nodes retrieved for intent routing, along with their associated identifiers. Next, it parses the core entities (such as "2023", "European market", "new energy vehicle business", "revenue") and interrogative words (such as "how much") from the text input by user A.
[0106] The evaluation agent checks whether the text of the Top-N nodes contains answer entities that support the interrogative word. In this example, the evaluation agent finds that the recalled "2023 Annual Report - Core Business Segment Analysis" section summary only mentions "growth in the European market for new energy vehicles" but does not include specific "revenue" figures.
[0107] The system receives the output identifier of the evaluation agent: Once the identifier is deemed sufficiently sufficient, the retrieval loop ends, and the system, along with the text input by user A, inputs this concise and highly relevant text fragment into the core large language model.
[0108] If the identification is insufficient, the system checks if the slice level has been reached. If the slice level has not been reached: the system extracts the identifier of the currently recalled text (e.g., if the content is determined to be in a document, extract document id = 'A123'). Document id = 'A123' is set as a metadata filter. The search target level is lowered to the document section level, and with this filter applied, a new search is initiated at the document section level. The user question and the recalled node text are input into the evaluation agent for a new round of evaluation.
[0109] If the slice layer has been reached but the information is still insufficient: output the standard message "Insufficient information found in the knowledge base" to prevent the model from forcibly fabricating answers.
[0110] Based on the judgment logic of the evaluation agent: The system is deemed insufficient and has not reached the lowest retrieval level: because the currently recalled chapter summary does not contain specific revenue figures, and the document chapter level is not the lowest level (the slice level is even lower), the evaluation agent determines that the current information is insufficient. The system extracts the identifiers of the top-N node texts currently recalled, for example, extracting the chapter ID of the chapter, and sets it as a metadata filter. This metadata filter is used to filter candidate documents in real time based on the document's structural attributes before or simultaneously with vector similarity search. Intent routing descends the retrieval level to the slice level, carrying this metadata filter, and continues to distribute the intent to the vector database corresponding to the slice level for similarity matching. This mechanism avoids the problem of traditional systems blindly submitting incomplete information to a large language model when information is insufficient, leading to the model generating illusions or inaccurate answers, while reducing the waste of token resources through precise drill-down.
[0111] Secondary retrieval and re-evaluation: At the slice layer, the system applies a metadata filter, performing similarity matching only on text blocks related to the previously recalled chapters, accurately recalling text blocks containing "new energy vehicles achieved revenue of 5.02 billion yuan in the European market." The evaluation agent receives this text block again and checks if it contains the answer entity "revenue amount." This time, the evaluation agent confirms that the text block contains the specific number "5.02 billion yuan," satisfying user A's query requirement.
[0112] The assessment agent determines that the current information is sufficient. The system, along with the text input by user A, inputs this concise and highly relevant text fragment into the core large language model.
[0113] 5. Final Answer Generation The core large language model receives user A's query and evaluates the intelligent agent to determine that the concise text fragment is "sufficient", and generates the final natural language answer to user A: "The specific revenue of the new energy vehicle business in the European market in 2023 was 5.02 billion yuan". Through the aforementioned progressive retrieval and question-answering method, the system achieves "on-demand retrieval, dynamic evaluation, and precise positioning." Compared to the traditional one-time flat retrieval mode, this method avoids unnecessary shallow or deep retrieval through an intent-adaptive starting point mechanism. It uses a multi-level retrieval gating mechanism to evaluate information sufficiency at each level and dynamically drills down to finer-grained levels when information is insufficient. Simultaneously, it utilizes metadata filters to precisely narrow the retrieval scope. This significantly reduces the context token consumption of large language model inference, fundamentally eliminating contextual interference and illusion problems caused by irrelevant information, and improving the accuracy and response efficiency of the question-answering system. When information cannot be found at the lowest level, the system generates a "not found" prompt to answer the user, preventing the model from forcibly fabricating answers.
[0114] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, the disclosure, and the appended claims in carrying out the claimed invention. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality. A single processor or other unit can implement several functions listed in the claims. While different dependent claims may recite certain measures, this does not mean that these measures cannot be combined to produce good results.
[0115] Although the invention has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely exemplary descriptions of the invention as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if such modifications and modifications of the invention fall within the scope of the claims and their equivalents, the invention is also intended to include such modifications and modifications.
Claims
1. A progressive retrieval and question-answering method, characterized in that, The method includes: S1. Establish a multi-level indexed database; S2. Vectorize the data in the multi-level index database to obtain numerical data that can be understood by computers, and store it in the vector database; S3. Configure the intent classifier, which includes intent routing. The intent classifier performs the following steps: S30. Receive text input by the user and identify the intent in the text; S31. Vectorize the intent and output the intent label; S32. Input the intent label into the intent route; the intent route determines the initial level of the retrieval hierarchy; S33. Intent routing distributes the intent to the vector database for similarity matching, and retrieves the text of the Top-N nodes and their associated identifiers; S4. Configure the evaluation agent, which performs the following steps: S40. Receive the text input by the user and the text of the Top-N nodes retrieved for intent routing, along with their associated identifiers; S41. Parse the core entities and interrogative words in the user's input text; S42. Check whether the text of the Top-N nodes contains answer entities that support the question word, and execute the judgment logic.
2. The progressive retrieval and question-answering method according to claim 1, characterized in that, Multi-level indexed databases include: S20. Configure the text data and parse the text data into a tree structure; S21. Construct multi-level indexes based on a tree structure; S22. A multi-level indexed database consists of text data and multi-level indexes.
3. The progressive retrieval and question-answering method according to claim 2, characterized in that, A multi-level index must have at least three levels.
4. The progressive retrieval and question-answering method according to claim 3, characterized in that, The multi-level index includes a document summary layer: storing the metadata and a global summary of the entire document, with each node having a unique document ID.
5. The progressive retrieval and question-answering method according to claim 3, characterized in that, The multi-level index includes a document chapter layer: storing the title and summary of each chapter, with each node having a document ID and a chapter ID.
6. The progressive retrieval and question-answering method according to claim 3, characterized in that, Multi-level indexes contain slice layers: storing specific text blocks, each text block with a document ID and chapter ID as metadata tags.
7. The progressive retrieval and question-answering method according to claim 1, characterized in that, The evaluation agent is an intelligent agent with autonomous perception, decision-making, execution, and interaction capabilities; it can autonomously identify whether the text block corresponding to the chapter ID sufficiently contains the text fragments required by the user.
8. The progressive retrieval and question-answering method according to claim 1, characterized in that, The evaluation logic for the agent is to check whether the text of the Top-N nodes contains answer entities that support the question word; When the check results do not contain the answer entity or contain only part of the answer entity, it is determined that the results are insufficient and have not reached the lowest level of the retrieval level. The document IDs of the currently recalled Top-N node texts are extracted, the document IDs are set as metadata filters, and the intent is further distributed to the vector database corresponding to the next level of the retrieval level for similarity matching. If the search results do not contain the answer entity or contain only part of the answer entity, it is determined to be insufficient and has reached the lowest level of the search hierarchy, and a "not found" message is generated to answer the user. If the check result contains the answer entity, it is deemed sufficient, and together with the text entered by the user, it is input into the core large language model to generate the final natural language answer to the user.
9. A progressive retrieval and question-answering method according to claim 8, characterized in that, Metadata filters are used to filter candidate documents in real time based on their structured attributes, either before or simultaneously with vector similarity searches.
10. A progressive retrieval question-answering system, executing the progressive retrieval question-answering method according to any one of claims 1 to 9, characterized in that, This includes a multi-level index database, a vector database, an intent classifier, and an evaluation agent; The system receives text input by the user, sets the initial retrieval level through intent routing, performs a retrieval, recalls the retrieval results and inputs them into the evaluation agent for judgment, performs a new round of evaluation based on the judgment results, and finally determines that the concise text fragment is sufficiently complete. This fragment, along with the text input by the user, is then input into the retrieval enhancement and generation model to generate the final natural language answer to the user.