Coarse-to-fine multi-document question and answer method based on search head
Through a two-stage method based on the search head, the background document is gradually filtered and attention is dynamically adjusted, which solves the problem of insufficient retrieval and reasoning capabilities of large language models when inputting long contexts, and realizes efficient multi-document question-and-answer task processing.
Patent Information
- Application Number
- CN202510408103.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-22
AI Technical Summary
When handling multi-document Q&A tasks, existing large language models have problems with insufficient retrieval and reasoning capabilities during long context input, especially in balancing search accuracy and recall, and it is difficult to effectively filter background documents and interfere documents, affecting model performance.
A two-stage method based on search heads is adopted, firstly, the background document is quickly eliminated through coarse-grained filtering, and the candidate document set is retained, and then the interfering document is further suppressed through fine-grained guidance. The dynamic attention bias mechanism is used to enhance attention to the golden evidence document, and the document input sequence is optimized by combining position reordering and attention bias injection.
It significantly improves the efficiency and accuracy of long context processing, solves the trade-off between accuracy and recall in the existing technology, adapts to different models and reduces noise interference, and improves the performance of multi-document Q&A tasks.
Smart Images

Figure CN120353892A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models (LLMs), and particularly to a coarse-to-fine multi-document question answering method based on a retrieval head. Background Art
[0002] In the prior art, when large language models (LLMs) handle multi-document question answering tasks, although the input context length and inference ability have been extended through improvements in the model architecture and optimization of training methods, they still face limitations in retrieval and inference capabilities. Existing methods mainly alleviate this problem through external prompting strategies and attention mechanisms. For example, some studies use external prompting strategies to guide the model to extract key information related to the question from long inputs, and then answer the question in combination with the original context. Other studies analyze the internal attention mechanism of the model and use retrieval heads to identify key information. These methods have improved the model's long-context processing ability to a certain extent, but there are still challenges in balancing retrieval precision and recall.
[0003] However, the prior art has the following disadvantages when dealing with long-context inputs: First, the effectiveness of external prompting strategies is limited by the instruction-following ability of LLMs, especially in long inputs, this limitation is more obvious. Second, although retrieval heads can be used to identify key information, these methods have difficulties in balancing retrieval precision and recall. Increasing recall will introduce more irrelevant information, affecting the inference process of the model, while increasing retrieval precision will reduce recall. In addition, existing methods are difficult to effectively filter background documents and interfering documents when dealing with multi-document question answering tasks, resulting in a significant decline in the retrieval and inference capabilities of the model in long-context inputs. Therefore, the prior art still has problems with insufficient retrieval and inference capabilities when dealing with long-context inputs, which affects the performance of the model in multi-document question answering tasks.
[0004] The information disclosed in this background art section is only intended to deepen the understanding of the overall background art of the present invention, and should not be regarded as an admission or any form of implication that this information constitutes prior art known to those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a coarse-to-fine multi-document question answering method based on a retrieval head to solve the technical problems existing in the prior art.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] The present invention provides a coarse-to-fine multi-document question answering method based on a retrieval head, including the following steps:
[0008] S1, Coarse-grained filtering: quickly eliminate a large number of background documents and retain a candidate document set;
[0009] S2, Fine-grained guidance: Further suppress interfering documents from the candidate set and enhance the attention to the gold evidence document.
[0010] Furthermore, the specific implementation process of step S1 is as follows:
[0011] S1.1. The key component includes a retrieval head selection module: Pre-select the Top-K retrieval heads through the validation set, and calculate the retrieval scores of each document, where α h (q, d i ) represents the attention weight between the query q and the document d, and β h (d i ) represents the score of the current retrieved document, H ret represents the set of retrieval heads in the first stage, and η(h) represents the retrieval score of the attention head h, and H represents the set of all attention heads;
[0012]
[0013] H ret = Top-K ~ η(h), h ∈ H;
[0014] S1.2. Screen and re-rank the documents through the document screening and re-ranking module, where γ h (d) represents the re-ranking score corresponding to the document d, and D′ represents the sorted document sequence;
[0015]
[0016] S1.3. Joint retrieval module: Select the Top-M1 documents for each retrieval head and take the union to form a candidate set;
[0017] S1.4. Position-based re-ranking: Place the high-score documents at the end of the input sequence according to the total document scores, and utilize the "end preference" characteristic of LLMs to improve the inference effect.
[0018] Furthermore, the specific implementation process of step S2 is as follows:
[0019] S2.1. The key component includes an interfering document identification module: Use another set of retrieval heads to screen out the Top-M2 candidate evidence documents, and D candidate represents the candidate documents in the second stage;
[0020]
[0021] S2.2. Perform attention bias injection through the attention bias injection module;
[0022] S2.3, Dynamic Bias Injection: During model inference, apply a positive attention bias to the Tokens of the candidate evidence documents to amplify their attention weights;
[0023] S2.4, softmax Reweighting: Automatically suppress the attention scores of non-critical documents through the softmax function, where A represents the attention score, Q represents the Query matrix, K represents the Key matrix, B represents the dynamic bias, and d represents the vector dimension;
[0024]
[0025] Adopting the above technical solutions, the present invention has the following beneficial effects:
[0026] The CAFE (Coarse-to-Fine Information Seeking) proposed in this patent systematically solves the precision-recall trade-off problem in long context processing through a two-stage retrieval and attention guidance mechanism, while avoiding modifications to the model architecture or reliance on large-scale pre-training. Its core innovation lies in transforming the human cognitive process of "gradual screening - focused reasoning" into a computable retrieval head collaboration and attention dynamic adjustment process, providing an efficient and scalable solution for multi-document question answering tasks.
[0027] The two-stage retrieval and attention guidance mechanism significantly optimizes the long context processing flow. The core differences from existing methods are reflected in:
[0028] 1. Phased processing flow: Divide the retrieval into two stages: coarse-grained filtering (background document elimination) and fine-grained guidance (distracting document attention suppression), gradually narrowing the retrieval scope.
[0029] 2. Dynamic attention adjustment: During the inference stage, directly intervene in the model's attention weights for key documents through post-hoc attention steering, rather than relying on external cues or single retrieval.
[0030] 3. Multi-retrieval head collaboration: Select different combinations of retrieval heads for different stages. For example, use high-recall retrieval heads in the coarse-grained stage and high-precision retrieval heads in the fine-grained stage to achieve complementary optimization. Description of the Drawings
[0031] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] Figure 1 This is the system architecture diagram of the coarse-to-fine multi-document question answering method based on the retrieval head provided by the embodiments of the present invention. Detailed implementation manners
[0033] Next, the technical solutions of the present invention will be described clearly and completely with reference to the accompanying drawings. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] The following will describe the specific implementation manners of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only for explaining and illustrating the present invention, and are not used to limit the present invention.
[0035] The architecture of the coarse-to-fine multi-document question answering method based on the retrieval head provided by the present invention is as Figure 1 shown (the documents in the accompanying drawings are classified into three colors: red, blue, and yellow), and includes the following core components and steps:
[0036] Stage 1: Coarse-Grained Filtering:
[0037] Objective: Quickly eliminate a large number of background documents (yellow) and retain the candidate document set (blue and red).
[0038] The key components include a retrieval head selection module: preselecting the top-K retrieval heads through a validation set and calculating the retrieval scores of each document; a document screening and re-ranking module; a joint retrieval module: selecting the top-M1 documents for each retrieval head and taking the union to form a candidate set.
[0039] Locality-Based Re-Ranking: Place the high-score documents at the end of the input sequence according to the total document scores, and utilize the "end preference" characteristic of LLMs to improve the inference effect.
[0040] Effect: Significantly reduce the input length (for example, from 100 documents to 10 documents), while retaining a candidate set with a high recall rate.
[0041] Stage 2: Fine-Grained Steering:
[0042] Objective: Further suppress the interfering documents (blue) from the candidate set and enhance the attention to the golden evidence documents (red).
[0043] The key components include the interference document identification module: using another set of retrieval heads to filter out the Top-M2 candidate evidence documents; the attention bias injection module (labeled as Attention Steering); dynamic bias injection: during model inference, a positive attention bias is applied to the token of the candidate evidence document, for example, setting δ=5 to amplify its attention weight; softmax reweighting: automatically suppressing the attention score of non-critical documents through the softmax function.
[0044] Effect: Without removing interfering documents, reduce their influence on answer generation while avoiding recall loss.
[0045] The order of steps:
[0046] Preprocessing: Preselect two-stage retrieval heads on the validation set;
[0047] Phase 1 execution: input long context → retrieval head screening → document reranking → output candidate document set.
[0048] Phase 2 execution: candidate document set → secondary retrieval head screening → attention bias injection → generate final answer.
[0049] Compared with the prior art, the objective improvements achieved by the present invention through differences include:
[0050] Balance precision and recall: The coarse-grained stage uses joint screening by multiple search heads (high recall), and the fine-grained stage uses attention bias to suppress interference (high precision). The two complement each other to solve the single-stage limitations of existing technologies.
[0051] Reducing noise interference: The position-based reordering and attention bias mechanism directly alleviates the "Lost-in-the-Middle" problem in long contexts.
[0052] Training independence: CAFE does not require fine-tuning of the model. It can adapt to different LLMs (such as Llama and Mistral) only through intervention in the inference stage, avoiding the problem of short text performance degradation caused by long-context pre-training in existing technologies.
[0053] The applicant has verified the actual application effect of this application through experiments, the specific situation is as follows:
[0054] In the HotpotQA-32K task, CAFE achieved a SubEM score of 68.5% on the Llama-3.1-8B model, which is 11.4% higher than the baseline method (such as Vanilla RAG's 61.5%). Its core advantages are:
[0055] 1. Long context adaptability: As the input length increases from 8K to 32K, the performance of CAFE only drops by 2.1%, while that of traditional methods (such as In-Context Retrieval) drops by 28%.
[0056] 2. Multi-model generalization: On the Mistral-3-7B and Phi-3.5-Mini models, the F1 scores of CAFE increase by 9.3% and 7.5% respectively, verifying its versatility.
[0057] The following further elaborates on the implementation process of the technical solution of this application in combination with specific cases:
[0058] For scenarios of processing and analyzing a large number of documents, the CAFE technology proposed in this patent can effectively filter out irrelevant background information and prompt the model to focus on individual important documents.
[0059] In the emergency scenario in the medical field, doctors often face the pressure of processing a large amount of medical data. A patient's five-year electronic medical record may contain hundreds of examination reports, imaging records, and medication histories, and rapid diagnosis requires accurately locating key clues from this information. The two-stage processing mechanism of CAFE technology demonstrates unique advantages in this scenario.
[0060] The first stage: Coarse-grained filtering - Intelligent cleaning of medical data:
[0061] When the system accesses the patient's full medical data, CAFE first initiates document-level screening.
[0062] Collect more than 5,000 annotated medical records (annotating duplicate examination items / irrelevant medical histories / key diagnoses) for the calibration and screening of the medical retrieval head. The conventional attention head (97%) maintains basic semantic understanding, while the medical retrieval head (3%) is used for the detection and analysis of the vertical domain, focusing on the association between medical data and medical terms.
[0063] The medical dedicated retrieval head deployed with calibration data (only accounting for 3% of the total attention heads) automatically identifies and eliminates duplicate conventional examination reports (such as stable numerical records in multiple blood routine tests), and at the same time filters out historical medical records irrelevant to the current chief complaint (such as dental visit records from many years ago). Compress the patient's historical medical record of more than 1,000 pages to the core 50 pages (compression rate 95%).
[0064] This process is not simply deleting data, but establishing the association relationships between multiple medical data for further analysis and diagnosis.
[0065] The second stage: Fine-grained guidance - Precise focus on key information:
[0066] On the compressed core data set, CAFE further refines and focuses on key information:
[0067] After completing the first-stage coarse-grained filtering, the system enables a second set of independently trained retrieval heads for in-depth semantic enhancement. These specifically optimized retrieval heads perform multi-dimensional feature analysis on the compressed document set, generating dynamic weight scores by calculating the inter-layer attention correlation between each document and the question. Adopting the Top-M screening enhancement strategy, the system automatically constructs a candidate evidence document set - this set is continuously optimized through iterative calculations to ensure coverage of over 90% of the key document information. Documents not included in the candidate set are marked as "potential interference sources", and their semantic influence will be effectively suppressed in subsequent processing.
[0068] Core feature extraction: Automatically capture important index descriptions in medical records (such as keywords like "continuous chest pain", "abnormal ST segment", etc.); Cross-document association: Dynamically match the descriptions in inspection reports, patient complaints, and medical guidelines; Risk priority: Automatically highlight in red and warn about contradictory information (such as inconsistent symptoms and test results).
[0069] Summary: Practical application features:
[0070] Intelligent filtering: Automatically hide duplicate inspection reports and irrelevant medical histories, quickly locating key clues scattered across hundreds of pages of medical records.
[0071] Dynamic focusing: Strengthen the citation weight of relevant medical guidelines according to symptom characteristics, avoiding typical misdiagnosis caused by information overload (such as confusing angina pectoris with pulmonary embolism).
[0072] Auxiliary decision-making: Generate a list of etiological speculations with an evidence chain (such as "high possibility of myocardial infarction: based on 1 / 2 / 3").
[0073] In summary, existing methods usually face the problem of difficulty in balancing retrieval precision and recall rate when dealing with long-context inputs, resulting in the model's inability to effectively distinguish key evidence documents and interference documents in multi-document question-answering tasks. In addition, most of these methods rely on external retrieval models or simple prompting strategies, failing to fully utilize the internal retrieval capabilities of the model and being easily negatively affected by background documents and interference documents under long-text inputs. To address the above deficiencies, this patent proposes an information retrieval method (CAFE) based on a two-stage combination of coarse and fine filtering, which significantly improves the model's retrieval and reasoning capabilities under long-context inputs by gradually filtering background documents and guiding the model to focus on key documents. In addition, existing methods usually only perform simple document filtering or directly use retrieval heads for information extraction when dealing with long contexts, but these methods do not fully optimize the input order of documents, resulting in the model's inability to fully utilize key information during the reasoning process. In contrast, this patent further improves the efficiency and accuracy of long-context processing by introducing a dynamic document rearrangement strategy based on retrieval heads, placing key documents at the end of the input sequence to better utilize the locality characteristics of the model.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A coarse-to-fine multi-document question answering method based on a retrieval head, characterized in that, It includes the following steps: S1. Coarse-grained filtering: quickly eliminate a large number of background documents and retain the candidate document set; S2. Fine-grained guidance: further suppress interfering documents from the candidate set and enhance the attention to the golden evidence documents.
2. The coarse-to-fine multi-document question answering method based on a retrieval head according to claim 1, wherein The specific implementation process of step S1 is as follows: S1.
1. The key component includes a retrieval head selection module: preselect the Top-K retrieval heads through the validation set, and calculate the retrieval scores of each document, where α h (q, d i ) represents the attention weight between the query q and the document d, β h (d i ) represents the score of the current retrieved document, H ret represents the set of retrieval heads in the first stage, and η(h) represents the retrieval score of the attention head h, and H represents the set of all attention heads; H ret = Top-K ~ η(h), h ∈ H; S1.
2. Screen and reorder the documents through the document screening and reordering module, where γ h (d) represents the reordering score corresponding to document d, and D′ represents the sorted document sequence; S1.
3. Joint retrieval module: select the Top-M1 documents for each retrieval head and take the union to form the candidate set; S1.
4. Position-based reordering: place the high-score documents at the end of the input sequence according to the total document score, and utilize the "end preference" characteristic of LLMs to improve the inference effect.
3. The coarse-to-fine multi-document question answering method based on a retrieval head according to claim 1, wherein The specific implementation process of step S2 is as follows: S2.
1. The key component includes an interference document recognition module: Use another set of retrieval headers to screen out the Top-M2 candidate evidence documents, D candidate which represents the candidate documents in the second stage; S2.
2. Perform attention bias injection through the attention bias injection module; S2.
3. Dynamic bias injection: apply a positive attention bias to the Tokens of the candidate evidence documents during model inference to amplify their attention weights; S2.
4. Softmax reweighting: automatically suppress the attention scores of non-critical documents through the softmax function, where A represents the attention score, Q represents the Query matrix, K represents the Key matrix, B represents the dynamic bias, and d represents the vector dimension;
Citation Information
Cited By
Multi-Word document question and answer method and system based on knowledge graph and RAG
CN121958631A