Retrieval enhancement generation method and system for optimizing language big model attention based on layer knowledge
By adding boundary markers to the language model and evaluating paragraph correlation using each layer embedding representation, combined with the attention mask based on layer knowledge selection and correlation guidance, the problems of insufficient information attention and inaccurate answers faced by the language model in knowledge-intensive tasks are solved, and more accurate answer generation is achieved.
Patent Information
- Application Number
- CN202411946656.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-03
AI Technical Summary
Existing language models face problems such as factual hallucinations, outdated knowledge and lack of domain-specific expertise when dealing with knowledge-intensive tasks, and in long-context scenarios, they are prone to insufficient attention to intermediate information due to ‘edge focus’ attention, resulting in inaccurate answers.
By adding boundary markers at the beginning and end positions of each paragraph, and using the embedding representation of each layer of the language model to evaluate the correlation and confidence of the paragraph and the problem, combining the entropy-based layer knowledge selection module to dynamically determine the applicability of knowledge to the paragraph, update the hierarchical correlation guidance weight, and construct a correlation-guided attention mask to achieve the attention pattern of intermediate focus.
It effectively alleviates the problems of distraction and position deviation, and can more accurately identify and utilize highly relevant search paragraphs to the question in open question-and-answer tasks, thereby generating more accurate answers.
Smart Images

Figure CN120086320A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of natural language processing, and in particular, relates to a retrieval-enhanced generation method and system for optimizing the attention of a language large model based on hierarchical knowledge. Background Art
[0002] In recent years, large language models (LLMs) have demonstrated remarkable performance, scalability, and adaptability in a variety of natural language processing tasks. Although LLMs can perform well with their large number of parameters and rich pre-trained knowledge, they face challenges in processing knowledge-intensive tasks, including factual hallucinations, knowledge obsolescence, and lack of domain-specific expertise.
[0003] To address the above problems, Retrieval-Augmented Language Models (RALMs) have become a mainstream solution. RALMs aim to enhance generation and improve the accuracy of answers by enabling LLMs to dynamically refer to additional information by integrating knowledge from external corpora. Specifically, RALMs first use information retrieval (IR) methods to retrieve several paragraphs from a web search engine or database, and then the retrieval fusion part concatenates these paragraphs as context into the LLMs to generate the final answer. In the implementation of RALMs, different paragraph fusion methods have been proposed to optimize the way the model processes the retrieved information. To effectively integrate the retrieved paragraph information, different RALM implementations have proposed a variety of paragraph fusion methods, aiming to improve the utilization efficiency of the retrieved content when the model generates answers. Currently, RALMs can be divided into three ways: query-based, logits-based, and latent representation-based fusion methods.
[0004] Query-based fusion method: By designing templates, the retrieved text segments are directly concatenated with the input question or features, enabling the LLMs to directly refer to the retrieved content when generating answers.
[0005] Logits-based fusion: Combines the output probability distributions of the input and the retrieved paragraphs, and optimizes the prediction results of the model by citing relevant examples.
[0006] Latent representation-based fusion: Integrates the retrieved paragraphs into the hidden state of the model through attention or weighted addition to enhance the internal representation. By deeply integrating the semantic information of the retrieved text segments into the hidden layer representation of the model, the model can dynamically refer to the retrieved content during generation, thereby enhancing the context relevance of the generation.
[0007] However, in the above method, the model tends to focus more on the beginning and ending positions of the context while ignoring the content in the middle; for long context scenarios, the model may pay insufficient attention to the key information in the middle of the context due to this "edge-focus" attention, resulting in "loss of middle information"; due to the limitations of information retrieval technology, the retrieved paragraphs are not always completely relevant to the given question. Therefore, the LLM may be interfered by irrelevant information when processing these paragraphs, leading to inaccurate answers. Summary of the Invention
[0008] In view of the deficiencies in the prior art, the present invention provides a retrieval-enhanced generation method and system for optimizing the attention of a large language model based on layer knowledge.
[0009] In a first aspect, the present invention provides a retrieval-enhanced generation method for optimizing the attention of a large language model based on layer knowledge, including:
[0010] Obtain a user question;
[0011] Retrieve multiple paragraphs from a knowledge base according to the user question;
[0012] Add boundary markers at the beginning and ending positions of each paragraph and mark the question part before and after the user question text;
[0013] Use the embedding representations of each layer of the LLM to evaluate the relevance score and confidence score of each paragraph to the question;
[0014] Use an entropy-based layer knowledge selection module to calculate the applicability of the knowledge of each layer of the LLM to the paragraph;
[0015] Update the hierarchical relevance guidance weights according to the relevance score, confidence score, and the applicability of the knowledge of each layer of the LLM to the paragraph;
[0016] Construct a relevance-guided attention mask according to the updated hierarchical relevance guidance weights to achieve relevance-aware paragraph fusion;
[0017] Generate a final answer through forward inference of the LLM according to the relevance-aware paragraph fusion result.
[0018] Optionally, the use of an entropy-based layer knowledge selection module to calculate the applicability of the knowledge of each layer of the LLM to the paragraph includes:
[0019] Calculate the applicability of the knowledge of each layer of the LLM to the paragraph according to the following formula:
[0020]
[0021] where H iThe applicability of the knowledge of each layer of the LLM to paragraph i; K is the scaling factor; n' is the total number of tokens in paragraph i; is the attention weight from the last token to the j-th token in paragraph i.
[0022] Optionally, updating the hierarchical relevance guidance weight according to the relevance score, confidence score, and the applicability of the knowledge of each layer of the LLM includes:
[0023] Calculating the updated relevance guidance weight of the q-th layer in the LLM for paragraph i according to the following formula
[0024]
[0025] where β is a hyperparameter used to balance the evaluation of the current layer and the previous layer of the LLM; H i is the applicability of the knowledge of each layer of the LLM to paragraph i; r i is the relevance between paragraph i and the question; c i is the confidence score between paragraph i and the question; is the relevance guidance weight of the (q - 1)-th layer in the LLM for paragraph i.
[0026] Optionally, the first aspect further includes:
[0027] Calculating the relevance loss L according to the following formula re :
[0028]
[0029] where n is the total number of paragraphs retrieved from the knowledge base according to the user's question; is the true relevance label of paragraph i; r i is the relevance score of paragraph i estimated by the LLM;
[0030] Calculating the confidence loss L according to the following formula co :
[0031]
[0032] where c i is the confidence score of paragraph i estimated by the LLM; K is the total number of external models; M k represents k different external models; I(·) is the exponential function; q' represents the user's question;
[0033] Calculating the diversity loss L according to the following formula di :
[0034]
[0035] Among them, H(·) is the entropy function;
[0036] The sum of the relevance loss, confidence loss, and diversity loss is used as the total loss of the hierarchical paragraph evaluator to evaluate the relevance score and confidence score of each paragraph to the question using the embedded representations transformed by each layer of the LLM.
[0037] In a second aspect, the present invention provides a retrieval-enhanced generation system based on layer knowledge to optimize the attention of a large language model, including:
[0038] An acquisition module for acquiring a user question;
[0039] A retrieval module for retrieving multiple paragraphs from a knowledge base according to the user question;
[0040] A marker addition module for adding boundary markers at the beginning and end positions of each paragraph and marking the question part before and after the user question text;
[0041] A first evaluation module for evaluating the relevance score and confidence score of each paragraph to the question using the embedded representations of each layer of the LLM;
[0042] A first calculation module for calculating the applicability of the knowledge of each layer of the LLM to the paragraph using an entropy-based layer knowledge selection module;
[0043] An update module for updating the hierarchical relevance guidance weight according to the relevance score, confidence score, and the applicability of the knowledge of each layer of the LLM to the paragraph;
[0044] A construction module for constructing a relevance-guided attention mask according to the updated hierarchical relevance guidance weight to achieve relevance-aware paragraph fusion;
[0045] An answer generation module for generating a final answer through forward inference of the LLM according to the relevance-aware paragraph fusion result.
[0046] Optionally, the first calculation module includes:
[0047] A first calculation unit for calculating the applicability of the knowledge of each layer of the LLM to the paragraph according to the following formula:
[0048]
[0049] where H i is the applicability of the knowledge of each layer of the LLM to paragraph i; K is a scaling factor; n' is the total number of tokens in paragraph i; is the attention weight from the last token to the j-th token in paragraph i.
[0050] Optionally, the update module includes:
[0051] A second calculation unit for calculating the updated relevance guidance weight of the q-th layer in the LLM for passage i according to the following formula
[0052]
[0053] where β is a hyperparameter for balancing the evaluation of the current layer and the previous layer of the LLM; H i is the applicability of the knowledge of each layer of the LLM to passage i; r i is the relevance between passage i and the question; c i is the confidence score between passage i and the question; is the relevance guidance weight of the (q - 1)-th layer in the LLM for passage i.
[0054] Optionally, the second aspect further includes:
[0055] A second calculation module for calculating the relevance loss L according to the following formula re :
[0056]
[0057] where n is the total number of passages retrieved from the knowledge base according to the user's question; is the true relevance label of passage i; r i is the relevance score of passage i estimated by the LLM;
[0058] A third calculation module for calculating the confidence loss L according to the following formula co :
[0059]
[0060] where c i is the confidence score of passage i estimated by the LLM; K is the total number of external models; M k represents k different external models; I(·) is the exponential function; q' represents the user's question;
[0061] A fourth calculation module for calculating the diversity loss L according to the following formula di :
[0062]
[0063] where H(·) is the entropy function;
[0064] The second evaluation module is used to take the sum of the relevance loss, confidence loss, and diversity loss as the total loss of the hierarchical paragraph evaluator, so as to evaluate the relevance score and confidence score of each paragraph to the question by using the embedded representations converted by each layer of the LLM.
[0065] In a third aspect, the present invention provides a computer device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the retrieval-enhanced generation method based on layer knowledge optimization of the language large model attention described in the first aspect are implemented.
[0066] In a fourth aspect, the present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the retrieval-enhanced generation method based on layer knowledge optimization of the language large model attention described in the first aspect are implemented.
[0067] The present invention provides a retrieval-enhanced generation method and system based on layer knowledge optimization of language large model attention. The method utilizes the diverse knowledge of each layer of the LLM to accurately predict relevance and confidence; since the knowledge of not all layers contributes equally to relevance evaluation, an entropy-based layer knowledge selection method is used to dynamically determine which layers of knowledge are applicable to the channels; in order to reduce the interference and position bias of irrelevant channels, a relevance-guided attention mask is adopted, enabling the question tokens to selectively focus on the retrieved channels to achieve an intermediate-focused attention pattern, which can more accurately identify and utilize the retrieved paragraphs highly relevant to the question in open-ended question answering tasks, thereby generating more accurate answers. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0069] Figure 1 It is a schematic flowchart of a retrieval-enhanced generation method based on layer knowledge optimization of the language large model attention provided by an embodiment of the present invention;
[0070] Figure 2 It is a graph showing the performance results of different model types on each dataset provided by an embodiment of the present invention;
[0071] Figure 3 It is a graph showing the experimental results of LKG-RALM under different ablation settings provided by an embodiment of the present invention;
[0072] Figure 4Performance trend chart of LKG-RALM and baseline model provided by embodiments of the present invention on different datasets;
[0073] Figure 5 Performance trend chart of LKG-RALM provided by embodiments of the present invention for noise and irrelevant information;
[0074] Figure 6 Comparison result chart of the efficiency and accuracy of LKG-RALM and baseline model provided by embodiments of the present invention;
[0075] Figure 7 Model performance comparison chart trained under different data scales provided by embodiments of the present invention;
[0076] Figure 8 Schematic structural diagram of a retrieval-enhanced generation system based on layer knowledge optimization of the attention of a large language model provided by embodiments of the present invention. Detailed implementation manners
[0077] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0078] Embodiment 1
[0079] As Figure 1 shown, this embodiment provides a retrieval-enhanced generation method based on layer knowledge optimization of the attention of a large language model, including:
[0080] Step 101, obtain a user question.
[0081] Step 102, retrieve multiple paragraphs from the knowledge base according to the user question.
[0082] In this step, after receiving the question input by the user, relevant paragraphs are retrieved from the knowledge base through a hybrid retriever. The hybrid retriever includes two important components: a BM25 retriever and a dense retriever. Among them, the BM25 retriever is implemented based on ElasticSearch and mainly relies on keyword matching for retrieval; the dense retriever is implemented based on the FAISS index and retrieves by calculating semantic similarity. The retriever will finally return the n paragraphs with the highest similarity [1, 2,..., n]. In practical applications, the performance of the retriever will directly affect the quality of the final answer. Therefore, it is necessary to improve the retrieval effect through means such as parameter tuning and index optimization. In addition, the number of returned paragraphs n needs to be reasonably set to balance coverage and computational efficiency.
[0083] Step 103: Add boundary markers at the beginning and end of each paragraph and mark the question part before and after the user question text.
[0084] The core of this step is to add explicit structured markers to the retrieved content to help the model better understand and process the input information. Specifically, add the [d] marker at the beginning of each retrieved paragraph and the [e_i] marker at the end, where i represents the paragraph number. At the same time, add the [q'] and [eq'] markers at the beginning and end of the user question q' respectively. The design of these special markers fully considers the requirements of subsequent processing. They can help the model clearly distinguish the boundaries of different paragraphs and identify the question part, thus providing important structured information for subsequent relevance evaluation. These markers will also gradually learn their special semantic meanings during the model training process.
[0085] Step 104: Use the embedding representations of each layer of the LLM to evaluate the relevance score and confidence score of each paragraph with respect to the question.
[0086] In this step, the hierarchical knowledge inside the LLM is used to evaluate the relevance of the retrieved paragraphs to the question. First, evaluate the relevance of each paragraph to the question through the embedding representations of each layer of the LLM. Specifically, add a trainable low-rank weight (LoRA) to each decoder layer to adapt to the relevance evaluation task. For each layer, use a blocked bidirectional attention mask to enhance the context understanding within the paragraph. Extract the hidden state from the last marker of each paragraph and the question as the sentence embeddings: e i represents the paragraph embedding, e q' represents the question embedding. These embeddings are processed through dropout and adapter layers to improve robustness, and then two cross-attention components are used to calculate the relevance score r i between paragraph i and the question, as well as the confidence score c i .
[0087] Step 105: Use an entropy-based layer knowledge selection module to calculate the applicability of the knowledge of each layer of the LLM to the paragraph.
[0088] To ensure the effective use of layer-specific knowledge in paragraph evaluation, this embodiment proposes entropy-based layer knowledge selection to identify the most informative and context-rich hierarchical representation in each paragraph. For each paragraph, calculate the attention distribution entropy H i from the last marker of paragraph i to other markers as the applicability of the knowledge of each layer of the LLM to paragraph i. A higher entropy value indicates that the sentence embedding captures more extensive context information.
[0089] Exemplarily, the applicability of the knowledge of each layer of the LLM to the paragraph is calculated according to the following formula:
[0090]
[0091] where H i is the applicability of the knowledge of each layer of the LLM to paragraph i; K is the scaling factor; n' is the total number of tokens in paragraph i; is the attention weight from the last token to the j-th token in paragraph i.
[0092] Step 106: Update the hierarchical relevance guidance weight according to the relevance score, confidence score, and the applicability of the knowledge of each layer of the LLM to the paragraph.
[0093] In this step, weights are selected to aggregate the hierarchical relevance and confidence scores to update the relevance guidance. To reduce numerical sensitivity, the logarithmic product is used in this embodiment. This method combines hierarchical relevance, confidence assessment, and entropy-based layer knowledge selection, and can make full use of diverse knowledge among different layers of the LLM, thus providing more robust guidance for the attention mechanism of LLMs.
[0094] Exemplarily, the updated relevance guidance weight of the q-th layer of the LLM for paragraph i is calculated according to the following formula
[0095]
[0096] where β is a hyperparameter used to balance the evaluation of the current layer and the previous layer of the LLM; H i is the applicability of the knowledge of each layer of the LLM to paragraph i; r i is the relevance between paragraph i and the question; c i is the confidence score of paragraph i with respect to the question; is the relevance guidance weight of the (q - 1)-th layer of the LLM for paragraph i.
[0097] Step 107: Construct a relevance-guided attention mask according to the updated hierarchical relevance guidance weight to achieve relevance-aware paragraph fusion.
[0098] In this step, a relevance-guided attention mask M q is constructed to control the allocation of attention. In specific implementation, the dimension of the mask matrix is [seq_len, seq_len], representing the attention relationship between all token pairs in the sequence. For the question token i' (i' ∈ q') accessing the paragraph token j' (j' ∈ p i) In this case, the system differentiates between two types of attention heads: middle-focused and other types. The middle-focused attention heads use relevance weights for weighting. These heads are typically located in the middle layers of the model (such as layers 8 - 16), and they have shown better ability to identify relevant content in previous analyses. Other types of attention heads (such as edge-focused or uniform types) keep the original weight of 1 unchanged, which ensures the LLM's perception of basic features such as text sequence positions. For the attention between tokens within a paragraph, the weight is kept at 1 to maintain semantic coherence. For tokens from different paragraphs, the weight is set to 0 to reduce interference. This hierarchical attention control mechanism enables the model to more effectively focus on relevant content while maintaining basic language understanding capabilities.
[0099] When i' belongs to the set of query token positions and j' belongs to the set of token positions in paragraph i (p i ), and it is a middle-focus attention head, the mask value is This value represents the relevance guidance coefficient of paragraph i in the q-th layer.
[0100] When i' belongs to the set of query token positions and j' belongs to the set of token positions in paragraph i (p i ), if it is an attention head of other types (non-middle focus), the mask value is 1, keeping the original attention weight unchanged.
[0101] When both i' and j' belong to the set of token positions in the same paragraph i (p i ), the mask value is 1, which allows full attention interaction between tokens within the same paragraph.
[0102] When i' belongs to the set of token positions in paragraph i (p i ) while j' belongs to the set of token positions in a different paragraph m (i ≠ m), the mask value is set to 0, which can prevent interference between different paragraphs.
[0103] Step 108, generate the final answer through the forward inference of the LLM based on the relevance-aware paragraph fusion result.
[0104] Based on the relevance-aware paragraph fusion result obtained in step 107, generate the final answer y through the forward inference of the LLM ansBefore this process, a standard language modeling loss function is used for joint fine-tuning of the model to ensure that the generated answers not only conform to the language expression norms but also accurately reflect the relevant information retrieved. In particular, when generating answers, the system focuses on ensuring that the answers are consistent with the relevant paragraphs retrieved, avoiding generating any false content that does not conform to the facts, and using a relevance guidance mechanism to ensure the accuracy of the answers. This generation strategy can significantly improve the reliability and accuracy of the answers.
[0105] In this embodiment, the relevance-guided attention mask is selectively applied to the head of the intermediate attention. At the same time, the functions of edge attention and uniform attention are retained. The LLM is jointly fine-tuned using the labeled language model loss: L LLM = -log(p(y ans |q′, 1, 2, …, n)); where p(·) represents the probability that the LLM outputs the answer y ans .
[0106] Traditional passage evaluators have a structure similar to BERT and usually have a low accuracy rate and fail to fully utilize the rich hierarchical-specific knowledge in LLMs. The hierarchical passage evaluator of this embodiment uses the knowledge of each layer of the LLM to evaluate the relevance score and confidence from multiple perspectives. In addition, in the hierarchical passage evaluator of this embodiment, an entropy-based layer-knowledge selector is introduced to determine the applicability of the knowledge of each layer to the passage by analyzing the attention distribution. By combining these comprehensive evaluations, the method of this embodiment provides reliable guidance for the attention mechanism of the LLM.
[0107] To optimize the hierarchical passage evaluator, this embodiment introduces three specialized loss functions, and each loss function corresponds to a key aspect of effective relevance evaluation:
[0108] Relevance Loss: To ensure that the LLM can accurately identify relevant passages, the relevance loss is used. The relevance loss function makes the evaluated relevance score closer to the true value, thereby improving the ability of the LLM to distinguish relevant and irrelevant passages.
[0109] Confidence Loss: Given that not all relevance predictions are equally reliable, the confidence loss is introduced to accurately reflect the credibility of the relevance predictions. This embodiment assumes that the LLM should show high confidence in samples that are easy to classify, while maintaining a lower confidence in samples that are more difficult and confusing. To this end, an external model (such as BGE) is used to assist in determining the difficulty of the samples.
[0110] Diversity Loss: To ensure the full utilization of layer-specific knowledge and avoid overly homogeneous relevance guidance, this embodiment introduces an entropy-based diversity loss.
[0111] Exemplarily, the retrieval-augmented generation method for optimizing the attention of a language large model based on layer knowledge provided in this embodiment further includes:
[0112] Calculate the relevance loss L according to the following formula re :
[0113]
[0114] where n is the total number of paragraphs retrieved from the knowledge base according to the user's question; is the true relevance label of paragraph i; r i is the relevance score of paragraph i estimated by the LLM;
[0115] Calculate the confidence loss L according to the following formula co :
[0116]
[0117] where c i is the confidence score of paragraph i estimated by the LLM; K is the total number of external models; M k represents k different external models; I(·) is the exponential function; q' represents the user's question;
[0118] Calculate the diversity loss L according to the following formula di :
[0119]
[0120] where H(·) is the entropy function.
[0121] Take the sum of the relevance loss, confidence loss, and diversity loss as the total loss of the hierarchical paragraph evaluator to evaluate the relevance score and confidence score of each paragraph to the question using the embedded representations transformed by each layer of the LLM.
[0122] By adding these loss functions, the hierarchical paragraph evaluator can learn to provide accurate, confident, and diverse relevance evaluations at different levels of the LLM. The relevance loss helps quickly narrow down to the most relevant paragraph range, while the diversity loss encourages a broader exploration of potential relevant information, increasing the probability of finding the correct answer. Although these two losses seem to be in opposition, their balanced combination makes the relevance evaluation more robust and comprehensive.
[0123] To evaluate the performance of the model under different data characteristics, this embodiment selects a series of representative RALM datasets, covering different aspects of question-answering tasks from factual questions to multi-hop reasoning and strategic questions. The following is a brief description of each dataset:
[0124] Natural Questions (NQ): Developed by Google Research, it contains real queries submitted by users to Google Search and high-quality manually annotated answers extracted from Wikipedia pages. This dataset includes both long and short answer formats, reflecting the real information-seeking behavior of users, covering a wide range of topics, and providing rich paragraphs and concise answer content. It is an ideal dataset for testing RALM.
[0125] TriviaQA: It contains a large number of question-answer pairs. The questions come from Trivia enthusiasts and are accompanied by supporting evidence obtained from Wikipedia and web searches. The main feature of this dataset is the high degree of lexical and syntactic variation between questions and answers, which requires the model to have strong retrieval and reasoning capabilities. By covering the web and Wikipedia domains, TriviaQA provides a comprehensive evaluation environment.
[0126] StrategyQA: Focuses on multi-hop reasoning questions that require implicit strategic thinking. The questions in StrategyQA are mostly common sense reasoning, and the answers are usually binary (yes / no), but they require complex cognitive processes. This dataset is designed to evaluate the strategic thinking ability of the model.
[0127] HotpotQA: Based on Wikipedia, it provides question-answer pairs that require reasoning between multiple supporting documents. This dataset contains sentence-level supporting facts for answer explanation and maintains a balance between different reasoning types (such as bridging and comparison). This structure makes HotpotQA particularly effective in evaluating multi-hop reasoning ability.
[0128] PopQA: Mainly focuses on questions related to popular culture, including movies, music, celebrities, and current events. This dataset is crucial for testing the model's ability to handle contemporary and rapidly evolving information, requiring the model to be able to handle ambiguity and context-dependence, reflecting the dynamic nature of real-world knowledge.
[0129] 2WikiMQA: A multi-hop open-domain question-answering dataset built on Wikipedia. This dataset contains questions that require reasoning across multiple Wikipedia pages and includes complex queries that cannot be answered by a single fact alone. This dataset is designed to test both the retrieval accuracy and advanced reasoning ability of RALM.
[0130] In this embodiment, the baseline models are divided into three categories: closed LLM without retrieval, LLM with retrieval, and robust RALM. The first two categories include LLAMA-3.1, Qwen-2.5, ChatGPT, GPT-4, and Claude-3-Sonnet. The third category includes REPLUGE, Self-RAG, RA-ISF, Robust-RALM, ChatQA-1.5, and RankRAG. The following is a brief description of each model:
[0131] LLAMA-3.1(2024): The latest version of the LLAMA series, trained on 1.5 trillion text data, becoming one of the best-performing open-source models currently, demonstrating excellent natural language processing capabilities.
[0132] Qwen-2.5(2024): Qwen-2.5 has made significant progress in multilingual processing, trained on 1.8 trillion data, reaching the state-of-the-art level in multiple tasks.
[0133] ChatGPT(2022): Developed by OpenAI, famous for its powerful dialogue capabilities and extensive knowledge base, suitable for question-and-answer tasks in various fields.
[0134] GPT-4(2023): A large-scale multimodal model developed by OpenAI, capable of processing image and text inputs and generating text outputs, showing near-human-level performance on multiple professional and academic benchmarks.
[0135] Claude-3-Sonnet(2024): One of the Claude 3 series models launched by Anthropic, known for its powerful performance in various tasks.
[0136] REPLUG(2023): A retrieval-augmented language modeling framework that treats the language model as a black box and pre-adds retrieved documents to the input to enhance its performance. The retrieval model is adjustable and applicable to frozen language models.
[0137] Self-RAG(2023): Improves the quality and authenticity of language models through retrieval and self-reflection. The framework uses a single language model that adaptively retrieves paragraphs as needed and generates and reflects on retrieved and generated content through special tokens.
[0138] RA-ISF(2024): A framework that iteratively decomposes tasks through three sub-modules, enhancing the model's factual reasoning ability and reducing hallucination generation.
[0139] Robust-RALM(2024): Focuses on enhancing the robustness of retrieval-augmented language models in irrelevant contexts. It proposes two methods: one is to filter out irrelevant paragraphs through a natural language inference (NLI) model, and the other is to automatically generate data for fine-tuning the language model.
[0140] ChatQA-1.5(2024): An upgraded version of the ChatQA model, introducing improvements that particularly enhance the performance of question-answering tasks in a conversational environment.
[0141] RankRAG(2024): An instruction-tuning framework that uses a single large language model (LLM) for the dual tasks of context ranking and answer generation, optimizing performance in retrieval-augmented tasks.
[0142] As Figure 2 shown, the performance differences of different model types on various datasets are recorded. Bold numbers represent the best scores among all models, and underlined numbers represent the best scores in each category. Although closed-form LLMs (without retrieval) are powerful, they have limited performance on datasets that require up-to-date or specialized knowledge. Retrieval-augmented LLMs generally outperform closed-form models, but some models (such as LLAMA-3.1-8B) experience a performance drop on the 2WikiMQA dataset, which may be due to positional bias and insufficient noise resistance. In contrast, robust retrieval-augmented language models (such as RankRAG) show significant improvements in leveraging retrieved paragraphs, demonstrating their effectiveness in improving model accuracy.
[0143] The layer knowledge-guided attention mechanism for RALM (LKG-RALM) performs excellently on various datasets, outperforming almost all baseline models. Compared with closed-form LLMs, LKG-RALM improves the accuracy by 12.3 percentage points on the Natural Questions (NQ) dataset and 10 percentage points on the 2WikiMQA dataset. Compared with retrieval-augmented LLMs, LKG-RALM also shows significant performance improvements, such as a 6.4 percentage point increase on the NQ dataset and an 8.8 percentage point increase on the 2WikiMQA dataset. In particular, LKG-RALM based on Qwen outperforms other robust RALM methods (such as RankRAG) on all datasets, with an improvement range of 2.8 to 7.3 percentage points.
[0144] This continuous performance improvement indicates that LKG-RALM can effectively alleviate the problems of attention dispersion and positional bias through the layer knowledge-guided attention mechanism. LKG-RALM performs particularly well on complex reasoning tasks and datasets that require up-to-date knowledge, demonstrating its ability to fully utilize external information and perform multi-hop reasoning. In addition, the scalability of LKG-RALM is also reflected in the performance gap between LLAMA-3.1-8B and LLAMA-3.1-70B, where the 70B model outperforms the 8B model on all datasets, with an improvement ranging from 2.3 to 5.7 percentage points. This shows that larger models can further enhance the advantages of LKG-RALM when dealing with more challenging tasks (such as PopQA).
[0145] In this embodiment, ablation experiments are further conducted on each component. As Figure 3 shown, the experimental results of LKG-RALM based on the LLAMA-3.1-8B base model under different ablation settings are presented. "LPR", "ELS", and "RPF" represent hierarchical paragraph relevance, entropy-based hierarchical knowledge selection, and relevance-aware channel fusion, respectively. Generally speaking, all the proposed components contribute significantly to the final performance of the model. The following is an analysis one by one:
[0146] Hierarchical Paragraph Evaluator (LPR): Replacing this component with a re-ranking model led to a significant decrease in performance on all tasks, with the average EM score decreasing by 3.96. This indicates that the hierarchical paragraph evaluator plays a key role in evaluating the relevance of retrieved paragraphs using the hierarchical knowledge of the LLM.
[0147] Entropy-based Layer Knowledge Selection (ELS): Removing this mechanism also resulted in a 1.52 decrease in the average EM score, highlighting the importance of dynamically selecting the most informative layer representations for each paragraph.
[0148] Relevance-guided Paragraph Fusion (RPF): Removing this component had an obvious negative impact on performance, with the average EM score decreasing by 22.22. This proves that the method of this embodiment can effectively alleviate the problems of attention dispersion and positional bias when dealing with multiple paragraphs compared with traditional attention mechanisms.
[0149] Auxiliary Loss: Adding auxiliary loss is also helpful for improving the performance of most tasks, with the average EM score increasing by 1.06. The auxiliary loss improves the performance of the model by guiding the model to explicitly focus on paragraph relevance, prediction confidence, and the diverse utilization of layer knowledge during the training process.
[0150] This embodiment also conducts experimental analysis on the robustness of the model, mainly considering two factors: the number of retrieved paragraphs and the proportion of irrelevant paragraphs.
[0151] To evaluate the scalability and efficiency of LKG-RALM in processing a large amount of retrieved information, relevant experiments were conducted, increasing the number of retrieved paragraphs from 0 to 50. As Figure 4 shown, it presents the performance trends of LKG-RALM and the baseline models on different datasets.
[0152] As Figure 4 can be seen, LKG-RALM demonstrated excellent scalability and high performance when dealing with a large number of retrieved paragraphs. As the number of retrieved paragraphs increased from 0 to 50, the EM score of LKG-RALM steadily rose from 24.8 to 61.5, with only a slight plateau effect when exceeding 35 paragraphs, indicating its ability to effectively utilize additional information without being affected by information overload. In contrast, the baseline model RankRAG reached a peak of 53.7 at 25 paragraphs and then slightly decreased to 54.2 at 50 paragraphs; Self-RAG reached a peak of 43.5 at 25 paragraphs but dropped sharply to 28.4 at 50 paragraphs. GPT-4 with retrieval capabilities showed stability, with its performance remaining almost unchanged (about 40.4) regardless of the number of paragraphs, indicating its strong internal knowledge but potential deficiencies in fully utilizing retrieved information. The EM score of LKG-RALM reached 61.0 at 50 paragraphs, 6.8 higher than RankRAG, demonstrating the effectiveness of its relevance-guided paragraph fusion mechanism in coping with a large amount of information interference.
[0153] LKG-RALM internally adopts a relevance-guided mechanism, which can focus on paragraphs relevant to the question in a noisy context, thereby improving the accuracy of evidence retrieval. To evaluate its robustness and noise resistance, this embodiment conducted an adversarial test, gradually replacing 50 retrieved paragraphs with irrelevant paragraphs, with the replacement ratio increasing from 0% to 100%.
[0154] As Figure 5As shown, LKG-RALM demonstrates remarkable resistance to noise and irrelevant information. As the proportion of irrelevant paragraphs increases to 100%, the EM score of LKG-RALM only gradually drops from 61.0 to 42.7, significantly outperforming other retrieval-enhanced models under the same conditions. Even when 80% of the paragraphs in the input are irrelevant, LKG-RALM still maintains a high EM score of 54.8. In contrast, models lacking explicit relevance modeling (such as RankRAG) experience a significant performance decline, from 54.2 to 15.6, in the case of complete irrelevance. GPT-4 with retrieval capabilities shows the highest robustness when faced with irrelevant paragraphs, with its performance only slightly decreasing from 40.4 to 31.5, but it is less effective than LKG-RALM in leveraging additional relevant information. The remarkable performance of LKG-RALM can be attributed to its explicit relevance modeling, enabling it to focus on important paragraphs while effectively ignoring the interference of extreme noise. By effectively utilizing the well-trained parametric knowledge of the model and external information, LKG-RALM achieves an optimal balance between leveraging retrieval knowledge and relying on the model's inherent capabilities.
[0155] To evaluate the efficiency of LKG-RALM, this embodiment compares its performance with existing models, and the results are as Figure 6 shown. Self-RAG and RA-ISF respectively require multiple rounds of retrieval and problem decomposition based on multi-round conversations. The lower parallelism results in their inference time being approximately 3.7 times that of LLAMA-3.1-8B. Among them, Self-RAG takes 3.07 seconds per query, and RA-ISF takes 3.44 seconds per query. Robust-RALM-8B and RankRAG-70B respectively use a BERT-based natural language inference (NLI) model and a re-ranking model to assist in paragraph filtering or ranking. This brings about an approximate 0.8% inference delay and 1.9% computational overhead for RankRAG-70B, increasing the processing time per query to 1.25 seconds and the computational cost per 1024 tokens to 145.4 TFLOPs. Similarly, LKG-RALM-70B uses the Qwen-2.5-7B model for paragraph relevance analysis, with the inference delay increasing by 2.4% (from 1.24 seconds to 1.27 seconds / query) and the computational cost increasing by 10.2% (from 142.6 TFLOPs to 157.2 TFLOPs). This low latency benefits from the high parallelism of the hierarchical paragraph evaluator during the LLM inference process.
[0156] To further explore the trade-off between efficiency and accuracy, experiments on the inter-layer paragraph evaluator were conducted using LLMs of different sizes. As Figure 6As shown, increasing the size of the evaluator from Qwen-2.5-500M to Qwen-2.5-7B led to a steady improvement in the EM scores for the benchmark models of LLAMA-3.1-8B and LLAMA-3.1-70B. For example, the EM score of LLAMA-3.1-8B increased from 53.6 to 55.3 as the evaluator size increased, while the processing time per query only slightly increased from 0.83 seconds to 0.91 seconds, and the computational cost increased from 18.1 TFLOPs to 30.9 TFLOPs. Notably, even with the largest Qwen-2.5-7B evaluator, LKG-RALM-70B achieved a higher EM score (61.0 vs. 54.2) while maintaining comparable competitive efficiency with RankRAG-70B (1.27 seconds / query vs. 1.25 seconds / query).
[0157] The flexible framework of LKG-RALM allows users to select an appropriate evaluator size according to their needs to balance accuracy and efficiency. For example, using Qwen-2.5-1.5B with LLAMA-3.1-70B increased the EM score by 0.7 compared to the 500M version, but the increase in query time and computational cost was very small.
[0158] This embodiment further analyzes the impact of the training data size on the model performance. Specifically, 30k, 60k, 90k, 160k, and 220k instances were randomly sampled from the original 440k training instances, and five variants of LKG-RALM-70B were fine-tuned using these data subsets. Then, the performance of these models on NQ and HotpotQA was compared with that of LKG-RALM fine-tuned using the full 440k dataset. At the same time, LLAMA-3.1-70B was also fine-tuned on the same data subsets as a baseline. As Figure 7 shown, the performance of the models trained with different data sizes is presented.
[0159] On the NQ dataset, the accuracy of LKG-RALM-70B gradually increased as the training data increased from 30k to 440k, rising from 50.5 to 61.0. In contrast, the improvement in LLAMA-3.1-FineTune was very small, increasing from 44.6 to 45.7, only an increase of 1.1. The performance gap between LKG-RALM-70B and LLAMA-3.1-FineTune widened significantly, from a 5.9-point gap on the 30k dataset to a 15.3-point gap on the 440k dataset.
[0160] Most of the capabilities of the model come from the good parameter space obtained during pre-training. Even with only 30k of fine-tuning data, LKG-RALM-70B shows strong performance, achieving an EM score of 50.5 on the NQ dataset and 37.1 on HotpotQA. This indicates that even a small amount of fine-tuning data is sufficient to stimulate the potential of LKG-RALM in passage evaluation and robustness. Even in the case of limited data, the performance gap between LKG-RALM-70B and LLAMA-3.1-FineTune is still significant, which demonstrates the effectiveness of the method in this embodiment. As the amount of training data increases, the accuracy of LKG-RALM-70B shows a continuous upward trend, but the growth rate gradually slows down after 220k instances. On NQ, the increase in data from 220k to 440k only brought an improvement of 0.7, while the increase in data from 160k to 220k improved by 1.8 points. This shows that although more data usually brings performance improvement, when the data volume reaches a certain level, the marginal benefit gradually decreases. The architecture and pre-training of the LLM play a key role in the effectiveness of the model, enabling it to achieve excellent results even with limited fine-tuning data, while the impact of additional training instances becomes less obvious after reaching a certain point.
[0161] In summary, this embodiment provides a retrieval-enhanced generation method for optimizing the attention of a large language model based on layer knowledge, which utilizes the diverse knowledge of each layer of the LLM to accurately predict relevance and confidence. Since the knowledge of not all layers contributes equally to relevance evaluation, an entropy-based layer knowledge selection method is used to dynamically determine which layer of knowledge is applicable to the passage. To reduce the interference and positional bias of irrelevant passages, a relevance-guided attention mask is adopted, enabling the question tokens to selectively focus on the retrieved passages to achieve an intermediate-focused attention pattern.
[0162] Embodiment 2
[0163] Based on the same inventive concept as Embodiment 1, this embodiment provides a retrieval-enhanced generation system for optimizing the attention of a large language model based on layer knowledge. Since the principle of the system for solving problems is similar to the aforementioned retrieval-enhanced generation method for optimizing the attention of a large language model based on layer knowledge, the implementation of the system can refer to the implementation of the retrieval-enhanced generation method for optimizing the attention of a large language model based on layer knowledge.
[0164] As Figure 8 shown, the retrieval-enhanced generation system for optimizing the attention of a large language model based on layer knowledge includes:
[0165] An acquisition module 10 for acquiring user questions.
[0166] A retrieval module 20 for retrieving multiple passages from the knowledge base according to the user questions.
[0167] A tagging addition module 30 for adding boundary tags at the beginning and end of each paragraph and tagging the question part before and after the user question text for differentiation.
[0168] A first evaluation module 40 for evaluating the relevance score and confidence score of each paragraph to the question using the embedding representations of each layer of the LLM.
[0169] A first calculation module 50 for calculating the applicability of the knowledge of each layer of the LLM to the paragraph using an entropy-based layer knowledge selection module.
[0170] An update module 60 for updating the hierarchical relevance guidance weights according to the relevance score, confidence score, and the applicability of the knowledge of each layer of the LLM to the paragraph.
[0171] A construction module 70 for constructing a relevance-guided attention mask according to the updated hierarchical relevance guidance weights to achieve relevance-aware paragraph fusion.
[0172] An answer generation module 80 for generating a final answer through LLM forward inference according to the relevance-aware paragraph fusion result.
[0173] Exemplarily, the first calculation module includes:
[0174] A first calculation unit for calculating the applicability of the knowledge of each layer of the LLM to the paragraph according to the following formula:
[0175]
[0176] where H i is the applicability of the knowledge of each layer of the LLM to paragraph i; K is a scaling factor; n' is the total number of tokens in paragraph i; is the attention weight from the last token to the j-th token in paragraph i.
[0177] Exemplarily, the update module includes:
[0178] A second calculation unit for calculating the updated relevance guidance weight of the q-th layer in the LLM for paragraph i according to the following formula
[0179]
[0180] where β is a hyperparameter for balancing the evaluation of the current layer and the previous layer of the LLM; H i is the applicability of the knowledge of each layer of the LLM to paragraph i; r i is the relevance of paragraph i to the question; c i is the confidence score of paragraph i to the question; It is the relevance guidance weight of the (q - 1)-th layer in the LLM for paragraph i.
[0181] Exemplarily, the retrieval-enhanced generation system for optimizing the attention of a language large model based on layer knowledge further includes:
[0182] A second calculation module, configured to calculate the relevance loss L according to the following formula re :
[0183]
[0184] where n is the total number of paragraphs retrieved from the knowledge base according to the user's question; is the true relevance label of paragraph i; r i is the relevance score of paragraph i estimated by the LLM.
[0185] A third calculation module, configured to calculate the confidence loss L according to the following formula co :
[0186]
[0187] where c i is the confidence score of paragraph i estimated by the LLM; K is the total number of external models; M k represents k different external models; I(·) is the exponential function; q' represents the user's question;
[0188] A fourth calculation module, configured to calculate the diversity loss L according to the following formula di :
[0189]
[0190] where H(·) is the entropy function.
[0191] A second evaluation module, configured to use the sum of the relevance loss, the confidence loss, and the diversity loss as the total loss of the hierarchical paragraph evaluator to evaluate the relevance score and the confidence score of each paragraph to the question by using the transformed embedding representations of each layer of the LLM.
[0192] For the more specific working processes of the above various modules, reference can be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated here.
[0193] Embodiment 3
[0194] This embodiment provides a computer device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the retrieval-enhanced generation method for optimizing the attention of a language large model according to Embodiment 1 are implemented.
[0195] For a more specific process of the above method, reference may be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated here.
[0196] Embodiment 4
[0197] This embodiment provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the retrieval-enhanced generation method for optimizing the attention of a language large model based on layer knowledge described in Embodiment 1 are implemented.
[0198] For a more specific process of the above method, reference may be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated here.
[0199] Embodiment 5
[0200] This embodiment provides a computer program product, including computer-executable instructions or a computer program, when the computer-executable instructions or the computer program are executed by a processor, the steps of the retrieval-enhanced generation method for optimizing the attention of a language large model based on layer knowledge described in Embodiment 1 are implemented.
[0201] For a more specific process of the above method, reference may be made to the corresponding content disclosed in Embodiment 1, which will not be elaborated here.
[0202] In this specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the systems, devices, storage media, and computer program products disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and the relevant parts can be referred to the descriptions in the method part.
[0203] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0204] In some embodiments, the computer-executable instructions can be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0205] By way of example, the computer-executable instructions may or may not correspond to a file in a file system, and may be stored as part of a file that holds other programs or data, for example, in one or more scripts stored in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or, stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).
[0206] By way of example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one site, or, on multiple electronic devices distributed at multiple sites and interconnected via a communication network.
[0207] The present invention has been described in detail above in connection with specific embodiments and exemplary examples, but these descriptions should not be construed as limiting the present invention. Those skilled in the art understand that, without departing from the spirit and scope of the present invention, various equivalent substitutions, modifications, or improvements can be made to the technical solutions of the present invention and their implementation manners, and all of these fall within the scope of the present invention. The protection scope of the present invention is subject to the appended claims.
Claims
1. A retrieval enhancement generation method based on layer knowledge optimization of language large model attention, characterized in that: include: Get user questions; Retrieve multiple paragraphs from the knowledge base based on user questions; Add boundary markers at the beginning and end of each paragraph and mark the question text before and after the user's question to distinguish the question part; The embedding representation of each layer of LLM is used to evaluate the relevance score and confidence score of each paragraph to the question; The entropy-based layer knowledge selection module is used to calculate the applicability of the knowledge of each layer of LLM to the paragraph; Update the hierarchical relevance guidance weights based on the relevance score, confidence score, and the applicability of the knowledge of each layer of LLM to the paragraph; Constructing relevance-guided attention masks based on the updated hierarchical relevance guidance weights to achieve relevance-aware paragraph fusion; According to the relevance-aware paragraph fusion results, the final answer is generated through LLM forward reasoning.
2. The search enhancement generation method according to claim 1, characterized in that: The method of calculating the applicability of knowledge of each layer of LLM to the paragraph by using the entropy-based layer knowledge selection module includes: The applicability of each layer of LLM knowledge to the paragraph is calculated according to the following formula: Among them, H i is the applicability of the knowledge of each layer of LLM to paragraph i; K is the scaling factor; n' is the total number of tokens in paragraph i; is the attention weight from the last token to the jth token in paragraph i.
3. The search enhancement generation method according to claim 1, characterized in that: The updating of the hierarchical relevance guidance weight according to the relevance score, the confidence score and the applicability of the knowledge of each layer of the LLM to the paragraph includes: The updated relevance guidance weight of paragraph i at the qth level in LLM is calculated according to the following formula Among them, β is a hyperparameter used to balance the evaluation of the current layer and the previous layer of LLM; H i is the applicability of the knowledge of each layer of LLM to paragraph i; r i is the relevance of paragraph i to the question; c i is the confidence score of paragraph i and the question; is the relevance guidance weight of the q-1th layer in LLM to paragraph i.
4. The search enhancement generation method according to claim 3, characterized in that: Also includes: The correlation loss L is calculated according to the following formula re : Where n is the total number of paragraphs retrieved from the knowledge base based on the user's question; is the true relevance label of paragraph i; r i The relevance score of paragraph i estimated by LLM; The confidence loss L is calculated according to the following formula: co : Among them, c i is the confidence score of paragraph i estimated by LLM; K is the total number of external models; M k represents k different external models; I(·) is an exponential function; q' represents the user question; The diversity loss L is calculated according to the following formula di : Among them, H(·) is the entropy function; The sum of relevance loss, confidence loss and diversity loss is taken as the total loss of the hierarchical paragraph evaluator to evaluate the relevance score and confidence score of each paragraph to the question using the embedded representation converted by each layer of the LLM.
5. A retrieval enhancement generation system based on layer knowledge optimization of language large model attention, characterized in that: include: Acquisition module, used to obtain user questions; A retrieval module, used to retrieve multiple paragraphs from the knowledge base according to user questions; A marker adding module, used for adding boundary markers at the beginning and end of each paragraph and marking the question parts before and after the user's question text; The first evaluation module is used to evaluate the relevance score and confidence score of each paragraph to the question using the embedding representation of each layer of LLM; A first calculation module is used to calculate the applicability of knowledge of each layer of LLM to the paragraph using an entropy-based layer knowledge selection module; An updating module, which is used to update the hierarchical relevance guidance weights according to the relevance score, confidence score and the applicability of the knowledge of each layer of LLM to the paragraph; A building module for constructing relevance-guided attention masks based on the updated hierarchical relevance guidance weights to achieve relevance-aware paragraph fusion; The answer generation module is used to generate the final answer through LLM forward reasoning based on the relevance-aware paragraph fusion results.
6. The search enhancement generation system according to claim 5, characterized in that: The first calculation module includes: The first calculation unit is used to calculate the applicability of the knowledge of each layer of LLM to the paragraph according to the following formula: Among them, H i is the applicability of the knowledge of each layer of LLM to paragraph i; K is the scaling factor; n' is the total number of tokens in paragraph i; is the attention weight from the last token to the jth token in paragraph i.
7. The search enhancement generation system according to claim 5, characterized in that: The update module includes: The second calculation unit is used to calculate the updated relevance guidance weight of paragraph i in the qth layer of LLM according to the following formula Among them, β is a hyperparameter used to balance the evaluation of the current layer and the previous layer of LLM; H i is the applicability of the knowledge of each layer of LLM to paragraph i; r i is the relevance of paragraph i to the question; c i is the confidence score of paragraph i and the question; is the relevance guidance weight of the q-1th layer in LLM to paragraph i.
8. The search enhancement generation system according to claim 7, characterized in that: Also includes: The second calculation module is used to calculate the correlation loss L according to the following formula re : Where n is the total number of paragraphs retrieved from the knowledge base based on the user's question; is the true relevance label of paragraph i; r i The relevance score of paragraph i estimated by LLM; The third calculation module is used to calculate the confidence loss L according to the following formula co : Among them, c i is the confidence score of paragraph i estimated by LLM; K is the total number of external models; M k represents k different external models; I(·) is an exponential function; q' represents the user question; The fourth calculation module is used to calculate the diversity loss L according to the following formula di : Among them, H(·) is the entropy function; The second evaluation module is used to take the sum of relevance loss, confidence loss and diversity loss as the total loss of the hierarchical paragraph evaluator to evaluate the relevance score and confidence score of each paragraph to the question using the embedded representation converted by each layer of the LLM.
9. A computer device, characterized in that: It includes a processor and a memory; wherein, when the processor executes the computer program stored in the memory, it implements the steps of the retrieval enhancement generation method based on layer knowledge optimization of language large model attention as described in any one of claims 1-4.
10. A computer-readable storage medium, characterized in that: Used to store computer programs; when the computer programs are executed by the processor, the steps of the retrieval enhancement generation method based on layer knowledge optimization of language large model attention as described in any one of claims 1-4 are implemented.
Citation Information
Cited By
Scientific research reasoning task processing method, equipment and medium
CN121480697A