Text retrieval enhancement generation method and device, medium and equipment
By constructing a golden segmentation vector database and using adversarial samples to train language big models, the problem of retrieval noise in complex retrieval environments in the prior art is solved, and the robustness of the model and the accuracy of generating answers are improved.
Patent Information
- Application Number
- CN202510546788.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-28
AI Technical Summary
When existing search enhancement generation technologies deal with complex real-world search environments, it is difficult to effectively identify and process different types of search noise, resulting in inaccurate or incorrect answers from the model.
By constructing a golden segmentation vector database containing key information annotations, and combining adversarial samples of relevant search noise, irrelevant search noise and counterfactual search noise, a fusion training set is formed. Then, a noise classification layer and an attention fusion layer are added before the output layer of the language big model, and the model is trained to identify and distinguish different types of noise.
It improves the robustness of the model in a noisy environment, enhances the model's adaptability to different noises, helps the model distinguish relevant and unrelated information, and thus generates more accurate and reliable answers.
Smart Images

Figure CN120104718A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of retrieval enhancement generation, and in particular to a text retrieval enhancement generation method, device, medium and equipment. Background Art
[0002] Retrieval-Augmented Generation (RAG) is an effective solution to alleviate the limitations of large language models (LLMs) in specific domains or tasks by integrating external database knowledge. Although RAG performs well in enhancing model generation capabilities, inappropriate retrieval paragraphs may potentially hinder the ability of LLMs to generate comprehensive and high-quality answers. For example, irrelevant information or low-quality content retrieved may introduce noise, causing the model to generate inaccurate or even wrong answers. However, RAG research on retrieval noise robustness fails to fully reflect the complex retrieval environment of the real world, thus limiting its effectiveness in practical applications. Summary of the invention
[0003] Based on this, in order to solve the technical problems in the prior art, the present invention provides a text retrieval enhancement generation method, device, medium and equipment.
[0004] The present invention provides a text retrieval enhancement generation method, comprising: Perform vectorization operations on known corpus to generate an initial vector database; perform screening, annotation and vectorization operations on known corpus to generate a golden section vector database containing key information annotations; Obtain training problems contained in a data set used to train a large language model, extract adversarial samples including relevant retrieval noise, irrelevant retrieval noise, and counterfactual retrieval noise from an external database based on the relevance to the training problems; recall the first corpus related to the training problems from the golden section vector database; fuse the first corpus with the adversarial samples to obtain a fused training set; A classification layer for classifying noise types and an attention fusion layer connected to the output end of the classification layer are added before the output layer of the large language model to obtain an improved large language model; the improved large language model is trained using the fusion training set; Obtain user questions and recall the second corpus related to the user questions from the initial vector database; input the user questions and the second corpus into the trained language model, and output the noise category of the second corpus through the classification layer; use the noise category output by the classification layer as a generation condition, input it into the attention fusion layer together with the user questions and the second corpus, perform feature fusion through the attention mechanism, and generate a fused context representation; the output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate an answer.
[0005] Furthermore, the screening of the known corpus specifically includes: Cut the known corpus into chunks, and extract key feature vectors including keywords, topics, entities, and sentence structures from the chunks; Based on the key feature vector, at least one evaluation index score of the corpus information density, source authority and relevance to the search topic is calculated to obtain the value score of each corpus; wherein, the evaluation index score of the corpus information density is obtained by measuring the distribution and importance of keywords in the corpus using the TF-IDF and BM25 algorithms, the evaluation index score of the source authority is obtained based on the number of citations of the source and the weight of the publication channel, and the evaluation index score of the relevance to the search topic is obtained by using the cosine similarity to calculate the degree of relevance to the search topic; The corpus blocks with value scores higher than the set threshold are regarded as high-value corpus blocks.
[0006] Furthermore, the extracting of adversarial samples including relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise from an external database based on the relevance to the training problem specifically includes: For each training question, a pre-trained retrieval model is used to retrieve multiple candidate contexts from an external knowledge base; Filter the contexts with the highest semantic similarity to the current training question but missing the answer field from the candidate contexts as relevant retrieval noise samples; Randomly select a candidate context that is irrelevant to the current training problem from the retrieval results of other training problems as an irrelevant retrieval noise sample; A candidate context containing the correct answer is randomly selected from the retrieval results of the current training question, and the answer entity is replaced with incorrect information to generate a counterfactual retrieval noise sample.
[0007] Furthermore, the step of recalling the second corpus related to the user question from the initial vector database specifically includes: Use natural language processing technology to analyze the vocabulary complexity and sentence structure of user questions and obtain the complexity labels of user questions: {simple, medium, complex}; Select a search strategy that matches the complexity of the user's question: select a search strategy based on the complexity of the user's question: for simple questions, use a keyword matching search strategy; for medium questions, use a similarity calculation search strategy based on a vector space model; for complex questions, use a semantic search strategy based on the deep learning model BERT; The second corpus related to the user's question is recalled from the initial vector database according to the retrieval strategy.
[0008] Furthermore, a classification layer for classifying noise types and an attention fusion layer connected to the output end of the classification layer are added before the output layer of the large language model; Among them, the classification layer is a multi-layer feedforward neural network structure, including an input layer that receives the hidden state of the last transformer layer of the language model, an intermediate layer composed of multiple fully connected layers, and an output layer that outputs the probability distribution of the noise type through the softmax activation function; the attention fusion layer uses multi-head attention mechanism, weighted fusion mechanism and residual connection to perform attention fusion on the noise.
[0009] Furthermore, the use of the fused training set to train and improve the language model uses the sum of the noise classification error and the answer relevance error as the loss function, and optimizes the parameters of the Prompt language model through gradient inversion.
[0010] Furthermore, the output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate an answer, specifically including: Recalling a conditional weight set matching the noise category from a pre-trained weight database, wherein the conditional weight set includes decoding parameter configurations for different noise characteristics; During the decoding process, the similarity scores between the context representation of the current time step and all possible text blocks are calculated; the similarity scores are multiplied by the corresponding conditional weights to adjust the selection probability of the text block associated with the context to generate an initial answer; The initial answer is semantically corrected, logically consistent and noise filtered to generate the final answer.
[0011] The present invention provides a text retrieval enhancement generation device, comprising: The corpus construction module is used to perform vectorization operations on known corpora to generate an initial vector database; to screen, annotate and vectorize known corpora to generate a golden section vector database containing key information annotations; The data augmentation module is used to obtain the training problems contained in the data set used to train the large language model, extract adversarial samples including relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise from the external database according to the relevance to the training problems; recall the first corpus related to the training problems from the golden section vector database; fuse the first corpus with the adversarial samples to obtain a fused training set; A training module is used to add a classification layer for classifying noise types and an attention fusion layer connected to the output end of the classification layer before the output layer of the language model to obtain an improved language model; and train the improved language model using the fusion training set; The answer generation module is used to obtain user questions and recall the second corpus related to the user questions from the initial vector database; input the user questions and the second corpus into the trained language model, and output the noise category of the second corpus through the classification layer; use the noise category output by the classification layer as the generation condition, input it into the attention fusion layer together with the user questions and the second corpus, perform feature fusion through the attention mechanism, and generate a fused context representation; the output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate an answer.
[0012] The present invention provides a computer-readable storage medium, wherein the storage medium stores a computer program, and when the computer program is executed by a processor, the text retrieval enhancement generation method is implemented.
[0013] The present invention provides a computer device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned text retrieval enhancement generation method when executing the program.
[0014] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects: In the text retrieval enhancement generation method provided by the present invention, firstly, a golden section vector database containing key information annotations is constructed by screening, annotating and vectorizing known corpora, so as to ensure that the first corpus recalled by the training question has higher quality, so as to guide the model to pay more attention to important features during the training process and improve the model's understanding and answering capabilities; then, the golden section vector database is fused with adversarial samples of relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise to form a fused training set, and a large language model with added noise classification tasks is trained through the fused training set, so that the model can recognize and distinguish different types of noise, enhance the model's adaptability to different noises, and help the model learn how to distinguish relevant and irrelevant information, so as to be more robust when dealing with practical problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0016] Figure 1 A flowchart of a text retrieval enhancement generation method provided by the present invention; Figure 2 A schematic diagram of a text retrieval enhancement generation step framework provided by the present invention; Figure 3 A schematic diagram of a text retrieval enhancement generation device provided by the present invention. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in combination with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0018] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0019] RAG research on retrieval noise robustness is often limited to a limited set of noise types, failing to fully reflect the complex retrieval environment of the real world, thus limiting its effectiveness in practical applications. Retrieval noise in the real world may include relevant but redundant information, completely irrelevant content, and even counterfactual information that contradicts the facts. These different types of noise place higher demands on the robustness of LLMs. Based on this, the present invention addresses the deficiencies of the above-mentioned prior art and provides a retrieval enhancement generation method based on adaptive adversarial training and noise classification. By introducing adversarial training and noise classification mechanisms, the robustness of the model in a noisy environment is improved, and the accuracy and reliability of the generated results are optimized.
[0020] Figure 1 and Figure 2 The text retrieval enhancement generation method flow and step framework shown in the figure describe in detail the technical solutions provided by each embodiment of the present application. Specifically, the following steps are included:
[0021] S1: Perform vectorization operations on known corpus to generate an initial vector database; perform screening, annotation and vectorization operations on known corpus to generate a golden section vector database containing key information annotations. Specifically include:
[0022] S11: Use the period as a segmentation mark to perform segmentation operation on the known corpus.
[0023] S12: Select any existing embedding model that can embed text to perform embedding operation on the corpus block, obtain a vector representation of each corpus block and store it in any database such as PostgreSQL.
[0024] S13: Construct a golden section database by screening high-quality corpus blocks and marking their relevance to form a high-quality initial database.
[0025] In a specific embodiment, the step of selecting high-quality corpus blocks in step S13 includes: S131: extracting key features from the corpus based on the initial vector database, the key features including but not limited to keywords, topics, entities, and sentence structures, and converting the key features into feature vectors.
[0026] S132: Calculate the value score for each corpus according to the predefined scoring criteria, wherein the scoring criteria include at least one evaluation index of the relevance of the corpus to the search topic, information density, and source authority. The evaluation index score of the corpus information density is obtained by measuring the distribution and importance of keywords in the corpus using the TF-IDF and BM25 algorithms, the evaluation index score of the source authority is obtained based on the number of citations of the source and the weight of the publication channel, and the evaluation index score of the relevance of the search topic is obtained by calculating the relevance to the search topic using methods such as cosine similarity or semantic matching.
[0027] S133: Setting a score threshold, and screening corpora with value scores higher than the score threshold as high-value corpora.
[0028] S2: Obtain the training problems contained in the data set used to train the large language model, extract adversarial samples including relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise from the external database according to the relevance to the training problems; recall the first corpus related to the training problems from the golden section vector database; fuse the first corpus with the adversarial samples to obtain a fused training set. Specifically including:
[0029] S21: Initialize the retrieval model and language model, using the pre-trained retrieval model DPR and language model LLaMA-2 as the infrastructure.
[0030] S22: Retrieve candidate contexts from an external database. For each query, a retrieval model is used to retrieve multiple candidate contexts from an external knowledge base.
[0031] S23: Generate relevant retrieval noise samples, select the context most relevant to the query (but not containing the correct answer) from the retrieved candidate contexts, and use it as the relevant retrieval noise sample; generate irrelevant retrieval noise samples, randomly select a context irrelevant to the current query from the retrieval results of other queries, and use it as the irrelevant retrieval noise sample; generate counterfactual retrieval noise samples, randomly select one from the retrieved context containing the correct answer and replace its answer entity with incorrect information, and generate counterfactual retrieval noise samples.
[0032] S24: Construct an adversarial sample set by combining the generated relevant retrieval noise samples, irrelevant retrieval noise samples, and counterfactual retrieval noise samples with the golden retrieval context (the context containing the correct answer) to form an adversarial sample set.
[0033] The step of recalling the first corpus related to the training problem from the golden section vector database includes: Building an index: In the golden section vector database, each corpus will be converted into a vector representation and indexed for fast retrieval.
[0034] Query vector representation: Convert the training question into a vector representation, which is usually done through a language model, such as using BERT, GPT and other models to obtain the embedding vector of the question.
[0035] Similarity calculation: Calculate the similarity between the training question vector and all the corpus vectors in the database. This can be done using cosine similarity, Euclidean distance, or other distance metrics.
[0036] Sorting and selection: Sort the corpora according to the similarity scores and select the most relevant N corpora as the results of "recall".
[0037] S3: Add a classification layer for classifying noise types before the output layer of the large language model and an attention fusion layer connected to the output end of the classification layer to obtain an improved large language model; use the fusion training set to train the improved large language model. Specifically include:
[0038] S31: Auxiliary tasks are designed through a multi-task learning mechanism to enable the model to recognize and distinguish different types of noise. This large language model adds a classification layer for classifying noise types and an attention fusion layer connected to the output of the classification layer before the output layer of the Prompt language large model.
[0039] Among them, the classification layer is a multi-layer feedforward neural network structure, including an input layer that receives the hidden state of the last transformer layer of the language model, an intermediate layer consisting of multiple (2-3) fully connected layers, each layer is followed by BatchNorm and ReLU activation functions, and an output layer that outputs the probability distribution of the noise type through the softmax activation function. Specific parameter design: First layer: map the model hidden state (assuming the dimension is d_model) to a smaller dimension (such as d_model / 2), intermediate layer: the dimension remains unchanged or gradually decreases, output layer: the dimension is equal to the predefined number of noise types N. Regularization technology: Dropout (0.1-0.3) is applied to each fully connected layer, and weight regularization (L2 regularization) prevents overfitting.
[0040] The attention fusion layer uses multi-head attention mechanism, weighted fusion mechanism and residual connection to perform attention fusion on noise. Among them, the multi-head attention mechanism: uses the classification layer output as the query vector; uses the final layer hidden state of the language model as the key and value vectors; the number of heads can be set to 4-8, and the dimension of each head is d_model / number of heads. Weighted fusion mechanism: based on the classification results, different attention weights are assigned to different types of noise; a gating mechanism is designed to dynamically adjust the attention allocation according to the noise classification probability. Residual connection design: Add a residual connection between the original language model output and the output after attention processing; contains a learnable parameter α to control the weight ratio of the original output to the processed output. Output layer normalization: Apply LayerNorm after the attention output to ensure a stable output distribution.
[0041] S32: Adopt an adaptive adversarial training strategy to inject the generated adversarial samples into the training process.
[0042] S33: Optimize model parameters through gradient reversal to make it robust in noisy environments.
[0043] In a specific embodiment, step S31 includes: S311: Build a retrieval enhancement model and use the retrieval enhancement framework DPR to retrieve contextual information related to the question from the external knowledge base.
[0044] S312: Classify the retrieved contexts into three types according to their relevance to the question: contexts that are relevant but do not contain the correct answer (relevant retrieval noise), contexts that are irrelevant to the question (irrelevant retrieval noise), and contexts that are relevant to the question topic but contain incorrect information (counterfactual retrieval noise).
[0045] S313: Train a classification model, using the labeled questions and retrieval context datasets, to train a classification model that can identify different types of retrieval noise.
[0046] S314: Input a new question, input the new question into the retrieval enhancement model, retrieve relevant context and classify it.
[0047] S315: Determine the retrieval context category corresponding to the question based on the output of the classification model, and provide classification information for subsequent generation tasks.
[0048] In a specific embodiment, step S32 includes: S321: Based on the golden section database, the adversarial sample generation module is used to dynamically generate relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise.
[0049] S322: Mix the generated adversarial samples with the original training data, adopt an adaptive adversarial training strategy, and optimize the model through gradient reversal or adversarial loss function.
[0050] S323: Design a multi-task learning mechanism and introduce a noise classification task as an auxiliary task, so that the model can identify and distinguish relevant retrieval noise, irrelevant retrieval noise, and counterfactual retrieval noise.
[0051] S324: By jointly optimizing the loss functions of the main task and the auxiliary task, that is, using the sum of the noise classification error and the answer correlation error as the loss function to adjust the model parameters, the robustness and generation accuracy of the model in a noisy environment are improved.
[0052] S4: Get the user's question, and recall the second corpus related to the user's question from the initial vector database; input the user's question and the second corpus into the trained language model, and output the noise category of the second corpus through the classification layer; use the noise category output by the classification layer as the generation condition, input it into the attention fusion layer with the user's question and the second corpus, perform feature fusion through the attention mechanism, and generate the fused context representation; the output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate the answer. Specifically including:
[0053] S41: Classify the questions input by the user and select the retrieval strategy according to the question type and complexity.
[0054] S42: Execute the search strategy and obtain relevant corpus.
[0055] S43: Input the retrieved corpus and user questions into the large language model, and design prompt words to guide the model to generate high-quality answers.
[0056] In a specific embodiment, step S41 includes: S411: Question complexity analysis uses natural language processing technology to analyze input questions, including vocabulary complexity and sentence structure, to assess the complexity of the question. The pre-trained classification model is used to mark the question as simple, medium or complex, and identify keywords.
[0057] S412: Retrieval strategy selection, select a retrieval strategy based on the complexity of the question. For simple questions, keyword matching is used. For medium questions, a vector space model is used for similarity calculation. For complex questions, the deep learning model BERT is called for semantic retrieval. BERT is a pre-trained natural language processing model that can understand the deep meaning and contextual relationships in language. The steps of using BERT for semantic retrieval include: pre-processing the questions raised by the user, inputting the pre-processed input questions into the BERT model, outputting the contextualized embedding vector of each word in the input question, and calculating the similarity between the contextualized embedding vector and the document vector pre-stored in the database to achieve recall.
[0058] S413: Dynamically adjust and provide relevance feedback to the search results. If the keyword matching effect is not ideal, switch to vector space model search.
[0059] S414: According to the selected strategy, relevant corpus is retrieved from the initial vector database and the golden section database. By searching the two databases, it is ensured that high-quality corpus matching the complexity of the question is obtained to provide support for subsequent text generation.
[0060] In a specific embodiment, step S43 includes: S431: For each user question, recall a second corpus related to the user question from the initial vector database.
[0061] S42: Input the user question and the second corpus into the trained language model, and identify the noise category to which the second corpus belongs through the noise classification layer in the model.
[0062] S433: The identified noise category is used as a conditional input, fused with the user question and the second corpus through the attention mechanism, and passed to the output layer of the large language model.
[0063] S434: Dynamically adjust the weights or parameters of the output layer according to the noise category to guide the generation of answers that are suitable for the noise category. In the process of generating answers, the model will dynamically adjust the weights or parameters of the output layer according to the noise category. Specifically, different noise categories correspond to different conditional weights, and a set of conditional weights matching the noise category is called from a pre-established weight database. The conditional weight set contains decoding parameter configurations for different noise characteristics. These conditional weights are used to dynamically adjust the decoding process during decoding: the output layer applies conditional weights at each time step of decoding to adjust the selection probability of the text block in the answer to generate an answer that is suitable for the noise category.
[0064] S435: After generating the answer, the generated text is post-processed based on the noise category, including spelling correction, semantic modification and logical consistency check, to improve the accuracy and contextual consistency of the answer.
[0065] S436: Based on the prompt word, the model is guided to generate a high-quality answer based on the retrieved corpus. The prompt word can be any prompt word that enables the large language model to answer the user's question based on the retrieved corpus. The following is an example:
[0066] "Below I will give you a user-entered question and some text blocks related to the question as reference information. Please answer the user question based on the text blocks provided to you. Please answer the user question strictly based on the reference information provided to you.
[0067] Let’s start: User Question: <question> Reference Information: Text Block:<Text Chunk> ”.
[0068] The above is a text retrieval enhancement generation method provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding Figure 3 The text retrieval enhancement generation device shown includes: The corpus construction module is used to perform vectorization operations on known corpora to generate an initial vector database; to screen, annotate and vectorize known corpora to generate a golden section vector database containing key information annotations.
[0069] The data enhancement module is used to obtain the training problems contained in the data set used to train the large language model, extract adversarial samples including relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise from the external database according to the relevance to the training problems; recall the first corpus related to the training problems from the golden section vector database; and fuse the first corpus with the adversarial samples to obtain a fused training set.
[0070] The training module is used to add a classification layer for classifying noise types and an attention fusion layer connected to the output end of the classification layer before the output layer of the language model to obtain an improved language model; and use the fusion training set to train the improved language model.
[0071] The answer generation module is used to obtain user questions and recall the second corpus related to the user questions from the initial vector database; input the user questions and the second corpus into the trained language model, and output the noise category of the second corpus through the classification layer; use the noise category output by the classification layer as the generation condition, input it into the attention fusion layer together with the user questions and the second corpus, perform feature fusion through the attention mechanism, and generate a fused context representation; the output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate an answer.
[0072] The specific definition of the text retrieval enhancement generation device can be found in the definition of the text retrieval enhancement generation method above, which will not be repeated here. Each module in the above-mentioned text retrieval enhancement generation device can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0073] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 Provided text retrieval enhancement generation method.
[0074] The present invention also provides a computer device structure. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 Provided text retrieval enhancement generation method.
[0075] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0076] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.< / question>
Claims
1. A text retrieval enhancement generation method, characterized in that: include: Perform vectorization operations on known corpus to generate an initial vector database; perform screening, annotation and vectorization operations on known corpus to generate a golden section vector database containing key information annotations; Obtain training problems contained in a data set used to train a large language model, extract adversarial samples including relevant retrieval noise, irrelevant retrieval noise, and counterfactual retrieval noise from an external database based on the relevance to the training problems; recall the first corpus related to the training problems from the golden section vector database; fuse the first corpus with the adversarial samples to obtain a fused training set; A classification layer for classifying noise types and an attention fusion layer connected to the output end of the classification layer are added before the output layer of the language model to obtain an improved language model. Improve the large language model using fusion training set training; Obtain user questions and recall the second corpus related to the user questions from the initial vector database; Input the user question and the second corpus into the trained language model, and output the noise category of the second corpus through the classification layer; The noise category output by the classification layer is used as the generation condition and input into the attention fusion layer together with the user question and the second corpus. The feature fusion is performed through the attention mechanism to generate a fused context representation. The output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate an answer.
2. The text retrieval enhancement generation enhancement method according to claim 1, characterized in that: The screening of the known corpus specifically includes: Cut the known corpus into corpus chunks, and extract key feature vectors including keywords, topics, entities, and sentence structures from the corpus chunks; Based on the key feature vector, at least one evaluation index score of the corpus information density, source authority and relevance to the search topic is calculated to obtain the value score of each corpus; wherein, the evaluation index score of the corpus information density is obtained by measuring the distribution and importance of keywords in the corpus using the TF-IDF and BM25 algorithms, the evaluation index score of the source authority is obtained based on the number of citations of the source and the weight of the publication channel, and the evaluation index score of the relevance to the search topic is obtained by using the cosine similarity to calculate the degree of relevance to the search topic; The corpus blocks with value scores higher than the set threshold are regarded as high-value corpus blocks.
3. The text retrieval enhancement generation enhancement method according to claim 1, characterized in that: The adversarial samples including relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise are extracted from the external database according to the relevance to the training problem, specifically including: For each training question, a pre-trained retrieval model is used to retrieve multiple candidate contexts from an external knowledge base; Filter the contexts with the highest semantic similarity to the current training question but missing the answer field from the candidate contexts as relevant retrieval noise samples; Randomly select a candidate context that is irrelevant to the current training problem from the retrieval results of other training problems as an irrelevant retrieval noise sample; A candidate context containing the correct answer is randomly selected from the retrieval results of the current training question, and the answer entity is replaced with incorrect information to generate a counterfactual retrieval noise sample.
4. The text retrieval enhancement generation enhancement method according to claim 1, characterized in that: The step of recalling the second corpus related to the user's question from the initial vector database specifically includes: Perform vocabulary complexity analysis and sentence structure analysis on user questions to obtain complexity labels of user questions: {simple, medium, complex}; Select a search strategy that matches the complexity of the user's question: select a search strategy based on the complexity of the user's question: for simple questions, use a keyword matching search strategy; for medium questions, use a similarity calculation search strategy based on a vector space model; for complex questions, use a semantic search strategy based on the deep learning model BERT; The second corpus related to the user's question is recalled from the initial vector database according to the retrieval strategy.
5. The text retrieval enhancement generation enhancement method according to claim 1, characterized in that: The method further comprises adding a classification layer for classifying noise types and an attention fusion layer connected to the output end of the classification layer before the output layer of the large language model; Among them, the classification layer is a multi-layer feedforward neural network structure, including an input layer that receives the hidden state of the last transformer layer of the language model, an intermediate layer composed of multiple fully connected layers, and an output layer that outputs the probability distribution of the noise type through the softmax activation function; the attention fusion layer uses multi-head attention mechanism, weighted fusion mechanism and residual connection to perform attention fusion on the noise.
6. The text retrieval enhancement generation enhancement method according to claim 5, characterized in that: The use of the fusion training set to train and improve the language model uses the sum of the noise classification error and the answer correlation error as the loss function, and optimizes the parameters of the Prompt language model through gradient reversal.
7. The text retrieval enhancement generation enhancement method according to claim 1, characterized in that: The output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate an answer, specifically including: Recalling a conditional weight set matching the noise category from a pre-trained weight database, wherein the conditional weight set includes decoding parameter configurations for different noise characteristics; During the decoding process, the similarity scores between the context representation of the current time step and all possible text blocks are calculated; the similarity scores are multiplied by the corresponding conditional weights to adjust the selection probability of the text block associated with the context to generate an initial answer; The initial answer is semantically corrected, logically consistent and noise filtered to generate the final answer.
8. A text retrieval enhancement generation device, characterized in that: include: The corpus construction module is used to perform vectorization operations on known corpora and generate an initial vector database; Screen, annotate and vectorize known corpus to generate a golden section vector database containing key information annotations; The data augmentation module is used to obtain the training problems contained in the data set used to train the large language model, extract adversarial samples including relevant retrieval noise, irrelevant retrieval noise and counterfactual retrieval noise from the external database according to the relevance to the training problems; recall the first corpus related to the training problems from the golden section vector database; fuse the first corpus with the adversarial samples to obtain a fused training set; A training module is used to add a classification layer for classifying noise types and an attention fusion layer connected to the output end of the classification layer before the output layer of the language model to obtain an improved language model; Improve the large language model using fusion training set training; The answer generation module is used to obtain the user's question and recall the second corpus related to the user's question from the initial vector database; input the user's question and the second corpus into the trained language model, and output the noise category of the second corpus through the classification layer; The noise category output by the classification layer is used as the generation condition and input into the attention fusion layer together with the user question and the second corpus. The feature fusion is performed through the attention mechanism to generate a fused context representation. The output layer dynamically adjusts the weight distribution or decoding strategy according to the noise category, and decodes the fused context through the dynamically adjusted weight distribution or decoding strategy to generate an answer.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Relationship extraction system, method and device
CN112131879A
Financial question and answer text generation method and device based on retrieval enhancement and storage medium
CN116662502A
Large model retrieval enhancement generation method and device
CN118467688A
Data processing method, system and device, computer program product and storage medium
CN118643196A
End-to-end large model fine tuning method and system based on retrieval enhancement generation
CN119719268A
Cited By
Enhanced retrieval-based agent rapid construction method and system
CN120407751A
Network response retrieval enhancement generation method fused with large language model, electronic equipment and computer readable storage medium
CN120705265A
Multi-modal large model identification method and device for leakage noise of water supply network
CN120766690A
Approximate patent retrieval method and device, electronic equipment and storage medium
CN120804294A
Retrieval model training method, reply obtaining method and device
CN120849541A