Multi-field adaptive RAG optimization method and system
Through the multi-domain adaptive RAG optimization method, the adaptability and answer quality of the RAG system in specific fields are improved, the problems of data acquisition difficulties and separate training of retrievers and generators are solved, and the adaptive optimization and robustness improvement of the model are achieved.
Patent Information
- Application Number
- CN202510689254.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-12
AI Technical Summary
The existing RAG system lacks adaptability, answer quality, model generalization and robustness in specific fields. In addition, data acquisition is difficult, labeling costs are high, and the separate training of the retriever and generator leads to performance bottlenecks.
Through a multi-domain adaptive RAG optimization method, including data collection, data screening, preference data generation and joint fine-tuning, and utilizing a chain thinking prompt mechanism and contrastive training loss function, the collaborative working ability of large language models and retrievers is improved.
It achieves adaptive optimization of the RAG system in specific fields, reduces data acquisition and annotation costs, improves answer quality and the generalization and robustness of the model in complex scenarios, and solves the performance bottleneck caused by separate training of the retriever and generator.
Smart Images

Figure CN120633754A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of deep learning technology, and more specifically to a multi-domain adaptive RAG optimization method and system. Background Art
[0002] With the development of large language model technology, RAG systems have been applied in many fields. However, existing RAG systems have many problems when dealing with tasks in specific fields, such as chemistry, biology, medicine, and finance. On the one hand, the insufficient domain knowledge reserves make it difficult to accurately understand and handle complex problems in specific fields, resulting in poor professionalism and accuracy in the generated content. On the other hand, data in specific fields is difficult to obtain and the cost of data annotation is high, which seriously limits the training effect and performance improvement of the model in specific fields. In addition, there is insufficient interaction between the retriever and the generator. The separate training of the two makes it impossible for the retrieval information to effectively support the generation process, and may even be misleading, resulting in performance bottlenecks. These problems hinder the in-depth application of RAG systems in specific fields.
[0003] How to improve the adaptability, answer quality, generalization and robustness of the RAG system in specific fields is a technical problem that needs to be solved. Summary of the Invention
[0004] The technical task of the present invention is to address the above shortcomings and provide a multi-domain adaptive RAG optimization method and system to solve the technical problem of how to improve the adaptability, answer quality, and generalization and robustness of the RAG system in specific fields.
[0005] In a first aspect, the present invention provides a multi-domain adaptive RAG optimization method, comprising the following steps:
[0006] Data collection: Generate initial questions and answers and their related context from the corpus of the target domain through a general large language model and retriever;
[0007] Data screening: Based on the prompt mechanism of chain thinking, the large language model is guided to evaluate and classify the retrieved context and generate a detailed reasoning process. The generated answers are then evaluated by the large language model to screen out training samples.
[0008] Preference data generation: Divide the screened training samples into preferred responses and non-preferred responses, and construct comparative training samples;
[0009] Joint fine-tuning: Jointly fine-tune the large language model and retriever based on the filtered preference data.
[0010] Preferably, data collection includes the following steps:
[0011] Corpus input: Receive a domain-specific corpus D containing multiple text paragraphs as input. Corpus D is a public domain-specific dataset or a private dataset provided by the user. D={p1,p2,...,p n}, p n Indicates the nth paragraph;
[0012] Paragraph selection: randomly select a basic paragraph p from the corpus D by random sampling;
[0013] Question generation: Based on the basic paragraph p, a related question q is generated through a general large language model;
[0014] Standard answer generation: After generating question q, the corresponding answer a is generated based on the basic paragraph p and question q through a universal large language model;
[0015] Contextual retrieval: Based on the generated question q, the K paragraphs with the highest relevance to question q are retrieved from the corpus D through a dense vector retriever. The K paragraphs serve as the contextual information of the answer, where the K paragraphs are represented as represents the kth paragraph;
[0016] Data storage: The generated question q, answer a, basic paragraph p and context form a training sample s, and the training sample s is stored in the training dataset. The corpus input, paragraph selection, question generation, standard answer generation and context retrieval operations are repeated until a sufficient number of training samples are generated.
[0017] Preferably, data screening includes the following steps:
[0018] Chaining Prompts: A prompt mechanism based on chaining thinking is set up to prompt the large language model to evaluate each context before generating an answer, and classify the evaluation results into three categories: relevant, irrelevant, and misleading;
[0019] Generate reasoning process and answer: The large language model generates a reasoning process based on the prompt mechanism ′ , and generates the final answer based on the evaluation of the context, the reasoning process e ′ It includes the classification results of the context and the reasoning process of the large language model. The context is divided into two types, namely C1 and C2. C2 does not include the basic paragraph p;
[0020] Answer evaluation and screening: a ′To check the correctness of the answer, the question q, the basic paragraph p, and the answer a are provided to the large language model, and the large language model is prompted to judge the correctness of the answer. If the answer is correct, the training sample is removed and only the training samples with incorrect answers are retained;
[0021] Standard reasoning process generation and data screening: For training samples with incorrect answers, the question q, context C, and answer a are provided to the large language model, and the large language model is prompted to generate the answer. The training sample is then screened based on the predetermined screening rules to obtain the final training sample, which is represented as a triple (q+C,e+a,e ′ +a ′ ), where q+C represents the input, including question q and context C, and e+a represents the preferred response, including the reasoning process e and the answer a, e ′ +a ′ Indicates a non-preferred response, including the reasoning process e ′ and answer a ′ ;
[0022] The filtering rules include the following:
[0023] Basic Paragraph Assessment: For Context If the base paragraph p is not classified as relevant in the inference process e, the training sample is removed;
[0024] Context quantity evaluation: for contextual paragraphs The context does not include the base paragraph p. If more than two paragraphs are classified as related during the inference process, the training sample is removed.
[0025] As a preference, when generating preference data, the inference process e from the generator predicts the output ′ and answer a ′ Extract the training signal from ′ When correct, the context C is classified as a relevant paragraph as a positive sample P + , and the paragraphs classified as misleading are negative samples P - , and when the answer a is generated ′ When incorrect, the paragraphs in context C that are classified as relevant are negative samples, and the base paragraph p is always a positive sample.
[0026] Preferably, the joint fine-tuning includes the following steps:
[0027] Generator fine-tuning: Model training is based on the ORPO method that combines supervised fine-tuning SFT and preference alignment. The loss function L during training is G Expressed as:
[0028]
[0029] Among them, x represents the input q+C, y ω Indicates the preferred response e+a,y l Indicates a non-preferred response ′ +a ′ , m is y ω The number of tokens, λ represents the hyperparameter, σ represents the standard deviation, P θ (y ω |x) means given input x, predict the preferred response y ω The probability, P θ (y l |x) means given input x, predict the non-preferred response y ω probability;
[0030] Retriever fine-tuning: Collect positive samples P of question q from the generator’s generated data + and negative samples P - , using the contrast ranking loss function to fine-tune the retriever, the loss function L R Expressed as:
[0031]
[0032] Among them, p + ∈P + , P + is a positive sample, P - is a negative sample; E(·) denotes the text editor and τ denotes the temperature hyperparameter.
[0033] In a second aspect, the present invention provides a multi-domain adaptive RAG optimization system for performing model optimization using a multi-domain adaptive RAG optimization method as described in any one of the first aspects, the system comprising a data acquisition module, a data screening module, a preference data generation module, and a joint fine-tuning module;
[0034] The data acquisition module is used to generate initial questions and answers and their related context from the corpus of the target domain through a general large language model and retriever;
[0035] The data screening module uses a prompt mechanism based on chain thinking to guide the large language model to evaluate and classify the retrieved context and generate a detailed reasoning process. The large language model then evaluates the generated answers and selects training samples.
[0036] The preference data generation module is used to divide the screened training samples into preferred responses and non-preferred responses, and to construct comparative training samples;
[0037] The joint fine-tuning module is used to jointly fine-tune the large language model and retriever based on the filtered preference data.
[0038] Preferably, the data acquisition module is used to perform the following operations:
[0039] Corpus input: Receive a domain-specific corpus D containing multiple text paragraphs as input. Corpus D is a public domain-specific dataset or a private dataset provided by the user. D={p1,p2,...,p n}, p n Indicates the nth paragraph;
[0040] Paragraph selection: randomly select a basic paragraph p from the corpus D by random sampling;
[0041] Question generation: Based on the basic paragraph p, a related question q is generated through a general large language model;
[0042] Standard answer generation: After generating question q, the corresponding answer a is generated based on the basic paragraph p and question q through a universal large language model;
[0043] Contextual retrieval: Based on the generated question q, the K paragraphs with the highest relevance to question q are retrieved from the corpus D through a dense vector retriever. The K paragraphs serve as the contextual information of the answer, where the K paragraphs are represented as represents the kth paragraph;
[0044] Data storage: The generated question q, answer a, basic paragraph p and context form a training sample s, and the training sample s is stored in the training dataset. The corpus input, paragraph selection, question generation, standard answer generation and context retrieval operations are repeated until a sufficient number of training samples are generated.
[0045] Preferably, the data screening module is used to perform the following operations including the following steps:
[0046] Chaining Prompts: A prompt mechanism based on chaining thinking is set up to prompt the large language model to evaluate each context before generating an answer, and classify the evaluation results into three categories: relevant, irrelevant, and misleading;
[0047] Generate reasoning process and answer: The large language model generates a reasoning process based on the prompt mechanism ′ , and generates the final answer based on the evaluation of the context, the reasoning process e ′ It includes the classification results of the context and the reasoning process of the large language model. The context is divided into two types, namely C1 and C2. C2 does not include the basic paragraph p;
[0048] Answer evaluation and screening: a ′ To check the correctness of the answer, the question q, the basic paragraph p, and the answer a are provided to the large language model, and the large language model is prompted to judge the correctness of the answer. If the answer is correct, the training sample is removed and only the training samples with incorrect answers are retained;
[0049] Standard reasoning process generation and data screening: For training samples with incorrect answers, the question q, context C, and answer a are provided to the large language model, and the large language model is prompted to generate the answer. The training sample is then screened based on the predetermined screening rules to obtain the final training sample, which is represented as a triple (q+C,e+a,e ′ +a ′ ), where q+C represents the input, including question q and context C, and e+a represents the preferred response, including the reasoning process e and the answer a, e ′ +a ′ Indicates a non-preferred response, including the reasoning process e ′ and answer a ′ ;
[0050] The filtering rules include the following:
[0051] Basic Paragraph Assessment: For Context If the base paragraph p is not classified as relevant in the inference process e, the training sample is removed;
[0052] Context quantity evaluation: for contextual paragraphs The context does not include the base paragraph p. If more than two paragraphs are classified as related during the inference process, the training sample is removed.
[0053] As an example, when preference data is generated, the preference data generation module is used to predict the inference process of the output from the generator. ′ and answer a ′ Extract the training signal from ′ When correct, the context C is classified as a relevant paragraph as a positive sample P + , and the paragraphs classified as misleading are negative samples P - , and when the answer a is generated ′ When incorrect, the paragraphs in context C that are classified as relevant are negative samples, and the base paragraph p is always a positive sample.
[0054] Preferably, the joint fine-tuning module is configured to perform the following operations:
[0055] Generator fine-tuning: Model training is based on the ORPO method that combines supervised fine-tuning SFT and preference alignment. The loss function L during training is G Expressed as:
[0056]
[0057] Among them, x represents the input q+C, y ω Indicates the preferred response e+a,y l Indicates a non-preferred response ′ +a ′ , m is y ω The number of tokens, λ represents the hyperparameter, σ represents the standard deviation, P θ (y ω |x) means given input x, predict the preferred response y ω The probability, P θ (y l |x) means given input x, predict the non-preferred response y ω probability;
[0058] Retriever fine-tuning: Collect positive samples P of question q from the generator’s generated data + and negative samples P - , using the contrast ranking loss function to fine-tune the retriever, the loss function L R Expressed as:
[0059]
[0060] Among them, p + ∈P + , P + is a positive sample, P - is a negative sample; E(·) denotes the text editor and τ denotes the temperature hyperparameter.
[0061] The multi-domain adaptive RAG optimization method and system of the present invention have the following advantages:
[0062] 1. It realizes the adaptive optimization and adjustment of the RAG system in specific fields without relying on external field-specific data or models, reducing the cost of data acquisition and annotation, and improving the flexibility and scalability of the model in different fields;
[0063] 2. The prompt mechanism based on Chain of Thought (CoT) can effectively improve the quality of training data, thereby improving the quality of the RAG system's answers in specific areas;
[0064] 3. By jointly fine-tuning the generator and retriever, the ability of the two to work together is enhanced, effectively solving the performance bottleneck caused by the separate training of the retriever and generator in existing technologies;
[0065] 3. Adopting strict data screening and optimization strategies, through multiple rounds of self-alignment and feedback mechanisms, not only improves the quality and consistency of training data, but also improves the generalization ability and robustness of the model in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0067] The present invention will be further described below with reference to the accompanying drawings.
[0068] Figure 1 This is a flowchart of a multi-domain adaptive RAG optimization method in Example 1. DETAILED DESCRIPTION
[0069] The present invention will be further described below with reference to the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments given are not intended to limit the present invention. Unless there is a conflict, the embodiments of the present invention and the technical features in the embodiments may be combined with each other.
[0070] The embodiments of the present invention provide a multi-domain adaptive RAG optimization method and system for solving the technical problem of how to improve the adaptability, answer quality, and generalization and robustness of the RAG system in specific fields.
[0071] Example 1:
[0072] The present invention provides a multi-domain adaptive RAG optimization method, which includes four steps: data collection, data screening, preference data generation and joint fine-tuning.
[0073] Step S100: Data collection: Generate initial questions and answers and their related contexts from the corpus of the target domain through a general large language model and retriever.
[0074] As a specific implementation of data collection, this step includes the following operations:
[0075] (1) Corpus input: A domain-specific corpus D containing multiple text paragraphs is received as input. Corpus D is a public domain-specific dataset or a private dataset provided by the user. D = {p1, p2, ..., p n}, p n Indicates the nth paragraph;
[0076] (2) Paragraph selection: To ensure the randomness and diversity of the training data, a basic paragraph p is randomly selected from the corpus D by random sampling. To avoid selecting repeated paragraphs, a certain sampling interval can be set or different random seeds can be used;
[0077] (3) Question generation: Based on the basic paragraph p, a related question q is generated by the general large language model. To guide the large language model LLM to generate high-quality questions, a prompt word is designed: Please generate a relevant and challenging question based on the following paragraph to guide the LLM to generate more in-depth and difficult questions, thereby improving the quality of training data;
[0078] (4) Standard answer generation: After generating question q, the corresponding answer a is generated through the general large language model based on the basic paragraph p and question q. The prompt words are designed: Please generate an accurate and detailed answer based on the following paragraph and question. This guides the LLM to generate more accurate and comprehensive answers.
[0079] (5) Contextual retrieval: In order to enable LLM to generate answers using external knowledge, based on the generated question q, the K paragraphs with the highest relevance to question q are retrieved from the corpus D through a dense vector retriever. The K paragraphs serve as the contextual information of the answer, where the K paragraphs are represented as To improve retrieval results, the following strategies can be adopted: 1) Multi-searcher fusion: Use multiple search engines of different types to perform retrieval and fuse their results to improve the comprehensiveness and accuracy of retrieval; 2) Reranking mechanism: Rerank the retrieved paragraphs according to their relevance to the question, and select the top-ranked paragraphs as context;
[0080] (6) Data storage: The generated question q, answer a, basic paragraph p and context form a training sample s, and the training sample s is stored in the training dataset. The corpus input, paragraph selection, question generation, standard answer generation and context retrieval operations are repeated until a sufficient number of training samples are generated, where:
[0081] Step S200: Data screening: Based on the prompt mechanism of chain thinking, the large language model is guided to evaluate and classify the retrieved context, generate a detailed reasoning process, and evaluate the generated answers through the large language model to screen out training samples.
[0082] As a specific implementation of data screening, this step includes the following operations:
[0083] (1) Chain thinking prompts: In order to enable the LLM to evaluate and classify the retrieved context, a prompt mechanism is set based on chain thinking. The prompt mechanism is used to prompt the large language model to evaluate each context before generating an answer, and the evaluation results are divided into three categories: relevant, irrelevant, and misleading;
[0084] (2) Generate reasoning process and answer: The large language model generates a reasoning process based on the prompt mechanism. ′ , and generates the final answer based on the evaluation of the context, the reasoning process e ′ It includes the classification results of the context and the reasoning process of the large language model. The context is divided into two types, namely C1 and C2. C2 does not include the basic paragraph p;
[0085] (3) Answer evaluation and screening: To evaluate the answers generated by large language models ′ To check the correctness of the answer, the question q, the basic paragraph p, and the answer a are provided to the large language model, and the large language model is prompted to judge the correctness of the answer. If the answer is correct, the training sample is removed and only the training samples with incorrect answers are retained;
[0086] (4) Standard reasoning process generation and data screening: For training samples with incorrect answers, the question q, context C, and answer a are provided to the large language model, and the large language model is prompted to generate the reasoning process and the classification results for each context paragraph. The training samples are screened based on the predetermined screening rules to obtain the final training samples. The training samples are represented as a triple (q+C,e+a,e ′ +a ′ ), where q+C represents the input, including question q and context C, and e+a represents the preferred response, including the reasoning process e and the answer a, e ′ +a ′ Indicates a non-preferred response, including the reasoning process e ′ and answer a ′ .
[0087] The screening rules include the following:
[0088] (1) Basic paragraph evaluation: for context That is, the context contains the basic paragraph p. If the basic paragraph p is not classified as relevant in the reasoning process e, the training sample is removed;
[0089] (2) Context quantity evaluation: for the upper and lower paragraphs The context does not include the base paragraph p. If more than two paragraphs are classified as related during the inference process, the training sample is removed.
[0090] Step S300: Preference data generation: dividing the screened training samples into preferred responses and non-preferred responses, and constructing comparative training samples.
[0091] This example addresses the shortcomings of existing general-purpose retrievers in matching specific domain knowledge, as well as the problem that the retriever and generator are trained separately, resulting in the retrieval information being unable to effectively support generation and even misleading. To address this issue, it is necessary to fine-tune the retriever to better align it with the generator's preferences, thereby better supporting retrieval-enhanced generation within specific domains.
[0092] When preference data is generated, the inference process e of predicting the output from the generator ′ and answer a ′ Extract the training signal from ′ When correct, the context C is classified as a relevant paragraph as a positive sample P + , and the paragraphs classified as misleading are negative samples P - , and when the answer a is generated ′ When incorrect, the paragraphs in context C that are classified as relevant are negative samples, and the base paragraph p is always a positive sample.
[0093] Step S400: Joint fine-tuning: Based on the filtered preference data, the large language model and the retriever are jointly fine-tuned.
[0094] As a specific implementation of joint fine-tuning, this step includes the following operations:
[0095] (1) Generator fine-tuning: Model training is based on the ORPO method that combines supervised fine-tuning SFT and preference alignment, so that the model learns the preferred response e+a and its specific output style, thereby preventing the model from learning non-preferred responses e ′ +a ′ , the loss function L during training G Expressed as:
[0096]
[0097] Among them, x represents the input q+C, y ω Indicates the preferred response e+a,y l Indicates a non-preferred response ′ +a ′ , m is y ω The number of tokens, λ represents the hyperparameter, σ represents the standard deviation, P θ (y ω |x) means given input x, predict the preferred response y ω The probability, P θ (y l |x) means given input x, predict the non-preferred response yω probability;
[0098] (2) Retrieval fine-tuning: Collect positive samples P of question q from the generated data of the generator + and negative samples P - , using the contrast ranking loss function to fine-tune the retriever, the loss function L R Expressed as:
[0099]
[0100] Among them, p + ∈P + , P + is a positive sample, P - is a negative sample; E(·) denotes the text editor and τ denotes the temperature hyperparameter.
[0101] Example 2:
[0102] The present invention provides a multi-domain adaptive RAG optimization system, which includes a data acquisition module, a data screening module, a preference data generation module and a joint fine-tuning module.
[0103] The data acquisition module is used to generate initial questions and answers and their related context from the corpus of the target domain through a general large language model and retriever.
[0104] As a specific implementation of the data acquisition module, this module is used to perform the following operations:
[0105] (1) Corpus input: A domain-specific corpus D containing multiple text paragraphs is received as input. Corpus D is a public domain-specific dataset or a private dataset provided by the user. D = {p1, p2, ..., p n}, p n Indicates the nth paragraph;
[0106] (2) Paragraph selection: To ensure the randomness and diversity of the training data, a basic paragraph p is randomly selected from the corpus D by random sampling. To avoid selecting repeated paragraphs, a certain sampling interval can be set or different random seeds can be used;
[0107] (3) Question generation: Based on the basic paragraph p, a related question q is generated by the general large language model. To guide the large language model LLM to generate high-quality questions, a prompt word is designed: Please generate a relevant and challenging question based on the following paragraph to guide the LLM to generate more in-depth and difficult questions, thereby improving the quality of training data;
[0108] (4) Standard answer generation: After generating question q, the corresponding answer a is generated through the general large language model based on the basic paragraph p and question q. The prompt words are designed: Please generate an accurate and detailed answer based on the following paragraph and question. This guides the LLM to generate more accurate and comprehensive answers.
[0109] (5) Contextual retrieval: In order to enable LLM to generate answers using external knowledge, based on the generated question q, the K paragraphs with the highest relevance to question q are retrieved from the corpus D through a dense vector retriever. The K paragraphs serve as the contextual information of the answer, where the K paragraphs are represented as To improve retrieval results, the following strategies can be adopted: 1) Multi-searcher fusion: Use multiple search engines of different types to perform retrieval and fuse their results to improve the comprehensiveness and accuracy of retrieval; 2) Reranking mechanism: Rerank the retrieved paragraphs according to their relevance to the question, and select the top-ranked paragraphs as context;
[0110] (6) Data storage: The generated question q, answer a, basic paragraph p and context form a training sample s, and the training sample s is stored in the training dataset. The corpus input, paragraph selection, question generation, standard answer generation and context retrieval operations are repeated until a sufficient number of training samples are generated, where:
[0111] The data screening module is used to guide the large language model to evaluate and classify the retrieved context based on the prompt mechanism of chain thinking, and generate a detailed reasoning process. The generated answers are evaluated by the large language model to screen out training samples.
[0112] As a specific implementation of the data filtering module, this module is used to perform the following operations:
[0113] (1) Chain thinking prompts: In order to enable the LLM to evaluate and classify the retrieved context, a prompt mechanism is set based on chain thinking. The prompt mechanism is used to prompt the large language model to evaluate each context before generating an answer, and the evaluation results are divided into three categories: relevant, irrelevant, and misleading;
[0114] (2) Generate reasoning process and answer: The large language model generates a reasoning process based on the prompt mechanism. ′ , and generates the final answer based on the evaluation of the context, the reasoning process e ′ It includes the classification results of the context and the reasoning process of the large language model. The context is divided into two types, namely C1 and C2. C2 does not include the basic paragraph p;
[0115] (3) Answer evaluation and screening: To evaluate the answers generated by large language models ′ To check the correctness of the answer, the question q, the basic paragraph p, and the answer a are provided to the large language model, and the large language model is prompted to judge the correctness of the answer. If the answer is correct, the training sample is removed and only the training samples with incorrect answers are retained;
[0116] (4) Standard reasoning process generation and data screening: For training samples with incorrect answers, the question q, context C, and answer a are provided to the large language model, and the large language model is prompted to generate the reasoning process and the classification results for each context paragraph. The training samples are screened based on the predetermined screening rules to obtain the final training samples. The training samples are represented as a triple (q+C,e+a,e ′ +a ′ ), where q+C represents the input, including question q and context C, and e+a represents the preferred response, including the reasoning process e and the answer a, e ′ +a ′ Indicates a non-preferred response, including the reasoning process e ′ and answer a ′ .
[0117] The screening rules include the following:
[0118] (1) Basic paragraph evaluation: for context That is, the context contains the basic paragraph p. If the basic paragraph p is not classified as relevant in the reasoning process e, the training sample is removed;
[0119] (2) Context quantity evaluation: for the upper and lower paragraphs The context does not include the base paragraph p. If more than two paragraphs are classified as related during the inference process, the training sample is removed.
[0120] The preference data generation module is used to divide the screened training samples into preferred responses and non-preferred responses, and to construct comparative training samples.
[0121] This example addresses the shortcomings of existing general-purpose retrievers in matching specific domain knowledge, as well as the problem that the retriever and generator are trained separately, resulting in the retrieval information being unable to effectively support generation and even misleading. To address this issue, it is necessary to fine-tune the retriever to better align it with the generator's preferences, thereby better supporting retrieval-enhanced generation within specific domains.
[0122] As a specific implementation, the preference data generation module is used to predict the inference process of the output from the generator. ′ and answer a ′ Extract the training signal from ′When correct, the context C is classified as a relevant paragraph as a positive sample P + , and the paragraphs classified as misleading are negative samples P - , and when the answer a is generated ′ When incorrect, the paragraphs in context C that are classified as relevant are negative samples, and the base paragraph p is always a positive sample.
[0123] The joint fine-tuning module is used to jointly fine-tune the large language model and retriever based on the filtered preference data.
[0124] As a specific implementation of the joint fine-tuning module, this module is used to perform the following operations:
[0125] (1) Generator fine-tuning: Model training is based on the ORPO method that combines supervised fine-tuning SFT and preference alignment, so that the model learns the preferred response e+a and its specific output style, thereby preventing the model from learning non-preferred responses e ′ +a ′ , the loss function L during training G Expressed as:
[0126]
[0127] Among them, x represents the input q+C, y ω Indicates the preferred response e+a,y l Indicates a non-preferred response ′ +a ′ , m is y ω The number of tokens, λ represents the hyperparameter, σ represents the standard deviation, P θ (y ω |x) means given input x, predict the preferred response y ω The probability, P θ (y l |x) means given input x, predict the non-preferred response y ω probability;
[0128] (2) Retrieval fine-tuning: Collect positive samples P of question q from the generated data of the generator + and negative samples P - , using the contrast ranking loss function to fine-tune the retriever, the loss function L R Expressed as:
[0129]
[0130] Among them, p + ∈P + , P + is a positive sample, P - is a negative sample; E(·) denotes the text editor and τ denotes the temperature hyperparameter.
[0131] The above is a detailed introduction to the multi-domain adaptive RAG optimization method and system provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A multi-domain adaptive RAG optimization method, characterized in that: The steps include: Data collection: Generate initial questions and answers and their related context from the corpus of the target domain through a general large language model and retriever; Data screening: Based on the prompt mechanism of chain thinking, the large language model is guided to evaluate and classify the retrieved context and generate a detailed reasoning process. The generated answers are then evaluated by the large language model to screen out training samples. Preference data generation: Divide the screened training samples into preferred responses and non-preferred responses, and construct comparative training samples; Joint fine-tuning: Jointly fine-tune the large language model and retriever based on the filtered preference data.
2. The multi-domain adaptive RAG optimization method according to claim 1, characterized in that: Data collection includes the following steps: Corpus input: Receive a domain-specific corpus D containing multiple text paragraphs as input. Corpus D is a public domain-specific dataset or a private dataset provided by the user. D={p1,p2,...,p n }, p n Indicates the nth paragraph; Paragraph selection: randomly select a basic paragraph p from the corpus D by random sampling; Question generation: Based on the basic paragraph p, a related question q is generated through a general large language model; Standard answer generation: After generating question q, the corresponding answer a is generated based on the basic paragraph p and question q through a universal large language model; Contextual retrieval: Based on the generated question q, the K paragraphs with the highest relevance to question q are retrieved from the corpus D through a dense vector retriever. The K paragraphs serve as the contextual information of the answer, where the K paragraphs are represented as represents the kth paragraph; Data storage: The generated question q, answer a, basic paragraph p and context form a training sample s, and the training sample s is stored in the training dataset. The corpus input, paragraph selection, question generation, standard answer generation and context retrieval operations are repeated until a sufficient number of training samples are generated.
3. The multi-domain adaptive RAG optimization method according to claim 1, characterized in that: Data screening includes the following steps: Chaining Prompts: A prompt mechanism based on chaining thinking is set up to prompt the large language model to evaluate each context before generating an answer, and classify the evaluation results into three categories: relevant, irrelevant, and misleading; Generate reasoning process and answer: The large language model generates a reasoning process based on the prompt mechanism ′ , and generates the final answer based on the evaluation of the context, the reasoning process e ′ It includes the classification results of the context and the reasoning process of the large language model. The context is divided into two types, namely C1 and C2. C2 does not include the basic paragraph p; Answer evaluation and screening: a ′ To check the correctness of the answer, the question q, the basic paragraph p, and the answer a are provided to the large language model, and the large language model is prompted to judge the correctness of the answer. If the answer is correct, the training sample is removed and only the training samples with incorrect answers are retained; Standard reasoning process generation and data screening: For training samples with incorrect answers, the question q, context C, and answer a are provided to the large language model, and the large language model is prompted to generate the answer. The training sample is then screened based on the predetermined screening rules to obtain the final training sample, which is represented as a triple (q+C,e+a,e ′ +a ′ ), where q+C represents the input, including question q and context C, and e+a represents the preferred response, including reasoning process e and answer a, e ′ +a ′ Indicates a non-preferred response, including the reasoning process e ′ and answer a ′ ; The filtering rules include the following: Basic Paragraph Assessment: For Context If the base paragraph p is not classified as relevant in the inference process e, the training sample is removed; Context quantity evaluation: for contextual paragraphs The context does not include the base paragraph p. If more than two paragraphs are classified as related during the inference process, the training sample is removed.
4. The multi-domain adaptive RAG optimization method according to claim 1, characterized in that: When preference data is generated, the inference process e of predicting the output from the generator ′ and answer a ′ Extract the training signal from ′ When correct, the context C is classified as a relevant paragraph as a positive sample P + , and the paragraphs classified as misleading are negative samples P - , and when the answer a is generated ′ When incorrect, the paragraphs in context C that are classified as relevant are negative samples, and the base paragraph p is always a positive sample.
5. The multi-domain adaptive RAG optimization method according to claim 1, characterized in that: Joint fine-tuning includes the following steps: Generator fine-tuning: Model training is based on the ORPO method that combines supervised fine-tuning SFT and preference alignment. The loss function L during training is G Expressed as: Among them, x represents the input q+C, y ω Indicates the preferred response e+a,y l Indicates a non-preferred response ′ +a ′ , m is y ω The number of tokens, λ represents the hyperparameter, σ represents the standard deviation, P θ (y ω |x) means given input x, predict the preferred response y ω The probability, P θ (y l |x) means given input x, predict the non-preferred response y ω probability; Retriever fine-tuning: Collect positive samples P of question q from the generator’s generated data + and negative samples P - , using the contrast ranking loss function to fine-tune the retriever, the loss function L R Expressed as: Among them, p + ∈P + , P + is a positive sample, P - is a negative sample; E(·) denotes the text editor and τ denotes the temperature hyperparameter.
6. A multi-domain adaptive RAG optimization system, characterized in that: Used to perform model optimization by a multi-domain adaptive RAG optimization method according to any one of claims 1 to 5, the system comprising a data acquisition module, a data screening module, a preference data generation module, and a joint fine-tuning module; The data acquisition module is used to generate initial questions and answers and their related context from the corpus of the target domain through a general large language model and retriever; The data screening module uses a prompt mechanism based on chain thinking to guide the large language model to evaluate and classify the retrieved context and generate a detailed reasoning process. The large language model then evaluates the generated answers and selects training samples. The preference data generation module is used to divide the screened training samples into preferred responses and non-preferred responses, and to construct comparative training samples; The joint fine-tuning module is used to jointly fine-tune the large language model and retriever based on the filtered preference data.
7. The multi-domain adaptive RAG optimization system according to claim 6, characterized in that: The data acquisition module is used to perform the following operations: Corpus input: Receive a domain-specific corpus D containing multiple text paragraphs as input. Corpus D is a public domain-specific dataset or a private dataset provided by the user. D={p1,p2,...,p n }, p n Indicates the nth paragraph; Paragraph selection: randomly select a basic paragraph p from the corpus D by random sampling; Question generation: Based on the basic paragraph p, a related question q is generated through a general large language model; Standard answer generation: After generating question q, the corresponding answer a is generated based on the basic paragraph p and question q through a universal large language model; Contextual retrieval: Based on the generated question q, the K paragraphs with the highest relevance to question q are retrieved from the corpus D through a dense vector retriever. The K paragraphs serve as the contextual information of the answer, where the K paragraphs are represented as represents the kth paragraph; Data storage: The generated question q, answer a, basic paragraph p and context form a training sample s, and the training sample s is stored in the training dataset. The corpus input, paragraph selection, question generation, standard answer generation and context retrieval operations are repeated until a sufficient number of training samples are generated.
8. The multi-domain adaptive RAG optimization system according to claim 6, characterized in that: The data filtering module is used to perform the following operations including the following steps: Chaining Prompts: A prompt mechanism based on chaining thinking is set up to prompt the large language model to evaluate each context before generating an answer, and classify the evaluation results into three categories: relevant, irrelevant, and misleading; Generate reasoning process and answer: The large language model generates a reasoning process based on the prompt mechanism ′ , and generates the final answer based on the evaluation of the context, the reasoning process e ′ It includes the classification results of the context and the reasoning process of the large language model. The context is divided into two types, namely C1 and C2. C2 does not include the basic paragraph p; Answer evaluation and screening: a ′ To check the correctness of the answer, the question q, the basic paragraph p, and the answer a are provided to the large language model, and the large language model is prompted to judge the correctness of the answer. If the answer is correct, the training sample is removed and only the training samples with incorrect answers are retained; Standard reasoning process generation and data screening: For training samples with incorrect answers, the question q, context C, and answer a are provided to the large language model, and the large language model is prompted to generate the answer. The training sample is then screened based on the predetermined screening rules to obtain the final training sample, which is represented as a triple (q+C,e+a,e ′ +a ′ ), where q+C represents the input, including question q and context C, and e+a represents the preferred response, including reasoning process e and answer a, e ′ +a ′ Indicates a non-preferred response, including the reasoning process e ′ and answer a ′ ; The filtering rules include the following: Basic Paragraph Assessment: For Context If the base paragraph p is not classified as relevant in the inference process e, the training sample is removed; Context quantity evaluation: for contextual paragraphs The context does not include the base paragraph p. If more than two paragraphs are classified as related during the inference process, the training sample is removed.
9. The multi-domain adaptive RAG optimization system according to claim 6, characterized in that: When preference data is generated, the preference data generation module is used to predict the inference process of the output from the generator. ′ and answer a ′ Extract the training signal from ′ When correct, the context C is classified as a relevant paragraph as a positive sample P + , and the paragraphs classified as misleading are negative samples P - , and when the answer a is generated ′ When incorrect, the paragraphs in context C that are classified as relevant are negative samples, and the base paragraph p is always a positive sample.
10. The multi-domain adaptive RAG optimization system according to claim 6, characterized in that: The joint fine-tuning module is used to perform the following operations: Generator fine-tuning: Model training is based on the ORPO method that combines supervised fine-tuning SFT and preference alignment. The loss function L during training is G Expressed as: Among them, x represents the input q+C, y ω Indicates the preferred response e+a,y l Indicates a non-preferred response ′ +a ′ , m is y ω The number of tokens, λ represents the hyperparameter, σ represents the standard deviation, P θ (y ω |x) means given input x, predict the preferred response y ω The probability, P θ (y l |x) means given input x, predict the non-preferred response y ω probability; Retriever fine-tuning: Collect positive samples P of question q from the generator’s generated data + and negative samples P - , using the contrast ranking loss function to fine-tune the retriever, the loss function L R Expressed as: Among them, p + ∈P + , P + is a positive sample, P - is a negative sample; E(·) denotes the text editor and τ denotes the temperature hyperparameter.
Citation Information
Cited By
RAG optimization method fusing bidirectional confusion blocking and adaptive selection
CN121478990A
RAG-based question and answer system optimization method and device
CN121681771A
A method and device for optimizing a RAG-based question answering system
CN121681771B