Retrieval enhancement generation optimization system under multi-domain framework

By introducing a search-enhanced generation optimization system under a multi-domain framework in natural language processing, a large language model is solved in the face of hallucinations when facing knowledge beyond corpus and the problem of insufficient matching degree in query and external knowledge, achieving more accurate, diversified and targeted answer generation.

CN120162407APending Publication Date: 2025-06-17TIANJIN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224147.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

When using large language models for natural language processing, the prior art will experience problems beyond the corpus or internal parameterized knowledge when encountering problems that cannot be understood by the corpus or internal parameterized knowledge, which will affect the credibility of the generated answers, and there are problems such as insufficient matching degree and insufficient support in the combination of query and external knowledge.

Method used

Introduce a search enhancement generation optimization system under the multi-domain framework, including consistency generators and multi-domain searchers. Generate consistency question groups through synonymous extensions, and use domain classifiers and candidate filters for question filtering and background document generation to ensure the accuracy and diversity of answers.

Benefits of technology

It significantly improves the matching degree between user query questions and their true intentions, reduces hallucinations, enhances support from multiple fields, improves the accuracy and coverage of generated answers, and improves the overall efficiency and user experience of the question-and-answer system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162407A_ABST
    Figure CN120162407A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval enhancement generation optimization system under a multi-domain framework, and the system comprises the following modules: a consistency generator which is used for carrying out multiple queries on a query problem through synonymous extension to obtain a plurality of generation problems, and enabling the query problem and the plurality of generation problems to form a consistency problem group, correcting the matching degree between the query problem and the user intention; the multi-field retriever is used for carrying out background document generation on the generation problem from the perspective of a plurality of fields; the domain classifier is used for classifying query questions and generation questions in the consistency question group according to domain distribution; and the candidate filter is used for screening candidate generation questions in the consistency question group and ensuring the diversity and intention pertinence of generated documents.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence, computer natural language processing and generation, and particularly relates to a retrieval enhanced generation optimization system under a multi-domain framework. Background Art

[0002] When large language models reach hundreds of billions or even trillions of parameters, they exhibit deeper language understanding and more accurate text generation, and can complete various complex natural language processing tasks, such as text generation, dialogue systems, machine translation, etc. However, prompt-based large language models (LLMs) will hallucinate when encountering problems that exceed the corpus or cannot be understood by internal parametric knowledge, which affects the credibility of the generated answers. Some researchers have proposed using LLMs to generate relevant contexts and then providing the context as an additional input when answering common sense questions. Introducing a retriever into the LLM and adding the retrieved context to the prompt can effectively alleviate hallucinations. Retrieval-Augmented Generation (RAG) combines information retrieval systems with these powerful models, leveraging the language understanding and generation capabilities of LLMs, as well as the rich context information provided by the retrieval system, to jointly generate more accurate and relevant answers. This approach not only improves the quality of the answers but also enables the model to access a wider range of knowledge and information.

[0003] According to the above-mentioned retrieval-augmented generation paradigm, researchers are more focused on how to use components such as retrievers and evaluators to optimize external knowledge represented by unstructured information, and the thinking about user intentions represented by queries only stays at a shallow semantic understanding. However, the query is almost involved in the entire RAG process, and it is the key to initiate the entire information retrieval and text generation process in retrieval-augmented generation. There are differences between the query and the knowledge required for retrieval, and they do not match exactly. Such differences will be amplified in the subsequent process, directly affecting the quality of the final generated result. Therefore, researchers have started to conduct more explorations on RAG with the query as the guidance. Xinbei Ma et al. conducted an in-depth horizontal analysis and decomposition of the problem and proposed a new framework - the rewrite-retrieve-read framework, which better aligns the query with the frozen module. Shuting Wang et al. conducted a vertical excavation of the problem and proposed RichRAG, which allows the document to fully cover various query aspects through a sub-aspect browser and a multi-aspect retriever, and knows the preferences of the generator, thus motivating the generator to generate rich and comprehensive responses for users. Harsh Trivedi et al. integrated the problem chain and the retriever, intertwined the retrieval with the steps in COT, used COT to guide the retrieval, and then used the retrieval results to improve COT. Soyeong Jeong et al. proposed a new adaptive QA framework that can dynamically select the most suitable strategy for (retrieval-augmented) LLMs from the simplest to the most complex order according to the complexity of the query. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a retrieval-augmented generation optimization system under a multi-domain framework, which conducts a more in-depth analysis of user intentions represented by queries and makes a more effective combination of them with external knowledge represented by unstructured information in the processes of retrieval, enhancement, and generation, and conducts more effective explorations on sub-domains. This patent introduces a multi-domain framework and combines retrieval-augmented generation technology so that the large model can give more targeted and reasonable answers when facing knowledge-intensive fields.

[0005] The purpose of the present invention is achieved through the following technical solutions:

[0006] A retrieval-augmented generation optimization system under a multi-domain framework includes a consistency generator and a multi-domain retriever. The multi-domain retriever includes a domain classifier and a candidate filter:

[0007] The consistency generator is used to perform multiple queries on the query problem through synonym expansion to obtain several generated problems, and jointly form a group of consistency problem groups with the query problem to correct the matching degree between the query problem and the user intention;

[0008] In the multi-domain retriever, the domain classifier is used to classify the query questions and generated questions within the consistency problem group according to the domain distribution; the candidate filter is used to screen the candidate generated questions in the consistency problem group to ensure the diversity and intent pertinence of the generated documents; finally, background documents are generated for the generated questions from the perspectives of several domains to ensure the accuracy and comprehensiveness of the answers.

[0009] Furthermore, the consistency generator further includes a module for calculating the similarity between the query question and the generated question, so as to screen the candidate generated questions.

[0010] Furthermore, the domain classifier is a classifier based on a large language model, which is used to classify the generated questions in the consistency problem group, so as to match the corresponding domain content for the query question and each generated question.

[0011] Furthermore, the candidate filter screens questions by calculating the similarity and category difference between the query question and the generated question, and generates the topk candidate generated questions according to the screening results.

[0012] Furthermore, the multi-domain retriever guides the generation of background text from the perspectives of several domains by assigning background documents in different domains to the query question and each generated question, and uses it as the input for subsequent generation.

[0013] Furthermore, the joint action of the consistency generator, the domain classifier and the candidate filter improves the efficiency and answer quality in the generation process.

[0014] Compared with the prior art, the beneficial effects brought by the technical solution of the present invention are:

[0015] 1. Improve the matching degree of query intent: The original query question is synonymously expanded by the Consistency-based Generator to generate multiple queries to obtain several generated questions, which can comprehensively understand and express the user's intent, and avoid the intent ambiguity caused by the colloquial or incomplete query question. This feature significantly improves the matching degree between the user's query question and its true intent, and thus provides a more accurate basis for the subsequent retrieval and generation processes.

[0016] 2. Optimize the combination of query questions and external knowledge: The consistency generator not only performs synonymous expansion, but also embeds the similarity information between the query question and the generated question. By screening the consistency problem group, it can ensure that the generated questions are highly relevant to the intent of the original query question. This method effectively reduces the Hallucination phenomenon, ensures that the generated background documents can reflect the real needs, and thus improves the accuracy of the generated answers.

[0017] 3. Enhance multi-domain support and improve generation diversity: Through the Multi-view Retriever, perspectives from different domains can simultaneously participate in the generation of background documents. This multi-domain support enables the generated answers to not only originate from a single domain but also integrate cross-domain comprehensive knowledge, ensuring that the answers are more comprehensive and rich. For knowledge-intensive domains, especially open-domain question-answering tasks, this feature can significantly enhance the coverage and depth of the answers.

[0018] 4. Efficiently screen candidate questions and reduce redundancy and invalid generation: Candidate filters screen the candidate generated questions by calculating the similarity between questions and the differences in category distributions, ensuring that the finally generated background documents maintain diversity while avoiding redundancy. This mechanism avoids the low generation efficiency caused by repeated or invalid questions and improves the overall efficiency of the generation process.

[0019] 5. Optimization based on domain classification to ensure the pertinence of background documents: The Domain classifier accurately classifies and matches background documents in the corresponding domain for each question by analyzing the question categories in the consistent question group. This feature ensures that background information in different domains can be appropriately and accurately incorporated into the generation process, further improving the pertinence of the generated content and making the answers more in line with the knowledge requirements of specific domains.

[0020] 6. Improve the quality of question answering and user experience: Through the collaborative effect of the above-mentioned modules, the present invention significantly improves the quality of the answers of the question-answering system. Especially when facing query questions in knowledge-intensive domains, users can obtain more targeted, in-depth, and diverse answers, avoiding the low-quality answers caused by insufficient knowledge coverage or inaccurate generated information in traditional question-answering systems. At the same time, this technology can automatically adjust the generation strategy according to the complexity of the query question, thereby enhancing the user experience.

[0021] 7. Solve the problems of query question diversity and unclear intent: The present invention particularly solves the problems caused by "query question diversity" and "unclear intent" in the prior art. Through the cooperation of the consistency generator and the multi-domain retriever, it can effectively identify the diverse query needs of users and provide more accurate and in-depth answers for users with appropriate background information. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the system working process.

[0023] Figure 2 It is a schematic diagram of the specific working process of the consistency generator and the multi-domain retriever. DETAILED DESCRIPTION OF THE INVENTION

[0024] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0025] In view of the problems that the current method lacks a finer-grained analysis of user intentions and a deeper coupling between user intentions and external knowledge, this embodiment proposes a multi-domain retrieval enhanced generation optimization system. The system mainly includes a Consistency-based Generator and a Multi-view Retriever; the Multi-view Retriever includes a Domain classifier and Candidate filters.

[0026] (1) Consistency-based Generator

[0027] The query problem often deviates from the true user intention. The Consistency-based Generator can perform multiple queries on the original query problem through synonym expansion to obtain several generated problems, thereby further correcting the matching degree between the query problem and the user intention. In this embodiment, a large language model is used for synonym generation, and similarity is embedded to construct the Consistency-based Generator. The large language model has strong semantic understanding ability and rich knowledge reserve, and can effectively explore the true intention of the query problem in different contexts.

[0028] Since there will be a certain degree of hallucination and factual errors in the output of the large model, in order to prevent biases from being inherited and amplified after entering the RAG process, supervision and review need to be strengthened. Therefore, the similarity between the query problem and the generated problem also needs to be embedded in the consistency problem group as the basis for screening.

[0029] (2) Domain classifier

[0030] Since a generated problem sometimes involves multiple domains, the consistency problem group often contains problems in multiple domains. This requires a classifier to classify problems in different domains. There are many choices for classifiers, such as Decision Tree Classifier, K-Nearest Neighbors Classifier, Neural Network Classifier, etc. However, since the quality of classification is strongly related to expert matching and requires strong understanding and analysis ability and flexibility. Therefore, a classifier based on a large language model is finally still selected.

[0031] (3) Candidate filters

[0032] In a group of consistency query-questions, there are several generated questions and an original query question. If background document generation is directly performed on all the generated questions in the group, there may be phenomena such as duplicate generation and ineffective generation, resulting in problems such as low generation efficiency. Therefore, a candidate filter is designed to screen the questions in the group. The screening principle is to ensure the diversity and intention pertinence of the context documents. To synthesize these two goals, the similarity and category difference within the consistency query-questions group are selected as the two indicators of the candidate filter. By combining weights, the scores of each generated question are obtained, and then sorted to perform top-k screening. Top-k screening means screening out the top k combined weights by sorting the combined weights.

[0033] Specifically, as Figure 1 shown, the working process of this multi-domain retrieval enhanced generation optimization system is as follows:

[0034] First, the original query question query is input into the Consistency-based Generator for synonym expansion to achieve multiple query rewriting. The query question and several generated questions together form a group of consistency query-questions, which can automatically complete the content of the query question and ask questions with diverse word orders. Given a dataset for a knowledge-intensive task (e.g., open-domain QA), D = {(x, y)i}, i = 0, 1, 2,..., N, where x is the query question and y is the expected response. Take a certain x as the original query question Q and the prompt P extend indicating the large language model LLM to predict and expand and generate as the input, and a group of consistency query-questions is obtained, that is, {Q, q1, q 2, ..., q n}:

[0035] {Q, q1, q2,..., q n} ~ LLMs(Q, P extend )

[0036] It is also necessary to embed the similarity between the query question and each generated question into the group of consistency query-questions as the basis for screening. First, the group of consistency query-questions needs to be converted into a vector group:

[0037] {vector 1 , vector 2 ,..., vector n} = Embedding{Q, q1, q2...q n},

[0038]

[0039] where vector i represents the vectorized representation for generating question q i , and is the vector parameter of the j-th dimension, and vector i is extended to l dimensions. q i are the respective generated questions;

[0040] Then, the cosine similarity between the query question within the consistency question group and each generated question is calculated through the vector group, {s1, s2,..., s n}, and s i is the similarity between Q and q i :

[0041]

[0042] {s1, s2,..., s n} = similarity({Q, q1, q2,..., q n )

[0043] The class distribution of the consistency question group plays a crucial role in screening the final topk, and this step is achieved through the domain classifier Domain classifier. The class distribution in the consistency question group is obtained through the domain classifier Domain classifier, and the distribution is input into the multi-domain retriever Multi-view Retriever to match the corresponding domain content for each generated question to generate background knowledge. The domain classifier takes the consistency question group and the prompt P category indicating the LLM for classification and restriction as inputs, and obtains a set of classification sets corresponding to the consistency question group, that is, {c1, c2,..., c n}:

[0044] {c1, c2,..., c n} ~ LLMs({Q, q1, q 2, ..., q n}, P category )

[0045] Different role generations often result in different contents, which can greatly enrich the domain diversity in the candidate library. However, if the background texts in the candidate library are not screened, it is easy to produce duplication and redundancy, affecting the accuracy and generation efficiency of the final response. Candidate filters filter the candidate library by calculating the cosine similarity and the variance of category distribution between the query question and each generated question, so that the finally formed context document can not only reflect the in-depth thinking of the query question, but also make a constructive response widely and comprehensively. The similarity obtained from the consistency generator is used to measure the intention pertinence, while the differential distribution of categories is reflected by the distribution of categories within the group:

[0046] {α1, α2,..., α n} = {s1, s2,..., s n} = similarity({Q, q1, q2,..., q n})

[0047]

[0048] where α represents the intention parameter, β represents the distribution parameter, and {β1, β2,..., β n} is obtained by the formula. And W represents the comprehensive weight corresponding to each question, which can reflect both the similarity within the consistency question group and the category difference.

[0049]

[0050] After obtaining the overall {W1, W2,..., W n} through formula calculation, topk generated questions are screened out for multi-role retrieval according to needs.

[0051] Get the filtered consistency question group: {Q, q1, q2,..., q k}, because the category distributions within each consistency question group are inconsistent, matching multiple different domain contents for each generated question to participate in the RAG process can more pertinently guide the "black box" large model to generate background knowledge texts in the corresponding domain. Take the filtered generated questions and the indicators indicating the LLM for background document retrieval as the input, and obtain a set of document sets corresponding to the consistency question group, that is, {doc1, doc2,..., doc k}:

[0052]

[0053] The present invention is not limited to the embodiments described above. The above description of the specific embodiments is intended to describe and illustrate the technical solutions of the present invention. The above specific embodiments are merely illustrative and not restrictive. Without departing from the spirit of the present invention and the scope protected by the claims, those of ordinary skill in the art can make many specific changes in form under the inspiration of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A retrieval enhancement generation optimization system under a multi-domain framework, characterized in that: Includes a consistency generator and a multi-domain retriever, which includes a domain classifier and candidate filters: The consistency generator is used to perform multiple queries on the query problem through synonym expansion to obtain a number of generated questions, and the query problem and the generated questions are combined into a set of consistent question groups to correct the matching degree between the query problem and the user's intention; The domain classifier in the multi-domain retriever is used to classify the query questions and generated questions in the consistency question group according to the domain distribution; Candidate filters are used to screen candidate generated questions in the consistency question group to ensure the diversity and targeted intent of the generated documents; finally, background documents are generated for the generated questions from the perspective of several fields to ensure the accuracy and comprehensiveness of the answers.

2. According to the multi-domain framework retrieval enhancement generation optimization system of claim 1, it is characterized by: The consistency generator also includes a module for calculating the similarity between the query question and the generated question so as to screen the candidate generated questions.

3. According to the multi-domain framework retrieval enhancement generation optimization system of claim 1, it is characterized in that: The domain classifier is a classifier based on a large language model, and is used to classify the generated questions in the consistency question group so as to match the corresponding domain content for the query question and each generated question.

4. According to the multi-domain framework retrieval enhancement generation optimization system of claim 1, it is characterized in that: The candidate filter performs question screening by calculating the similarity and category difference between the query question and the generated question, and generates topk candidate generated questions according to the screening results.

5. According to the multi-domain framework retrieval enhancement generation optimization system of claim 1, it is characterized in that: The multi-domain retriever guides the generation of background text from the perspectives of several domains by assigning background documents in different domains to the query question and each generated question, and uses the background text as the input for subsequent generation.

6. According to the multi-domain framework retrieval enhancement generation optimization system of claim 1, it is characterized by: The consistency generator, domain classifier and candidate filter work together to improve the efficiency and answer quality of the generation process.