Retrieval enhancement generation optimization method and device based on adaptive classification and autonomous reasoning

By constructing a proprietary domain knowledge vector database and combining adaptive classification and autonomous reasoning, the problem of inaccurate answers in professional domain knowledge tasks is solved, and the answers are achieved with high accuracy and reliability.

CN120429449APending Publication Date: 2025-08-05ZHEJIANG UNIV +2

Patent Information

Application Number
CN202510474694.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

Large language models cannot effectively respond to the latest information when facing professional domain knowledge tasks, and external knowledge sources may provide inaccurate answers.

Method used

By building a proprietary domain knowledge vector database, combining adaptive classification and autonomous inference, using semantics and word frequency similarity calculation to filter documents, the adaptive classification module performs document classification, the autonomous inference module corrects errors, generates answers, and verifies the correctness of the answers through the adaptive classification module.

Benefits of technology

Improve the accuracy of answers, reduce the interference of error messages on answers, and ensure the reliability of generated results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429449A_ABST
    Figure CN120429449A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval enhancement generation optimization method and device based on adaptive classification and autonomous reasoning. The method comprises the following steps: constructing a domain knowledge vector database; inputting a proprietary domain problem, and finding a plurality of matched knowledge documents with the highest score in the vector database from two dimensions of semantics and word frequency; inputting the matched knowledge document into an adaptive classification module to obtain a semantic attribute category of the document; the document and the corresponding semantic attribute category are input into an autonomous reasoning module, uncertain items and error items are eliminated, and enhanced knowledge is obtained; the enhanced knowledge is sent to the intelligent question and answer module to generate a response answer; the response answers are sent to the self-adaptive classification module again, and semantic attribute categories of the response answers are checked; if the answer is not judged to be correct, the autonomous reasoning module and the subsequent steps are executed again on the response answer; and if the answer is judged to be correct, displaying the answer to the user. The problem that in the prior art, an external knowledge source has negative effects on a generation result can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a retrieval enhancement generation optimization method and device based on adaptive classification and autonomous reasoning. Background Art

[0002] In recent years, with the rapid development of artificial intelligence and natural language processing technologies, large-scale pre-trained language models (big models) have demonstrated outstanding performance in many tasks, such as text generation, text classification, and knowledge question answering. However, for knowledge-intensive tasks, big models themselves cannot effectively handle tasks that require the latest information or specialized domain knowledge because the model parameters themselves cannot be updated.

[0003] Retrieval Augmented Generation (RAG) is a common method to solve the knowledge deficiency of large models. It applies knowledge retrieval technology to large models, retrieves information related to the user's query (such as documents, tables, conversations, etc.) from external knowledge bases, and uses contextual learning technology to allow the large model to answer the user's query based on the retrieved information, thereby alleviating the hallucination problem of large models.

[0004] For example, Chinese patent document with publication number CN119312917A discloses a retrieval enhancement generation method, device and medium based on a large model; Chinese patent document with publication number CN118394890A discloses a knowledge retrieval enhancement generation method and system based on a large language model.

[0005] However, retrieval enhancement generation only retrieves all relevant information based on vector similarity, without determining whether there is erroneous information in these documents. When erroneous information exists, the large model will be misled by the erroneous information related to the query and output an answer that is faithful to the erroneous information, which seriously affects the correctness of the answer output by the large model. Summary of the Invention

[0006] The present invention provides a retrieval enhancement generation optimization method and device based on adaptive classification and autonomous reasoning, which can solve the problem in the prior art that external knowledge sources have a negative impact on generation results.

[0007] A retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning includes the following steps:

[0008] (1) Collect domain-specific knowledge and build a vector database of domain-specific knowledge;

[0009] (2) Input relevant domain-specific questions. The vector retrieval module calculates the correlation between the two dimensions, semantics and word frequency, from the vector database and finds the highest-scoring matching knowledge documents after mixing them with parameter weights.

[0010] (3) Input the domain-specific questions and the matched knowledge documents into the adaptive classification module. Based on the document classification model deconstructed by the single decoder, the semantic attribute category corresponding to each knowledge document and the relevance label and correctness label corresponding to the classification are obtained;

[0011] (4) Input the knowledge document and the corresponding semantic attribute categories into the autonomous reasoning module, use the large model to eliminate the uncertainties and errors in the knowledge document, and infer the enhanced knowledge;

[0012] (5) The enhanced knowledge is fed into the intelligent question-answering module, which uses the large language model to generate responses to questions;

[0013] (6) The response answer generated in step (5) is used as a knowledge document and is re-sent to the adaptive classification module to check the semantic attribute category of the response answer; if the response answer is not judged as the correct answer by the adaptive classification module, the response answer is used as input and step (4) and subsequent steps are re-executed; if the response answer is judged as the correct answer by the adaptive classification module, the response answer is displayed to the user.

[0014] The specific process of step (2) is:

[0015] (2-1) Convert the input domain-specific problem q into a set of vectors e q ;

[0016] (2-2) The vector e of the domain-specific problem q Vector Database The domain knowledge vector e in D Correspondingly, calculate the semantic similarity of the two vectors mean (q, D) and word frequency similarity word (q, D), after calculation, the text similarity result of the document is obtained by mixing the weight parameters;

[0017] (2-3) Re-rank the documents from high to low according to the results of text similarity calculation, with the documents at the top representing higher similarity;

[0018] (2-4) Take the domain knowledge with the highest similarity as the result of vector retrieval and recall.

[0019] In step (2-2), the text similarity calculation formula of the document is as follows:

[0020] scorei =α×similarity mean (q,D i )+β×similarity word (q,D i )

[0021] Among them, q is the query domain question, D is the knowledge document in the vector database; similarity mean (q,D) function represents the semantic similarity between q and D, similarity word (q, D) represents the word frequency similarity between q and D; α and β are weight parameters;

[0022] Semantic similarity mean The calculation formula for (q,D) is as follows:

[0023]

[0024] Among them, ||e q || represents the vector e q The modulus length of ||e D || represents the vector e D Length of the module;

[0025] Word frequency similarity word The calculation formula for (q,D) is as follows:

[0026]

[0027] Where n is the total number of words, f(q i ,D) represents the vocabulary q i The number of occurrences in the knowledge document D, k and b are the hyperparameters for adjustment, length avg is the average length of knowledge documents.

[0028] The similarity calculation method of the present invention calculates from two dimensions, semantics and word frequency, and mixes them through parameter weights, which solves the pain point that traditional methods may ignore the frequency differences of important words and reduces the interference of noise and irrelevant words on similarity.

[0029] In step (3), the adaptive classification module builds a deep learning model based on a single decoder structure, which is responsible for classifying the domain-specific question q and the knowledge document D according to the two dimensions of relevance and correctness. The model uses adaptive task-aware normalization to dynamically adjust the normalization behavior, and uses hierarchical bidirectional attention to decouple and separate the features of the two modalities. The specific working process is as follows:

[0030] (3-1) The input is the domain-specific question q and the knowledge document D. First, the embedding layer is calculated to generate a vector representation E q and E D , and after generating the vector, the rotation position encoding is calculated and added to the vector embedding as follows:

[0031] E q =Embedding(q)+P rot (q)

[0032] E D =Embedding(D)+P rot (D)

[0033] Among them, E q and E D It represents the vector representation of the domain-specific question q and the knowledge document D after the embedding layer. Embedding() represents the word embedding operation. rot () indicates the rotation position encoding calculation;

[0034] (3-2) Vector representation E q and E D Each encoder layer calculates the bidirectional attention score of the input through the attention mechanism, and then adapts it to the current classification dimension through adaptive task-aware normalization.

[0035] Adaptive task-aware normalization builds a lightweight fully connected network using a single-layer MLP to generate independent normalization parameters γ for classification tasks of different dimensions. task and β task , the input of MLP is the global pooling result output by the Encoder part, and the calculation formula is as follows:

[0036] x=concat(Pooling(H),t task )

[0037] [γ task ,β task ]=MLP(x)

[0038] Among them, Pooling() is the global pooling result, H is the attention score vector output by the encoder, and t task is a one-hot vector used to determine whether the current classification dimension is correctness or relevance; the final normalized calculation formula is as follows:

[0039]

[0040] (3-3) At the top level of the encoder, bidirectional attention is used to decouple the unimodal features (internal information of each) and cross-modal features (interactive information) of the domain question q and the knowledge document D. First, the domain question q and the knowledge document D calculate their own self-attention results, and then the interaction between the domain question q and the knowledge document D is modeled through bidirectional attention. Finally, the classification result is obtained through two different gating networks. The calculation formula is as follows:

[0041] D att =Self-Attention(x D ),Q att =Self-Attention(x q )

[0042] I att =Bi-Attention(D att ,Q att )

[0043] Among them, Self-Attention() represents the calculation result of self-attention, Bi-Attention() represents the calculation result of bidirectional attention; Q att and D att Represents the domain problem vector x q and knowledge document or response question vector x D The self-attention calculation results, I att Q att and D att The calculation results of bidirectional attention.

[0044] Finally, the attention results of the three levels are concatenated and assigned to different classification dimensions through the task gating network:

[0045] F rel =Gate rel ([D att ;Q att ;I att ])∈{True,False}

[0046] F cor =Gate cor ([D att ;Q att ;I att ])∈{True,False}

[0047] Among them, Gate() represents the gated network, which is a lightweight fully connected layer; F rel Represents the classification result of the correlation dimension, F cor Represents the classification results of the correctness dimension.

[0048] In the adaptive classification module, adaptive task-aware normalization constructs a lightweight fully connected network by dynamically generating task-specific normalization parameters γ task and β task , combined with the task identification vector t task And the global pooling results are combined to achieve adaptive normalization for different tasks, which solves the problem that fixed normalization parameters in multi-task learning cannot adapt to different task distributions and introduces a task-aware dynamic adjustment mechanism.

[0049] Bidirectional attention decoupling decouples the attention mechanism into unimodal self-attention (capturing the internal semantics of the document or query) and cross-modal bidirectional attention (modeling the associative semantics of the document and query). It solves the problem of insufficient task differentiation caused by unidirectional interaction or mixed features in the traditional attention mechanism, explicitly separates unimodal and cross-modal features, and improves the flexibility of feature selection for multiple tasks (such as relevance and correctness).

[0050] In step (3), the training process of the adaptive classification module is as follows:

[0051] Collect query documents from various fields and positive samples (q, D, y rel ,y cor ), and construct a dataset D for training the classification model r * , dataset D r * The negative samples in are obtained by random sampling; where q represents the query, D represents the document, and y rel represents the category of document D divided according to its relevance to query q, y cor Represents the categories of document D divided according to the correctness of query q;

[0052] In the classification dataset D r * The classification model is trained above, and the training goal is to minimize the multi-task joint cross entropy loss function:

[0053]

[0054] Among them, w rel and w cor is the loss function score weight of the relevance dimension and the correctness dimension, y rel and y cor They are the relevance dimension and correctness dimension labels actually corresponding to the document, and They are the document relevance dimension and correctness dimension labels predicted by the model respectively.

[0055] The specific process of step (4) is:

[0056] (4-1) For label y rel =True and y cor = True, the autonomous reasoning module regards the document as the answer document. When it is subsequently fed into the intelligent question-answering module, the autonomous reasoning module will re-sort the documents from high to low confidence based on the confidence of the classification results of these documents, and give them a positive weight at the attention level;

[0057] (4-2) For label y rel =False and y cor = False, the autonomous reasoning module regards the document as an interference document. When it is subsequently sent to the intelligent question-answering module, the document is marked as an error document for the large model to judge on its own and is given a negative weight at the attention level.

[0058] (4-3) For label y rel =True or y cor = True, the autonomous reasoning module regards the document as an uncertain document and trains the autonomous reasoning model to decompose the knowledge document D into multiple atomic facts {f1,f2,...,f n}, for each atomic fact f i The fact is re-sent to the adaptive classification module to check its relevance; facts that are still irrelevant are removed from the list, otherwise the trained autonomous reasoning model is used to determine whether the error in the document can be corrected based on the knowledge of the model itself; atomic facts that cannot be corrected will be removed from the list to reduce the impact on subsequent outputs.

[0059] In step (4), the training process of the autonomous reasoning module is as follows:

[0060] Collect atomic fact question-answer pairs (f, t) that meet the error correction requirements, where f represents the atomic fact and t represents the corrected true fact.

[0061] Construct a dataset D for training autonomous reasoning models * For each positive sample (f, t), the model uses prompt word engineering to make the model think about the authenticity of fact f when inputting f, and attempts to output the corrected result t' after the model thinks about it; at the same time, negative samples are randomly sampled, and for negative samples, the model outputs a special token when inputting to indicate that the model cannot correct this knowledge;

[0062] In the autonomous reasoning dataset D * The large language model LLM is trained above, and the training goal is to minimize the cross entropy loss function, that is:

[0063]

[0064] Among them, y ij is the category j, y corresponding to the i-th token ij ' is the probability that the i-th token is predicted to be of category j.

[0065] The specific process of step (6) is as follows:

[0066] (6-1) The generated response answer ans is re-sent to the adaptive classification module to obtain the relevance and correctness label of the response answer ans relative to the question q;

[0067] (6-2) If the classification result is y rel =False or y cor = False, it means that the generated answer is wrong. In this case, the response answer is sent back to the autonomous reasoning module ans for deconstruction. The deconstructed answer fragment is sent back to the adaptive classification module to check the relevance and correctness of the fragment. Any irrelevant or incorrect answer fragment will be removed.

[0068] (6-3) Combine the remaining answer fragments with the original context to regenerate a new response answer; repeat (6-1) to (6-3) until the correct response answer is generated.

[0069] The present invention addresses the pain point in traditional methods where large language models may generate incorrect answers due to hallucinations by re-examining the corresponding answers generated by the model and iteratively eliminating hallucinations in the response answers.

[0070] A retrieval enhancement generation optimization device based on adaptive classification and autonomous reasoning includes a memory and one or more processors. The memory stores executable code. When the one or more processors execute the executable code, they are used to implement the above-mentioned retrieval enhancement generation optimization method.

[0071] Compared with the prior art, the present invention has the following beneficial effects:

[0072] The present invention can classify the retrieved documents to avoid the phenomenon that the answer generation of the large language model is interfered with by the documents that are highly relevant to the answer but do not contain the answer content. It can effectively improve the accuracy of the answer text generated by retrieval enhancement and reduce the interference of impure corpora in the real world on the generation results of the large model. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 This is a flow chart of a retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to an embodiment of the present invention.

[0074] Figure 2This is a structural diagram of a document classification model based on single decoder deconstruction in an embodiment of the present invention. DETAILED DESCRIPTION

[0075] The present invention will be described in further detail below with reference to the accompanying drawings and examples. It should be noted that the following examples are intended to facilitate understanding of the present invention and do not have any limiting effect on the present invention.

[0076] like Figure 1 As shown, a retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning includes the following steps:

[0077] S01: Collect a large amount of diverse text corpora from various scenarios to train the classifiers used in the framework. The text corpus contains multiple texts, covering common world knowledge in both Chinese and English, as well as various forms of language organization, including question-and-answer pairs, books, web pages, papers, and code. The collected corpus is cleaned and filtered, including removing noise, duplicate data, irrelevant data, and harmful data. The text corpus is then labeled and classified to construct a training dataset.

[0078] The purpose of making labels is to train a classifier for a query question q and a related knowledge document D. The goal is to output the correct category y rel and y cor Used to determine the correctness and relevance of documents relative to questions.

[0079] S02, training a classifier based on a single encoder deep learning model structure as an adaptive classification module Specifically, we use the collected text from various fields to construct a training dataset D r * , the dataset format is query document pair (q,D,y rel ,y cor ), where q represents the query question, D represents the document, and y rel and y cor The categories representing the relevance and correctness of the query and document D.

[0080] In the classification dataset D r * The classification model is trained above, and the training goal is to minimize the multi-task joint cross entropy loss function:

[0081]

[0082] Among them, w rel and w cor is the loss function score weight of the relevance dimension and the correctness dimension, yrel and y cor They are the relevance dimension and correctness dimension labels actually corresponding to the document, and They are the document relevance dimension and correctness dimension labels predicted by the model respectively.

[0083] S03, training the reasoning model in the autonomous reasoning module Collect atomic fact question-answer pairs (f, t) that meet the error correction requirements, where f represents the atomic fact and t represents the corrected true fact. Note that f is not guaranteed to be an incorrect fact, that is, it is possible that f itself is a correct fact, in which case f is equivalent to t. Then use the atomic fact question-answer pairs (f, t) to construct a dataset D for training the autonomous reasoning model. * Specifically, for each (f, t), the model uses prompt word engineering to let the model think about the authenticity of fact f when it is input, and tries to output the corrected result t' after the model thinks about it.

[0084] In the autonomous reasoning dataset D * Lora fine-tuning is performed on the large language model (LLM). Lora fine-tuning is a method for efficiently fine-tuning large language models. It is implemented by adding two low-rank matrices next to the original model weight matrix. During fine-tuning, only the parameters of the low-rank matrices need to be updated, without adjusting all the parameters of the original matrix. This significantly reduces computational costs and ensures effective adaptation to new tasks.

[0085] The training goal is to minimize the cross entropy loss function, that is:

[0086]

[0087] Among them, y ij is the category j, y corresponding to the i-th token ij ' is the probability that the i-th token is predicted to be of category j.

[0088] S04, use the embedding model to generate feature vectors for the collected domain-specific knowledge documents and create a vector database The database is used to calculate text similarity when a question is input and to find semantically similar documents as candidate documents.

[0089] For the embedding model, select an open-source embedding model (such as BGE or Contrivier) to generate sentence vectors in the feature space for the corpus documents. BGE is a new embedding model proposed by the Beijing Zhiyuan Artificial Intelligence Research Institute. It uses the RetroMAE architecture, an asymmetric encoder-decoder structure. This model can handle inputs of different granularities, such as "sentences" or "documents," and supports both dense and lexical retrieval.

[0090] Dense retrieval is given to a text encoder that transforms the input query q into a hidden state H q , and uses a normalized hidden state with a special tag [CLS] to represent the embedding representation vector e of the input query q =norm(H q0 ), similarly, for a document D, its embedding representation vector e can also be obtained D =norm(H D0 ). The semantic relevance score can be measured by the inner product of two embedding representation vectors, calculated as follows:

[0091]

[0092] Among them, similarity mean The (q,D) function represents the semantic similarity between q and D, where ||e q || represents the vector e q The mold length.

[0093] Lexical retrieval typically uses the TF-IDF formula to evaluate a document's importance to a corpus. A document's importance increases proportionally with the number of times its words appear in the document set, and decreases inversely with the frequency of its appearance in the corpus. The key idea is that if a word or phrase appears frequently in one article and rarely in other articles, it is considered to have good classificatory power and is suitable for classification. The calculation formula is as follows:

[0094]

[0095] Among them, similarity word (q, D) represents the similarity between q and D in terms of word frequency, where n is the total number of words, f(q i ,D) represents the vocabulary q i The number of occurrences in D, k and b are hyperparameters for adjustment, length avg is the average length of the documents.

[0096] S05, when question q is asked, the embedding model is used to generate the feature vector e of the question q , and in the vector database Calculate the text similarity of the question and find the documents with the highest scores as candidates. The formula for calculating text similarity is as follows:

[0097] score i =α×similarity mean (q,D i)+β×similarity word (q,D i )

[0098] t=argmax(score i ),

[0099] Among them, q is the query question, D is the vector database The knowledge documents in . The final similarity of the two dimensions is weighted by the parameters α and β. t represents the index of the knowledge with the largest dot product calculation result.

[0100] S06, traverse each candidate document D i , input the document into the adaptive classification module, the classifier model Responsible for predicting document D i For the query q, the categories of semantic correctness and semantic relevance Represents the result of document classification.

[0101] like Figure 2 As shown in the figure, the adaptive classification module builds a deep learning model based on a single decoder structure, adopts adaptive task-aware normalization to dynamically adjust the normalization behavior to enhance adaptability to different tasks, and adopts hierarchical bidirectional attention to decouple and separate modal features. The specific process is as follows:

[0102] For a given input query q and document D, the embedding layer is first calculated to generate a vector representation E q and E D , and after generating the vector, the rotation position encoding is calculated and added to the vector embedding as follows:

[0103] E q =Embedding(q)+P rot (q)

[0104] E D =Embedding(D)+P rot (D)

[0105] Among them, E q and E D Represents the vector representation generated by query q and document D x after passing through the embedding layer, Embedding() represents the word embedding operation, P rot () indicates the rotation position encoding calculation.

[0106] Vector representation E q and E DEach encoder layer calculates the bidirectional attention score of the input through the attention mechanism, and then adapts it to the current classification dimension through adaptive task-aware normalization. The attention calculation formula is as follows:

[0107]

[0108] Among them, x is the input feature, QKV is the query, key and value matrix obtained by linear transformation, d k is the dimension of the key matrix.

[0109] Adaptive task-aware normalization builds a lightweight fully connected network and uses a single-layer MLP to generate independent normalization parameters γ for classification tasks of different dimensions. task and β task , the input of MLP is the global pooling result output by the Encoder part, and the calculation formula is as follows:

[0110] x=concat(Pooling(H),t task )

[0111] [γ task ,β task ]=MLP(x)

[0112] Among them, Pooling() is the global pooling result, H is the attention score vector output by the encoder, and t task is a one-hot vector used to determine whether the current classification dimension is "correctness" or "relevance". The final normalized calculation formula is as follows:

[0113]

[0114] At the top level of the encoder, bidirectional attention is used to decouple the unimodal features (internal information of each) and cross-modal features (interaction information) of query q and document D. First, the self-attention results of query q and document D are calculated. Then, the interaction between query q and document D is modeled through bidirectional attention. Finally, the classification result is obtained through two different gating networks. The calculation formula is as follows:

[0115] D att =Self-Attention(x D ),Q att =Self-Attention(x q )

[0116] I att =Bi-Attention(D att ,Q att )

[0117] Among them, Self-Attention() represents the self-attention calculation result, and Bi-Attention() represents the bidirectional attention calculation result.

[0118] Finally, the attention results of the three levels are concatenated and assigned to different classification dimensions through the task gating network:

[0119] F rel =Gate rel ([D att ;Q att ;I att ])∈{True,False}

[0120] F cor =Gate cor ([D att ;Q att ;I att ])∈{True,False}

[0121] Among them, Gate() represents the gated network, which is a lightweight fully connected layer.

[0122] S07: Process the document based on the result in S06. Specifically, it includes the following steps:

[0123] For the classification result y rel =True and y cor = True, the autonomous reasoning module regards the document as the answer document. When it is subsequently sent to the question-answering module, the autonomous reasoning module will re-sort the documents from high to low according to the confidence of the classification results of these documents, and will give these documents a positive offset bias in subsequent calculations. + , the offset size is sorted from high to low according to the order of the documents, that is, the model will give more attention to this part of the documents;

[0124] For the classification result y rel =False and y cor = False, the document will be regarded as an interference document that is completely irrelevant to the query. When it is subsequently sent to the question-answering module, the document will be marked as an error document for the large model to judge on its own. At the same time, a negative offset bias will be given to this part of the document in subsequent calculations. - , the offset size is sorted from high to low according to the order of the documents, that is, the model will give less attention to these documents;

[0125] For the classification result y rel =True or y cor= True, the document is considered partially relevant to the query. This means that some of the content in these documents may be helpful in answering the user's query, but they may also contain some distracting or erroneous information. For these documents, a large language model is used to decompose them into multiple atomic facts. Each atomic fact is then re-entered into the classification module to check its relevance. Irrelevant facts are removed from the list. Otherwise, the trained inference model is used to determine whether the errors in the document can be corrected based on the model's knowledge. Atomic facts that cannot be corrected are removed from the list to minimize their impact on subsequent output.

[0126] Specifically, for document D, generate multiple atomic facts {f1,f2,...,f n}, where f i Represents the i-th atomic fact generated by document D, satisfying the following formula:

[0127] {f1,f2,...,f n}=Decompose(D)

[0128]

[0129] The Decompose() function represents the deconstruction of document D. Represents a merger of atomic facts.

[0130] S08, traverse each atomic fact f i , using the trained inference model Determine the correctness of each atomic fact and output the correct document after the model is corrected.

[0131] In this step, you can use the chain of thought reasoning method to let the large model think about the correctness of the fact. If the fact is wrong, then determine whether the fact can be corrected by the large model through its own parameterized knowledge. If it can, drive the large model to correct it with its own parameterized knowledge, otherwise discard the document. The main principle of the CoT (Chain of Thought) chain of thought is to gradually decompose complex problems and use the chain reasoning method to enable the large language model to think and summarize step by step. The core of this method is to break down complex problems into multiple simple sub-problems, and through step-by-step reasoning and solving, each step is based on the result of the previous step, and finally get a solution to the entire problem. This step-by-step reasoning process can help the large language model better understand the context and details of the problem, avoid errors in the intermediate steps, and thus improve the accuracy and reliability of the overall answer. The details are as follows:

[0132]

[0133] where f' iFor the revised atomic facts, This is the inference model trained in S03. is the revised set of atomic facts.

[0134] S09, the atomic facts after processing S08 And the original query q is fed into the big model Let it answer the question q based on these processed atomic facts and generate the corresponding response answer ans, as follows:

[0135]

[0136] Where q is the original query, is the revised set of atomic facts, Acting as an answer generator for arbitrarily large models.

[0137] S10, the response answer ans generated by S09 is sent to the adaptive classification module to check the correctness of the response answer ans to prevent the generation of untrue response answers due to hallucination of the large model. Specifically:

[0138]

[0139] If the adaptive classification module determines that ans is an incorrect response, it takes ans as input and re-executes step S07 and subsequent steps S08, S09, and S10, repeating the cycle until ans is determined to be a correct response in S10;

[0140] If the adaptive classification module determines that ans is the correct response answer, the response answer ans is returned to the user.

[0141] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning, characterized in that: The following steps are involved: (1) Collect domain-specific knowledge and build a vector database of domain-specific knowledge; (2) Input relevant domain-specific questions. The vector retrieval module calculates the correlation between the two dimensions, semantics and word frequency, from the vector database and finds the highest-scoring matching knowledge documents after mixing them with parameter weights. (3) Input the domain-specific questions and the matched knowledge documents into the adaptive classification module. Based on the document classification model deconstructed by the single decoder, the semantic attribute category corresponding to each knowledge document and the relevance label and correctness label corresponding to the classification are obtained; (4) Input the knowledge document and the corresponding semantic attribute categories into the autonomous reasoning module, use the large model to eliminate the uncertainties and errors in the knowledge document, and infer the enhanced knowledge; (5) The enhanced knowledge is fed into the intelligent question-answering module, which uses the large language model to generate responses to questions; (6) The response answer generated in step (5) is used as a knowledge document and is re-sent into the adaptive classification module to check the semantic attribute category of the response answer; If the response answer is not judged as the correct answer by the adaptive classification module, the response answer is used as input and step (4) and subsequent steps are re-executed; if the response answer is judged as the correct answer by the adaptive classification module, the response answer is displayed to the user.

2. The retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to claim 1 is characterized in that: The specific process of step (2) is: (2-1) Convert the input domain-specific problem q into a set of vectors e q ; (2-2) The vector e of the domain-specific problem q Vector Database The domain knowledge vector e in D Correspondingly, calculate the semantic similarity of the two vectors mean (q, D) and word frequency similarity word (q, D), after calculation, the text similarity result of the document is obtained by mixing the weight parameters; (2-3) Re-rank the documents from high to low according to the results of text similarity calculation, with the documents at the top representing higher similarity; (2-4) Take the domain knowledge with the highest similarity as the result of vector retrieval and recall.

3. The retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to claim 2 is characterized in that: In step (2-2), the text similarity calculation formula of the document is as follows: score i =α×similarity mean (q,D i )+β×similarity word (q,D i ) Among them, q is the domain question of the query, and D is the knowledge document in the vector database; similarity mean (q,D) function represents the semantic similarity between q and D, similarity word (q, D) represents the word frequency similarity between q and D; α and β are weight parameters; Semantic similarity mean The calculation formula for (q,D) is as follows: Among them, ||e q || represents the vector e q The modulus length of ||e D || represents the vector e D Length of the module; Word frequency similarity word The calculation formula for (q,D) is as follows: Where n is the total number of words, f(q i ,D) represents the vocabulary q i The number of occurrences in the knowledge document D, k and b are the hyperparameters for adjustment, length avg is the average length of knowledge documents.

4. The retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to claim 1 is characterized in that: In step (3), the working process of the adaptive classification module is: (3-1) The input is the domain-specific question q and the knowledge document D. First, the embedding layer is calculated to generate a vector representation E q and E D , and after generating the vector, the rotation position encoding is calculated and added to the vector embedding as follows: E q =Embedding(q)+P rot (q) E D =Embedding(D)+P rot (D) Among them, E q and E D It represents the vector representation of the domain-specific question q and the knowledge document D after the embedding layer. Embedding() represents the word embedding operation. rot () indicates the rotation position encoding calculation; (3-2) Vector representation E q and E D Each encoder layer calculates the bidirectional attention score of the input through the attention mechanism, and then adapts it to the current classification dimension through adaptive task-aware normalization. Adaptive task-aware normalization builds a lightweight fully connected network using a single-layer MLP to generate independent normalization parameters γ for classification tasks of different dimensions. task and β task , the input of MLP is the global pooling result output by the Encoder part, and the calculation formula is as follows: x=concat(Pooling(H),t task ) [γ task ,β task ]=MLP(x) Among them, Pooling() is the global pooling result, H is the attention score vector output by the encoder, and t task is a one-hot vector used to determine whether the current classification dimension is correctness or relevance; the final normalized calculation formula is as follows: (3-3) At the top level of the encoder, bidirectional attention is used to decouple and separate the unimodal features and cross-modal features of the domain question q and the knowledge document D. First, the domain question q and the knowledge document D calculate their own self-attention results, and then the interaction between the domain question q and the knowledge document D is modeled through bidirectional attention. Finally, the classification result is obtained through two different gating networks. The calculation formula is as follows: D att =Self-Attention(x D ),Q att =Self-Attention(x q ) I att =Bi-Attention(D att ,Q att ) Among them, Self-Attention() represents self-attention calculation, Bi-Attention() represents bidirectional attention calculation; Q att and D att Represents the domain problem vector x q and knowledge document vector x D The self-attention calculation results, I att Q att and D att The calculation results of the two-way attention; Finally, the attention results of the three levels are concatenated and assigned to different classification dimensions through the task gating network: F rel =Gate rel ([D att ;Q att ;I att ])∈{True,False} F cor =Gate cor ([D att ;Q att ;I att ])∈{True,False} Among them, Gate() represents the gated network, which is a lightweight fully connected layer; F rel Represents the classification result of the correlation dimension, F cor Represents the classification results of the correctness dimension.

5. The retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to claim 1 is characterized in that: In step (3), the training process of the adaptive classification module is as follows: Collect query documents from various fields and positive samples (q, D, y rel ,y cor ), and construct a dataset D for training the classification model r * , dataset D r * The negative samples in are obtained by random sampling; where q represents the query, D represents the document, and y rel represents the category of document D divided according to its relevance to query q, y cor Represents the categories of document D divided according to the correctness of query q; In the classification dataset D r * The classification model is trained above, and the training goal is to minimize the multi-task joint cross entropy loss function: Among them, w rel and w cor is the loss function score weight of the relevance dimension and the correctness dimension, y rel and y cor They are the relevance dimension and correctness dimension labels actually corresponding to the document, and They are the document relevance dimension and correctness dimension labels predicted by the model respectively.

6. The retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to claim 1 is characterized in that: The specific process of step (4) is: (4-1) For label y rel =True and y cor = True, the autonomous reasoning module regards the document as the answer document. When it is subsequently fed into the intelligent question-answering module, the autonomous reasoning module will re-sort the documents from high to low confidence based on the confidence of the classification results of these documents, and give them a positive weight at the attention level; (4-2) For label y rel =False and y cor = False, the autonomous reasoning module regards the document as an interference document. When it is subsequently sent to the intelligent question-answering module, the document is marked as an error document for the large model to judge on its own and is given a negative weight at the attention level. (4-3) For label y rel =True or y cor = True, the autonomous reasoning module regards the document as an uncertain document and trains the autonomous reasoning model to decompose the knowledge document D into multiple atomic facts {f1,f2,...,f n }, for each atomic fact f i The fact is re-sent to the adaptive classification module to check its relevance; facts that are still irrelevant are removed from the list, otherwise the trained autonomous reasoning model is used to determine whether the error in the document can be corrected based on the knowledge of the model itself; atomic facts that cannot be corrected will be removed from the list to reduce the impact on subsequent outputs.

7. The retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to claim 6 is characterized in that: In step (4), the training process of the autonomous reasoning module is as follows: Collect atomic fact question-answer pairs (f, t) that meet the error correction requirements, where f represents the atomic fact and t represents the corrected true fact. Construct a dataset D for training autonomous reasoning models * For each positive sample (f, t), the model uses prompt word engineering to make the model think about the authenticity of fact f when inputting f, and attempts to output the corrected result t' after the model thinks about it; at the same time, negative samples are randomly sampled, and for negative samples, the model outputs a special token when inputting to indicate that the model cannot correct this knowledge; In the autonomous reasoning dataset D * The large language model LLM is trained above, and the training goal is to minimize the cross entropy loss function, that is: Among them, y ij is the category j, y corresponding to the i-th token ij ' is the probability that the i-th token is predicted to be of category j.

8. The retrieval enhancement generation optimization method based on adaptive classification and autonomous reasoning according to claim 1 is characterized in that: The specific process of step (6) is as follows: (6-1) The generated response answer ans is re-sent to the adaptive classification module to obtain the relevance and correctness label of the response answer ans relative to the question q; (6-2) If the classification result is y rel =False or y cor = False, it means that the generated answer is wrong. In this case, the response answer is sent back to the autonomous reasoning module ans for deconstruction. The deconstructed answer fragment is sent back to the adaptive classification module to check the relevance and correctness of the fragment. Any irrelevant or incorrect answer fragment will be removed. (6-3) Combine the remaining answer fragments with the original context to regenerate a new response answer; repeat (6-1) to (6-3) until the correct response answer is generated.

9. A retrieval enhancement generation optimization device based on adaptive classification and autonomous reasoning, characterized in that: The system comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, the system is used to implement the search enhancement generation optimization method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Knowledge retrieval enhancement generation method and system based on large language model

    CN118394890A

  • Retrieval enhancement generation method and device based on large model and medium

    CN119312917A

Cited By

  • Knowledge reasoning method and device for integrated circuit wafer process and medium

    CN120744140A

  • Request processing method, electronic equipment, storage medium and program product

    CN121255160A

  • Retrieval enhancement generation method suitable for building structure field

    CN121301589A