Retrieval enhancement generation method based on multistage semantics

By introducing multi-level semantic computing and professional field knowledge base in the RAG system, the problem of poor retrieval results in professional field is solved, more accurate retrieval and generation results are achieved, and the advantages of RAG and LLM are fully utilized.

CN120011535APending Publication Date: 2025-05-16SHANGHAI MEDIWORKS PRECISION INSTR CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202411986263.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing search enhancement generation (RAG) system has poor retrieval effect in the professional field and cannot fully combine the advantages of RAG and large language model (LLM), resulting in poor generation results.

Method used

A search enhancement generation method based on multi-level semantics is adopted, and a knowledge base in professional fields is constructed, and a semantic vector model and keyword embedding model is used to vectorize and extract information, calculate the semantic similarity between word level and sentence level, improve the search accuracy, and decide whether to use RAG external knowledge or LLM internal knowledge based on the search whitelist.

Benefits of technology

It improves the search accuracy in the professional field, fully combines the advantages of RAG and LLM, and the generated results are more accurate and reliable.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011535A_ABST
    Figure CN120011535A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-level semantic-based retrieval enhancement generation method, which comprises the following steps of: 1, constructing a knowledge base by using corpora in a professional field based on LLM (Logistics Language Model); 2, obtaining a to-be-processed text, and obtaining a vector representation and an information extraction result of the to-be-processed text; processing the to-be-processed text by using a key information extraction algorithm to obtain a keyword list; 3, according to the vector representation of the to-be-processed text and an information extraction result, calculating the similarity between information in the knowledge base and the to-be-processed text by using multistage semantics to obtain related original text block content and a corresponding document source; step 4, searching a white list selection answer mainly based on RAG external knowledge or LLM self knowledge; according to the method, the problems that the RAG system is difficult to effectively retrieve the vertical field with higher specialty and cannot ensure the efficient use of the priori knowledge of the LLM itself are solved, the retrieval accuracy of the RAG for the professional field is improved, and the advantages of the RAG and the LLM are fully combined and played.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a retrieval enhancement generation technology, and in particular to a retrieval enhancement generation method based on multi-level semantics. Background Art

[0002] With the development of artificial intelligence and generative technology, intelligent dialogue systems based on large language models (LLM) have been widely used in e-commerce customer service, government Q&A, emotional companionship, and medical consultation. Due to the hallucination problem of large models, a large number of intelligent dialogue systems in vertical fields have adopted retrieval-augmented generation (RAG) technology. RAG mainly retrieves fragments related to the user input text from an external knowledge base, and allows LLM to use the retrieved relevant information and user input information to generate the final result.

[0003] Existing technical features:

[0004] The existing RAG system has a good retrieval effect on general knowledge, but when applied to vertical fields such as finance, medicine and law where professional terms are very different from general knowledge, the retrieval effect of RAG will be very poor. The inaccurate external knowledge retrieved will further lead to poor LLM generation results. At the same time, most of the existing RAG systems only use external knowledge, which seriously limits the generation ability of LLM itself and cannot fully combine the advantages of RAG and LLM.

[0005] Specifically, the existing RAG system finds it difficult to effectively search highly professional vertical fields, and cannot guarantee the efficient use of LLM's own prior knowledge.

[0006] from Figure 1 From the schematic diagram of a common retrieval enhancement generation method shown in FIG. 1 , we can see that: (1) in a common RAG system, the vector representation of the segmented text is directly used to calculate the similarity with the encoding vector of the question, and the top-k related text blocks are returned accordingly; (2) in addition, after obtaining the top-k text blocks, they are directly used as context information and combined with a prompt template 1 to construct a prompt text. The content of the prompt template 1 used is similar to: "Use context information to answer the following question, \nContext information: {text block} \nQuestion: {user question}"

[0007] For general fields, step (1) is feasible enough. However, for professional fields such as medicine, for example, when a user asks a question like "Tell me about diabetic retinopathy", searching only by the vector similarity between the text block and the question will result in the text block with the "diabetes" field having a higher similarity to the question than the text block with the "diabetic retinopathy" field, resulting in the retrieved text block fragment not matching the question. Therefore, searching by similarity based on vector index alone cannot meet the semantic matching requirements of professional fields such as medicine.

[0008] At the same time, regarding step (1), since the text block contains a lot of invalid information, directly vectorizing the entire text block will easily cause important information to be affected by the invalid information during the encoding process, which will lead to poor results when calculating similarity. Therefore, directly vectorizing the entire original text block without processing will also lead to poor retrieval results.

[0009] At the same time, regarding step (1), since it is necessary to balance the size of the text block granularity and the retrieval speed, the length of a single text block may be around 500 words. Therefore, a single text block is likely to contain more than one topic at the same time. If this is not restricted when calculating the similarity, the content of the retrieved text block will be inconsistent, which will affect the generation effect of LLM.

[0010] At the same time, regarding step (1), existing vectorization models are often trained based on large general corpora, so the effect of text vectorization on professional fields is naturally inaccurate. If such vectorization models are directly fine-tuned to adapt to professional fields, a large amount of resources will be consumed and the generalization ability of the vectorization model may also be damaged. Therefore, it is unreasonable to directly fine-tune the existing vectorization models.

[0011] In addition, for the user's question, if the content involved is in the knowledge base, then step (2) is feasible enough. If it involves content outside the knowledge base, for example, for an ophthalmology RAG system, if the user asks "What are the precautions and suggestions for fitting glasses?" If we only rely on the RAG knowledge base, since there is very little content related to glasses fitting in the current medical textbooks, papers and other corpora on ophthalmology, we cannot answer this question well and need to rely on the generation capabilities of LLM itself. Summary of the invention

[0012] In view of the problems that the RAG system has difficulty in effectively searching highly professional vertical fields and cannot ensure the efficient use of LLM's own prior knowledge, a retrieval enhancement generation method based on multi-level semantics is proposed to improve the accuracy of RAG's retrieval in professional fields and fully combine and give play to the advantages of both RAG and LLM.

[0013] The technical solution of the present invention is:

[0014] A retrieval enhancement generation method based on multi-level semantics includes the following steps:

[0015] Step 1: Based on LLM, use the corpus of professional fields to build a knowledge base; the knowledge base includes a summary vector library and a keyword embedding library, which are stored in blocks to obtain the original text block content and the corresponding document source, as well as summary vector association and keyword association information;

[0016] Step 2: Get the text to be processed, and get its vector representation and information extraction results; the text to be processed is the question input by the user. When it involves professional fields, use the semantic vector model to vectorize the text to be processed to obtain the encoding vector of the question; use the key information extraction algorithm to process the text to be processed and obtain a keyword list containing no less than 1 element;

[0017] Step 3: Based on the vector representation and information extraction results of the text to be processed, use multi-level semantics to calculate the similarity between the information in the knowledge base and the text to be processed, and obtain the relevant original text block content and the corresponding document source;

[0018] Step 4: Choose whether the answer is based on RAG external knowledge or LLM's own knowledge, depending on whether the source of the relevant reference text is from the retrieval whitelist. If the document source is in the retrieval whitelist, RAG's external knowledge will be used as the main answer. If the document source is not in the retrieval whitelist, LLM will be used to generate a reply first, supplemented by external knowledge retrieved by RAG, and then LLM will be used to improve the reply based on the existing RAG external knowledge.

[0019] Furthermore, in step 1, the corpus of a professional field includes documents in the field that can be parsed to obtain text content; when parsing these corpus texts, the texts need to be divided into pieces instead of parsing all the texts at once;

[0020] For the summary vector library, each text block is sent to the prompt word prompt template 2 to obtain the corresponding prompt text, and then sent to the LLM to obtain the summary text. The summary text is vectorized using the semantic vector model, i.e., the embedding model 1, to obtain the corresponding summary vector result;

[0021] For the keyword library, each text block is sent to the prompt word prompt template 3 to obtain the corresponding prompt text, and then sent to the LLM to obtain the keyword list. Each word in the list is vectorized using the word embedding model, that is, the embedding model 2, to obtain the corresponding word embedding list result;

[0022] The content of prompt template 2 is as follows: "Summarize the text below, text content: {text block}"; where "{text block}" is a detailed text paragraph;

[0023] The content of prompt template 3 is as follows: "Extract keywords from the following text and optimize it for retrieval, text content: {text block}"; where "{text block}" is a detailed text paragraph;

[0024] For embedding model 1, it is the m3e or bge general vector model;

[0025] For embedding model 2, it is a word2vec or fastText word embedding model trained based on professional domain corpus.

[0026] Furthermore, in step 3, for each target keyword in the keyword list obtained by performing the key information extraction algorithm on the text to be processed, the similarity algorithm 1 is used to calculate the word-level semantic similarity between the target keyword and each keyword in the keyword list corresponding to each text block in the keyword library;

[0027] For the encoding vector of the question obtained by vectorizing the text to be processed, use similarity algorithm 2 to calculate the sentence-level semantic similarity between it and the summary vector corresponding to each text block in the summary vector library;

[0028] For similarity algorithm 1, cosine similarity is used;

[0029] For similarity algorithm 2, cosine similarity is used;

[0030] For the word-level semantic similarity and sentence-level semantic similarity obtained above for the text to be processed,

[0031] After obtaining the semantic similarity measurement results, the relevance is sorted accordingly, and the semantic similarity values ​​are used to sort from large to small, and the top N similarity Top-N groups of text blocks and information containing source documents, that is, the original text block content and the corresponding document source, are returned.

[0032] Further, in step 4, regarding the retrieval whitelist, the list of files specified by the user according to his / her needs;

[0033] For the document source corresponding to the relevant reference text retrieved from the pending question, if the document source is in the retrieval whitelist, the external knowledge of RAG is mainly used, prompt template 1 is used to construct the prompt text and sent to LLM to generate the final answer; if the document source is not in the retrieval whitelist, LLM is first used to generate a reply, supplemented by the external knowledge retrieved by RAG, and then LLM is used to improve the reply based on the existing RAG external knowledge, and prompt template 4 is used to construct the prompt text and sent to LLM to generate the final answer;

[0034] The content of prompt template 1 is as follows: "Answer the following question using context information, \nContext information: {text block} \nQuestion: {user question}}";

[0035] The content of prompt template 4 is as follows: "User's original question: {User's question}\nThis is the answer that currently exists: {LLM's first reply}\nThere is an opportunity to improve the existing answer based on the following context (only if needed): {External knowledge retrieved by RAG}; If the context is not useful, return the current answer;".

[0036] Furthermore, in step 3, the multi-level semantic similarity between the text block and all text blocks in the knowledge base is calculated according to the following mathematical model:

[0037] Sim i =(1-a)Word i +aSen i

[0038]

[0039] Among them, Sim i Indicates the semantic similarity between the text to be processed and the i-th text block; a is used to control the weight of the word-level similarity and sentence-level similarity in the overall similarity calculation, which is a decimal between [0,1]; Word i Indicates the word-level semantic similarity between the text to be processed and the i-th text block; Sen i represents the sentence-level semantic similarity between the text to be processed and the i-th text block; T represents the length of the keyword list extracted from the text to be processed, C represents the length of the keyword list corresponding to the i-th text block in the knowledge base, word tc Represents the word-level semantic similarity between the tth target keyword and the cth search keyword; Ref i Represents the semantic similarity between all keywords in the i-th text block; word klrepresents the word-level semantic similarity between any two different words in the same set, k represents the kth keyword, l represents the lth keyword, and the range of l and k are both 1 to C; Target represents the semantic similarity between all keywords in the text to be processed; it is assumed here that the text to be processed comes from the user, so it has a high semantic similarity. Therefore, the consistency of the text to be processed is taken as the benchmark, and the semantic similarity of the knowledge base text block relative to the semantic similarity of the text to be processed is used as the consistency weight of the word-level semantic similarity of the text block. The expression is

[0040] Preferably, in step 2, the key information extraction algorithm is a named entity recognition algorithm obtained based on professional domain corpus training, and also includes a keyword extraction algorithm based on professional domain corpus training, a topic extraction algorithm based on professional domain corpus training, and a named entity recognition algorithm, a keyword extraction algorithm and a topic extraction algorithm based on professional domain corpus training are used simultaneously.

[0041] Preferably, the LLM model includes an open source chatGLM class model and a closed source chatGPT model.

[0042] Preferably, the text segmentation strategy uses a delimiter-based segmentation strategy while retaining complete sentences and containing some overlapping content. The segmentation result will try to meet the segmentation strategy of the set block size upper limit. You can also choose to directly segment according to the block size or other segmentation strategies; the block size is generally in the range of [300,500].

[0043] Preferably, for similarity algorithm 1, inner product and Euclidean distance function can be selected; for similarity algorithm 2, inner product and Euclidean distance function can be selected.

[0044] Preferably, considering the limitation of LLM input length, N generally takes a value not exceeding 3.

[0045] The beneficial effects of the present invention are:

[0046] By using multi-level semantic similarity at the word level and sentence level, the semantic information contained in the professional domain corpus and the problem text to be processed is fully utilized to improve the accuracy of the text blocks retrieved in the professional domain corpus:

[0047] 1. When calculating semantic similarity at the word level, the internal consistency of the keyword list was used to measure the impact of questions that may contain more than one topic in a text block;

[0048] 2. When calculating the semantic similarity at the sentence level, the text block summary is used to alleviate the problem of irrelevant information contained in the text block;

[0049] 3. When performing vectorization and key information extraction, word embedding models and key information extraction models trained with professional domain corpora are used to make up for the shortcomings of general vector models and embed professional domain information.

[0050] By setting up a search whitelist, we can determine the source of the relevant text blocks and decide whether to base our search on RAG knowledge or LLM prior knowledge, so as to fully combine and give play to the respective advantages of RAG external knowledge and LLM internal prior knowledge. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A schematic diagram of a common search enhancement generation method in the background art;

[0052] Figure 2 A flowchart of a retrieval enhancement generation method based on multi-level semantics of the present invention;

[0053] Figure 3 A schematic diagram of a retrieval enhancement generation method based on multi-level semantics according to the present invention. DETAILED DESCRIPTION

[0054] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.

[0055] The present invention proposes a retrieval enhancement generation method based on multi-level semantics, such as Figure 2 , 3 As shown, its implementation includes the following steps:

[0056] Step 1: Based on LLM, use the corpus of professional fields to build a knowledge base.

[0057] Large Language Model (LLM) is a natural language processing model based on deep learning, usually composed of hundreds of millions to hundreds of billions of parameters. They can generate and understand natural language by training on large-scale text data. Common LLM models include chatGPT, chatGLM, Wenxin Yiyan, Tongyi Qianwen. The chatGLM model is used in this invention.

[0058] The corpus in a professional field includes any document that can be parsed to obtain text content, such as papers, books, patents, web pages, and textbooks in the field. When parsing these corpus texts, it is necessary to divide the text into pieces instead of parsing all the texts at once.

[0059] The text segmentation strategy uses line breaks and other delimiters to segment text, while retaining complete sentences and containing some overlapping content. The segmentation result will try to meet the segmentation strategy of the set block size limit; you can also choose to segment directly according to the block size or other segmentation strategies. There is no limit on the segmentation strategy, but the block size is generally in the range of [300,500].

[0060] For the knowledge base, it includes a summary vector library and a keyword embedding library. In addition, it also stores the original text block content obtained by segmentation and the corresponding document source, as well as summary vector association and keyword association information.

[0061] For the summary vector library, each text block is sent to the prompt template 2 to obtain the corresponding prompt text, and then sent to the LLM to obtain the summary text. The summary text is vectorized using the semantic vector model (i.e., the embedding model 1) to obtain the corresponding summary vector result;

[0062] For the keyword library, each text block is sent to the prompt template 3 to obtain the corresponding prompt text, and then sent to the LLM to obtain the keyword list. Each word in the list is vectorized using the word embedding model (i.e., embedding model 2) to obtain the corresponding word embedding list result.

[0063] The content of prompt template 2 is as follows: "Summarize the text below, text content: {text block}". "{text block}" is a detailed text paragraph.

[0064] The content of prompt template 3 is as follows: "Extract keywords from the following text and optimize it for retrieval, text content: {text block}". "{text block}" is a detailed text paragraph.

[0065] For embedding model 1, it is a general vector model such as m3e (Moka Massive Mixed Embedding) or bge (BAAI general embedding);

[0066] For embedding model 2, it is a word embedding model such as word2vec or fastText trained based on professional domain corpus.

[0067] For the LLM model used, the open source chatGLM (General Language Model, GLM) model is adopted. Optionally, the closed source chatGPT (Generative Pre-trained Transformer, GPT) and other models can also be used.

[0068] Step 2: Get the text to be processed, and obtain its vector representation and information extraction results.

[0069] The text to be processed is a question input by the user, and its content may be related to a professional field, such as medical, legal or financial, etc. Among them, the professional field refers to medical, legal or financial fields that are different from daily professional fields.

[0070] In the field of natural language processing (NLP), an encoding vector (also called an embedding vector or feature vector) refers to the result of converting text (such as words, sentences, or documents) into a digital representation (usually a high-dimensional vector). Such vectors are usually generated by some pre-trained model or algorithm, with the purpose of capturing the semantic information of the text so that the machine can effectively understand and process the language. The core of the encoding vector is that it can represent the semantics of the text in a continuous and dense way, so that it can be used for various subsequent tasks, such as text classification, clustering, question answering, information retrieval, etc.

[0071] Use the semantic vector model (embedding model 1) to vectorize the text to be processed to obtain the encoding vector of the question;

[0072] In the field of natural language processing (NLP) and information retrieval, a keyword list usually refers to a collection of words or phrases with important meanings extracted from a text. These keywords are usually words closely related to the subject, content or key information of the text, and are used to summarize the core meaning of the text or help subsequent processing (such as retrieval, classification, summarization, etc.).

[0073] The key information extraction algorithm is used to process the text to be processed to obtain a keyword list containing no less than 1 element.

[0074] The key information extraction algorithm is a named entity recognition algorithm trained based on professional domain corpus. Optionally, it can be a keyword extraction algorithm trained based on professional domain corpus, or a topic extraction algorithm trained based on professional domain corpus. Alternatively, multiple key information extraction algorithms such as a named entity recognition algorithm trained based on professional domain corpus, a keyword extraction algorithm, and a topic extraction algorithm can be used simultaneously.

[0075] Step 3: Based on the vector representation and information extraction results of the text to be processed, use the similarity between the information in the multi-level semantic computing knowledge base (abstract vector library and keyword library) and the text to be processed to obtain the relevant original text block content and the corresponding document source.

[0076] For the keyword list obtained by performing the key information extraction algorithm on the processed text, for each target keyword, similarity algorithm 1 is used to calculate the word-level semantic similarity between it and each keyword in the keyword list corresponding to each text block in the keyword library (similar to the keywords in the keyword library).

[0077] For the encoding vector of the question obtained by vectorizing the text to be processed, similarity algorithm 2 is used to calculate the sentence-level semantic similarity between it and the summary vector corresponding to each text block in the summary vector library.

[0078] For similarity algorithm 1, cosine similarity is used, and optionally, distance functions such as inner product and Euclidean distance are also used; for similarity algorithm 2, cosine similarity is used, and optionally, distance functions such as inner product and Euclidean distance are also used, and they may be different.

[0079] For the word-level semantic similarity and sentence-level semantic similarity obtained above for the text to be processed, the multi-level semantic similarity with all text blocks in the knowledge base is calculated according to the following mathematical model:

[0080] Sim i =(1-a)Word i +aSen i

[0081]

[0082] Among them, Sim i Indicates the semantic similarity between the text to be processed and the i-th text block; a is used to control the weight of the word-level similarity and sentence-level similarity in the overall similarity calculation, which is a decimal between [0,1]; Word i Indicates the word-level semantic similarity between the text to be processed and the i-th text block; Sen i represents the sentence-level semantic similarity between the text to be processed and the i-th text block. T represents the length of the keyword list extracted from the text to be processed, C represents the length of the keyword list corresponding to the i-th text block in the knowledge base, and word tc Represents the word-level semantic similarity between the tth target keyword and the cth search keyword; Ref i Represents the semantic similarity between all keywords in the i-th text block; word klrepresents the word-level semantic similarity between any two different words in the same set, k represents the kth keyword, l represents the lth keyword, and the range of l and k are both 1 to C; Target represents the semantic similarity between all keywords in the text to be processed. Here we assume that the text to be processed comes from the user, so it has a high semantic similarity. Therefore, we take the consistency of the text to be processed as the benchmark, and use the semantic similarity of the knowledge base text block relative to the semantic similarity of the text to be processed as the consistency weight of the word-level semantic similarity of the text block. The expression is

[0083] After obtaining the semantic similarity measurement results, the relevance is sorted accordingly, and the semantic similarity values ​​are used to sort from large to small, and the top N similarity (Top-N) groups of text blocks and information containing source documents (original text block content and corresponding document sources) are returned. Considering the limitation of LLM input length, N is generally not more than 3.

[0084] Step 4: Depending on whether the source of the relevant reference text is from the retrieval whitelist, choose whether the answer is based on RAG external knowledge or LLM's own knowledge.

[0085] Regarding the retrieval whitelist, it is a list of files specified by the user according to his or her own needs, such as: product manuals, rules and regulations, documents in a certain professional field, etc.

[0086] For the document source corresponding to the relevant reference text retrieved from the question to be processed, if the document source is in the retrieval whitelist, the external knowledge of RAG is mainly used, and prompt template 1 is used to construct the prompt text and send it to LLM to generate the final answer; if the document source is not in the retrieval whitelist, LLM is first used to generate a reply, supplemented by the external knowledge retrieved by RAG, and then LLM is used to improve the reply based on the existing RAG external knowledge, and prompt template 4 is used to construct the prompt text and send it to LLM to generate the final answer.

[0087] The content of prompt template 1 is as follows: "Answer the following question using context information, \nContext information: {text block} \nQuestion: {user question}}".

[0088] The content of prompt template 4 is as follows: "Original question of the user: {User's question}\nHere is the answer that currently exists: {LLM's first reply}\nThere is an opportunity to improve the existing answer (only if needed) based on the following context: {External knowledge retrieved by RAG}. If the context is not useful, return the current answer.".

[0089] References:

[0090] 1. CN 117573815 B, a retrieval enhancement generation method based on vector similarity matching optimization;

[0091] 2. CN 118227771 B, a knowledge question answering processing method and device;

[0092] 3. CN 117669717 A, a large model question answering method, device, equipment and medium based on knowledge enhancement;

[0093] 4.CN 118394946 B, a retrieval enhancement generation method and system based on multi-view clustering.

[0094] The above-mentioned embodiment only expresses one implementation mode of the present invention, and its description is relatively specific and detailed, but it cannot be understood as limiting the scope of the invention patent. It should be pointed out that for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be based on the attached claims.

Claims

1. A retrieval enhancement generation method based on multi-level semantics, characterized in that: The following steps are involved: Step 1: Based on LLM, use the corpus of professional fields to build a knowledge base; the knowledge base includes a summary vector library and a keyword embedding library, which are stored in blocks to obtain the original text block content and the corresponding document source, as well as summary vector association and keyword association information; Step 2: Get the text to be processed, and get its vector representation and information extraction results; the text to be processed is the question input by the user. When it involves professional fields, use the semantic vector model to vectorize the text to be processed to obtain the encoding vector of the question; use the key information extraction algorithm to process the text to be processed and obtain a keyword list containing no less than 1 element; Step 3: Based on the vector representation and information extraction results of the text to be processed, use multi-level semantics to calculate the similarity between the information in the knowledge base and the text to be processed, and obtain the relevant original text block content and the corresponding document source; Step 4: Choose whether the answer is based on RAG external knowledge or LLM's own knowledge, depending on whether the source of the relevant reference text is from the retrieval whitelist. If the document source is in the retrieval whitelist, RAG's external knowledge will be used as the main answer. If the document source is not in the retrieval whitelist, LLM will be used to generate a reply first, supplemented by external knowledge retrieved by RAG, and then LLM will be used to improve the reply based on the existing RAG external knowledge.

2. The multi-level semantics-based retrieval enhancement generation method according to claim 1, characterized in that: In step 1, the corpus of a professional field includes documents in the field that can be parsed to obtain text content; when parsing these corpus texts, the texts need to be divided into pieces; For the summary vector library, each text block is sent to the prompt word prompt template 2 to obtain the corresponding prompt text, and then sent to the LLM to obtain the summary text. The summary text is vectorized using the semantic vector model, i.e., the embedding model 1, to obtain the corresponding summary vector result; For the keyword library, each text block is sent to the prompt word prompt template 3 to obtain the corresponding prompt text, and then sent to the LLM to obtain the keyword list. Each word in the list is vectorized using the word embedding model, that is, the embedding model 2, to obtain the corresponding word embedding list result; The content of prompt template 2 is as follows: "Summarize the text below, text content: {text block}"; where "{text block}" is a detailed text paragraph; The content of prompt template 3 is as follows: "Extract keywords from the following text and optimize it for retrieval, text content: {text block}"; where "{text block}" is a detailed text paragraph; For embedding model 1, it is the m3e or bge general vector model; For embedding model 2, it is a word2vec or fastText word embedding model trained based on professional domain corpus.

3. The multi-level semantics-based retrieval enhancement generation method according to claim 2, characterized in that: In step 3, for each target keyword in the keyword list obtained by performing a key information extraction algorithm on the text to be processed, the similarity algorithm 1 is used to calculate the word-level semantic similarity between the target keyword and each keyword in the keyword list corresponding to each text block in the keyword library; For the encoding vector of the question obtained by vectorizing the text to be processed, use similarity algorithm 2 to calculate the sentence-level semantic similarity between it and the summary vector corresponding to each text block in the summary vector library; For similarity algorithm 1, cosine similarity is used; For similarity algorithm 2, cosine similarity is used; For the word-level semantic similarity and sentence-level semantic similarity obtained above for the text to be processed, After obtaining the semantic similarity measurement results, the relevance is sorted accordingly, and the semantic similarity values ​​are used to sort from large to small, and the top N similarity Top-N groups of text blocks and information containing source documents, that is, the original text block content and the corresponding document source, are returned.

4. The multi-level semantics-based retrieval enhancement generation method according to claim 3, characterized in that: In step 4, regarding the search whitelist, the file list is specified by the user according to his or her needs; For the document source corresponding to the relevant reference text retrieved from the question to be processed, if the document source is in the retrieval whitelist, the external knowledge of RAG is mainly used, prompt template 1 is used to construct the prompt text and sent to LLM to generate the final answer; If the document source is not in the search whitelist, LLM is used to generate a response first, supplemented by the external knowledge retrieved by RAG, and then LLM is used to improve the response based on the existing RAG external knowledge, and prompt template 4 is used to construct the prompt text and sent to LLM to generate the final answer; The content of prompt template 1 is as follows: "Use context information to answer the following questions, \nContext information: {text block} \nQuestion: {user question}}"; The content of prompt template 4 is as follows: "User's original question: {User's question}\nThis is the answer that currently exists: {LLM's first reply}\nThere is an opportunity to improve the existing answer (only if needed) based on the following context: {External knowledge retrieved by RAG}; If the context is not useful, return the current answer;".

5. The multi-level semantics-based retrieval enhancement generation method according to claim 3, characterized in that: In step 3, the multi-level semantic similarity between the text block and all text blocks in the knowledge base is calculated according to the following mathematical model: Sim i 1(1-a)Word i +aSen i Among them, Sim i Indicates the semantic similarity between the text to be processed and the i-th text block; a is used to control the weight of the word-level similarity and sentence-level similarity in the overall similarity calculation, which is a decimal between [0,1]; Word i Indicates the word-level semantic similarity between the text to be processed and the i-th text block; Sen i represents the sentence-level semantic similarity between the text to be processed and the i-th text block; T represents the length of the keyword list extracted from the text to be processed, C represents the length of the keyword list corresponding to the i-th text block in the knowledge base, word tc Represents the word-level semantic similarity between the tth target keyword and the cth search keyword; Ref i Represents the semantic similarity between all keywords in the i-th text block; word kl represents the word-level semantic similarity between any two different words in the same set, k represents the kth keyword, l represents the lth keyword, and the range of l and k are both 1 to C; Target represents the semantic similarity between all keywords in the text to be processed; it is assumed here that the text to be processed comes from the user, so it has a high semantic similarity. Therefore, the consistency of the text to be processed is taken as the benchmark, and the semantic similarity of the knowledge base text block relative to the semantic similarity of the text to be processed is used as the consistency weight of the word-level semantic similarity of the text block. The expression is 6. The multi-level semantics-based retrieval enhancement generation method according to claim 1, characterized in that: In step two, the key information extraction algorithm is a named entity recognition algorithm trained based on professional domain corpus, and also includes a keyword extraction algorithm trained based on professional domain corpus, a topic extraction algorithm trained based on professional domain corpus, and a named entity recognition algorithm, a keyword extraction algorithm and a topic extraction algorithm trained based on professional domain corpus are used simultaneously.

7. The multi-level semantics-based retrieval enhancement generation method according to claim 1, characterized in that: The LLM model includes the open source chatGLM model and the closed source chatGPT model.

8. The multi-level semantics-based retrieval enhancement generation method according to claim 1, characterized in that: The text segmentation strategy uses delimiter-based segmentation while retaining complete sentences and containing some overlapping content. The segmentation result will try to meet the segmentation strategy of the set block size upper limit. You can also choose to directly segment according to the block size or other segmentation strategies; the block size is generally in the range of [300,500].

9. The multi-level semantics-based retrieval enhancement generation method according to claim 2, characterized in that: For similarity algorithm 1, you can select inner product or Euclidean distance function; for similarity algorithm 2, you can select inner product or Euclidean distance function.

10. The multi-level semantics-based retrieval enhancement generation method according to claim 3, characterized in that: Considering the limitation of LLM input length, N is usually set to a value not exceeding 3.

Citation Information

Cited By

  • Knowledge construction method and system based on large model and RAG technology

    CN120296111A

  • Knowledge retrieval method based on multilayer index

    CN120633843A

  • A knowledge retrieval method based on multi-level indexing

    CN120633843B

  • Hierarchical knowledge network construction and retrieval method for intelligent electric charge questions and answers

    CN120744073A

  • Laboratory quality management document intelligent generation method and system based on retrieval enhancement

    CN120951955A