Enhanced retrieval method based on language large model

By introducing enhanced retrieval methods into the large language model, including text vectorization and user problem rewriting, the knowledge limitations and hallucinations of the large language model under specific business needs are solved, the model's knowledge processing ability and data security are improved, and the accuracy and consistency are achieved.

CN120234429APending Publication Date: 2025-07-01AEROSPACE INTERNET OF THINGS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510317568.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Large language models have knowledge limitations, hallucination problems, and data security problems when dealing with specific business needs, resulting in a decrease in accuracy and consistency of model answers.

Method used

The enhanced search method based on the language model is adopted to enhance the model's knowledge processing capabilities and data security by obtaining target domain file data, text vectorization, user problem rewriting, post-retrieval rearrangement and context compression.

Benefits of technology

It improves the model's knowledge processing ability in specific fields, reduces the impact of irrelevant noise information, improves recall and accuracy, and optimizes the search effect and information reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234429A_ABST
    Figure CN120234429A_ABST
Patent Text Reader

Abstract

The invention discloses an enhanced retrieval method based on a language large model. The method comprises the following steps: acquiring target domain file data; the target domain file data is input into a text vectorization model, file vectors are obtained, the text vectorization model is obtained through training of a training set, and the training set is documents related to query and documents unrelated to query; based on the file vector, obtaining target document information in combination with a user query problem; compressing the target document information to obtain prompt information; and inputting the prompt information into a large language model to obtain a text result. According to the method, the knowledge processing capability of a specific field is enhanced in a targeted manner, the influence of irrelevant noise information is reduced, and the recall rate and the accuracy rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large models, and particularly relates to an enhanced retrieval method based on a language large model. Background Art

[0002] Currently, large models provide services based on models that have been pre-trained with large-scale data. They can answer user queries without accessing any external memory, and their knowledge base is static and cannot be updated in real time. However, there are three problems with general large models in meeting specific business needs: First, the limitation of knowledge. The knowledge of the model comes from pre-training data, and it cannot obtain some real-time or non-public data. Second, the hallucination problem. Facing some knowledge and questions that the large model does not possess, it may generate content that does not conform to the actual situation or contains incorrect information. Third, data security. For enterprises, since their private domain data will not be uploaded to a third-party platform for training, relying solely on the capabilities of general large models cannot solve specific data problems, and retrieval-augmented generation is a better solution.

[0003] Retrieval Augmented Generation (RAG) is a currently popular application solution for enhancing model capabilities. It can mainly solve the following problems: Make up for the deficiencies of large models in specific data. Reduce the hallucination problem and improve the accuracy of model generation. RAG is mainly divided into two stages: Retrieval stage: First, given an input query or question, RAG technology will use an efficient retrieval system to find the few paragraphs or sentences most relevant to the input from a large number of documents or databases. Generation stage: Then, the retrieved relevant information and the original input are used as the input of the large model. The large model will use this additional information to generate a more accurate, relevant, and detailed response.

[0004] When large models are put into practical applications, according to different requirements, when constructing RAG only using general vectorization models and a single processing process, the accuracy and consistency of model answers will both decrease. Therefore, how to improve the performance of the model is a topic that most current applications need to study and solve. Summary of the Invention

[0005] To solve the above technical problems, the present invention proposes an enhanced retrieval method based on a language large model to solve the hallucination problem of large language models and improve domain adaptability.

[0006] To achieve the above object, the following technical solutions are provided:

[0007] The present invention provides an enhanced retrieval method based on a language large model, including:

[0008] Obtain the file data of the target domain;

[0009] Input the file data of the target domain into the text vectorization model to obtain a file vector, where the text vectorization model is obtained by training with a training set, and the training set is relevant and irrelevant documents related to the query;

[0010] Based on the file vector and combined with the user's query problem, obtain the target document information;

[0011] Compress the target document information to obtain a prompt message;

[0012] Input the prompt message into the large language model to obtain a text result.

[0013] Optionally, obtaining the file data of the target domain includes:

[0014] Obtain the domain file data;

[0015] Preprocess the file data to obtain the file data of the target domain.

[0016] Optionally, preprocessing the file data includes: screening, cleaning, and splitting the files.

[0017] Optionally, based on the file vector and combined with the user's query problem, obtaining the target document information includes:

[0018] Rewrite the user's query problem to obtain a target query problem;

[0019] Based on the file vector, retrieve the target query problem to obtain the target document information.

[0020] Optionally, rewriting the user's query to obtain a target query problem includes:

[0021] Create a query rewriting chain according to the user's query problem;

[0022] Rewrite the user's query problem according to the rewriting chain to obtain the target query problem.

[0023] Optionally, based on the file vector, retrieving the target query problem to obtain the target document information includes:

[0024] Based on the file vector, retrieve the target query problem to obtain a retrieval result;

[0025] Input the retrieval result into the bge-reranker model to obtain the target document information, where the bge-reranker model is used to rearrange the document information according to the similarity to ensure the relevance of the document information.

[0026] Optionally, the general cross - similarity method is used to calculate the correlation between two sentences.

[0027] Optionally, the target document information is compressed to obtain prompt information, including:

[0028] The target document information is input into the abstract extraction model to obtain the prompt information, where the abstract extraction model is used to extract the abstract of the target document information.

[0029] Compared with the prior art, the present invention has the following advantages and technical effects:

[0030] Based on the large - model, the present invention adds links such as text vector adjustment, user question rewriting, and post - retrieval re - ranking. By fine - tuning the model with file data in a specific field, it more specifically enhances the knowledge processing ability in a specific field, reduces the influence of irrelevant noise information, and improves the recall rate and precision rate.

[0031] The present invention defines a rewriting prompt template to guide the pre - trained language model to query and rewrite the input question, making the method more systematic and refined, thus optimizing the retrieval effect.

[0032] The present invention improves the reliability of the retrieved information through post - retrieval re - ranking. By placing the most relevant information at the front, it helps the model better understand and respond to users. By combining multiple factors such as text length and keyword matching features for scoring, the reliability of the information is improved.

[0033] The present invention reduces the cost of computing resources through context compression, and at the same time makes the model focus on key information, thus improving the accuracy of the model's answer. Using a specific abstract extraction model for compression makes the method more specific and practical. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0035] Figure 1 is a specific flowchart of an enhanced retrieval method based on a large language model according to an embodiment of the present invention;

[0036] Figure 2 is a schematic framework diagram of an enhanced retrieval method based on a large language model according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will describe this application in detail with reference to the accompanying drawings and in combination with the embodiments.

[0038] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0039] This embodiment proposes an enhanced retrieval method based on a large language model, as Figure 1 shown, which specifically includes the following steps:

[0040] Obtain the target domain file data;

[0041] Input the target domain file data into a text vectorization model to obtain file vectors, where the text vectorization model is obtained by training with a training set, and the training set is relevant and irrelevant documents related to the query;

[0042] Based on the file vectors, in combination with the user's query problem, obtain the target document information;

[0043] Compress the target document information to obtain prompt information;

[0044] Input the prompt information into the large language model to obtain a text result.

[0045] Specifically, load the files in a specific domain, read the file text information, and cut the text information for subsequent text vectorization.

[0046] Vectorize the cut information through the text vectorization model and store it in the vector library.

[0047] The user inputs a query problem, uses a pre-trained language model to query and rewrite the input problem, and vectorize the rewritten problem.

[0048] According to the vectorized user problem, retrieve in the vector library to obtain relevant information.

[0049] Calculate the relevance of the retrieved information through the bge-reranker model and reorder it.

[0050] Context compression compresses the relevant information through abstract generation and keyword extraction.

[0051] Add the matched text information and the problem together to the prompt, submit it to the large language model, and answer the user's problem.

[0052] More specifically, fine-tuning the embedding model:

[0053] Extract document information data, load the embedding model bge-large-zh using the sentence-transformers library, and use the loss function to train the model to improve the accuracy of subsequent document vectorization. The data format is query-answer.

[0054] The sentence-transformers library used is based on the Hugging Face's transformers library and provides many pre-trained models that can conveniently convert text into high-quality vector representations. These vectors can be used for various natural language processing tasks such as semantic similarity calculation and retrieval. Among them, bge-large-zh is a pre-trained language model developed by BAAI, specifically for generating high-quality Chinese text embedding vectors, and this model is optimized for Chinese. The loss function uses MultipleNegativesRankingLoss for fine-tuning training. Its core idea is based on triplet loss and is optimized by using multiple negative samples instead of a single negative sample. When there are only positive samples in the dataset, batch n-1 negative samples will be added to each sample in the loss function, which can effectively improve the quality and relevance of sentence embeddings.

[0055] Furthermore, obtaining target domain file data includes:

[0056] Obtain domain file data;

[0057] Preprocess the file data to obtain target domain file data.

[0058] Specifically, text reading and segmentation:

[0059] Load files in a specific domain, read the text information of the files, and segment the text information for subsequent text vectorization. The files include pdf, doc, docx, xls, xlsx, ppt, pptx, etc.

[0060] Text segmentation uses semantic-based text segmentation to ensure the fine-grainedness of text segmentation. Specifically, the pre-trained language model BERT is used to calculate the semantic relationships between sentences and paragraphs. To load various types of data sources, the present invention utilizes SQLAlchemy to process structured database data, which is a powerful Python SQL toolkit and object-relational mapping framework. The present invention uses the ORM in SQLAlchemy to define models and construct queries, converts the data into the required format using the json package, and indexes the database to optimize performance. For unstructured data such as pdf, the present invention uses the pdfplumber library for processing, parsing documents, extracting text, and indexing embeddings.

[0061] Text vectorization:

[0062] The segmented file information is vectorized using the fine-tuned vector model and stored in the Faiss vector database.

[0063] Information vectorization can improve the accuracy of retrieving relevant context information. Faiss is an open-source vector database that provides an efficient index structure and fast search capabilities. The present invention uses IndexFlatL2 to index the vectors and uses the Faiss database methods faiss.write_index and faiss.read_index to write the vectors to a file and load and read them from the file.

[0064] Further, the preprocessing of the file data includes: screening, cleaning, and segmenting the files.

[0065] Further, based on the file vectors and combined with the user's query problem, obtaining the target document information includes:

[0066] Rewriting the user's query problem to obtain the target query problem;

[0067] Based on the file vectors, retrieving the target query problem to obtain the target document information.

[0068] Further, rewriting the user's query to obtain the target query problem includes:

[0069] Creating a query rewriting chain according to the user's query problem;

[0070] According to the rewriting chain, rewriting the user's query problem to obtain the target query problem.

[0071] Specifically, the user inputs a question query, creates a query rewriting chain using langchain, and uses a pre-trained language model to rewrite the user input to generate a more accurate query. Using LangChain to create a query rewriting chain is an effective method to improve the performance of an information retrieval system, where a rewriting prompt template is defined to guide the pre-trained language model to rewrite the user input query. The pre-trained language model calls the qwen-plus pre-trained language model of OpenAI to rewrite the user query and uses the LLMChain method to create a rewriting chain. Rewriting the user's question can make the query statement entered by the user more precise and specific, helping the model to better understand the user's intention and thus return more relevant results.

[0072] Furthermore, based on the file vectors, retrieve the target query question to obtain the target document information including:

[0073] Based on the file vectors, retrieve the target query question to obtain the retrieval result;

[0074] Input the retrieval result into the bge-reranker model to obtain the target document information. Among them, the bge-reranker model is used to re-rank the document information according to the similarity to ensure the relevance of the document information.

[0075] Specifically, use the bge-reranker model to calculate the relevance of the retrieved information. Here, it is because only the top-n retrieval result information was obtained previously, and it is impossible to determine whether it is relevant to the user input question. Irrelevant information may affect the performance of the model.

[0076] bge-reranker is a model specifically used to re-rank retrieval results. It can help the present invention more accurately determine which documents are most relevant to the user query. Input the retrieved documents into bge-reranker, calculate the relevance score between the query and the documents, and combine the text length and keyword matching features for scoring, so that the relevance of the documents can be evaluated more comprehensively.

[0077] Furthermore, compress the target document information to obtain the prompt information including:

[0078] Input the target document information into the abstract extraction model to obtain the prompt information. Among them, the abstract extraction model is used to extract the abstract of the target document information.

[0079] Specifically, use the abstract extraction model to further compress the retrieved information to reduce the input size of the model while maintaining the quality of the generated content.

[0080] Load the pre-trained abstract extraction model t5-base and perform abstract extraction on the retrieved documents, which can effectively reduce resource consumption and improve the quality of the incoming data.

[0081] More specifically, answer generation:

[0082] Add the matched text information and the question to the prompt and submit it to the large language model to answer the user's question.

[0083] Combine the matched text information with the user's query question to form a comprehensive prompt. This prompt can clearly indicate the task that the model needs to complete and contains sufficient context information to ensure that the generated answer is accurate. Input the constructed prompt into the pre-trained large language model, and the model can understand natural language and generate coherent answers based on the provided context.

[0084] The following combines the attached Figure 2 Elaborate on this embodiment in detail:

[0085] This embodiment provides an enhanced retrieval method based on a large language model, as Figure 2 shown, the implementation steps of this embodiment are as follows:

[0086] 1. Prepare specific domain file data:

[0087] Preparing specific domain file data is the primary step of the present invention, and its purpose is to create a specific domain and real-time resource library for subsequent fine-tuning, retrieval, and analysis. This process mainly includes the following two key steps.

[0088] (1) Collection of specific domain data:

[0089] Obtain data from the official websites of specific domains and files provided by domain experts. These files should contain relevant titles and detailed content. At the same time, for the files provided by experts, the data form should be query-answer structure.

[0090] (2) Preprocessing of data files:

[0091] The collected files should be preprocessed, including file screening, text cleaning (removing irrelevant information, unifying document formats), and text segmentation to adapt to subsequent vectorization processing.

[0092] 2. Fine-tuning of the vector model:

[0093] The purpose of vector model fine-tuning is to improve the accuracy of specific domain text vectorization for subsequent user queries and specific domain document vectorization. This process mainly includes the following process.

[0094] (1) Define the fine-tuning training task:

[0095] Define the MultipleNegativesRankingLoss task to minimize the distance between positive samples (documents relevant to the query) and the query, while maximizing the distance between negative samples (irrelevant documents) and the query, in order to improve the performance of the embedding model in processing documents in a specific domain.

[0096] (2) Evaluation and adjustment:

[0097] Evaluate the model performance on the test set to ensure that the model is not overfitting. At the same time, adjust the hyperparameters to find the optimal parameter configuration of the model.

[0098] 3. User query preprocessing:

[0099] Build a user query preprocessing module to optimize the natural language queries entered by users so that they can be effectively processed by the system. The main task of this module is to utilize the user's query intent and rewrite it into a format suitable for subsequent retrieval. This process mainly includes the following steps.

[0100] (1) User interface development:

[0101] Design a user-system interaction interface that allows users to simply enter query questions. The interface should be as intuitive and reasonable as possible.

[0102] (2) User input rewriting:

[0103] Design a rewriting chain to rewrite the user's query using other large language models. This process will modify both the content and format of the user's query to meet the requirements of subsequent models.

[0104] 4. Develop document retrieval:

[0105] The purpose of developing document retrieval is to meet the needs of users to query specific domain knowledge. Therefore, it is necessary to process a large amount of relevant file data and find the document information relevant to the user's query. This process involves the following key steps.

[0106] (1) Vectorization of the specific domain knowledge base:

[0107] After preparing the specific domain knowledge base in step 1, use the vectorization model fine-tuned in step 2 to represent the documents in vector form. These vectors can represent the main content and semantic information of the documents. At the same time, store these vectors in a database suitable for fast retrieval and provide an index for them to ensure fast similarity matching during subsequent retrieval.

[0108] (2) Document retrieval and re-ranking:

[0109] To retrieve the knowledge base according to the user's query, it is first necessary to use a similarity calculation method to obtain the similarity between the user's query and each document, then sort the similarity results, and finally obtain the document information most relevant to the query.

[0110] Through the above steps, the most relevant information can be found from the knowledge base, but the relevance of the information still cannot be guaranteed. Therefore, document re-ranking is introduced to further screen the results to improve the quality of the final generated answer. This process of document re-ranking usually uses a more complex model to re-evaluate and adjust the order of the preliminary retrieval results, such as bge-reranker.

[0111] (3) Document compression:

[0112] To enhance the input quality of the large model, the re-ranked information is further processed through document compression. The main methods include abstract extraction and keyword extraction methods, and common abstract extraction models such as t5-base.

[0113] 5. Implementing question-answer generation:

[0114] The purpose of generating questions and answers is to generate detailed and in-depth answers based on the user's query and the relevant information retrieved from the database. This process not only requires the text to be closely related to the user's query, but also requires providing accurate domain information and in-depth analysis. Therefore, generating questions and answers requires combining advanced large language models to provide efficient and accurate questions and answers for users. This process involves the following key steps.

[0115] (1) Selecting a suitable large language model:

[0116] Select a suitable model, such as the open-source model qwen of Alibaba. Considering the specific application scenario, a model suitable for processing relevant languages should be selected.

[0117] (2) Implementing input integration:

[0118] Effectively integrate the retrieved target document information and the user's query into the input of the generation model. Design and construct a prompt template, fill the user input and the retrieved information into the designed prompt template, and finally pass the prompt into the large model.

[0119] (3) Text generation:

[0120] Based on the integrated prompt, use the selected large language model to generate text. The text includes specific descriptions of the user's query, relevant analysis, result explanations, etc.

[0121] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present application should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. An enhanced retrieval method based on a large language model, characterized in that: include: Obtain target domain file data; Input the target domain file data into a text vectorization model to obtain a file vector, wherein the text vectorization model is obtained by training a training set, and the training set is documents related to the query and documents not related to the query; Based on the document vector and in combination with the user query question, obtaining target document information; Compressing the target document information to obtain prompt information; The prompt information is input into a large language model to obtain a text result.

2. The enhanced retrieval method based on a large language model according to claim 1, characterized in that: Obtaining target domain file data includes: Get domain file data; The file data is preprocessed to obtain the target domain file data.

3. The enhanced retrieval method based on a large language model according to claim 2, characterized in that: The preprocessing of the file data includes: screening, cleaning and segmenting the file.

4. The enhanced retrieval method based on a large language model according to claim 1, characterized in that: Based on the document vector and in combination with the user query question, obtaining target document information includes: Rewrite the user query question to obtain a target query question; Based on the document vector, the target query question is retrieved to obtain target document information.

5. The enhanced retrieval method based on a large language model according to claim 4, characterized in that: Rewrite the user query to obtain the target query question including: Creating a query rewrite chain according to the user query question; The user query question is rewritten according to the rewriting chain to obtain the target query question.

6. The enhanced retrieval method based on a large language model according to claim 4, characterized in that: Based on the document vector, searching the target query question to obtain target document information includes: Based on the document vector, searching the target query question to obtain a search result; The retrieval results are input into the bge-reranker model to obtain the target document information, wherein the bge-reranker model is used to re-rank the document information according to similarity to ensure the relevance of the document information.

7. The enhanced retrieval method based on a large language model according to claim 6, characterized in that: The general cross-similarity method is used to calculate the relevance between two sentences.

8. The enhanced retrieval method based on a large language model according to claim 1, characterized in that: Compressing the target document information and obtaining prompt information includes: The target document information is input into a summary extraction model to obtain the prompt information, wherein the summary extraction model is used to extract a summary of the target document information.