Document dialogue method, system and equipment based on large language model and medium

By constructing a vector database of documents and using large language models for multi-step adaptation methods, the problems of relatively scattered knowledge and hallucinations in the existing technology are solved, and more accurate and comprehensive answer results are achieved.

CN120011521APending Publication Date: 2025-05-16XIAN YANGU TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510341372.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When faced with problems with relatively scattered knowledge or issues with a particular word, the existing search enhancement generation methods have poor answers and serious hallucinations, resulting in inaccurate, incomplete or misleading output.

Method used

By constructing a vector database of documents, using the questions to be searched to search, the most relevant sentences in the questions to be searched and the related corpus blocks are extracted, and prompt words are designed, and multi-step adaptation is combined with the large language model to comprehensively answer the questions to be searched.

Benefits of technology

It improves the comprehensiveness and accuracy of knowledge, especially for questions with relatively scattered answers, reduces the hallucinations of the model and obtains more accurate results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011521A_ABST
    Figure CN120011521A_ABST
Patent Text Reader

Abstract

The invention discloses a document dialogue method, system, equipment and medium based on a large language model, and relates to the technical field of retrieval enhancement generation, and the method comprises the following steps: obtaining a vector database of a to-be-retrieved document after coding, and obtaining a to-be-retrieved question; converting the to-be-retrieved question into a vector, and retrieving a plurality of related corpus blocks in a vector database by using the vector; the to-be-retrieved question and the related corpus block are input into the large language model, the most related sentence in the to-be-retrieved question and the related corpus block is extracted, a cue word is designed, the large language model obtains the related position of the sentence and the to-be-retrieved question through the cue word, and the to-be-retrieved question is comprehensively answered. The related corpus blocks are retrieved in the vector database through the to-be-retrieved problem, corpus block retrieval of various knowledge of the document is achieved, comprehensiveness and accuracy of the knowledge are enhanced, and a more accurate result is obtained through multi-step adaptation of the large language model to the related corpus blocks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of retrieval enhancement generation technology, and in particular to a document dialogue method, system, device and medium based on a large language model. Background Art

[0002] In recent years, the rapid development of large language models (LLMs) has brought convenience to people. However, in the era of information explosion, massive documents and papers are published every day, and the training data of large language models is time-sensitive (training data is before a certain moment). Therefore, when asking for some knowledge that the large language model itself does not have, hallucinations may occur (that is, answers that are not asked or answers that are irrelevant). Especially for some industries that need to keep up with current affairs, this hallucination phenomenon is intolerable, so other methods are needed to make up for this shortcoming of large language models.

[0003] Retrieval-Augmented Generation (RAG) uses documents and papers that closely follow current events as external knowledge bases. In the process of Q&A with users, it uses the questions entered by users to retrieve relevant paragraphs in the knowledge base, and uses LLM to complete the Q&A based on the user's questions. This method can achieve good Q&A performance without training or fine-tuning LLM, and has been very popular in recent years.

[0004] Existing retrieval-enhanced generation methods can usually only answer questions where the knowledge in the document is relatively concentrated. For questions where the knowledge is relatively dispersed or the questions focus more on a certain word, the answer results are often poor. At the same time, there is a serious hallucination phenomenon, and when faced with certain inputs, inaccurate, incomplete or misleading outputs are produced. Summary of the invention

[0005] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and to provide a document dialogue method, system, device and medium based on a large language model, so as to solve the problem in the prior art that for problems with more dispersed knowledge or problems that focus on a certain word, the answer results are often poor, and there is a serious hallucination phenomenon, which results in inaccurate, incomplete or misleading output when faced with certain inputs.

[0006] The present invention specifically provides the following technical solutions: A document dialogue method based on a large language model comprises the following steps: Obtain the vector database of the documents to be retrieved after encoding, and obtain the questions to be retrieved; Convert the question to be searched into a vector, and use the vector to retrieve multiple relevant corpus blocks in a vector database; The question to be retrieved and the related corpus are input into the large language model, the most relevant sentences in the question to be retrieved and the related corpus are extracted, and a prompt word is designed. The large language model obtains the relevance between the sentence and the question to be retrieved through the prompt word, and comprehensively answers the question to be retrieved.

[0007] Preferably, the step of obtaining a vector database of documents encoded to be retrieved includes: Get the corpus information in the document to be retrieved; The corpus block information is cut into blocks according to certain rules to obtain corpus blocks, and the corpus blocks are encoded by an encoder to obtain vectors of the corpus blocks, and a vector database is constructed using the vectors.

[0008] Preferably, the question to be searched is converted into a vector, and the vector is used to retrieve multiple relevant corpus blocks in a vector database, including: Encode the input question to be retrieved through the encoder to obtain the vector of the question to be retrieved; wherein the encoder here has the same output vector length as the encoder used when constructing the vector database; The vector of the question to be searched is used to retrieve multiple corpus blocks with similarity higher than a threshold in the database.

[0009] Preferably, after using the vector to retrieve a plurality of related corpus blocks in the vector database, the method further includes: Extract keywords of the questions to be searched, perform data cleaning on the corpus blocks based on the keywords, and reorder the corpus blocks to obtain keywords related to the questions to be searched; Based on relevant keywords, the corpus block is compressed to remove redundant information that is irrelevant or not closely related to the keywords, and the prompt words are adjusted to optimize the structure of the text to obtain a final corpus block after removing irrelevant redundant information in the corpus block. The final corpus block is used to input a large language model.

[0010] Preferably, data cleaning of the corpus blocks is performed based on the keywords, and the corpus blocks are reordered, including: Design relevant prompt words, and use the prompt words to retain the keywords in the question to be searched; Traverse the retrieved corpus blocks and determine whether each corpus block contains the extracted keywords. If it does not contain any keywords, delete the corresponding corpus block; The query and the remaining corpus blocks are segmented, and the word frequency is used to obtain the relevance of each remaining corpus block to the query, and the corpus blocks are rearranged according to the relevance.

[0011] The present invention provides a document dialogue system based on a large language model, comprising: The acquisition module is used to obtain the vector database of the documents to be retrieved after encoding and obtain the questions to be retrieved; A conversion module, used to convert the question to be searched into a vector, and use the vector to retrieve multiple relevant corpus blocks in a vector database; The dialogue module is used to input the question to be retrieved and the related corpus into the large language model, extract the most relevant sentences from the question to be retrieved and the related corpus, and design a prompt word. The large language model obtains the relevance between the sentence and the question to be retrieved through the prompt word and comprehensively answers the question to be retrieved.

[0012] The present invention provides a computer device, including a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of the above-mentioned document dialogue method based on a large language model.

[0013] The present invention provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned document dialogue method based on a large language model are implemented.

[0014] Compared with the prior art, the present invention has the following significant advantages: The present invention builds a vector database of documents, and searches for relevant corpora in the vector database through questions to be searched, thereby realizing corpus retrieval of various knowledge in documents, thereby enhancing the comprehensiveness and accuracy of knowledge. The present invention is particularly suitable for questions with scattered answers. At the same time, through multi-stage question-answering, the most relevant sentences in the questions to be searched and the final corpus are obtained, and the relevance of the sentences to the questions to be searched is answered by a large language model through prompt words. A more accurate result is obtained through multi-step adaptation of the large language model to the relevant corpora, thereby reducing the illusion phenomenon of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 The present invention is a general flow chart of a document dialogue method based on a large language model; Figure 2 The present invention is a partial flow chart of a document dialogue method based on a large language model. DETAILED DESCRIPTION

[0016] The following is a clear and complete description of the technical solutions of the embodiments of the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0017] like Figure 1 and Figure 2 As shown, the present invention provides a document dialogue method based on a large language model, which specifically includes the following steps: Step S1: Obtain a vector database of the encoded documents to be retrieved and obtain the questions to be retrieved.

[0018] Preparation stage: Firstly, the existing document parsing method is used to obtain the corpus information in the document to be retrieved.

[0019] Then the corpus information is cut into pieces (document cutting) according to certain rules to obtain the corpus, and the corpus is encoded using the (Embedding) encoder to obtain the vector of the corpus, and a vector database is constructed based on this vector.

[0020] Step S2: Convert the question to be searched into a vector, and use the vector to search for multiple relevant corpus blocks in the vector database.

[0021] Question and answer session: The input question to be retrieved is extracted into triples and encoded through the Transformer-based (Embedding) encoder to obtain the vector of the question to be retrieved. This encoder can be understood as a small module in the model. For example, when building a database, the encoder will encode a sentence into a vector of length 4096. In the question-answering stage, the question input by the user also needs to be encoded into a vector of 4096 using the same encoder. Among them, since different models encode different vectors after encoding the same sentence, the encoder here is consistent with the encoder used to build the vector database, and the encoder output vector length is consistent, otherwise, the retrieval effect will be worse. The consistency of vector length is mainly to facilitate the calculation of the similarity between user questions and information in the database. The same model is used mainly because the vector lengths output by encoders of different models may be the same, but the vectors encoded by encoders of different models for the same sentence are very different. After encoding, the question to be retrieved can be represented as a denser and more detailed vector (the length of the vector is determined by the encoder. The more parameters a model has, the longer the length of the vector after encoding, and the stronger the ability to represent the sentence in the end).

[0022] After the vector of the question to be retrieved is recalled, multiple corpus blocks with similarity higher than the threshold are retrieved from the database. The vector here is the query vector.

[0023] This step S2 also includes: Using a large language model generally requires instructions. You need to clearly tell the large language model what you want to do and what effect you want to achieve. At this time, you need prompt words. Different prompt words need to be designed depending on the functions that the model needs to complete.

[0024] Extract keywords of the questions to be searched, clean the corpus based on the keywords, and reorder the corpus to obtain keywords related to the questions to be searched. These keywords are the basis for text analysis and compression.

[0025] Based on relevant keywords, the corpus blocks are compressed to remove redundant information that is irrelevant or not very relevant to the keywords to improve the quality and relevance of the corpus blocks. While compressing, the prompt words are adjusted to optimize the structure of the text to ensure clear logic and distinct levels. The final corpus blocks are obtained after removing irrelevant redundant information in the corpus blocks. The final corpus blocks are used to input the large language model. The compressed corpus blocks have higher information density, that is, they contain more useful information in less text. This is very important for quickly obtaining key information and conducting in-depth analysis.

[0026] Clean the corpus blocks based on keywords and reorder the corpus blocks, including: Design relevant prompt words, such as: "Next, I will give you a paragraph. Please extract the key words in this sentence, especially nouns, restrictive words, numbers, words expressing negation, etc." The prompt words retain the key words in the question to be searched, and the extracted words fully summarize the meaning of the original question.

[0027] Using the keywords extracted from the original question in the previous step, perform data cleaning (corpus cleaning) on ​​the retrieved corpus blocks. The specific operation steps are: traverse the retrieved corpus blocks, and determine whether each corpus block contains the extracted keywords in turn. If it does not contain any keywords, delete the corresponding corpus block.

[0028] Using the existing word segmentation library Jieba, we segmented the query questions and the remaining corpus blocks, compressed the corpus, and rearranged the corpus blocks using the bm25 rearrangement algorithm. The main principle is to use the word frequency to obtain the relevance of each remaining corpus block to the query questions, and rearrange the corpus blocks according to the relevance. The specific expression of relevance is:

[0029] ; in is the number of words, Representative words, IDF represents the inverse document frequency, f express In the documentation D The number of times it appears in len ( D ) indicates a document D Length, avg_len represents the average length of the document, k 1 and bis a hyperparameter, usually set to 1.5 and 0.75.

[0030] Step S3: Input the question to be retrieved and the related corpus into the large language model LLM, extract the most relevant sentences in the question to be retrieved and the related corpus, and design a prompt word. The large language model obtains the relevance between the sentence and the question to be retrieved through the prompt word and comprehensively answers the question to be retrieved.

[0031] After the command is input into the large language model (LLM), based on the original question and the information of the compressed corpus, the large language model can quickly filter out the sentences most relevant to the original question from these compressed corpus by designing appropriate prompt words. This process not only improves the efficiency of information retrieval, but also ensures the quality and relevance of the selected information. In this way, queries can be responded to more accurately, providing users with more accurate and useful answers.

[0032] After extracting the sentences most relevant to the original question, the next key step is to design appropriate prompt words that will guide the model to conduct in-depth analysis of these sentences to determine their relevance to the original question. The design of prompt words needs to be carefully considered, and they should be able to stimulate the model to fully understand the content of the sentence, including the topic, context, and logical connection between it and the original question.

[0033] After completing the in-depth analysis of the original question and the efficient screening of the information compression corpus, a key stage is to comprehensively answer the original question. This stage requires the model to not only understand the core points of the question, but also to accurately grasp the relevant sentences extracted from the corpus and synthesize this information to form an answer to the original question.

[0034] The purpose of this search operation is to alleviate the hallucination problem of the large language model (when asking questions to the large language model, the large language model itself has not learned it, and the model will answer randomly, that is, the answer is wrong and seriously inconsistent with reality). However, it is not certain whether the model has been trained on the questions asked by the user, so it is necessary to search for some reference content, and then let the model combine its own knowledge and the retrieved content to answer.

[0035] Based on the above method, the present invention provides a document dialogue system based on a large language model, including: a collection module, a conversion module and a dialogue module.

[0036] Among them, the acquisition module is used to obtain the vector database after encoding the documents to be retrieved, and obtain the questions to be retrieved; the conversion module is used to convert the questions to be retrieved into vectors, and use the vectors to retrieve multiple related corpora in the vector database; the dialogue module is used to input the questions to be retrieved and the related corpora into the large language model, extract the most relevant sentences in the questions to be retrieved and the related corpora, and design a prompt word. The large language model obtains the relevance between the sentence and the question to be retrieved through the prompt word, and comprehensively answers the question to be retrieved.

[0037] The present invention also provides a computer device, including a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a document dialogue method based on a large language model.

[0038] According to the disclosed embodiments, a computing device may communicate with one or more external devices (e.g., keyboards, pointing devices, Bluetooth communications, etc.), or with any device (e.g., routers, modems, etc.) that enables a computing device to communicate with one or more other computing devices.

[0039] The present invention also provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of a document dialogue method based on a large language model are implemented.

[0040] According to the disclosed embodiments, the storage medium may be a non-volatile computer-readable storage medium, such as but not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0041] The above content is a further detailed description of the present invention in combination with a specific preferred embodiment. For technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as belonging to the protection scope of the present invention.

Claims

1. A document dialogue method based on a large language model, characterized in that: The steps include: Obtain the vector database of the documents to be retrieved after encoding, and obtain the questions to be retrieved; Convert the question to be searched into a vector, and use the vector to retrieve multiple relevant corpus blocks in a vector database; The question to be retrieved and the related corpus are input into the large language model, the most relevant sentences in the question to be retrieved and the related corpus are extracted, and a prompt word is designed. The large language model obtains the relevance between the sentence and the question to be retrieved through the prompt word, and comprehensively answers the question to be retrieved.

2. A document dialogue method based on a large language model as claimed in claim 1, characterized in that: The step of obtaining a vector database of documents to be retrieved after being encoded includes: Get the corpus information in the document to be retrieved; The corpus block information is cut into blocks according to certain rules to obtain corpus blocks, and the corpus blocks are encoded by an encoder to obtain vectors of the corpus blocks, and a vector database is constructed using the vectors.

3. A document dialogue method based on a large language model as claimed in claim 2, characterized in that: The question to be searched is converted into a vector, and the vector is used to retrieve multiple relevant corpus blocks in the vector database, including: Encode the input question to be retrieved through the encoder to obtain the vector of the question to be retrieved; wherein the encoder here has the same output vector length as the encoder used when constructing the vector database; The vector of the question to be searched is used to retrieve multiple corpus blocks with similarity higher than a threshold in the database.

4. A document dialogue method based on a large language model as claimed in claim 1, characterized in that: After the vector is used to retrieve a plurality of related corpus blocks in the vector database, the method further includes: Extract keywords of the questions to be searched, perform data cleaning on the corpus blocks based on the keywords, and reorder the corpus blocks to obtain keywords related to the questions to be searched; Based on relevant keywords, the corpus block is compressed to remove redundant information that is irrelevant or not closely related to the keywords, and the prompt words are adjusted to optimize the structure of the text to obtain a final corpus block after removing irrelevant redundant information in the corpus block. The final corpus block is used to input a large language model.

5. A document dialogue method based on a large language model as claimed in claim 4, characterized in that: The data of the corpus blocks are cleaned based on the keywords, and the corpus blocks are reordered, including: Design relevant prompt words, and use the prompt words to retain the keywords in the question to be searched; Traverse the retrieved corpus blocks and determine whether each corpus block contains the extracted keywords. If it does not contain any keywords, delete the corresponding corpus block; The query and the remaining corpus blocks are segmented, and the word frequency is used to obtain the relevance of each remaining corpus block to the query, and the corpus blocks are rearranged according to the relevance.

6. A document dialogue system based on a large language model, characterized in that: include: The acquisition module is used to obtain the vector database of the documents to be retrieved after encoding and obtain the questions to be retrieved; A conversion module, used to convert the question to be searched into a vector, and use the vector to retrieve multiple relevant corpus blocks in a vector database; The dialogue module is used to input the question to be retrieved and the related corpus into the large language model, extract the most relevant sentences from the question to be retrieved and the related corpus, and design a prompt word. The large language model obtains the relevance between the sentence and the question to be retrieved through the prompt word and comprehensively answers the question to be retrieved.

7. A computer device, characterized in that: It comprises a memory and a processor, wherein a program is stored in the memory, and when the program is executed by the processor, the processor executes the steps of a document dialogue method based on a large language model as described in any one of claims 1 to 5.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a document dialogue method based on a large language model as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Knowledge retrieval enhancement generation method and system based on large language model

    CN118394890A

  • Method, system and equipment for retrieval enhancement generation based on large language model and medium

    CN118964387A

  • Large language model question and answer method, device and equipment based on retrieval enhancement and medium

    CN119597874A

Cited By

  • Retrieval enhancement generation method and system, computer equipment and storage medium

    CN121210644A