Long text question and answer collaborative reasoning method based on large language model

Through the collaborative framework of the integrated inference module and the retrieval enhancement generation module, the context window size and middle information loss of large language models in long text Q&A is solved, and the accuracy and user experience of long text Q&A are improved.

CN120354945APending Publication Date: 2025-07-22EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510481698.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Large language models are limited by context window size and middle information loss problems in long text Q&A tasks, resulting in poor performance.

Method used

A collaborative reasoning framework with a comprehensive inference module and a search enhancement generation module is adopted. By combining global text inference with local text inference, a search enhancement generation strategy is introduced to alleviate the context window size limitation and the loss of middle information, and improve the accuracy of long text Q&A.

Benefits of technology

Enhanced the accuracy of the large language model in long text Q&A, ensures the complete and accurate answers, and improves user experience and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354945A_ABST
    Figure CN120354945A_ABST
Patent Text Reader

Abstract

The invention discloses a long text question and answer collaborative reasoning method based on a large language model, which is characterized in that a long text question and answer reasoning framework with reasoning collaboration is generated by adopting comprehensive reasoning and retrieval enhancement constructed by a prompt project, and two answers obtained by reasoning are integrated into a structured prompt word by utilizing the large language model; according to the comprehensive reasoning, global text reasoning and local text reasoning are fused through Gaussian distribution weighting, and an answer is obtained; the retrieval enhancement generation reasoning uses a large language model to answer questions according to retrieved contents to obtain another answer. Compared with the prior art, the method has the advantages that the accuracy of the model in the long text question-answering task is enhanced, the negative influence of information loss in the long text on the long text question-answering performance of the large language model is relieved, the limitation of the model context window size on long text processing is relieved, the answer content is complete and accurate, the user experience and satisfaction degree are improved, and the user experience is improved. Good application prospects are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a collaborative inference method for long text question answering based on a large language model. Technical Background

[0002] With the development of large language models, they are widely used in processing long text tasks such as In-Context Learning and Tool Learning. However, large language models still have deficiencies in long text processing: on the one hand, large language models cannot accommodate overly long texts. The long text question answering of large language models is limited by the size of the context window, that is, when the length of the text input to the model exceeds the length set during the model's pre-training, it is difficult for the model to adapt to longer position encodings, thus causing the "Out-of-Distribution" (OOD) problem. Due to the quadratic complexity of the self-attention mechanism of large language models, they face high costs during training and inference, which limits their use of longer texts for pre-training. On the other hand, large language models have difficulty understanding overly long texts. Some studies have pointed out that large language models have the "Lost-in-Middle" phenomenon, that is, when dealing with long text document question answering, if the relevant text information is in the middle of the document, the accuracy of the model will drop significantly. Even models such as Deepseek have broadened the length of the text that can be accommodated through training, but they may still perform poorly in long text processing tasks due to the "middle information" loss phenomenon. For example, the context window size of DeepSeek-R1 is 128k, but there is still room for improvement in its performance when dealing with inference tasks with text lengths from 32k to 128k. Therefore, the performance of large language models in long text question answering tasks still needs to be optimized.

[0003] In summary, the performance of large language models in the prior art in long text question answering tasks is limited by the context window size and the middle information loss problem. Therefore, a new technical solution is needed to alleviate the above problems and improve the performance of large language models in long text question answering tasks. Summary of the Invention

[0004] The object of the present invention is to provide a long text Q&A collaborative reasoning method based on a large language model in view of the deficiencies of the prior art. It adopts a reasoning framework for long text Q&A collaboration with an integrated reasoning module and a retrieval-augmented generation reasoning module architecture to improve the long text Q&A performance of the model. This method uses the collaborative reasoning framework constructed by prompt engineering. Through the text local information supplement reasoning mechanism, it can effectively alleviate the negative impact of the loss of middle information in long texts on the long text Q&A performance of the large language model. By introducing the retrieval-augmented generation strategy, it can effectively alleviate the limitation of the model context window size on long text processing. By improving the accuracy of long text Q&A and combining with prompt engineering, the accuracy of the large language model in long text Q&A is greatly enhanced, the answer content can be complete and accurate, the user experience and satisfaction are improved, and it has good application prospects.

[0005] The object of the present invention is achieved as follows: A long text Q&A collaborative reasoning method based on a large language model, characterized by adopting a collaborative reasoning framework of text local information supplement reasoning and retrieval-augmented generation constructed by prompt engineering. Through the text local information supplement reasoning mechanism and by introducing the retrieval-augmented generation strategy, it alleviates the negative impact of the loss of middle information in long texts on the long text Q&A performance of the large language model and the limitation of the model context window size on long text processing, so as to improve the accuracy of the large language model in long text Q&A. The The collaborative reasoning framework of text local information supplement reasoning and retrieval-augmented generation constructed by prompt engineering specifically includes the following steps: Step 1: Use the large language model to obtain two reasoning representations respectively through two methods of global text reasoning and local text reasoning according to the text content. The specific operation process is as follows: 1-1: Concatenate the question input by the user and the relevant long text document; 1-2: For global text reasoning, input the concatenated text content completely into the first-layer decoder of the large language model for reasoning, and the tensor output by the first-layer decoder is used as the global reasoning representation; 1-3: For local text reasoning, divide the concatenated text into N equal-length chunks, and input each chunk into the first-layer decoder of the large language model for reasoning to obtain N tensors output by the first-layer decoder; 1-4: Concatenate all the N tensors along the first dimension, and sum the concatenated tensors along the first dimension; 1-5: Divide the summed tensor by the value of the first dimension of the concatenated tensor to obtain the local reasoning representation.

[0006] Step 2: Use the large language model to fuse and reason the two obtained reasoning representations through Gaussian distribution weighting based on Euclidean distance correction to obtain A alternative answers. The specific operation process is as follows: 2-1: Generate weights for each position in the first dimension of the global inference representation according to the Gaussian distribution; 2-2: Calculate the Euclidean distance between each row tensor in the first dimension of the global inference representation and the local inference representation; 2-3: Calculate the Euclidean distance weights through the activation function; 2-4: Combine the Gaussian distribution weights and the Euclidean distance weights, and update each row tensor in the first dimension of the global inference representation in a weighted manner; 2-5: The updated global inference representation will continue to reason using the remaining decoders of the large language model (i.e., all decoders except the first-layer decoder), and finally obtain Option A as the answer.

[0007] Step 3: Process the relevant long text documents, and then use the sparse method, dense method, or tree structure recursive clustering and summarization model to construct a document index library. The sparse method uses a specific symbol as a delimiter to split all the spliced documents into multiple index units, and constructs a document index library through an inverted index for the processed text; the dense method uses a specific symbol as a delimiter, and constructs a document index library based on the approximate nearest neighbor index for the processed text using a pre-trained language model.

[0008] Step 4: Encode the user's question using an inverted index or a pre-trained language model embedding representation, retrieve relevant documents from the document index library according to the encoded question, and calculate the relevance between the question and the inverted index through BM25 for the document index library constructed by the sparse method; for the document index library constructed by the dense method, calculate the relevance between the question and the document vector through cosine similarity, and finally retrieve the top-K text segments most relevant to the question.

[0009] Step 5: Obtaining Option B as the answer 5-1: Integrate the filtered documents and the user's question through structured prompts and input them into the large language model for reasoning to obtain an alternative answer; 5-2: Place the user's question and the retrieved relevant text segments in a question-answering prompt template and input them into the large language model for reasoning to obtain another alternative answer; 5-3: Integrate the above two alternative answers into the structured prompt, and use the large language model to screen out the final answer as Option B.

[0010] Step 6: Place the Option A obtained in Step 2 and the Option B obtained in Step 5 in the structured prompt template and input them into the large language model. The large language model selects one from Option A and Option B according to the prompt instructions to obtain the final answer to the user's question.

[0011] The present invention has the following beneficial technical effects and remarkable technological progress compared with the prior art: 1) By integrating global text reasoning and local text reasoning, the accuracy of the large language model in long text Q&A tasks is enhanced, effectively alleviating the negative impact of information loss in the middle of long texts on the long text Q&A performance of the large language model; 2) The retrieval-augmented generation strategy is introduced, effectively alleviating the limitation of the model context window size on long text processing; 3) The prompt engineering is used to construct a collaborative reasoning strategy for the above two methods to further improve the long text Q&A performance of the model.

[0012] 4) By improving the accuracy of long text Q&A and combining with prompt engineering, the answer content can be complete and accurate, improving the user experience and satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 is a flowchart of the present invention; Figure 2 is a flowchart of the comprehensive reasoning module; Figure 3 is a flowchart of the retrieval-augmented generation reasoning module. DETAILED DESCRIPTION OF THE INVENTION

[0014] The specific implementation manners of the present invention will be described in detail below with reference to the accompanying drawings.

[0015] Refer to Figure 1 , a collaborative reasoning framework for long text Q&A based on a large language model, specifically including the following steps: Step 1, using the large language model to obtain two reasoning representations respectively through global text reasoning and local text reasoning according to the text content.

[0016] Step 2, using the large language model to fuse and reason the two obtained reasoning representations through Gaussian distribution weighting based on Euclidean distance correction to obtain alternative answers.

[0017] Step 3, processing the relevant long text documents to construct a document index library.

[0018] Step 4, after encoding the user's question, retrieving relevant documents from the document index library according to the question.

[0019] Step 5, integrating the screened documents and the question through structured prompts and then inputting them into the large language model for reasoning to obtain another alternative answer.

[0020] Step 6, integrating the two alternative answers into the structured prompt and using the large language model to screen out the final answer from them.

[0021] Refer toFigure 2 In step 2, preliminary reasoning is performed from both the global and local perspectives of the long text content to obtain the global reasoning representation and the local reasoning representation. Finally, the two reasoning representations obtained are fused and reasoned through Gaussian distribution weighting corrected based on the Euclidean distance to obtain an alternative answer.

[0022] The global text reasoning in step 1 includes the following steps: 1-1-1: Concatenate the question input by the user and the relevant long text document.

[0023] 1-1-2: Input the concatenated text content completely into the first-layer decoder of the large language model for reasoning. The output tensor is used as the global reasoning representation.

[0024] The local text reasoning in step 1 includes the following steps: 1-2-1: The local text reasoning divides the concatenated text into N equally long chunks.

[0025] 1-2-2: Each chunk is respectively input into the first-layer decoder of the large language model for reasoning to obtain the N tensors output by the first-layer decoder.

[0026] 1-2-3: Concatenate all the N tensors along the first dimension, and sum the concatenated tensors along the first dimension.

[0027] 1-2-4: Divide the summed tensor by the value of the first dimension of the concatenated tensor to obtain the local reasoning representation.

[0028] The Gaussian distribution weighting corrected based on the Euclidean distance in step 2 includes the following steps: 2-1: Generate weights for each position in the first dimension of the global reasoning representation according to the Gaussian distribution.

[0029] 2-2: Calculate the Euclidean distance between each row tensor in the first dimension of the global reasoning representation and the local reasoning representation.

[0030] 2-3: Calculate the Euclidean distance weights for the values of the several Euclidean distances obtained in the previous step through the activation function Sigmoid.

[0031] 2-4: Update each row tensor in the first dimension of the global reasoning representation in a weighted manner by combining the Gaussian distribution weights and the Euclidean distance weights.

[0032] 2-5: The updated global reasoning representation will continue to be reasoned using the remaining decoders of the large language model (i.e., all decoders except the first-layer decoder), and finally obtain alternative answer A.

[0033] The document index library in step 3 can be established mainly by the sparse method or the dense method, or a tree structure recursive clustering and summarization model can be used to complete the establishment of the document index library.

[0034] The sparse method includes the following steps: 3-1-1: Use a specific symbol "\n\n" as a delimiter to split all the concatenated documents into multiple index units.

[0035] 3-1-2: After processing, the text constructs a document index library through an inverted index.

[0036] The dense method includes the following steps: 3-2-1: Use a specific symbol "\n\n" as a delimiter to split the document content.

[0037] 3-2-2: Use a pre-trained language model to construct a document index library for the processed text based on approximate nearest neighbor indexing.

[0038] In step 4, after encoding the user's question, relevant documents are retrieved from the document index library according to the question. The sparse method calculates the relevance between the question and the inverted index through BM25; the dense method calculates the relevance between the question and the document vector through cosine similarity, and finally retrieves the top-K text segments most relevant to the question.

[0039] See Figure 3 , this retrieval-enhanced generation inference process shows the process of constructing a long text document index library, then retrieving relevant documents from the index library according to the question, and finally using a large language model to answer the question based on the retrieved content to obtain another alternative answer.

[0040] The structured prompt in step 5 needs to include the "question" raised by the user, the "question-related text" retrieved, and an instruction requiring the model to answer according to the specific requirements of the question, and place the question raised by the user and the retrieved relevant text segments in the Q&A prompt template and input them into the large language model to infer the alternative answer B.

[0041] In step 6, the two alternative answers A and B are placed into the structured prompt template and input into the large language model. The large language model selects one alternative answer from the two alternative answers according to the prompt instruction, that is, a binary choice, and the obtained answer is the answer to the user's question. The structured prompt in this step needs to include the "question" raised by the user, the "long text document", the "alternative answers" obtained in steps 2 and 5, and the instruction requiring the model to select the final answer from the alternative answers according to the specific requirements of the question. The structured prompt method is a general method in the field of prompt engineering. Generally, through the above 6 steps, an answer to a complete and accurate long text question can be generated. The F1 Score is a commonly used evaluation metric for the question-and-answer task. As a statistical metric for measuring the accuracy of a binary classification model, it is calculated by the harmonic mean of precision and recall, and its value ranges from 0 to 1. The closer the value is to 1, the better the performance of the model. The experimental results are shown in the comparison of the F1 Score of the long text question-and-answer system, the original large language model, and the traditional RAG system in Table 1 below: Table 1: Comparison of F1 Score of the present invention, the original large language model, and the traditional RAG system

[0042] It can be seen from Table 1 above that the effect of the present invention is obvious compared with the original large language model and the RAG method, improving the accuracy of long text question and answer, making the answer content complete and accurate, and improving the user experience and satisfaction.

[0043] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A long text Q&A collaborative reasoning method based on a large language model, characterized in that This method uses a comprehensive reasoning and retrieval-enhanced generation reasoning collaborative long text Q&A reasoning framework constructed by prompt engineering, and uses a large language model to integrate the inferred answer into the structured prompt words, and filters out the final answer from them. The specific process of collaborative reasoning includes the following steps: Step 1: Concatenate the question input by the user and the relevant long text document; Step 2: Use the large language model to perform global text reasoning and local text reasoning on the concatenated text content respectively to obtain two reasoning representations; Step 3: Use the large language model to fuse and reason the two reasoning representations through Gaussian distribution weighting corrected based on the Euclidean distance to obtain alternative answer A; Step 4: Slice the relevant long text document to construct a document index library. The slicing process is to slice the relevant long text document into multiple equal-length text segments; Step 5: Encode the user's question and input it into the document index library to retrieve the documents related to the question. The encoding uses the inverted index method or the pre-trained language model embedding representation method; Step 6: Screen the retrieved relevant documents and integrate them with the user's question through structured prompts, and input them into the large language model for reasoning to obtain alternative answer B; Step 7: Integrate alternative answers A and B into the structured prompt, and use the large language model to filter out one answer from them as the answer to the user's question.

2. The long text question-answering collaborative reasoning method based on a large language model according to claim 1, wherein The global text reasoning in step 2 inputs the concatenated text content into the first-layer decoder of the large language model for reasoning, and the tensor output by the first-layer decoder is used as the global reasoning representation; the local text reasoning slices the concatenated text into N equal-length language chunks, and each language chunk is respectively input into the first-layer decoder of the large language model for reasoning to obtain the N tensors output by the first-layer decoder, and the N tensors are concatenated along the first dimension for all tensors, and the concatenated tensors are summed along the first dimension, and the summed tensors are divided by the value of the first dimension of the concatenated tensors to obtain the local reasoning representation.

3. The long text question-answering collaborative reasoning method based on a large language model according to claim 1, characterized in that, The reasoning based on the Gaussian distribution weighting corrected by the Euclidean distance in step 3 is as follows: 3-1: Generate weights for each position in the first dimension of the global reasoning representation according to the Gaussian distribution; 3-2: Calculate the Euclidean distance between each row tensor in the first dimension of the global reasoning representation and the local reasoning representation; 3-3: Calculate the Euclidean distance weight through the activation function, and update each row tensor in the first dimension of the global reasoning representation in a weighted manner by combining the Gaussian distribution weight and the Euclidean distance weight; 3-4: The updated global reasoning representation uses the remaining decoders of the large language model, that is, all decoders other than the first-layer decoder, to continue reasoning to obtain alternative answer A.

4. The long text question-answering collaborative reasoning method based on a large language model according to claim 1, wherein, The document index library in step 4 is established by the sparse method or the dense method. The sparse method uses a specific symbol as a separator to slice all the concatenated documents into multiple index units, and constructs a document index library through an inverted index for the processed text; the dense method uses a specific symbol as a separator, and uses a pre-trained language model to construct a document index library based on approximate nearest neighbor indexing for the processed text.

5. The long text Q&A collaborative reasoning method based on a large language model according to claim 1, wherein In step 5, retrieve the top-K text fragments most relevant to the question through the document index library. If the document index library is established using the sparse method, calculate the relevance between the question and the inverted index through BM25; If the document index library is established using the dense method, calculate the relevance between the question and the document vector through cosine similarity.

6. The long text Q&A collaborative reasoning method based on the large language model according to claim 1, characterized in that In step 6, place the question raised by the user and the retrieved relevant text fragments into the Q&A prompt template and input them into the large language model to infer B alternative answers.

7. The long text question answering collaborative reasoning method based on a large language model according to claim 1, wherein In step 7, place the two alternative answers A and B into the structured prompt template and input them into the large language model. The large language model selects one of them according to the prompt instructions as the final answer to the user's question.

8. The long text question-answering collaborative reasoning method based on a large language model according to claim 1 or claim 4, characterized in that, The document index library may be established using a tree structure recursive clustering and summarization model.