Context optimization method for retrieval enhancement generation, question and answer processing method and equipment

By sorting the initial context text in correlation and compressing the secondary statements, dynamically adjusting the k value to optimize the context text, solving the problem of insufficient answer quality caused by information loss and N value fixation in the prior art, achieving higher information retention and answer accuracy.

CN120216639APending Publication Date: 2025-06-27SHANGHAI STAR MAP BIT INFORMATION TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510285296.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When processing context text, the existing QA scheme based on RAG fixedly retains the first N tokens and discards the rest of the content, resulting in the loss of key information and the N value during truncation is fixed, which is not suitable for the complexity and context importance of different query contents, affecting the quality of answers.

Method used

By sorting the statements in the initial context text relevance, the k main statements with the highest correlation with the query content are determined, and the remaining secondary statements are compressed to form the optimized target context text. Dynamically adjust the k value through a reinforcement learning algorithm based on Q-Learning, ensuring that information is fully retained without increasing computational costs.

Benefits of technology

It effectively avoids the loss of key information, improves the relevance of context text and the accuracy of answers, and dynamically adjusts the k value to balance the calculation cost and answer quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216639A_ABST
    Figure CN120216639A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a context optimization method for retrieval enhancement generation, and a question and answer processing method and equipment. According to the scheme, when an initial context text is compressed, a mode of reserving first N tokens is not simply adopted, relevancy sorting is performed on statements, k main statements with the highest relevancy are reserved, and meanwhile, the initial context text is compressed; the secondary statements with low relevancy are not directly discarded, but are compressed to retain important information of the secondary statements and then spliced with the primary statements to obtain the optimized target context text, so that information is reserved as much as possible, loss of key information is avoided, and the accuracy of answering after the target context text is input into LLM is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and in particular, to a method for optimizing context of retrieval-augmented generation, a method for question-answering processing, and a device. Background Art

[0002] The QA (Question Answering) system is one of the most important applications of large language models and is widely used in various fields. In some QA applications in specific fields, since the LLM (Large Language Model) has blind spots for the latest technologies or professional terms, it is necessary to combine with the RAG (Retrieval-augmented Generation) technology to obtain a large amount of relevant context information and then input it into the LLM to provide more accurate and reliable answers.

[0003] Currently, the existing RAG-based OA solutions include the following steps:

[0004] 1. Determine the maximum context length: Preset an appropriate value of N, usually matching the context window size of the LLM.

[0005] 2. Perform text preprocessing: Tokenize the context text obtained through the RAG solution, calculate the number of tokens obtained, and determine whether the number of tokens corresponding to the context text exceeds the set value of N.

[0006] 3. Execute the truncation strategy: If the number of tokens corresponding to the following text ≤ N, directly input it into the LLM. If the number of tokens corresponding to the following text > N, retain the first N tokens and discard the rest.

[0007] 4. Input into the LLM for inference: Use the context text after executing the truncation strategy as the input of the LLM to obtain the corresponding answer.

[0008] It can be seen that the existing technology has the following problems: Since only the first N tokens are fixedly retained when the context text is too long and the rest of the content is directly discarded, it may cause the loss of key information in the context text and it is difficult to ensure that the information most relevant to the user's query content is retained. At the same time, since the set value of N for truncation is fixed and will not be adjusted according to the complexity of the query content or the importance of the context, it is difficult to dynamically adapt to different query contents. This will affect the answer quality and result in insufficient accuracy of the answer. Summary of the Invention

[0009] An object of the present application is to provide a context optimization method, a question-answering processing method and a device for retrieval-augmented generation, so as to solve the problem of insufficient answer accuracy in existing RAG-based QA solutions.

[0010] To achieve the above object, an embodiment of the present application provides a context optimization method for retrieval-augmented generation, and the method includes:

[0011] Retrieve according to the query content to obtain an initial context text;

[0012] Sort each statement in the initial context text to determine the k main statements with the highest relevance to the query content;

[0013] Compress the secondary statements other than the k main statements in the initial context text to obtain the compressed secondary statements;

[0014] Concatenate the main statements and the compressed secondary statements into a target context text.

[0015] Further, the method further includes:

[0016] Based on the Q-Learning reinforcement learning algorithm, dynamically determine the k main statements with the highest relevance to the query content.

[0017] Further, dynamically determining the k main statements with the highest relevance to the query content through the Q-Learning reinforcement learning algorithm includes:

[0018] Vectorize the query content and the initial context text respectively to obtain a query vector and a context vector;

[0019] Calculate the difference between the context vector and the query vector as the current given state;

[0020] According to the current given state, calculate the optimal action corresponding to the current given state through the optimal policy obtained by training with the Q-Learning reinforcement learning algorithm, where the optimal policy is used to determine the optimal action corresponding to the given state, and the optimal action is used to determine the number of main statements with the highest relevance to the query content in the given state.

[0021] Further, the optimal policy is obtained by training in the following manner:

[0022] Train through the Q-Learning reinforcement learning algorithm to obtain a Q-value table, and the Q-value table is used to determine the corresponding Q value obtained after context optimization by performing the corresponding action in the given state;

[0023] According to the Q-value table, the optimal policy π*(state) is determined in the following manner:

[0024] π*(state) = argmaxQ*(state, action)

[0025] where Q*(state, action) is the Q-value obtained through training by a reinforcement learning algorithm based on Q-Learning, state represents a given state, action represents the corresponding action executed in the given state, and argmax represents selecting the action that maximizes the Q-value in a given state.

[0026] Furthermore, through training by a reinforcement learning algorithm based on Q-Learning, a Q-value table is obtained, including:

[0027] The Q-value table is iteratively updated in the following manner:

[0028] Q*(state, action) = Q(state, action) + 1 / n(reward - Q(state, action))

[0029] where Q(state, action) represents the Q-value before context optimization when executing the corresponding action in the current given state, Q*(state, action) represents the updated Q-value after context optimization when executing the corresponding action in the current given state, the reward is the reward function, which represents the actual reward after context optimization when executing the corresponding action in the current given state, and is related to the context compression ratio and answer accuracy before and after context optimization when executing the corresponding action, and 1 / n is the learning rate, which is used to control the step size of updating the Q-value.

[0030] Furthermore, the reward is determined using the following formula:

[0031] R = -(1 - α)·τ + α·(2r - r*)

[0032] where τ is the context compression ratio before and after context optimization when executing the corresponding action, r is the ROUGE score corresponding to the target context text after using the corresponding action for context optimization, r* is the ROUGE score corresponding to the initial context text before using the corresponding action for context optimization, and α is a preset trade-off coefficient.

[0033] Furthermore, the main statement and the compressed secondary statements are concatenated into the target context text, including:

[0034] Concatenate the main statement and the compressed secondary statements in the original order corresponding to the initial context text to obtain the target context text.

[0035] An embodiment of the present application also provides a question-and-answer processing method, and the method includes:

[0036] Obtain the query content;

[0037] According to the query content, adopt the context optimization method generated by retrieval enhancement to obtain the target context text;

[0038] Input the query content and the target context text into a large language model to obtain an answer content regarding the query content.

[0039] Some embodiments of the present application also provide a computing device, where the device includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the device is triggered to execute the aforementioned context optimization method or question-and-answer processing method generated by retrieval enhancement.

[0040] Some other embodiments of the present application also provide a computer-readable medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the context optimization method or question-and-answer processing method generated by retrieval enhancement.

[0041] Compared with the prior art, an embodiment of the present application provides a context optimization scheme generated by retrieval enhancement. After retrieving according to the query content to obtain the initial context text, first sort each statement in the initial context text to determine the k main statements with the highest relevance to the query content, then compress the secondary statements other than the k main statements in the initial context text to obtain the compressed secondary statements, and then concatenate the main statements and the compressed secondary statements into the target context text. When compressing the initial context text, this scheme does not simply retain the first N tokens, but sorts the statements by relevance. While retaining the k main statements with the highest relevance, for the secondary statements with lower relevance, they are not directly discarded, but compressed to retain their important information and then concatenated with the main statements to obtain the optimized target context text, so as to retain as much information as possible and avoid loss of key information.

[0042] In addition, in a question-and-answer processing scheme provided by an embodiment of the present application, the aforementioned context optimization scheme generated by retrieval enhancement is adopted to optimize the context text. Therefore, the optimized target context text can retain as much information as possible and avoid loss of key information, thereby effectively improving the accuracy of the answer after the target context text is input into the LLM. Brief Description of the Drawings

[0043] Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments with reference to the following drawings:

[0044] Figure 1 It is a processing flowchart of a context optimization method for retrieval augmented generation provided by an embodiment of the present application;

[0045] Figure 2 It is a schematic diagram before compressing the context text in an embodiment of the present application;

[0046] Figure 3 It is a schematic diagram after compressing the context text in an embodiment of the present application;

[0047] Figure 4 It is a processing flowchart of an OA system processing a user query by adopting the solution in an embodiment of the present application;

[0048] The same or similar reference numerals in the drawings represent the same or similar components. Detailed Description of the Embodiments

[0049] The present application will be further described in detail below with reference to the drawings.

[0050] In a typical configuration of the present application, devices of the terminal and the service network both include one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0051] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of the computer-readable medium.

[0052] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer program instructions, data structures, program devices, or other data. Examples of the computer storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0053] The embodiment of the present application provides a context optimization method for retrieval enhancement generation. After searching according to the query content and obtaining the initial context text, the method first sorts the sentences in the initial context text, determines the k main sentences with the highest relevance to the query content, and then compresses the secondary sentences other than the k main sentences in the initial context text to obtain the compressed secondary sentences, and then splices the main sentences and the compressed secondary sentences into the target context text. When compressing the initial context text, this scheme does not simply retain the first N tokens, but sorts the sentences by relevance, retains the k main sentences with the highest relevance, and does not directly discard the secondary sentences with lower relevance, but compresses them to retain their important information and then splices them with the main sentences to obtain the optimized target context text, thereby retaining as much information as possible and avoiding the loss of key information.

[0054] In actual scenarios, the execution subject of the method can be a user device, a network device, or a device formed by integrating a user device and a network device through a network, or it can also be an application running on the above device. The user device includes but is not limited to various terminal devices such as computers, mobile phones, and tablet computers; the network device includes but is not limited to network hosts, single network servers, multiple network server sets, or cloud computing-based computer sets. Here, the cloud is composed of a large number of hosts or network servers based on cloud computing (Cloud Computing), where cloud computing is a type of distributed computing, a virtual computer composed of a group of loosely coupled computer sets.

[0055] Figure 1 The processing flow of a context optimization method for retrieval enhancement generation provided by an embodiment of the present application is shown, and the method at least includes the following processing steps:

[0056] Step S101, searching according to the query content to obtain the initial context text.

[0057] The query content refers to the question that needs to be input into the OA system. In the OA system based on RAG, the query content will be searched first to preliminarily obtain the context text related to the query content, that is, the initial context text in this solution. Since the initial context text generally includes invalid redundant information, and the LLM of the OA system usually has a certain context window limit, it is necessary to continue to perform subsequent steps to optimize the initial context text and reduce the length of the initial context text while minimizing the impact on the accuracy of the answer.

[0058] Step S102: Sort each statement in the initial context text to determine the top k main statements with the highest relevance to the query content.

[0059] In an actual scenario, the initial context text may include multiple statements. For example, for the initial context text C1, it may include 100 statements. When sorting each statement in the initial context text, the relevance between these statements and the query content can be considered. That is, the higher the relevance to the query content, the higher the ranking of the statement. After completing the sorting, the top k statements with the highest relevance to the query content can be determined according to the sorting result.

[0060] Among them, the relevance between a statement and the query content can be represented by various numerical values for measuring similarity. For example, in this embodiment, cosine similarity can be used to represent the relevance between a statement and the query content. After vectorizing each statement in the initial context text and the query content respectively, the cosine similarity between the embedding vectors of each statement and the query content is calculated to represent the relevance between the two, and then the top k main statements with the highest relevance to the query content can be selected according to the sorting result.

[0061] In this embodiment, the k main statements determined in this step can be defined as Top-k statements. The Top-k statements (Top-k Sentences) can be represented by the following formula:

[0062]

[0063] where, v q represents the embedding vector of the query content, represents the embedding vector of the i-th statement in the initial context text, represents sorting the cosine similarities in the set V in descending order.

[0064] Step S103: Compress the secondary statements other than the k main statements in the initial context text to obtain the compressed secondary statements.

[0065] Among them, compression means using a preset text compression algorithm to retain part of the information in the secondary statements, so that the important information can still be retained while reducing the unimportant information. This reduces the length of the context text and avoids completely discarding all information in the secondary statements, resulting in the loss of key information. In an actual scenario, the preset text compression algorithm can use any open-source text compression method, such as algorithms like BART, GPT-3 / 4, TextRank, etc.

[0066] For example, for the 100 sentences sent1 to sent100 included in the initial context text text1 in the foregoing embodiment, if k is set to 40, then the 40 sentences with the highest relevance to the query content will be selected as the main sentences, and the remaining 60 sentences will be used as the secondary sentences in this solution. When compressing these 60 secondary sentences, any of the foregoing text compression algorithms can be used to process them to obtain the compressed secondary sentences. In an actual scenario, the number of compressed secondary sentences can be less than or equal to 60. If a secondary sentence does not contain important information, the secondary sentence can be deleted. If a secondary sentence contains important information, then other information except the important information can be deleted, and only the required important information is retained, thereby shortening the length of the secondary sentence, and thus realizing the compression of the secondary sentence.

[0067] Step S104, splice the main sentence and the compressed secondary sentence into a target context text.

[0068] After completing the compression of the secondary sentences, the main sentence and the compressed secondary sentence can be spliced to complete the initial context text and obtain an optimized target context text. Since this solution does not simply retain the first N tokens when compressing the initial context text, but sorts the sentences by relevance, retains the k main sentences with the highest relevance, and does not directly discard the secondary sentences with lower relevance, but compresses them to retain their important information and then splices them with the main sentences to obtain an optimized target context text, so as to retain as much information as possible and avoid the loss of key information.

[0069] In some embodiments of the present application, when splicing the main sentence and the compressed secondary sentence, the main sentence and the compressed secondary sentence can be spliced in the original corresponding order in the initial context text to obtain a target context text. For example, taking 5 sentences s1 to s5 in the initial context text C2 as an example, where sentences s1, s3, and s5 are the main sentences, and s2 and s4 are the secondary sentences, as Figure 2 shown. After compressing the secondary sentences, the compressed secondary sentences s2' and s4' can be obtained. When splicing, the main sentences s1, s3, s5 and the compressed secondary sentences s2', s4' can be spliced in the original corresponding order in the initial context text to obtain a target context text C2', as Figure 3 shown. Thus, during the entire optimization process, the original order of the sentences is maintained, which ensures the integrity of the semantics and does not affect the meaning of the context due to rearranging the sentences, which is beneficial to improving the accuracy of subsequent answers.

[0070] In the solution of the embodiment of the present application, the core optimization goal is to reduce the length of the initial context text to reduce the computational cost during the LLM processing of the OA system, while also minimizing the impact on the accuracy of the OA system's answers after text compression. Specifically, assuming the number of tokens in the context C is T, and after being processed by the solution of the embodiment of the present application, the new context is C', and the number of tokens is t. The optimization goal is to reduce the token ratio τ before and after optimization and keep the accuracy of the answers as unchanged as possible. Thus, the optimization goal can be expressed by the following formula:

[0071] min((1-α)×τ+α×∣acc-acc*∣)

[0072] where τ = t / T represents the token ratio before and after optimization, α is a trade-off coefficient used to control the balance between the token ratio and the accuracy, acc is the accuracy of the answer to the context text before optimization, and acc* is the accuracy of the answer to the context text after optimization.

[0073] In some embodiments of the present application, based on the Q-Learning reinforcement learning algorithm, k main sentences with the highest relevance to the query content are dynamically determined. In this solution, a Q-learning-based reinforcement learning algorithm is used to dynamically select the appropriate value of k. According to the specific situation of each query content problem and the initial context text through reinforcement learning, the size of k is dynamically adjusted, so as to achieve the balance between the optimal computational cost and the answer accuracy.

[0074] Specifically, when the solution of the embodiment of the present application dynamically determines k main sentences with the highest relevance to the query content through the Q-Learning reinforcement learning algorithm, the query content and the initial context text can be vectorized respectively, and the query vector and the context vector can be obtained respectively. Among them, the query vector is the embedding vector of the query content, which is defined as v q , and the context vector is the embedding vector of the initial context text, which is defined as v c .

[0075] In this solution, the difference between the query vector and the context vector is used to define the current given state. Thus, the difference between the context vector and the query vector can be calculated as the current given state, that is, state = v c -v q .

[0076] Then, according to the current given state, the optimal action corresponding to the current given state can be calculated through the optimal policy obtained by training with the Q-Learning reinforcement learning algorithm. Among them, the optimal policy is used to determine the optimal action corresponding to the given state.

[0077] In the solution of the embodiment of the present application, the optimal policy can be trained and obtained in the following manner:

[0078] Train through a reinforcement learning algorithm based on Q-Learning to obtain a Q-value table, where the Q-value table is used to determine the corresponding Q-value obtained after context optimization by performing the corresponding action in a given state;

[0079] Determine the optimal policy π*(state) according to the Q-value table in the following manner:

[0080] π*(state) = argmaxQ*(state,action)

[0081] Among them, π*(state) is the optimal policy, indicating the optimal action taken in a certain given state (state). The Q*(state,action) is the Q-value obtained by training through a reinforcement learning algorithm based on Q-Learning. State represents the given state, action represents the corresponding action executed in the given state, and argmax represents selecting the action that maximizes the Q-value in a given state. The action in this solution can determine the specific number of Top-k statements to be retained dynamically in the given state of a specific query content and the initial context text. In this processing process, the specific number k of Top-k statements can be adaptive, which means that this solution will dynamically adjust according to each query content and the initial context text. The optimal action among them is the action that can maximize the Q-value in the given state.

[0082] For each query content-initial context text data pair, the number k of Top-k statements to be retained can be calculated through the current action value (action value) a in the aforementioned reinforcement learning algorithm and the total number of statements (n) in the initial context text. The specific calculation formula can be:

[0083] k = a × n.

[0084] Among them, a is the current action value calculated by the reinforcement learning algorithm, representing the currently selected compression degree. For example, in an actual scenario, it can be set to a value between 0.05 and 0.4, and n is the total number of statements in the initial context text.

[0085] Furthermore, in the process of training through a reinforcement learning algorithm based on Q-Learning to obtain a Q-value table, the Q-value table can be iteratively updated in the following manner:

[0086] Q*(state, action) = Q(state, action) + 1 / n * (reward - Q(state, action))

[0087] Among them, Q(state, action) represents the Q value before context optimization when performing the corresponding action under the current given state, and Q*(state, action) represents the updated Q value after context optimization when performing the corresponding action under the current given state. The reward is the reward function, which represents the actual reward after context optimization when performing the corresponding action under the current given state, and is related to the context compression ratio and answer accuracy before and after context optimization when performing the corresponding action. 1 / n is the learning rate, which is used to control the step size of updating the Q value. It can be seen from this that during the iterative update process, each action will be evaluated through the reward function reward, and the output value of the reward function determines the update of the Q value table. When reward > Q(state, action), it means that the current Q value underestimates the reward, and iterative update needs to be performed by increasing the Q value; when reward < Q(state, action), it means that the current Q value overestimates the reward, and iterative update needs to be performed by decreasing the Q value. By using the Q-Learning-based reinforcement learning algorithm to update the Q value table, an appropriate compression strategy can be selected for each query content-initial context text data pair, ensuring that the optimal action can be taken under different given states, thereby improving the optimization effect.

[0088] In some embodiments of the present application, the reward function reward adopts the following formula:

[0089] R = -(1 - α)·τ + α·(2r - r*)

[0090] Among them, R represents the output value of the reward function, τ is the context compression ratio before and after context optimization when performing the corresponding action, which can be the token ratio before and after optimization, r is the ROUGE (Recall-Oriented Understudy for Gisting Evaluation) score corresponding to the target context text after context optimization when performing the corresponding action, r* is the ROUGE score corresponding to the initial context text before context optimization when performing the corresponding action, and α is a preset trade-off coefficient, which is used to control the balance between the token ratio and accuracy and can be set according to the needs of the actual application scenario, such as set to 0.8, 0.9, etc.

[0091] As can be seen from the above, the reward function can be defined based on the token ratio and accuracy, and combines the ROUGE score. Its goal is to minimize the token ratio before and after optimization while keeping the accuracy of the answer corresponding to the optimized target context text as close as possible to that of the initial context text. Through this design, the accuracy of the answer content and the compression degree of the context are balanced, which has unique innovation.

[0092] In addition, the embodiment of the present application also provides a question-answering processing method. This method can first obtain the query content, and then according to the query content, use the above-mentioned context optimization method of retrieval-enhanced generation to obtain the optimized target context text. Then, the query content and the target context text are input into the large language model to obtain the answer content about the query content. Since the above-mentioned context optimization scheme of retrieval-enhanced generation is used to optimize the context text, the optimized target context text can retain as much information as possible and avoid the loss of key information, thus effectively improving the accuracy of the answer after input into the LLM.

[0093] Figure 4 The processing flow of the OA system using the solution in the embodiment of the present application to process user queries is shown, which at least includes the following steps:

[0094] Step S401, obtain the initial context text C: Retrieve based on the query content input by the user to obtain the relevant initial context text C.

[0095] Step S402, calculate the state: Calculate the difference between the embedding vector of the query content and the embedding vector of the initial context text to obtain the current state.

[0096] Step S403, select an action: Based on the current state, use the reinforcement learning algorithm based on Q-Learning to select an appropriate action and decide how many Top-k statements should be retained.

[0097] Step S404, generate the optimized target context text: According to the selected action, retain the most relevant k main statements, compress the redundant information in the secondary statements and splice them with the main statements to obtain the compressed target context text C'.

[0098] Step S405, call the API (Application Programming Interface) of the LLM to obtain the corresponding answer: Use the optimized target context text C' and the query content to generate the corresponding answer content through the LLM.

[0099] Based on another aspect of the present application, embodiments of the present application further provide a computing device, which includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the device is triggered to execute the context optimization method or the question-and-answer processing method of the aforementioned retrieval-augmented generation.

[0100] Specifically, the methods and / or embodiments in the embodiments of the present application can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium. The computer program contains program codes for executing the methods shown in the flowcharts. When the computer program is executed by a processing unit, the above functions defined in the methods of the present application are executed.

[0101] It should be noted that the computer-readable medium described in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device.

[0102] In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.

[0103] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).

[0104] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0105] As another aspect, this application also provides a computer-readable medium, which can be included in the devices described in the above embodiments; or it can exist separately and not be assembled into the device. The above computer-readable medium carries one or more computer program instructions, and the computer program instructions can be executed by a processor to implement the methods and / or technical solutions of the multiple embodiments of this application described above.

[0106] It should be noted that the present application can be implemented in software and / or a combination of software and hardware. For example, it can be implemented using an application specific integrated circuit (ASIC), a general purpose computer, or any other similar hardware device. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of the present application can be implemented using hardware, for example, as a circuit that cooperates with the processor to execute each step or function.

[0107] For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present application, the present application can be implemented in other specific forms. Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within the present application. Any reference signs in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims can also be implemented by one unit or device through software or hardware. The words such as first and second are used to denote names and do not represent any particular order.

Claims

1. A context optimization method for retrieval enhancement generation, characterized in that: The method comprises: Search according to the query content to obtain the initial context text; Sorting each sentence in the initial context text to determine k main sentences with the highest relevance to the query content; Compressing the secondary sentences other than the k main sentences in the initial context text to obtain compressed secondary sentences; The main sentence and the compressed secondary sentence are concatenated into a target context text.

2. The method according to claim 1, characterized in that The method further comprises: Based on the Q-Learning reinforcement learning algorithm, k main sentences with the highest relevance to the query content are dynamically determined.

3. The method according to claim 2, characterized in that Through a Q-Learning-based reinforcement learning algorithm, the k main sentences with the highest relevance to the query content are dynamically determined, including: Vectorize the query content and the initial context text respectively to obtain a query vector and a context vector respectively; Calculate the difference between the context vector and the query vector as the current given state; According to the current given state, the optimal action corresponding to the current given state is calculated by the optimal strategy obtained through training of the Q-Learning-based reinforcement learning algorithm, wherein the optimal strategy is used to determine the optimal action corresponding to the given state, and the optimal action is used to determine the number of main sentences with the highest relevance to the query content in the given state.

4. The method according to claim 3, characterized in that The optimal strategy is obtained by training in the following way: Training is performed by a Q-Learning-based reinforcement learning algorithm to obtain a Q-value table, wherein the Q-value table is used to determine a corresponding Q-value obtained after performing context optimization for a corresponding action in a given state; According to the Q value table, the optimal strategy π*(state) is determined in the following manner: π*(state)=argmaxQ*(state,action) Among them, the Q*(state, action) is the Q value obtained by training through the reinforcement learning algorithm based on Q-Learning, state represents a given state, action represents the corresponding action performed in a given state, and argmax represents selecting the action that maximizes the Q value in a given state.

5. The method according to claim 4, characterized in that Through training based on Q-Learning reinforcement learning algorithm, the Q value table is obtained, including: The Q value table is iteratively updated in the following way: Q*(state,action)=Q(state,action)+1 / n(reward-Q(state,action)) Among them, the Q(state, action) represents the Q value before executing the corresponding action for context optimization in the current given state, Q*(state, action) represents the updated Q value after executing the corresponding action for context optimization in the current given state, the reward is the reward function, which represents the actual reward after executing the corresponding action for context optimization in the current given state, and is related to the context compression ratio before and after executing the corresponding action for context optimization and the answer accuracy, and 1 / n is the learning rate, which is used to control the step size of updating the Q value.

6. The method according to claim 5, characterized in that The reward is determined using the following formula: R=-(1-α)·τ+α·(2r-r*) Among them, τ is the context compression ratio before and after the corresponding action is performed for context optimization, r is the ROUGE score corresponding to the target context text after the corresponding action is performed for context optimization, r* is the ROUGE score corresponding to the initial context text before the corresponding action is performed for up and down optimization, and α is the preset trade-off coefficient.

7. The method according to claim 1, characterized in that The main sentence and the compressed secondary sentence are concatenated into a target context text, including: The main sentence and the compressed secondary sentence are concatenated according to the corresponding original order in the initial context text to obtain the target context text.

8. A question-answering processing method, characterized in that: The method comprises: Get the query content; According to the query content, adopt the method described in any one of claims 1 to 7 to obtain the target context text; The query content and target context text are input into a large language model to obtain answer content related to the query content.

9. A computing device, wherein: The device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute the method according to any one of claims 1 to 8.

10. A computer readable medium having computer program instructions stored thereon, wherein the computer program instructions can be executed by a processor to implement the method according to any one of claims 1 to 8.