Question-answering method and apparatus based on large language model
By extracting knowledge of historical dialogue data and building memory chain data, the problems of insufficient knowledge refining and information entanglement in the use of dialogue data in the existing technology are solved, and the quality improvement and user experience optimization of answers generated by large language models are achieved.
Patent Information
- Application Number
- PCT/CN2024/128032
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-10
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-15
AI Technical Summary
In the prior art, there is insufficient knowledge refinement in the utilization of historical dialogue data, and the original dialogue data is directly vectorized, resulting in information entanglement between data of different topics, affecting the semantic accuracy of vector representation, and thus affecting the quality of answers generated by big models.
By extracting knowledge from historical dialogue data, building memory chain data, memory chain data assists large language models to generate answers, and improving user experience. The specific implementation includes classifying and marking historical Q&A pairs, filtering out highly generalized dialogue content to build memory chain nodes, filtering low generalized dialogue data, and achieving context content optimization for a long time offline.
With the assistance of memory chain data, the quality and semantic accuracy of the answers generated by large language models are improved, the user experience is improved, and the optimization of long-term dialogue memory is achieved.
Smart Images

Figure CN2024128032_15052025_PF_FP_ABST
Abstract
Description
A question-answering method and device based on a large language model
[0001] This application claims priority to Chinese patent application number 202311510489.2, filed on November 10, 2023, entitled “A Question Answering Method and Device Based on a Large Language Model,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a question-answering method and device based on a large language model. Background Art
[0003] The advent of ChatGPT is revolutionary for search. Conversational search will gradually replace traditional search (webpage relevance search) and become the new search method of the future. The industry generally agrees that historical conversation data is a key basis for improving conversational search systems. Leveraging this data not only helps clarify user intent and generate better answers, but also expands and refines training data, continuously iterating on the effectiveness of large models.
[0004] However, existing technologies have various problems with the use of historical conversation data. For example, the original conversation data is directly vectorized without knowledge extraction, and all historical conversation data is encoded using vectors. This leads to information entanglement between different topic data, affecting the semantic accuracy of vector representation and, in turn, the quality of answers generated by large models.
[0005] Summary of the Invention
[0006] An embodiment of the present application provides a question-answering method based on a large language model. By extracting knowledge from historical conversation data and constructing memory chain data, the memory chain data assists the large language model in generating higher-quality answers, thereby improving the user experience.
[0007] In a first aspect, the present application provides a question-answering method based on a large language model, comprising obtaining a current conversation state, the current conversation state including a query text in the current conversation; retrieving target memory chain data from a memory chain database based on the query text, the memory chain database including multiple memory chain data, the multiple memory chain data being obtained based on knowledge extraction of historical conversation data; obtaining a prompt text based on the query text and the target memory chain data; using the prompt text as input to a large language model (LLM), and outputting an answer corresponding to the query text.
[0008] The question-answering method based on a large language model provided in this application constructs memory chain data by extracting knowledge from historical conversation data. The memory chain data assists the large language model in generating better answers and improving the user experience.
[0009] In one possible implementation, a specific implementation of constructing each memory chain data in multiple memory chain data is: classifying each historical question-answer pair in each historical dialogue data to obtain the category of each historical question-answer pair; selecting historical question-answer pairs of a target category from multiple historical question-answer pairs in the dialogue data, where the generalization of the historical question-answer pairs of the target category is greater than the generalization of historical question-answer pairs of other categories; topic-tagging the historical question-answer pairs of the target category to obtain the topic of the historical question-answer pairs of the target category; and constructing initial memory chain data corresponding to each historical dialogue data based on historical question-answer pairs with the same topic.
[0010] In this possible implementation, by classifying and topic-tagging historical question-answer pairs, conversation content with high generalization value is screened out to form memory chain nodes, and conversation data with low generalization value is filtered out, which significantly reduces the amount of data and realizes the optimization of contextual content under a longer timeline.
[0011] In another possible implementation, each historical question-answer pair in each historical conversation data is classified, and a specific implementation of obtaining the category of each historical question-answer pair is: the historical query text in each historical question-answer pair is used as the input of the first labeling model, and the category label is output, where the category label indicates the category of the historical question-answer pair.
[0012] Optionally, the first labeling model is trained based on the transformer-based BERT model, and the historical query text is used as the input of the first labeling model, and the category corresponding to the historical query text is output. For example, the output category can be one of the knowledge category, scenario category, creation category, chat category, and role-playing category.
[0013] By determining the type of question-answer pair based on the query text, we can quickly filter out low-value conversation data from redundant historical conversation data, increase the amount of prompt information sent to the LLM, and achieve long-term conversation memory optimization.
[0014] It's understandable that low-value conversation data refers to conversation data with low generalizability. For example, conversation data about the weather, due to its timeliness, has low generalizability. Knowledge-based conversation data, such as "Who were the Seven Warring States?" and "The Seven Warring States were...", has high generalizability and can be considered high-value conversation data. Conversation data with low generalizability cannot be used in similar conversations, so its value is low. Conversation data with high generalizability, on the other hand, can be used to assist LLM in generating answers when similar conversations occur again, so its value is high.
[0015] In another possible implementation, historical question-answer pairs of the target category are topic-tagged, and a specific implementation of obtaining the topic of the historical question-answer pairs of the target category is: the historical question-answer pairs of the target category are used as the input of the second labeling model, and the topic label is output, where the topic label indicates the topic information of the historical question-answer pairs of the target category.
[0016] Optionally, the first labeling model and the second labeling model can be the same labeling model, for example, both can be trained by the transformer-based BERT model, and the query text is used as the input of the labeling model, and the category corresponding to the query text is output; the question and answer pair is used as the input of the labeling model, and the topic corresponding to the question and answer pair is output.
[0017] In one example, the categories of the historical question-answer pairs include knowledge and scenario categories, as well as one or more of creation, chatting, and role-playing categories; and the target category includes knowledge and / or scenario categories.
[0018] In another possible implementation, the initial memory chain data includes a topic node and several knowledge nodes; wherein the topic node records topic information, and the topic information indicates the topics of the several knowledge nodes; the several knowledge nodes record historical question-answer pairs with the same topic in the order of the generation time of the historical question-answer pairs.
[0019] When building a memory chain, memory chain data is constructed around memory node topics of similar or identical topics to achieve fast search and reduction of similar memory chain data.
[0020] In another possible implementation, obtaining the plurality of memory chain data based on knowledge extraction of the historical conversation data further includes: performing reduction processing on the plurality of initial memory chain data corresponding to the plurality of historical conversation data to obtain the plurality of first memory chain data.
[0021] For example, the initial memory chain data with similar paths are reduced, including operations such as expansion, deletion, and knowledge update, to generate new memory chain data.
[0022] In another possible implementation, each memory chain data in the multiple memory chains is stored in a chain storage structure. The chain storage structure can effectively store high-value knowledge extracted from historical conversation data, thereby helping LLM better select appropriate memories.
[0023] In another possible implementation, based on the query text, a specific implementation of retrieving the target memory chain data from the memory chain database is: retrieving historical query texts similar or identical to the query text from the memory chain database to obtain the target historical query text; recalling the memory chain data where the target historical query text is located to obtain the target memory chain data.
[0024] Optionally, the target memory chain data can be obtained by searching the memory chain database through vector search or inverted search. This application does not limit the specific search method used, and you can choose a suitable search method for search as needed.
[0025] In another possible implementation, the question-answering method based on a large language model provided by the present application also includes: displaying summary data of the target memory chain data in a dialog box, the summary data including summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded by each knowledge node.
[0026] For example, in the question-and-answer interface, the memory chain search results are displayed in a highlighted form, that is, the summary of the target memory chain data is highlighted, so that answers can be generated with basis and the explainability of answer generation is increased.
[0027] In another possible implementation, the question-answering method based on a large language model provided in the present application also includes obtaining feedback information of the answer output by the large language model, where the feedback information includes a target segment selected in the answer and feedback on the target segment; and adjusting the answer based on the target segment and the feedback.
[0028] Through the feedback mechanism, the answers generated by the large language model can be corrected in real time according to the user's intention, making the answers more accurate. At the same time, it allows the selection of some fragments in the answer, and feedback on these fragments is provided to achieve sentence-level error feedback.
[0029] In another possible implementation, a specific implementation of obtaining feedback information of the answer output by the large language model is: detecting that a target segment is selected, displaying multiple operators for the target segment; obtaining the target segment and the target operator, where the target operator is the operator selected from the multiple operators.
[0030] In this possible implementation, the present application defines multiple operation elements. After the user selects the target fragment with problems in the answer (for example, the user selects the fragment that the user thinks is problematic by sliding), multiple operation elements pop up for the selected target fragment. The user determines the direction in which the target fragment needs to be improved by selecting the operation element.
[0031] Optionally, the multiple operation elements include one or more of correction, rewriting, expansion, guidance, and drop-down arrow; wherein the drop-down arrow is used to display an input box for the user to enter feedback when selected.
[0032] Exemplarily, the user selects the correction operation element, which indicates that the user believes that part of the description of the selected target segment in the answer is incorrect and needs to be partially corrected; the user selects the rewrite operation element, which indicates that the user believes that all the descriptions of the selected target segment in the answer are incorrect and need to be regenerated; the user selects the expansion operation element, which indicates that the user believes that the description of the selected target segment in the answer is too simple and needs to be expanded, such as giving the reason why the target segment is described in this way; the user selects the guidance operation element, which indicates that the user believes that the description of the selected target segment in the answer is not detailed enough and needs to be described in more detail based on the target segment; the user selects the drop-down menu operation element, which indicates that the user believes that the several given operation elements (i.e., correction, rewrite, expansion and guidance) cannot well describe the problems of the target segment. At this time, the user clicks the drop-down menu and a text box pops up for the user to enter a description of the problems of the target segment.
[0033] In another possible implementation, the question-answering method based on a large language model provided by the present application further includes recording error notes, which include feedback information and adjusted answers.
[0034] In another possible implementation, the question-answering method based on a large language model provided by the present application also includes, after the current conversation ends, retrieving third memory chain data from the memory chain database based on the error notes, where the third memory chain data contains error node data; and updating the error node data based on the error notes.
[0035] Through user feedback recorded in error notes, knowledge updates are performed on error nodes in the memory chain data to ensure the correctness of the memory chain data.
[0036] In another possible implementation, the question-answering method based on a large language model provided in the present application also includes, after the current conversation ends, performing knowledge extraction on the conversation data of the current conversation to obtain knowledge node data corresponding to the current conversation data; based on the large language model, determining second memory chain data from the memory chain database, where the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current conversation data; and merging the knowledge node data corresponding to the current conversation data and the second memory chain data to obtain third memory chain data.
[0037] By merging the new knowledge node data newly extracted from the current conversation data with the memory chain data with similar viewpoints, for example, after extracting the new knowledge node, the memory chain data related to its topic is retrieved from the memory chain database, and the knowledge node is added to the end of the memory chain data to form new memory chain data, thereby further reducing redundant data.
[0038] In another possible implementation, the question-answering method based on a large language model provided in the present application also includes determining fourth memory chain data from a memory chain database based on the large language model, where the fourth memory chain data contains knowledge node data that is inconsistent with the knowledge node data corresponding to the current conversation data; verifying the inconsistent knowledge node data to obtain correct knowledge node data; and obtaining fifth memory chain data based on the correct knowledge node data and the third memory chain data.
[0039] For example, the logical reasoning ability of LLM can be used to determine whether the new knowledge nodes extracted from the current conversation data are contradictory to the old knowledge nodes in the memory chain database. If so, they are verified in conjunction with the search engine to retain the knowledge nodes with correct opinions. This can reduce the illusion of contradiction in knowledge reduction.
[0040] In a second aspect, the present application provides a question-and-answer device based on a large language model, comprising an acquisition module, a retrieval module, a prompt text generation module, and an answer generation module, wherein the acquisition module is used to acquire the current dialogue state, which includes the query text in the current dialogue; the retrieval module is used to retrieve target memory chain data from a memory chain database based on the query text, the memory chain database including multiple memory chain data, and the multiple memory chain data are obtained based on knowledge extraction of historical dialogue data; the prompt text generation module is used to obtain a prompt text based on the query text and the target memory chain data; the answer generation module is used to use the prompt text as input to the large language model and output the answer corresponding to the query text.
[0041] In one possible implementation, the present application provides a question-answering device based on a large language model, which also includes a knowledge extraction module and a memory chain construction module, wherein the knowledge extraction module is used to classify each historical question-answer pair in each historical dialogue data to obtain the category of each historical question-answer pair; select historical question-answer pairs of a target category from multiple historical question-answer pairs in the dialogue data, and the generalization of the historical question-answer pairs of the target category is greater than the generalization of historical question-answer pairs of other categories; perform topic tagging on the historical question-answer pairs of the target category to obtain the topic of the historical question-answer pairs of the target category; the memory chain construction module is used to construct initial memory chain data corresponding to each historical dialogue data based on historical question-answer pairs with the same topic.
[0042] In another possible implementation, the knowledge extraction module is specifically configured to use the historical query text in each historical question-answer pair as input to the first labeling model, and output a category label, where the category label indicates the category of the historical question-answer pair.
[0043] In another possible implementation, the knowledge extraction module is specifically configured to take the historical question-answer pairs of the target category as input to the second labeling model, and output a topic label, where the topic label indicates topic information of the historical question-answer pairs of the target category.
[0044] In another possible implementation, the categories of historical question-answer pairs include knowledge and scenario categories, as well as one or more of creation, chat, and role-playing categories; and the target category includes knowledge and / or scenario categories.
[0045] In another possible implementation, the initial memory chain data includes a topic node and several knowledge nodes; wherein the topic node records topic information, and the topic information indicates the topics of the several knowledge nodes; the several knowledge nodes record historical question-answer pairs with the same topic in the order of the generation time of the historical question-answer pairs.
[0046] In another possible implementation, the memory chain building module is further configured to reduce the multiple initial memory chain data corresponding to the multiple historical conversation data to obtain multiple first memory chain data.
[0047] Optionally, each memory chain data in the plurality of memory chain data is stored in a chain storage structure.
[0048] In another possible implementation, the retrieval module is specifically used to retrieve historical query texts similar or identical to the query text from the memory chain database to obtain target historical query texts; and recall the memory chain data where the target historical query texts are located to obtain target memory chain data.
[0049] In another possible implementation, the question-and-answer device based on a large language model provided in the present application also includes a display module, which is used to display summary data of the target memory chain data in a dialog box, and the summary data includes summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded by each knowledge node.
[0050] In another possible implementation, the acquisition module is also used to obtain feedback information of the answer output by the large language model, and the feedback information includes the target segment selected in the answer and the feedback on the target segment; the question-answering device based on the large language model provided in this application also includes an answer adjustment module, which is used to adjust the answer based on the target segment and the feedback.
[0051] In another possible implementation, the display module is further used to detect that the target segment is selected and display multiple operators for the target segment; the acquisition module is specifically used to acquire the target segment and the target operator, and the target operator is the selected operator among the multiple operators.
[0052] In another possible implementation, the multiple operation elements include one or more of correction, rewriting, expansion, guidance, and a drop-down arrow; wherein the drop-down arrow is used to display an input box for the user to enter feedback when selected.
[0053] In another possible implementation, the question-answering device based on the large language model provided by the present application also includes an error note recording module, which is used to record error notes, and the error notes include feedback information and adjusted answers.
[0054] In another possible implementation, the retrieval module is also used to retrieve the third memory chain data from the memory chain database based on the error notes after the current conversation ends, and the third memory chain data contains error node data; the question and answer device based on the large language model provided in this application also includes a node update module, which is used to update the error node data based on the error notes.
[0055] In another possible implementation, the knowledge extraction module is further used to extract knowledge from the conversation data of the current conversation after the current conversation ends, so as to obtain knowledge node data corresponding to the current conversation data; the question-answering device based on the large language model provided in the present application also includes a determination module, which is used to determine the second memory chain data from the memory chain database based on the large language model, where the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current conversation data; the memory chain construction module is further used to merge the knowledge node data corresponding to the current conversation data with the second memory chain data to obtain third memory chain data.
[0056] In another possible implementation, the determination module is further used to determine fourth memory chain data from the memory chain database based on the large language model, where the fourth memory chain data contains knowledge node data that is inconsistent with the knowledge node data corresponding to the current dialogue data; the question-answering device based on the large language model provided in this application also includes a verification module, which is used to verify the inconsistent knowledge node data and obtain the correct knowledge node data; the memory chain construction module is used to obtain the fifth memory chain data based on the correct knowledge node data and the third memory chain data.
[0057] In a third aspect, an embodiment of the present application provides a computing device comprising a memory and a processor, wherein the memory stores instructions, and when the instructions are executed by the processor, the method described in the first aspect is implemented.
[0058] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the first aspect is implemented.
[0059] In a fifth aspect, an embodiment of the present application further provides a computer program or a computer program product, which includes instructions that, when executed, cause a computer to execute the method described in the first aspect.
[0060] In a sixth aspect, an embodiment of the present application further provides a chip comprising at least one processor and a communication interface, wherein the processor is used to execute the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] In order to more clearly illustrate the technical solutions of the multiple embodiments disclosed in this specification, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings described below are only the multiple embodiments disclosed in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0062] The following is a brief introduction to the drawings required for describing the embodiments or prior art.
[0063] FIG1 is a schematic diagram of a dialogue interaction of a DST-based dialogue system;
[0064] FIG2 is a schematic diagram of a feedback interface in related art 2;
[0065] FIG3 is a schematic diagram of a feedback interface in related art 3;
[0066] FIG4 shows a system architecture diagram for implementing the large language model-based question-answering method provided in an embodiment of the present application;
[0067] FIG5 is a flow chart of a question-answering method based on a large language model provided in an embodiment of the present application;
[0068] FIG6 shows a schematic diagram of a memory chain;
[0069] FIG7 shows a flow chart of an implementation of a conversational marking method;
[0070] FIG8 shows a schematic diagram of a question-answering interface to which the question-answering method based on a large language model provided in an embodiment of the present application is applied;
[0071] FIG9 shows a schematic diagram of real-time correction of answers based on user feedback;
[0072] FIG10 is a schematic diagram showing how the large language model-based question-answering method provided in an embodiment of the present application processes conversation data after the conversation ends;
[0073] FIG11 is a schematic diagram of the structure of a question-answering device based on a large language model provided in an embodiment of the present application;
[0074] FIG12 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0075] The term "and / or" as used herein describes an association relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " in this document indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0076] Throughout the specification and claims herein, the terms "first" and "second" are used to distinguish between different objects, rather than to describe a specific order of objects. For example, "first memory chain data" and "second memory chain data" are used to distinguish between different memory chain data, rather than to describe a specific order of the memory chain data.
[0077] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0078] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0079] To facilitate understanding of the solutions of the embodiments of the present application, the technical terms involved in this document are first explained below.
[0080] Large language models (LLMs) are artificial intelligence models designed to understand and generate human language. They are trained on large amounts of text data and can perform a wide range of tasks.
[0081] Dialogue systems: A dialogue system is a computer system that simulates humans and aims to engage in coherent conversations with them. Due to the inherent complexity of natural language, dialogue systems involve numerous natural language processing (NLP) subtasks. They enable machines to understand and process human language through conversation.
[0082] Session: In a dialogue system, a session typically refers to what a user or system says during an interaction. It is the basic unit in a dialogue system, used to represent the information conveyed by a user or system during an interaction.
[0083] Prompt: A prompt is an input or query provided by the user to the LLM to guide the model to produce a specific response. Effective prompt design is crucial to achieving the expected results of the LLM.
[0084] Related art LLMs have various issues with their use of historical conversation data and feedback mechanisms. For example, related art 1 utilizes short-term conversation memory. When a user engages in a conversation with a trained LLM, the LLM temporarily memorizes the user's input and its generated output to predict subsequent outputs. However, once the model has finished generating output, it "forgets" the previous user input and output. This approach works by storing a summary or vector of the historical conversation and retrieving the top-K relevant memories when needed. It also includes a Dialogue State Tracking (DST) module, a key module in task-oriented dialogue systems (TODS). Its goal is to monitor user goals hidden in the conversation history and represent them as a dialogue state consisting of a series of domain, slot, and value triplets. The DST maintains the dialogue system state (the value and probability corresponding to each slot) and updates the dialogue state based on the current conversation turn. This helps the system better understand user intent and respond accordingly. As shown in Figure 1, DST uses a tuple consisting of slot values and probabilities to identify the conversation state. It then uses historical conversation information to modify the query and generate the conversation state for the current round. In Figure 1, U represents the user and S represents the dialogue system. The top row illustrates the conversation process, while the bottom row corresponds to the conversation states of different rounds.
[0085] However, this technical solution does not perform knowledge extraction on historical conversation data, but directly vectorizes the original conversation data. Using vectors to encode all conversation data will lead to information entanglement between different topic data, affecting the semantic accuracy of vector representation. In addition, slot data generation is difficult and difficult to implement. The generation of slot data relies on complete knowledge, which makes it difficult to implement in practice.
[0086] Related technology 2 is a case in which after the conversation ends, users are provided with a way to evaluate the results of the conversation, such as buttons for "good review" and "bad review" (see Figure 2), but they are unable to write comments or regenerate answers based on user feedback.
[0087] In this technical solution, after the conversation ends, thumbs-down or thumbs-up cannot affect the model's answer generation logic, and answer correction requires the user to restart a new round of conversation, lacking the ability to provide rapid feedback and real-time correction of answers.
[0088] Related technology three, after the conversation ends, provides a text input box for users to fill in their feedback on the results of the conversation (see Figure 3), but it is unable to locate errors at the sentence level.
[0089] In this technical solution, the feedback method is cumbersome and highly dependent on the user's ability to describe the problem. The user can only feedback rough error information but cannot feedback the error at the sentence granularity, making it difficult for the search system to fix the problem in a targeted manner.
[0090] To address the above issues, the present application provides a question-answering method and device based on a large language model. By extracting knowledge from historical conversation data, the extracted knowledge points are processed to obtain memory chain data, and the memory chain data is stored in a memory chain database. When a user performs a query or question and triggers a similar memory path, the memory chain data of the similar memory path is recalled from the memory chain database to assist the large language model in generating a more comprehensive answer.
[0091] The technical solution of the present application is further described in detail below through the accompanying drawings and examples.
[0092] FIG4 shows a system architecture diagram that can implement the question-answering method based on a large language model provided in an embodiment of the present application.
[0093] In the offline stage, knowledge is extracted from each historical conversation data, the extracted knowledge points are processed to obtain memory chain data, and the memory chain data is written into the memory chain database.
[0094] In the online question-and-answer stage, the query text is obtained, and the memory chain data corresponding to the query text is retrieved from the memory chain database based on the query text. A prompt text is generated based on the memory chain data and the query text. The large language model generates an answer based on the prompt text, so that the memory chain data can assist the large language model in generating more comprehensive and accurate answers.
[0095] FIG5 is a flow chart of a large language model-based question-answering method provided in an embodiment of the present application. The method can be executed by any device, equipment, platform, or device cluster with computing capabilities. The embodiment of the present application does not specifically limit the specific computing device for executing the method, and an appropriate computing device can be selected as needed. As shown in FIG5 , the large language model-based question-answering method includes at least steps S501 to S504.
[0096] In step S501, the current dialog state is obtained, where the current dialog state includes the query text in the current dialog.
[0097] The user may interact with the computing device in any suitable manner and input query text into the dialog box. For example, the user may interact with the computing device through voice, image, text, etc. The computing device converts the non-text input into text and inputs it into the dialog box.
[0098] When it is detected that the current conversation includes a preset number of query texts, the current conversation state is obtained, which includes the query texts in the current conversation. For example, the preset number of query texts can be set according to actual needs. For example, the preset number of query texts can be set to 1. That is, when it is detected that the user has entered a query text, the current conversation state is obtained, and the user-entered query text is obtained. At this point, it can be determined that the user has asked question 1, for example, question 1 can be: "Why is it called the Warring States Period?" The memory chain database is then searched based on question 1 to obtain target memory chain data (hereinafter referred to as memory chain for ease of description). When the user asks the question for the first time, the target memory chain LLM can be used to generate a more comprehensive answer.
[0099] Of course, the preset number of items can also be set to other values, for example, 2. That is, when the user enters two query texts, the current conversation state is obtained, and the two query texts entered by the user are obtained. At this point, it can be determined that the user has asked questions 1 and 2. For example, question 1 can be: "Why is it called the Warring States Period?" and question 2 can be: "Who are the Seven Heroes of the Warring States Period?" Then, based on questions 1 and 2, the memory chain database is searched to obtain the target memory chain. The retrieved target memory chain is more accurate, and the target memory chain LLM generates a more comprehensive answer.
[0100] In another example, the current conversation state may also include the query text in the current conversation and the answer corresponding to the query text (i.e., the response output of the LLM). For example, when it is detected that the current conversation has gone through a preset number of rounds, the current conversation state is obtained, and the multi-round conversation information included in the current conversation is obtained, that is, the query text in the current conversation and the answer corresponding to the query text. The preset number of rounds can be set as needed. For example, the preset number of rounds can be 2. When it is detected that the current conversation has gone through 2 rounds, the current conversation state is obtained, and the query text 1 and the answer 1 corresponding to the query text 1 of the current conversation, as well as the query text 2 and the answer 2 corresponding to the query text 2 are obtained. In other words, question-answer pair 1 (including question 1 and answer 1) and question-answer pair 2 (including question 2 and answer 2) are obtained. Then, based on question-answer pair 1 and question-answer pair 2, the memory chain database is searched to obtain the target memory chain, and a more comprehensive answer is generated with the help of the target memory chain LLM.
[0101] In step S502, based on the query text, target memory chain data is retrieved from the memory chain database.
[0102] The memory chain database includes multiple memory chains, which are derived from knowledge extraction from historical conversation data. For example, as shown in Figure 4, after the current conversation ends, knowledge is automatically extracted from the current conversation data. The extracted knowledge is processed (for example, filtering for highly generalizable knowledge) to construct a memory chain.
[0103] As you can understand, the Memory Chain is a storage data structure that reprocesses historical conversation data. Its primary goal is to refine the conversation history data and optimize the quality of LLM prompt text. Optionally, the Memory Chain data is stored in a chain storage structure, which facilitates retrieval and helps quickly select appropriate memories for LLM.
[0104] The nodes in the memory chain store highly generalizable conversational knowledge. To obtain this knowledge, this application proposes a conversation tagging technology. After the conversation ends, the query texts in each round of conversation in the conversation data are classified, and the query texts in the highly generalizable categories and the corresponding answers are extracted to construct a memory chain.
[0105] In another example, in order to more effectively organize conversation knowledge, conversation tagging technology also includes topic tagging of question-answer pairs in conversation data, obtaining the topic of each question-answer pair, and constructing conversation knowledge with the same topic in highly generalized conversation knowledge into a memory chain.
[0106] In real-life conversations, people often discuss a topic. Therefore, building a memory chain around a topic is in line with people's conversational habits and helps the memory chain assist LLM in generating answers that meet people's expectations.
[0107] A memory chain consists of topic nodes and knowledge nodes. The topic node records the theme of the memory chain, while the knowledge node records conversational knowledge related to the same theme. The topic node is a label derived from the conversation data, representing the topic of the discussion. For example, if a conversation revolves around the Warring States Period in history, the topic node of the memory chain constructed for that conversation data would be "History - Warring States Period," and the knowledge node would record the question and answer pairs under that theme.
[0108] FIG6 shows a schematic diagram of a memory chain.
[0109] As shown in Figure 6, a memory chain consists of topic nodes and conversational knowledge related to that topic. For example, a conversation revolves around the topic "History - Warring States" and spans multiple rounds. In the first round, the user enters the query "Why is it called the Warring States?", and the LLM generates the answer "The origin of the Warring States" based on this query. In the second round, the user continues by entering the query "Who were the Seven Heroes of the Warring States?", and the LLM generates the answer "Introduction to the Seven Heroes..." based on this query. In the third round, the user continues by entering the query "How did the Qin Dynasty conquer the six states?", and the LLM generates the answer "The process and results of the Qin Dynasty's conquest of the six states..." Based on this conversation data, we can conclude that the topic is "History - Warring States." Extracting conversational knowledge from this conversation data yields three question-answer pairs: Q&A pair 1: "Q: Why is it called the Warring States? A: The origin of the Warring States...", Q&A pair 2: "Q: Who were the Seven Heroes of the Warring States? A: Introduction to the Seven Heroes...", and Q&A pair 3: "Q: How did the Qin Dynasty conquer the six states? A: The process and results of the Qin Dynasty's conquest of the six states..." Based on the topic of the summarized conversation data and the conversation knowledge under the topic (i.e., question-answer pairs), a memory chain is constructed (see Example 1 in Figure 6).
[0110] In one example, the knowledge nodes in the memory chain are constructed in the order of the question-answer pairs' generation. For example, in a conversation involving multiple rounds of conversations, the user enters the query "What models does B have?", and the LLM generates the answers "B1, B2, B3..." based on this query. In the second round, the user continues by entering the query "What is the price / performance ratio of B3?", and the LLM generates the answers "B3's environmental performance and performance..." based on this query. Knowledge extracted from this conversation data yields two question-answer pairs: Q&A pair 1, "Q: What models does B have? A: B1, B2, B3...", and Q&A pair 2, "Q: What is the price / performance ratio of B3? A: B3's environmental performance and performance...". Since Q&A pair 1 was generated before Q&A pair 2, the knowledge nodes for Q&A pair 1 precede the knowledge nodes for Q&A pair 2 during the memory chain construction process, as shown in Example 2 in Figure 6.
[0111] The memory chain is constructed in the order of the time when question-answer pairs are generated in the conversation, reflecting the usual user conversation logic. For example, on a certain topic, the user will first ask question 1 and then question 2. The memory chain is constructed according to the order in which the user asks questions, which preserves the user's conversation logic information. At the same time, the memory chain makes the generated answers more logical when assisting the LLM in generating answers.
[0112] Of course, in some other examples, the order of knowledge nodes in the memory chain can also be constructed with other logics. For example, in historical topics, the memory chain can be constructed according to the chronological order of historical events. For example, the knowledge nodes with earlier occurrence times of historical events are placed earlier in the order, and the knowledge nodes with later occurrence times of historical events are placed later in the order.
[0113] The following introduces the specific implementation scheme of the dialogue tagging technology and the memory chain construction scheme based on the tagging technology.
[0114] FIG7 shows a flowchart of an implementation of a conversational tagging method.
[0115] As shown in Figure 7, a complete historical conversation consists of multiple rounds of conversation, each of which generates a question-answer pair. For example, the historical conversation data in Figure 7 includes N conversations, consisting of N question-answer pairs: question1-answer1, question2-answer2, …, questionN-answerN. First, each question-answer pair in the conversation is categorized and assigned a specific category. Then, the conversations are filtered to retain those with high generalization. These highly generalized question-answer pairs are then categorized and assigned a specific topic. A memory chain is then constructed based on the topic information and the question-answer pairs.
[0116] In one example, a first labeling model can be trained and used to infer the category of the query text in a question-answer pair, thereby obtaining the category of the question-answer pair in which the query text resides. For example, for a certain historical conversation data, the query text of each conversation round is used as the input of the labeling model. The labeling model infers the category of the query text and outputs the category of the query text. For example, the category of the query text can be any of the following categories: knowledge, scenario, creative, casual chat, or role-playing. In this way, the category of each question-answer pair in the historical conversation data is obtained, and then question-answer pairs with high generalization categories are filtered out. For example, question-answer pairs in the knowledge category and / or scenario category are question-answer pairs with high generalization categories and need to be retained for subsequent memory chain construction. Question-answer pairs in the creative, casual chat, and role-playing categories are question-answer pairs with low generalization and are filtered out to streamline the data volume of the historical conversation data.
[0117] Understandably, when similar conversations recur, the answers to question-and-answer pairs with low generalizability cannot be used, thus having low value. Therefore, filtering is necessary. For example, a question-and-answer pair asking about the weather, such as "Will it rain tomorrow?" and "A: It will be cloudy tomorrow, no rain," is time-sensitive and therefore has low generalizability. When the LLM receives a similar question, such as "Will it rain tomorrow?" a few days later, it cannot use the answers from previous question-and-answer pairs. This shows that question-and-answer pairs with low generalizability cannot effectively help the LLM generate answers. Therefore, filtering out these question-and-answer pairs can reduce the amount of data in the memory chain database, increase retrieval efficiency, and thus speed up answer generation.
[0118] For question-and-answer pairs with high generalizability, the answers can be used to assist the LLM in generating answers when similar conversations occur again. Therefore, they are highly valuable and should be retained. For example, consider a knowledge-based question-and-answer pair like "Who were the Seven Heroes of the Warring States Period?" and "The Seven Heroes of the Warring States Period are..." When the LLM receives a similar question again, such as "Please introduce the Seven Heroes of the Warring States Period," it can reuse the previous answer without having to infer again, speeding up answer generation and achieving a memory function similar to that of the human brain. For example, when solving a problem they have already solved, people will directly use the remembered answer instead of re-inferring the answer, which greatly speeds up problem-solving.
[0119] After selecting highly generalizable question-answer pairs, these pairs are then topic-labeled to obtain the topics of each pair. For example, this can be accomplished by training a second labeling model and using the trained second labeling model to infer the topics of the question-answer pairs.
[0120] The second labeling model uses the question-answer pair as input and outputs the topic label for the question-answer pair. For example, the question-answer pair "Q: How many kings were there in Qin? A: There were six kings in Qin" is used as input for the labeling model. The labeling model infers the topic of the question-answer pair and outputs the topic label "History-Qin".
[0121] Optionally, the first labeling model and the second labeling model are trained separately by using a transformer-based BERT model.
[0122] In another example, the first labeling model and the second labeling model can be the same labeling model, for example, both can be trained by the transformer-based BERT model, and the query text is used as the input of the labeling model, and the category corresponding to the query text is output; the question and answer pair is used as the input of the labeling model, and the topic corresponding to the question and answer pair is output.
[0123] Through category tagging, conversation data can be quickly screened, conversation data with low generalization value can be filtered out, the amount of data can be significantly reduced, and contextual content optimization can be achieved under longer timelines. Through topic tagging, the accuracy of memory chain construction can be improved. The tagging model performs topic tagging based on deep semantic features. When building the memory chain in the next stage, the memory chain can be quickly searched and reduced based on the topic of the memory chain.
[0124] LLM is very powerful and can implement different functions based on the design of different prompt texts. Therefore, in some other examples, dialogue labeling technology can also be implemented through LLM without the need for other means (for example, by training a labeling model, saving training resources for training additional AI models).
[0125] For example, if the prompt text "Please identify the category of the following question-answer pair: question1-answer1" is input to the LLM, the LLM infers that the category of question1-answer1 is a knowledge category, and thus the category of the question-answer pair is obtained. If the prompt text "Please identify the topic of the following question-answer pair: question1-answer1" is input to the LLM, the LLM infers that the topic of question1-answer1 is "History - Qin State", and thus the topic of the question-answer pair is obtained.
[0126] For a historical conversation data, after conversation labeling, highly generalized conversation knowledge (i.e., highly generalized question-answer pairs) and conversation topics are obtained. Then, based on the highly generalized conversation knowledge and conversation topics, the initial memory chain of the historical conversation data is constructed, and then it is reduced with the memory chains with similar paths in the memory chain database to generate a new memory chain.
[0127] Similar paths can also be referred to as similar thinking paths. For example, the theme of memory chain 1 is Topic 10, and the knowledge nodes include question-answer pair 11, question-answer pair 12, question-answer pair 13, question-answer pair 14, and question-answer pair 15. The theme of memory chain 2 is Topic 10, and the knowledge nodes include question-answer pair 11 and question-answer pair 12. In this case, memory chains 1 and 2 can be referred to as memory chains with similar paths and similar thinking paths. Reducing memory chains with similar paths, for example, reducing memory chains 1 and 2, yields the following reduced memory chain: Topic 10 - Question-answer pair 11 - Question-answer pair 12 - Question-answer pair 13 - Question-answer pair 14 - Question-answer pair 15.
[0128] For example, memory chain 3 is: Topic 20 - Question and Answer Pair 21 - Question and Answer Pair 22 - Question and Answer Pair 23; memory chain 4 is: Topic 20 - Question and Answer Pair 22 - Question and Answer Pair 23 - Question and Answer Pair 24 - Question and Answer Pair 25. It can be determined that memory chains 3 and 4 also have similar paths. After reducing memory chains 3 and 4, the reduced memory chain is: Topic 20 - Question and Answer Pair 21 - Question and Answer Pair 22 - Question and Answer Pair 23 - Question and Answer Pair 24 - Question and Answer Pair 25.
[0129] From the above, we can see that the conditions for judging whether two memory chains have similar paths include whether the two memory chains have the same subject nodes and whether the two memory chains have all or part of the similar or identical knowledge nodes.
[0130] By building a memory chain, we can achieve information reduction in long-term user conversations, extract high-quality common user thinking paths, and improve the comprehensiveness of answers.
[0131] Memory chains are constructed by extracting and mining knowledge from historical conversation data. The constructed memory chains are stored in a memory chain database. When the user interacts with the LLM in the conversation interface, the current conversation state is obtained, and the memory chain database is retrieved based on the current conversation state. Memory chains with similar paths are recalled as target memory chains.
[0132] As described above, the current dialog state may include a query text that is asked a preset number of times by the user, and then a target memory chain with the same or similar memory path is retrieved from the memory chain database based on the query text.
[0133] For example, when it is detected that the user enters a query text in the dialogue interface, the target memory chain is retrieved from the memory chain database based on the query text. The target memory chain assists LLM in generating a more comprehensive answer, reducing the number of interaction rounds between the user and LLM.
[0134] For another example, when it is detected that the user enters query text twice in the dialogue interface, that is, asks the second question, the target memory chain is retrieved from the memory chain database based on the query text entered twice by the user. In this way, the retrieved target memory chain is more accurate.
[0135] A target memory chain with a similar path to the query text can be understood as one in which the knowledge nodes in the target memory chain have the same or similar thinking paths as the query text (similar here means that the query text and the query text on the knowledge node are close in meaning. For example, if the query text is: Will it rain tomorrow? The query text on the knowledge node is: What's the weather like tomorrow?, then the query text and the query text on the knowledge node are considered similar). For example, in the current conversation, the user asks questions 1 and 2, and multiple knowledge nodes in the memory chain have questions 1 and 2 recorded in that order. This memory chain is said to have a similar path to the query text, and it is recalled as the target memory chain. For example, in the current conversation, the user asks questions 1 and 2 to the LLM successively; memory chain 1 is: topic 1-question-answer pair 1 (including question 1-answer 1)-question-answer pair 2-question-answer pair 3-question-answer pair 4; then memory chain 1 is retrieved from the memory chain database according to questions 1 and 2, and memory chain 1 assists LLM in generating the answer corresponding to question 2, such as answer 2, answer 3 and answer 4, to generate a more comprehensive answer, without the user having to ask questions 3 and 4 again, effectively reducing the number of interaction rounds between the user and LLM and improving the user experience.
[0136] In one example, the subject of the query text may be extracted first, and then memory chains irrelevant to the subject in the memory chain database may be filtered based on the subject, the search scope may be shortened, and then the search may be performed from the memory chains under the subject to speed up the search.
[0137] It should be pointed out that the target memory chain data can be obtained by searching the memory chain database through vector search or inverted search. The specific search method used in this application is not limited, and the appropriate search method can be selected for search as needed.
[0138] When the target memory chain is not retrieved from the memory chain database, the LLM generates an answer based on conventional formal reasoning, that is, based on the query text, it infers the answer corresponding to the query text and then responds to the answer of the input query text.
[0139] In step S503, a prompt text is obtained based on the query text and the target memory chain data.
[0140] The query text and the target memory chain database are combined to generate prompt text. For example, after retrieving the target memory chain from the memory chain database, prompt text can be designed to assist the LLM in generating a more comprehensive answer. For example, if the query text is Question 1 and Question 2, and the target memory chain is Memory Chain 1, the prompt text may be: "The user asked Questions 1 and 2 in sequence. Please refer to Memory Chain 1 to generate the answer to Question 2."
[0141] By constructing memory chains, the embodiments of this application can reduce information from long-term user conversations, extract high-quality, common user thought paths, and improve the comprehensiveness of answers. Furthermore, given the limited length of LLM prompt text input, a shorter prompt text input can convey a greater amount of conversation information to the LLM. Furthermore, the memory chains formed by conversation reduction can be used to mine user interests, enabling personalized interest analysis and enhancing the LLM's ability to infer personalized answers.
[0142] In step S504, the prompt text is used as input to the large language model, and the answer corresponding to the query text is output.
[0143] Help LLMs generate more comprehensive answers by including reminder text for memory chains.
[0144] FIG8 shows a schematic diagram of a question-and-answer interface that uses the large language model-based question-and-answer method provided in an embodiment of the present application. As shown in FIG8 , when a user enters the question "Who are the Seven Kingdoms of the Warring States Period?" in the question-and-answer interface, the LLM responds to the question and outputs the corresponding answer. When a user enters the question "Who won in the end?", a memory chain with a similar path is searched from the memory chain database, and the target memory chain is recalled. The target memory chain is the memory chain shown in FIG8 : The Seven Kingdoms of the Warring States Period - The Battle of Qin's Conquest of the Six Kingdoms - Qin Shi Huang's Unification of the Six Kingdoms. This memory chain assists the LLM in generating a more comprehensive answer.
[0145] In one example, in order to increase the explainability of the generated answers, summary data of the target memory chain data is displayed in a dialog box for the answers generated by the memory chain assisted LLM. The summary data includes summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded by each knowledge node.
[0146] Furthermore, in the question-and-answer interface, the memory chain search results are displayed in a highlighted form, that is, the summary of the target memory chain data is highlighted, so that answers can be generated based on evidence and the explainability of answer generation is increased.
[0147] Continuing with FIG8 , for the answers generated by the memory chain assisted LLM, there is also a text annotation "Based on the analysis of the memory chain of historical dialogues, we will expand on the historical events of this period for you", so that users can clearly understand which answers are generated with the assistance of the memory chain.
[0148] It should be explained that this application does not make any specific limitations on the specific implementation of LLM. For example, LLM can be the Pangu large language model, the Wenyan Yixin large language model, and the ChatGPT large language model, etc. The question-answering system can select a suitable large language model as needed to implement the question-answering method based on the large language model provided in the embodiment of this application.
[0149] In some other examples, the large language model-based question-answering method provided in the examples of this application also includes a user feedback mechanism. This mechanism collects user feedback on the answers output by the large language model. For example, the feedback information may include a selected target segment in the answer and feedback on the target segment. The answer is then adjusted based on the target segment and the feedback.
[0150] That is to say, users are allowed to provide feedback on the problematic parts of some or all sentences in the answers generated by LLM, and provide feedback to the question-answering system, so that LLM can make real-time corrections to the answers based on the user feedback to ensure that the answers generated are satisfactory to the users.
[0151] Current LLMs are limited by the amount of training data, fine-tuning, and the need for greater human-like accuracy. The answers they output may contain errors, which are easily identifiable by humans. If a user discovers an error in an LLM's output, they can mark the error and provide feedback on the issue.
[0152] For another example, part of the answer output by LLM may be too brief. The user can mark this part and then provide feedback that the marked content is too brief and needs to be generated with more detailed content.
[0153] After obtaining the target segment with problems in the answer output by LLM and the feedback on the target segment, LLM can regenerate the answer based on the target segment and feedback and correct the problems in the answer in real time.
[0154] Optionally, LLM can use search engines to correct the target segment provided by the user. For example, a user enters the query "Please tell me about the history of the Qin State?" and LLM responds with the following answer: "...The Qin State had seven kings..." The user notices that the description "The Qin State had seven kings" in the LLM's output is incorrect and marks it as incorrect, providing feedback. Based on this feedback, LLM then searches online using a search engine, for example, searching for "How many kings were there in the Qin State?" and correcting the incorrect description provided by the user based on the search results, revising the answer to "...The Qin State had six kings..."
[0155] In order to facilitate users to provide feedback on the answers generated by LLM, this application also defines a variety of operators. After the user selects the target fragment with problems in the answer generated by LLM, multiple operators will pop up for the target fragment. The user can provide feedback on the problems with the target fragment by selecting the corresponding operator, reducing the tediousness of the feedback steps.
[0156] FIG9 shows a schematic diagram of real-time correction of answers based on user feedback. As shown in FIG9 , when it is detected that the user has selected (for example, by sliding the user to select) a target segment with problems in the answer generated by the LLM, multiple operation elements pop up for the target segment. The multiple operation elements include, for example, correction, rewriting, expansion, and guidance. The user feedbacks the problems with the target segment or the direction that needs improvement by selecting (for example, clicking to select) the corresponding operation element. When the target operation element selected by the user is detected, a prompt text is determined based on the target segment and the target operation element. For example, when it is detected that segment 1 is selected in the answer generated by the LLM and the correction operation element is selected, a prompt text is generated based on the selected segment 1 and the correction operation element. The prompt text can be: "There is a problem with segment 1 and correction processing is required. Please correct segment 1." Then, the LLM corrects segment 1 according to the prompt text and regenerates the answer description for the selected segment 1 in the answer.
[0157] It can be understood that when the user selects the correction operation element, it means that the user believes that part of the description of the selected target segment in the answer is incorrect and needs to be partially corrected; when the user selects the rewrite operation element, it means that the user believes that all the descriptions of the selected target segment in the answer are incorrect and a new description needs to be generated; when the user selects the expansion operation element, it means that the user believes that the description of the selected target segment in the answer is too simple and needs to be expanded, such as giving the reason why the target segment is described in this way; when the user selects the guidance operation element, it means that the user believes that the description of the selected target segment in the answer is not detailed enough and needs to be described in more detail based on the target segment.
[0158] The correction, rewriting, expansion, and guidance operation elements described in the embodiments of the present application are only examples and do not constitute a limitation on the operation elements. More or fewer operation elements can be defined according to actual needs. For example, a drop-down operation element can also be defined. The drop-down operation element is used when the user believes that the given operation elements (i.e., correction, rewriting, expansion, and guidance) cannot well describe the problems of the target segment. The user can select the drop-down operation element to pop up a text box, in which the user can describe the problems of the target segment or the directions that need improvement in the text box.
[0159] In another example, feedback can also be used to update the memory chain. For example, the user's feedback on the LLM-generated answer will be recorded in the error note. For example, the error note will record the target segment selected by the user, the feedback on the target segment (including the selected operator or the description of the opinion entered in the text box by selecting the drop-down arrow), and the answer regenerated by the LLM based on the feedback. This error note can be used in the memory chain reduction to update the memory chain. Reduction can include operations such as expanding, deleting, and updating knowledge of the memory chain.
[0160] For example, the erroneous knowledge nodes in the memory chain nodes can be updated and corrected based on the erroneous notes. Exemplarily, after the current conversation is completed, based on the erroneous notes, the memory chain database is searched to see if there are knowledge nodes with the same or similar errors as in the erroneous notes, and then the errors of the knowledge nodes are corrected. For example, the erroneous notes record that "There are seven kings of Qin in Qin State" and need to be corrected. After correction, the note content is "There are six kings of Qin in Qin State". Then, the memory chain of the "History-Qin State" theme in the memory chain database is searched to see if there is a knowledge node with "There are seven kings of Qin in Qin State" or similar "There are x kings of Qin in Qin State". If so, the memory chain is updated, and the knowledge of "There are seven kings of Qin in Qin State" on the knowledge node is updated to "There are six kings of Qin in Qin State". In this way, the knowledge accuracy of the memory chain is guaranteed, so that the correct answer can be generated in the subsequent auxiliary LLM answer generation.
[0161] Figure 10 shows a schematic diagram of how the large language model-based question-answering method provided in an embodiment of the present application processes conversation data after a conversation ends. As shown in Figure 10, after the current conversation ends, the current conversation data becomes historical conversation data. The conversation data is processed according to the method for processing historical conversation data described above, such as by tagging the conversations, extracting highly generalizable knowledge (i.e., question-answer pairs) and knowledge topics, and obtaining an initial memory chain for the conversation data. This initial memory chain includes multiple knowledge nodes, each of which records the new knowledge extracted from the conversation data. Of course, if no highly generalizable knowledge is extracted from the conversation data, no subsequent steps are required.
[0162] Then, a reduction operation is performed on the initial memory chain and the memory chains in the memory chain data. For example, topic aggregation is first performed. From the memory chain database, memory chains with the same topic as the initial memory chain in the current conversation data are retrieved and recalled. Memory chains similar to the initial memory chain are then merged with the initial memory chain and its similar memory chains.
[0163] Optionally, the LLM can be used to determine the similarity between the initial memory chain and the memory chains of the same subject in the memory chain database. For example, a prompt text can be designed to enable the LLM to determine the similarity between two memory chains. For example, the prompt text could be: "Please determine the similarity between memory chain a and memory chain b." This prompt text is input into the LLM, and the LLM responds by outputting the similarity between memory chain a and memory chain b. If the similarity between memory chain a and memory chain b is greater than a preset threshold (e.g., 0.9), memory chain a and memory chain b are determined to be similar.
[0164] Here, the meaning of "memory chain a" and "memory chain b" being similar can be understood as meaning that the knowledge viewpoints at the knowledge nodes of memory chain a and memory chain b are similar or identical. For example, memory chain a is: Topic 30 - Question and Answer Pair 31 - Question and Answer Pair 32 - Question and Answer Pair 33 - Question and Answer Pair 34; memory chain b is: Topic 30 - Question and Answer Pair 31` - Question and Answer Pair 32` - Question and Answer Pair 33`. Among these, the viewpoints of question and answer pairs 31 and 31` are similar or identical, the viewpoints of question and answer pair 32 and 32` are similar or identical, and the viewpoints of question and answer pair 33 and 33` are similar or identical. Therefore, memory chain a and memory chain b can be judged similar. Memory chain a and memory chain b can be merged to obtain a new memory chain c: Topic 30 - Question and Answer Pair 31` - Question and Answer Pair 32` - Question and Answer Pair 33` - Question and Answer Pair 34.
[0165] The meaning of similar or identical knowledge viewpoints is that the facts or content described in the answers on the knowledge nodes are similar or identical. For example, the answer on knowledge node 31 is "There were six kings in the Qin Dynasty", and the answer on knowledge node 31' is "There were six kings in the history of the Qin Dynasty, namely...", then knowledge node 31 and knowledge node 31' can be said to have similar viewpoints.
[0166] In another example, if no similar memory chain is retrieved from the memory chain database, the memory chain database is further retrieved for a data chain that contradicts the initial data chain viewpoint of the current conversation data.
[0167] LLM can be used to determine whether the initial memory chain and the memory chains on the same topic in the memory chain database have conflicting viewpoints. For example, by designing prompt text, LLM can determine whether the two memory chains have conflicting knowledge nodes. For example, the prompt text could be: "Please determine whether memory chains c and d have conflicting knowledge nodes." This prompt text is input into the LLM, and the LLM responds with an output indicating whether memory chains c and d have conflicting knowledge nodes. If there are conflicting knowledge nodes, the LLM, combined with a search engine, verifies the two conflicting knowledge nodes and retains the correct knowledge node.
[0168] The meaning of a knowledge node contradiction here can be understood as: the content described by the answer at the knowledge node is contradictory. For example, memory chain c is: Topic 40 - Question and Answer Pair 41 - Question and Answer Pair 42 - Question and Answer Pair 43; memory chain d is: Topic 40 - Question and Answer Pair 41` - Question and Answer Pair 42` - Question and Answer Pair 43`. If the answer to question and answer pair 41 is "The Qin Dynasty had six kings," and the answer to question and answer pair 41` is "The Qin Dynasty had seven kings," then question and answer pair 41 and question and answer pair 41` are considered contradictory, and memory chains c and d are therefore contradictory. If the correct answer is found through a search engine, "The Qin Dynasty had six kings," the correct knowledge node is retained, and a new memory chain f is obtained: Topic 40 - Question and Answer Pair 41` - Question and Answer Pair 42` - Question and Answer Pair 43`.
[0169] The embodiment of the present application uses a rich interactive mechanism for feedback on error notes to correct answers and historical conversations in real time based on user intentions, update error notes, and assist in conversation knowledge reduction. By merging viewpoints and identifying contradictions based on LLM, the illusion of contradiction in knowledge reduction can be reduced. The prompt framework designed in conjunction with the prompt text generator constrains LLM's generation of contradictory results. By summarizing historical query knowledge and generalizing the thinking process, LLM can be assisted in forming a block-like knowledge system, enabling LLM to not only provide fragmented answers, but also reduce large-grained knowledge gathered in massive user conversations, present it systematically, and improve professionalism and comprehensiveness.
[0170] Based on the same concept as the aforementioned embodiment of a large language model-based question-answering method, the present embodiment also provides a large language model-based question-answering device 1100. This large language model-based question-answering device 1100 can be deployed on any device, equipment, platform, or device cluster with computing capabilities to implement the large language model-based question-answering method provided in the embodiment of the present application, so as to extract knowledge from historical conversation data and construct memory chain data. The memory chain data assists the large language model in generating more comprehensive and high-quality answers, thereby improving the user experience. The large language model-based question-answering device 1100 includes units or modules for implementing each step of the large language model-based question-answering method shown in Figures 4-10.
[0171] Figure 11 is a schematic diagram of the structure of a large language model-based question-answering device provided in an embodiment of the present application. As shown in Figure 11, the large language model-based question-answering device 1100 includes at least an acquisition module 1101, a retrieval module 1102, a prompt text generation module 1103, and an answer generation module 1104. The acquisition module 1101 is used to acquire the current conversation state, which includes the query text in the current conversation; the retrieval module 1102 is used to retrieve target memory chain data from a memory chain database based on the query text. The memory chain database includes multiple memory chain data, each of which is obtained by knowledge extraction from historical conversation data; the prompt text generation module 1103 is used to obtain a prompt text based on the query text and the target memory chain data; and the answer generation module 1104 is used to use the prompt text as input to the large language model and output an answer corresponding to the query text.
[0172] In one possible implementation, the present application provides a question-answering device 1100 based on a large language model, which also includes a knowledge extraction module 1105 and a memory chain construction module 1106, wherein the knowledge extraction module 1105 is used to classify each historical question-answer pair in each historical dialogue data to obtain the category of each historical question-answer pair; select historical question-answer pairs of a target category from multiple historical question-answer pairs in the dialogue data, and the generalization of the historical question-answer pairs of the target category is greater than the generalization of historical question-answer pairs of other categories; perform topic tagging on the historical question-answer pairs of the target category to obtain the topic of the historical question-answer pairs of the target category; the memory chain construction module 1106 is used to construct initial memory chain data corresponding to each historical dialogue data based on historical question-answer pairs with the same topic.
[0173] In another possible implementation, the knowledge extraction module 1105 is specifically configured to use the historical query text in each historical question-answer pair as the input of the first labeling model, and output a category label, where the category label indicates the category of the historical question-answer pair.
[0174] In another possible implementation, the knowledge extraction module 1105 is specifically configured to take the historical question-answer pairs of the target category as input to the second labeling model and output a topic label, where the topic label indicates topic information of the historical question-answer pairs of the target category.
[0175] In another possible implementation, the categories of historical question-answer pairs include knowledge and scenario categories, as well as one or more of creation, chat, and role-playing categories; and the target category includes knowledge and / or scenario categories.
[0176] In another possible implementation, the initial memory chain data includes a topic node and several knowledge nodes; wherein the topic node records topic information, and the topic information indicates the topics of the several knowledge nodes; the several knowledge nodes record historical question-answer pairs with the same topic in the order of the generation time of the historical question-answer pairs.
[0177] In another possible implementation, the memory chain building module 1106 is further configured to perform reduction processing on the multiple initial memory chain data corresponding to the multiple historical conversation data to obtain multiple first memory chain data.
[0178] Optionally, each memory chain data in the plurality of memory chain data is stored in a chain storage structure.
[0179] In another possible implementation, the retrieval module 1102 is specifically used to retrieve historical query texts that are similar or identical to the query text from the memory chain database to obtain a target historical query text; and recall the memory chain data where the target historical query text is located to obtain the target memory chain data.
[0180] In another possible implementation, the question-and-answer device 1100 based on a large language model provided in the present application also includes a display module 1107, which is used to display summary data of the target memory chain data in a dialog box, and the summary data includes summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded by each knowledge node.
[0181] In another possible implementation, the acquisition module 1101 is also used to obtain feedback information of the answer output by the large language model, and the feedback information includes the target segment selected in the answer and the feedback on the target segment; the question-answering device 1100 based on the large language model provided in this application also includes an answer adjustment module 1108, which is used to adjust the answer based on the target segment and the feedback.
[0182] In another possible implementation, the display module 1107 is further used to detect that the target segment is selected and display multiple operators for the target segment; the acquisition module is specifically used to acquire the target segment and the target operator, and the target operator is the operator selected from the multiple operators.
[0183] In another possible implementation, the multiple operation elements include one or more of correction, rewriting, expansion, guidance, and a drop-down arrow; wherein the drop-down arrow is used to display an input box for the user to enter feedback when selected.
[0184] In another possible implementation, the question-answering device 1100 based on a large language model provided in the present application further includes an error note recording module 1109, which is used to record error notes, and the error notes include feedback information and adjusted answers.
[0185] In another possible implementation, the retrieval module 1102 is also used to retrieve the third memory chain data from the memory chain database based on the error notes after the current conversation ends, and the third memory chain data contains error node data; the question-and-answer device based on the large language model provided in this application also includes a node update module, which is used to update the error node data based on the error notes.
[0186] In another possible implementation, the knowledge extraction module 1105 is further used to extract knowledge from the conversation data of the current conversation after the current conversation ends, to obtain knowledge node data corresponding to the current conversation data; the question-answering device 1100 based on the large language model provided in this application also includes a determination module 1110, which is used to determine second memory chain data from the memory chain database based on the large language model, where the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current conversation data; the memory chain construction module is further used to merge the knowledge node data corresponding to the current conversation data with the second memory chain data to obtain third memory chain data.
[0187] In another possible implementation, the determination module 1110 is further used to determine fourth memory chain data from the memory chain database based on the large language model, where the fourth memory chain data contains knowledge node data that is inconsistent with the knowledge node data corresponding to the current dialogue data; the question-answering device 1100 based on the large language model provided in this application also includes a verification module 1111, which is used to verify the inconsistent knowledge node data and obtain the correct knowledge node data; the memory chain construction module is used to obtain the fifth memory chain data based on the correct knowledge node data and the third memory chain data.
[0188] According to the embodiment of the present application, the question-answering device 1100 based on a large language model can correspond to executing the method described in the embodiment of the present application, and the above-mentioned and other operations and / or functions of each module in the question-answering device 1100 based on a large language model are respectively for realizing the corresponding processes of each method in Figures 4-10. For the sake of brevity, they will not be repeated here.
[0189] It should be explained that the question-answering device 1100 based on a large language model adopted in the embodiment of the present application can be implemented by software or by hardware. For example, the question-answering device based on a large language model can act on the question-answering system as a plug-in to implement the question-answering method based on a large language model provided in the embodiment of the present application.
[0190] An embodiment of the present application also provides a computing device, including at least one processor, a memory, and a communication interface, wherein the processor is configured to execute the methods described in Figures 4-10.
[0191] FIG12 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application.
[0192] As shown in Figure 12, the computing device 1200 includes at least one processor 1201, a memory 1202, and a communication interface 1203. The processor 1201, the memory 1202, and the communication interface 1203 are communicatively connected, and the communication connection can be achieved through a wired manner (such as a bus) or a wireless manner. The communication interface 1203 is used to send and / or receive data sent by other devices; the memory 1202 stores computer instructions, and the processor 1201 executes the computer instructions to execute the question-answering method based on the large language model in the aforementioned method embodiment, so as to realize the construction of memory chain data by extracting knowledge from historical conversation data. The memory chain data assists the large language model in generating better answers and improving the user experience.
[0193] It should be understood that in the embodiment of the present application, the processor 1201 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0194] The memory 1202 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1201. The memory 1202 may also include a nonvolatile random access memory.
[0195] The memory 1202 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0196] It should be understood that the computing device 1200 according to the embodiment of the present application can execute the method shown in Figures 4-10 in the embodiment of the present application. The detailed description of the implementation of the method is given above and will not be repeated here for the sake of brevity.
[0197] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the above-mentioned method is implemented.
[0198] An embodiment of the present application provides a chip, which includes at least one processor and an interface, wherein the at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the method mentioned above.
[0199] An embodiment of the present application provides a computer program or a computer program product, which includes instructions. When the instructions are executed, the computer is caused to perform the above-mentioned method.
[0200] Those skilled in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0201] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0202] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of this application. It should be understood that the above description is only the specific implementation methods of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application should be included in the scope of protection of this application.
Claims
1. A question-answering method based on a large language model, characterized in that: include: Acquire a current dialog state, wherein the current dialog state includes a query text in the current dialog; Based on the query text, target memory chain data is retrieved from a memory chain database, where the memory chain database includes a plurality of memory chain data, and the plurality of memory chain data are obtained based on knowledge extraction of historical conversation data; Based on the query text and the target memory chain data, obtaining a prompt text; The prompt text is used as the input of the large language model, and the answer corresponding to the query text is output.
2. The method according to claim 1, characterized in that Each memory chain data in the plurality of memory chain data is constructed by the following steps: Classifying each historical question-answer pair in each historical conversation data to obtain a category of each historical question-answer pair; Selecting a historical question-answer pair of a target category from a plurality of historical question-answer pairs in the dialogue data, wherein the generalizability of the historical question-answer pairs of the target category is greater than the generalizability of the historical question-answer pairs of other categories; Performing topic tagging on historical question-answer pairs of the target category to obtain topics of the historical question-answer pairs of the target category; Based on historical question-answer pairs with the same topic, initial memory chain data corresponding to each historical dialogue data is constructed.
3. The method according to claim 2, characterized in that The classifying each historical question-answer pair in each historical conversation data to obtain the category of each historical question-answer pair includes: The historical query text in each historical question-answer pair is used as the input of the first labeling model, and a category label is output, where the category label indicates the category of the historical question-answer pair.
4. The method according to claim 2 or 3, characterized in that: The topic tagging of the historical question-answer pairs of the target category to obtain the topic of the historical question-answer pairs of the target category includes: The historical question-answer pairs of the target category are used as input of the second labeling model, and a topic label is output, where the topic label indicates topic information of the historical question-answer pairs of the target category.
5. The method according to any one of claims 2 to 4, characterized in that: The categories of the historical question-answer pairs include knowledge category and scenario category, and one or more of creation category, small talk category and role-playing category; The target category includes the knowledge category and / or the scenario category.
6. The method according to any one of claims 2 to 5, characterized in that: The initial memory chain data includes a subject node and several knowledge nodes; wherein the subject node records subject information, and the subject information indicates the subject of the several knowledge nodes; the several knowledge nodes record the historical question-answer pairs with the same subject in the order of the generation time of the historical question-answer pairs.
7. The method according to any one of claims 2 to 6, characterized in that: The plurality of memory chain data are obtained based on knowledge extraction of historical conversation data, and further include: A plurality of initial memory chain data corresponding to the plurality of historical conversation data are reduced to obtain a plurality of first memory chain data.
8. The method according to any one of claims 1 to 7, characterized in that: Each memory chain data in the plurality of memory chain data is stored in a chain storage structure.
9. The method according to any one of claims 1 to 8, characterized in that: The step of retrieving target memory chain data from a memory chain database based on the query text includes: Retrieving historical query texts similar to or identical to the query text from the memory chain database to obtain target historical query texts; The memory chain data where the target historical query text is located is recalled to obtain the target memory chain data.
10. The method according to any one of claims 1 to 9, characterized in that: Also includes: The summary data of the target memory chain data is displayed in a dialog box, wherein the summary data includes summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded in each knowledge node.
11. The method according to any one of claims 1 to 10, characterized in that: Also includes: Acquire feedback information of the answer output by the large language model, wherein the feedback information includes a target segment selected from the answer and feedback on the target segment; The answer is adjusted based on the target segment and the feedback.
12. The method according to claim 11, characterized in that The obtaining feedback information of the answer output by the large language model includes: detecting that the target segment is selected, and displaying a plurality of operation elements for the target segment; The target fragment and the target operator are obtained, where the target operator is an operator selected from the multiple operators.
13. The method according to claim 12, characterized in that The multiple operation elements include one or more of correction, rewriting, expansion, guidance, and drop-down arrow; Wherein, the drop-down arrow is used to display an input box for the user to input the feedback when it is selected.
14. The method according to any one of claims 11 to 13, characterized in that: Also includes: An error note is recorded, wherein the error note includes the feedback information and the adjusted answer.
15. The method according to claim 14, characterized in that Also includes: After the current conversation ends, based on the error note, third memory chain data is retrieved from the memory chain database, and the third memory chain data contains error node data; Based on the error note, the error node data is updated.
16. The method according to any one of claims 1 to 15, characterized in that: Also includes: After the current conversation ends, knowledge extraction is performed on the conversation data of the current conversation to obtain knowledge node data corresponding to the current conversation data; Based on the large language model, determining second memory chain data from the memory chain database, wherein the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current dialogue data; The knowledge node data corresponding to the current conversation data and the second memory chain data are merged to obtain third memory chain data.
17. The method according to claim 16, characterized in that Also includes: Based on the large language model, determining fourth memory chain data from the memory chain database, wherein the fourth memory chain data contains knowledge node data that is inconsistent with the knowledge node data corresponding to the current dialogue data; Verify the contradictory knowledge node data and obtain the correct knowledge node data; Based on the correct knowledge node data and the third memory chain data, the fifth memory chain data is obtained.
18. A question-answering device based on a large language model, characterized in that: include: An acquisition module, used to acquire a current dialog state, wherein the current dialog state includes a query text in the current dialog; A retrieval module, configured to retrieve target memory chain data from a memory chain database based on the query text, wherein the memory chain database includes a plurality of memory chain data, and the plurality of memory chain data are obtained based on knowledge extraction of historical conversation data; A prompt text generation module, used to obtain a prompt text based on the query text and the target memory chain data; The answer generation module is used to use the prompt text as the input of the large language model and output the answer corresponding to the query text.
19. A computing device comprising a memory and a processor, characterized in that: Instructions are stored in the memory, and when the instructions are executed by the processor, the method according to any one of claims 1 to 17 is implemented.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 17 is implemented.
Citation Information
Patent Citations
Intelligent dialogue method, electronic device and storage medium
CN109814831A
Text generation method, device and equipment and readable storage medium
CN115146050A
Question and answer method and device, electronic equipment and readable storage medium
CN116662518A
Knowledge question and answer method, device and equipment and storage medium
CN116680384A
Cited By
Innovation and entrepreneurship coaching question and answer matching method and system based on semantic understanding
CN120256590A
Question and answer task processing method, related device, equipment and storage medium
CN120337943A
Retrieval question and answer method and device for table, medium, equipment and program product
CN120448407A
Multi-round reasoning question answering method and system based on query graph driving
CN120687579A
Intelligent question and answer method, device and equipment, medium and program product
CN120744051A