Question and answer method and device based on large language model
By extracting knowledge of historical dialogue data and building memory chain data, the problem of information entanglement in the existing technology is solved, and the quality and user experience of the answers generated by large language models are improved.
Patent Information
- Application Number
- CN202311510489.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-10
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, the failure to effectively use historical dialogue data for knowledge refining leads to information entanglement between different topic data, affecting the semantic accuracy of vector representation, and thus affecting the quality of answers generated by large language models.
By extracting knowledge of historical dialogue data, building memory chain data and storing it in memory chain database, helping the large language model to generate answers. The specific implementation includes classifying and marking historical Q&A pairs, and filtering out highly generalized dialogue content to build memory chain nodes.
With the assistance of memory chain data, the quality and user experience of answers generated by large language models are improved, the amount of data is significantly reduced, and the context content optimization is achieved for a long time offline.
Smart Images

Figure CN119990300A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence (AI) technology, and in particular to a question-answering method and device based on a large language model. Background Art
[0002] The advent of ChatGPT is revolutionary for search. Conversational search will gradually replace traditional search (webpage relevance search) and become a new way of searching in the future. At present, the industry generally agrees that historical conversation data is an important basis for improving conversational search systems. Making full use of this data can not only clarify user intentions and generate better answers, but also expand and correct training data and continuously iterate the effects of large models.
[0003] However, in the existing technology, there are various problems in the use of historical conversation data. For example, the original conversation data is directly vectorized without knowledge extraction from the historical conversation data, and all historical conversation data is encoded using vectors, which leads to information entanglement between different topic data, affects the semantic accuracy of the vector representation, and further affects the quality of the answers generated by the large model. Summary of the invention
[0004] The embodiments of the present application provide a question-answering method based on a large language model, which constructs memory chain data by extracting knowledge from historical conversation data. The memory chain data assists the large language model in generating better answers and improving the user experience.
[0005] In a first aspect, the present application provides a question-answering method based on a large language model, including obtaining a current dialogue state, the current dialogue state including a query text in the current dialogue; based on the query text, retrieving target memory chain data from a memory chain database, the memory chain database including multiple memory chain data, the multiple memory chain data being obtained based on knowledge extraction of historical dialogue data; based on the query text and the target memory chain data, obtaining a prompt text; using the prompt text as an input of a large language model (LLM), and outputting an answer corresponding to the query text.
[0006] The question-answering method based on a large language model provided in this application constructs memory chain data by extracting knowledge from historical conversation data. The memory chain data assists the large language model in generating better answers and improving the user experience.
[0007] In a possible implementation, a specific implementation of constructing each memory chain data in the multiple memory chain data is: classifying each historical question and answer pair in each historical dialogue data to obtain the category of each historical question and answer pair; selecting historical question and answer pairs of a target category from multiple historical question and answer pairs in the dialogue data, the generalization of the historical question and answer pairs of the target category is greater than the generalization of historical question and answer pairs of other categories; topic tagging the historical question and answer pairs of the target category to obtain the topic of the historical question and answer pairs of the target category; based on the historical question and answer pairs with the same topic, constructing the initial memory chain data corresponding to each historical dialogue data.
[0008] In this possible implementation, by classifying and topic-tagging historical question-answer pairs, the conversation contents with high generalization value are screened out to form memory chain nodes, and the conversation data with low generalization value is filtered out, which significantly reduces the amount of data and realizes the optimization of contextual content under a longer timeline.
[0009] In another possible implementation, each historical question-answer pair in each historical conversation data is classified, and a specific implementation of obtaining the category of each historical question-answer pair is: using the historical query text in each historical question-answer pair as the input of the first labeling model, and outputting a category label, wherein the category label indicates the category of the historical question-answer pair.
[0010] Optionally, the first labeling model is trained by a transformer-based BERT model, and the historical query text is used as the input of the first labeling model to output a category corresponding to the historical query text. For example, the output category may be one of knowledge, scenario, creation, chat, and role-playing.
[0011] By determining the type of question-answer pairs based on the query text, it is possible to quickly filter out low-value conversation data from redundant historical conversation data, increase the amount of prompt information sent to the LLM, and achieve long-term conversation memory optimization.
[0012] It can be understood that low-value conversation data refers to conversation data with low generalization, such as conversation data about the weather, which has low generalization due to its timeliness; knowledge-based conversation data, such as "Who are the Seven Heroes of the Warring States Period?" and "The Seven Heroes of the Warring States Period are...", has high generalization and can be called high-value conversation data. Conversation data with low generalization cannot be used when similar conversations occur again, so the value is low, while conversation data with high generalization can be used to assist LLM in generating answers when similar conversations occur again, so the value is high.
[0013] In another possible implementation, historical question-answer pairs of the target category are topic-tagged, and a specific implementation of obtaining the topic of the historical question-answer pairs of the target category is: taking the historical question-answer pairs of the target category as the input of the second labeling model, and outputting the topic label, wherein the topic label indicates the topic information of the historical question-answer pairs of the target category.
[0014] Optionally, the first labeling model and the second labeling model can be the same labeling model, for example, both can be trained by the transformer-based BERT model, and the query text is used as the input of the labeling model, and the category corresponding to the query text is output; the question and answer pair is used as the input of the labeling model, and the topic corresponding to the question and answer pair is output.
[0015] In one example, the categories of the historical question-answer pairs include knowledge and scenario categories, and one or more of creation, chat, and role-playing categories; and the target category includes knowledge and / or scenario categories.
[0016] In another possible implementation, the initial memory chain data includes a topic node and several knowledge nodes; wherein the topic node records topic information, and the topic information indicates the topics of several knowledge nodes; and the several knowledge nodes record historical question-answer pairs with the same topic in the order of the generation time of the historical question-answer pairs.
[0017] When building a memory chain, the memory chain data is built around memory node topics of similar or identical topics to achieve fast search and reduction of similar memory chain data.
[0018] In another possible implementation, the multiple memory chain data are obtained based on knowledge extraction of historical conversation data, and further include: reducing the multiple initial memory chain data corresponding to the multiple historical conversation data to obtain multiple first memory chain data.
[0019] For example, the initial memory chain data with similar paths are reduced, including operations such as expansion, deletion, and knowledge update, to generate new memory chain data.
[0020] In another possible implementation, each memory chain data in the plurality of memory chain data is stored in a chain storage structure. The chain storage structure can effectively store high-value knowledge extracted from the historical conversation data, thereby helping LLM to better select appropriate memory.
[0021] In another possible implementation, based on the query text, a specific implementation of retrieving the target memory chain data from the memory chain database is: retrieving the historical query text similar or identical to the query text from the memory chain database to obtain the target historical query text; recalling the memory chain data where the target historical query text is located to obtain the target memory chain data.
[0022] Optionally, the memory chain database may be searched by vector search or inverse search to obtain target memory chain data. The present application does not limit the specific search method used, and an appropriate search method may be selected for search as needed.
[0023] In another possible implementation, the question-answering method based on a large language model provided by the present application also includes: displaying summary data of the target memory chain data in a dialog box, the summary data including summaries of multiple knowledge nodes, the summary of each knowledge node in the summaries of the multiple knowledge nodes indicating a summary of the content recorded by each knowledge node.
[0024] Exemplarily, in the question-and-answer interface, the memory chain search results are displayed in a highlighted form, that is, the summary of the target memory chain data is highlighted, so that answers can be generated with basis, increasing the explainability of answer generation.
[0025] In another possible implementation, the question-answering method based on a large language model provided in the present application also includes obtaining feedback information of the answer output by the large language model, wherein the feedback information includes a selected target segment in the answer and feedback on the target segment; and adjusting the answer based on the target segment and the feedback.
[0026] Through the feedback mechanism, the answers generated by the large language model can be corrected in real time according to the user's intention, making the answers more accurate. At the same time, it allows the selection of some fragments in the answer, provides feedback on some fragments, and realizes error feedback at the sentence granularity.
[0027] In another possible implementation, a specific implementation of obtaining feedback information of the answer output by the large language model is: detecting that a target segment is selected, displaying multiple operators for the target segment; obtaining the target segment and the target operator, the target operator being the operator selected from the multiple operators.
[0028] In this possible implementation, the present application defines multiple operation elements. After the user selects a problematic target segment in the answer (for example, the user selects a segment that the user thinks is problematic by sliding), multiple operation elements pop up for the selected target segment. The user determines the direction in which the target segment needs to be improved by selecting the operation element.
[0029] Optionally, the multiple operation elements include one or more of correction, rewriting, expansion, guidance, and a drop-down arrow; wherein the drop-down arrow is used to display an input box for the user to input feedback when selected.
[0030] Exemplarily, the user selects the correction operation element, which indicates that the user believes that part of the description of the selected target segment in the answer is incorrect and needs to be partially corrected; the user selects the rewrite operation element, which indicates that the user believes that all the descriptions of the selected target segment in the answer are incorrect and needs to be regenerated; the user selects the expansion operation element, which indicates that the user believes that the description of the selected target segment in the answer is too simple and needs to be expanded, such as giving the reason why the target segment is described in this way; the user selects the guidance operation element, which indicates that the user believes that the description of the selected target segment in the answer is not detailed enough and needs to be described in more detail based on the target segment; the user selects the drop-down menu operation element, which indicates that the user believes that the several given operation elements (i.e., correction, rewrite, expansion and guidance) cannot well describe the problems with the target segment. At this time, the user clicks the drop-down menu and a text box pops up for the user to enter a description of the problems with the target segment.
[0031] In another possible implementation, the question-answering method based on a large language model provided by the present application also includes recording error notes, which include feedback information and adjusted answers.
[0032] In another possible implementation, the question-answering method based on the large language model provided by the present application also includes, after the current conversation ends, retrieving third memory chain data from the memory chain database based on the error notes, the third memory chain data contains error node data; and updating the error node data based on the error notes.
[0033] Through user feedback recorded in error notes, the knowledge of error nodes on the memory chain data is updated to ensure the correctness of the memory chain data.
[0034] In another possible implementation, the question-answering method based on a large language model provided in the present application also includes, after the current conversation ends, performing knowledge extraction on the conversation data of the current conversation to obtain knowledge node data corresponding to the current conversation data; based on the large language model, determining second memory chain data from the memory chain database, the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current conversation data; merging the knowledge node data corresponding to the current conversation data with the second memory chain data to obtain third memory chain data.
[0035] By merging the new knowledge node data newly extracted from the current conversation data with the memory chain data with similar viewpoints, for example, after extracting the new knowledge node, retrieving the memory chain data related to its topic from the memory chain database, and adding the knowledge node to the end of the memory chain data to form new memory chain data, redundant data can be further reduced.
[0036] In another possible implementation, the question-answering method based on a large language model provided in the present application also includes determining fourth memory chain data from a memory chain database based on the large language model, the fourth memory chain data containing knowledge node data that is inconsistent with the knowledge node data corresponding to the current dialogue data; verifying the inconsistent knowledge node data to obtain correct knowledge node data; and obtaining fifth memory chain data based on the correct knowledge node data and the third memory chain data.
[0037] For example, the logical reasoning ability of LLM is used to determine whether the new knowledge nodes extracted from the current conversation data are contradictory to the old knowledge nodes in the memory chain database. If so, they are verified in conjunction with the search engine to retain the knowledge nodes with correct opinions, thereby reducing the illusion of contradiction in knowledge reduction.
[0038] In a second aspect, the present application provides a question-and-answer device based on a large language model, comprising an acquisition module, a retrieval module, a prompt text generation module and an answer generation module, wherein the acquisition module is used to acquire a current dialogue state, and the current dialogue state includes a query text in the current dialogue; the retrieval module is used to retrieve target memory chain data from a memory chain database based on the query text, and the memory chain database includes multiple memory chain data, and the multiple memory chain data are obtained based on knowledge extraction of historical dialogue data; the prompt text generation module is used to obtain a prompt text based on the query text and the target memory chain data; the answer generation module is used to use the prompt text as an input of the large language model and output an answer corresponding to the query text.
[0039] In one possible implementation, the present application provides a question-and-answer device based on a large language model, which also includes a knowledge extraction module and a memory chain construction module, wherein the knowledge extraction module is used to classify each historical question-and-answer pair in each historical dialogue data to obtain the category of each historical question-and-answer pair; select historical question-and-answer pairs of a target category from multiple historical question-and-answer pairs in the dialogue data, and the generalization of the historical question-and-answer pairs of the target category is greater than the generalization of historical question-and-answer pairs of other categories; perform topic tagging on the historical question-and-answer pairs of the target category to obtain the topic of the historical question-and-answer pairs of the target category; and the memory chain construction module is used to construct initial memory chain data corresponding to each historical dialogue data based on historical question-and-answer pairs with the same topic.
[0040] In another possible implementation, the knowledge extraction module is specifically used to use the historical query text in each historical question-answer pair as the input of the first labeling model, and output a category label, where the category label indicates the category of the historical question-answer pair.
[0041] In another possible implementation, the knowledge extraction module is specifically used to take the historical question-answer pairs of the target category as the input of the second labeling model, and output a topic label, where the topic label indicates the topic information of the historical question-answer pairs of the target category.
[0042] In another possible implementation, the categories of historical question-answer pairs include knowledge categories and scene categories, as well as one or more of creation categories, chat categories, and role-playing categories; and the target categories include knowledge categories and / or scene categories.
[0043] In another possible implementation, the initial memory chain data includes a topic node and several knowledge nodes; wherein the topic node records topic information, and the topic information indicates the topics of several knowledge nodes; and the several knowledge nodes record historical question-answer pairs with the same topic in the order of the generation time of the historical question-answer pairs.
[0044] In another possible implementation, the memory chain building module is further used to reduce the multiple initial memory chain data corresponding to the multiple historical conversation data to obtain multiple first memory chain data
[0045] Optionally, each memory chain data in the plurality of memory chain data is stored in a chain storage structure.
[0046] In another possible implementation, the retrieval module is specifically used to retrieve historical query texts that are similar or identical to the query text from the memory chain database to obtain target historical query texts; and recall the memory chain data where the target historical query texts are located to obtain target memory chain data.
[0047] In another possible implementation, the question-and-answer device based on a large language model provided by the present application also includes a display module, which is used to display summary data of the target memory chain data in a dialog box, wherein the summary data includes summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded by each knowledge node.
[0048] In another possible implementation, the acquisition module is also used to obtain feedback information of the answer output by the large language model, and the feedback information includes the selected target segment in the answer and the feedback on the target segment; the question-answering device based on the large language model provided in the present application also includes an answer adjustment module, which is used to adjust the answer based on the target segment and the feedback.
[0049] In another possible implementation, the display module is further used to detect that the target segment is selected, and display multiple operation elements for the target segment; the acquisition module is specifically used to acquire the target segment and the target operation element, and the target operation element is the selected operation element among the multiple operation elements.
[0050] In another possible implementation, the multiple operation elements include one or more of correction, rewriting, expansion, guidance, and a drop-down arrow; wherein the drop-down arrow is used to display an input box for the user to input feedback when selected.
[0051] In another possible implementation, the question-answering device based on the large language model provided by the present application also includes an error note recording module, which is used to record error notes, and the error notes include feedback information and adjusted answers.
[0052] In another possible implementation, the retrieval module is also used to retrieve the third memory chain data from the memory chain database based on the error notes after the current conversation ends, and the third memory chain data contains error node data; the question and answer device based on the large language model provided in the present application also includes a node update module, which is used to update the error node data based on the error notes.
[0053] In another possible implementation, the knowledge extraction module is further used to extract knowledge from the conversation data of the current conversation after the current conversation ends, so as to obtain knowledge node data corresponding to the current conversation data; the question-answering device based on the large language model provided in the present application also includes a determination module, which is used to determine the second memory chain data from the memory chain database based on the large language model, and the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current conversation data; the memory chain construction module is also used to merge the knowledge node data corresponding to the current conversation data with the second memory chain data to obtain third memory chain data.
[0054] In another possible implementation, the determination module is also used to determine fourth memory chain data from the memory chain database based on the large language model, and the fourth memory chain data contains knowledge node data that is inconsistent with the knowledge node data corresponding to the current dialogue data; the question-and-answer device based on the large language model provided in the present application also includes a verification module, which is used to verify the inconsistent knowledge node data and obtain the correct knowledge node data; the memory chain construction module is used to obtain the fifth memory chain data based on the correct knowledge node data and the third memory chain data.
[0055] In a third aspect, an embodiment of the present application provides a computing device, including a memory and a processor, wherein the memory stores instructions, and when the instructions are executed by the processor, the method described in the first aspect is implemented.
[0056] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the method described in the first aspect is implemented.
[0057] In a fifth aspect, an embodiment of the present application further provides a computer program or a computer program product, wherein the computer program or the computer program product comprises instructions, which, when executed, cause a computer to execute the method described in the first aspect.
[0058] In a sixth aspect, an embodiment of the present application further provides a chip, comprising at least one processor and a communication interface, wherein the processor is used to execute the method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A schematic diagram of a dialogue interaction of a DST-based dialogue system;
[0060] Figure 2 It is a schematic diagram of the feedback interface in the related technology 2;
[0061] Figure 3 It is a schematic diagram of the feedback interface in the related art three;
[0062] Figure 4 A system architecture diagram is shown that can implement the question-answering method based on a large language model provided in an embodiment of the present application;
[0063] Figure 5 A flowchart of a question-answering method based on a large language model provided in an embodiment of the present application;
[0064] Figure 6 A schematic diagram of a memory chain is shown;
[0065] Figure 7 A flowchart of implementing a dialog marking method is shown;
[0066] Figure 8 A schematic diagram of a question-answering interface to which the question-answering method based on a large language model provided in an embodiment of the present application is applied is shown;
[0067] Fig. 9 A schematic diagram showing real-time correction of answers based on user feedback is shown;
[0068] Fig.10 A schematic diagram showing the processing of dialogue data by the question-answering method based on a large language model provided in an embodiment of the present application after the dialogue ends;
[0069] Fig.11 A schematic diagram of the structure of a question-answering device based on a large language model provided in an embodiment of the present application;
[0070] Fig.12 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0071] The term "and / or" mentioned in this article is a kind of relationship that describes the association of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The symbol " / " in this article indicates that the associated objects are in an or relationship, for example, A / B means A or B.
[0072] The terms "first" and "second" in the specification and claims herein are used to distinguish different objects rather than to describe a specific order of the objects. For example, first memory chain data and second memory chain data are used to distinguish different memory chain data rather than to describe a specific order of the memory chain data.
[0073] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.
[0074] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more than two. For example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.
[0075] To facilitate understanding of the solutions of the embodiments of the present application, the technical terms involved in this document are first explained below.
[0076] Large language models (LLMs) are artificial intelligence models designed to understand and generate human language. They are trained on large amounts of text data and can perform a wide range of tasks.
[0077] Dialogue system: A dialogue system is a computer system that simulates humans and aims to form a coherent conversation with humans. Due to the inherent complexity of natural language, the dialogue system involves a large number of natural language processing (NLP) subtasks. It allows machines to understand and process human language through dialogue.
[0078] Session: In a dialogue system, a session usually refers to what a user or system says in an interaction. It is the basic unit in a dialogue system and is used to represent the information conveyed by a user or system in an interaction.
[0079] Prompt: It is the input or query provided by the user to the LLM to guide the model to produce a specific response. Effective prompt design is critical to achieving the expected results of LLM.
[0080] The LLM in the related art has more or less problems in using historical dialogue data and feedback mechanisms. For example, related technology 1 uses short-term dialogue memory. When a user has a dialogue with a trained LLM, the LLM will temporarily remember the user's input and the output it has generated in order to predict the subsequent output. After the model output is completed, it will "forget" the previous user's input and its output. The principle is to store the summary or vector of the historical dialogue and retrieve the Top-K related memories when used. It also includes a Dialogue State Tracking (DST) module, which is a key module in the task-oriented dialogue system (TODS). Its goal is to monitor the user goals hidden in the dialogue history and represent them as a dialogue state consisting of a series of domain, slot, and slot value triplets. DST is responsible for maintaining the dialogue system state (the value corresponding to each slot and the corresponding probability) and updating the dialogue state according to the current round of dialogue. This helps the system better understand the user's intentions and respond accordingly. As Figure 1 As shown in the figure, DST uses a tuple consisting of slot values and probabilities to identify the dialogue state, modifies the query based on historical conversation information, and then generates the dialogue state of the current round. Figure 1 In the figure, U represents the user, S represents the dialogue system, the top row demonstrates the process of the dialogue, and the bottom row corresponds to the dialogue status of different rounds.
[0081] However, this technical solution does not perform knowledge extraction on historical conversation data, but directly vectorizes the original conversation data. Using vectors to encode all conversation data will lead to information entanglement between different topic data, affecting the semantic accuracy of vector representation; and the generation of slot data is difficult and difficult to implement. The generation of slot data relies on complete knowledge, which makes it difficult to implement in practice.
[0082] Related technology 2: After the conversation ends, the user is provided with a case where the conversation result is evaluated, such as two buttons for good and bad reviews (see Figure 2 ), but cannot write comments or regenerate answers based on user feedback.
[0083] In this technical solution, after the conversation ends, thumbs down or thumbs up cannot affect the model's answer generation logic, and answer correction requires the user to start a new round of conversation, lacking the ability to provide rapid feedback and correct answers in real time.
[0084] Related technology three, after the conversation ends, a text input box is provided for the user to fill in the feedback on the result of the conversation (see Figure 3 ), but it is not possible to locate errors at the sentence level.
[0085] In this technical solution, the feedback method is cumbersome and highly dependent on the user's problem description ability. The user can only feedback rough error information but cannot feedback errors at the sentence granularity, which makes it difficult for the search system to fix the problem in a targeted manner.
[0086] In response to the above problems, the present application provides a question-answering method and device based on a large language model. By performing knowledge extraction on historical conversation data, the extracted knowledge points are processed to obtain memory chain data. The memory chain data is stored in a memory chain database. When a user performs a query or question and triggers a similar memory path, the memory chain data of the similar memory path is recalled from the memory chain database to assist the large language model in generating a more comprehensive answer.
[0087] The technical solution of the present application is further described in detail below through the accompanying drawings and embodiments.
[0088] Figure 4 A system architecture diagram is shown that can implement the question-answering method based on a large language model provided in an embodiment of the present application.
[0089] In the offline stage, knowledge is extracted from each historical conversation data, the extracted knowledge points are processed to obtain memory chain data, and the memory chain data is written into the memory chain database.
[0090] In the online question-answering stage, the query text is obtained, and the memory chain data corresponding to the query text is retrieved from the memory chain database based on the query text. A prompt text is generated based on the memory chain data and the query text. The large language model generates an answer based on the prompt text, so that the memory chain data can assist the large language model in generating more comprehensive and accurate answers.
[0091] Figure 5 The flowchart of a question-answering method based on a large language model provided in the embodiment of the present application is shown below. The method can be executed by any device, equipment, platform or device cluster with computing capabilities. The embodiment of the present application does not specifically limit the specific computing device for executing the method, and a suitable computing device can be selected as needed. Figure 5 As shown, the question-answering method based on a large language model at least includes steps S501 to S504.
[0092] In step S501, the current dialog state is obtained, and the current dialog state includes the query text in the current dialog.
[0093] The user may interact with the computing device in any suitable manner and input query text into the dialog box, for example, by voice, image, text, etc. The computing device converts the non-text input into text and inputs it into the dialog box.
[0094] When it is detected that the current conversation includes a preset number of query texts, the current conversation state is obtained, and the current conversation state includes the query texts in the current conversation. For example, the preset number can be set according to actual needs, for example, the preset number can be set to 1, that is, when it is detected that the user has entered a query text, the current conversation state is obtained, and 1 query text entered by the user is obtained. At this time, it can be determined that the user has asked question 1, for example, question 1 can be: "Why is it called the Warring States?" Then, according to question 1, the memory chain database is searched to obtain the target memory chain data (hereinafter, for the convenience of description, the memory chain data is referred to as the memory chain). When the user asks a question for the first time, a more comprehensive answer can be generated with the help of the target memory chain LLM.
[0095] Of course, the preset number of items can also be set to other values, for example, the preset number of items can also be set to 2, that is, when it is detected that the user enters the query text twice, the current dialogue state is obtained, and the two query texts entered by the user are obtained. At this time, it can be determined that the user has raised questions 1 and 2. For example, question 1 can be: "Why is it called the Warring States Period?", and question 2 can be: "Who are the Seven Heroes of the Warring States Period?" Then, according to questions 1 and 2, the memory chain database is searched to obtain the target memory chain. The retrieved target memory chain is more accurate, and a more comprehensive answer is generated with the help of the target memory chain LLM.
[0096] In another example, the current dialogue state may also include the query text in the current dialogue and the answer corresponding to the query text (i.e., the response output of the LLM). For example, when it is detected that the current dialogue has conducted a preset round of conversations, the current dialogue state is obtained, and the multi-round conversation information included in the current dialogue is obtained, that is, the query text in the current dialogue and the answer corresponding to the query text are obtained. The preset rounds can be set as needed. For example, the preset rounds can be 2. When it is detected that the current dialogue has conducted a 2-round conversation, the current dialogue state is obtained, and the query text 1 of the current dialogue and the answer 1 corresponding to the query text 1, as well as the query text 2 and the answer 2 corresponding to the query text 2 are obtained. In other words, question-answer pair 1 (including question 1 and answer 1) and question-answer pair 2 (including question 2 and answer 2) are obtained. Then, according to question-answer pair 1 and question-answer pair 2, the memory chain database is retrieved to obtain the target memory chain, and a more comprehensive answer is generated with the help of the target memory chain LLM.
[0097] In step S502, based on the query text, target memory chain data is retrieved from the memory chain database.
[0098] The memory chain database includes multiple memory chains, which are obtained based on knowledge extraction from historical conversation data. Figure 4 As shown, after the current conversation ends, knowledge is automatically extracted from the current conversation data, the extracted knowledge is processed (for example, highly generalizable knowledge is screened), and a memory chain is constructed.
[0099] It can be understood that the memory chain is a storage data structure that reprocesses historical conversation data. The main goal is to refine the conversation history data and optimize the quality of the LLM prompt text. Optionally, the memory chain data is stored in a chain storage structure, which is easy to search and find, thereby helping to quickly select appropriate memories for LLM.
[0100] The nodes in the memory chain store highly generalized conversation knowledge. In order to obtain highly generalized conversation knowledge, this application proposes a conversation tagging technology. After the conversation ends, the query texts in each round of conversation in the conversation data are classified, and the query texts of the highly generalized categories and the answers corresponding to the query texts are extracted to build a memory chain.
[0101] In another example, in order to organize conversation knowledge more effectively, conversation labeling technology also includes topic tagging of question-answer pairs in conversation data to obtain the topic of each question-answer pair, and constructing conversation knowledge with the same topic in highly generalized conversation knowledge into a memory chain.
[0102] In real life, people often discuss a topic. Therefore, building a memory chain around a topic is in line with people's conversational habits and helps the memory chain assist LLM in generating answers that meet people's expectations.
[0103] The memory chain includes topic nodes and knowledge nodes. The topic node records the topic of the memory chain, and the knowledge node records the conversation knowledge with the same topic. The topic node is the topic content label obtained by summarizing the conversation data, that is, the topic of the conversation discussion. For example, if a conversation revolves around the historical Warring States period, the topic node of the memory chain constructed for the conversation data is "History-Warring States period", and the knowledge node records the question and answer pairs under the topic.
[0104] Figure 6 A schematic diagram of a memory chain is shown.
[0105] like Figure 6As shown in Figure 1, the memory chain includes topic nodes and conversation knowledge under the topic. For example, a conversation revolves around the topic of "History-Warring States" for multiple rounds. In the first round of conversation, the user inputs the query text "Why is it called the Warring States?", and LLM generates the answer "The origin of the Warring States" based on the query text; in the second round of conversation, the user continues to input the query text "Who are the seven heroes of the Warring States?", and LLM generates the answer "Introduction to the seven heroes of the Warring States..." based on the query text; in the third round of conversation, the user continues to input the query text "How did the Qin Dynasty destroy the six countries?", and LLM generates the answer "The process and results of Qin's destruction of the six countries..." based on the query text. According to the conversation data, it can be concluded that the topic of the conversation data is "History-Warring States". The conversation knowledge extracted from the conversation data is obtained to obtain three question-answer pairs, namely, question-answer pair 1 is "Q: Why is it called the Warring States? A: The origin of the Warring States...", question-answer pair 2 is "Q: Who are the seven heroes of the Warring States? A: Introduction to the seven heroes of the Warring States...", and question-answer pair 3 is "Q: How did the Qin Dynasty destroy the six countries? A: The process and results of Qin's destruction of the six countries...". Based on the topic of the inductive dialogue data and the dialogue knowledge under the topic (i.e., question-answer pairs), a memory chain is constructed (see Figure 6 Example 1).
[0106] In one example, the knowledge nodes in the memory chain are constructed in the order of the generation time of the question-answer pairs during the construction process. For example, in a conversation, multiple rounds of conversations were conducted around the topic of car-B car. In the first round of conversation, the user input the query text "What models does B car have?", and LLM generated the answer "B1, B2, B3..." based on the query text; in the second round of conversation, the user continued to input the query text "What is the cost-effectiveness of B3?", and LLM generated the answer "B3 in environmental protection and performance..." based on the query text. The knowledge extracted from the conversation data is used to obtain two question-answer pairs, namely, question-answer pair 1 is "Q: What models does B car have? A: B1, B2, B3...", and question-answer pair 2 is "Q: What is the cost-effectiveness of B3? A: B3 in environmental protection and performance...". Since question-answer pair 1 is generated before question-answer pair 2, in the process of constructing the memory chain, the knowledge node of question-answer pair 1 is before the knowledge node of question-answer pair 2. Figure 6 As shown in Example 2.
[0107] The memory chain is constructed in the order of the time when the question-answer pairs are generated in the conversation, reflecting the user's usual conversation logic. For example, for a certain topic, the user will first ask question 1 and then question 2. The memory chain is constructed in the order of the user's questions, which retains the user's conversation logic information. At the same time, the memory chain makes the generated answers more logical when assisting LLM in generating answers.
[0108] Of course, in some other examples, the order of knowledge nodes in the memory chain can also be constructed with other logics. For example, in historical topics, the memory chain is constructed according to the chronological order of historical events. For example, knowledge nodes with earlier occurrences of historical events are placed earlier in the order, and knowledge nodes with later occurrences of historical events are placed later in the order.
[0109] The following introduces the specific implementation scheme of the dialogue tagging technology and the memory chain construction scheme based on the tagging technology.
[0110] Figure 7 A flow chart for implementing a dialog marking method is shown.
[0111] like Figure 7 As shown in , a complete historical conversation data includes multiple rounds of conversations, and each round of conversation will generate a question-answer pair, for example Figure 7 The historical conversation data in includes N conversations, including N question-answer pairs, question1-answer1, question2-answer2, ..., question N-answer N. First, each question-answer pair in the conversation is categorized to obtain the category of each question-answer pair, and then the conversation is filtered to retain the question-answer pairs with high generalization categories; then the highly generalized question-answer pairs are topic-tagged to obtain the topics of the highly generalized question-answer pairs, and a memory chain is constructed based on the topic information and question-answer pairs.
[0112] In one example, the category of the query text in the question-answer pair can be obtained by training the first labeling model and using the trained first labeling model to infer the category of the query text in the question-answer pair. For example, for a certain historical dialogue data, the query text of each conversation round is used as the input of the labeling model, and the labeling model infers the category of the query text and outputs the category of the query text. For example, the category of the query text can be any one of the knowledge category, the scene category, the creation category, the chat category, and the role-playing category. In this way, the category of each question-answer pair in the historical dialogue data is obtained, and then the question-answer pairs with high generalization categories are screened out. For example, the question-answer pairs of the knowledge category and / or the question-answer pairs of the scene category are question-answer pairs of high generalization categories, which need to be retained for subsequent memory chain construction, and the question-answer pairs of the creation category, the chat category, and the role-playing category are question-answer pairs with low generalization, which are filtered to reduce the data volume of the historical dialogue data.
[0113] It is understandable that when a similar conversation occurs again, the answers to the question-answer pairs with low generalization cannot be used, so they have low value. Therefore, they need to be filtered. For example, the question-answer pairs asking about the weather, Q: Will it rain tomorrow? A: It will be cloudy tomorrow, and there will be no rain. Due to its timeliness, the generalization is low. When LLM receives similar questions again after a few days, such as "Will it rain tomorrow?", it cannot use the answers in the historical question-answer pairs. It can be seen that the question-answer pairs with low generalization cannot provide effective help for LLM's answer generation. Therefore, the question-answer pairs with low generalization can be filtered out to reduce the amount of data in the memory chain database, increase retrieval efficiency, and thus speed up the answer generation.
[0114] For question-and-answer pairs with high generalization, when similar conversations occur again, the answers can be used to assist LLM in generating answers, so they are of high value and need to be retained. For example, for knowledge-based question-and-answer pairs, Q: "Who are the Seven Heroes of the Warring States Period?", A: "The Seven Heroes of the Warring States Period are...", when LLM receives similar questions again, such as "Please introduce who the Seven Heroes of the Warring States Period are?", it can reuse the previous answers without having to reason again, speeding up the answer generation time and achieving a memory function similar to that of the human brain. For example, for questions that people have done before, they will directly use the answers in their memory to solve the questions instead of reasoning again to solve the questions, which greatly speeds up the problem-solving speed.
[0115] After the question-answer pairs of high generalization categories are screened out, the highly generalized question-answer pairs are subject-labeled to obtain the topics of each question-answer pair. For example, the topics of each question-answer pair can be obtained by training a second labeling model and using the trained second labeling model to infer the topics of the question-answer pairs.
[0116] The question-answer pair is used as the input of the second labeling model, and the topic label of the question-answer pair is output. For example, the question-answer pair "Q: How many kings were there in Qin? A: There were six kings in Qin" is used as the input of the labeling model. After inferring the topic of the question-answer pair, the labeling model outputs the topic label "History-Qin".
[0117] Optionally, the first labeling model and the second labeling model are respectively trained by using a transformer-based BERT model.
[0118] In another example, the first labeling model and the second labeling model can be the same labeling model, for example, both can be trained by the transformer-based BERT model, and the query text is used as the input of the labeling model, and the category corresponding to the query text is output; the question and answer pair is used as the input of the labeling model, and the topic corresponding to the question and answer pair is output.
[0119] Through category labeling, conversation data can be quickly screened, conversation data with low generalization value can be filtered out, the amount of data can be significantly reduced, and contextual content optimization can be achieved under a longer timeline. Through topic tagging, the accuracy of memory chain construction can be improved. The labeling model performs topic tagging based on deep semantic features. When building the memory chain in the next stage, the memory chain can be quickly searched and reduced based on the topic of the memory chain.
[0120] LLM has very powerful functions and can implement different functions according to the design of different prompt texts. Therefore, in some other examples, dialogue labeling technology can also be implemented through LLM without the help of other means (for example, by training the labeling model, saving training resources for training additional AI models).
[0121] For example, the prompt text "Please identify the category of the following question-answer pair: question1-answer1" is input to the LLM, and the LLM infers that the category of question1-answer1 is the knowledge category, and the category of the question-answer pair is obtained. The prompt text "Please identify the topic of the following question-answer pair: question1-answer1" is input to the LLM, and the LLM infers that the topic of question1-answer1 is "History-Qin State", and the topic of the question-answer pair is obtained.
[0122] For a historical conversation data, after conversation labeling, highly generalized conversation knowledge (i.e., highly generalized question-answer pairs) and conversation topics are obtained. Then, based on the highly generalized conversation knowledge and conversation topics, an initial memory chain of the historical conversation data is constructed. It is then reduced with the memory chains with similar paths in the memory chain database to generate a new memory chain.
[0123] Among them, similar paths can also be called similar thinking paths. For example, the theme of memory chain 1 is theme 10, and the knowledge nodes include question-answer pair 11, question-answer pair 12, question-answer pair 13, question-answer pair 14, and question-answer pair 15. The theme of memory chain 2 is theme 10, and the knowledge nodes include question-answer pair 11 and question-answer pair 12. At this time, memory chain 1 and memory chain 2 can be called memory chains with similar paths and similar thinking paths. Reduce the memory chains with similar paths, for example, after reducing memory chain 1 and memory chain 2, the reduced memory chain is obtained: theme 10-question-answer pair 11-question-answer pair 12-question-answer pair 13-question-answer pair 14-question-answer pair 15.
[0124] For another example, memory chain 3 is: topic 20-question-answer pair 21-question-answer pair 22-question-answer pair 23; memory chain 4 is: topic 20-question-answer pair 22-question-answer pair 23-question-answer pair 24-question-answer pair 25. It can be judged that memory chain 3 and memory chain 4 are also memory chains with similar paths. After reducing memory chain 3 and memory chain 4, the reduced memory chain is obtained: topic 20-question-answer pair 21-question-answer pair 22-question-answer pair 23--question-answer pair 24-question-answer pair 25.
[0125] From the above, we can see that the conditions for judging whether two memory chains have similar paths include whether the two memory chains have the same subject nodes, and whether the two memory chains have all or part of the similar or identical knowledge nodes.
[0126] By constructing memory chains, we can achieve information reduction in long-term user conversations, extract high-quality common user thinking paths, and improve the comprehensiveness of answers.
[0127] Memory chains are constructed by extracting and mining knowledge from historical conversation data. The constructed memory chains are stored in a memory chain database. When the user interacts with the LLM in the conversation interface, the current conversation state is obtained, and the memory chain database is retrieved according to the current conversation state. Memory chains with similar paths are recalled as target memory chains.
[0128] As described above, the current dialog state may include a query text with a preset number of questions by the user, and then a target memory chain with the same or similar memory path is retrieved from the memory chain database according to the query text.
[0129] For example, when it is detected that the user enters a query text in the dialogue interface, the target memory chain is retrieved from the memory chain database based on the query text. The target memory chain assists LLM in generating a more comprehensive answer, reducing the number of interaction rounds between the user and LLM.
[0130] For another example, when it is detected that the user enters query text twice in the dialogue interface, that is, asks the second question, the target memory chain is retrieved from the memory chain database based on the query text entered twice by the user. In this way, the retrieved target memory chain is more accurate.
[0131] The target memory chain with a similar path to the query text can be understood as the knowledge nodes in the target memory chain having the same or similar thinking paths as the query text (the similarity here means that the query text and the query text on the knowledge node are close in meaning, for example, if the query text is: Will it rain tomorrow? The query text on the knowledge node is: What will the weather be like tomorrow?, then the query text is judged to be similar to the query text on the knowledge node). For example, in the current conversation, the user asked questions 1 and 2 successively, and multiple knowledge nodes in the memory chain have questions 1 and 2 recorded successively, then the memory chain is called to have a similar path to the query text, and it is recalled as the target memory chain. For example, in the current conversation, the user asks LLM questions 1 and 2 successively; memory chain 1 is: topic 1-question and answer pair 1 (including question 1-answer 1)-question and answer pair 2-question and answer pair 3-question and answer pair 4; then memory chain 1 is retrieved from the memory chain database according to questions 1 and 2, and memory chain 1 assists LLM in generating the answer corresponding to question 2, such as answer 2, answer 3 and answer 4, to generate a more comprehensive answer, without the user having to ask questions 3 and 4 again, effectively reducing the number of interaction rounds between the user and LLM and improving the user experience.
[0132] In one example, the subject of the query text may be extracted first, and then memory chains irrelevant to the subject in the memory chain database may be filtered according to the subject, the search scope may be shortened, and then the search may be performed from the memory chains under the subject to speed up the search.
[0133] It should be pointed out that the target memory chain data can be obtained by searching the memory chain database through vector search or inverse search. The specific search method used in this application is not limited, and the appropriate search method can be selected for search as needed.
[0134] When the target memory chain is not retrieved from the memory chain database, the LLM generates an answer based on conventional formal reasoning, that is, based on the query text, it infers the answer corresponding to the query text, and then responds to the answer of the input query text.
[0135] In step S503, a prompt text is obtained based on the query text and the target memory chain data.
[0136] The query text and the target memory chain database are spliced to obtain the prompt text. For example, after retrieving the target memory chain from the memory chain database, the prompt text is designed to assist LLM in generating a more comprehensive answer. For example, if the query text is question 1 and question 2, and the target memory chain is memory chain 1, the prompt text is: the user asked questions 1 and 2 in sequence, please refer to memory chain 1 to generate the answer to question 2.
[0137] By constructing a memory chain, the embodiment of the present application can achieve information reduction of long-term user conversations, extract high-quality common user thinking paths, and improve the comprehensiveness of answers. In addition, when the length of the prompt text input of LLM is limited, a larger amount of conversation information can be passed to LLM with a shorter prompt text input. At the same time, the memory chain formed by conversation reduction can be used to mine user interests, realize personalized interest analysis of users, and enhance the ability of LLM to infer personalized answers.
[0138] In step S504, the prompt text is used as the input of the large language model, and the answer corresponding to the query text is output.
[0139] Help LLMs generate more comprehensive answers by including reminder texts for memory chains.
[0140] Figure 8 A schematic diagram of a question-answering interface using the question-answering method based on a large language model provided in an embodiment of the present application is shown. Figure 8 As shown in Figure 1, when the user enters the question "Who are the seven major powers in the Warring States Period?" in the question-answering interface, LLM outputs the corresponding answer based on the response to the question. When the user enters the question "Who won in the end?", the memory chain database searches for memory chains with similar paths and recalls the target memory chain. The target memory chain is as follows: Figure 8 The memory chain shown in the video is: The Seven Kingdoms of the Warring States Period - The Qin Dynasty's Conquest of the Six Kingdoms - The First Emperor of Qin's Unification of the Six Kingdoms. This memory chain helps the LLM generate a more comprehensive answer.
[0141] In one example, in order to increase the explainability of the generated answer, summary data of the target memory chain data is displayed in a dialog box for the answer generated by the memory chain assisted LLM, and the summary data includes summaries of multiple knowledge nodes, and the summary of each knowledge node in the summary of the multiple knowledge nodes indicates the summary of the content recorded by each knowledge node.
[0142] Furthermore, in the question-and-answer interface, the memory chain search results are displayed in a highlighted form, that is, the summary of the target memory chain data is highlighted, so that answers can be generated based on evidence, increasing the explainability of answer generation.
[0143] Continue to see Figure 8 For the answers generated by the memory chain assisted LLM, it is also annotated in text form: "Based on the analysis of the memory chain of historical dialogues, we will expand the introduction of historical events of this period for you", so that users can clearly understand which answers are generated with the assistance of the memory chain.
[0144] It should be explained that the present application does not make any specific limitation on the specific implementation of LLM. For example, LLM can be the Pangu big language model, the Wenyan Yixin big language model, and the ChatGPT big language model, etc. The question-answering system can select a suitable big language model as needed to implement the question-answering method based on the big language model provided in the embodiment of the present application.
[0145] In some other examples, the question-answering method based on the large language model provided in the examples of the present application also includes a user feedback mechanism. That is, the user's feedback information on the answer output by the large language model is collected, for example, the feedback information may include the target segment selected in the answer, and the feedback on the target segment; then the answer is adjusted based on the target segment and the feedback.
[0146] That is to say, users are allowed to provide feedback on problematic parts of some or all sentences in the answers generated by LLM, and provide feedback to the question-answering system, so that LLM can correct the answers in real time based on the user feedback to ensure that the answers that satisfy the users are generated.
[0147] The current LLM is limited by the amount of training data, the quality of fine-tuning, or the need for more anthropomorphism. The answers output by LLM may contain some errors, which can be recognized by humans. When users find that the answers output by LLM are wrong, they can mark the wrong content in the error and then feedback the problems with the marked content.
[0148] For another example, part of the answer output by LLM may be too brief. The user can mark this part and then provide feedback that the marked content is too brief and needs to be generated with more detailed content.
[0149] After obtaining the target segment with problems in the answer output by the LLM response and the feedback on the target segment, the LLM can regenerate the answer based on the target segment and the feedback, and correct the problems in the answer in real time.
[0150] Optionally, LLM can correct the target segment fed back by the user by combining with the search engine. For example, the user inputs the query text "Please tell me about the history of the Qin State?" LLM responds and outputs the answer as follows "... There were seven Qin kings in the Qin State, ..." The user finds that the description of "There were seven Qin kings in the Qin State" in the answer output by LLM is wrong, marks it out, and gives feedback that the description is wrong. Based on the feedback, LLM searches for "How many Qin kings were there in the Qin State?" through the search engine online, and then corrects the wrong description fed back by the user according to the search results, and corrects the answer to "... There were six Qin kings in the Qin State, ...".
[0151] In order to facilitate users to provide feedback on the answers generated by LLM, this application also defines a variety of operation elements. After the user selects the target fragment with problems in the answer generated by LLM, multiple operation elements will pop up for the target fragment. The user can provide feedback on the problems with the target fragment by selecting the corresponding operation element, thereby reducing the tediousness of the feedback steps.
[0152] Fig. 9 A schematic diagram of real-time correction of answers based on user feedback is shown. Fig. 9 As shown, when it is detected that the user has selected (for example, the user selects by sliding) a target segment with problems in the answer generated by LLM, multiple operation elements pop up for the target segment, and the multiple operation elements include, for example, correction, rewriting, expansion, and guidance. The user feedbacks the problems with the target segment or the direction that needs to be improved by selecting (for example, clicking to select) the corresponding operation element. The target operation element selected by the user is detected, and the prompt text is determined according to the target segment and the target operation element. For example, it is detected that segment 1 is selected in the answer generated by LLM, and the correction operation element is selected, then a prompt text is generated according to the selected segment 1 and the correction operation element. The prompt text can be: "There is a problem with segment 1 and correction processing is required. Please correct segment 1." Then LLM corrects segment 1 according to the prompt text and regenerates the answer description for the selected segment 1 in the answer.
[0153] It can be understood that when the user selects the correction operation element, it means that the user believes that part of the description of the selected target segment in the answer is incorrect and needs to be partially corrected; when the user selects the rewrite operation element, it means that the user believes that all the descriptions of the selected target segment in the answer are incorrect and a new description needs to be regenerated; when the user selects the expansion operation element, it means that the user believes that the description of the selected target segment in the answer is too simple and needs to be expanded, such as giving the reason why the target segment is described in this way; when the user selects the guidance operation element, it means that the user believes that the description of the selected target segment in the answer is not detailed enough and needs a more detailed description based on the target segment.
[0154] The correction, rewrite, expansion, and guidance operation elements described in the embodiments of the present application are only examples and do not constitute a limitation on the operation elements. More or fewer operation elements can be defined according to actual needs. For example, a drop-down operation element can also be defined. The drop-down operation element is used when the user believes that the given operation elements (i.e., correction, rewrite, expansion, and guidance) cannot well describe the problems existing in the target segment. The user can select the drop-down operation element to pop up a text box, in which the user can describe the problems existing in the target segment or the directions that need improvement in the text box.
[0155] In another example, feedback can also be used to update the memory chain. For example, the user's feedback on the LLM generated answer will be recorded in the error note. For example, the error note will record the target segment selected by the user, the feedback on the target segment (including the selected operation element or the description of the opinion entered in the text box by selecting the drop-down arrow), and the answer regenerated by the LLM based on the feedback. The error note can be applied to the memory chain reduction to update the memory chain. Reduction can include operations such as expansion, deletion, and knowledge update of the memory chain.
[0156] For example, the erroneous knowledge nodes in the memory chain nodes can be updated and corrected according to the error notes. Exemplarily, after the current conversation is completed, according to the error notes, the memory chain database is searched to see if there are knowledge nodes with the same or similar errors as those in the error notes, and then the errors of the knowledge nodes are corrected. For example, the error notes record that "There are seven kings of Qin in Qin" and need to be corrected. After correction, the note content is "There are six kings of Qin in Qin". Then, the memory chain of the "History-Qin" theme in the memory chain database is searched to see if there are knowledge nodes such as "There are seven kings of Qin in Qin" or similar to "There are x kings of Qin in Qin". If so, the memory chain is updated, and the knowledge of "There are seven kings of Qin in Qin" on the knowledge node is updated to "There are six kings of Qin in Qin". In this way, the accuracy of the knowledge of the memory chain is guaranteed, so as to generate the correct answer in the subsequent auxiliary LLM answer generation.
[0157] Fig.10 FIG. 1 is a schematic diagram showing a method for processing conversation data based on a large language model provided in an embodiment of the present application after the conversation ends. Fig.10 As shown, after the current conversation is completed, the current conversation data becomes historical conversation data. The conversation data is processed according to the method for processing historical conversation data described above, such as conversation tagging, extracting knowledge with high generalization (i.e., question-answer pairs) and knowledge topics, and obtaining the initial memory chain of the conversation data. The initial memory chain includes multiple knowledge nodes, and the multiple knowledge nodes record the new knowledge extracted from the conversation data. Of course, if no highly generalized knowledge is extracted from the conversation data, there is no need to perform subsequent steps.
[0158] Then, the initial memory chain and the memory chains in the memory chain data are reduced. For example, topic aggregation is first performed, and memory chains similar to the initial memory chain are retrieved from the memory chain database that have the same topic as the initial memory chain of the current conversation data, and then the initial memory chain is merged with the memory chains similar to it.
[0159] Optionally, LLM is used to judge the similarity between the initial memory chain and the memory chain of the same subject in the memory chain database. For example, LLM can be used to judge the similarity of two memory chains by designing a prompt text. For example, the prompt text can be: "Please determine the similarity between memory chain a and memory chain b". The prompt text is input into LLM, and LLM responds and outputs the similarity between memory chain a and memory chain b. If the similarity between memory chain a and memory chain b is greater than a preset threshold (such as 0.9), it is determined that memory chain a and memory chain b are similar.
[0160] Here, the meaning of the similarity between memory chain a and memory chain b can be understood as that the knowledge viewpoints on the knowledge nodes of memory chain a and memory chain b are similar or identical. For example, memory chain a is: Topic 30-Question and answer pair 31-Question and answer pair 32-Question and answer pair 33-Question and answer pair 34; memory chain b is: Topic 30-Question and answer pair 31 ` - Question and answer pair 32 ` - Question and answer pair 33 ` ; Among them, question and answer pair 31 and question and answer pair 31 ` The views of question and answer pair 32 and question and answer pair 32` are similar or the same, and the views of question and answer pair 33 and question and answer pair 33` are similar or the same. Then we can judge that memory chain a and memory chain b are similar, and merge memory chain a and memory chain b to get a new memory chain c: Topic 30-Question and answer pair 31 ` - Question and answer pair 32 ` -Question and answer to 33`-Question and answer to 34.
[0161] The meaning of similar or identical knowledge viewpoints is that the facts or contents described by the answers on the knowledge nodes are similar or identical. For example, the answer on knowledge node 31 is "There were six kings in the Qin Dynasty", and the answer on knowledge node 31` is "There were six kings in the Qin Dynasty, namely...", then knowledge nodes 31 and 31` are similar or identical. ` It can be called similarity of views.
[0162] In another example, if no similar memory chain is retrieved from the memory chain database, the memory chain database continues to be retrieved for a data chain that contradicts the initial data chain view of the current conversation data.
[0163] LLM can be used to judge whether the initial memory chain and the memory chain of the same topic in the memory chain database have contradictory views. For example, by designing a prompt text, LLM can be used to judge whether there are contradictory knowledge nodes in the two memory chains. For example, the prompt text can be: "Please determine whether there are knowledge nodes with contradictory views in memory chain c and memory chain d". The prompt text is input into LLM, and LLM responds and outputs whether there are contradictory knowledge nodes in memory chain c and memory chain d. If there are contradictory knowledge nodes, LLM combines with the search engine to verify the two contradictory knowledge nodes and retain the correct knowledge nodes.
[0164] The meaning of knowledge node contradiction here can be understood as: the content described by the answer on the knowledge node is contradictory. For example, memory chain c is: topic 40-question and answer pair 41-question and answer pair 42-question and answer pair 43; memory chain d is: topic 40-question and answer pair 41`-question and answer pair 42`-question and answer pair 43`; if the answer to question and answer pair 41 is "There were six kings in the Qin Dynasty", question and answer pair 41 ` The answer is "There were seven kings in the Qin Dynasty". At this time, question and answer pair 41 and question and answer pair 41` are judged to be contradictory, and then memory chain c and memory chain d are judged to be contradictory. Then the correct answer is found through the search engine, which is "There were six kings in the Qin Dynasty". The correct knowledge node is retained to obtain a new memory chain f: Topic 40-Question and answer pair 41 ` - Question and answer pair 42 ` - Question and answer pair 43 ` .
[0165] The embodiment of the present application can correct answers and historical conversations in real time according to user intentions through a rich error note feedback interactive mechanism, update error notes, and assist in conversation knowledge reduction; based on LLM, point of view merging and contradiction discrimination can reduce the illusion of contradiction in knowledge reduction, and combine the prompt text generator to design a prompt framework to constrain LLM's generation of contradictory results. By summarizing historical query knowledge and summarizing the thinking process, LLM can be assisted in forming a block-like knowledge system, so that LLM can not only give fragmented answers, but also reduce the large-grained knowledge gathered in massive user conversations, and systematically display it, thereby improving professionalism and comprehensiveness.
[0166] Based on the same concept as the aforementioned embodiment of the question-answering method based on a large language model, the embodiment of the present application also provides a question-answering device 1100 based on a large language model. The question-answering device 1100 based on a large language model can be deployed on any device, equipment, platform or device cluster with computing power to implement the question-answering method based on a large language model provided in the embodiment of the present application, so as to realize the construction of memory chain data by extracting knowledge from historical conversation data. The memory chain data assists the large language model in generating more comprehensive and high-quality answers, thereby improving the user experience. The question-answering device 1100 based on a large language model includes a device for realizing Figure 4-10 The units or modules of each step in the question answering method based on a large language model are shown.
[0167] Fig.11 The structure diagram of a question-answering device based on a large language model provided in an embodiment of the present application is shown in FIG. Fig.11 As shown, the question-answering device 1100 based on the large language model at least includes an acquisition module 1101, a retrieval module 1102, a prompt text generation module 1103 and an answer generation module 1104, wherein the acquisition module 1101 is used to obtain the current dialogue state, and the current dialogue state includes the query text in the current dialogue; the retrieval module 1102 is used to retrieve the target memory chain data from the memory chain database based on the query text, and the memory chain database includes a plurality of memory chain data, and the plurality of memory chain data are obtained based on knowledge extraction of historical dialogue data; the prompt text generation module 1103 is used to obtain the prompt text based on the query text and the target memory chain data; the answer generation module 1104 is used to use the prompt text as the input of the large language model and output the answer corresponding to the query text.
[0168] In one possible implementation, the present application provides a question-and-answer device 1100 based on a large language model, which also includes a knowledge extraction module 1105 and a memory chain construction module 1106, wherein the knowledge extraction module 1105 is used to classify each historical question-and-answer pair in each historical dialogue data to obtain the category of each historical question-and-answer pair; select historical question-and-answer pairs of a target category from multiple historical question-and-answer pairs in the dialogue data, and the generalization of the historical question-and-answer pairs of the target category is greater than the generalization of historical question-and-answer pairs of other categories; perform topic tagging on the historical question-and-answer pairs of the target category to obtain the topic of the historical question-and-answer pairs of the target category; and the memory chain construction module 1106 is used to construct initial memory chain data corresponding to each historical dialogue data based on historical question-and-answer pairs with the same topic.
[0169] In another possible implementation, the knowledge extraction module 1105 is specifically configured to use the historical query text in each historical question-answer pair as the input of the first labeling model, and output a category label, where the category label indicates the category of the historical question-answer pair.
[0170] In another possible implementation, the knowledge extraction module 1105 is specifically configured to use the historical question-answer pairs of the target category as input to the second labeling model, and output a topic label, where the topic label indicates topic information of the historical question-answer pairs of the target category.
[0171] In another possible implementation, the categories of historical question-answer pairs include knowledge categories and scene categories, as well as one or more of creation categories, chat categories, and role-playing categories; and the target categories include knowledge categories and / or scene categories.
[0172] In another possible implementation, the initial memory chain data includes a topic node and several knowledge nodes; wherein the topic node records topic information, and the topic information indicates the topics of several knowledge nodes; and the several knowledge nodes record historical question-answer pairs with the same topic in the order of the generation time of the historical question-answer pairs.
[0173] In another possible implementation, the memory chain building module 1106 is further configured to perform reduction processing on multiple initial memory chain data corresponding to multiple historical conversation data to obtain multiple first memory chain data.
[0174] Optionally, each memory chain data in the plurality of memory chain data is stored in a chain storage structure.
[0175] In another possible implementation, the retrieval module 1102 is specifically used to retrieve historical query texts that are similar or identical to the query text from the memory chain database to obtain target historical query texts; and recall the memory chain data where the target historical query texts are located to obtain target memory chain data.
[0176] In another possible implementation, the question-and-answer device 1100 based on a large language model provided in the present application also includes a display module 1107, which is used to display summary data of the target memory chain data in a dialog box, the summary data including summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded in each knowledge node.
[0177] In another possible implementation, the acquisition module 1101 is also used to obtain feedback information of the answer output by the large language model, and the feedback information includes the target segment selected in the answer and the feedback on the target segment; the question-answering device 1100 based on the large language model provided in the present application also includes an answer adjustment module 1108, which is used to adjust the answer based on the target segment and the feedback.
[0178] In another possible implementation, the display module 1107 is further used to detect that the target segment is selected and display multiple operation elements for the target segment; the acquisition module is specifically used to acquire the target segment and the target operation element, and the target operation element is the selected operation element among the multiple operation elements.
[0179] In another possible implementation, the multiple operation elements include one or more of correction, rewriting, expansion, guidance, and a drop-down arrow; wherein the drop-down arrow is used to display an input box for the user to input feedback when selected.
[0180] In another possible implementation, the question-answering device 1100 based on a large language model provided in the present application further includes an error note recording module 1109, which is used to record error notes, and the error notes include feedback information and adjusted answers.
[0181] In another possible implementation, the retrieval module 1102 is also used to retrieve the third memory chain data from the memory chain database based on the error notes after the current conversation ends, and the third memory chain data contains error node data; the question and answer device based on the large language model provided in the present application also includes a node update module, which is used to update the error node data based on the error notes.
[0182] In another possible implementation, the knowledge extraction module 1105 is also used to extract knowledge from the dialogue data of the current dialogue after the current dialogue ends, so as to obtain knowledge node data corresponding to the current dialogue data; the question-answering device 1100 based on the large language model provided in the present application also includes a determination module 1110, which is used to determine the second memory chain data from the memory chain database based on the large language model, and the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current dialogue data; the memory chain construction module is also used to merge the knowledge node data corresponding to the current dialogue data with the second memory chain data to obtain third memory chain data.
[0183] In another possible implementation, the determination module 1110 is also used to determine fourth memory chain data from the memory chain database based on the large language model, and the fourth memory chain data contains knowledge node data that is inconsistent with the knowledge node data corresponding to the current dialogue data; the question-answering device 1100 based on the large language model provided in the present application also includes a verification module 1111, which is used to verify the inconsistent knowledge node data and obtain the correct knowledge node data; the memory chain construction module is used to obtain the fifth memory chain data based on the correct knowledge node data and the third memory chain data.
[0184] The large language model-based question-answering device 1100 according to the embodiment of the present application may correspond to executing the method described in the embodiment of the present application, and the above and other operations and / or functions of each module in the large language model-based question-answering device 1100 are respectively to implement Figure 4-10 For the sake of brevity, the corresponding processes of each method in are not repeated here.
[0185] It should be explained that the question-answering device 1100 based on a large language model adopted in the embodiment of the present application can be implemented by software or by hardware. For example, the question-answering device based on a large language model can be used as a plug-in in the question-answering system to implement the question-answering method based on a large language model provided in the embodiment of the present application.
[0186] The present application also provides a computing device, including at least one processor, a memory and a communication interface, wherein the processor is used to execute Figure 4-10 The method described.
[0187] Fig.12 A schematic diagram of the structure of a computing device provided in an embodiment of the present application.
[0188] like Fig.12 As shown, the computing device 1200 includes at least one processor 1201, a memory 1202 and a communication interface 1203. The processor 1201, the memory 1202 and the communication interface 1203 are connected in communication, and the communication connection can be realized by wired means (such as a bus) or by wireless means. The communication interface 1203 is used to send and / or receive data sent by other devices; the memory 1202 stores computer instructions, and the processor 1201 executes the computer instructions to execute the question-answering method based on the large language model in the aforementioned method embodiment, so as to realize the construction of memory chain data by extracting knowledge from historical conversation data, and the memory chain data assists the large language model to generate better answers and improve the user experience.
[0189] It should be understood that in the embodiment of the present application, the processor 1201 may be a central processing unit CPU, and the processor 1201 may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0190] The memory 1202 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1201. The memory 1202 may also include a nonvolatile random access memory.
[0191] The memory 1202 may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. Among them, the nonvolatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).
[0192] It should be understood that the computing device 1200 according to the embodiment of the present application can execute the implementation of the embodiment of the present application. Figure 4-10 The method shown, the detailed description of the implementation of this method can be found above, for the sake of brevity, it will not be repeated here.
[0193] An embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer instructions are executed by a processor, the above-mentioned method is implemented.
[0194] An embodiment of the present application provides a chip, which includes at least one processor and an interface, wherein the at least one processor determines program instructions or data through the interface; the at least one processor is used to execute the program instructions to implement the method mentioned above.
[0195] An embodiment of the present application provides a computer program or a computer program product, wherein the computer program or the computer program product comprises instructions, and when the instructions are executed, the computer is caused to execute the above-mentioned method.
[0196] Those of ordinary skill in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0197] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented by hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0198] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the present application in detail. It should be understood that the above description is only the specific implementation method of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.
Claims
1. A question-answering method based on a large language model, characterized in that: include: Acquire a current dialog state, wherein the current dialog state includes a query text in the current dialog; Based on the query text, target memory chain data is retrieved from a memory chain database, where the memory chain database includes a plurality of memory chain data, and the plurality of memory chain data are obtained based on knowledge extraction of historical conversation data; Based on the query text and the target memory chain data, obtaining a prompt text; The prompt text is used as the input of the large language model, and the answer corresponding to the query text is output.
2. The method according to claim 1, characterized in that Each memory chain data in the plurality of memory chain data is constructed by the following steps: Classifying each historical question-answer pair in each historical conversation data to obtain a category of each historical question-answer pair; Selecting a historical question-answer pair of a target category from a plurality of historical question-answer pairs in the dialogue data, wherein the generalizability of the historical question-answer pairs of the target category is greater than the generalizability of the historical question-answer pairs of other categories; Performing topic tagging on historical question-answer pairs of the target category to obtain topics of the historical question-answer pairs of the target category; Based on historical question-answer pairs with the same topic, initial memory chain data corresponding to each historical dialogue data is constructed.
3. The method according to claim 2, characterized in that The classifying each historical question-answer pair in each historical conversation data to obtain the category of each historical question-answer pair includes: The historical query text in each historical question-answer pair is used as the input of the first labeling model, and a category label is output, where the category label indicates the category of the historical question-answer pair.
4. The method according to claim 2 or 3, characterized in that: The topic tagging of the historical question-answer pairs of the target category to obtain the topic of the historical question-answer pairs of the target category includes: The historical question-answer pairs of the target category are used as input of the second labeling model, and a topic label is output, where the topic label indicates topic information of the historical question-answer pairs of the target category.
5. The method according to any one of claims 2 to 4, characterized in that: The categories of the historical question-answer pairs include knowledge category and scenario category, and one or more of creation category, small talk category and role-playing category; The target category includes the knowledge category and / or the scenario category.
6. The method according to any one of claims 2 to 5, characterized in that: The initial memory chain data includes a subject node and several knowledge nodes; wherein the subject node records subject information, and the subject information indicates the subject of the several knowledge nodes; the several knowledge nodes record the historical question-answer pairs with the same subject in the order of the generation time of the historical question-answer pairs.
7. The method according to any one of claims 2 to 6, characterized in that: The plurality of memory chain data are obtained based on knowledge extraction of historical conversation data, and further include: A plurality of initial memory chain data corresponding to the plurality of historical conversation data are reduced to obtain a plurality of first memory chain data.
8. The method according to any one of claims 1 to 7, characterized in that: Each memory chain data in the plurality of memory chain data is stored in a chain storage structure.
9. The method according to any one of claims 1 to 8, characterized in that: The step of retrieving target memory chain data from a memory chain database based on the query text includes: Retrieving historical query texts similar to or identical to the query text from the memory chain database to obtain target historical query texts; The memory chain data where the target historical query text is located is recalled to obtain the target memory chain data.
10. The method according to any one of claims 1 to 9, characterized in that: Also includes: The summary data of the target memory chain data is displayed in a dialog box, wherein the summary data includes summaries of multiple knowledge nodes, and the summary of each knowledge node in the summaries of the multiple knowledge nodes indicates a summary of the content recorded in each knowledge node.
11. The method according to any one of claims 1 to 10, characterized in that: Also includes: Acquire feedback information of the answer output by the large language model, wherein the feedback information includes a target segment selected from the answer and feedback on the target segment; The answer is adjusted based on the target segment and the feedback.
12. The method according to claim 11, characterized in that The obtaining feedback information of the answer output by the large language model includes: detecting that the target segment is selected, and displaying a plurality of operation elements for the target segment; The target fragment and the target operator are obtained, where the target operator is an operator selected from the multiple operators.
13. The method according to claim 12, characterized in that The multiple operation elements include one or more of correction, rewriting, expansion, guidance, and drop-down arrow; Wherein, the drop-down arrow is used to display an input box for the user to input the feedback when it is selected.
14. The method according to any one of claims 11 to 13, characterized in that: Also includes: An error note is recorded, wherein the error note includes the feedback information and the adjusted answer.
15. The method according to claim 14, characterized in that Also includes: After the current conversation ends, based on the error note, third memory chain data is retrieved from the memory chain database, and the third memory chain data contains error node data; Based on the error note, the error node data is updated.
16. The method according to any one of claims 1 to 15, characterized in that: Also includes: After the current conversation ends, knowledge extraction is performed on the conversation data of the current conversation to obtain knowledge node data corresponding to the current conversation data; Based on the large language model, determining second memory chain data from the memory chain database, wherein the second memory chain data contains knowledge node data similar to the knowledge node data corresponding to the current dialogue data; The knowledge node data corresponding to the current conversation data and the second memory chain data are merged to obtain third memory chain data.
17. The method according to claim 16, characterized in that Also includes: Based on the large language model, determining fourth memory chain data from the memory chain database, wherein the fourth memory chain data contains knowledge node data that is inconsistent with the knowledge node data corresponding to the current dialogue data; Verify the contradictory knowledge node data and obtain the correct knowledge node data; Based on the correct knowledge node data and the third memory chain data, the fifth memory chain data is obtained.
18. A question-answering device based on a large language model, characterized in that: include: An acquisition module, used to acquire a current dialog state, wherein the current dialog state includes a query text in the current dialog; A retrieval module, configured to retrieve target memory chain data from a memory chain database based on the query text, wherein the memory chain database includes a plurality of memory chain data, and the plurality of memory chain data are obtained based on knowledge extraction of historical conversation data; A prompt text generation module, used to obtain a prompt text based on the query text and the target memory chain data; The answer generation module is used to use the prompt text as the input of the large language model and output the answer corresponding to the query text.
19. A computing device comprising a memory and a processor, characterized in that: Instructions are stored in the memory, and when the instructions are executed by the processor, the method according to any one of claims 1 to 17 is implemented.
20. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 17 is implemented.
Citation Information
Cited By
Image-text question and answer and person swimming detection method and system
CN120356139A
Data archiving method, data retrieval method, data archiving device, data retrieval device, medium, equipment and product
CN120705274A
Machine question and answer dialogue method and device
CN120705284A
Large model memory processing method and device, equipment, storage medium and product
CN121144976A
Memory management method of multi-agent interaction system, related equipment and program product
CN121434926A