Question and answer method and system based on large language model

By reconstructing the original question in sentences and combining key information in historical dialogue records, the problems of slow response speed and high computing resource consumption in multiple rounds of question-and-answer are solved, achieving a more efficient question-and-answer process and a more anthropomorphic user experience.

CN120146188APending Publication Date: 2025-06-13ZHONGKE DINGFU BEIJING TECH DEV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510211332.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In multiple rounds of Q&A, the prior art requires the user's complete historical dialogue records as input to the next round of Q&A of the large language model, resulting in slow response speed and increased computing resource consumption.

Method used

By reconstructing the original problem in sentences, generating new problems, and combining key information in historical dialogue records, the completeness and accuracy of the problem are improved, thereby reducing the amount of information that needs to be processed.

Benefits of technology

It improves the response speed of large language models in multiple rounds of question-and-answer, reduces the consumption of computing resources, and provides more anthropomorphic dialogue and interaction experience and convenient information consultation services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146188A_ABST
    Figure CN120146188A_ABST
Patent Text Reader

Abstract

The invention provides a question and answer method and system based on a large language model, which can improve the response speed of the large language model while realizing a multi-round question and answer function based on the large model. The method comprises the steps of obtaining an original question input by a user; if the original question supports multiple rounds of questions and answers, according to a historical dialogue record corresponding to the original question, performing sentence reconstruction on the original question to improve the integrity and / or accuracy of the question so as to generate a new question; respectively carrying out recall processing on the original question and the new question, and respectively determining a first candidate text fragment associated with the original question and a corresponding similarity thereof, and a second candidate text fragment associated with the new question and a corresponding similarity thereof; determining the text segment with the highest similarity in the first candidate text segment and the second candidate text segment as a first target text segment; and inputting the first target text fragment and the corresponding question into the first large language model through a prompt instruction to obtain an answer output by the first large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a question-answering method and system based on a large language model. Background Art

[0002] In the technical field of question-answering based on large language models, there are single-round question-answering and multi-round question-answering. Single-round question-answering refers to a single-round natural language interaction between a machine and a user, and multi-round question-answering refers to a continuous and multi-round natural language interaction between a machine and a user. This interaction method not only requires the machine to understand the user's intention, but also requires the machine to make intelligent responses and feedback based on context information.

[0003] Currently, usually, after understanding the user's question by using a large language model, an answer is generated to achieve multi-round question-answering. Specifically, after obtaining the user's question, text fragments associated with the question are retrieved from a database according to the user's question. Then, the retrieved text fragments, the user's question, and the historical conversation record are input into the large language model together, and the large language model itself is relied on to complete the reply to the question raised by the user in multi-round question-answering.

[0004] However, in the above method of performing multi-round question-answering, the complete historical conversation record of the user needs to be used as the input for the next round of question-answering by the large language model, resulting in a large amount of information that the large language model needs to process, making the response speed of the large language model slower and increasing the consumption of computing resources. Therefore, there is an urgent need for a question-answering solution that can improve the response speed of the large language model in multi-round question-answering. Summary of the Invention

[0005] This application provides a question-answering method and system based on a large language model, which can improve the response speed of the large language model and reduce the consumption of computing resources while achieving the continuous dialogue interaction effect with the user, providing a more anthropomorphic dialogue interaction experience for the user, and a more convenient information consultation service function.

[0006] In a first aspect, a question-answering method based on a large language model is provided, including:

[0007] Obtain the original question input by the user;

[0008] If the original question supports multi-round question-answering, according to the historical conversation record corresponding to the original question, reconstruct the original question to generate a new question for the purpose of improving the integrity and / or accuracy of the question;

[0009] Perform a recall process on the original question based on the database to determine the first candidate text fragment associated with the original question and its corresponding similarity;

[0010] Recall the new question based on each data source cited in the historical answers in the historical conversation record, determine the second candidate text segment associated with the new question and its corresponding similarity;

[0011] Determine the text segment with the highest similarity among the first candidate text segment and the second candidate text segment as the first target text segment;

[0012] Input the first target text segment and the corresponding question into the first large language model through a prompt instruction to obtain the answer output by the first large language model.

[0013] In a feasible design, according to the historical conversation record corresponding to the original question, sentence reconstruction is performed on the original question for the purpose of improving the integrity and / or accuracy of the question, including:

[0014] Call the second large language model to summarize the historical conversation record corresponding to the original question to obtain the first key information of the historical conversation record;

[0015] Call the third large language model to perform sentence reconstruction on the original question for the purpose of improving the integrity and / or accuracy of the question according to the first key information and the historical questions in the historical conversation record. In a feasible design, sentence reconstruction is performed on the original question for the purpose of improving the integrity and / or accuracy of the question, including:

[0016] Supplement the missing sentence components of the original question to generate a new question.

[0017] In a feasible design, sentence reconstruction is performed on the original question for the purpose of improving the integrity and / or accuracy of the question, including:

[0018] Identify and correct the typos in the original question to generate a new question.

[0019] In a feasible design, sentence reconstruction is performed on the original question for the purpose of improving the integrity and / or accuracy of the question, including:

[0020] Identify and delete the redundant words in the original question to generate a new question.

[0021] In a feasible design, the method further includes:

[0022] If the original question is not the first question in this multi-round Q&A, obtain the data source types cited by the first large language model when generating each historical answer in the historical conversation record;

[0023] If the data source types corresponding to each historical answer do not support multi-round Q&A, determine that the original question does not support multi-round Q&A;

[0024] Otherwise, determine that the original question supports multi-round Q&A.

[0025] In a feasible design, the method further includes:

[0026] If the original question does not support multi-round Q&A, or the original question is the first question of this multi-round Q&A, perform a recall process on the original question based on the database to determine a second target text segment associated with the original question;

[0027] Input the second target text segment and the original question into the first large language model through a prompt instruction to obtain an answer output by the first large language model.

[0028] In a feasible design, the recall process includes:

[0029] Extract at least one second key information from the target data, where the second key information includes the keyword and summary information of the target data, and the target data is data with information related to the field of the original question;

[0030] Determine at least one first target data in the target data according to the second key information, and each first target data corresponds to at least one second key information;

[0031] Based on the comparison result of the semantic vectors corresponding to the target data and the question to be processed, determine at least one second target data, where the second target data and the first target data are data in the target data;

[0032] If at least one association relationship is extracted from the target data, according to the extracted association relationship, determine the target association data of the first target data and the second target data, where the target association data is data having an association relationship with the first target data and the second target data, and the association relationship is the data association information in the target data;

[0033] Calculate the target similarity between the first target data, the second target data, the target association data and the question to be processed respectively according to the text similarity and semantic vector similarity between the first target data, the second target data, the target association data and the question to be processed;

[0034] Determine the third target data corresponding to the highest m target similarities as the candidate text segment associated with the question to be processed, where the third target data is data in the first target data, the second target data and the target association data.

[0035] In a feasible design, calculating the target similarity between the first target data, the second target data, the target association data and the question to be processed respectively according to the text similarity and semantic vector similarity between the first target data, the second target data, the target association data and the question to be processed includes:

[0036] Construct a set based on the first target data, the second target data, and the target association data. The set includes all the first target data, the second target data, the target association data, and the corresponding text similarity and semantic vector similarity between the first target data, the second target data, the target association data, and the problem to be processed.

[0037] Perform normalization processing based on the text similarity between the first target data, the second target data, and the target association data in the set and the highest and lowest values of the text similarity and semantic vector similarity in the set to obtain a normalized text similarity value.

[0038] Calculate the target similarity based on the normalized text similarity value and the semantic vector similarity.

[0039] In a second aspect, a question-answering system based on a large language model is provided, including:

[0040] A question acquisition module for acquiring the original question input by the user.

[0041] A question reconstruction module for, if the original question supports multi-turn question answering, reconstructing the original question into a new question for the purpose of improving the completeness and / or accuracy of the question according to the historical conversation record corresponding to the original question.

[0042] A recall processing module for performing recall processing on the original question based on the database to determine the first candidate text segment associated with the original question and its corresponding similarity.

[0043] The recall processing module is also used to perform recall processing on the new question based on each data source cited in the historical answer in the historical conversation record to determine the second candidate text segment associated with the new question and its corresponding similarity.

[0044] A decision-making module for determining the text segment with the highest similarity among the first candidate text segment and the second candidate text segment as the first target text segment.

[0045] An answer generation module for inputting the first target text segment and the corresponding question into the first large language model through a prompt instruction to obtain the answer output by the first large language model.

[0046] After obtaining the original question in the embodiment of the present application, by judging whether the original question supports multi-turn question answering, it is possible to determine whether to enter multi-turn question answering, thereby reducing the resource consumption in the overall question answering process. After entering the multi-turn question answering session, by reconstructing the original question in combination with the historical conversation record, more background information can be supplemented or ambiguous expressions can be corrected, thereby improving the integrity and / or accuracy of the question, so that the text fragments recalled by the new question can more accurately match the user's needs and improve the multi-turn question answering ability of the question answering solution. Therefore, it can solve the problem that when there are defects in the user's question and effective data cannot be retrieved or the retrieval effect is poor, it will directly affect the final answer effect, and may lead to incomplete answers, wrong answers and other situations. It realizes the continuous dialogue interaction effect with the user and provides a more anthropomorphic dialogue interaction experience for the user. Among them, since the original question supports multi-turn question answering, it means that each data source cited in the historical answer in the corresponding historical conversation record contains relevant document materials, that is, the answer to the original question can probably be obtained through these data sources. Therefore, compared with recalling new questions based on the database, recalling new questions based on each data source cited in the historical answer in the historical conversation record can further improve the recall rate of new questions.

[0047] Further, the present application determines the target text fragment most relevant to the question by comparing the similarity between the first candidate text fragment obtained based on the original question and the second candidate text fragment obtained based on the new question. It realizes integrating the recall results of both and selecting the text fragment with the highest similarity, which can avoid the limitations that may be brought by a single data source and cover the information related to the question more comprehensively. Therefore, inputting the accurate target text fragment and the corresponding question into the first large language model through the prompt instruction can enable the first large language model to output an answer that can solve the problem as much as possible, thereby improving the accuracy of the multi-turn question answering of the question answering solution of the present application. Compared with the current solution of inputting the complete historical conversation record and question into the first large language model, the present application only inputs the first target text fragment and the corresponding question into the first large language model, so that the first large language model can give an answer without understanding the above information, improving the response speed of the system and providing a more convenient information consultation service function for the user. And because the amount of data processed by the first large language model is reduced, the consumption of GPU resources can also be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] In order to more clearly illustrate the technical solutions of the present application, the drawings required for the embodiments will be briefly introduced below. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0049] Figure 1It is a schematic diagram of a question-answering system based on a large language model provided by an exemplary embodiment of the present application;

[0050] Figure 2 It is a schematic diagram of the data processing flow of a dialogue management module provided by an exemplary embodiment of the present application;

[0051] Figure 3 It is a schematic diagram of the data processing flow of a question reconstruction module provided by an exemplary embodiment of the present application;

[0052] Figure 4 It is a schematic diagram of the data processing flow of a question-answering system based on a large language model provided by an exemplary embodiment of the present application;

[0053] Figure 5 It is a schematic flowchart of a question-answering method based on a large language model provided by an exemplary embodiment of the present application. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0055] Currently, two common methods are usually adopted for multi-round question answering in the commonly used question-answering solutions. One is to utilize the understanding ability of the large language model (LLM) itself, store all the dialogue contexts, and use them as the model input to realize the multi-round question-answering function; the other is based on traditional natural language processing (NLP) technology, refine and extract the data in the industry and field, form structured (similar to a table) data, and realize the multi-round question-answering function through directional query and classification. The following are the specific contents, advantages, and disadvantages of these two methods:

[0056] I. Method 1 - Question-answering system based on traditional industry knowledge base

[0057] 1. Implementation steps

[0058] (1) Preprocess the relevant knowledge data required for question answering, including data cleaning, deduplication, data augmentation, etc.

[0059] (2) Carry out knowledge engineering work according to the requirements of the business scenario, further refine the preprocessed data, and extract relevant knowledge from the data. For example, for a news article describing a car accident, extract relevant information such as the people in the news, the type of event, the time of the event, and the location of the event. The specific knowledge engineering work needs to be designed according to the requirements of the business scenario, and there may be differences in different scenarios.

[0060] (3) Based on the knowledge base processed in the previous step, construct corresponding question-and-answer pair data and use this question-and-answer pair data to train an intent recognition model and an entity recognition model. This model is used to determine the specific question category to which the user's question belongs when providing question-and-answer services later and finally obtain the corresponding question answer from the database according to this category. How to construct the specific intent recognition model needs to be formulated according to the requirements.

[0061] 2. Advantages:

[0062] (1) The accuracy of question answering is relatively high.

[0063] (2) The operation efficiency is relatively high and the hardware requirements are easy to meet.

[0064] 3. Disadvantages:

[0065] (1) It is extremely dependent on data. In the absence of real question-and-answer data, using the method of manually constructing data is likely to lead to a decline in performance.

[0066] (2) Knowledge engineering is complex and labor-intensive. In the process of extracting knowledge from data, not only corresponding business knowledge is required as support, but also manual review is needed after using the corresponding algorithm to complete the extraction to ensure the accuracy of the extracted knowledge. The whole process not only has high requirements for professional knowledge and algorithms, but also requires a large amount of manpower.

[0067] (3) It is difficult to reuse in different business scenarios. Since this solution requires refined extraction of knowledge, it is almost impossible to reuse the trained model in different scenarios.

[0068] (4) The support scope of multi-round question answering is limited, only limited to the processed data scope, and cannot perform open-ended or semi-open-ended question answering processing.

[0069] II. Solution 2 - Question Answering System Based on Large Language Model (LLM)

[0070] 1. Implementation steps:

[0071] (1) Slice the business data, that is, slice a document into several short text fragments with a text length not exceeding 300 words, and use a text vectorization model to convert the text into vectors. Finally, store the text fragments and their corresponding vectors in the database.

[0072] (2) Build a knowledge retrieval system to retrieve text fragments related to the user's input question. Usually, this step uses text vector retrieval or a combination of text vector and text content retrieval to recall relevant content.

[0073] (3) Input the retrieved relevant materials, the user's question, and the historical conversation records into the LLM, and rely on the capabilities of the LLM itself to complete the response to the user's question.

[0074] 2. Advantages:

[0075] (1) Convenient data processing and a wide range of applicable scenarios.

[0076] (2) It can support a wide range of question-and-answer types, including open-ended and semi-open-ended.

[0077] (3) High accuracy in question answering.

[0078] 3. Disadvantages:

[0079] (1) High resource consumption.

[0080] (2) Slow response speed.

[0081] (3) Extremely dependent on the accuracy of knowledge retrieval. When the user's question cannot retrieve valid data or the retrieval effect is poor, it will directly affect the final answer effect, and may lead to incomplete answers, wrong answers, etc.

[0082] (4) It is necessary to use the user's complete historical conversation records as the input for the next round of question and answer, which greatly increases the consumption of Graphics Processing Unit (GPU) resources.

[0083] To solve the respective defects of the above two methods in the multi-round question-and-answer session, improve the implementation efficiency in the application link and reduce the manual input cost, as Figure 1 shown, this application provides a question-and-answer system based on a large language model. The system includes an input question acquisition module, a dialogue management module, a question reconstruction module, a recall processing module, a decision-making module, and an answer generation module. The functions of each module are introduced below:

[0084] (1) Question acquisition module, which is used to acquire the original question input by the user into the system. When users have doubts and hope to get answers from the system, they will enter the content they want to ask on the corresponding interaction interface. The question acquisition module can capture these questions input by the users and input the original questions into the dialogue management module.

[0085] (2) The dialogue management module is used to store and manage the historical dialogue records of different users, and refine and summarize the historical dialogue records of users by calling the second LLM. At the same time, it saves the data sources and data source types in the recall results cited when the LLM generates answers for each round of dialogue, and determines whether the current original question supports multi-round Q&A. Specifically, as Figure 2 shown, the dialogue management module obtains and saves the questions input by the user for each round, the data sources and data source types of the final recall results output by the recall processing module according to the questions, and the answers generated by the answer generation module. After obtaining this information, the dialogue management module can generate historical dialogue records based on the historical questions and historical answers, and then call the second LLM to refine and summarize the historical dialogue records of the user to obtain the first key information of the historical dialogue records.

[0086] Through research, it is found in this application that some data sources do not support multi-round Q&A or have a poor support effect for multi-round Q&A. Therefore, another function of this module is to determine whether the current input question supports multi-round Q&A when a new question is input by the user by combining the data source types involved in the user's historical dialogue records. If the input question supports multi-round Q&A, it means that the system can enable the subsequent multi-round Q&A steps to implement the multi-round Q&A function. If the input question does not support multi-round Q&A, it means that the system does not need to enable the subsequent multi-round Q&A steps and directly performs recall processing according to the input question.

[0087] Among them, it can be set according to actual needs which data source types support multi-round Q&A and which do not. For example, in the field of government affairs Q&A, the data source types that the multi-round Q&A system needs to cover may include policy document types, news types, service types, and government-citizen interaction letter types. Among them, the questions corresponding to the data source of the government-citizen interaction letter type are usually questions in the form of Frequently Asked Questions (FAQ). In the Q&A mode involving this type of data source, after the user asks a question, the system can give an answer by querying this type of data source to solve the user's problem. Usually, the user will not continue to ask questions about this data source, and the next question is likely to point to another data source. Therefore, this type of data source cannot support multi-round Q&A well, so in the actual use process, this type of data source does not need to be connected to the multi-round Q&A link. That is to say, if the data source types involved in the historical dialogue record corresponding to the input question do not support multi-round Q&A, then the input question does not support multi-round Q&A, and there is no need to enable the multi-round Q&A steps subsequently.

[0088] Similarly, most data sources in the Q&A field of customer service for e-commerce trading platforms also contain FAQ-style questions. For example, data sources of the return processing flow type, data sources for answering product inquiries, etc. In the Q&A mode involving this type of data source, after a user usually asks a question and gets an answer, they then ask another question that has no relation to the previous question and its answer. For example, the user asks, "What is the quality of this piece of clothing?" After getting the answer, the user then asks, "What is the return address?" Both of the above questions are FAQ-style questions, and the corresponding data sources do not support multi-turn Q&A and do not need to be incorporated into the multi-turn Q&A process. That is to say, the second question does not support multi-turn Q&A.

[0089] It can be seen that questions that do not support multi-turn Q&A are independent of the questions in the historical conversation record, have no relation to the questions in the historical conversation record, and do not generate logical guidance or association. Therefore, if the data source types corresponding to the questions in the historical conversation record do not support multi-turn Q&A, it means that when generating an answer to the current user's input question, there is no need to search for the answer from the data sources in the historical conversation record, and the answer to the question can be found from the data sources in the database other than those involved in the historical conversation record.

[0090] Correspondingly, if the data source types corresponding to the questions in the historical conversation record include data source types that support multi-turn Q&A, it means that there is likely a logical connection between the current user's input question and the questions in the historical conversation record, that is, the current input question supports multi-turn Q&A, and the system needs to enable subsequent multi-turn Q&A steps for this input question. That is to say, when generating an answer to the current input question, it is very likely that the answer can be found from the data sources in the historical conversation record. Compared with searching for answers from all the data sources in the database every time, it can greatly reduce the time for the system to generate answers. Data source types that support multi-turn Q&A include, for example, data sources of the process handling type, data sources of document types with recorded information, etc. Taking the data source of the process handling type as an example, the user asks, "How to apply for an ID card?" After getting the answer, the user then asks, "Where to apply?" Among the first two questions, the second question has a logical connection with the first question, and the data source involved is the same as that of the first question, which means that when generating an answer to the second question, it can be found from the data sources in the historical conversation record. Therefore, in actual use, this type of data source can be incorporated into the multi-turn Q&A process. That is to say, the second question supports multi-turn Q&A.

[0091] (3) The question reconstruction module, which is responsible for reconstructing the current input question using the third LLM in combination with the user's current input question, the historical questions in the historical conversation record, and the first key information, with the aim of improving the integrity and / or accuracy of the question. Specifically, as Figure 3As shown, if the dialogue management module determines that the original question input by the current user supports multi-round Q&A, the question reconstruction module obtains the original question input by the user and the summarized first key information from the data saved by the dialogue management module. The first key information is used to improve the completeness of the question.

[0092] Completeness refers to the degree of completeness of the sentence components. The completeness of a sentence can be the ratio of the number of components contained in the sentence to the number of set complete components. For example, if it is set that a complete sentence component needs to include a subject, a predicate, and an object, when the sentence only has a predicate and an object, the completeness of the sentence is considered to be the ratio of the number of components contained in the sentence (2) to the number of complete components (3), that is, 66.7%.

[0093] Accuracy refers to the degree of accuracy of a sentence. For example, the accuracy of a sentence is the ratio of the number of accurate characters to the total number of characters in the sentence. For example, the question is "Where to handle the identity increase?", where "increase" is a misspelling and should be changed to "certificate". Therefore, the accuracy of this question is 87.5%.

[0094] The following is an example to illustrate the reconstruction of questions:

[0095] For example, the first question is "My ID card is lost. How to handle it?", and the second question is "Where to handle it?", at this time, the corresponding object (i.e., the ID card) is lacking in this question sentence. After question sentence reconstruction, the second question will be rewritten as "Where to handle the ID card?". The reconstructed question has a more complete grammar and is more convenient for subsequent steps such as recall.

[0096] (4) Recall processing module, the main purpose of this module is to find out each text fragment related to the question from the database according to the question input by the user, and score and sort according to the relevance degree of each text fragment to the question, and finally obtain the text fragment with the highest score as the final candidate.

[0097] (5) Decision-making module, this module is used to make a final decision on the result output by the recall processing module. Since after question sentence reconstruction, there will be a new question, and the recall processing module will synchronously use the new question and the original question for recall, and obtain two recall results. Each recall result includes m candidate text fragments associated with the corresponding question, and the similarity between each candidate text fragment and the corresponding question. The decision-making module sorts the similarities in these two recall results, and finally selects a candidate text fragment with the highest similarity as the final model input.

[0098] (6) Answer generation module, which is used to splice the finally obtained relevant text fragments and corresponding questions output by the decision-making module into corresponding Prompt instruction sets, input them into the first LLM, and then call the first LLM to generate answers. The above-mentioned modules cooperate with each other to perform the process of multi-turn Q&A based on large language models as Figure 4 shown.

[0099] The multi-turn Q&A system based on large language models provided by this application has all the advantages of the Q&A system solution based on LLM, such as comprehensive Q&A coverage and support for open-ended Q&A, etc. It overcomes the disadvantages of the cumbersome and complex data preprocessing work in the Q&A system solution based on traditional industry knowledge bases, and reduces the construction cost of the database.

[0100] In the implementation of the multi-turn Q&A function of the Q&A solution, this application abandons the conventional LLM implementation solution (that is, using all or part of the historical conversation records and adding them into the LLM as input to enable the model to obtain the above information). Instead, it uses the question reconstruction module to reconstruct the current question, and introduces the summary result of the historical conversation record during the reconstruction process, so that the first large language model can effectively judge whether it belongs to the continuation of the previous question according to the reconstructed new question, so as to generate answers more accurately. Since the historical conversation records are not input into the large language model in the solution of this application, there is no need for the large language model to understand the above text through the historical conversation records, which greatly reduces the information that the large language model needs to process, improves the response speed of the large language model, and reduces the consumption of computing resources.

[0101] In addition, in the scenario of complex data, such as the situation where some data sources do not support multi-turn Q&A, the Q&A solution used in this application can effectively distinguish various types of data sources involved in the historical records, and does not enable the corresponding multi-turn Q&A function when all the data sources included in the historical conversation records do not support multi-turn Q&A, effectively avoiding resource waste. Through the design of reducing resource consumption and improving response speed in many aspects, this application provides a lightweight Q&A system.

[0102] Based on the above system embodiments, as Figure 5 shown, this application also provides a Q&A method based on large language models, including:

[0103] S110, obtaining the original question input by the user.

[0104] In a feasible design, the following method is used to determine whether the original question supports multi-turn Q&A:

[0105] If the original question is not the first question of this multi-turn Q&A, obtain the data source types cited by the first large language model when generating each historical answer in the historical conversation record;

[0106] If the data source types corresponding to each historical answer do not support multi-turn Q&A, it is determined that the original question does not support multi-turn Q&A;

[0107] Otherwise, it is determined that the original question supports multi-turn Q&A.

[0108] Exemplarily, the data sources in the database are labeled with types, and it can be determined whether the data source type supports multi-turn Q&A by querying the setting records of whether each data source type supports multi-turn Q&A.

[0109] In the above example, if the original question is not the first question of the current multi-turn Q&A, it means that it has corresponding historical conversation records. Then, it can be determined whether the original question supports multi-turn Q&A according to the data source types corresponding to each historical answer, so as to determine whether to execute the multi-turn Q&A steps subsequently. According to the description in the system embodiment, if the data source types corresponding to each historical answer do not support multi-turn Q&A, it means that the current original question is very likely not related to the data sources cited in the historical conversation records. Therefore, if the original question does not support multi-turn Q&A, the multi-turn Q&A steps do not need to be executed subsequently. If the data source types involved in the historical conversation records include types that support multi-turn Q&A, then the original question supports multi-turn Q&A, and the multi-turn Q&A steps need to be started subsequently.

[0110] It should be understood that if the original question is the first question of the current multi-turn Q&A, it means that it has no corresponding historical conversation records, and the multi-turn Q&A steps do not need to be started either. By determining whether the original question supports multi-turn Q&A, it can affect whether the subsequent multi-turn Q&A steps are enabled, effectively avoiding resource waste.

[0111] S120, if the original question supports multi-turn Q&A, according to the historical conversation records corresponding to the original question, sentence reconstruction is performed on the original question for the purpose of improving the completeness and / or accuracy of the question to generate a new question.

[0112] In a feasible design, it is implemented in the following way. Sentence reconstruction is performed on the original question for the purpose of improving the completeness and / or accuracy of the question according to the historical conversation records corresponding to the original question, including:

[0113] Call the second large language model to summarize the historical conversation records corresponding to the original question to obtain the first key information of the historical conversation records;

[0114] Call the third large language model to perform sentence reconstruction on the original question for the purpose of improving the completeness and / or accuracy of the question according to the first key information and the historical questions in the historical conversation records to generate a new question.

[0115] Among them, the first key information can be used to improve the integrity and / or accuracy of the original question. For example, it includes the constituent elements of a sentence such as the subject, predicate, and object of the sentence. It can also include person-related information (such as identity information, role relationship information, etc.), event-related information (such as event subject and action information, event time information, location information, and result information, etc.), thing attribute-related information (such as product characteristic information, location feature information, etc.), and view and attitude-related information (such as personal view and attitude information, group view and attitude information, etc.).

[0116] Through the above example, by invoking the second large language model, the historical conversation records can be summarized and refined to obtain a brief summary of the historical conversation records, so as to stabilize the resource consumption and speed up the process of question reconstruction.

[0117] It should be noted that this application does not limit the network frameworks of the first large language model, the second large language model, and the third large language model.

[0118] It should be understood that the first large language model, the second large language model, and the third large language model in this application can be the same model.

[0119] In a feasible design, it is achieved in the following way. The original question is reconstructed into a new question for the purpose of improving the integrity and / or accuracy of the question:

[0120] Generate a new question by supplementing the missing constituent elements of the original question.

[0121] For example, taking the original question input by the user as "Where to handle" and the question in the historical conversation record as "My ID card is lost. How to handle it" as an example. By analysis, it can be determined that the object of the original question sentence is missing, and the object of the question can be completed through the question in the historical conversation record, and the reconstructed question is "Where to handle the ID card".

[0122] The above example can improve the integrity of the question by supplementing the missing constituent elements of the original question, so as to accurately perform the recall work subsequently.

[0123] In a feasible design, it is achieved in the following way. The original question is reconstructed into a new question for the purpose of improving the integrity and / or accuracy of the question:

[0124] Identify and correct the typos in the original question to generate a new question.

[0125] The above example can improve the accuracy of the question by correcting the typos in the original question, so as to accurately perform the recall work subsequently.

[0126] In a feasible design, it is achieved by the following method: The original problem is reconstructed into a new problem for the purpose of improving the integrity and / or accuracy of the problem:

[0127] Identify and delete redundant words in the original problem to generate a new problem.

[0128] The above example can improve the accuracy of the problem by deleting redundant words in the original problem, so as to facilitate the subsequent accurate recall work.

[0129] S130, perform a recall process on the original problem based on the database to determine the first candidate text segment associated with the original problem and its corresponding similarity.

[0130] In a feasible design, the recall process includes the following steps:

[0131] S131, extract at least one second key information from the target data, where the second key information includes the keyword and summary information of the target data.

[0132] Among them, when performing a recall process on the original problem, the target data is the data in the database with information related to the field of the original problem. For example, the field of the original problem is the government affairs field. When performing a recall process on the new problem, the target data is each data source cited in the historical answers in the historical conversation record, and each data source contains one or more files.

[0133] The second key information is the keyword and / or sentence in the target data, such as the chapter name and brief description information, etc. The association relationship is the association relationship or reference relationship between a certain file and other files in the target data, such as the reference documents in a paper, etc.

[0134] When extracting the association relationship, in the current file, the associated files before and after the trigger word can be extracted through the pre-set trigger word, and stored in the database in the form of a quadruple of "current file - trigger word - relationship type - associated file". In some examples, the trigger word can be recognized by means of Name Entity Recognition (NER), the associated files can be extracted through relationship extraction, and then the names of the associated files can be compared with the names of the standard files in the database by means of regular matching. The names of the standard files with low matching degrees and multiple matching results are filtered out, and the names of the standard files with high matching degrees are retained, which are the names of the associated files. For example, if there is content such as "in accordance with the 'Regulations on Water Supply in City X'" in the current file, the trigger word "in accordance with" is recognized by NER, and the "Regulations on Water Supply in City X" is extracted through relationship extraction. Then, it is compared with the names of the corresponding standard files in the database by means of regular matching. If the matching degree is higher than the preset matching degree threshold and there is only one matching result, the name of the corresponding standard file is recorded in the quadruple and stored for subsequent extraction of the corresponding content in the file with the corresponding name.

[0135] It should be noted that the association relationship may not necessarily be extracted from the target data, which may specifically be manifested as an empty result of the association relationship extraction, that is, no quadruple corresponding to the association relationship is stored in the database after extraction.

[0136] Before extracting at least one second key information and the association relationship from the target data, it is generally necessary to preprocess the target data, including the following steps:

[0137] Perform data cleaning and deduplication on the target data; perform structured parsing on the data after data cleaning and deduplication to obtain at least one hierarchical data, where the hierarchical data is any structural hierarchical data in the data after data cleaning and deduplication;

[0138] Store the preprocessed data in the database.

[0139] The above example removes the duplicate data, abnormal symbols, html tags and other non-text-related content in the target data to ensure the overall neatness and effectiveness of the processed text data. The target data can be data with information related to the government affairs field (hereinafter referred to as government affairs data), such as policy documents, notices, news, etc.

[0140] By analyzing the structural characteristics of different types of government affairs data, corresponding structured parsing schemes are formulated, that is, different types of government affairs data correspond to different structured parsing schemes. In this way, when actually applied, the appropriate structured parsing scheme can be selected according to the file type of the processed text data to complete the structured parsing work of this type of file. This step can restore the hierarchical relationships originally existing in the data, such as multi-level tags, etc. The question-and-answer data sources in the government affairs field are usually government documents, and such documents often have specific file formats. The hierarchical structures such as chapters and sections in the documents can be separately parsed according to this specific format, which is convenient for subsequent data retrieval.

[0141] According to the formulated structured parsing scheme, the target data is structurally parsed to obtain text fragment data at different levels and converted into corresponding semantic vectors. If the amount of target data is small, the structured parsing may not be performed, and the target data is directly converted into corresponding semantic vectors.

[0142] The text fragment data at each level and the corresponding semantic vectors are stored in the corresponding database.

[0143] Exemplarily, the key information extraction and association relationship extraction of the target data are realized in the following ways:

[0144] Extract keywords and summary information from the hierarchical data other than the hierarchical data corresponding to the lowest level.

[0145] Extract the data association information of all hierarchical data.

[0146] After the preprocessing is completed, taking the text fragment data at different levels (excluding the lowest level) obtained after the structured parsing as the main body, extract its second key information, such as keywords and summary information, and each hierarchical data corresponds to one piece of second key information. In some examples, the text structure after the structured parsing includes three levels: chapter, section, and article. Then, in this step, the corresponding extraction work is carried out for the content of the chapter and section levels. The text content to be extracted at the "section" level is the content of all "articles" under this section (when extracting the second key information of the penultimate level, it is to extract the second key information of all the lowest-level data included under it, so there is no need to extract the lowest-level data separately). Similarly, the extraction of the "chapter" is completed. The main purpose of this step is to facilitate the implementation of subsequent hierarchical recall.

[0147] Taking the text fragment data at all levels after the structured parsing as the main content, extract the information of the other text fragment data that may be mentioned in each text fragment data, that is, the association relationship between the text fragment data existing in the text, and save it.

[0148] S132. Determine at least one first target data from the target data according to the second key information, where each first target data corresponds to at least one second key information.

[0149] After the user's question, using all levels of text fragment data after structured parsing as the main body, according to the keywords and summary information corresponding to each level of text fragment data, in the order of the level height, determine whether there is text fragment data that meets the answer corresponding to the user's question in each level of text fragment data, that is, the first target data.

[0150] Exemplarily, the following method is used to determine at least one first target data according to the extracted key information:

[0151] According to the level height of the hierarchical data, determine at least one first target data from the high level to the low level in turn.

[0152] In the above example, since the text structure after structured parsing includes three levels: chapter, section, and article, first determine whether there is first target data according to the keywords and summary information of the text fragment data at the chapter level, and then determine whether there is first target data according to the keywords and summary information of the text fragment data at the section level. If there is, recall the corresponding text fragment data, that is, determine at least one first target data.

[0153] Compared with finding text fragment data related to the user's question data in the full text, this step performs content recall in a multi-level and gradually detailed manner, with higher accuracy and higher running efficiency than the conventional recall method.

[0154] S133. Determine at least one second target data based on the comparison result of the semantic vectors corresponding to the target data and the problem to be processed. The second target data and the first target data are data in the target data.

[0155] Among them, the problem to be processed is the original problem or the newly reconstructed problem. The field to which the problem to be processed belongs is, for example, the government affairs field.

[0156] Exemplarily, the following method is used to determine at least one second target data based on the comparison result of the semantic vectors corresponding to the target data and the problem to be processed:

[0157] Convert the hierarchical data into a semantic text vector;

[0158] Convert the question data into a question semantic text vector;

[0159] Compare the semantic text vector with the question semantic text vector;

[0160] Determine the hierarchical data corresponding to the k semantic text vectors with the highest vector comparison similarity as the second target data, where k≥1.

[0161] In the above example, the problem to be processed is vectorized, and then the semantic vector of the problem to be processed is compared with the semantic vectors of all text segment data in the database to select the k semantic vectors of the text segment data with the highest vector comparison similarity, k≥1, and recall the text segment data corresponding to the k vectors. Specifically, the vector comparison similarity can be obtained by calculating the cosine similarity between the semantic vector of the problem to be processed and the semantic vector of the text segment data.

[0162] S134. According to the extraction result in the foregoing step S131, determine whether there is an association relationship between the text segments recalled in the foregoing steps S132 and S133.

[0163] Judge whether there is a quadruple corresponding to the association relationship in the database. If it exists, there is an association relationship; if not, there is no association relationship.

[0164] If there is an association relationship, execute steps S1341 to S1344; if not, execute steps S1345 to S1347.

[0165] The following describes steps S1341 to S1344:

[0166] S1341. According to the extracted association relationship, determine the target association data of the first target data and the second target data.

[0167] According to the association relationship in the form of quadruples stored in the database, determine the text segment data associated with the first target data and the second target data among all the text segment data, that is, the target association data. The association relationship is the data association information in the target data.

[0168] S1342. Calculate the text similarity (similarity of the text itself) and semantic vector similarity (similarity of the meaning expressed by the text) between the first target data, the second target data, the target association data and the question data respectively;

[0169] It is easy to understand that since the text similarity has been calculated when determining the first target data, only the semantic vector similarity between the first target data and the question data needs to be calculated; the same applies to the second target data.

[0170] In some examples, the BM25 algorithm can be used to calculate the text similarity, and the cosine similarity method in step S133 can be used for the semantic vector similarity. It should be noted that other algorithms can also be used, and the embodiments of the present application do not limit the specific algorithm.

[0171] S1343. Calculate the target similarities between the first target data, the second target data, the target associated data and the problem to be processed respectively according to the text similarity and the semantic vector similarity.

[0172] S1344. Select the third target data corresponding to the m target similarities with the highest scores, that is, the text segment data corresponding to the m target similarities with the highest scores, and determine the text segment data corresponding to the m target similarities with the highest scores as the candidate text segments associated with the problem to be processed, where the third target data is the data among the first target data, the second target data and the target associated data.

[0173] After determining the existence of the association relationship, use the content associated with the text segments recalled by the hierarchical retrieval and the text segments recalled by the semantic vector as the intermediate candidate texts, and sort and score the three recall results (including the text segments recalled by the hierarchical retrieval, the text segments recalled by the semantic vector and the intermediate candidate texts) to select the most effective reference text as the input of the large language model.

[0174] The steps of calculating the target similarity include:

[0175] (1) Construct a set according to the first target data, the second target data and the target associated data. The set includes all the first target data, the second target data, the target associated data, and the text similarities and semantic vector similarities between the corresponding first target data, second target data, target associated data and the problem to be processed.

[0176] (2) Perform normalization processing according to the text similarities between the first target data, the second target data and the target associated data in the set and the highest and lowest values among all the text similarities and semantic vector similarities in the set to obtain the normalized text similarity value.

[0177] The formula for performing the normalization processing is as shown in the following formula (1):

[0178]

[0179] where S i is the normalized text similarity value, S i0 is any text similarity in the set, S max is the highest value among the text similarities and semantic vector similarities in the set, and S min is the lowest value among the text similarities and semantic vector similarities in the set.

[0180] (3) Calculate the target similarity according to the normalized text similarity value and the semantic vector similarity.

[0181] The formula for calculating the target similarity is as shown in the following formula (2):

[0182] F = (1 + αS i )V, Equation (2);

[0183] where F is the target similarity, α is a hyperparameter with a value in [0, 1] used to control the importance degree of the text similarity normalization value S i and V is the semantic vector similarity.

[0184] The following is an explanation of steps S1345 to S1347:

[0185] S1345, calculate the text similarity and semantic vector similarity between the first target data, the second target data and the problem to be processed respectively;

[0186] S1346, calculate the target similarity between the first target data, the second target data and the problem to be processed respectively according to the text similarity and semantic vector similarity;

[0187] S1347, select the third target data corresponding to the highest m target similarities, that is, the text segment data corresponding to the highest m target similarities, and determine the text segment data corresponding to the highest m target similarities as the candidate text segments associated with the problem to be processed

[0188] In step S1347, the third target data is the data in the first target data and the second target data.

[0189] The calculation methods in steps S1345 to S1347 are the same as those in steps S1341 to S1344, and will not be elaborated here.

[0190] In practice, the processing of information in some fields into knowledge in the corresponding fields through natural language processing technology, the construction of the corresponding knowledge base, and subsequent maintenance often require a large amount of human resources. For example, in the government affairs field, due to the particularity of government affairs information, processing government affairs-related information into knowledge in the government affairs field through natural language processing technology and constructing the corresponding knowledge base not only requires a large amount of human resources but also requires frequent manual intervention in subsequent maintenance. This is because some information and knowledge in the government affairs field are prone to changes and alterations, such as changes in policy orientation, etc. When such situations occur, manual maintenance and update of the knowledge base are required. For the intelligent Q&A in the government affairs field based on large language model technology, although it can effectively save the process of knowledge base construction, the simple data slicing method cannot well handle the data characteristics of government affairs data. The reason is that some data in the government affairs field, such as government official documents, policy documents, etc., have different file formats from conventional texts, and conventional segmentation methods may lead to information damage; and some policy documents and government official documents have strong correlations, that is, one document has an associated relationship with another document, so it is necessary to obtain some relevant information from the associated documents.

[0191] In the recall processing solution provided by this application, the data processing is relatively simple and does not require complex data knowledge engineering. Through the combination of hierarchical retrieval recall, semantic vector recall, and associated recall, it can extract more comprehensively the content related to the user's question, with a comprehensive coverage of the Q&A scope, and can solve the problems of incomplete and fragmented knowledge segments in the reply content, effectively improving the Q&A accuracy, reply integrity, and efficiency.

[0192] S140, perform a recall process on the new question based on each data source cited in the historical answer in the historical conversation record, and determine the second candidate text segment associated with the new question and its corresponding similarity.

[0193] Among them, the method for performing a recall process on the new question refers to the method for performing a recall process on the question to be processed in S130, which will not be elaborated here.

[0194] Exemplarily, the data source used for performing a recall process on the new question is the data source of the type supporting multi-round Q&A in the historical conversation record. Since there is likely no associated relationship between the new question and the data source in the historical conversation record that does not support multi-round Q&A, this example can reduce the number of data sources and improve the speed of the recall process.

[0195] S150, determine the text segment with the highest similarity among the first candidate text segments and the second candidate text segments as the first target text segment.

[0196] Among them, the number of text segments in the first candidate text segments may be one or more, and the number of text segments in the second candidate text segments may be one or more.

[0197] In most cases where the original question has incomplete grammar, typos, or redundant words, the similarity between the original question and the correct document materials is relatively low, while the similarity between the reconstructed new question and the correct document materials is relatively high. For example, the original question is "Where to handle", the reconstructed question is "Where to apply for an ID card", and the correct document materials are "ID card application locations". Since the original question lacks the object to be handled, in the recall results for the original question, there will be results for the application locations of various documents such as "ID card application locations" and "driver's license application locations". That is to say, the recall results determined according to the original question are ambiguous, which means that the similarity between the original question and these document materials is relatively low. However, the object to be handled in the reconstructed question is accurate, enabling more precise recall, and the similarity with the document materials of "ID card application locations" is relatively high and higher than the similarity corresponding to each document material in the recall results of the original question.

[0198] It can be inferred that the higher the similarity between the candidate text segment and the question, the higher the probability that it is the correct document material. Therefore, the first target text segment can be determined by comparing the similarities of the first candidate text segment and the second candidate text segment.

[0199] Of course, if the original question does not have incomplete grammar, typos, or redundant words, and the reconstructed new question has little difference from the original question, then it is uncertain which candidate text segment corresponding to the question in the recall results has a higher similarity. Therefore, compared with blindly taking the recall results corresponding to the reconstructed new question as the first target text segment, by comparing the similarities of the first candidate text segment and the second candidate text segment, the first target text segment can be accurately determined.

[0200] S160, Input the first target text segment and the corresponding question into the first large language model through a prompt instruction to obtain the answer output by the first large language model.

[0201] It should be noted that the above S120 - S160 are multi-round Q&A steps.

[0202] In a feasible design, for the original question that does not need to enter the multi-round Q&A session, the Q&A solution can be implemented in the following way:

[0203] S170, If the original question does not support multi-round Q&A, or the original question is the first question of this multi-round Q&A, perform a recall process on the original question based on the database to determine the second target text segment associated with the original question.

[0204] Exemplarily, if the original question does not support multi-turn Q&A, the following method is used to perform recall processing on the original question based on the database to further improve the recall processing rate:

[0205] Perform recall processing on the original question based on the data sources in the database other than the data sources cited in the historical answers.

[0206] S180, input the second target text segment and the original question into the first large language model through a prompt instruction to obtain the answer output by the first large language model.

[0207] For the method of performing recall processing on the original question, refer to the description in S130 and will not be elaborated here.

[0208] In the above example, if the original question does not support multi-turn Q&A, it means that it is not related to the data sources cited in the historical answers, then recall needs to be performed based on the data sources in the database.

[0209] After obtaining the original question in the embodiment of the present application, by judging whether the original question supports multi-turn Q&A, it is possible to determine whether to enter multi-turn Q&A, thereby reducing the resource consumption in the overall Q&A process. After entering the multi-turn Q&A session, by combining the historical conversation record to reconstruct the original question, more background information can be supplemented or ambiguous expressions can be corrected, thereby improving the integrity and / or accuracy of the question, making the text segments recalled by the new question more accurately match the user's needs, and improving the multi-turn Q&A ability of the Q&A solution. Therefore, it can solve the problem that when there are defects in the user's question and effective data cannot be retrieved or the retrieval effect is poor, it will directly affect the final answer effect, and may lead to situations such as incomplete answers and wrong answers. It realizes the continuous dialogue interaction effect with the user and provides a more anthropomorphic dialogue interaction experience for the user. Among them, since the original question supports multi-turn Q&A, it means that each data source cited in the historical answer in its corresponding historical conversation record contains relevant document materials, that is, the answer to the original question can probably be obtained through these data sources. Therefore, compared with performing recall processing on the new question based on the database, performing recall processing on the new question based on each data source cited in the historical answer in the historical conversation record can further improve the recall rate of the new question.

[0210] Furthermore, the present application determines the target text segment most relevant to the question by comparing the similarity between the first candidate text segment obtained based on the original question and the second candidate text segment obtained based on the new question. It realizes the integration of the recall results of both and selects the text segment with the highest similarity, which can avoid the limitations that may be brought by a single data source and more comprehensively cover the information related to the question. Therefore, inputting the accurate target text segment and the corresponding question into the first large language model through a prompt instruction can enable the first large language model to output an answer that can solve the question as much as possible, thereby improving the accuracy of the multi-round question-and-answer solution of the present application. Compared with the current solution of inputting the complete historical conversation record and the question into the first large language model, the present application only inputs the first target text segment and the corresponding question into the first large language model, enabling the first large language model to give an answer without understanding the above information, improving the response speed of the system, and being able to provide a more convenient information consultation service function for users. Moreover, since the amount of data processed by the first large language model is reduced, the consumption of GPU resources can also be reduced.

[0211] Combined with the above embodiments, the present application also provides a question-and-answer system based on a large language model, including:

[0212] A question acquisition module for acquiring the original question input by the user;

[0213] A question reconstruction module for, if the original question supports multi-round question-and-answer, reconstructing the original question into a new question for the purpose of improving the integrity and / or accuracy of the question according to the historical conversation record corresponding to the original question;

[0214] A recall processing module for performing a recall process on the original question based on the database to determine the first candidate text segment associated with the original question and its corresponding similarity;

[0215] The recall processing module is further configured to perform a recall process on the new question based on each data source cited in the historical answer in the historical conversation record to determine the second candidate text segment associated with the new question and its corresponding similarity;

[0216] A decision-making module for determining the text segment with the highest similarity among the first candidate text segment and the second candidate text segment as the first target text segment;

[0217] An answer generation module for inputting the first target text segment and the corresponding question into the first large language model through a prompt instruction to obtain the answer output by the first large language model

[0218] In a feasible design, the system further includes a dialogue management module, which is used to call the second large language model to summarize the historical dialogue record corresponding to the original question to obtain the first key information of the historical dialogue record; and send the first key information and the historical question in the historical dialogue record to the question reconstruction module;

[0219] The question reconstruction module realizes sentence reconstruction of the original question to generate a new question for the purpose of improving the integrity and / or accuracy of the question based on the historical dialogue record corresponding to the original question in the following way:

[0220] Call the third large language model to perform sentence reconstruction on the original question for the purpose of improving the integrity and / or accuracy of the question based on the first key information and the historical question in the historical dialogue record to generate a new question.

[0221] In a feasible design, the question reconstruction module realizes sentence reconstruction of the original question to generate a new question for the purpose of improving the integrity and / or accuracy of the question in the following way:

[0222] Generate a new question by supplementing the missing sentence elements of the original question.

[0223] In a feasible design, the question reconstruction module realizes sentence reconstruction of the original question to generate a new question for the purpose of improving the integrity and / or accuracy of the question in the following way, including:

[0224] Identify and correct the typos in the original question to generate a new question.

[0225] In a feasible design, the question reconstruction module realizes sentence reconstruction of the original question to generate a new question for the purpose of improving the integrity and / or accuracy of the question in the following way, including:

[0226] Identify and delete the redundant words in the original question to generate a new question.

[0227] In a feasible design, the dialogue management module is also used to determine whether the original question supports multi-round Q&A in the following way:

[0228] If the original question is not the first question of this multi-round Q&A, obtain the data source type cited by the first large language model when generating each historical answer in the historical dialogue record;

[0229] If the data source type corresponding to each historical answer does not support multi-round Q&A, determine that the original question does not support multi-round Q&A;

[0230] Otherwise, determine that the original question supports multi-round Q&A.

[0231] In a feasible design, the recall processing module is further configured to, if the original question does not support multi-turn Q&A, or the original question is the first question of the current multi-turn Q&A, perform recall processing on the original question to determine a second target text segment associated with the original question;

[0232] The answer generation module is further configured to input the second target text segment and the original question into the first large language model through a prompt instruction to obtain an answer output by the first large language model.

[0233] In a feasible design, the recall processing module implements recall of questions in the following manner:

[0234] Extract at least one second key information from the target data, where the second key information includes the keyword and summary information of the target data, and the target data is data with information related to the field of the original question;

[0235] Determine at least one first target data in the target data according to the second key information, and each first target data corresponds to at least one second key information;

[0236] Based on the comparison result of the semantic vectors corresponding to the target data and the question to be processed, determine at least one second target data, where the second target data and the first target data are data in the target data;

[0237] If at least one association relationship is extracted from the target data, according to the extracted association relationship, determine the target association data of the first target data and the second target data, where the target association data is data having an association relationship with the first target data and the second target data, and the association relationship is the data association information in the target data;

[0238] According to the text similarity and semantic vector similarity between the first target data, the second target data, the target association data and the question to be processed, calculate the target similarity between the first target data, the second target data, the target association data and the question to be processed respectively;

[0239] Determine the third target data corresponding to the highest m target similarities as the candidate text segments associated with the question to be processed, where the third target data is data in the first target data, the second target data and the target association data.

[0240] In a feasible design, the recall processing module is implemented in the following manner. According to the text similarity and semantic vector similarity between the first target data, the second target data, the target association data and the question to be processed, calculate the target similarity between the first target data, the second target data, the target association data and the question to be processed respectively:

[0241] Construct a set based on the first target data, the second target data, and the target association data. The set includes all the first target data, the second target data, the target association data, and the corresponding text similarity and semantic vector similarity between the first target data, the second target data, the target association data, and the problem to be processed.

[0242] Perform normalization processing based on the text similarity between the first target data, the second target data, and the target association data in the set and the highest and lowest values of the text similarity and semantic vector similarity in the set to obtain a normalized text similarity value.

[0243] Calculate the target similarity based on the normalized text similarity value and the semantic vector similarity.

[0244] For other implementation manners and effects of the above system, refer to the descriptions in the embodiments of the question-answering method based on the large language model, which will not be elaborated here.

[0245] The basic principles of the present application are described above in combination with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations. It cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the above-disclosed specific details are only for the purposes of illustration and easy understanding, rather than limitations. The above details do not limit the present application to necessarily adopt the above specific details to implement.

[0246] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps is not strictly limited in order, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. Their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0247] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms meaning "including but not limited to" and can be used interchangeably with each other. The words "or" and "and" used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0248] It should also be noted that in the apparatuses, equipment, and methods of this application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this application.

[0249] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0250] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A question-answering method based on a large language model, characterized in that: include: Get the original question entered by the user; If the original question supports multiple rounds of question answering, based on the historical conversation records corresponding to the original question, reconstruct the original question to generate a new question for the purpose of improving the completeness and / or accuracy of the question; Recalling the original question based on the database to determine a first candidate text segment associated with the original question and its corresponding similarity; Recalling the new question based on the data sources referenced by the historical answers in the historical conversation records, and determining a second candidate text segment associated with the new question and its corresponding similarity; Determine the text segment with the highest similarity between the first candidate text segment and the second candidate text segment as the first target text segment; The first target text segment and the corresponding question are input into the first large language model through the prompt instruction to obtain the answer output by the first large language model.

2. The method according to claim 1, characterized in that The step of reconstructing the original question to generate a new question for the purpose of improving the completeness and / or accuracy of the question according to the historical conversation record corresponding to the original question includes: Calling the second largest language model to summarize the historical conversation record corresponding to the original question to obtain the first key information of the historical conversation record; The third language model is called to reconstruct the original question according to the first key information and the historical questions in the historical conversation record to generate a new question for the purpose of improving the completeness and / or accuracy of the question.

3. The method according to claim 2, characterized in that The sentence reconstruction of the original question to generate a new question for the purpose of improving the completeness and / or accuracy of the question includes: A new question is generated by supplementing the missing sentence constituent elements of the original question.

4. The method according to claim 2, characterized in that: The sentence reconstruction of the original question to generate a new question for the purpose of improving the completeness and / or accuracy of the question includes: A new question is generated by identifying and correcting typos in the original question.

5. The method according to claim 2, characterized in that: The sentence reconstruction of the original question to generate a new question for the purpose of improving the completeness and / or accuracy of the question includes: Redundant words in the original question are identified and deleted to generate a new question.

6. The method according to any one of claims 1 to 3, characterized in that The method further comprises: If the original question is not the first question of this multi-round question-answering, obtaining the data source type referenced by the first language model when generating each historical answer in the historical conversation record; If the data source type corresponding to each historical answer does not support multiple rounds of question answering, it is determined that the original question does not support multiple rounds of question answering; Otherwise, it is determined that the original question supports multi-round question and answer.

7. The method according to any one of claims 1 to 3, characterized in that The method further comprises: If the original question does not support multiple rounds of question answering, or the original question is the first question of this multiple rounds of question answering, recalling the original question based on the database to determine a second target text segment associated with the original question; The second target text segment and the original question are input into the first large language model through a prompt instruction to obtain an answer output by the first large language model.

8. The method according to any one of claims 1 to 3, characterized in that The recall process includes: Extracting at least one second key information from the target data, wherein the second key information includes keywords and summary information of the target data, and the target data is data having relevant information of the field to which the original question belongs; Determine at least one first target data in the target data according to the second key information, each of the first target data corresponds to at least one of the second key information; Determine at least one second target data based on a comparison result between the target data and the semantic vector corresponding to the problem to be processed, wherein the second target data and the first target data are data in the target data; If at least one association relationship is extracted from the target data, target association data of the first target data and the second target data are determined according to the extracted association relationship, the target association data being data that has an association relationship with the first target data and the second target data, and the association relationship being data association information in the target data; Calculate target similarities between the first target data, the second target data, the target-related data and the problem to be processed, respectively, according to text similarities and semantic vector similarities between the first target data, the second target data, the target-related data and the problem to be processed; The third target data corresponding to the m target similarities with the highest scores is determined as the candidate text segment associated with the problem to be processed, wherein the third target data is data among the first target data, the second target data and the target associated data.

9. The method according to claim 8, characterized in that The step of calculating target similarities between the first target data, the second target data, the target associated data and the problem to be processed respectively according to the text similarities and semantic vector similarities between the first target data, the second target data, the target associated data and the problem to be processed comprises: Constructing a set according to the first target data, the second target data and the target-related data, the set including all the first target data, the second target data, the target-related data and the text similarity and semantic vector similarity of the first target data, the second target data, the target-related data and the problem to be processed; Performing normalization processing according to the text similarity between the first target data, the second target data, and the target-related data in the set and the problem to be processed, and the highest value and the lowest value of the text similarity and the semantic vector similarity in the set, to obtain a normalized value of text similarity; The target similarity is calculated according to the normalized value of the text similarity and the semantic vector similarity.

10. A question-answering system based on a large language model, characterized in that: include: The question acquisition module is used to obtain the original question input by the user; A question reconstruction module is used to reconstruct the original question to generate a new question for the purpose of improving the completeness and / or accuracy of the question, if the original question supports multiple rounds of question answering, based on the historical dialogue records corresponding to the original question; A recall processing module, configured to perform recall processing on the original question based on a database, and determine a first candidate text segment associated with the original question and its corresponding similarity; The recall processing module is also used to perform recall processing on the new question based on the data sources referenced by the historical answers in the historical conversation records, and determine the second candidate text segment associated with the new question and its corresponding similarity; A decision module, configured to determine the text segment with the highest similarity between the first candidate text segment and the second candidate text segment as a first target text segment; The answer generation module is used to input the first target text segment and the corresponding question into the first large language model through the prompt instruction to obtain the answer output by the first large language model.