Method and device for generating answers
By performing user problem rewriting and related text retrieval concurrently, and retrieving when preset conditions are not met, the existing system solves the problems of fuzzy expressions, complex problems and low retrieval accuracy, and improves the efficiency of answer generation.
Patent Information
- Application Number
- CN202510122527.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-02
AI Technical Summary
When the existing knowledge base question and answer system deals with fuzzy expressions, complex problems and low retrieval accuracy, there are problems such as missing recalls and irrelevant answers.
Concurrently performs two processes: user problem rewriting and related text retrieval. If the retrieved related text does not meet the preset conditions, re-retrieve based on the rewritten user problem to avoid waiting for the user problem to rewritten.
It significantly improves the efficiency of answer generation, reduces the system's running time, and is suitable for scenarios with high efficiency requirements.
Smart Images

Figure CN119917634A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification belong to the field of artificial intelligence, and more particularly, to a method and device for generating answers. Background Art
[0002] Retrieval-augmented generation (RAG) technology has become an important strategy to address existing challenges and improve model performance, especially showing great potential in dealing with knowledge-intensive tasks that require high accuracy and reliability. RAG technology not only promotes the real-time updating of knowledge and the seamless integration of information in specific fields, but also effectively alleviates the problems of hallucinations and reliance on outdated information that may occur in the model by increasing transparency and traceability.
[0003] RAG technology can be applied to a variety of scenarios such as knowledge base question and answer. Taking the knowledge base question and answer scenario as an example, it can first receive user questions, then retrieve documents related to the user questions from the knowledge base (or corpus), and pass these documents as context to the language big model (hereinafter referred to as the big model). Finally, the big model uses these documents to generate answers to user questions.
[0004] It should be noted that although the knowledge base question-answering system implemented by the above-mentioned RAG technology has a high operating efficiency, it still has problems such as missed recall and irrelevant answer content of the large model. Summary of the invention
[0005] The purpose of the present invention is to provide a method and device for generating answers, so as to improve the efficiency of generating answers.
[0006] A first aspect of this specification provides a method for generating an answer, comprising:
[0007] Receive user questions;
[0008] Based on the user question, a first process and a second process are performed in parallel; the first process includes rewriting the user question; the second process includes retrieving one or more first texts matching the user question;
[0009] Determining whether the one or more first texts meet a preset condition;
[0010] If the judgment result indicates that the preset condition is not met, one or more second texts are obtained based on the rewritten user question, and the target answer to the user question is generated based on the one or more second texts using the first large model.
[0011] A second aspect of the present specification provides a device for generating an answer, comprising:
[0012] A receiving unit, used for receiving user questions;
[0013] An execution unit, configured to execute a first process and a second process in parallel based on the user question; the first process includes rewriting the user question; the second process includes retrieving one or more first texts matching the user question;
[0014] A judging unit, used to judge whether the one or more first texts meet a preset condition;
[0015] A generating unit is used to obtain one or more second texts based on the rewritten user question if the judgment result indicates that the preset condition is not met, and to generate a target answer to the user question based on the one or more second texts using the first large model.
[0016] A third aspect of the present specification provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method described in the first aspect.
[0017] A fourth aspect of the specification provides a computing device, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method described in the first aspect is implemented.
[0018] The fifth aspect of this specification provides a computer program product, including a computer program / instruction, which implements the steps of the method described in the first aspect when executed by a processor.
[0019] The method and device for generating answers provided in one or more embodiments of the present specification concurrently execute two processes, namely, rewriting the user question and retrieving related texts. This allows the user to directly search for the rewritten user question if the retrieved related texts do not satisfy preset conditions. There is no need to wait for the user question to be rewritten, thereby greatly improving the efficiency of answer generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0021] Figure 1 It is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification;
[0022] Figure 2 A flowchart of a method for generating an answer according to an embodiment of the present specification is shown;
[0023] Figure 3A flowchart showing a method of generating an answer in one example of this specification;
[0024] Figure 4 A schematic diagram of a device for generating answers according to an embodiment of the present specification is shown. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.
[0026] As mentioned above, RAG technology can be applied to knowledge base question answering scenarios. However, the knowledge base question answering system implemented using RAG technology has the following disadvantages:
[0027] (1) Knowledge recall is easily missed: Since the current system lacks the ability to handle user questions, the recalled text fragments are generally limited in length for user questions that are vague or contain typos, which can easily cut off key knowledge and result in incomplete content input into the large model.
[0028] (2) Unable to answer difficult or complex questions: The current system uses raw queries for retrieval and cannot retrieve all the dependent knowledge for all complex questions at once, making it impossible to answer complex questions.
[0029] (3) Low retrieval accuracy: The current system does not have sufficient understanding of the semantics of user questions, which results in the retrieved text fragments being irrelevant to the user questions.
[0030] Based on this, the inventor of this solution proposes to execute the two processes of user question rewriting and related text retrieval concurrently. If the retrieved related text does not meet the preset conditions, the retrieval can be directly based on the rewritten user question without waiting for the user question to be rewritten. This can greatly improve the efficiency of answer generation.
[0031] Figure 1 It is a schematic diagram of an implementation scenario of an embodiment disclosed in this specification. Figure 1In the method, after receiving a user question, a first process and a second process are performed in parallel based on the user question, wherein the first process includes rewriting the user question and the second process includes retrieving text related to the user question. Afterwards, it is determined whether the retrieved related text meets a preset condition. If the determination result indicates that the preset condition is met, a target answer to the user question is generated based on the related text using the large model. If the determination result indicates that the preset condition is not met, a re-retrieval is performed based on the rewritten user question, and a target answer to the user question is generated based on the re-retrieved text using the large model.
[0032] Figure 2 The flowchart of the method for generating an answer according to one embodiment of the present specification is shown. The method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 2 As shown, the method may include the following steps:
[0033] Step S202: receiving user questions.
[0034] Among them, the user questions here may only include one question, such as, "What should I do if I can't pay back the money for Huabei?", "How to repay Huabei?" and "What is the profit of Yu'ebao?" etc.
[0035] Of course, in practice, the above user questions may also include two or more questions, such as, "What interesting people did you meet during your travels? What unforgettable things did you experience? What unique foods did you taste?" and "What fun places are there in Chengdu? Are there any specially recommended foods suitable for family gatherings?" etc.
[0036] Step S204: Based on the user question, the first process and the second process are executed in parallel.
[0037] In practice, before executing the first process and the second process in parallel, the user's question may be first identified for user intent (i.e., understanding the actual needs or purposes of the user input) to obtain the target intent. Afterwards, if the identified target intent is not an irrelevant question type, the first process and the second process are executed in parallel.
[0038] The above-mentioned identification of user intent for user questions can be achieved in any of the following ways:
[0039] The first one is to input the user question into a classification model based on several intent categories and determine the target intent based on the output of the classification model.
[0040] The intent categories described in this specification may be designed according to business scenarios. In one example, the above-mentioned several intent categories may include irrelevant question categories, self-introduction categories, and knowledge retrieval categories.
[0041] The above classification model can be implemented as Naive Bayes, support vector machine, etc., or as BERT or a Transformer-based neural network model, etc.
[0042] The second method is to calculate the similarity between the user question and each sample question with known intent categories, and determine the intent category of the sample question with the largest similarity as the target intent of the user question.
[0043] The similarity mentioned above may be, for example, cosine similarity, Euclidean distance, Manhattan distance, Pearson correlation coefficient, and the like.
[0044] In a more specific embodiment, for each word in the user's question, a pre-trained word vector model (such as Word2Vec, etc.) can be used to convert it into a word vector, and then these word vectors are combined (added or averaged) to obtain a semantic vector of the user's question. Finally, based on the semantic vector of the user's question and the semantic vector of the sample question, the similarity between the user's question and each sample question is calculated.
[0045] The third method is to construct a prompt text based on the user question and several intent categories, which indicates to determine the intent category matching the user question from several intent categories. The constructed prompt text is input into the large model to obtain the target intent.
[0046] The above-mentioned large model can be a BERT model, a GPT series (for example, GPT-2, GPT-3, GPT-4, GPT-4V or GPT-4o) model, etc.
[0047] The above prompt text will indicate the user question and several intent categories. Of course, in practice, the prompt text will also indicate the term explanation and output format of the domain terms contained in the user question, so that the large model can more accurately identify the user intent.
[0048] Returning to step S204 , the first processing in step S204 may include rewriting the user question.
[0049] In this solution, the purpose of rewriting user questions is to convert them into a more accurate or targeted form to improve search efficiency. Common forms of user question rewriting include: extracting keywords or entities, deleting stop words, correcting typos, replacing synonyms, correcting spelling errors, adjusting phrase structures, and extracting multiple target questions from user questions.
[0050] Among them, extracting multiple target questions from user questions generally refers to user questions that contain two or more questions. For example, for the above user questions: "What interesting people did you meet during your trip? What unforgettable things did you experience? What unique foods did you taste?", the multiple target questions extracted are: "What interesting people did you meet during your trip?", "What unforgettable things did you experience?", and "What unique foods did you taste?".
[0051] In one embodiment, some regular matching rules may be used to rewrite the user question.
[0052] In another embodiment, the user question may also be rewritten using a BERT model or a generative large model.
[0053] Of course, in practice, natural language processing technology can also be used to analyze user questions, and user questions can be expanded, simplified, or reconstructed based on contextual information or historical data to better match the document content in the database. For example, for the user question "How to make coffee at home", after rewriting it, the following user question can be obtained: "Home coffee brewing method". Generally speaking, the rewritten user question is more in line with the user's real needs, so more relevant results can usually be obtained based on the rewritten user question.
[0054] In addition, the user question may also be rewritten in combination with the target intention of the user question, which is not limited in this specification.
[0055] The second process in step S204 may include retrieving one or more texts txt1 matching the user question.
[0056] In this solution, the above-mentioned one or more texts txt1 can be retrieved based on a multi-channel recall mechanism, that is, multiple retrieval strategies are used to retrieve the above-mentioned one or more texts txt1 from different data sources, so as to integrate data from different channels and ensure that the recalled text txt1 is as comprehensive as possible and contains all the required reference information.
[0057] Among them, the above-mentioned multiple search strategies may include but are not limited to BM25 search (a search method based on term frequency inverse document frequency (TF-IDF) and text length), vector search, metadata search, multi-query search, etc.
[0058] The above-mentioned different data sources may include but are not limited to document knowledge base, standard question and answer base, web page address or link, etc. Among them, the document knowledge base records multiple text fragments obtained by segmenting multiple documents. In a specific embodiment, the length of the text fragments in the document knowledge base is generally short, for example, generally about 256 / 512 bits, because too long text fragments will introduce too much noise. The standard question and answer base records multiple standard questions and their corresponding standard answers.
[0059] When the data source is a document knowledge base, the similarity between the user question and each text fragment in the document knowledge base can be calculated to retrieve the text fragment matching the user question from the document knowledge base, and the one or more texts txt1 can be determined based on the retrieved text fragments.
[0060] Taking the search strategy of vector search as an example, the similarity can be calculated based on the user question and the semantic vector of each text segment. Here, the method for determining the semantic vector of each text segment is similar to the method for determining the semantic vector of the user question, that is, it is obtained by combining the word vectors of each word contained therein.
[0061] The text segments that match the user question may refer to text segments whose corresponding similarities are greater than a preset threshold, or may refer to k text segments that are ranked top in terms of similarity.
[0062] When the data source is a standard question and answer library, the similarity between the user question and each standard question in the standard question and answer library can be calculated to retrieve a standard question matching the user question from the standard question and answer library, and one or more texts txt1 can be determined based on the standard answer corresponding to the standard question.
[0063] Taking the search strategy of vector search as an example, the similarity can be calculated based on the semantic vectors of the user question and each standard question. The method for determining the semantic vector of each standard question is similar to the method for determining the semantic vector of the user question.
[0064] The standard questions matching the user questions may refer to standard questions whose corresponding similarities are greater than a preset threshold, or may refer to k standard questions ranked top according to the similarities.
[0065] When the data source is a web page address, a text segment matching the user question is selected from the web page contents corresponding to the multiple web page addresses, and one or more texts txt1 are determined based thereon.
[0066] In one example, the webpage content may be divided into multiple text segments according to paragraphs, and then the similarities between the user question and the multiple text segments are calculated, and then the text segment matching the user question is selected.
[0067] Of course, in practice, the text segments may also be divided according to titles, and this specification does not limit this.
[0068] It should be noted that one or more data sources such as a document knowledge base, a standard question and answer base, and a web page address can be selected to retrieve one or more texts txt1 that match the user's question.
[0069] Step S206, determining whether the retrieved one or more texts txt1 meet a preset condition.
[0070] In one embodiment, the user question and one or more texts txt1 can be input into the target scoring model to obtain the matching score between the one or more texts txt1 and the user question. It is determined whether there is a text txt1 with a corresponding matching score greater than a preset threshold in the one or more texts txt1. If there is, the preset condition is met; if not, the preset condition is not met.
[0071] In a more specific embodiment, the target scoring model includes a cross encoder and a classifier;
[0072] Among them, after any text txt1 is input into the target scoring model, a cross encoder can be used to encode the combination of the user question and the any text txt1 to obtain a text embedding vector, and a classifier can be used to determine the matching score between the any text txt1 and the user question based on the text embedding vector.
[0073] Of course, in practice, the above-mentioned determination of whether the retrieved one or more texts txt1 meet the preset conditions may also include determining whether there is a txt1 whose corresponding intent category is the target intent in the one or more txt1.
[0074] In addition, whether one or more texts txt1 meet preset conditions may also be determined based on factors such as content quality or security, which is not limited in this specification.
[0075] Step S208: If the judgment result indicates that the preset condition is not met, one or more texts txt2 are obtained based on the rewritten user question, and the target answer to the user question is generated based on the one or more texts txt2 using the large model.
[0076] It should be understood that when judging whether one or more texts txt1 meet the preset conditions based on the matching score, if the judgment result indicates that the preset conditions are not met, it means that the matching score between the one or more texts txt1 and the user question is less than the preset threshold, that is, the matching score is low, and thus it can be considered that the user question is more difficult, for example, it may contain two or more questions, and the retrieved one or more texts txt1 are not comprehensive enough, so the matching score is low.
[0077] In this solution, for user questions with greater difficulty (also called difficult questions or complex questions), re-retrieval can be performed based on the rewritten user questions.
[0078] It should be understood that since the user question is rewritten during the process of retrieving text txt1, when it is necessary to re-retrieve based on the rewritten user question, there is no need to wait for the user question to be rewritten, which can greatly improve the efficiency of answer generation.
[0079] Taking the case where the user question includes two or more questions as an example, the rewritten user question is a plurality of target questions extracted from the original user question, so that the one or more texts txt2 are retrieved based on the plurality of target questions.
[0080] In this solution, the retrieval method of text txt2 is similar to that of txt1 in this article, that is, it can also be retrieved based on a multi-way recall mechanism. The specific retrieval process is described above and will not be repeated here.
[0081] It should be noted that in order to avoid introducing too much noise, the length of text fragments in the document knowledge base is generally short, which may lead to the loss of important information. In order to ensure that sufficient information can be provided for the large model, this solution will expand one or more of the above-mentioned texts txt2.
[0082] In which, when the one or more texts txt2 are selected from the document or web page content, the one or more texts txt2 can be expanded based on the document or web page content to obtain one or more expanded texts.
[0083] In a specific embodiment, for any text txt2, the text txt2 can be extended based on the context and the context of the text txt2 of a predetermined length in the document or webpage content. It should be understood that the length of the text txt2 extended in this way is generally within the predetermined range.
[0084] In another specific embodiment, for any text txt2, the text txt2 can be expanded based on the content text in the document or webpage content that belongs to the same title as the text txt2.
[0085] Of course, in practice, for any text txt2, similar text fragments of the text txt2 can be obtained by calculating the similarity between the user question and each text fragment in the document or web page content where the arbitrary text txt2 is located, and then the text txt2 can be expanded based on the similar text fragment.
[0086] It should be understood that when one or more texts txt2 are also expanded, a large model can be used to generate a target answer to the user's question based on the one or more expanded texts.
[0087] In addition, for the above-mentioned extended text, its length / content can also be compressed to reduce the noise in the input large model, avoid the illusion of the large model, and avoid the problem of excessively long input text.
[0088] Finally, when the number of the above-mentioned text txt2 is 1, the user question and the text txt2 can be input into the large model. When the number of the above-mentioned text txt2 is multiple, based on the matching scores of the multiple text txt2 and the user question, k text txt2 with the highest matching scores can be screened out from the multiple text txt2, and the k text txt2 and the user question can be input into the large model.
[0089] Of course, when the text txt2 is also expanded or expanded and compressed, the corresponding expanded text or compressed expanded text is input into the large model.
[0090] The large model in the above step S208 can be a BERT model, a GPT series (for example, GPT-2, GPT-3, GPT-4, GPT-4V or GPT-4o) model, etc.
[0091] Specifically, a prompt text can be constructed based on the user question and the finally screened text txt2 (or the expanded text or the compressed expanded text), wherein the prompt text is indicated to use the text txt2 (or the expanded text or the compressed expanded text) as a context to generate an answer to the user question. The prompt text is input into the large model to obtain the target answer to the user question.
[0092] Of course, in practice, the above prompt text may also indicate customized content such as generation rules, speech techniques, and refusal to answer words, which is not limited in this specification.
[0093] In summary, the large model can generate coherent and targeted answers to user questions based on the final filtered text xt2 (or extended text or compressed extended text) and its own knowledge. This design enables it to provide detailed and appropriate answers even when faced with complex questions.
[0094] The above is an explanation of the judgment result indicating that the preset conditions are not met. When the judgment result indicates that the preset conditions are met, the large model can be used to generate the target answer to the user's question based on the above one or more texts txt1.
[0095] In the case where the number of the above-mentioned text txt1 is 1, the user question and the text txt1 can be input into the big model. In the case where the number of the above-mentioned text txt1 is multiple, based on the matching scores of the multiple text txt1 and the user question, k text txt1 with the highest matching scores can be screened out from the multiple text txt1, and the k text txt1 and the user question can be input into the big model, so as to obtain the target answer to the user question.
[0096] It should be understood that when judging whether one or more texts txt1 meet the preset conditions based on the matching score, if the judgment result indicates that the preset conditions are met, it means that the matching score between the one or more texts txt1 and the user question is greater than the preset threshold, that is, the matching score is high, so it can be considered that the user question is relatively simple.
[0097] It can be seen that in the solution, simple problems and complex problems will be handled differently. Specifically, for simple problems, a simple query method is used to avoid consuming more resources. For complex problems, a combination of multiple strategies is used to increase the recall rate of complex problems and avoid losing erroneous information.
[0098] Figure 3 A flow chart showing a method for generating an answer in one example of the present specification. Figure 3 In the process, for the received user question, the intent of the user question is first identified to obtain the target intent. When the target intent is an irrelevant question class, the process ends directly. When the target intent is not an irrelevant question class, the first process and the second process can be performed in parallel, wherein the first process includes rewriting the user question. The second process includes retrieving one or more texts txt1 that match the user question, and sorting the multiple texts txt1 when there are multiple texts txt1. Afterwards, it can be determined whether the one or more texts txt1 meet the preset conditions, and when the preset conditions are met, the target answer to the user question is generated based on the one or more texts txt1 using the large model. When the preset conditions are not met, the search can be performed again based on the rewritten user question, and the retrieved one or more texts txt2 can be sorted, and the final several texts txt2 can be selected from them for expansion. Finally, the target answer to the user question is generated based on several extended texts using the large model.
[0099] To sum up, this solution can generate more accurate and informative answers to user questions.
[0100] The following describes the shortcomings that this solution can overcome and the technical effects achieved compared to the existing solutions.
[0101] Disadvantages that can be overcome:
[0102] 1. This solution pre-searches and re-arranges the original user questions, and sets up judgment strategies to classify the difficulty of user questions. For simple questions, a simple query method is used to avoid consuming more resources. For difficult or complex questions, a combination of multiple strategies is used to increase the recall rate of difficult or complex questions and avoid losing erroneous information.
[0103] 2. This solution performs two concurrent requests, namely, rewriting and retrieving user questions. If it is difficult to determine the user question, there is no need to wait for the user question to be rewritten, thus saving the time of the overall system.
[0104] Technical effects achieved:
[0105] 1. This solution is applicable to all kinds of intelligent question-and-answer scenarios, supports multiple data sources such as documents, standard question-and-answer libraries, web pages, etc., and has greatly improved the effect compared with existing solutions.
[0106] 2. This solution can reduce the system's operating time by about 40% and is suitable for use in scenarios with more stringent efficiency requirements.
[0107] The following is an explanation of the innovative features of this solution:
[0108] (1) Dynamic query processing mechanism: The original user questions are pre-retrieved and re-ranked, and the complexity of the user questions is classified based on this. This approach allows the system to adopt different processing strategies for questions of different difficulty: simple questions are responded to quickly, while complex questions are answered through more sophisticated methods to ensure the quality of the answers. This approach can effectively balance the relationship between resource utilization and answer quality.
[0109] (2) Concurrent execution optimization: The system can start the user question rewriting and direct retrieval process in parallel, rather than waiting for each step to complete in sequence. This approach significantly reduces the overall response time and improves the speed of service, which is especially important for application scenarios that require instant feedback.
[0110] (3) Multi-source data support and performance improvement: Compared with existing solutions, this solution can better adapt to various types of data sources (such as documents, standard question and answer libraries, web pages, etc.), and significantly reduce operating costs while maintaining or exceeding the existing technical level.
[0111] Corresponding to the above-mentioned method for generating an answer, an embodiment of the present specification also provides a device for generating an answer, such as Figure 4 As shown, the device may include:
[0112] The receiving unit 402 is used to receive user questions.
[0113] The execution unit 404 is used to execute a first process and a second process in parallel based on the user question, wherein the first process includes rewriting the user question, and the second process includes retrieving one or more first texts matching the user question.
[0114] The judging unit 406 is configured to judge whether one or more first texts meet a preset condition.
[0115] The generating unit 408 is used to obtain one or more second texts based on the rewritten user question if the judgment result indicates that the preset condition is not met, and generate a target answer to the user question based on the one or more second texts using the first large model.
[0116] In one embodiment, the execution unit 404 is specifically configured to execute one or more of the following:
[0117] Retrieving text fragments matching the user's question from a document knowledge base, and determining one or more first texts based thereon; the document knowledge base records a plurality of text fragments obtained by segmenting a plurality of documents;
[0118] Retrieving a standard question matching the user's question from a standard question-and-answer database, and determining one or more first texts based on a standard answer corresponding to the standard question; the standard question-and-answer database records a plurality of standard questions and their corresponding standard answers;
[0119] A text segment matching the user question is selected from the web page contents corresponding to the multiple web page addresses, and one or more first texts are determined based thereon.
[0120] In one embodiment, the apparatus further comprises:
[0121] An input unit 410 is used to input the user question and one or more first texts into a target scoring model to obtain a matching score between the one or more first texts and the user question;
[0122] The determination unit 406 is specifically used for:
[0123] It is determined whether there is a first text with a corresponding matching score greater than a preset threshold in the one or more first texts.
[0124] In a more specific embodiment, the target scoring model includes: a cross encoder and a classifier; the input unit 410 is specifically used for:
[0125] Using a cross encoder, encode the combination of the user question and any first text to obtain a text embedding vector;
[0126] Using a classifier, based on the text embedding vector, a matching score between the first text and the user question is determined.
[0127] In one embodiment, the apparatus further comprises:
[0128] The identification unit 412 is used to identify the user's intention for the user's question and obtain the target intention;
[0129] The execution unit 404 is specifically used for:
[0130] In the case where the target intent is not an irrelevant question class, the first process and the second process are performed in parallel.
[0131] In a more specific embodiment, the identification unit 412 is specifically configured to:
[0132] Input the user question into a classification model based on several intent categories, and determine the target intent based on the output of the classification model; or
[0133] Calculate the similarity between the user question and each sample question with known intent categories, and determine the intent category of the sample question with the greatest similarity as the target intent; or
[0134] A prompt text is constructed based on the user question and several intent categories, wherein an intent category matching the user question is determined from among several intent categories, and the prompt text is input into the second largest model to obtain the target intent.
[0135] In one embodiment, the apparatus further comprises:
[0136] An expansion unit 414 is used for, when the one or more second texts are selected from the document or webpage content, expanding the one or more second texts based on the document or webpage content to obtain one or more expanded texts;
[0137] The generating unit 408 is specifically used for:
[0138] A target answer to the user's question is generated based on one or more extended texts using the first model.
[0139] In one embodiment, the apparatus further comprises:
[0140] A compression unit 416, used for compressing one or more extended texts;
[0141] The generating unit 408 is specifically used for:
[0142] A target answer to the user's question is generated based on the compressed one or more expanded texts using the first large model.
[0143] The execution unit 404 is further configured to execute one or more of the following:
[0144] Extract keywords or entities, remove stop words, correct typos, replace synonyms, and extract multiple target questions from user questions.
[0145] In one embodiment, the generating unit 408 is further configured to generate a target answer to the user's question based on one or more first texts using the first large model if the judgment result indicates that a preset condition is met.
[0146] The functions of the functional units of the device in the above-mentioned embodiment of this specification can be implemented through the steps of the above-mentioned method embodiment. Therefore, the specific working process of the device provided by one embodiment of this specification will not be repeated here.
[0147] An answer generation device provided in an embodiment of the present specification can improve the efficiency of answer generation.
[0148] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute a combination of Figure 2 The method described.
[0149] According to another embodiment of the present invention, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the Figure 2 The method described.
[0150] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the medium or device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0151] The steps of the method or algorithm described in conjunction with the disclosure of this specification can be implemented in hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, register, hard disk, mobile hard disk, CD-ROM or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a server. Of course, the processor and the storage medium can also be present in the server as discrete components.
[0152] In the 1990s, improvements to a technology could be clearly distinguished as hardware improvements (for example, improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the method flow). However, with the development of technology, many improvements to the method flow today can be regarded as direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement in a method flow cannot be implemented using a hardware entity module. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask a chip manufacturer to design and produce a dedicated integrated circuit chip. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.
[0153] The controller can be implemented in any appropriate manner, for example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) that can be executed by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26k20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that in addition to implementing the controller in a purely computer-readable program code manner, the controller can be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method steps. Therefore, this controller can be considered as a hardware component, and the devices included therein for implementing various functions can also be regarded as structures within the hardware component. Or even, the devices for implementing various functions can be regarded as both software modules for implementing the method and structures within the hardware component.
[0154] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a server system. Of course, the present application does not exclude that with the development of computer technology in the future, the computer that implements the functions of the above embodiments may be, for example, a personal computer, a laptop computer, a vehicle-mounted human-computer interaction device, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0155] Although one or more embodiments of the present specification provide method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such a process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. For example, if the words first, second, etc. are used to represent the name, they do not represent any specific order.
[0156] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more of the present specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0157] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0158] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0160] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0161] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0162] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage, graphene storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0163] It should be understood by those skilled in the art that one or more embodiments of the present specification may be provided as a method, system or computer program product. Therefore, one or more embodiments of the present specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, one or more embodiments of the present specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0164] One or more embodiments of the present specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. One or more embodiments of the present specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0165] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. In the description of this specification, the description of the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.
[0166] The above description is only an example of one or more embodiments of the present specification and is not intended to limit one or more embodiments of the present specification. For those skilled in the art, one or more embodiments of the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims.
Claims
1. A method for generating an answer, comprising: Receive user questions; Based on the user question, executing a first process and a second process in parallel; The first processing includes rewriting the user question; The second processing includes retrieving one or more first texts matching the user question; Determining whether the one or more first texts meet a preset condition; If the judgment result indicates that the preset condition is not met, one or more second texts are obtained based on the rewritten user question, and the target answer to the user question is generated based on the one or more second texts using the first large model.
2. The method according to claim 1, wherein: The retrieving one or more first texts matching the user question includes performing one or more of the following: Retrieving text segments matching the user question from a document knowledge base, and determining the one or more first texts based thereon; the document knowledge base records a plurality of text segments obtained by segmenting a plurality of documents; Retrieving a standard question matching the user question from a standard question and answer library, and determining the one or more first texts based on a standard answer corresponding to the standard question; the standard question and answer library records a plurality of standard questions and their corresponding standard answers; A text segment matching the user question is selected from web page contents corresponding to a plurality of web page addresses, and the one or more first texts are determined based on the text segment.
3. The method according to claim 1, further comprising: Inputting the user question and the one or more first texts into a target scoring model to obtain a matching score between the one or more first texts and the user question; The determining whether the one or more first texts meet a preset condition includes: It is determined whether there is a first text with a corresponding matching score greater than a preset threshold among the one or more first texts.
4. The method according to claim 3, wherein: The target scoring model includes: a cross encoder and a classifier; the obtaining of the matching score between the one or more first texts and the user question includes: Using the cross encoder, encode the combination of the user question and any first text to obtain a text embedding vector; The classifier is used to determine a matching score between any first text and the user question based on the text embedding vector.
5. The method according to claim 1, wherein: The parallel execution of the first process and the second process comprises: Performing user intent recognition on the user question to obtain the target intent; In a case where the target intention is not an irrelevant question class, the first process and the second process are performed in parallel.
6. The method according to claim 5, wherein: The identifying the user intention of the user question includes: Input the user question into a classification model based on several intent categories, and determine the target intent according to the output of the classification model; or Calculate the similarity between the user question and each sample question with known intent categories, and determine the intent category of the sample question with the greatest similarity as the target intent; or A prompt text is constructed based on the user question and several intent categories, wherein the prompt text indicates to determine an intent category matching the user question from the several intent categories; the prompt text is input into the second largest model to obtain the target intent.
7. The method according to claim 1, wherein: Generating a target answer to the user question includes: In a case where the one or more second texts are selected from a document or webpage content, expanding the one or more second texts based on the document or webpage content to obtain one or more extended texts; A target answer to the user question is generated based on the one or more extended texts using the first large model.
8. The method according to claim 7, further comprising: Compressing the one or more extended texts; The step of using the first large model to generate a target answer to the user question based on the one or more extended texts includes: A target answer to the user question is generated based on the compressed one or more expanded texts using the first large model.
9. The method according to claim 1, wherein: Rewriting the user question includes performing one or more of the following: Extract keywords or entities, remove stop words, correct typos, replace synonyms, and extract multiple target questions from user questions.
10. The method according to claim 1, further comprising: If the judgment result indicates that the preset condition is met, the first large model is used to generate a target answer to the user question based on the one or more first texts.
11. A device for generating an answer, comprising: A receiving unit, used for receiving user questions; An execution unit, configured to execute a first process and a second process in parallel based on the user question; The first processing includes rewriting the user question; the second processing includes retrieving one or more first texts matching the user question; A judging unit, used to judge whether the one or more first texts meet a preset condition; A generating unit is used to obtain one or more second texts based on the rewritten user question if the judgment result indicates that the preset condition is not met, and to generate a target answer to the user question based on the one or more second texts using the first large model.
12. A computing device comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
RAG knowledge question-answering method and device based on fusion vector and keyword retrieval
CN117951274A
Steel industry intelligent question and answer method and device based on large model and medium
CN118051587A
Document question and answer method and device, electronic equipment and storage medium
CN118170887A
Text generation method and device, computer program product, electronic equipment and medium
CN118396123A
Knowledge question and answer method and related device
CN118484524A