Response method and electronic equipment

By combining iterative prediction with question-answering models, more accurate reply information is generated, which solves the problem of insufficient understanding of user intentions in existing technologies, improves the quality and efficiency of reply information, and reduces the number of user interactions and the waste of computing resources.

CN120633848APending Publication Date: 2025-09-12LENOVO (BEIJING) LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510726306.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing retrieval augmentation technology (RAG) lacks a deep understanding of user intent and subsequent conversation content when generating reply information, resulting in the output reply information being difficult to meet user needs. Users need to interact with the model repeatedly, wasting computing resources and affecting the user experience.

Method used

By generating the first predicted reply information corresponding to the user input information, performing iterative prediction operations, combining the question-answering model and the predicted feedback information to generate the target reply information, using the prediction model with a smaller number of parameters for preliminary rewriting, and the question-answering model with a larger number of parameters for precise recall, the user's multiple input processes are simulated to improve the recall effect.

Benefits of technology

It improves knowledge recall and accuracy, reduces the number of interactions between users and models, saves computing resources and time, and improves user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633848A_ABST
    Figure CN120633848A_ABST
Patent Text Reader

Abstract

The invention provides a response method which can be applied to the technical field of artificial intelligence. The response method comprises the following steps: according to target input information from a user, generating first prediction response information corresponding to the target input information; performing an iterative prediction operation based on the first prediction reply information to obtain prediction feedback information; and generating target reply information corresponding to the target input information based on prompt information including the target input information, the first predicted reply information and the predicted feedback information by using the question and answer model. The invention further provides an electronic device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, and more particularly to a response method and electronic device. Background Art

[0002] Retrieval-Augmented Generation (RAG) technology can generate responses to user queries by recalling relevant knowledge blocks from the knowledge base. To optimize recall and generate more accurate responses, the user's query is often rewritten. However, these rewriting methods are often limited to directly replacing synonyms or adjusting the grammar of the query. They lack a deep understanding of user intent and subsequent conversation content, resulting in responses that fail to meet user needs. Summary of the Invention

[0003] In view of the above problems, the present disclosure provides a response method and an electronic device.

[0004] According to the first aspect of the present disclosure, a response method is provided, including: generating first predicted reply information corresponding to the target input information based on target input information from a user; performing an iterative prediction operation based on the first predicted reply information to obtain predicted feedback information; and utilizing a question-answering model to generate target reply information corresponding to the target input information based on prompt information including the target input information, the first predicted reply information, and the predicted feedback information.

[0005] Optionally, the first predicted reply information and the predicted feedback information are generated by a prediction model, and the number of parameters of the prediction model is smaller than the number of parameters of the question-answering model.

[0006] Optionally, the prediction feedback information includes: prediction input information and second prediction reply information corresponding to the prediction input information; performing an iterative prediction operation based on the first prediction reply information to obtain the prediction feedback information includes: performing a first round of prediction operation on the first prediction reply information to obtain the first round of prediction input information and the first round of second prediction reply information corresponding to the first round of prediction input information; performing a second round of prediction operation on the first round of second prediction reply information to obtain the second round of prediction input information and the second round of second prediction reply information corresponding to the second round of prediction input information; iteratively performing the prediction operation to obtain multiple rounds of prediction input information corresponding to multiple rounds of prediction operations and multiple rounds of second prediction reply information corresponding to the multiple rounds of prediction input information.

[0007] Optionally, performing a round of prediction operation includes: generating K+1 round prediction input information based on the K round second prediction reply information obtained from the K round prediction operation, where K is a positive integer; and generating K+1 round second prediction reply information based on the K+1 round prediction input information.

[0008] Optionally, generating the K+1th round second prediction reply information based on the K+1th round prediction input information includes: retrieving the first context information from the user's local knowledge base based on the K+1th round prediction input information; and generating the K+1th round second prediction reply information based on the first context information.

[0009] Optionally, the first context information includes at least one of: predicted question association information and user background information; wherein, the user background information includes at least one of the following: historical query information, user business information, and professional field information.

[0010] Optionally, retrieving the first context information from the user's local knowledge base according to the K+1th round of prediction input information includes: retrieving the first context information from the user's local knowledge base according to the prompt word.

[0011] Optionally, utilizing a question-answering model, based on prompt information including target input information, first predicted reply information, and predicted feedback information, generating target reply information corresponding to the target input information includes: utilizing the question-answering model to retrieve second context information related to the prompt information; and outputting the target reply information based on the second context information.

[0012] Optionally, the prediction feedback information includes: multiple first prediction reply information corresponding to multiple prediction operations; performing a prediction operation includes: when it is determined that the Lth first prediction reply information obtained by the Lth prediction operation does not meet the target condition, generating the L+1th first prediction reply information according to the target input information, wherein the target condition includes at least one of the following: the number of predictions is equal to the threshold, and the satisfaction value of the first prediction reply information meets the target range.

[0013] The second aspect of the present disclosure provides a response device, including: a first generation module, used to generate first predicted reply information corresponding to the target input information based on the target input information from the user; an iterative prediction module, used to perform an iterative prediction operation based on the first predicted reply information to obtain predicted feedback information; a second generation module, used to utilize a question-answering model to generate target reply information corresponding to the target input information based on prompt information including the target input information, the first predicted reply information and the predicted feedback information.

[0014] A third aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more computer programs, wherein the one or more processors execute the one or more computer programs to implement the steps of the above method.

[0015] The fourth aspect of the present disclosure further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above method when the computer program or instructions are executed by a processor.

[0016] The fifth aspect of the present disclosure further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The above contents and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0018] Figure 1 A schematic diagram schematically illustrates a related rewriting method;

[0019] Figure 2 Schematically illustrates an application scenario diagram of a response method, apparatus, device, medium, and program product according to an embodiment of the present disclosure;

[0020] Figure 3 The following schematically shows a flow chart of a response method according to an embodiment of the present disclosure;

[0021] Figure 4 The following schematically shows a flow chart of a response method according to another embodiment of the present disclosure;

[0022] Figure 5 A schematic diagram schematically illustrates a response method according to an embodiment of the present disclosure;

[0023] Figure 6 Schematically shows a structural block diagram of a response device according to an embodiment of the present disclosure; and

[0024] Figure 7 A schematic block diagram of an example electronic device that can be used to implement the method of the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0027] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0028] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0029] When a user enters a query (e.g., a "query"), the RAG system needs to retrieve relevant knowledge blocks (e.g., chunks) from the knowledge base to generate a response. However, due to the diversity and complexity of user-entered queries, such as multiple expressions of the same intent or a mix of Chinese and English, it is difficult to accurately retrieve the knowledge that meets the user's needs. To optimize the recall effect, the user-entered query needs to be rewritten.

[0030] Figure 1 A schematic diagram of the related rewriting method is schematically shown.

[0031] like Figure 1 As shown in the figure, taking the query information input by the user as a query question as an example, after the query question is rewritten, it will be converted into a vector form through embedding processing, and the top N knowledge blocks with high similarity to the above vector will be recalled from the knowledge base. Then, through prompt word assembly, multiple prompt elements are integrated together to guide the large model to output the response information. The large model includes, for example, a large language model (LLM). Relevant rewriting methods include related method 1 and related method 2.

[0032] A related approach is to rewrite the query into multiple sub-problems, such as using the Step-Back method. However, for simple queries, rewriting the query into multiple sub-problems may be unnecessary and may result in a large model and reduced query efficiency.

[0033] A related approach is to rewrite the user's query into multiple similar questions, such as by replacing synonyms or adjusting word order. However, this rewriting approach lacks diversity and fails to fully capture the different aspects of the user's query, resulting in overly similar responses.

[0034] Furthermore, the aforementioned rewriting methods all rely on simple rewriting based on text similarity. These methods perform linear transformations within the same semantic space and fail to capture the user's underlying needs. This lack of a deep understanding of user intent and subsequent conversation content forces users to repeatedly interact with the large model based on its output, often repeatedly asking the model more in-depth questions to obtain satisfactory responses. This repeated interaction wastes computing power and negatively impacts the user experience.

[0035] In view of this, an embodiment of the present disclosure provides a response method, including: generating first predicted reply information corresponding to the target input information based on the target input information from the user; performing an iterative prediction operation based on the first predicted reply information to obtain predicted feedback information; and utilizing a question-answering model to generate target reply information corresponding to the target input information based on prompt information including the target input information, the first predicted reply information and the predicted feedback information.

[0036] Figure 2 The application scenario diagram of the response method, apparatus, device, medium and program product according to the embodiments of the present disclosure is schematically shown.

[0037] like Figure 2 As shown, the application scenario 200 according to this embodiment may include a first terminal device 201, a second terminal device 202, a third terminal device 203, a network 204, and a server 205. The network 204 is used as a medium for providing a communication link between the first terminal device 201, the second terminal device 202, the third terminal device 203, and the server 205. The network 204 may include various connection types, such as wired or wireless communication links or optical fiber cables, etc.

[0038] A user may use a first terminal device 201, a second terminal device 202, or a third terminal device 203 to interact with a server 205 via a network 204 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 201, the second terminal device 202, or the third terminal device 203, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).

[0039] The first terminal device 201 , the second terminal device 202 , and the third terminal device 203 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0040] Server 205 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using first terminal device 201, second terminal device 202, and third terminal device 203. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0041] For example, a user can input information through the first terminal device 201, the second terminal device 202, and the third terminal device 203. In response to receiving the target input information from the user, the server 205 can generate first predicted reply information corresponding to the target input information based on the target input information from the user; perform an iterative prediction operation based on the first predicted reply information to obtain predicted feedback information; utilize the question-answering model to generate target reply information corresponding to the target input information based on prompt information including the target input information, the first predicted reply information, and the predicted feedback information, and display the target reply information to the user.

[0042] It should be noted that the response method provided in the embodiment of the present disclosure can generally be executed by the server 205. Accordingly, the response device provided in the embodiment of the present disclosure can generally be set in the server 205. The response method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 205 and can communicate with the first terminal device 201, the second terminal device 202, the third terminal device 203 and / or the server 205. Accordingly, the response device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 205 and can communicate with the first terminal device 201, the second terminal device 202, the third terminal device 203 and / or the server 205.

[0043] It should be understood that Figure 2 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0044] The following will be based on Figure 2 The scene described by Figures 3 to 5 The response method of the embodiment of the present disclosure is described in detail.

[0045] Figure 3The flowchart of the response method according to the embodiment of the present disclosure is schematically shown.

[0046] like Figure 3 As shown, the response of this embodiment includes operations S310 to S330.

[0047] In operation S310 , first predicted answer information corresponding to the target input information is generated based on target input information from a user.

[0048] For example, the target input information may represent the information requested by the user. A prediction may be performed on the target input information to obtain first predicted answer information. For example, the target input information may include question 1 input by the user to the model. A prediction may be performed based on question 1 to obtain predicted answer information 1 corresponding to question 1.

[0049] In operation S320, an iterative prediction operation is performed based on the first predicted answer information to obtain prediction feedback information.

[0050] Due to the diversity and complexity of the input information from the user, for example, the input information may not be expressed accurately, contain typos or omissions, or be a mixture of Chinese and English, etc., the first predicted reply information generated directly based on the target input information from the user may not be accurate enough. Therefore, the final reply information output to the user directly based on the target input information and / or the first predicted reply information cannot meet the user's needs, and the user may need to input the model multiple times to obtain satisfactory reply information.

[0051] Performing an iterative prediction operation based on the first predicted reply information can simulate multiple user input processes through the iterative prediction operation. The iterative prediction operation can include iteratively predicting the user's possible input information, iteratively predicting the possible reply information to be generated, etc. Accordingly, the predicted feedback information can include the predicted possible user input information, the possible reply information to be generated, etc.

[0052] In operation S330, target answer information corresponding to the target input information is generated based on prompt information including the target input information, the first predicted answer information, and the predicted feedback information using the question-answering model.

[0053] The target input information input by the user, the first predicted reply information predicted for the target input information, and the predicted feedback information obtained by the iterative prediction operation can be input into the question-answering model as prompt information, so that the question-answering model can recall relevant knowledge blocks from the knowledge base more comprehensively at one time based on multiple information such as the target input information, the first predicted reply information, and the predicted feedback information, so that more accurate target reply information can be generated based on the relevant knowledge blocks, and then the question-answering model will display the target reply information to the user.

[0054] For example, the user inputs question 1. A prediction is performed based on question 1 to obtain the possible generated answer information 1 corresponding to question 1. An iterative prediction operation is performed based on answer information 1. The iterative prediction operation simulates the process of the user asking the model multiple questions. For example, the user is not satisfied with answer information 1, asks the model again, is also not satisfied with the answer information output by the model for the second time, and asks the model a third time. After obtaining the predicted feedback information through the iterative prediction operation, question 1, answer information 1, and predicted feedback information can be collectively input into the question-answering model as prompt information. The question-answering model can retrieve knowledge blocks with a high degree of similarity to question 1, answer information 1, and predicted feedback information based on the above prompt information, and generate target answer information based on the above knowledge blocks and display it to the user.

[0055] By iteratively predicting the model and simulating the process of multiple user inputs, we can gain a deeper understanding of the user's query intent and subsequent needs, and obtain predictive feedback information that better meets their needs. By using the target input information, the first predicted response information, and the prompt information for the predicted feedback information as prompt information and inputting it into the question-answering model, the question-answering model can more comprehensively and accurately match relevant knowledge blocks based on this prompt information, generating more accurate, relevant, and in-depth responses, thereby improving knowledge recall and accuracy. As a result, users no longer need to repeatedly input into the model; instead, they can obtain responses that better meet their needs with just a single input. This reduces the number of interactions between the user and the model, improves response efficiency, reduces the model's computational workload, and increases user satisfaction, saving users time and effort.

[0056] According to an embodiment of the present disclosure, the first predicted reply information and the predicted feedback information are generated by a prediction model, and the number of parameters of the prediction model is smaller than the number of parameters of the question-answering model.

[0057] Figure 4 The flowchart of the response method according to another embodiment of the present disclosure is schematically shown.

[0058] like Figure 4 As shown, the response method of the embodiment of the present disclosure may include operations S410 to S450.

[0059] In operation S410 , a user may input target input information to a prediction model.

[0060] In operation S420 , in response to receiving the target input information, the prediction model generates first predicted answer information corresponding to the target input information based on the target input information from the user.

[0061] In operation S430 , the prediction model may perform an iterative prediction operation based on the first prediction reply information to obtain prediction feedback information.

[0062] In operation S440, the prediction model inputs prompt information including the target input information, the first predicted answer information, and the predicted feedback information into the question-answering model.

[0063] In operation S450 , in response to receiving the prompt information, the question-answering model may generate target answer information based on the prompt information including the target input information, the first predicted answer information, and the predicted feedback information.

[0064] Depend on Figure 4 It can be seen that users do not need to repeatedly input into the model. Instead, the prediction model simulates the user's multiple input process through iterative prediction operations, and then the question-answering model generates the target reply information displayed to the user based on the first predicted reply information and predicted feedback information. Therefore, users only need to input into the model once to obtain reply information that better meets their needs.

[0065] Furthermore, the number of parameters of the prediction model can be smaller than the number of parameters of the question-answering model. For example, the number of parameters of the prediction model can be 7B, and the number of parameters of the question-answering model can be 70B. The above parameter amounts of the prediction model and question-answering model are only for illustrative purposes and are not limited here.

[0066] Models with large parameters have better complex reasoning and long context understanding capabilities, and can generate higher-quality response information, but the cost is high and the reasoning delay is high. The rewriting process does not require complex reasoning capabilities. Therefore, if the rewriting is directly performed by a model with large parameters, the parameters and computing resources cannot be fully utilized, resulting in resource waste, long reasoning time, slow response speed, and affecting user experience.

[0067] Models with a small number of parameters are less expensive and have faster responses. Since the rewriting process does not require complex reasoning capabilities, the prediction model with a small number of parameters can be used first to perform the rewriting to generate the first predicted reply information and predicted feedback information. This fully utilizes the prediction model's computing resources and can quickly generate the first predicted reply information and predicted feedback information. The question-answering model then generates the target reply information based on the prompt information, fully utilizing the question-answering model's reasoning capabilities to generate higher-quality reply information. Therefore, by combining the prediction model with a small number of parameters and the question-answering model with a large number of parameters, the respective advantages of both models can be fully utilized to quickly generate reply information that better meets user needs. This can be well applied to scenarios such as intelligent customer service where users may input frequently.

[0068] According to an embodiment of the present disclosure, the prediction feedback information includes: prediction input information and second prediction reply information corresponding to the prediction input information.

[0069] Performing an iterative prediction operation based on the first prediction reply information to obtain prediction feedback information includes: performing a first round of prediction operation on the first prediction reply information to obtain first round of prediction input information and first round of second prediction reply information corresponding to the first round of prediction input information; performing a second round of prediction operation on the first round of second prediction reply information to obtain second round of prediction input information and second round of second prediction reply information corresponding to the second round of prediction input information; iteratively performing prediction operations to obtain multiple rounds of prediction input information corresponding to multiple rounds of prediction operations and multiple rounds of second prediction reply information corresponding to the multiple rounds of prediction input information.

[0070] Figure 5 The figure schematically shows a response method according to an embodiment of the present disclosure.

[0071] like Figure 5 As shown, a prediction can be made based on the target input information to obtain first predicted reply information.

[0072] For example, a user inputs the question "How to improve model training speed" to the model, and based on this question, the model generates the first predicted answer information "Appropriately set hyperparameters." The first predicted answer information indicates that the model may output the above predicted answer information based on the user input question.

[0073] A first round of prediction operations can be performed for the first predicted reply information. The first round of prediction operations can simulate the process of the user re-inputting information into the model in response to the first predicted reply information. Specifically, the first round of predicted input information obtained from the first round of prediction operations can simulate the user re-inputting a question into the model in response to the first predicted reply information, and the first round of second predicted reply information obtained from the first round of prediction operations can simulate the reply information generated in response to the question re-input into the model by the user.

[0074] For example, the first prediction round for the first prediction answer, "Set hyperparameters appropriately," yields prediction input 1, "How to set hyperparameters appropriately." This indicates that after receiving the answer, "Set hyperparameters appropriately," the user might ask again, "How to set hyperparameters appropriately." The first prediction round also yields second prediction answer 1, "Optimize relevant parameters," indicating that the model might output this answer to the question, "How to set hyperparameters appropriately."

[0075] A second round of prediction operation may be performed for the first round of second predicted reply information, to indicate a process in which the user again inputs information into the model for the first round of second predicted reply information after obtaining the first round of second predicted reply information.

[0076] For example, a second round of prediction operation is performed for the second prediction reply information 1 "Optimize related parameters", generating prediction input information 2 "Which related parameters need to be optimized", and second prediction reply information 2 "Related parameters include learning rate, number of iterations and other parameters", which is used to simulate that after receiving the above-mentioned first round of second prediction reply information, the user may again ask the model the question "Which related parameters need to be optimized", and the model may output the reply information "Related parameters include learning rate, number of iterations and other parameters".

[0077] By iteratively executing the prediction operation, each round of prediction input information is a further question about the difficult knowledge points in the second prediction reply information of the previous round, so that the user input information can be understood in depth layer by layer, and the user's deep needs can be inferred to obtain reply information that better meets the user's needs.

[0078] For example, the iterative operation can be stopped when a preset maximum number of iterations is reached. Alternatively, the iterative operation can be stopped when a preset total iteration duration is reached. For example, the iterative prediction operation can be stopped after the user inputs information into the model and the iterative prediction operation reaches a preset total iteration duration, and then prediction feedback information is obtained. Alternatively, the iterative operation can be stopped when a predetermined indicator meets a preset indicator condition, for example, when the similarity between the predicted input information and the target input information falls below a preset similarity threshold.

[0079] By setting the stopping condition for stopping the iterative operation, the problem of wasting computing resources and taking a long time due to too many iterative operations can be avoided.

[0080] According to an embodiment of the present disclosure, performing a round of prediction operations includes: generating the K+1th round of prediction input information based on the Kth round second prediction reply information obtained from the Kth round of prediction operations, where K is a positive integer; and generating the K+1th round second prediction reply information based on the K+1th round of prediction input information.

[0081] For example, after obtaining the second predicted answer information for round K, the user may be predicted to submit a question to the model based on the content of the second predicted answer information for round K, thereby obtaining the predicted input information for round K+1. For example, keywords in the second predicted answer information for round K may be identified, and the predicted input information for round K+1 may be generated based on the keywords.

[0082] For example, after generating the K+1th round prediction input information, information related to the K+1th round prediction input information can be retrieved from the knowledge base based on the K+1th round prediction input information, and the K+1th round second prediction reply information can be generated based on the similarity with the K+1th round prediction input information that meets the preset similarity condition, such as information greater than or equal to the preset similarity threshold.

[0083] According to an embodiment of the present disclosure, based on the K+1 round prediction input information, the K+1 round second prediction reply information is generated, including: based on the K+1 round prediction input information, retrieving the first context information from the user's local knowledge base; based on the first context information, generating the K+1 round second prediction reply information.

[0084] The user's local knowledge base may include a collection of knowledge stored in the user's local device, such as a computer, server, or other device. The first context information may include information in the user's local knowledge base whose similarity to the predicted input information in the K+1 round meets a preset similarity condition. For example, it may include the top N knowledge blocks in the user's local knowledge base that are ranked in terms of similarity to the predicted input information in the K+1 round.

[0085] The user's local knowledge base can store information such as the user's historical input information (such as historical search terms, high-frequency search terms), the frequency of use of historical reply information (such as the number of times the historical reply information has been collected and the number of times it has been exported), etc. This information can accurately reflect the user's needs. Therefore, by retrieving the first context information from the user's local knowledge base, a knowledge block that is more closely matched with the user's needs can be obtained, thereby generating a more accurate second predicted reply information.

[0086] In embodiments of the present disclosure, the user's consent or authorization may be obtained before obtaining the user's information. For example, before retrieving the first context information from the user's local knowledge base, a request to obtain the user's information may be issued to the user. If the user consents or authorizes the acquisition of the user's information, the operation of retrieving the first context information from the user's local knowledge base is performed.

[0087] According to an embodiment of the present disclosure, the first context information includes at least one of: predicted question association information and user background information; wherein, the user background information includes at least one of the following: historical query information, user business information, and professional field information.

[0088] Prediction question-related information may include information that is highly similar to the prediction input information for the K+1 round, such as data belonging to the same professional field as the prediction input information for the K+1 round. Historical query information in user background information may include, for example, historical input information, the frequency of use of historical response information, etc. User business information may include relevant information about the business scenario associated with the user. For example, if the user serves in the financial field, the user business information may include business information related to the financial field. Professional field information may include, for example, relevant information about the professional field associated with the user. For example, if the user is engaged in computer-related work in the financial field, the professional field information may include computer-related information.

[0089] Since the first context information includes predicted question-related information and user background information, the first context information is more consistent with the target input information from the user and user needs. The second predicted reply information generated based on the first context information can also be more consistent with the target input information from the user and better meet user needs.

[0090] According to an embodiment of the present disclosure, retrieving the first context information from the user's local knowledge base according to the K+1th round of prediction input information includes: retrieving the first context information from the user's local knowledge base according to the prompt word.

[0091] Prompt words can be input into the prediction model in advance, and the prompt words are used to guide the prediction model to retrieve the first context information from the user's local knowledge base, and can also be used to guide the prediction model to generate the K+1th round of second prediction reply information based on the K+1th round of prediction input information.

[0092] For example, the prompt word may be in the form of "You are a prediction expert. You can predict potential questions that subsequent users may ask based on the first question entered by the user. I will provide you with a plug-in for the user's own local database. You need to generate answers and predict the next question based on the plug-in database. This ensures that the answers to the questions are more accurate, and the predicted questions are more inclined to the user's intentions." The above prompt word form is only an example and the specific form of the prompt word is not limited. It can guide the prediction model to retrieve the first context information from the user's local knowledge base and generate the K+1th round of second predicted answer information based on the K+1th round of prediction input information.

[0093] By setting prompt words, the prediction model can be guided to accurately retrieve the first context information and improve the accuracy of the second predicted reply information.

[0094] According to an embodiment of the present disclosure, utilizing a question-answering model, based on prompt information including target input information, first predicted reply information, and predicted feedback information, generating target reply information corresponding to the target input information includes: utilizing the question-answering model, retrieving second context information related to the prompt information; and outputting the target reply information based on the second context information.

[0095] For example, the question-answering model can retrieve information from a knowledge base that has a high similarity to the prompt information, and use the information in the knowledge base whose similarity is greater than or equal to a preset similarity threshold as the second context information. For example, the question-answering model can retrieve information from a knowledge base that is respectively related to the target input information, the first predicted reply information, and the predicted feedback information, and use the information in the knowledge base that has a similarity greater than or equal to a preset similarity threshold as the second context information. The knowledge base may include a user's local knowledge base, a public knowledge base, etc. After obtaining the second context information, the second context information can be integrated and converted into natural language form to obtain the target reply information.

[0096] By retrieving the second context information related to the target input information, the first predicted reply information and the predicted feedback information, rather than only retrieving the context information related to the information input by the user, the context information can be obtained comprehensively and accurately. The context information is more relevant to the user needs, thereby generating target reply information that is more accurate and can better meet the user needs.

[0097] For example, after obtaining the prompt information, the question-answering model can extract keyword information from the prompt information, or convert the prompt information into a dense vector and a sparse vector respectively. The question-answering model can retrieve context information related to the above-mentioned keyword information, and / or dense vector, and / or sparse vector respectively based on the keyword information, and / or the dense vector corresponding to the prompt information, and / or the sparse vector corresponding to the prompt information, to obtain the second context information. It should be noted that there is no limitation on the method of extracting keyword information from the prompt information and converting the prompt information into a dense vector and a sparse vector, and it can be any method that can extract keywords from the prompt information and convert the prompt information into a dense vector and a sparse vector.

[0098] Dense vectors can reflect the deeper semantic information in the prompt information, sparse vectors can reflect the shallower semantic information in the prompt information, and keyword information can reflect the core semantic information in the prompt information. By searching based on keyword information, dense vectors, sparse vectors and other information, multi-angle and more comprehensive contextual information can be obtained, thereby improving the accuracy of the reply information.

[0099] The question-answering model can integrate the second context information, convert the integrated second context information into natural language form, and obtain the target answer information.

[0100] Alternatively, the question-answering model may input the second context information into the prediction model, which then filters the second context information based on its similarity to the target input information to obtain the target second context information. The prediction model may also input the target second context information into the question-answering model, which then outputs the target response information based on the target second context information. For example, the question-answering model may convert the integrated target second context information into natural language to obtain the target response information.

[0101] Optionally, based on the similarity between the second context information and the target input information, filtering the second context information may include taking information in the second context information with a similarity greater than or equal to a preset threshold as the target second context information.

[0102] Since the context information retrieved based on keyword information, dense vectors, and sparse vectors is large in quantity and rich in content, there may be useless noise information. The prediction model can filter out the second context information that is highly similar to the information input by the user, thereby filtering out the noise information. The question-answering model then outputs the response based only on the second context information with higher similarity, which can save computing resources and improve response efficiency.

[0103] According to an embodiment of the present disclosure, the prediction feedback information includes: a plurality of first prediction reply information corresponding to a plurality of prediction operations.

[0104] For example, one prediction operation may include generating first predicted answer information corresponding to the target input information based on the target input information from the user. By iteratively performing multiple prediction operations, multiple first predicted answer information corresponding to the multiple prediction operations may be generated.

[0105] Specifically, performing a prediction operation includes: when it is determined that the Lth first prediction reply information obtained by the Lth prediction operation does not meet the target condition, generating the L+1th first prediction reply information according to the target input information.

[0106] For example, a prediction operation may include determining whether the first predicted response information generated based on the target input information previously met the target condition. If the target condition is not met, the first predicted response information for this prediction operation may be generated again based on the target input information. If the target condition is met, the iterative prediction operation may be stopped, and the first predicted response information generated by the executed prediction operation may be used as the prediction feedback information.

[0107] Optionally, the target condition includes at least one of the following: the number of predictions is equal to a threshold, and the satisfaction value of the first predicted reply information meets a target range.

[0108] For example, "the number of predictions equal to the threshold" may include the number of executions of the prediction operation reaching a preset prediction threshold. To avoid wasting computing resources and increasing response time due to excessive prediction operations, a prediction threshold may be set. When the number of executions of the prediction operation reaches the threshold, the iterative prediction operation is stopped.

[0109] For example, the satisfaction value can represent whether the first predicted reply information can meet the needs of the user, and the satisfaction value of the first predicted reply information can be predicted by the model. The satisfaction value of the predicted reply information meeting the target range can include the satisfaction value being greater than or equal to a preset satisfaction value range. When the satisfaction value of the first predicted reply information meets the target range, it means that the generated first predicted reply information has largely met the user's needs and there is no need to further perform the prediction operation. Therefore, the iterative prediction operation can be stopped when the satisfaction value of the first predicted reply information meets the target range.

[0110] By judging whether the first predicted reply information meets the target conditions, and continuing to perform the prediction operation if the target conditions are not met to generate the next first predicted reply information, more accurate first predicted reply information can be generated, and multiple prediction operations can be avoided, saving computing resources and improving reply efficiency.

[0111] Based on the above response method, the present disclosure also provides a response device. Figure 6 The device is described in detail.

[0112] Figure 6 The following schematically shows a structural block diagram of a response device according to an embodiment of the present disclosure.

[0113] like Figure 6 As shown, the response device 600 of this embodiment includes a first generation module 610 , an iterative prediction module 620 and a second generation module 630 .

[0114] The first generating module 610 is used to generate first predicted answer information corresponding to the target input information according to the target input information from the user. In one embodiment, the first generating module 610 can be used to perform the operation S310 described above, which will not be repeated here.

[0115] The iterative prediction module 620 is configured to perform an iterative prediction operation based on the first predicted reply information to obtain prediction feedback information. In one embodiment, the iterative prediction module 620 may be configured to perform the operation S320 described above, which will not be described in detail herein.

[0116] The second generation module 630 is configured to utilize the question-answering model to generate target response information corresponding to the target input information based on prompt information including the target input information, the first predicted response information, and the predicted feedback information. In one embodiment, the second generation module 630 may be configured to perform operation S330 described above, which will not be further described herein.

[0117] Optionally, the prediction feedback information includes: prediction input information and second prediction reply information corresponding to the prediction input information. The iterative prediction module includes a prediction operation submodule, which is used to perform a first round of prediction operation on the first prediction reply information to obtain the first round of prediction input information and the first round of second prediction reply information corresponding to the first round of prediction input information; perform a second round of prediction operation on the first round of second prediction reply information to obtain the second round of prediction input information and the second round of second prediction reply information corresponding to the second round of prediction input information; iteratively perform prediction operations to obtain multiple rounds of prediction input information corresponding to multiple rounds of prediction operations and multiple rounds of second prediction reply information corresponding to the multiple rounds of prediction input information.

[0118] Optionally, the prediction operation submodule includes a first generation unit and a second generation unit.

[0119] The first generation unit is used to generate the K+1th round prediction input information based on the Kth round second prediction reply information obtained from the Kth round prediction operation, where K is a positive integer; the second generation unit is used to generate the K+1th round second prediction reply information based on the K+1th round prediction input information.

[0120] Optionally, the second generating unit includes a retrieval subunit and a generating subunit.

[0121] The retrieval subunit is used to retrieve the first context information from the user's local knowledge base according to the K+1th round prediction input information; the generation subunit is used to generate the K+1th round second prediction reply information according to the first context information.

[0122] Optionally, the retrieval subunit includes a retrieval component, and the retrieval component is used to retrieve the first context information from the user's local knowledge base according to the prompt word.

[0123] Optionally, the second generation module includes a retrieval submodule and an output submodule.

[0124] The retrieval submodule is used to retrieve second context information related to the prompt information using the question-answering model; the output submodule is used to output target answer information based on the second context information.

[0125] Optionally, the prediction feedback information includes: a plurality of first prediction reply information corresponding to the plurality of prediction operations; and the iterative prediction module includes a generation submodule. The generation submodule is configured to generate L+1th first prediction reply information based on the target input information if it is determined that the Lth first prediction reply information obtained from the Lth prediction operation does not meet a target condition, wherein the target condition includes at least one of the following: the number of predictions is equal to a threshold, and a satisfaction value of the first prediction reply information meets a target range.

[0126] According to embodiments of the present disclosure, any multiple modules among the first generation module 610, the iterative prediction module 620, and the second generation module 630 can be combined into a single module, or any one of these modules can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present disclosure, at least one of the first generation module 610, the iterative prediction module 620, and the second generation module 630 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the first generation module 610, the iterative prediction module 620, and the second generation module 630 can be at least partially implemented as a computer program module that, when executed, can perform the corresponding functionality.

[0127] The present disclosure also provides an electronic device, a readable storage medium, and a computer program product. Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement methods according to embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein. Device 700 includes a computing unit 701 that performs various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or loaded from a storage unit 707 into a random access memory (RAM) 703. Various programs and data required for the operation of device 700 may also be stored in RAM 703. Computing unit 701, ROM 702, and RAM 703 are connected via a bus 704. An input / output (I / O) interface 705 is connected to bus 704.

[0128] Multiple components in electronic device 700 are connected to I / O interface 705, including input unit 706, such as a keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as a magnetic disk, optical disk, etc.; and communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0129] The computing unit 701 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the application execution method. For example, in some embodiments, the application execution method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the application execution method described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the application execution method via any other suitable means (e.g., via firmware).

[0130] Various implementations of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Implementations can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] The program code for implementing the disclosed method may be written in any combination of one or more programming languages. The program code may be provided to a processor or controller of a general-purpose or special-purpose computer or other programmable data processing device, such that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0134] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), a computing system that includes middleware components (e.g., an application server), a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0135] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. This client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service system that addresses the management difficulties and poor business scalability of traditional physical hosts and VPS services ("Virtual Private Servers"). The server may also be a server in a distributed system or a server integrated with blockchain. The various forms of the process shown above may be used, with steps reordered, added, or deleted. The steps described in this disclosure may be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of this disclosure are achieved, and this is not intended to limit the scope of protection of this disclosure. The above-described specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure are intended to be included within the scope of protection of this disclosure.

Claims

1. A response method, comprising: generating, based on target input information from a user, first predicted answer information corresponding to the target input information; performing an iterative prediction operation based on the first prediction reply information to obtain prediction feedback information; Utilizing the question-answering model, target answer information corresponding to the target input information is generated based on prompt information including the target input information, the first predicted answer information, and the predicted feedback information.

2. The method according to claim 1, wherein The first predicted reply information and the predicted feedback information are generated by a prediction model, and the number of parameters of the prediction model is smaller than the number of parameters of the question-answering model.

3. The method according to claim 1, wherein The prediction feedback information includes: prediction input information and second prediction answer information corresponding to the prediction input information; The performing an iterative prediction operation based on the first prediction reply information to obtain prediction feedback information includes: performing a first round of prediction operation on the first predicted reply information to obtain first round of prediction input information and first round of second predicted reply information corresponding to the first round of prediction input information; performing a second round of prediction operation on the first round of second predicted reply information to obtain second round of prediction input information and second round of second predicted reply information corresponding to the second round of prediction input information; The prediction operation is iteratively performed to obtain multiple rounds of prediction input information corresponding to the multiple rounds of prediction operations and multiple rounds of second prediction answer information corresponding to the multiple rounds of prediction input information.

4. The method according to claim 3, wherein: Performing a round of prediction operations includes: Generate the K+1th round of prediction input information based on the Kth round of second prediction reply information obtained from the Kth round of prediction operation, where K is a positive integer; Generate K+1 round second prediction reply information according to the K+1 round prediction input information.

5. The method according to claim 4, wherein Generating the K+1th round second prediction reply information according to the K+1th round prediction input information includes: Retrieving first context information from the user's local knowledge base according to the K+1th round prediction input information; The K+1th round second predicted answer information is generated according to the first context information.

6. The method according to claim 5, wherein: The first context information includes at least one of: predicted question related information and user background information; The user background information includes at least one of the following: historical query information, user business information, and professional field information.

7. The method according to claim 5, wherein: The retrieving first context information from the user's local knowledge base according to the K+1th round of prediction input information includes: The first context information is retrieved from the local knowledge base of the user according to the prompt word.

8. The method according to claim 1, wherein The generating of target response information corresponding to the target input information by using the question-answering model based on prompt information including the target input information, the first predicted response information, and the predicted feedback information includes: Retrieving second context information related to the prompt information using the question-answering model; The target reply information is output according to the second context information.

9. The method according to claim 1, wherein The prediction feedback information includes: a plurality of first prediction reply information corresponding to a plurality of prediction operations; Performing a prediction operation includes: When it is determined that the Lth first prediction reply information obtained by the Lth prediction operation does not meet the target condition, the L+1th first prediction reply information is generated according to the target input information, wherein the target condition includes at least one of the following: the number of predictions is equal to the threshold, and the satisfaction value of the first prediction reply information meets the target range.

10. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Control method and device, equipment and computer readable storage medium

    CN113110736A

  • Self-inspiration intelligent question answering implementation method and system based on Scogla bottom type question asking

    CN117786091A

  • Dialogue question and answer processing method and device, equipment and medium

    CN118779434A

  • Problem processing method and system, intelligent terminal and computer readable storage medium

    CN119938819A

  • Intelligent question and answer method and device, equipment and medium

    CN120011484A