Method and apparatus for dialog interaction, and device and storage medium

By generating requests and extracting knowledge records for the target entity, the problem of decreased retrieval efficiency and accuracy caused by excessive knowledge base data in traditional dialogue systems is solved, thereby improving the accuracy and efficiency of dialogue system responses.

WO2025236158A1PCT designated stage Publication Date: 2025-11-20BEIJING YOUZHUJU NETWORK TECH CO LTD

Patent Information

Application Number
PCT/CN2024/092950
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-13
Publication Date
2025-11-20

AI Technical Summary

Technical Problem

Traditional dialogue systems experience a decline in retrieval efficiency and accuracy as the amount of knowledge base data increases, and they are unable to proactively acquire external knowledge to improve the accuracy and efficiency of responses.

Method used

By generating requests for the target entity, utilizing the information provided by the target entity and generating knowledge records, and combining the first model and the second model to extract relevant knowledge, the knowledge base is updated.

Benefits of technology

It improves the accuracy and efficiency of the dialogue system's responses, reduces the amount of data in the knowledge base, and enhances the model's reflective ability and domain knowledge acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024092950_20112025_PF_FP_ABST
    Figure CN2024092950_20112025_PF_FP_ABST
Patent Text Reader

Abstract

On the basis of the embodiments of the present disclosure, provided are a method and apparatus for dialog interaction, and a device and a storage medium. The method comprises: on the basis of a first user input indicative of a first task, using a first model to generate a request for a target entity, so as to instruct the target entity to provide information related to the first task; on the basis of a response from the target entity to the request, providing an execution result of the first task as a reply to the first user input; and on the basis of the response and the first user input, generating a knowledge record for storage in a knowledge base, wherein the knowledge base is used by the first model. Therefore, a task indicated by a user input can be executed by means of a target entity. An improvement in the dialog interaction capability of a dialog system is facilitated, thereby improving the accuracy and efficiency of the dialog system in providing a reply.
Need to check novelty before this filing date? Find Prior Art

Description

Method, apparatus, device and storage medium for dialogue interaction TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, apparatus, device and computer readable storage medium for dialogue interaction. BACKGROUND

[0002] With the development of information technology, various terminal devices can provide people with various services in work and life, etc. Applications and / or systems providing services can be deployed in the terminal devices. The terminal devices present corresponding content through the user interfaces of the applications / systems and implement interactions with users to meet various needs of the users. In some cases, a user can initiate a dialogue interaction request in an application / system. The application / system can provide a corresponding reply to the user based on a user input from the user. Therefore, how to improve the accuracy and efficiency of generating the reply is a problem of concern.

[0003] SUMMARY

[0004] In a first aspect of the present disclosure, a dialogue interaction method is provided. The method comprises: based on a first user input indicating a first task, generating, by using a first model, a request for a target entity to provide information related to the first task; providing, based on a response of the target entity to the request, an execution result of the first task as a reply to the first user input; and generating, based on the response and the first user input, a knowledge record for storage into a knowledge base used by the first model.

[0005] In a second aspect of the present disclosure, an apparatus for dialogue interaction is provided. The apparatus comprises: a request generation module configured to, based on a first user input indicating a first task, generate, by using a first model, a request for a target entity to provide information related to the first task; a reply providing module configured to, based on a response of the target entity to the request, provide an execution result of the first task as a reply to the first user input; and a record generation module configured to, based on the response and the first user input, generate a knowledge record for storage into a knowledge base used by the first model.

[0006] In a third aspect of the present disclosure, an electronic device is provided. The electronic device comprises: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.

[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program which, when executed by a processor, causes the processor to implement the method according to the first aspect of the present disclosure.

[0008] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method according to the first aspect of the present disclosure.

[0009] It should be understood that the contents described in this section are not intended to limit the key features or important features of the implementations of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:

[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;

[0012] FIG. 2 shows a schematic diagram of an architecture for dialog interaction according to some embodiments of the present disclosure;

[0013] FIG. 3 shows a schematic diagram of an example for dialog interaction according to some embodiments of the present disclosure;

[0014] FIG. 4 shows a flowchart of a process for dialog interaction according to some embodiments of the present disclosure;

[0015] FIG. 5 shows a schematic structural block diagram of an apparatus for dialog interaction according to some embodiments of the present disclosure; and

[0016] FIG. 6 shows a block diagram of an electronic device in which one or more embodiments of the present disclosure can be implemented. DETAILED DESCRIPTION

[0017] It can be understood that, before using the technical solutions disclosed by the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scenario of use, etc. should be informed to the user and the authorization of the user should be obtained in accordance with relevant laws and regulations.

[0018] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will need to acquire and use personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium performing the operation of the technical solution of the present disclosure according to the prompt information.

[0019] As an optional but non-limiting implementation, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, and the prompt information may, for example, be presented in the pop-up window in the form of text. In addition, the pop-up window may, for example, also carry a selection control for the user to select “agree” or “disagree” to provide personal information to the electronic device.

[0020] It can be understood that the above notification and acquisition of user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation of the present disclosure.

[0021] It can be understood that the data involved in the technical solution (including but not limited to the data itself, acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0022] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, rather, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0023] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or a different section / subsection in any manner.

[0024] In this document, unless explicitly stated otherwise, performing a step “in response to A” does not mean that the step is performed immediately after A, but can include one or more intermediate steps.

[0025] In the description of embodiments of the disclosure, the term "includes" and its synonyms shall be understood as open-ended terms (i.e., "comprising but not limited to"). The term "based on" shall be understood as "based, at least in part, on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The term "some embodiments" shall be understood as "at least some embodiments". Other explicit or implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or same objects. Other explicit and implicit definitions can also be included below.

[0026] As used herein, the term "model" can learn the relationship between the corresponding input and output from the training data, so that the corresponding output can be generated for a given input after the training is completed. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. The neural network model is an example of a model based on deep learning. In this paper, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", which are used interchangeably in this paper.

[0027] A "neural network" is a machine learning network based on deep learning. Neural networks are capable of processing inputs and providing corresponding outputs, which generally include input layers and output layers and one or more hidden layers between the input layers and the output layers. Neural networks used in deep learning applications usually include many hidden layers, thereby increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also known as processing nodes or neurons), each of which processes input from the previous layer.

[0028] Generally, machine learning can roughly include three stages, namely training stage, testing stage and application stage (also known as inference stage). In the training stage, a given model can be trained using a large amount of training data, and the parameter values are updated iteratively until the model can obtain consistent inference from the training data that meets the expected target. Through training, the model can be considered to learn the relationship between input and output (also known as input to output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, so as to determine the performance of the model. The testing stage can sometimes be integrated into the training stage. In the application or inference stage, the trained model can be used to process the actual model input based on the parameter values obtained by training to determine the corresponding model output.

[0029] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. The environment 100 involves a dialog system 110, which can support interactions with a user 145. For example, the dialog system 110 can support the user 145 to ask questions in natural language, and can provide answers to the questions to the user 145. The user 145 can be referred to as an end user of the dialog system 110.

[0030] In some embodiments, the dialog system 110 can include or be implemented as a digital assistant 122. The digital assistant 122 can be configured to have intelligent dialog. As an example, the digital assistant 122 can be configured as a standalone application, such as a web application or other types of applications. As another example, the digital assistant 122 can be configured within an application as a part of the application. In such an example, the digital assistant 122 and the application can be considered as the same application. The digital assistant 122 is provided to assist users in various task handling needs in different applications, scenarios. During the interaction with the digital assistant 122, the user inputs an interaction message, and the digital assistant 122 provides a reply message in response to the user input. Generally, the digital assistant 122 is capable of supporting the user to input questions in natural language, and perform tasks and provide replies based on the understanding of the natural language input and logical reasoning capability. The digital assistant 122 or at least a part thereof can be implemented as an agent. In this document, the digital assistant and the agent are used interchangeably.

[0031] For each user 145, a client of the dialog system 110 can present an interaction window 142 of the digital assistant 122 in a client interface, such as a dialog window with the digital assistant 122. The user 145 can input a message in the dialog window, and the dialog system 110 can determine a reply message of the digital assistant 122 and present to the user 145 in the interaction window 142. In some embodiments, the interaction message of the dialog system 110 can include messages in multi-modal forms, such as text messages (e.g., natural language text), voice messages, image messages, video messages, and the like.

[0032] The dialog system 110 can be deployed locally at each user's 145 terminal device, and / or can be supported by a server device. For example, a user's 145 terminal device can run a client of the dialog system 110, which can support the user's interaction with parts of the service provided by the server. In the case that the dialog system 110 runs locally at the user's terminal device, the user 145 can directly interact with the local dialog system 110 using the terminal device. In the case that the dialog system 110 runs at a server device, the server device can implement the service provision to the client running at the terminal device based on the communication connection between the server and the terminal device. The dialog system 110 can present a corresponding interface to the user 145 based on the user's 145 operation, to output and / or receive relevant information from the user 145.

[0033] In some embodiments, the implementation of at least part of the functions of the dialog system 110, and / or the implementation of at least part of the functions of the digital assistant 122 in the dialog system 110 can be based on models. During the running of the dialog system 110, one or more models, e.g., the model 155, can be invoked. The digital assistant 122 in the dialog system 110 can utilize the model 155 to understand the user input, and provide a reply to the user based on the output of the model 155.

[0034] Although shown as being independent of the dialog system 110, the model(s) 155 can run on the dialog system 110, or other remote servers. In some embodiments, the model 155 can be a machine learning model, a deep learning model, a learning model, a neural network, etc. In some embodiments, the model can be based on a language model (LM). A language model can be capable of question answering by learning from a large amount of corpus. The model 155 can also be based on other suitable models.

[0035] The dialog system 110 can run on a suitable electronic device. The electronic device here can be any type of device with computing capability, including a terminal device or a server device. The terminal device can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination of the aforementioned, including accessories and peripherals of these devices or any combination thereof. The server device can include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, etc. In some embodiments, the dialog system 110 can be implemented based on a cloud service.

[0036] It should be understood that the structure and functionality of the environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the present disclosure.

[0037] Conventionally, a dialog system can only provide dialog services to a user based on the knowledge or data it has already acquired, and it cannot acquire external knowledge actively. The knowledge that the dialog system has already acquired can be, for example, the knowledge stored in a knowledge base. With the increase of the acquired knowledge, the content in the knowledge base will also become more and more. Conventionally, the dialog system tends to add historical information directly to the knowledge base, which can result in a large amount of knowledge base data and affect the retrieval efficiency and accuracy.

[0038] In view of this, according to an embodiment of the present disclosure, an improved solution for dialog interaction is proposed. In this solution, based on a first user input indicating a first task, a request for a target entity is generated by a first model to instruct the target entity to provide information related to the first task. Based on the response of the target entity to the request, an execution result of the first task is provided as a reply to the first user input. Based on the response and the first user input, a knowledge record for storage into a knowledge base is generated, which is used by the first model.

[0039] In this way, the task indicated by the user input can be executed with the help of the target entity. This helps to improve the dialog interaction capability of the dialog system and improve the accuracy and efficiency of the dialog system in providing replies. Further, the response provided by the target entity can be used to summarize knowledge for addition to the knowledge base for subsequent use. This can compensate for the lack of self-reflection capability of the model on the one hand, and on the other hand, it can really enable the model to acquire the domain knowledge hidden behind the answer.

[0040] Some example embodiments of the present disclosure will be described below with continued reference to the drawings.

[0041] Example architecture

[0042] FIG. 2 shows a schematic diagram of an architecture 200 for dialog interaction, according to some embodiments of the present disclosure. The architecture 200 can be implemented at the dialog system 110. For ease of discussion, the architecture 200 will be described with reference to the environment 100 of FIG. 1.

[0043] As shown in FIG. 2, the architecture 200 includes a module 210, which can be used to implement the dialog system 110 or the digital assistant 122, for example. The module 210 is configured to receive a user input 201 from a user. The user input 201 can be any suitable type of user input, including but not limited to text, voice, image, gesture, etc. The module 201 can receive the user input 201 in any suitable manner. For example, the user input 201 can be received via an input box in the interaction window 142.

[0044] The user input 201 can indicate any appropriate task. For example, if the user input 201 includes the text "generate an image", the user input 201 can indicate an image generation task. If the user input 201 includes a question, the user input can indicate a question-answering task. The module 210 can generate, based on the user input 201 (which can also be referred to as a first user input) indicating a first task, a request 220 for the target entity 230 to provide information related to the first task, by using a first model 211.

[0045] The request 220 can indicate, for example, the target entity 230 to complete the entire first task. For example, if the first task is a question-answering task for a question, the request 220 can indicate the target entity 230 to provide an answer to the question. The request 220 can also indicate, for example, the target entity 230 to provide partial information for the first task. For example, if the first task is a question-answering task for a question, the request 220 can indicate the target entity 230 to provide relevant information for the question.

[0046] As to the specific way of generating the request 220, in some embodiments, the module 210 can generate, based on the first user input, a first prompt for the first model 211. The module 210 can provide the first prompt to the first model 211. The first model 211 can generate, based on the first prompt, a model output for the first prompt. The model output can directly include the request 220, for example. The module 210 can obtain the request 220 from the first model 211.

[0047] Alternatively or additionally, in some embodiments, the model output of the first model 211 can include, for example, a first query for a target database. The module 210 can obtain the first query from the first model 211 and provide the first query to the target database. Alternatively or additionally, the first query can also be provided directly from the first model to the target database, for example. The target database can be a general database or a specific database for the first task. The module 210 can obtain a result for the first query from the target database.

[0048] The module 210 can generate the request 220 based on the first query and / or the first user input, in response to not obtaining the result for the first query from the target database. The module 210 can indicate, for example, the first model 211 to generate the request 220. The module 210 can generate the request 220 in any other appropriate way, such as by using an executor or by using another model, etc. The present disclosure does not limit the specific way of generating the request.

[0049] The target entity 230 can include a predetermined user 231 and / or a third model 232 different from the first model. If the first task belongs to a first domain, the predetermined user 231 can be a professional for the first domain, for example. For example, if the first task is a question-answering task for the vehicle maintenance domain, the predetermined user 231 can be a professional for the vehicle maintenance domain. Similarly, the third model 232 can be a model trained for the first domain. For example, the third model 232 can be a specific model for the vehicle maintenance domain. Alternatively or additionally, the third model 232 can also be a model trained for the type of the first task. For example, if the first task is an image generation task, the type of the first task is the type of image generation, and the third model can be an image generation model. The third model 232 can also be a model with a larger number of parameters, for example. For example, the third model 232 can have a larger number of parameters than the first model, the third model can be more complex than the first model, and the third model can perform a more difficult task than the first model.

[0050] In some embodiments, the target entity 230 can provide a corresponding response 240 for the received request 220. The response 240 can include information related to the first task, for example. For example, if the first task is a question-answering task for a question, the response 240 can include an answer to the question and an analysis of the answer. The module 210 can provide an execution result of the first task as a reply to the first user input based on the response 240, for example. For example, if the first task is a question-answering task for a question, the module 210 can provide an answer to the question as a reply to the first user input. If the first task is an image generation task, the module 210 can provide a generated image as a reply to the first user input. In some embodiments, the module 210 can provide the reply to the user via the interaction window 142.

[0051] In some embodiments, the module 210 can also generate a knowledge record 213 for storage into a knowledge base 212 based on the response 240 and the first user input. The knowledge base 212 can also be a memory base for storing historical records. The knowledge base 212 is used by the first model 211. The first model 211 can execute the first user input based on the content stored in the knowledge base 212, for example. The first model 211 can generate the request 220 or generate a query to the target database only if it is determined that the first user input cannot be executed based on the knowledge base 212.

[0052] In some embodiments, the module 210 can generate the knowledge record 213 in any suitable manner. For example, the module 210 can generate the knowledge record 213 based on a predetermined rule or algorithm. For another example, the module 210 can generate the knowledge record 213 by utilizing a second model 214. The second model 214 can be the same model as the first model 211 or a different model from the first model 211. The second model 214 can also be referred to as a reflection model. The module 210 can generate second prompt information for the second model 214 based on the response 240 and the first user input. The second prompt information can indicate that the second model 214 extracts knowledge related to the first task from the response 240. The module 210 can then provide the second prompt information to the second model 214 to obtain the knowledge record 213. In this way, the knowledge record can be extracted based on the content obtained from the target entity 230, and the first model 211 can subsequently perform the task based on its own knowledge base if it encounters the same type of task again. This can enable the dialog system 110 / digital assistant 122 to continuously obtain external information and continuously improve its dialog interaction capability. In addition, the knowledge record is extracted, which avoids the case of storing all content as a knowledge record in the knowledge base, can reduce the amount of data to be stored, and improve the efficiency of the first model in retrieving the knowledge base.

[0053] It should be noted that if the module 210 can obtain the result for the first query from the target knowledge base, the module 210 can directly provide the execution result of the first task as the reply to the first user input. The module 210 does not need to generate the request 220, let alone obtain the response 240 and obtain the knowledge record 213.

[0054] As for the training manner of the first model 211, the first model 211 can be trained at an electronic device where the dialog system 110 runs. The first model 211 can also be trained at an electronic device other than the electronic device where the dialog system 110 runs. In this case, the electronic device 110 trained by the dialog system 110 can obtain the trained first model 211, or perform dialog interaction by means of the trained first model 211 deployed at the other electronic device via a communication connection between the two electronic devices. For ease of description, the electronic device that trains the first model 211 is referred to as a training device hereinafter.

[0055] The training device can obtain a training sample, which includes reference user input indicating a reference task and a reference execution result of the reference task. The training device utilizes the first model 211 to generate a processing result for the reference user input, and determines a feedback score for the training sample based on the reference execution result and the processing result.

[0056] As to the specific manner of determining the feedback score, the training device can assign a first score as the feedback score in response to the processing result including the execution result of the reference task and the included execution result matching the reference execution result. The training device can assign a second score as the feedback score in response to the processing result including the execution result of the reference task and the included execution result not matching the reference execution result. The training device can also assign a third score as the feedback score in response to the processing result indicating that information related to the reference task is requested from the target entity. The first score here is higher than the third score, and the third score is higher than the second score. For example, the first score can be 1, the third score can be 0.8, and the second score can be 0.

[0057] Illustratively, if the reference task is a question and answer task, the training device can determine that the corresponding feedback score is 1 in response to the first model 211 being able to directly determine the answer to the question and the answer being correct. The training device can determine that the corresponding feedback score is 0 in response to the answer determined by the first model 211 being incorrect. The training device can determine that the corresponding feedback score is 0.8 in response to the first model 211 needing to generate a request to determine the answer to the question with the help of the target entity or the target database.

[0058] The training device 110 can obtain the feedback scores of the first model 211 for one or more training samples, and update the model parameters of the first model according to the feedback scores. The update goal is to improve the feedback scores as much as possible. For example, if there are a total of 100 training samples, the update goal is to make the sum of the feedback scores of the 100 training samples as close to 100 as possible.

[0059] In addition to the above training manner, the training device can also use other training manners to train the first model. For example, the training device can also use the manner of traditional supervised training to train the first model. The present disclosure does not limit the specific training manner.

[0060] FIG. 3 shows a schematic diagram of an example 300 for a dialog interaction, according to some embodiments of the present disclosure. The example 300 can be performed at the dialog system 110 or the digital assistant 122. The following is an illustrative example taking the example 300 being performed at the dialog system 110 as an example.

[0061] At block 301, the dialog system 110 receives a user input from a user.

[0062] At block 302, the dialog system 110 starts a new dialog in response to receiving the user input. The dialog system 110 may, for example, start the new dialog in the interaction window 142.

[0063] At block 303, the dialog system 110 can retrieve a knowledge base based on the user input.

[0064] If the dialog system 110 can retrieve the content associated with the task indicated by the user input from the knowledge base, the dialog system 110 performs block 304. At block 304, the dialog system 110 provides a reply to the user input based on the content retrieved from the knowledge base.

[0065] If the dialog system 110 cannot retrieve the content associated with the task indicated by the user input from the knowledge base, the dialog system 110 can determine whether to perform block 305 or block 306. In some embodiments, the dialog system 110 can receive a user configuration in advance, and determine the priority of each of block 305 and block 306 based on the user configuration. If the priority of block 305 is higher, the dialog system 110 can preferentially perform block 305. Similarly, if the priority of block 306 is higher, the dialog system 110 can preferentially perform block 306. It can be appreciated that the dialog system 110 can determine whether to perform block 305 or block 306 based on any appropriate manner, and the present disclosure does not limit the specific determination manner.

[0066] At block 305, the dialog system 110 can generate a request for the target entity, and determine the reply to the user input by means of the target entity. For example, the dialog system 110 can generate the request for the target entity based on the user input, by means of the first model, and provide the reply to the user input based on the response of the target entity to the request.

[0067] At block 306, the dialog system 110 can retrieve the database. Specifically, the dialog system 110 can generate prompt information for the first model based on the user input. The dialog system 110 can provide the prompt information to the first model to obtain a query for the database output by the first model. The dialog system 110 can obtain the result for the query from the database.

[0068] At block 307, the dialog system 110 can provide the reply to the user input based on the result obtained from the database in response to obtaining the result for the query from the database.

[0069] At block 308, the dialog system 110 can generate a request for the target entity based on at least one of the query or the user input in response to failing to obtain the result for the first query from the target database. Similarly to block 305, the dialog system 110 can determine the reply to the user input by means of the target entity.

[0070] Thus, it can be determined whether to perform the task by means of the target entity or the target database based on the retrieval result of the knowledge base. This can enable the dialog system 110 to have the ability to automatically ask for help from the outside world (e.g., from the target entity or the target database), and help to improve the dialog interaction capability of the dialog system.

[0071] In summary, according to embodiments of the present disclosure, a task indicated by user input can be performed with the aid of a target entity. This helps to improve the dialog interaction capability of the dialog system, and improve the accuracy and efficiency of the dialog system in providing a reply. Further, the response provided by the target entity can be used to summarize knowledge for addition to the knowledge base for subsequent use. In addition, the amount of data included in the summarized knowledge is less than the amount of data of the response itself, which helps to reduce the amount of data stored in the knowledge base and improve the efficiency and accuracy of retrieving the knowledge base. This can compensate for the lack of self-reflection capability of the model on the one hand, and on the other hand, it can truly enable the model to obtain the domain knowledge hidden behind the answer.

[0072] Example process

[0073] FIG. 4 shows a flowchart of a process 400 for dialog interaction, according to some embodiments of the present disclosure. The process 400 can be implemented at the dialog system 110. For ease of discussion, the process 400 will be described with reference to the environment 100 of FIG. 1.

[0074] At block 410, the dialog system 110 generates, based on a first user input indicating a first task, a request for a target entity to provide information related to the first task, using a first model.

[0075] At block 420, the dialog system 110 provides, based on a response of the target entity to the request, an execution result of the first task as a reply to the first user input.

[0076] At block 430, the dialog system 110 generates, based on the response and the first user input, a knowledge record for storage into a knowledge base used by the first model.

[0077] In some embodiments, generating the request for the target entity includes generating, based on the first user input, first prompt information for the first model, and providing the first prompt information to the first model to obtain the request output by the first model.

[0078] In some embodiments, generating the request for the target entity includes generating, based on the first user input, first prompt information for the first model, providing the first prompt information to the first model to obtain a first query for the target database output by the first model, and in response to not obtaining a result of the first query from the target database, generating the request based on at least one of the first query or the first user input.

[0079] In some embodiments, generating the knowledge record for storage into the knowledge base comprises: based on the response and the first user input, generating second prompt information for a second model, the second prompt information indicating the second model to extract knowledge related to the first task from the response; and providing the second prompt information to the second model to obtain the knowledge record.

[0080] In some embodiments, the target entity comprises at least one of: a third model different from the first model, or a predetermined user.

[0081] In some embodiments, the third model satisfies at least one of: a parameter amount of the third model is greater than the first model, or the third model is trained for a type of the first task.

[0082] In some embodiments, the process 400 further comprises: based on second user input indicating a second task, generating, by the first model, a second query for the target database; and based on a result for the second query obtained from the target database, providing an execution result of the second task as a reply to the second user input.

[0083] In some embodiments, the first model is trained by: obtaining a training sample, the training sample comprising reference user input indicating a reference task and a reference execution result of the reference task; generating, by the first model, a processing result for the reference user input; determining a feedback score for the training sample based on the reference execution result and the processing result; and updating a model parameter of the first model according to the feedback score.

[0084] In some embodiments, determining the feedback score for the processing result comprises: in response to the processing result comprising an execution result of the reference task and the included execution result matching the reference execution result, assigning a first score as the feedback score; in response to the processing result comprising an execution result of the reference task and the included execution result not matching the reference execution result, assigning a second score as the feedback score; and in response to the processing result indicating requesting the target entity for information related to the reference task, assigning a third score as the feedback score, wherein the first score is higher than the third score, and the third score is higher than the second score.

[0085] Example apparatus and device

[0086] FIG. 5 shows a schematic structural block diagram of an apparatus 500 for dialog interaction, according to some embodiments of the present disclosure. The apparatus 500 can be implemented as or included in the dialog system 110. Various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0087] As shown, the apparatus 500 includes a request generation module 510 configured to generate, based on a first user input indicative of a first task, a request for a target entity using a first model to instruct the target entity to provide information related to the first task. The apparatus 500 further includes a reply providing module 520 configured to provide, based on a response of the target entity to the request, an execution result of the first task as a reply to the first user input. The apparatus 500 further includes a record generation module 530 configured to generate, based on the response and the first user input, a knowledge record for storage into a knowledge base used by the first model.

[0088] In some embodiments, the request generation module 510 includes a first information generation module configured to generate, based on the first user input, first prompt information for the first model, and a first request generation module configured to provide the first prompt information to the first model to obtain the request output by the first model.

[0089] In some embodiments, the request generation module 510 includes a second information generation module configured to generate, based on the first user input, first prompt information for the first model, a first query generation module configured to provide the first prompt information to the first model to obtain a first query for a target database output by the first model, and a second request generation module configured to generate, in response to a result of the first query not being obtained from the target database, the request based on at least one of the first query or the first user input.

[0090] In some embodiments, the record generation module 530 includes a third information generation module configured to generate, based on the response and the first user input, second prompt information for a second model, the second prompt information instructing the second model to extract knowledge related to the first task from the response, and a first record acquisition module configured to provide the second prompt information to the second model to obtain the knowledge record.

[0091] In some embodiments, the target entity includes at least one of a third model different from the first model, or a predetermined user.

[0092] In some embodiments, the third model satisfies at least one of a following condition: a parameter amount of the third model is greater than the first model, or the third model is trained for a type of the first task.

[0093] In some embodiments, the apparatus 500 further includes a second query generation module configured to generate, based on a second user input indicative of a second task, a second query for the target database using the first model, and a first reply providing module configured to provide, based on a result of the second query obtained from the target database, an execution result of the second task as a reply to the second user input.

[0094] In some embodiments, the first model is trained by obtaining training samples, the training samples comprising reference user input indicative of a reference task and a reference execution result of the reference task; generating, by the first model, a processing result for the reference user input; determining, based on the reference execution result and the processing result, a feedback score for the training sample; and updating, according to the feedback score, a model parameter of the first model.

[0095] In some embodiments, determining the feedback score for the processing result comprises: in response to the processing result comprising an execution result of the reference task and the included execution result matching the reference execution result, assigning a first score as the feedback score; in response to the processing result comprising an execution result of the reference task and the included execution result not matching the reference execution result, assigning a second score as the feedback score; and in response to the processing result indicating requesting the target entity for information related to the reference task, assigning a third score as the feedback score, wherein the first score is higher than the third score, and the third score is higher than the second score.

[0096] The units and / or modules included in the apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 500 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SOCs), complex programmable logic devices (CPLDs), etc.

[0097] It should be understood that one or more steps in the above methods can be performed by a suitable electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices may, for example, include the dialog system 110 in FIG. 1.

[0098] FIG. 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 600 shown in FIG. 6 is merely an example and should not be construed to limit the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG. 6 can be used to implement the dialog system 110 of FIG. 1 or the apparatus 500 of FIG. 5.

[0099] As shown in FIG. 6, electronic device 600 is in the form of a general-purpose electronic device. Components of electronic device 600 can include, but are not limited to, one or more processors or processing units 610, memory 620, storage 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. Processing unit(s) 610 can be actual or virtual processors and capable of executing various processing in accordance with programs stored in memory 620. In a multi-processing system, multiple processing units execute computer-executable instructions in parallel to improve the processing power of electronic device 600.

[0100] Electronic device 600 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Memory 620 can be volatile (such as, for example, registers, cache, random access memory (RAM)), non-volatile (such as, for example, read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage 630 can be removable or non-removable and can include machine-readable media, such as, for example, flash drives, disks, or any other media capable of storing information and / or data and accessible by electronic device 600.

[0101] Electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile media slot (such as a "floppy" disk). In such cases, each drive can be connected to the bus (not shown) by one or more data media interfaces. Memory 620 can include a computer program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0102] Communication unit(s) 640 enable communication with other electronic devices via communication media. Additionally, functionality of components of electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.

[0103] The input device 650 can be one or more input devices such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc., one or more devices that enable a user to interact with the electronic device 600, or any devices (e.g., a network card, a modem, etc.) that enable the electronic device 600 to communicate with one or more other electronic devices, as desired, via the communication unit 640. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0104] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0105] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0106] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium. The instructions stored on the computer readable storage medium can be used to program a computer, a programmable data processing apparatus, and / or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable data processing apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0107] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0108] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, an optical disc (e.g., a DVD), a flash drive, a memory stick, or a hard disk drive, having such instances of the software. These instances (or a compressed version thereof) can be presented to a processor of a user device by providing a usable device, such as a floppy disk drive, a CD-ROM drive, an optical drive, a flash drive, a memory stick, or a hard disk drive.

[0109] The implementations of the disclosure have been described above with the intent to be illustrative rather than limiting. Although the implementations of the disclosure have been described with regard to one or more implementations, it will be recognized that a variety of modifications and changes can be made to these implementations without departing from the broader spirit and scope of the implementations as set forth in the preceding disclosure. For example, certain aspects of the implementations can be performed using hardware, software, and / or firmware, or any combination thereof. The preceding description is intended to be illustrative, and not to limit the scope of the present disclosure. Other arrangements, methods, or modifications can be devised without departing from the scope of the present disclosure, the described implementations being illustrative. Numerous specific details are described to provide a thorough understanding of the implementations. However, in certain instances, well-known methods, procedures, components, and circuits have not been described in detail so as not to obscure the implementations. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components of the described implementations can be combined or configured in various ways. It will be appreciated that the features and components

Claims

1. A method for conversational interaction, comprising: generating, based on a first user input indicative of a first task, a request for a target entity to provide information related to the first task, using a first model; providing an execution result of the first task as a reply to the first user input, based on a response of the target entity to the request; and generating, based on the response and the first user input, a knowledge record for storage into a knowledge base used by the first model.

2. The method of claim 1, wherein generating the request for the target entity comprises: generating, based on the first user input, a first prompt information for the first model; and providing the first prompt information to the first model to obtain the request output by the first model.

3. The method of claim 1, wherein generating the request for the target entity comprises: generating, based on the first user input, a first prompt information for the first model; providing the first prompt information to the first model to obtain a first query for a target database output by the first model; and in response to no result of the first query being obtained from the target database, generating the request based on at least one of the first query or the first user input.

4. The method of claim 1, wherein generating the knowledge record for storage into the knowledge base comprises: generating, based on the response and the first user input, a second prompt information for a second model, the second prompt information instructing the second model to extract knowledge related to the first task from the response; and providing the second prompt information to the second model to obtain the knowledge record.

5. The method of claim 1, wherein the target entity comprises at least one of: a third model different from the first model, or a predetermined user.

6. The method of claim 5, wherein the third model satisfies at least one of: a parameter amount of the third model is greater than the first model, or the third model is trained for a type of the first task.

7. The method of claim 1, further comprising: generating, based on a second user input indicative of a second task, a second query for a target database, using the first model; and providing an execution result of the second task as a reply to the second user input, based on a result of the second query obtained from the target database.

8. The method of claim 1, wherein the first model is trained by: obtaining a training sample comprising a reference user input indicative of a reference task and a reference execution result of the reference task; generating, using the first model, a processing result for the reference user input; determining a feedback score for the training sample, based on the reference execution result and the processing result; and updating a model parameter of the first model according to the feedback score. ​ ​ ​ ​ ​ ​ 9. The method of claim 8, wherein determining a feedback score for the processing result comprises: in response to the processing result including an execution result of the reference task and the included execution result matching the reference execution result, assigning a first score as the feedback score; in response to the processing result including an execution result of the reference task and the included execution result not matching the reference execution result, assigning a second score as the feedback score; and in response to the processing result indicating requesting the target entity for information related to the reference task, wherein the first score is higher than the third score, and the third score is higher than the second score.

10. An apparatus for dialog interaction, comprising: a request generation module configured to generate, based on a first user input indicating a first task, a request for a target entity to indicate the target entity to provide information related to the first task using a first model; a reply providing module configured to provide, based on a response of the target entity to the request, an execution result of the first task as a reply to the first user input; and a record generation module configured to generate, based on the response and the first user input, a knowledge record for storage into a knowledge base used by the first model.

11. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1-10.

12. A computer-readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1-10.

13. A computer program product comprising a computer program, wherein the computer program, when executed by a processor, implements the method according to any one of claims 1-10. ​ ​ ​

Citation Information

Patent Citations

  • Knowledge question and answer model training method and device, knowledge question and answer method and device and computer equipment

    CN115062134A

  • Task-based dialogue system and implementation method thereof

    CN116911312A

  • Live broadcast question and answer method and device, electronic equipment and computer readable storage medium

    CN117251577A

  • Processing method for intelligent customer service dialogue and related product

    CN117312521A

  • Interaction method and device, equipment and storage medium

    CN117908736A

Cited By

  • Model training method, dialogue processing method and dialogue system

    CN121960792A