Method, device and equipment for dialogue interaction and storage medium
By selecting the retrieval or generated dialogue mode based on user input and target objects in the interaction between users and digital assistants, the problem of unreasonable response of machine learning models in the prior art is solved, and a more flexible and efficient dialogue interaction is achieved.
Patent Information
- Application Number
- CN202311734948.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-20
AI Technical Summary
In the existing user-based machine learning model interactions with digital assistants, the responses of the machine learning model are not completely reasonable, and it is difficult to flexibly select conversation patterns to adapt to different types of problems.
Replies are generated by responding to user input, determining the target object, and selecting the appropriate mode based on the retrieved and generated dialogue patterns. Specifically, for open and closed questions, the dialogue mode based on generation and retrieval is selected respectively to provide more reasonable and realistic responses.
It realizes that the dialogue system flexibly selects dialogue mode during the interaction process, provides more reasonable and practical replies, and improves the quality and efficiency of user interaction.
Smart Images

Figure CN120179761A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for dialogue interaction. Background Art
[0002] With the development of information technology, various terminal devices can provide various services to people in aspects such as work and life. Applications that provide services can be deployed in the terminal devices. The terminal devices present corresponding content through the user interfaces of the applications and implement interactions with users to meet various needs of the users. Therefore, the demand for the flexibility of implementing the interaction between users and digital assistants based on machine learning models is also increasing. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for dialogue interaction is provided. The method includes: in response to receiving a first user input, determining a first target object to be queried indicated by the first user input; based on the first target object, selecting at least one dialogue mode for the first user input from a retrieval-based dialogue mode and a generation-based dialogue mode; and based on query results of the at least one dialogue mode for the first target object, determining a first reply to the first user input.
[0004] In a second aspect of the present disclosure, a device for dialogue interaction is provided. The device includes: a first target object determination module configured to, in response to receiving a first user input, determine a first target object to be queried indicated by the first user input; a dialogue mode selection module configured to, based on the first target object, select at least one dialogue mode for the first user input from a retrieval-based dialogue mode and a generation-based dialogue mode; and a reply determination module configured to, based on query results of the at least one dialogue mode for the first target object, determine a first reply to the first user input.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the electronic device executes the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.
[0007] It should be understood that the content described in this section is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, like or similar reference numerals denote like or similar elements, where:
[0009] Figure 1 FIG. shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0010] Figure 2 FIG. shows a flowchart of an example process for dialogue interaction according to some embodiments;
[0011] Figure 3 FIG. shows a flowchart of a process for dialogue interaction according to some embodiments of the present disclosure;
[0012] Figure 4 FIG. shows a schematic structural block diagram of a device for dialogue interaction according to some embodiments of the present disclosure; and
[0013] Figure 5 FIG. shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0015] In the description of the embodiments of the present disclosure, the term "including" and its like shall be understood as an open inclusion, that is, "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The term "some embodiments" shall be understood as "at least some embodiments". There may be other explicit and implicit definitions hereinafter.
[0016] In this document, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.
[0017] It is understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of the data) shall comply with the requirements of the corresponding laws, regulations and related provisions.
[0018] It is understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the information involved in the present disclosure shall be informed to the relevant users and the authorization of the relevant users shall be obtained through appropriate means in accordance with the relevant laws and regulations. Among them, the relevant users may include any type of right subject, such as individuals, enterprises, and groups.
[0019] For example, when receiving an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested by the user will require obtaining and using the information of the relevant user, so that the relevant user can autonomously choose whether to provide information to software or hardware such as an electronic device, an application program, a server or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0020] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the relevant user in response to receiving an active request from the relevant user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide information to the electronic device.
[0021] It is understood that the above notification and user authorization process is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure. The enabling of the digital assistant-related functions, the acquired data, the data processing and storage manners, etc. in the embodiments of the present disclosure shall all obtain the prior authorization of the user and other right subjects associated with the user, and shall comply with the agreements and rules between the relevant laws, regulations and right subjects.
[0022] As used herein, the term "model" can learn the corresponding association relationship between the input and the output from the training data, so that after the training is completed, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes the input and provides the corresponding output by using multiple processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably in this article.
[0023] Figure 1FIG. 0 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Environment 100 relates to a dialogue system 110, which can support interactions with a user 145. For example, the dialogue system 110 can support the user 145 in asking questions in natural language and can provide answers to the user 145. The user 145 can be referred to as an end user of the dialogue system 110.
[0024] In some embodiments, the dialogue system 110 can include or be implemented as a digital assistant 122. The digital assistant 122 can be configured to have intelligent conversations. As an example, the digital assistant 122 can be configured as a stand-alone application, such as a web application or other types of applications. As another example, the digital assistant 122 can be configured within an application as part of the application. In such an example, the digital assistant 122 and the application can be regarded as the same application. The digital assistant 122 is provided to assist users in handling various task requirements in different applications and scenarios. During the interaction with the digital assistant 122, the user inputs an interaction message, and the digital assistant 122 provides a reply message in response to the user input. Generally, the digital assistant 122 can support the user in inputting questions in natural language and perform tasks and provide replies based on the understanding of the natural language input and logical reasoning capabilities.
[0025] For each user 145, the client of the dialogue system 110 can present an interaction window 142 of the digital assistant 122 in the client interface, such as a dialogue window with the digital assistant 122. The user 145 can input a message in the dialogue window, and the dialogue system 110 can determine the reply message of the digital assistant 122 and present it to the user 145 in the interaction window 142. In some embodiments, the interaction messages of the dialogue system 110 can include messages in multimodal forms, such as text messages (e.g., natural language text), voice messages, image messages, video messages, and so on.
[0026] The dialogue system 110 can be deployed locally on the terminal device of each user 145, and / or can be supported by a server device. For example, the terminal device of the user 145 can run a client of the dialogue system 110, which can support the interaction of the user with a part provided by the server. When the dialogue system 110 runs locally on the user's terminal device, the user 145 can directly interact with the local dialogue system 110 using the terminal device. When the dialogue system 110 runs on the server device, the server device can, based on the communication connection with the terminal device, realize the service supply to the client running on the terminal device. The dialogue system 110 can present a corresponding interface to the user 145 based on the operation of the user 145 to output and / or receive relevant information from the user 145.
[0027] In some embodiments, the implementation of at least some functions of the dialogue system 110, and / or the implementation of at least some functions of the digital assistant 122 in the dialogue system 110 can be implemented based on a model. During the operation of the dialogue system 110, one or more models can be invoked, such as model 155. In the dialogue system 110, the digital assistant 122 can utilize the model 155 to understand user input and provide a response to the user based on the output of the model 155.
[0028] Although shown as independent of the dialogue system 110, one or more models 155 can run on the dialogue system 110 or other remote servers. In some embodiments, the model 155 can be a machine learning model, a deep learning model, a learning model, a neural network, etc. In some embodiments, the model can be based on a language model (LM). By learning from a large amount of corpus, the language model can possess the ability to answer questions. The model 155 can also be based on other suitable models.
[0029] The dialogue system 110 can run on a suitable electronic device. The electronic device here can be any type of device with computing capabilities, including terminal devices or server devices. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, game devices or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server device can include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and so on. In some embodiments, the dialogue system 110 can be implemented based on cloud services.
[0030] It should be understood that the structure and functions of the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure.
[0031] Currently, with the rapid development of machine learning models, it is becoming more and more universal to implement the interaction between users and digital assistants based on machine learning models. However, although the interaction between users and digital assistants can be achieved based on machine learning models, sometimes the responses of machine learning models are not entirely reasonable and may not conform to the actual situation.
[0032] In other cases, the answer to the question may be closed, that is, it has a fixed answer. For such questions, conventionally, the user's question can be answered by retrieval, so that a fixed answer can be obtained for the user's question. In this case, for different types of questions, it is expected that the dialogue system can flexibly select the dialogue mode, and then provide a more reasonable and more practical answer.
[0033] To this end, embodiments of the present disclosure propose an improved solution for dialogue interaction. In this solution, in response to receiving user input, the target object to be queried indicated by the user input is determined. Based on the target object, at least one dialogue mode for the user input is selected from a retrieval-based dialogue mode and a generation-based dialogue mode. Then, based on the query result of the target object in the selected dialogue mode, a reply to the user input is determined. Thus, in the interaction with the user, the dialogue system flexibly selects the dialogue mode according to the content to be queried, so as to provide a more reasonable and more practical answer.
[0034] Continue to refer to Figure 2 Describe the exemplary embodiments of the present disclosure in detail. Figure 2 FIG. 10 shows a flowchart of an exemplary process 200 for dialogue interaction according to some embodiments. Through process 200, the dialogue system 110 can flexibly reply to the input of the user 145 through the digital assistant 122.
[0035] Some or all of process 200 can be implemented by the dialogue system 110 or can be implemented by other devices independent of the dialogue system 110. For example, it can be implemented by other remote devices with computing capabilities (in a terminal device or a service device). Hereinafter, for the convenience of discussion, the execution of process 200 is described from the perspective of the dialogue system 110, but this is only exemplary.
[0036] In block 210, the dialogue system 110 interacts with the user 145 through the digital assistant 122. For example, the digital assistant 122 and the user 145 can interact via a language user interface (LUI). The user 145 can interact with the digital assistant 122 in natural language (including text or speech). In the interaction, the dialogue system 110 receives the user input 230 from the user 145.
[0037] At block 220, the dialogue system 110 understands the user input 230 of the user, and then determines the target object that the user needs to query. For example, the dialogue system 110 can utilize the machine learning model 155 to understand the user input 230. The target object described herein can be used to indicate what the user 145 needs to query. For example, the user 145 may need to query the vacation policy within the company, may need to query the rules within the organization, may need to query the functions of the digital assistant 122, and so on.
[0038] Continuing with process 200, the dialogue system 110 can select at least one dialogue mode from the retrieval-based dialogue mode 201 and the generation-based dialogue mode 202 according to the target object, for replying to the user input 230 of the user 145. The retrieval-based dialogue mode can be referred to as retrieval dialogue, in which the query result is obtained by retrieving subsequent answers from a data source (e.g., a database). The generation-based dialogue mode can also be referred to as generative dialogue, in which answers are created from scratch. For example, templates, rules, or machine learning models can be utilized to generate answers.
[0039] In some embodiments, the dialogue system 110 can determine whether the query for the target object is an open-ended query according to the type of the target object. If the dialogue system 110 determines, according to the type of the target object, that the query for the target object is not an open-ended query, then the retrieval-based dialogue mode 201 can be selected.
[0040] For example, if the target object to be queried is the self-introduction of the digital assistant 122, the dialogue system 110 can determine that the query for this target object is not an open-ended query. The dialogue system 110 can select the retrieval-based dialogue mode 202 to retrieve the pre-configured self-introduction about the digital assistant 122. For example, the pre-configured self-introduction can be retrieved from the configuration information of the digital assistant 122.
[0041] For another example, if the target object to be queried is the regulations within an organization, the dialogue system 110 determines that the query for this target object is not an open-ended query. The dialogue system 110 can select the retrieval-based dialogue mode 201 to retrieve the regulations in the database about the organization.
[0042] In some embodiments, if the dialogue system 110 determines, according to the type of the target object, that the query for the target object is an open-ended query, then the dialogue system 110 can only select the generation-based dialogue mode 202. Alternatively, in some embodiments, in this case, the dialogue system 110 can also select the retrieval-based dialogue mode 201 and the generation-based dialogue mode 202 to reply to the user input 230 of the user 145. In this way, the sources of the reply can be enriched.
[0043] For example, if the target object to be queried is the brand of an electric vehicle, the dialogue system 110 determines that the query for the target object is an open-ended query. Correspondingly, the dialogue system 110 can select the retrieval-based dialogue mode 201 and the generation-based dialogue mode 202 to generate a query result for the brand of the electric vehicle.
[0044] Proceed to process 200. After selecting the dialogue mode, the dialogue system 110 can obtain the query result of the selected dialogue mode 201 for the target object. If the retrieval-based dialogue mode 201 is selected, the dialogue system 110 can use the retrieval-based dialogue mode 201 to generate a first query result 211 for the target object. For example, the dialogue system 110 can select an appropriate data source according to the target object to retrieve the required data from it. If the generation-based dialogue mode 202 is selected, the dialogue system 110 can use the generation-based dialogue mode 202 to generate a second query result 212 for the target object. For example, the dialogue system 110 can use the model 155 to generate the second query result 212.
[0045] In some embodiments, when using the generation-based dialogue mode 202, the context of the dialogue, that is, the historical interaction record, needs to be considered. Correspondingly, the dialogue system 110 determines the historical user input associated with the user input 230 of the user 145 and the corresponding historical response to the historical user input. For example, the dialogue system 110 can use all the dialogue information during the interaction between the user 145 and the digital assistant 122 as the historical interaction record. The dialogue system 110 generates a second query result 212 for the user input 230 according to the user input 230 of the first user, the historical user input, and the historical response corresponding to the historical user input, and by using the generation-based dialogue mode 202.
[0046] Then, the dialogue system 110 can generate a response 231 to the user input 230 based on the query result for the target object. In some embodiments, the query result generated using the selected dialogue mode can be directly used as the response 231. In some embodiments, the obtained query result can also be further adjusted or combined. Example embodiments are described below.
[0047] In some embodiments, generative dialogue can be used to adjust the query result of retrieval-based dialogue. Exemplarily, if the retrieval-based dialogue mode 201 is selected, the dialogue system 110 can use the retrieval-based dialogue mode 201 to generate a first query result 211 for the target object. Further, the dialogue system 110 can use the generation-based dialogue mode 202 to adjust the first query result 211 to generate the response 231. Such adjustments can include, for example, rewriting, a certain degree of expansion, language conversion, text style conversion, and so on.
[0048] For example, if the target object to be queried is the self-introduction of the digital assistant 122, the dialogue system 110 may select the retrieval-based dialogue mode 201 and generate a first query result 211. For example, the query result is the text "Hello, I am the XX assistant. I can help you with the following things..." expressed in Chinese. In this case, if the dialogue system 110 determines that the user 145 is accustomed to expressing in English, the dialogue system 110 may use the machine learning model 155 to convert the above query result into English. For example, the dialogue system 110 may determine the user's habitual language based on the set system language, the language used by the user during the conversation, etc. The embodiments of the present disclosure do not limit this.
[0049] Again, for example, if the target object to be queried is the self-introduction of the digital assistant 122, the dialogue system 110 may select the retrieval-based dialogue mode 201 and generate a first query result 211. For example, the query result is the text "Hello! I am your XX assistant. Here are the things I can help you with..." expressed in a lively style. If the dialogue system 110 determines that the user 145 is accustomed to using a formal language expression, the dialogue system 110 may use the generation-based dialogue mode 202 to convert the above query result into a formal language expression.
[0050] In some embodiments, in the generation-based dialogue mode 202, the machine learning model 155 may be used. In this case, the dialogue system 110 may generate a first prompt word based on the first query result 211 and the adjustment target. The first prompt word generated by the dialogue system 110 is used to indicate the adjustment of the first query result 211. The adjustment target may be, for example, rewriting to make the language more standard, converting the text style, etc. Again, for example, the adjustment target may be to expand the scope of the first query result 211 to a certain extent. Again, for example, the adjustment target may be language conversion, converting the first query result 211 from one language to another language.
[0051] Furthermore, the dialogue system 110 may provide the first prompt word to the machine learning model 155 and obtain the first output of the machine learning model 155. The dialogue system 110 may determine a first reply 231 based on the first output of the machine learning model 155.
[0052] In some embodiments, both the retrieval-based dialogue mode 201 and the generation-based dialogue mode 202 can be selected. In such an embodiment, the dialogue system 110 can utilize the retrieval-based dialogue mode 201 to generate a first query result 211. The dialogue system 110 can utilize the generation-based dialogue mode 202 to generate a second query result 212. The dialogue system 110 can combine the first query result 211 and the second query result 212 to form a reply 231. The reply 231 is used to reply to the user input 230 of the user 145.
[0053] Continuing with the example of the electric vehicle brand mentioned above. If the target object to be queried is the electric vehicle brand, the dialogue system 110 can generate a first query result 211 through the retrieval-based dialogue mode 201. For example, the first query result 211 can be "The electric vehicle brands include the first electric vehicle, the second electric vehicle, the third electric vehicle...". At the same time, the dialogue system 110 can generate a second query result 212 through the generation-based dialogue mode 202. For example, the second query result 212 can be "The electric vehicle brands include the second electric vehicle, whose advantages are...; the fifth electric vehicle, whose advantages are...". The dialogue system 110 can combine the first query result 211 and the second query result 212 to form a reply 231. In this way, a more comprehensive and all-round reply can be provided.
[0054] In some embodiments, a machine learning model can be used to combine these two query results. Exemplarily, the dialogue system 110 can generate a second prompt word based on the first query result 211, the second query result 212, and the combination objective. The second prompt word generated by the dialogue system 110 can be used to indicate combining the first query result 211 and the second query result 212. The combination objective can be, for example, which of the first query result 211 and the second query result 212 is the main and which is the auxiliary, or which result's writing style is the main, etc. The dialogue system 110 can provide the second prompt word to the machine learning model 155 to obtain the second output of the machine learning model 155. The dialogue system 110 can then determine the first reply 231 according to the second output of the machine learning model 155.
[0055] In some embodiments, the dialogue with the user can be a multi-round dialogue. In this case, after the reply 231 is provided to the user 145, the dialogue system 110 can also receive another user input, also referred to as the second user input. The dialogue system 110 determines the second target object to be queried for the second user input at least based on the previously provided reply and the second user input. The dialogue system 110 can select at least one dialogue mode from the retrieval-based dialogue mode 201 and the generation-based dialogue mode 202 according to the second target object for replying to the second user input.
[0056] Furthermore, the dialogue system 110 can obtain the query results of at least one selected dialogue mode for the second target object. In this way, the dialogue system 110 can reply to the second user input based on the query results for the second target object to generate a second reply.
[0057] For example, in the multi-round dialogue between the user 145 and the digital assistant, for each round of questions from the user 145, the dialogue system 110 will re-select whether to use the generated dialogue mode 202 or the retrieved-based dialogue mode 201, instead of simply using the dialogue mode selected in the previous round. In this way, it can be ensured that in each round of the multi-round dialogue, a flexible and accurate reply can be provided.
[0058] Figure 3 FIG. shows a flowchart of a process 300 for dialogue interaction according to some embodiments of the present disclosure. The process 300 can be implemented at the dialogue system 110. The following refers to Figure 1 Describe the process 300.
[0059] At block 310, in response to receiving a first user input, the dialogue system 110 determines a first target object to be queried indicated by the first user input.
[0060] At block 320, the dialogue system 110 selects at least one dialogue mode for the first user input from a retrieved-based dialogue mode and a generated-based dialogue mode based on the first target object.
[0061] At block 330, the dialogue system 110 determines a first reply to the first user input based on the query results of the at least one dialogue mode for the first target object.
[0062] In some embodiments, selecting at least one dialogue mode for the first user input includes: determining whether a query for the target object is an open-ended query based on the type of the target object; and in response to determining that the query for the target object is not an open-ended query, selecting a retrieved-based dialogue mode.
[0063] In some embodiments, the process 300 further includes: in response to determining that the query for the target object is an open-ended query, selecting a retrieved-based dialogue mode and a generated-based dialogue mode.
[0064] In some embodiments, when a retrieved-based dialogue mode is selected for the first user input, and determining the first reply to the first user input includes: using the retrieved-based dialogue mode to generate a first query result for the first target object; and using the generated-based dialogue mode to adjust the first query result to a first reply.
[0065] In some embodiments, adjusting the first query result to a first response includes: generating a first prompt word indicating adjusting the first query result based on the first query result and an adjustment objective; providing the first prompt word to a machine learning model to obtain a first output of the machine learning model; and determining the first response based on the first output.
[0066] In some embodiments, both a retrieval-based dialogue mode and a generation-based dialogue mode are selected for a first user input, and determining a first response to the first user input includes: generating a first query result for a first target object using the retrieval-based dialogue mode; generating a second query result for the first target object using the generation-based dialogue mode; and combining the first query result and the second query result into the first response.
[0067] In some embodiments, combining the first query result and the second query result into the first response includes: generating a second prompt word indicating combining the first query result and the second query result based on the first query result, the second query result, and a combination objective; providing the second prompt word to a machine learning model to obtain a second output of the machine learning model; and determining the first response based on the second output.
[0068] In some embodiments, the generation-based dialogue mode is selected for the first user input, and the method further includes: determining a historical user input associated with the first user input and a corresponding historical response to the historical user input; and generating a second query result for the first user input using the generation-based dialogue mode based on the first user input, the historical user input, and the corresponding historical response.
[0069] In some embodiments, process 300 further includes: in response to receiving a second user input after the first response is provided, determining a second target object to be queried indicated by the second user input based at least on the first response and the second user input; selecting at least one dialogue mode for the second user input from the retrieval-based dialogue mode and the generation-based dialogue mode based on the second target object; and determining a second response to the second user input based on a query result for the second target object of the at least one dialogue mode for the second user input.
[0070] Figure 4 FIG. 400 shows a schematic structural block diagram of a device 400 for digital assistant creation according to some embodiments of the present disclosure. The device 400 may be implemented in or included in, for example, a dialogue system 110. Each module / component in the device 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0071] As shown in the figure, the device 400 includes a target object determination module 410 configured to determine a first target object to be queried indicated by a first user input in response to receiving the first user input. The device 400 further includes a dialogue mode selection module 420 configured to select at least one dialogue mode for the first user input from a retrieval-based dialogue mode and a generation-based dialogue mode based on the first target object. The device 400 further includes a reply determination module 430 configured to determine a first reply to the first user input based on the query result of the at least one dialogue mode for the first target object.
[0072] In some embodiments, the dialogue mode selection module 420 is further configured to: determine whether the query for the target object is an open-ended query based on the type of the target object; and select the retrieval-based dialogue mode in response to determining that the query for the target object is not an open-ended query.
[0073] In some embodiments, the target object determination module 410 is further configured to select the retrieval-based dialogue mode and the generation-based dialogue mode in response to determining that the query for the target object is an open-ended query.
[0074] In some embodiments, the retrieval-based dialogue mode is selected for the first user input, and the reply determination module 430 is further configured to: generate a first query result for the first target object using the retrieval-based dialogue mode; and adjust the first query result to a first reply using the generation-based dialogue mode.
[0075] In some embodiments, the reply determination module 430 is further configured to: generate a first prompt word indicating adjusting the first query result based on the first query result and an adjustment target; provide the first prompt word to a machine learning model to obtain a first output of the machine learning model; and determine the first reply based on the first output.
[0076] In some embodiments, both the retrieval-based dialogue mode and the generation-based dialogue mode are selected for the first user input, and the reply determination module 430 is further configured to: generate a first query result for the first target object using the retrieval-based dialogue mode; generate a second query result for the first target object using the generation-based dialogue mode; and combine the first query result and the second query result into a first reply.
[0077] In some embodiments, the reply determination module 430 is further configured to generate a second prompt word indicating combining the first query result and the second query result based on the first query result, the second query result, and a combination target; provide the second prompt word to a machine learning model to obtain a second output of the machine learning model; and determine the first reply based on the second output.
[0078] In some embodiments, a generated dialogue pattern is selected for a first user input, and the response determination module 430 is further configured to: determine historical user inputs associated with the first user input and corresponding historical responses to the historical user inputs; and generate a second query result for the first user input based on the first user input, the historical user inputs, and the corresponding historical responses, using the generated dialogue pattern.
[0079] In some embodiments, the target object determination module 410 is further configured to: in response to receiving a second user input after a first response is provided, determine a second target object to be queried indicated by the second user input, based at least on the first response and the second user input.
[0080] In some embodiments, the dialogue pattern selection module 420 is further configured to: based on the second target object, select at least one dialogue pattern for the second user input from a retrieved-based dialogue pattern and a generated-based dialogue pattern.
[0081] In some embodiments, the response determination module 430 is further configured to: determine a second response for the second user input based on query results for the second target object using at least one dialogue pattern for the second user input.
[0082] Figure 5 A block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 5 the illustrated electronic device 500 is merely exemplary and should not impose any limitation on the functions and scope of the embodiments described herein. Figure 5 The illustrated electronic device 500 may include or be implemented as Figure 1 the dialogue system 110, or Figure 4 the device 400.
[0083] As Figure 5 shown, the electronic device 500 is in the form of a general-purpose electronic device. The components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, a memory 520, a storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. The processing unit 510 may be an actual or virtual processor and is capable of performing various processes according to programs stored in the memory 520. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 500.
[0084] An electronic device 500 generally includes multiple computer storage media. Such media can be any accessible media to which the electronic device 500 can gain access, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be removable or non-removable media and can include machine-readable media, such as flash drives, magnetic disks, or any other media that can be capable of storing information and / or data and can be accessed within the electronic device 500.
[0085] The electronic device 500 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 5 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 520 can include a computer program product 525 having one or more program modules that are configured to perform various methods or actions of the various embodiments of the present disclosure.
[0086] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.
[0087] The input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 560 can be one or more output devices, such as a display, speaker, printer, etc. The electronic device 500 can also communicate with one or more external devices (not shown) as needed via the communication unit 540, external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 500, or communicate with any device that enables the electronic device 500 to communicate with one or more other electronic devices (such as a network card, modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).
[0088] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above.
[0089] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0090] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, which causes a computer, a programmable data processing device, and / or other devices to operate in a specific manner, so that the computer-readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0091] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operational steps are performed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0092] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.
[0093] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field of the present technology without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the field of the present technology to understand the various implementation manners disclosed herein.
Claims
1. A dialogue interaction method, comprising: In response to receiving a first user input, determine a first target object to be queried indicated by the first user input; Based on the first target object, select at least one conversation mode for the first user input from a retrieval-based conversation mode and a generation-based conversation mode; And Based on the query result of the at least one conversation mode for the first target object, determine a first reply to the first user input.
2. The method according to claim 1, wherein selecting at least one dialogue mode for the first user input comprises: Based on the type of the target object, determine whether the query for the target object is an open-ended query; And In response to determining that the query for the target object is not an open-ended query, select the retrieval-based conversation mode.
3. The method according to claim 2, further comprising: In response to determining that the query for the target object is an open-ended query, select the retrieval-based conversation mode and the generation-based conversation mode.
4. The method according to claim 1, wherein the retrieval-based dialogue mode is selected for the first user input, and determining a first response to the first user input comprises: Use the retrieval-based conversation mode to generate a first query result for the first target object; And Use the generation-based conversation mode to adjust the first query result to the first reply.
5. The method according to claim 4, wherein adjusting the first query result to the first response comprises: Based on the first query result and an adjustment target, generate a first prompt word indicating adjusting the first query result; Provide the first prompt word to a machine learning model to obtain a first output of the machine learning model; And Based on the first output, determine the first reply.
6. The method according to claim 1, wherein both the retrieval-based dialogue mode and the generation-based dialogue mode are selected for the first user input, and determining a first response to the first user input comprises: Use the retrieval-based conversation mode to generate a first query result for the first target object; Use the generation-based conversation mode to generate a second query result for the first target object; And Combine the first query result and the second query result into the first reply.
7. The method according to claim 6, wherein combining the first query result and the second query result into the first response comprises: Based on the first query result, the second query result, and a combination target, generate a second prompt word indicating combining the first query result and the second query result; Provide the second prompt word to a machine learning model to obtain a second output of the machine learning model; And Based on the second output, determine the first reply.
8. The method according to claim 1, wherein the generation-based dialogue mode is selected for the first user input, and the method further comprises: Determine historical user inputs associated with the first user input and corresponding historical replies to the historical user inputs; And Based on the first user input, the historical user inputs, and the corresponding historical replies, use the generation-based conversation mode to generate a second query result for the first user input.
9. The method according to claim 1, further comprising: In response to receiving a second user input after the first reply is provided, determine a second target object to be queried indicated by the second user input based at least on the first reply and the second user input; Based on the second target object, select at least one conversation mode for the second user input from the retrieval-based conversation mode and the generation-based conversation mode; And Based on the query result of the at least one conversation mode for the second user input for the second target object, determine a second reply to the second user input.
10. A device for dialogue interaction, comprising: Determine a target object module, configured to, in response to receiving a first user input, determine a first target object to be queried indicated by the first user input; A selection dialogue mode module, configured to select, based on the first target object, at least one dialogue mode for the first user input from a retrieval-based dialogue mode and a generation-based dialogue mode; And A determination reply module, configured to determine a first reply to the first user input based on query results of the at least one dialogue mode for the first target object.
11. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to execute the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, on which a computer program is stored, and the computer program can be executed by a processor to implement the method according to any one of claims 1 to 9.