Method, device and equipment for dialogue interaction and storage medium

By using machine learning models in the dialogue system to determine the dialogue state, the problem of difficulty in managing multiple rounds of dialogue in the prior art is solved, and the dialogue system can more accurately understand user input and provide services.

CN120179813APending Publication Date: 2025-06-20BEIJING YOUZHUJU NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311735931.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively manage multiple rounds of conversations, resulting in machine learning models being unable to assist the dialogue system in providing the correct services.

Method used

By utilizing a machine learning model, the conversation status of the conversation is determined based on user input and dialogue context information, at least the target request corresponding to the user input is indicated, and a reply is provided according to the dialogue status.

Benefits of technology

It realizes effective tracking and management of dialogue status, and can understand user input more accurately, thereby providing better services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179813A_ABST
    Figure CN120179813A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a dialogue interaction method and device, equipment and a storage medium. The method comprises the following steps: receiving a first user input in a conversation; based on the first user input and the context information of the dialogue, determining a dialogue state of the dialogue by using a machine learning model, the dialogue state at least indicating a target request corresponding to the first user input; and providing a first reply for the first user input according to the dialogue state. In this way, a machine learning model can be utilized to track a dialogue state to provide better services through a dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for dialogue interaction. Background Art

[0002] With the development of information technology, various terminal devices can provide various services to people in aspects such as work and life. Applications that provide services can be deployed in terminal devices. The terminal device presents corresponding content through the user interface of the application and realizes interaction with the user to meet various needs of the user. In some cases, the terminal device can use, for example, a digital assistant to talk to the user, thereby providing services to the user. Summary of the Invention

[0003] In a first aspect of the present disclosure, a dialogue interaction method is provided. The method includes: receiving a first user input in a dialogue; based on the first user input and the context information of the dialogue, using a machine learning model to determine the dialogue state of the dialogue, where the dialogue state at least indicates a target request corresponding to the first user input; and providing a first reply to the first user input according to the dialogue state.

[0004] In a second aspect of the present disclosure, an apparatus for dialogue interaction is provided. The apparatus includes: a user input receiving module configured to receive a first user input in a dialogue; a dialogue state determining module configured to, based on the first user input and the context information of the dialogue, use a machine learning model to determine the dialogue state of the dialogue, where the dialogue state at least indicates a target request corresponding to the first user input; and a reply providing module configured to provide a first reply to the first user input according to the dialogue state.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When the instructions are executed by the at least one processing unit, the electronic device executes the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the medium, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0007] It should be understood that the content described in this part is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Brief Description of the Drawings

[0008] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:

[0009] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;

[0010] Figure 2 A schematic diagram showing an example process for dialogue interaction according to some embodiments of the present disclosure;

[0011] Figure 3 A schematic diagram showing an example task configuration representing a dialogue state according to some embodiments of the present disclosure;

[0012] Figure 4 A flowchart showing a process for session interaction according to some embodiments of the present disclosure;

[0013] Figure 5 A schematic structural block diagram showing a device for session interaction according to some embodiments of the present disclosure; and

[0014] Figure 6 A block diagram showing an electronic device that can implement one or more embodiments of the present disclosure. Detailed Embodiments

[0015] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0016] In the description of the embodiments of the present disclosure, the term "including" and its like shall be understood as an open inclusion, that is, "including but not limited to". The term "based on" shall be understood as "at least partially based on". The term "one embodiment" or "the embodiment" shall be understood as "at least one embodiment". The term "some embodiments" shall be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter.

[0017] In this document, unless otherwise specified, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.

[0018] It is understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition, use, storage or deletion of data) should comply with the requirements of relevant laws, regulations and related provisions.

[0019] It is understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, scope of use, usage scenarios, etc. of the information involved in the present disclosure should be informed to relevant users and the authorization of relevant users should be obtained through appropriate means in accordance with relevant laws and regulations. Among them, relevant users may include any type of rights holders, such as individuals, enterprises, and groups.

[0020] For example, when responding to an active request from a user, a prompt message is sent to the relevant user to clearly prompt the relevant user that the operation requested by the user will require obtaining and using the information of the relevant user, so that the relevant user can autonomously choose whether to provide information to software or hardware such as an electronic device, application program, server or storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0021] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the relevant user in response to receiving an active request from the relevant user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to choose "agree" or "disagree" to provide information to the electronic device.

[0022] It is understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure. The enabling of the digital assistant-related functions in the embodiments of the present disclosure, the data obtained, the processing and storage methods of the data, etc. should all obtain the prior authorization of the user and other rights holders associated with the user, and should comply with the agreements and rules between relevant laws, regulations and rights holders.

[0023] As used herein, the term "model" can learn the corresponding association relationship between input and output from training data, so that after training, for a given input, a corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple processing units. A neural network model is an example of a model based on deep learning. In this article, "model" can also be referred to as "machine learning model", "learning model", "machine learning network" or "learning network", and these terms can be used interchangeably in this article.

[0024] Figure 1FIG. 0 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. Environment 100 relates to a dialogue system 110, which can support interactions with a user 145. For example, the dialogue system 110 can support the user 145 in asking questions in natural language and can provide answers to the user 145. The user 145 can be referred to as an end user of the dialogue system 110.

[0025] In some embodiments, the dialogue system 110 can include or be implemented as a digital assistant 122. The digital assistant 122 can be configured to have intelligent conversations. As an example, the digital assistant 122 can be configured as a stand-alone application, such as a web application or other types of applications. As another example, the digital assistant 122 can be configured within an application as part of the application. In such an example, the digital assistant 122 and the application can be regarded as the same application. The digital assistant 122 is provided to assist users with various task processing requirements in different applications and scenarios. During the interaction with the digital assistant 122, the user inputs an interaction message, and the digital assistant 122 provides a reply message in response to the user input. Generally, the digital assistant 122 can support the user in inputting questions in a natural language manner and can perform tasks and provide replies based on the understanding of the natural language input and logical reasoning capabilities.

[0026] For each user 145, the client of the dialogue system 110 can present an interaction window 142 of the digital assistant 122 in the client interface, such as a dialogue window with the digital assistant 122. The user 145 can input a message in the dialogue window, and the object system 110 can determine the reply message of the digital assistant 122 and present it to the user 145 in the interaction window 142. In some embodiments, the interaction messages of the dialogue system 110 can include messages in multimodal forms, such as text messages (e.g., natural language text), voice messages, image messages, video messages, and so on.

[0027] The dialogue system 110 can be deployed locally on the terminal device of each user 145, and / or can be supported by a server device. For example, the terminal device of the user 145 can run a client of the dialogue system 110, which can support the user's interaction with a part provided by the server. When the dialogue system 110 runs locally on the user's terminal device, the user 145 can directly use the terminal device to interact with the local dialogue system 110. When the dialogue system 110 runs on the server device, the server device can, based on the communication connection with the terminal device, implement the service supply to the client running on the terminal device. The dialogue system 110 can present a corresponding interface to the user 145 based on the operation of the user 145 to output and / or receive relevant information from the user 145.

[0028] In some embodiments, the implementation of at least some functions of the dialogue system 110, and / or the implementation of at least some functions of the digital assistant 122 in the dialogue system 110 can be implemented based on a model. During the operation of the dialogue system 110, one or more models, such as model 155, can be invoked. In the dialogue system 110, the digital assistant 122 can utilize the model 155 to understand the user input and provide a response to the user based on the output of the model 155.

[0029] Although shown as independent of the dialogue system 110, one or more models 155 can run on the dialogue system 110 or other remote servers. In some embodiments, the model 155 can be a machine learning model, a deep learning model, a learning model, a neural network, etc. In some embodiments, the model can be based on a language model (LM). By learning from a large amount of corpus, the language model can possess the ability to answer questions. The model 155 can also be based on other suitable models.

[0030] The dialogue system 110 can run on a suitable electronic device. The electronic device here can be any type of device with computing capabilities, including terminal devices or server devices. The terminal device can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. The server device can, for example, include a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and so on. In some embodiments, the dialogue system 110 can be implemented based on cloud services.

[0031] It should be understood that the structure and functions of the environment 100 are described only for exemplary purposes and do not imply any limitation on the scope of the present disclosure.

[0032] As mentioned above, in the process of the dialogue system providing services through conversations with users, it may be necessary to utilize machine learning models. In particular, in some cases, it may be necessary to have multiple rounds of conversations with users to obtain the necessary information required to provide services to users. However, machine learning models generally do not have the ability to manage multi-round conversations. This may result in the machine learning model being unable to assist the dialogue system in providing correct services in cases where multiple rounds of conversations are required.

[0033] To this end, embodiments of the present disclosure provide a solution for dialogue interaction. In this solution, a machine learning model is utilized to implement at least a part of dialogue management. For example, for a user input in a dialogue, based on the user input and the context information of the dialogue, a machine learning model is used to determine the dialogue state of the dialogue. The dialogue state at least indicates a target request corresponding to the user input, such as a target task or a target query. According to the dialogue state, a reply to the user input is provided. According to embodiments of the present disclosure, a machine learning model can be used to track the dialogue state. In this way, a machine learning model can be used to manage the information in the dialogue, so as to more accurately understand the user input. Thus, better services can be provided through dialogue.

[0034] Next, continue to refer to Figure 2 Describe the exemplary embodiments of the present disclosure in detail. Figure 2 A flowchart of an exemplary process 200 for dialogue interaction according to some embodiments is shown. Through process 200, the dialogue system 110 can flexibly reply to the input of the user 145 through the digital assistant 122.

[0035] Part or all of process 200 can be implemented by the dialogue system 110 or can be implemented by other devices independent of the dialogue system 110. For example, it can be implemented by other remote devices with computing capabilities (in a terminal device or a service device). In the following, for the convenience of discussion, the execution of process 200 is described from the perspective of the dialogue system 110, but this is only exemplary.

[0036] As Figure 2 shown, in the dialogue with the user 145, the dialogue system 110 receives a user input 201 from the user 145. The dialogue system 110 can utilize the model 155 to determine the current dialogue state of the dialogue based on the user input 201 and the context information of the dialogue. The dialogue state at least indicates a target request corresponding to the user input 201.

[0037] Exemplarily, at block 210, the dialogue system 110 can perform user input understanding by using the model 155. Through the understanding of the user input 201, a target request corresponding to the user input 201 can be determined.

[0038] The target request can be a service that the user 145 expects to obtain through the dialogue. In some embodiments, the target request can include a target task, such as ordering food, buying tickets, etc. In some embodiments, the target request can include a target query, such as querying the weather, querying the company's rules and regulations, etc.

[0039] If the user input 201 is the input from user 145 in the first-round conversation, the target request is determined based on the user input 201. If the user input 201 is not an input in the first-round conversation and the target request has been determined based on the previous conversation, the user input 201 can still be determined to correspond to the previously determined target request according to the context information.

[0040] Continue process 200. At block 220, the dialogue system 110 can perform dialogue state tracking using the model 155 to determine the dialogue state of the current dialogue, which at least indicates the target request. In some embodiments, if the target request is a target query, the dialogue state can indicate that user 145 is requesting a query for a certain object (also referred to as the target object).

[0041] In some embodiments, if the target request is a target task, the dialogue state can include a target task configuration corresponding to the target task, which can include at least one task parameter. For example, the model 155 can be used to obtain the filled target task configuration corresponding to the target task. The model 155 can fill the parameters in the task configuration of the target task based on the user input 201 and the context of the dialogue. For parameters that cannot be filled according to the current ongoing dialogue, they can remain vacant. In this way, the dialogue system 110 can determine the dialogue state based on the filling status of the target task configuration. For example, the currently filled parameters and the missing parameters of the target task configuration can represent the dialogue state.

[0042] To this end, in some embodiments, the model 155 can be pre-trained to have the ability to fill task configurations. In some embodiments, the model 155 can fill the task configuration through prompt words. For example, the dialogue system 110 can generate prompt words based on multiple task configurations corresponding to multiple predetermined tasks, each of which includes at least one task parameter. The dialogue system 110 can provide the prompt words to the model 155. In this way, after the model 155 determines the target task, it can select the target task configuration corresponding to the target task from multiple task configurations and fill the target task configuration according to the user input 201.

[0043] As an example, for various predetermined tasks, corresponding task templates can be set. The model 155 can fill the task template corresponding to the target task according to the dialogue context. Refer to Figure 3 Describe an example. Figure 3Shows the filled task configuration 300 corresponding to the "ordering food" task. For example, the user input 201 can be "Help me reserve XX Restaurant at noon tomorrow". Based on such user input, the target request can be determined as the "ordering food" task. Accordingly, the task configuration corresponding to the "ordering food" task can be filled. This task configuration includes multiple parameters (corresponding to slots), specifically "restaurant name", "time", and "number of people" in this example. Based on the above user input, the two parameters "restaurant name" and "time" can be filled, while the "number of people" parameter is vacant.

[0044] In some embodiments, in addition to task parameters, the task configuration can also include other information. As an example, the task configuration can also indicate the function used to execute the target task. In Figure 3 the example of, for the "ordering food" task, the task configuration indicates to call the "ordering food" function.

[0045] Continue to refer to Figure 2 . After obtaining the dialogue state, at block 230, the dialogue system 110 performs a dialogue decision. Depending on various factors such as the type of the target request, the dialogue state, and the strategy used for the dialogue decision, the dialogue system 110 can make different decisions to generate a reply 202 to the user input 201. The following describes exemplary embodiments in different situations.

[0046] In some embodiments, if the target request includes a target task and the dialogue state indicates a lack of task parameters for the target task, the decision made by the dialogue system 110 can be to reply to the user input 201 to instruct the user 145 to provide the missing task parameters. In this case, through the dialogue decision at block 230, the dialogue system 110 can generate a reply trigger 231 to trigger the reply generation performed at block 240. In this embodiment, the generated reply 202 can instruct the user 145 to provide the missing task parameters. Then, the reply 202 can be provided to the user 145, for example, presented in the dialogue interface. Continue Figure 3 the example of. In this example, the provided reply 202 is, for example, "Excuse me, how many people's seats do you want to reserve?".

[0047] In some embodiments, the dialogue with the user 145 continues, thereby updating the dialogue state. Exemplarily, if after the reply 202 is provided, another user input in the dialogue is received, also referred to as the second user input. The dialogue system 110 can update the dialogue state based on the second user input and using the model 155. Then, according to the updated dialogue state, the dialogue system 110 can provide a second reply to the second user input.

[0048] Continue Figure 3For example, after receiving the response "May I ask how many seats you would like to reserve?", the user can provide further user input, such as "A table for five". Based on this further user input, the model 155 can determine that the target request is the previously determined "ordering food" task. Accordingly, the missing task parameters in the task configuration can be filled according to this further user input. In this example, the parameter "number of people" is filled, as Figure 3 shown.

[0049] Through such multi-round conversations, the dialogue system 110 can obtain the necessary information required to execute the target task in order to execute the target task.

[0050] Continuing to refer to Figure 2 to describe other exemplary embodiments of dialogue decision-making. In some embodiments, if the dialogue state indicates that the execution conditions of the target request are met, the decision made by the dialogue system 110 can be to execute the target request. Accordingly, the dialogue system 110 can provide a response 202 based on the execution result of the target request.

[0051] In some embodiments, if the target request includes a target task, the decision made by the dialogue system 110 can be to call a function 232 (also referred to as the target function) to execute the target task. For example, when making a dialogue decision, the dialogue system 110 can determine the target function for the target task. If a suggested function is indicated in the dialogue state (for example, Figure 3 example), the dialogue system 110 can decide to call this function or can determine to call other functions according to its own strategy. The embodiments of the present disclosure are not limited in this regard.

[0052] Subsequently, the dialogue system 110 can call the function 232 to process the target task. After the target task is processed, the process 200 can proceed to block 240. At block 240, the dialogue system 110 can generate a response 202 based on the processing result of the target task (for example, provided by the function 232). For example, for the "ordering food" task, such a response can include information on whether the food order is successful. In the case of a successful food order, the response can also include specific reservation information, etc.

[0053] In this embodiment, a machine learning model is used to track the dialogue state, thereby completing the task corresponding to the user input.

[0054] In some embodiments, the target request can include a query for a target object, also referred to as a target query. For example, the user 145 may want to query the company's internal articles of association. In some embodiments, the dialogue system 110 can combine multiple dialogue modes to provide a query result for the target query.

[0055] In some embodiments, the dialogue system 110 may select at least one dialogue mode from a retrieval-based dialogue mode and a generation-based dialogue mode based on the target object to be queried, for replying to the user input 201. The retrieval-based dialogue mode may be referred to as retrieval-based dialogue, in which the query result is obtained by retrieving subsequent answers from a data source (e.g., a database). The generation-based dialogue mode may also be referred to as generation-based dialogue, in which the answer is created from scratch. For example, templates, rules, or machine learning models may be utilized to generate the answer.

[0056] In some embodiments, the dialogue system 110 may determine whether the query for the target object is an open-ended query according to the type of the target object. If the dialogue system 110 determines, according to the type of the target object, that the query for the target object is not an open-ended query, it may select the retrieval-based dialogue mode.

[0057] For example, if the target object to be queried is the self-introduction of the digital assistant 122, the dialogue system 110 may determine that the query for the target object is not an open-ended query. The dialogue system 110 may select the retrieval-based dialogue mode to retrieve the pre-configured self-introduction about the digital assistant 122. For example, the pre-configured self-introduction may be retrieved from the configuration information of the digital assistant 122.

[0058] In some embodiments, if the dialogue system 110 determines, according to the type of the target object, that the query for the target object is an open-ended query, the dialogue system 110 may only select the generation-based dialogue mode. Alternatively, in some embodiments, in this case, the dialogue system 110 may also select the retrieval-based dialogue mode and the generation-based dialogue mode to reply to the user input 230 of the user 145. In this way, the sources of the reply can be enriched.

[0059] For example, if the target object to be queried is the brand of an electric vehicle, the dialogue system 110 determines that the query for the target object is an open-ended query. Accordingly, the dialogue system 110 may select the retrieval-based dialogue mode 201 and the generation-based dialogue mode 202 to generate the query result for the brand of the electric vehicle.

[0060] Continue with process 200. The selection of the dialogue mode may be performed, for example, at block 230. After selecting the dialogue mode, the dialogue system 110 may invoke the prompt engine 233 to generate prompt information for the at least one selected dialogue mode. Then process 200 proceeds to block 240. At block 240, the dialogue system 110 may, based on the prompt information, utilize the at least one selected dialogue mode to determine the query result for the target query. Then a reply 202 may be provided according to the query result.

[0061] In some embodiments, if a retrieval-based dialogue mode is selected, the generated prompt information may be a retrieval condition constructed according to the target query, such as Structured Query Language. In such an embodiment, at block 240, a search engine may be invoked to retrieve results for the target query.

[0062] In some embodiments, if a generation-based dialogue mode is selected, the generated prompt information may be a prompt word provided to a machine learning model (e.g., model 155). In such an embodiment, at block 240, the prompt word may be provided to the machine learning model to obtain an output of the machine learning model, and based on the output, a query result for the target query may be determined.

[0063] At block 240, if query results provided by at least one selected dialogue mode are obtained, a response 202 may be generated based on the obtained query results. In some embodiments, the obtained query results may be directly used as the response 202. In some embodiments, the obtained query results may also be adjusted. For example, if a retrieval-based dialogue mode is selected and a first query result is generated, the first query result may be adjusted using a machine learning model, such as paraphrasing, expanding, language conversion, style conversion, etc. The adjusted first query result may be used to generate the response 202. Another example is that if both a retrieval-based dialogue mode and a generation-based dialogue mode are selected, and a first query result and a second query result are generated respectively, the first query result and the second query result may be combined using a machine learning model. The combined query results may be used to generate the response 202.

[0064] In such an embodiment, the dialogue system flexibly selects a dialogue mode according to the content to be queried, so as to provide a more reasonable and more practical response.

[0065] Figure 4 A flowchart of a process 400 for dialogue interaction according to some embodiments of the present disclosure is shown. The process 400 may be implemented at the dialogue system 110. The following is referenced Figure 1 to describe the process 400.

[0066] At block 310, the dialogue system 110 receives a first user input in the dialogue.

[0067] At block 320, the dialogue system 110 determines a dialogue state of the dialogue based on the first user input and context information of the dialogue, using a machine learning model. The dialogue state at least indicates a target request corresponding to the first user input.

[0068] At block 330, the dialogue system 110 provides a first response to the first user input according to the dialogue state.

[0069] In some embodiments, determining the dialogue state of a conversation includes: using a machine learning model based on a first user input and context information of the conversation to determine that a target request includes a target task; using the machine learning model to obtain a populated target task configuration corresponding to the target task, the target task configuration including at least one task parameter; and determining the dialogue state based on the population status of the target task configuration.

[0070] In some embodiments, obtaining a populated task configuration information corresponding to a target task includes: generating a prompt based on a plurality of task configurations respectively corresponding to a plurality of predetermined tasks, each of the plurality of task configurations including at least one task parameter; and providing the prompt to a machine learning model to obtain a populated target task configuration.

[0071] In some embodiments, providing a first response to the first user input includes: generating a dialogue message instructing the user to provide the missing task parameter in response to the target request including a target task and the dialogue state also indicating a lack of task parameters for the target task; and providing the dialogue message as the first response.

[0072] In some embodiments, process 400 further includes: after the first response is provided, receiving a second user input in the conversation; using a machine learning model based on the second user input to update the dialogue state; and providing a second response to the second user input according to the updated dialogue state.

[0073] In some embodiments, providing a first response to the first user input includes: providing the first response based on the execution result of the target request in response to the dialogue state indicating that the execution condition of the target request is satisfied.

[0074] In some embodiments, the target request includes a target query for a target object, and providing the first response based on the execution result of the target request includes: selecting at least one dialogue mode for the first user input from a retrieval-based dialogue mode and a generation-based dialogue mode based on the target object; and providing the first response based on the query result of the target object for the at least one dialogue mode.

[0075] In some embodiments, the query result is determined in the following manner: generating prompt information for at least one dialogue mode; and determining the query result using the at least one dialogue mode based on the prompt information.

[0076] In some embodiments, the target request includes a target task, and providing the first response based on the execution result of the target request includes: determining a target function for the target task; invoking the target function to process the target task; and providing the first response based on the processing result of the target task.

[0077] In some embodiments, the target function is indicated by the dialogue state.

[0078] Figure 5 FIG. 4 shows a schematic structural block diagram of a device 500 for dialogue interaction according to some embodiments of the present disclosure. The device 500 can be implemented in or included in, for example, a dialogue system 110. Each module / component in the device 500 can be implemented by hardware, software, firmware, or any combination thereof.

[0079] As shown in the figure, the device 500 includes a user input receiving module 510 configured to receive a first user input in the dialogue. The device 500 further includes a dialogue state determination module 520 configured to determine the dialogue state of the dialogue based on the first user input and the context information of the dialogue, using a machine learning model, where the dialogue state at least indicates a target request corresponding to the first user input. The device 500 further includes a reply providing module 530 configured to provide a first reply to the first user input according to the dialogue state.

[0080] In some embodiments, the dialogue state determination module 520 is further configured to: determine, based on the first user input and the context information of the dialogue, using a machine learning model, that the target request includes a target task; obtain, using a machine learning model, a filled target task configuration corresponding to the target task, where the target task configuration includes at least one task parameter; and determine the dialogue state based on the filling state of the target task configuration.

[0081] In some embodiments, the dialogue state determination module 520 is further configured to: generate a prompt word based on a plurality of task configurations respectively corresponding to a plurality of predetermined tasks, where each of the plurality of task configurations includes at least one task parameter; and provide the prompt word to the machine learning model to obtain a filled target task configuration.

[0082] In some embodiments, the reply providing module 530 is further configured to: generate a dialogue message indicating that the user provides the missing task parameter in response to the target request including a target task and the dialogue state further indicating a lack of task parameters of the target task; and provide the dialogue message as the first reply.

[0083] In some embodiments, the user input receiving module 510 is further configured to receive a second user input in the dialogue after the first reply is provided; the dialogue state determination module 520 is further configured to update the dialogue state based on the second user input, using a machine learning model; and the reply providing module 530 is further configured to provide a second reply to the second user input according to the updated dialogue state.

[0084] In some embodiments, the response providing module 530 is further configured to: in response to the dialogue state indicating that the execution condition of the target request is satisfied, provide a first response based on the execution result of the target request.

[0085] In some embodiments, the target request includes a target query for a target object, and the response providing module 530 is further configured to: based on the target object, select at least one dialogue mode for the first user input from a retrieval-based dialogue mode and a generation-based dialogue mode; and provide a first response based on the query result of the target object for the at least one dialogue mode.

[0086] In some embodiments, the query result is determined in the following manner: generating prompt information for the at least one dialogue mode; and determining the query result based on the prompt information and using the at least one dialogue mode.

[0087] In some embodiments, the target request includes a target task, and the response providing module 530 is further configured to: determine a target function for the target task; call the target function to process the target task; and provide a first response based on the processing result of the target task.

[0088] In some embodiments, the target function is indicated by the dialogue state.

[0089] Figure 6 The block diagram of an electronic device 600 in which one or more embodiments of the present disclosure can be implemented is shown. It should be understood that Figure 6 the shown electronic device 600 is merely exemplary and should not constitute any limitation to the functions and scope of the embodiments described herein. Figure 6 The shown electronic device 600 may include or be implemented as Figure 1 the dialogue system 110, or Figure 5 the device 500.

[0090] As Figure 6 shown, the electronic device 600 is in the form of a general-purpose electronic device. The components of the electronic device 600 may include but are not limited to one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 may be an actual or virtual processor and be capable of performing various processes according to the programs stored in the memory 620. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 600.

[0091] An electronic device 600 generally includes multiple computer storage media. Such media can be any accessible media that the electronic device 600 can access, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 620 can be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 630 can be removable or non-removable media and can include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be capable of storing information and / or data and can be accessed within the electronic device 600.

[0092] The electronic device 600 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 6 it, a disk drive for reading from or writing to a removable, non-volatile magnetic disk (such as a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk can be provided. In these cases, each drive can be connected to a bus (not shown) by one or more data media interfaces. The memory 620 can include a computer program product 625 having one or more program modules that are configured to perform the various methods or actions of the various embodiments of the present disclosure.

[0093] The communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 600 can be implemented by a single computing cluster or multiple computer machines that can communicate via a communication connection. Thus, the electronic device 600 can operate in a networked environment using a logical connection with one or more other servers, network personal computers (PCs), or another network node.

[0094] The input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. The output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. The electronic device 600 can also communicate with one or more external devices (not shown) as needed via the communication unit 640, external devices such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the electronic device 600, or communicate with any device that enables the electronic device 600 to communicate with one or more other electronic devices (such as a network card, a modem, etc.). Such communication can be performed via an input / output (I / O) interface (not shown).

[0095] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.

[0096] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0097] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is produced that implements the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other devices to work in a specific manner. Thus, the computer-readable medium storing the instructions includes a manufacture that includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0098] The computer-readable program instructions can be loaded onto a computer, other programmable data processing device, or other device, such that a series of operational steps are performed on the computer, other programmable data processing device, or other device to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing device, or other device implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0099] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0100] The implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art in the field to understand the various implementations disclosed herein.

Claims

1. A method for dialogue interaction, comprising: Receive a first user input in the conversation; Based on the first user input and the context information of the conversation, use a machine learning model to determine the conversation state of the conversation, where the conversation state at least indicates a target request corresponding to the first user input; And According to the conversation state, provide a first reply to the first user input.

2. The method according to claim 1, wherein determining the dialogue state of the dialogue comprises: Based on the first user input and the context information of the conversation, use the machine learning model to determine that the target request includes a target task; Use the machine learning model to obtain a filled target task configuration corresponding to the target task, where the target task configuration includes at least one task parameter; And Based on the filling state of the target task configuration, determine the conversation state.

3. The method according to claim 2, wherein obtaining the filled task configuration information corresponding to the target task comprises: Generate a prompt word based on multiple task configurations corresponding to multiple predetermined tasks respectively, where each of the multiple task configurations includes at least one task parameter; And Provide the prompt word to the machine learning model to obtain the filled target task configuration.

4. The method according to claim 1, wherein providing a first reply to the first user input comprises: In response to the target request including a target task and the conversation state also indicating a lack of task parameters for the target task, generate a conversation message indicating that the user provides the missing task parameters; and Provide the conversation message as the first reply.

5. The method according to claim 4, further comprising: After the first reply is provided, receive a second user input in the conversation; Based on the second user input, use the machine learning model to update the conversation state; And According to the updated conversation state, provide a second reply to the second user input.

6. The method according to claim 1, wherein providing a first reply to the first user input comprises: In response to the conversation state indicating that the execution condition of the target request is satisfied, provide the first reply based on the execution result of the target request.

7. The method according to claim 6, wherein the target request comprises a target query for a target object, and providing the first reply based on the execution result of the target request comprises: Based on the target object, select at least one conversation mode for the first user input from a retrieval-based conversation mode and a generation-based conversation mode; And Provide the first reply based on the query result of the at least one conversation mode for the target object.

8. The method according to claim 7, wherein the query result is determined in the following manner: Generating prompt information for the at least one dialogue mode; and Based on the prompt information, using the at least one dialogue mode to determine the query result.

9. The method according to claim 6, wherein the target request comprises a target task, and providing the first reply based on the execution result of the target request comprises: Determine a target function for the target task; Invoke the target function to process the target task; And Provide the first reply based on the processing result of the target task.

10. The method according to claim 9, wherein the target function is indicated by the dialogue state.

11. An apparatus for dialogue interaction, comprising: A user input receiving module, configured to receive a first user input in the conversation; A conversation state determining module, configured to based on the first user input and the context information of the conversation, use a machine learning model to determine the conversation state of the conversation, where the conversation state at least indicates a target request corresponding to the first user input; And A reply providing module, configured to according to the conversation state, provide a first reply to the first user input.

12. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit causing the electronic device to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium, having stored thereon a computer program, which can be executed by a processor to implement the method according to any one of claims 1 to 10.