Method implemented by automated assistant and related storage medium

By allowing users to create customized conversation routines, the automated assistant system solves the problem of handling non-standard commands, improves user experience and resource utilization efficiency, and enhances the system's flexibility and task execution capabilities.

CN112801626BActive Publication Date: 2025-07-11GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110154214.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2017-10-03
Filing Date
2018-10-02
Publication Date
2025-07-11
Estimated Expiration
2039-03-08

AI Technical Summary

Technical Problem

Existing automation assistant systems are difficult to handle user non-standard commands and require users to remember and learn specification commands to perform tasks, resulting in poor user experience and waste of resources.

Method used

Allows users to create customized conversation routines through free-form natural language input, automating assistant learning and storing mappings of commands and tasks, which users can then call to execute tasks and process unknown commands through clarification mechanisms.

Benefits of technology

提高了用户记住和使用自动化助理的能力,减少了计算资源的浪费,支持用户创建和共享定制命令,增强了系统的灵活性和任务执行效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112801626B_ABST
    Figure CN112801626B_ABST
Patent Text Reader

Abstract

The present disclosure relates to methods and related storage media implemented by an automated assistant. The technology involves allowing a user to program an automated assistant with customized routines or "dialogue routines" using voice-based human-machine conversations, and the customized routines or "dialogue routines" can be later invoked to complete tasks. In various embodiments, a first free-form natural language input can be received from the user, which identifies a command to be mapped to a task and slots to be filled with values to perform the task. A dialogue routine can be stored, which includes the mapping between the command and the task, and the dialogue routine accepts values for filling the slots as input. A subsequent free-form natural language input can be received from the user to (i) invoke the dialogue routine based on the mapping, and / or (ii) identify values for filling the slots. Data indicating at least the values can be sent to a remote computing device for performing the task.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Division Explanation

[0002] This application is a divisional application of Chinese Patent Application No. 201880039314.9 with an application date of October 2, 2018. Technical Field

[0003] The present disclosure relates to methods for implementing an automated assistant and related storage media. Background Art

[0004] Humans can engage in human-machine conversations using an interactive software application referred to herein as an "automated assistant" (also known as a "chatbot", "interactive personal assistant", "intelligent personal assistant", "personal voice assistant", "conversational agent", etc.). For example, humans (who may be referred to as "users" when they interact with the automated assistant) can use free-form natural language input that may include spoken words that are converted to text and then processed and / or typed free-form natural language input to provide commands, queries, and / or requests (collectively referred to herein as "queries").

[0005] Typically, an automated assistant is configured to perform various tasks in response to various predetermined canonical commands to which tasks are mapped. These tasks can include things like ordering items (e.g., food, products, services, etc.), playing media (e.g., music, video), modifying a shopping list, performing home control (e.g., controlling a thermostat, controlling one or more lights, etc.), answering questions, booking tickets, and the like. While natural language analysis and semantic processing enable users to issue slight variations of canonical commands, these variations can only deviate to the extent that natural language analysis and semantic processing can determine which task to perform. In short, despite many advances in natural language and semantic analysis, task-oriented dialogue management remains relatively rigid. Additionally, users often do not know or forget canonical commands and may therefore be unable to invoke the automated assistant to perform many tasks that it can do. Furthermore, adding new tasks requires third-party developers to add new canonical commands, and the automated assistant typically needs to spend time and resources learning acceptable variations of those canonical commands. Summary of the Invention

[0006] Techniques are described herein for allowing a user to program an automated assistant with custom routines or "dialogue routines" using voice-based human-machine conversations, which can later be invoked to perform tasks. In some embodiments, the user can cause the automated assistant to learn new dialogue routines by providing free-form natural language input including commands for performing tasks. If the automated assistant cannot interpret the command, the automated assistant can solicit clarification from the user regarding the command. For example, in some embodiments, the automated assistant can prompt the user to identify one or more slots that are to be filled with values to fulfill the task. In other embodiments, the user can proactively identify the slots without being prompted by the automated assistant. In some embodiments, the user can provide, for example, in response to an automated assistant request or proactively, an enumerated list of possible values for filling one or more of these slots. The automated assistant can then store the dialogue routine, which includes a mapping between the command and the task, and the dialogue routine accepts one or more values for filling one or more slots as input. The user can later invoke the dialogue routine using free-form natural language input including the command or some syntactic / semantic variation thereof.

[0007] Once the dialogue routine is invoked and the slots of the dialogue routine are filled with values by the user, the automated assistant can take various actions. In some embodiments, the automated assistant can send data indicating at least the slots provided by the user, the slots themselves, and / or data indicating the command / task to a remote computing system. In some cases, such transmission can cause the remote computing system to output, for example, a natural language output or other data indicating the value / slot / command / task to another person. Such natural language output can be provided to another person in various ways (which may not require a third-party software agent installed or configured by another person to process the request) (e.g., via email, text message, automated phone call, etc.). The other person can then perform the task.

[0008] Additionally or alternatively, in some embodiments, various aspects of the dialogue routine (such as slots, potential slot values, commands, etc.) can be compared to similar components of multiple known candidate tasks (e.g., for which the user does not know the canonical command). A mapping can be generated between the best-matching candidate task and the user's command such that future use of the user's command (or its syntactic and / or semantic variations) for the automated assistant will invoke the dialogue routine and, ultimately, the best-matching candidate task. If multiple candidate tasks match the dialogue routine equally, the user can be prompted to select a task, or other signals such as the user's context, prior application usage, etc. can be used to break the tie.

[0009] Suppose the user engages an automated assistant in the following dialogue:

[0010] User: "I want a pizza (I want a pizza)"

[0011] AA: "I don't know how to order a pizza (I don't know how to order a pizza)"

[0012] User: "To order a pizza, you need to know the type of crust and Toppings (To order a pizza, you need to know the type of crust and the list of toppings)"

[0013] AA: "What are the possible pizza crust types? (What are the possible pizza crust types?)"

[0014] User: "Thin crust or thick crust (Thin crust or thick crust)"

[0015] AA: "What are the possible toppings? (What are the possible toppings?)"

[0016] User: "Here are the possible values (Here are the possible values)"

[0017] AA: "Okay, ready to order a pizza (Okay, ready to order a pizza?)"

[0018] User: "Yes, get me a thin crust pizza with a tomato topping (Yes, get me a thin crust pizza with a tomato topping)"

[0019] The command in this scenario is "I want a pizza (I want a pizza)", and the task is to order a pizza. The user-defined slots to be filled to fulfill the task include the type of crust and the list of toppings.

[0020] In some embodiments, the task of ordering a pizza can be accomplished by providing natural language output to a pizza store (which the user can specify or which can be automatically selected, for example, based on distance, rating, price, known user preferences, etc.) via, for example, email, text message, automated phone call, etc. An employee of the pizza store can receive the natural language output via the output of one or more computing devices (e.g., a computer terminal in the store, the employee's phone, a speaker in the store, etc.), and the natural language output can say something like " <user>language such as "the <user> would like to order a <crust_style> pizza with <topping 1, topping 2, ...>".

[0021] In some embodiments, it may be required that a pizza store employee confirm the user's request, for example, by pressing "1" or by saying "OK", "I accept", etc. Once the confirmation is received, in some embodiments, the automated assistant for the requesting user may or may not provide a confirmatory output, such as "your pizza is on the way". In some embodiments, the natural language output provided at the pizza store may also convey other information, such as payment information, the user's address, etc. Such other information may be obtained from the requesting user while creating the dialogue routine or may be automatically determined, for example, based on the user's profile.

[0022] In other embodiments where the command is mapped to a predetermined third-party software agent (e.g., a third-party software agent for a specific pizza store), the task of ordering a pizza may be automatically completed via the third-party software agent. For example, information indicating slot / value may be provided to the third-party software agent in various forms. Assuming that all required slots are filled with appropriate values, the third-party software agent may perform the task of placing a pizza order for the user. If for some reason the third-party software agent requires additional information (e.g., additional slot values), it may interface with the automated assistant to prompt the user for the requested additional information.

[0023] The techniques described herein can yield a variety of technical advantages. As noted above, task-based dialogue management is currently primarily handled by manually creating and mapping to a predefined task's canonical commands. This is limited in its scalability because it requires third-party developers to create these mappings and notify users of them. Similarly, it requires users to learn the canonical commands and remember them for later use. For these reasons, users with limited ability to provide input to complete a task (such as users with physical disabilities and / or users engaged in other tasks (e.g., driving)) may have trouble getting an automated assistant to perform the task. Additionally, when a user attempts to invoke a task with an unexplained command, additional computing resources are needed to disambiguate the user's request or otherwise seek clarification. By allowing users to create their own dialogue routines that are invoked using custom commands, users are more likely to remember the commands and / or be able to successfully and / or more quickly complete the task via the automated assistant. This can conserve computing resources that might otherwise be necessary for the aforementioned disambiguation / clarification. Additionally, in some embodiments, the user-created dialogue routines can be shared with other users, enabling the automated assistant to be more responsive to "long tail" commands from individual users that might be used by others.

[0024] In some embodiments, a method performed by one or more processors is provided, the method comprising: receiving, at one or more input components of a computing device, a first free-form natural language from a user, wherein the first free-form natural language input includes a command for performing a task; performing semantic processing on the free-form natural language input; determining, based on the semantic processing, that the automated assistant cannot interpret the command; providing, at one or more output components of the computing device, an output soliciting clarification from the user regarding the command; receiving, at one or more of the input components, a second free-form natural language input from the user, wherein the second free-form natural language input identifies one or more slots that are required to be filled with values in order to perform the task; storing a dialogue routine that includes a mapping between the command and the task, and the dialogue routine accepts one or more values for filling the one or more slots as input; receiving, at one or more of the input components, a third free-form natural language input from the user, wherein the third free-form natural language input invokes the dialogue routine based on the mapping; identifying, based on the third free-form natural language input or additional free-form natural language input, one or more values to be used for filling the one or more slots that are required to be filled with values in order to perform the task; and sending data indicating at least one or more of the values to be used for filling the one or more slots to a remote computing device, wherein the sending causes the remote computing device to perform the task.

[0025] These and other embodiments of the technology disclosed herein may optionally include one or more of the following features.

[0026] In various embodiments, the method may further include: comparing the dialogue routine with a plurality of candidate tasks executable by the automated assistant; and based on the comparison, selecting, from the plurality of candidate tasks, the task to which the command is mapped. In various embodiments, the task to which the command is mapped includes a third-party agent task, wherein the sending causes the remote computing device to perform the third-party agent task using the one or more values for populating the one or more slots. In various embodiments, the comparison may include comparing the one or more slots required to be populated to fulfill the task with the one or more slots associated with each of the plurality of candidate tasks.

[0027] In various embodiments, the method may further include: receiving, at one or more of the input components, a fourth free-form natural language input from the user prior to the storage. In various embodiments, the fourth free-form natural language input may include a user-provided enumerated list of possible values for populating one or more of the slots. In various embodiments, the comparison may include, for each of the plurality of candidate tasks, comparing the user-provided enumerated list of possible values with an enumerated list of possible values for populating the one or more slots of the candidate task.

[0028] In various embodiments, the data indicating at least the one or more values may further include one or both of an indication of the command or an indication of the task to which the command is mapped. In various embodiments, the data indicating at least the one or more values may take the form of a natural language output requesting execution of the task based on the one or more values, and the sending causes the remote computing device to provide the natural language as an output.

[0029] In another closely related aspect, a method may include: receiving a first free-form natural language input from a user at one or more input components, wherein the first free-form natural language input identifies a command that the user intends to be mapped to a task, and one or more slots that are required to be filled with values ​​in order to perform the task; storing a dialog routine, the dialog routine including a mapping between the command and the task, and the dialog routine accepting as input one or more values ​​for filling the one or more slots; receiving a second free-form natural language input from the user at one or more of the input components, wherein the second free-form natural language input invokes the dialog routine based on the mapping; identifying one or more values ​​to be used to fill the one or more slots that are required to be filled with values ​​in order to perform the task based on the second free-form natural language input or additional free-form natural language input; and sending data indicating at least one or more values ​​to be used to fill the one or more slots to a remote computing device, wherein the sending causes the remote computing device to perform the task.

[0030] In addition, some embodiments include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause the performance of any of the foregoing methods. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to perform any of the foregoing methods.

[0031] It should be appreciated that all combinations of the above-described concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is a block diagram of an example environment in which implementations disclosed herein may be implemented.

[0033] Figure 2 One example of how data generated during an invocation of a dialog routine may flow between various components is schematically depicted in accordance with various embodiments.

[0034] Figure 3 One example of how data may be exchanged between various components when a dialog routine is invoked in accordance with various embodiments is schematically demonstrated.

[0035] Figure 4 Depicted is a flow chart illustrating an example method according to implementations disclosed herein.

[0036] Figure 5 Illustrative example architecture of a computing device. Detailed implementation

[0037] Now turning to Figure 1 , an example environment in which the techniques disclosed herein may be implemented is illustrated. The example environment includes a plurality of client computing devices 106 1-N . Each client device 106 may execute a respective instance of an automated assistant client 118. One or more cloud-based automated assistant components 119, such as a natural language processor 122, may be implemented on one or more computing systems (collectively referred to as "cloud" computing systems), and the one or more computing systems are communicatively coupled to the client devices 106 via one or more local area networks and / or wide area networks generally indicated at 110 (e.g., the Internet). 1-N .

[0038] In some embodiments, an instance of the automated assistant client 118, through its interaction with one or more cloud-based automated assistant components 119, may form what appears to the user to be a logical instance of an automated assistant 120 with which the user may engage in a human-machine conversation. Two instances of such an automated assistant 120 are depicted in Figure 1 . The first automated assistant 120A enclosed by the dashed line serves a first user (not depicted) operating the first client device 1061 and includes the automated assistant client 1181 and one or more cloud-based automated assistant components 119. The second automated assistant 120B enclosed by the double-dashed line serves a second user (not depicted) operating another client device 106 N and includes the automated assistant client 118 N and one or more cloud-based automated assistant components 119. Thus, it should be understood that in some embodiments, each user engaging with the automated assistant client 118 executing on the client device 106 may actually engage with a logical instance of his or her own automated assistant 120. For the sake of brevity and simplicity, the term "automated assistant" as used herein, such as "serving" a particular user, will refer to the combination of the automated assistant client 118 executing on the client device 106 operated by the user and one or more cloud-based automated assistant components 119 (which may be shared among multiple automated assistant clients 118). It should also be understood that in some embodiments, the automated assistant 120 may respond to requests from any user, regardless of whether that user is actually "served" by that particular instance of the automated assistant 120.

[0039] Client device 106 1-N may include, for example, one or more of the following: a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of a user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a stand-alone interactive speaker, a smart home appliance such as a smart TV, and / or a wearable device of the user that includes a computing device (e.g., a user's watch having a computing device, the user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices may be provided.

[0040] In various embodiments, each of the client computing devices 106 1-N may operate a variety of different applications, such as a corresponding one of the plurality of messaging clients 107 1-N among them. The messaging client 107 1-N may take various forms and these forms may vary across the client computing devices 106 1-N and / or may operate multiple forms on a single client computing device among the client computing devices 106 1-N In some embodiments, one or more of the messaging clients 107 1-N may take the form of one or more of the following: a Short Message Service ("SMS") and / or Multimedia Message Service ("MMS") client, an online chat client (e.g., an instant messenger, Internet Relay Chat or "IRC", etc.), a messaging application associated with a social network, a personal assistant messaging service dedicated to conversations with the automated assistant 120, etc. In some embodiments, one or more of the messaging clients 107 1-N may be implemented via a web page or other resource presented by a web browser (not depicted) or other application of the client computing device 106.

[0041] As described in more detail herein, the automated assistant 120 participates in a human-machine conversation session with one or more users via the user interface input and output devices of one or more client devices 106 1-N In some embodiments, the automated assistant 120 may participate in a human-machine conversation session with a user in response to a user interface input provided by the user via one or more user interface input devices of one of the client devices 106 1-N In some of those embodiments, the user interface input is explicitly directed to the automated assistant 120. For example, the messaging client 107 1-N One of them can be a personal assistant messaging service dedicated to conversations with the automated assistant 120, and user interface inputs provided via the personal assistant messaging service can be automatically provided to the automated assistant 120. Additionally, for example, user interface inputs can be explicitly directed to the messaging client 107 based on specific user interface inputs indicating that the automated assistant 120 is to be invoked. 1-N The automated assistant 120 among one or more of them. For example, specific user interface inputs can be one or more typed characters (e.g., @AutomatedAssistant), user interactions with hardware buttons and / or virtual buttons (e.g., tap, long tap), verbal commands (e.g., "Hey Automated Assistant"), and / or other specific user interface inputs.

[0042] In some embodiments, even when user interface inputs are not explicitly directed to the automated assistant 120, the automated assistant 120 can participate in a conversation session in response to the user interface inputs. For example, the automated assistant 120 can examine the content of the user interface inputs and participate in the conversation session in response to the presence of certain items in the user interface inputs and / or based on other cues. In many embodiments, the automated assistant 120 can perform interactive voice response ("IVR") such that the user can speak commands, searches, etc., and the automated assistant can use natural language processing and / or one or more grammars to convert the utterance into text and respond accordingly to the text. In some embodiments, the automated assistant 120 can additionally or alternatively respond to the utterance without converting the utterance into text. For example, the automated assistant 120 can convert the voice input into an embedding, into an entity representation (which indicates one or more entities present in the voice input), and / or other "non - text" representations and operate on such non - text representations. Thus, embodiments described herein as operating based on text converted from voice input can additionally and / or alternatively operate directly on the voice input and / or other non - text representations of the voice input.

[0043] Client computing device 106 1-N Each of the client computing device 106 and the computing device operating the cloud - based automated assistant component 119 can include one or more memories for storing data and software applications, one or more processors for accessing the data and executing the applications, and other components facilitating communication over the network. Operations performed by one or more of the client computing device 106 1-N and / or by the automated assistant 120 can be distributed across multiple computer systems. The automated assistant 120 can be implemented as a computer program running on one or more computers coupled to each other via a network in one or more locations.

[0044] As noted above, in various embodiments, each of the client computing devices 106 1-N may operate an automated assistant client 118. In various embodiments, each automated assistant client 118 may include a corresponding voice capture / text-to-speech ("TTS") / STT module 114. In other embodiments, one or more aspects of the voice capture / TTS / STT module 114 may be implemented separately from the automated assistant client 118.

[0045] Each voice capture / TTS / STT module 114 may be configured to perform one or more functions: for example, capture a user's voice via a microphone (which may include a presence sensor 105 in some cases); convert the captured audio to text (and / or to other representations or embeddings); and / or convert text to voice. For example, in some embodiments, because the client device 106 may be relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the voice capture / TTS / STT module 114 local to each client device 106 may be configured to convert a limited number of different spoken phrases - particularly phrases that invoke the automated assistant 120 - to text (or to other forms, such as reduced-dimension embeddings). Other voice inputs may be sent to the cloud-based automated assistant component 119, which may include a cloud-based TTS module 116 and / or a cloud-based STT module 117.

[0046] The cloud-based STT module 117 may be configured to use the virtually unlimited resources of the cloud to convert the audio data captured by the voice capture / TTS / STT module 114 to text (which may then be provided to the natural language processor 122). The cloud-based TTS module 116 may be configured to use the virtually unlimited resources of the cloud to convert text data (e.g., a natural language response formulated by the automated assistant 120) to computer-generated voice output. In some embodiments, the TTS module 116 may provide the computer-generated voice output to the client device 106 for direct output, for example, using one or more speakers. In other embodiments, the text data (e.g., natural language response) generated by the automated assistant 120 may be provided to the voice capture / TTS / STT module 114, which may then convert the text data to computer-generated voice for local output.

[0047] The automated assistant 120 (and in particular the cloud-based automated assistant component 119) can include a natural language processor 122, the aforementioned TTS module 116, the aforementioned STT module 117, a dialogue state tracker 124, a dialogue manager 126, and a natural language generator 128 (which in some embodiments may be combined with the TTS module 116). In some embodiments, one or more of the engines and / or modules of the automated assistant 120 may be omitted, combined, and / or implemented in components separate from the automated assistant 120.

[0048] In some embodiments, the automated assistant 120 generates response content in response to various inputs generated by a user of one of the client devices 106 1-N during a human-machine dialogue session with the automated assistant 120. The automated assistant 120 can provide the response content (e.g., via one or more networks when separate from the user's client device) for presentation to the user as part of the dialogue session. For example, the automated assistant 120 can generate response content in response to free-form natural language input provided via one of the client devices 106 1-N as used herein, free-form natural language input is input formulated by the user and not limited to a set of options presented for the user to select.

[0049] As used herein, a "dialogue session" can include a logically self-contained exchange of one or more messages between the user and the automated assistant 120 (and in some cases, other human participants) and / or the execution of one or more response actions by the automated assistant 120. The automated assistant 120 can distinguish multiple dialogue sessions with the user based on various signals such as the passage of time between sessions, changes in the user context between sessions (e.g., location, before / during / after scheduling a meeting, etc.), detection of one or more intermediate interactions between the user and the client device rather than the dialogue between the user and the automated assistant (e.g., the user temporarily switches applications, the user walks away and then later returns to a voice-activated product), locking / sleeping of the client device between sessions, changes in the client device for docking with one or more instances of the automated assistant 120, etc.

[0050] The natural language processor 122 of the automated assistant 120 processes inputs from the user via the client device 106 1-N The generated free-form natural language input and, in some embodiments, may generate an annotated output for use by one or more other components of the automated assistant 120. For example, the natural language processor 122 may process natural language free-form input generated by a user via one or more user interface input devices of the client device 1061. The generated annotated output includes one or more annotations of the natural language input and optionally one or more (e.g., all) of the items of the natural language input.

[0051] In some embodiments, the natural language processor 122 is configured to identify and annotate various types of syntactic information in the natural language input. For example, the natural language processor 122 may include a part-of-speech tagger (not depicted) that is configured to annotate items with their syntactic roles. For example, the part-of-speech tagger may tag each item with a part of speech such as "noun", "verb", "adjective", "pronoun", etc. Additionally, for example, in some embodiments the natural language processor 122 may additionally and / or alternatively include a dependency parser (not depicted) that is configured to determine syntactic relationships between items in the natural language input. For example, the dependency parser may determine which items modify other items, the subject and verb of a sentence, etc. (e.g., a parse tree) - and may annotate such dependencies.

[0052] In some embodiments, the natural language processor 122 may additionally and / or alternatively include an entity tagger (not depicted) configured to annotate entity references in one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), and the like. In some embodiments, data about entities may be stored in one or more databases, such as in a knowledge graph (not depicted). In some embodiments, the knowledge graph may include nodes representing known entities (and, in some cases, entity attributes) and edges connecting the nodes and representing relationships between the entities. For example, a "banana" node may be connected (e.g., as a child node) to a "fruit" node, which may in turn be connected (e.g., as a child node) to a "produce" and / or "food" node. As another example, a restaurant called "Hypothetical café" may be represented by nodes that also include attributes such as its address, the types of food served, business hours, contact information, and the like. The "Hypothetical café" node may be connected (e.g., by an edge representing a child-parent relationship) to one or more other nodes, such as a "restaurant" node, a "business" node, a node representing the city and / or state in which the restaurant is located, and the like.

[0053] The entity tagger of the natural language processor 122 may annotate references to entities at a high granularity level (e.g., such that all references to entity classes such as people can be identified) and / or at a lower granularity level (e.g., such that all references to specific entities such as a particular person can be identified). The entity tagger may rely on the content of the natural language input to resolve specific entities and / or may optionally communicate with a knowledge graph or other entity database to resolve specific entities.

[0054] In some embodiments, the natural language processor 122 may additionally and / or alternatively include a coreference resolver (not depicted) configured to group or "cluster" references to the same entity based on one or more context clues. For example, a coreference resolver may be used to resolve the term "there" in the natural language input "I liked Hypothetical Café last time we ate there" to "Hypothetical Café".

[0055] In some embodiments, one or more components of the natural language processor 122 may rely on annotations from one or more other components of the natural language processor 122. For example, in some embodiments, a named entity tagger may rely on annotations from a coreference resolver and / or a dependency parser when annotating all mentions of a particular entity. Additionally, for example, in some embodiments a coreference resolver may rely on annotations from a dependency parser when clustering references to the same entity. In some embodiments, when processing a particular natural language input, one or more components of the natural language processor 122 may use relevant prior inputs and / or other relevant data in addition to the particular natural language input to determine one or more annotations.

[0056] In the context of task-oriented dialogue, the natural language processor 122 may be configured to map free-form natural language inputs provided by a user at each turn of a dialogue session to semantic representations that may be referred to herein as "dialogue acts". The semantic representations, whether dialogue acts generated from user input or other semantic representations of automated assistant utterances, may take various forms. In some embodiments, the semantic representations may be modeled as discrete semantic frames. In other embodiments, the semantic representations may be formed as vector embeddings, such as in a continuous semantic space.

[0057] In some embodiments, a dialogue act (or more generally, a semantic representation) can in particular indicate one or more slot / value pairs corresponding to parameters of certain actions or tasks that the user may be attempting to perform via the automated assistant 120. For example, assume the user provides a free-form natural language input in the form: "Suggest an Indian restaurant for dinner tonight". In some embodiments, the natural language processor 122 can map the user input to a dialogue act that includes parameters such as, for example: intent (find restaurant); inform (cuisine = Indian, meal = dinner, time = tonight). Dialogue acts can appear in various forms, such as "greeting" (e.g., invoking the automated assistant 120), "inform" (e.g., providing parameters for slot filling), "intent" (e.g., finding an entity, ordering something), "request" (e.g., requesting specific information about an entity), "confirm", "affirm", and "thank you" (optionally, can close the dialogue session and / or be used as positive feedback and / or indicate that a positive reward value should be provided). These are just examples and are not intended to be limiting.

[0058] The dialogue state tracker 124 can be configured to track the "dialogue state", which includes, for example, the belief state of the user's goal (or "intent") during (and / or across multiple) human-computer dialogue sessions. In determining the dialogue state, some dialogue state trackers can attempt to determine the most likely values for the slots to be instantiated in the dialogue based on the user and system utterances in the dialogue session. Some techniques utilize a fixed ontology that defines a set of slots and a set of values associated with those slots. Additionally or alternatively, some techniques can be customized for individual slots and / or domains. For example, some techniques may require training a model for each slot type in each domain.

[0059] The dialogue manager 126 can be configured to map the current dialogue state provided, for example, by the dialogue state tracker 124 to one or more "response actions" among a plurality of candidate response actions then performed by the automated assistant 120. Depending on the current dialogue state, the response actions can occur in various forms. For example, the initial and intermediate dialogue states corresponding to the rounds of the dialogue session that occurred before the last round (e.g., when the task desired by the end user is performed) can be mapped to various response actions including the automated assistant 120 outputting additional natural language dialogue. This response dialogue can include, for example, a request for the user to provide parameters for an action (i.e., filling a slot) that the dialogue state tracker 124 believes the user intends to perform.

[0060] In some embodiments, the dialogue manager 126 can include a machine learning model such as a neural network. In some such embodiments, the neural network can take the form of, for example, a feedforward neural network having two hidden layers followed by a softmax layer. However, other configurations of neural networks and other types of machine learning models can be employed. In some embodiments where the dialogue manager 126 employs a neural network, the input to the neural network can include, but is not limited to, user actions, the previous response action (i.e., the action performed by the dialogue manager in the previous round), the current dialogue state (e.g., a binary vector provided by the dialogue state tracker 124 indicating which slots have been filled), and / or other values.

[0061] In various embodiments, the dialogue manager 126 can operate at the semantic representation level. For example, the dialogue manager 126 can receive new observations in the form of a semantic dialogue framework (which can include, for example, dialogue acts provided by the natural language processor 122 and / or the dialogue state provided by the dialogue state tracker 124) and randomly select a response action from among a plurality of candidate response actions. The natural language generator 128 can be configured to map the response action selected by the dialogue manager 126 to one or more utterances provided as output to the user at the end of each round of the dialogue session.

[0062] As noted above, in various embodiments, a user may be able to create a custom "dialogue routine" that the automated assistant 120 can later be able to effectively reformulate to perform various user-defined or user-selected tasks. In various embodiments, a dialogue routine may include a mapping between a command (e.g., a free-form natural language utterance converted to text or a reduced-dimension embedding, a free-form natural language input typed, etc.) and a task that will be performed, in whole or in part, by the automated assistant 120 in response to the command. Additionally, in some cases, a dialogue routine may include one or more user-defined "slots" (also referred to as "parameters" or "attributes") that are required to be filled with values (also referred to herein as "slot values") in order to perform the task. In various embodiments, once created, a dialogue routine may accept one or more values as input to fill one or more slots. In some embodiments, a dialogue routine may also include one or more user-enumerated values that can be used to fill a slot for one or more slots associated with the dialogue routine, but this is not required.

[0063] In various embodiments, when one or more required slots are filled with values, the task associated with the dialogue routine may be performed by the automated assistant 120. For example, assume that a user invokes a dialogue routine that requires two slots to be filled with values. If, during the invocation, the user provides values for the two slots, the automated assistant 120 may use those provided slot values to perform the task associated with the dialogue routine without asking the user for additional information. Thus, it is possible that a dialogue routine, when invoked, involves only a single "turn" of the dialogue (assuming the user has provided all necessary parameters in advance). On the other hand, if the user fails to provide a value for at least one required slot, the automated assistant 120 may automatically provide a natural language output that requests a value for the required but unfilled slot.

[0064] In some embodiments, each client device 106 may include a local dialogue routine index 113 configured to store one or more dialogue routines created by one or more users at the device. In some embodiments, each local dialogue routine index 113 may store dialogue routines created by any user at the corresponding client device 106. Additionally or alternatively, in some embodiments, each local dialogue routine index 113 may store dialogue routines created by a particular user who operates the client device 106's coordinated "ecosystem". In some cases, each client device 106 of the coordinated ecosystem may store dialogue routines created by a controlling user. For example, assume a user creates a dialogue routine at a first client device (e.g., 1061) in the form of a stand-alone interactive speaker. In some embodiments, the dialogue routine may be propagated to other client devices 106 (e.g., smart phones, tablet computers, another speaker, smart TVs, vehicle computing systems, etc.) that form part of the same coordinated ecosystem of client devices 106 and stored in the local dialogue routine index 113 of the other client devices 106.

[0065] In some embodiments, dialogue routines created by individual users may be shared among multiple users. To this end, in some embodiments, the global dialogue routine engine 130 may be configured to store dialogue routines created by multiple users in a global dialogue routine index 132. In some embodiments, the dialogue routines stored in the global dialogue routine index 132 may be available to selected users based on permissions granted by the creator (e.g., via one or more access control lists). In other embodiments, the dialogue routines stored in the global dialogue routine index 132 may be available to all users for free. In some embodiments, a dialogue routine created by a particular user at one client device 106 of the coordinated ecosystem of client devices may be stored in the global dialogue routine index 132 and may thereafter be available to the particular user at other client devices of the coordinated ecosystem (e.g., for optional download or online use). In some embodiments, the global dialogue routine engine 130 may be able to access both the globally available dialogue routines in the global dialogue routine index 132 and the locally available dialogue routines stored in the local dialogue routine index 113.

[0066] In some embodiments, a conversation routine may be limited to invocation by its creator. For example, in some embodiments, voice recognition technology may be used to assign a newly created conversation routine to the voice profile of its creator. When the conversation routine is later invoked, the automated assistant 120 may compare the speaker's voice to the voice profile associated with the conversation routine. If there is a match, the speaker may be authorized to invoke the conversation routine. If the speaker's voice does not match the voice profile associated with the conversation routine, in some cases, the speaker may not be permitted to invoke the conversation routine.

[0067] In some embodiments, a user may create a custom conversation routine that effectively overrides an existing canonical command and associated task. Suppose a user creates a new conversation routine to perform a user-defined task and uses a canonical command that was previously mapped to a different task to invoke the new conversation routine. In the future, when that particular user invokes the conversation routine, the user-defined task associated with the conversation routine may be fulfilled instead of the different task to which the canonical command was previously mapped. In some embodiments, the user-defined task may be executed in response to the canonical command only if it is the creator-user who is invoking the conversation routine (e.g., this may be determined by matching the speaker's voice to the voice profile of the creator of the conversation routine). If another user issues or otherwise provides the canonical command, a different task that is conventionally mapped to the canonical command may instead be executed.

[0068] Referring again to Figure 1 , in some embodiments, the task switchboard 134 may be configured to route data generated when a dialogue routine is invoked by a user to one or more appropriate remote computing systems / devices, e.g., such that tasks associated with the dialogue routine can be fulfilled. Although the task switchboard 134 is described separately from the cloud-based automated assistant component 119, this is not intended to be restrictive. In various embodiments, the task switchboard 134 may form an integral part of the automated assistant 120. In some embodiments, the data routed by the task switchboard 134 to the appropriate remote computing device may include one or more values to be used to populate one or more slots associated with the invoked dialogue routine. Additionally or alternatively, depending on the nature of the remote computing system / device, the data routed by the task switchboard 134 may include other information, such as the slots to be populated, data indicating the invocation command, data indicating the task to be performed (e.g., the user's perceived intent), etc. In some embodiments, once the remote computing systems / devices perform their roles in fulfilling the task, they may return response data directly and / or via the task switchboard 134 to the automated assistant 120. In various embodiments, the automated assistant 120 may then generate (e.g., via the natural language generator 128) a natural language output to be provided to the user, e.g., via one or more audio and / or visual output devices of the client device 106 operated by the invoking user.

[0069] In some embodiments, the task switchboard 134 may be operatively coupled to a task index 136. The task index 136 may store a plurality of candidate tasks that can be performed in whole or in part by the automated assistant 120. In some embodiments, the candidate tasks may include third-party software agents configured to automatically respond to orders, participate in human-machine conversations (e.g., as chatbots), etc. In various embodiments, these third-party software agents may interact with the user via the automated assistant 120, with the automated assistant 120 acting as a mediator. In other embodiments, particularly where the third-party agent itself is a chatbot, the third-party agent may be directly connected to the user, e.g., via the automated assistant 120 and / or the task switchboard 134. Additionally or alternatively, in some embodiments, the candidate tasks may include aggregating information provided by the user into a specific form, e.g., by populating specific slots, and presenting the information (e.g., in a predetermined format) to a third party, such as a human. In some embodiments, the candidate tasks may additionally or alternatively include tasks that do not necessarily need to be submitted to a third party, in which case the task switchboard 134 may not route the information to a remote computing device.

[0070] Suppose a user creates a new conversation routine to map a custom command to a task that is yet to be determined. In various implementations, the task switch 134 (or one or more components of the automated assistant 120) can compare the new conversation routine with multiple candidate tasks in the task index 136. For example, one or more user-defined slots associated with the new conversation routine can be compared with the slots associated with the candidate tasks in the task index 136. Additionally or alternatively, one or more user-enumerated values that can be used to fill the slots of the new conversation routine can be compared with the enumerated values that can be used to fill the slots associated with one or more of the multiple candidate tasks. Additionally or alternatively, other aspects of the new conversation routine (such as the command to be mapped, one or more other trigger words included in the user's invocation, etc.) can be compared with the various attributes of the multiple candidate tasks. Based on the comparison, a task to which the command is to be mapped can be selected from the multiple candidate tasks.

[0071] Suppose a user creates a new conversation routine invoked by the command "I want to order tacos". Further suppose this new conversation routine is intended to place a food order at a Mexican restaurant yet to be determined (perhaps the user is relying on the automated assistant 120 to guide the user to make the best choice). The user can define various slots associated with this task, such as shell type (e.g., crispy, soft, flour, corn, etc.), meat selection, type of cheese, type of sauce, toppings, etc., for example, by participating in a natural language conversation with the automated assistant 120. In some implementations, these slots can be compared with the slots to be filled in existing third-party food ordering applications (i.e., third-party agents) to determine which third-party agent is the most suitable. There may be multiple third-party agents configured to receive orders for Mexican food. For example, a first software agent can accept orders for pre-determined menu items (e.g., without options for custom toppings). A second software agent can accept custom taco orders and can thus be associated with slots such as toppings, shell type, etc. The new taco ordering conversation routine (including its associated slots) can be compared with the first software agent and the second software agent. Since the second software agent has slots that more closely align with those defined by the user in the new conversation routine, the second software agent can be selected, for example, by the task switch 134, for mapping the command "I want to order tacos" (or sufficiently syntactically / semantically similar utterances).

[0072] When a dialogue routine defines one or more slots that need to be filled in order to complete a task, the user is not required to proactively fill in these slots when initially invoking the dialogue routine. Instead, in various embodiments, when the user invokes a dialogue routine, with respect to the user not providing values for the required slots during the invocation, the automated assistant 120 can cause an output (e.g., audible, visual) to be provided, such as a natural language output soliciting these values from the user. For example, in the case of the taco ordering dialogue routine above, assume the user later provides the utterance "I want to order tacos". Since this dialogue routine has slots that need to be filled, the automated assistant 120 can respond by prompting the user for the values to be filled in any missing slots (e.g., shell type, toppings, meat, etc.). On the other hand, in some embodiments, the user can proactively fill in slots when invoking a dialogue routine. Assume the user issues the phrase "I want to order some fish tacos with hard shells". In this example, the slots for shell type and meat have already been filled with the corresponding values "hard shells" and "fish". Thus, the automated assistant 120 can only prompt the user for any missing slot values, such as toppings. Once all required slots have been filled with values, in some embodiments, the task switchboard 134 can take action to cause the task to be executed.

[0073] Figure 2 Depicts an example of how free-form natural language input provided by the user (the "FFNLI" in Figure 2 and elsewhere) can be used to invoke a dialogue routine and how data aggregated by the automated assistant 120 as part of implementing the dialogue routine can be propagated to various components for task fulfillment. The user provides the FFNLI (during one or more turns of a human-machine dialogue session) to the automated assistant 120 in typed form or as a spoken utterance. The automated assistant 120 interprets and parses the FFNLI into various semantic information, such as the user's intent, one or more slots to be filled, one or more values to be used to fill the slots, etc., for example, through a natural language processor 122 (not depicted in Figure 2 and / or a dialogue state tracker 124 (also not depicted in Figure 2 ).

[0074] The automated assistant 120, for example, through a dialogue manager 126 (in Figure 2 (not depicted in the figure) can consult the dialogue routine engine 130 to identify a dialogue routine that includes a mapping between commands and tasks contained in the FFNLI provided by the user. In some embodiments, the dialogue routine engine 130 can consult one or both of the local dialogue routine index 113 or the global dialogue routine index 132 of the computing device operated by the user. Once the automated assistant 120 selects a matching dialogue routine (e.g., including the dialogue routine that is semantically / syntactically most similar to the commands contained in the user's FFNLI), the automated assistant 120 can prompt the user for values to fill all unfilled and required slots for the dialogue routine, if necessary.

[0075] Once all necessary slots are filled, the automated assistant 120 can provide data indicating at least the values used to fill the slots to the task switch 134. In some cases, the data can also identify the slots themselves and / or one or more tasks mapped to the user's command. The task switch 134 can then select something that will be referred to herein as a "service" for facilitating the execution of the task. For example, in Figure 2 it, the services include a public switched telephone network ("PSTN") service 240, a service 242 for handling SMS and MMS messages, an email service 244, and one or more third-party software agents 246. As indicated by the ellipsis, any other number of additional services may or may not be utilized by the task switch 134. These services can be used to route data indicating the invoked dialogue routine or simply a "task request" to one or more remote computing devices.

[0076] For example, the PSTN service 240 can be configured to receive data indicating the invoked dialogue routine (including the values used to fill any required slots) and provide the data to a third-party client device 248. In this scenario, the third-party client device 248 can take the form of a computing device configured to receive a phone call, such as a cellular phone, a conventional phone, an Internet Protocol voice ("VOIP") phone, a computing device configured to make / receive phone calls, etc. In some embodiments, the information provided to such a third-party client device 248 can include a natural language output, such as generated by the automated assistant 120 (e.g., via the natural language generator 128) and / or by the PSTN service 240. This natural language output can include, for example, computer-generated utterances conveying the task to be performed and the parameters associated with the task (i.e., the values of the required slots), and / or enable the recipient to participate in a limited dialogue designed to enable the fulfillment of the user's task (e.g., much like a harassing phone call). This natural language output can be presented by the third-party computing device 248 as a human-perceivable output 250, for example, auditorily, visually, as tactile feedback, etc.

[0077] Suppose a dialogue routine is created to place a pizza order. Further suppose that the task identified for the dialogue routine (e.g., by the user or by the task switch 134) is to provide the user's pizza order to a particular pizza store that lacks its own third-party software agent. In some such embodiments, in response to a call to the dialogue routine, the PSTN service 240 can place a phone call to the phone at the particular pizza store. When an employee at the particular pizza store answers the phone, the PSTN service 240 can initiate an automated (e.g., IVR) dialogue notifying the pizza store employee that the user wishes to order a pizza with a crust type and toppings specified by the user when the user invoked the dialogue routine. In some embodiments, the pizza store employee may be required to confirm that the pizza store will fulfill the user's order, e.g., by pressing "1", providing an oral confirmation, etc. Once this confirmation is received, it can be provided, for example, to the PSTN service 240, which can in turn forward the confirmation information (e.g., via the task switch 134) to the automated assistant 120, which can then notify the user that the pizza is on the way (e.g., using an audible and / or visual natural language output such as "your pizza is on the way"). In some embodiments, the pizza store employee may be able to request additional information (e.g., a slot not specified during the creation of the dialogue routine) that the user may not have specified when invoking the dialogue routine.

[0078] The SMS / MMS service 242 can be used in a similar manner. In various embodiments, data indicating the invoked dialogue routine, such as one or more slot / value pairs, can be provided to the SMS / MMS service 242, for example, via the task switchboard 134. Based on this data, the SMS / MMS service 242 can generate text messages in various formats (e.g., SMS, MMS, etc.) and send the text message to a third-party client device 248, which can again be a smart phone or another similar device. The person operating the third-party client device 248 (e.g., a pizza shop employee) can then use the text message (e.g., read it aloud, have it read aloud, etc.) as a human-perceivable output 250. In some embodiments, the text message can request the person to provide a response, such as "REPLY ‘1’ IF YOU CAN FULFILL THIS ORDER. REPLY ‘2’ IF YOU CANNOT". In this way, similar to the example described above in the case of the PTSN service 240, the first user invoking the dialogue routine can exchange data asynchronously with the second user operating the third-party device 248 so that the second user can assist in fulfilling the task associated with the invoked dialogue routine. The email service 244 can operate similarly to the SMS / MMS service 242, except that the email service 244 utilizes email-related communication protocols (such as IMAP, POP, SMTP, etc.) to generate and / or exchange emails with the third-party computing device 248.

[0079] Services 240-244 and the task switchboard 134 enable users to create dialogue routines to interact with third parties while reducing the requirements for third parties to implement complex software services with which they can interact. However, at least some third parties may prefer to build or have the ability to build third-party software agents 246 that are configured to automatically interact with remote users (e.g., via the automated assistants 120 engaged by those remote users). Thus, in various embodiments, one or more third-party software agents 246 can be configured to interact with the automated assistants 120 and / or the task switchboard 134 such that users can create dialogue routines that can match these third-party agents 246.

[0080] Suppose a user creates a dialogue routine that matches a specific third-party agent 246 (as described above) based on slots, enumerated potential slot values, other information, etc. When invoked, the dialogue routine can cause the automated assistant 120 to send data indicating the dialogue routine (including slot values provided by the user) to the task switchboard 134. The task switchboard 134 can in turn provide this data to the matching third-party software agent 246. In some embodiments, the third-party software agent 246 can perform tasks associated with the dialogue routine and return results (e.g., success / failure messages, natural language output, etc.) to the task switchboard 134, for example.

[0081] As indicated by the arrow from the third-party agent 246 directly to the automated assistant 120, in some embodiments, the third-party software agent 246 can interface directly with the automated assistant 120. For example, in certain embodiments, the third-party software agent 246 can provide data (e.g., status data) to the automated assistant 120 that enables the automated assistant 120 to generate natural language output, for example, via the natural language generator 128, and the natural language output is then presented to the user who invoked the dialogue routine as an audible and / or visual output, for example. Additionally or alternatively, the third-party software agent 246 can generate its own natural language output, which is then provided to the automated assistant 120, and the automated assistant 120 in turn outputs the natural language output to the user.

[0082] As indicated by Figure 2 the other arrows among the various arrows in, the above examples are not intended to be limiting. For example, in some embodiments, the task switchboard 134 can provide data indicating the invoked dialogue routine to one or more services 240-244, and these services can in turn provide this data (or modified data) to one or more third-party software agents 246. Some of these third-party software agents 246 can be configured to receive, for example, text messages or emails and automatically generate responses that can be returned to the task switchboard 134 and continue to be returned to the automated assistant 120.

[0083] The conversation routines configured according to selected aspects of the present disclosure are not limited to tasks performed / fulfilled remotely from the client device 106. Instead, in some embodiments, a user can access the automated assistant 120 to create conversation routines that perform various tasks locally. As a non-limiting example, a user can create a conversation routine that uses a single command to configure multiple settings of a mobile device, such as a smart phone, at once. For example, a user can create a conversation routine that receives all Wi-Fi settings, Bluetooth settings, and hotspot settings as input at once and changes those settings accordingly. As another example, a user can create a conversation routine that is invoked when the user says "I’m gonna be late". The user can instruct the automated assistant 120 that this command should cause the automated assistant 120 to notify another person, such as the user's spouse, that the user will be late to a certain destination, for example, using a text message, an email, etc. In some cases, the slot for such a conversation routine can include a predicted time at which the user will reach the user's predetermined destination, which can be filled by the user or automatically predicted by the automated assistant 120, for example, based on location coordinate data, calendar data, etc.

[0084] In some embodiments, a user may be able to configure a conversation routine to use preselected slot values in specific slots such that the user does not need to provide those slot values and will not be prompted for those values when the user does not provide them. Assume that a user creates a pizza ordering conversation routine. Further assume that the user always prefers a thin crust. In various embodiments, unless the user specifies otherwise, the user can instruct the automated assistant 120 that when this specific conversation routine is invoked, the slot "crust type" should be automatically filled with the default value "thin crust". Thus, if the user occasionally wants to order a different crust type (e.g., the user has a visitor who prefers a thick crust), then in addition to the user being able to specifically request a different type of crust (e.g., "Hey assistant,order me a hand-tossed pizza"), the user can invoke the conversation routine as usual. If the user simply says "Hey assistant,order me a pizza", then the automated assistant 120 may have assumed a thin crust and prompted the user for the other required slot values. In some embodiments, the automated assistant 120 can "learn" over time which slot values the user prefers. Later, when the user invokes a conversation routine without explicitly providing those learned slot values, for example, if the user has provided those slot values more than a predetermined number of times or more than a specific threshold frequency of invoking the conversation routine, then the automated assistant 120 can assume those values (or ask the user to confirm those slot values).

[0085] Figure 3 Depict an example process flow that may occur when a user invokes a pizza ordering dialog routine according to various implementations. At 301, the user invokes the pizza ordering dialog routine by, for example, issuing the invocation phrase "Order a thincrust pizza" to the automated assistant client 118. At 302, the automated assistant client 118 provides the invocation phrase to the cloud-based automated assistant component ("CBAAC") 119, such as as a recording, transcribed text segment, dimensionality-reduced embedding, etc. At 303, various components of the CBAAC 119 (such as the natural language processor 122, the dialogue state tracker 124, the dialogue manager 126, etc.) can use various cues (such as dialogue context, verb / noun dictionary, canonical utterances, thesaurus (e.g., lexicon), etc.) to process the request as described above to extract information such as the object "pizza" and the attribute (or "slot value") of "thin crust".

[0086] At 304, this extracted data can be provided to the task switchboard 134. In some embodiments, at 305, the task switchboard 134 can consult the dialogue routine engine 130 to identify, for example, a dialogue routine that matches the user's request based on the data extracted at 303 and received at 304. As Figure 3 shown, the dialogue routine identified in this example includes the action of "order" (which can itself be a slot), the object "pizza" (which can also be a slot in some cases), the attribute (or "slot") of "crust" (which is required), another attribute (or slot) of "topping", and the so-called "implementor" of "order_service". Depending on how the user creates the dialogue routine and / or whether the dialogue routine matches a specific task (e.g., a specific third-party software agent 246), the "implementor" can be, for example Figure 2 any of the services 240 - 244 and / or one or more third-party software agents 246.

[0087] At 306, it can be determined, for example by the task switchboard 134, that one or more required slots for the dialogue routine have not been filled with values. Thus, the task switchboard 134 can notify a component such as the automated assistant 120 (e.g., Figure 3 the automated assistant client 118 in, but it could be another component such as one or more CBAACs 119) one or more slots are to be filled with slot values. In some embodiments, the task switchboard 134 can generate the necessary natural language output to prompt the user for these unfilled slots (e.g., "what topping?"), and the automated assistant client 118 can simply provide this natural language output to the user, e.g., at 307. In other embodiments, the data provided to the automated assistant client 118 can provide a notification of missing information, and the automated assistant client 118 can interface with one or more components of the CBAAC 119 to generate the natural language output that is presented to the user to prompt the user for the missing slot values.

[0088] Although not shown for sake of brevity and completeness in Figure 3 it, the slot values provided by the user can be returned to the task switchboard 134. At 308, in the case where all the required slots are filled with the slot values provided by the user, the task switchboard 134 can then be able to formulate a complete task. This complete task can be provided, for example, by the task switchboard 134 to the appropriate implementer 350, which as noted above can be one or more services 240-244, one or more third-party software agents 246, etc.

[0089] Figure 4 is a flowchart illustrating an example method 400 according to embodiments disclosed herein. For convenience, the operations of the flowchart are described with reference to a system that performs the operations. This system can include various components of various computer systems, such as one or more components of a computing system that implements the automated assistant 120. Additionally, although the operations of method 400 are shown in a particular order, this is not intended to be limiting. One or more operations can be reordered, omitted, or added.

[0090] At block 402, the system can receive a first free-form natural language input from the user, e.g., at one or more input components of the client device 106. In various embodiments, the first free-form natural language input can include a command to perform a task. As a working example, assume the user provides the spoken utterance "I want a pizza".

[0091] At block 404, the system can perform semantic processing on free-form natural language input. For example, one or more CBAACs 119 can compare the user's utterance (or its lower-dimensional embedding) with one or more canonical commands, various lexicons, etc. The natural language processor 122 can perform various aspects of the above analysis to identify entities, perform anaphora resolution, tag parts of speech, etc. At step 406, the system can determine that the automated assistant 120 cannot interpret the command based on the semantic processing of block 404. In some embodiments, at block 408, the system can provide an output at one or more output components of the client device 106 that solicits clarification from the user about the command, such as outputting the natural language output: "I don't know how to order a pizza".

[0092] At block 410, the system can receive a second free-form natural language input from the user at one or more of the input components. In various embodiments, the second free-form natural language input can identify one or more slots that are required to be filled with values to fulfill the task. For example, the user can provide a natural language input such as "to order a pizza, you need to know the type of crust and a list of toppings". This particular free-form natural language input identifies two slots: the type of crust and the list of toppings (which may be any number of slots technically depending on how many toppings the user desires).

[0093] As indirectly mentioned above, in some embodiments, a user may be able to enumerate a list of potential or candidate slot values for a given slot of a conversation routine. In some embodiments, this may effectively limit the slot to one or more values from the enumerated list. In some cases, enumerating the possible values for a slot may enable the automated assistant 120 to determine which slot will be filled with a particular value and / or determine that a provided slot value is invalid. For example, assume that the user invokes a conversation routine by the phrase "order me a pizza with thick crust, tomatoes, and tires". The automated assistant 120 may match "thick crust" with the slot "crust type" based on the fact that "thick crust" is one of the enumerated list of potential values. This also applies to "tomatoes" and the slot "toppings". However, since "tires" is not likely to be in the enumerated list of potential toppings, the automated assistant 120 may ask the user to correct the specified topping of tires. In other embodiments, the enumerated list provided by the user may simply include non-restrictive potential slot values that can be used by the automated assistant 120 (e.g., as suggestions to be provided to the user during future invocations of the conversation routine). This can be beneficial in contexts such as pizza ordering where the list of possible pizza toppings may be large and may vary significantly across pizza companies and / or over time (e.g., a pizza store may offer different toppings at different times of the year depending on seasonal products).

[0094] Continuing with the working example, the automated assistant 120 may ask questions such as "what are the possible pizza crust types?" or "what are the possible toppings?" The user may respond to each such question by providing an enumerated list of possibilities and indicating whether the enumerated list is intended to be restrictive (i.e., slot values outside of those enumerated lists are not allowed) or simply exemplary. In some cases the user may respond that a given slot is not limited to specific values, such that the automated assistant 120 is not restricted and may fill the slot with any slot value provided by the user.

[0095] Returning to Figure 4 , once the user has finished defining any required / optional slots and / or enumerating a list of potential slot values, at block 412, the system (e.g., the dialogue routine engine 130) can store a dialogue routine that includes a mapping between the command provided by the user and the task. The created dialogue routine can be configured to accept one or more values for populating one or more slots as input and cause the task associated with the dialogue routine to be performed, for example, at the remote computing device as previously described. The dialogue routine can be stored in various formats, and it is not critical which format is used in the context of the present disclosure.

[0096] In some embodiments, particularly when the user explicitly requests the automated assistant 120 to generate a dialogue routine rather than in the case where the automated assistant 120 first fails to interpret what the user says, various operations such as operations 402 - 408 can be omitted. Figure 4 For example, the user can simply say a phrase such as the following to the automated assistant 120 to trigger the creation of a dialogue routine: "Hey assistant, I want to teach you a new trick" or something to that effect. This can trigger a portion of method 400 starting, for example, at block 410. Of course, many users may not be aware that the automated assistant 120 is capable of learning dialogue routines. Therefore, it can be beneficial for the automated assistant 120 to guide the user through the process as described above with respect to blocks 402 - 408 when the user issues a command or request that the automated assistant 120 cannot interpret.

[0097] At a later time, at block 414, the system can receive subsequent free - form natural language input from the user at one or more input components of the same client device 106 or a different client device 106 (e.g., another client device in the same coordinated ecosystem of client devices). The subsequent free - form natural language input can include a command or some syntactic and / or semantic variation thereof that can invoke the dialogue routine based on the mapping stored at block 412.

[0098] At block 416, the system can identify one or more values to be used for populating one or more slots based on the subsequent free - form natural language input or additional free - form natural language input (e.g., solicited from a user who failed to provide one or more required slot values upon invocation of the dialogue routine), where the one or more slots are required to be populated with values in order to perform the task associated with the dialogue routine. For example, if the user simply invokes the dialogue routine without providing values for any required slots, the automated assistant 120 can solicit slot values from the user, for example, one at a time, in batches, etc.

[0099] In some embodiments, at block 418, the system can send data indicating at least one or more values to be used to populate one or more slots to a remote computing device such as a third-party client device 248 and / or to a third-party software agent 246, for example, via one or more of the task switcher 134 and / or services 240 - 244. In various embodiments, the sending can cause the remote computing device to perform a task. For example, if the remote computing device operates a third-party software agent 246, receiving data from the task switcher 134, for example, can trigger the third-party software agent 246 to perform the task using the slot values provided by the user.

[0100] The techniques described herein can be used to effectively "glue together" tasks that can be performed by various different third-party software applications (e.g., third-party software agents). In fact, it is entirely possible to create a single conversation routine in which multiple tasks are performed by multiple parties. For example, a user can create a conversation routine that is invoked by a phrase such as "Hey assistant, I want to take my wife to dinner and a movie". The user can define slots associated with multiple tasks (such as making a dinner reservation and purchasing movie tickets) within a single conversation routine. Slots for making a dinner reservation can include, for example, the restaurant (assuming the user has selected a specific restaurant), type of cuisine (if the user has not selected a restaurant), price range, time range, review range (e.g., above three stars), etc. Slots for purchasing movie tickets can include, for example, the movie, theater, time range, price range, etc. Later, when the user invokes this "dinner and a movie" reservation, the automated assistant 120 can solicit such values from the user in cases where the user has not proactively provided slot values to populate the respective slots. Once the automated assistant has slot values for all the required slots for each task of the conversation routine, the automated assistant 120 can send data to various remote computing devices as previously described to have each task performed. In some embodiments, the automated assistant 120 can publish to the user which tasks have been performed and which are still pending. In some embodiments, the automated assistant 120 can notify the user when all tasks have been performed (or in cases where one or more of these tasks cannot be performed).

[0101] In some cases (whether or not multiple tasks are glued together in a single conversation booking), the automated assistant 120 can prompt the user for a specific slot value by first searching for potential slot values (e.g., movies at a cinema, show times, available dinner bookings, etc.) and then presenting these potential slot values to the user (e.g., as suggestions or as an enumerated list of possibilities). In some embodiments, the automated assistant 120 can utilize various aspects of the user, such as the user's preferences, past user activities, etc., to narrow the scope of such a list. For example, if the user (and / or the user's spouse) prefers a particular type of movie (e.g., highly rated, comedy, horror, action, drama, etc.), then the automated assistant 120 can narrow the scope of the list of potential slot values before presenting them to the user.

[0102] The automated assistant 120 can take various approaches regarding payment that may be necessary to fulfill a particular task (e.g., ordering a product, making a booking, etc.). In some embodiments, the automated assistant 120 can access, if necessary, payment information provided by the user that the automated assistant 120 can provide, for example, to a third-party software agent 246 (e.g., one or more credit cards). In some embodiments, when the user creates a conversation routine to fulfill a task that requires payment, the automated assistant 120 can prompt the user for payment information and / or for permission to use payment information that has already been associated with the user's profile. In some embodiments where the data indicating the invoked conversation routine (including one or more slot values) is provided to a third-party computing device (e.g., 248) to be output as a natural language output, the user's payment information may or may not be provided. In the case where the user's payment information is not provided, for example, when ordering food, the food vendor can simply request payment from the user when delivering the food to the user's doorstep.

[0103] In some embodiments, the automated assistant 120 can "learn" new conversation routines by analyzing the user's interactions with one or more applications running on one or more client computing devices to detect patterns. In various embodiments, the automated assistant 120 can provide a natural language output to the user, for example, proactively during an existing human-computer conversation or as another type of notification asking the user if they would like to assign a sequence of actions / tasks that they typically perform to a spoken command (e.g., a pop-up card, a text message, etc.), effectively building and recommending conversation routines without the user explicitly asking for them.

[0104] As an example, assume that a user repeatedly visits a single food ordering website (e.g., associated with a restaurant), views web pages associated with a menu, and then opens a separate phone application to place a call to a phone number associated with the same food ordering website using user actions. The automated assistant 120 can detect this pattern and generate a conversation routine for recommendation to the user. In some embodiments, the automated assistant 120 can scrape the menu web page to obtain potential slots and / or potential slot values that can be incorporated into the conversation routine, and map one or more commands (which can be suggested by the automated assistant 120 or provided by the user) to the food ordering task. In this case, the food ordering task can include calling a phone number as described above with respect to the PSTN 240 and outputting a natural language message (e.g., a harassing call) to an employee of the food ordering website.

[0105] Other sequences of actions for ordering food (or generally performing other tasks) can also be detected. For example, assume that a user typically opens a third-party client application to order food, and the third-party client application is a GUI-based application. The automated assistant 120 can detect this and, for example, determine that the third-party client application interfaces with a third-party software agent (e.g., 246). In addition to interacting with the third-party client application, this third-party software agent 246 may also have been configured to interact with the automated assistant. In such a scenario, the automated assistant 120 can generate a conversation routine to interact with the third-party software agent 246. Alternatively, assume that the third-party software agent 246 is currently unable to interact with the automated assistant. In some embodiments, the automated assistant can determine what information the third-party client application provides for each order, and can use that information to generate slots for the conversation routine. When the user later invokes the conversation routine, the automated assistant 120 can populate the required slots and then, based on these slots / slot values, generate data that is compatible with the third-party software agent 246.

[0106] Figure 5 is a block diagram of an example computing device 510 that can optionally be utilized to perform one or more aspects of the techniques described herein. In some embodiments, one or more of the client computing device and / or other components can include one or more components of the example computing device 510.

[0107] The computing device 510 generally includes at least one processor 514 that communicates with a number of peripheral devices via a bus subsystem 512. These peripheral devices can include a storage subsystem 524 (including, for example, a memory subsystem 525 and a file storage subsystem 526), a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices allow a user to interact with the computing device 510. The network interface subsystem 516 provides an interface to an external network and is coupled to a corresponding interface device in other computing devices.

[0108] The user interface input device 522 can include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated into the display, audio input devices such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into the computing device 510 or onto a communication network.

[0109] The user interface output device 520 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), flat panel devices such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display via, for example, an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways to output information from the computing device 510 to a user or to another machine or computing device.

[0110] The storage subsystem 524 stores programming and data constructs that provide some or all of the functionality described in the modules herein. For example, the storage subsystem 524 can include logic for performing Figure 4 selected aspects of the methods and for implementing the various components depicted in Figures 1 to 3 .

[0111] These software modules are generally executed by the processor 514, alone or in combination with other processors. The memory 525 used in the storage subsystem 524 can include a number of memories, including a main random access memory (RAM) 530 for storing instructions and data during program execution and a read-only memory (ROM) 532 storing fixed instructions. The file storage subsystem 526 can provide persistent storage for program and data files and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functionality of certain embodiments can be stored in the storage subsystem 524 by the file storage subsystem 526 or in other machines accessible by the processor 514.

[0112] The bus subsystem 512 provides a mechanism for enabling the various components and subsystems of the computing device 510 to communicate with each other as expected. Although the bus subsystem 512 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0113] The computing device 510 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 510 depicted in Figure 5 is only intended as a specific example for the purpose of illustrating some embodiments. Many other configurations of the computing device 510 may have more or fewer components than the computing device depicted in Figure 5 .

[0114] In cases where certain embodiments discussed herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, information about the user's time, the user's biometric information, and the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use user personal information only when explicit authorization to do so is received from the relevant user.

[0115] For example, the user is provided with control over whether a program or feature collects user information about a particular user or other users related to the program or feature. Each user for whom personal information is to be collected is presented with one or more options to allow control over the information collection related to that user, to provide permission or authorization regarding whether information is collected and regarding which portions of the information are to be collected. For example, one or more such control options can be provided to the user via a communication network. In addition, certain data may be processed in one or more ways before it is stored or used such that personally identifiable information is removed. As an example, a user's identity can be processed such that personally identifiable information cannot be determined. As another example, a user's geographical location can be generalized to a larger area such that the user's specific location cannot be determined.

[0116] Although several embodiments have been described and illustrated herein, various other means and / or structures may be utilized for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each such variation and / or modification is to be regarded as being within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend upon the one or more specific applications for which the teachings are used. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific embodiments described herein. Accordingly, it is to be understood that the above-described embodiments are presented by way of example only, and that embodiments may be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. Embodiments of the present disclosure are directed to each and every separate feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.< / user>

Claims

1. A method implemented by one or more processors, comprising: Receiving, at one or more input components of a computing device, one or more voice inputs from a user, the one or more voice inputs directed to an automated assistant executed by one or more of the processors, wherein the one or more voice inputs define a custom voice command and identify a task to be performed in response to the automated assistant receiving the custom voice command and one or more slots that are required to be filled with values in order to perform the task; Identifying the custom voice command, the task, and the one or more slots based on a speech recognition output generated from the one or more voice inputs; and Creating and storing a custom dialogue routine, the custom dialogue routine including a mapping between the custom voice command and the task, and the custom dialogue routine accepting one or more values for filling the one or more slots as inputs, wherein subsequent utterances of the custom command cause the automated assistant to participate in the custom dialogue routine.

2. The method according to claim 1, further comprising: Receiving, by the automated assistant, a subsequent voice input, wherein the subsequent voice input includes the custom command; Identifying the custom command based on a subsequent speech recognition output generated from the subsequent voice input; and Causing the automated assistant to participate in the custom dialogue routine based on identifying the custom command in the subsequent voice input and based on the mapping.

3. The method according to claim 2 further comprises: Identifying the one or more values to be used for filling the one or more slots based on an additional subsequent speech recognition output generated from the subsequent voice input or an additional subsequent voice input, the one or more slots being required to be filled with values in order to perform the task.

4. The method according to claim 3, wherein, The task to which the command is mapped includes a third-party agent task, and the method further comprises: sending data indicating at least the one or more values to be used for filling the one or more slots to a remote computing device, wherein the sending causes a third-party agent executing on the remote computing device to perform the third-party agent task.

5. The method according to claim 2 further comprises: Generating a natural language output as part of the automated assistant participating in the custom dialogue routine, the natural language output requesting performance of the task based on one or more values provided to fill one or more of the slots.

6. The method according to claim 5 further comprises: Sending the natural language output as part of a Short Message Service (SMS) text message to a remote computing device.

7. The method according to claim 5 further comprises: Sending the natural language output to a remote device to cause the remote device to audibly output the natural language output.

8. The method according to claim 1, wherein The one or more voice inputs include slot-filling voice inputs from the user, wherein the slot-filling voice inputs include an enumerated list of possible values provided by the user to fill one or more of the slots.

9. A system comprising one or more processors and a memory storing instructions that, in response to execution of the instructions by the one or more processors, cause the one or more processors to perform the following operations: Receiving, at one or more input components of a computing device, one or more voice inputs from a user, the one or more voice inputs directed to an automated assistant executed by one or more of the processors, wherein, The one or more voice inputs define custom voice commands and identify tasks to be performed in response to the custom voice commands being received by the automated assistant, and one or more slots that are required to be filled with values in order to perform the tasks; Identify the custom voice commands, the tasks, and the one or more slots based on a speech recognition output generated from the one or more voice inputs; And Create and store a custom conversation routine that includes a mapping between the custom voice commands and the tasks, and the custom conversation routine accepts one or more values for filling the one or more slots as input, wherein subsequent utterances of the custom command cause the automated assistant to participate in the custom conversation routine.

10. The system according to claim 9, further comprising instructions for performing the following operations: Receiving subsequent voice input by the automated assistant, wherein, The subsequent voice input includes the custom command; Identify the custom command based on a subsequent speech recognition output generated from the subsequent voice input; And Cause the automated assistant to participate in the custom conversation routine based on identifying the custom command in the subsequent voice input and based on the mapping.

11. The system according to claim 10, further comprising instructions for performing the following operations: Identify the one or more values to be used for filling the one or more slots that are required to be filled with values in order to perform the tasks, based on an additional subsequent speech recognition output generated from the subsequent voice input or an additional subsequent voice input.

12. The system according to claim 11, wherein, The tasks to which the commands are mapped include third-party agent tasks, and the system further comprises instructions for performing the following operations: Send data indicating at least the one or more values to be used for filling the one or more slots to a remote computing device, wherein the sending causes a third-party agent executing on the remote computing device to perform the third-party agent task.

13. The system according to claim 10, further comprising instructions for performing the following operations: Generate a natural language output as part of the automated assistant's participation in the custom conversation routine, the natural language output requesting performance of the task based on one or more values provided to fill one or more of the slots.

14. The system according to claim 13, further comprising instructions for performing the following operations: Send the natural language output as part of a Short Message Service (SMS) text message to a remote computing device.

15. The system according to claim 13, further comprising instructions for performing the following operations: Send the natural language output to a remote device to cause the remote device to aurally output the natural language output.

16. The system according to claim 9, wherein, The one or more voice inputs include slot-filling voice inputs from the user, wherein the slot-filling voice inputs include an enumerated list of possible values provided by the user to fill one or more of the slots.

17. At least one non-transitory computer-readable medium including instructions that, when executed by one or more processors, cause the one or more processors to perform the following operations: Receiving, at one or more input components of a computing device, one or more voice inputs from a user, the one or more voice inputs being directed to an automated assistant executed by one or more of the processors, wherein, The one or more voice inputs define custom voice commands and identify tasks to be performed in response to the custom voice commands received by the automated assistant, as well as one or more slots that are required to be filled with values to fulfill the tasks; Identify the custom voice commands, the tasks, and the one or more slots based on a speech recognition output generated from the one or more voice inputs; And Create and store a custom conversation routine that includes a mapping between the custom voice commands and the tasks, and the custom conversation routine accepts one or more values for filling the one or more slots as inputs, wherein subsequent utterances of the custom commands cause the automated assistant to participate in the custom conversation routine.

18. The at least one non-transitory computer-readable medium according to claim 17, further comprising instructions for performing the following operations: Receiving subsequent voice input by the automated assistant, wherein, The subsequent voice input includes the custom command; Identify the custom command based on a subsequent speech recognition output generated from the subsequent voice input; And Cause the automated assistant to participate in the custom conversation routine based on identifying the custom command in the subsequent voice input and based on the mapping.

19. The at least one non-transitory computer-readable medium according to claim 18, further comprising instructions for performing the following operations: identify the one or more values to be used for filling the one or more slots, which are required to be filled with values to fulfill the tasks, based on an additional subsequent speech recognition output generated from the subsequent voice input or an additional subsequent voice input.

20. The at least one non-transitory computer-readable medium according to claim 19, wherein, The tasks to which the commands are mapped include third-party agent tasks, and the at least one non-transitory computer-readable medium further comprises instructions for performing the following operations: send data indicating at least the one or more values to be used for filling the one or more slots to a remote computing device, wherein the sending causes a third-party agent executed on the remote computing device to perform the third-party agent task.

Citation Information

Patent Citations

  • Training an at least partial voice command system

    CN105027197A

  • Parameter collection and automatic dialog generation in dialog systems

    CN108701454A