Recommendation to include automatic assistant actions in automatic assistant routine
By recommending and automatically adding actions in the automatic assistant routine, the problem of users consuming resources for a long time of interaction is solved, and the effect of simplifying human-computer interaction and improving efficiency is achieved.
Patent Information
- Application Number
- CN202510362146.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2018-05-07
- Filing Date
- 2019-05-04
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing automatic assistant system, users need to perform actions through long-term interaction, which consumes a large amount of computer and network resources, and the action addition process is cumbersome.
By recommending automatic assistant actions in existing automatic assistant routines and automatically adding them after user confirmation, it reduces user input and resource consumption, and uses shortcut commands and user interface output prompt actions to add.
It simplifies human-computer interaction, reduces user input, saves computing and network resources, and improves the efficiency and user experience of the automatic assistant system.
Smart Images

Figure CN120407632A_ABST
Abstract
Description
Division Explanation
[0001] This application is a divisional application of Chinese Patent Application No. 201980037318.8 with an application date of May 4, 2019. Technical Field
[0002] The present disclosure relates to recommending including automated assistant actions in automated assistant routines. Background Art
[0003] Humans can participate in human-machine conversations with interactive software applications referred to herein as "automated assistants" (also known as "chatbots", "interactive personal assistants", "intelligent personal assistants", "personal voice assistants", "conversation agents", etc.). For example, a human (who can be referred to as a "user" when they interact with an automated assistant) can provide commands, queries, and / or requests using spoken natural language input (i.e., utterances) that can be converted into text and then processed in some cases and / or by providing text (e.g., typed) natural language input.
[0004] An automated assistant can execute routines of multiple actions in response to, for example, receiving a specific command (such as a shortcut command). For example, in response to receiving the spoken utterance "Good Night", the automated assistant can cause a series of actions to be performed, such as turning off a networked lighting fixture, rendering a weather forecast for tomorrow, and rendering the user's agenda for tomorrow. Automated assistant routines can be specific to a user and / or the ecosystem of a client device, and can provide the user with control to manually add certain actions to certain routines. For example, a first user can have: a "Good Morning" automated assistant routine with a first set of good morning automated assistant actions defined by the first user, a "good night" automated assistant routine with a first set of good night automated assistant actions defined by the first user, and an additional automated assistant routine with additional automated assistant actions defined by the first user. A separate second user can have: a "good morning" automated assistant routine with a different set of good morning automated assistant actions defined by the second user, a "Good Night" routine with a different set of good night automated assistant actions defined by the second user, etc. Summary of the Invention
[0005] As described herein, various assistant routines can be initiated in response to detecting a shortcut command in a spoken or typed user interface input of a user, in response to a user's interaction with a virtual or hardware element at a client device, in response to detecting a user gesture, and / or in response to other abbreviated user interface inputs. The user interface inputs that are abbreviated for initiating assistant routines are abbreviated because they require less user input and / or less processing of user input compared to the user input that would be required to cause the performance of an action of an assistant routine without the abbreviation. For example, a shortcut command for an automated routine that causes a series of actions to be performed can be abbreviated because the length of the shortcut command is shorter than the length of the command that would need to be spoken / typed to cause the assistant to perform the set of actions without the compressed command. Additionally or alternatively, various assistant routines can be automatically initiated after one or more conditions occur and optionally without requiring explicit user interface input. This automatic initiation of assistant routines may also require less user input and / or less processing of user input compared to the user input that would be required to cause the performance of an action of an assistant routine.
[0006] Despite the presence of assistant routines, there are still various assistant actions that are not incorporated into any of the user's assistant routines but are instead performed by the user through long interactions with an assistant and / or through long interactions with other computer applications. Each occurrence of such a long interaction can consume a significant amount of computer and / or network resources. For example, to turn on a networked light in a kitchen via an assistant, the user may be required to speak "turn on the kitchen lights" or a similar spoken utterance to an assistant interface of a client device. The audio data corresponding to the spoken utterance can be transmitted from the client device to a remote system. The remote system can process the audio data (e.g., speech-to-text processing, natural language processing of the text, etc.) to determine an appropriate command for the spoken utterance, can transmit the command to an agent to cause the kitchen light to be turned on, can receive a confirmation that the kitchen light has been turned on from the agent, and can transmit data to the client device to cause the client device to render a notification that the light has been turned on.
[0007] In view of these and other considerations, embodiments disclosed herein relate to recommending an automated assistant action for inclusion in a user's existing automated assistant routine, where the existing automated assistant routine includes one or more pre-existing automated assistant actions among pre-existing automated assistant actions. If the user confirms the recommendation via an affirmative user interface input, the automated assistant action can be automatically added to the existing automated assistant routine. Thereafter, when the automated assistant routine is initialized, the pre-existing automated assistant actions of the routine are performed, as well as the automated assistant action that was automatically added to the routine in response to the affirmative user interface input received in response to the recommendation.
[0008] This scenario eliminates the need for the user to alternatively perform other resource-intensive persistent interactions to have the action performed. Instead, the action is performed as one of the multiple actions of the automated assistant routine to which it is added. This scenario can result in improved human-computer interaction, for example, by reducing the number of inputs that need to be provided by the user (e.g., to an automated assistant interface) to cause the execution of the action. In addition to improving human-computer interaction, this scenario can directly result in various computer and / or network efficiencies as compared to, for example, alternatively requiring the user to provide another user interface input to have the action performed. Further, as described above and elsewhere herein, the initialization of the automated assistant routine can be responsive to abbreviated user interface inputs, which can be efficiently processed and / or can be automatic after the occurrence of conditions that do not require the processing of user interface inputs. Still further, recommending an automated assistant action for inclusion in an existing routine and automatically adding the action to the routine in response to an affirmative input of a single touch and / or single utterance (such as "yes, add") provides an effective addition to the existing routine. For example, automatically adding an action in response to a simplified user interface input can be more computationally efficient than requiring the user to alternatively engage in a long interaction to manually add the action (such as a long interaction that requires the user to manually select the routine to be modified, specify the action to be added to the routine, confirm that the action should be added to the routine, etc.).
[0009] The embodiments disclosed herein can identify assistant actions potentially recommended for inclusion in one or more assistant routines of a user. Additionally, the assistant actions can be compared to each of a plurality of assistant routines stored in association with the user, and based on the comparison (e.g., if one or more criteria are met), a subset of the assistant routines (e.g., one of the assistant routines) can be selected. Graphical and / or audio user interface output can be rendered via the user's client device, where the user interface output prompts the user to add the action to the subset of selected routines. If an affirmative user interface input is received in response to the user interface output, the assistant action can be automatically added (e.g., without the need for another user interface input) to one of the selected routines indicated by the affirmative user interface input. Thereafter, when one of the selected routines is initiated, automatic execution of the pre-existing assistant actions of the routine and automatic execution of the added selected action of the selected actions can be initialized.
[0010] Various techniques can be used to determine assistant actions for potential recommendation for inclusion in one or more assistant routines of a user. In some embodiments, determining an assistant action is based on the assistant action being initiated by one or more instances of user interface input provided by the user. For example, the "adjust the smart thermostat to 72°" assistant action can be determined based on one or more previous instances of the user initiating the assistant action via speech and / or other user interface input provided to the assistant. In some embodiments, determining an assistant action is based on different but related assistant actions being initiated by one or more instances of user interface input provided by the user. For example, the "grilling tips" assistant action that renders different grilling-related tips each time it is executed can be determined based on one or more past instances of the user providing user interface input related to grilling (e.g., "how long do I cook chicken on the grill"). In some embodiments, determining an assistant action is based on determining that the assistant action is related to one or more non-assistant interactions of the user. For example, the "adjust the smart thermostat to 72°" assistant action can be determined based on detecting one or more occurrences of the user manually adjusting the smart thermostat (e.g., via direct interaction with the smart thermostat) or adjusting the smart thermostat via a non-assistant application that involves controlling the smart thermostat. For example, the detection can be based on a state change indication that is transmitted to the assistant and indicates that the smart thermostat has been adjusted to 72°. The state change indication can be transmitted to the assistant by an agent controlling the smart thermostat and can be pushed or provided in response to a status request from the assistant. In some embodiments, determining an assistant action is based on determining that the action can improve human-assistant interaction without having to determine that the user has executed the assistant action and / or related actions.
[0011] Selecting a subset of (multiple) auto - assistant routines for a user to whom an auto - assistant action is recommended to be added based on comparing the action with the (multiple) routines may increase the likelihood that the (multiple) correct routines (if any) are recommended for the action, thereby reducing the risk of providing recommendations that will be ignored or cancelled. Additionally, instead of rendering the entire set of the user's routines, rendering a subset of the auto - assistant routines can save resources utilized during rendering. In some embodiments, the subset of the auto - assistant routines is selected based on comparing the time - property of the past occurrences of the routines with the time - property of the (multiple) past occurrences of the auto - assistant action (or an action related to the auto - assistant action). For example, if the past occurrences of the auto - assistant action for the user occurred only on weekday mornings, then the user's routines that are frequently initiated on weekday mornings are more likely to be selected for the subset than the user's routines that are frequently initiated on weekday evenings or the user's routines that are frequently initiated on weekend mornings.
[0012] In some embodiments, the subset of the auto - assistant routines is selected based on additionally or alternatively comparing the (multiple) device - topology properties of the (multiple) devices utilized in the routine with the (multiple) device - topology properties of the (multiple) devices utilized in the auto - assistant action. For example, if the auto - assistant action involves controlling a television with a device - topology property assignment of "location = living room", then a routine that includes actions for controlling (multiple) other devices (such as living - room lights, living - room speaker lights) with a device - topology property assignment of "location = livingroom" is more likely to be selected for the subset than a routine that lacks any action for controlling a device with a "location = living room" device - topology property. As another example, if the auto - assistant action requires displaying content via a device with a device - topology property assignment of "capable of displaying", then a routine that includes actions for controlling (multiple) devices located in the area of the device with a "capable of displaying" device - topology property is more likely to be selected for the subset than a routine that lacks any action for controlling devices in the area of the device with a "capable of displaying" device - topology property.
[0013] In some embodiments, a subset of the automated assistant routines is selected based on additionally or alternatively excluding any routines that include a contradiction to an automated assistant action from the subset. For example, for the automated assistant action of "adjust the smart thermostat to 72°", a given routine may be excluded based on its having a conflicting action of "adjust the smart thermostat to 75°". Additionally, for example, for an automated assistant action of docking with an agent to perform a service that is only available at night, a given routine may be excluded based on its occurring automatically in the morning rather than at night.
[0014] In some embodiments, multiple competing actions for addition to an automated assistant routine may be considered, and a subset of the competing actions (e.g., one of the competing actions) may be selected for outputting an actual recommendation via a user interface for addition to the routine. In some of those embodiments, each of the multiple competing actions may be compared to the routine, and the selected action may be the one of the multiple competing actions that most closely conforms to the routine, lacks any conflict with the routine, and / or meets other criteria. In some additional or alternative embodiments, the multiple competing actions may include a first action performed by a first agent and a second action performed by a second agent. For example, the first action may cause a first agent that streams music to render jazz music, while the second action may cause jazz music to be rendered by a second agent that streams music. In some of those embodiments, one of the first agent and the second agent may be selected based on one or more measurements associated with the first agent and the second agent. Such measurements may include, for example, the frequency of use of the first agent and / or the second agent by the user; the frequency of use of the first agent and / or the second agent by a user group; the ratings assigned to the first agent and / or the second agent (by the user and / or the user group); and / or data provided by or on behalf of the first agent and / or the second agent that indicates the ability and / or desirability of the corresponding agent(s) to perform the action.
[0015] In some embodiments, a user interface output that prompts the user to add an auto-assistant action to at least one selected auto-assistant routine may be provided at the completion of the execution of the selected routine or at the completion of the execution of the auto-assistant action. For example, after executing the "Good Morning" routine, the auto-assistant may provide a graphical and / or audible prompt "by the way, it looks like Action X would be good to add to this routine, want me to add it", and if a positive input is received in response, "Action X" may be added to the "Good Morning" routine. Additionally, for example, after executing "Action X" in response to a user interface input from the user (such as a speech input of "Assistant, perform Action X"), the auto-assistant may provide a graphical and / or audible prompt of "by the way, it looks like this would be good to add to your Good Morning routine, want me to add it". If a positive input is received in response, "Action X" may be added to the "Good Morning" routine. In these and other ways, the presented user interface output facilitates human-auto-assistant interaction and, if a positive input is received in response, reduces the amount of user input required for both pre-existing actions of a routine to be executed in the future and actions added to the routine. For example, "Action X" may control a connected device, and the presented user interface input may facilitate human-auto-assistant interaction by enabling the user to provide a positive input in response and thus causing the control of the connected device to occur automatically in response to another initiation of the corresponding routine.
[0016] In some embodiments, when adding an assistant action to a routine, the execution location of the assistant action among the pre-existing actions of the routine can be determined. In some versions of those embodiments, the location can be determined based on the duration required for the assistant action of the user interface output (if any) and the duration(s) required for the pre-existing actions of the routine of the user interface output (if any). For example, if the assistant action is controlling a smart device and does not require a user interface output (or a very brief "device controlled" output), the assistant action can be positioned such that it is executed before one or more actions that require rendering a more persistent user interface output. In some additional or alternative versions, the location of execution can be last and can optionally occur with a time delay. The time delay can be based, for example, on the time delay(s) of the past execution of the assistant action (in response to the user's past input(s)) relative to the past execution of the routine. For example, assume that the assistant action is starting a brewing cycle for a smart coffee maker and is added to the "Good Morning" routine. Further assume that, after the completion of the Good Morning routine, the assistant action was previously executed three times (before being added to the "Good Morning" routine) in response to the spoken utterance "Assistant, brew my coffee". If the three spoken utterances occurred two minutes, three minutes, and four minutes after the completion of the Good Morning routine, the assistant action can be added to the Good Morning routine with a time delay (such as a three-minute (average of the three delays) time delay).
[0017] As described above, some automated assistant routines can be initialized in response to detecting a shortcut phrase or command in a spoken or typed natural language input of a user. The shortcut commands provide abbreviated commands for causing the automated assistant to optionally perform a set of actions in a particular order. Providing a shortcut command to cause the performance of a set of actions instead of a longer command for the set of actions can enable fewer user inputs to be provided (and transmitted and / or processed), thereby saving computing and network resources. Additionally, thereafter, automatically adding an automated assistant action as an additional action of a routine in response to an existing shortcut command enables the additional action of the routine and the previously existing actions to be performed in response to the shortcut command. Thus, thereafter, the user can speak or type the shortcut command instead of speaking both the command for the additional action and the shortcut command. As an example of an abbreviated command for an automated assistant routine, when a user wakes up in the morning, the user can trigger the "Good Morning" routine by providing a spoken utterance to a kitchen assistant device (i.e., a client computing device located in the kitchen). The spoken utterance can be, for example, "Good Morning", which can be processed by the assistant device and / or a remote assistant device (communicating with the assistant device) for initializing the "Good Morning" routine. For example, the assistant device and / or the remote device can process audio data corresponding to the spoken utterance to convert the spoken utterance to text, and can also determine that the text "Good Morning" is an automated assistant action set assigned by the user to be performed in response to the spoken utterance of "Good Morning". Adding a new action to the "Good Morning" routine according to the embodiments described herein also assigns the new action to the user for the text "Good Morning".
[0018] While various automated assistant routines can be initialized in response to spoken or typed shortcut commands, in some embodiments, automated assistant routines can additionally or alternatively be initialized in response to a user pressing a virtual or hardware element at a client device or peripheral device, performing a gesture detected via a client device's sensor(s), providing other tactile input(s) at the client device, and / or providing any other type of computer-readable user interface input. For example, a graphical user interface (GUI) can be presented at the client device with a selectable icon that provides the user with a suggestion to initialize an automated assistant routine. When the user selects a selectable icon (e.g., a GUI button that says "Good Morning"), the automated assistant can, in response, initialize the corresponding automated assistant routine. Additionally or alternatively, for example, an automated assistant routine can be automatically initialized in response to the automated assistant detecting the presence of a user (e.g., using voice authentication and / or facial recognition to detect a particular user), the location of the user (e.g., using voice authentication and / or facial recognition to detect a particular user in a location such as at home, in a car, in a living room, and / or additional location(s)), disarming an alarm (e.g., a wake-up alarm set on an associated phone or other device), opening an app, and / or other user actions that can be recognized by the automated assistant (e.g., based on signals from one or more client devices). For example, a "Good Morning" routine can be initialized when a user is detected using facial recognition in the kitchen at "the morning" (i.e., at a particular time, during a time range, and / or additional time(s) associated with the routine).
[0019] The example "Good Morning" routine can include actions such as causing the user's schedule to be rendered, causing a particular appliance to be turned on, and causing a podcast to be rendered. Again, by enabling the auto-assistant to respond to shortcut commands, the user does not necessarily need to provide a string of commands in order for the auto-assistant to perform the corresponding actions (e.g., the user does not need to recite the words: "Assistant, read me my schedule, turn on my appliance, and play my podcast"). Instead, the auto-assistant can respond to the shortcut command, and the auto-assistant can process the shortcut command to identify the actions corresponding to the shortcut command. In some embodiments, the routine can be personalized such that a particular shortcut command or other input can cause the routine to execute to cause the auto-assistant to perform a particular set of actions for one user, while the same input can cause the auto-assistant to perform a different set of actions for a different user. For example, a particular user can specifically configure the auto-assistant to perform a first set of actions in response to a shortcut command, and the spouse of the particular user can configure the auto-assistant to perform a second set of actions in response to the same shortcut command. The auto-assistant can distinguish between users providing the shortcut command using one or more sensor inputs and / or one or more determined characteristics such as voice signature, facial recognition, image feed, motion characteristics, and / or other data. Additionally, as described herein, the auto-assistant can distinguish between users when determining the auto-assistant actions to recommend to a given user, determining user-specific auto-assistant routines, and determining when to provide a prompt to add an auto-assistant action to the user's routine in the user interface output (e.g., a user-specific user interface output can be provided only when it is determined that the user is interacting with the auto-assistant).
[0020] When the user provides a shortcut command (such as "Good Morning") to an assistant device (such as a kitchen assistant device), content corresponding to one or more actions of the routine can initially be rendered by the kitchen assistant device since the spoken words pertain to the kitchen assistant device. For example, although other devices recognize the shortcut command, the content can initially be rendered exclusively at the kitchen assistant device (i.e., not simultaneously at any other client device). For example, multiple devices can recognize the shortcut command received at their respective auto-assistant interfaces; however, the device that receives the loudest and / or least distorted shortcut command can be designated as the device at which the auto-assistant routine will be initialized.
[0021] If the user performs an action in the office via the office assistant device, the automated assistant can still suggest adding an action to the "Good Morning" routine. For example, if the user moves to their office after completing their "Good Morning" routine in the kitchen and then asks the office assistant device for a spoken utterance regarding news headlines, the automated assistant can suggest adding an action to render news headlines to the "Good Morning" routine being performed in the kitchen via the kitchen assistant device. In some embodiments, the office assistant device can suggest adding an action performed in the office (i.e., rendering news headlines) to the "Good Morning" routine even if the action is performed on a different device from the "Good Morning" routine. In some embodiments, when the user is in the office, after the user asks the office assistant for news headlines, the office device can be used to suggest adding an action to render news headlines to the "Good Morning" routine. In other embodiments, the kitchen assistant device can suggest adding a new action to render news headlines to the "Good Morning" routine after the "Good Morning" routine is completed while the user is still near the kitchen device. Additionally or alternatively, in some embodiments, when the user is detected near a kitchen appliance before executing the "Good Morning" routine but after the user requests the routine through a shortcut command for the kitchen appliance, the automated assistant can suggest that the user add actions that are historically performed in the office.
[0022] In some embodiments, restrictions can be imposed on adding actions to an auto - assistant routine. For example, a user living with a spouse may have a "Good Morning" routine, while the spouse may have their own "Good Morning" routine. The user may not want to add actions to the spouse's good - morning routine (and vice versa). In some embodiments, the auto - assistant can determine whether the user or their spouse is performing an action through sensor inputs and / or one or more determined characteristics such as voice signature, face recognition, image feed, motion characteristics, and / or other data. In some embodiments, the default setting on an auto - assistant routine in a multi - user household can be that only the user of the routine has the right to modify their personal routine. In some embodiments, the user can give permission to the auto - assistant to allow other users (such as their spouse) to add actions to their personal routine. Additionally, in some embodiments, if appropriate consent is obtained, the auto - assistant can check for actions for other members of the household with a similar routine before suggesting a routine for adding an action. For example, before suggesting adding an action to turn on a device (such as a smart coffee maker) to the user's good - morning routine, the auto - assistant can check the spouse's good - morning routine and determine whether the action has already been performed in the household in the morning. In other words, when determining the routine for which an action is to be added, the auto - assistant can check the user's routine and compare it with the routines associated with other members of the household to, for example, prevent duplicate actions (such as both users turning on the coffee maker in the morning) as part of their "Good Morning" routine.
[0023] In some embodiments, when executing a routine, the auto - assistant docks with one or more local and / or remote agents. For example, for a routine that includes three actions, the auto - assistant can dock with a first agent when executing the first action, a second agent when executing the second action, and a third agent when executing the third action. As used herein, "agent" refers to one or more computing devices and / or software utilized by the auto - assistant. In some cases, the agent can be separate from the auto - assistant and / or can communicate with the auto - assistant through one or more communication channels. In some of these cases, the auto - assistant can transmit data (such as agent commands) from a first network node to a second network node that implements all or various aspects of the agent's functionality. In some cases, the agent can be a third - party (3P) agent because it is managed by a party separate from the party that manages the auto - assistant. In some other cases, the agent can be a first - party (1P) agent because it is managed by the same party that manages the auto - assistant.
[0024] The agent is configured to receive call requests and / or other agent commands from the automated assistant (e.g., via a network and / or via an API). In response to receiving an agent command, the agent generates response content based on the agent command and transmits the response content for providing a user interface output based on the response content. For example, the agent may transmit the response content to the automated assistant for the automated assistant to provide an output based on the response content. As another example, the agent itself may provide the output. For example, a user may interact with the automated assistant via a client device (e.g., the automated assistant may be implemented on the client device and / or communicate with the client device over a network), and the agent may be an application installed on the client device or an application that is executable away from the client device but "streamable" on the client device. When the application is invoked, it may be executed by the client device and / or brought to the front of the client device (e.g., its content may take over the display of the client device).
[0025] The above is provided as an overview of various embodiments disclosed herein. Additional details regarding those various embodiments and additional embodiments are provided herein.
[0026] In some embodiments, a method implemented by one or more processors is provided and includes determining an action initiated by an automated assistant. The action is initiated by the automated assistant in response to one or more instances of user interface input provided by a user via one or more automated assistant interfaces that interact with the automated assistant. The method also includes identifying a plurality of automated assistant routines stored in association with the user. Each of the automated assistant routines defines a plurality of corresponding actions that will be automatically performed via the automated assistant in response to the initialization of the automated assistant routine, and the action is an action other than the actions of the automated assistant routine. The method also includes comparing the action with the plurality of automated assistant routines stored in association with the user. The method also includes selecting a subset of the automated assistant routines stored in association with the user based on the comparison. The method also includes, based on the action being initiated in response to user interface input provided by the user and based on selecting the subset of automated assistant routines: rendering a user interface output via the user's client device. The user interface output prompts the user to add the action to one or more of the subset of automated assistant routines. The method also includes receiving affirmative user interface input in response to the user interface output, where the affirmative user interface input indicates a given routine of the subset of automated assistant routines. The method also includes automatically adding the action as an additional action to be automatically performed in response to the initialization of the given routine among the plurality of corresponding actions of the given routine in response to receiving the affirmative user interface input.
[0027] These and other embodiments of the techniques disclosed herein may include one or more of the following features.
[0028] In some embodiments, one or more instances of user interface input include one or more spoken utterances that, when converted to text, are first text of a first length. In those embodiments, in response to a given spoken utterance occurring, an initialization of a given routine occurs, the given spoken utterance being second text of a second length when converted to text, and the second length being shorter than the first length.
[0029] In some embodiments, rendering the user interface output via the user's client device is also based on executing a given routine and occurring at the completion of executing the given routine.
[0030] In some embodiments, one or more instances of user interface input that initiate an action include a spoken utterance. In some versions of those embodiments, the method further includes identifying a user profile based on one or more voice characteristics of the spoken utterance. In those versions, identifying a plurality of auto assistant routines stored in association with the user includes identifying the plurality of auto assistant routines based on being stored in association with the user profile. Additionally, in those versions, executing a given routine can occur in response to an additional spoken utterance from the user, and rendering the user interface output via the user's client device can also be based on determining that the additional spoken utterance has one or more voice characteristics corresponding to the user profile.
[0031] In some embodiments, rendering the user interface output via the user's client device occurs at the completion of an action initiated by an auto assistant in response to one or more instances of user interface input.
[0032] In some embodiments, the action includes providing a command to change at least one state of a connected device managed by the auto assistant.
[0033] In some embodiments, the action includes providing a command to an agent controlled by a third party, the third party being different from the party controlling the auto assistant.
[0034] In some embodiments, selecting a subset of auto assistant routines includes: comparing the action with each corresponding action among a plurality of corresponding actions in each auto assistant routine; based on the comparison and for each auto assistant routine, determining whether the action includes one or more contradictions with any of the corresponding actions in the plurality of corresponding actions in the auto assistant routine; and selecting a subset of auto assistant routines based on determining that the plurality of corresponding actions of the subset of auto assistant routines lack one or more contradictions with the action.
[0035] In some embodiments, one or more contradictions include temporal contradictions and / or device-incompatibility contradictions. A temporal contradiction can include an inability to complete an action during the time frame of each corresponding action among a plurality of corresponding actions. A device-incompatibility contradiction can include an action requiring a device with certain capabilities and each corresponding action among a plurality of corresponding actions not being associated with any device having the certain capabilities.
[0036] In some embodiments, automatically adding an action includes determining a location for performing the action within a given routine, where the location is relative to a plurality of corresponding actions of the given routine. In some versions of those embodiments, the location for performing the action is before the plurality of corresponding actions, after the plurality of corresponding actions, or between two corresponding actions among the plurality of corresponding actions. In some of those versions, the location is after the plurality of corresponding actions and includes a time delay. The time delay can be based on one or more past time delays of the user's past initiation of the action after one or more past initiations of the given routine by the user.
[0037] In some embodiments, one or more instances of user interface input for initiating an action are received via an additional client device that is linked to but separate from the client device, and a user interface output is rendered via the client device.
[0038] In some embodiments, the user interface output includes graphical output, and the method further includes selecting the additional client device for providing the user interface output based on determining that the client device lacks display capabilities and the additional client device includes display capabilities.
[0039] In some embodiments, one or more instances of user interface input for initiating an action include speaking a phrase, and the method further includes identifying a profile of the user based on one or more voice characteristics of the spoken phrase. In some of those embodiments, identifying a plurality of auto-assistant routines stored in association with the user is based on the plurality of auto-assistant routines being stored in association with the profile of the user that is identified based on the voice characteristics of the spoken phrase.
[0040] In some embodiments, a method implemented by one or more processors is provided and includes determining an assistant action and identifying a plurality of assistant routines stored in association with a user. Each of the assistant routines defines a plurality of corresponding actions that will be automatically performed via an assistant in response to initialization of the assistant routine. The determined assistant action may be an action other than the plurality of corresponding actions of the assistant routine. The method further includes comparing the assistant action with the plurality of assistant routines stored in association with the user. The method further includes selecting a subset of the assistant routines stored in association with the user based on the comparison. The method further includes determining that a user is interacting with the client device based on sensor data from one or more sensors of the client device. The method further includes rendering, via the client device, a user interface output based on determining that the user is interacting with the client device and based on selecting the subset of assistant routines, wherein the user interface output prompts the user to add the assistant action to one or more of the assistant routines of the subset. The method further includes receiving a positive user interface input in response to the user interface output, wherein the positive user interface input indicates a given routine of the subset of assistant routines. The method further includes, in response to receiving the positive user interface input: adding the assistant action as an additional action among the plurality of corresponding actions of the given routine to be automatically performed in response to initialization of the given routine.
[0041] In some embodiments, the addition is performed in response to receiving the positive user interface input and without any additional user interface input.
[0042] In some embodiments, the sensor data includes voice data based on one or more microphones of the client device, and determining that a user is interacting with the client device includes: determining one or more voice characteristics based on the voice data; and identifying the user based on the voice characteristics matching the user's profile.
[0043] In some embodiments, the assistant action includes controlling a particular connected device. In some versions of those embodiments, comparing the assistant action with the plurality of assistant routines stored in association with the user includes: determining, based on the device topology of the user's connected devices, that a particular connected device and an additional connected device controlled by at least one of the plurality of corresponding actions of a given routine are assigned to the same physical location in the device topology. In some of those versions, selecting a given routine for inclusion in the subset based on the comparison includes selecting a given routine for inclusion in the subset based on determining that the particular connected device and the additional connected device are assigned to the same physical location. The same physical location may be, for example, a semantic identifier assigned by the user to a room in the user's house.
[0044] In some embodiments, the automated assistant actions include controlling a particular connected device, and determining the automated assistant actions includes: determining that the particular connected device has been controlled in a particular manner in response to a non-automated assistant interaction based on the reported state of the particular connected device; and determining that the automated assistant action causes the particular connected device to be controlled in a particular manner.
[0045] In some embodiments, comparing the automated assistant action with a plurality of automated assistant routines stored in association with a user includes: determining at least one action time characteristic of the automated assistant action based on one or more past executions by the user of the automated assistant action or related actions; determining at least one routine time characteristic of the automated assistant routine based on one or more past executions of the automated assistant routine; and comparing the at least one action time characteristic with the at least one routine time characteristic.
[0046] In some embodiments, the automated assistant actions include an interaction with a third-party agent controlled by a particular third party, and determining the automated assistant actions includes determining past executions by the user of a separate interaction with a separate third party.
[0047] In some embodiments, determining the automated assistance action is based on determining the relationship between the separate interaction and the automated assistance action.
[0048] Additionally, some embodiments include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in an associated memory, and wherein the instructions are configured to cause the execution of any one of the foregoing methods. Some embodiments also include one or more non-transitory computer-readable storage media that store computer instructions executable by one or more processors to perform any of the foregoing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a block diagram illustrating an example environment in which various embodiments may be implemented.
[0050] Figure 2 is a diagram illustrating an example interaction between a user and an automated assistant.
[0051] Figure 3 is a diagram illustrating another example interaction between a user and an automated assistant.
[0052] Figure 4 is a flowchart illustrating an example process in accordance with embodiments disclosed herein.
[0053] Figure 5 is a block diagram illustrating an example architecture of a computing device. DETAILED DESCRIPTION
[0054] The embodiments disclosed herein relate to methods, apparatuses, and computer-readable media (transitory and non-transitory) for determining whether to recommend that an automated assistant action be incorporated into a user's automated assistant routine. In some of those embodiments, when it is determined to recommend incorporating an action into a routine, a user interface output is rendered for presentation to the user. The user interface output can prompt the user as to whether he / she wishes to add the action to the routine. Additionally, in some versions of those embodiments, the action is added to the routine in response to a positive user interface input being received in response to the user interface output. For example, the action can be automatically added to the routine in response to a positive user interface input being received.
[0055] An example routine can be a morning routine, where the automated assistant sequentially performs a plurality of different actions in the morning to prepare the user for their day. For example, the morning routine can involve the automated assistant audibly rendering to the user, via a client device, a schedule for a particular day (e.g., the current day), the automated assistant causing a device (e.g., smart lighting) to turn on, and then causing a podcast to be audibly rendered via the client device when the user is ready. After completion of the morning routine, the user can cause an action that can be included in the morning routine to be performed. For example, after the morning routine is completed, the user can provide an uttered statement to an assistant interface of a second client device, where the uttered statement causes the automated assistant to turn on an appliance associated with the automated assistant (e.g., a smart coffee maker). At least partially based on the temporal proximity of the execution of the morning routine and the user providing the uttered statement that causes the automated assistant to turn on the appliance, the "turn on the appliance" action can be recommended to be included in the morning routine. If the user accepts the recommendation, the "turn on the appliance" action can be automatically added to the routine. Thereafter, instead of the user having to give the assistant a separate command to turn on the appliance, the automated assistant can cause the appliance to turn on in response to the initiation of the morning routine.
[0056] Turning now to the drawings, Figure 1 illustrates an example environment 100 in which various embodiments can be implemented. The example environment 100 includes one or more client computing devices 102. Each client device can execute a respective instance of an automated assistant client 112. One or more cloud-based automated assistant components 116, such as a natural language processor 122 and / or a routine module 124, can be implemented on one or more computing systems (collectively referred to as "cloud" computing systems) that are communicatively coupled to the client devices 102 via one or more local area networks and / or wide area networks, typically indicated as 114 (e.g., the Internet).
[0057] In various embodiments, an instance of the Auto-Assistant Client 108 can, through its interaction with one or more cloud-based Auto-Assistant components 116, form something that appears to the user as a logical instance of the Auto-Assistant 112 with which the user can engage in a conversation. In Figure 1 An instance of such an Auto-Assistant 112 is depicted in dashed lines in. Thus, it should be understood that each user who engages with the Auto-Assistant Client 112 executing on the client device 102 can effectively engage with their own logical instance of the Auto-Assistant 112. For simplicity and brevity, the term "Auto-Assistant" used herein to refer to "serving" a particular user can often refer to the combination of the client device 108 operated by the user and one or more cloud-based Auto-Assistant components 116 (which can be shared among multiple Auto-Assistant Clients 108). It should also be understood that in some embodiments, the Auto-Assistant 112 can respond to requests from any user, regardless of whether that particular instance of the Auto-Assistant 112 actually "serves" that user.
[0058] The client device 102 can include, for example, one or more of the following: a desktop computing device, a laptop computing device, a tablet computing device, a touch-sensitive computing device (e.g., a computing device that can receive input via a touch from a user), a mobile phone computing device, a computing device in the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, an in-vehicle navigation system), a standalone interactive speaker, a smart appliance such as a smart TV, and / or a wearable device of the user that includes a computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device). Additional and / or alternative client computing devices can be provided.
[0059] In various embodiments, the client device 102 can include one or more sensors (not depicted) that can take various forms. The sensors can sense different types of input to the Auto-Assistant 112, such as speech, text, graphics, physical (e.g., touch on a display device including a touch-sensitive projector and / or a touch-sensitive screen of a computing device), and / or vision-based (e.g., gesture) input. Some client devices 102 can be equipped with one or more digital cameras configured to capture and provide a signal indicative of movement detected in the field of view. Additionally or alternatively, some client devices can be equipped with sensors that detect sound waves (or pressure waves), such as one or more microphones.
[0060] The client device 102 and / or the cloud-based auto-assistant component 116 can communicate with one or more devices 104. The devices 104 can include any of a variety of devices such as the Internet of Things including, for example, smart electrical, smart thermostats, smart coffee makers, smart lights, smart toilets, and so on. The devices 104 are linked to the client device 102 (and / or a particular user of the client device 102) and to each other. For example, the devices 104 can be linked to a profile assigned to the client device 102 (and optionally other client devices) and / or can be linked to a profile assigned to a user of the client device 102. The client device 102, other client devices, and the devices 104 can together define a coordinated ecosystem of devices. In various embodiments, the devices are linked to each other via a device topology representation that can be created by the user and / or automatically created and that can define various auxiliary client devices, various smart devices, an identifier for each device, and / or an attribute of each device. For example, the identifier of a device can specify the room (e.g., living room, kitchen) (and / or other area) of the structure in which the device is located and / or can specify a nickname and / or an alias of the device (e.g., couch light, front door lock, bedroom speaker, kitchen assistant, etc.). In this way, the identifier of a device can be the name, alias, and / or location of the corresponding device that a user might associate with the corresponding device. Such identifiers can be utilized in the various embodiments disclosed herein. For example, such identifiers can be utilized when determining whether an auto-assistant action conforms to an auto-assistant routine.
[0061] Device 104 can be directly controlled by the automatic assistant 112, and / or device 104 can be controlled by one or more third-party agents 106 hosted by (multiple) remote devices (such as another cloud-based component). In addition, one or more third-party agents 106 can also perform (multiple) functions other than controlling device 104 and / or controlling other hardware devices. For example, the automatic assistant 112 can interact with the third-party agent 106 to cause a service to be performed, a transaction to be initiated, etc. For example, the user voice command "order a large pepperoni pizza from Agent X" can cause the automatic assistant client 108 (or (multiple) cloud-based automatic assistant components 116) to send an agent command to the third-party agent "Agent X". The agent command can include, for example, an intent value indicating the "order" intent determined from the voice command and optional slot values such as "type=pizza", "toppings=pepperoni", and "size=large". In response, the third-party agent can cause an order for a large pepperoni pizza to be initiated and provide content indicating that the order has been successfully initiated to the automatic assistant 112. The content (or its transformation) can then be rendered to the user via the (multiple) output devices (such as (multiple) speakers and / or (multiple) displays) of the client device 102. As described herein, some embodiments can recommend adding the "order a large pepperoni pizza from Agent X" action and / or similar actions to the user's existing routine. When added to the user's existing routine, the "order a large pepperoni pizza from Agent X" automatic assistant action can be initiated in response to the initiation of the existing routine, without the user having to separately provide the spoken utterance "order a large pepperoni pizza from Agent X". Thus, the transmission and / or processing of audio data capturing such spoken utterances can be avoided, saving computer and / or network resources.
[0062] In many embodiments, the automated assistant 112 participates in a conversation session with one or more users via the user interface input and output devices of one or more client devices 102. In some embodiments, in response to user interface input provided by a user via one or more user interface input devices of one of the client devices 102, the automated assistant 112 can participate in a conversation session with the user. In some of those embodiments, the user interface input is explicitly directed directly to the automated assistant 112. For example, the user can say a predetermined invocation phrase, such as "OK, Assistant" or "Hey, Assistant", to cause the automated assistant 112 to start actively listening. In many embodiments, the user can say a predetermined shortcut phrase to start running a routine such as "OK Assistant, Good Morning" to start running the good morning routine.
[0063] In some embodiments, even when the user interface input is not explicitly directed directly to the automated assistant 112, the automated assistant 112 can participate in a conversation session in response to the user interface input. For example, the automated assistant 112 can examine the content of the user interface input in response to certain terms present in the user interface input and / or based on other cues and participate in a conversation session. In many embodiments, the automated assistant 112 can utilize speech recognition to convert the user's utterance into text and respond to the text, for example, by providing visual information, by search results, by providing general information, and / or by taking one or more response actions (e.g., playing media, starting a game, ordering food, etc.). In some embodiments, the automated assistant 112 can additionally or alternatively respond to the utterance without converting the utterance into text. For example, the automated assistant 112 can convert the voice input into an embedding, into an entity representation (indicating one or more entities present in the voice input), and / or other "non-text" representations, and operate on such non-text representations. Thus, embodiments described herein as operating based on text converted from voice input can additionally and / or alternatively operate directly on the voice input and / or other non-text representations of the voice input.
[0064] Each of the client computing devices 102 and the computing device operating the cloud-based automated assistant component 116 can include one or more memories for storing data and software applications, one or more processors for accessing the data and executing the applications, and other components that facilitate communication over a network. Operations performed by one or more of the computing devices 102 and / or by the automated assistant 112 can be distributed across multiple computer systems. The automated assistant 112 can be implemented as, for example, a computer program running on one or more computers at one or more locations coupled to each other via a network.
[0065] As described above, in various embodiments, the client computing device 102 may operate an auto-assistant client 108. In various embodiments, each auto-assistant client 108 may include a corresponding voice capture / text-to-speech (“TTS”) speech-to-text (“STT”) module 110. In other embodiments, one or more aspects of the voice capture / TTS / STT module 110 may be implemented separately from the auto-assistant client 108.
[0066] Each voice capture / TTS / STT module 110 may be configured to perform one or more functions: for example, capture a user's voice via a microphone (which may include sensors in the client device 102 in some cases); convert the captured audio to text (and / or other representations or embeddings); and / or convert text to speech. For example, in some embodiments, because the client device 102 may be relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the voice capture / TTS / STT module 110 local to each client device 102 may be configured to convert a limited number of different spoken phrases - particularly phrases that invoke the auto-assistant 112 - to text (or other forms, such as lower-dimensional embeddings). Other voice inputs may be sent to a cloud-based auto-assistant component 116, which may include a cloud-based TTS module 118 and / or a cloud-based STT module 120.
[0067] The cloud-based STT module 120 may be configured to use the virtually unlimited resources of the cloud to convert the audio data captured by the voice capture / TTS / STT module 110 to text (which may then be provided to the natural language processor 122). The cloud-based TTS module 118 may be configured to use the virtually unlimited resources of the cloud to convert text data (e.g., a natural language response formulated by the auto-assistant 112) to computer-generated voice output. In some embodiments, the TTS module 118 may provide the computer-generated voice output to the client device 102 for direct output using, for example, one or more speakers. In other embodiments, the text data (e.g., natural language response) generated by the auto-assistant 112 may be provided to the voice capture / TTS / STT module 110, which may then convert the text data to computer-generated voice for local output.
[0068] The automated assistant 112 (e.g., a cloud-based assistant component 116) can include a natural language processor 122, the aforementioned TTS module 118, the aforementioned STT module 120, and other components, some of which are described in more detail below. In some embodiments, one or more engines and / or modules of the automated assistant 112 can be omitted, combined, and / or implemented in a component separate from the automated assistant 112. In some embodiments, for privacy protection, at least a portion of one or more components of the automated assistant 112, such as the natural language processor 122, the voice capture / TTS / STT module 110, the routine module 124, etc., can be implemented at least partially on the client device 102 (e.g., excluded from the cloud).
[0069] In some embodiments, the automated assistant 112 generates response content in response to various inputs generated by a user of the client device 102 during a human-machine conversation session with the automated assistant 112. The automated assistant 112 can provide the response content as part of the conversation session to the user (e.g., via one or more networks when separated from the user's client device). For example, the automated assistant 112 can generate response content in response to a free-form natural language input provided via the client device 102. As used herein, a free-form input is formulated by the user and is not limited to a set of options presented to the user for selection.
[0070] The natural language processor 122 of the automated assistant 112 processes the natural language input generated by the user via the client device 102 and can generate an annotated output for use by one or more components of the automated assistant 112. For example, the natural language processor 122 can process a free-form natural language input generated by the user via one or more user interface input devices of the client device 102. The generated annotated output includes one or more annotations of the natural language input and optionally includes one or more (e.g., all) terms of the natural language input.
[0071] In some embodiments, the natural language processor 122 is configured to identify and annotate various types of syntactic information in the natural language input. For example, the natural language processor 122 can include a part-of-speech tagger configured to annotate lexical items with their syntactic roles. Additionally, for example, in some embodiments, the natural language processor 122 can additionally and / or alternatively include a dependency parser (not shown) configured to determine syntactic relationships between lexical items in the natural language input.
[0072] In some embodiments, the natural language processor 122 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references in one or more paragraphs, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), and the like. The entity tagger of the natural language processor 122 may annotate references to entities at a higher granularity level (e.g., enabling identification of all references to entity classes such as people) and / or at a lower granularity level (e.g., enabling identification of all references to specific entities such as a particular person). The entity tagger may rely on the content of the natural language input to resolve specific entities and / or may optionally communicate with a knowledge graph or other entity database to resolve specific entities.
[0073] In some embodiments, the natural language processor 122 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more context clues. For example, a coreference resolver may be utilized to resolve the term "there" in the natural language input "I liked Hypothetical Café last time we ate there." to "Hypothetical Café".
[0074] In many embodiments, one or more components of the natural language processor 122 may rely on annotations from one or more other components of the natural language processor 122. For example, in some embodiments, when annotating all mentions of a specific entity, the named entity tagger may rely on annotations from the coreference resolver and / or the dependency parser. Similarly, for example, in some embodiments, when clustering references to the same entity, the coreference resolver may rely on annotations from the dependency parser. In many embodiments, when processing a particular natural language input, one or more components of the natural language processor 122 may use relevant previous inputs and / or other relevant data outside of the particular natural language input to determine one or more annotations.
[0075] The routine module 124 of the automated assistant 112 can determine candidate actions for addition to a user's automated assistant routine, can determine a subset of the user's available routines that are recommended for adding the candidate actions, can provide a recommendation for adding the action to one of the (multiple) routines in the subset, and can add the action to one of the routines based on an affirmative input indicating one of the routines received in response to the recommendation. The routine module 124 can be implemented by the cloud-based automated assistant component 116 as depicted, and / or by the automated assistant client 108 of the client device 102.
[0076] In some embodiments, the routine module 124 determines automated assistant actions for potentially recommending for addition to a user's routine based on the action being initiated by the user through interaction with one or more instances of the automated assistant 112. In some embodiments, the routine module 124 additionally or alternatively determines automated assistant actions for potentially recommending for addition to a user's routine based on a related but different action being initiated by the user through interaction with one or more instances of the automated assistant 112. In some embodiments, the routine module 124 additionally or alternatively determines automated assistant actions for potentially recommending for addition to a user's routine based on data associated with the user indicating that the user may have performed the action or a related action. Data associated with the user can include data based on the user's interaction with other non-automated assistant applications, data based on the state of smart devices in response to direct user control of those smart devices, data indicating bits and / or transactions, etc.
[0077] In some embodiments, the routine module 124 can access a stored list of routines for the user and determine a subset of the (multiple) routines that are recommended for adding the determined automated assistant action. In some of those embodiments, the routine module 124 selects the subset of automated assistant routines based on comparing the time attribute of the past occurrences of the routine with the time attribute of the (multiple) past occurrences of the automated assistant action (or an action related to the automated assistant action). For example, if the automated assistant action is the "decrease smart thermostat set point 3°" action determined based on past occurrences of the user manually adjusting the smart thermostat, then those past occurrences can be determined based on the time from the state updates of the smart thermostat reflecting those decreases. Additionally, routines of the user that were frequently initiated near those times (e.g., within one hour) are more likely to be selected for the subset than routines of the user that were rarely or never initiated near those times.
[0078] In some of those embodiments, the routine module 124 additionally or alternatively selects a subset of the auto-assistant routines based on comparing the (multiple) device topology attributes of the (multiple) devices to be utilized in the routine with the (multiple) device topology attributes of the (multiple) devices utilized in the auto-assistant action. For example, if the auto-assistant action involves controlling a component of a user's vehicle with an assigned device topology attribute of "location=vehicle", a routine that includes actions for controlling the (multiple) other devices with an assigned device topology attribute of "location=vehicle" may be more likely to be selected for the subset than a routine that lacks any actions for controlling devices with the "location=vehicle" device topology attribute.
[0079] In some of those embodiments, the routine module 124 additionally or alternatively selects a subset of the auto-assistant routines that excludes any routines that include actions conflicting with the auto-assistant action. For example, for an auto-assistant action of "dim the livingroom lights to 50%", a given routine may be excluded based on its having a conflicting action of "turn off the living room lights". Additionally, for example, for an auto-assistant action of docking with an agent to perform a service that results in delivering an item to the user's home on demand, a given routine may be excluded based on the given routine being assigned to the user's "work" location (e.g., explicitly or based on being associated only with linked devices having a "work" device topology attribute).
[0080] In some embodiments, the routine module considers multiple competing actions for addition to the auto-assistant routine, and selects a subset of the competing actions (e.g., one of the competing actions) that can be output via the user interface as an actual recommendation for addition to the routine. In some of those embodiments, each of the multiple competing actions may be compared with the routine, and the (multiple) selected actions may be one or more of the competing actions that most closely conform to the routine, lack any conflicts with the routine, and / or satisfy other criteria.
[0081] Regardless of the technique used to determine the action(s) and / or the routine(s) to which the action(s) can be added, the routine module 124 can cause a user interface output to be added that prompts the user as to whether to add the action(s) to the routine(s). In some embodiments, when recommending adding an action to a routine, the routine module 124 causes the corresponding user interface output to be rendered in response to the user's interaction with the auto assistant to perform the action (e.g., after the action has been completed in response to the user interaction). In some embodiments, when recommending adding an action to a routine, the routine module 124 causes the corresponding user interface output to be rendered in response to the initiation of the routine (e.g., before any action of the routine is performed or after all actions of the routine are completed).
[0082] If an affirmative user interface input is received in response to the recommendation provided by the routine module 124, the routine module 124 can cause the corresponding action to be added to the corresponding routine. For example, the routine module 124 can automatically add the action to the routine without any additional user interface input from the user. In some embodiments, when adding an auto assistant action to a routine, the routine module 124 can determine the execution location of the auto assistant action among the pre-existing actions of the routine. In some versions of those embodiments, the location can be determined based on the duration required for the auto assistant action of the user interface output (if any) and the duration(s) required for the pre-existing actions of the routine of the user interface output (if any). In some additional or alternative versions, the location of execution can be last and can optionally occur with a time delay.
[0083] In Figure 2 FIG. illustrates an example of a user interacting with a client device to add an action to an auto assistant routine. Image 200 includes a scene of a room that includes user 202 and client device 204. User 202 interacts with client device 204 via the spoken utterance indicated in dialog box 206. The spoken utterance in dialog box 206 requests the auto assistant to perform an action that increases the set point of the smart thermostat linked to client device 204 by five degrees. The auto assistant associated with client device 204 can process the spoken utterance to determine the appropriate command to increase the set point of the smart thermostat by five degrees. In some embodiments, the auto assistant associated with client device 204 can directly interface with the smart thermostat to change the temperature. For example, the auto assistant can provide the command directly to the smart thermostat. Additionally or alternatively, the auto assistant can interface with a third-party agent associated with the smart thermostat to change the temperature. For example, the auto assistant can provide the command to the third-party agent, which in turn generates the corresponding command and transmits the corresponding command to the smart thermostat to effect a temperature change.
[0084] In some embodiments, as shown in dialog box 208, the client device 204 may render an audible user interface input confirming that the temperature has increased and may also recommend adding a five-degree temperature increase to the user's "Good Morning" routine. The recommendation may include a prompt asking the user if they want to add the action to the recommended routine. In some embodiments, the recommendation may be made if the user typically increases the temperature in the morning after completing their "Good Morning" routine. In some embodiments, the recommendation may be made even if the user executes the "Good Morning" routine using a first client device and the user changes the temperature using a second client device.
[0085] Before making the recommendation, the auto-assistant associated with the client device 204 may select the "Good Morning" routine from among the user's multiple other routines as the routine to which the "temperature increase" auto-assistant action should be recommended to be added. Various techniques may be used to select the "Good Morning" routine to exclude the user's other routines, such as the various techniques described in detail herein.
[0086] The user may decide whether they want to add the action to the recommended routine and respond to the prompt via another spoken utterance indicated by dialog box 210. If the user responds affirmatively to the prompt, as indicated by the spoken utterance of dialog box 210, the auto-assistant automatically adds the action to the user's "Good Morning" routine. Additionally, the auto-assistant causes another audible user interface output to be provided via the client device 204, as indicated by dialog box 212, where the other audible output confirms that the five-degree temperature increase has been added to the "Good Morning" routine. However, in some embodiments, the confirmation may be optional. Additionally or alternatively, the client device 204 may confirm that the action has been added to the routine the next time the user runs the routine. In various embodiments where the user has given permission for example to their spouse to add the action to the user's routine, this may provide the user with additional confirmation that they want to add the action to a particular routine.
[0087] In some embodiments, user 202 may provide input to client device 204 to cause the client device to interact with one or more third-party agents to effect an action that is not an action to control the smart device. For example, instead of requesting a temperature change, the user may use client device 204 to place an online order from a third-party agent associated with "Local Coffee Shop", and then pick up their ordered coffee on their way to work. The user may interact with an auto-assistant associated with client device 204 when placing such an order, or with a separate application of client device 204. In many embodiments, an auto-assistant associated with client 204 (such as the one described above in Figure 1The routine module 124) described therein can determine a coffee ordering action and determine whether to recommend adding the action to the user's "Good Morning" routine and / or one or more other auto-assistant routines. For example, if the user orders coffee from the "Local Coffee Shop" using the client device 204 immediately after the "Good Morning" routine is completed every morning, the auto-assistant action can determine to recommend adding the action to the "Good Morning" routine at least in part based on the temporal proximity of the action to the "Good Morning" routine. In some implementations, the user can order coffee from the "Local Coffee Shop" using the client device 204, for example, with an average delay of 25 minutes after the "Good Morning" routine is completed on the client device 204 (i.e., the user places the order when they leave home, which occurs on average 25 minutes after the "Good Morning" routine is completed), so that there is a hot coffee waiting for them when they arrive at the "Local Coffee Shop". In some of those implementations, the coffee order from the "Local Coffee Shop" can still be added to the "Good Morning" routine so that it is automatically executed in response to the execution of the "Good Morning" routine and together with the other pre-existing actions of the "Good Morning" routine. However, the auto-assistant can optionally include a delay (e.g., 25 minutes after the last action in the pre-existing "Good Morning" routine is completed) before ordering coffee from the "Local Coffee Shop" for the user via a third-party agent in the "Good Morning" routine. Additionally or alternatively, the auto-assistant can add it to the "Good Morning" routine in such a way that a prompt (e.g., "ready to order your coffee") is provided after a delay after the last action in the pre-existing "Good Morning" routine is completed, and the order is only initiated if an affirmative input response is received in response to the prompt.
[0088] Additionally or alternatively, in some embodiments, if the user only makes one coffee order from "LocalCoffee Shop" using the client device 204, the automated assistant can check the actions in the "Good Morning" routine. For example, if the routine includes turning on the coffee machine, it is possible that the user only ran out of coffee that morning, and the automated assistant should not recommend adding the coffee order to the "Good Morning" routine. However, in some embodiments, if the user frequently visits "Local Coffee Shop" in the morning even without using the client device to pre-order coffee, the automated assistant can decide to recommend adding the action of making a coffee order to the "Good Morning" routine.
[0089] Furthermore, even if the user has not previously interacted with a third party, the automated assistant can suggest adding an automated assistant action to a routine, and the action will be performed via that third party. For example, the user may not be aware of the new coffee shop "New CoffeeShop", which opened last month and is conveniently located on the route between the user's home and work. The action of having coffee ordered at "New Coffee Shop" can be recommended for addition to the user's "Good Morning" routine, which currently lacks any coffee-ordering actions. Additionally or alternatively, in some embodiments, the action of having coffee ordered at "New Coffee Shop" can be recommended for addition to an existing routine and replacing an existing action of having coffee ordered at another coffee shop in the routine.
[0090] In some embodiments, more than one action for potentially adding to a routine can be presented to a user, where each of the actions is associated with a different third-party agent. For example, a user who makes a daily morning coffee order from "Local Coffee Shop" via client device 204 can be presented with a recommendation to add any one of three separate actions to an existing "Good Morning" routine. Each of the three actions can cause a coffee order to be initiated, but the first action can cause it to be ordered from "Local Coffee Shop", the second action can cause it to be ordered from "New Coffee Shop", and the third action can cause it to be ordered from "Classic Coffee Shop". In some embodiments, corresponding additional information can be provided for each of the actions, such as: the amount of time by which accessing the corresponding coffee shop will increase and / or decrease the user's commute, the corresponding price of the coffee order, the corresponding rewards program and / or other information, to help the user make an informed decision about which coffee shop to add to their "Good Morning" routine. In some embodiments, one or more of the recommended actions can be recommended based on the corresponding party providing a monetary consideration for its inclusion. In some embodiments, the amount of the required monetary consideration can be related to how well the action matches the routine to which the action is proposed to be added, how frequently the user performs the routine, and / or other criteria.
[0091] Although the spoken utterances of user 202 and the audible output of client device 204 are illustrated in Figure 2 , additional or alternative user input can be provided by user 202 and / or by content rendered by client device 204 via additional or alternative modalities. For example, user 202 can additionally or alternatively provide typed input, touch input, gesture input, etc. Additionally, for example, the output can be additionally or alternatively rendered graphically by client device 204. As a specific instance, client device 204 can graphically render a prompt for adding an action to a routine, and optional interface elements that can be selected via a touch interaction by user 202 to cause the action to be automatically added to the routine.
[0092] In Figure 3Additional examples are illustrated in which a user interacts with a client device to add an action to an automated assistant routine. Image 300 includes a scene of a room that includes user 302 and client device 304. User 302 may interact with client device 304 to start a routine by providing an uttered speech that includes a predefined shortcut phrase. For example, in dialog box 306, user 302 starts the "Good Morning" routine by providing an uttered speech that is detected by client device 304 and includes the phrase "Good Morning". The automated assistant associated with client device 304 may process the audio data capturing the uttered speech, determine that the uttered speech includes the shortcut phrase for the "Good Morning" routine, and in response, cause each action in the "Good Morning" routine to be executed. In some embodiments, the automated assistant may ask user 302, via client device 304, whether the user would like to add an action to the just-completed routine after the routine is completed. For example, dialog box 308 indicates that client device 304 has executed the user's "Good Morning" routine and then asks the user whether the user would like to add lighting control to their "Good Morning" routine. Additionally or alternatively, in some embodiments, client device 304 may ask the user whether the user would like to add lighting control to their "Good Morning" routine before running the routine and including the new action as part of the running routine.
[0093] In many embodiments, the automated assistant associated with client device 304 may determine the recommendation to add a new action to the just-completed routine (in this case, the "Good Morning" routine) based on the user's historical actions. For example, the recommendation may be determined based on the user always using client device 304 to turn on their smart lights after the "Good Morning" routine is completed. Additionally or alternatively, in some embodiments, the user may have smart lights in their house but be unaware of the ability to control the smart lights using the automated assistant and thus have not previously used the automated assistant to control the smart lights. In such an example, the automated assistant may make a recommendation based on determining that the user has not performed a lighting control action in an attempt to add functionality to the user's routine that the user may otherwise be unaware of. Additionally or alternatively, in some embodiments, the automated assistant may make a recommendation based on determining that control of the smart lights is an action that other users commonly have in a related routine. For example, if more than 90% of the users who have the devices required for lighting control (as determined via the user's device topology) also have a lighting control action in their "Good Morning" routine, the assistant may make a recommendation to add lighting control to the "Good Morning" routine.
[0094] The user can decide whether to add the suggested action to the routine. For example, dialog box 310 indicates that the user affirmatively decides to add the suggested action (i.e., lighting control) to their "Good Morning" routine. However, the user does not need to add the recommended action to the suggested routine. In some embodiments, the client device 304 associated with the auto-assistant can confirm that the recommended action has been added to the routine. For example, dialog box 312 indicates that the lighting control has been added to the user's "Good Morning" routine. However, in various embodiments, the client device 304 can confirm to the user that the action has been added to a specific routine the next time the user runs the routine. In some embodiments, this can provide additional backup to confirm that the user still wants the new action in the routine and has not changed their mind the next time the routine runs. Additionally, in some embodiments, the user can give permission for others to modify their routines. In some such embodiments, confirming that the new action has been added before the user runs the routine can allow the user to confirm actions added to the user's routine by a different user (such as the user's spouse).
[0095] Figure 4 A process for determining whether to add an action to an auto-assistant routine is illustrated in accordance with various embodiments. Process 400 can be performed by one or more client devices and / or any other device capable of interacting with the auto-assistant. The process includes determining (402) an auto-assistant action. In some embodiments, the action can be determined based on having been performed one or more times in response to the user's interaction with the auto-assistant. In some embodiments, the action can include controlling a physical device (e.g., controlling a smart thermostat, turning on a smart coffee maker, controlling a smart speaker, etc.) and / or interacting with a remote agent (e.g., a calendar agent that can provide daily agents for the user, a weather agent that can provide the day's weather for the user, etc.). In some embodiments, the action can be determined based on having been performed by the user a plurality of times historically. For example, an action for controlling certain smart lights can be determined based on the user controlling certain smart lights every night before going to bed through interaction with the auto-assistant, through manual interaction, and / or through interaction with an application dedicated to controlling the smart lights. The action can be determined based on having been performed at least a threshold number of times. In some embodiments, an auto-assistant action can be determined even if the user has never caused the action to be performed. For example, an action related to controlling a smart lock can be determined in response to determining that the smart lock has just been installed by the user and / or added to the topology of link devices that can be controlled by the auto-assistant. For example, the user may have just installed a new smart lock, but the smart lock is not associated with any routine.
[0096] The automated assistant can identify (404) the user's routines. The user can have many routines, some of which can overlap. For example, the user can have a first "Good Morning" routine during the work week and a second "Good Morning" routine on the weekend. Additionally or alternatively, the user can have several routines to choose from when coming home from work. For example, the user can have a "Veg Out" routine and a "Read a Book" routine. In some implementations, the user's profile can be identified, and the set of routines assigned to the user's profile can be the identified set of routines associated with that user. In some of those implementations, the set of routines for the user's profile is identified for the action determined at (402) based on the determination that the action at (402) is being performed for the user's profile. As described above with respect to Figure 1 Any one of the various sensors described above can be used to identify the user profile associated with an input to the automated assistant (e.g., voice matching based on audio data) and the automated assistant routines associated with that user profile. In various implementations, the set of routines can be shared among members of the same household. For example, a user and their spouse can enjoy cooking together and have a shared "Grilling Night" routine, where both users interact with one or more automated assistant clients and one or more client devices. For example, the Grilling Night routine can include actions such as: turning on the lights in the kitchen, turning on the outdoor grill to start heating it, and rendering speakers in both the kitchen and the outdoor grilling area to play the same music. Both users can interact with the devices in the kitchen and at the grill at various points in time. In some implementations, since the routine is a shared routine, the "Grilling Night" routine can be included in the user profiles for both users, and should be considered for both users when potentially adding new actions to the new routine. Allowing one user to add an action to a shared routine can also share computing resources, as both users will not have to manually go through the process of adding a new action to the shared routine.
[0097] The set of automated assistant routines can be selected (406) from the set of actions identified at (404), potentially adding the action identified at (402) to the set of automated assistant routines. For example, if the user is performing an action when arriving home from work, the "Veg Out" routine and the "Read a Book" routine are potential automated assistant routines to add to the action. However, the automated assistant can compare the action with those routine(s) and / or other routines to select a set of automated assistant routines that is a subset of all available automated assistant routines. For example, the "Veg Out" routine can include turning on the living room TV and turning off the living room lights. In contrast, the "Read a Book" routine can include setting customized lighting for book reading in the study and playing soft jazz in the study. If the user action is to request information related to the current TV program being played that night, then in many implementations, the automated assistant will select the routine that includes the TV for inclusion in the set that potentially adds the action. Additionally, for example, using the device topology of the linked device for the user, the automated assistant can determine that there is no device in the study that is capable of displaying, the study being the room associated with all actions for the "Book Reading" routine. Thus, since the action involves pushing information to a device capable of displaying and the "Book Reading" routine is currently only associated with a room that has no device capable of displaying, the automated assistant can exclude the "Book Reading" routine from the set that potentially adds the action. In some implementations, the automated assistant can analyze the action and compare the action with each action in each routine for the user and / or location to generate a set of routines that potentially add the action.
[0098] In some embodiments, potential actions can be compared to a number of routines before a routine (or set of routines) is selected to potentially add an action. For example, a user may have four routines: "Party Time", "Good Morning", "Good Night", and "Drive Home". The "Party Time" routine may include the following actions: rendering party music to speakers in the living room, adjusting the networked lights in the living room to flash in a random color pattern, and ordering pizza for delivery from the "Local Pizza Shop" via a third-party agent. The "Good Morning" routine may include the following actions: rendering the user's schedule for the day, turning on the networked lights in the bedroom, rendering the weather for the day for the user, and rendering news headlines for the user. The user's "Good Night" routine may include turning off all networked lights in the house, rendering the user's schedule for tomorrow, rendering the weather for tomorrow for the user, and rendering two hours of white noise for the user. Additionally, the user's "Drive Home" routine may include the following actions: rendering the user's evening schedule, notifying the user's spouse on another client device that the user is driving home, rendering information about a set of stock prices determined by the user, and rendering a podcast for the remainder of the drive.
[0099] As an example, the automated assistant can examine four routines of a user and select which routine to suggest adding an action of ordering coffee from "Local Coffee Shop". In various implementations, the automated assistant can analyze various information of each routine, including: actions included in the routine compared to the new action, historical relationship (if any) between the new action and the routine, temporal relationship (if any) between the new action and the routine, whether other users typically have actions in a similar routine, etc. With respect to the "Party Time" routine, the automated assistant can compare the new action of ordering coffee with the routine. Other users may typically order coffee before a party (i.e., so they are awake during the party). "Party Time" has an inherent time constraint and needs to run when "Local Pizza Shop" is open. The user may historically order coffee around the same time they run the "Party Time" routine (i.e., the user orders afternoon coffee while working from home, but the user also has a party around the same afternoon), so historical data on when the user runs the "Party Time" routine can indicate this should be a suggestion to the user. However, the action of ordering pizza from "Local Pizza Shop" for delivery conflicts with leaving the house to pick up a cup of coffee (i.e., if the user has left the house to pick up the coffee order, the user may not be home when the pizza is delivered). Thus, the automated assistant generally does not recommend adding the action of ordering a cup of coffee from "Local Coffee Shop" to the "Party Time" routine. Additionally or alternatively, in some implementations, the automated assistant can look at the user's historical actions after running the "Party Time" routine.
[0100] Similarly, the automated assistant can analyze the "Good Morning" routine compared to the new coffee ordering action. Other users typically order coffee in a similar routine. During the time the user runs their "Good Morning" routine, "Local CoffeeShop" is likely to be open. Additionally, there are no pre-existing actions in the "Good Morning" routine that conflict with ordering coffee (e.g., no coffee machine is turned on that is included in the "Good Morning" routine). Further, the action of the user ordering coffee in the morning around the same time the user runs the "GoodMorning" routine can all be indicators to suggest adding the new action of ordering coffee from "Local CoffeeShop" to the "Good Morning" routine.
[0101] Conversely, when "Local Coffee Shop" closes, the "Good Night" routine can be executed, which can indicate to the assistant that adding an action is not recommended. However, in some embodiments, if the user's preferred coffee shop closes, the automated assistant can recommend another coffee shop to order coffee from. Most users do not include ordering coffee in a similar night routine, which can indicate to the automated assistant not to recommend that action. Additionally, the actions of the user in the "Good Night" routine, such as turning off the networked lighting throughout the house, followed by the action of playing white noise for two hours, are contradictory to leaving the house to get a cup of coffee. The action of turning off the lights in the house itself is not contradictory to getting a cup of coffee, but that action paired with other actions in the "Good Night" routine is contradictory to leaving the house to get coffee.
[0102] In addition, the "Drive Home" routine can be compared with the new coffee ordering action. The user is already driving, so there is no contradiction between any actions in the "Drive Home" routine and going to get a cup of coffee. However, the automated assistant typically checks many factors before suggesting adding a new action to a routine. The automated assistant can compare with similar routines of other users and see that most users do not order coffee when driving home from work at night. This situation can indicate to the automated assistant not to recommend adding an action to the "Drive Home" routine. Additionally, the historical actions of the user ordering coffee can be analyzed, and the automated assistant can see that the user has not ordered coffee since 3 pm, and the user typically starts the "Drive Home" routine between 5 pm and 6 pm. The time when the user executes the "Drive Home" routine is contradictory to the historical data regarding when the user orders coffee from the "Local Coffee Shop". This situation can also indicate to the automated assistant not to recommend adding the new action of ordering coffee to the "Drive Home" routine.
[0103] Therefore, after analyzing all routines for the user, the automated assistant in this example should select the "Good Morning" routine to recommend adding the new action of ordering a cup of coffee from the "Local Coffee Shop".
[0104] The automated assistant can render a user interface output (408) that prompts whether an action should be added to any of the routines in a potential set of automated assistant routines. In some embodiments, the user can add the same action to several routines. In some embodiments, the user may want to add an action to a single routine. Additionally, in some embodiments, the user can execute something similar to what is described above Figure 2After the action, it is asked whether the action should be added to the routine. Additionally or alternatively, it can be asked whether the action should be added to the routine after running a specific routine similar to the one described above Figure 3 after the specific routine.
[0105] If an affirmative user interface input is received (410) in response to the output at (408), the automatic assistant can add the action to the corresponding automatic assistant routine. For example, if the output at (408) prompts whether the action should be added to a specific routine, a spoken utterance of "yes" can cause the action to be added to the specific routine. Additionally, for example, if the output at (408) prompts whether the action should be added to the first specific routine or the second specific routine, a spoken utterance of "the first one" can cause the action to be added to the first specific routine rather than the second specific routine; a spoken utterance of "the second one" can cause the action to be added to the second specific routine rather than the first specific routine; and a spoken utterance of "both" can cause the action to be added to both the first and second specific routines. In many implementations, the automatic assistant can prompt the user for the name of the specific routine, and the user can add the action to the routine by speaking the name of the specific routine. For example, if the output (408) prompts the user to suggest adding a routine to the Good Morning routine and the Drive to Work routine, a spoken utterance of the name of the routine (e.g., "Drive to Work") can cause the action to be added to the "Drive to Work" routine. In some implementations, the automatic assistant can automatically add the action to the routine and can notify the user once the action has been added to the routine. In some implementations, similar to the discussion above Figure 2 the automatic assistant can determine where to add the new action in the action sequence within the pre-existing routine.
[0106] Figure 5 is a block diagram of an example computing device 510. The computing device 510 generally includes at least one processor 514 that communicates with a number of peripheral devices via a bus subsystem 512. These peripheral devices can include a storage subsystem 524 (including, for example, a memory 525 and a file storage subsystem 526), a user interface input device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices allow the user to interact with the computer system 510. The network interface subsystem 516 provides an interface to an external network and is coupled to a corresponding interface device in other computer systems.
[0107] The user interface input device 522 may include a keyboard, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touch screen incorporated into a display, audio input devices such as a voice recognition system, microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways to input information into the computer system 510 or onto a communication network.
[0108] The user interface output device 520 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, for example, via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways to output information from the computer system 510 to a user or to another machine or computer system.
[0109] The storage subsystem 524 stores programming and data constructs that provide some or all of the functionality described herein. For example, the storage subsystem 524 may include Figure 1 selected aspects and / or processes, any operations discussed herein, and / or implement the logic for one or more of the cloud-based automated component 116, automated assistant 112, client device 102, and / or any other devices or applications discussed herein.
[0110] These software modules are typically executed by the processor 514, alone or in combination with other processors. The memory 525 used in the storage subsystem 524 may include a number of memories, including a main random access memory (RAM) 530 for storing instructions and data during program execution and a read-only memory (ROM) 532 storing fixed instructions. The file storage subsystem 526 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of certain embodiments may be stored by the file storage subsystem 524 in the storage subsystem 426 or in other machines accessible by the processor 514.
[0111] The bus subsystem 512 provides a mechanism for enabling the various components and subsystems of the computer system 510 to communicate with each other as expected. Although the bus subsystem 512 is schematically shown as a single bus, alternative embodiments of the bus subsystem may use multiple buses.
[0112] The computer system 510 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computer system 510 depicted in Figure 5 is provided only as a specific example for the purpose of illustrating some embodiments. Many other configurations of the computer system 510 may have more or fewer components than the computer system depicted in Figure 5 .
[0113] In cases where the systems described herein collect personal information about a user (or what is often referred to herein as a "participant") or can make use of personal information, the user may be provided with an opportunity to control whether programs or features collect user information (e.g., information about the user's social network, social activities or actions, occupation, user preferences, or the user's current geographic location), or to control whether and / or how content more relevant to the user is received from a content server. Similarly, certain data may be processed in one or more ways before storage or use to remove personal identity information. For example, a user's identity may be processed so that no personal identifiable information can be determined for the user, or the user's geographic location may be generalized where location information is obtained (e.g., at city, ZIP code, or state level) so that the user's specific geographic location cannot be determined. Thus, the user can control how information about the user is collected and / or used.
Claims
1. A method implemented by one or more processors, the method comprising: Identifying one or more occurrences of an action initiated by an automated assistant, each of the occurrences in response to user interface input provided by a user via one or more automated assistant interfaces that interact with the automated assistant; Identifying an automated assistant routine stored in association with the user, wherein the automated assistant routine defines a plurality of existing actions that are automatically performed via the automated assistant in response to initialization of the automated assistant routine, and wherein the action is an action in addition to the plurality of existing actions of the automated assistant routine; Determining a correlation between the automated assistant routine stored in association with the user and the action initiated by the automated assistant in the one or more occurrences; Based on determining the correlation between the automated assistant routine and the action: Causing a user interface output to be rendered via the user's client device, wherein the user interface output prompts the user to add the action to the automated assistant routine; Receiving a positive user interface input in response to the user interface output; In response to receiving the positive user interface input: Automatically adding the action to the automated assistant routine such that the action is automatically performed via the automated assistant in response to initialization of the automated assistant routine, wherein automatically adding the action to the automated assistant routine includes: Automatically adding the action with a determined position for performing the action in the automated assistant routine, the position relative to the position of performing the plurality of existing actions in the automated assistant routine; and In response to subsequent initialization of the automated assistant routine: Causing the plurality of existing actions and the action added to the automated assistant routine to be performed, including causing the action with the determined position to be performed.
2. The method according to claim 1, further comprising: Determining the determined position of the action based on a duration of the user interface output required to perform the action.
3. The method according to claim 2, wherein Determining the determined position of the action is further based on an additional duration of the user interface output required to perform an existing action among the plurality of existing actions.
4. The method according to claim 3, wherein Determining the determined position includes: Based on the duration being a duration less than the additional duration, determining that the determined position is before an additional position of the existing action.
5. The method according to claim 1, further comprising: Determining the determined position of the action based on whether user interface output is required when performing the action.
6. The method according to claim 1, further comprising: Determining the determined position of the action based on one or more temporal relationships, each of the temporal relationships being between a corresponding one of the one or more occurrences initiating the action and a corresponding occurrence of the automated assistant routine.
7. The method according to claim 6, wherein Determining the determined location of the action includes determining that the determined location is after the location of performing the plurality of existing actions based on each of the temporal relationships being after the corresponding occurrence completion of the auto assistant routine.
8. The method according to claim 7, wherein The determined location of the action is after the location of performing the plurality of existing actions and has a time delay after performing the plurality of existing actions.
9. The method according to claim 8, further comprising determining the time delay as a function of the temporal relationship.
10. The method according to claim 1, Among them, The corresponding user interface input includes one or more spoken utterances, which, when converted to text, include a first text of a first length; and wherein, the initialization of the auto assistant routine occurs in response to a given spoken utterance, which, when converted to text, includes a second text of a second length, wherein the second length is shorter than the first length.
11. The method according to claim 1, wherein, Rendering the user interface output via the user's client device is further based on executing the auto assistant routine and occurs at the completion of executing the auto assistant routine.
12. The method according to claim 1, wherein, The corresponding user interface input includes a corresponding spoken utterance and further includes: Identifying a profile of the user based on one or more voice characteristics of the corresponding spoken utterance; wherein, identifying the plurality of auto assistant routines stored in association with the user includes identifying the plurality of auto assistant routines based on the plurality of auto assistant routines stored in association with the profile of the user.
13. A method implemented by one or more processors, the method comprising: Identifying one or more occurrences of an action initiated by an auto assistant, wherein each of the occurrences is in response to corresponding user interface input provided by a user via one or more auto assistant interfaces interacting with the auto assistant, and wherein the action includes controlling a device having one or more device topology attributes; Identifying auto assistant routines stored in association with the user, wherein the auto assistant routines define a plurality of existing actions that are automatically performed via the auto assistant in response to initialization of the auto assistant routines, wherein the plurality of existing actions include controlling a routine device having one or more routine device topology attributes, and wherein the action is an action in addition to the plurality of existing actions of the auto assistant routines; Determining a correlation between the auto assistant routines stored in association with the user and the action initiated by the auto assistant in the one or more occurrences, wherein determining the correlation is based on a comparison of: The one or more device topology attributes, with The one or more routine device topology attributes; Based on determining the correlation between the auto assistant routine and the action: Rendering a user interface output via the user's client device, wherein the user interface output prompts the user to add the action to the auto assistant routine; Receiving affirmative user interface input in response to the user interface output; In response to receiving the affirmative user interface input: automatically add the action to the auto - assistant routine to be automatically executed via the auto - assistant in response to initialization of the auto - assistant routine; and in response to subsequent initialization of the auto - assistant routine: cause the plurality of existing actions and the action added to the auto - assistant routine to be executed.
14. The method according to claim 13, wherein, The one or more device topology attributes include the location of the device, and wherein the one or more routine device topology attributes include the routine location of the routine device.
15. The method according to claim 14, wherein, Determine the relevance based on: determining that the location matches the routine location based on the comparison.
16. The method according to claim 13, wherein, The plurality of existing actions further include controlling additional routine devices having one or more additional routine device topology attributes, wherein determining the relevance is further based on comparing the one or more device topology attributes with the one or more additional routine device topology attributes.
17. The method according to claim 13, Among them, The corresponding user interface input includes one or more spoken utterances, which when converted to text, include first text of a first length; and wherein initialization of the auto - assistant routine occurs in response to a given spoken utterance, which when converted to text, includes second text of a second length, wherein the second length is shorter than the first length.
18. The method according to claim 13, wherein, Causing the user interface output to be rendered via the user's client device is further based on executing the auto - assistant routine and occurs at the end of the execution of the auto - assistant routine.
19. A method implemented by one or more processors, the method comprising: identifying one or more occurrences of an initiating action, each occurrence in response to corresponding user interface input provided by a user via one or more computer interfaces; identifying an auto - assistant routine stored in association with the user, wherein the auto - assistant routine defines a plurality of existing actions that are automatically executed via the auto - assistant in response to initialization of the auto - assistant routine, and wherein the action is an action in addition to the plurality of existing actions of the auto - assistant routine; determining a relevance between the auto - assistant routine stored in association with the user and the action in the one or more occurrences, wherein determining the relevance is based on a comparison of: at least one time attribute of a past occurrence of the action with at least one corresponding time attribute of a past occurrence of the auto - assistant routine; Based on determining the relevance between the auto - assistant routine and the action: rendering a user interface output via the user's client device, wherein the user interface output prompts the user to add the action to the auto - assistant routine; receiving affirmative user interface input in response to the user interface output; In response to receiving the affirmative user interface input: automatically add the action to the auto - assistant routine so that the action is automatically executed in response to initialization of the auto - assistant routine.
20. The method according to claim 19, further comprising: In response to a subsequent initialization of the automatic assistant routine, cause a plurality of corresponding actions of a given routine to be performed, including causing an action added as an additional action among the plurality of corresponding actions of the given routine to be performed.
21. The method according to claim 19, wherein Automatically adding the action to the automatic assistant routine includes: Adding the action with a determined position for performing the action in the automatic assistant routine, the position relative to the positions of the plurality of existing actions in the automatic assistant routine.
22. The method according to claim 21, further comprising: Determining the determined position of the action based on a duration of a user interface output required to perform the action.
23. The method according to claim 22, wherein, Determining the determined position of the action is further based on an additional duration of a user interface output required to perform an existing action among the plurality of existing actions.
24. The method according to claim 23, wherein, Determining the determined position includes: Based on the duration being a duration less than the additional duration, determining that the determined position is before an additional position of the existing action.
25. The method according to claim 21, further comprising: Determining the determined position of the action based on whether a user interface output is required when performing the action.
26. The method according to claim 21, further comprising: Determining the determined position of the action based on at least one time attribute of a past occurrence of the action and based on at least one corresponding time attribute of a past occurrence of the automatic assistant routine.
27. The method according to claim 21, wherein Determining the determined position of the action includes determining that the determined position is after the positions of the plurality of existing actions.
28. The method according to claim 21, wherein The determined position of the action is after the positions of the plurality of existing actions and has a time delay after the plurality of existing actions are performed.
29. The method according to claim 19, Among them, The corresponding user interface input includes one or more spoken utterances, the one or more spoken utterances including, when converted to text, a first text of a first length; and wherein the initialization of the automatic assistant routine occurs in response to a given spoken utterance, the given spoken utterance including, when converted to text, a second text of a second length, wherein the second length is shorter than the first length.
30. The method according to claim 19, wherein, Causing the user interface output to be rendered via the user's client device is further based on performing the automatic assistant routine and occurs at the end of performing the automatic assistant routine.
31. The method according to claim 19, wherein, The corresponding user interface input includes a corresponding spoken utterance and further includes: Identifying a profile of the user based on one or more voice characteristics of the corresponding spoken utterance; wherein identifying the automatic assistant routine stored in association with the user includes identifying the automatic assistant routine based on the automatic assistant routine stored in association with the profile of the user.
32. A method implemented by one or more processors, the method comprising: Determining an action initiated by an automatic assistant, the action being initiated by the automatic assistant in response to one or more instances of user interface input provided by a user via one or more automatic assistant interfaces for interacting with the automatic assistant; Identify multiple natural language (NL) text descriptions of auto-assistant routines stored in association with the user, each of the auto-assistant routines defining multiple corresponding actions to be automatically performed via the auto-assistant in response to initialization of the auto-assistant routine, where the actions are additional actions beyond those of the auto-assistant routine; Compare the NL text descriptions of the actions with the multiple NL text descriptions of the auto-assistant routines stored in association with the user; Select a subset of the NL text descriptions of the auto-assistant routines stored in association with the user based on the comparison; Based on the actions being initiated in response to user interface input provided by the user and based on selecting the subset of the NL text descriptions of the auto-assistant routines: Cause a user interface output to be rendered via the user's client device, where the user interface output prompts the user to add the actions to one or more of the auto-assistant routines in the subset; and Receive a positive user interface input in response to the user interface output, where the positive user interface input indicates a given routine in the subset of the auto-assistant routines; and In response to receiving the positive user interface input: Automatically add the action as an additional action to be automatically performed in response to initialization of the given routine among the multiple corresponding actions of the given routine.
33. The method according to claim 32, wherein, The one or more instances of the user interface input include one or more spoken utterances that, when converted to text, include first text of a first length; and where initialization of the given routine occurs in response to a given spoken utterance that, when converted to text, includes second text of a second length, where the second length is shorter than the first length.
34. The method according to claim 32, wherein Causing the user interface output to be rendered via the user's client device is also based on executing the given routine and occurs at the completion of executing the given routine.
35. The method according to claim 34, wherein, The one or more instances of the user interface input that initiate the action include spoken utterances and also include: Identify the user's profile based on one or more voice characteristics of the spoken utterance; Where identifying the multiple NL text descriptions of the auto-assistant routines stored in association with the user includes identifying the multiple NL text descriptions of the auto-assistant routines based on the multiple NL text descriptions of the auto-assistant routines being stored in association with the user's profile; Where execution of the given routine occurs in response to an additional spoken utterance from the user; and Where causing the user interface output to be rendered via the user's client device is also based on determining that the additional spoken utterance has one or more voice characteristics corresponding to the user's profile.
36. The method according to claim 32, wherein, Causing the user interface output to be rendered via the user's client device occurs at the completion of the action initiated by the auto-assistant in response to the one or more instances of the user interface input.
37. The method according to claim 32, wherein The action includes providing a command to change at least one state of a connected device managed by the automated assistant.
38. The method according to claim 32, wherein, The action includes providing a command to an agent controlled by a third party, where the third party is different from the party controlling the automated assistant.
39. The method according to claim 32, wherein, Selecting the subset of the NL text descriptions of the automated assistant routines includes: comparing the NL text description of the action with the NL text descriptions of each of the multiple corresponding actions in each of the automated assistant routines in the automated assistant routines; and for each of the automated assistant routines in the automated assistant routines, based on the comparison, determining whether the NL text description of the action includes one or more contradictions with any of the NL text descriptions of the multiple corresponding actions in the automated assistant routines; and selecting the subset of the NL text descriptions of the automated assistant routines based on determining that the NL text descriptions of the multiple corresponding actions of the subset of the automated assistant routines lack one or more contradictions with the action.
40. The method according to claim 39, wherein, The one or more contradictions include temporal contradictions.
41. The method according to claim 40, wherein, The temporal contradiction includes that the action cannot be completed during the time frame of each of the multiple corresponding actions.
42. The method according to claim 39, wherein, The one or more contradictions include device incompatibility contradictions.
43. The method according to claim 42, wherein, The device incompatibility contradiction includes that the action requires a device with certain capabilities and none of the multiple corresponding actions is associated with any device having the certain capabilities.
44. The method according to claim 32, wherein, Automatically adding the action includes determining a position in the given routine for performing the action, the position relative to the multiple corresponding actions of the given routine.
45. The method according to claim 44, wherein, The position for performing the action is before the multiple corresponding actions, after the multiple corresponding actions, or between two of the multiple corresponding actions.
46. The method according to claim 45, wherein, The position is after the multiple corresponding actions and includes a time delay, where the time delay is based on one or more past time delays of past initiations of the action by the user after one or more past initiations of the given routine by the user.
47. The method according to claim 32, wherein, The one or more instances of the user interface input for initiating the action are received via an additional client device that is linked to but separate from the client device, and the user interface output is rendered via the client device.
48. The method according to claim 47, wherein, The user interface output includes graphical output and further includes: selecting the additional client device for providing the user interface output based on determining that the client device lacks display capabilities and the additional client device includes display capabilities.
49. The method according to claim 32, wherein, Among them, the multiple NL text descriptions that identify the automated assistant routines stored in association with the user include multiple NL text descriptions that are stored in association with the user's profile based on the multiple NL text descriptions of the automated assistant routines to identify the automated assistant routines.
50. A system, comprising: a memory that stores instructions; one or more processors operable to execute the instructions to perform the method of any of the preceding claims.