Providing a specific rationale for implementing assistant commands

By processing user input through ASR and NLU models to provide rationales for assistant actions, the solution addresses user confusion and enhances data security in automated assistants, reducing interaction time and improving privacy settings.

JP7731989B2Active Publication Date: 2025-09-01GOOGLE LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023537164
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-22
Filing Date
2021-11-29
Publication Date
2025-09-01
Estimated Expiration
2041-11-29

Smart Images

  • Figure 0007731989000001
    Figure 0007731989000001
  • Figure 0007731989000002
    Figure 0007731989000002
  • Figure 0007731989000003
    Figure 0007731989000003
Patent Text Reader

Abstract

Implementations described herein relate to providing a particular rationale for why an automated assistant performed (or did not perform) a particular and / or another implementation of an assistant command. For example, implementations may receive user input including an assistant command, process the user input to determine data to use to perform the particular or another implementation of the assistant command, and have the automated assistant use the data to perform the particular or another implementation of the assistant command. In certain implementations, output including the particular rationale can be provided for presentation to a user in response to additional user input requesting the particular rationale. In certain implementations, a selectable element can be visually rendered and, upon selection by a user, output including the particular rationale can be provided for presentation to a user.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans may engage in human-computer interactions in interactive software applications referred to herein as “automated assistants” (also referred to as “chatbots,” “interactive personal assistants,” “self-controlled personal assistants,” “personal voice assistants,” “conversational agents,” etc.). For example, a human (who may also be referred to as a “user” when interacting with an automated assistant) may provide spoken natural language input (i.e., spoken utterances) to the automated assistant, which utterances are optionally converted to text and then processed, and / or by providing textual (e.g., keyed) natural language input or touch input. Typically, the automated assistant responds to such user input, including assistant commands, by providing responsive user interface output (e.g., auditory and / or visual user interface output), controlling smart network devices, and / or performing other actions.

[0002] Typically, automated assistants rely on a pipeline of components to interpret and respond to such user input, including assistant commands. For example, an automatic speech recognition (ASR) engine can process audio data corresponding to a user's utterance and generate an ASR output, such as a transcription of the utterance (i.e., a series of terms and / or tokens). Further, a natural language interpretation (NLU) engine can process the ASR output to generate an NLU output, such as the user's intent in providing the utterance and, possibly, parameter slot values ​​associated with the intent. Further, a fulfillment engine can be used to process the NLU output and generate a fulfillment output, such as a structured request to obtain response content for the utterance, and / or to perform an action responsive to the utterance.

[0003] In some cases, a user may not understand why an automated assistant provides particular response content and / or performs a particular action in response to receiving user input that includes an assistant command. This misunderstanding may be exacerbated when the automated assistant does not provide specific response content and / or does not perform a particular action as intended by the user's user input. For example, if a given user provides user input directed to an automated assistant requesting that the automated assistant play music, but the automated assistant instead provides music playback using an undesired software application or provides search results in response to the user input, the given user may be confused as to why the automated assistant is not playing the music as desired. As a result, the given user may provide additional user input, including other instances of the same or different user input, to play music using the desired application, thereby increasing the human-computer interaction time between the given user and the automated assistant. Furthermore, even if the automated assistant plays music using the desired software application, the given user may have concerns about the automated assistant because they do not understand how the automated assistant is able to play music using the desired application or because they are concerned about the security of their data. For this reason, it is beneficial to provide a mechanism for a given user to understand why an automated assistant is performing a particular implementation of an assistant command. Summary of the Invention [Means for solving the problem]

[0004] Implementations disclosed herein are directed to providing a particular rationale for why an automated assistant performed (or did not perform) a particular and / or another implementation of an assistant command. For example, implementations may receive user input including an assistant command, process the user input to determine data to use to perform a particular or another implementation of the assistant command, and cause the automated assistant to use the data to perform the particular or another implementation of the assistant command. In some implementations, output including the particular rationale can be provided for presentation to a user in response to additional user input requesting the particular rationale. In some implementations, one or more selectable elements are visually renderable, and when a given one of the one or more selectable elements is selected by a user, output including the particular rationale can be provided for presentation to a user.

[0005] For example, suppose a user of a client device provides the utterance, "Play rock music." In this example, the automated assistant can process audio data capturing the utterance with an automatic speech recognition (ASR) model to generate ASR outputs such as a speech hypothesis predicted to correspond to the utterance, predicted phonemes predicted to correspond to the utterance, and / or other ASR outputs, and possibly ASR metrics associated with each of the speech hypotheses, predicted phonemes, and / or other ASR outputs (e.g., the likelihood that a given speech hypothesis or a given predicted phoneme corresponds to the utterance). Additionally, the automated assistant may process the ASR output using a natural language understanding (NLU) model to generate NLU outputs, such as one or more predicted intents of the user when providing the utterance, one or more slot values ​​of corresponding parameters associated with each of the one or more predicted intents, and / or other NLU outputs, and possibly NLU metrics associated with each of the intents, slot values, and / or other NLU outputs (e.g., indicating the likelihood that a given intent and / or given slot value corresponds to an actual intent and / or desired slot value when providing the utterance). In this example, the automated assistant may infer one or more slot values, such as an artist slot value of an artist parameter associated with the music intent, a song slot value of a song parameter associated with the music intent, or a software application or streaming service slot value of a software application or streaming service parameter associated with the music intent, because the user specifies only a particular genre of music (e.g., rock). Variations in the inferred slot values ​​may result in one or more interpretations of the utterance. In various implementations, assuming the automated assistant has access to one or more user profiles for the user of the client device, the automated assistant can leverage the user profile data to infer one or more slot values.Otherwise, the automated assistant may utilize one or more default slot values.

[0006] Furthermore, the automated assistant can process the NLU output using one or more realization rules and / or realization models to generate a realization output, such as one or more structured requests, that are sent to one or more fulfillers (e.g., software applications, servers, etc.) that can satisfy the utterance. Upon sending the one or more structured requests, the one or more fulfillers can generate one or more candidate realizations and send the one or more candidate realizations back to the automated assistant. In response to receiving the one or more candidate realizations, the automated assistant can generate realization metrics associated with each of the one or more candidate realizations (e.g., likelihood, indicating that a given one of the one or more candidate realizations will satisfy the utterance when implemented) based on user profile data, assuming the automated assistant has access to one or more user profiles of a user of the client device. The automated assistant can rank the one or more candidate realizations based on ASR metrics, NLU metrics, and / or realization metrics and select a particular candidate realization based on the ranking. Furthermore, the automated assistant can cause a selected particular candidate realization to be implemented in an attempt to satisfy the utterance.

[0007] For example, in this example, assume that the automated assistant has determined a first interpretation of the utterance "play rock music," where the utterance has an artist slot value of "Artist 1" for an artist parameter associated with the musical intent, a song slot value of "Song 1" for a song parameter associated with the musical intent, and a software application or streaming service slot value of "Application 1" for a software application or streaming service parameter associated with the musical intent. Further, assume that the automated assistant has determined a first interpretation of the utterance "play rock music," where the utterance has an artist slot value of "Artist 1" for the artist parameter associated with the musical intent, a song slot value of "Song 1" for the song parameter associated with the musical intent, and a software application or streaming service slot value of "Application 2" for the software application or streaming service parameter associated with the musical intent. In this example, "Application 1" and "Application 2" can be considered one or more realizers capable of satisfying the utterance. Thus, the automated assistant can send one or more structured requests to "Application 1" and "Application 2" (and possibly other realizing parties capable of satisfying the utterance) to obtain one or more candidate realizations. Further, the automated assistant can rank one or more candidate realizations and select a particular candidate realization to perform a particular realization in response to the utterance. This example further assumes that the automated assistant selects a candidate realization associated with the first interpretation. Thus, the automated assistant can cause "Song 1" by "Artist 1" to be played as a particular realization of the utterance using "Application 1" through a speaker on the client device (or an additional client device in communication with the client device).

[0008] In some implementations, following the automated assistant executing a particular implementation, the user of the client device may provide additional user input. The additional user input requests the automated assistant to provide a specific rationale for why a particular implementation was performed and / or why another implementation was not performed. In some variations of these implementations, the request for a specific rationale may be a general request for a specific rationale (e.g., "Why did you do that?"), while in other implementations, the request for a specific rationale may be a specific request for a specific rationale (e.g., "Why did you play music in application 1?", "Why didn't you use application 2?", "Why did you choose artist 1?", etc.). For example, assume that the user provides yet another utterance, "Why did you do that?" In this example, the request is a general request for a specific rationale, and the automated assistant can determine additional data associated with the first interpretation of the utterance to generate output in response to the general request (e.g., "You most often use application 1 to listen to music, you have listened to artist 1 in the past, and song 1 is artist 1's most popular song," etc.). Meanwhile, assume that the user provides yet another utterance, "Why didn't you use Application 2?" In this example, the request is a specific request for a particular justification, and the automated assistant can determine additional data associated with the first interpretation and / or the second interpretation of the utterance to generate an output that is responsive to the general request (e.g., "You use Application 1 more often to listen to music than Application 2"). However, in this example, it is assumed that the user did not grant the automated assistant access to "Application 2" for the specific request.In this example, the automated assistant may additionally or alternatively determine recommendation data associated with the recommended action and generate a prompt (e.g., "You do not have permission to access using application 2. Can you grant me access?") based on the recommendation data that includes the recommended action. Thus, not only may the automated assistant provide a particular rationale for a particular aspect of implementation, but the automated assistant may also prompt the user to adapt current and / or future implementations in response to receiving user input that includes assistant commands.

[0009] In additional or alternative implementations, following execution of a particular implementation by the automated assistant, the automated assistant may actively provide one or more selectable elements associated with a particular rationale for presentation to the user via a display on the client device. For example, a first selectable element associated with a general request for a particular rationale can be provided for presentation to the user and, when selected, cause the automated assistant to provide the particular rationale in response to the general request. Furthermore, a second selectable element associated with a first specific request for a particular rationale can additionally or alternatively be provided for presentation to the user and, when selected, cause the automated assistant to provide the particular rationale in response to the first specific request. Furthermore, a third selectable element associated with a second specific request for a particular rationale can additionally or alternatively be provided for presentation to the user and, when selected, cause the automated assistant to provide the particular rationale in response to the second specific request. In some variations of these implementations, the automated assistant can provide one or more selectable elements for presentation to the user in response to a determination that the ASR metrics, NLU metrics, and / or realization metrics did not meet a threshold that indicates the automated assistant was not highly confident about a particular realization of the assistant command. In some variations of these implementations, the automated assistant can provide one or more selectable elements for presentation to the user regardless of the ASR metrics, NLU metrics, and / or realization metrics.

[0010] While the above example describes why the automated assistant selected a particular software application to play music (e.g., “Application 1”) in relation to providing a particular rationale, it should be understood that this is intended to be illustrative and not limiting. As described herein, the techniques described herein can be used to provide a particular rationale for any implementation aspect, such as why a particular computing device was selected and utilized to implement the assistant command, why particular slot values ​​were selected for corresponding parameters, or why the automated assistant failed to perform another implementation and / or other aspects described herein. Furthermore, although the recommended action described in the above example includes the user granting the automated assistant access to a particular software application (e.g., “Application 2”), it should be understood that this is intended to be illustrative and not limiting. The techniques described herein can be used to provide any recommended action for adapting the implementation of an assistant command, such as downloading a software application at a client device, communicatively connecting an additional client device to the client device via a network, and / or any other recommended action described herein.

[0011] By using the techniques described herein, one or more technical benefits are achieved. As a non-limiting example, the techniques described herein enable an automated assistant to provide a specific rationale (or lack thereof) for a particular aspect of implementation, thereby enabling a user of a client device to understand when and how their data is being used. Furthermore, the techniques described herein enable an automated assistant to quickly and effectively adapt user data privacy settings, reducing the amount of user input by eliminating the need for the user to manually change the user data privacy settings by navigating between various interfaces. As a result, the security of user data can be improved and the computing resources of the client device can be protected. As another non-limiting example, by providing recommended actions to be performed and continuing the human-computer interaction, an automated assistant can reclaim user interaction that would otherwise be unused. For example, if a user provides the utterance "turn on the lights" to control one or more smart lights, but the user has not authorized the automated assistant to access the software application to control the smart lights, the automated assistant can prompt the user to grant access to the software application to control the smart lights rather than simply indicating that the lights cannot be controlled at that particular instance in time. As a result, computational and / or network resources can be protected using the techniques described herein.

[0012] The above description is provided as a summary of only certain implementations disclosed herein. Such implementations and other implementations are described in further detail herein.

[0013] It should be appreciated that any combination of the above concepts and additional concepts described in more detail herein is considered part of the presently disclosed subject matter, for example, any combination of claimed subject matter at the end of this disclosure is considered part of the presently disclosed subject matter. [Brief explanation of the drawings]

[0014] [Figure 1] 1A-1C show block diagrams of exemplary embodiments illustrating various aspects of the present disclosure, in which implementations disclosed herein can be implemented. [Figure 2] 1 shows a flowchart illustrating an example method for, in various implementations, causing a user input included in the automated assistant to execute a particular realization of an assistant command directed to the automated assistant and providing a particular rationale for why the automated assistant executed the particular realization of the assistant command. [Figure 3] A flowchart illustrating an example method for, in various implementations, determining that a particular realization of an assistant command contained in user input and directed to an automated assistant is infeasible and having the automated assistant provide a particular rationale as to why the particular rationale cannot be implemented is shown. [Figure 4A] 10A-10C illustrate non-limiting examples that provide a specific rationale for realizing assistant commands according to various implementations. [Figure 4B] 10A-10C illustrate non-limiting examples that provide a specific rationale for realizing assistant commands according to various implementations. [Figure 4C] 10A-10C illustrate non-limiting examples that provide a specific rationale for realizing assistant commands according to various implementations. [Figure 5A] Additional non-limiting examples are provided that provide a particular rationale for realizing assistant commands according to various implementations. [Figure 5B]Additional non-limiting examples are provided that provide a particular rationale for realizing assistant commands according to various implementations. [Figure 6] 1 illustrates an exemplary architecture of a computing device, according to various implementations. DETAILED DESCRIPTION OF THE INVENTION

[0015] 1, a block diagram of an example embodiment illustrating various aspects of the present disclosure, in which implementations disclosed herein may be implemented, is shown. The example embodiment includes a client device 110, one or more cloud-based automated assistant components 115, one or more first-party servers 191, and one or more third-party servers 192.

[0016] Client device 110 is capable of executing automated assistant client 113. Automated assistant client 113 may be an application separate from (e.g., installed "on top of") the operating system of client device 110, or alternatively may be implemented directly by the operating system of client device 110. As further described below, automated assistant client 113 may also interact with one or more cloud-based automated assistant components 115 in some cases in response to various requests received by user interface component 112 of client device 110. Additionally, and as described below, other engines of client device 110 may interact with one or more cloud-based automated assistant components 115 in some cases.

[0017] One or more cloud-based automated assistant components 115 may be implemented on one or more computing systems (e.g., servers, collectively referred to as "cloud" or "remote" computing systems) communicatively connected to client device 110 via one or more local area networks ("LANs," including Wi-Fi LANs, Bluetooth networks, near-field communication networks, mesh networks, etc.), wide area networks ("WANs," including the Internet, etc.), and / or other networks. The communicative connection of cloud-based automated assistant component 115 with client device 110 is generally shown as 1991 in FIG. 1 . In some implementations, client device 110 may also be communicatively connected to other client devices (not shown) described herein via one or more networks (e.g., LANs and / or WANs).

[0018] Additionally, one or more cloud-based automated assistant components 115 may be communicatively connected to one or more first party servers 191 and / or one or more third party servers 192 via one or more networks (e.g., a LAN, a WAN, and / or other networks). The communicative connection of the cloud-based automated assistant component 115 with one or more first party servers 191 is generally indicated by 1992 in FIG. 1 . Additionally, the communicative connection of the cloud-based automated assistant component 115 with one or more third party servers 192 is generally indicated by 1993 in FIG. 1 . In some implementations, although not explicitly shown in FIG. 1 , the client device 110 is additionally or alternatively communicatively connected to one or more first party servers 191 and / or one or more third party servers 192 via one or more networks (e.g., a LAN, a WAN, and / or other networks). Additionally, the one or more networks 1991, 1992, and 1993 are hereinafter collectively referred to as “network 199” for brevity.

[0019] Automated assistant client 113, through its interaction with one or more cloud-based automated assistant components 115, may form what appears from the user's perspective as a logical instance of automated assistant 120, with which a user of client device 110 may engage in a human-computer interaction. For example, the instance of automated assistant 120 enclosed within the dashed line includes the automated assistant client 113 on client device 110 and one or more cloud-based automated assistant components 115. In this manner, it should be understood that each user interacting with an automated assistant client 113 running on client device 110 may in fact be interacting with their own logical instance of automated assistant 120 (or a logical instance of automated assistant 120 shared among a household or other group of users and / or among multiple automated assistant clients 113). While only client device 110 is shown in FIG. 1 , it should be understood that one or more cloud-based automated assistant components 115 may additionally serve many additional groups of client devices. Additionally, although FIG. 1 illustrates a cloud-based automated assistant component 115, it should be understood that in various implementations, the automated assistant 120 may be implemented solely on the client device 110.

[0020] As used herein, a first party device or system (e.g., one or more first party servers 191, one or more first party software applications, etc.) refers to a system controlled by the same party that controls the automated assistant 120 referenced herein. For example, one or more first party servers 191 may refer to a system hosting a search engine service, a communication service (e.g., email, SMS messages, etc.), a navigation service, a music service, a document editing or sharing service, and / or other services that are controlled by the same party that controls the automated assistant 120 referenced herein. In contrast, a third party device or system (e.g., one or more third party servers 192, one or more third party software applications, etc.) refers to a system controlled by a party different from the party that controls the automated assistant 120 referenced herein. For example, one or more third party servers 192 may refer to a system hosting the same services, but each service is controlled by a party different from the party that controls the automated assistant 120 referenced herein.

[0021] The client devices 110 may include, for example, one or more of a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device in a user's vehicle (e.g., an in-vehicle computing device, an in-vehicle entertainment system, an in-vehicle navigation system), an interactive standalone speaker (e.g., with or without a display), a smart appliance, a smart networked device such as a smart TV, smart lighting, or a smart washer / dryer, a wearable device with the user's computing device (e.g., a user's watch with a computing device, a user's glasses with a computing device, a virtual or augmented reality computing device), and / or any IoT device capable of receiving user input directed to the automated assistant 120. Additional and / or alternative client devices may be provided.

[0022] In various implementations, client device 110 may include one or more presence sensors 111 configured to provide a signal indicating the detection of a presence, particularly the presence of a person, upon acknowledgement from a user of client device 110. In some such implementations, automated assistant 120 may be able to identify client device 110 (or another computing device associated with the user of client device 110) and fulfill utterances (or other inputs directed to automated assistant 120) based at least in part on the presence of the user at client device 110 (or at another computing device associated with the user of client device 110). The utterance (or other input directed to automated assistant 120) is satisfied by rendering responsive content (auditory and / or visual) at client device 110 and / or other computing devices associated with the user of client device 110, causing client device 110 and / or other computing devices associated with the user of client device 110 to be controlled, and / or causing client device 110 and / or other computing devices associated with the user of client device 110 to perform other actions to satisfy the utterance (or other input directed to automated assistant 120). As described herein, automated assistant 120 can utilize data determined based on presence sensor 111 in determining client device 110 (or other computing device) based on whether the user is or was recently in proximity, and can provide corresponding commands only to client device 110 (or other computing device).In additional or alternative implementations, the automated assistant 120 can utilize data determined based on the presence sensor 111 in determining whether any user (any user or a specific user) is currently in proximity to the client device 110 (or other computing device), and in some cases can suppress the provision of data to and / or from the client device 110 (or other computing device) based on the user's proximity to the client device 110 (or other computing device).

[0023] The presence sensor 111 may take various forms. For example, the client device 110 may include one or more vision components (e.g., a digital camera and / or other vision components) configured to capture and provide signals indicative of detected movement within its field of view. Additionally or alternatively, the client device 110 may include other types of light-based presence sensors 111, such as a passive infrared (“PIR”) sensor that measures infrared (“IR”) light emitted from objects within its field of view. Additionally or alternatively, the client device 110 may include a presence sensor 111 that detects acoustic (or pressure) waves, such as one or more microphones.

[0024] Additionally or alternatively, in some implementations, presence sensor 111 may be configured to detect other phenomena associated with the presence of a person or a device. For example, in some embodiments, client device 110 may include presence sensor 111 that detects various types of wireless signals (e.g., radio, ultrasonic, electromagnetic, or other wave motions) emitted by, for example, other computing devices (e.g., mobile devices, wearable computing devices, etc.) held / operated by a user and / or other computing devices. For example, client device 110 may be configured to emit waves, such as ultrasonic or infrared waves, that are imperceptible to people but can be detected by other computing devices (e.g., via an ultrasonic / infrared receiver such as an ultrasonic-sensing microphone).

[0025] Additionally or alternatively, client device 110 may emit other types of waves, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.), that are undetectable to a person but detectable by other computing devices (e.g., mobile devices, wearable computing devices, etc.) carried / operated by the user, and used to identify the user's specific location. In some implementations, GPS and / or Wi-Fi triangulation may be used to detect a person's location, for example, based on GPS and / or Wi-Fi signals to / from client device 110. In other implementations, other wireless signal characteristics, such as time of flight, signal strength, etc., may be used, alone or in combination, by client device 110 to determine the location of a particular individual based on signals emitted by other computing devices carried / operated by the user.

[0026] Additionally or alternatively, in some implementations, client device 110 may recognize a voice and recognize the user from the voice. For example, some instances of automated assistant 120 may be configured to match the voice with a user profile, e.g., for purposes of providing / restricting access to various resources. In some implementations, the speaker's movement may then be determined, for example, by presence sensor 111 (possibly a GPS sensor and / or an accelerometer) of client device 110. In some implementations, the user's location may be predicted based on such detected movement, and if any content is rendered on client device 110 and / or other computing devices based at least in part on the proximity of client device 110 and / or other computing devices to the user's location, this location may be considered the user's location. In some implementations, the user may simply be considered to be at the most recent location where the user engaged with automated assistant 120, especially if not much time has passed since the last engagement.

[0027] Additionally, client device 110 includes a user interface component 112, which may include one or more user interface input devices (e.g., a microphone, a touchscreen, a keyboard, and / or other input devices), and / or one or more user interface output devices (e.g., a display, speakers, a projector, and / or other output devices). Additionally, client device 110 and / or any other computing device may include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components for facilitating communication over network 199. In some implementations, operations performed by client device 110, other computing devices, and / or automated assistant 120 may be distributed across multiple computing devices, while in other implementations, the operations described herein may be performed solely at client device 110 or at a remote system. Automated assistant 120 may be implemented, for example, as a computer program running on one or more computers at one or more locations connected to each other via a network (e.g., network 199 of FIG. 1 ).

[0028] As noted above, in various implementations, client device 110 may operate automated assistant client 113. In various embodiments, automated assistant client 113 may include speech capture / automatic speech recognition (ASR) / natural language understanding (NLU) / text-to-speech (TTS) / realization module 114. In other implementations, one or more aspects of speech capture / ASR / NLU / TTS / realization module 114 may be implemented separately from automated assistant client 113 (e.g., by one or more of cloud-based automated assistant components 115).

[0029] The speech capture / ASR / NLU / TTS / realization module 114 may be configured to perform one or more functions, including, for example, capturing a user's speech (speech capture, e.g., through a respective microphone (which may optionally include one or more presence sensors 111)), converting the captured speech into recognized text and / or other representations or embeddings using ASR models stored in the machine learning (ML) model database 120A, parsing and / or annotating the recognized text using NLU models stored in the ML model database 120A, and / or determining realization data to be used in generating a structured request, obtaining the data and / or performing an action in response to the user's utterance using one or more realization rules and / or realization models stored in the ML model database 120A. Additionally, speech capture / ASR / NLU / TTS / realization module 114 may be configured to convert text to speech using TTS models stored in ML model database 120A, and synthesized speech audio data capturing the synthesized speech based on the text to speech conversion may be provided for audible presentation to a user of client device 110 through a speaker of client device 110. Instances of these ML models may be stored locally at client device 110 and / or accessible to client device 110 via network 199 of FIG. 1 . In some implementations, because client device 110 may be relatively constrained in terms of computational resources (e.g., processor cycles, memory, battery, etc.), speech capture / ASR / NLU / TTS / realization module 114 local to client device 110 may be configured to convert a finite number of different spoken phrases into text (or into other formats, such as lower-dimensional embeddings) using speech recognition models.

[0030] Any speech input may be sent to one or more cloud-based automated assistant components 115, which may include cloud-based ASR module 116, cloud-based NLU module 117, cloud-based TTS module 118, and / or cloud-based realization model 119. These cloud-based automated assistant components 115 can utilize the virtually limitless resources of the cloud and can perform the same or similar functions as described with respect to local speech capture / ASR / NLU / TTS / realization module 114 for client device 110, although it should be noted that speech capture / ASR / NLU / TTS / realization module 114 can perform this function locally at client device 110 without interaction with cloud-based automated assistant component 115.

[0031] 1 illustrates a single client device for a single user, it should be understood that this is for purposes of illustration and not limitation. For example, one or more additional client devices for a user can also implement the techniques described herein. These additional client devices may communicate with client device 110 (e.g., via network 199). As another example, client device 110 may be available to multiple users in a shared setting (e.g., a group of users, a household, a hotel room, a shared space at work).

[0032] In some implementations, client device 110 may further comprise various engines utilized to ensure that a particular rationale for implementing an assistant command included in user input and directed to automated assistant 120 is provided for presentation to a user of client device 110 in response to additional user input requesting a particular rationale. For example, as shown in FIG. 1 , client device 110 may further comprise request engine 130 and inference engine 140. Client device 110 may further comprise on-device memory including user profile database 110A, ML model database 120A, and metadata database 140A. In some implementations, these various engines are executable solely on client device 110. In additional or alternative implementations, one or more of these various engines are executable remotely from client device 110 (e.g., as part of cloud-based automated assistant component 115). For example, in implementations in which assistant commands are implemented locally on client device 110, on-device instances of these various engines are available to perform the operations described herein. However, in implementations where assistant commands are implemented remotely from the client device 110 (e.g., as part of the cloud-based automated assistant component 115), remote instances of these various engines are available to perform the operations described herein.

[0033] In some implementations, following execution of an assistant command directed to automated assistant 120 contained in user input detected via client device user interface component 112, request engine 130 can execute additional user input (e.g., using one or more speech capture / ASR / NLU / TTS / realization modules 114) to determine whether the additional user input includes a request. For example, a user of client device 110 may provide the utterance "Turn on the lights" to switch lights in the user's residence of client device 110 from an off state to an on state. As described in further detail with reference to FIGS. 2 and 3 , audio data capturing the utterance can be processed with an ASR model stored in ML model database 120A to generate ASR output (possibly including ASR metrics), the ASR output can be processed with an NLU model stored in ML model database 120A to generate NLU output (possibly including NLU metrics), and the NLU output can be processed with realization rules and / or realization models stored in ML model database 120A to generate realization output. The structure request associated with the realization output can be sent to one or more realization parties, such as various software applications running locally on the client device 110 and / or remotely on the first party server 191 and / or third party server 192, and one or more candidate realizations can be generated in response to the structure request (each realization metric can be associated with one or more corresponding realizations). The automated assistant 120 can rank the one or more candidate realizations based on ASR metrics, NLU metrics, and / or realization metrics. The ASR metrics, NLU metrics, realization metrics, and / or any other data associated with the realization of the utterance can be stored in metadata database 140A.This data can then be accessed by the automated assistant 120 to determine data associated with providing a particular rationale for implementing the assistant commands described herein (see, e.g., Figures 2, 3, 4A-4C, and 5A-5B).

[0034] However, in this example, it is assumed that the user of client device 110 has not authorized automated assistant 120 to access software applications or services associated with lighting control at the client device user's residence. Thus, one or more realization candidates in this example may indicate that no software application is accessible at client device 110 or at a server capable of fulfilling the utterance (e.g., first-party server 191 and / or a third-party server) due to an inability to determine the data to be used to turn on the lights due to a lack of access to at least one or more of these realization parties. As a result, automated assistant 120 can determine alternative data to be used to notify the user of client device 110 that automated assistant 120 is unable to fulfill the utterance. Thus, automated assistant 120 can use speech capture / ASR / NLU / TTS / realization module 114 to generate synthesized speech audio data based on the alternative data, including, for example, synthesized speech saying, "Sorry, I can't turn on the lights," which can be provided for auditory presentation through the speaker of client device 110.

[0035] In some implementations, it is assumed that the user of client device 110 provides additional user input requesting automated assistant 120 to provide a specific rationale for why automated assistant 120 should execute a specific implementation of the assistant command, and request engine 130 can determine whether the request is a general request for a specific rationale for the implementation or a specific request for a specific rationale for the implementation. Request engine 130 can determine whether the request is a general request for a specific rationale for the implementation or a specific request for a specific rationale for the implementation based at least on NLU output generated based on processing the additional user input, and inference engine 140 can adapt the additional data determined to be used to provide the specific rationale based on the type of request (e.g., as described with reference to FIGS. 2, 3, 4A-4C, and 5A-5B). In additional or alternative implementations, automated assistant 120 can obtain recommendation data for determining a recommended action that, when executed, enables automated assistant 120 to produce a specific implementation of the assistant command. In this example, the recommended action may include granting the automated assistant 120 access to a software application or service that enables lighting control. To this end, the automated assistant 120 may utilize the voice capture / ASR / NLU / TTS / realization module 114 to generate additional synthetic voice audio data based on the recommendation data, including, for example, a synthetic voice stating, "If you grant me access to the lighting application, I can control the lights," and have the synthetic voice provided for audible presentation via the speaker of the client device 110. It should be understood that the above description is provided for purposes of illustration and not limitation, and further description of the techniques described herein is provided below with reference to FIGS. 2, 3, 4A-4C, and 5A-5B.

[0036] Referring to FIG. 2, a flowchart illustrating an example method 200 for causing a particular implementation of an assistant command included in user input and directed to an automated assistant is shown. For convenience, the operations of method 200 are described with reference to a system that performs those operations. The system of method 200 includes one or more processors, memory, and / or other components of a computing device (e.g., client device 110 of FIG. 1 , client device 410 of FIGS. 4A-4C , client device 510 of FIGS. 5A-5B , computing device 610 of FIG. 6 , one or more servers, and / or other computing devices). Furthermore, although the operations of method 200 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.

[0037] In block 252, the system receives user input from a user of the client device, the user input including an assistant command, directed to the automated assistant. In some implementations, the user input may correspond to speech captured in audio data generated by a microphone of the client device. In additional or alternative implementations, the user input may correspond to touch input or key input received through a display of the client device or other input device of the client device (e.g., a keyboard and / or mouse).

[0038] In block 254, the system processes the user input to determine data to utilize in executing a particular realization of the assistant command. In implementations in which the user input corresponds to speech, the audio data capturing the speech can be processed using an ASR mode to generate ASR output (e.g., speech hypotheses, phonemes, and / or other ASR output) and, in some cases, ASR metrics associated with the ASR output. Further, the ASR output can be processed using an NLU model to generate NLU output (e.g., an intent determined based on the ASR output, slot values ​​for parameters associated with the intent determined based on the ASR output, etc.) and, in some cases, NLU metrics associated with the NLU output. Further, the NLU output can be processed using realization rules and / or realization modes to generate realization output used to generate an outgoing request, to obtain data used to execute the realization of the assistant command, and / or to cause an action to be performed based on the realization output (e.g., sent to first party server 191 of FIG. 1 , third party server 192 of FIG. 1 , a first party software application implemented locally on the client device, a third party software application implemented locally on the client device, etc.) and, in some cases, realization metrics associated with the realization data. In an implementation in which the user input corresponds to touch input or key input, the text corresponding to the touch input or key input can be processed using the NLU model to generate the NLU output and, in some cases, NLU metrics associated with the NLU output. Further, the NLU output can be processed using realization rules and / or realization models to generate realization output used to generate an outgoing request, to obtain data used to execute the realization of the assistant command, and / or to cause an action to be performed based on the realization output in the execution of the realization of the assistant command.

[0039] In block 256, the system utilizes the data to cause the automated assistant to perform a particular realization of the assistant command. In particular, the realization output may include data sent to one or more of first party server 191 of FIG. 1, third party server 192 of FIG. 1, a first party software application implemented locally on the client device, a software application implemented locally on the client device, etc., resulting in one or more candidate realizations. The system can select a particular candidate realization from the one or more candidate realizations and perform a particular realization of the assistant command based, for example, on ASR metrics, NLU metrics, and / or realization metrics. For example, suppose a user of a client device provides the utterance, "Play rock music." In this example, audio data capturing the utterance may be processed using an ASR model to generate, as ASR output, a first speech hypothesis "play rock music" associated with a first ASR metric (e.g., a probability, a binary value, a log-likelihood, or other likelihood that the first speech hypothesis corresponds to a term and / or phrase contained in the utterance), a second speech hypothesis "play some Bach music" associated with a second ASR metric, and / or other speech hypotheses and corresponding metrics. Each speech hypothesis may then be processed using an NLU model to generate, as NLU data, a first intent "play music" with a slot value of "rock" for a genre parameter associated with a first NLU metric (e.g., a probability, a binary value, a log-likelihood, or other likelihood that the first intent and slot value correspond to the user's desired intent), and a second intent "play music" with a slot value of "Bach" for an artist parameter associated with a second NLU metric.

[0040] Further, the one or more realization candidates in this example may include, for example, a first realization candidate that plays rock music using a first party media application associated with a first realization metric (e.g., returned to the system by the first party media application in response to a realization request from the system), a second realization candidate that plays rock music using a third party media application associated with a second realization metric (e.g., returned to the system by the third party media application in response to a realization request from the system), a third realization candidate that plays a Bach piece using a first party media application associated with a third realization metric (e.g., returned to the system by the first party media application in response to a realization request from the system), a fourth realization candidate that plays a Bach piece using a third party media application associated with a fourth realization metric (e.g., returned to the system by the third party media application in response to a realization request from the system), and / or other realization candidates. In this example, assuming that the ASR metrics, NLU metrics, and / or realization metrics indicate that a first candidate realization of playing rock music using the first party's media application is most likely to satisfy the utterance, the automated assistant causes the first party's media application to begin playing rock music as a particular realization of the assistant command. In this example, the system can infer slot values ​​(e.g., as part of the NLU output) for other parameters associated with the intent of "play music" and for two interpretations.For example, for a first interpretation of "play rock music," the system may infer an artist slot value for the artist parameter (e.g., the rock artist the user listens to most), a software application slot value for the software application parameter (e.g., the application the user most often uses to listen to music), a song slot value for the song slot parameter (e.g., the rock song the user listens to most, or the Bach composition they most often listen to), etc., based on user profile data (e.g., stored in user profile database 110A of client device 110 of FIG. 1), if accessible by the system. Otherwise, the system may utilize default slot values ​​for one or more of these parameters. In other examples, the user may specify slot values ​​for one or more of these parameters, such as a particular artist, a particular song, a particular software application, etc.

[0041] In some implementations, block 256 may include sub-block 256A. If included, in sub-block 256A, the system causes one or more selectable elements associated with a particular rationale for a particular realization to be provided for presentation to the user. The particular rationale for a particular realization may include, for example, one or more reasons why a particular candidate realization was selected from one or more candidate realizations in response to user input. In response to receiving a user selection of one or more selectable elements from a user of a client device, the system may cause the particular rationale to be provided for presentation to the user of the client device. In some variations of these implementations, the system can provide one or more selectable elements associated with a particular rationale for a particular realization in response to determining that ASR metrics, NLU metrics, and / or realization metrics associated with a particular candidate realization did not satisfy a metric threshold. In other words, in response to the system determining that a particular selected candidate realization is most likely what the user intended, but that the system is not highly confident about the particular selected candidate realization, the system can provide one or more selectable elements for presentation to the user that are associated with a particular rationale for the particular realization. In other types of these implementations, the system can provide one or more selectable elements for presentation to the user that are associated with a particular rationale for the particular realization, regardless of the ASR metrics, NLU metrics, and / or realization metrics associated with the particular selected candidate realization. As described herein (e.g., see FIG. 5B ), one or more selectable elements can be associated with a general request or one or more corresponding specific requests.

[0042] In block 258, the system determines whether a request for a particular rationale for why the automated assistant performed a particular realization has been received by the client device. In some implementations, the request for a particular rationale for why the automated assistant performed a particular realization may be included in additional user input received by the client device. In some variations of these implementations, the system can process the additional user input using an ASR model to generate ASR output, an NLU model to generate NLU data, and / or a realization rule or model to generate realization output in the same or similar manner as described with reference to block 252, and can determine whether the additional user input includes a request for a particular rationale. For example, the additional user input, whether spoken or keyed, can be processed to generate an NLU output, and the system can determine whether the additional user input includes a request for a particular rationale for why the automated assistant performed a particular realization based on the NLU input (e.g., an intent associated with the request for a particular rationale). In additional or alternative versions of such implementations, a request for a particular rationale for why the automated assistant performed a particular implementation may be included in a user selection of one or more selectable elements received at the client device (e.g., as described above with reference to sub-block 256A). If, in an iteration of block 258, the system determines that a request for a particular rationale for why the automated assistant performed a particular implementation has not been received at the client device, the system continues to monitor for a request for a particular rationale for why the automated assistant performed a particular implementation at block 258 (possibly for a threshold time (e.g., 5 seconds, 10 seconds, 15 seconds, and / or any threshold time) after the particular implementation was performed).If, in the iteration of block 258, the system determines that a request has been received by the client device for a particular justification as to why the automated assistant performed a particular implementation, the system proceeds to block 260.

[0043] In block 260, the system processes the additional user input, including the request, to determine additional data to use in providing a particular rationale. The additional data determined to be used in providing a particular rationale may be based, for example, on the type of request included in the additional user input. Thus, in block 262, the system determines the type of request for a particular rationale. The type of request may be, for example, a general request for a particular rationale for why the automated assistant performed a particular implementation (e.g., "Why did you do that?") or a specific request for a particular rationale for why the automated assistant performed a particular implementation (e.g., "Why did you select the first party music application?", "Why didn't you select the third party music application?", "Why did you select that artist?", "Why did you select the artist's best-known song?", and / or other specific requests). For example, additional user input can be processed to generate NLU output, and the system can determine whether the additional user input includes a request for a particular justification as to why the automated assistant performed a particular realization based on the NLU input (e.g., an intent associated with a general request for a particular justification and / or an intent associated with a specific request for a particular justification).

[0044] If, in the iteration of block 262, the system determines that the type of request for a particular rationale is a general request for a particular rationale, the system proceeds to block 264. In block 264, the system determines the first additional data as additional data in providing the particular rationale. Continuing with the above example of "play rock music," the general request for a particular rationale can be embodied, for example, by the request "Why did you do that?", where "so" indicates a selected particular realization, such as playing rock music using the first party's media application with a particular rock artist and a selected particular rock song. In response to determining that the request type is a general request, the system can obtain first data associated with a particular candidate realization as additional data corresponding to the output of, for example, "You are telling us about your application usage and you selected the first party's application because you use it most often to listen to music," "You are telling us about your music preferences and you selected a certain artist because you listen to that artist the most," "You chose a certain song because it is the artist's most well-known song," and / or other reasoning associated with why the system selected a particular candidate realization (or inferred particular slot value) in response to the user's input.

[0045] If, in the iteration of block 262, the system determines that the type of request for a particular rationale is a specific request for a particular rationale, the system proceeds to block 266. In block 266, the system determines the second additional data as additional data in providing the particular rationale. Continuing with the above example of "play rock music," the specific request for a particular rationale may be embodied, for example, by requests such as "Why did you choose the first party's music application?", "Why not a third party's music application?", "Why did you choose that artist?", "Why did you choose that artist's best-known song?", and / or other specific requests. In response to determining that the request type is a specific request, the system can obtain second data associated with a particular candidate realization or another candidate realization included in one or more candidate realizations as additional data, for example, "You are telling us about your application usage and you selected the first party application because you use it most for listening to music," "You did not select a third party application because it does not provide access to third party applications," "You are telling us about your music preferences and you selected an artist because you listen to that artist the most," "You chose a song by an artist because it is their most well-known song," and / or other rationalization output associated with why the system selected a particular candidate realization or did not select a particular other candidate realization in response to user input.

[0046] In block 268, the system causes the automated assistant to utilize the additional data to provide output including the particular rationale for presentation to the user. In some implementations, the output including the particular rationale may include synthesized speech audio data, including synthesized speech incorporating the particular rationale characterized by the additional data. In some variations of these implementations, the system may process text corresponding to the particular rationale, generated based on metadata associated with one or more candidate realizations, with a TTS model (e.g., stored in ML model database 120A of FIG. 1 ), to generate synthesized speech audio data that is audibly playable for presentation to the user through a speaker of the client device or an additional client device in communication with the client device. In additional or alternative implementations, the output including the particular rationale may include text or other graphical content that is rendered for presentation to the user by a display of the client device or an additional client device in communication with the client device. The system returns to block 252 to perform additional iterations of the method 200 of FIG. 2 in response to receiving additional user input directed to the automated assistant, including additional assistant commands.

[0047] Referring to FIG. 3 , a flowchart illustrating an example method 300 for determining that a particular implementation of an assistant command included in user input and directed to an automated assistant is infeasible is shown. For convenience, the operations of method 300 are described with reference to a system that performs those operations. The system of method 300 includes one or more processors, memory, and / or other components of a computing device (e.g., client device 110 of FIG. 1 , client device 410 of FIGS. 4A-4C , client device 510 of FIGS. 5A-5B , computing device 610 of FIG. 6 , one or more servers, and / or other computing devices). Furthermore, although the operations of method 300 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, and / or added.

[0048] In block 352, the system receives user input from a user of the client device, including an assistant command, directed to the automated assistant. In some implementations, the user input may correspond to speech captured in audio data generated by a microphone on the client device. In additional or alternative implementations, the user input may correspond to touch or key input received through a display on the client device or other input devices on the client device (e.g., a keyboard and / or a mouse).

[0049] In block 354, the system determines whether data to be utilized in executing a particular realization of the assistant command can be determined. In implementations in which the user input corresponds to speech, audio data capturing the speech can be processed using an ASR mode to generate ASR output (e.g., speech hypotheses, phonemes, and / or other ASR output) and, in some cases, ASR metrics associated with the ASR output. The ASR output can then be processed using an NLU model to generate NLU output (e.g., an intent determined based on the ASR output, slot values ​​for parameters associated with the intent determined based on the ASR output, etc.) and, in some cases, NLU metrics associated with the NLU output. The NLU output can then be processed using realization rules and / or a realization model to generate realization data to be utilized in executing a realization of the assistant command and, in some cases, realization metrics associated with the realization data. In implementations where user input corresponds to touch input or key input, the text corresponding to the touch input or key input can be processed using an NLU model to generate NLU output and, in some cases, NLU metrics associated with the NLU output. Further, the NLU output can be processed using realization rules and / or a realization model to generate realization output used to generate an outgoing request, obtain data used in executing the realization of an Assistant command, and / or perform an action based on the realization output in executing the realization of an Assistant command.

[0050] The system can determine whether data to be used in executing a particular realization of the assistant command is determined based on data received in response to the transmitted realization output, obtain data to be used in executing a particular realization of the assistant command, and / or take an action to be performed based on the realization output in executing the realization of the assistant command. As described above with reference to FIG. 2, in response to the candidate realizations, the obtainable data can include one or more candidate realizations. The system can select a particular candidate realization from the one or more candidate realizations and execute a particular realization of the assistant command based on, for example, ASR metrics, NLU metrics, and / or realization metrics associated with the user input. For example, assume that a user of a client device provides the utterance "Turn on the lights." In this example, audio data capturing the utterance can be processed using an ASR model to generate, as an ASR output, a speech hypothesis for "Turn on the lights" associated with an ASR metric (e.g., a probability, a binary value, a likelihood, such as a log-likelihood, that the first speech hypothesis corresponds to a term and / or phrase included in the utterance). Further, each speech hypothesis can be processed with an NLU model to generate, as NLU data, the intent "turn on the lights" associated with NLU metrics (e.g., likelihood that the first intent and slot values ​​correspond to the user's desired intent, such as probability, binary value, log-likelihood, etc.). Further, the realized output can include, for example, a request to one or more software applications (e.g., first party and / or third party software applications capable of turning on the lights). However, in this example, it is assumed that the automated assistant does not have access to the software application used to control the lights.Thus, in this example, the system need not be able to send the fulfillment output to any software application; one or more fulfillment candidates may include only a null fulfillment candidate, indicating that the utterance is not fulfillable because the automated assistant is unable to interact with the software application that controls the lights. On the other hand, assuming that the automated assistant does not have access to the software application used to control the lights, the system may determine that the automated assistant is able to interact with the software application that controls the lights and therefore can determine data to use for a particular fulfillment of the assistant command; one or more fulfillment candidates may include one or more assistant commands that, when executed, cause the lights to be controlled.

[0051] If, in an iteration of block 354, the system determines that the data to be used in executing a particular implementation of the assistant command can be determined, the system proceeds to block 256 of FIG. 2 and continues the iteration of method 200 of FIG. 2. For example, in the example described above where the automated assistant has access to a software application used to control the lights, the system can proceed to block 256 of FIG. 2 and continue the iteration of method 200 of FIG. 2 described above from block 256, causing a particular implementation of the assistant command to be executed and providing a particular rationale for the particular implementation if requested as described above with reference to FIG. 2. If, in an iteration of block 354, the system determines that the data to be used in executing a particular implementation of the assistant command cannot be determined, the system proceeds to block 356. For example, in the example described above where the automated assistant has access to a software application used to control the lights, the system can proceed to block 356.

[0052] In block 356, the system determines whether alternative data to be used to execute an alternative realization of the assistant command can be determined. For example, the system can analyze one or more candidate realizations to determine whether there are one or more alternative realizations. If, in a repeat of block 356, the system determines that alternative data to be used to execute an alternative realization of the assistant command cannot be determined, the system proceeds to block 358. For example, continuing with the example above where a user of a client device provides the utterance "turn on the lights" and the automated assistant does not have access to any software applications that control lights, the system can determine that one or more candidate realizations do not include alternative realizations (e.g., only empty candidate realizations). In this example, the system can decide to proceed to block 358.

[0053] In block 358, the system processes the user input to determine recommendation data to be used to generate a recommended action for how the automated assistant can perform a particular implementation. Continuing with the example above, the system may determine that a particular implementation of lighting control can be performed in response to the utterance "Turn on the lights," but the user has not actually granted the automated assistant access to the software application used to control the lights. Thus, in this example, the recommended action may include content indicating that the user should grant the automated assistant access to the software application used to control the lights. As another example, assuming that a software application is not installed on a client device that controls the lights, the system may determine that a particular implementation of lighting control can be performed in response to the utterance "Turn on the lights," but the user has not actually installed the software application used to control the lights and the user must grant the automated assistant access to the software application used to control the lights.

[0054] In block 360, the system has the automated assistant utilize the recommendation data to provide an output including a recommended action for presentation to the user. The output including the recommended action can be played audibly and / or visually for presentation to the user (e.g., as described above with reference to block 268 of FIG. 2). In some implementations, the output including the recommended action can include a prompt that allows the user to provide additional input that causes the automated assistant to automatically perform the recommended action. Continuing with the example above, the system may generate the output, "I can't turn on the light right now, but I can if you allow me access to the software application used to control the light. Do you allow me access?" Thus, for presentation to the user, the output including the recommended action can indicate one or more of the specific implementations of the assistant command that the automated assistant can perform (e.g., "I can't turn on the light right now, but I can if you allow me access to the software application used to control the light.") and prompt the user to perform a specific implementation of the assistant command (e.g., "Do you allow me access?"). In additional or alternative implementations, the output including the recommended action may include step-by-step instructions for the user to follow to enable the automated assistant to perform a particular realization of the assistant command (e.g., "(1) Open Settings, (2) Open Software Application Share Settings, (3) Share Software Application Settings for Lighting Application"). The system returns to block 352 and performs additional iterations of method 300 of FIG. 3 in response to receiving yet another user input directed to the automated assistant, including yet another assistant command.

[0055] If, in an iteration of block 356, the system determines that alternative data can be determined to utilize in executing an alternative realization of the assistant command, the system proceeds to block 362. Meanwhile, in the example in which the user provides the utterance "Turn on the lights," it is assumed that the user also provides the utterance "Play rock music in Application 2," which is received in block 352. It is further assumed that the user has not granted the automated assistant access to "Application 2." Thus, in this example, a particular realization of playing rock music in "Application 2" is not executable in the instance of block 354. However, in the instance of block 356, unlike the previous example, the system may determine that there is an alternative realization candidate. For example, in this example, it is further assumed that the user has granted the automated assistant access to "Application 1," so that the automated assistant can alternatively utilize "Application 1" to play rock music.

[0056] In block 362, the system processes the user input to determine additional data to utilize in executing another realization of the assistant command. Continuing with the above example in which the user provides the utterance, "Play rock music in Application 2," a realization output including a structured request to play rock music can be initially sent to at least "Application 2." In some implementations, in response to sending the realization output to "Application 2," the system may receive content for an empty realization candidate because the user has not authorized the automated assistant to access "Application 2." In additional or alternative implementations, the system may determine that the user has not authorized the automated assistant to access "Application 2," and the system may refrain from sending a request to "Application 2" and determine an empty realization candidate because the user has not authorized the automated assistant to access "Application 2." However, in attempting to execute the realization of the request to play music, the system may send the realization output to "Application 1" (possibly in response to determining that "Application 2" is associated with the empty realization candidate) and determine additional data associated with another realization candidate to play rock music in "Application 1" because the user has authorized the automated assistant to access "Application 1." Thus, even if the alternative realization candidate is not a particular realization of the assistant command included in the user input (for example, because the alternative realization candidate is associated with "Application 1" rather than "Application 2" as requested by the user), the system can attempt to satisfy the assistant command using the alternative realization candidate.

[0057] In block 364, the system causes the automated assistant to execute another realization of the assistant command using the other data. Continuing with the example above, the system may cause “Application 1” to begin playing rock music through the speaker of the client device or through the speaker of another computing device communicating with the client device (e.g., a smart speaker communicating with the client device, other client devices, etc.). In some implementations, similar to sub-block 256A of FIG. 2 , the system may provide one or more selectable elements associated with a particular rationale for presentation to the user. However, unlike the above-described operations of sub-block 256A of FIG. 2 , a particular rationale may be provided for why another realization was executed (e.g., why rock music was played using “Application 1”) or why a particular realization was not executed (e.g., why rock music was not played using “Application 2”). In these implementations, a particular rationale for the alternative realization may include, for example, one or more reasons why the alternative realization was selected from one or more alternative realizations in response to user input.

[0058] In block 366, the system determines whether a request for a particular rationale for why the automated assistant performed another realization has been received by the client device. In some implementations, the request for a particular rationale for why the automated assistant performed another realization may be included in additional user input received by the client device. In some variations of these implementations, the system can process the additional user input using an ASR model to generate ASR output, an NLU model to generate NLU data, and / or a realization rule or model to generate realization output in the same or similar manner as described with reference to block 252, and can determine whether the additional user input includes a request for a particular rationale. For example, the additional user input, whether spoken or keyed, can be processed to generate an NLU output, and the system can determine whether the additional user input includes a request for a particular rationale for why the automated assistant performed another particular realization based on the NLU input (e.g., the intent associated with the request for a particular rationale). In additional or alternative versions of such implementations, a request for a particular rationale for why the automated assistant executed another realization may be included in the user selection of one or more selectable elements received at the client device (e.g., as described above with reference to subblock 256A). If, in an iteration of block 366, the system determines that a request for a particular rationale for why the automated assistant executed another realization has not been received at the client device, the system continues to monitor for a request for a particular rationale for why the automated assistant executed another realization at block 366 (possibly for a threshold time (e.g., 15 seconds, 20 seconds, 30 seconds, and / or any other threshold time) after the execution of a particular realization).If, in the iteration of block 366, the system determines that a request for a specific rationale for why the automated assistant executed another realization was received by the client device, the system proceeds to block 368. In some implementations, similar to block 262 of FIG. 2, the system can determine the type of request for a specific rationale (e.g., a general request, a first specific request, a second specific request, etc.).

[0059] In block 368, the system processes additional user input, including the request, to determine additional data to use in providing a particular rationale. Continuing with the "Play rock music on Application 2" example, a general request for a particular rationale can be embodied, for example, by the request "Why did you do that?", where "Yes" indicates a selected alternative realization, such as playing rock music using "Application 1" (instead of "Application 2" at the user's request) with a particular rock artist and a selected particular rock song (e.g., an inferred slot value, as described above with reference to FIG. 2). In response to determining that the request type is a general request, the system can obtain additional data associated with the additional data corresponding to the output in response to the general request (e.g., as described above with reference to FIG. 2). Continuing with the "Play rock music" example, a specific request for a particular rationale can be embodied, for example, by the request "Why did you use Application 1 instead of the requested Application 2?", "Why did you choose that artist?", "Why did you choose the artist's best-known song?", and / or other specific requests. In response to determining that the request type is a specific request, the system can obtain additional data corresponding to output in response to the specific request (eg, as described above with reference to FIG. 2).

[0060] In block 370, the system causes the automated assistant to utilize the additional data to provide output including the particular rationale for presentation to the user. In some implementations, the output including the particular rationale may include synthesized speech audio data, including synthesized speech incorporating the particular rationale characterized by the additional data. In some variations of these implementations, the system may process text corresponding to the particular rationale, generated based on metadata associated with one or more candidate realizations, with a TTS model (e.g., stored in ML model database 120A of FIG. 1 ), to generate synthesized speech audio data that is audibly playable for presentation to the user through a speaker of the client device or an additional client device in communication with the client device. In additional or alternative implementations, the output including the particular rationale may include text or other graphical content rendered for presentation to the user on a display of the client device or an additional client device in communication with the client device. The system returns to block 352 to perform additional iterations of the method 300 of FIG. 3 in response to receiving additional user input directed to the automated assistant that includes additional assistant commands.

[0061] 4A-4C , various non-limiting examples are shown that provide certain rationales for implementing assistant commands. Client device 410 (e.g., an instance of client device 110 in FIG. 1 ) may include various user interface components, including, for example, a microphone that generates audio data based on speech and / or other auditory input and / or a speaker that audibly reproduces synthesized speech and / or other auditory output. While client device 410 shown in FIGS. 4A-4C is a standalone speaker without a display, it should be understood that this is for illustrative purposes only and not meant to be limiting. For example, client device 410 may be a standalone speaker with a display, a mobile phone (e.g., as described above with reference to FIGS. 5A-5B ), a home automation device, an in-vehicle system, a laptop, a desktop computer, and / or any other device capable of running an automated assistant and engaging in a human-computer interaction session with user 401 of client device 410.

[0062] 4A , assume that user 401 of client device 410 provides utterance 452A, "Assistant, play rock music." In response to receiving utterance 452A, the automated assistant can use an ASR model to process audio data capturing utterance 452A to generate ASR output, including, for example, one or more predicted speech hypotheses (e.g., term hypotheses and / or transcription hypotheses) corresponding to utterance 452A, one or more predicted phonemes corresponding to utterance 452A, and / or other ASR outputs. In generating the ASR output, the ASR model can optionally generate ASR metrics associated with each of the one or more speech hypotheses, predicted phonemes, and / or other ASR outputs, the ASR metrics indicating the likelihood that the one or more speech hypotheses, predicted phonemes, and / or other ASR outputs correspond to utterance 452A. Further, the ASR output can be processed using an NLU model to generate NLU output including, for example, one or more intents determined based on the ASR output, one or more slot values ​​for one or more corresponding parameters associated with each of the one or more intents determined based on the ASR output, and / or other NLU output. In generating the NLU output, the NLU model can optionally generate NLU metrics associated with each of the one or more intents, one or more slot values ​​for the corresponding parameters associated with the intents, and / or other NLU output, where the NLU metrics indicate the likelihood that the one or more intents, one or more slot values ​​for the corresponding parameters associated with the intents, and / or other NLU output correspond to the actual intent of user 401 in providing utterance 452A.

[0063] In particular, the automated assistant may infer one or more slot values ​​for corresponding parameters associated with each of one or more intents, resulting in one or more interpretations of utterance 452A, each of which includes at least one unique slot value for a given corresponding parameter. Thus, in the example of FIG. 4A , a first interpretation may include a “play music” intent, which may include a slot value of “Application 1” for an application parameter associated with the “play music” intent, a slot value of “Artist 1” for an artist parameter associated with the “play music” intent, and a slot value of “Song 1” for a song parameter associated with the “play music” intent. A second interpretation may include a “play music” intent, which may include a slot value of “Application 2” for an application parameter associated with the “play music” intent, a slot value of “Artist 1” for an artist parameter associated with the “play music” intent, and a slot value of “Song 1” for a song parameter associated with the “play music” intent. A third interpretation may include a "play music" intent, which may include a slot value of "application 1" in the application parameters associated with the "play music" intent, a slot value of "artist 2" in the artist parameters associated with the "play music" intent, and a slot value of "song 2" in the song parameters associated with the "play music" intent, and similarly for other interpretations.

[0064] Further, the automated assistant can process the NLU output using realization rules and / or realization models to generate realization outputs. The realization outputs can include, for example, one or more structuring requests, which are generated based on multiple interpretations (e.g., determined based on the NLU output) and sent to one or more realizing parties, such as first party server 191 of FIG. 1, third party server 192 of FIG. 1, a first party software application accessible to client device 410, a third party software application accessible to client device 410, and / or any other realizing party capable of realizing utterance 452A. In the example of FIG. 4A, the automated assistant can send corresponding structuring requests to at least “Application 1” and “Application 2” based on the software applications identified as capable of satisfying utterance 452A indicated by the NLU output. In response to sending these structuring requests, the automated assistant can receive one or more candidate realizations from “Application 1” and “Application 2.” For example, the automated assistant can receive one or more candidate realizations from "Application 1" indicating whether "Application 1" can realize one or more of the structured requests generated based on the multiple interpretations, and can receive one or more candidate realizations from "Application 2" indicating whether "Application 2" can realize one or more of the structured requests generated based on the multiple interpretations. The one or more candidate realizations may optionally include realization metrics indicating how likely each of the one or more candidate realizations is to satisfy utterance 452A.

[0065] The automated assistant can rank one or more realization candidates based on ASR metrics, NLU metrics, and / or realization metrics and select a particular realization candidate from the one or more realization candidates based on the ranking. For example, in the example of FIG. 4A , as described above, the automated assistant selects a particular realization candidate associated with a first interpretation including the intent to “play music” based on the ranking, where the intent to “play music” is assumed to have a slot value of “Application 1” for the application parameter, a slot value of “Artist 1” for the artist parameter, and a slot value of “Song 1” for the song parameter. Furthermore, it is assumed that user 401 has not authorized the automated assistant to access “Application 2,” and as a result, one or more realization candidates determined based on one or more structured requests and sent to “Application 1” are empty realization candidates. In additional or alternative implementations, the automated assistant may refrain from sending structured requests to “Application 2” to conserve computer and / or network resources. This is because the automated assistant knows that user 401 has not granted the automated assistant access to “Application 2” and has automatically determined null realization candidates for any structured request that may be sent to “Application 2.” In response to selection of a particular realization candidate associated with the first interpretation, the automated assistant may cause synthesized speech 454A1, “Okay, play rock music in Application 1,” to be audibly presented to user 401 via a speaker on client device 410, and may cause an assistant command determined based on the first interpretation associated with the particular realization candidate to be implemented as shown in 452A (e.g., “Play Song 1 by Artist 1 in Application 1”) to satisfy utterance 452A.

[0066] However, it is further assumed that user 401 provides additional utterance 456A, such as, "Why did you do that?" In response to receiving additional utterance 456A, the automated assistant can process the audio data incorporating additional utterance 456A using an ASR model to generate an ASR output the same as or similar to that described above for processing utterance 452A. Further, the ASR output can be processed using an NLU model to generate an NLU output the same as or similar to that described above for processing utterance 452A. The automated assistant can determine, based on the ASR output and / or the NLU output, whether additional utterance 456A includes a request to the automated assistant to provide a specific rationale for why a specific implementation of an assistant command included in utterance 452A was performed. In some implementations, the automated assistant can additionally or alternatively determine, based on the ASR output and / or the NLU output, whether the request for a specific rationale is a general request for a specific rationale or one or more specific requests for a specific rationale.

[0067] For example, in the example of FIG. 4A , the automated assistant may determine that the request for a particular rationale is a general request based on the ASR output and / or NLU output generated based on processing of additional utterance 456A. ​​This is because the user is not inquiring about a particular aspect of a particular realization (e.g., as described below with reference to FIG. 4B ). Instead, additional utterance 456A generally asks the automated assistant to explain why it is having “Application 1” play “Song 1” by “Artist 1.” Thus, the automated assistant can obtain metadata associated with the particular selected realization and determine that the additional data is to be used to generate output in response to additional utterance 456A. ​​Based on the additional data, the automated assistant can provide additional synthesized speech 458A1, such as, “You are telling me about your use of applications, you seem to use Application 1 most often for listening to music, you have listened to Artist 1 in the past, and Song 1 is a new release by Artist 1,” for audible presentation to user 401 via the speaker of client device 410. In some implementations, the automated assistant can optionally provide a prompt 458A2 (e.g., determined based on recommendation data) for auditory presentation to user 401 through a speaker on client device 410, such as, for example, "Is there anything else I can do?" This request prompts user 401 to provide an additional utterance if user 401 wants the automated assistant to execute a different implementation, such as by requesting user 401 to provide an additional utterance 460A, such as, "Use application 2, but only for rock music."

[0068] In particular, another additional utterance 460A may cause the automated assistant to access and utilize “Application 2” in response to utterance 452A and may implicitly authorize future instances of utterances that include an assistant command that causes the automated assistant to play rock music. In some implementations, another additional utterance 460A may implicitly authorize the automated assistant to access and utilize “Application 2” only to play rock music. In additional or alternative implementations, another additional utterance 460A may implicitly authorize the automated assistant to access and utilize “Application 2” to play music of any genre. In some implementations, the automated assistant can transition from playing “Song 1” by “Artist 1” using “Application 1” (e.g., an assistant command based on the first interpretation described above) to playing “Song 1” by “Artist 1” using “Application 2” (e.g., an assistant command based on the second interpretation described above) in response to receiving another utterance 460A that authorizes the automated assistant to access “Application 2.”

[0069] 4B , unlike the general request for a particular rationale described with reference to FIG. 4A , it is also contemplated that the user may provide the same utterance 452B, "Assistant, play rock music," and the automated assistant may provide synthesized speech 454B1, "Okay, play rock music in Application 1," for auditory presentation to user 401 through the speaker of client device 410, and have the automated assistant implement an assistant command determined based on a first interpretation associated with a particular candidate realization, as indicated by 454B2 (e.g., "Play song 1 by artist 1 in Application 1"), thereby satisfying utterance 454B2. However, in the example of FIG. 4B , it is contemplated that user 401 provides an additional utterance, "Why did you use Application 1?" In this example, the automated assistant may determine that the request for a particular rationale is a specific request based on the ASR output and / or NLU output generated based on processing the additional utterance 456B. This is because the user is inquiring about a particular aspect of a particular realization (e.g., why did the automated assistant select "Application 1" to play "Song 1" by "Artist 1"). Thus, the automated assistant can obtain metadata associated with the particular selected realization and determine that additional data is to be used to generate output in response to additional utterance 456A. ​​Notably, in the example of FIG. 4B , the additional data may differ from the additional data in the example of FIG. 4A in that the additional data used in the example of FIG. 4B may be adapted or tailored to the particular request of user 401, particularly to inquire about why "Application 1" was selected. Based on the additional data, the automated assistant can provide additional synthesized speech 458B1, "You're telling me about your application usage, and it seems like you use Application 1 most for listening to music," for audible presentation to user 401 through the speaker of client device 410.In some implementations, the automated assistant can optionally provide prompt 458B2 (e.g., determined based on recommendation data) for auditory presentation to user 401 through a speaker of client device 410, such as "Is there anything else I can do?", which prompts user 401 to provide another, additional utterance if user 401 wants the automated assistant to execute some other possible implementation, such as by requesting user 401 to provide another, additional utterance 460B, such as "Use application 2, but only with rock music." Similar to what was discussed above with reference to FIG. 4A , another, additional utterance 460B may cause the automated assistant to access and utilize "application 2" in response to utterance 452B, and implicitly authorize future instances of utterances that include an assistant command that causes the automated assistant to play rock music.

[0070] 4C , unlike in FIGS. 4A and 4B , it is assumed that the user provides the same utterance 452C, “Assistant, play rock music.” However, in the example of FIG. 4C , it is assumed that user 401 has not authorized the automated assistant to access any software application or server capable of playing music (e.g., a streaming service implemented by one or more of first-party server 191 of FIG. 1 and / or third-party server 192 of FIG. 1 ). As such, any realization candidate included in one or more realization candidates associated with “Application 1,” “Application 2,” or other software applications or services may be associated with null realization candidates. Nevertheless, in an attempt to satisfy utterance 452C, the automated assistant may send a structured request, for example, to a web browser, and may retrieve content responsive to utterance 452C, such as search results for “rock music,” to avoid wasting computer resources on interaction. As shown in FIG. 4C, the automated assistant can provide synthesized speech 454C, "Rock music is a broad genre of popular music derived from 'rock and roll'...", for auditory presentation to user 401 through the speaker of client device 410.

[0071] However, it is further assumed that user 401 provides additional utterance 456C, such as, "Why didn't you play music?" In response to receiving additional utterance 456C, the automated assistant can process the audio data incorporating additional utterance 456C using an ASR model to generate an ASR output the same as or similar to that described above for processing utterance 452A in FIG. 4A. Further, the ASR output can be processed using an NLU model to generate an NLU output the same as or similar to that described above for processing utterance 452A in FIG. 4A. The automated assistant can determine, based on the ASR output and / or the NLU output, whether additional utterance 456C includes a request to provide the automated assistant with a particular rationale for why a particular implementation of the assistant command included in utterance 452A was performed. Notably, in the example of FIG. 4C, rather than user 401 inquiring about why the automated assistant performed a particular implementation, user 401 is inquiring about why the automated assistant did not perform a particular implementation.

[0072] Thus, in the example of FIG. 4C , the automated assistant may determine that a request for a particular rationale is a specific request based on the ASR output and / or NLU output generated based on processing additional utterance 456A. ​​This is because the user is inquiring about a particular aspect of a particular realization (e.g., why the automated assistant didn't play music). Therefore, the automated assistant may obtain metadata associated with one or more alternative realizations that were not selected to determine that additional data be utilized to generate the output in response to additional utterance 456C. Notably, in the example of FIG. 4C , the additional data may differ from the additional data in the examples of FIGS. 4A and 4B in that the additional data utilized in the example of FIG. 4C may be adapted or tailored to the specific request of user 401, particularly to inquire about why music wasn't played. Based on the additional data, the automated assistant can provide additional synthesized speech 458C1, such as, "You haven't authorized access to some applications or services, so they can't be used to play music," for audible presentation to user 401 through the speaker of client device 410. In some implementations, the automated assistant can optionally provide prompt 458C2 (e.g., determined based on the recommendation data) for audible presentation to user 401 through the speaker of client device 410, such as, "Do you want to allow access to the application or service?", which prompt requests user 401 to provide another additional utterance, such as, if user 401 wants the automated assistant to enable one or more software applications or services for music playback, user 401 providing additional utterance 460C, such as, "Yes, use application 2, but only for rock music."In the example of FIG. 4C, another additional utterance 460C may explicitly allow the automated assistant to access and utilize “Application 2” in response to utterance 452C and future instances of utterances that include an assistant command to cause the automated assistant to play rock music.

[0073] 5A-5B, various additional non-limiting examples are shown that provide a particular rationale for implementing an assistant command. Client device 510 (e.g., the example of client device 110 in FIG. 1) may include various user interface components, including, for example, a microphone that generates audio data based on speech and / or other auditory input, a speaker that audibly plays synthesized speech and / or other auditory output, and / or a display 580 that visually renders visual output. Additionally, display 580 of client device 510 may include various system interface elements 581, 582, and 583 (e.g., hardware and / or software interface elements) that interact with a user of client device 510 and cause client device 510 to perform one or more actions. The display 580 of the client device 510 allows a user to interact with content rendered on the display 580 through touch input (e.g., by directing user input to the display 580 or portions thereof (e.g., into a text entry box (not shown), into a keyboard (not shown), or into other portions of the display 580)) and / or through speech input (e.g., by selecting the microphone interface element 584, or by simply speaking at the client device 510 without having to select the microphone interface element 584 (i.e., the automated assistant may monitor one or more terms or phrases, gestures, gaze, mouth movements, lip movements, and / or other conditions to activate speech input)). It should be understood that the client device 510 shown in FIGS. 5A-5B is a mobile phone, but this is for purposes of illustration and not limitation.For example, client device 510 may be a standalone speaker with a display, a standalone speaker without a display (e.g., as described above with reference to Figures 4A-4C), a home automation device, an in-vehicle system, a laptop, a desktop computer, and / or any other device that runs an automated assistant and engages in a human-computer interaction session with a user of client device 510.

[0074] 5A , it is assumed that a user of client device 510 provides utterance 552A, "Play some rock music." In response to receiving utterance 552A, the automated assistant can process audio data capturing utterance 552A using an ASR model to generate an ASR output, process the ASR output using an NLU output to generate an NLU output, and process the NLU output using implementation rules and / or implementation models to generate the same or similar realized output. It is further assumed that the automated assistant determines that the user's client device 510 is communicatively connected to a smart speaker (e.g., a speaker in the living room), which has more stable speakers than client device 510 and is capable of playing rock music. Thus, based on processing utterance 552A, the automated assistant may determine that rock music is to be played on the living room speakers, and synthesized speech 554A, "OK, I'll play rock music on the living room speakers," may be provided for auditory presentation to the user through the speakers of client device 510 and / or for visual presentation to the user on display 580 of client device 510, and rock music may be played on the living room speakers.

[0075] However, it is further assumed that the user of client device 510 provides additional utterance 556A, such as, "Why did you decide to play music on the living room speakers?" In this example, the automated assistant may determine, based on the ASR output and / or NLU output generated based on processing additional utterance 556A, that the additional utterance contains a request for a particular rationale and is a specific request because the user is inquiring about a particular aspect of a particular realization (e.g., why did the automated assistant select "living room speakers" to play rock music). In this example, the automated assistant can obtain metadata associated with the selected candidate realization and determine additional data to use in providing a particular rationale for why the automated assistant decided to play music on the living room speakers rather than client device 510 or another computing device communicatively connected to client device 510 and capable of playing music (e.g., kitchen speakers, den speakers, etc.). Based on the additional data, the automated assistant can provide additional synthesized speech 558A1, such as "The speakers in the living room are more stable than the speakers on your phone" (possibly based on detecting the presence of the user of client device 510 in the living room (e.g., by presence sensor 111 of client device 110 in FIG. 1 )), for audible presentation to the user through the speakers of client device 510 and / or for visual presentation to the user on display 580 of client device 510.In some implementations, the automated assistant can optionally provide prompt 558A2 (e.g., determined based on recommendation data) for audible presentation to the user through the speaker of client device 510 and / or for visual presentation to the user on display 580 of client device 510, such as, for example, "Do you want to play rock music on your phone or another device?", with the prompt providing additional speech to the user if the user wants the automated assistant to perform some other possible implementation, such as playing rock music using client device 510, playing music using another software application accessible to client device 510, having the automated assistant switch artists / songs, etc. If the user provides some other input, the automated assistant can adapt the music playback accordingly.

[0076] In additional or alternative implementations, rather than waiting for the user of client device 510 to provide some additional user input 556A that includes a request for a particular rationale, the automated assistant can proactively provide one or more selectable elements associated with a particular rationale. For example, with reference to FIG. 5B , it is again assumed that the user of client device 510 provides utterance 552B, "Play some rock music," in response to receiving utterance 552. In response to receiving utterance 552B, the automated assistant can process audio data incorporating utterance 552B with an ASR model to generate an ASR output, process the ASR output with an NLU output to generate an NLU output, and process the NLU output with realization rules and / or realization models to generate the same or similar realization output. It is further assumed that the automated assistant determines that the user's client device 510 is communicatively connected to a smart speaker (e.g., a living room speaker), which again is assumed to have a more robust speaker than the client device 510 and is capable of playing rock music. Thus, based on processing utterance 552B, the automated assistant may determine that rock music is to be played on the living room speaker, and synthesized speech 554B of "OK, I'll play rock music on the living room speaker" may be provided for auditory presentation to the user through the speaker of client device 510 and / or for visual presentation to the user on display 580 of client device 510, and rock music may be played on the living room speaker.

[0077] 5B example, however, it is further contemplated that the automated assistant will proactively provide one or more selectable elements associated with a particular rationale without the user of client device 510 providing additional speech or other user input. For example, as shown in FIG. 5B, the automated assistant can create a first selectable element 556B1, "Why did you do that?" associated with a general request as to why the automated assistant performed a particular implementation of an assistant command in general, and / or a second selectable element 556B2, "Why did you use the living room speaker?" associated with a specific request as to why the automated assistant performed a particular implementation of an assistant command using the living room speaker. In response to a user selection of first selectable element 556B1 from a user of a client device (e.g., via touch input or speech input), the automated assistant can retrieve metadata associated with a particular realization of the assistant command based on the user selection (e.g., user selection of first selectable element 556B1 or user selection of second selectable element 556B2) to determine additional data to use in providing a particular rationale. For example, in response to receiving a user selection of first selectable element 556B1, the automated assistant can determine additional data about why the automated assistant decided to select a particular application that plays rock music (e.g., "Application 1" vs. "Application 2," as described above with reference to FIGS. 4A-4C), select a particular artist (e.g., as described above with reference to FIGS. 4A-4C), select a song by a particular artist (e.g., as described above with reference to FIGS. 4A-4C), play music on living room speakers (e.g., as described above with reference to FIG. 5A), and / or other aspects of a particular realization.For example, in response to receiving a user selection of second selectable element 556B2, the automated assistant can determine additional data regarding why the automated assistant is playing music on the living room speakers rather than on client device 510 or another computing device (e.g., kitchen speakers, study speakers, etc.) that is communicatively connected to client device 510 and capable of playing music.

[0078] 5B , it is assumed that the user of client device 510 provides a user selection of second selectable element 556B2. Based on additional data determined based on the user selection of second selectable element 556B2, the automated assistant can provide additional synthesized speech 558B1, "The speakers in the living room are more stable than the speakers on your phone" (possibly based on detecting the presence of the user of client device 510 in the living room (e.g., by presence sensor 111 of client device 110 in FIG. 1 )), for audible presentation to the user through the speaker of client device 510 and / or for visual presentation to the user on display 580 of client device 510. In some implementations, the automated assistant can optionally provide prompt 558B2 (e.g., determined based on recommendation data) for auditory presentation to the user through the speaker of client device 510 and / or for visual presentation to the user on display 580 of client device 510, such as, for example, "Do you want to play rock music on your phone or another device?", where the prompt requests the user to provide another additional utterance if the user wants the automated assistant to perform another possible implementation, such as playing rock music using client device 510, playing music using another software application accessible to client device 510, having the automated assistant switch artists / songs, etc. If the user provides another input, the automated assistant can adapt the music playback accordingly.

[0079] While the above-described examples of Figures 4A-4C and 5A-5B are described above with the automated assistant causing an implementation to be performed based on a particular utterance for a media application or media service and providing a particular rationale for the implementation in response to an additional particular utterance, it should be understood that this is for illustrative purposes only and not limiting. For example, the techniques described herein can be used to provide a particular rationale for any aspect of any implementation performed by an automated assistant and / or for any candidate implementation selected or not selected by the automated assistant. Furthermore, while the above-described examples of Figures 4A-4C and 5A-5B are described above with the automated assistant providing a particular recommended action determined based on certain recommendation data (e.g., allowing the automated assistant to access a particular software application for media playback), it should be understood that this is also for illustrative purposes only and not limiting. As a non-limiting example, the recommended actions described herein may include allowing the automated assistant to access any software applications, any user accounts, any computing devices associated with the user, query activity history, and / or any other user data that may be utilized to determine how to implement any assistant command by the automated assistant.

[0080] 6, a block diagram of an exemplary computing device 610 is shown that may be utilized to perform one or more aspects of the techniques described herein. In some implementations, one or more client devices, cloud-based automated assistant components, and / or other components may comprise one or more components of the exemplary computing device 610.

[0081] Typically, computing device 610 includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624, including a memory subsystem 625 and a file storage subsystem 626, user interface output devices 620, user interface input devices 622, and a network interface subsystem 616. The input / output devices allow a user to interact with computing device 610. Network interface subsystem 616 provides an interface to external networks and connects to corresponding interface devices in other computing devices.

[0082] The user interface input devices 622 may include keyboards, pointing devices such as a mouse, trackball, touchpad, or graphics tablet, scanners, touchscreens integrated into displays, voice recognition systems, audio input devices such as microphones and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways of inputting information into the computing device 610 or computer network.

[0083] The user interface output device(s) 620 may include a display subsystem, a printer, a fax, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projector device, or some other mechanism for creating visual images. The display subsystem may also provide non-visual displays such as an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and manners for outputting information from the computing device 610 to a user or to other devices or computing devices.

[0084] Storage subsystem 624 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 624 may include logic that performs selected aspects of the methods disclosed herein and implements the various components shown in Figures 1 and 2.

[0085] These software modules are typically executed by the processor 614, alone or in combination with other processors. The memory 625 used in the storage subsystem 624 may comprise several memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 containing fixed instructions. The file storage subsystem 626 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular implementation may be stored by the file storage subsystem 626, within the storage subsystem 624, or in other devices accessible by the processor 614.

[0086] Bus subsystem 612 provides a mechanism that allows the various components and subsystems of computer device 610 to communicate with each other as intended. Although bus subsystem 612 is shown schematically as a single bus, alternative implementations of bus subsystem 612 may use multiple buses.

[0087] The computing device 610 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 610 shown in Figure 6 is intended only as a specific example to illustrate some implementations. Many other configurations of the computing device 610 are possible, including those with more or fewer components than the computing device shown in Figure 6.

[0088] In situations where the systems described herein may collect or monitor personal information about a user or utilize personal and / or monitored information, the user may be provided with an opportunity to control whether a program or feature collects user information (e.g., the user's social network, social behavior or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how to receive content from content servers that may be more relevant to the user. Also, certain data may be processed in one or more ways before being stored or used to remove personally identifiable information. For example, the user's identity may be processed such that personally identifiable information cannot be identified about the user, or the user's geographic location from which geographic location information is obtained may be generalized (to the city, zip code, or state level) so that the user's specific geographic location cannot be identified. In this manner, the user may have control over how information is collected about and / or used.

[0089] In some implementations, a method implemented by one or more processors is provided that includes receiving user input from a user of a client device, the user input including an assistant command and directed to an automated assistant executing at least partially on the client device; processing the user input to determine data to utilize in executing a particular implementation of the assistant command; causing the automated assistant to utilize the data to execute the particular implementation of the assistant command; receiving additional user input from the user of the client device, the additional user input including a request to the automated assistant to provide a particular rationale as to why the automated assistant executed the particular implementation of the assistant command; processing the additional user input to determine additional data to utilize in providing the particular rationale as to why the automated assistant executed the particular implementation of the assistant command; and causing the automated assistant to utilize the additional data to provide output for presentation to the user of the client device, the output including the particular rationale as to why the automated assistant executed the particular implementation of the assistant command.

[0090] These and other implementations of the disclosed technology may optionally include one or more of the following features.

[0091] In some implementations, user input including assistant commands and directed to the automated assistant may be captured in audio data generated by one or more microphones of the client device. In some variations of these implementations, processing the user input to determine data to use in performing a particular implementation of the assistant command may include processing the audio data capturing the user input including the assistant command using an automatic speech recognition (ASR) model to generate ASR output, processing the ASR output using a natural language understanding (NLU) model to generate NLU output, and determining the data to use in performing a particular implementation of the assistant command based on the NLU output.

[0092] In some implementations, user input, including assistant commands directed to the automated assistant, may be captured in keystrokes detected via a display on the client device. In some variations of these implementations, processing the user input to determine data to use in performing a particular implementation of the assistant command may include processing the keystrokes using a natural language understanding (NLU) model to generate NLU output, and generating the data to use in performing a particular implementation of the assistant command based on the NLU output.

[0093] In some implementations, the request to the automated assistant to provide a particular rationale for why the automated assistant executed a particular implementation of the assistant command may include a specific request to the automated assistant to provide a particular rationale for why the automated assistant selected a particular software application from multiple heterogeneous software applications used to execute the particular implementation. In some variations of these implementations, processing the additional user input to determine additional data to use in providing a particular rationale for why the automated assistant executed a particular implementation of the assistant command may include obtaining metadata associated with the particular software application used to execute the particular implementation, and determining, based on the metadata associated with the particular software application, the additional data to use in providing a particular rationale for why the automated assistant executed the particular implementation of the assistant command.

[0094] In some implementations, the request to the automated assistant to provide a particular rationale for why the automated assistant performed a particular implementation of the assistant command may include a specific request to the automated assistant to provide a particular rationale for why the automated assistant selected a particular interpretation of the user input from multiple disparate interpretations of the user input to use in performing the particular implementation. In some variations of these implementations, processing the additional user input to determine additional data to use in providing a particular rationale for why the automated assistant performed a particular implementation of the assistant command may include obtaining metadata associated with a particular interpretation of the user input to use in performing the particular implementation, and determining the additional data to use in providing a particular rationale for why the automated assistant performed the particular implementation of the assistant command based on the metadata associated with the particular interpretation of the user input.

[0095] In some implementations, the request to the automated assistant to provide a particular rationale for why the automated assistant executed a particular implementation of the assistant command may include a specific request to the automated assistant to provide a particular rationale for why the automated assistant selected an additional client device of the user instead of the user's client device to use in executing the particular implementation. In some variations of these implementations, processing the additional user input to determine additional data to use in providing a particular rationale for why the automated assistant executed a particular implementation of the assistant command may include obtaining metadata associated with the additional client device to use in executing the particular implementation, and determining the additional data to use in providing a particular rationale for why the automated assistant executed the particular implementation of the assistant command based on the metadata associated with the additional client device.

[0096] In some implementations, the request to the automated assistant to provide a particular rationale for why the automated assistant executed a particular implementation of the assistant command may include a general request to the automated assistant to provide a particular rationale for why the automated assistant executed a particular implementation. In some variations of these implementations, processing the additional user input to determine additional data to use in providing a particular rationale for why the automated assistant executed a particular implementation of the assistant command may include obtaining corresponding metadata associated with one or more of: (i) a particular software application among a plurality of heterogeneous software applications used to execute a particular implementation; (ii) a particular interpretation of a user input among a plurality of heterogeneous interpretations of a user input used to execute a particular implementation; or (iii) an additional client device of the user in place of the user's client device used to execute a particular implementation; and determining, based on the corresponding metadata, the additional data to use in providing a particular rationale for why the automated assistant executed a particular implementation of the assistant command.

[0097] In some implementations, having the automated assistant utilize the additional data to provide output for presentation to a user of the client device that includes a particular rationale for why the automated assistant performed a particular implementation of the assistant command may include processing the additional data using a text-to-speech (TTS) model to generate synthesized speech audio data that includes synthesized speech corresponding to the particular rationale for why the automated assistant performed a particular implementation of the assistant command.

[0098] In some implementations, having the automated assistant utilize the additional data to provide output for presentation to a user of the client device that includes a particular rationale for why the automated assistant performed a particular implementation of the assistant command may include causing the output that includes the particular rationale for why the automated assistant performed a particular implementation of the assistant command to be visually rendered on a display of the client device.

[0099] In some implementations, processing the additional user input to determine additional data to use in providing a particular rationale for why the automated assistant performed a particular implementation of the assistant command may include selecting the additional data from multiple disparate instances of pre-generated data based on a request included in the additional user input.

[0100] In some implementations, processing the additional user input to determine additional data to use in providing a particular rationale for why the automated assistant performed a particular implementation of an assistant command may include generating the additional data based on a request included in the additional user input.

[0101] In some implementations, a method is provided that is implemented by one or more processors, the method including receiving user input from a user of a client device, the user input including an assistant command, directed to an automated assistant that is at least partially executing on the client device; determining whether data utilized to execute a particular realization of the assistant command is determinable; in response to determining that the data utilized to execute a particular realization of the assistant command is not determinable, processing the user input to determine different data utilized to execute a different realization of the assistant command; causing the automated assistant to execute the different realization of the assistant command using the different data; and informing the automated assistant why the automated assistant did not execute the assistant command. The method includes receiving additional user input from a user of the client device, the additional user input including a request to provide a particular rationale for why the automated assistant executed an alternative implementation of the assistant command instead of a particular implementation of the assistant command; processing the additional user input to determine additional data to use in providing the particular rationale for why the automated assistant executed the alternative implementation of the assistant command instead of the particular implementation of the assistant command; and having the automated assistant use the additional data to provide output for presentation to the user of the client device, the output including the particular rationale for why the automated assistant executed the alternative implementation of the assistant command instead of the particular implementation of the assistant command.

[0102] These and other implementations of the disclosed technology may optionally include one or more of the following features.

[0103] In some implementations, processing the additional user input to determine additional data used to provide a particular rationale for why the automated assistant executed another implementation of the assistant command instead of a particular implementation of the assistant command may further include processing the additional user input to generate recommended data used to generate a recommended action on how the automated assistant can execute the particular implementation of the assistant command. In some variations of these implementations, the output may further include a recommended action on how the automated assistant can enable the particular execution of the assistant command. In another version of these implementations, the recommended action may include a prompt that, when selected, causes the automated assistant to execute the recommended action.

[0104] In some implementations, a method implemented by one or more processors is provided, the method including: receiving user input from a user of a client device, the user input including an assistant command and directed to an automated assistant executing at least partially on the client device; determining whether data to utilize in executing a particular realization of the assistant command is determinable; in response to determining that the data to utilize in executing a particular realization of the assistant command is not determinable, processing the user input to determine recommended data to use in generating recommended actions on how the automated assistant can execute the particular realization of the assistant command; having the automated assistant utilize the recommended data to provide, for presentation to the user of the client device, an output including recommended actions on how the automated assistant can enable the particular execution of the assistant command, and including a prompt that, when selected, causes the automated assistant to execute the recommended action; and in response to receiving additional user input from the user of the client device, including user selection of the prompt, causing the automated assistant to execute the recommended action to enable execution of the particular realization of the assistant command.

[0105] These and other implementations of the disclosed technology may optionally include one or more of the following features.

[0106] In some implementations, processing user input to determine recommendation data to use in generating recommended actions for how the automated assistant can perform a particular realization of the assistant command may be responsive to determining that there is no alternative realization of the assistant command.

[0107] In some implementations, a method implemented by one or more processors is provided, the method including the steps of receiving user input from a user of a client device, the user input including an assistant command and directed to an automated assistant executing at least partially on the client device; processing the user input to determine data to utilize in executing a particular implementation of the assistant command; causing the automated assistant to utilize the data to execute the particular implementation of the assistant command; receiving additional user input from the user of the client device including a request to the automated assistant to provide a particular rationale as to why the automated assistant did not execute another implementation of the assistant command instead of the particular implementation of the assistant command; processing the additional user input to determine additional data to utilize in providing the particular rationale as to why the automated assistant did not execute the other implementation of the assistant command instead of the particular implementation of the assistant command; and causing the automated assistant to utilize the additional data to provide output for presentation to the user of the client device, the output including the particular rationale as to why the automated assistant did not execute the other implementation of the assistant command instead of the particular implementation of the assistant command. In some implementations, a method implemented by one or more processors is provided, the method including: receiving user input from a user of a client device, the user input including an assistant command, directed to an automated assistant executing at least partially on the client device; processing the user input to determine data to utilize in executing a particular implementation of the assistant command; causing the automated assistant to execute the particular implementation of the assistant command using the data; and, as the automated assistant executes the particular implementation of the assistant command, causing the automated assistant to visually render for presentation to the user of the client device a selectable element that, when selected, causes the automated assistant to provide a particular rationale for why the automated assistant executed the particular implementation of the assistant command; and, in response to receiving additional user input from the user of the client device including a user selection of the selectable element, processing the additional user input to determine additional data to utilize in providing the particular rationale for why the automated assistant did not execute another implementation of the assistant command instead of the particular implementation of the assistant command; and causing the automated assistant to provide an output for presentation to the user of the client device that includes the particular rationale for why the automated assistant executed the particular implementation of the assistant command.

[0108] These and other implementations of the disclosed technology may optionally include one or more of the following features.

[0109] In some implementations, user input including an assistant command and directed to the automated assistant may be captured in audio data generated by one or more microphones on the client device. In some variations of these implementations, processing the user input to determine data to utilize in executing a particular implementation of the assistant command may include processing the audio data capturing the user input including the assistant command using an automatic speech recognition (ASR) model to generate an ASR output, processing the ASR output using a natural language understanding (NLU) model to generate an NLU output, and determining the data to utilize in executing a particular implementation of the assistant command based on the NLU output. In another version of these implementations, causing the automated assistant to visually render for presentation to a user of the client device a selectable element that, when selected, causes the automated assistant to provide a particular rationale for why the automated assistant executed a particular implementation of the assistant command may be in response to determining that an NLU metric associated with the NLU output does not satisfy an NLU metric threshold.

[0110] Further, some implementations include one or more processors (e.g., central processing units (CPUs), graphics processing units (GPUs), and / or tensor processing units (TPUs) of one or more computing devices), operable to execute instructions stored in associated memory, the instructions configured to perform any of the methods described above. Some implementations also include one or more non-transitory computer-readable storage media having computer instructions stored thereon that are executable by the one or more processors to perform any of the methods described above. Some implementations also include computer program products that include instructions that are executable by the one or more processors to perform any of the methods described above. [Explanation of symbols]

[0111] 110 Client device 111 Presence Sensor 112 User Interface Components 113 Automated Assistant Client 114 Voice Capture / ASR / NLU / TTS / Realization 115 Cloud-based Automated Assistant Components 119 Realization 120 Automated Assistants 130 Request Engine 140 Inference Engine 191 First Party Server 192 Third-party servers 199 Network 401 Users 410 Client Device 510 Client Device 580 Display 584 Microphone Interface Elements 614 processor 616 Network User Interface 620 User Interface Output Device 622 User Interface Input Device 624 Memory Subsystem 625 Memory Subsystem 626 File Storage Subsystem 110A User Profile 120A ML model 140A Metadata 452A Speech 452B Speech 454A1 Synthetic Speech 454B1 Synthetic speech 454C Synthetic Speech 456A Speech 458A1 Synthetic Speech 458A2 Prompt 458B1 Synthetic Speech 458B2 Prompt 458C1 Synthetic Speech 458C2 Prompt 460A Speech 552A Speech 552B Speech 554A Synthetic Speech 554B Synthetic Speech 556B1 First selectable element 556B2 Second selectable element 558A1 Synthetic Speech 558A2 Prompt 558B1 Synthetic Voice 55B2 Prompt

Claims

1. 1. A method implemented by one or more processors, the method comprising: receiving user input from a user of a client device, the user input including an assistant command, directed to an automated assistant executing at least partially on the client device; processing the user input to determine data to utilize in executing a particular implementation of the assistant command; causing the automated assistant to utilize the data to perform the particular realization of the assistant command; receiving additional user input from the user of the client device, the additional user input including a request to the automated assistant to provide a specific rationale as to why the automated assistant executed the specific implementation of the assistant command, the request to the automated assistant to provide a specific rationale as to why the automated assistant executed the specific implementation of the assistant command including a specific request to the automated assistant to provide the specific rationale as to why the automated assistant selected an additional client device of the user instead of the client device of the user to use to execute the specific implementation; processing the additional user input to determine additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command; and causing the automated assistant to utilize the additional data to provide output for presentation to the user of the client device that includes the particular rationale for why the automated assistant performed the particular implementation of the assistant command.

2. 10. The method of claim 1, wherein the user input, including the assistant command, directed to the automated assistant is captured in audio data generated by one or more microphones on the client device.

3. The step of processing the user input to determine data to utilize in performing a particular implementation of the assistant command comprises: processing the audio data incorporating the user input, including the assistant command, with an automatic speech recognition (ASR) model to generate an ASR output; processing the ASR output with a natural language understanding (NLU) model to generate an NLU output; determining the data to use in performing a particular realization of the assistant command based on the NLU output; 3. The method of claim 2, comprising:

4. 4. The method of claim 1, wherein the user input including the assistant command directed to the automated assistant corresponds to a touch input detected via a display of the client device or a key input received through another input device of the client device.

5. The step of processing the user input to determine data to utilize in performing a particular implementation of the assistant command comprises: processing the keystrokes using a natural language understanding (NLU) model to generate an NLU output; generating the data to be used to perform a particular realization of the assistant command based on the NLU output; 5. The method of claim 4, comprising:

6. 6. The method of claim 1, wherein the request to the automated assistant to provide a particular rationale as to why the automated assistant performed the particular implementation of the assistant command comprises a specific request to the automated assistant to provide the particular rationale as to why the automated assistant selected a particular software application from multiple disparate software applications used to perform the particular implementation.

7. processing the additional user input to determine the additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command, obtaining metadata associated with the particular software application utilized in executing the particular implementation; and determining, based on the metadata associated with the particular software application, additional data to be used to provide the particular rationale for why the automated assistant executed the particular implementation of the assistant command. The method of claim 6, comprising:

8. 8. The method of claim 1, wherein the request to the automated assistant to provide a particular rationale for why the automated assistant performed the particular implementation of the assistant command comprises a specific request to the automated assistant to provide the particular rationale for why the automated assistant selected a particular interpretation of the user input from multiple disparate interpretations of the user input to use in performing the particular implementation.

9. processing the additional user input to determine the additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command, obtaining metadata associated with the particular interpretation of the user input utilized in executing the particular implementation; determining additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command based on the metadata associated with the particular interpretation of the user input; 9. The method of claim 8, comprising:

10. processing the additional user input to determine the additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command, obtaining metadata associated with the additional client devices utilized in executing the particular implementation; determining additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command based on the metadata associated with the additional client device; 2. The method of claim 1, comprising:

11. 11. The method of claim 1, wherein the request to the automated assistant to provide a particular rationale as to why the automated assistant performed the particular implementation of the assistant command comprises a general request to the automated assistant to provide the particular rationale as to why the automated assistant performed the particular implementation.

12. processing the additional user input to determine the additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command, obtaining corresponding metadata associated with one or more of: (i) a particular software application among a plurality of disparate software applications utilized in execution of the particular implementation; (ii) a particular interpretation of the user input among a plurality of disparate interpretations of the user input utilized in execution of the particular implementation; or (iii) an additional client device of the user in place of the user's client device utilized in execution of the particular implementation; determining additional data to be used to provide the particular rationale for why the automated assistant executed the particular implementation of the assistant command based on the corresponding metadata; 12. The method of claim 11, comprising:

13. causing the automated assistant to utilize the additional data to provide output for presentation to the user of the client device including the particular rationale for why the automated assistant performed the particular implementation of the assistant command, processing the additional data using a text-to-speech (TTS) model to generate synthetic speech audio data, the synthetic speech corresponding to the particular rationale for why the automated assistant executed the particular implementation of the assistant command; 13. The method according to any one of claims 1 to 12.

14. causing the automated assistant to utilize the additional data to provide output for presentation to the user of the client device including the particular rationale for why the automated assistant performed the particular implementation of the assistant command, causing the output, including the particular rationale for why the automated assistant executed the particular implementation of the assistant command, to be visually rendered on a display of the client device.

14. The method according to any one of claims 1 to 13.

15. processing the additional user input to determine the additional data to use in providing the particular rationale for why the automated assistant executed the particular implementation of the assistant command, generating the additional data based on the request contained in the additional user input; 15. The method according to any one of claims 1 to 14.

16. 1. A method implemented by one or more processors, the method comprising: receiving user input from a user of a client device, the user input including an assistant command, directed to an automated assistant executing at least partially on the client device; determining whether data to be utilized in executing a particular realization of the assistant command can be determined; In response to determining that the data to utilize in performing a particular realization of the assistant command is undeterminable, processing the user input to determine additional data to utilize in executing additional realizations of the assistant command; causing the automated assistant to utilize the different data to perform a different realization of the assistant command; receiving additional user input from the user of the client device, the additional user input including a request to the automated assistant to provide a particular rationale for why the automated assistant executed the alternative implementation of the assistant command instead of the particular implementation of the assistant command; processing the additional user input to determine additional data used to provide the particular rationale for why the automated assistant executed the alternative implementation of the assistant command instead of the particular implementation of the assistant command, wherein processing the additional user input to determine the additional data used to provide the particular rationale for why the automated assistant executed the alternative implementation of the assistant command instead of the particular implementation of the assistant command includes processing the additional user input to generate recommendation data used to generate a recommended action for how the automated assistant can execute the particular implementation of the assistant command; causing the automated assistant to utilize the additional data to provide output for presentation to the user of the client device, the output including the particular rationale for why the automated assistant executed the alternative implementation of the assistant command instead of the particular implementation of the assistant command; A method comprising:

17. The method of claim 16 , wherein the output further includes the recommended action on how the automated assistant can enable a particular execution of the assistant command.

18. The method of claim 17 , wherein the recommended action includes a prompt that, when selected, causes the automated assistant to perform the recommended action.

19. 1. A method implemented by one or more processors, the method comprising: receiving user input from a user of a client device, the user input including an assistant command, directed to an automated assistant executing at least partially on the client device; processing the user input to determine data to utilize in executing a particular implementation of the assistant command; causing the automated assistant to utilize the data to perform a particular realization of the assistant command; receiving additional user input from the user of the client device, the additional user input including a request to the automated assistant to provide a particular rationale for why the automated assistant did not execute another implementation of the assistant command instead of the particular implementation of the assistant command; processing the additional user input to determine additional data to use in providing the particular rationale for why the automated assistant did not execute the other implementation of the assistant command instead of the particular implementation of the assistant command; causing the automated assistant to utilize the additional data to provide output for presentation to the user of the client device, the output including the particular rationale for why the automated assistant did not execute the other implementation of the assistant command instead of the particular implementation of the assistant command; A method comprising:

20. 1. A method implemented by one or more processors, the method comprising: receiving user input from a user of a client device, the user input including an assistant command, directed to an automated assistant executing at least partially on the client device; processing the user input to determine data to utilize in executing a particular implementation of the assistant command; causing the automated assistant to utilize the data to perform a particular realization of the assistant command; When the automated assistant executes a particular implementation of the assistant command, causing the automated assistant to visually render for presentation to the user of the client device a selectable element that, when selected, causes the automated assistant to provide a particular rationale as to why the automated assistant performed the particular implementation of the assistant command; in response to receiving additional user input from the user of the client device, the user input including a user selection of the selectable element; processing the additional user input to determine additional data to use in providing the particular rationale for why the automated assistant did not execute another implementation of the assistant command instead of the particular implementation of the assistant command; causing the automated assistant to provide, for presentation to the user of the client device, an output including the particular rationale for why the automated assistant executed the particular implementation of the assistant command; A method comprising:

21. 21. The method of claim 20, wherein the user input, including the assistant command, directed to the automated assistant is captured in audio data generated by one or more microphones on the client device.

22. The step of processing the user input to determine data to be utilized in executing a particular implementation of the assistant command comprises: processing the audio data incorporating the user input, including the assistant command, with an automatic speech recognition (ASR) model to generate an ASR output; processing the ASR output with a natural language understanding (NLU) model to generate an NLU output; determining the data to use in performing a particular realization of the assistant command based on the NLU output; 22. The method of claim 21, comprising:

23. 23. The method of claim 22, wherein causing the automated assistant to visually render for presentation to the user of the client device the selectable element that, when selected, causes the automated assistant to provide the particular rationale for why the automated assistant executed the particular implementation of the assistant command is in response to determining that an NLU metric associated with the NLU output does not satisfy an NLU metric threshold.

24. at least one processor; a memory storing instructions that, when executed, cause said at least one processor to perform the operations of any one of claims 1 to 23; A system equipped with

25. 24. A non-transitory computer-readable storage medium having stored thereon instructions that, when executed, cause at least one processor to perform the operations of any one of claims 1 to 23.

Citation Information

Patent Citations

  • Vehicle-mounted dialogue system and method for information service

    JP2005121526A

  • Method for executing service for explaining reason by using interactive robot, device and program thereof

    JP2007011674A

  • Contextual voice user interface

    US20200118564A1

  • Method and apparatus for capability-based processing of voice queries in a multi-assistant environment

    US20200135200A1