Passive ambiguation resolution of assistant commands

The automated assistant addresses ambiguous user inputs by initiating a primary action and offering selectable alternatives, reducing latency and resource waste through efficient disambiguation and direct user selection.

JP7833021B2Active Publication Date: 2026-03-18GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2026-03-18

AI Technical Summary

Technical Problem

Existing automated assistants often misinterpret ambiguous user inputs, leading to prolonged interactions and wastage of computing resources as users are prompted to clarify their intentions, especially when multiple interpretations of a command are possible.

Method used

An automated assistant that automatically initializes a primary action based on user input while simultaneously providing selectable elements for alternative actions, allowing users to pivot quickly to the intended action without additional clarification.

Benefits of technology

Reduces latency and conserves computing resources by enabling quick and efficient disambiguation of user commands, allowing users to select alternative actions directly, thus shortening human-computer interaction duration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007833021000001
    Figure 0007833021000001
  • Figure 0007833021000002
    Figure 0007833021000002
  • Figure 0007833021000003
    Figure 0007833021000003
Patent Text Reader

Abstract

To passively disambiguate assistant commands.SOLUTION: An automated assistant can initialize execution of an assistant command associated with an interpretation that is predicted to be responsive to user input, while simultaneously providing suggestions for alternative assistant command(s) associated with alternative interpretation(s) that is / are also predicted to be responsive to the user input. The alternative assistant command(s) can be selectable such that, when selected, the automated assistant can pivot from executing the assistant command to initializing execution of the selected alternative assistant command(s). Further, the alternative assistant command(s) can be partially fulfilled prior to any user selection thereof. Accordingly, implementations set forth herein can enable the automated assistant to quickly and efficiently pivot between assistant commands that are predicted to be responsive to the user input.SELECTED DRAWING: Figure 1B
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans may engage in human-computer interaction using interactive software applications (also known as “digital agents,” “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “assistant applications,” “conversational agents,” etc.) referred to herein as “automated assistants.” For example, a person (who may be referred to as a “user” when interacting with an automated assistant) may give commands and / or requests to an automated assistant by using oral natural language input (i.e., utterances) which may, in some cases, be converted to text and then processed, and / or by providing textual (e.g., typed) natural language input.

[0002] Often, interacting with an automated assistant can present many opportunities to misinterpret ambiguous user input, including the assistant's requests and / or commands. For example, assume a user gives a request to play media content (e.g., a song) that is available through multiple different media applications. In this example, rather than the automated assistant automatically selecting a particular application in response to the request and causing immediate playback of the media content, the automated assistant may provide an output (e.g., an audible and / or visual output) that requests the user to select the particular application to be used for playback of the media content. Further, assume that there are multiple media content items with the same name. The automated assistant may provide an additional or alternative output (e.g., an audible and / or visual output) that requests the user to select the particular media content item with the same name, rather than the automated assistant selecting a particular media content in response to the request and causing playback of the particular media content on a particular application. As a result, the interaction between the automated assistant and the user is prolonged, thereby wasting the computing resources of the client device utilized to conduct the interaction between the automated assistant and the user and wasting the user's time.

[0003] [[ID=ģ]] In some cases, even assuming that an automated assistant automatically selects a specific application to play media content and / or specific media content in response to a request, there is often no efficient mechanism for the user to pivot to an alternative application or alternative media content that the request might otherwise fulfill. For example, suppose a request to play media content is a request to play a song titled "Crazy" that is available for playback by several different media applications. Let's further assume that there are multiple songs titled "Crazy" by different artists, including at least one by a first artist and one by a second artist. In this example, if the user intended to trigger playback of "Crazy" by the second artist, but the automated assistant automatically selects and triggers playback of "Crazy" by the first artist, the user may be asked to provide further user input to cancel playback of "Crazy" by the first artist and to provide further additional user input (often improved to include the name of the second artist) to provide playback of "Crazy" by the second artist. As a result, the interaction between the automated assistant and the user becomes prolonged, thereby wasting the computing resources of the client device used to conduct the interaction between the automated assistant and the user. [Overview of the project] [Means for solving the problem]

[0004] Some implementations described herein relate to an automated assistant that automatically initializes the execution of at least a first action to fulfill an assistant command contained in an oral utterance given by a user. Furthermore, while the first action is being performed by the automated assistant, the automated assistant simultaneously provides the user with selectable elements relating to corresponding alternative actions related to the alternative performance of the assistant command contained in the oral utterance. Thus, if the user selects an selectable element, the automated assistant terminates the execution of the first action and automatically initializes the execution of the corresponding alternative action related to the alternative performance of the assistant command contained in the oral utterance. In these and other ways described herein, the automated assistant can passively remove ambiguity in an oral utterance so that the user does not subsequently have to send another oral utterance or any clarification oral utterance when the first action does not correspond to a particular action intended by the user. In other words, the automated assistant can initialize the execution of a given action that is expected to respond to an oral utterance and can quickly and efficiently pivot to an alternative action that is also considered to respond to the oral utterance. As used herein, having an automated assistant simultaneously provide selectable elements to the user while the automated assistant is performing a first action may include having the automated assistant provide selectable elements to the user at the same time the first action is automatically initialized by the automated assistant, and / or within a threshold duration before and / or after the time the first action is automatically initialized by the automated assistant.

[0005] In some versions of these implementations, assistant input data characterizing multiple interpretations predicted to respond to oral utterances may be generated based on the oral utterances. In these implementations, each of the multiple interpretations may include a corresponding intent, one or more corresponding parameters associated with the corresponding intent, and one or more corresponding slot values ​​for one or more of the corresponding parameters associated with the corresponding intent. Furthermore, metric data characterizing the predicted degree of correspondence between each of the multiple interpretations and the assistant commands contained in the oral utterances may be generated based on the assistant input data. In some implementations, the metric data may include confidence levels related to processing oral utterances using various components, such as ASR metrics related to ASR outputs generated based on processing oral utterances using an automatic speech recognition (ASR) model, NLU metrics related to NLU outputs generated based on processing ASR outputs using a natural language understanding (NLU) model, performance metrics related to performance outputs generated based on processing NLU outputs using performance models and / or rules, and / or other metrics related to processing oral utterances. In additional or alternative implementations, metric data may be based on user profile data of the user who provided the oral utterance, user profile data of other users similar to the user who provided the oral utterance, aggregate increase of oral utterances including requests across specific geographical regions, and / or other data. In some versions of those implementations, metric data may be generated using one or more machine learning models and / or heuristic processes that process these various signals to generate metric data. Based on the assistant input data and metric data, the automated assistant can be automatically initialized to perform an assistant command contained in the oral utterance, relating to a first interpretation of multiple interpretations, and can be provided to present selectable elements relating to other interpretations of multiple interpretations to the user.

[0006] For example, suppose a user gives the verbal utterance "Play the song 'Crazy'". An automated assistant can process the verbal utterance using various components (e.g., ASR components, NLU components, performance components, and / or other components) to generate assistant input data that characterizes multiple interpretations. In this example, each of the multiple interpretations may be associated with a musical intent, based on the determination that the user intends for a song titled "Crazy" to be played in response to the verbal utterance. However, when giving the verbal utterance, the user specified a slot value for a song parameter associated with the musical intent (e.g., "Crazy"), but did not specify a slot value associated with an artist parameter associated with the musical intent, a slot value for an application parameter associated with the musical intent, or any other parameter that may be associated with the musical intent. Nevertheless, when generating assistant input data that characterizes multiple interpretations that are expected to respond to the verbal utterance, the automated assistant can infer various slot values ​​to generate multiple interpretations. For example, the first interpretation may include the musical intent, the slot value "Crazy" for the song parameter, the slot value "Artist 1" for the artist parameter, and the slot value "Application 1" for the application parameter; the second interpretation may include the musical intent, the slot value "Crazy" for the song parameter, the slot value "Artist 2" for the artist parameter, and the slot value "Application 1" for the application parameter; the third interpretation may include the musical intent, the slot value "Crazy" for the song parameter, the slot value "Artist 1" for the artist parameter, and the slot value "Application 2" for the application parameter, and so on.

[0007] The above example illustrates multiple interpretations that share the same intent (e.g., a musical intent), but it should be understood that this is merely an example and not intended to be limiting. In contrast to the above example, suppose the user instead gives the verbal utterance "Play 'The Floor is Lava'." Similarly, an automated assistant can process a verbal utterance using various components (e.g., ASR components, NLU components, performance components, and / or other components) to generate assistant input data that characterizes multiple interpretations. However, in this example, the multiple interpretations may be associated with different intents. For example, the first interpretation may include the musical intent, the slot value of the song parameter "The Floor is Lava", the slot value of the artist parameter "Artist 1", and the slot value of the application parameter "Application 1", the second interpretation may include the musical intent, the slot value of the song parameter "The Floor is Lava", the slot value of the artist parameter "Artist 2", and the slot value of the application parameter "Application 1", the third interpretation may include the television program intent, the slot value of the video parameter "The Floor is Lava", and the slot value of the application parameter "Application 2", the fourth interpretation may include the game intent, the slot value of the game parameter "The Floor is Lava", and the slot value of the application parameter "Application 3", and so on.

[0008] Furthermore, in these examples, the automated assistant can automatically initialize a given action associated with one of several interpretations to perform an assistant command contained in a spoken utterance, based on the assistant input data and on metric data generated for each of several interpretations when processing the spoken utterance (e.g., ASR metric data, NLU metric data, and / or performance metric data). For example, in the first example, the automated assistant can automatically initialize the execution of a first action associated with the first interpretation, including musical intent, the slot value "Crazy" for the song parameter, the slot value "Artist 1" for the artist parameter, and the slot value "Application 1" for the application parameter, assuming that the metric data indicates that the first interpretation is most likely to correspond to an assistant command in the spoken utterance. Nevertheless, in this example, the automated assistant can render to present the user with selectable elements associated with alternative interpretations while the first action is automatically initialized and executed (e.g., while "Crazy" by "Artist 1" is played by "Application 1"). For example, while "Crazy" by "Artist 1" is being played by "Application 1", corresponding selectable elements related to a second interpretation (e.g., musical intent, song parameter slot value "Crazy", artist parameter slot value "Artist 2", and application parameter slot value "Application 1") and corresponding selectable elements related to a third interpretation (e.g., musical intent, song parameter slot value "Crazy", artist parameter slot value "Artist 1", and application parameter slot value "Application 2") may be rendered for the user to see.

[0009] Therefore, in response to a user selection of one of the corresponding selectable elements (for example, by touch selection targeting the corresponding selectable element or voice selection of the corresponding selectable element), the automated assistant can pivot away from performing the first action to automatically initialize and perform an alternative action associated with an alternative action to carry out an assistant command contained in the spoken utterance. For example, in response to a user selection of the corresponding selectable element related to a second interpretation, the automated assistant can terminate the execution of the first action (for example, the playback of "Crazy" by "Artist 1" in "Application 1") and automatically initialize the execution of the second action related to the second interpretation (for example, initializing the playback of "Crazy" by "Artist 2" in "Application 1"). As a result, the automated assistant can pivot quickly and efficiently between different interpretations of the utterance without having to generate and render prompts for the user to actively remove the ambiguity of these different interpretations of the spoken utterance before initializing any action. More precisely, as described herein, by enabling passive disambiguation of assistant commands contained in oral utterances, the automated assistant can prioritize reducing latency when performing assistant commands contained in oral utterances, and can prioritize concluding human-computer interactions between the automated assistant and the user more quickly and efficiently.

[0010] Furthermore, while the above examples are illustrated in relation to assistant commands that control media content, please understand that they are examples only and not intended to be limiting. For example, suppose the user gives the oral utterance "Translate the light to 20 percent." Similarly, an automated assistant can process the oral utterance using various components (e.g., ASR component, NLU component, performance component, and / or other components) to generate assistant input data that characterizes multiple interpretations. In this example, suppose the oral utterance is intended by the user to reduce the brightness of one or more smart lights controllable via the automated assistant to 20%. However, since the user typically does not need to use this local language to adjust the brightness settings of their lights, the first interpretation of the oral utterance may include a translation intent with the slot value "the light to 20 percent" as a term or phrase parameter to be translated from the first language to one or more second languages. Still, the second interpretation of the oral utterance may include a translation intent with the slot value "20 The intention of the spoken utterance may include a lighting control intent with a slot value of "20 percent" and a lighting position parameter slot value of "kitchen," and the third interpretation of the spoken utterance may include a lighting control intent with a slot value of "20 percent" and a lighting position parameter slot value of "living room," and so on.As a result, the automated assistant can not only automatically initialize a given action related to one of several interpretations in order to perform an assistant command contained in a spoken utterance (for example, providing a translation of "the light to 20 percent" from a first language into one or more second languages), but can also provide selectable elements related to alternative interpretations of the spoken utterance, allowing the user to adjust one or more smart lights by touch selection targeting the selectable elements or voice selection of the selectable elements when selected.

[0011] In various implementations, when a user provides voice selections of selectable elements through additional verbal utterances, the processing of these additional verbal utterances may be biased towards content related to the selectable elements being rendered for presentation to the user. For example, suppose, as described above, the user gives the verbal utterance "Play the song 'Crazy'". Furthermore, the automated assistant automatically initializes the execution of a first action related to a first interpretation, which includes the musical intent, the slot value "Crazy" for the song parameter, the slot value "Artist 1" for the artist parameter, and the slot value "Application 1" for the application parameter. Furthermore, corresponding selectable elements related to a second interpretation, including the musical intent, the song parameter slot value "Crazy," the artist parameter slot value "Artist 2," and the application parameter slot value "Application 1," are rendered for presentation to the user. In this example, the ASR and / or NLU processing of additional oral utterances can be biased towards the slot values ​​related to these alternative interpretations, or the ASR and / or NLU output generated based on the processing of additional oral utterances can be biased towards the slot values ​​related to these alternative interpretations.

[0012] In various implementations, selectable elements related to alternative interpretations may be rendered to present to the user for a threshold duration after the automated assistant, which is responsible for executing assistant commands contained in verbal utterances, has automatically initialized itself. In some versions of these implementations, the threshold duration can be static. For example, selectable elements related to alternative interpretations may be rendered to present for 10 seconds, 15 seconds, or any other static threshold duration after the automated assistant, which is responsible for executing assistant commands contained in verbal utterances, has automatically initialized itself. In other versions of these implementations, the threshold duration can be dynamic, for example, based on one or more intentions out of multiple interpretations, or one or more types of actions performed based on one or more of multiple actions. For example, in an implementation where multiple actions are associated with media content, selectable elements related to alternative interpretations may be rendered to present to the user for only 10 seconds. However, in an implementation where multiple actions are associated with controlling the user's smart device, selectable elements related to alternative interpretations may be rendered to present to the user for only 30 seconds.

[0013] In various implementations, execution data for alternative actions related to selectable elements may be pre-cached on the client device used by the user to interact with the automated assistant in order to reduce latency when performing the alternative actions. Continuing with the above example, where the user gives the verbal utterance "Play the song 'Crazy'", and the automated assistant automatically initializes the execution of a first action related to a first interpretation, including the musical intent, the song parameter slot value "Crazy", the artist parameter slot value "Artist 1", and the application parameter slot value "Application 1", the automated assistant can then queue "Crazy" by "Artist 2" in "Application 1" for a second interpretation, and "Crazy" by "Artist 1" in "Application 2" for a third interpretation. In other words, the automated assistant can establish connections with various applications in the background of the client device, queue various content for playback on the client device, queue various content for display on the client device, and / or otherwise pre-cached the execution on the client device.

[0014] In various implementations, selectable elements rendered for presentation to the user may be included in GUI data that characterizes the verbal utterance-responsive assistant graphical user interface (GUI). GUI data may include, for example, one or more control elements related to controlling a first action automatically initialized by the automated assistant, selectable elements related to alternative actions, and / or any other data related to how and / or when content is rendered for visual presentation to the user. For example, GUI data may include prominence data indicating the corresponding area of ​​the client device's display used to display the selectable elements related to alternative actions, and the corresponding size of the corresponding area of ​​the client device's display used to display the selectable elements related to alternative actions. Continuing with the above example where the user gives the verbal utterance "Play the song 'Crazy'", and the automated assistant automatically initializes the execution of a first action related to a first interpretation, including the musical intent, the slot value of the song parameter "Crazy", the slot value of the artist parameter "Artist 1", and the slot value of the application parameter "Application 1", the GUI data in this example may indicate that the selectable elements related to the second interpretation should be displayed more prominently than the selectable elements related to the third interpretation, for example, by having the selectable elements related to the second interpretation occupy a larger area of ​​the client device's display, or by having the selectable elements related to the second interpretation appear above the selectable elements related to the third interpretation. In this example, the GUI data may indicate that the selectable elements related to the second interpretation should be displayed more prominently than the selectable elements related to the third interpretation based on metric data related to the second interpretation that indicates the second interpretation is more likely to respond to the verbal utterance compared to the third interpretation.Furthermore, in this example, the GUI data may indicate that only the selectable elements related to the second interpretation should be displayed, without displaying the selectable elements related to the third interpretation. The GUI data may indicate, for example, that the metric data related to the second interpretation meets the metric threshold, while the metric data related to the third interpretation does not meet the metric threshold, that the size of the client device's display can only display one selectable element, and / or other criteria should indicate that only the selectable elements related to the second interpretation should be displayed, without displaying the selectable elements related to the third interpretation.

[0015] In various implementations, the user may provide one or more inputs to cause additional selectable elements associated with additional interpretations and not initially provided for presentation to the user. For example, the user may provide a swipe in a specific direction on the client device's display (e.g., up swipe, down swipe, left swipe, and / or right swipe) or other touch input (e.g., long tap, hard tap) to provide additional selectable elements for presentation to the user. Another example is the user may provide an automated assistant with additional verbal utterances, including one or more specific words or phrases (e.g., "More," "Show me more"), to cause the automated assistant, when detected, to provide additional selectable elements for presentation to the user.

[0016] In various implementations, when an automated assistant is automatically initialized to perform a given action that is expected to respond to a spoken utterance, the automated assistant may consider one or more costs associated with the given action. These costs may include, for example, whether the assistant command associated with the given action utilizes the user's financial information, whether the assistant command causes electronic communication (e.g., phone calls, text messages, email messages, social media messages, etc.) to be initiated and / or sent from the user's client device to further client devices of other users, whether the assistant command consumes a threshold amount of computing resources, whether the assistant command delays other processes, whether an inaccurate action would affect the user or another person in any way, and / or other costs. For example, the cost associated with using financial information or initiating and / or sending electronic communication from a client device may be considered much greater than the cost associated with providing media content for playback to the user.

[0017] By using the techniques described herein, various technical advantages can be realized. As one non-limiting example, the techniques described herein enable an automated assistant to passively remove ambiguity in assistant commands contained in verbal utterances given by a user. For example, a user may give a verbal utterance containing a request, and the automated assistant can automatically initialize the execution of an assistant command that is expected to fulfill the request based on a first interpretation of the request, while simultaneously providing selectable elements that are also expected to fulfill the request and are associated with other interpretations of the request. Thus, if the user gives a user selection of one of the selectable elements, the automated assistant can quickly and efficiently pivot to automatically initializing the execution of an assistant command associated with the other interpretations. In particular, the automated assistant does not need to generate and render prompts before initializing the execution of an assistant command. As a result, latency in fulfilling requests can be reduced, computing resources on the client device can be saved, and the duration of human-computer interaction can be shortened. Furthermore, when the user provides a choice among the selectable elements, the execution data related to these other interpretations can be pre-cached on the client device to reduce latency when switching execution to one of these alternative assistant commands.

[0018] The above description is given as an overview of some implementations of this disclosure. Further descriptions of those and other implementations are provided below in more detail. [Brief explanation of the drawing]

[0019] [Figure 1A] This diagram illustrates a scenario where a user invokes an automated assistant that can simultaneously suggest alternative interpretations in response to user input, while also executing a specific interpretation predicted to be most relevant to the user input. [Figure 1B] This diagram illustrates a scenario where a user invokes an automated assistant that can simultaneously suggest alternative interpretations in response to user input, while also executing a specific interpretation predicted to be most relevant to the user input. [Figure 1C] This diagram illustrates a scenario where a user invokes an automated assistant that can simultaneously suggest alternative interpretations in response to user input, while also executing a specific interpretation predicted to be most relevant to the user input. [Figure 1D] This diagram illustrates a scenario where a user invokes an automated assistant that can simultaneously suggest alternative interpretations in response to user input, while also executing a specific interpretation predicted to be most relevant to the user input. [Figure 2A] This figure shows a scenario in which a user invokes an automated assistant that, in response to user input, can initialize the execution of a specific interpretation while simultaneously providing selectable suggestions on the display interface of a computing device. [Figure 2B] This figure shows a scenario in which a user invokes an automated assistant that, in response to user input, can initialize the execution of a specific interpretation while simultaneously providing selectable suggestions on the display interface of a computing device. [Figure 2C] This figure shows a scenario in which a user invokes an automated assistant that, in response to user input, can initialize the execution of a specific interpretation while simultaneously providing selectable suggestions on the display interface of a computing device. [Figure 3] This figure shows a system that can simultaneously perform a specific interpretation predicted to be most relevant to the user input, and invoke an automated assistant to suggest alternative user interpretations in response to the user input. [Figure 4A] This figure illustrates how to operate an automated assistant to provide alternative interpretations in response to oral utterances, which may be interpreted in various ways and / or may be missing certain parameters. [Figure 4B] A diagram showing a method for operating an automated assistant to provide alternative interpretations in response to spoken utterances that may be interpreted in various different ways and / or may be missing certain parameters. [Figure 5] A block diagram of an exemplary computer system.

Best Mode for Carrying Out the Invention

[0020] FIGS. 1A, 1B, 1C, and 1D show scenarios 100, 120, 140, and 160 in which a user 102 invokes an automated assistant that can propose alternative interpretations in response to a user input, simultaneously with the execution of a particular interpretation predicted to be most relevant to the user input. The proposed alternative interpretations can be presented to the user 102 for selection via the display interface 106 of the computing device 104 and / or other interfaces if the user 102 determines that the interpretation being executed is not what the user intended. For example, as shown in FIG. 1A, the user 102 can give a spoken utterance 108 such as "Assistant, 'Play Science Rocks'" to the automated assistant via the audio interface of the computing device 104, referring to the song the user 102 wants to listen to. In response to receiving the spoken utterance 108, the automated assistant can initiate the execution of an assistant command associated with a particular interpretation of the spoken utterance predicted to have the greatest degree of correspondence with the spoken utterance 108 given by the user 102.

[0021] As shown in FIG. 1B, in response to the spoken utterance 108, the automated assistant can cause the computing device 104 to render a graphical element 122 on the display interface 106 of the computing device 104. The graphical element 122 can act as an interface for controlling one or more actions corresponding to the intent in execution. For example, in response to receiving the spoken utterance 108, the graphical element 122 can be rendered on the display interface 106, and can indicate that a music application accessible to the computing device 104 is playing the song "Science Rocks". In some implementations, since the user 102 did not specify a particular music application for the song to be rendered or a particular artist of the song "Science Rocks", the automated assistant can infer one or more different slot values of corresponding parameters related to the music intent, such as one or more different music applications where the user might have intended to interact with the automated assistant, and can infer one or more different artists who the user might have intended to be used by the automated assistant for playing the song "Science Rocks". For example, the automated assistant may select the slot value of the corresponding parameter predicted to have the greatest degree of correspondence with the spoken utterance 108, and the automated assistant can use those slot values, such as the slot value of the first music application for the application parameter and the slot value of Pop Class for the artist parameter as shown by the graphical element 122 in FIG. 1B, to automatically initialize the playback of the song "Science Rocks".

[0022] In some implementations, in addition to performing this initial interpretation, the automated assistant may generate one or more other alternative interpretations that may be proposed to the user 102 via one or more selectable elements, such as a first suggestion element 124 and a second suggestion element 126. In some implementations, additional suggested interpretations may be provided to the user 102 when the initial interpretation does not have a predicted degree of correspondence (i.e., correspondence metric) with the user request and / or oral utterance 108 that satisfies the degree of correspondence threshold (i.e., metric threshold). For example, the first suggestion element 124 may suggest that the requested song be played using a different music application (e.g., a second music application in the application parameter), and the second suggestion element 126 may suggest that the initially predicted music application plays a song by a different artist (e.g., "Music Group," as shown in Figure 1B) that has the same name (e.g., "Science Rocks"). Furthermore, the first suggestion element 124 and the second suggestion element 126 may be rendered simultaneously with the automated assistant executing the assistant command related to the initial interpretation. This may allow the automated assistant to perform an initial interpretation that has the greatest degree of correspondence with the requested command, while also providing suggestions in case the automated assistant is wrong about the initial interpretation. In particular, the first suggestion element 124 and the second suggestion element 126 may be provided to be presented to the user only for the duration of a threshold.

[0023] In various implementations, as will be explained in more detail with respect to Figure 3, the interpretation may be generated based on an automated assistant in the computing device 104 processing the oral utterance 108 using various available components. For example, an automatic speech recognition (ASR) component may be used to process audio data capturing the oral utterance 108 to generate an ASR output. The ASR output may include, for example, a speech hypothesis predicted to correspond to the oral utterance 108, ASR metrics associated with the speech hypothesis, phonemes predicted to correspond to the oral utterance 108, and / or other ASR outputs. Furthermore, a natural language understanding (NLU) component may be used to process the ASR output to generate an NLU output. The NLU output may include, for example, one or more intentions predicted to satisfy the oral utterance 108 (e.g., the musical intention in the examples in Figures 1A to 1D), corresponding parameters associated with each of the one or more intentions, corresponding slot values ​​for the corresponding parameters, NLU metrics associated with one or more intentions, corresponding parameters, and / or corresponding slot values, and / or other NLU outputs. Furthermore, a performance component may be used to process the NLU output and generate performance data. The performance data, when executed, can correspond to a variety of actions that cause the automated assistant to execute corresponding assistant commands in an attempt to perform the oral utterance, and may optionally be associated with performance metrics indicating how likely it is that the execution of a given assistant command is predicted to satisfy the oral utterance. In particular, each of the interpretations described herein may include various combinations of intentions, corresponding parameters, and corresponding slot values, and the automated assistant can generate assistant input data characterizing these various combinations of intentions, corresponding parameters, and corresponding slot values.

[0024] In some implementations, the automated assistant can automatically initialize the execution of actions related to an initial interpretation based on corresponding metrics for each of multiple interpretations. In some versions of these implementations, corresponding metrics may be generated based on ASR metrics, NLU metrics, and performance metrics related to each of multiple interpretations. In additional or alternative implementations, corresponding metrics may be generated based on user profile data of the user who gave the oral utterance 108 (e.g., user preferences, history of user interactions with various applications accessible on computing device 104, user search history, user purchase history, user calendar information, and / or any other information about the user on computing device 104), user profile data of other users similar to the user who gave the oral utterance 108, an increase in the total number of oral utterances including requests across a specific geographical area, and / or other data.

[0025] In some versions of these implementations, metric data may be generated using one or more machine learning models and / or heuristic processes that process these various signals to generate metric data. For example, one or more machine learning models may be trained to process these signals to determine a correspondence metric that indicates the predicted degree of correspondence between each of the multiple interpretations and the oral utterance 108, for each of the multiple interpretations. In some of these implementations, one or more machine learning models may be trained on-device based on locally generated data on the computing device 104 so that one or more machine learning models are personalized for the user 102 of the computing device 104. One or more machine learning models can be trained based on multiple training instances, each of which may include a training instance input and a training instance output. The training instance input may include any combination of these signals and / or examples of interpretations of the oral utterance, and the training instance output may include a ground truth output that shows the ground truth interpretation of the training instance input. A given training case input can be applied as input across one or more machine learning models to generate predicted outputs containing corresponding metrics for each interpretation example, and these predicted outputs can be compared to ground truth outputs to generate one or more losses. Furthermore, one or more losses can be used to update one or more machine learning models (e.g., by backpropagation). One or more machine learning models can be deployed after sufficient training (e.g., based on processing a threshold amount of training cases, based on training over a threshold duration, based on the performance of one or more machine learning models during training). As another example, one or more heuristic-based processes or rules can be used to process these signals and / or the multiple interpretations to determine, for each of the multiple interpretations, a corresponding metric indicating the predicted degree of correspondence between each of the multiple interpretations and the oral utterance 108.Based on assistant input data and metric data, the automated assistant can automatically initialize a first action related to a first interpretation among multiple interpretations to perform an assistant command contained in the spoken utterance, and can provide selectable elements related to other interpretations among multiple interpretations to the user.

[0026] For example, an initial interpretation may be selected from among several interpretations as the one most likely to satisfy the oral utterance, based on the fact that the initial interpretation has the greatest predicted degree of correspondence to satisfy the oral utterance, and can be automatically initialized by an automated assistant in response to the oral utterance. For example, as shown in Figure 1B, the automated assistant can use a first music application to render an audio output in the first music application that corresponds to the song "Science Rocks" by artist "Pop Class". Nevertheless, one or more other interpretations having the next greatest degree of correspondence to the oral utterance may form the basis for a first selectable element 124 (e.g., the song "Science Rocks" by artist "Pop Class," but using a second music application instead of a first music application) and a second selectable element 126 (e.g., the song "Science Rocks" by artist "Music Class" instead of "Pop Class," but using the first music application). In this way, in some implementations, user 102 may select one or more selectable suggestion elements simultaneously with or after the execution of an action related to an initial interpretation that is expected to have the highest degree of correspondence with the spoken utterance.

[0027] In some implementations, selectable suggestions may be selected using assistant input, such as another verbal utterance and / or other input gestures to the computing device and / or any other device associated with the automated assistant. For example, as shown in Figure 1C, user 102 may give another verbal utterance 142, such as "Play on the second music application." In some implementations, the automated assistant may provide an indication 128 that the audio interface of the computing device 104 remains initialized to receive input directed to one or more of the selectable suggestions (for example, shown in Figure 1B). As shown in Figures 1C and 1D, in response to receiving the other verbal utterance 142, the automated assistant may determine that user 102 is giving a request for the automated assistant to stop performing the action associated with the automatically initialized initial interpretation and initialize the performance of an alternative action associated with an alternative interpretation corresponding to the first selectable element 124. In response to receiving other oral utterances 142, the automated assistant can automatically initialize the execution of alternative actions related to alternative interpretations corresponding to the first selectable element 124. Alternatively, in some implementations, the automated assistant may cause the graphical user interface 106 of the computing device 104 to render a graphical element 162 to indicate that an alternative action related to an alternative interpretation has been performed.

[0028] In some implementations, the automated assistant can cause the automated assistant and / or other applications to bias one or more ASR and / or NLU components toward content associated with the first selectable element 124 and the second selectable element 126, such as one or more slot value words or phrases corresponding to one or more of the selectable elements. For example, in the example in Figure 1C, the automated assistant can bias one or more processing of the oral utterance 142 toward words such as "second application" and / or "Music Group". In this way, when the automated assistant leaves the audio interface initialized (as shown by the graphical element 128), the user 202 may choose to give another oral utterance such as "second application". In response, the processing of the audio data capturing the oral utterance 142 may be biased toward any words associated with the first selectable element 124 and the second selectable element 126.

[0029] In some implementations, the prominence and / or area of ​​each selectable suggestion and / or graphical element compared to other selectable suggestions and / or graphical elements may be based on the predicted degree of correspondence of its respective interpretation of the request from user 102. For example, as shown in Figure 1B, the first area of ​​the display interface 106 associated with the first selectable element 124 and the second area of ​​the display interface 106 associated with the second selectable element 126 may occupy the same area in the display interface 106 when their respective degrees of correspondence with the oral utterance 108 are the same. However, when the predicted degree of correspondence related to the alternative interpretation of the first selectable element 124 is greater than that of the alternative interpretation of the second selectable element 126, the first area of ​​the display interface 106 associated with the first selectable element 124 may be larger than the second area of ​​the display interface 106 associated with the second selectable element 126. In some implementations, optional suggestions corresponding to some or all of the alternative interpretations may be omitted from the display interface 106 of the computing device 104, based on the metrics.

[0030] Figures 2A, 2B, and 2C show scenes 200, 220, and 240, respectively, in which user 202 invokes an automated assistant that can provide selectable suggestions on the display interface 206 of computing device 204, while simultaneously initializing the execution of a specific interpretation of the user input in response to the user input. The selectable suggestions may be generated based on the natural language content of the user input, which may have multiple different interpretations. For example, as shown in Figure 2A, user 202 may give an oral utterance 208 such as "Assistant, translate to 20 percent." As an example throughout Figures 2A–2C, suppose user 202 gives the oral utterance 208 to adjust the brightness level of one or more lights in user 202's house to 20%. However, since the oral utterance 208 may have multiple different interpretations, the automated assistant can process the oral utterance 208 to determine whether other suggestions should be rendered to present to user 202.

[0031] For example, in response to receiving an oral utterance 208, the automated assistant may render a first selectable element 222 and a second selectable element 224 on the display interface 206 of the computing device 204, as shown in Figure 2B. The first selectable element 222 may correspond to a first interpretation that the automated assistant or other application is expected to most likely satisfy the request embodied in the oral utterance 208. The second selectable element 224 may correspond to a second interpretation that is expected to be less likely to satisfy the request than the first interpretation associated with the first selectable element 222. In some implementations, the automated assistant may execute an assistant command related to the first interpretation in response to the oral utterance while simultaneously rendering a second selectable element 224 that allows the user 202 to initialize the execution of an alternative assistant command related to the second interpretation.

[0032] Alternatively or additionally, the automated assistant may be able to automatically execute an assistant command related to a first interpretation in response to receiving a verbal utterance 208. The automatic execution of an assistant command and / or an alternative assistant command may occur when one or more “costs” of executing the assistant command and / or alternative assistant command meet a threshold. For example, a value may be estimated for the amount of processing and / or time expected to be consumed to perform the proposed user intent. When the value meets a certain threshold, the automated assistant may be able to execute an assistant command and / or an alternative assistant command in response to the verbal utterance 208. For example, as shown in Figure 2B, the automated assistant may be able to automatically execute an assistant command related to a first selectable element 222, thereby enabling an IoT home application to adjust the brightness of the lights in the user's home (e.g., “Kitchen Lights”) from 50% to 20%. Furthermore, the automated assistant may also be able to enable a translation application to translate a portion of the verbal utterance 208 (e.g., “To 20%”). In this case, both the assistant command and the alternative assistant command can be executed without consuming a lot of processing and / or time, so both the assistant command and the alternative assistant command can be executed. However, if one of the assistant commands requires an amount of processing and / or time that exceeds a threshold, or otherwise would negatively affect the user or additional users, neither the assistant command nor the alternative assistant command may be executed.

[0033] In some implementations, an automated assistant may respond to a verbal utterance 208 by causing the display interface 206 of the computing device 204 to render specific selectable elements, but the user 202 may want additional alternative interpretations for selection. For example, as shown in Figure 2C, the user 202 may give the computing device 204 an input gesture 250 to cause the display interface 206 to render a third selectable element 242 and a fourth selectable element 246. The prominence of each selectable element may depend on the expected relevance of each element to the user input. For example, the area of ​​the third selectable element 242 and the area of ​​the fourth selectable element 246 may be smaller than the areas of the first selectable element 222 and the second selectable element 224, respectively. This may be based in part on the predicted degree of correspondence between the third interpretation (controllable via the third selectable element 242) and the fourth interpretation (controllable via the fourth selectable element 246) of the oral utterance 208. In other words, since the first interpretation associated with the first selectable element 222 and the second interpretation associated with the second selectable element 224 are predicted to have a higher probability of satisfying the requirements embodied in the oral utterance 208 than the third and fourth interpretations, the display areas of the third selectable element 242 and the fourth selectable element 246 may be smaller than those of the first selectable element 222.

[0034] In some implementations, the input gesture 250 can cause an automated assistant to select one or more other slot values ​​relating to the first and / or second interpretations. These selected slot values ​​may form the basis for the third and fourth selectable elements 242 and 246. For example, the first interpretation may include a specific slot value that identifies "Kitchen Lights" with respect to the lighting position parameter of the first interpretation, and the third interpretation may include another slot value that identifies "Basement Lights" with respect to the lighting position parameter of the third interpretation. Furthermore, the fourth interpretation may include a different slot value that identifies "Hallway Lights" with respect to the lighting position parameter of the fourth interpretation.

[0035] Figure 3 shows a system 300 that can invoke an automated assistant 304 to suggest alternative interpretations in response to user input, while simultaneously executing a specific interpretation predicted to be most relevant to the user input. The automated assistant 304 can operate as part of an assistant application provided on one or more computing devices, such as computing device 302 (e.g., computing device 104 in Figures 1A–1D, computing device 204 in Figures 2A–2C, and / or other computing devices such as server devices). The user can interact with the automated assistant 304 via an assistant interface 320, which can be a microphone, camera, touchscreen display, user interface, and / or any other device capable of providing an interface between the user and the automated assistant. For example, the user can initialize the automated assistant 304 by providing verbal, text, and / or graphical input to the assistant interface 320 to cause the automated assistant 304 to initialize one or more actions (e.g., providing data, controlling a peripheral device, accessing an agent, generating input and / or output, etc.). Alternatively, the automated assistant 304 may be initialized based on processing of context data 336 using one or more trained machine learning models. The context data 336 may characterize one or more features of the environment accessible to the automated assistant 304, and / or one or more features of users who are expected to interact with the automated assistant 304. The computing device 302 may include a display device which is a display panel that includes a touch interface for receiving touch input and / or gestures to enable a user to control the application 334 of the computing device 302 via the touch interface.In some implementations, the computing device 302 may not have a display device, thereby providing an audible user interface output without providing a graphical user interface output. Furthermore, the computing device 302 may provide a user interface such as a microphone for receiving oral natural language input from the user. In some implementations, the computing device 302 may include a touch interface, or it may not have a camera or other visual component, but may optionally include one or more other sensors.

[0036] Computing device 302 and / or other third-party client devices may optionally communicate with a server device via a network such as the Internet to implement system 300. Furthermore, computing device 302 and any other computing devices may communicate with each other via a local area network (LAN), such as a Wi-Fi network. In some implementations, computing device 302 may offload computing tasks to a server device to conserve computing resources of computing device 302. For example, the server device may host an automated assistant 304, and / or computing device 302 may send inputs received in one or more assistant interfaces 320 to the server device. However, in some additional or alternative implementations, the automated assistant 304 may be hosted locally on computing device 302, and various processes that may be associated with the operation of the automated assistant may be executed on computing device 302.

[0037] In various implementations, all or some aspects of the automated assistant 304 may be implemented on the computing device 302. In some of these implementations, aspects of the automated assistant 304 may interface with a server device that is implemented by the computing device 302 and can implement other aspects of the automated assistant 304. The server device may optionally serve multiple users and their associated assistant applications through multiple threads. In implementations where all or some aspects of the automated assistant 304 are implemented by the computing device 302, the automated assistant 304 can be an application separate from the operating system of the computing device 302 (e.g., installed "on top of" the operating system) -- or alternatively, it can be implemented directly by the operating system of the computing device 302 (e.g., considered an application of the operating system but integrated with the operating system).

[0038] In some implementations, the automated assistant 304 may include an input processing engine 306, which may employ several different modules to process inputs and / or outputs for the computing device 302 and / or the server device. For example, the input processing engine 306 may include a speech processing engine 308 that can process audio data capturing oral utterances received at the assistant interface 320 in order to identify the text embodied in the oral utterances. The audio data may be sent from the computing device 302 to the server device, for example, to conserve the computing resources of the computing device 302, and the server device may send the text embodied in the oral utterances back to the computing device 302. Additionally or alternatively, the audio data may be processed exclusively at the computing device 302.

[0039] The process for converting audio data to text may include a speech recognition algorithm that can employ neural networks and / or statistical models to identify groups of audio data corresponding to words or phrases (e.g., using an ASR model). The text converted from the audio data may be parsed by a data analysis engine 310 and made available to the automated assistant 304 as text data that can be used to generate and / or identify user-specified command phrases, intentions, actions, slot values, and / or any other content (e.g., using an NLU model). In some implementations, the output data provided by the data analysis engine 310 may be provided to a parameter engine 312 to determine whether the user has provided input corresponding to specific intentions, actions, and / or routines that can be performed by the automated assistant 304 and / or applications or agents that can be accessed by the automated assistant 304 (e.g., using an execution model and / or rules). For example, assistant data 338 may be stored in a server device and / or computing device 302 and may include data defining one or more actions that can be performed by the automated assistant 304 and the parameters required to perform the actions. The parameter engine 312 can generate one or more parameters relating to intent, action, and / or slot values, and provide one or more parameters to the output generation engine 314. The output generation engine 314 can use one or more parameters to communicate with the assistant interface 320 to provide output to the user, and / or with one or more applications 334 to provide output to one or more applications 334.Output to the user may include, for example, visual output that can be visually rendered for presentation to the user via the display interface of computing device 302 or the display of an additional computing device communicating with computing device 302; audible output that can be rendered to be heard for presentation to the user via the speaker of computing device 302 or the speaker of an additional computing device communicating with computing device 302; smart device control commands that control one or more networked smart devices communicating with computing device 304; and / or other outputs.

[0040] An automated assistant application may include, and / or have access to, on-device ASR, on-device NLU, and on-device execution. For example, on-device ASR may be performed using an on-device ASR module that processes audio data (detected by the microphone) using an ASR model stored locally on computing device 302. The on-device ASR module generates ASR outputs based on processing the audio data, such as one or more speech hypotheses corresponding to recognized text of oral utterances (if any) present in the audio data end-to-end, or predicted phonemes that are expected to correspond to the oral utterances, and the speech hypotheses corresponding to the recognized text may be generated based on the predicted phonemes. Alternatively, for example, an on-device NLU module may process the speech hypotheses corresponding to recognized text generated by the ASR module using an NLU model to generate NLU data. The NLU data may include predicted intentions corresponding to oral utterances and, optionally, slot values ​​of parameters associated with the intentions. Furthermore, for example, an on-device execution module may process NLU data using execution models and / or rules, and optionally other local data, to determine assistant commands to execute in order to perform the anticipated intent of a verbal utterance. These assistant commands may include, for example, obtaining local and / or remote responses (e.g., answers) to the verbal utterance, interactions with locally installed applications to execute based on the verbal utterance, commands to send (directly or via a corresponding remote system) to an Internet of Things (IoT) device based on the verbal utterance, and / or other resolution actions to execute based on the verbal utterance. The on-device execution may then initiate local and / or remote performance / execution of the determined actions to fulfill the verbal utterance.

[0041] In various implementations, remote ASR, remote NLR, and / or remote execution may be used, at least selectively. For example, speech hypotheses corresponding to recognized text may be sent, at least selectively, to a remote automated assistant component for remote NLU and / or remote execution. For example, speech hypotheses corresponding to recognized text may optionally be sent for remote execution in parallel with on-device execution, or in response to failure of on-device NLU and / or on-device execution. However, on-device ASR, on-device NLU, on-device execution, and / or on-device execution may be preferred, at least because of the reduced latency they provide when fulfilling oral utterances (by eliminating the need for client-server round trips to fulfill oral utterances). Furthermore, on-device functionality may be the only functionality available in situations where network connectivity is absent or limited.

[0042] In some implementations, the computing device 302 may include one or more applications 334 that may be provided by a first-party entity which is the same entity that provided the computing device 302 and / or the automated assistant 304, and / or by a third-party entity which is different from the entity that provided the computing device 302 and / or the automated assistant 304. The application state engine of the automated assistant 304 and / or the computing device 302 can access application data 330 to determine one or more actions that may be performed by one or more applications 334, as well as the state of each application of one or more applications 334, and / or the state of each device associated with the computing device 302. The device state engine of the automated assistant 304 and / or the computing device 302 can access device data 332 to determine one or more actions that may be performed by the computing device 302 and / or one or more devices associated with the computing device 302 (for example, one or more networked smart devices that communicate with the computing device 302). Furthermore, application data 330 and / or any other data (e.g., device data 332) can be accessed by an automated assistant 304 to generate context data 336, which can characterize the context in which a particular application 334 and / or device is running, as well as the context in which a particular user is accessing the computing device 302, application 334, and / or any other device or module.

[0043] While one or more applications 334 are running on the computing device 302, device data 332 can characterize the current operational state of each application 334 running on the computing device 302 and / or remotely running on the computing device 302 (for example, one or more streaming applications). Furthermore, application data 330 can characterize one or more features of the running applications 334, such as the contents of one or more graphical user interfaces being rendered at the direction of one or more applications 334. Alternatively or additionally, application data 330 can characterize action schemas, which can be updated by each application and / or an automated assistant 304 based on the current operational state of each application. Alternatively or additionally, one or more action schemas for one or more applications 334 may remain static but be accessed by the application state engine to determine a suitable action to initialize via the automated assistant 304.

[0044] The computing device 302 may further include an assistant invocation engine 322 that can process application data 330, device data 332, context data 336, and / or any other data accessible to the computing device 302 using one or more trained machine learning models. The assistant invocation engine 322 can determine whether to wait for the user to explicitly speak an invocation phrase to invoke the automated assistant 304 (and whether the user has spoken the invocation phrase to invoke the automated assistant 304), or—instead of requiring the user to explicitly speak the invocation phrase—process the data to determine that the data indicates the user's intention to invoke the automated assistant. For example, one or more trained machine learning models may be trained using instances of training data based on scenarios in which the user is in an environment where multiple devices and / or applications exhibit various operating states. Instances of training data may be generated to capture training data that characterizes at least audio data containing an invocation phrase, and / or contexts in which the user invokes the automated assistant, and other contexts in which the user does not invoke the automated assistant. When one or more trained machine learning models are trained on these instances of the training data, the assistant invocation engine 322 can cause the automated assistant 304 to detect or restrict detection of verbal invocation phrases from the user based on contextual and / or environmental characteristics. Additionally or alternatively, the assistant invocation engine 322 can cause the automated assistant 304 to detect or restrict detection of one or more assistant commands from the user based on contextual and / or environmental characteristics, without requiring the user to provide any invocation phrases.

[0045] In some implementations, system 300 may include a suggestion generation engine 316 that can assist in processing user input to determine whether a particular interpretation of a spoken utterance should be performed and / or to provide one or more suggestions regarding alternative interpretations of the spoken utterance. In response to receiving assistant input, the suggestion generation engine 316 may identify several different interpretations that have a certain degree of correspondence with the request embodied in the assistant input, such as different combinations of slot values ​​for predicted intent and / or parameters related to the predicted intent. The degree of correspondence of a particular interpretation may be characterized by a metric that can have a value that can be compared to one or more different thresholds. For example, a particular interpretation of a spoken utterance may be identifiable in response to user input, and the metric for that particular interpretation may be determined to meet a target threshold. Based on this determination, the automated assistant 304 may automatically initialize the execution of the action associated with the particular interpretation and omit the rendering of suggestion elements corresponding to other interpretations associated with lower value metrics.

[0046] In some implementations, if a particular interpretation has the highest metric value for a particular user input, but the metric value does not meet a target threshold, the automated assistant 304 may decide to render one or more additional suggestion elements. For example, the automated assistant may still automatically initialize the execution of the action associated with a particular interpretation, but may also render selectable elements corresponding to other alternative interpretations that have the highest metric value. These selectable elements may be rendered on the display interface of the computing device 302 without any additional user input requesting the selectable elements. In some implementations, the prominence of one or more of these selectable elements may be based on the respective metric values ​​associated with each alternative interpretation of the selectable element. In some implementations, the metric value associated with each selectable suggestion element may be compared to a suggestion threshold to determine whether a corresponding selectable element among the selectable elements should be rendered. For example, in some implementations, an alternative interpretation associated with a metric value that does not meet the suggestion threshold may still be rendered in response to subsequent user input (e.g., a swipe gesture to reveal additional suggestions).

[0047] In some implementations, the system 300 may include a proposed feature engine 318 that can generate assistant GUI data for rendering one or more GUI elements in response to receiving user input. The assistant GUI data may feature the size of a particular selectable element based on corresponding metrics associated with that particular selectable element; the placement of the particular selectable element on the display based on corresponding metrics associated with that particular selectable element (e.g., the horizontal and / or vertical displacement of the particular selectable element is characterized by the corresponding placement data of the particular selectable element); display data of the particular selectable element on the display that characterizes the corresponding display characteristics of the particular selectable element (e.g., bold characteristics of the particular selectable element, fill characteristics of the particular selectable element against a background, and / or any other display characteristics related to visually rendering the particular selectable element); the amount of the particular selectable element displayed based on the size of the display interface of the computing device 302 and / or based on the amount of alternative descriptive interpretations having corresponding metrics that satisfy the proposed thresholds; and / or other GUI data. In particular, GUI data can also characterize control elements for actions associated with initial interpretations that are expected to satisfy oral utterances. Control elements can be based on the type of action and may include, for example, media control elements when the action involves playing media content, slider elements for adjusting parameters of IoT devices when the action involves controlling IoT devices (e.g., thermostat temperature, smart lighting brightness, etc.), and / or other control elements based on the type of action.

[0048] In some implementations, system 300 may include an execution engine 326 that can determine, in response to user input, whether to initialize the execution and / or partial execution of one or more user intentions identified by an automated assistant 304. For example, the execution engine 326 may, in response to user input, perform an action related to a particular interpretation of a spoken utterance when a metric related to that particular interpretation meets a target threshold. Alternatively or additionally, if the identified interpretation of a spoken utterance does not meet the target threshold, the execution engine 326 may perform a particular interpretation associated with the highest metric (e.g., the greatest degree of correspondence with user input). In some implementations, one or more alternative interpretations with the next highest metric (e.g., the next greatest degree of correspondence with user input) may be at least partially performed and / or executed by the execution engine 326. For example, data may be retrieved by the automated assistant 304 and / or another application to facilitate the at least partially performed alternative interpretation with the next highest metric.

[0049] In some implementations, System 300 may include a training data engine 324 capable of generating training instances for initially training one or more machine learning models as described herein (for example, used when generating corresponding metrics as described in relation to Figure 1), and generating training instances for updating one or more machine learning models based on how the user interacts with the automated assistant 304. Initial training of one or more machine learning models based on multiple training instances is described above in relation to Figure 1, and multiple training instances may be generated using the training data engine 324. Furthermore, the selection or non-selection of any selectable elements related to alternative interpretations provided to the user for presentation as described herein may be used to generate additional training instances for updating one or more machine learning models. For example, suppose the user does not select any of the selectable elements related to alternative interpretations. In this example, signals and / or interpretations related to oral utterances may be used as input to a training instance, and a ground truth output indicating the initially selected interpretation may be used to positively reinforce the initial interpretation selection. In contrast, suppose the user selects one of the selectable elements related to alternative interpretations. In this example, signals and / or interpretations associated with oral utterances can be used as input for training cases, and ground truth outputs indicating alternative interpretations can be used to positively reinforce the selection of the alternative interpretations (while simultaneously negatively reinforcing the selection of the initial interpretation). Thus, one or more machine learning models can be updated over time so that the automated assistant 304 can improve its selection of the initial interpretation. For example, if the selected initial interpretation is associated with a first application for media playback, but the user generally chooses an alternative interpretation associated with a second application for media playback, one or more machine learning models can be updated over time to reflect the user's preference for the second application for media playback.

[0050] Figures 4A and 4B illustrate methods 400 and 420 for operating an automated assistant to provide alternative interpretations in response to oral utterances that may be interpreted in various ways and / or may lack certain parameters. Methods 400 and 420 may be performed using one or more applications, computing devices, and / or any other devices or modules that can interface with the automated assistant. Method 400 may include an action 402 that determines whether an assistant input has been detected by the computing device. The assistant input can be an oral utterance containing a request from a user to cause the automated assistant to execute an assistant command. In response to receiving an oral utterance, the automated assistant may perform one or more actions (e.g., an ASR action, an NLU action, a perform action, and / or other actions) to generate multiple interpretations of the oral utterance, such as an oral utterance like "Assistant, play Congolese." A user might intend to instruct an automated assistant to play Congolese music through spoken utterances, but the automated assistant can respond to receiving the spoken utterance by automatically initializing certain actions related to the first interpretation, and further identify other actions to suggest to the user based on alternative interpretations.

[0051] Method 400 can proceed from operation 402 to operation 404 when assistant input is detected. Otherwise, the automated assistant can continue to detect user input in operation 402. Operation 404 may include generating input data that can identify different interpretations of the oral utterance. For example, in response to the aforementioned oral utterance, the automated system may generate assistant input data that characterizes different interpretations of the oral utterance, such as playing Congolese music by a first application (e.g., playMusic("Congolese", first application)), translating the speech into a language spoken in Congo (e.g., translateSpeech("play", English, Swahili)), and / or playing Congolese music by a second application (e.g., playMusic("Congolese", second application)).

[0052] Method 400 can proceed from operation 404 to operation 406, which may include generating metric data characterizing the predicted degree of correspondence between user requests contained in oral utterances and several different interpretations of the oral utterances. In other words, for each interpretation, metrics may be generated to characterize the degree of correspondence between each interpretation and the oral utterances. For example, to advance the above example, the interpretation for playing music in the first application may have the greatest degree of correspondence with the user's request compared to the other interpretations. The interpretation for playing music in the second application may have the next greatest degree of correspondence, and another interpretation for translating speech may have a particular degree of correspondence that is smaller than the other two interpretations for playing music.

[0053] Method 400 can proceed from action 406 to action 408, which may include determining whether a metric for a particular interpretation meets a metric threshold. The metric threshold can be the threshold to which the automated assistant determines that a corresponding interpretation is expected to satisfy the oral utterance to such an extent that additional suggested interpretations may not be accepted by the automated assistant (e.g., 75%, 60%, or any other value). Thus, when a metric for a particular interpretation meets the metric threshold, the automated assistant can initialize the execution of the action associated with the interpretation without rendering any other suggested interpretations. However, when none of the metrics for any of the identified interpretations meet the metric threshold, the automated assistant may initialize the execution of the action associated with the interpretation corresponding to the metric that is closest to meeting the metric threshold, and simultaneously propose other interpretations to the user.

[0054] When the metric for a particular interpretation meets the metric threshold, method 400 can proceed from operation 408 to operation 410. Otherwise, when none of the metrics for any particular interpretation meet the metric threshold, method 400 can proceed from operation 408 to operation 412, which is described below. Operation 410 may include causing a computing device to initialize an execution of a particular interpretation corresponding to a metric that meets the metric threshold. For example, when the metric corresponding to the interpretation of playing music by a first application meets the metric threshold, the automated assistant can cause a computing device to initialize an execution of an operation associated with that particular interpretation. In such a case, the automated assistant may be objectively confident that the execution of an operation associated with the particular interpretation will satisfy the user's request, but the automated assistant may optionally offer the user an operation associated with an alternative interpretation.

[0055] For example, method 400 can proceed from operation 410 to operation 422 provided in method 420 of Figure 4B via continuation element "A". Operation 422 may include determining whether input has been received and is intended to identify an alternative interpretation. For example, a user may perform a swipe gesture on the display interface of a computing device to cause an automated assistant to provide additional suggestions of other actions relating to other interpretations that may be performed in response to a verbal utterance. These other interpretations may be identified, for example, in operation 404, but did not have a corresponding metric that met the threshold of the metric. For example, when a computing device is playing music by a first application, a user may give the computing device an input gesture or other assistant input (for example, swiping on the display interface of the computing device) to reveal one or more selectable elements. Each of the one or more selectable elements may correspond to an action relating to an alternative interpretation of a verbal utterance. For example, when the user swipes the display interface while the first application is playing music, a first selectable element corresponding to the action of playing music in the second application may be rendered, as indicated by the alternative interpretation.

[0056] When it is determined in operation 422 that user input has been received, method 420 may proceed from operation 422 to operation 416, described below, via continuation element "B". When in operation 408 none of the metrics for any interpretation meet the metric threshold, method 400 may proceed from operation 408 to operation 412. Operation 412 may include identifying a particular interpretation that has the greatest degree of correspondence with a user request contained in an oral utterance. For example, the metric for an interpretation to play music in the first application may not meet the metric threshold, but nevertheless the metric may have the greatest degree of correspondence with a user request contained in an oral utterance. In other words, the execution of an interpretation may be predicted to be at least the most likely to satisfy the request compared to other interpretations identified based on the oral utterance.

[0057] Method 400 can proceed from operation 412 to operation 414, operation 414 may include causing a computing device to automatically initialize the execution of an operation related to a particular interpretation. According to the example above, this operation related to a particular interpretation may cause a first application to play Congolese music. Thus, the computing device can render Congolese music through the first application and the computing device's audio interface. Method 400 can proceed from operation 414 to operation 416, operation 416 may include identifying one or more other interpretations that may satisfy the requirements contained in the oral utterance. For example, operation 416 may include identifying alternative interpretations related to user data generated in operation 404.

[0058] Alternatively or additionally, the automated assistant may identify one or more other alternative interpretations that may not have been identified by the input data generated in action 404. For example, based on the actions associated with a particular interpretation performed in action 414, and / or contextual data associated with user requests contained in the oral utterance, the automated assistant may identify other content and / or actions associated with other interpretations to suggest to the user. For example, the automated assistant may identify actions that have interpretations that may be similar to the user's intent currently being performed, but may include other parameters, or other slot values ​​for parameters or other parameters. For example, when an action associated with a particular interpretation of playing music is being performed, other actions associated with suggested alternative interpretations may include slot values ​​based on their similarity to the speech characteristics of the oral utterance during the ongoing processing (e.g., phonemes detected). For example, the automated assistant may determine that the word "Portuguese" has a degree of correspondence and / or similarity to the oral word "Congolese" spoken by the user in the oral utterance. Based on this determination, the automated assistant can generate suggestion data characterizing selectable suggestions for playing Portuguese music by the first application as an alternative interpretation.

[0059] Method 400 can proceed from operation 416 to operation 418, operation 418 which may include causing a computing device to render one or more selectable elements corresponding to one or more other interpretations. One or more selectable elements may be rendered concurrently with specific selectable elements corresponding to a particular interpretation being performed by an automated assistant. For example, an audio control interface for controlling audio playback of content by a first application may be rendered concurrently with one or more selectable elements identified in operation 416.

[0060] Method 400 can proceed from operation 418 to operation 424 via continuation element "C", as shown in method 420 in Figure 4B. Operation 424 may include determining whether a user selection has been received for selecting another selectable element. If it is determined that user input has been received for selecting another selectable element from one or more selectable elements, Method 420 can proceed to operation 426. Operation 426 may include causing a computing device to initialize the execution of an alternative operation related to a particular alternative interpretation corresponding to the selected selectable element. Alternatively or additionally, in response to a selection of a selectable element, the automated assistant may cause the execution of an alternative operation to cancel and / or pause the ongoing execution of the first selected operation related to a particular interpretation (for example, playing Congo music by the first application).

[0061] In some implementations, method 420 may optionally proceed from action 426 to any action 428. Alternatively, if no input is received from the selection of another selectable element, method 420 may proceed from action 424 to any action 428. Action 428 may include training (or updating) one or more machine learning models in response to the selection or non-selection of one or more selectable elements. For example, if another selectable element corresponds to a different interpretation than the intent of the user being executed, the trained machine learning models may then be trained (or updated) so that subsequent interpretations are biased towards the selected interpretation in response to the receipt of subsequent instances of the request. As a result, an alternative interpretation (e.g., playing Congo music in a second application) may be initialized in response to the receipt of subsequent instances of the request.

[0062] Figure 5 is a block diagram 500 of an exemplary computer system 510. Generally, the computer system 510 includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices may include, for example, a storage subsystem 524 including memory 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with the computer system 510. The network interface subsystem 516 provides an interface to an external network and is coupled to a corresponding interface device of another computer system.

[0063] The user interface input device 522 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, audio input devices such as scanners, touchscreens integrated into displays, speech recognition systems, microphones, and / or other types of input devices. In general, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computer system 510 or a communication network.

[0064] The user interface output device 520 may include non-visual displays such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include flat panel devices such as cathode ray tubes (CRTs) or liquid crystal displays (LCDs), projection devices, or any other mechanism for generating visible images. The display subsystem may also provide non-visual displays, such as through an audio output device. In general, the use of the term “output device” is intended to include all possible types of devices and methods for outputting information from the computer system 510 to a user or another machine or computer system.

[0065] The storage subsystem 524 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 524 may include logic for performing selected embodiments of Method 400 and Method 420, and / or for implementing one or more of the System 300, computing device 104, computing device 204, automated assistant, and / or any other applications, devices, apparatus, and / or modules considered herein.

[0066] These software modules are generally executed by processor 514 alone or in combination with other processors. The memory 525 used in the storage subsystem 524 may include several memories, including a main random access memory (RAM) 530 for storing program instructions and data, and a read-only memory (ROM) 532 for storing fixed instructions. The file storage subsystem 526 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing the functionality of a particular implementation may be stored by the file storage subsystem 526 within the storage subsystem 524, or on other machines that can be accessed by processor 514.

[0067] The bus subsystem 512 provides a mechanism for various components and subsystems of the computer system 510 to communicate with each other as intended. Although the bus subsystem 512 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0068] The computer system 510 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, client devices (for example, computing device 104 in Figures 1A–1D, computing device 302 in Figure 3, and / or other client devices), or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computer system 510 shown in Figure 5 is intended only as a specific example to illustrate several implementations. Many other configurations of the computer system 510 are possible, having more or fewer components than the computer system shown in Figure 5.

[0069] Where the systems described herein may collect or use personal information about a user (or more often referred to herein as “Participant”), the user may be given the opportunity to control whether the program or feature collects user information (for example, information about the user’s social networks, social behavior or activities, occupation, user preferences, or the user’s current geographical location) or whether and / or how content that may be more relevant to the user should be received from the content server. Furthermore, certain data may be processed in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user’s identity may be processed so that personally identifiable information cannot be determined about the user, or if geographical location information is obtained, the user’s geographical location may be generalized (to the level of city, zip code, or state, for example) so that the user’s specific geographical location cannot be determined. Thus, the user may be able to control how and / or how information is collected about them.

[0070] While several implementations are described and illustrated herein, various other means and / or structures may be used to perform the functions described herein and / or to obtain one or more of the results and / or benefits, and each of such changes and / or modifications shall be considered within the scope of the implementations described herein. More broadly, all parameters, dimensions, materials and configurations described herein are intended to be illustrative, and actual parameters, dimensions, materials and / or configurations shall depend on the particular one or more applications in which the teaching is used. A person skilled in the art can recognize or determine many equivalents of the particular implementations described herein by means of ordinary experimentation alone. Thus, it should be understood that the implementations described herein are presented merely as examples, and within the scope of the appended claims and their equivalents, implementations may be carried out in ways different from those specifically described and claimed. Implementations of this disclosure shall apply to each individual feature, system, item, material, kit and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included in the scope of this disclosure, provided that such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

[0071] In some implementations, a method is provided which is performed by one or more processors, and in a computing device, the steps are: receiving a user's oral utterance directed to an automated assistant, the oral utterance including an assistant command to be performed by the automated assistant; generating assistant input data based on the oral utterance, characterizing a plurality of interpretations that are expected to respond to the oral utterance, each interpretation including a corresponding intent, one or more corresponding parameters associated with the corresponding intent, and one or more corresponding slot values ​​for each of the one or more corresponding parameters, and each interpretation including at least one unique corresponding slot value; and generating a plurality based on the assistant input data The process includes the steps of: generating metric data characterizing the predicted degree of correspondence between each interpretation and an assistant command contained in an oral utterance; automatically initializing an automated assistant to perform a first action associated with a first interpretation of a plurality of interpretations for carrying out an assistant command contained in an oral utterance, based on the metric data and assistant input data; and rendering one or more selectable suggestion elements on the display interface of a computing device, based on the metric data and assistant input data, wherein each of the one or more selectable suggestion elements is associated with a corresponding alternative interpretation of a plurality of interpretations for carrying out an assistant command contained in an oral utterance. A user selection of a given selectable suggestion element from one or more selectable suggestion elements causes the automated assistant to initialize the performance of a corresponding alternative action associated with the given selectable suggestion element.

[0072] These and other implementations of the technologies disclosed herein may optionally include one or more of the following features:

[0073] In some implementations, the step of automatically initializing the automated assistant to perform a first action to carry out an assistant command contained in a verbal utterance may involve the first application creating an instance of specific content.

[0074] In some implementations, the method may further include a step of biasing an automated speech recognition (ASR) process or a natural language understanding (NLU) process toward content related to one or more selectable suggestion elements, in response to automatically initializing an automated assistant to perform a first action for carrying out an assistant command contained in a verbal utterance.

[0075] In some implementations, the method may further include the step of allowing the automated assistant to access application data to facilitate its preparation for performing a corresponding alternative action related to one or more selectable suggestion elements, based on metric data and assistant input data.

[0076] In some implementations, the method may further include the step of determining, based on the oral utterance, that one or more of the corresponding slot values ​​of one or more corresponding parameters associated with a plurality of interpretations were not specified by the user in the oral utterance. The automated assistant may infer one or more specific slot values ​​of the corresponding parameters associated with the first interpretation. In some versions of those implementations, the method may further include the step of inferring the alternative specific slot value for each of the corresponding alternative interpretations based on the oral utterance. The user's selection of a given selectable suggestion element may cause the alternative behavior to be initialized using the alternative specific slot value. In some further versions of those implementations, the specific slot value may identify a first application for rendering specific content, and the alternative specific slot value may identify a different second application for rendering the alternative specific content. In some further additional or alternative versions of those implementations, the specific slot value may identify a reference to a first entity for rendering specific content, and the alternative specific slot value may identify a reference to a different second entity for rendering the alternative specific content.

[0077] In some implementations, the step of rendering one or more selectable suggestion elements on the computing device's display interface may include, after automatically initializing an automated assistant to perform a first action to carry out an assistant command contained in a verbal utterance, rendering one or more selectable suggestion elements on the computing device's display interface for a duration of a threshold.

[0078] In some implementations, a method is provided, performed by one or more processors, that a computing device receives a user's oral utterance directed to an automated assistant, the oral utterance comprising an assistant command to be performed by the automated assistant; generates metric data that identifies a first metric characterizing the degree to which a first action is expected to satisfy the assistant command, and a second metric characterizing another degree to which a second action is expected to satisfy the assistant command, based on the oral utterance; generates GUI data that characterizes an assistant graphical user interface (GUI) that responds to the oral utterance, based on the first and second actions; causes the automated assistant to automatically initialize the execution of the first action in response to the receipt of the oral utterance; and causes the display interface of the computing device to render the assistant GUI according to the GUI data and metric data. The GUI data is generated to identify a first selectable element and a second selectable element, the first selectable element being selectable to control the execution of the first action, and the second selectable element being selectable to automatically initialize the execution of the second action.

[0079] These and other implementations of the technologies disclosed herein may optionally include one or more of the following features:

[0080] In some implementations, when the degree to which the first action is expected to satisfy an assistant command is greater than another degree to which the second action is expected to satisfy an assistant command, the first selectable element may be positioned more prominently than the second selectable element in the assistant GUI.

[0081] In some implementations, the step of causing the display interface to render the assistant GUI according to GUI data and metric data may include positioning a first selectable element adjacent to a second selectable element. When the degree to which the first action is expected to satisfy an assistant command is greater than another degree to which the second action is expected to satisfy an assistant command, the first area of ​​the first selectable element in the assistant GUI may be larger than the second area of ​​the second selectable element.

[0082] In some implementations, the step of causing the display interface to render the assistant GUI according to GUI data and metric data may include positioning a first selectable element adjacent to a second selectable element. In the assistant GUI, the first selectable element may be positioned according to first positioning data that is unique to the first selectable element and characterizes the corresponding first position of the first selectable element in the assistant GUI, and the second selectable element may be positioned according to second positioning data that is unique to the second selectable element and characterizes the corresponding second position of the second selectable element in the assistant GUI.

[0083] In some implementations, in the assistant GUI, a first selectable element may be displayed based on corresponding first display data that is specific to the first selectable element and characterizes the corresponding first display characteristic of the first selectable element in the assistant GUI, and a second selectable element may be displayed based on corresponding second display data that is specific to the second selectable element in the assistant GUI and characterizes the corresponding second display characteristic of the second selectable element in the assistant GUI.

[0084] In some implementations, the first action may be a specific action performed on a computing device, and the second action may be a specific action performed on another computing device. Furthermore, the user's selection of the second selectable element may cause the specific action to be initialized on the other computing device.

[0085] In some implementations, the second selectable element may be selectable by touch input on the display interface of the computing device while the first operation is being performed.

[0086] In some implementations, the second selectable element may be selectable by an additional oral utterance that occurs at the same time as the first action is performed.

[0087] In some implementations, the prominence of rendering a first selectable element compared to a second selectable element in the assistant GUI may be based on the difference between the first and second metrics. In some versions of those implementations, if the difference between the first and second metrics does not meet the proposed threshold, the second selectable element does not need to be rendered to the display interface, and a specific touch gesture given by the user to the computing device's display interface may cause the second selectable element to be rendered to the computing device's display interface.

[0088] In some implementations, a method is provided, performed by one or more processors, that a computing device receives a user's oral utterance directed to an automated assistant, the oral utterance comprising an assistant command to be performed by the automated assistant; in response to receiving the oral utterance, identify a first action that can be initialized by the automated assistant to perform the assistant command; in response to the oral utterance, automatically initialize the execution of the first action; based on the degree to which the first action is expected to respond to the oral utterance, identify at least a second action that can be initialized by the automated assistant to perform the assistant command; and based on the oral utterance, cause the display interface of the computing device to render at least a second selectable element related to the second action, the second selectable element causing the automated assistant to initialize the second action instead of the first action when selected.

[0089] These and other implementations of the technologies disclosed herein may optionally include one or more of the following features:

[0090] In some implementations, the step of identifying a second action that can be initialized by an automated assistant to perform an assistant command may include generating a metric that characterizes the degree to which the first action is expected to respond to a verbal utterance, and determining whether the metric meets a metric threshold. The automated assistant may identify at least a second action to suggest to the user when the metric does not meet the metric threshold.

[0091] In some implementations, the step of identifying a second action that can be initialized by an automated assistant to perform an assistant command may include determining whether the first action is of a particular type. The automated assistant may decide to identify at least a second action if the first action is of a particular type, and the identified second action may be of a different type than the particular type of action of the first action. In some versions of those implementations, the particular type of action may include a communication action that involves communicating with a different user.

[0092] In some implementations, the method may further include the step of terminating the execution of the first operation in response to receiving a user selection of a second selectable element.

[0093] In some implementations, the step of identifying a second action that can be initialized by an automated assistant to perform an assistant command may include determining the amount of additional action to identify for suggestion to the user via the computing device's display interface. In some versions of those implementations, determining the amount of additional action to identify for suggestion to the user via the computing device's display interface may be based on the size of the computing device's display interface. In some additional or alternative versions of those implementations, determining the amount of additional action to identify for suggestion to the user via the computing device's display interface may be based on a corresponding metric that characterizes the corresponding degree to which each of the additional actions is expected to respond to an oral utterance.

[0094] Other implementations may include a non-temporary computer-readable storage medium that stores instructions that can be executed by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform one or more of the methods described above and / or elsewhere in this specification. Further other implementations may include a system of one or more computers, each containing one or more processors capable of operating to execute the stored instructions to perform one or more of the methods described above and / or elsewhere in this specification.

[0095] It should be understood that all combinations of the concepts described above and any additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of claims appearing at the end of this disclosure are considered to be part of the subject matter disclosed herein. [Explanation of symbols]

[0096] 100 Scenes 102 users 104 Computing Devices 106 Display Interfaces 108 Oral speech 120 Scenes 122 Graphical Elements 124 First proposed element, first selectable element 126 Second proposed element, second optional element 128 Indication 140 Scenes 142 Another oral utterance 160 Scenes 162 Graphical Elements 200 scenes 202 users 204 Computing Devices 206 Display Interfaces 220 Scenes 222 First selectable element 224 Second selectable element 240 Scenes 242 Third Selectable Element 246 Fourth selectable element 250 Input Gestures 300 Systems 302 Computing Devices 304 Automated Assistant 306 Input Processing Engine 308 Speech Processing Engine 310 Data Analysis Engines 312 Parameter Module 314 Output Generation Engine 316 Proposal generation engine 318 Proposed Feature Engine 320 Assistant Interface 322 Assistant Call Engine 324 Training Data Engine 326 Execution Engine 330 Application Data 332 Device Data 334 Applications 336 Context Data 338 Assistant Data 400 ways 420 method 510 Computer Systems 512 Bus Subsystem 514 processors 516 Network Interface Subsystem 520 User Interface Output Devices 522 User Interface Input Devices 524 Storage Subsystems 525 memory 526 File Storage Subsystem 530 Main Random Access Memory (RAM) 532 Read-only memory (ROM)

Claims

1. A method carried out by one or more processors, A computing device comprising the step of receiving a user's oral utterance directed to an automated assistant, wherein the oral utterance includes an assistant command to be performed by the automated assistant, The steps include generating metric data that, based on the oral utterance, identifies a first metric that characterizes the degree to which a first action is expected to satisfy the assistant command, and a second metric that characterizes another degree to which a second action is expected to satisfy the assistant command, A step of generating GUI data that characterizes an assistant graphical user interface (GUI) that responds to oral speech, based on the first and second operations, The GUI data is generated to identify the first selectable element and the second selectable element. The first selectable element is selectable to control the execution of the first operation, and the second selectable element is selectable to automatically initialize the execution of the second operation. While the first operation is being performed, by touch input on the display interface of the computing device, and / or A selectable step and, through additional oral utterances performed simultaneously with the execution of the first action, The steps include: in response to receiving the verbal utterance, causing the automated assistant to automatically initialize the execution of the first action; A method comprising the step of causing the display interface of the computing device to render the assistant GUI according to the GUI data and the metric data.

2. The method according to claim 1, wherein, when the degree to which the first action is expected to satisfy the assistant command is greater than the other degree to which the second action is expected to satisfy the assistant command, the first selectable element is arranged in the assistant GUI to be more prominent than the second selectable element.

3. The step of causing the display interface to render the assistant GUI according to the GUI data and the measurement reference data is: The first selectable element is positioned adjacent to the second selectable element, The method according to claim 1 or 2, wherein, when the degree to which the first action is expected to satisfy the assistant command is greater than the other degree to which the second action is expected to satisfy the assistant command, the assistant GUI is arranged such that the first area of ​​the first selectable element is larger than the second area of ​​the second selectable element.

4. The step of causing the display interface to render the assistant GUI according to the GUI data and the measurement reference data is: The first selectable element is positioned adjacent to the second selectable element, The method according to any one of claims 1 to 3, wherein the assistant GUI is configured such that the first selectable element is configured in accordance with first configuration data that is specific to the first selectable element and characterizes a corresponding first position of the first selectable element in the assistant GUI, and the second selectable element is configured in accordance with second configuration data that is specific to the second selectable element and characterizes a corresponding second position of the second selectable element in the assistant GUI.

5. The method according to any one of claims 1 to 4, wherein in the assistant GUI, the first selectable element is displayed based on corresponding first display data that is specific to the first selectable element and characterizes the corresponding first display characteristic of the first selectable element in the assistant GUI, and the second selectable element is displayed based on corresponding second display data that characterizes the corresponding second display characteristic of the second selectable element in the assistant GUI.

6. The first operation is a specific operation performed on the computing device, and the second operation is the same specific operation performed on another computing device. The method according to any one of claims 1 to 5, wherein the user's selection of the second selectable element causes the particular operation to be initialized on the other computing device.

7. The method according to any one of claims 1 to 6, wherein the prominence of the rendering of the first selectable element compared to the second selectable element in the assistant GUI is based on the difference between the first criterion and the second criterion.

8. When the difference between the first criterion and the second criterion does not meet the proposed threshold, the second selectable element is not rendered to the display interface. The method according to claim 7, wherein a specific touch gesture given by the user to the display interface of the computing device causes the second selectable element to be rendered on the display interface of the computing device.

9. At least one processor, A system including a memory that stores instructions causing at least one processor to perform an operation corresponding to any one of claims 1 to 8 when executed.

10. A computer-readable storage medium that stores instructions causing at least one processor to perform an operation corresponding to any one of claims 1 to 8 when executed.

Citation Information

Patent Citations

  • Mobile terminal and its menu control method

    JP2009252238A

  • Refinement of voice query interpretation

    US20200301657A1

  • Supplementing voice inputs to an automated assistant according to selected suggestions

    WO2020139408A1