Dynamic adaptation of graphical user interface elements by an automated assistant when a user repeatedly provides a voice utterance or a sequence of voice utterances

The automated assistant adapts GUI elements based on voice utterances to facilitate efficient completion of commands, reducing user input and computational waste by dynamically tailoring interface elements.

JP7804741B2Active Publication Date: 2026-01-22GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024189020
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-11-22
Filing Date
2024-10-28
Publication Date
2026-01-22
Estimated Expiration
2041-11-29

AI Technical Summary

Technical Problem

Automated assistants often fail to respond accurately to incomplete voice commands, requiring users to repeat themselves, wasting computational resources and extending interaction sessions.

Method used

An automated assistant dynamically adapts graphical user interface elements based on a user's voice utterances, using ASR and NLU models to render generic elements that evolve into tailored ones, allowing users to complete requests efficiently via touch input.

Benefits of technology

This approach reduces user input duration, preserves computational resources, and minimizes latency by enabling seamless adaptation of GUI elements, thus completing interactions quickly and accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007804741000001
    Figure 0007804741000001
  • Figure 0007804741000002
    Figure 0007804741000002
  • Figure 0007804741000003
    Figure 0007804741000003
Patent Text Reader

Abstract

To provide dynamic adaptation of graphical user interface elements by an automated assistant when iteratively providing a spoken utterance or a sequence of spoken utterances.SOLUTION: Implementations described herein relate to an automated assistant that iteratively renders various GUI elements as a user iteratively provides a spoken utterance, or a sequence of spoken utterances. In some implementations, a generic container graphical element associated with candidate intents can be initially rendered at a display interface and dynamically adapted with tailored container graphical elements as a particular intent is determined while the user iteratively provides the spoken utterance. In additional or alternative implementations, the tailored container graphical elements can include a current status of one or more settings associated with a computing device such that the user can view the current status while completing the spoken utterance.SELECTED DRAWING: Figure 2B
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans may engage in human-computer interactions with interactive software applications referred to herein as "automated assistants" (also referred to as "digital agents," "chatbots," "interactive personal assistants," "intelligent personal assistants," "assistant applications," "conversational agents," etc.). For example, humans (sometimes referred to as "users" when interacting with an automated assistant) may provide commands and / or requests to the automated assistant using vocal natural language input (i.e., utterances), which in some cases may be converted to text and then processed, and / or by providing textual (e.g., typed) natural language input.

[0002] In many cases, automated assistants may be invoked by users who may not have a complete command phrase in mind. For example, assume a user provides a voice utterance such as "Assistant, set ...," which includes a portion of a request that the automated assistant responds to. In this example, the portion of the request may indicate that the user intends to set the volume on a smart speaker, set the temperature on a smart thermostat, set the brightness level on a smart light bulb, and so on. However, in many such examples, if the user does not utter the complete command phrase within a certain time window, the automated assistant may not respond based on the request because it is too vague, or the automated assistant may respond based on the request and perform some action that the user did not intend. Thus, the user may need to re-invoke the automated assistant and again provide the voice utterance with the complete command phrase, thereby extending one or more interaction sessions between the user and the automated assistant and increasing the amount of user input received at the computing device.

[0003] In some cases, an automated assistant may be invoked by a user who has a complete command phrase in mind, but may not have a specific slot value associated with the command phrase in mind. For example, assume a user provides a spoken utterance of "Assistant, set the volume to ...," which includes a portion of a request that the automated assistant responds to. In this example, the request portion may indicate that the user intends to set the volume for the speaker as a specific slot value associated with the command to set the volume for the speaker. However, in many of these examples, the user may not know the current state of the speaker volume and, as a result, may not know how to modify the speaker volume for the current status. Thus, after providing the first portion of a voice utterance (e.g., "Assistant, set the volume to..."), the user may pause to think about how to modify the speaker volume. Similar to the above example, if the user does not pronounce a particular slot value within a certain time window, the automated assistant may not act on the request because it is too vague, or the automated assistant may act on the request and perform some action that the user did not intend. Again, the user may need to re-invoke the automated assistant and again provide the voice utterance with the complete command phrase and the particular slot value, thereby extending one or more interaction sessions between the user and the automated assistant and increasing the amount of user input received at the computing device. As a result, in such an example, processing the incomplete voice utterance may waste computational resources and require the user to re-engage in an interaction session with the automated assistant. Summary of the Invention [Means for solving the problem]

[0004] Implementations described herein relate to an automated assistant that can dynamically adapt graphical user interface (GUI) elements based on a user repeatedly providing a voice utterance or sequence of voice utterances that includes a request directed to an automated assistant that executes at least partially on the user's computing device. The GUI elements may characterize portions of an incoming request from the user and / or provide suggestions that may help the user more efficiently and accurately explain the request to the automated assistant. In some implementations, candidate intents may be determined based on processing portions of the request, and generic container graphical elements associated with the candidate intents may be rendered in a display interface of the computing device before the user completes the request. Furthermore, specific intents may be determined from the candidate intents based on processing additional portions of the request, and the generic container graphical elements may be dynamically adapted to specific, tailored container graphical elements associated with the specific intent without rendering a different display interface on the computing device. In additional or alternative implementations, specific words or phrases included in portions of the request may be directly mapped to generic container graphical elements without the need to determine candidate intents. In additional or alternative implementations, a particular adjusted container graphical element may include a current state of one or more settings of the computing device and / or additional computing devices communicating with the computing device (e.g., slot values ​​associated with the current state of the one or more settings) in response to determining that a portion of the request relates to modifying a current state of one or more settings of the computing device and / or additional computing devices.

[0005] For example, assume a user begins providing the voice utterance "Assistant, set ...," which includes a request portion asking the automated assistant to adjust a device state such as the volume for a smart speaker, the temperature for a smart thermostat, or the brightness level for a smart light bulb. As the user provides the request portion, the automated assistant may use a streaming automatic speech recognition (ASR) model to process the stream of audio data capturing the request portion and generate an ASR output. Additionally, the automated assistant may use a natural language understanding (NLU) model to process the ASR output and generate an NLU output. A generic container graphical element may be rendered in a display interface of a computing device based on the ASR output (e.g., indicating that the request portion includes "set" or another specific word or phrase) and / or the NLU output (e.g., indicating that the request portion includes a candidate intent associated with the generic container graphical element).

[0006] Further, assume that the user continues to provide the voice utterance "... the volume for the speakers ..." (or as an additional voice utterance following the voice utterance), which includes an additional portion of the request for the automated assistant to adjust the state of the device. Similarly, as the user provides the additional portion of the request, the automated assistant may use the streaming ASR model to process the stream of audio data that also captures the additional portion of the request and generate an additional ASR output. Furthermore, the automated assistant may use the NLU model to process the additional ASR output and generate an additional NLU output. Based on the additional ASR output and / or the additional NLU output, the automated assistant may determine that the user wants to set the volume for the smart speaker. Thus, the generic container graphical element may be dynamically adapted to an adjusted container graphical element specific to setting the volume for the smart speaker. For example, a tailored container graphical element specific to setting the volume for a smart speaker may include the current state of the volume for the smart speaker, a volume control graphical element that allows the user to set the volume using touch input, media content indicating that the volume is being set for the smart speaker, a device identifier associated with the smart speaker, and / or any other content associated with the smart speaker.

[0007] On the other hand, if the user indicates "... the temperature ..." while continuing to provide the voice utterance, the generic container graphical element may be dynamically adapted to a tailored container graphical element specific to setting a temperature for a smart thermostat that is separate from the tailored container graphical element specific to setting a volume for a smart speaker. For example, the tailored container graphical element specific to setting a temperature for a smart thermostat may include the current state of the temperature, media content indicating that the temperature is being set for the smart thermostat, a temperature control graphical element that allows the user to set the temperature using touch input, a device identifier associated with the smart thermostat, and / or any other content associated with the smart thermostat. Nevertheless, in both of these examples, the same generic container graphical element may be dynamically adapted to these various tailored container graphical elements without rendering any additional user interface.

[0008] The generic container graphical element may act as a placeholder for any one of a plurality of disparate tailored container graphical elements, each associated with a corresponding one of a plurality of disparate intents or directly mapped to a specific word or phrase. Thus, as the user continues to provide requests to the automated assistant, the generic container graphical element may dynamically and seamlessly adapt to a specific tailored graphical element associated with a specific intent determined based on processing additional portions of the request. For example, in the above example, in response to the portion of the request, "Assistant, set...," the generic container graphical element may initially be rendered in the display interface. The generic container graphical element may include, for example, an array of graphical elements (e.g., an array of dot shapes) to indicate a range of values. Subsequently, when the user provides additional portions of the request (e.g., "...speaker volume...," "...brightness...," "...temperature...," etc.), the automated assistant may adapt the arrangement of graphical elements based on the additional portions of the request. For example, based on the user providing an additional portion of the request of "...speaker volume...", the array of graphical elements may be adapted to reflect a range of values ​​associated with the volume of the smart speaker and to include a current state of the volume of the smart speaker to assist the user in determining how to modify the volume. Further, for example, based on the user providing an additional portion of the request of "...brightness...", the array of graphical elements may be adapted to reflect a range of values ​​associated with the brightness of the smart light bulb and to include a current state of the brightness of the smart light bulb to assist the user in determining how to modify the brightness.

[0009] In some implementations, the automated assistant may process the candidate intents to identify specific devices and / or applications that the user may be attempting to control when providing the request. When the automated assistant identifies a specific device and / or application, the automated assistant may cause an arrangement of graphical elements within the generic container graphical element to represent the current state of the specific application and / or device, resulting in an adjusted container graphical element. For example, the arrangement of graphical elements may include seven filled circles followed by three empty circles, thereby indicating that a smart light bulb associated with the specific device and / or application is currently at 70% of its maximum brightness level as the current state of the smart light bulb's brightness setting. Alternatively or additionally, the automated assistant may identify an icon representing the specific device and / or application to which the user is predicted to be referring (e.g., an icon representing a kitchen light). The automated assistant may include an icon for the adjusted container graphical element to identify the specific device and / or application that the automated assistant selected to control in response to the request from the user. In this way, the user may avoid providing another portion of the request via voice utterance to specify a particular application and / or device (e.g., to change brightness from 70% to 50%) and instead choose to utilize touch input, thereby preserving computational resources that would otherwise be consumed when processing the voice utterance or additional voice utterances.

[0010] In some implementations, a user may have an automated assistant control a particular device and / or application based on witnessing the current state indicated by the arrangement of graphical elements contained within the adjusted container graphical element by completing the request via one or more additional voice utterances. For example, by witnessing the arrangement of graphical elements contained within the adjusted container graphical element, the user may consider the final portion of the user's request. The user may provide a final voice utterance such as "... to 30%," thereby instructing the automated assistant to control the particular device and / or application to adjust the brightness level from 70% to 30%. Alternatively or additionally, the user may tap the portion of the arrangement of graphical elements corresponding to the "30% dot" to similarly cause the automated assistant to adjust the brightness level from 70% to 30%.

[0011] In some implementations, in response to a first portion of a request being provided by a user, the automated assistant may cause multiple adjusted container elements to be rendered. For example, when a user provides the first portion of a request via a voice utterance such as, "Assistant, play [song title 1] by ...," the automated assistant may cause multiple different adjusted container elements to be rendered in a display interface of a computing device. Each of the adjusted container graphical elements may correspond to a different action and / or interpretation that may be associated with the request. For example, a first adjusted container graphical element may correspond to an action to play "[song title 1]" by "[artist 1]" on the computing device or the additional computing device, and a second adjusted container graphical element may correspond to another action to play "[song title 1]" by "[artist 2]" on the computing device or the additional computing device. In some implementations, additionally or alternatively, each of the adjusted container graphical elements includes the current state of the computing device or the additional computing device (e.g., what is currently playing on the first device and / or the second device). The user may complete the request (e.g., via voice utterances and / or touch input), and the automated assistant may have the request fulfilled accordingly.

[0012] In some implementations, the ASR output generated using the streaming ASR model may include, for example, predicted speech hypotheses predicted to correspond to various portions of the request, predicted phonemes predicted to correspond to various portions of the request, predicted ASR measures indicating how likely the predicted speech hypotheses and / or predicted phonemes correspond to various portions of the request, and / or other ASR outputs. Furthermore, the NLU output generated using the NLU model may include, for example, candidate intents predicted to correspond to the user's actual intent in providing the various portions of the request, one or more slot values ​​for corresponding parameters associated with the candidate intents, and / or other NLU outputs. Furthermore, one or more structured requests may be generated based on the NLU output and processed by various devices and / or applications to generate fulfillment data for the requests. The fulfillment data, when implemented, may cause an automated assistant to fulfill a request provided by a user.

[0013] In some implementations, the generic container graphical element and / or the adjusted container graphical element described herein may only be rendered in a display interface of a computing device in response to determining that a user has paused from providing a request. The automated assistant may determine that a user has paused from providing a request, for example, based on NLU data and / or audio-based characteristics associated with a portion of the request received at the computing device. The audio-based characteristics associated with a portion of the request may include one or more of intonation, tone, stress, rhythm, tempo, pitch, and prolonged syllables. For example, assume that a user provides a request, "Assistant, set the volume...," contained within a voice utterance directed to the automated assistant. In this example, the automated assistant may determine that the user has paused based, for example, on a threshold duration elapsed since the aforementioned "to" and an NLU output indicating that the user has not provided a slot value for a volume parameter associated with a predicted intent to change the volume of the smart speaker. In response, the automated assistant may cause a volume container graphical element for the smart speaker to be rendered in a display interface of the computing device. A volume container graphical element for the smart speaker rendered in the display interface of the computing device may include the current volume of the smart speaker to assist the user in determining how to modify the volume relative to the current volume. Alternatively or additionally, assume further that the user includes a drawn-out syllable when providing the "to" (e.g., "Assistant, set the volume toooooo ..."). In this example, the automated assistant may determine that the user has paused based on audio-based characteristics that reflect uncertainty about how to modify the volume relative to the current volume, for example, based at least on the drawn-out syllable when providing the request.Thus, the volume container graphical element may assist the user in determining how to modify the volume relative to the current volume.

[0014] By using the techniques described herein, various technical advantages may be achieved. As a non-limiting example, the techniques described herein enable an automated assistant to dynamically adapt various GUI elements from a generic GUI element into an adjusted GUI element while a user provides a voice utterance or a sequence of voice utterances. For example, a user may provide a voice utterance including a portion of a request, and the automated assistant may render a generic GUI element, which is then adapted to an adjusted GUI element based on processing an additional portion of the voice utterance or the additional voice utterance. Such an adjusted GUI element may assist the user in completing the request, thereby more quickly and efficiently completing an interaction session between the user and the automated assistant and reducing the amount of user input received at the computing device. Furthermore, cases in which an automated assistant fails because a user does not complete a request within a certain time window may be mitigated. As a result, computational resources at the computing device may be preserved and latency in fulfilling requests may be reduced.

[0015] The above description is provided as a summary of some implementations of the present disclosure. More detailed descriptions of those implementations, as well as other implementations, are described in more detail below. [Brief explanation of the drawings]

[0016] [Figure 1A] 1A-1C illustrate a user iteratively providing example requests to an automated assistant, and the automated assistant iteratively rendering graphical elements corresponding to portions of the example requests, according to various implementations. [Figure 1B] 1A-1C illustrate a user iteratively providing example requests to an automated assistant, and the automated assistant iteratively rendering graphical elements corresponding to portions of the example requests, according to various implementations. [Figure 1C] 1A-1C illustrate a user iteratively providing example requests to an automated assistant, and the automated assistant iteratively rendering graphical elements corresponding to portions of the example requests, according to various implementations. [Figure 2A] 10A-10C illustrate a user iteratively providing additional example requests to an automated assistant, and the automated assistant iteratively rendering graphical elements corresponding to portions of the additional example requests, according to various implementations. [Figure 2B] 10A-10C illustrate a user iteratively providing additional example requests to an automated assistant, and the automated assistant iteratively rendering graphical elements corresponding to portions of the additional example requests, according to various implementations. [Figure 2C] 10A-10C illustrate a user iteratively providing additional example requests to an automated assistant, and the automated assistant iteratively rendering graphical elements corresponding to portions of the additional example requests, according to various implementations. [Figure 3] 1A-1C illustrate systems that provide an automated assistant that iteratively builds graphical elements as a user provides requests directed to the automated assistant, according to various implementations. [Figure 4A] 1A-1C illustrate methods for controlling an automated assistant to repeatedly present graphical elements in a display interface when a user provides requests directed to the automated assistant, according to various implementations. [Figure 4B] 1A-1C illustrate methods for controlling an automated assistant to repeatedly present graphical elements in a display interface when a user provides requests directed to the automated assistant, according to various implementations. [Figure 5] FIG. 1 is a block diagram of an exemplary computer system according to various implementations. DETAILED DESCRIPTION OF THE INVENTION

[0017] 1A, 1B, and 1C show views 100, 120, and 140, respectively, in which a user 102 repeatedly provides an example request to an automated assistant and the automated assistant repeatedly renders graphical elements corresponding to portions of the example request. The request may be included in a voice utterance directed to the automated assistant or a sequence of voice utterances directed to the automated assistant. For example, in the example of FIG. 1A, assume that user 102 provides the voice utterance “Assistant, set…” which includes a first portion of request 108 directed to the automated assistant accessible via a computing device 104 in the user’s home. User 102 may provide the voice utterance to prompt the automated assistant to perform a specific assistant action and / or execute a specific intent associated with the first portion of request 108. The automated assistant may use a streaming automatic speech recognition (ASR) model to process a stream of audio data capturing the first portion of request 108 to generate an ASR output. Additionally, the automated assistant may process the ASR output using a natural language understanding (NLU) model to generate the NLU output. In particular, these operations may be performed while the user 102 is providing a voice utterance that includes the first portion of the request 108.

[0018] In some implementations, the automated assistant may, based on processing the first portion of the request 108, determine a candidate intent associated with the first portion of the request 108 based on the NLU output. Further, the automated assistant may, based on the candidate intent, determine a generic container graphical element 110 to be rendered in the display interface 106 of the computing device 104. For example, the generic container graphical element 110 may be selected to represent a graphical user interface (GUI) element that can be dynamically adapted to a particular tailored container graphical element associated with a particular intent selected from the candidate intents based on processing additional portions of the request (e.g., as shown in FIGS. 1A and 1B). In additional or alternative implementations, the automated assistant may determine, based on the ASR output (and without considering the NLU output), that the first portion of the request indicates that it includes a particular word or phrase (e.g., "set"). For example, when a particular word or phrase is detected in the ASR output, the particular word or phrase may be mapped to a generic container graphical element 110 in on-device memory of the computing device 104 such that the generic container graphical element 110 may be rendered on the display interface 106 of the computing device 104. The generic container graphical element 110 may similarly be dynamically adapted to a particular tailored container graphical element associated with a particular intent based on processing additional portions of the request (e.g., as shown in FIGS. 1B and 1C ).In other words, because the first portion of the request 108 is predicted to correspond to a request to control an application and / or device involving "settings" (e.g., based on the first portion of the request 108 including the words "set," "change," and / or other specific words or phrases), which may be represented by an array of graphical elements (e.g., an array of empty or filled circles as shown in FIG. 1A), the generic container graphical element 110 may be rendered in the display interface 106 of the computing device 104.

[0019] In some implementations, in response to determining that the user 102 has paused providing a speech utterance including a first portion of the request 108, the generic container graphical element 110 may be rendered on the display interface 106 of the computing device 104. The automated assistant may determine that the user 102 has paused providing a speech utterance including a first portion of the request 108 based on, for example, NLU output generated when processing the speech utterance, audio-based characteristics determined based on processing the speech utterance, and / or a threshold duration that has elapsed since the user 102 provided the first portion of the request 108. For example, the automated assistant may determine that the user 102 has paused providing a speech utterance (e.g., the user 102 said “set it” but failed to provide any indication of what to “set it”) based on a threshold duration that has elapsed since the user 102 provided the first portion of the request 108 and based on an NLU output indicating that a slot value for a corresponding parameter associated with a candidate intent is unknown. In some versions of such implementations, the automated assistant may also consider one or more terms or phrases surrounding the predicted pause (e.g., whether the pause occurred after a preposition or a disfluent speech (e.g., uhmmm, uhhh, etc.)). Alternatively or additionally, the audio-based characteristics may indicate that the manner in which the user 102 provided the first part of the request 108 indicates that the user 102 paused to consider how to phrase the natural language provided to the automated assistant to complete the request.

[0020] To facilitate providing a complete request to the automated assistant, assume further that user 102 continues to provide the voice utterance or provides an additional voice utterance that includes a second portion of request 122 by providing "...lights..." as shown in overview 120 of FIG. 1B. In this example, the stream of audio data capturing the second portion of request 122 may be processed to generate an additional ASR output and an additional NLU output. Based on the additional NLU output, the automated assistant may select a specific intent from among the candidate intents described with respect to FIG. 1A that indicates to the automated assistant that user 102 wishes to modify the current status of the smart light bulbs in user 102's home. Thus, in response to receiving the second portion of request 122, the automated assistant may dynamically adapt generic container graphical element 110 of FIG. 1A to content associated with the specific intent determined based on processing the second portion of request 122, resulting in a specific, tailored graphical container element 112 in a seamless manner.

[0021] The particular adjusted graphical container element 112 may be one of multiple heterogeneous adjusted graphical container elements to which the generic graphical container element 110 may be dynamically adapted, and may be specific to a particular intent selected based on processing the second portion of the request 122. In other words, the particular adjusted graphical container element 112 shown in FIG. 1B may be specific to a particular intent associated with modifying the current state of the smart light bulb, and other intents may be associated with adjusted graphical container elements (e.g., modifying a temperature intent, modifying a speaker volume intent, etc.) that are different from the particular adjusted graphical container element 112 shown in FIG. 1B. The particular tailored graphical container element 112 may include, for example, media content 124 indicating a device and / or application that the automated assistant may associate with the second portion of the request 122 (e.g., the light bulb icon shown in FIG. 1B), a current state 120 indicating the current state of the device and / or application that the automated assistant may associate with the second portion of the request 122 and that is dynamically adapted from the arrangement of graphical elements shown in association with the generic container graphical element 110 of FIG. 1A (e.g., the seven filled circles shown in FIG. 1B to indicate that the smart light bulb is at 70% brightness), one or more control elements 144 that, when selected via touch input by the user 102, cause the current state 120 of the device and / or application to be controlled (e.g., as described in association with FIG. 1C), device identifiers associated with the various smart light bulbs to be controlled, and / or other content. In particular, when the user 102 provides a request, the generic container graphical element 110 is dynamically adapted to the specific tailored container graphical element 112, so that the specific tailored container graphical element 112 can be rendered on the display interface 106 of the computing device 104 before the user 102 completes the request and / or before the automated assistant completes fulfillment of the request.

[0022] In contrast to the example shown in FIG. 1B , assume that the user 102 continues to provide the voice utterance or provides an additional voice utterance that includes a second portion of the request 122, such as “…speaker volume….” In this example, the second portion of the request 122 may be processed in the same or similar manner as described above. However, the resulting particular adjusted container graphical element will differ from the particular adjusted container graphical element 112 shown in FIG. 1B . For example, the particular adjusted graphical container element in this contrasting example may include, for example, media content (e.g., a speaker icon) indicating a device and / or application that the automated assistant may associate with the second portion of the request 122, a current state indicating the current state of the device and / or application that the automated assistant may associate with the second portion of the request 122 and that is dynamically adapted from the arrangement of graphical elements shown in association with the generic container graphical element 110 of FIG. 1A (e.g., a filled circle may correspond to a speaker volume level), one or more control elements that, when selected via touch input by the user 102, cause the current state of the device and / or application to be controlled, and / or other content. Thus, the same generic container graphical element 110 shown in FIG. 1B may be adapted differently based on the second part of the request 122 .

[0023] Further assume that user 102 provides another additional voice utterance including a third portion of request 142 by providing "... to 30 percent," as shown in overview 140 of FIG. 1C , to complete the voice utterance, provide the additional voice utterance, or facilitate providing a complete request to the automated assistant. In this example, a stream of audio data capturing the third portion of request 142 may be processed to generate another additional ASR output and another additional NLU output. Based on the another additional NLU output, the automated assistant may generate a structured request to be utilized in fulfilling the voice utterance of user 102. For example, in response to receiving the third portion of request 142 from user 102, the automated assistant may send one or more structured requests to a smart light bulb and / or an application associated with the smart light bulb accessible at computing device 104 to modify the current state 122 of the smart light bulb from 70% brightness to 30% brightness. Further, in response to receiving the third part of the request 142, the automated assistant may reflect the change in brightness of the smart light bulb by adapting the current state 122 to reflect that the smart light bulb has changed from 70% brightness to 30% brightness (e.g., the seven filled circles shown in FIG. 1B to indicate that the smart light bulb is at 70% brightness changes to three filled circles in FIG. 1C to indicate that the smart light bulb is now at 30%).

[0024] 1C , the user 102 may direct a touch input to a given one of the one or more control graphical elements 144 (e.g., the third circle). Similarly, in response to the user 102 directing a touch input to a given one of the one or more control graphical elements 144, the automated assistant may generate a structured request to be utilized in fulfilling the user's 102's voice utterance to modify the current state 122 of the smart light bulb from 70% brightness to 30% brightness (assuming the user 102 directs the touch input to the third circle) and reflect the change in brightness of the smart light bulb by adapting the current state 122 to reflect that the smart light bulb has changed from 70% brightness to 30% brightness, as described above.

[0025] 1A, 1B, and 1C are described in connection with a request to change the brightness of a smart light bulb in various portions (e.g., first portion 108, second portion 122, and third portion 142), it should be understood that this is for purposes of illustration and not limitation. For example, even if user 102 does not pause while providing the request, the automated assistant may dynamically adapt GUI elements as described above to indicate to user 102 that the automated assistant is listening to user 102 and fulfilling the request, even if user 102 does not pause between providing various portions of the request. As a result, instances in which user 102 repeats one or more portions of the request may be mitigated because user 102 knows that the automated assistant is fulfilling the request based on the dynamic GUI elements described herein. However, if user 102 pauses in providing the request, the dynamic GUI elements described herein may assist user 102 in completing the request by at least providing the current state 122 of the smart light bulb to present to user 102 when the request is provided. As a result, an interactive session between the user 102 and the automated assistant may be concluded more quickly and efficiently. Additionally, while the examples of Figures 1A, 1B, and 1C are described in connection with a request to change the brightness of a smart light bulb, it should be understood that this is for purposes of illustration and not limitation. Rather, it should be understood that the techniques described herein may be utilized to provide dynamic GUI elements for any request directed to an automated assistant.

[0026] 2A, 2B, and 2C show views 200, 220, and 240, respectively, in which a user 202 repeatedly provides additional exemplary requests to an automated assistant and the automated assistant repeatedly renders graphical elements corresponding to portions of the additional exemplary requests. The requests can be included within a voice utterance directed to the automated assistant or a sequence of voice utterances directed to the automated assistant. For example, in the example of FIG. 2A, assume that user 202 provides the voice utterance "Assistant, set ...," which includes a first portion of request 208 directed to the automated assistant that is accessible via computing device 204 in the user's home. User 202 can provide the voice utterance to prompt the automated assistant to perform a specific assistant action and / or execute a specific intent associated with the first portion of request 208. 1A , the automated assistant may use a streaming automatic speech recognition (ASR) model to process the stream of audio data capturing the first portion of the request 208 to generate an ASR output, and may use a natural language understanding (NLU) model to process the ASR output to generate an NLU output. In response to receiving the first portion of the request 208, the automated assistant may cause the display interface 206 of the computing device 204 to render the generic container graphical element 210 in the same or similar manner as described above in connection with FIG. 1A .

[0027] Again, to facilitate providing a complete request to the automated assistant, assume further that user 102 continues to provide the voice utterance or provides an additional voice utterance that includes the second portion of request 226 by providing "...lights..." as shown in overview 220 of FIG. 2B. Similarly, the stream of audio data capturing the second portion of request 226 may be processed to generate an additional ASR output and an additional NLU output. Based on the additional NLU output, the automated assistant may select a specific intent that indicates to the automated assistant that user 202 wishes to modify the current state of the smart light bulbs in user 202's home. Thus, in response to receiving the second portion of request 226. However, in contrast to the example of FIG. 1A, assume that user 202 has multiple groups of smart light bulbs grouped together throughout user 202's home (e.g., a "kitchen" group of smart light bulbs, a "basement" group of smart light bulbs, and a "hallway" group of smart light bulbs). Thus, in the example of FIG. 2A, the automated assistant may dynamically adapt multiple instances of generic container graphical element 210 of FIG. 2A to content associated with a specific intent determined based on processing the second part of request 226 and based on multiple groups of smart light bulbs, resulting in multiple specific tailored graphical container elements 222A, 222B, and 222C.

[0028] For example, a first adjusted graphical container element 222A may be associated with a "kitchen" group of smart light bulbs and include a device identifier 224 of "kitchen" and other content (e.g., the current state of the "kitchen" lights at 50% brightness, indicated by five filled circles, control graphical elements associated with the "kitchen" group of smart light bulbs, media content, and / or other content), and a second adjusted graphical container element 222B may be associated with a "basement" group of smart light bulbs and include a device identifier 228 of "basement" and other content (e.g., nine filled circles). a current state of the "Basement" lights at 90% brightness, indicated by the seven filled circles, control graphical elements, media content, and / or other content associated with the "Basement" group of smart light bulbs, and a third adjusted graphical container element 222C associated with the "Hallway" group of smart light bulbs and including a device identifier 230 of "Hallway" and other content (e.g., a current state of the "Hallway" lights at 70% brightness, indicated by the seven filled circles, control graphical elements, media content, and / or other content associated with the "Hallway" group of smart light bulbs). In some implementations, the automated assistant may retrieve the current state of each of the group of smart light bulbs from an application associated with the smart light bulbs, while in other implementations, the automated assistant may retrieve the current state of each of the group of smart light bulbs directly from the smart light bulbs. Thus, in the example time that user 202 provides the second part of request 226, the automated assistant may dynamically adapt multiple instances of generic container graphical element 210, resulting in multiple distinct, tailored container graphical elements that dynamically provide instructions to user 202 on how to control the smart light bulbs based on the current status of each of the group of smart light bulbs.

[0029] In some implementations, when the user 202 provides a request, natural language content 232 corresponding to the request from the user 202 may be rendered in the display interface 206. For example, the natural language content 232 may include a streaming transcription (e.g., “Assistant, set the lights …,” as shown in the display interface 206 of the computing device 204 in FIG. 2B ) determined based on an ASR output generated using a streaming ASR model in processing a stream of audio data capturing the first portion of the request 208 and the second portion of the request 226. In particular, the natural language content may be rendered simultaneously with the plurality of coordinated container graphical elements 222A, 222B, and 222C being rendered in the display interface 206 of the computing device 204. In these and other ways, the automated assistant may invoke various graphical elements (e.g., coordinated container graphical elements 222A, 222B, and 222C, natural language content 232, and / or other graphical elements) in an attempt to assist the user 202 in completing the request quickly and efficiently, thereby reducing the duration of the interaction session between the user 202 and the automated assistant.

[0030] To facilitate providing a complete request to the automated assistant, further assume that user 202 completes the request by providing a third portion of request 242, "... in the basement to 30 percent," as shown in overview 240 of FIG. 2C . In this example, a stream of audio data capturing the third portion of request 242 may be processed to generate an ASR output and an NLU output. Based on the NLU output, the automated assistant may generate a structured request to be utilized in fulfilling the voice utterance of user 202. For example, in response to receiving the third portion of request 242 from user 202, the automated assistant may send one or more structured requests to smart light bulbs in the "Basement" group and / or applications associated with smart light bulbs accessible at computing device 204 to modify the current state of the smart light bulbs in the "Basement" group from 90% brightness to 30% brightness. Further, in response to receiving the third part of request 242, the automated assistant may reflect the change in brightness of the smart light bulbs by adapting the current state to reflect that the smart light bulbs have changed from 90% brightness to 30% brightness (e.g., the nine filled circles shown in FIG. 2B to indicate that the smart light bulbs in the “Basement” group are at 90% brightness change to three filled circles in FIG. 2C to indicate that the smart light bulbs in the “Basement” group are now at 30%).

[0031] In some implementations, as shown in FIG. 2C , when the ASR output indicates that the user selected a particular group of smart light bulbs (e.g., the “basement” group in the example of FIG. 2C ), the automated assistant may remove other adjusted container graphical elements 222A and 222C from display interface 206 of computing device 204 such that adjusted container graphical element 222B, associated with the particular intent of user 202, is the only remaining adjusted container graphical element. In some implementations, the ASR and / or NLU processes performed via the automated assistant may be biased according to the content of the container graphical elements. For example, in response to receiving the third part of request 242, the automated assistant may bias ASR processing and / or ASR output toward “kitchen,” “basement,” and “hallway.”

[0032] 3 illustrates a system 300 that provides an automated assistant 304 that iteratively builds graphical elements when a user provides a request directed to the automated assistant 304. The automated assistant 304 may operate as part of an assistant application provided on one or more computing devices, such as a computing device 302 (e.g., an example of computing device 104, an example of computing device 204, etc.) and / or a server device. A user may interact with the automated assistant 304 through an assistant interface 320, which may be a microphone, a camera, a touchscreen display, a user interface, and / or any other device capable of providing an interface between a user and an application. For example, a user may initialize the automated assistant 304 by providing verbal, textual, and / or graphical input to the assistant interface 320 to cause the automated assistant 304 to initiate one or more actions (e.g., provide data, control peripheral devices, access agents, generate input and / or output, etc.). Alternatively, the automated assistant 304 may be initialized based on processing of contextual data 336 using one or more trained machine learning models. Context data 336 may characterize one or more features of an environment to which automated assistant 304 has access and / or one or more characteristics of a user predicted to intend to interact with automated assistant 304. Computing device 302 may include a display device, which may be a display panel including a touch interface (e.g., display interface 106 of computing device 104, display interface 206 of computing device 204) for receiving touch inputs and / or gestures to enable a user to control application 334 of computing device 302 via the touch interface.In some implementations, the computing device 302 may lack a display device, thereby providing an audible user interface output without providing a graphical user interface output. Additionally, the computing device 302 may provide a user interface, such as a microphone, for receiving speech natural language input from a user. In some implementations, the computing device 302 may include a touch interface and may lack a camera, but may optionally include one or more other sensors.

[0033] The computing device 302 and / or other third-party client devices may be in communication with a server device over a network, such as the Internet. Additionally, the computing device 302 and any other computing devices may be in communication with each other over a local area network (LAN), such as a Wi-Fi network. The computing device 302 may offload computational tasks to the server device to conserve computational resources at the computing device 302. For example, the server device may host the automated assistant 304, and / or the computing device 302 may send inputs received at one or more assistant interfaces 320 to the server device. However, in some implementations, the automated assistant 304 may be hosted locally at the computing device 302, and various processes related to automated assistant operation may be performed exclusively at the computing device 302.

[0034] In various implementations, all or less than all aspects of the automated assistant 304 may be implemented on the computing device 302. In some such implementations, aspects of the automated assistant 304 are implemented via the computing device 302 and may interface with a server device, which may implement other aspects of the automated assistant 304. Optionally, the server device may serve multiple users and their associated assistant applications via multiple threads. In implementations in which all or less than all aspects of the automated assistant 304 are implemented via the computing device 302, the automated assistant 304 may be an application separate from (e.g., installed “on”) the operating system of the computing device 302, or alternatively, may be implemented directly by (e.g., considered an operating system application but integrated with) the operating system of the computing device 302.

[0035] In some implementations, the automated assistant 304 may include an input processing engine 306, which may utilize multiple different modules for processing input and / or output for the computing device 302 and / or a server device. For example, the input processing engine 306 may include a speech processing engine 308 that utilizes a streaming ASR model that may process a stream of audio data received at the assistant interface 320 to generate ASR output, such as text embodied in the stream of audio data. Further, for example, the input processing engine 306 may use an audio-based machine learning model and / or a heuristics-based approach to determine audio-based characteristics associated with any voice utterances / requests captured within the stream of audio data. In some implementations, the stream of audio data may be transmitted, for example, from the computing device 302 to a server device to conserve computational resources at the computing device 302. Additionally or alternatively, the stream of audio data may be processed exclusively at the computing device 302.

[0036] The process for converting audio data to text may include a speech recognition algorithm that may utilize a neural network and / or a statistical model (e.g., the streaming ASR model described herein) for identifying groups of audio data that correspond to words or phrases. The text converted from the audio data may be parsed by a data parsing engine 310 that utilizes an NLU model and made available to the automated assistant 304 as text data that can be used to generate NLU output, such as a command phrase, an intent, an action, a slot value, and / or any other content specified by the user. In some implementations, the output data provided by the data parsing engine 310 may be provided to a parameter engine 312 that may determine whether the user has provided input that corresponds to a particular intent, action, and / or routine that can be performed by the automated assistant 304 and / or an application or agent that can be accessed via the automated assistant 304. For example, assistant data 338 may be stored on the server device and / or computing device 302 and may include data defining one or more actions that can be performed by the automated assistant 304, as well as parameters necessary to perform the action. The parameter engine 312 may generate one or more parameters for the intent, action, and / or slot value and provide the one or more parameters to the output generation engine 314. The output generation engine 314 may use the one or more parameters to communicate with the assistant interface 320 to provide output (e.g., visual output and / or audible output) to the user and / or to communicate with one or more applications 334 to provide output to the one or more applications 334.

[0037] In some implementations, as mentioned above, the automated assistant 304 may be an application that may be installed “on” the operating system of the computing device 302 and / or may itself form part of (or the entire) of the operating system of the computing device 302. The automated assistant application includes and / or has access to on-device ASR, on-device NLU, and on-device fulfillment. For example, the on-device ASR may be implemented using a streaming ASR model that processes a stream of audio data (detected by a microphone) using an end-to-end streaming ASR model stored locally on the computing device 302. The on-device speech recognition generates an ASR output, such as recognized text, for a voice utterance present (if any) in the stream of audio data. Furthermore, on-device NLU may be implemented to generate an NLU output, for example, using an on-device NLU model that processes the ASR output generated using the streaming ASR model and, optionally, contextual data. The NLU output may include candidate intents corresponding to the voice utterance and, optionally, slot values ​​for corresponding parameters associated with the candidate intent.

[0038] Using on-device fulfillment models or fulfillment rules that utilize the NLU output and, optionally, other local data, on-device fulfillment can be performed to determine a structured request for determining actions to take to resolve the intent of the voice utterance (and, optionally, slot values ​​for corresponding parameters associated with the candidate intent). This can include determining local and / or remote responses (e.g., answers) to the voice utterance and / or request, interactions with locally installed applications to perform based on the voice utterance and / or request, commands to send to Internet of Things (IoT) devices (directly or via corresponding remote systems) based on the voice utterance and / or request, and / or other resolution actions to perform based on the voice utterance and / or request. The on-device fulfillment then initiates local and / or remote implementation / execution of the determined actions to resolve the voice utterance and / or request.

[0039] In various implementations, remote speech processing, remote NLU, and / or remote fulfillment may be utilized at least selectively to conserve computational resources at the computing device 302. For example, recognized text may be at least selectively sent to a remote automated assistant component remote NLU and / or remote fulfillment. For example, recognized text may optionally be sent for remote fulfillment in parallel with on-device fulfillment or in response to a failure of on-device NLU and / or on-device fulfillment. However, on-device speech processing, on-device NLU, on-device fulfillment, and / or on-device execution may be prioritized at least for the reduced latency achieved when resolving a speech utterance (because a client-server round trip is not required to resolve the speech utterance). Furthermore, on-device functionality may be the only functionality available in situations where there is no network connectivity or where network connectivity is limited.

[0040] In some implementations, the computing device 302 may have access to one or more applications 334, which may be provided by a third-party entity different from the entity that provided the computing device 302 and / or automated assistant 304, or by a first-party entity that is the same entity that provided the computing device 302 and / or automated assistant 304. The automated assistant 304 and / or computing device 302 may access application data 330 to determine one or more actions that may be performed by the one or more applications 334, as well as the state of each application of the one or more applications 334 and / or the state of each device associated with the computing device 302. Additionally, the automated assistant 304 and / or computing device 302 may access device data 332 to determine one or more actions that may be performed by the computing device 302 and / or one or more devices associated with the computing device 302. Additionally, the application data 330 and / or any other data (e.g., device data 332) may be accessed by the automated assistant 304 to generate context data 336 that may characterize the context in which a particular application 334 and / or device is running and / or the context in which a particular user is accessing the computing device 302, the application 334, and / or any other device or module.

[0041] While one or more applications 334 are executing on the computing device 302, the device data 332 may characterize the current operational state of each application 334 executing on the computing device 302. Additionally, the application data 330 may characterize one or more features of the executing applications 334, such as the content of one or more graphical user interfaces being rendered at the direction of the one or more applications 334. Alternatively or additionally, the application data 330 may characterize action schemas, which may be updated by the respective applications and / or the automated assistant 304 based on the respective applications' current operational status. Alternatively or additionally, the one or more action schemas for one or more applications 334 may remain static but may be accessed by the application state engine to determine appropriate actions to initialize via the automated assistant 304.

[0042] Computing device 302 may further include an assistant invocation engine 322 that may use one or more trained machine learning models to process application data 330, device data 332, context data 336, and / or any other data accessible to computing device 302. Assistant invocation engine 322 may process this data to determine whether to wait for the user to explicitly speak an invocation phrase to invoke automated assistant 304, or to consider the data to indicate an intent by the user to invoke the automated assistant instead of requiring the user to explicitly speak an invocation phrase. For example, the one or more trained machine learning models may be trained using example training data based on a scenario in which a user is in an environment in which multiple devices and / or applications exhibit various operating states.

[0043] Examples of training data may be generated to capture training data characterizing contexts in which a user invokes an automated assistant and other contexts in which a user does not invoke an automated assistant. When one or more trained machine learning models are trained according to these examples of training data, the assistant invocation engine 322 may cause the automated assistant 304 to detect or limit detection of a voice invocation phrase from the user based on contextual and / or environmental features. Additionally or alternatively, the assistant invocation engine 322 may cause the automated assistant 304 to detect or limit detection of one or more assistant commands from the user based on contextual and / or environmental features. In some implementations, the assistant invocation engine 322 may be disabled or limited based on the computing device 302 detecting an assistant suppression output from another computing device. In this manner, when the computing device 302 detects an assistant suppression output, the automated assistant 304 will not be invoked based on the contextual data 336; otherwise, the automated assistant 304 would be invoked if an assistant suppression output was not detected.

[0044] In some implementations, the system 300 may include a candidate intent engine 316 for determining one or more candidate intents that may be associated with one or more portions of a request provided by a user to the automated assistant 304. For example, when a user provides a voice utterance such as "Assistant, play ...," the candidate intent engine 316 may identify one or more intents that may be associated with the voice utterance based on the aforementioned NLU output. Alternatively or additionally, the candidate intent engine 316 may eliminate some candidate intents that may not be associated with the voice utterance provided by the user.

[0045] In some implementations, the system 300 may include a generic container engine 318 that may retrieve and / or generate one or more generic container graphical elements based on one or more candidate intents identified by the candidate intent engine 316. The generic container graphical element may be retrieved or generated (e.g., from on-device memory of the computing device 302) in response to receiving a voice utterance including a request that the automated assistant 304 is to fulfill. The generic container graphical element may be assigned other elements and / or features associated with one or more requests that the user is predicted to be providing via the voice utterance. For example, an initial voice utterance such as "Assistant, play..." does not identify a particular slot value (e.g., song, artist, movie, streaming service, etc.), but the generic container engine 318 may identify a type of slot value predicted to be associated with the request. Based on this type of slot value, the generic container engine 318 may retrieve and / or generate a generic container graphical element that may be dynamically adapted based on the portion of the request yet to be received at the computing device 302. For example, an initial voice utterance including the term "play" may be associated with a type of slot value for controlling media playback. Thus, a generic container graphical element associated with the "media playback" feature may be selected by the generic container engine 318 and then dynamically adapted to the feature associated with media playback. Alternatively, an initial voice utterance including the term "turn" may be associated with a slot value type for controlling an application and / or device output level. Thus, a generic container graphical element associated with adjusting an application and / or device setting may be selected by the generic container engine 318 and then dynamically adapted to the feature associated with an application and / or device output level control setting.

[0046] In some implementations, the system 300 may include an adjusted container engine 326, which may obtain and / or generate control elements for assigning dynamically adapting generic container graphical elements, resulting in adjusted container graphical elements. For example, as the user provides additional portions of a request to the automated assistant 304, the adjusted container engine 326 may iteratively assign and / or remove control elements in the generic container graphical element. Each control element may be associated with a slot value and / or a type of slot value determined to correspond to the request the user is predicted to be providing to the automated assistant 304. For example, when the automated assistant 304 predicts that the user is requesting to modify a device setting that may have a range of numeric values, the adjusted container engine 326 may select a “sliding” GUI element (or any other element suitable for controlling a device setting) to assign to the container graphical element. Alternatively, when the user is predicted to provide the automated assistant 304 with a request to control playback of media content (e.g., audio and / or video), the tuned container engine 326 may select to assign one or more media playback control elements (e.g., a pause button, a play button, a skip button, etc.) to the container graphical element.

[0047] In some implementations, when the user provides additional portions of the request to the automated assistant 304, the system 300 may include a state engine 324 that may retrieve and / or remove the current state of the application and / or device for the adjusted container graphical element. Additionally, the state engine may determine an updated state for the application and / or device based on processing voice utterances and / or touch inputs associated with the request. For example, when the automated assistant 304 predicts that the user is requesting to modify settings on another computing device, the state engine 326 may determine the current state of the settings on the other computing device. Based on this current state, the state engine 326 may identify a state GUI element that may characterize the status of the device and incorporate the state GUI element into the adjusted container graphical element. For example, when the state of a setting corresponds to a value within a numerical range, the state engine 326 may generate a state GUI element that characterizes the numerical range and highlights the current state of the setting. This state GUI element may be incorporated into the adjusted container graphical element currently rendered in the interface of the computing device 302. In this way, as the user continues to provide additional portions of the request, the user can be informed of the current state of the settings, thereby refreshing their memory of any current state and assisting the user in determining the desired updated state relative to the current state.

[0048] 4A and 4B show methods 400 and 420 for controlling an automated assistant to repeatedly provide graphical elements in a display interface when a user provides a request directed to the automated assistant. Methods 400 and 420 may be implemented by one or more applications, computing devices, and / or any other application or module capable of interacting with an automated assistant. Method 400 may include an operation 402 for determining whether a user has provided a request to the automated assistant. The request may be included in a voice utterance, such as, for example, "Assistant, adjust ...," where the voice utterance may refer to a first portion of the request to be fulfilled by the automated assistant.

[0049] When the automated assistant determines that at least a portion of the request has been received, method 400 may proceed from operation 402 to operation 404. Operation 404 may include determining whether the request corresponds to a complete request or an incomplete request. In other words, the automated assistant may determine whether the user has provided sufficient information for the automated assistant to initiate fulfillment of the request. Following the above example, when the user provides the voice utterance “Assistant, adjust…,” the automated assistant may determine that the request corresponds to an incomplete request. Based on this determination, method 400 may proceed from operation 404 to operation 406. Otherwise, if the automated assistant determines that the request corresponds to a complete request, method 400 may proceed from operation 404 to operation 424 via continuation element “B,” as shown in and described in connection with FIGS. 4A and 4B. In various implementations, method 400 may proceed to operation 406 even when the request is determined to be a complete request at operation 404. Therefore, it should be understood that the methods 400 and 420 are illustrative and not meant to be limiting.

[0050] Operation 406 may include rendering the generic container graphical element in a display interface of the computing device. The generic container graphical element may act as a placeholder for other graphical elements to which the generic container graphical element may be dynamically adapted. For example, the generic container graphical element may be a graphical rendering of a shape having a body that includes sufficient area to allocate other graphical elements. The other graphical elements may include, but are not limited to, control elements for controlling one or more applications and / or devices, status elements for indicating a current state of one or more applications and / or devices, device identifiers for one or more applications and / or devices, media elements based on media content, and / or any other type of element that may be rendered in a display interface. Method 400 may proceed from operation 406 to operation 408, which may include determining whether additional portions of the incomplete request have been received by the automated assistant.

[0051] The additional portion of the request may be a vocal utterance, such as "...temperature...". When it is determined that the additional portion of the incomplete request has been received, method 400 may proceed from operation 408 to operation 410. Otherwise, when it is determined that another portion of the incomplete request has not been received, method 400 may proceed from operation 408 to optional operation 424 of method 420 via continuation element "A", as shown in and described in connection with FIGS. 4A and 4B. Optional operation 424, shown in FIG. 4B, may include rendering one or more selectable suggestions in a display interface. The one or more selectable suggestions may be based on one or more portions of the incomplete request received from the user by the automated assistant. In this way, even though the user did not provide a complete request, the automated assistant may still provide one or more selectable suggestions that are predicted to correspond to one or more intents the user may be attempting to communicate. Thereafter, the method 420 may optionally return to operation 402 via continuation element "C" as shown in Figures 4A and 4B.

[0052] When it is determined that additional portions of the incomplete request have been received at operation 408, method 400 may proceed from operation 408 to operation 410. Operation 410 may include determining that the incomplete request corresponds to a particular intent. In some examples, operation 410 may be performed after a user provides one or more additional inputs to complete the incomplete request. For example, when a user provides a first portion of a request, “Assistant, adjust ...,” followed by a second portion of the request (e.g., contained within the same voice utterance or an additional voice utterance following the voice utterance), “...temperature...,” the automated assistant may determine that the user is requesting to modify the temperature setting of an application and / or device. In some implementations, a particular intent may have slot values ​​for corresponding parameters associated with the particular intent. In some implementations, a request from a user may be deemed incomplete based on a predicted probability, as determined by the automated assistant or other application. For example, the predicted probability may indicate the likelihood that the user is requesting that the particular intent be executed. When the predicted probability meets a probability threshold, the request from the user may be deemed complete. Slot values ​​for the corresponding parameters can then be assigned to specific intents based on additional input from the user and / or data available to the automated assistant.

[0053] Method 400 may proceed from operation 410 to operation 412, which may include dynamically adapting a generic container graphical element, resulting in a tailored container graphical element specific to a particular intent. The tailored container graphical element may include one or more control elements for controlling a particular assistant action, one or more status elements indicating the status of an application and / or device, and / or other content described herein. For example, based on portions of the request (e.g., "Assistant, adjust ..."), the automated assistant may dynamically adapt various graphical elements to the generic container graphical element for indicating the current temperature setting of a particular device (e.g., a hallway thermostat). Alternatively or additionally, the automated assistant may have another graphical element assigned to the generic container graphical element for adjusting the temperature setting of the particular device. In this way, the user will be able to see the current status of the particular device and even options for controlling the particular device. By repeatedly assigning graphical elements to the container graphical element, the automated assistant and the computing device may conserve time and resources that would otherwise be consumed waiting for the user to provide a completed request.

[0054] Method 400 may proceed from operation 412 to operation 414, which may include determining whether input for initializing a particular assistant action has been received. The input may be, for example, another portion of the request included within the same voice utterance or an additional voice utterance following the voice utterance (e.g., "...to 72 degrees..."), and / or touch input in an area of ​​the display interface that is rendering a particular adjusted container graphical element. When it is determined that input for initializing fulfillment of the request has been received, method 400 proceeds from operation 414 via continuation element "B" to operation 422 of method 420, as shown in and described in connection with FIG. 4A and FIG. 4B. Otherwise, method 400 may proceed from operation 414 via continuation element "A" to optional operation 424 and / or operation 402.

[0055] Operation 422 may include initializing a fulfillment based on a request from a user. The fulfillment may correspond to executing a particular intent to fulfill the request. For example, the user may provide an additional vocal utterance, such as "...to 72 degrees...," based on a particular adjusted container graphical element indicating that the device's current state is 67 degrees. Thus, the user's memory may be refreshed by the information conveyed in the particular adjusted container graphical element. Method 420 may proceed from operation 422 to optional operation 426, which may include rendering a response output based on the fulfillment. For example, the particular adjusted container graphical element may be assigned additional content based on the implementation of the fulfillment. Following the foregoing example, the container graphical element may be assigned additional graphical content to indicate that the device's temperature setting has been successfully adjusted or modified from a current state of 67 degrees to an updated state of 72 degrees. Thereafter, method 420 may return to operation 402 via continuation element "C."

[0056] 5 is a block diagram 500 of an exemplary computer system 510. The computer system 510 typically includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. Such peripheral devices may include a storage subsystem 524, including, for example, memory 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with the computer system 510. The network interface subsystem 516 provides an interface to external networks and is coupled to corresponding interface devices in other computer systems.

[0057] The user interface input devices 522 may include pointing devices such as keyboards, mice, trackballs, touchpads, graphics tablets, scanners, touchscreens integrated into displays, voice recognition systems, audio input devices such as microphones, and / or other types of input devices. In general, use of the term "input device" is intended to include all possible types of devices and ways to input information into the computer system 510 or onto a communications network.

[0058] The user interface output devices 520 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for producing a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. In general, use of the term "output device" is intended to include all possible types of devices and ways for outputting information from the computer system 510 to a user or to another machine or computer system.

[0059] Storage subsystem 524 stores programming and data configurations that provide the functionality of some or all of the modules described herein. For example, storage subsystem 524 may include logic for performing selected aspects of method 400 and / or implementing one or more of system 300, computing device 104, computing device 204, an automated assistant, and / or any other applications, devices, apparatuses, and / or modules discussed herein.

[0060] Such software modules are typically executed by the processor 514 alone or in combination with other processors. The memory 525 used in the storage subsystem 524 may include several memories, including a main random access memory (RAM) 530 for storing instructions and data during program execution, and a read-only memory (ROM) 532 in which fixed instructions are stored. The file storage subsystem 526 may provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of some implementations may be stored in the storage subsystem 524 by the file storage subsystem 526 or in other machines accessible by the processor 514.

[0061] Bus subsystem 512 provides the mechanism that allows the various components and subsystems of computer system 510 to communicate with each other in the intended manner. Although bus subsystem 512 is shown generally as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0062] The computer system 510 can be of various types, including a workstation, a server, a computing cluster, a blade server, a server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computer system 510 shown in Figure 5 is intended only as a specific example to illustrate some implementations. Many other configurations of the computer system 510 are possible, having more or fewer components than the computer system shown in Figure 5.

[0063] In situations where the systems described herein may collect or utilize personal information about users (or, as often referred to herein, “participants”), users may be provided with the opportunity to control whether a program or feature collects user information (e.g., information about the user's social network, social actions or activities, occupation, user preferences, or the user's current geographic location) or whether and / or how content that may be more relevant to the user is received from a content server. Additionally, certain data may be processed in one or more ways to remove personally identifiable information before being stored or used. For example, the user's identity may be processed so that personally identifiable information about the user cannot be determined, or if geographic location information is obtained (e.g., to the city, zip code, state level), the user's geographic location may be generalized so that the user's specific geographic location cannot be determined. Thus, users may have control over how information is collected and / or used about them.

[0064] While several implementations have been described and illustrated herein, various other means and / or structures for performing the functions and obtaining the results and / or one or more of the advantages described herein may be utilized, and each such variation and / or modification is deemed to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend on the particular application or applications in which the teachings are used. Those skilled in the art will recognize and be able to ascertain, using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, it should be understood that the foregoing implementations are presented merely by way of example, and that, within the scope of the appended claims and their equivalents, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

[0065] In some implementations, a method implemented by one or more processors is provided, including receiving, at a computing device, a first portion of a request from a user contained in a voice utterance. The voice utterance is directed to an automated assistant accessible via the computing device. The method further includes determining one or more candidate intents associated with the first portion of the request based on processing the first portion of the request, and rendering a generic container graphical element in a display interface of the computing device based on the one or more candidate intents. The generic container graphical element can be dynamically adapted to any one of a plurality of heterogeneous tailored container graphical elements, each of the plurality of heterogeneous tailored container graphical elements being associated with a corresponding one of the one or more candidate intents. The method further includes receiving, at the computing device, a second portion of the request from the user that is also included in the voice utterance or that is included in an additional voice utterance received after the voice utterance, determining that the request corresponds to a particular intent among the one or more candidate intents based on processing the second portion of the request, and dynamically adapting, based on the particular intent, a generic container graphical element rendered in the display interface to a particular tailored container graphical element among the plurality of heterogeneous tailored container graphical elements.

[0066] Optionally, these and other implementations of the techniques described herein include one or more of the following features.

[0067] In some implementations, a particular tailored container graphical element may feature a slot value for a corresponding parameter associated with a particular intent, and the first and second parts of the request may not specify a slot value.

[0068] In some implementations, the method may further include, in response to receiving the first portion of the request, causing a display interface of the computing device to visually render natural language content characterizing the first portion of the request. The display interface may render the natural language content of the first portion of the request simultaneously with rendering the generic container graphical element.

[0069] In some implementations, the method may further include determining that a threshold duration has elapsed after receiving the first portion of the request. Rendering the generic container graphical element in the display interface may be performed based on the elapse of the threshold duration.

[0070] In some implementations, the particular adjusted container graphical element may include a particular geographic control element associated with a current state of one or more settings of the computing device or one or more additional computing devices in communication with the computing device. In some versions of such implementations, the method may further include receiving, at the computing device, a third portion of the request from the user also included within the voice utterance, the additional voice utterance, or a further voice utterance received after the voice utterance or the further voice utterance, the third portion of the request including an updated state for the one or more settings, and causing, by the automated assistant, the one or more settings of the computing device or one or more of the additional computing devices to change from the current state to the updated state.

[0071] In some implementations, dynamically adapting a generic container graphical element rendered in the display interface to the particular tailored container graphical element may include selecting the particular tailored container graphical element from among a plurality of disparate tailored container graphical elements based on a type of slot value identified in the second speech utterance. The type of slot value may correspond to a numeric value limited to a range of numeric values.

[0072] In some implementations, the generic container graphical element may be rendered in a display interface of the computing device before receiving the second part of the request.

[0073] In some implementations, the generic container graphical element may be rendered in a display interface of the computing device while the second part of the request is received.

[0074] In some implementations, determining one or more candidate intents associated with the first portion of the request based on processing the first portion of the request may include: processing a stream of audio data generated by one or more microphones of the computing device using a streaming automatic speech recognition (ASR) model to generate an ASR output, where the stream of audio data captures the first portion of the request; processing the ASR output using a natural language understanding (NLU) model to generate the NLU output; and determining one or more candidate intents associated with the first portion of the request based on the NLU output. In some versions of such implementations, determining that the request corresponds to a particular intent of the one or more candidate intents based on processing a second portion of the request may include processing the stream of audio data using a streaming ASR model to generate an additional ASR output, where the stream of audio data also captures the second portion of the request; processing the additional ASR output using the NLU model to generate the additional NLU output; and selecting a particular intent from the one or more candidate intents based on the additional NLU output.

[0075] In some implementations, a method implemented by one or more processors is provided, the method including: receiving, at a computing device, a portion of a request submitted by a user, the portion of the request being included in a voice utterance directed to an automated assistant accessible via the computing device; determining, by the automated assistant, the portion of the request being associated with modifying a current state of one or more settings of the computing device or one or more additional computing devices in communication with the computing device via the automated assistant; determining, based on the current state of the one or more settings, adjusted container graphical element data characterizing the current state of the one or more settings; and, based on the adjusted container graphical element data, modifying the one or more settings. causing a display interface of the computing device to render one or more adjusted container graphical elements indicating a current state of the one or more settings; and in response to causing a display interface of the computing device to render one or more adjusted container graphical elements indicating the current state of the one or more settings, receiving at the computing device an additional portion of a request submitted by the user, the additional portion of the request being included in a voice utterance or an additional voice utterance received after the voice utterance, the additional portion of the request including an updated state for the one or more settings; and causing, by the automated assistant, to change the one or more settings of the computing device or one or more of the additional computing devices from the current state to the updated state.

[0076] Optionally, these and other implementations of the techniques described herein include one or more of the following features.

[0077] In some implementations, each of the one or more adjusted container graphical elements may include a graphical icon to represent the current state of one or more settings.

[0078] In some implementations, the method may further include, in response to receiving the portion of the request, causing a display interface of the computing device to visually render natural language content that characterizes the portion of the request. The one or more coordinated graphical container elements may be rendered simultaneously with the display interface rendering the natural language content.

[0079] In some implementations, the request portion may not include the current state of one or more settings.

[0080] In some implementations, causing the display interface of the computing device to render one or more adjusted container graphical elements indicating a current state of the one or more settings based on the adjusted container graphical element data may include rendering a first adjusted container graphical element of the one or more adjusted container graphical elements that indicates a first setting of the one or more settings of a first computing device of the one or more additional computing devices that is separate from the computing device, and rendering a second adjusted container graphical element of the one or more adjusted container graphical elements that indicates a second setting of the one or more settings of a second computing device of the one or more additional computing devices that is separate from the computing device. In some versions of such implementations, the first setting of the first computing device may correspond to a volume setting of the first computing device, and the second setting of the second computing device may correspond to a volume setting of the second computing device. In additional or alternative versions of such implementations, the first setting of the first computing device may correspond to a brightness setting of the first computing device, and the second setting of the second computing device may correspond to a brightness setting of the second computing device.

[0081] In some implementations, the method may further include determining, based on processing the portion of the request, that the user paused in providing the request. Rendering one or more adjusted container graphical elements on a display interface of the computing device that indicate a current state of the one or more settings may be responsive to determining that the user paused in providing the request. In some versions of such implementations, determining, based on processing the portion of the request, that the user paused in providing the request may include determining, based on processing the portion of the request, that the user paused after providing a particular word or phrase in providing the request. In additional or alternative versions of such implementations, rendering one or more adjusted container graphical elements on a display interface of the computing device that indicate a current state of the one or more settings may be responsive to determining that the user paused for a threshold duration in providing the request. In additional or alternative versions of such implementations, determining that the user paused in providing the request based on processing the portion of the request may include processing a stream of audio data generated by one or more microphones of the computing device using a streaming automatic speech recognition (ASR) model to generate an ASR output, where the stream of audio data captures the portion of the request; processing the ASR output using a natural language understanding (NLU) model to generate the NLU output; and determining that the user paused in providing the request based on the NLU output. In additional or alternative versions of such implementations, determining that the user paused in providing the request based on processing the portion of the request may include determining audio-based characteristics associated with the portion of the request based on processing the portion of the request; and determining that the user paused in providing the request based on the audio-based characteristics associated with the portion of the request.

[0082] In some implementations, a method implemented by one or more processors is provided, the method including receiving, at a computing device, a first portion of a request from a user contained within a voice utterance. The voice utterance is directed to an automated assistant accessible via the computing device. The method further includes, based on processing the first portion of the request, determining that the portion of the request includes a specific word or phrase associated with controlling the computing device or one or more additional computing devices in communication with the computing device, and rendering a generic container graphical element in a display interface of the computing device based on the first portion of the request including the specific word or phrase. The generic container graphical element can be dynamically adapted to any one of a plurality of heterogeneous tailored container graphical elements, each of the plurality of heterogeneous tailored container graphical elements being associated with a corresponding intent determined based on processing the first portion of the voice utterance. The method further includes receiving, at the computing device, a second portion of the request from the user that is also included in the voice utterance or that is included in an additional voice utterance received following the voice utterance, determining that the request corresponds to a particular intent among the one or more candidate intents based on processing the second portion of the request, and dynamically adapting, based on the particular intent, a generic container graphical element rendered in the display interface to a particular tailored container graphical element among the plurality of heterogeneous tailored container graphical elements.

[0083] Optionally, these and other implementations of the techniques described herein include one or more of the following features.

[0084] In some implementations, determining, based on processing the first portion of the request, that the portion of the request includes a specific word or phrase associated with controlling the computing device or one or more additional computing devices in communication with the computing device may include processing a stream of audio data generated by one or more microphones of the computing device using a streaming automatic speech recognition (ASR) model to generate an ASR output, wherein the stream of audio data captures the first portion of the request, and determining, based on the ASR output, that the portion of the request includes a specific word or phrase associated with controlling the computing device or one or more additional computing devices.

[0085] In some versions of such an implementation, rendering a generic container graphical element on a display interface of a computing device based on a first portion of a request that includes a particular word or phrase may include determining that the particular word or phrase maps to a generic container graphical element in on-device memory of the computing device, and in response to determining that the particular word or phrase maps to the generic container graphical element, rendering the generic container graphical element on the display interface of the computing device without processing the ASR output using a natural language understanding (NLU) model.

[0086] Another implementation may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) to perform a method such as one or more of the methods described above and / or elsewhere herein. Yet another implementation may include a system of one or more computers including one or more processors operable to execute the stored instructions to perform a method such as one or more of the methods described above and / or elsewhere herein.

[0087] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are contemplated as being part of the subject matter disclosed herein, for example, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the subject matter disclosed herein. [Explanation of symbols]

[0088] 102 users 104 Computing Devices 106 Display Interface 108 requests 110 Generic Container Graphical Elements 112 Graphical Container Elements 120 Current Status 122 request 124 Media Content 142 Request 144 Control Elements 202 users 204 Computing Devices 206 Display Interface 208 Request 210 Generic Container Graphical Elements 222A Graphical Container Elements 222B Graphical Container Elements 222C Graphical Container Elements 224 Device Identifier 226 Request 228 Device Identifier 230 Device Identifier 232 Natural Language Content 242 request 300 System 302 Computing Devices 304 Automated Assistant 306 Input Processing Engine 308 Audio Processing Engine 310 Data Parsing Engine 312 Parameter Engine 314 Output Generation Engine 316 Candidate Intent Engine 318 General-purpose Container Engine 320 Assistant Interface 322 Assistant Call Engine 324 State Engine 326 State Engine 330 Application Data 332 Device Data 334 Specific Applications 336 Context Data 510 Computer Systems 512 Bus Subsystem 514 processor 516 Network Interface Subsystem 520 User Interface Output Device 522 User Interface Input Devices 524 Memory Subsystem 526 File Storage Subsystem 525 memory 526 File Storage Subsystem 530 Main Random Access Memory (RAM) 532 Read-Only Memory (ROM)

Claims

1. 1. A method implemented by one or more processors, comprising: receiving, at a computing device, a first portion of a request from a user, the first portion of the request received from the user being included in a user input directed to the computing device; determining that the first portion of the request received from the user includes a particular word or phrase that is mapped to a generic container graphical element; in response to determining that the first portion of the request received from the user includes a particular word or phrase that is mapped to the generic container graphical element and determining that the first portion of the request corresponds to an incomplete request, causing the generic container graphical element to be rendered on a display of the computing device, the generic container graphical element being capable of being dynamically adapted to any one of a plurality of heterogeneous tailored container graphical elements; receiving, at the computing device, a second portion of the request from the user, the second portion of the request received from the user being included within the user input directed to the computing device that includes the first portion of the request; dynamically adapting the generic container graphical element rendered on the display of the computing device to one or more particular coordinated container graphical elements among the plurality of heterogeneous coordinated container graphical elements based on a second portion of the request; A method comprising:

2. determining that the second portion of the request received from the user corresponds to a particular intent among one or more candidate intents; 2. The method of claim 1, wherein the one or more specific tailored container graphical elements utilized to dynamically adapt to the generic container graphical element rendered on the display of the computing device are associated with the specific intent.

3. The method of claim 2 , wherein the one or more particular adjusted container graphical elements feature slot values ​​for corresponding parameters associated with the particular intent.

4. 10. The method of claim 1, further comprising causing the display of the computing device to render natural language content characterizing the first portion of the request simultaneously with rendering the generic container graphical element.

5. The method of claim 4 , wherein the natural language content characterizing the first portion of the request is rendered in another portion of the display relative to the generic container graphical element.

6. 2. The method of claim 1, wherein each of the one or more particular coordinated container graphical elements includes a corresponding particular graphical control element that enables the user to interact with each of the one or more particular coordinated container graphical elements.

7. receiving, at the computing device, a further user input directed to a given one of the corresponding specific graphical control elements; further dynamically adapting the generic container graphical element rendered on the display of the computing device based on the further user input; 7. The method of claim 6, further comprising:

8. 1. A method implemented by one or more processors, comprising: receiving, at a computing device, a first portion of a request from a user contained within a voice utterance, the voice utterance being directed to an automated assistant accessible via the computing device; determining one or more candidate intentions associated with the first portion of the request based on processing the first portion of the request, wherein processing the first portion of the request includes processing a stream of audio data using a streaming automatic speech recognition (ASR) model to generate a first streaming transcription corresponding to first natural language content for the first portion of the request, and determining the one or more candidate intentions based on processing the first streaming transcription corresponding to the first natural language content for the first portion of the request; in response to determining that the first portion of the request corresponds to an incomplete request, causing, at a display interface of the computing device based on the one or more candidate intents, to render a generic container graphical element along with the first streaming transcription corresponding to the first natural language content for the first portion of the request; the generic container graphical element may be dynamically adapted to any one of a plurality of heterogeneous tailored container graphical elements; each of the plurality of heterogeneous coordinated container graphical elements being associated with a corresponding one of the one or more candidate intents; receiving, at the computing device, a second portion of the request from the user contained within the voice utterance or contained within an additional voice utterance received subsequent to the voice utterance; determining that the request corresponds to a particular intent among the one or more candidate intents based on processing the second portion of the request, wherein processing the second portion of the request includes processing a stream of audio data using a streaming ASR model to generate a second streaming transcription corresponding to second natural language content for the second portion of the request, and determining that the request corresponds to the particular intent is based on processing the second streaming transcription corresponding to the second natural language content for the second portion of the request; dynamically adapting the generic container graphical element rendered in the display interface to a particular coordinated container graphical element of the plurality of heterogeneous coordinated container graphical elements based on the particular intent; A method comprising:

9. the particular adjusted container graphical element characterizes a slot value for a corresponding parameter associated with the particular intent; The method of claim 8 , wherein the first portion of the request and the second portion of the request do not specify the slot value.

10. responsive to receiving the first portion of the request, causing the display interface of the computing device to visually render natural language content that characterizes the first portion of the request; The method of claim 8 , wherein the display interface renders the natural language content of the first portion of the request simultaneously with rendering the generic container graphical element.

11. determining that a threshold duration has elapsed after receiving the first portion of the request; The method of claim 8 , wherein the step of causing the generic container graphical element to be rendered in the display interface is performed based on the lapse of the threshold duration.

12. said step of dynamically adapting said generic container graphical element rendered in said display interface to said particular tailored container graphical element comprises: selecting the particular adjusted container graphical element from among the plurality of heterogeneous adjusted container graphical elements based on a type of slot value identified in the second speech utterance; said type of slot value corresponds to a numeric value limited to a range of numeric values; 9. The method of claim 8, comprising:

13. The method of claim 8 , wherein the generic container graphical element is rendered in the display interface of the computing device before receiving the second portion of the request.

14. The method of claim 8 , wherein the generic container graphical element is rendered on the display interface of the computing device while the second portion of the request is received.

15. determining the one or more candidate intents based on processing the first streaming transcription corresponding to the first natural language content for the first portion of the request, processing the first streaming transcription corresponding to the first natural language content using a natural language understanding (NLU) model to generate a first NLU output; determining the one or more candidate intents associated with the first portion of the request based on the first NLU output; 9. The method of claim 8, comprising:

16. determining that the request corresponds to the particular intent based on processing the second streaming transcription corresponding to the second natural language content for the second portion of the request, processing the second streaming transcription corresponding to the second natural language content for the second portion of the request using the NLU model to generate a second NLU output; selecting the particular intent based on the second NLU output; 16. The method of claim 15, comprising:

17. 9. The method of claim 8, wherein the particular adjusted container graphical element comprises a particular graphical control element associated with a current state of one or more settings of the computing device or one or more additional computing devices in communication with the computing device.

18. receiving, at the computing device, a third portion of the request from the user contained within the voice utterance, the additional voice utterance, or a further voice utterance received after the voice utterance or the additional voice utterance, the third portion of the request including an updated state for the one or more settings; causing the automated assistant to change the one or more settings of the computing device or the one or more additional computing devices from the current state to the updated state; 18. The method of claim 17, further comprising:

19. 1. A method implemented by one or more processors, comprising: Receiving, at a computing device, a portion of a request submitted by a user, the portion of the request being included in a voice utterance directed to an automated assistant accessible via the computing device; determining, by the automated assistant, based on processing the portion of the request, that the portion of the request is associated with modifying a current state of one or more settings of the computing device or one or more additional computing devices in communication with the computing device via the automated assistant, wherein processing the portion of the request includes processing a stream of audio data using a streaming automatic speech recognition (ASR) model to generate a streaming transcription corresponding to natural language content for the portion of the request, and determining that the portion of the request is associated with modifying the current state of the one or more settings is based on processing the streaming transcription corresponding to the natural language content for the portion of the request; determining adjusted container graphical element data characterizing the current state of the one or more settings based on the current state of the one or more settings; in response to determining that the portion of the request corresponds to an incomplete request, causing a display interface of the computing device to render one or more adjusted container graphical elements indicating the current state of the one or more settings based on the adjusted container graphical element data and the streaming transcription corresponding to the natural language content for the portion of the request; in response to causing the display interface of the computing device to render the one or more adjusted container graphical elements indicating the current state of the one or more settings. receiving, at the computing device, an additional portion of the request submitted by the user, the additional portion of the request being included within the voice utterance or an additional voice utterance received after the voice utterance; determining, by the automated assistant, based on processing the additional portion of the request, that the additional portion of the request includes an updated state for the one or more settings, wherein processing the additional portion of the request includes processing a stream of audio data using a streaming ASR model to generate an additional streaming transcription corresponding to additional natural language content for the additional portion of the request, and determining that the additional portion of the request includes an updated state for the one or more settings is based on processing the additional streaming transcription corresponding to the additional natural language content for the additional portion of the request; causing the automated assistant to change the one or more settings of the computing device or one or more of the additional computing devices from the current state to the updated state; A method comprising:

20. 20. The method of claim 19, wherein each of the one or more adjusted container graphical elements includes a graphical icon to represent the current state of the one or more settings.

21. responsive to receiving the portion of the request, causing the display interface of the computing device to visually render natural language content that characterizes the portion of the request; The method of claim 19 , wherein the one or more coordinated graphical container elements are rendered simultaneously as the display interface renders the natural language content.

22. 20. The method of claim 19, wherein the portion of the request does not include the current state of the one or more settings.

23. causing the display interface of the computing device to render the one or more adjusted container graphical elements indicating the current state of the one or more settings based on the adjusted container graphical element data, causing a first adjusted container graphical element of the one or more adjusted container graphical elements to be rendered that indicates a first configuration of the one or more configurations of a first computing device of the one or more additional computing devices that is different from the computing device; causing a second adjusted container graphical element of the one or more adjusted container graphical elements to be rendered that indicates a second configuration of the one or more configurations of a second computing device of the one or more additional computing devices that is different from the computing device; 20. The method of claim 19, comprising:

24. the first setting of the first computing device corresponds to a volume setting of the first computing device; 24. The method of claim 23, wherein the second setting of the second computing device corresponds to a volume setting of the second computing device.

25. the first setting of the first computing device corresponds to a brightness setting of the first computing device; 24. The method of claim 23, wherein the second setting of the second computing device corresponds to a brightness setting of the second computing device.

26. determining, based on processing the portion of the request, that the user paused in providing the request; 20. The method of claim 19, wherein causing the display interface of the computing device to render the one or more adjusted container graphical elements indicating the current state of the one or more settings is in response to determining that the user paused in providing the request.

27. 1. A method implemented by one or more processors, comprising: receiving, at a computing device, a first portion of a request from a user contained within a voice utterance, the voice utterance being directed to an automated assistant accessible via the computing device; determining, based on processing the first portion of the request, that the first portion of the request includes a specific word or phrase associated with controlling the computing device or one or more additional computing devices in communication with the computing device, wherein processing the first portion of the request includes processing a stream of audio data using a streaming automatic speech recognition (ASR) model to generate a first streaming transcription corresponding to first natural language content for the first portion of the request, and determining that the first portion of the request includes a specific word or phrase based on processing the first streaming transcription corresponding to the first natural language content for the first portion of the request; in response to determining that the first portion of the request corresponds to an incomplete request, causing a generic container graphical element to be rendered at a display interface of the computing device based on the first portion of the request including the particular word or phrase, along with the first streaming transcription corresponding to the first natural language content for the first portion of the request; the generic container graphical element may be dynamically adapted to any one of a plurality of heterogeneous tailored container graphical elements; Associating each of the plurality of disparate coordinated container graphical elements with a corresponding intent determined based on processing the first portion of the speech utterance; receiving, at the computing device, a second portion of the request from the user contained within the voice utterance or contained within an additional voice utterance received subsequent to the voice utterance; determining, based on processing the second portion of the request, that the request corresponds to a particular one of the one or more candidate intentions for controlling the computing device or one or more additional computing devices in communication with the computing device, wherein processing the second portion of the request includes processing a stream of audio data using a streaming ASR model to generate a second streaming transcription corresponding to second natural language content for the second portion of the request, and determining that the request corresponds to the particular intention is based on processing the second streaming transcription corresponding to the second natural language content for the second portion of the request; dynamically adapting the generic container graphical element rendered in the display interface to a particular coordinated container graphical element of the plurality of heterogeneous coordinated container graphical elements based on the particular intent; A method comprising:

28. at least one processor; a memory storing instructions that, when executed, cause said at least one processor to perform operations corresponding to any one of claims 1 to 27; A system comprising:

29. 28. A non-transitory computer-readable storage medium storing instructions that, when executed, cause at least one processor to perform operations corresponding to any one of claims 1 to 27.

Citation Information

Patent Citations

  • Providing composite graphical assistant interfaces for controlling various connected devices

    EP3783867A1

  • Intelligent Digital Assistant in a Multitasking Environment

    JP2019522250A

  • Method for operating speech recognition service and electronic device supporting the same

    US20180285070A1

  • Methods, systems, and apparatus for providing composite graphical assistant interfaces for controlling connected devices

    WO2019216874A1

  • Display control device for selecting item on basis of speech

    WO2020137607A1