Supplementing voice input to the automated assistant with selected suggestions.
Visual suggestions help users complete ambiguous spoken commands, reducing resource waste and latency by guiding automated assistants to perform intended actions effectively.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2025-01-08
- Publication Date
- 2026-05-19
AI Technical Summary
Existing automated assistants often fail to understand incomplete or poorly formulated spoken utterances, leading to resource wastage and increased latency in performing intended actions, as they attempt to process and respond to ambiguous requests, resulting in users needing to reformulate their commands.
Providing real-time visual suggestions on a display to complete or clarify spoken utterances, allowing users to easily integrate additional text or actions to form complete requests, thereby guiding the assistant to perform the intended action.
Reduces computational resource consumption and interaction time by improving speech-to-text accuracy and enabling users to quickly formulate complete requests, thus enhancing the efficiency and responsiveness of automated assistants.
Smart Images

Figure 0007862618000001 
Figure 0007862618000002 
Figure 0007862618000003
Abstract
Description
[Background technology]
[0001] Humans may engage in human-computer dialogue with conversational software applications, which are referred to herein as “Automated Assistants” (also known as “Digital Agents,” “Chatbots,” “Conversational Personal Assistants,” “Intelligent Personal Assistants,” “Assistant Applications,” “Conversational Agents,” etc.). For example, a human (who may be referred to as a “User” when interacting with an Automated Assistant) may provide commands and / or requests to the Automated Assistant by using oral natural language input (i.e., utterances), which may in some cases be converted to text and then processed, and / or by providing natural language input in text form (e.g., typed). The Automated Assistant responds to requests by providing response user interface output, which may include audible and / or visual user interface output. Thus, the Automated Assistant may provide a voice-based user interface.
[0002] In some cases, a user may provide a spoken utterance that is intended by the user to trigger the performance of an automated assistant action, but does not result in the performance of the intended automated assistant action. For example, the spoken utterance may be provided in a syntax that is not understandable by the automated assistant and / or may lack essential parameters for the automated assistant action. As a result, the automated assistant may not be able to fully process the spoken utterance and determine that it is a request for an automated assistant action. This may lead to the automated assistant not providing a response to the spoken utterance or providing an error and / or error tone in a response such as "Sorry, I can't help with that." Despite the automated assistant failing to perform the intended action of the spoken utterance, various computer and / or network resources are consumed in an attempt to process the spoken utterance and resolve the appropriate action. For example, audio data corresponding to the spoken utterance may be transmitted, undergo speech-to-text processing, and / or undergo natural language processing. Such consumption of resources is wasted because the intended automated assistant action is not performed, and the user is likely to attempt to provide a reformulated spoken utterance when subsequently requesting the performance of the intended automated assistant action. Further, such reformulated spoken utterances will also have to be processed. Additionally, this results in latency in the performance of the automated assistant action compared to when the user initially provided a suitable spoken utterance to the automated assistant instead. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0003] The implementations described herein relate to providing suggestions via a display modality for completing spoken utterances directed to an automated assistant, and, optionally, continuously providing updated suggestions based on various factors while the spoken utterance is in progress. For example, suggestions may be provided to facilitate the development of actionable dialogue phrases that lead the automated assistant to comprehensively perform the automated assistant action initially intended by the user, while also streamlining the dialogue between the user and the automated assistant. Additionally or alternatively, suggestions may be provided to facilitate encouraging the user to engage the automated assistant to perform actions that can reduce the frequency and / or duration of the user's participation in the current and / or subsequent dialogue sessions with the automated assistant.
[0004] In some implementations, during a dialogue session between the user and the automated assistant, the user may provide oral utterances to facilitate the automated assistant performing one or more actions. Oral utterances may be captured by audio data and speech-to-text processing performed on the audio data to generate one or more text segments, each of which is a corresponding interpretation of the oral utterance. One or more text segments may be processed to determine whether the oral utterance corresponds to a complete request and / or an actionable request. As used herein, a “complete request” or “actionable request” is one that, when fully processed by the automated assistant, triggers the performance of one or more corresponding automated assistant actions, where the performed automated assistant actions are not default error type actions (e.g., not an audible response such as “Sorry, I can’t help you with that”). For example, an actionable request could trigger the control of one or more smart devices, or the rendering of specific audible and / or graphical content, etc. When it is determined that the oral utterance corresponds to an incomplete request, one or more text segments may be further processed to determine and provide one or more suggestions for completing the request. Each of one or more suggestions may contain additional text that, when combined with a corresponding text segment, provides a complete request. In other words, if the user continues their dialogue session by repeating the text of a particular suggestion (or alternative text that generally fits the text based on a particular suggestion), the automated assistant will perform the corresponding automated assistant action. Thus, the generation of suggestions assists the user in performing the task. While the user is typing the request, the user is provided with information that guides them through their interaction with the assistant.The information presented to the user is based on the user's existing input, and therefore, objective analysis is used to provide the user with information. Thus, the user is provided with objectively relevant information to assist them in performing the task.
[0005] In some implementations, a computing device can generate different interpretations of a spoken utterance, and variations in those interpretations can give rise to corresponding variations in suggestions provided in response to the utterance. For example, a spoken utterance such as "Assistant, can you change this...?" may correspond to multiple interpretations by the computing device. The aforementioned spoken utterance may be interpreted as a request to change IoT device settings via an automated assistant, or as a request to change application settings via an automated assistant. Each interpretation may be processed to generate a corresponding suggestion that can be presented to the user via a display panel communicating with the computing device. For example, the computing device may cause the display panel to show the initial text of the spoken utterance (e.g., "Assistant, can you change this...?") along with a list of suggestions based on both interpretations of the utterance. The suggestions may include at least one suggestion based on one interpretation and at least one other suggestion based on another interpretation.
[0006] For example, at least one suggestion based on the interpretation "change IoT device settings" could be text that complements the original spoken utterance when spoken by the user, and could be "set the thermostat to 72." As an example, the display panel could respond to a user speaking the text of at least one suggestion by presenting the original spoken utterance and at least one suggestion together. In other words, the display panel could present the following text, namely, "Assistant, can you change the thermostat to 72?". The user can see the presented text and identify that the part (i.e., "set the thermostat to 72") is the suggested part and was not included in their original spoken utterance (i.e., "Assistant, can you change it..."). Once the user identifies the suggested part, they can repeat the text of the suggested part, triggering the automated assistant to act to facilitate the completion of an action, such as changing the user's thermostat setting to 72. In some implementations, the display panel can present text as a suggestion, containing only the suggested, yet unspoken portion (e.g., "set the thermostat to 72" without presenting "Assistant, can you do that for me?"). This allows for an easy glance at the suggestion, enabling the user to quickly see the suggestion and confirm further verbal input that can be provided to construct an actionable request, such as "set the thermostat to 72" or "set the thermostat to 70" (which is different from the suggestion but verifiable based on the suggestion).
[0007] In various implementations, the text portion of a suggestion may be presented in combination with icons or other graphical elements that indicate the action to be taken as a result of providing further verbal input that conforms to the text portion of the suggestion. For example, a suggestion may include the text "Change the thermostat to 72" along with a graphical indicator of a thermostat and up and down arrows to indicate that further verbal input conforming to the suggestion will result in a change in the thermostat's setpoint. Another example is a suggestion that may include the text "Play a TV station on the living room TV" along with an outline of a television with a play icon inside to indicate that further verbal input conforming to the suggestion will result in playing streaming content on the television. Yet another example is a suggestion that may include "Turn off the living room lights" along with a light bulb icon with an X superimposed on it to indicate that further verbal input conforming to the suggestion will result in turning off the lights. Such icons can make suggestions easier to glance at, allowing users to quickly see the icon first, and then, if the icon matches the user's desired action, view the text portion of the suggestion. This allows users to quickly glance at the graphical elements of multiple simultaneously presented suggestions (without first reading the corresponding text portion), identify those that match their intent, and then read only the corresponding text portion. This reduces the latency of providing further verbal input when a user is presented with multiple suggestions, and consequently reduces the latency of performing the resulting action.
[0008] In some implementations, spoken utterances can be interpreted differently as a result of how the user explicitly states and / or pronounces certain words. For example, a computing device receiving the above-mentioned spoken utterance, "Assistant, can you change it...", may determine with X% certainty that the utterance contains the word "change", and with Y% certainty (where Y is less than X) that the utterance contains the word "organize", and therefore refers to a request to perform the "organize" function. For example, speech-to-text processing of the spoken utterance may result in two distinct interpretations ("change" and "organize") with different degrees of certainty. As a result, an automated assistant tasked with responding to spoken utterances may present suggestions based on each interpretation ("change" and "organize"). For example, the automated assistant can graphically present a first suggestion, such as "...set the thermostat to 72," which corresponds to the interpretation of "change," and a second suggestion, such as "...my desktop file folder," which corresponds to the interpretation of "organize." Furthermore, if the user later repeats the content of the first suggestion by repeating "set the thermostat to 72" (or similar content such as "set the thermostat to 80"), the user can trigger the "change" function to be performed via the automated assistant (and based on the user's further repetition). However, if the user later repeats the content of the second suggestion by repeating "my desktop file folder" (or similar content such as "my documents folder"), the user can trigger the "organize" function to be performed via the automated assistant (and based on the user's further repetition). In other words, the automated assistant will utilize the interpretation ("change" or "organize") that fits the repeated suggestion.For example, even if "organize" is predicted with only 20% probability and "change" is predicted with 80% probability, "organize" may still be used if the user further says "My Desktop Files Folder." As another example, and as variations of the "change" and "organize" examples, a first suggestion "Please change the thermostat to 75" and a second suggestion "Please organize My Desktop Files Folder" may be presented. If the user selects "Please organize My Desktop Files Folder" via touch input or provides further verbal input such as "Please organize My Desktop Files Folder" or "the second one," the "organize" interpretation will be used instead of the "change" interpretation. In other words, even if "change" is selected first as the correct interpretation based on a higher probability, "organize" may nevertheless replace that selection based on the user's selection or further verbal input directed towards the "organize" suggestion. In these and other ways, speech-to-text accuracy can be improved, and interpretations with lower certainty may be used in some situations. This can conserve computational resources that might otherwise be spent if an inaccurate interpretation of the initial spoken utterance is selected, forcing the user to repeat their initial utterance in an attempt to obtain an accurate interpretation. Furthermore, this can reduce the amount of time the user spends engaging with the automated assistant, as the user will repeat themselves less and will wait until the automated assistant renders an audible response.
[0009] In some implementations, provided suggestion elements can be used as a basis for biasing speech-to-text processing of further oral utterances received in response to the suggestion. For example, speech-to-text processing may be biased towards the terminology of the provided suggestion element and / or towards terminology that fits the expected category or other type of content corresponding to the suggestion. For example, with a provided suggestion of "[Musician's Name]", speech-to-text processing may be biased towards the musician's name (generally or towards musicians in the user's library); with a provided suggestion of "at 2:00", speech-to-text processing may be biased towards "2:00", or more generally towards the time; and with a provided suggestion of "[Smart Device]", speech-to-text processing may be biased towards the name of the user's smart device (for example, as identified from the user's stored device topology). Speech-to-text processing can be biased towards certain terms by, for example, selecting a specific speech-to-text model (from multiple candidate speech-to-text models) to use when performing speech-to-text processing, increasing the score of candidate translations (generated during speech-to-text processing) based on candidate translations corresponding to a term, increasing the number of paths in the state decoding graph (used in speech-to-text processing) corresponding to a term, and / or by additional or alternative speech-to-text biasing techniques. As an example, a user might provide an initial verbal utterance such as "Assistant, set up calendar events for ~," and in response to receiving the initial verbal utterance, the automated assistant might cause several suggestion elements to appear on the computing device's display panel. The suggestion elements might be, for example, "August...September...[current month]...," and the user can select a specific suggestion element by providing a subsequent verbal utterance such as "August."In response to a decision that a suggestion element includes January through December, the automated assistant can bias the speech-to-text processing of subsequent spoken utterances toward January through December. As another example, in response to the provision of a suggestion element that includes (or indicates) a numerical input, speech-to-text processing can be biased toward the numerical input. In these and other ways, the speech-to-text processing of subsequent utterances can be biased to prioritize the provided suggestion element, thereby increasing the probability that the text predicted from speech-to-text processing is accurate. This reduces the opportunity for further dialogue that would otherwise be required if the predicted text was inaccurate (for example, based on further input from the user to correct an inaccurate interpretation).
[0010] In some implementations, the timing of providing suggestion elements to an automated assistant to complete a specific request may be based on the context in which the user initiated the request. For example, if a user provides at least part of a request via verbal utterance while driving their vehicle, the automated assistant may determine that the user is driving and delay displaying suggestions to complete the request. In this way, the user is less likely to be distracted by graphics being presented on the vehicle's display panel. Additionally or alternatively, the content of the suggestion elements (e.g., suggested text segments) may also be based on the context in which the user initiated the verbal utterance. For example, with prior permission from the user, the automated assistant may determine that the user provided a verbal utterance such as "Assistant, turn on..." from their living room. In response, the automated assistant may cause the computing device to generate a suggestion element that identifies a specific device in the user's living room. The suggestion element may include natural language content such as "Living room TV...Living room lights...Living room stereo." The user can make a selection from a list of suggestion elements by providing a subsequent verbal utterance such as "living room light," and in response, the automated assistant can turn on the living room light. In some implementations, the devices identified by the suggestion elements may be based on device topology data that characterizes the relationships between various devices associated with the user and / or the placement of computing devices within the user's location, such as home or office. Device topology data may include identifiers for various areas within the location, as well as device identifiers that characterize whether a device is located within each area.For example, device topology data can identify areas such as "living room" and devices within that living room, such as "TV," "lights," and "stereo." Therefore, in response to the user providing the aforementioned utterance, "Assistant, turn it on," the device topology can be accessed and compared to the area the user is in in order to provide suggestions corresponding to specific devices within that area (for example, identifying the "living room" based on the fact that the utterance is provided via devices defined by the device topology data as being located in the "living room").
[0011] In some implementations, actions corresponding to at least a portion of the received request may be performed while suggestion elements are being presented to the user to complete the request. For example, a user might provide a verbal utterance such as, "Assistant, create calendar events for...". In response, the automated assistant may trigger the creation of a default calendar with default content, and may also trigger the creation of suggestion elements to complete the request for creating calendar events. For example, the suggestion elements might include "Tonight...Tomorrow...Saturday...", and the suggestion elements may be presented before, during, or after the default calendar event is created. When the user selects one of the suggestion elements, the previously created default calendar event may be modified according to the suggestion element selected by the user.
[0012] In some implementations, an automated assistant can provide suggestions to reduce the amount of speech processing and / or the amount of assistant interaction time, which is typically associated with several user inputs. For example, when a user completes a dialogue session with an automated assistant to create a calendar event, the automated assistant can nevertheless provide further suggestions. For instance, suppose the user provides a verbal utterance such as, "Assistant, create a calendar event for Matthew's birthday next Monday." In response, the automated assistant can cause the calendar event to be generated and also present suggestion elements such as "and repeat every year" to encourage the user to select a suggestion element, thereby causing the calendar event to repeat every year. This reduces the expected number of interactions between the user and the automated assistant, as the user no longer needs to create that calendar event repeatedly every year. This can also allow the user to be informed that "and repeat every year" may be provided via voice in the future, thereby enabling the user to provide future verbal inputs that include "and repeat every year," thereby making such future verbal appointments more efficient. Furthermore, in some implementations, suggestion elements may be suggested as abridged versions of previous spoken utterances provided by the user. In this way, computing resources such as power and network bandwidth can be conserved by reducing the total amount of time the user interacts with their automated assistant.
[0013] In some implementations, suggestion elements presented to the user may be provided to avoid certain verbal utterances the user is already familiar with. For example, if a user has a history of verbal utterances such as "Assistant, dim my lights," the automated assistant may not offer the suggestion "my lights" in response to the user later saying "Assistant, dim...". Instead, the user may be presented with other suggestions that they may not be familiar with, such as "my monitor... my tablet screen... the amount of [color] in my lights." As an alternative or addition, the aforementioned suggestions may be presented to the user in situations where the user has provided an incomplete verbal utterance (e.g., "Assistant, play...") and in other situations where the user has provided a complete verbal utterance (e.g., "Assistant, play some music").
[0014] In various implementations, the user will no longer explicitly indicate when they consider their spoken utterance to be complete. For example, the user will not press the “Submit” button or say “Finish,” “Execute,” “Done,” or other concluding phrases when they consider their spoken utterance to be complete. Therefore, in these various implementations, graphical suggestions may be presented in response to spoken utterances already made, and the automated assistant will need to decide whether to give the user more time to provide further spoken utterances that fit one of the suggestions and are a continuation of the spoken utterances already provided (or to select one of the suggestions via touch input), or instead act only on the spoken utterances already provided.
[0015] In some of these implementations, the automated assistant will only act on already provided verbal utterances in response to detecting the duration of a lack of verbal input from the user. The duration can be determined dynamically in various implementations. For example, the duration may depend on whether the already provided verbal utterance is an "incomplete" or "complete" request, and / or on the characteristics of the "complete request" of the already provided verbal utterance. As mentioned above, a "complete request," when fully processed by the automated assistant, triggers the execution of one or more corresponding automated assistant actions, where the executed automated assistant actions are not default error type actions (e.g., not an audible response such as "Sorry, I can't help you"). An "incomplete request," when fully processed by the automated assistant, triggers the execution of a default error type action (e.g., an error sound and / or a default verbal response such as "Sorry, I can't help you").
[0016] In various implementations, the duration can be longer when an already made verbal utterance is an incomplete request compared to when it is a complete request. For example, the duration might be 5 seconds when an already made verbal utterance is an incomplete request, but shorter when an already made verbal utterance is a complete request. For example, the suggestions "a calendar entry for X:00 on [date]" and "a reminder to do X at [time] or [location]" may be presented in response to the verbal utterance "create a". Since the verbal utterance "create a" is incomplete, the automated assistant can wait 5 seconds for further verbal utterances, and only after 5 seconds have elapsed without further verbal input will the automated assistant fully process the incomplete verbal utterance and perform the default error type action. Also, for example, the suggestion "for X:00 on [date]" may be presented in response to the verbal utterance "create a calendar entry that says 'call Mom'". The verbal utterance, "Create a calendar entry that says 'Call Mom'," is complete (i.e., it will lead to the creation of a calendar entry with a title, and possibly further prompts for date and time details), so the automated assistant can wait for 3 seconds or other shortened duration for further verbal utterances. Only after 3 seconds have elapsed without further verbal input will the automated assistant fully process the verbal utterance already made, create the calendar entry, and possibly further prompt the user to provide the date and / or time for the calendar entry.
[0017] As another example, the duration may be shorter for a complete request that includes all required parameters compared to the duration for a complete request that does not include all required parameters (and would result in further prompting for required parameters). For example, a complete request to create a calendar entry may include the required parameters "Title," "Date," and "Time." The verbal utterance, "Create a calendar entry that says 'Call Mom at 2 PM today'," includes all required parameters, but nevertheless, suggestions for optional parameters (e.g., the suggestion "And repeat every day") may be provided. Because all required parameters are included in the verbal utterance, the duration may be shorter than, for example, if the verbal utterance did not include all required parameters (and would result in suggestions such as "@[Date] at X:00"), then the duration may be shorter. As yet another example, the duration may be shorter for a complete request that results in the invocation of a specific automated assistant agent rather than a general search agent that issues a general search based on the request and displays the search results. For example, the verbal utterance "What is time?" is a complete request that generates a general search and a provision of search results that provide information related to the concept of "time." The suggestion "In [city]" may appear in response to the verbal utterance and may appear within two seconds or other durations. On the other hand, the verbal utterance "Turn on the kitchen lights" is a complete request that invokes a specific automated assistant agent that causes the "kitchen lights" to be turned on. The suggestion "To X% brightness" may appear in response to the verbal utterance but may appear within only one second (or other shortened durations) because the verbal utterance already provided is complete and invokes a specific automated assistant agent (rather than causing a general search to be performed).
[0018] The above description is provided as an overview of some implementations of this disclosure. These implementations, and other implementations, are described in more detail below.
[0019] Other implementations may include a non-temporary computer-readable storage medium that stores instructions executable by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) in order to carry out one or more of the methods described above and / or elsewhere in this specification. Other implementations may also include a system of one or more computers and / or one or more robots, each including one or more processors operable to execute the stored instructions in order to carry out one or more of the methods described above and / or elsewhere in this specification.
[0020] Please understand that all combinations of the above concepts and additional concepts described in more detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter disclosed herein. [Brief explanation of the drawing]
[0021] [Figure 1A] This diagram shows a user providing verbal utterances, and an automated assistant providing suggestions to complete those utterances. [Figure 1B] This diagram shows a user providing verbal utterances, and an automated assistant providing suggestions to complete those utterances. [Figure 2A] This diagram shows a user receiving category suggestions in response to their spoken utterances and / or based on one or more actions of the automated assistant. [Figure 2B]A diagram showing that a user receives category suggestions in response to an oral utterance and / or based on one or more actions of an automatic assistant. [Figure 3] A diagram showing that instead of an audible prompt being rendered by an automatic assistant, a user receives suggestions for completing an oral utterance. [Figure 4] A diagram of a system that provides suggestions via a display modality and completes an oral utterance via an automatic assistant. [Figure 5] A diagram of a method for providing one or more suggestion elements for completing and / or supplementing an oral utterance provided by a user to an automatic assistant, a device, an application, and / or any other device or module. [Figure 6] A block diagram of an exemplary computer system.
Best Mode for Carrying Out the Invention
[0022] Figures 1A and 1B illustrate Figures 100 and 120, where a user provides a verbal utterance and an automated assistant provides suggestions to complete the verbal utterance. User 112 can provide a verbal utterance 116 to the automated assistant interface of computing device 114, but the verbal utterance 116 may be incomplete, at least to the automated assistant. For example, the user may provide a verbal utterance 116 such as, "Assistant, could you put it down...", which is processed in computing device 114 and / or a remote computing device such as a server device and may be graphically rendered on the display panel 110 of computing device 114. The graphical rendering of the verbal utterance 116 may be presented as a graphical element 106 in the assistant interactive module 104 of computing device 114. In response to receiving the verbal utterance 116, the computing device may determine that the verbal utterance 116 is an incomplete verbal utterance. An incomplete oral utterance may be one in which at least one action does not result in the action being performed by the automated assistant and / or any other application or assistant agent. Alternatively or additionally, an incomplete oral utterance may be one or more parameter values that are necessary for a function to be performed or otherwise controlled. For example, an incomplete oral utterance may be one in which an assistant function is necessary for it to be performed by the automated assistant and therefore requires further or supplementary input from the user or other source.
[0023] In response to a determination that the spoken utterance 116 is an incomplete spoken utterance, the automatic assistant can cause one or more suggestion elements 108 to be presented in the first user interface 102. However, in some implementations, one or more suggestion elements 108, and / or their content, may be rendered audibly to a user who has a visual impairment or otherwise cannot easily view the first user interface 102. The suggestion elements 108 may include natural language content that characterizes candidate text segments that the user 112 can then select, the selection being by tapping on the first user interface 102, orally repeating the text of the candidate text segment, and / or by explaining it when at least one suggestion element 108 is presented in the first user interface 102. For example, rather than providing an additional spoken utterance such as "sound", the user 112 can provide a different additional spoken utterance such as "the first one". Alternatively, if the user desired to select a different suggestion element 108, the user 112 could provide a touch input to the suggestion element 108 labeled "light", tap on the first user interface 102 at a location corresponding to the desired suggestion element 108, and / or provide a different additional spoken utterance such as "the second one". In various implementations, one or more of the suggestion elements 108 may be presented with icons and / or other graphical elements to indicate actions that will be performed in response to their selection. For example, "sound" may be presented with a speaker icon, "light" may be presented with a light icon, and / or "garage door" may be presented with a garage door icon.
[0024] The natural language content of the text segment provided with each suggestion element 108 may be generated based on whether the combination of the incomplete oral utterance 116 and the candidate text segment would cause the automated assistant to perform an action. For example, each candidate text segment may be generated based on whether the combination of the incomplete oral utterance 116 and the candidate text segment, when spoken together by user 112, would cause an action to be performed via the automated assistant. In some implementations, when user 112 provides a complete oral utterance, the automated assistant, computing device 114, and / or associated server device can bypass generating the suggestion element 108. Alternatively, when user 112 provides a complete oral utterance, the automated assistant, computing device 114, and / or associated server device can generate the suggestion element for provisioning in the first user interface 102.
[0025] In some implementations, device topology data can be used as a basis for generating candidate text segments for suggestion element 108. For example, a computing device 114 may have access to device topology data that identifies one or more devices associated with user 112. Furthermore, the device topology data may indicate the location of each of the one or more devices. User 112's location, along with prior authorization from user 112, may be used to determine user 112's context, at least because the context characterizes user 112's location. Device topology data may be compared with location data to determine devices that may be near user 112. Furthermore, once nearby devices are identified, the status of each nearby device may be determined to compare the natural language content of the incomplete oral utterance 116 with the status of each nearby device. For example, as shown in Figure 1A, user 112 has provided a request for the assistant to “put something down,” so devices with a status that can be modified through the “put something down” action may be used as a basis for generating suggestion element 108. Therefore, if a speaker is playing music nearby, a suggestion "sound" may be provided. Furthermore, if a light is on near the user, a suggestion "light" may be provided. Additionally or alternatively, device topology data can characterize devices that may not be near user 112 but may nevertheless be associated with an incomplete verbal utterance 116. For example, user 112 may be in their bedroom, but the incomplete verbal utterance 116 may be associated with a garage door in user 112's home garage. Therefore, the automated assistant may also provide a suggestion element 108 that allows the user to "lower the garage door" even though the user is not in the garage.
[0026] Figure 1B shows Figure 120 in which user 112 provides an additional oral utterance 128 to facilitate the selection of a suggestion element 108 provided in the first user interface 102 shown in Figure 1A. By providing an additional oral utterance 128, such as "garage door," user 112 can select one of the previously displayed suggestion elements 108. In response to the receipt of the oral utterance in the automated assistant interface of the computing device 114, candidate text corresponding to the selected suggestion element 108 may be presented adjacent to the natural language content of the incomplete oral utterance 116. For example, as shown in Figure 1B, the completed candidate text 122 may be presented in the second user interface 126 as a combination of the natural language content from the incomplete oral utterance 116 and the selected suggestion element 108 (e.g., the natural language content of the additional oral utterance 128).
[0027] The computing device 114, the automated assistant, and / or the server device can determine that an additional oral utterance 128 combined with the initial oral utterance 116 has resulted in a complete oral utterance. In response, one or more actions may be performed via the automated assistant. For example, when the completed oral utterance corresponds to a completed candidate text 122, the automated assistant may cause the garage door at user 112's home to be electromechanically lowered to the closed position. In some implementations, when a complete oral utterance is formed, and / or when an incomplete oral utterance is formed, additional suggestion elements 124 may be provided in the second user interface 126. For example, even though the completed candidate text 122 is actionable by the automated assistant, the computing device 114, the automated assistant, and / or the server device may cause the additional suggestion elements 124 to be presented in the second user interface 126. As described herein, in various implementations, if the duration of time elapses without any further verbal input or selection of any of the additional suggestion elements 124, the automated assistant may take action based on the completed candidate text 122 (i.e., cause the garage door to be lowered) without any selection of any of the additional suggestion elements 124. As also described herein, the duration of that time may be shorter than the duration of time during which the automated assistant would wait in Figure 1A. This shorter duration of time may be based, for example, on the completed candidate text 122 being a "complete request" while candidate text 106 (Figure 1A) is an "incomplete request".
[0028] In some implementations, a decision may be made as to whether there are any additional features or actions that can be performed and are related to the completed candidate text 122. Alternatively or additionally, to limit the amount of computing and / or network resources that may be spent in any subsequent amount of interaction between user 112 and the automated assistant, the decision may be based on whether the additional features or actions can reduce such interaction. For example, in some implementations, the content of the completed candidate text 122 may be compared with the content of previous interactions between user 112 and the automated assistant, and / or previous interactions between user 112 and the automated assistant, along with prior permission from user 112 and one or more other users. Furthermore, the interaction time between users and their respective automated assistants may be compared when participating in interactions related to the candidate text 122. For example, another user may only interact with another user's respective automated assistant by requesting that a recurring action be performed, such as requesting that an action be performed in response to daily activities. Therefore, the amount of time that one user spends interacting with another user's respective automated assistant may be relatively less than that of user 112, at least because the interaction relates to the subject of the candidate text 122 and does not require regular and / or conditional actions to be performed. Thus, the second user interface 126 may be generated with additional suggestion elements 124 so that the user can select one or more of the additional suggestion elements 124, thereby reducing the consumption of computing and network resources. For example, if user 112 selects the first additional suggestion element 124 by speaking the phrase 130, "Every morning, after I leave for work," it no longer needs to be displayed to user 112. Furthermore, the spoken utterance 128 no longer needs to be processed, and / or the completed candidate text 122 no longer needs to be generated for display on the computing device 114.Furthermore, user 112 may be aware of additional conditional statements, such as "every night at 9 PM" and "every night when my spouse and I are at home," which can be used to reduce the consumption of computing resources associated with other commands that the user frequently requests. For example, by seeing additional suggestion elements 124, user 112 may remember them and use them later when providing verbal utterances such as "Assistant, could you turn on the home security system?"
[0029] Figures 2A and 2B provide Figures 200 and 220 illustrating that user 214 receives category suggestions in response to an oral utterance 218 and / or based on one or more actions of the automated assistant. As provided in Figure 2A, user 214 can provide an oral utterance 218 to the automated assistant interface of computing device 216. The oral utterance 218 could be, for example, "Assistant, can you turn it on?" In response to receiving the oral utterance 218, computing device 216, a remote computing device communicating with computing device 216, and / or the automated assistant can cause a first user interface 204 to be presented on the display panel 212 of computing device 216. The first user interface 204 may include an assistant interactive module 206 that can present various graphical elements based on one or more actions of the automated assistant. For example, the interpretation of the oral utterance 218 may be presented in the first user interface 204 as an incomplete text segment 208.
[0030] The first user interface 204 may also include one or more suggestion elements 210, which may include content that characterizes candidate categories for additional input to complete an incomplete text segment 208. The candidate categories characterized by the content of the suggestion elements 210 may be based on data from one or more different sources. For example, the content of the suggestion element 210 may be based on the context in which the user provided the oral utterance 218. The context may be characterized by available contextual data in the computing device 216, remote devices (such as server devices) communicating with the computing device 216, an automated assistant, and / or any other application or device associated with the user 214. For example, device topology data may indicate that the computing device 216 is paired with or otherwise communicating with one or more other devices. Based on this device topology data, one or more of the suggestion elements 210 may provide a candidate category such as "[device name]". Alternatively or additionally, data from one or more sources may include application data indicating media recently accessed by user 214. When recent media includes movies, the suggestion element 210 may be generated to include content such as "[Movie Title]". Alternatively or additionally, when recent media includes music, the suggestion element 210 may be generated to include content such as "[Song and Artist Name]".
[0031] Figure 2B shows Figure 220 in which user 214 provides an additional verbal utterance 228 to select a candidate category from the suggestion element 210 provided in Figure 2A. As shown, user 214 may provide the additional verbal utterance 228, "kitchen lights". In response to receiving the additional verbal utterance 228, an additional suggestion element 226 may be presented to user 214 via the display panel 212. The additional suggestion element 226 may be presented to user 214 in response to the user completing the initial verbal utterance 218 by providing the additional verbal utterance 228, or when the additional verbal utterance 228 does not complete the initial verbal utterance 218. In some implementations, the additional suggestion element 226 may contain content such as candidate categories, which may be in a similar format to the suggestion element 210. As an addition or alternative, the additional suggestion element 226 may include content such that, when the user speaks, the content of the voice may be used to facilitate the completion, modification, and / or initialization of one or more actions via the automated assistant. Thus, the assistant interactive module 222 may provide various formats for the suggestion text segment, which may include a suggestion category (e.g., provided in suggestion element 210), a suggestion request (e.g., provided in the additional suggestion element 226), and / or any other content that may be used to include in commands provided to the automated assistant and / or other applications or devices.
[0032] In some implementations, the provisioning of one or more suggestion elements, and / or the order in which one or more suggestion elements are presented, may be based on a priority assigned to certain content, actions, verbal utterances, candidate text segments, and / or any other data from which the suggestion may be based. For example, in some implementations, the provisioning of certain content within suggestion element 226 may be based on historical data characterizing previous interactions between user 214 and the automated assistant, between user 214 and the computing device 216, between user 214 and one or more other devices, between user 214 and one or more other applications, and / or between one or more other users and one or more devices and / or applications. For example, in response to receiving a selection of suggestion element 210 with the content "[device name]", the computing device 216 and / or the server device may access historical data characterizing previous interactions between user 214 and one or more devices associated with user 214. History data can be used as a basis for generating content for suggestion element 226. For example, even if the user selects "kitchen lights," history data can be used to generate suggestions about other devices the user might be interested in controlling, and / or other settings (other than "on") the user might be interested in modifying for devices not explicitly selected by the user.
[0033] As an example, and with prior permission from user 214, the history data may characterize one or more previous interactions in which user 214 turned off the lights in their kitchen a few minutes before going to bed. Based on these interactions characterized by the history data, the automated assistant may assign the highest priority to such interactions because they relate to a complete request 224 such as "Assistant, could you turn on the kitchen lights?". Furthermore, the automated assistant may generate content for suggestion elements 226, such as "until I go to sleep". Based on the content of the history data characterizing the highest priority interactions, the automated assistant may provide the suggestion element 226 "until I go to sleep" as the highest priority or highest-positioned suggestion element 226 in the second user interface 232, since at least its content relates to a completed request 224.
[0034] Alternatively or additionally, if the history data characterizes one or more previous interactions in which user 214 requested the automated assistant to set the kitchen light to half brightness or red, but those requests occurred less frequently than the aforementioned request to turn off the light before going to bed, the content of those requests may be assigned a priority lower than the priority previously assigned to the content of the aforementioned requests. In other words, because user 214 requested more frequently that the kitchen light be turned off when going to bed, the suggestion element "until I go to bed" will take precedence over and / or have a higher position than other suggestion elements such as "and at half brightness" and / or "and to red." In some implementations, the content of a suggestion element may be based on device or application features that user 214 has not used before, and / or one or more updates pushed to the application or device by a third-party entity. The term "third-party entity" may refer to an entity distinct from the manufacturer of the computing device 216, the automated assistant, and / or the server device communicating with the computing device 216. Therefore, when the manufacturer of the kitchen lights (i.e., the third-party entity) pushes an update that enables the kitchen lights to operate according to a red setting, the automated assistant can access the data characterizing this update and, based on the update provided by the third-party manufacturer (i.e., the third-party entity), generate a suggestion element 226 that reads "and to red setting."
[0035] As an alternative or addition, the content of a suggestion element 226 based on an update may be assigned a priority lower than another priority assigned to the content corresponding to another suggestion element 226 based on one or more actions that user 214 frequently requests to be performed. However, the content based on an update may be assigned a higher priority than content that user 214 has previously presented via the display panel 212 and / or via the automated assistant, but which user 214 has chosen not to select. In this way, the suggestion elements provided to the user may be cycled according to their assigned priorities in order to adapt the suggestions to the user. Furthermore, when those suggestions relate to reducing interaction X between user 214 and the automated assistant, computing and network resources may be protected by providing more suggestions that reduce such interaction, while limiting suggestions that do not reduce such interaction or are otherwise irrelevant to the user.
[0036] Figure 3 shows Figure 300 in which, instead of audible prompts being rendered by the automated assistant, user 312 receives suggestions to complete an oral utterance 316. User 312 may provide an oral utterance 316 such as, "Assistant, can you create a calendar entry for me?" In response, an assistant interactive module 304, which may be controlled via the automated assistant, may provide a graphical representation 306 of the oral utterance 316, which may be presented on the display panel 320 of the computing device 314.
[0037] Oral utterances 316 may respond to requests for a function to be performed, and that function may have one or more required parameters that must be specified by the user 312 or otherwise satisfied by one or more available values. To prompt the user to provide one or more values to satisfy one or more required parameters, the automated assistant may, in some cases, provide a graphical prompt 308 requesting values for specific parameters. For example, prompt 308 may contain content such as "What will the title be?", thereby indicating that the user 312 must provide a title for a calendar entry. However, in order to protect the power and computing resources of the computing device 314, the automated assistant may cause the computing device 314 to bypass rendering an audible prompt for the user 312. Furthermore, this can eliminate latency because the user 312 does not have to wait until the audible prompt is fully rendered.
[0038] Instead of the automated assistant providing audible prompts, it may generate one or more suggestion elements 310 that can be obtained based on one or more required parameters for a function requested by user 312. For example, if the function corresponds to generating a calendar entry and the parameter to be satisfied is the "title" of the calendar entry, then the suggestion element 310 may include title suggestions for the calendar entry. For example, the content of the suggestion may include "for my mother's birthday," "to take out the trash," and "to pick up the dry cleaning tomorrow."
[0039] To select one of the suggestion elements 310, user 312 may provide an additional verbal utterance 318 that identifies the content of at least one suggestion element 310. For example, user 312 may provide the additional verbal utterance 318, "To pick up the dry cleaning tomorrow." As a result, and from the perspective of the computing device 314 and / or the automated assistant, only the dialogue rendered during the interaction between user 312 and the automated assistant will be rendered by user 312. Thus, the rendering of audio will not be necessary or, otherwise, will not be employed by the automated assistant. Furthermore, in response to the user providing an additional verbal utterance 318, different prompts may be provided on the display panel 320 to request user 312 to provide values for other parameters if necessary for the requested function. Furthermore, when another parameter is required for the function, different suggestion elements may be presented on the display panel 320 with different content. Furthermore, subsequent prompts and suggestion elements may be presented without any audio being rendered on the computing device 314 or otherwise via the automated assistant. Thus, user 312 will, superficially, be providing ongoing verbal utterances while simultaneously viewing the display panel 320 to identify content that further completes the original function requested by the user.
[0040] Figure 4 shows a system 400 that provides suggestions for completing spoken utterances via a display modality through an automated assistant 408. In some implementations, suggestions may be provided to the user to complete a spoken utterance or at least supplement the user's previous spoken utterance. In some implementations, suggestions may be provided to reduce the frequency and / or length of time the user will be required to participate in subsequent dialogue sessions with the automated assistant.
[0041] The automated assistant 408 can operate as part of an automated assistant application provided on one or more computing devices, such as a client device 434 and / or a server device 402. A user can interact with the automated assistant 408 through one or more assistant interfaces 436, which may include one or more microphones, cameras, touchscreen displays, user interfaces, and / or any other devices capable of providing an interface between the user and the application. For example, a user can initialize the automated assistant 408 by providing verbal, text, and / or graphical input to the assistant interface, causing the automated assistant 408 to perform functions (e.g., providing data, controlling a device (e.g., an IoT device 442), accessing an agent, modifying settings, controlling an application, etc.). The client device 434 and / or the IoT device 442 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures, enabling a user to control applications on the client device 434 and / or the server device 402 via the touch interface. Touch input and / or other gestures (e.g., verbal utterances) can also enable the user to interact with the automated assistant 408, automated assistant 438, and / or IoT device 442 via client device 434.
[0042] In some implementations, the client device 434 and / or IoT device 442 may lack a display device but include an audio interface (e.g., a speaker and / or microphone) to provide an audible user interface output without providing a graphical user interface output, as well as a user interface such as a microphone to receive oral natural language input from the user. For example, in some implementations, the IoT device 442 may include one or more tactile input interfaces, such as one or more buttons, and omit a display panel from which graphical data from a graphics processing unit (GPU) would be provided. In this way, energy and processing resources may be saved compared to a computing device that includes a display panel and a GPU.
[0043] The client device 434 may be communicating with the server device 402 over a network 446, such as the Internet. The client device 434 can offload computing tasks to the server device 402 to protect computing resources on the client device 434 and / or the IoT device 442. For example, the server device 402 may host an automated assistant 408, and the client device 434 may transmit inputs received at one or more assistant interfaces and / or the user interface 444 of the IoT device 442 to the server device 402. However, in some implementations, the automated assistant 408 may be hosted on the client device 434. In various implementations, all or less of the embodiments of the automated assistant 408 may be implemented on the server device 402 and / or the client device 434. In some of these implementations, embodiments of the automated assistant 408 are implemented via the local automated assistant 438 on the client device 434, which interfaces with the server device 402, which can implement other embodiments of the automated assistant 408. The server device 402 may, in some cases, serve multiple users and their associated assistant applications via multiple threads. In implementations in which all or less of the Auto Assistant 408 are implemented via the local Auto Assistant 438 of the client device 434, the local Auto Assistant 438 may be an application separate from the operating system of the client device 434 (for example, installed "on top of" the operating system)—or alternatively, it may be implemented directly by the operating system of the client device 434 (for example, an application of the operating system, but considered to be integrated with the operating system).
[0044] In some implementations, the automated assistant 408 and / or the automated assistant 438 may include an input processing engine 412, which may employ multiple different engines to process inputs and / or outputs for the client device 434, the server device 402, and / or one or more IoT devices 442. For example, the input processing engine 412 may include a speech processing engine 414 that can process audio data received at the assistant interface 436 to identify text embedded in the audio data. The audio data may be sent from the client device 434 to the server device 402, for example, to conserve computing resources on the client device 434.
[0045] The process for converting audio data to text may include a speech recognition algorithm, which may employ a neural network, a word2vec algorithm, and / or a statistical model for identifying groups of audio data corresponding to words or phrases. The text converted from the audio data may be parsed by a data parsing engine 416 and made available to the automated assistant 408 as text data that can be used to generate and / or identify command phrases from the user. In some implementations, the output data provided by the data parsing engine 416 may be provided to an action engine 418 to determine whether the user has provided input corresponding to specific actions and / or routines that can be performed by the automated assistant 408, and / or applications, agents, and / or devices that can be accessed by the automated assistant 408. For example, assistant data 422 may be stored as client data 440 in the server device 402 and / or client device 434, and may include data defining one or more actions that can be performed by the automated assistant 408, as well as parameters related to performing those actions.
[0046] When the input processing engine 412 determines that the user has requested that a specific action or routine be performed, the action engine 418 may determine one or more parameters for that specific action or routine, and then the output generation engine 420 may provide output to the user based on the specific action, routine, and / or one or more parameters. For example, in some implementations, in response to user input such as a gesture directed to the assistant interface 436, the automated assistant 438 may cause data characterizing the gesture to be sent to the server device 402 to determine the action that the user intends the automated assistant 408 and / or automated assistant 438 to perform.
[0047] The client device 434 and / or IoT device 442 may each include one or more sensors capable of providing an output in response to user input from a user. For example, one or more sensors may include an audio sensor (i.e., an audio response device) that provides an output signal in response to audio input from a user. The output signal may be transmitted to the client device 434, server device 402, and IoT device 442 via a communication protocol, such as Bluetooth, LTE, Wi-Fi, and / or any other communication protocol. In some implementations, the output signal from the client device 434 may be converted into user input data by one or more processors in the client device 434. The user input data may be transmitted to the IoT device 442 (e.g., a Wi-Fi enabled light bulb) via a communication protocol and processed in the client device 434, or / or transmitted to the server device 402 via the network 446. For example, the client device 434 and / or server device 402 may determine, based on the user input data, the function of the IoT device 442 that will be controlled via the user input data.
[0048] The user may provide input, such as spoken utterances, which are received at the assistant interface 436 and provided by the user to facilitate control of the IoT device 442, client device 434, server device 402, the automated assistant, the agent assistant, and / or any other device or application. In some implementations, the system 400 may provide the user with suggestions on how to complete the spoken utterances, supplement them, and / or otherwise assist the user. For example, the user may provide incomplete spoken utterances along with the intention to cause the IoT device 442 to perform a particular function. The incomplete spoken utterances may be received at the assistant interface 436 and converted into audio data, which may be sent to the server device 402 for processing. The server device 402 may include a suggestion engine 424 that receives the audio data or other data based on the incomplete spoken utterances and can determine one or more suggestions for completing or supplementing the incomplete spoken utterances.
[0049] In some implementations, the timing for providing one or more suggestion elements, as well as the display panels of the client device 434 and / or IoT device 442, may be determined by the suggestion engine 424. The suggestion engine 424 may include a timing engine 428, which can determine when one or more suggestions for additional verbal utterances should be presented to the user. In some implementations, one or more thresholds may be characterized by action data 444 available to the suggestion engine 424. For example, the action data 444 may characterize thresholds corresponding to the amount of time that may be delayed after the user has provided a portion of a verbal utterance. If the user delays after providing an incomplete verbal utterance, the amount of delay time may be compared to the threshold characterized by the action data 444. When the delay satisfies the threshold, the timing engine 428 may cause one or more suggestion elements to be presented to the user to complete the incomplete verbal utterance.
[0050] In some implementations, the timing engine 428 can adapt a value for the delay time between the user providing an incomplete oral utterance and the user receiving a suggestion to complete the oral utterance. For example, the suggestion engine 424 can determine the user's context based on context data incorporated in the assistant data 422, action data 444, and / or client data 440, and use the context as the basis for determining the value for the delay time. For example, when the user is in a first context characterized by being at home and watching TV, the timing engine 428 can set the delay time to a first value. However, when the user is in a second context characterized by being driving and interacting with a vehicle computing device, the timing engine 428 can set the delay time to a second value greater than the first value. In this way, suggestion elements for completing oral utterances can be delayed relatively to limit distractions to the user while driving.
[0051] In some implementations, the value for the delay time may be proportional to, or adapted to, how frequently the user provides incomplete verbal utterances over time. For example, when the user first begins interacting with the automated assistant, the user may provide X number of incomplete verbal utterances over a period of Y. Some time later, the user may have a history of providing Z number of incomplete utterances Y. If Z is less than X, the timing engine 428 can increase the value for the delay time threshold over time. Thus, as the delay time threshold increases, computing resources on the server device 402 and / or client device 434 can be conserved because there are fewer changes to what is presented on the display panel over time, thereby reducing the load on the GPU responsible for processing the graphical data for the display panel.
[0052] In some implementations, the suggestion engine 424 may include a speech-to-text modeling engine 430, which can select a speech-to-text model from several different speech-to-text models according to the expected oral utterance from the user. Alternatively, the speech-to-text modeling engine 430 may select a speech-to-text model from several different speech-to-text models according to the expected type of oral utterance from the user. Alternatively, the speech-to-text modeling engine 430 may bias the selected speech-to-text model based on the expected oral utterance from the user. In some implementations, when the suggestion element provided to the user includes content with at least some numerical text, the speech-to-text modeling engine 430 may select a speech-to-text model adapted to process oral utterances characterizing numerical information, based on the content. As an alternative or addition, when the suggestion elements provided to the user include content containing at least proper nouns, such as the names of cities or people, the speech-to-text model engine 430 may bias the selected speech-to-text model to more easily interpret the proper nouns. For example, the selected speech-to-text model may be biased to reduce the error rate that would otherwise result from processing proper nouns. This can eliminate latency that would otherwise occur and / or reduce the frequency of irrelevant suggestions from the suggestion engine 424.
[0053] In some implementations, the content of each suggestion among one or more suggestion elements provided to the user in response to an incomplete or complete oral utterance may be based on the user's context. For example, the user's context may be characterized by data generated by the device topology engine 426. The device topology engine 426 can identify the device that received the incomplete or complete oral utterance from the user and determine information associated with that device, such as other devices connected to the receiving device, the location of the receiving device relative to the other devices, and an identifier for the receiving device, one or more different users associated with the receiving device, the functionality of the receiving device, the functionality of one or more other devices paired with or otherwise communicating with the receiving device, and / or any other information that may be associated with the device that received the oral utterance. For example, if the receiving device is a standalone speaker device located in the user's living room, the device topology engine 426 can determine that the standalone speaker device is in the living room and can also determine other devices located in the living room.
[0054] The suggestion engine 424 can use information such as descriptions of other devices in the living room to generate suggestion elements that will be presented to the user. For example, in response to receiving an incomplete oral utterance from the user via a standalone speaker device, the device topology engine 426 may determine that a television is in the living room together with the standalone speaker device. Based on the determination that the television is in the same room, the suggestion engine 424 can generate content for a suggestion element to complete the incomplete oral utterance. For example, if the incomplete oral utterance is "Assistant, change it," the content of the suggestion element may include "TV channel." Alternatively or additionally, the device topology engine 426 may also determine that the television is in the same room as the standalone speaker device and that the standalone speaker device is not capable of graphically representing the suggestion element, but the television is capable of graphically representing the suggestion element. Therefore, in response to an incomplete verbal utterance, the suggestion engine 424 can cause the television display panel to present one or more suggestion elements, including content such as "Change my TV channel." When the user repeats the content of the suggestion element, "Change my TV channel," the automated assistant can cause the TV channel to change, while also generating additional suggestion elements that may be presented simultaneously with the rendering of the new TV channel on the television.
[0055] In some implementations, data such as assistant data 422, action data 444, and / or client data 440 can characterize past interactions between the user and the automated assistant 408 and / or the automated assistant 438. The suggestion engine 424 can use the aforementioned data to generate content for suggestion elements that are presented to the user in response to the provision of incomplete or complete oral utterances to the assistant interface 436. In this way, if the user temporarily forgets some preferred commands, natural language text may be provided via the suggestion elements to remind the user of certain preferred commands. Such preferred commands may be identified by comparing the content with previously received incomplete or complete oral utterances and by comparing the content with previous natural language inputs received from the user. The most frequently occurring natural language inputs may be considered to correspond to preferred commands. Thus, when the content of complete or incomplete oral utterances corresponds to such natural language inputs, the suggestion elements may be rendered and / or generated based on the content of those natural language inputs. Furthermore, by identifying the most frequently processed natural language inputs and / or identifying preferred commands for a particular user, the suggestion engine 424 can provide the user, via the automated assistant, with suggestions based on the most frequently and successfully executed commands. As a result of providing such suggestions, the processing of the least successful commands is reduced, thus conserving computing resources.
[0056] Figure 5 illustrates a method 500 for providing one or more suggestion elements to complete and / or supplement an oral utterance provided by a user to an automated assistant, device, application, and / or any other device or module. Method 500 may be implemented by one or more computing devices, applications, and / or any other device or module that may be associated with an automated assistant. Method 500 may include an action 502 that determines that the user has provided at least a portion of an oral utterance. The oral utterance may be oral natural language input provided audibly by the user to one or more automated assistant interfaces of one or more computing devices. The oral utterance may consist of one or more words, one or more phrases, or any other components of oral natural language. For example, the automated assistant interface may include a microphone, and the computing device may be a smart device including a touch display panel that can also act as another automated assistant interface. In some implementations, the computing device, or a server device communicating with the computing device, may determine that the user has provided an oral utterance.
[0057] Method 500 may further include an action 504 that determines whether a spoken utterance is complete. Determining whether a spoken utterance is complete may include determining whether each of one or more parameters of a function has an assigned value. For example, when the spoken utterance is “Assistant, lower,” the user may be requesting that an action be performed to lower the output modality of the device. The action may correspond to a function such as “Lower Volume” or “Lower Brightness,” which may require the user to specify a device name. Alternatively or additionally, determining whether a spoken utterance is complete may include determining whether the automated assistant can perform an action in response to receiving the spoken utterance.
[0058] In some cases, operation 504 may include determining whether certain categories of actions can be performed in response to a verbal utterance in order to determine whether the verbal utterance is complete. For example, an action such as asking the automated assistant to repeat what the user said for clarity may not indicate that the user provided a complete verbal utterance, even though the automated assistant is responding to the verbal utterance. However, actions such as performing a web query, controlling another device, controlling an application, and / or requesting the automated assistant to perform an intended action may be considered a complete verbal utterance. For example, in some implementations, a verbal utterance may be considered complete if at least one request can be identified from the verbal utterance and the request can be performed through the automated assistant.
[0059] When it is determined that the oral utterance is complete (for example, "Assistant, turn down the TV"), method 500 may proceed to optional action 512, which renders and / or generates suggestions for the user based on the actions to be performed, and / or action 514, which performs the actions corresponding to the oral utterance. For example, a computing device that receives the oral utterance via an automated assistant interface may initialize the performance of one or more actions specified via the oral utterance. However, when it is determined that the oral utterance is incomplete, method 500 may proceed to action 506.
[0060] Operation 506 may be an optional operation that includes determining whether the delay threshold duration occurred after an incomplete spoken utterance. For example, whether the duration of spoken input silence meets the threshold may be determined based on the user providing the spoken utterance, the user's context, and / or any other information associated with the incomplete spoken utterance, along with permission from the user. If the amount of spoken input silence meets the delay threshold duration, method 500 may proceed to operation 508. However, if the amount of spoken input silence does not meet the delay threshold duration, method 500 may proceed to operation 510.
[0061] Method 500 may, in some cases, proceed from operation 504 to operation 508, or in some cases, from operation 504 to operation 506. In operation 506, when it is determined that the delay threshold duration is met, Method 500 may proceed to operation 508, which renders one or more suggestions for completing the oral utterance. The content of one or more suggestions for completing the oral utterance may be based on the user who provided the incomplete oral utterance, the content of the incomplete oral utterance, the context in which the oral utterance was provided, data representing the device topology associated with the device that received the oral utterance, dialogue history data characterizing previous interactions between the user and the automated assistant (or one or more users and the automated assistant), and / or any other information from which suggestions may arise. For example, when the incomplete verbal utterance is "Assistant, lower," the content of the suggestion provided to the user may include "television," which may be based on device topology data indicating that the television and / or the user are located in the same room as the computing device that received the incomplete verbal utterance. Alternatively, when the incomplete verbal utterance is "Assistant, send a message to ~," the content of the suggestion provided to the user may be based on dialogue history data that characterizes a previous dialogue session in which the user requested the automated assistant to send a message to the user's spouse and siblings. Thus, the content of the suggestion may include "my sibling" and "my wife." In some implementations, the content of one or more suggestion elements may be determined and / or generated after action 502 but before optional action 506 and before action 508.
[0062] Method 500 may proceed from action 506 or action 508 to action 510, which determines whether another oral utterance has been provided by the user. For example, the user may provide an additional oral utterance after providing an incomplete oral utterance. In such a case, Method 500 may proceed from action 510 back to action 504. For example, the user may provide an additional oral utterance corresponding to the content of a suggestion element, thereby indicating to the automated assistant that the user has selected a suggestion element. This may result in an oral utterance edit, which may consist of the initial incomplete oral utterance and the additional oral utterance corresponding to the content of the suggestion element selected by the user. This oral utterance edit may undergo action 504, which determines whether the oral utterance edit corresponds to a complete oral utterance. In some implementations, in operation 510, the system may wait for another oral utterance for the duration, and if no other oral utterance is received during the duration, it may proceed to operation 514 (or 512). If another oral utterance is received during the duration, the system may proceed to return to operation 504. In some of these implementations, the duration may be determined dynamically. For example, the duration may depend on whether the oral utterance already provided is an "incomplete" request or a "complete request," and / or on the characteristics of the "complete request" of the oral utterance already provided.
[0063] When the edited oral utterance corresponds to a complete oral utterance, method 500 may proceed to optional actions 512 and / or 514. However, if the edited oral utterance is not determined to be a complete oral utterance, method 500 may proceed to action 506 and / or action 508. For example, when method 500 again proceeds to action 508, suggestions for completing the oral utterance may be based at least on the content of the edited oral utterance. For example, if the user selects the suggestion "my brother", the combination of the incomplete oral utterance "assistant, send a message to ~" and "my brother" may be used as the basis for further suggestions provided. In this way, a cycle of suggestions may be provided as the user continues to select suggestions for completing each of their oral utterances. For example, subsequent suggestions might include "I'm on my way" based on conversation history data indicating that the user has previously sent the same message to the user through an application separate from the automated assistant.
[0064] In some implementations, method 500 may proceed from operation 510 to an optional operation 512 that renders suggestions for the user based on the action to be performed. For example, when the automated assistant is requested to perform an action such as turning on an alarm system for the user, the action, and / or any information associated with the action, may be used as a basis for providing additional suggestion elements for the user to choose from. For example, if the complete oral utterance to be performed through the automated assistant is "Assistant, turn on my alarm system," then one or more other suggestion elements may be rendered in operation 512 and include content such as "And set my alarm for ~," "Turn off all my lights," "And turn on my audiobook," and / or any other suitable suggestion content for the user. In some implementations, in operation 512, the system proceeds to operation 514 in response to the reception of further oral input or in response to the non-reception of other oral utterances in the duration. In some of those implementations, the duration may be determined dynamically. For example, the duration may depend on whether the oral utterance already provided constitutes an "incomplete" request or a "complete request," and / or on the characteristics of the "complete request" of the oral utterance already provided.
[0065] In some implementations, suggestions may be rendered and / or generated to reduce the amount of interaction between the user and the automated assistant, thereby saving power, computing resources, and / or network resources. For example, the content of a supplemental suggestion element rendered in operation 512 may include "Every night, after 10 p.m." In this way, if it is determined that the user has selected the supplemental suggestion element (e.g., via another verbal utterance), as determined in operation 510, method 500 can proceed to operation 504, and then to operation 514, causing the automated assistant to generate a setting that the alarm system will be turned on every night, after 10 p.m. Thus, during the following nights, the user does not need to provide the same initial verbal utterance, "Assistant, turn it on...", but rather can simply rely on the setting adopted via any other suggestion element selected.
[0066] When a user has not provided any additional verbal utterances after the initial complete verbal utterance, or has indicated a preference for one or more commands corresponding to an edited verbal utterance to be executed, method 500 may proceed from action 510 to optional actions 512 and / or action 514. For example, when a user provides a verbal utterance such as "Assistant, change it" and later selects a suggestion element containing content such as "TV channel", the user may choose for the request to be executed even though they have not specified a particular channel to change the TV. As a result and action 514, the TV channel may be changed even though the user has not selected any supplementary suggestion elements to further narrow the resulting instruction in the command (for example, the channel change may be based on the user's learned preferences derived from dialogue history data).
[0067] Figure 6 is a block diagram of an exemplary computer system 610. The computer system 610 typically includes at least one processor 614 that communicates with several peripheral devices via a bus subsystem 612. These peripheral devices may include, for example, a storage subsystem 624 including memory 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices enable user interaction with the computer system 610. The network interface subsystem 616 provides an interface to an external network and is coupled to a corresponding interface device in another computer system.
[0068] The user interface input device 622 may include a keyboard, and pointing devices such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen integrated into a display, and audio input devices such as a speech recognition system or microphone, and / or other types of input devices. Generally, the use of the term “input device” shall include all possible types of devices and methods for inputting information into the computer system 610 or onto a communication network.
[0069] The user interface output device 620 may include non-visual displays such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include flat panel devices such as cathode ray tubes (CRTs), liquid crystal displays (LCDs), projection devices, or any other mechanism for creating visible images. The display subsystem may also provide non-visual displays, such as via an audio output device. In general, the use of the term “output device” shall include all possible types of devices and methods for outputting information from the computer system 610 to a user or to another machine or computer system.
[0070] The storage subsystem 624 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 624 may include logic for implementing selected embodiments of Method 500 and / or of Computing Device 114, Server Device, Computing Device 216, Assistant Interaction Module, Client Device 434, Computing Device 314, Server Device 402, IoT Device 442, Auto Assistant 408, Auto Assistant 438, Auto Assistant, and / or any other devices, apparatus, applications, and / or modules described herein.
[0071] These software modules are generally executed by processor 614 alone or in combination with other processors. The memory 625 used within the storage subsystem 624 may include several memories, including main random access memory (RAM) 630 for storing instructions and data during program execution, and read-only memory (ROM) 632 for storing fixed instructions. The file storage subsystem 626 can provide persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing several functional implementations may be stored by the file storage subsystem 626 within the storage subsystem 624, or in other machines accessible by processor 614.
[0072] The bus subsystem 612 provides a mechanism for various components and subsystems of the computer system 610 to communicate with each other as intended. Although the bus subsystem 612 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0073] Computer system 610 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing systems or computing devices. Due to the ever-changing nature of computers and networks, the description of computer system 610 shown in Figure 6 is merely a specific example illustrating several possible implementations. Numerous other configurations of computer system 610 are possible, having more or fewer components than the computer system shown in Figure 6.
[0074] Where the systems described herein may collect or use personal information about a user (or, more often referred to herein as “Participant”), the user may be provided with the opportunity to control whether the program or feature collects user information (for example, information about the user’s social networks, social actions or activities, occupation, user preferences, or the user’s current geographical location), or whether and / or how content from a content server that may be more relevant to the user is received. Furthermore, some data may be handled in one or more ways before being stored or used so that personally identifiable information is removed. For example, a user’s identity may be handled in such a way that personally identifiable information cannot be determined about the user, or a user’s geographical location may be generalized, in which case the geographical location information is obtained in such a way that the user’s specific geographical location cannot be determined (e.g., down to the city, zip code, or state level). Thus, a user may have control over how information is collected and / or used about them.
[0075] In some implementations, a method is provided, performed by one or more processors, which includes performing speech-to-text processing on data characterizing a user-provided oral utterance. The oral utterance includes natural language content and is received via an automated assistant interface of a computing device connected to a display panel. The method determines whether the oral utterance is complete or incomplete based on performing speech-to-text processing on the data characterizing the oral utterance, and further includes determining whether the automated assistant can be prompted to perform one or more actions based on the natural language content. When the oral utterance is determined to be incomplete, the method further includes causing the display panel of the computing device to provide one or more suggestion elements in response to the determination that the oral utterance is incomplete. The one or more suggestion elements include specific suggestion elements that, when spoken by the user to the automated assistant interface, provide via the display panel other natural language content that causes the automated assistant to act to facilitate the completion of an action. The method further includes determining, after the display panel of a computing device provides one or more suggestion elements, that the user has provided another spoken utterance associated with other natural language content of a particular suggestion element. The method further includes determining, in response to the determination that the user has provided another spoken utterance, whether the combination of the spoken utterance and the other spoken utterance is complete. When it is determined that the combination of the spoken utterance and the other spoken utterance is complete, the method further includes causing one or more actions to be performed via an automated assistant based on the natural language content and the other spoken utterance.
[0076] These and other implementations of the technologies disclosed herein may include one or more of the following features:
[0077] In some implementations, performing speech-to-text processing on data characterizing user-provided oral utterances includes generating one or more candidate text segments from the data, and the method further includes causing the display panel of a computing device to provide a graphical representation of at least one of the one or more candidate text segments when the oral utterance is determined to be incomplete. In some versions of those implementations, the method further includes causing the display panel of a computing device to bypass providing a graphical representation of at least one of the one or more candidate text segments and to bypass providing one or more suggestion elements when the oral utterance is determined to be complete and to include all necessary parameters for one or more actions. In some other versions of these implementations, the method further includes identifying a natural language command, generating one or more other suggestion elements that characterize the natural language command based on the identification of the natural language command, when spoken to an automated assistant interface by one or more users, resulting in a reduced amount of speech-to-text processing compared to the amount of speech-to-text processing associated with the spoken utterance from the user, when the spoken utterance is determined to be complete, causing the automated assistant to act to facilitate the completion of the action, and causing the display panel of the computing device to provide one or more other suggestion elements in at least response to the determination that the spoken utterance is complete.
[0078] In some implementations, the method further includes selecting a speech-to-text processing model from several different speech-to-text processing models based on the type of speech utterance expected from the user, when it is determined that the oral utterance is incomplete, and when at least one suggestion element is provided via the display panel.
[0079] In some implementations, the method further includes biasing the speech-to-text processing towards one or more terms of one or more suggestion elements and / or towards one or more expected types of content corresponding to one or more suggestion elements when the oral utterance is determined to be incomplete.
[0080] In some implementations, the method further includes determining whether a threshold duration of oral input silence followed an oral utterance. In some of these implementations, the display panel of the computing device provides one or more suggestion elements, which further responds to the determination that at least a threshold duration of oral input silence followed an oral utterance.
[0081] In some implementations, performing speech-to-text processing involves determining a first candidate text segment and a second candidate text segment, where the first and second candidate text segments correspond to different interpretations of the spoken utterance. In some of these implementations, a specific suggestion element is determined based on the first candidate text segment, and at least one other suggestion element is determined based on the second candidate text segment. In some versions of these implementations, causing a display panel connected to a computing device to provide one or more suggestion elements involves causing the display panel connected to the computing device to graphically represent the first candidate text segment adjacent to a specific suggestion element, and to graphically represent the second candidate text segment adjacent to at least one other suggestion element. In some of those versions, determining that a user has provided another spoken utterance associated with other natural language content of a particular suggestion element involves determining, based on that other spoken utterance, whether the user identified a first suggestion text segment or a second suggestion text segment.
[0082] In some implementations, the method further includes, when it is determined that an oral utterance is incomplete, in response to the determination that the oral utterance is incomplete, generating other natural language content based on historical data characterizing one or more previous dialogues in which a previous action was achieved via an automated assistant and the user identified at least a portion of other natural language content during one or more previous dialogues between the user and the automated assistant.
[0083] In some implementations, the method, when it is determined that an oral utterance is incomplete, responds to the determination that the oral utterance is incomplete by generating other natural language content based on device topology data that characterizes the relationships between various devices associated with the user, further comprising certain suggestion elements that are based on device topology data and identify one or more devices among the various devices associated with the user.
[0084] In some implementations, controlling a device involves causing one or more actions to be performed via an automated assistant based on natural language content and other spoken utterances.
[0085] In some implementations, a particular suggestion element may further include a graphical element indicating an action.
[0086] In some implementations, the method, when it is determined that an oral utterance is incomplete, determines a specific duration for waiting for another oral utterance after provisioning one or more suggestion elements, wherein the specific duration is determined on the basis that the oral utterance is incomplete. In some versions of those implementations, the method, when it is determined that an oral utterance is complete, further includes generating one or more other suggestion elements based on a complete oral utterance, causing the display panel of a computing device to provide one or more other suggestion elements, and determining an alternative specific duration for waiting for further oral utterances after provisioning one or more other suggestion elements. The alternative specific duration is shorter than the specific duration, and the alternative specific duration is determined on the basis that the oral utterance is complete.
[0087] In some implementations, a method is provided, performed by one or more processors, which includes performing speech-to-text processing on data characterizing a user-provided oral utterance to facilitate the automated assistant performing an action. The oral utterance includes natural language content and is received via an automated assistant interface of a computing device connected to a display panel. The method further includes determining whether the oral utterance is complete based on performing speech-to-text processing on the data characterizing the oral utterance, and further includes determining whether the natural language content contains one or more parameter values for controlling a function associated with the action. When the oral utterance is determined to be incomplete, the method further includes causing the display panel connected to the computing device to provide one or more suggestion elements in response to the determination that the oral utterance is incomplete. The one or more suggestion elements include specific suggestion elements that, when spoken by the user to the automated assistant interface, provide via the display panel other natural language content that causes the automated assistant to act to facilitate the completion of an action. The method further includes determining that the user has selected a particular suggestion element from one or more suggestion elements via another oral utterance received in the automated assistant interface. The method further includes causing an action to be performed based on the natural language content of the oral utterance and the other natural language content of the particular suggestion element in response to the determination that the user has selected a particular suggestion element. The method further includes, when the oral utterance is determined to be complete, causing an action to be performed based on the natural language content of the oral utterance in response to the determination that the oral utterance is complete.
[0088] These and other implementations of the technologies disclosed herein may include one or more of the following features:
[0089] In some implementations, the method further includes causing the priority associated with other natural language content to be modified from the previous priority associated with the other natural language content before the oral utterance was determined to be incomplete, in response to the decision that the user has selected a particular suggestion element. In some of those implementations, the order in which a particular suggestion element is presented on the display panel relative to at least one other suggestion element is based at least in part on the priority assigned to the other natural language content.
[0090] In some implementations, a method is provided, performed by one or more processors, which includes performing speech-to-text processing on data characterizing user-provided oral utterances to facilitate triggering an automated assistant to perform an action. The oral utterances include natural language content and are received via an automated assistant interface of a computing device connected to a display panel. The method determines whether the oral utterances are complete based on performing speech-to-text processing on the data characterizing the oral utterances, and further includes determining whether the natural language content contains one or more parameter values for controlling features associated with the action. The method, when it is determined that an oral utterance is incomplete, includes determining contextual data, based on contextual data accessible via a computing device, that characterizes the context in which the user provided the oral utterance to the automated assistant interface; determining a time to present one or more suggestions via a display panel, based on the contextual data, wherein the contextual data indicates that the user had previously provided a separate oral utterance to the automated assistant in the context; and causing the display panel of the computing device to present one or more suggestions via the display panel, based on the determination of a time to present one or more suggestions via the display panel. The one or more suggestion elements include specific suggestion elements that, when spoken by the user to the automated assistant interface, cause the automated assistant to act to facilitate the completion of an action, by providing other natural language content via the display panel.
[0091] These and other implementations of the technologies disclosed herein may include one or more of the following features:
[0092] In some implementations, determining the timing for presenting one or more suggestions involves comparing contextual data with other contextual data that characterizes previous instances in which the user, or one or more other users, provided one or more other oral utterances while in that context.
[0093] In some implementations, the method further includes identifying natural language commands that, when spoken to an automated assistant interface by one or more users, result in a reduced amount of speech-to-text processing compared to the amount of speech-to-text processing associated with the user's speech utterance, when the oral utterance is determined to be incomplete, causing the automated assistant to act in order to facilitate the completion of the action. The natural language commands are incorporated by the content of one or more suggestion elements, one of which is a suggestion element.
[0094] In some implementations, a method is provided, performed by one or more processors, which includes performing speech-to-text processing on audio data characterizing a user-provided oral utterance. The oral utterance includes natural language content and is received via an automated assistant interface of a computing device connected to a display panel. The method further includes causing a first action to be performed based on the natural language content, based on the performance of speech-to-text processing on the audio data characterizing the oral utterance. The method further includes causing the display panel of the computing device to provide one or more suggestion elements in response to the determination that the user has provided an oral utterance containing natural language content. The one or more suggestion elements include specific suggestion elements that, when spoken by the user to the automated assistant interface, cause the automated assistant to act to facilitate the completion of an action, by providing other natural language content via the display panel. The method further includes determining that the user has selected a particular suggestion element from one or more suggestion elements via subsequent oral utterances received in an automated assistant interface, and causing a second action to be performed based on other natural language content identified by the particular suggestion element in response to the determination that the user has selected a particular suggestion element.
[0095] These and other implementations of the technologies disclosed herein may include one or more of the following features:
[0096] In some implementations, the execution of a second action is a modification of other action data resulting from the execution of a first action, and / or produces action data that supplements other action data resulting from the execution of a first action. In some of these implementations, the method is to generate other natural language content in response to a decision that the user has provided an oral utterance, further comprising the other natural language content identifying at least one suggested value for a parameter used during the execution of the second action.
[0097] In some implementations, causing the first action to be performed involves causing data based on predetermined default content associated with the user to be provided on the computing device. In some of those implementations, the method further includes modifying the default content to facilitate the availability of the modified default content when a subsequent incomplete request is received from the user, based on the user selecting a particular suggestion element.
[0098] In some implementations, a method is provided, performed by one or more processors, that generates text based on speech-to-text processing of audio data capturing user-provided oral utterances. The oral utterances are received via an automated assistant interface of a computing device connected to a display panel. The method further includes determining, based on the text, whether the oral utterance is complete or incomplete. The method further includes determining a specific duration based on whether the oral utterance is determined to be complete or incomplete. The specific duration is shorter when the oral utterance is determined to be complete than when it is determined to be incomplete. The method further includes generating one or more suggestion elements based on the text, wherein each of the one or more suggestion elements, when combined with the text, indicates corresponding additional text that causes the automated assistant to act to facilitate a corresponding action. The method further includes causing the display panel of the computing device to provide one or more suggestion elements, and, after provisioning the one or more suggestion elements, monitoring for further oral input for the determined duration. The method further includes generating an Auto Assistant command based on candidate text and additional text generated from performing speech-to-text processing on additional audio data capturing the further verbal input, when further verbal input is received within the duration, and causing the Auto Assistant to execute the Auto Assistant command. The method further includes causing the Auto Assistant to execute an alternative command based solely on candidate text when no further verbal input is received within the duration.
[0099] These and other implementations of the technologies disclosed herein may include one or more of the following features:
[0100] In some implementations, determining a specific duration based on whether an oral utterance is determined to be complete or incomplete further includes determining that a specific duration is a first specific duration when the oral utterance is determined to be complete and directed to a specific automated assistant agent that is not a general search agent, and determining that a specific duration is a second specific duration when the oral utterance is determined to be complete and directed to a general search agent. The second specific duration is longer than the first specific duration.
[0101] In some implementations, determining whether an oral utterance is complete or incomplete further includes determining that a specific duration is a first specific duration when the oral utterance is determined to be complete and include all required parameters, and determining that a specific duration is a second specific duration when the oral utterance is determined to be complete but lacks all required parameters. The second specific duration is longer than the first specific duration.
[0102] In each implementation, the oral utterance may be determined to be incomplete, and associated processing may be performed. In other implementations, the oral utterance may be determined to be complete, and associated processing may be performed. [Explanation of symbols]
[0103] 102, 204 First User Interface 104, 206, 222, 304 Assistant Interactive Modules 106 Graphical elements, suggested text 108, 210, 310 Suggestion elements 110, 212, 320 display panels 112, 214, 312 users 114, 216, 314 computing devices 116 Oral utterance, incomplete oral utterance, first oral utterance 122 Completed candidate text, candidate text 124 Additional suggestion elements, 1st additional suggestion element 126, 232 Second User Interface 128 Additional oral utterances, oral utterances 130 phrases 208 Incomplete text segments 218 Oral utterance, first oral utterance 224 Complete requirements, finished requirements 226 Additional suggestion elements, suggestion elements 228, 318 Additional oral utterances 230 "Until I fall asleep." 306 Graphical representation 308 Graphical prompt, prompt 316 Oral speech 400 System 402 Server Device 408 Automated Assistant 412 Input Processing Engine 414 Speech Processing Engine 416 Data Parsing Engine 418 Action Engine 420 Power Generation Engine 422 Assistant Data 424 Suggestion Engine 426 Device Topology Engine 428 Timing Engine 430 Speech-Text Modeling Engine 432 Action Engine 434 Client Devices 436 Assistant Interface 438 Automated Assistant, Local Automated Assistant 440 Client Data 442 IoT devices 444 User Interface, Action Data 446 Network 610 Computer Systems 612 Bus Subsystem 614 Processors 616 Network Interface Subsystem 620 User Interface Output Devices 622 User Interface Input Devices 624 Memory subsystem 625 memory 626 File Storage Subsystem 630 Main Random Access Memory (RAM) 632 Read-only memory (ROM)
Claims
1. A method carried out by one or more processors, A step of performing speech-to-text processing on audio data characterizing spoken utterances provided by a user, The oral utterance includes natural language content and is received via an automated assistant interface of a computing device connected to a display panel, step and A step of determining whether the oral utterance is complete or incomplete based on the text obtained by the aforementioned speech-to-text processing, A step in which, based on the step of performing speech-to-text processing on the audio data characterizing the oral utterance, a first action is caused to be performed based on the natural language content, In response to a decision that the user has provided the oral utterance including the natural language content, the step of causing the display panel of the computing device to provide one or more suggestion elements, A step including a specific suggestion element that provides, via the display panel, other natural language content that causes the automated assistant to act to facilitate the completion of an action when one or more suggestion elements are spoken by the user to the automated assistant interface, The steps include: after provisioning one or more suggestion elements, monitoring for further verbal input for a specific duration; When the aforementioned further oral input is received within the specified duration, The steps include determining that the user has selected a particular suggestion element from among the one or more suggestion elements via subsequent oral utterances received in the automated assistant interface, A step in which, in response to the decision that the user has selected the particular suggestion element, a second action is performed based on the other natural language content identified by the particular suggestion element, When the oral utterance is complete and no further oral input is received within the specified duration, A method comprising the step of causing the second action to be carried out based solely on the natural language content.
2. The method according to claim 1, wherein the execution of the second action is a modification of other action data resulting from the execution of the first action, and / or produces action data that supplements the other action data resulting from the execution of the first action.
3. A step of generating other natural language content in response to a decision that the user has provided the oral utterance, wherein the other natural language content identifies at least one suggested value for a parameter used during the performance of the second action. The method according to claim 2, further comprising:
4. A step of modifying the priority associated with each of the one or more suggestion elements in response to a decision that the user has selected a particular suggestion element from among the one or more suggestion elements, Each priority indicates whether the corresponding suggestion element will be presented on the display panel in response to a subsequent request for the automated assistant to perform the action. The method according to claim 1, further comprising:
5. The method according to claim 1, wherein the step causing the first action to be performed includes the step causing data based on predetermined default content associated with the user to be provided on the computing device.
6. When a subsequent incomplete request is received from the user based on the user selecting the particular suggestion element, the step of modifying the default content in order to facilitate the availability of the modified default content. The method according to claim 5, further comprising:
7. The method according to claim 1, wherein the other natural language content is based on a previous conversation between the user and the automated assistant that occurred before at least the first action was performed via the automated assistant.
8. The method described above is: A step of determining the specific duration, wherein the specific duration is determined to be shorter when the oral utterance is determined to be complete than when the oral utterance is determined to be incomplete. The method according to claim 1, including the method described in claim 1.
9. The step of determining the specific duration based on whether the oral utterance is determined to be complete or incomplete is: When it is determined that the oral utterance is complete and directed to a specific automated assistant agent that is not a general search agent, the step of determining that the specific duration is a first specific duration, When it is determined that the oral utterance is complete and directed to the general search agent, the step is to determine that the particular duration is a second particular duration, The second specific duration is longer than the first specific duration, and The method according to claim 8, further comprising:
10. The step of determining the specific duration based on whether the oral utterance is determined to be complete or incomplete is: When it is determined that the oral utterance is complete and includes all required parameters, the step of determining that the particular duration is a first particular duration, When it is determined that the oral utterance is complete but lacks all of the aforementioned essential parameters, the step of determining that the particular duration is a second particular duration, The second specific duration is longer than the first specific duration, and The method according to claim 8, further comprising:
11. A computer program comprising instructions that, when executed by one or more processors, cause the one or more processors to perform the method described in any one of claims 1 to 10.
12. A computer-readable storage medium comprising instructions, when executed by one or more processors, causing the one or more processors to perform the method according to any one of claims 1 to 10.
13. A system comprising one or more processors for performing the method described in any one of claims 1 to 10.