Automated assistants perform non-assistant application actions in response to user input that may be limited to parameters.
Selectable GUI elements and machine learning-based action identification in automated assistants streamline user interactions by allowing parameter specification without intent declaration, improving interaction efficiency and clarity.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2025-01-07
- Publication Date
- 2026-04-22
AI Technical Summary
Existing automated assistants require users to specify both intent and parameters in verbal utterances and may not correctly interpret whether the utterance is intended for application control or general response, leading to inefficient and unclear interactions.
Automated assistants provide selectable GUI elements that allow users to specify parameters without explicitly stating intent, using microphone and camera inputs after permission, and utilize machine learning to identify compatible actions, reducing the need for explicit invocation phrases and complex interactions.
This approach enables more concise verbal interactions, reduces processing duration, and ensures clear intent recognition, allowing users to control applications through verbal commands without direct application input, enhancing user experience and efficiency.
Smart Images

Figure 0007850295000001 
Figure 0007850295000002 
Figure 0007850295000003
Abstract
Description
[Background technology]
[0001] Humans can engage in human-computer interaction using interactive software applications referred to herein as “automated assistants” (also known as “digital agents,” “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “conversational agents,” etc.). For example, a human (sometimes referred to as a “user” when interacting with an automated assistant) can give commands and / or requests by using oral natural language input (i.e., utterances), which may in some cases be converted to text and then processed, and / or by providing textual natural language input (e.g., typed).
[0002] In some cases, an automated assistant can offer a variety of features that can be initialized even when the user is interacting with a separate application in the foreground of their computing device. For example, a user can use the automated assistant to perform a search within a separate foreground application. For instance, in response to a verbal utterance directed at the automated assistant, such as "Search for Asian fusion," while a separate restaurant review application is in the foreground, the automated assistant can interact with the foreground application (e.g., directly and / or by emulating user input) to submit a search for "Asian fusion" using the restaurant review application's search interface. Also, as an example, in response to a verbal utterance directed at the automated assistant, such as "Add a calendar entry for a patent meeting tomorrow at 2:00," while a separate calendar application is in the foreground, the automated assistant can interact with the foreground application to create a calendar entry titled "Patent Meeting" for "Tomorrow" at 2:00.
[0003] However, when interacting with a foreground application using an automated assistant, the user must specify both an intent (e.g., "search for" or "add a calendar entry" in the previous examples) and parameters regarding the intent (e.g., "Asian fusion" or "patent meeting at 2:00 tomorrow) in their verbal utterance. Also, depending on the situation, the user may have to provide a verbal call phrase to invoke the automated assistant, or other automated assistant invocation input, before providing the verbal utterance.
[0004] Furthermore, depending on the situation, the user may not be aware that the automated assistant is capable of interacting with the foreground application in response to the user's verbal utterance as desired. Thus, instead, the user may utilize a greater number of inputs and / or inputs over a longer duration directed to the application when interacting directly with the application. For example, assume that the user is not aware that saying "add a calendar entry for a patent meeting at 2:00 tomorrow" to the automated assistant causes the automated assistant to interact with the calendar application. In such a situation, to add the corresponding calendar entry in the calendar application, the user may instead locate and tap on the position of the "add calendar entry" interface element of the calendar application that presents the entry interface, and then click in sequence on the date field, time field, and title field of the entry interface and fill in those fields (e.g., using a virtual keyboard and / or selection menu).
[0005] Furthermore, in some situations, the automated assistant may not be able to correctly determine whether a verbal utterance is attempting to control the foreground application, or whether it is simply requesting a general response from the automated assistant that is generated independently of the foreground application and without any control over it. For example, consider a verbal utterance directed at the automated assistant, "Search for Asian fusion," while a separate restaurant review application is in the foreground. In such an example, it may not be clear whether the user is asking the assistant to perform a search for "Asian fusion" restaurants within the restaurant review application, or whether they simply want the automated assistant to perform a general search (independent of the restaurant review application) and return a general description of what "Asian fusion" cuisine is. [Overview of the Initiative] [Means for solving the problem]
[0006] The implementations described herein relate to an automated assistant that provides selectable GUI elements when a user interacts with an application that can be controlled via the automated assistant. The selectable GUI elements can be rendered when the automated assistant determines that the application interface identifies an action (e.g., a search function) that can be initialized or otherwise controlled via the automated assistant. The selectable GUI elements may include content such as text and / or graphical content that identifies the action and / or prompts the user to provide one or more parameters about the action. When the selectable GUI elements are rendered, the microphone and / or camera can be activated with prior permission from the user to allow the user to identify one or more action parameters without explicitly identifying the automated assistant or intent / action. When one or more action parameters are provided by the user, the automated assistant can control the application to perform the action (e.g., a search function) using those one or more action parameters (e.g., search terms).
[0007] In the above and other forms, the interaction between the application and the automated assistant can be performed using reduced and / or more concise user input. For example, the user's verbal utterance can specify only parameters about intent or action, without specifying the intent or action itself. As a result, verbal utterances become more concise, and the processing of verbal utterances by the automated speech recognition component and / or other components is reduced accordingly. In addition, the user does not need to provide a clear calling phrase (e.g., "Assistant..."), thereby further reducing the duration of verbal utterances and the overall duration of human / assistant interaction. Furthermore, the user's intent becomes clear through the user's selection of GUI elements, thereby preventing the automated assistant from misinterpreting verbal utterances as general assistant requests rather than requests for the assistant to control the foreground application. Furthermore, through the presentation of GUI elements, users become aware of their ability to control the foreground application through verbal utterances directed at the automated assistant rather than through more complex, direct interactions with the foreground application, and / or users become more likely to control the foreground application more frequently through verbal utterances (e.g., provided after selecting a GUI element).
[0008] In some implementations, the Auto Assistant can determine whether an application interface contains features corresponding to each action that is compatible with the Auto Assistant and / or the Assistant's actions. In some cases, multiple different compatible actions may be identified for a single application interface, causing the Auto Assistant to render one or more selectable GUI elements for each action. The type of selectable GUI element rendered by the Auto Assistant may depend on the corresponding action identified by the Auto Assistant. For example, when a user accesses a home control application that includes a dial GUI element for controlling the temperature of a house, the Auto Assistant may render a selectable GUI element that identifies a command phrase for adjusting the temperature. In some implementations, the selectable GUI element may contain text such as "Set the temperature to _____" that indicates that the selectable GUI element corresponds to the action of setting the temperature of a house.
[0009] A blank or placeholder area (e.g., "____") in a selectable GUI element can prompt the user for a verbal utterance or other input to identify parameters for completing a command phrase and / or initializing the execution of the corresponding action, and / or can otherwise provide an indicator that the user can provide such verbal utterance or other input. For example, the user may tap on the selectable GUI element to complete a command phrase written in the text of the selectable GUI element, and / or subsequently provide a verbal utterance such as "65 degrees". In response to receiving the verbal utterance, the automated assistant can control the application to adjust the application's temperature setting to "65" degrees. In some implementations, when the selectable GUI element is rendered by the automated assistant, the automated assistant can also activate the computing device's audio interface (e.g., one or more microphones). Therefore, instead of tapping on selectable GUI elements, users can provide verbal utterances that specify a parameter value (e.g., "65 degrees") without specifying the action they want to perform (e.g., "change the temperature") or the assistant (e.g., "assistant").
[0010] In some implementations, selectable GUI elements can be rendered by an automated assistant in the foreground of the computing device's display interface for a threshold duration. The duration can be selected according to one or more characteristics related to user-application interaction. For example, when the application's home screen is rendered on the display interface and the user is not providing any input to the application, the selectable GUI element can be rendered for a static duration (e.g., 3 seconds). However, when the selectable GUI element is rendered on the application interface while the user is interacting with the application (e.g., scrolling through the application interface), the selectable GUI element can be rendered for a duration based on how often the user provides input to the application. Alternatively, or in addition to this, the rendering duration of a selectable GUI element can be based on the amount of time the corresponding application interface element is rendered or is expected to be rendered on the application interface. For example, if a user typically provides application input that moves the application from the home screen to the login screen within the time t spent viewing the home screen, the selectable GUI element can be rendered on the home screen for a duration based on time t.
[0011] In some implementations, the selection of the types of selectable GUI elements to render can be based on a heuristic process and / or one or more trained machine learning models. For example, an automated assistant and / or operating system of a computing device may process data related to an application interface to identify one or more actions that can be initialized through user interaction with the application interface. This data, with prior permission from the user, may include screenshots of the application interface, links corresponding to the graphical elements of the interface, library data and / or other functional data related to the application and / or interface, and / or any other information that can indicate actions that can be initialized through the application interface.
[0012] Depending on one or more actions identified for an application interface, the Auto Assistant can select and / or generate selectable GUI elements corresponding to each action. Selectable GUI elements can be selected to provide indicators that each action can be controlled via the Auto Assistant. For example, the Auto Assistant may determine that a magnifying glass icon (e.g., a search icon) placed on or adjacent to a blank text field (e.g., a search field) in an application interface can indicate that the application interface can control the application's search behavior. Based on this determination, the Auto Assistant may render selectable GUI elements that include the same or different magnifying glass icons and / or one or more natural language words synonymous with the word "search" (e.g., "search for ______"). In some implementations, when a GUI element is selectable, the user can select it by providing a verbal utterance that specifies the search parameter (e.g., "restaurants nearby"), or by tapping the selectable GUI element and then providing a verbal utterance that specifies the search parameter.
[0013] In some implementations, the microphone of the computing device rendering the selectable GUI element may remain active after the user has selected the selectable GUI element, with prior permission from the user. Alternatively, or in addition to this, when the application interface changes in response to the selection of a selectable GUI element and / or verbal utterance, the automated assistant may select another selectable GUI element to render. The automated assistant may select the other selectable GUI element to render based on the next application interface to which the application will transition. For example, when a user issues search parameters directed to a selectable GUI element, the application may render a list of search results. The search results from the search results list may be selectable by the user to cause the application to perform a specific action. The automated assistant may determine that this specific action is compatible with an action of the assistant (e.g., an action that can be performed by the automated assistant) and may render another selectable GUI element (e.g., a hand with an index finger extended toward the corresponding search result) on or adjacent to the corresponding search result. Alternatively, or in addition to this, this other selectable GUI element may contain a text string that identifies the word corresponding to the search result (e.g., "Time Four Thai Restaurant"). When a user provides an oral utterance containing one or more words that identify a corresponding search result (e.g., "Thai restaurants"), the automated assistant can select the corresponding search result without the user explicitly identifying a "select" action or the automated assistant. In this way, the automated assistant continues to identify appropriate actions at each interface of the application, allowing the user to navigate the interface by providing parameter values (e.g., "Nearby restaurants... Thai restaurants... Menu...").In some implementations, the user can initiate interaction by first instructing the automated assistant to open a specific application via a first verbal utterance (e.g., "Assistant, open my recipes application..."). Then, once the automated assistant identifies a suitable application action, the user can provide another command via a second verbal utterance, allowing the automated assistant to control this particular application according to parameters (e.g., the user can say "Pad Thai" to have the automated assistant search for "Pad Thai" within their recipes application). The user can then continue navigating this particular application using these concise verbal utterances until the automated assistant recognizes at least one or more application actions as suitable and / or controllable through one or more of the automated assistant's actions.
[0014] The above description is provided as an overview of some implementations of this disclosure. Further descriptions of those implementations and other implementations are provided below in more detail.
[0015] Other implementations may include a non-temporary computer-readable storage medium that stores instructions for performing methods such as one or more of the methods described above and / or elsewhere in this specification, which can be executed by one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)). Further implementations may include a system of one or more computers, including one or more processors capable of operating to execute the stored instructions for performing methods such as one or more of the methods described above and / or elsewhere in this specification.
[0016] It should be understood that any combination of the aforementioned concepts and any further concepts described in more detail herein is intended to be part of the subject matter disclosed herein. For example, any combination of the claimed subject matter appearing at the end of this disclosure is intended to be part of the subject matter disclosed herein. [Brief explanation of the drawing]
[0017] [Figure 1A] This diagram shows a view in which a user interacts with an automated assistant that recognizes and suggests assistant-appropriate behaviors within a third-party application interface. [Figure 1B] This diagram shows a view in which a user interacts with an automated assistant that recognizes and suggests assistant-appropriate behaviors within a third-party application interface. [Figure 1C] This diagram shows a view in which a user interacts with an automated assistant that recognizes and suggests assistant-appropriate behaviors within a third-party application interface. [Figure 2A] This diagram shows a view in which a user interacts with an application that performs one or more actions that can be controlled by an automated assistant. [Figure 2B] This diagram shows a view in which a user interacts with an application that performs one or more actions that can be controlled by an automated assistant. [Figure 2C] This diagram shows a view in which a user interacts with an application that performs one or more actions that can be controlled by an automated assistant. [Figure 3]This figure shows a system with an automated assistant that can provide selectable action intent suggestions, including indicators that the user should provide one or more parameters to control the behavior when the user is accessing a third-party application that can be controlled via the automated assistant. [Figure 4] This figure shows a method for providing selectable GUI elements within a computing device interface when application behavior suitable for an automated assistant is executable via that interface. [Figure 5] This is a block diagram of an exemplary computer system. [Modes for carrying out the invention]
[0018] Figures 1A, 1B, and 1C show views 100, 120, and 140, respectively, of user 102 interacting with an automated assistant that recognizes and proposes assistant-compatible actions in a third-party application interface. When an assistant-compatible action is detected in the application interface, the automated assistant can provide an indicator that the user can provide one or more parameters to initialize the execution of the compatible action. In this way, the user can avoid providing an assistant invocation phrase and / or other inputs that would normally be required to control the automated assistant and / or third-party applications.
[0019] For example, user 102 may interact with an assistant-enabled device such as a computing device 104 to change the settings of the home control application 108. First, user 102 may initialize the home control application 108 in response to user input, which may be touch input provided by user 102's hand 116. When the home control application 108 is launched, user 102 may choose to control specific devices in user 102's home, such as the home's heating, ventilation, and air conditioning (HVAC) system. To control the HVAC system, the home control application 108 may render an application interface 110 which may include a thermostat GUI 112.
[0020] When the application interface 110 is rendered in the display interface 118 of the computing device 104, the automatic assistant can identify one or more assistant-compliant operations 114. When an assistant-compliant operation is identified as being associated with a selectable GUI element, the automatic assistant can cause one or more graphical elements to be rendered in the display interface 118. For example, as presented in view 120 of FIG. 1B, the automatic assistant can cause a selectable GUI element 122 and / or a proposed element 126 to be rendered in the display interface 118. In some implementations, the selectable GUI element 122 can be rendered in the foreground of the home control application 108 and / or over the home control application 108. The selectable GUI element 122 can provide the user with an indication that the automatic assistant is currently initialized and that a wake phrase is not necessarily required when the selectable GUI element 122 is rendered. Alternatively or in addition, the proposed element 126 can provide natural language content characterizing a possible verbal utterance provided to the automatic assistant to control the home control application 108. Further, the presence of the proposed element 126 can indicate that the automatic assistant is initialized and that a wake phrase is not necessarily required prior to recognizing verbal input. Alternatively or in addition, the natural language content of the proposed element 126 can include at least a portion of the proposed verbal utterance and a blank space that can serve as a placeholder for one or more parameters for the assistant-compliant operation.
[0021] For example, as shown in Figure 1B, the automated assistant can render a selectable GUI element 122 on the home control application 108 to indicate that it can receive temperature input. In some implementations, an icon corresponding to an assistant-adaptive operation can be selected to be rendered using the selectable GUI element 122. The selection of the icon can be based on parameters that can be provided by the user 102 to control the thermostat GUI 112. For example, a thermometer icon can be selected to indicate that the user can specify a temperature value to control the application interface 110 and adjust the thermostat GUI 112. The user 102 can then provide a verbal utterance 124, such as "65 degrees," thereby indicating to the automated assistant that they want the user's automated assistant to change the temperature setting of the home control application 108 from 72 degrees to 65 degrees.
[0022] In response to a verbal utterance 124, the automated assistant can initialize an assistant-adapted operation 144. For example, the automated assistant can generate a request to the home control application 108 to modify the current setting of the thermostat from 72 degrees to 65 degrees. In some implementations, an application programming interface (API) can be used to interface between the automated assistant and the home control application 108. The home control application 108 can process the request from the automated assistant and modify the thermostat setting accordingly. In addition, to indicate to the user 102 that the automated assistant and the home control application 108 have successfully performed the operation, an updated thermostat GUI 142 may be rendered as an update in the application interface 110 146. In some implementations, selectable GUI elements 122 and / or suggestion elements 126 may be removed from the display interface 118 after a threshold duration has elapsed and / or regardless of whether the user 102 has interacted with the selectable GUI elements 122 or suggestion elements 126. For example, a selectable GUI element 122 can be rendered for a threshold duration from the time the home control application 108 is first rendered on the display interface 118. If user 102 does not interact with the selectable GUI element 122 during the threshold duration, the automated assistant can provide a notification that the selectable GUI element 122 will no longer be rendered on the display interface 118 and / or that the selectable GUI element 122 will be removed after a certain period of time.
[0023] Figures 2A, 2B, and 2C each show view 200, view 220, and view 240 in which user 202 interacts with an application that can perform one or more operations controlled by an automated assistant. For example, user 202 can access messaging application 208 using computing device 204. In response to user 202 launching messaging application 208, an automated assistant accessible via computing device 204 can determine whether features of messaging application 208 can be controlled through the automated assistant. In some implementations, the automated assistant can determine that one or more operations can be performed using specific parameters specified by user 202. For example, the automated assistant can determine that to reply to a message, checkbox 212 must be selected for a particular message and then reply icon 218 must be selected. In some implementations, the automated assistant can make this determination based on an exploratory process and / or one or more trained machine learning models. For example, training data for one or more trained machine learning models can characterize cases where reply icon 218 is selected without checkbox 212 being selected, and other cases where reply icon 218 is selected when checkbox 212 is selected. In some implementations, one or more trained machine learning models can be used to process screenshots and / or other images of an application interface with prior permission from the user to render specific selectable GUI elements and / or selectable suggestions on computing device 204.
[0024] When the computing device 204 initializes the messaging application 208 in response to user input 206, the automated assistant can identify available assistant-compatible actions through the application interface 210 of the messaging application 208. When one or more assistant-compatible actions are identified, the automated assistant can cause the computing device 204 to render one or more selectable GUI elements 222 and / or one or more selectable suggestions 224 214. The selectable GUI elements 222 can be rendered to indicate that the automated assistant and / or audio interface have been initialized, and that the user 202 can identify parameters to allow the automated assistant to control the messaging application 208 using those parameters.
[0025] For example, a selectable GUI element 222 may include a graphical representation of a person or contact, thereby indicating that user 202 should identify the name of the person to whom they wish to send a message. Alternatively, or in addition to this, a selectable suggestion 224 may include a text identifier and / or graphical content that identifies a command that can be issued to the automated assistant but lacks one or more parameters. For example, a selectable suggestion 224 may include the phrase "Reply to message from _____", which indicates that the automated assistant can reply to a message identified within the application interface 210, provided user 202 identifies the contact associated with that particular message. User 202 can then provide a verbal utterance 226 that identifies parameters for this assistant-adapted behavior. In response to the verbal utterance 226, the automated assistant can make selections in checkboxes 212 corresponding to the parameters identified by user 202 (e.g., "Linus"). In addition, in response to a verbal utterance 226, the automated assistant may select a reply icon 218 to cause the messaging application 208 to reply to a message from a contact identified by user 202. Alternatively, or in addition to that, as a backend process, the automated assistant may communicate an API call to the messaging application 208 to initialize a reply to a message from a contact identified by user 202.
[0026] In response to a verbal utterance 226 from user 202, the automated assistant can cause the messaging application 208 to process a request to reply to a message from a contact identified by user 202 (e.g., Linus). When the messaging application 208 receives the request from the automated assistant, it can render an updated application interface 248. The application interface 248 can accommodate a draft reply message that can be modified by user 202. The automated assistant can process the content of the application interface 248 and / or other data stored in relation to the application interface 248 to determine whether to provide additional suggestions to user 202. For example, the automated assistant can cause one or more additional selectable GUI elements 242 to be rendered in the foreground of the application interface 248. The selectable GUI elements 242 can be rendered in step 244 to indicate to user 202 that the automated assistant is active and that user 202 can provide a verbal utterance detailing the text of the reply message. For example, when a selectable GUI element 242 is being rendered, the user 202 can provide another verbal utterance 246 as a message text, such as "Yes, I'll see you then," without explicitly specifying an action and / or an automated assistant.
[0027] In response, the automated assistant can communicate another request to the messaging application 208 to cause it to perform one or more actions to input the text "Yes, see you then" into the message body. The user 202 can then provide another verbal utterance (e.g., "Send") directed to the automated assistant and a separate selectable GUI element 250. In this way, the user 202 can have the messaging application 208 send a message without explicitly identifying the automated assistant and without providing touch input to the computing device 204 and the messaging application 208. This can reduce the number of inputs that need to be directly provided by the user 202 to the third-party application. Furthermore, the user 202 can utilize the automated assistant when interacting with the vast majority of other applications that may not use a trained machine learning model trained on actual interactions with the user 202.
[0028] Figure 3 shows a system 300 with an automated assistant 304 that can provide selectable action intent suggestions, including an indicator that the user should provide one or more parameters to initialize an action when the user is accessing a third-party application controllable via the automated assistant. One or more actions associated with the action intent can be initialized in response to the user providing input (e.g., verbal utterance) specifying one or more parameters without necessarily specifying the automated assistant 304 or the third-party application. The automated assistant 304 can operate as part of an assistant application provided in one or more computing devices, such as computing device 302 and / or server devices. The user can interact with the automated assistant 304 via an assistant interface 320, which can be a microphone, camera, touchscreen display, user interface, and / or any other device capable of providing an interface between the user and the application. For example, a user can initialize the automated assistant 304 by providing verbal, text, and / or graphical inputs to the assistant interface 320 to cause the automated assistant 304 to initialize one or more actions (e.g., providing data, controlling a peripheral device, accessing an agent, generating input and / or output). Alternatively, the automated assistant 304 may be initialized based on processing context data 336 using one or more trained machine learning models. The context data 336 can characterize one or more features of the environment accessible to the automated assistant 304, and / or one or more features of users who are expected to interact with the automated assistant 304.The computing device 302 may include a display device, which may be a display panel including a touch interface, which is for receiving touch input and / or gestures that allow the user to control the application 334 of the computing device 302 via the touch interface. In some implementations, the computing device 302 may lack a display device and therefore provide an audible user interface output without providing a graphical user interface output. Furthermore, the computing device 302 may provide a user interface such as a microphone for receiving oral natural language input from the user. In some implementations, the computing device 302 may include a touch interface and may not have a camera, but the computing device 302 may optionally include one or more other sensors.
[0029] Computing device 302 and / or other third-party client devices can communicate with a server device over a network such as the Internet. In addition, computing device 302 and any other computing devices can communicate with each other over a local area network (LAN), such as a Wi-Fi network. Computing device 302 can offload computing tasks to the server device to conserve computing resources on computing device 302. For example, the server device can host an automated assistant 304, and / or computing device 302 can send inputs received at one or more assistant interfaces 320 to the server device. However, in some implementations, the automated assistant 304 can be hosted on computing device 302, and various processes that can be associated with the operation of the automated assistant can be performed on computing device 302.
[0030] In various implementations, all or some aspects of the automated assistant 304 can be implemented on the computing device 302. In some of these implementations, aspects of the automated assistant 304 can interface with a server device that is implemented by the computing device 302 and can implement other aspects of the automated assistant 304. The server device can optionally serve multiple users and their associated assistant applications via multiple threads. In implementations where all or some aspects of the automated assistant 304 are implemented by the computing device 302, the automated assistant 304 can be a separate application from the operating system of the computing device 302 (e.g., installed "on top of" the operating system), or alternatively, it can be implemented directly by the operating system of the computing device 302 (e.g., integrated with the operating system but considered an application of the operating system).
[0031] In some implementations, the automated assistant 304 may include an input processing engine 306, which can process inputs and / or outputs of the computing device 302 and / or the server device using multiple different modules. For example, the input processing engine 306 may include a speech processing engine 308 that can process audio data received at the assistant interface 320 to perform speech recognition and / or identify text specifically expressed within that audio data. To conserve computing resources in the computing device 302, the audio data can be sent, for example, from the computing device 302 to the server device. Alternatively, the audio data can be processed entirely within the computing device 302.
[0032] The process for converting audio data to text may include a speech recognition algorithm, which can use a neural network and / or statistical model to identify groups of audio data corresponding to words or phrases. The text converted from the audio data can be parsed by the data parsing engine 310 and made available to the automated assistant 304 as text data that can be used to generate and / or identify command phrases, intents, actions, slot values, and / or other arbitrary content specified by the user. In some implementations, the output data provided by the data parsing engine 310 can be provided to the parameter engine 312 to determine whether the user has provided input corresponding to a particular intent, action, and / or routine that can be performed by the automated assistant 304 and / or an application or agent accessible through the automated assistant 304. For example, assistant data 338 can be stored in the server device and / or computing device 302, and the assistant data 338 may include data defining one or more actions that can be performed by the automated assistant 304, as well as the parameters necessary to perform those actions. The parameter engine 312 can generate one or more parameters relating to intent, action, and / or slot values, and provide these one or more parameters to the output generation engine 314. The output generation engine 314 can use these one or more parameters to communicate with the assistant interface 320 to provide output to the user, and / or with one or more applications 334 to provide output to one or more applications 334.
[0033] In some implementations, the automated assistant 304 can be an application that can be installed "on top of" the operating system of the computing device 302, and / or it can form part (or all) of the operating system of the computing device 302 itself. The automated assistant application includes and / or can access on-device speech recognition, on-device natural language understanding, and on-device implementation. For example, on-device speech recognition can be implemented using an on-device speech recognition module that processes audio data (detected by the microphone) using an end-to-end speech recognition machine learning model stored locally on the computing device 302. On-device speech recognition generates recognized text for any oral utterances present in the audio data. Alternatively, for example, on-device natural language understanding (NLU) can be implemented using an on-device NLU module that processes the recognized text generated using on-device speech recognition, and optionally contextual data, to generate NLU data.
[0034] NLU data can include the intent corresponding to a verbal utterance and optionally parameters related to that intent (e.g., slot values). On-device performance can be carried out using an on-device performance module that utilizes the NLU data (from the on-device NLU) and optionally other local data to determine the actions to be taken to resolve the intent (and optionally parameters related to that intent) of the verbal utterance. This may include determining local and / or remote responses (e.g., answers) to the verbal utterance, interactions with locally installed applications to be performed based on the verbal utterance, commands to be sent to an Internet of Things (IoT) device (directly or via a corresponding remote system) based on the verbal utterance, and / or other resolution actions to be performed based on the verbal utterance. On-device performance can then initiate local and / or remote implementation / execution of the determined actions to resolve the verbal utterance.
[0035] In various implementations, remote speech processing, remote NLU, and / or remote execution may be used, at least selectively. For example, recognized text may be sent, at least selectively, to a remote automated assistant component for remote NLU and / or remote execution. For instance, recognized text may optionally be sent for remote execution in parallel with on-device execution, or in response to failure of on-device NLU and / or on-device execution. However, on-device speech processing, on-device NLU, on-device execution, and / or on-device execution may be preferred, at least for the reduced latency they introduce when resolving oral utterances (i.e., because client-server round trips are not required to resolve oral utterances). Furthermore, on-device functionality may be the only functionality available in situations with no or limited network connectivity.
[0036] In some implementations, the computing device 302 may include one or more applications 334 that may be provided by a third-party entity different from the entity that provided the computing device 302 and / or the automated assistant 304. The application state engine of the automated assistant 304 and / or the computing device 302 may access application data 330 to determine one or more actions that may be performed by one or more applications 334, and the state of each application of one or more applications 334 and / or the state of each device associated with the computing device 302. The device state engine of the automated assistant 304 and / or the computing device 302 may access device data 332 to determine one or more actions that may be performed by the computing device 302 and / or one or more devices associated with the computing device 302. Furthermore, application data 330 and / or other arbitrary data (e.g., device data 332) can be accessed by the automated assistant 304 to generate context data 336, which can characterize the context in which a particular application 334 and / or device is running, as well as the context in which a particular user is accessing the computing device 302, the application 334, and / or any other arbitrary device or module.
[0037] While one or more applications 334 are running on the computing device 302, device data 332 can characterize the current operational state of each application 334 running on the computing device 302. Furthermore, application data 330 can characterize one or more features of the running applications 334, such as the content of one or more graphical user interfaces being rendered at the direction of one or more applications 334. Alternatively or in addition, application data 330 can characterize action schemas that can be updated by each application and / or by the automated assistant 304 based on the current operational state of each application. Alternatively or in addition, one or more action schemas for one or more applications 334 may remain fixed but can be accessed by the application state engine to determine the appropriate action to be initialized via the automated assistant 304.
[0038] The computing device 302 may further include an assistant invocation engine 322, which can process application data 330, device data 332, context data 336, and / or any other data accessible to the computing device 302 using one or more trained machine learning models. The assistant invocation engine 322 may process this data to determine whether the data should be considered to indicate the user's intention to invoke the automated assistant, instead of waiting for the user to explicitly speak an invocation phrase to invoke the automated assistant 304 or requiring the user to explicitly speak an invocation phrase. For example, one or more trained machine learning models can be trained using training data instances based on a scenario in which the user is in an environment where multiple devices and / or applications exhibit various operating states. Training data instances can be generated to take up training data characterizing the contexts in which the user invokes the automated assistant and other contexts in which the user does not invoke the automated assistant.
[0039] When one or more trained machine learning models have been trained according to these training data instances, the assistant invocation engine 322 can cause the automated assistant 304 to detect or limit the detection of verbal invocation phrases from the user based on contextual and / or environmental characteristics. In addition to or instead, the assistant invocation engine 322 can also cause the automated assistant 304 to detect or limit the detection of one or more assistant commands from the user based on contextual and / or environmental characteristics. In some implementations, the assistant invocation engine 322 can be disabled or restricted based on the fact that computing device 302 has detected an assistant suppression output from another computing device. In this way, when computing device 302 has detected an assistant suppression output, the automated assistant 304 will not be invoked based on the context data 336 that would normally cause the automated assistant 304 to be invoked if no assistant suppression output had been detected.
[0040] In some implementations, the system 300 may include an action detection engine 316 capable of identifying one or more actions that can be executed by an application 334 and controlled via an automated assistant 304. For example, the action detection engine 316 may process application data 330 and / or device data 332 to determine whether the application is running on the computing device 302. The automated assistant 304 may determine whether the application can be controlled via the automated assistant 304 and may identify one or more application actions that can be controlled via the automated assistant 304. For example, application data 330 may identify one or more application GUI elements rendered on the interface of the computing device 302, and this application data 330 may be processed to identify one or more actions that can be controlled via the application GUI elements. When an action is identified as suitable for the automated assistant 304, the action detection engine 316 may communicate with a GUI element content engine 318 to generate a selectable GUI element corresponding to that action.
[0041] The GUI element content engine 318 can identify that one or more actions determined by the automated assistant 304 are suitable for the automated assistant 304, and can generate one or more selectable GUI elements based on those actions. For example, when it is determined that a search icon and / or a search text field are available to the application, and the application search action is suitable for the automated assistant 304, the GUI element content engine 318 can generate content for rendering on the display interface of the computing device 302. The content may include text content (e.g., natural language content) and / or graphical content that can be based on the suitable action (e.g., the application search action). In some implementations, a command phrase can be generated and rendered to instruct the automated assistant 304 to initialize the action in order to inform the user about the identified suitable action. Alternatively, or in addition to that, the command phrase may be a partial command phrase that omits one or more parameters about the action, thereby indicating to the user that the user can provide one or more parameters to the automated assistant 304 to initialize the action. Text content and / or graphical content can be rendered on the display interface of the computing device 302 at the same time that the application renders one or more additional GUI elements. The user can initialize the execution of adaptive behavior by tapping the display interface to select a selectable GUI element and / or by providing the automated assistant 304 with a verbal utterance specifying one or more parameters.
[0042] In some implementations, the system 300 may include a GUI element duration engine 326 that can control the duration for which a selectable GUI element is rendered in the display interface by the automated assistant 304. In some implementations, the amount of time a selectable GUI element is rendered may be based on the amount of interaction between the user and the automated assistant 304, and / or the amount of interaction between the user and the application associated with the selectable GUI element. For example, the GUI element duration engine 326 may establish a longer display duration for a selectable GUI element when the selectable GUI element has been rendered and the user has not yet provided input to the application. This longer duration may be longer than the display duration for a selectable GUI element that is being rendered when the user provides input to the application (not to the automated assistant 304). Alternatively, or in addition to that, the display duration for a selectable GUI element may be longer than for a selectable GUI element that the user has previously interacted with in the past. This longer duration can be compared to the duration of other selectable GUI elements that were previously presented to the user but which the user had not previously interacted with or otherwise shown interest in.
[0043] In some implementations, the system 300 may include an action execution engine 324 that can initialize one or more actions of an identified application in response to a user specifying one or more parameters for that action. For example, when selectable GUI elements are being rendered on the application interface by an automated assistant 304, the user can provide a verbal utterance to select a selectable GUI element and / or specify parameters. The action execution engine 324 can then process this selection and / or verbal utterance and generate one or more requests to the application based on the one or more parameters specified by the user. For example, a verbal utterance processed by an input processing engine 306 may specify one or more specific parameter values. These parameter values can be used by the action execution engine 324 to generate one or more requests to the application corresponding to the selectable GUI element identified by the user. For example, a request generated by the automated assistant 304 may identify the action to be performed, one or more parameters specified by the automated assistant 304, and / or one or more parameters specified by the user. In some implementations, the automated assistant 304 can select one or more parameters for an operation in order for the operation to be initialized, and the user can specify one or more additional parameters. For example, when the application is a travel booking application, the automated assistant 304 can assume a date parameter (e.g., the month "January"), and the user can specify the destination city (e.g., "Nairobi") via verbal utterance.Based on this data and the corresponding selectable GUI elements rendered in the display interface, the action execution engine 324 can generate a request to the travel booking application (e.g., Application.Travel.com[search.setCity("Nairobi"), search.setTime("January")]) to initialize the execution of an action. This request can be received by the travel booking application from the automated assistant 304, and in response, the travel booking application can render another application interface containing the results of the action (e.g., a list of available hotels in Nairobi in January).
[0044] Figure 4 illustrates Method 400 for providing selectable GUI elements on a computing device interface when application behaviors suitable for an automated assistant are executable via that interface. Method 400 can be implemented by one or more applications, devices, and / or any other apparatus or module capable of interacting with the automated assistant. In some implementations, Method 400 may include a step 402 for determining whether a non-assistant application is running on the computing device interface. The computing device may enable access to the automated assistant, which can respond to natural language input from the user to control multiple different applications and / or devices. The automated assistant may process data indicating whether a particular application is running on the computing device. For example, data based on content rendered on the interface may be processed by the automated assistant to identify application behaviors that can be initialized via the interface. When it is determined that a non-assistant application is running on an interface such as the computing device's display interface, Method 400 can proceed from step 402 to step 404. Otherwise, the automated assistant can continue to determine whether the application is running on the computing device interface.
[0045] Step 404 may include determining whether the application behavior is compatible with the automated assistant. In other words, the automated assistant may determine whether any behavior that can be performed by the application is compatible with or otherwise initialized by the automated assistant. For example, if the application is a home control application and the application interface includes a control dial GUI, the automated assistant may determine that any behavior controlled by the control dial GUI is compatible with one or more functions of the automated assistant. Thus, the automated assistant can operate to control the control dial GUI and / or the corresponding application behavior. If it is determined that the application behavior is compatible with the automated assistant, method 400 may proceed from step 404 to step 406. Otherwise, the automated assistant may continue to determine whether any other application behavior is compatible with the automated assistant, or whether any other non-assistant application is running on the computing device or a separate computing device.
[0046] Step 406 may include rendering a selectable GUI element in the interface and activating an audio interface on the computing device. The selectable GUI element can provide an indicator that it is active so that the automated assistant can receive one or more input parameters. In some implementations, the selectable GUI element may include text and / or graphical content based on the application behavior identified in step 404. In this way, the user can be informed that the automated assistant can receive input specifying one or more parameters for a particular application behavior, at least while the selectable GUI is being rendered in the interface. In some implementations, the graphical and / or text content of the selectable GUI element may indicate that the microphone is active so that it can receive user input from the user. For example, the selectable GUI element may have a dynamic property indicating that one or more sensors associated with the computing device are active. Alternatively or in addition, the text content of the selectable GUI element may identify one or more partial assistant command phrases lacking one or more parameters that should be specified for one or more respective application behaviors to be performed.
[0047] When a selectable GUI element is rendered in the interface, method 400 may move from step 406 to an optional step 408, which includes determining whether the user has provided touch input or other input directed to the selectable GUI element. If it is determined that the user has provided input directed to the selectable GUI element, method 400 may move from step 408 to step 410. Step 410 may include initializing the detection of audio data corresponding to parameters about application behavior. For example, an automated assistant may identify one or more speech processing models for identifying one or more parameters related to application behavior. In some cases, when application behavior includes one or more numbers as possible parameters, speech processing models may be used to identify numbers of varying magnitudes. Alternatively or in addition, when application behavior includes one or more appropriate names as possible parameters, speech processing models may be used to identify appropriate names in speech.
[0048] Method 400 can move from step 410 or step 408 to step 412, which can determine whether the user has provided the automated assistant with input parameters related to application behavior. For example, the user may provide input related to application behavior by specifying a value in the control dial GUI. Alternatively, or in addition to that, the user may provide input related to application behavior by specifying one or more other values that can be used as one or more parameters for the application behavior. For example, when the application behavior is controllable via the application's control dial GUI, the user may provide the automated assistant with a verbal utterance, such as "10 percent." This verbal utterance indicates that the user is specifying "10 percent" as a parameter for the application behavior and that the automated assistant must initialize the application behavior based on this specified parameter. When the application behavior corresponds to, for example, the brightness of lighting in the user's home, the user can specify a parameter value, allowing the automated assistant to adjust the brightness of the lighting via the application (e.g., an IoT application that controls Wi-Fi enabled light bulbs).
[0049] When it is determined that the user has provided input specifying one or more parameters about the operation of the application, method 400 may move from step 412 to step 414. Step 414 may include having the automated assistant control a non-assistant application according to the input parameters specified by the user. For example, when the user provides a verbal utterance such as "10 percent", the automated assistant may control the non-assistant application to adjust one or more lights associated with the non-assistant application to a brightness level of 10%. This can be done without the user explicitly specifying the assistant or non-assistant application in a verbal utterance. This can conserve computational resources and limit the possibility that certain interferences (e.g., background noise) may affect the audio data captured by the automated assistant. When the user does not provide input specifying parameters within a threshold duration, method 400 may move from step 412 to step 416, which may include having the selectable GUI element removed from the interface after a threshold duration. Method 400 can move from step 414 to step 416, and then method 400 can return to step 402 or another step.
[0050] Figure 5 is a block diagram 500 of an exemplary computer system 510. The computer system 510 typically includes at least one processor 514 that communicates with several peripheral devices via a bus subsystem 512. These peripheral devices may include, for example, a storage subsystem 524 including memory 525 and a file storage subsystem 526, a user interface output device 520, a user interface input device 522, and a network interface subsystem 516. The input and output devices enable user interaction with the computer system 510. The network interface subsystem 516 provides an interface to an external network and is coupled to a corresponding interface device in another computer system.
[0051] The user interface input device 522 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touchscreens integrated into displays, voice recognition systems, microphones, and / or other types of input devices. Generally, the use of the term “input device” is intended to include all possible types of devices and means for inputting information within the computer system 510 or onto a communication network.
[0052] The user interface output device 520 may include non-visual displays such as a display subsystem, printer, fax machine, or audio output device. The display subsystem may include flat panel devices such as cathode ray tubes (CRTs) or liquid crystal displays (LCDs), projection devices, or any other mechanism for producing visible images. The display subsystem may also provide non-visual displays, such as through an audio output device. In general, the use of the term “output device” is intended to include all possible types of devices and means for outputting information from the computer system 510 to the user or to another machine or computer system.
[0053] The storage subsystem 524 stores programming structures and data structures that provide the functionality of some or all of the modules described herein. For example, the storage subsystem 524 may include logic for implementing selected embodiments of method 400 and / or for implementing one or more of the system 300, computing device 104, computing device 204, automated assistant, and / or any other applications, devices, apparatus, and / or modules discussed herein.
[0054] These software modules are generally executed by the processor 514 alone or in combination with other processors. The memory 525 used within the storage subsystem 524 may include several types of memory, including a main random access memory (RAM) 530 for storing instructions and data during program execution, and a read-only memory (ROM) 532 for storing fixed instructions. The file storage subsystem 526 can provide persistent storage for program files and data files and may include hard disk drives, floppy disk drives and associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Modules implementing a particular implementation of a function can be stored by the file storage subsystem 526 within the storage subsystem 524 or on other machines accessible by the processor 514.
[0055] The bus subsystem 512 provides a mechanism for various components and subsystems of the computer system 510 to communicate with each other as intended. Although the bus subsystem 512 is schematically shown as a single bus, multiple buses can be used in alternative implementations of the bus subsystem.
[0056] The computer system 510 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the constantly changing nature of computers and networks, the description of the computer system 510 depicted in Figure 5 is intended only as a specific example to illustrate several implementation forms. Many other configurations of the computer system 510 are possible, having more or fewer components than the computer system depicted in Figure 5.
[0057] Wherever the systems described herein collect or can use personal information about a user (or, as often referred to herein, “Participant”), the user may be given the opportunity to control whether the program or features collect user information (e.g., information about the user’s social networks, social behavior or activities, occupation, preferences, or current geographical location), or whether and / or how they receive content from a content server that may be more relevant to the user. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, user identification information may be processed so that personally identifiable information cannot be identified to that user, or, if geographical location information is obtained, the user’s geographical location may be generalized (to the city level, zip code level, or state level, etc.) so that the user’s specific geographical location cannot be identified. Thus, users may have control over how information about them is collected and / or used.
[0058] While several implementations have been described and illustrated herein, a variety of other means and / or structures can be used to perform the function and / or to obtain one or more of the results and / or advantages described herein, and such variations and / or modifications are each considered to fall within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended to be illustrative, and the actual parameters, dimensions, materials, and / or configurations will depend on one or more specific applications in which this teaching is used. Those skilled in the art will recognize many equivalents of the specific implementations described herein, or can verify them using only ordinary experiments. Therefore, it should be understood that the aforementioned implementations are presented as examples only, and that implementations can be carried out in ways other than those specifically described and claimed, within the scope of the appended claims and their equivalents. The implementations of this disclosure cover each individual feature, system, article, material, tool, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, sets of tools, and / or methods is included within the scope of this disclosure, provided that such features, systems, articles, materials, sets of tools, and / or methods are not inconsistent with each other.
[0059] In some implementations, a method implemented by one or more processors is described as including steps such as determining that the assistant's actions are compatible with an application running on the computing device, and that the application is separate from the automated assistant accessible through the computing device. The method may further include rendering, based on the compatibility of the assistant's actions with the application, selectable graphical user interface (GUI) elements on the computing device's display interface, the selectable GUI elements identifying the assistant's actions, and rendering in the foreground of the computing device's display interface. The method may further include the automated assistant detecting a user's selection of the selectable GUI elements via the computing device's display interface. The method may further include performing speech recognition on audio data capturing oral utterances provided by the user after the selection of the selectable GUI elements and received at the computing device's audio interface, the oral utterances specifying specific values for parameters of the assistant's actions without explicitly identifying the assistant's actions. The method may further include a step of having the automated assistant control the application based on the assistant's actions and specific parameter values in response to verbal utterances from the user.
[0060] In some implementations, rendering selectable GUI elements on the computing device's display interface includes generating content rendered using the selectable GUI elements, the content including a text identifier or graphical representation of the assistant's actions and a placeholder area indicating that the user can specify parameter values. In some implementations, rendering selectable GUI elements on the computing device's display interface includes rendering the selectable GUI elements on the application interface of an application for a threshold duration, the threshold duration being based on the amount of interaction between the user and the application. In some implementations, the automated assistant does not respond to verbal utterances provided by the user after the threshold duration and after the selectable GUI elements are no longer rendered on the computing device's display interface.
[0061] In some implementations, determining that an assistant's actions are compatible with an application running on the computing device includes determining that additional selectable GUI elements rendered in the application interface of the application correspond to possible application actions that can be performed in response to the initialization of the assistant's actions. In some implementations, determining that an assistant's actions are compatible with an application running on the computing device includes determining that additional selectable GUI elements include a search icon or search field, and that the application action corresponds to a search action. In some implementations, allowing the automated assistant to control the application based on the assistant's actions and specific parameter values includes allowing the automated assistant to provide search results based on specific parameter values specified in a user-verbal utterance to the automated assistant.
[0062] In other implementations, a method implemented by one or more processors is described as including steps such as determining that a user has provided a first verbal utterance to an automated assistant accessible via a computing device, and that the first verbal utterance includes a request to initialize an application which is separate from the automated assistant. The method may further include causing the application to initialize in response to the first verbal utterance and causing the application to render an application interface in the foreground of the computing device's display interface, wherein the rendering of the application interface includes content that identifies an action that can be controlled via the automated assistant. The method may further include causing selectable GUI elements to render on the application interface of the application, based on the fact that the action can be controlled via the automated assistant, wherein the rendering of the selectable GUI elements includes a text identifier or graphical representation of an action that can be controlled by the automated assistant. The method may further include determining that a user has provided an automated assistant with a second verbal utterance, the second verbal utterance identifying parameters that may be used by the application during the execution of an action, and the second verbal utterance not explicitly specifying an action. The method may further include, in response to the second verbal utterance, causing the automated assistant to initialize an execution of the action, via the application, using the parameters identified in the second verbal utterance.
[0063] In some implementations, rendering selectable GUI elements on the application interface of an application includes rendering text identifiers using command phrases, wherein the command phrases include words that identify an action and whitespace indicating that user-identifiable parameters are omitted from the command phrases. In some implementations, the method may further include causing the audio interface of a computing device to be initialized so that it can receive specific verbal utterances from a user, based on the fact that the action is controllable via an automated assistant, and when the audio interface is initialized, the user can provide specific verbal utterances to control the automated assistant without explicitly identifying the automated assistant. In some implementations, rendering selectable GUI elements on the application interface of an application includes generating content that is rendered using the selectable GUI elements, wherein the content includes a graphical representation of the assistant's actions, which are selectable via touch input to the display interface of the computing device.
[0064] In some implementations, rendering a selectable GUI element onto the application interface of an application includes rendering the selectable GUI element onto the application interface of an application for a threshold duration, the threshold duration being based on the amount of interaction between the user and the automated assistant since the selectable GUI element was rendered onto the application interface. In some implementations, the automated assistant does not respond to additional verbal utterances provided by the user after the selectable GUI element is no longer rendered onto the application interface. In some implementations, the method may further include the step of initializing the audio interface of a computing device to be able to detect another verbal utterance that identifies one or more parameters about an action, based on the fact that the action is controllable via the automated assistant.
[0065] In yet another implementation, a method implemented by one or more processors is described as including steps such as determining that the assistant's actions are compatible with an application running on a computing device, and that the application is separate from the automated assistant accessible through the computing device. The method may further include rendering, based on the compatibility of the assistant's actions with the application, a selectable graphical user interface (GUI) element on the computing device's display interface, the selectable GUI element identifying the assistant's actions and rendering in the foreground of the computing device's display interface. The method may further include determining, while the selectable GUI element is being rendered on the computing device's display interface, that the user has provided an oral utterance directed to the automated assistant, and that the oral utterance specifies a particular value for a parameter of the assistant's actions without explicitly identifying the assistant's actions. The method may further include, in response to the oral utterance from the user, causing the automated assistant to control the application based on the assistant's actions and specific parameter values.
[0066] In some implementations, rendering selectable GUI elements on the computing device's display interface includes generating content rendered using the selectable GUI elements, the content including icons that are selected based on the assistant's actions and are selectable via touch input to the computing device's display interface. In some implementations, rendering selectable GUI elements on the computing device's display interface includes generating content rendered using the selectable GUI elements, the content including natural language content that characterizes a partial command phrase with one or more parameter values omitted for the assistant's actions. In some implementations, determining that the assistant's actions are suitable for an application running on the computing device includes determining that additional selectable GUI elements rendered by the application control application actions that can be initialized by the automated assistant. In some implementations, allowing the automated assistant to control an application based on the assistant's actions and specific parameter values includes allowing the application to render another application interface that is generated by the application based on specific parameter values. In some implementations, rendering selectable GUI elements on the computing device's display interface includes rendering selectable GUI elements simultaneously with the application rendering one or more application GUI elements of the application. [Explanation of Symbols]
[0067] 100 views 102 users 104 Computing Devices 106 Initialize the home control application 108 in response to user input. 108 Home Control Applications 110 Application Interfaces 112 Thermostat GUI 114 Identify one or more assistant-compatible actions. 116 hands 11 Interfaces Display Interfaces 120 views 122 Selectable GUI elements 124 Oral speech 126 Proposal Elements 140 views 142 Thermostat GUI 144 Initialize Assistant Adaptive Operation 146 The updated thermostat GUI 142 is rendered as an update in the application interface 110. 200 views 202 users 204 Computing Devices 206 Initialized messaging application 208 in response to user input. 208 Messaging applications 210 Application Interfaces 212 checkboxes 214 Render one or more selectable GUI elements 222 and / or one or more selectable suggestions 224 on the computing device 204. 218 Reply icon 220 views 222 Selectable GUI elements 224 selectable proposals 226 Oral speech 240 views 242 selectable GUI elements 244 process 246 Oral speech 248 Application Interfaces 250 selectable GUI elements 300 Systems 302 Computing Devices 304 Automated Assistant 306 Input Processing Engine 308 Speech Processing Engine 310 Data Parsing Engine 312 Parameter Engine 314 Output Generation Engine 316 Motion detection engine 318 GUI element content engine 320 Assistant Interface 322 Assistant Call Engine 324 Operation Execution Engine 326 GUI element duration engine 330 Application Data 332 Device Data 334 applications 336 Context Data 338 Assistant Data 400 ways 500 Block Diagram 510 Computer Systems 512 Bus Subsystem 514 processors 516 Network Interface Subsystem 520 User Interface Output Devices 522 User Interface Input Devices 524 Storage Subsystems 525 memory 526 File Storage Subsystem 530 Main Random Access Memory (RAM) 532 Read-only memory (ROM)
Claims
1. A method implemented by one or more processors, A step of determining whether the assistant's actions are appropriate for an application running on a computing device, The aforementioned application is separate from the automated assistant accessible via the computing device. Steps and A step of rendering selectable elements on the display interface of the computing device based on the fact that the operation of the assistant is suitable for the application, wherein the selectable elements include a text identifier for the operation of the assistant and a placeholder area indicating that the user can specify parameters for the operation of the assistant; The steps include: detecting a touch selection of the selectable element via the display interface of the computing device using the automated assistant; The step of performing speech recognition on audio data capturing spoken utterances received by the audio interface of the computing device after the touch selection of the selectable elements, The verbal utterance specifies a particular value for the parameter of the assistant's action without explicitly identifying the assistant's action. Steps and The step includes, in response to the verbal utterance, causing the automated assistant to control the application based on the assistant's actions and the specific values of the parameters, When controlling the application based on the assistant's actions and the specific value, the step of causing the automated assistant to use the assistant's actions is based on the touch selection of the selectable element and the selectable element including the text identifier of the assistant's actions, When controlling the application based on the assistant's actions and the specific value, the step of causing the automated assistant to use the assistant's actions is based on specifying the specific value and the verbal utterance provided after the touch selection of the selectable element. method.
2. The step of rendering the selectable elements on the display interface of the computing device is: The step includes rendering the selectable elements on the application interface of the application for a threshold duration. The method according to claim 1.
3. The process further includes the step of determining the threshold duration based on the amount of interaction between the user and the application. The method according to claim 2.
4. The method according to claim 3, wherein the automated assistant does not respond to the verbal utterance when the verbal utterance is provided by the user after the threshold duration and after the selectable elements are no longer rendered on the display interface of the computing device.
5. The step of determining whether the operation of the assistant is compatible with the application running on the computing device is: The step of determining that additional selectable elements rendered in the application interface of the application correspond to possible application actions that can be performed in response to the initialization of the assistant's actions. including The method according to claim 1.
6. The step of determining whether the operation of the assistant is compatible with the application running on the computing device is: The step of determining that the additional selectable element includes a search icon or a search field, and that the application action corresponds to a search action. The method according to claim 5, including the method described in claim 5.
7. The step of causing the automated assistant to control the application based on the assistant's actions and the specific values of the parameters is: The step involves causing the automated assistant to cause the application to perform a search based on the specific value of the parameter specified in the verbal utterance from the user to the automated assistant. The method according to claim 6, including the method described in claim 6.
8. A method implemented by one or more processors, A step of determining that a user has provided a first verbal utterance to an automated assistant accessible via a computing device, The first oral utterance includes a request to initialize an application that is separate from the automated assistant, Steps and A step of initializing the application in response to the first oral utterance and causing the application to render an application interface in the foreground of the display interface of the computing device, The application interface includes content that identifies actions that can be controlled via the automated assistant, Steps and A step of rendering selectable elements on the display interface, based on the fact that the operation can be controlled via the automatic assistant, The selectable element includes a text identifier for the action that can be controlled by the automated assistant, and a placeholder area indicating that the user can specify parameters for the action. Steps and The step of determining to the automated assistant that the user provided a touch selection of the selectable elements before the second verbal utterance, The second oral utterance identifies a specific value for the parameter that can be used by the application during the execution of the operation, and The second oral utterance does not clearly specify the action, Steps and A step of causing the automated assistant to initialize, in response to the second oral utterance, the execution of the action via the application using the specific value identified in the second oral utterance; Includes, The step of causing the automatic assistant to use the specific value when initializing the execution of the action using the parameter is based on the touch selection of the selectable element and the selectable element including the text identifier of the action, The step of causing the automated assistant to use the specific value when initializing the execution of the operation using the parameter includes identifying the specific value and based on the verbal utterance provided after the touch selection of the selectable element. method.
9. The step of rendering the selectable elements on the display interface includes the step of rendering the selectable elements on the application interface. The method according to claim 8.
10. A step of initializing the audio interface of the computing device so that it can receive specific verbal utterances from the user, based on the fact that the operation described above can be controlled via the automated assistant, When the audio interface is initialized, the user can provide the specific verbal utterance to control the automated assistant without explicitly identifying the automated assistant. The method according to claim 8, further comprising the step.
11. The step of rendering the selectable elements on the display interface is The step includes having the application render one or more application elements of the application, while simultaneously rendering the selectable elements, The method according to claim 8, including the step.
12. The step of rendering the selectable elements on the display interface is The selectable elements are rendered on the application interface of the application for a threshold duration. The method according to claim 8.
13. If additional verbal utterances are provided by the user after the selectable elements are no longer rendered on the application interface, the automated assistant will not respond to the additional verbal utterances. The method according to claim 12.
14. A method implemented by one or more processors, A step of determining whether the assistant's actions are appropriate for an application running on a computing device, The aforementioned application is separate from the automated assistant accessible via the computing device. Steps and A step of rendering selectable elements on the display interface of the computing device, based on the operation of the assistant being suitable for the application, The selectable element includes a text identifier for the assistant's action and a placeholder area indicating that the user can specify parameters for the assistant's action. Steps and The step of determining that the user has provided an oral utterance directed to the automated assistant when the selectable elements are being rendered on the display interface of the computing device, The verbal utterance specifies a particular value for the parameter of the assistant's action without explicitly identifying the assistant's action. Steps and The steps include determining that the specific value specified in the verbal utterance is associated with the parameter of the assistant's behavior, which is identified by the selectable element rendered in the foreground of the display interface, The steps include: causing the automated assistant to control the application based on the assistant's actions and the specific values of the parameters, in response to the verbal utterance from the user and the determination that the specific values are associated with the parameters of the assistant's actions as identified by the selectable elements; Methods that include...
15. The step of determining whether the operation of the assistant is compatible with the application running on the computing device is: The step of determining that additional selectable elements rendered by the application control application behavior that can be initialized by the automated assistant. including, The method according to claim 14.
16. The step of causing the automated assistant to control the application based on the assistant's actions and the specific values of the parameters is: The step of causing the application to render another application interface that is generated by the application based on the specific value of the parameter. The method according to claim 14, including the method described in claim 14.
17. The step of rendering the selectable elements on the display interface of the computing device is: The step of rendering the selectable elements at the same time that the application renders one or more application elements of the application. The method according to claim 14, including the method described in claim 14.
18. A computing device comprising a display interface, an audio interface, a memory for storing instructions, and one or more processors, wherein the one or more processors are capable of executing the instructions that implement the method according to any one of claims 1 to 17. Computing device.
Citation Information
Patent Citations
Voice control method and electronic equipment
CN109584879A
Speech recognition device and its setup method
JP2007127813A
Initiating a conversation with an automated agent through selectable graphical elements
JP2020520507A
Initializing a conversation with an automated agent via selectable graphical element
US20180324115A1
Initializing non-assistant background actions, via an automated assistant, while accessing a non-assistant application
US20200395018A1