Providing context-based automated assistant action suggestions (multiple suggestions possible) via vehicle computing devices.

An automated assistant in vehicles suggests actions to control applications, addressing manual interaction waste and distraction by enabling hands-free operation.

JP7897324B2Inactive Publication Date: 2026-07-29GOOGLE LLC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2022-07-01
Publication Date
2026-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Users interacting with vehicle computing devices often manually control applications, wasting resources and increasing distraction by diverting attention from driving.

Method used

An automated assistant provides context-based suggestions for controlling applications via vehicle and mobile computing devices, reducing manual input and distraction by offering hands-free interaction.

Benefits of technology

Reduces computing resource consumption and user distraction by allowing hands-free control of applications through automated assistant suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007897324000001
    Figure 0007897324000001
  • Figure 0007897324000002
    Figure 0007897324000002
  • Figure 0007897324000003
    Figure 0007897324000003
Patent Text Reader

Abstract

The embodiments described herein relate to an automated assistant that may provide suggestions for a user to interact with the automated assistant to control an application while in a vehicle. Suggestions may be provided to encourage hands-free interaction with the application by suggesting assistant inputs that invoke the automated assistant to act as an interface between the user and the application. The assistant suggestions may be based on the user's context and / or the vehicle's context, such as the content of a display interface of a device that the user is accessing while in the vehicle. For example, the automated assistant may identify that an action performed by the user using the application can be initiated more safely and / or in less time by utilizing a particular assistant input. The particular assistant input may then be rendered in an interface of the vehicle computing device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the provision of context-based automatic assistant action suggestions (multiple possible) via a vehicle computing device.

Background Art

[0002] Humans can interact with computers using an interactive software application, which is referred to herein as an "automatic assistant" (also referred to as a "digital agent", "chatbot", "interactive personal assistant", "intelligent personal assistant", "assistant application", "conversation agent", etc.). For example, a human (who may be referred to as a "user" when interacting with an automatic assistant) can provide commands and / or requests to an automatic assistant using natural language input by voice (i.e., speech) and / or by providing natural language input by text (e.g., typed input). The natural language input by voice can, in some cases, be processed after being converted to text.

[0003] In some cases, automated assistants can be accessed within the vehicle via an integrated vehicle computing device, which can also provide access to other applications. While automated assistants can offer many benefits to users, users may not be fully aware of all of their capabilities. Specifically, automated assistants can provide functionality that assists users by allowing hands-free control of other applications available through the vehicle computing device. However, users may only invoke and utilize the automated assistant when trying to control functions that they perceive as being exclusive to the automated assistant, and may manually control the functions of other applications without invoking or utilizing the automated assistant. Manually controlling the functions of these applications in this way can increase the amount of user input received via the vehicle computing device, thereby wasting the computing resources of the vehicle computing device. Furthermore, manually controlling the functions of these applications in this way can also increase user distraction while driving, as manual interaction with the vehicle computing device while driving can divert the user's attention from driving.

[0004] For example, a user making a phone call via a vehicle phone application may need to stare at and tap a graphical keypad rendered on the vehicle computing device's display interface to input each digit of the phone number or each letter of the contact's name. Similarly, a user attempting to stream media from the internet may need to navigate the media application by staring at and tapping its GUI elements to begin playback of the desired media. Thus, manually controlling the functionality of these applications not only unnecessarily wastes the vehicle computing device's computing resources but also increases user distraction. [Overview of the project] [Means for solving the problem]

[0005] Embodiments described herein relate to an automated assistant, which may provide suggestions for assistant inputs to be provided to the automated assistant to control a particular other application when the user is in a vehicle and is currently controlling or is expected to control that particular other application. The vehicle may include a vehicle computing device that can provide access to a variety of different applications, including an automated assistant application associated with the automated assistant. Furthermore, an automated assistant accessible via the vehicle computing device may correspond to another automated assistant accessible via one or more other devices, such as a mobile computing device (e.g., a mobile phone). In some embodiments, the automated assistant may operate as an interface between the user and a given vehicle application on the vehicle computing device, and / or as an interface between the user and a given mobile application on a mobile computing device. In certain cases, the automated assistant may provide suggestions for assistant inputs that the user may submit to the automated assistant via the vehicle computing device, and when submitted, the assistant inputs may cause the automated assistant to control one or more functions of these applications. This can streamline certain interactions between the user and applications by reducing the amount of input that may be used to control the functions of a particular application through manual interaction with that application. Furthermore, this can reduce distraction for the user while driving or riding in a vehicle by allowing them to control various applications by relying on hands-free interaction with the automated assistant.

[0006] For example, a user can ride while interacting with an application. The application could be, for instance, a medical application, allowing the user to interact with it to schedule an appointment with their primary care physician. The vehicle may include a vehicle computing device, which may provide access to the automated assistant application and the medical application, and may include a display interface. The automated assistant may also be accessible via the user's mobile phone, which may include an automated assistant application corresponding to the automated assistant accessible via the vehicle computing device. Furthermore, the medical application may also be accessible via the user's mobile phone. Therefore, with prior permission from the user, the automated assistant can identify, via the vehicle computing device and / or the user's mobile phone, that the user is interacting with the medical application while riding, process contextual data, and provide context-based automated assistant suggestion actions to the user via the vehicle computing device's display interface. These context-based automated assistant suggestion actions could, for example, teach the user how to interact with the medical application hands-free via the automated assistant.

[0007] In this embodiment, context data may include, for example, context data associated with the user, context data associated with the user's mobile phone, and / or context data associated with the vehicle computing device, such as screenshots of a medical application in which the user is interacting via the user's mobile phone and / or vehicle computing device, the vehicle's destination, the vehicle's current location, the user's identifier (with prior permission from the user), data characterizing the functionality of recent interactions between the user and one or more applications, and / or any other relevant data that may be processed by the automated assistant. Furthermore, the context data may be processed to identify one or more actions that the user is currently engaged in and / or is expected to initiate via the medical application. For example, a screenshot of the medical application may include graphics and / or text characterizing a graphical user interface (GUI) for scheduling an appointment with a doctor. The data characterizing this screenshot may be processed using one or more heuristic processes and / or one or more trained machine learning models to identify one or more actions that can be controlled via the GUI (e.g., creating an appointment, canceling an appointment, rescheduling an appointment, etc.). Based on this identification, one or more specific actions may be selected as the basis for generating proposed data, which may characterize assistant inputs that the user may submit to an automated assistant to control one or more specific actions of the medical application, instead of the user manually interacting with the medical application.

[0008] In this embodiment, the suggestion data generated by the automated assistant may characterize utterances, such as, "Try saying, 'Assistant, schedule an appointment with Dr. Chow for Tuesday at 3 p.m.'" The automated assistant may visually render the text of the utterance as a graphic element on the display interface of the vehicle computing device, simultaneously with and / or after the user interacts with the medical application. In response to the user providing the rendered utterance, the graphic element may be displayed before, during, and / or after the user interacts with the medical application, with respect to the current interaction with the medical application, and / or in place of future interactions with the medical application. The automated assistant can then notify the user that it can perform a specific action (e.g., scheduling an appointment). For example, the user may provide a utterance to the automated assistant when they look at a suggestion rendered on the display interface of the vehicle computing device. The voice embodying the utterance is received by the audio interface of the vehicle computing device and / or the mobile phone (e.g., one or more microphones), thereby allowing an instance of the automated assistant to interact with a medical application. This allows the automated assistant to schedule the user's appointment in the medical application without the user having to manually interact with the mobile phone and / or the vehicle computing device while in the vehicle (e.g., directly tapping the mobile phone's touch interface).

[0009] In some embodiments, suggestions for assistant input may be rendered after the user has interacted with the application. For example, in the above embodiment, it is assumed that the user manually interacts with a medical application to schedule an appointment. Furthermore, it is assumed that through this manual interaction the user successfully schedules the appointment, and the appointment confirmation is displayed via a mobile phone and / or vehicle computing device. In this case, the user may receive a suggestion from the automated assistant while in the vehicle, when the confirmation is displayed. A suggestion such as, "Next time, you can say, 'Assistant, schedule an appointment with Dr. Chow for Tuesday at 3 p.m.'" may characterize the utterance in such a way that, when the utterance is provided to the automated assistant in the future, the automated assistant will control the functions of the medical application that the user may have previously used while in the vehicle.

[0010] In some embodiments, suggestions for assistant input may be rendered before the user interacts with another application. For example, a user who interacts with a particular application while driving on one day may receive a suggestion from the automated assistant while driving on another day. The suggestion may characterize the utterance so that, once the utterance is provided to the automated assistant, the automated assistant controls the functionality of an application that the user may have previously used while driving. For example, if a user previously accessed a podcast application and played "medical podcasts" while driving, the automated assistant may make a suggestion such as "Assistant, play medical podcasts" while driving. Alternatively or additionally, applications accessed directly via the vehicle computing device may be controlled by the automated assistant with prior permission from the user and thus subject to assistant suggestions. For example, a user expected to interact with a vehicle maintenance application on the vehicle computing device may receive suggestions via the vehicle computing device interface regarding control of the functionality of the vehicle maintenance application. For example, a user may typically access the vehicle maintenance application a few minutes after starting a drive that exceeds a threshold distance (e.g., 100 miles) to check if there is a charging station near their destination. Based on this contextual data, the next time the user chooses to navigate to a destination that is beyond a threshold distance (for example, more than 100 miles away), the automated assistant can render suggestions on the vehicle computing device, such as "Assistant, show me charging stations near my destination," before the user accesses the vehicle maintenance application.

[0011] In some embodiments, a given suggestion for a given assistant input is rendered to present to a given user a threshold number of times, which can reduce the amount of computational resources spent generating suggestions and alleviate user discomfort. For example, the suggestion in the above embodiment, "Next time, you can say 'Assistant, schedule an appointment with Dr. Chow for Tuesday at 3 p.m.'" may be presented to the user only once to train the user regarding the automated assistant's functionality. Therefore, if the user subsequently initiates interaction with the medical application via the vehicle computing device to schedule the next appointment, the suggestion may not be provided. However, if the user subsequently initiates interaction with the medical application via the mobile phone and / or vehicle computing device to cancel a previously scheduled medical appointment, an additional suggestion, "Next time, you can say 'Assistant, cancel an appointment with Dr. Chow'," may be generated and presented to the user in the same or similar manner as described above.

[0012] In additional or alternative embodiments, a given suggestion for a given application may be rendered to present to a given user a threshold number of times, thereby reducing the amount of computational resources spent generating suggestions and mitigating user discomfort. For example, the suggestion in the above embodiment, "Next time, you can say, 'Assistant, schedule an appointment with Dr. Chow for Tuesday at 3 p.m.'" may be presented to the user only once to train the user regarding the functionality of the automated assistant. Furthermore, the suggestion may additionally include other suggestions regarding the medical application. For example, the suggestion may additionally include the suggestion, "You can also say, 'Assistant, cancel the appointment' or 'Assistant, reschedule the appointment'," or any further functionality that the automated assistant can perform with respect to the medical application. In this example, the automated assistant actively trains the user regarding multiple assistant inputs that may be provided to allow the automated assistant to control various functions of the medical application.

[0013] By using the techniques described herein, various technical advantages can be achieved. In a non-limiting embodiment, the techniques described herein enable a system to provide context-relevant automated assistant action suggestions in a vehicle environment to reduce the consumption of computing resources and / or reduce driver distraction. For example, the techniques described herein can detect interactions between the user's vehicle computing device and / or the user's mobile computing device while the user is inside the vehicle. Furthermore, the techniques described herein can identify contextual information associated with the interaction. When the user completes an interaction and / or when an interaction is anticipated to begin, the system can generate and provide suggestions based on the interaction and / or the contextual information associated with the interaction, enabling hands-free initiation and completion of the interaction. As a result, computing resources can be saved by reducing the amount of user input required to achieve the interaction, and user distraction can be reduced by eliminating the need for user input.

[0014] The above description is provided as an overview of some embodiments of this disclosure. Further descriptions of these embodiments and other embodiments will be given in more detail later. [Brief explanation of the drawing]

[0015] [Figure 1A] The diagram shows a user receiving suggestions for assistant inputs that can be used to give instructions to an automated assistant in order to safely interact with the application while the user is inside the vehicle, in various embodiments. [Figure 1B] The diagram shows a user receiving suggestions for assistant inputs that can be used to give instructions to an automated assistant in order to safely interact with the application while the user is inside the vehicle, in various embodiments. [Figure 1C]The diagram shows a user receiving suggestions for assistant inputs that can be used to give instructions to an automated assistant in order to safely interact with the application while the user is inside the vehicle, in various embodiments. [Figure 1D] The diagram shows a user receiving suggestions for assistant inputs that can be used to give instructions to an automated assistant in order to safely interact with the application while the user is inside the vehicle, in various embodiments. [Figure 2] The present invention relates to a system that provides an automated assistant in various embodiments, wherein the automated assistant can be invoked to generate suggestions for controlling applications that may distract the user's attention from driving a vehicle. [Figure 3] This document describes a method for providing an automated assistant that can generate suggestions for assistant inputs that a user may provide, in order to promote the control of a separate application by the automated assistant while the user is in the vehicle, through various embodiments. [Figure 4] This is a block diagram of an exemplary computer system in various embodiments. [Modes for carrying out the invention]

[0016] Figures 1A, 1B, 1C, and 1D illustrate scenarios 100, 120, 140, and 160, respectively, in which user 102 receives suggestions for assistant inputs available to give instructions to the automated assistant in order to safely interact with the application while inside the vehicle 108. In some embodiments, user 102 can interact with the application via a mobile computing device 104 or a portable computing device 104 (e.g., user 102's mobile phone) and / or a vehicle computing device 106 in the vehicle 108. The application may be separate from the automated assistant, which may be accessible via the portable computing device 104 and / or the vehicle computing device 106 by respective instances of the automated assistant application running on them. User 102 can access each instance of the application via the portable computing device 104 while User 102 is driving and / or riding in the vehicle 108, and / or via the vehicle computing device 106 while User 102 is driving and / or riding in the vehicle 108. The automated assistant acts as an interface between User 102 and the application instances to facilitate safe driving and can also reduce the number of inputs that may be required to control each instance of the application.

[0017] For example, as shown in Figure 1A, when user 102 is driving vehicle 108, user 102 can access an instance of an Internet of Things (IoT) application via a portable computing device 104. In this embodiment, user 102 can interact with the instance of the IoT application via the application interface of the IoT application to change the temperature setting of a residential automatic temperature controller. Before and / or during the interaction between user 102 and the portable computing device 104, the automated assistant may determine that user 102 is inside vehicle 108 and interacting with an instance of the IoT application. The automated assistant may make this determination by processing data from various different sources. For example, with prior permission from user 102, the presence of the user inside vehicle 108 may be determined by using data based on one or more sensors inside vehicle 108 (e.g., using facial recognition, voice signature, touch input, etc.). Alternatively or additionally, with prior permission from user 102, data characterizing the screen content 110 of the portable computing device 104 may be processed to identify one or more functions of an IoT application and / or instances of one or more applications running in the background of the portable computing device 104 that user 102 can control and / or attempt to control.

[0018] For example, screen content 110 may include a GUI element for controlling the temperature of user 102's living room. Using one or more heuristic processes and / or one or more trained machine learning models, screen content 110 may be processed to identify one or more different controllable functions of screen content 110. Using the identified functions, suggestions for assistant inputs that user 102 can provide to the automated assistant may be generated to invoke the automated assistant to control the identified functions(s) of the instance of the IoT application. In some implementations, suggestions may be generated using an application programming interface (API) with respect to data communicated between the instance of the IoT application and the automated assistant. Alternatively or additionally, suggestion data may be generated using content accessible through one or more interfaces of the instance of the IoT application. For example, the functions of screen content 110 may be selectable GUI elements, which may be rendered in relation to natural language content (e.g., "72 degrees"). In some embodiments, a GUI element may be identified as selectable based on data available to the automated assistant, such as HTML code, XML code, Document Object Model (DOM) data, Application Programming Interface (API) data, and / or any other information that may indicate whether a particular feature of the application interface of an instance of an IoT application is selectable.

[0019] The portion of screen content 110 occupied by selectable GUI elements is identified along with natural language content and can be used when generating suggestions for assistant inputs presented to user 102. In this way, once user 102 has provided the assistant inputs included in the suggestions back to the automated assistant, the automated assistant may interact with the application interface of the IoT application to control the suggested function and / or communicate API data to the IoT application to control the suggested function. For example, as shown in Figure 1B, before user 102 accesses an instance of the IoT application, while user 102 is accessing an instance of the IoT application, and / or after user 102 has performed some action through an instance of the IoT application, the automated assistant may render an assistant suggestion 124 on the display interface 122 of the vehicle computing device 106. For example, suppose user 102 manually interacts with an instance of the IoT application to change the temperature of user 102's house. After user 102 manually interacts with an instance of the IoT application, the automated assistant may render an assistant suggestion 124 containing natural language content corresponding to the suggested assistant inputs. A suggested assistant input (i.e., a suggested command phrase) could be, "Next time, try saying, 'Assistant, lower the temperature in the home control application'," which could refer to a request to the automated assistant to automatically change the temperature value managed by an instance of the IoT application (i.e., the "home control application"). For example, assistant suggestion 124 might be stored in relation to a command submitted to the IoT application by the automated assistant, such as "lowerTemp(HomeControlApp, Assistant, [0, -3 degrees])."The automated assistant may additionally or alternatively offer other suggestions associated with the instance of the IoT application, such as saying "(1) 'Assistant, open the garage' and (2) 'Assistant, deactivate the alarm system'," and / or offer any other actions that are not directly related to the user 102's manual interaction with the instance of the IoT application, but which the automated assistant can perform by interfaceing with the instance of the IoT application on behalf of the user 102.

[0020] In some embodiments, as shown in Figure 2, Proposal 124 may be presented to User 102 on the display interface 122 of the vehicle computing device 106. In additional or alternative embodiments, Proposal 124 may be presented additionally or alternatively on the display interface of the portable computing device 104. In particular, in some embodiments, User 102 may not immediately utilize Proposal 124. For example, when User 102 is in the vehicle 108 and looking at the portable computing device 104, User 102 may look at the Assistant Proposal 124 provided on the display interface 122 of the vehicle computing device 106, recognize the ability of the automated assistant to control IoT applications, and then put down their phone. In this embodiment, User 102 may utilize Proposal 124 the next time they want to use the automated assistant to control one or more IoT devices. Alternatively or additionally, user 102 may see the assistant suggestion 124 during their first trip in vehicle 108 (e.g., a long trip), but may not utilize the suggestion 124 until their second or subsequent trip in vehicle 108. In other words, user 102 may drive to their first destination during their first trip and be presented with the assistant suggestion 124, but may not utilize the suggestion 124 until their second trip. Thus, user 102 can then use the natural language content of the suggestion 124 to have the automated assistant perform an action on their behalf, rather than manually interacting with the various applications described herein.

[0021] For example, as shown in FIG. 1C, user 102 may provide utterance 142 corresponding to assistant suggestion 124 during the same movement as shown in FIGS. 1A and 1B or during a subsequent movement (e.g., during a movement on the day following the days of FIGS. 1A and 1B). Utterance 142 may be "Assistant, lower the temperature in the living room", which may embody a call phrase (e.g., "Assistant") and a request or command phrase (e.g., "lower the temperature in the living room"). In particular, utterance 142 provided by user 102 is similar to assistant suggestion 124 that included a call phrase (e.g., "Assistant") and a request or command phrase (e.g., "lower the temperature of the home control application") as shown in FIG. 1B. In some embodiments, user 102 provides an utterance that is similar but not identical to the natural language content of assistant suggestion 124, but can cause an operation corresponding to assistant suggestion 124 to be performed. For example, when user 102 provides utterance 142, context data associated with the instance can be processed to generate a similarity value that can quantify the similarity between the previous instance (i.e., the previous context) when assistant suggestion 124 was previously rendered and the context data.

[0022] When the similarity values ​​satisfy the similarity threshold, the automated assistant may respond to the utterance 142 by initiating the execution of one or more actions corresponding to the assistant suggestion 124 (e.g., lowering the temperature in the living room of user 102 associated with the home control application). For example, as shown in Figure 1D, the automated assistant may respond to the utterance 142 by initiating the execution of one or more actions that advance the performance of the request or command phrase, and optionally by rendering an assistant output 162 indicating that one or more of the actions have been performed. The assistant output 162 may be an audible and / or text output that informs the user that the automated assistant has performed the request or command phrase. For example, the assistant output 162 may be text and / or audio that embodies the message “The living room temperature has been changed via the home control application,” as shown in Figure 1C. In some embodiments, training data may be generated based on this interaction between user 102, the automated assistant, and / or instances of the IoT application to advance the training of one or more trained machine learning models. One or more trained machine learning models can then be used in subsequent contexts to process contextual data, potentially providing user 102 with suggestions to safely streamline interactions with automated assistants and / or other applications.

[0023] The embodiments described above with respect to FIGS. 1A, 1B, 1C, and 1D are described with regard to user 102 interacting with an IoT application via a portable computing device 104, but this is for example purposes and is not meant to be limiting. For example, manual interaction with an instance of an IoT application may be performed via a vehicle computing device 106 rather than the described portable computing device 104. In this case, proposal 124 may be provided via the vehicle computing device 106 in the same or similar manner as described above. Also, for example, user 102 may additionally or alternatively interact with other applications via the portable computing device 104 and / or the vehicle computing device 106, and the proposal may be generated for these other applications in the same or similar manner. For example, user 102 may interact with an instance of a media application via the vehicle computing device 106 to play a song. In some of these embodiments, when user 102 opens an instance of a music application, a proposal such as "Tell me the song or artist you want to hear" may be presented to user 102 audibly and / or visually. Thus, rather than navigating through an instance of a media application to play music, user 102 may simply provide an indication of the song or artist they want to hear. In other embodiments, the proposal may not be provided until the user manually selects the song or artist they want to hear and that song or artist is played.

[0024] Figure 2 shows a system 200 providing an automated assistant 204, which may offer suggestions to invoke the automated assistant 204 to control an application that the user can access from the vehicle (e.g., via the portable computing device 104 in Figures 1A–1D and / or via the vehicle computing device 106). The automated assistant 204 may operate as part of an automated assistant application provided on one or more computing devices, such as computing device 202 and / or a server device. The user may interact with the automated assistant 204 via an automated assistant interface(s) 220, which may be a microphone, camera, touchscreen display, user interface, and / or any other device capable of providing an interface between the user and the application. For example, the user may initiate the automated assistant 204 by providing verbal, text, gesture, and / or graphic input to the assistant interface 220 to cause the automated assistant 204 to initiate one or more actions (e.g., providing data, controlling peripheral devices, accessing agents, generating input and / or output, and / or other actions). Alternatively, the automated assistant 204 may be initiated based on processing context data 236 using one or more trained machine learning models and / or heuristic processes. The context data 236 may characterize one or more features of the environment accessible to the automated assistant 204 (e.g., one or more features of the computing device 202 and / or one or more features of the vehicle in which the user is located) and / or one or more features in which the user is expected to intend to interact with the automated assistant 204. The computing device 202 may include a display device, which may be a display panel including a touch interface for receiving touch input and / or gestures that enable the user to control the application 234 of the computing device 202 via a touch interface.In some embodiments, the computing device 202 may lack a display device and therefore does not provide a graphical user interface output, but rather an audible user interface output. Furthermore, the computing device 202 may provide a user interface, such as a microphone, to receive natural language input spoken by the user. In some embodiments, the computing device 202 may include a touch interface and may lack a camera, but may optionally include one or more other sensors.

[0025] The computing device 202 and / or other third-party client devices (e.g., third-party client devices provided by the entity in addition to the entity providing the computing device 202 and / or the automated assistant 204) may communicate with the server device via a network such as the Internet. Furthermore, the computing device 202 and any other computing devices may communicate with each other via a local area network (LAN), such as a Wi-Fi network or a Bluetooth network. The computing device 202 may offload computing tasks to the server device in order to conserve computing resources on the computing device 202. For example, the server device may host the automated assistant 204, and / or the computing device 202 may transmit inputs received by one or more assistant interfaces 220 to the server device. However, in some embodiments, the automated assistant 204 may be hosted on the computing device 202, and various processes that may be associated with the automated assistant operation may be executed on the computing device 202.

[0026] In various embodiments, all or fewer embodiments of the automated assistant 204 may be implemented on the computing device 202. In some of these embodiments, embodiments of the automated assistant 204 may be implemented via the computing device 202 and interface with a server device, which may implement other embodiments of the automated assistant 204. The server device may optionally provide services to multiple users and their associated automated assistant applications via multiple threads. In embodiments in which all or fewer embodiments of the automated assistant 204 are implemented via the computing device 202, the automated assistant 204 may be an application separate from the operating system of the computing device 202 (e.g., installed "on top of" the operating system), or it may be implemented directly by the operating system of the computing device 202 (e.g., an application of the operating system, but may be considered integrated with the operating system).

[0027] In some embodiments, the automated assistant 204 may include an input processing engine 206, which may process inputs and / or outputs of the computing device 202 and / or server device using a number of different modules. For example, the input processing engine 206 may include a speech processing engine 208, which may process audio data received by the assistant interface 220 to identify text corresponding to utterances embodied in the audio data. To conserve computing resources in the computing device 202, the audio data may be sent, for example, from the computing device 202 to the server device. Additionally or alternatively, the audio data may be processed only in the computing device 202.

[0028] The process for converting audio data to text may include an automatic speech recognition (ASR) algorithm, which may use a neural network and / or a statistical model to identify groups of audio data corresponding to words or phrases. The text converted from the audio data (or text received as text input) may be parsed by the data analysis engine 210 and made available to the automated assistant 204 as text data that can be used to generate and / or identify command phrases, intents, actions, slot values, and / or any other content specified by the user. In some embodiments, the output data provided by the data analysis engine 210 may be provided to the parameter engine 212 to determine whether the user has provided input corresponding to specific intents, actions, and / or routines that can be executed by the automated assistant 204 and / or applications or agents accessible through the automated assistant 204. For example, assistant data 238 may be stored in a server device and / or computing device 202 and may include data defining one or more actions that the automated assistant 204 can perform, as well as the parameters required to perform the actions. The parameter engine 212 may generate one or more parameters relating to intents, actions, and / or slot values, and provide one or more parameters to the output generation engine 214. The output generation engine 214 may use one or more parameters to communicate with the assistant interface 220 to provide output to the user, and / or with one or more applications 234 to provide output to one or more applications 234.

[0029] An automated assistant application may include and / or have access to on-device ASR, on-device natural language understanding (NLU), and on-device fulfillment. For example, on-device ASR may be performed using an on-device ASR module, which processes audio data (detected by the microphone(s)) using an end-to-end speech recognition machine learning model stored locally on computing device 202, for example. The on-device ASR module generates recognized text with respect to the utterances (if any) present in the audio data. Furthermore, on-device NLU may be performed using an on-device NLU module, which processes the recognized text generated using on-device speech recognition and optionally context data to generate NLU data. The NLU data may include intent(s) corresponding to the utterances and optionally parameters(s) of the intent(s) (e.g., slot values). Furthermore, on-device fulfillment may be performed using an on-device fulfillment module, which utilizes NLU data (from the on-device NLU module) and optionally other local data to identify the actions (and optionally parameters) to take to resolve the intent(s) of the utterance. This may include identifying local and / or remote responses (e.g., answers) to the utterance, interactions (multiple) to perform with locally installed applications (multiple) based on the utterance, commands (multiple) to send (directly or via corresponding remote systems (multiple)) to Internet of Things (IoT) devices (multiple) based on the utterance, and / or other resolution actions (multiple) to perform based on the utterance. The on-device fulfillment module may then initiate local and / or remote performance / execution of the identified actions (multiple) to resolve the utterance.

[0030] In various embodiments, remote ASR, remote NLU, and / or remote fulfillment may be used, at least selectively. For example, recognition text generated by an on-device ASR module may be sent, at least selectively, to a remote automated assistant component(s) for remote NLU and / or remote fulfillment. For example, recognition text generated by an on-device ASR module may optionally be sent for remote performance in parallel with on-device performance, or in response to failures of on-device NLU and / or on-device fulfillment. However, on-device ASR, on-device NLU, on-device fulfillment, and / or on-device execution may be preferred, at least because they reduce latency when resolving utterances (by eliminating the need for round trips between client and server to resolve utterances). Furthermore, on-device functionality may be the only functionality available in situations where network connectivity is absent or limited.

[0031] In some embodiments, the computing device 202 may include one or more applications 234 that can be provided by the same first-party entity that provided the computing device 202 and / or the automated assistant 204, and / or one or more applications 234 that can be provided by a third-party entity different from the entity that provided the computing device 202 and / or the automated assistant 204. The application state engine (not shown) of the automated assistant 204 and / or the computing device 202 may access application data 230 to identify one or more actions that can be performed by one or more applications 234, as well as the state of each application 234 and / or the state of each device associated with one or more applications 234. The device state engine (not shown) of the automated assistant 204 and / or the computing device 202 may access device data 232 to identify one or more actions that can be performed by the computing device 202 and / or one or more devices associated with the computing device 202 and / or one or more applications 234. Furthermore, in order to generate context data 236, application data 230 and / or any other data (e.g., device data 232) may be accessed by the automated assistant 204, and the context data 236 may characterize the context in which a particular application and / or device from one or more applications 234 is running, and / or the context in which a particular user is accessing a particular application and / or any other device or module from one or more applications 234 that is accessing the computing device 202.

[0032] While one or more applications 234 are running on computing device 202, device data 232 may characterize the current operational state of each of the one or more applications 234 running on computing device 202. Furthermore, application data 230 may characterize one or more functions of the running one or more applications 234, such as the content of one or more graphical user interfaces rendered at the direction of one or more applications 234. Alternatively or additionally, application data 230 may characterize action schemas, which may be updated by each of the one or more applications 234 and / or by the automated assistant 204 based on the current operational state of each application. Alternatively or additionally, one or more action schemas of the one or more applications 234 may be static states but may be accessed by the application state engine to identify actions suitable for initiating via the automated assistant 204.

[0033] The computing device 202 may further include an assistant invocation engine 222, which may use one or more trained machine learning models to process inputs received via the assistant interface 220, application data 230, device data 232, context data 236, and / or any other data accessible by the computing device 202. The assistant invocation engine 222 may process this data to determine whether to wait for the user to explicitly utter an invocation phrase to invoke the automated assistant 204 via the assistant interface 220, or whether to consider the data to indicate the user's intent to invoke the automated assistant, instead of requiring the user to explicitly utter an invocation phrase to invoke the automated assistant 204. For example, one or more trained machine learning models may be trained using instances of training data based on a scenario in which the user is in an environment where multiple devices and / or applications exhibit various operating states. Instances of training data may be generated to capture training data that characterizes the context in which the user invokes the automated assistant and other contexts in which the user does not invoke the automated assistant 204. When one or more trained machine learning models are trained according to instances of this training data, the assistant invocation engine 222 allows the automated assistant 204 to detect, or limit, invocation phrases uttered by the user based on the context and / or the capabilities of the environment.

[0034] In some embodiments, the system 200 may further include an interaction analysis engine 216, which may process various data to determine whether an interaction between the user and the application should be subject to an assistant suggestion. For example, the interaction analysis engine 216 may process data indicating the number of user inputs provided by the user to the application 234 so that the application 234 can perform a specific action while the user is in their vehicle. Based on this processing, the interaction analysis engine 216 may determine whether the automated assistant 204 can perform a specific action even with less user input. For example, a user who switches a corresponding instance of a given application on a portable computing device to render a specific medium on a vehicle computing device may be able to cause the automated assistant 204 to render that specific medium by uttering a specific utterance (e.g., one input).

[0035] Furthermore, based on the interaction analysis engine 216's determination that a particular action can be initiated via the automated assistant 204 with less input, the vehicle context engine 218 may determine when to render a corresponding assistant suggestion to the user. For example, the vehicle context engine 218 may process data from one or more different sources to determine the appropriate time, place, and / or computing device to render an assistant suggestion to the user. For instance, it may process application data 230 to determine whether the user has switched applications on the vehicle computing device and / or portable computing device to prompt the rendering of a particular medium. In response to this determination, the vehicle context engine 218 may determine that this instance of the application being switched to is suitable for rendering an assistant suggestion, which relates to invoking an automated assistant to interact with a specific application (for example, an application that renders media via the suggestion, "Next time, just say 'Open media application'") to perform a specific action. Alternatively or additionally, data from one or more different sources may be processed using one or more heuristic processes and / or one or more trained machine learning models to identify suitable times, places, and / or computing devices for rendering a particular assistant suggestion. Data may be processed periodically and / or responsively to generate embeddings or other low-dimensional representations, which may be mapped into a latent space to which other existing embeddings have been previously mapped. When a generated embedding is identified as being a threshold distance away from an existing embedding in the latent space, an assistant suggestion corresponding to that existing embedding may be rendered to the user.

[0036] In some embodiments, the system 200 may further include an assistant suggestion engine 226 that can generate assistant suggestions based on interactions between the user and one or more applications and / or devices. For example, the assistant suggestion engine 226 may process application data 230 to identify one or more actions that can be initiated by a particular application among one or more applications 234. Alternatively or additionally, interaction data may be processed to determine whether the user has previously initiated the execution of one or more actions in a particular application among one or more applications 234. The assistant suggestion engine 234 may, in response to assistant input sent from the user to the automated assistant 204, determine whether the automated assistant 204 can initiate one or more actions. For example, the API of a particular application among one or more applications 234 may be accessed by the automated assistant 204 to determine whether the automated assistant 204 can use API commands to control a particular action of that particular application among one or more applications 234. When the automated assistant 204 identifies an API command to control a specific action that the user may be interested in, the assistant suggestion engine 226 may generate a request or command phrase and / or other assistant input that can be rendered to the user as an assistant suggestion. For example, based on the automated assistant's identification that an API call (e.g., playMedia(Podcast Application, 1, resume_playback())) is available to control a specific application function of interest to the user, the automated assistant 204 may generate the suggestion "Assistant, play recent podcasts," in which "Assistant" corresponds to the invocation phrase for the automated assistant 204, and "Play recent podcasts," when detected, corresponds to a request or command phrase that causes the automated assistant 204 to perform the action on behalf of the user.

[0037] In some embodiments, while a user interacts with one or more applications 234, before the user interacts with one or more applications 234, in response to the completion of a specific application action via one or more applications 234, and / or before the completion of a specific application action via one or more applications 234, the automated assistant 204 may render assistant suggestions. In some embodiments, assistant suggestions may be generated and / or rendered based on whether the user is the vehicle owner, the vehicle driver, a passenger in the vehicle, a vehicle renter, and / or any other person that may be associated with the vehicle. The automated assistant 204 may determine whether a user in the vehicle belongs to one or more of these categories using any known technology (e.g., speaker recognition, facial recognition, password recognition, fingerprint recognition, and / or other technologies). For example, when the owner is driving the vehicle, the automated assistant 204 may render more personalized assistant suggestions to the owner, and to passengers and / or renters of the vehicle, other assistant suggestions that may be useful to a wider audience. The automated assistant 204 may operate in such a way that the user's level of familiarity with various automated assistant functions is indirectly proportional to the frequency to which assistant suggestions are rendered to that user. For example, if a given suggestion has already been presented to the vehicle's user, the suggestion may not be presented to that user afterward, but may be presented to another user of the vehicle (e.g., a vehicle borrower).

[0038] In some embodiments, the system 200 may further include a suggestion training engine 224, which may be used to generate training data that can be used to train one or more different machine learning models, which may be used to provide assistant suggestions to different users related to the vehicle. Alternatively or additionally, the suggestion training engine 224 may generate training data based on whether a particular user has interacted with a rendered assistant suggestion. In this way, one or more models may be further trained to provide assistant suggestions that the user is likely to be interested in, so as not to distract the user with an excessive number of suggestions that the user may not be of strong interest to and / or that the user may already be aware of.

[0039] In some embodiments, an assistant suggestion may embody a command phrase, and upon detection of the command phrase by the automated assistant 204, the automated assistant 204 interacts with one or more applications 234 to perform one or more different actions and / or routines. Alternatively or additionally, multiple different assistant suggestions may be rendered simultaneously for a particular application among the one or more applications 234, and / or for multiple applications among the one or more applications 234, thereby enabling the user to learn multiple different command phrases when controlling one or more applications 234. In some embodiments, more assistant suggestions may be rendered for a vehicle borrower compared to the vehicle owner. The decision to render more assistant suggestions for borrowers and / or other guests may be based on how often the vehicle owner uses the automated assistant 204 via the vehicle computing device. In other words, the assistant suggestion engine 226 may display fewer assistant suggestions to users who frequently use the functions of the automated assistant 204, and more assistant suggestions to users who do not frequently use the functions of the automated assistant 204 (users who have received prior permission from the user). As described above, vehicle owners can be distinguished from other users using various technologies.

[0040] Figure 3 illustrates a method 300 for providing an automated assistant that can generate suggestions for assistant inputs that a user may provide, in order to facilitate the automated assistant controlling an application while the user is in a vehicle. Method 300 may be performed by one or more computing devices, applications, and / or any other devices or modules that may be associated with the automated assistant. Method 300 may include an action 302 for determining whether a user is in a vehicle. The vehicle may be, for example, a vehicle computing device that may provide access to the automated assistant. The automated assistant may determine that a user is in a vehicle based on data from one or more different sources, such as the vehicle computing device, a portable computing device (e.g., a computer that the user may bring into the vehicle), and / or any other device that may communicate with the automated assistant. For example, device data may indicate that a user has entered a vehicle, at least based on device data indicating that a portable computing device owned by the user (e.g., a mobile phone) has been paired with the vehicle computing device (e.g., via the Bluetooth protocol). Alternatively or additionally, the automated assistant may determine that a user is in a vehicle based on contextual data that may characterize the user's context at a particular time. For example, when an automated assistant determines whether a user is inside a vehicle, a calendar entry might indicate that the user is driving to a class. Based on this instruction, the automated assistant may use one or more heuristic processes and / or one or more trained machine learning models to determine (e.g., with a threshold probability) whether a user is likely to be inside the vehicle. Alternatively or additionally, one or more sensors in the vehicle connected to the vehicle's vehicle computing device may generate sensor data (e.g., via an occupancy sensor) indicating that a user is inside the vehicle.

[0041] Method 300 may proceed from operation 302 to operation 304, operation 304 may include processing context data associated with the user and / or vehicle. In some embodiments, the context data may include user identifiers (e.g., username, user account, and / or any other identifiers), vehicle identifiers (e.g., vehicle type, vehicle name, vehicle original equipment manufacturer (OEM), etc.), and / or vehicle location. Alternatively or additionally, the context data may include device data, application data, and / or any other data that may indicate devices and / or applications that the user has recently interacted with, is currently interacting with, and / or is expected to interact with. Alternatively or additionally, the application data may indicate that the user has recently established a habit of accessing the real estate application every morning to view the "Recent Listings" application page.

[0042] Method 300 may proceed from action 304 to action 306, which determines whether the user is interacting with or is expected to interact with a particular application and / or application function. If the user is not identified as interacting with or is not expected to interact with a particular application, Method 300 may proceed from action 306 to any action 308 that updates the training data. The training data may be updated to reflect the state in which the user is not interacting with the Auto Assistant and / or a particular application function while inside the vehicle. Subsequent processing of the context data using one or more trained machine learning models trained on the updated training data may provide the user with more relevant and / or more useful suggestions. Alternatively, if the user is identified as interacting with or is expected to interact with a particular application function, Method 300 may proceed from action 306 to action 310.

[0043] Operation 310 may include generating suggestion data that characterizes an assistant input for controlling a specific application function. For example, suggestion data may be generated using an Application Programming Interface (API), which may enable the automated assistant to submit executable requests to and / or receive executable requests from a particular application. Alternatively or additionally, the automated assistant may process HTML code, XML code, Document Object Model (DOM) data, Application Programming Interface (API) data, and / or any other information that may indicate whether a particular function of an application is controllable. Based on this determination, the automated assistant may generate one or more suggestions for an utterance, which, once submitted by the user, causes the automated assistant to control one or more functions of a particular application. For example, the automated assistant may identify that accessing the "Recent Listings" page of a real estate application includes opening the real estate application and selecting a selectable GUI element labeled "Recent Listings". Based on this identification, the automated assistant may generate one or more actionable requests to a real estate application and text content corresponding to an utterance, which, once provided by the user, prompts the automated assistant to submit actionable requests to the real estate application. The one or more requests and / or text content may then be stored as suggestion data, which may be used to render suggestions for one or more users. Once the user provides an utterance, the automated assistant provides one or more requests to the real estate application. The real estate application may respond by generating a "Recent Lists" page, which the automated assistant may render on the vehicle computing device's display interface. Alternatively or additionally, the automated assistant may render an audible output characterizing the content of the "Recent Lists" page.

[0044] Method 300 may proceed from operation 310 to operation 312, which may include causing the text of the utterance to be rendered on the display interface of the vehicle computing device and / or another computing device associated with the user (e.g., a portable computing device associated with the user). For example, according to the above embodiment, the automated assistant may cause a suggestion GUI element (optionally selected) to be rendered on the display interface of the vehicle computing device while the user is in the vehicle. The suggestion GUI element may be rendered by natural language content that characterizes the utterance, such as "Assistant, show me the recent list." In some embodiments, the automated assistant may render request suggestions without a specific identifier of the application to be controlled. Rather, the suggestions may be generated in such a way that the user can correlate the suggestion with the context of the suggestion to infer the application to be controlled. In some embodiments, the user may initiate communication of a request from the automated assistant to a real estate application by providing an utterance and / or tapping the touch interface of the vehicle computing device. For example, the display interface of a vehicle computing device may respond to touch input, and / or one or more buttons (e.g., buttons on the steering wheel), switches, and / or other interfaces of a vehicle computing device may respond to touch input. One or more of these types of touch inputs may be used to initiate the execution of one or more actions in response to one or more suggestions from an automated assistant.

[0045] Method 300 may proceed from action 312 to action 314, which determines whether the user has provided assistant input to facilitate control of an application function. For example, the user may provide an utterance such as "Assistant, show me the recent list" to facilitate control of the real estate application by the automated assistant. If the automated assistant receives assistant input, Method 300 may proceed from action 314 to action 316, which may include facilitating control of a specific application function by the automated assistant. Alternatively, if the user does not provide assistant input in response to a suggestion from the automated assistant, Method 300 may proceed from action 314 to any action 308.

[0046] In some embodiments, when a user provides an assistant input corresponding to a suggestion, method 300 may proceed from action 316 to an optional action 318. The optional action 318 may include updating training data based on the user using the automated assistant to control a specific application function. For example, training data may be generated based on the user providing assistant input, and this training data may be used to train one or more trained machine learning models. One or more trained machine learning models may then be used to process contextual data when it is identified that the user or another user is inside the vehicle. This further training of one or more trained machine learning models enables the automated assistant to provide more relevant and / or effective suggestions for the user to invoke the automated assistant to control one or more separate applications, while also promoting safe driving habits.

[0047] Figure 4 is a block diagram 400 of an exemplary computer system 410. The computer system 410 typically includes at least one processor 414 that communicates with a number of peripheral devices via a bus subsystem 412. These peripheral devices may include, for example, a storage subsystem 424 including memory 425 and a file storage subsystem 426, a user interface output device 420, a user interface input device 422, and a network interface subsystem 416. The input and output devices enable user interaction with the computer system 410. The network interface subsystem 416 provides an interface to an external network and connects to a corresponding interface device in another computer system.

[0048] User interface input devices 422 may include keyboards, pointing devices (e.g., mice, trackballs, touchpads, or graphic tablets), scanners, touchscreens integrated into displays, audio input devices such as speech recognition systems, microphones, and / or other types of input devices. In general, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computer system 410 or communication network.

[0049] The user interface output device 420 may include non-visual displays such as display subsystems, printers, fax machines, or audio output devices. The display subsystem may include flat-panel devices such as cathode ray tubes (CRTs) and liquid crystal displays (LCDs), projection devices, or other mechanisms for creating visible images. The display subsystem may also provide non-visual displays via audio output devices, etc. In general, the use of the term “output device” is intended to include all possible types of devices and methods for outputting information from the computer system 410 to the user or to another machine or computer system.

[0050] The storage subsystem 424 stores programming and data structures that provide some or all of the functionality of the modules described herein. For example, the storage subsystem 424 may include logic for performing a selected embodiment of method 300, and / or logic for performing one or more of the system 200, the vehicle computing device 106, the portable computing device 104, and / or any other applications, assistants, devices, apparatus, and / or modules discussed herein.

[0051] These software modules are typically executed on processor 414 alone or in combination with other processors. The memory 425 used by the storage subsystem 424 may include a large amount of memory, including main random access memory (RAM) 430 for storing instructions and data during program execution, and read-only memory (ROM) 432 for storing fixed instructions. The file storage subsystem 426 may provide persistent storage for program files and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules that implement the functionality of a particular embodiment may be stored by the file storage subsystem 426 within the storage subsystem 424, or on another machine accessible to the processor 414(s).

[0052] The bus subsystem 412 provides a mechanism that enables various components and subsystems of the computer system 410 to communicate with each other as intended. Although the bus subsystem 412 is schematically shown as a single bus, multiple buses may be used in alternative embodiments of the bus subsystem.

[0053] The computer system 410 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing systems or computing devices. Because computers and networks are constantly evolving, the description of the computer system 410 shown in Figure 4 is intended to be merely a specific example illustrating several embodiments. Many other configurations of the computer system 410 may have more or fewer components than the computer system shown in Figure 4.

[0054] Where the systems described herein collect or use personal information relating to a user (or often referred to herein as “Participant”), the user may be given the opportunity to control whether the program or features collect user information (e.g., information about the user’s social networks, social actions or activities, occupation, user preferences, or current geographical location), or to control whether and / or how they receive content that may be more relevant to the user from a content server. Furthermore, certain data may be processed in one or more ways so that personally identifiable information is removed before it is stored or used. For example, a user’s identity may be processed so that personally identifiable information does not identify the user, or, if geographical location information (such as city, zip code, or state level) is available, the user’s geographical location may be generalized so that the user’s specific geographical location cannot be identified. Thus, the user can control how information about them is collected and / or used.

[0055] While several embodiments are described and illustrated herein, various other means and / or structures may be used to perform the function and / or obtain one or more of the results and / or benefits described herein, and each of such variations and / or modifications is considered to fall within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials and configurations described herein are intended for illustrative purposes only, and actual parameters, dimensions, materials and / or configurations will depend on the specific application in which the teaching(plural) is used. A person skilled in the art will recognize or be able to confirm many equivalents to the specific embodiments described herein using only routine experimentation. Thus, it should be understood that the embodiments described herein are presented for illustrative purposes only, and embodiments may be practiced in ways other than those specifically described and claimed, within the scope of the appended claims and their equivalents. The embodiments of this disclosure cover the individual functions, systems, products, materials, kits and / or methods described herein. Furthermore, any combination of two or more such functions, systems, products, materials, kits, and / or methods is included in the scope of this disclosure, provided that they do not conflict with each other.

[0056] In some embodiments, a method is provided which is implemented by one or more processors, and the method includes identifying that a user is engaged in interaction with a given mobile application via a mobile computing device present in the vehicle with the user. The given mobile application is separate from an automated assistant application, which is accessible via the mobile computing device and the vehicle's vehicle computing device. The method further includes generating proposed data characterizing a command phrase based on the interaction with the given mobile application, and when the command phrase is submitted to the automated assistant application by the user, the automated assistant application causes the automated assistant application to control a particular action of a given vehicle application. The given vehicle application corresponds to a given mobile application. The method further includes arranging the command phrase to be visually rendered in the foreground of the vehicle computing device's display interface based on the proposed data.

[0057] These and other embodiments of the technology disclosed herein may optionally include one or more of the following functions:

[0058] In some embodiments, generating the proposed data may be performed on a mobile computing device, and the method may further include providing the proposed data to a vehicle computing device in response to interaction with a given mobile application.

[0059] In some embodiments, generating the proposed data may be performed on a vehicle computing device, and the method may further include receiving interaction data from a given mobile application that characterizes the interaction between the user and the given mobile application. The proposed data may be generated based further on the interaction data.

[0060] In some embodiments, command phrases may be visually rendered as selectable graphical user interface (GUI) elements on a display interface, and selectable GUI elements may be selectable via touch input received in the area of ​​the display interface corresponding to the selectable GUI element.

[0061] In some embodiments, generating proposed data characterizing a command phrase may involve generating a command phrase based on one or more application actions that can be initiated through direct interaction between a user and a given mobile application's GUI interface. The GUI interface may be rendered on a mobile computing device, and the one or more application actions may include specific actions.

[0062] In some embodiments, generating proposed data characterizing a command phrase may involve generating a command phrase based on one or more application actions initiated by the user via a given mobile application on a mobile computing device during one or more previous instances while the user was inside a vehicle. These one or more application actions may include specific actions.

[0063] In some embodiments, methods are provided that are carried out by one or more processors, the method comprising generating predictive data indicating that a user is expected to interact with the application interface of a given application and control the functionality of the given application via the display interface of a vehicle computing device. The given application is separate from an automated assistant application accessible via the vehicle computing device of a vehicle. The method further comprises generating, based on the predictive data, suggestive data characterizing a command phrase, the command phrase, when submitted to the automated assistant application by a user, causing the automated assistant application to control the functionality of the given application; ensuring that at least the command phrase is rendered on the display interface of the vehicle computing device before the user interacts with the functionality of the application; and, in response to at least the command phrase being rendered on the display interface of the vehicle computing device, receiving an assistant input from the user to the automated assistant application, which includes at least the command phrase, and, based on having received the assistant input, causing the automated assistant application to control the functionality of the given application based on the assistant input.

[0064] These and other embodiments of the technology disclosed herein may optionally include one or more of the following functions:

[0065] In some embodiments, command phrases may be rendered as selectable graphical user interface (GUI) elements on the display interface, and assistant input may be touch input received in the area of ​​the display interface corresponding to the selectable GUI element. In some versions of these embodiments, generating predictive data may include identifying that the vehicle is being driven toward a location where the user is expected to interact with the application interface and control the application's functions. The proposed data may further be based on the location the vehicle is heading toward.

[0066] In some embodiments, the method may further include identifying that the application interface of a given application is rendered on the display interface of a vehicle computing device. In response to identifying that the application interface is rendered on the display interface of a vehicle computing device, it may be performed to generate predictive data.

[0067] In some embodiments, the method may further include identifying that the user is currently inside a vehicle and identifying that, during a previous instance while the user was inside the vehicle, the user accessed the application interface of a given application. Based on the identification that the user is currently inside a vehicle and that the user previously accessed the application interface of a given application, it may be performed to generate predictive data.

[0068] In some embodiments, the method may further include processing contextual data using one or more trained machine learning models. The contextual data may characterize one or more features of the user's context, and based on having processed the contextual data, it may be performed to generate predictive data. In some versions of these embodiments, one or more trained machine learning models may be trained using data generated during one or more previous instances in which one or more other users accessed features of a given application while in their respective vehicles.

[0069] In some embodiments, a method is provided which is carried out by one or more processors, and the method includes the vehicle's vehicle computing device identifying that the user is engaged in interaction with a given application on the vehicle computing device while the user is inside the vehicle. The given application is separate from an automated assistant application accessible via the vehicle computing device. The method further includes the vehicle computing device generating proposal data characterizing a command phrase based on the interaction with the given application, the command phrase, when submitted by the user to the automated assistant application, causes the given application to perform a specific action associated with the interaction between the user and the given application, and the vehicle computing device causing the proposal data to be visually rendered on the vehicle computing device's display interface. The command phrase is rendered in the foreground of the vehicle computing device's display interface.

[0070] These and other embodiments of the technology disclosed herein may optionally include one or more of the following functions:

[0071] In some embodiments, the command phrase may include an invocation phrase for invoking an automated assistant application and natural language content characterizing a request to perform a specific action. In some versions of these embodiments, the method may further include: the vehicle computing device receiving an utterance embodying a request to the automated assistant application; the vehicle computing device identifying that the utterance includes a request rendered on the vehicle computing device's display interface; and, in response to receiving the utterance, causing the given application to perform a specific action associated with the interaction between the user and the given application. In some further versions of these embodiments, the command phrase may be rendered when the user is in the vehicle during the first long trip, and the utterance may be received when the user is in the vehicle during the second long trip.

[0072] In some embodiments, the method may further include identifying that a particular action is controllable via selectable content rendered on the display interface of the vehicle computing device. Generating proposed data characterizing a command phrase may be based on the identification that a particular action is controllable via selectable content rendered on the display interface of the vehicle computing device. In some versions of these embodiments, the application may be a communications application, and the selectable content includes one or more selectable elements for specifying a phone number and / or contact to make a call. In additional or alternative versions of these embodiments, the application may be a media application, and the selectable content includes one or more selectable elements for specifying media to be rendered visually via the display interface of the vehicle computing device and / or audibly via one or more speakers of the vehicle computing device.

[0073] In some embodiments, rendering proposed data on the display interface of a vehicle computing device may be in response to identifying that the user has completed an interaction with a given application on the vehicle computing device.

[0074] In some embodiments, generating the proposed data characterizing the command phrase may further be based on the context of the user involved in the interaction with a given application on the vehicle computing device. In some versions of these embodiments, generating the proposed data characterizing the command phrase may further be based on the display context of the display interface of the vehicle computing device.

[0075] In some embodiments, the method may further include, after the command phrase has been rendered on the vehicle computing device's display interface, the vehicle computing device identifying, via the vehicle computing device's display interface, that the user is engaged in a separate interaction with a further application of the vehicle while the user is inside the vehicle. The further application may be separate from the Auto Assistant application and separate from a given application. The method may further include, by the vehicle computing device, generating further suggestion data characterizing the further command phrase based on the separate interaction with the further application, the further command phrase, when submitted by the user to the Auto Assistant application, causing the Auto Assistant application to control further actions of the further application, and by, by, the vehicle computing device, causing the further suggestion data to be visually rendered on the vehicle computing device's display interface. In some versions of these embodiments, causing the further suggestion data to be rendered on the vehicle computing device's display interface may be done in response to the identification that the user has completed a separate interaction with a further application of the vehicle computing device.

[0076] Other embodiments may include a non-temporary computer-readable storage medium for storing instructions that can be executed by one or more processors (e.g., a central processing unit (CPU)(or more), a graphics processing unit (GPU)(or more), and / or a tensor processing unit (TPU)(or more)) in order to perform one or more of the methods described above and / or elsewhere in this Spec. Further embodiments may include a system of one or more computers, including one or more processors capable of operating to execute the stored instructions, in order to perform one or more of the methods described above and / or elsewhere in this Spec.

[0077] It should be understood that all combinations of the aforementioned concepts and any further concepts described in more detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter disclosed herein.

Claims

1. A method carried out by one or more processors, In response to the execution of an automated assistant application configured to cooperate with (i) a mobile computing device located in a vehicle and having a touchscreen interface, and (ii) a vehicle computing device having a display interface, by one or more of the aforementioned processors, An IoT application provided to the mobile computing device, wherein it identifies that one of the GUI elements has been selected by manual interaction by a user in the vehicle via the touchscreen interface of the mobile computing device, where one or more selectable GUI elements for identified functions of an instance of the IoT application are displayed. To generate proposed data characterizing command phrases for controlling the functionality of an instance of an IoT application associated with the selected GUI element, To cause the command phrase to be visually rendered in the foreground of the display interface of the vehicle computing device, the proposed data is transmitted to the vehicle computing device, In response to an utterance made by the user in the vehicle based on the rendered command phrase, the system controls the functionality of the instance relating to the selected GUI element by interacting with the application interface (API) of the IoT application or by communicating API data to the IoT application. A method of doing something that includes

2. One or more processors, Memory for storing instructions, A system equipped with, When the instruction is executed, it causes one or more processors to perform the method according to claim 1. system.

3. A non-temporary computer-readable storage medium for storing instructions, wherein, when the instructions are executed, the non-temporary computer-readable storage medium causes one or more processors to perform the method according to claim 1.