Transient Personalization Mode for Guest Users of the Automation Assistant

Through the collaborative work of the host automation assistant and the guest automation assistant, biometric signature and embedded data authentication are used to solve the problem of limited automation assistant functions in user instant interaction, achieving seamless personalized response and resource conservation.

CN115769202BActive Publication Date: 2025-07-08GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080102329.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-14
Filing Date
2020-12-14
Publication Date
2025-07-08
Estimated Expiration
2040-12-14

AI Technical Summary

Technical Problem

On computing devices where users only interact instantly, the automation assistant cannot enter the login mode, resulting in limited functions or data security issues, and multiple inputs are cumbersome and delayed user experience.

Method used

Through the collaboration between the host automation assistant and the guest automation assistant, the biometric signature and embedded data are used for authentication, and instantaneous personalization mode is realized, allowing the guest user to receive personalized responses without logging in.

Benefits of technology

It realizes that visitors and users seamlessly receive personalized responses on the computing device, save computing resources, improve user experience and enhance data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115769202B_ABST
    Figure CN115769202B_ABST
Patent Text Reader

Abstract

The embodiments described herein relate to an automated assistant that is capable of operating in an instant personalization mode and / or assisting a separate automated assistant to provide an output according to the instant personalization mode. The instant personalization mode can allow a guest user of a device that enables the assistant to receive a personalized response from the device that enables the assistant—even without logging in to the device that enables the assistant. A host automated assistant of the device that enables the assistant can securely communicate with the guest user's automated assistant through a backend process. In this way, an input query from the guest user to the host automated assistant can be personalized according to the guest automated assistant—without the guest user directly using their own personal device.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Humans can use interactive software applications to conduct conversations between humans and computers. Interactive software applications are referred to herein as "automation assistants" (also known as "digital agents", "chatbots", "interactive personal assistants", "intelligent personal assistants", and "dialogue agents", etc.). For example, humans (who can be referred to as "users" when interacting with the automation assistant) can provide commands and / or requests using spoken natural language inputs (i.e., utterances) and / or by providing text (e.g., typed) natural language inputs. In some cases, the spoken natural language inputs can be converted into text and then processed.

[0002] In some cases, the automation assistant can be available to the user via each of a plurality of different automation assistant devices (i.e., computing devices that each provide access to the automation assistant), and these devices are each in the user's login mode. In the login mode, the computing device can utilize the user's credentials so that the automation assistant accessible via the computing device can at least selectively access (e.g., in response to the user's speaker verification and / or facial verification) various data specific to the user. In addition, the automation assistant can utilize such data when processing user requests submitted to the automation assistant via the computing device. For example, when performing speech recognition of an utterance from the user (e.g., when selecting a speech recognition language, when biasing towards (certain) terms, etc.), when determining the underlying content for responding to the utterance (e.g., determining content from such data or using such data to identify content), and / or when determining which speech-synthesized voice (e.g., a voice easily understood by the user) can be audibly presented as a response, such data can be utilized. Therefore, using the automation assistant in the login mode provides various technical benefits, such as ensuring accurate speech processing of the user's requests, generating responses related to the requests, and / or presenting the responses in a manner easily understood by the user.

[0003] However, for a given computing device, it may often be necessary for multiple user device interactions to at least selectively be in the user's login mode. These interactions can include multiple touch inputs to the automation assistant application to add the user as an authorized user of the computing device. In addition, for a computing device where the user is not an administrator, the user may need to interact with the administrator to have the administrator add the user as an authorized user. In addition, for a given computing device that is only transiently utilized by the user, data security issues can arise when the user operates in the login mode.

[0004] Given these and other considerations, there are multiple benefits to operating in a logged-in mode for a user's personal computing device(s) and / or computing device(s) with which the user persistently interacts (e.g., computing devices in the user's home). However, for computing devices with which the user only interacts transiently (e.g., only a limited number of interactions and / or for a limited duration), the user may not be able to be in a logged-in mode (e.g., a given user may lack authorization to be added as a logged-in user). Additionally or alternatively, the multiple inputs required to add a user as a logged-in user may not be guaranteed for transient interactions, and - moreover, providing the multiple inputs would require a delay in the transient interaction. As an example, when a user is using a computing device in a friend's home or in an enterprise (e.g., a hotel), the user may only be able to operate the automated assistant of the computing device in a guest mode. The functionality of the automated assistant can be limited in the guest mode, and / or various benefits of the logged-in mode may not be available in the guest mode. SUMMARY OF THE INVENTION

[0005] Embodiments presented herein relate to various techniques for instantaneously adjusting the handling of (a)utomated assistant requests based on user personal data - particularly when (a) request(s) are provided by a user at an automated assistant device where the user is not a logged-in / authenticated user. This instant adaptation is sometimes referred to herein as operating in an instant personalization mode. Operating in the instant personalization mode allows the use of user personal data to process guest user requests received at a host automated assistant device, even though the user is not authenticated on the host automated assistant device. This can include, for example, using the data when performing speech recognition if the request is speech, using the data when determining the underlying content of a response to the request, and / or using the data when determining which voice synthesis voice to audibly present the response in. Although in some cases, the guest user has no prior interaction with the host automated assistant device, some embodiments enable instant personalization.

[0006] As used herein, "host automated assistant" will be used to refer to an instance of an automated assistant accessible by a host automated assistant device where the guest user of the host automated assistant device is not a logged-in user of the automated assistant. As used herein, "guest automated assistant" will be used to refer to an instance of an automated assistant accessible by a guest automated assistant device where the guest user is a logged-in user thereof. In other words, the guest user is not an authenticated user of the host device, and thus the host automated assistant device cannot be used to directly access the user's personal automated assistant data. On the other hand, the guest user is an authenticated user of the guest automated assistant device, and thus the guest automated assistant device can provide direct access to the guest user's personal and / or automated assistant data stored in association with the guest user's account.

[0007] In some embodiments, for the host automation assistant to operate in an instant personalization mode for a guest user, the host automation assistant can determine that the guest user is associated with a guest automation assistant. For example, various users can have assistant accounts associated with their respective automation assistants (i.e., guest users can have their own personal automation assistants). However, when a particular user is considered a guest user with respect to a host automation assistant (e.g., an automation assistant accessible via a host device), this host automation assistant can determine that the user has an account established with a guest automation assistant (e.g., an automation assistant accessible via the user's personal computing device).

[0008] In some embodiments, before operating in an instant personalization mode, the host automation assistant can ensure that there is a correlation between the guest user and a specific input. For example, the determination of the guest user's correlation can be initialized in response to the host automation assistant device receiving an input from a guest user who may be away from work. The input can be a utterance, such as "Assistant, what's on my calendar?", which can be provided by the guest user to, for example, the host automation assistant device in a hotel room. In response to receiving the utterance, the host automation assistant can initially determine whether the source of the utterance corresponds to an existing authenticated user (e.g., the hotel owner). For example, the host automation assistant device or another network device can determine whether the biometric signature (e.g., voice, face, fingerprint, pupil, etc.) of the person providing the utterance matches the biometric signature of any existing authenticated user (e.g., an employee of the hotel). Based on the host automation assistant determining that the utterance is provided by an unauthenticated user (e.g., does not match any logged-in user of the device), the host automation assistant can identify nearby devices associated with the user who provided the utterance or other input to the host automation assistant.

[0009] For example, in some embodiments, the host automation assistant is capable of confirming that the utterance corresponds to a user within the vicinity of the host automation assistant device. The host automation assistant is capable of generating: a voice embedding and / or voice vector based on the voice signature embodied in the utterance, a face embedding and / or face vector based on one or more images, a fingerprint embedding and / or fingerprint vector based on a user finger scan, and / or any other information that can be used for biometric authentication with the prior permission of the user. The voice embedding can be used to encrypt an authentication value (e.g., a secret string or other data), and the encrypted value can be shared with one or more nearby devices. For example, one or more devices including the guest device can receive the encrypted authentication value via Bluetooth, ultrasonic, local area network (LAN), wide area network (WAN), the Internet, intranet, and / or Wi-Fi connection. In some embodiments, the devices eligible to receive the encrypted authentication value can be restricted to certain devices within a threshold distance from the host device. In response, the guest device can attempt to decrypt the encrypted authentication value using the same or a similar voice embedding accessible to the guest device. Since the host device and the guest device have each received the utterance from the guest user, their respective embeddings can have a similar arrangement in the latent space. Thus, the guest device having a voice embedding corresponding to the guest user providing the utterance will be able to decrypt the encrypted authentication value. In this way, the host device can ensure that the utterance corresponds to a nearby user and nearby devices, thereby reserving the transient personalization mode for those users who are truly close to the host device.

[0010] In some embodiments, when the guest device decrypts the encrypted authentication value, the guest device can transmit the authentication value back to the host device to indicate to the host device that the guest user has been authenticated on the guest device. In response to receiving the correct authentication value, the host device can transmit the utterance to the guest device. For example, the host device can generate encrypted query data embodying the utterance and can transmit the encrypted query data to the guest device. The transmitted query data can include audio data, text data (e.g., the text of the speech-to-text processing performed at the host device), and / or natural language processing data (e.g., an identifier of the action intent and / or parameters of the action intent). Then, the guest device can generate response data based on the encrypted query data and share the response data with the host device. Alternatively or additionally, the host device can provide the encrypted authentication value to the encrypted query data so that only the guest device having the correct voice embedding can decrypt the assistant query and the authentication value. Then, the response data along with the authentication value can be provided back to the host device, and the host device can present an output based on the response data.

[0011] According to the foregoing example, the visitor device can decrypt the encrypted query data to determine that the visitor user is requesting the host automation assistant to tell the visitor user what is scheduled on the visitor user's calendar. Based on this determination, the visitor device (e.g., the visitor user's mobile phone) can enable the visitor automation assistant or a separate application to access the visitor user's calendar application to generate response data for presentation by the host automation assistant. When the visitor device and / or the associated device generates response data that can correspond to a description of a scheduled event (e.g., "At 6:00 PM this afternoon, you have an arrangement of 'dinner with dad'"), the visitor device can transmit the response data to the host device. Alternatively or additionally, the visitor device can transmit one or more user preferences of the visitor user, such as a preferred voice profile of the automation assistant. The host device can optionally receive the response data as encrypted response data. Then, the host device can process the response data to present a corresponding output at one or more interfaces of the host device. For example, as a result of this process, the host device can provide an audible response to the visitor user, such as "According to your calendar, at 6:00 PM this afternoon, you have an arrangement of 'dinner with dad'". In this way, the visitor user does not have to specifically rely on their personal device to receive personalized responses from the automation assistant. This can allow the visitor user to conserve the computing resources of their personal device, such as battery life and network usage, when away from home.

[0012] The host automation assistant can determine that a statement is suitable for a personalized response based on, for example, determining that the statement includes content that can only be accessed by someone with access to the calendar application managed by the visitor user. Alternatively or additionally, the automation assistant can determine that the statement is suitable for a personalized response based on determining that the topic of the statement (e.g., calendar) is related to user-customizable information and / or the statement includes a possessive pronoun (e.g., "my"). Alternatively or additionally, one or more trained machine learning models can be used to determine whether the statement includes a query suitable for a personalized response. Alternatively or additionally, the host automation assistant can omit determining whether the statement is suitable for a personalized response and instead determine whether the visitor user is associated with the visitor automation assistant. As used herein, the visitor automation assistant can be (i) another automation assistant provided by the same entity that provides the host automation assistant, (ii) another automation assistant provided by a different entity, and / or (iii) associated with a specific automation assistant that can be accessed via an application programming interface (API) available to the host automation assistant.

[0013] When the host automation assistant determines that the utterance includes a query suitable for a personalized response, and / or when the host automation assistant determines that the user is associated with a separate automation assistant, the host automation assistant can initiate operation in the transient personalization mode. However, prior to transitioning to the transient personalization mode, the host automation assistant can initially confirm whether the utterance is related to nearby users and / or nearby assistant-enabled devices. In some embodiments, when the host device receives an utterance that includes a personal query but the host device cannot authenticate any nearby devices, the host automation assistant can provide a non-personalized response. Alternatively or additionally, the host automation assistant can provide a response that explicitly states that the response from the host automation assistant is not personalized for the visitor user who provided the personal query and / or that the host automation assistant cannot identify an account and / or device associated with the visitor user. This can make certain visitor users aware that although they may be aware that they can receive personalized results from the host automation assistant, the response they are currently receiving is not personalized for them. In these cases, such notifications can eliminate miscommunications with any host automation assistant that is capable of operating in the transient personalization mode.

[0014] In some embodiments, the user can provide permission for the host automation assistant and the visitor automation assistant to coordinate personalized responses prior to the host automation assistant processing a query from the user. Alternatively or additionally, the user can limit the permissions of the host automation assistant based on time, context, topic, and / or any other parameter suitable for restricting the responsiveness of the automation assistant. For example, when a visitor user initially provides a personal query to the host automation assistant, the host automation assistant can request that the visitor automation assistant process the personal query. In response to receiving the request from the host automation assistant, the visitor automation assistant can present a prompt to the visitor user to obtain permission for the visitor automation assistant to coordinate a personalized response with the host automation assistant. Alternatively or additionally, the visitor automation assistant or another application can prompt the visitor user as to whether the visitor user wishes to limit the transient personalization mode of the host automation assistant. In response, the visitor user can choose to limit the transient personalization mode of the host automation assistant to a specific time period (e.g., the next 24 hours), a specific location (e.g., when the visitor user is near a threshold of the host automation assistant device), and / or a specific context (e.g., when the visitor user's calendar indicates that the visitor user is on a business trip).

[0015] In some embodiments, when a guest user has given permission for the host automated assistant to provide a personalized response, the host automated assistant is also able to operate to provide personalized suggestions to the guest user. For example, when a guest user is staying in a hotel room that includes the host automated assistant device and the user has given permission to receive personalized responses, the host automated assistant can present specific content based on the user's personal preferences. For example, when a guest user provides utterances, or regardless of whether the guest user provides a query to the automated assistant, the guest device can share the user preferences with the host automated assistant when the guest device has been authorized to share such permissions. Using this user preference data, the host automated assistant can select and / or organize certain search results in order to present personalized content to the user. For example, user preferences can characterize the user's language preferences, food preferences, music preferences, event preferences, and / or any other preferences that can be characterized in the data. In this way, when, for example, the host automated assistant at the host device in a hotel room is presenting restaurant suggestions to a guest user, the host automated assistant will be able to filter the suggested content based on the user preferences identified by the guest automated assistant. Alternatively or additionally, when the host device is processing an utterance from the guest user, the host device can use the automatic speech recognition (ASR) model employed by the guest automated assistant to perform the processing. Alternatively or additionally, when the host device presents an audible output in response to an utterance from the guest user, the host device can present the audible output according to a preferred text-to-speech (TTS) profile selected by the guest automated assistant.

[0016] The above description is provided as an overview of some embodiments of the present disclosure. Further descriptions of these embodiments and other embodiments will be described in more detail below.

[0017] Other embodiments may include a non-transitory computer-readable storage medium storing instructions executable by one or more processors (e.g., one or more central processing units (CPUs), one or more graphics processing units (GPUs), and / or one or more tensor processing units (TPUs)) to perform methods such as one or more of the methods described above and / or elsewhere herein. Other embodiments may include a system of one or more computers including one or more processors operable to execute the stored instructions to perform methods such as one or more of the methods described above and / or elsewhere herein.

[0018] It should be understood that all combinations of the above concepts and additional concepts described in more detail herein are considered to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are considered to be part of the subject matter disclosed herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1A and Figure 1B respectively illustrate views of a user interacting with a host automation assistant that can invoke a guest automation assistant when operating in an ephemeral personalization mode for a guest user.

[0020] Figure 2A and Figure 2B illustrate a view of a user interacting with a host automation assistant that can adopt guest user preferences when operating in an ephemeral personalization mode for a guest user.

[0021] Figure 3 illustrate a system for providing an automation assistant that can operate in an ephemeral personalization mode and / or communicate with another automation assistant operating in an ephemeral personalization mode.

[0022] Figure 4 illustrate a method for handling requests from a host automation assistant when the host automation assistant attempts to operate in an ephemeral personalization mode.

[0023] Figure 5 illustrate a method for operating an automation assistant in an ephemeral personalization mode when one or more guest users interact with the automation assistant.

[0024] Figure 6 is a block diagram of an example computer system. DETAILED DESCRIPTION

[0025] Figure 1A and Figure 1B respectively illustrate views 100 and 120 of a user 102 interacting with a host automation assistant that can invoke a guest automation assistant when operating in an ephemeral personalization mode for a guest user. For example, user 102 can travel outside of their corresponding country and stay in a particular hotel room 118. User 102 can arrive at hotel room 118 with their personal device 110, which can be a portable computing device such as a cellular phone. Additionally, hotel room 118 can include one or more assistant-enabled devices such as host device 108 and host television 106.

[0026] Initially, when user 102 arrives at hotel room 118, host device 108 and host TV 106 can operate according to an account corresponding to an entity separate from user 102, such as a business hotel. Thus, initially, host device 108 and host TV 106 will not have access to a different account corresponding to user 102 and may therefore initially be unable to provide a personalized response to user 102. For example, personal device 110 owned by user 102 can provide access to a visitor automated assistant that can provide a personalized response to user 102 based on previous interactions with user 102 and / or other data. However, although host device 108 and host TV 106 can provide access to a host automated assistant, the host automated assistant may not be able to provide personalized information to user 102 without interacting with the visitor automated assistant.

[0027] To interact with the visitor automated assistant, the host automated assistant can operate in an instantaneous personalization mode. This mode can allow the host automated assistant to provide a personalized response to a visitor user associated with another automated assistant. For example, user 102 can provide utterance 104 to host device 108, such as "Assistant, what are my favorite restaurants here?" In response to receiving utterance 104, the host automated assistant accessible via host device 108 can optionally determine whether utterance 104 includes one or more assistant queries that can have a personalized response. For example, the host automated assistant can determine whether the utterance embodies at least one assistant query that can be personalized using data currently inaccessible to the host automated assistant. Alternatively or in addition, the host automated assistant can omit determining whether utterance 104 embodies a query that can have a personalized response and instead determine whether the source of utterance 104 is associated with another automated assistant.

[0028] For example, in some embodiments, the host device 108 can provide a host relevance request 112 to the user 102's personal device 110 before or after receiving the utterance 104. The host relevance request 112 can be a request for the personal device 110 or the guest automation assistant to provide information to the host automation assistant that indicates that the guest automation assistant is associated with the user 102 who provided the utterance 104 and / or that the device enabling the guest automation assistant is associated with the area near the operation of the guest automation system. In some embodiments, the host device 108 or an associated device can generate embedded data or other authentic data and use this data to encrypt secret data that will be accessible to the personal device 110, but not to any other device without specific permission from the guest user. The embedded data can be, for example, a voice embedding or voice vector based on at least some amount of audio captured when the user 102 provided the utterance 104. In this way, because the guest automation assistant has previously received the utterance from the user 102, the guest automation assistant will be able to decrypt the secret data using the same embedding or a similar embedding. For example, when the personal device 110 receives the host relevance request 112, the personal device 110 or another associated personal device can decrypt the host relevance request to identify the secret data. Then, the personal device 110 can generate a guest relevance response 114 that identifies the secret data or is otherwise based on the secret data. As Figure 1A illustrated, an indication that the secret data has been successfully decrypted by the personal device 110 can be embodied in the guest relevance response 114 and provided back to the host device 108 via a network connection (e.g., Wi-Fi, Bluetooth, ultrasonic connection, ZigBee, etc.).

[0029] When the host device 108 determines that a nearby personal device 110 is associated with the user 102, the host device 108 can provide host query data 122 to the personal device 110. Alternatively or additionally, the host query data 122 can be provided to the personal device 110 together with the host relevance request 112. In some embodiments, the host device 108 can provide the original audio data of the utterance provided by the user 102. Alternatively or additionally, the host device 108 can provide encrypted audio data that can be decrypted by the personal device 110. Alternatively or additionally, the host device 108 can provide natural language understanding (NLU) data that characterizes one or more actions requested by the user 102. Alternatively or additionally, the host device 108 can provide a text record of one or more portions of the utterance 104 to the personal device 110.

[0030] In response to receiving host query data 122, personal device 110 and / or the guest automation assistant can generate guest query response data 124. The guest query response data 124 can characterize one or more automation assistant outputs in response to one or more queries embodied in utterance 104 from user 102. In some embodiments, the guest query response data 124 can be encrypted in a manner that allows the host device 108 to decrypt the automation assistant output. In some embodiments, the guest query response data 124 can include natural language content characterizing the output 128 to be presented by the host automation assistant. For example, when the host device 108 receives the guest query response data 124 from the personal device 110, the host device 108 can use the guest query response data 124 to present an audible output 128. For example, the host automation assistant of the host device 108 can present natural language content such as "These are some personalized results for you on the TV."

[0031] Alternatively or additionally, the guest query response data 124 can characterize data in response to utterance 104 but not embodied in natural language sentence format. For example, the guest query response data 124 can include a list 126, and the host device 108 can cause the list 126 to be presented at the host TV 106. In this way, user 102 can interact seamlessly with the host device to receive personalized responses without the user having to go through a dedicated extended authentication process.

[0032] In some embodiments, the personal device 110 can prompt user 102 as to whether user 102 wishes for the host device 108 to no longer use the personal device 110 for the transient personalization mode. Alternatively or additionally, the personal device 110 and / or the host device 108 can prompt the user as to whether user 102 wishes to limit the transient personalization mode to a certain time period, a certain location, and / or any other identifiable limitation. In this way, user 102 can allow the host device 108 to operate strictly in the transient personalization mode during the user's vacation without having to constantly confirm approval for the host device 108 to operate in the transient personalization mode. This can save computational resources that might be consumed during interactions where user 102 repeats certain permissions to the host automation assistant.

[0033] Figure 2A and Figure 2B Views 200 and 220 illustrate the interaction of user 202 with the host automation assistant, and the host automation assistant can adopt guest user preferences when operating in the transient personalization mode of a guest user. In some embodiments, Figure 2A and Figure 2B The illustrated interaction can be Figure 1A and Figure 1B a continuation of the interaction between the illustrated user 102 and the host device 108. Additionally, regardingFigure 1A and Figure 1B the functionality described can be applied to Figure 2A and Figure 2B the features illustrated.

[0034] In some embodiments, user 202 can travel and stay in guest room 218 of a host device or host devices that provide access to a host automated assistant. For example, the one or more host devices can include host device 208 and host television 206. When user 202 is outside their home, they can carry their personal device 210, which can be a cellular phone or other device that provides access to a guest automated assistant, or— that is, an automated assistant with prior permission to access user 202's account.

[0035] In some embodiments, since user 202 is traveling and host device 208 may not be personalized for user 202, host device 208 can request user preference data from one or more devices and / or applications associated with user 202. Such a request can be provided in response to user 202 uttering words 204 such as "Assistant, I'm going to sleep now". In response to receiving words 204, the host automated assistant accessible via host device 208 can determine that words 204 embody a request for the automated assistant to perform one or more actions and / or routines. Alternatively or additionally, the host automated assistant can determine that words 204 embody one or more queries suitable for a personalized response.

[0036] In response to receiving words 204, host device 208 and / or the host automated assistant can provide host relevance request 212, which can be based on one or more embodiments discussed with respect to host relevance request 112. Additionally, in accordance with one or more embodiments discussed with respect to Figure 1A and Figure 1B guest relevance response 114, personal device 210 can provide guest relevance response 214. Based on successfully receiving guest relevance response 214, host device 208 and / or the host automated assistant can provide host query data 222 to personal device 210. Host query data 222 can include a request for personal device 210 and / or the guest automated assistant to provide data that can be used to generate a response to words 204.

[0037] For example, the requested data can include user preference data, ASR data, TTS data, one or more trained machine learning models, and / or any other information that can be used to generate a response to utterance 204. For example, personal device 210 and / or the visitor automation assistant can provide visitor assistant data 224 to host device 208. The visitor assistant data 224 can indicate one or more user preferences associated with one or more queries embodied in utterance 204. For example, since utterance 204 refers to one or more assistant actions that will assist user 202 (e.g., a routine of one or more assistant actions performed by the visitor automation assistant in response to user 202 saying "I'm going to sleep" at night), the user preferences identified in the visitor assistant data 224 can include one or more preferred parameters for the host automation assistant to use for the user when performing one or more assistant actions.

[0038] For example, one or more assistant actions can include setting a thermostat and playing certain specific music or other audio. Thus, in this case, the visitor assistant data 224 can identify a specific temperature setting for the thermostat and a specific radio station to play. In response to receiving utterance 204 and based on the visitor assistant data 224, the host automation assistant can provide an output 228, such as "Okay, I'll play some sounds of nature and set the temperature to 70 degrees." Additionally, based on the visitor assistant data 224, the host automation assistant can cause the thermostat in room 218 to change the temperature setting to 70 degrees and can also present additional audio from the sounds of nature radio station. In this way, computational resources can be conserved when the user can bypass entering certain preferences directly into each assistant device that the user wishes to temporarily personalize. Bypassing such operations can reduce the amount of audio processing or reduce the amount of other input processing otherwise performed to enable the host automation assistant to capture all of the preferences of the visitor user.

[0039] Figure 3FIG. illustrates a system 300 for providing an automated assistant 304 that is capable of operating in an instant personalization mode and / or assisting another automated assistant operating in an instant personalization mode. The automated assistant 304 can operate as part of an assistant application provided at one or more computing devices such as computing device 302 and / or server device. A user can interact with the automated assistant 304 via the (multiple) assistant interfaces 320, which can be a microphone, a camera, a touch screen display, a user interface, and / or any other device capable of providing an interface between the user and the application. For example, a user can initialize the automated assistant 304 by providing verbal, text, and / or graphical input to the assistant interface 320 to cause the automated assistant 304 to initiate one or more actions (e.g., provide data, control a peripheral device, access an agent, generate input and / or output, etc.). Optionally, the automated assistant 304 can be initialized based on the processing of context data 336 using one or more trained machine learning models. The context data 336 can characterize one or more features of the environment accessible to the automated assistant 304 and / or one or more features of a user predicted to intend to interact with the automated assistant 304.

[0040] The computing device 302 can include a display device that can be a display panel including a touch interface for receiving touch input and / or gestures to allow a user to control an application 334 of the computing device 302 via the touch interface. In some embodiments, the computing device 302 can lack a display device, thereby providing an audible user interface output without providing a graphical user interface output. Additionally, the computing device 302 can provide a user interface such as a microphone for receiving spoken natural language input from a user. In some embodiments, the computing device 302 can include a touch interface and can lack a camera, but can optionally include one or more other sensors.

[0041] The computing device 302 and / or other third-party client devices can communicate with the server device via a network such as the Internet. Additionally, the computing device 302 and any other computing device can communicate with each other via a local area network (LAN) such as a Wi-Fi network. The computing device 302 can offload computing tasks to the server device to conserve computing resources at the computing device 302. For example, the server device can host the automated assistant 304, and / or the computing device 302 can transmit input received at one or more assistant interfaces 320 to the server device. However, in some embodiments, the automated assistant 304 can be hosted at the computing device 302 and can perform various processes associated with automated assistant operation at the computing device 302.

[0042] In various embodiments, all or less than all aspects of the automated assistant 304 can be implemented on the computing device 302 (e.g., at a client computing device or a server computing device). Such embodiments can be based on whether the response from the automated assistant 304 corresponds to data not stored at the client computing device and / or whether the response corresponds to an operation to be performed by a separate computing device. In some of these embodiments, aspects of the automated assistant 304 are implemented via the computing device 302 and can be connected to a server device that can implement other aspects of the automated assistant 304. The server device can optionally serve multiple users and their associated assistant applications via multiple threads. In embodiments where all or less than all aspects of the automated assistant 304 are implemented via the computing device 302, the automated assistant 304 can be an application separate from the operating system of the computing device 302 (e.g., installed “on top of” the operating system) — or alternatively can be directly implemented by the operating system of the computing device 302 (e.g., considered an application of the operating system but integrated with the operating system).

[0043] In some embodiments, the automated assistant 304 can include an input processing engine 306 that can employ multiple different modules to process the input and / or output of the computing device 302 and / or the server device. For example, the input processing engine 306 can include a speech processing engine 308 that can process audio data received at the assistant interface 320 to identify the text embodied in the audio data. The audio data can be transmitted from, for example, the computing device 302 to the server device to conserve computing resources at the computing device 302. Additionally or alternatively, the audio data can be specifically processed at the computing device 302.

[0044] The process of converting audio data to text can include a speech recognition algorithm that can use neural networks and / or statistical models to identify audio data groups corresponding to words or phrases. The text converted from the audio data can be parsed by a data parsing engine 310 and made available as text data to an automation assistant 304, which can be used to generate and / or identify one or more command phrases, one or more intents, one or more actions, one or more slot values, and / or any other content specified by the user. In some embodiments, the output data provided by the data parsing engine 310 can be provided to a parameter engine 312 to determine whether the user has provided input corresponding to a particular intent, action, and / or routine that can be performed by the automation assistant 304 and / or an application or agent accessible via the automation assistant 304. For example, assistant data 338 can be stored at the server device and / or the computing device 302 and can include data defining one or more actions that can be performed by the automation assistant 304, as well as the parameters required to perform the actions. The parameter engine 312 can generate one or more parameters for the intent, action, and / or slot value and provide the one or more parameters to an output generation engine 314. The output generation engine 314 can use the one or more parameters to communicate with an assistant interface 320 to provide output to the user and / or communicate with one or more applications 334 to provide output to the one or more applications 334.

[0045] In some embodiments, the automation assistant 304 can be an application that can be installed "above" the operating system of the computing device 302 and / or can itself form part (or all) of the operating system of the computing device 302. The automation assistant application includes and / or can access on-device speech recognition, on-device natural language understanding, and on-device fulfillment. For example, on-device speech recognition can be performed using an on-device speech recognition module that uses an end-to-end speech recognition machine learning model locally stored at the computing device 302 to process audio data (detected by one or more microphones). The on-device speech recognition generates recognized text of the utterance(s) present in the audio data, if any. Additionally, for example, on-device natural language understanding (NLU) can be performed using an on-device NLU module that processes the recognized text generated using on-device speech recognition and optionally context data to generate NLU data.

[0046] NLU data can include the (multiple) intents corresponding to the utterance and optionally the (multiple) parameters of the (multiple) intents (e.g., time slot values). Device-side fulfillment can be performed using a device-side implementation module that utilizes the NLU data (from device-side NLU) and optionally other local data to determine the (multiple) actions to take to resolve the (multiple) intents of the utterance (and optionally the (multiple) parameters of the intents). This can include determining local and / or remote responses (e.g., answers) to the utterance, (multiple) interactions with (multiple) locally installed applications based on the utterance execution, commands transmitted to (multiple) Internet of Things (IoT) devices based on the utterance (either directly or via (multiple) corresponding remote systems), and / or (multiple) other parsing actions based on the utterance execution. Then, device-side fulfillment can initiate the local and / or remote execution / implementation of the determined (multiple) actions to resolve the utterance.

[0047] In various embodiments, remote speech processing, remote NLU, and / or remote fulfillment can be utilized at least selectively. For example, the recognized text can be transmitted at least selectively to the (multiple) remote automation assistant components for remote NLU and / or remote fulfillment. For example, the recognized text can be transmitted for remote execution optionally in parallel with the device-side execution or in response to a failure of the device-side NLU and / or device-side implementation. However, device-side speech processing, device-side NLU, device-side fulfillment, and / or device-side implementation can be prioritized at least because of the reduced latency they provide in resolving the utterance (due to the absence of the need for (multiple) client-server round trips to resolve the utterance). Additionally, the device-side functions can be the only functions available in the absence of a network connection or with a limited network connection.

[0048] In some embodiments, computing device 302 can include one or more applications 334, which can be provided by a third-party entity different from the entity providing computing device 302 and / or automated assistant 304. The application state engine of automated assistant 304 and / or computing device 302 can access application data 330 to determine one or more actions that can be performed by one or more applications 334, as well as the state of each of the one or more applications 334 and / or the state of the corresponding device associated with computing device 302. The device state engine of automated assistant 304 and / or computing device 302 can access device data 332 to determine one or more actions that can be performed by computing device 302 and / or one or more devices associated with computing device 302. Additionally, application data 330 and / or any other data (e.g., device data 332) can be accessed by automated assistant 304 to generate context data 336, which can characterize the context in which a particular application 334 and / or device is operating, and / or the context in which a particular user is accessing computing device 302, accessing application 334, and / or any other device or module.

[0049] When one or more applications 334 are executing at computing device 302, device data 332 can characterize the current operating state of each application 334 executing at computing device 302. Additionally, application data 330 can characterize one or more features of the executed application 334, such as the content of one or more graphical user interfaces presented under the direction of one or more applications 334. Alternatively or additionally, application data 330 can characterize an action mode, which can be updated by the corresponding application and / or by automated assistant 304 based on the current operating state of the corresponding application. Alternatively or additionally, one or more action modes of one or more applications 334 can remain static but can be accessed by the application state engine to determine appropriate actions to be initiated via automated assistant 304.

[0050] The computing device 302 can also include an assistant invocation engine 322 that can use one or more trained machine learning models to process application data 330, device data 332, context data 336, and / or any other data accessible to the computing device 302. The assistant invocation engine 322 can process this data to determine whether to wait for the user to explicitly utter an invocation phrase to invoke the automated assistant 304 or to consider that the data indicates the user's intent to invoke the automated assistant—rather than requiring the user to explicitly utter an invocation phrase. For example, one or more trained machine learning models can be trained using instances of training data based on scenarios where the user is in an environment where multiple devices and / or applications exhibit various operational states. Instances of training data can be generated to capture training data that characterizes contexts in which the user invokes the automated assistant and other contexts in which the user does not invoke the automated assistant.

[0051] When training one or more trained machine learning models based on these instances of training data, the assistant invocation engine 322 can cause the automated assistant 304 to detect or limit the detection of an uttered invocation phrase from the user based on characteristics of the context and / or environment or the user's non-verbal activity. Additionally or alternatively, the assistant invocation engine 322 can cause the automated assistant 304 to detect or limit the detection of one or more assistant commands from the user based on characteristics of the context and / or environment. In some embodiments, the assistant invocation engine 322 can be disabled or limited based on the computing device 302 detecting an assistant suppression output from another computing device. In this way, when the computing device 302 detects an assistant suppression output, the automated assistant 304 will not be invoked based on the context data 336—otherwise the context data 236 might cause the automated assistant 304 to be invoked if no assistant suppression output is detected.

[0052] In some embodiments, system 300 can include a visitor correlation engine 316. The visitor correlation engine 316 can be used to perform one or more operations to determine whether the user providing input to the automated assistant 304 is a visitor user or a host user. Alternatively or additionally, when a visitor user provides input to the automated assistant 304, either indirectly or directly, the visitor correlation engine 316 can determine whether the visitor user is within a threshold vicinity of the computing device 302 or an associated computing device. For example, the visitor correlation engine 316 can determine that the voice signature or face embedding associated with the user who has provided input does not correspond to the user logged into the automated assistant 304 or otherwise has (a) specific access permission(s) to the automated assistant 304. Then, the visitor correlation engine 316 can conclude that the user is a visitor user. When the visitor correlation engine 316 determines that a visitor user is using the automated assistant 304, either directly or indirectly, the visitor correlation engine 304 can invoke a visitor signature engine 318 to identify another assistant device associated with the visitor user interacting with the automated assistant 304.

[0053] The visitor signature engine 318 can use a true signature and / or embedding associated with the visitor user to identify one or more other devices that may be associated with the visitor user. For example, the visitor signature engine 318 can use a voice embedding to encrypt communications that can be sent to one or more other devices. A device that can decrypt the communications and indicate to the automated assistant 304 that the device has successfully decrypted the communications can be considered associated with the visitor user. For example, a visitor device can decrypt the communications using the same or a similar voice embedding generated from one or more previous interactions between the visitor device and the visitor user. Alternatively or additionally, the visitor signature engine 318 can identify a secret that only certain devices can access (e.g., a personal identification code presented at a user interface for pairing purposes), and the secret can be used to associate a specific visitor device with the visitor user. When the automated assistant 304 determines that a visitor device is associated with the visitor user providing the input, the automated assistant 304 can further communicate with the visitor device to enable a visitor automated assistant associated with the visitor user to assist in processing the input received from the visitor user. Then, the visitor device can provide response data in response to a request from the host automated assistant 304.

[0054] In some embodiments, the automated assistant 304 can include a pattern preference engine 324 that can determine one or more preferences of a guest user or an acquaintance of the guest user interacting with the host automated assistant. For example, the automated assistant 304 can receive requests or provide requests to identify one or more preferences that a user may have when interacting with their respective automated assistants. Such preferences can include preferences explicitly identified by the user or preferences that are appropriate for the user over time. For example, the automated assistant can provide preference data that identifies one or more trained machine learning models that can be used when processing input from or output to the user. For example, the trained machine learning models can include ASR models, speech-to-text models, text-to-speech models, and / or any other type of trained computer learning model that can be used during one or more operations of the automated assistant. This can allow the host automated assistant to provide responses that can be more easily interpreted by the guest user because the responses can be pronounced, for example, in a manner that the host automated assistant does not typically pronounce to the host user.

[0055] In some embodiments, the automated assistant 304 can include a personal query engine 326 that can determine whether an input from a user is associated with information that can be personalized for a specific user. For example, the personal query engine 326 can use one or more trained machine learning models to determine whether an input to the automated assistant 304 and / or other interactions with the automated assistant 304 are associated with information that can be personalized for a specific user. In some embodiments, the personal query engine 326 can be optional and can optionally cause the automated assistant 304 to transition to an instantaneous personalization mode when a guest user provides an input that is determined to be associated with personalized information. Alternatively or additionally, when the personal query engine 326 determines that an input or interaction is not associated with personal information (e.g., the input is a request that can be satisfied using public data not associated with a specific user account), the personal query engine 326 can omit causing the automated assistant 304 to transition to the instantaneous personalization mode.

[0056] Figure 4 Illustrated is a method 400 for processing requests from a host automated assistant when the host automated assistant attempts to operate in an instantaneous personalization mode. The method 400 can be performed by one or more applications, devices, and / or any other device or module capable of performing operations associated with the automated assistant. The method 400 can include an operation 402 of determining whether a relevance request has been received from the host automated assistant. This determination can be made at a guest device that provides access to a guest automated assistant that can be associated with a user in the vicinity of another assistant-enabled device.

[0057] When receiving a relevance request from a host automation assistant, method 400 can proceed from operation 402 to operation 404, which can include determining whether a guest user can be relevant to the input of the host automation assistant. In some embodiments, the guest device can receive encrypted data from the host device and can encrypt the encrypted data using a value generated based on a unique input from the user. For example, the value can be a voice vector or a voice embedding, which is based on the user's voice feature(s) when the user provides a spoken input to the host automation assistant. In this way, since the guest automation assistant has previously received the utterance from the guest user, the guest automation assistant will be able to decrypt the encrypted data transmitted from the host automation assistant.

[0058] When the host automation assistant determines that the guest device or the guest automation assistant is associated with the user providing the input to the host automation assistant, method 400 can proceed to operation 406. Otherwise, method 400 can return to operation 402. Operation 406 can be an optional operation that includes transmitting an authentication value to the host automation assistant. The authentication value can be, for example, a secret generated by the host automation assistant, and it is expected that only the guest device logged in by the user can decrypt the encrypted data and identify the authentication value. Alternatively or additionally, query data representing one or more requests embodied in the input from the user can be received by the guest automation assistant and can operate without transmitting the authentication value back to the host device.

[0059] Method 400 can proceed from operation 404 or operation 406 to operation 408, and operation 408 can include processing a request identifying one or more assistant queries from the user. One or more assistant queries can be embodied in the utterance from the user to the host automation assistant. However, the host automation assistant can transmit the request representing one or more assistant queries to the guest automation assistant. In response to receiving the request, the guest automation assistant or the guest device can generate response data based on the one or more assistant queries. For example, the guest automation assistant can process the queries as if the user had provided the queries directly to the guest automation assistant. As a result, the guest automation assistant can generate response data that can represent the output and / or other data processed by the host automation assistant to fulfill the input from the user to the host automation assistant.

[0060] Method 400 can proceed from operation 410 to operation 412, and operation 412 can include causing the host automation assistant to present an output based on the response data. For example, the response data can represent natural language content that can be presented at one or more interfaces of the host device. The natural language content can be in response to the utterance provided by the user to the host automation assistant. In this way, when the user is outside their home, the user can quickly personalize nearby automation assistants that can operate in an instantaneous personalization mode.

[0061] Figure 5 FIG. illustrates method 500 for operating an automated assistant in an instantaneous personalization mode when one or more guest users interact with the automated assistant. Method 500 can be performed by one or more applications, devices, and / or any other device or module capable of providing access to the automated assistant. Method 500 can include an operation 502 of determining whether an input is received from a guest user at a host automated assistant. A guest user can be a person who is not logged in to the host automated assistant and / or the owner account of the device currently authorized to access the host automated assistant device that provides access to the host automated assistant. When it is determined that an input has been received from the guest user, method 500 can proceed from operation 502 to operation 504. Otherwise, the host automated assistant can continue to determine whether the guest user has provided an input.

[0062] Operation 504 can include providing a relevance request to a guest device operating in the vicinity of the host device. The relevance request can be a request for a nearby device to indicate that the device is associated with the guest user who provided the input to the host automated assistant. Method 500 can proceed from operation 504 to operation 506, which can include determining whether the guest device can be relevant to the input from the guest user. In some embodiments, the guest device can be relevant to the input when the guest device can decrypt an authentication value that encrypts information using the input from the guest user. For example, the authentication value can be encrypted using a face embedding, voice embedding, image embedding, video embedding, and / or any other signature of the guest user. Thus, when the guest device can decrypt the authentication value using a similar embedding and transmit the authentication value back to the host device, method 500 can proceed to operation 510. Otherwise, method 500 can proceed to operation 508, which can include responding to the guest user without relying on the guest automated assistant.

[0063] Operation 510 can include providing a request based on one or more assistant queries embodied in the input from the user. For example, in some embodiments, the host automated assistant can transmit input data to the guest automated assistant so that the guest automated assistant can generate response data based on the input data. Alternatively or additionally, the host automated assistant can transmit a request to the guest automated assistant to obtain user preferences for responding to one or more assistant queries. In some embodiments, the user preferences can include, but are not limited to, a voice profile or pronunciation that the host automated assistant should adopt when presenting a response to the guest user so that the guest user can more easily interpret the output from the host automated assistant.

[0064] Method 500 can proceed from operation 510 to operation 512, and operation 512 can include processing response data based on one or more assistant queries. For example, in some embodiments, the response data can embody audio data, text data, natural language processing (NLP) data, such as action intents and / or parameters, and / or any other data that can be used as a basis for generating an automated assistant response. Method 500 can proceed from operation 512 to operation 514, and operation 514 can include causing a host automated assistant to present an output based on the response data. For example, when the host automated assistant receives NLP data, the host automated assistant can use any parameters also identified in the NLP data to perform one or more actions identified by the NLP data.

[0065] Figure 6 is a block diagram 600 of an example computer system 610. Computer system 610 generally includes at least one processor 614 communicating with a plurality of peripheral devices via a bus subsystem 612. These peripheral devices can include a storage subsystem 624, which includes, for example, a memory 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow a user to interact with computer system 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computer systems.

[0066] The user interface input device 622 can include a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen incorporated in a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and ways of inputting information into computer system 610 or onto a communication network.

[0067] The user interface output device 620 can include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem can include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display, such as via an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and ways of outputting information from computer system 610 to a user or another machine or computer system.

[0068] Storage subsystem 624 stores programming and data structures that provide the functionality of some or all of the modules described herein. For example, storage subsystem 624 may include logic to perform method 400, selected aspects of method 500, and / or implement one or more of host device 108, personal device 110, host television 106, host device 208, personal device 210, host television 206, system 300, and / or any other application, device, apparatus, and / or module discussed herein.

[0069] These software modules are usually executed by processor 614 alone or in combination with other processors. The memory 625 used in storage subsystem 624 can include multiple memories, including a main random access memory (RAM) 630 for storing instructions and data during program execution and a read-only memory (ROM) 632 storing fixed instructions. File storage subsystem 626 can provide persistent storage for program and data files, and can include a hard disk drive, a floppy disk drive and associated removable media, a CD-ROM drive, an optical drive, or a removable media box. Modules that implement the functions of certain embodiments can be stored in storage subsystem 624 by file storage subsystem 626, or stored in other machines accessible to processor (s) 614.

[0070] The bus subsystem 612 provides a mechanism for the various components and subsystems of the computer system 610 to communicate with each other as intended. Although the bus subsystem 612 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple busses.

[0071] Computer system 610 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, Figure 6 The description of the computer system 610 depicted in FIG. 6 is intended only as a specific example for illustrating some embodiments. Many other configurations of the computer system 610 may have more Figure 6 The computer system may have more or fewer components than those depicted.

[0072] In the systems described herein, in cases where personal information about a user (or “participant” as often referred to herein) is collected or personal information may be used, the user may be provided with the opportunity to control whether the program or feature collects user information (e.g., information about the user's social network, social behavior or activities, occupation, user preferences or the user's current geographic location), or to control whether and / or how content that may be more relevant to the user is received from a content server. Additionally, before storing or using certain data, the data may be processed in one or more ways so as to remove personally identifiable information. For example, the user's identity may be processed so that personally identifiable information cannot be determined for the user, or where geographic location information is obtained, the user's geographic location may be generalized (such as to city, zip code, or state level) so that the user's specific geographic location cannot be determined. Accordingly, the user may control how information about the user is collected and / or used.

[0073] Although several implementations have been described and illustrated herein, various other devices and / or structures may be used for performing the functions and / or obtaining the results and / or one or more advantages described herein, and each of these variations and / or modifications is considered to be within the scope of the implementations described herein. More generally, all of the parameters, dimensions, materials, and configurations described herein are exemplary, and the actual parameters, dimensions, materials, and / or configurations will depend upon one or more specific applications of the teachings. Those skilled in the art will recognize, or be able to ascertain using only routine experimentation, many equivalents to the specific implementations described herein. Accordingly, it is to be understood that the above-described implementations are presented by way of example only, and that the implementations may be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. The implementations of the present disclosure are directed to each separate feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods that are not mutually inconsistent is included within the scope of the present disclosure.

[0074] In some embodiments, a method implemented by one or more processors is set forth as including operations such as: receiving, at a first computing device, a request for the first computing device to process an utterance submitted by a user to a second computing device, where each of the first computing device and the second computing device is located in a common environment and provides access to a respective automated assistant, and where the second computing device encrypts the request using signature data generated by the second computing device using a biometric signature corresponding to the user. The operations can further include: processing, by the first computing device, the request from the second computing device to identify one or more assistant requests embodied in the request. The operations can further include: generating, by the first computing device, assistant response data that characterizes one or more automated assistant responses in response to the one or more assistant requests. The operations can further include: causing, by the first computing device, the second computing device to present the one or more automated assistant responses to the user using the assistant response data.

[0075] In some embodiments, processing the request from the second computing device includes: accessing, by the first computing device, other signature data associated with the user and identifying an authentication value embodied in the request from the second computing device or other data using the other signature data. In some embodiments, causing the second computing device to present the one or more automated assistant responses includes: providing, from the first computing device to the second computing device, the authentication value, where the authentication value is generated by the second computing device in response to receiving an utterance from the user. In some embodiments, generating the assistant response data includes: accessing, by the first computing device, stored content not stored at the second computing device when the second computing device receives an utterance from the user.

[0076] In some embodiments, generating the assistant response data includes: accessing content associated with the user's account, where the second computing device is not authenticated to directly access the user's account. In some embodiments, causing the second computing device to present the one or more automated assistant responses includes: transmitting the assistant response data from the first computing device to the second computing device via a local area network, a Bluetooth connection, or a wide area network, where transmitting the assistant response data causes the second computing device to present the one or more automated assistant responses. In some embodiments, the method can further include the operation of: providing, at an interface of the first computing device and in response to receiving a request from the second computing device, a prompt that allows the user to select whether to permit the first computing device to respond to the request or a subsequent request from the second computing device. In some embodiments, the method can further include the operation of: providing, at an interface of the first computing device and in response to receiving a request from the second computing device, a prompt that allows the user to restrict when to permit the first computing device to respond to the request or a subsequent request from the second computing device.

[0077] In other embodiments, a method implemented by one or more processors is described as including operations such as: receiving an utterance from a user associated with a first computing device, where the utterance is received at a second computing device in a common environment with the first computing device and the user, and where each of the first computing device and the second computing device provides access to a respective automated assistant. The operations can also include: the second computing device providing a first request to the first computing device to confirm that the user has been authenticated on the first computing device, where the first request embodies an authentication value accessible by one or more devices that can be authenticated by the user. The operations can also include: the second computing device receiving the authentication value, which indicates to the second computing device that the first computing device can access the authentication value. The operations can also include: the second computing device and based on the authentication value, providing a second request for the first computing device to respond to one or more assistant requests embodied in the utterance. The operations can also include: the second computing device and in response to providing the second request, receiving assistant response data in response to the one or more assistant requests embodied in the utterance. The operations can also include: the second computing device causing one or more interfaces of the second computing device to present an automated assistant output based on the assistant response data.

[0078] In some embodiments, the operations can also include: the second computing device identifying the user's true signature; the second computing device generating the first request by encrypting the authentication value using the true signature. In some embodiments, the operations can also include: the second computing device using the true signature to process the assistant response data, where the assistant response data is encrypted by the first computing device using the true signature. In some embodiments, the user's true signature corresponds to an audio-based signature or an image-based signature. In some embodiments, the operations can also include: in response to receiving the utterance, determining that the utterance embodies one or more requests to access content that the second computing device is not currently permitted to access. In some embodiments, providing the second request for the first computing device to respond to one or more assistant requests includes: providing audio data or text data to the first computing device that characterizes one or more portions of the utterance provided by the user to the second computing device. In some embodiments, providing the second request for the first computing device to respond to one or more assistant requests includes: in response to the user providing an utterance to the second computing device, providing action data to the first computing device that characterizes one or more automated assistant actions to be performed by the automated assistant.

[0079] In yet other embodiments, a method implemented by one or more processors is described as including operations such as: receiving an utterance from a user associated with a first computing device, where the utterance is received at a second computing device that is in a common environment with the first computing device and the user, and where each of the first computing device and the second computing device provides access to a respective automated assistant. The operations can further include: the second computing device providing to the first computing device a first request for the first computing device to confirm that the user is authenticated on the first computing device, where the first request embodies an authentication value that can be accessed by one or more devices that can be authenticated by the user. The operations can further include: when the first computing device can access the authentication value: the second computing device receiving authentication data that indicates to the second computing device that the first computing device can access the authentication value. The operations can further include: the second computing device, and based on the first computing device being able to access the authentication value, providing a second request for the first computing device to provide user preference data in response to one or more assistant requests embodied in the utterance. The operations can further include: the second computing device, and in response to providing the second request, receiving user preference data that identifies one or more user preferences to be employed by the automated assistant of the second computing device when responding to one or more assistant requests submitted by the user. The operations can further include: the second computing device causing one or more interfaces of the second computing device to present an automated assistant output based on the user preference data.

[0080] In some embodiments, the method can further include operations such as: generating automated assistant output data on which the automated assistant output is further based, based on the user preference data, where the user preference data identifies one or more automatic speech recognition models to be used when processing an utterance from the user. In some embodiments, the operations can further include: generating automated assistant output data on which the automated assistant output is further based, based on the user preference data, where the user preference data identifies one or more text-to-speech models to be used when presenting the automated assistant output to the user. The operations can further include: generating automated assistant output data in response to one or more assistant requests, based on the user preference data, where the user preference data identifies a content ranking of candidate content identified by the second computing device when generating the automated assistant output data. The operations can further include: when the first computing device cannot access the authentication value: the second computing device causing one or more interfaces of the second computing device to present a different automated assistant output that is not based on the user preference data.

Claims

1. A method implemented by one or more processors, the method comprising: Receiving, at a first computing device, an encryption request for the first computing device to process utterances submitted by a user to a second computing device, wherein each of the first computing device and the second computing device is located in a common environment and provides access to a respective automated assistant, wherein the second computing device, in response to the user submitting the utterances: Encrypts the encryption request using signature data generated by the second computing device using a biometric signature corresponding to the user, and Provides the encrypted request to the first computing device in the common environment, and wherein each of the first computing device and the second computing device is a client computing device; Processing, by the first computing device, the encrypted request from the second computing device to identify one or more assistant requests embodied in the encrypted request; Generating, by the first computing device, assistant response data that characterizes one or more automated assistant responses in response to the one or more assistant requests; and Causing, by the first computing device, the second computing device to present the one or more automated assistant responses to the user using the assistant response data.

2. The method according to claim 1, wherein, Processing the encrypted request received from the second computing device includes: The first computing device accessing other signature data associated with the user, and Using the other signature data to identify an authentication value embodied in the encrypted request or other data from the second computing device.

3. The method according to claim 2, wherein, Causing the second computing device to present the one or more automated assistant responses includes: Providing the authentication value from the first computing device to the second computing device, wherein the authentication value is generated by the second computing device in response to the second computing device receiving the utterances from the user.

4. The method according to claim 1, wherein, Generating the assistant response data includes: When the second computing device receives the utterances from the user, the first computing device accessing stored content not stored at the second computing device.

5. The method according to claim 1, wherein Generating the assistant response data includes: Accessing content associated with the user's account, wherein the second computing device is not authenticated to directly access the user's account.

6. The method according to claim 1, wherein Causing the second computing device to present the one or more automated assistant responses includes: Transmitting the assistant response data from the first computing device to the second computing device via a local area network, a Bluetooth connection, or a wide area network, wherein transmitting the assistant response data causes the second computing device to present one or more automated assistant responses.

7. The method according to claim 1, further comprising: At an interface of the first computing device and in response to receiving the encrypted request from the second computing device, providing a prompt that allows the user to select whether to permit the first computing device to respond to the encrypted request or subsequent encrypted requests from the second computing device.

8. The method according to any one of claims 1-7, further comprising: At an interface of the first computing device and in response to receiving the encryption request from the second computing device, provide a prompt that allows the user to restrict when the first computing device is permitted to respond to the encryption request or subsequent encryption requests from the second computing device.

9. A system for transiently adjusting the handling of automated assistant requests, the system comprising: one or more computers; and one or more storage devices storing instructions that are operative and, when executed by the one or more computers, cause the one or more computers to perform operations, the operations including: receiving, at a first computing device, an encryption request for the first computing device to process utterances submitted by a user to a second computing device, wherein each of the first computing device and the second computing device is located in a common environment and provides access to a respective automated assistant, wherein the second computing device, in response to the user submitting the utterance: encrypts the encryption request using signature data generated by the second computing device using a biometric signature corresponding to the user, and provides the encrypted request to the first computing device in the common environment, and wherein each of the first computing device and the second computing device is a client computing device; processing, by the first computing device, the encrypted request from the second computing device to identify one or more assistant requests embodied in the encrypted request; generating, by the first computing device, assistant response data that characterizes one or more automated assistant responses in response to the one or more assistant requests; and causing, by the first computing device, the second computing device to present the one or more automated assistant responses for the user using the assistant response data.

10. The system according to claim 9, wherein Processing the encrypted request received from the second computing device includes: accessing, by the first computing device, other signature data associated with the user, and identifying, using the other signature data, an authentication value embodied in the encrypted request or other data from the second computing device.

11. The system according to claim 10, wherein, Causing the second computing device to present the one or more automated assistant responses includes: providing, from the first computing device to the second computing device, the authentication value, wherein the authentication value is generated by the second computing device in response to the second computing device receiving the utterance from the user.

12. The system according to claim 9, wherein, Generating the assistant response data includes: accessing, by the first computing device, stored content not stored at the second computing device when the second computing device receives the utterance from the user.

13. The system according to claim 9, wherein, Generating the assistant response data includes: accessing content associated with the user's account, wherein the second computing device is not authenticated to directly access the user's account.

14. The system according to claim 9, wherein, Causing the second computing device to present the one or more automated assistant responses includes: transmitting the assistant response data from the first computing device to the second computing device via a local area network, a Bluetooth connection, or a wide area network, Transmitting the assistant response data causes the second computing device to present one or more automated assistant responses.

15. The system according to claim 9, wherein, The operation further includes: At an interface of the first computing device and in response to receiving the encrypted request from the second computing device, providing a prompt that allows the user to select whether to permit the first computing device to respond to the encrypted request or a subsequent encrypted request from the second computing device.

16. The system according to claim 9, wherein The operation further includes: At an interface of the first computing device and in response to receiving the encrypted request from the second computing device, providing a prompt that allows the user to restrict when to permit the first computing device to respond to the encrypted request or a subsequent encrypted request from the second computing device.

17. A non-transitory computer-readable storage medium including instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including: Receiving, at a first computing device, an encrypted request for the first computing device to process utterances submitted by a user to a second computing device, wherein each of the first computing device and the second computing device is located in a common environment and provides access to a respective automated assistant, wherein the second computing device, in response to the user submitting the utterance: Encrypts the encrypted request using signature data generated by the second computing device using a biometric signature corresponding to the user, and Provides the encrypted request to the first computing device in the common environment, and wherein each of the first computing device and the second computing device is a client computing device; Processing the encrypted request received from the second computing device by the first computing device to identify one or more assistant requests embodied in the encrypted request; Generating, by the first computing device, assistant response data that characterizes one or more automated assistant responses in response to the one or more assistant requests; and Causing, by the first computing device, the second computing device to present the one or more automated assistant responses for the user using the assistant response data.

18. The non-transitory computer-readable storage medium according to claim 17, wherein, Processing the encrypted request received from the second computing device includes: The first computing device accessing other signature data associated with the user, and Using the other signature data to identify an authentication value embodied in the encrypted request or other data from the second computing device.

19. The non-transitory computer-readable storage medium according to claim 18, wherein, Causing the second computing device to present the one or more automated assistant responses includes: Providing the authentication value from the first computing device to the second computing device, wherein the authentication value is generated by the second computing device in response to receiving the utterance from the user.

20. The non-transitory computer-readable storage medium according to claim 17, wherein, Generating the assistant response data includes: When the second computing device receives the utterance from the user, the first computing device accessing stored content not stored at the second computing device.

Citation Information

Patent Citations

  • Virtual assistant operation in multi-device environments

    US20190371315A1

  • Centralized gateway server for providing access to services

    US20200045041A1