Methods and systems for pausing conversations between users and automated assistants
By storing the conversation state on the client device and generating prompts using a non-assistant platform, the conversation is resumed asynchronously, solving the problems of resource waste and poor experience in user-automated assistant conversations, and achieving efficient conversation management and user interaction processing.
Patent Information
- Application Number
- CN202180062903.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-11-10
- Filing Date
- 2021-09-17
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2041-09-17
AI Technical Summary
In conversations between users and automated assistants, existing technologies cannot effectively manage the conversation state when pausing and resuming, resulting in wasted resources and a poor user experience.
By storing the conversation session state on the client device and generating prompts to complete user interactions using a non-assistant platform, the conversation session is resumed asynchronously. Tokens are used to manage session recovery, reducing resource consumption and improving the user experience.
It enables efficient pausing and resuming of conversations, reduces resource consumption on client devices, improves user experience, and supports asynchronous processing of complex user interactions.
Smart Images

Figure CN116076062B_ABST
Abstract
Description
Background Technology
[0001] Humans can participate in human-computer dialogue sessions with interactive software applications, referred to herein as “automated assistants” (also known as “chatbots,” “interactive personal assistants,” “intelligent personal assistants,” “personal voice assistants,” “conversational agents,” etc.). For example, a human (who may be referred to as a “user” when interacting with an automated assistant) can provide input (e.g., commands, queries, and / or requests) to the automated assistant, which can cause the automated assistant to generate and provide response outputs to control one or more Internet of Things (IoT) devices and / or to perform one or more other functions. Input provided by the user can be, for example, spoken natural language input (i.e., speech) that can be converted into text (or other semantic representations) and then further processed in some cases, and / or typed natural language input.
[0002] In some cases, an automated assistant may include an automated assistant client, executed locally by a client device and directly interacted with by the user; and a cloud-based counterpart that leverages virtually unlimited cloud resources to assist the automated assistant client in responding to user input. For example, the automated assistant client may provide the cloud-based counterpart with an audio recording of the user's spoken words (or their text-based equivalent), and optionally with data indicating the user's identity (e.g., credentials). The cloud-based counterpart may perform various processes on the query to return the results to the automated assistant client, which can then provide the corresponding output to the user.
[0003] Many users can use multiple client devices to interact with the automation assistant. For example, some users may have a coordinated "ecosystem" of client devices, such as one or more smartphones, one or more tablets, one or more vehicle computing systems, one or more wearable computing devices, one or more smart TVs, and / or one or more independent interactive speakers. Users can use any of these client devices to participate in a human-machine dialogue session with the automation assistant (assuming the automation assistant client is installed). In some cases, a given dialogue session can be interrupted to perform other tasks, but when the given dialogue session is resumed, the automation assistant may not be aware of the context of these other tasks. Summary of the Invention
[0004] The implementation described herein relates to suspending the conversation between the user and the automation assistant at the client device and via the automation assistant platform in response to determining that user input received during a conversation session requires user interaction with a non-assistant platform different from the automation assistant platform, and asynchronously resuming the conversation between the user and the automation assistant at the client device or an attached client device based on the stored state of the conversation session and based on the result of the user's interaction with the non-assistant platform. When the conversation session is suspended, the automation assistant may cause the state of the conversation session to be stored in memory (e.g., at the client device) and / or (e.g., at the client device and / or a remote server) or one or more databases, and may transmit a request to the non-assistant platform. The non-assistant platform may generate one or more prompts to complete the user interaction in response to receiving the request, and may render one or more prompts at the client device or an attached client device and via the non-assistant platform. Additional user input may be received in response to rendering one or more prompts to complete the user interaction, and one or more tokens associated with the user interaction may be generated. The automation assistant may then resume the conversation session at the client device or an attached client device based on the stored state of the conversation session and based on the generated one or more tokens associated with the user interaction.
[0005] By pausing and asynchronously resuming a conversation session according to the techniques described herein, the conversation session can be resumed via the automation assistant platform using the context of the stored conversation session. This eliminates the need for the user to initiate a new conversation session before pausing to achieve the same conversation state, thereby reducing user input received at the client device and saving computational resources. Furthermore, pausing and storing this state allows the automation assistant at the client device to be used in other automation assistant conversation sessions without affecting the stored state. This state can optionally be stored in local and / or remote non-volatile memory and cleared from limited volatile memory upon storage, making the volatile memory available for and / or used by other processes. Additionally, the conversation session can be resumed via the automation assistant platform based on one or more tokens associated with user interactions completed via a non-automation assistant platform, allowing the stored conversation session to advance beyond its stored state and to another state affected by both the stored state and the tokens.
[0006] An automated assistant can determine, based on processing user input, that the user input requires interaction with different non-assistant platforms. For example, suppose the user input is spoken utterance. A speech-to-text module can be used to process the spoken utterance to generate discriminative text, and a natural language processor can be used to parse the user's intent included in the spoken utterance to process the discriminative text. Further suppose the user's parsed intent requires interaction with a non-assistant platform to authenticate and / or verify the user's account, enter credit card information, verify credit card information, etc., to continue the conversation. In response to determining that the user input requires user interaction with a non-assistant platform, the automated assistant can store the state of the conversation and pause the conversation. The stored state of the conversation can include, for example, user input provided by the user during the conversation, responses provided by the automated assistant during the conversation, contextual information associated with the conversation (e.g., user location, the time the conversation was initiated and / or paused, the duration of the conversation, etc.), the user's current intent and / or the currently parsed slot values of parameters associated with the current intent, the user's past intents and / or the past parsed slot values of parameters associated with past intents, and / or other data associated with the conversation. Furthermore, the automation assistant can display an indication that the conversation has been paused on the client device, allowing the user to complete the interaction. Notably, while the conversation is paused, the user can participate in other conversations with the automation assistant and / or have the automation assistant perform assistant-based actions.
[0007] In some implementations, the non-assistant platform may be a first-party platform that shares a public publisher with the automation assistant. For example, a first-party platform may include an email platform, a navigation platform, an Internet of Things (IoT) device platform, a web-based platform, a software application platform, and / or other platforms that share a public publisher with the automation assistant. In some additional or alternative implementations, the non-assistant platform may be a third-party platform that does not share a public publisher with the automation assistant. For example, a third-party platform may include platforms similar to the first-party platforms listed above, but which do not share a public publisher with the automation assistant. In other words, the platform may be a third-party platform because it is controlled by a third-party entity that is different from the first-party entity controlling the assistant platform, and the first party does not have any direct control over that third-party entity. It is noteworthy that the non-assistant platform can be utilized by the automation assistant when performing various actions. However, according to the techniques described herein, the automation assistant is not utilized when initiating or performing user interactions with the non-assistant platform.
[0008] In some implementations, one or more prompts for completing a user interaction may be generated by a non-assistant platform in response to receiving a request for user interaction initiated via the non-assistant platform. One or more prompts may be generated based on user input (or intent determined based on user input) provided by the user before pausing the conversation and / or a response provided by an assistant before pausing the conversation. Furthermore, one or more prompts may be transmitted as, for example, electronic communications associated with the non-assistant platform (e.g., text messages, instant messages, emails, etc.), software application notifications from software applications associated with the non-assistant platform, as part of data sent to an application programming interface associated with the non-assistant platform, and / or other representations.
[0009] In some versions of those implementations, one or more prompts may be sent to the same client device in which the conversation session is paused, while in other versions, one or more prompts may be sent to an additional client device different from the client device in which the conversation session is paused. One or more prompts may be sent to the client device and / or the additional client device based on device capabilities. For example, if user interaction requires a client device that includes a display, but the client device in which the conversation session is paused is a standalone speaker device lacking a display, one or more prompts may be sent to an additional client device associated with the user that includes a display (e.g., the user's mobile device). However, if the client device in which the conversation session is paused includes a display, one or more prompts may be sent back to the client device in which the conversation session is paused.
[0010] In some implementations, one or more prompts may be rendered in response to receiving one or more prompts at a client device or attached client device and via a non-assistant platform. The one or more prompts may request additional user input in response to the prompts, and the user interaction may be completed at the client device or attached client device. For example, the one or more prompts may request additional user input to confirm a user account linked to an account on the client device, enter new credit card information, verify current credit card information, confirm an address, check into accommodations, prompt a service provider, purchase game credits for gaming, and / or authenticate and / or verify other user information or transactions that may require the user to specify one or more values. In various implementations, the one or more prompts may additionally or alternatively require the user to provide biometric information in addition to additional user input, or biometric information in lieu of additional user input. Biometric information may include, for example, fingerprint verification via a fingerprint scanner associated with the user, voice verification via one or more microphones associated with the user, facial verification via one or more visual components associated with the user, and / or other biometric information.
[0011] One or more tokens may be, for example, data objects comprising the result of a user interaction completed via a non-assistant platform, one or more values associated with the user interaction, and / or other information provided by the non-assistant platform. In various embodiments, the result of a user interaction on which one or more tokens are generated may not contain any data from the user interaction. For example, if one or more prompts require fingerprint recognition, the result of the interaction may indicate that the user's fingerprint has been recognized, but the actual fingerprint information may not be transmitted. Furthermore, in various embodiments, one or more tokens may be encrypted to ensure user privacy. In some embodiments, one or more tokens may be generated by a non-assistant platform and transmitted back to a client device or an additional client device based on additional user input. In some additional or alternative embodiments, additional user input may be transmitted to a client device and / or a remote computing device, and the client device and / or the remote computing device may generate one or more tokens based on the additional user input. In some embodiments, one or more tokens may be stored in association with the stored state of a conversation session. The stored state of the conversation session and the one or more tokens stored in association with the stored state of the conversation may be accessible by multiple client devices over one or more networks.
[0012] The stored state of a conversation can be loaded on the client device or an attached client device, and the conversation can be resumed based on one or more tokens. One or more tokens may be required to resume a conversation, and the resumption of a conversation may vary depending on the tokens used. For example, if a user is checking into a hotel and the conversation prompts them to verify their user account, the user may also be prompted to purchase a specific TV package for their stay at the hotel, including a choice between a normal TV package and a premium TV package. If the user purchases a normal TV package, the automated assistant may present a first TV guide associated with the normal TV package when the conversation resumes. However, if the user purchases a premium TV package, the automated assistant may then present a second TV guide associated with the premium TV package when the conversation resumes. It is worth noting that there may be a time delay between the first time the conversation is paused and the second time it is resumed, thus the conversation is resumed asynchronously. This time delay can range from a relatively short duration (e.g., seconds) to a relatively long duration (e.g., hours, days, weeks). It should be noted that the time delay between the pause and resumption of the conversation may be based on the time it takes for the user to provide additional user input in response to one or more prompts.
[0013] In some implementations, a conversation session can be automatically resumed at the client device in response to the receipt of a token. For example, a conversation session can be paused at the client device in response to receiving user input that requires interaction with a user on a non-assistant platform, and remain paused until the user interaction is completed at the client device. When the user interaction is completed at the client device (or an additional client device as described above), the paused conversation session can be resumed at the client device. For example, the receipt of a token can cause the stored conversation state to be loaded into the memory at the client device, and the conversation session is resumed, and optionally, the conversation session is resumed to another state depending on the stored state and the received token. In some versions of those implementations, a conversation session can only be resumed at the client device when the user is present in the vicinity of the client device. In some other versions of those implementations, the client device resuming the conversation session can be an additional client device that is different from the client device that paused the conversation session, and / or different from the additional client device that is rendered at its location if one or more prompts are rendered at the additional client device. For example, if a user is located in an ecosystem of multiple client devices that can resume a conversation, the conversation is loaded on the given client device that is closest to the user, and the conversation can be resumed on the given client device based on one or more tokens.
[0014] In some additional or alternative implementations, the conversation session may be resumed at the client device (or an additional client device) in response to receiving user input after a token is received to resume the conversation session after the token is received. For example, the conversation session may be paused at the client device in response to receiving user input that requires interaction with a user on a non-assistant platform, and remain paused until the user interaction is completed at the client device. When the user interaction is completed at the client device (or the additional client device as described above), one or more client devices in an ecosystem with multiple client devices may provide an indication that one or more tokens have been received (e.g., audible indications and / or graphical elements), and may resume the conversation session based on one or more tokens. However, the conversation session may not be resumed until user input to resume the conversation is received. For example, in implementations where the client device includes a display, optional graphical elements may be displayed, and when the user selects the graphical element, the conversation session is resumed at the client device. As another example, the client device may receive spoken utterances to resume the conversation session (e.g., “resume,” “continue,” etc.). In these examples, if a token has already been received at the client device, the speech processing of the spoken utterances may be biased towards the spoken utterances to resume the conversation.
[0015] The foregoing description is provided as an overview of only a few embodiments of this disclosure. Further descriptions of those and other embodiments are provided herein in more detail. As a non-limiting example, various embodiments are described in more detail in the claims and detailed description included herein. Attached Figure Description
[0016] Figure 1 This is a block diagram of an example environment in which the implementation methods disclosed herein can be carried out.
[0017] Figure 2 Example state diagrams according to various implementation methods are illustrated.
[0018] Figure 3 Another example state diagram according to various implementation methods is illustrated.
[0019] Figure 4A , Figure 4B and Figure 4C The illustration shows example dialogue sessions between a user and an automation assistant according to various implementation methods.
[0020] Figure 5A , Figure 5B and Figure 5C The illustration shows another example dialogue session between a user and an automation assistant, according to various implementation methods.
[0021] Figure 6 The illustration shows an example architecture of a computing device. Detailed Implementation
[0022] Go to Figure 1 This paper describes an example environment in which the techniques disclosed herein can be implemented. The example environment includes multiple client computing devices 106 1-N (Hereinafter referred to as “client device 106”), a conversation state database 115, one or more cloud-based automated assistant components 119, one or more token engines 130, one or more non-assistant platforms 140, and a user information database 135. Client device 106 may be communicatively coupled to each other, and to the conversation state database 115, non-assistant platforms 140, and / or other resources (e.g., the Internet) via one or more networks 1101—such as one or more wired or wireless local area networks (“LANs”, including Wi-Fi LANs, mesh networks, Bluetooth, near-field communication, etc.) and / or wide area networks (“WANs”, including the Internet).
[0023] Client device 106 may be one or more of the following: desktop computing devices, laptop computing devices, tablet computing devices, mobile phone computing devices, computing devices in a user's vehicle (e.g., in-vehicle communication systems, in-vehicle entertainment systems, in-vehicle navigation systems), independent interactive speakers (optionally with displays), smart devices such as smart TVs, and / or wearable devices of the user that include computing devices (e.g., a user's watch with computing devices, a user's glasses with computing devices, virtual or augmented reality computing devices). Additional and / or replaceable client computing devices may be provided.
[0024] At least one of the client devices 106 can execute the automation assistant client 118. In some embodiments, each of the client devices 106 can execute, for example... Figure 1 A corresponding instance of the depicted automation assistant client 118. The instance of automation assistant client 118 may be an application separate from the operating system of each of the client devices 106 (e.g., installed "on top" of the operating system), or it may alternatively be implemented directly by the operating system of each of the client devices 106. In other embodiments, a subset of client devices 106 may omit automation assistant client 118. For example, a first client device 1061 may include a corresponding instance of automation assistant client 118, but a second client device 1062 may omit a corresponding instance of automation assistant client 118.
[0025] One or more cloud-based automation assistant components 119 may be implemented on one or more computing systems (collectively referred to as "cloud" or "remote" computing systems) communicatively coupled to client device 106 via one or more LANs and / or WANs. The communication coupling between the cloud-based automation assistant component 119 and client device 106 is typically handled by… Figure 1 The 1102 indicates this. For example, the cloud-based automation assistant component 119 can be implemented by one or more servers that communicate with the client device 106.
[0026] A corresponding instance of the automation assistant client 118 (and optionally through its interaction with the cloud-based automation assistant component 119) can form a logical instance of the automation assistant 120 that appears from the user's perspective as if the user could participate in a human-computer dialogue with it. These three instances of the automation assistant 120... Figure 1The following are depicted and are generally referred to herein as "Automation Assistant 120". The first Automation Assistant 120A, enclosed by a dashed line, includes an Automation Assistant Client 1181 of Client Device 1061 and an optional cloud-based Automation Assistant Component 119. The second Automation Assistant 120B, enclosed by a dashed line, includes an Automation Assistant Client 1182 of Client Device 1062 and an optional cloud-based Automation Assistant Component 119. The third Automation Assistant 120C, enclosed by a dashed line, includes a Client Device 106... N Automated Assistant Client 118 N And an optional cloud-based automation assistant component 119. Therefore, it should be understood that each user interacting with the automation assistant client 118 running on client device 106 can actually interact with his or her own logical instance of automation assistant 120 (or a logical instance of automation assistant 120 shared within a home or other user group). For the sake of brevity and simplicity, the term "automation assistant" as used herein will refer to the automation assistant client 118 running on client device 106 and / or the cloud-based automation assistant component 119 (which can be shared among multiple automation assistant clients 118). Although in Figure 1 Only a number of associated client devices 106 are shown in the illustration, but it should be understood that the cloud-based automation assistant component 119 can also serve multiple additional groups of associated client devices.
[0027] In various implementations, the client device 106 may include a corresponding presence sensor 105. 1-N (Also referred to herein as "Presence Sensor 105"), which is configured to provide a signal indicating the presence—particularly human presence—after approval by the corresponding user. Presence Sensor 105 can take various forms. Some client devices 106 may be equipped with one or more digital cameras configured to capture and provide signals indicating movement detected in their field of view. Additionally or alternatively, some client devices 106 may be equipped with other types of light-based presence sensors 105, such as passive infrared ("PIR") sensors that measure infrared ("IR") light radiated from objects within their field of view. Additionally or alternatively, some client devices 106 may be equipped with presence sensors 105 that detect acoustic (or pressure) waves, such as one or more microphones.
[0028] Additionally or alternatively, in some embodiments, the presence sensor 105 may be configured to detect other phenomena associated with human presence. For example, in some embodiments, the client device 106 may be equipped with the presence sensor 105, which detects various types of waves (e.g., radio, ultrasonic, electromagnetic, etc.) emitted by, for example, a mobile device carried and / or operated by a particular user. For example, some client devices 106 may be configured to emit waves imperceptible to humans, such as ultrasonic or infrared waves, which can be detected by other client devices 106 (e.g., via an ultrasonic / infrared receiver, such as a microphone with ultrasonic capabilities).
[0029] Additionally or alternatively, one or more of the client devices 106 may emit other types of waves imperceptible to humans, such as radio waves (e.g., Wi-Fi, Bluetooth, cellular, etc.), which can be detected by other client devices 106 and used to determine the specific location of the operating user. In some embodiments, Wi-Fi triangulation can be used, for example, to detect a person's location based on Wi-Fi signals to / from client devices 106. In other embodiments, other wireless signal characteristics—such as time-of-flight, signal strength, etc.—can be used individually or collectively by the individual client devices 106 to determine the location of a specific person based on signals emitted by the client devices 106 carried by them.
[0030] Additionally or alternatively, in some implementations, one or more client devices 106 may perform voice recognition to identify an individual from their voice. For example, for the purpose of providing / restricting access to various resources, some automation assistants 120 may be configured to match voice with a user's profile. In some implementations, it may be simply assumed that the individual was in his or her last contact with the automation assistant 120, especially if not much time has passed since the last contact.
[0031] Each of the client devices 106 also includes a corresponding user interface component 107. 1-N (Also referred to herein as "User Interface Component 107"), each of which may include one or more user interface input devices (e.g., microphone, touchscreen, keyboard) and / or one or more user interface output devices (e.g., display, speaker, projector). User interface components may vary between client devices. As an example, user interface component 1071 may include only a speaker and microphone, while user interface component 1072 may include a speaker, touchscreen, and microphone. As another example, user interface component 1071 may include only a speaker and microphone, while user interface component 1072 may include a fingerprint sensor.
[0032] Furthermore, each of the client device 106 and / or the cloud-based automation assistant component 119 may include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components that facilitate communication via the network 110. Operations performed by one or more client computing devices 106 and / or by the automation assistant 120 can be distributed across multiple computer systems. The automation assistant 120 may be implemented as, for example, a computer program running on one or more computers in one or more locations, coupled to each other via a network.
[0033] As described above, in various embodiments, one or more of the client devices 106 can operate the automation assistant client 118. In various embodiments, each automation assistant client 118 may include a corresponding voice capture / text-to-speech (TTS) / speech-to-text (STT) module 114. In other embodiments, one or more aspects of the voice capture / TTS / STT module 114 may be implemented separately from the automation assistant client 118.
[0034] Each voice capture / TTS / STT module 114 can be configured to perform one or more functions, including: capturing a user's voice (voice capture, e.g., via a microphone, which in some cases may include one or more presence sensors 105)); converting the captured audio into text and / or other representations or embeddings (STT); and / or converting text to speech (TTS). In some implementations, because each client device 106 may be relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the voice capture / TTS / STT module 114 local to each client device 106 can be configured to convert a limited number of different spoken phrases into text (or into other forms, such as lower-dimensional embeddings). Other voice inputs can be sent to a cloud-based automation assistant component 119, which may include a cloud-based TTS module 116 and / or a cloud-based STT module 117.
[0035] The cloud-based STT module 117 can be configured to utilize virtually unlimited cloud resources to convert audio data captured by the voice capture / TTS / STT module 114 into text (which can then be provided to the natural language processor 122). The cloud-based TTS module 116 can be configured to utilize virtually unlimited cloud resources to convert text data (e.g., text specified by the automation assistant 120) into computer-generated speech output. In some embodiments, the TTS module 116 can provide the computer-generated speech output to a corresponding one of the client devices 106 for direct output, for example, using one or more speakers. In other embodiments, text data generated by the automation assistant 120 can be provided to the voice capture / TTS / STT module 114, which can then locally convert the text data into computer-generated speech rendered via a local speaker.
[0036] The automation assistant 120 (and particularly the cloud-based automation assistant component 119) may include a natural language processor 122, the aforementioned TTS module 116, the aforementioned STT module 117, and / or other components. In some embodiments, one or more engines and / or modules of the automation assistant 120 may be omitted, combined, and / or implemented in components separate from the automation assistant 120.
[0037] In some implementations, the automation assistant 120 generates responsive content in response to various inputs generated by a user on a client device 106 during a human-computer dialogue session with the automation assistant 120. The automation assistant 120 may provide responsive content (e.g., when disconnected from the user's client device, on one or more networks) for presentation to the user as part of the dialogue session. For example, the automation assistant 120 may generate responsive content in response to receiving free-form natural language input provided via a client device 106. As used herein, free-form input is input constrained by a set of options presented by the user and not used for the user's selection.
[0038] The natural language processor 122 of the automation assistant 120 processes natural language input generated by a user via client device 106 and can generate annotated output for use by one or more other components of the automation assistant 120. For example, the natural language processor 122 can process free-form natural language input generated by a user via one or more user interface input components 1071 of client device 1061. The generated annotated output includes one or more annotations to the natural language input, and optionally includes one or more (e.g., all) terms from the natural language input.
[0039] In some embodiments, the natural language processor 122 is configured to recognize and annotate various types of grammatical information in the natural language input. For example, the natural language processor 122 may include a speech annotator configured to annotate a portion of a term with grammatical roles. In some embodiments, the natural language processor 122 may additionally and / or alternatively include an entity annotator (not depicted) configured to annotate entity references in one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and fictional), etc. In some embodiments, data about entities may be stored in one or more databases, such as in a knowledge graph (not depicted). In some embodiments, the knowledge graph may include nodes representing known entities (and in some cases, entity attributes), and edges connecting nodes and representing relationships between entities.
[0040] The entity annotator of the natural language processor 122 can annotate references to entities at a high-granularity level (e.g., enabling the identification of all references to entity classes such as people), and / or annotate references to entities at a lower-granularity level (e.g., enabling the identification of all references to a specific entity such as a particular person). The entity annotator may rely on the content of the natural language input to parse specific entities and / or may optionally communicate with a knowledge graph or other entity database to parse specific entities.
[0041] In some implementations, the natural language processor 122 may additionally and / or alternatively include a coreference parser (not depicted) configured to group or “cluster” references to the same entity based on one or more contextual clues. For example, based on “theatre tickets” mentioned in a client device notification rendered immediately before receiving the natural language input “buy them”, the coreference parser can be used to parse the term “them” into “buy theatre tickets” in the natural language input “buy them”.
[0042] In some implementations, one or more components of the natural language processor 122 may rely on annotations from one or more other components of the natural language processor 122. For example, in some implementations, a named entity annotator may rely on annotations from a coreference parser and / or a dependency parser to annotate all references to a particular entity. Furthermore, for example, in some implementations, in clustered references to the same entity, the coreference parser may rely on annotations from the dependency parser. In some implementations, when processing a particular natural language input, one or more components of the natural language processor 122 may use relevant data beyond the specific natural language input to determine one or more annotations.
[0043] As shown below (for example, regarding) Figure 2 , Figure 3 , Figures 4A-4C and Figures 5A-5C As described in more detail, the automation assistant 120 can determine, based on processing user input, that user input received during a given human-computer dialogue session requires user interaction with one of the non-assistant platforms 140 (e.g., as described with respect to 114, 117, 112). The non-assistant platform 140 may include, for example, an email platform, a navigation platform, an Internet of Things (IoT) device platform, a web-based platform, a software application platform, and / or other platforms. If the automation assistant 120 determines that user input received during a given dialogue session requires user interaction with a given non-assistant platform among the non-assistant platforms 140, the automation assistant 120 may store the state of the dialogue in a dialogue state database 115 and may pause the dialogue session. Furthermore, the automation assistant 120 may be able to identify the given non-assistant platform among the non-assistant platforms 140 that will be used to complete the user interaction. The given non-assistant platform among the non-assistant platforms 140 may be one or more first-party platforms 141 or one or more third-party platforms 142. First-party platforms 141 include platforms that share a public publisher with the automation assistant 120, while third-party platforms 142 do not share a public publisher with the automation assistant 120. Furthermore, the automation assistant 120 can generate a request to be transmitted to a given non-assistant platform 140 to complete user interaction with the user. This request can be generated based on user input and / or other dialogue between the user and the automation assistant 120 provided during the conversation session. The automation assistant 120 can then transmit the request to the designated non-assistant platform 140.
[0044] The non-assistant platform 140 may generate one or more prompts to complete the user interaction in response to receiving a request for user interaction initiated via the non-assistant platform 140. The one or more prompts may be generated based on user input (or intent determined based on user input) provided by the user before pausing the conversation and / or a response provided by an assistant before pausing the conversation, and may request additional user input to complete the user interaction. Furthermore, the one or more prompts may be transmitted as, for example, electronic communications associated with the non-assistant platform (e.g., text messages, instant messages, emails, etc.), software application notifications from software applications associated with the non-assistant platform, as part of data sent to an application programming interface associated with the non-assistant platform, and / or other representations.
[0045] Additional user input may be received in response to rendering one or more prompts. For example, additional user input may be requested to complete a user interaction. User interactions may include, for example, confirming a user account linked to an account on client device 106, entering new credit card information, verifying current credit card information, confirming an address, checking into accommodation, prompting a service provider, purchasing game credits for gaming, and / or authenticating and / or verifying other user information or transactions that may require the user to specify one or more values. In various implementations, one or more prompts may additionally or alternatively require the user to provide biometric information in addition to or in lieu of additional user input. Biometric information may include, for example, fingerprint recognition via a fingerprint scanner associated with the user, voice recognition via one or more microphones associated with the user, facial recognition via one or more visual components associated with the user, and / or other biometric information. Additional user input received in response to one or more prompts rendered by non-assistant platform 140 may be transmitted to token engine 130.
[0046] The communication coupling between the cloud-based automated assistant component 119 and the token engine 130 and / or the non-assistant platform 140 is typically handled by... Figure 1 Instruction 1103. For example, token engine 130 and / or non-assistant platform 140 may be implemented by one or more servers communicating with client device 106 and / or cloud-based automated assistance component 119. Although in Figure 1 Described separately, however, token engine 130 can be implemented by one or more of the following: client device 106, cloud-based automated assistant component 119, non-assistant platform 140 (e.g., as per [reference]). Figure 3The token engine 130 may generate one or more tokens based on the results of user interactions with the non-assistant platform 140. The one or more tokens may be, for example, data objects comprising the results of user interactions performed via the non-assistant platform 140, one or more values associated with the user interactions, and / or other information provided by the non-assistant platform 140. In various embodiments, the results of user interactions on which one or more tokens are generated may not contain any data from the user interactions, and the one or more tokens may be encrypted to ensure user privacy. In some embodiments, the one or more tokens may be generated by the non-assistant platform 140 and transmitted back to the client device 106, and may be stored in association with the stored state of the conversation (e.g., in the conversation state database 115). The stored state of the conversation and the one or more tokens stored in association with the stored state of the conversation may be accessible by one or more client devices 106 via one or more networks 1101.
[0047] The stored state of a conversation session can be loaded at a given client device in client device 106, and the conversation session can be resumed based on one or more tokens (e.g., as described below). Figures 4A-4C and Figures 5A-5C (Detailed description). In some implementations, one or more tokens may be required to resume the conversation session, and the resumption of the conversation session can be influenced based on one or more tokens. For example, if a user is checking in for accommodation and the conversation is prompted to verify the user's user account, the user may also be prompted to purchase a specific TV package for the stay at the accommodation, including a choice between a normal TV package and a premium TV package. If the user purchases a normal TV package, the automated assistant can then present a first TV guide associated with the normal TV package when the conversation session resumes. However, if the user purchases a premium TV package, the automated assistant can then present a second TV guide associated with the premium TV package when the conversation session resumes. It is worth noting that there may be a time delay between the first time the conversation session is paused and the second time the conversation session is resumed, so the conversation session is resumed asynchronously. This time delay can range from a relatively short duration (e.g., a few seconds) to a relatively long duration (e.g., hours, days, weeks). It should be noted that the time delay between the pause and resumption of the conversation session can be based on the time it takes for the user to provide additional user input in response to one or more prompts.
[0048] In some implementations, a conversation session can be automatically resumed at a given client device in client device 106 in response to receiving a token from token engine 130 and / or non-assistant platform 140. For example, a conversation session can be paused at client device 1061 in response to receiving user input requiring user interaction with non-assistant platform 140, and remain paused until user interaction is completed via non-assistant platform 140 (and at another client device in client device 1061 or client device 106). When user interaction is completed, the paused conversation session can be resumed at client device 1061 (or another client device in client device 106). For example, receiving a token can cause the stored conversation state to be loaded into memory at client device 1061 (or another client device in client device 106), and the conversation session is resumed, and optionally, the conversation session is resumed to another state depending on the stored state and the received token. In some versions of those implementations, a conversation session can only be resumed at client device 1061 when the user is in the vicinity of client device 1061. In some other versions of those implementations, the client device that resumes the conversation session can be an additional client device (e.g., 1062 or 106...). N ), which is different from the client device 1061 that pauses the dialogue session and / or different from the additional client device where one or more prompts are rendered if rendered at the additional client device.
[0049] In some additional or alternative implementations, the conversation session may resume at client device 1061 (or another client device within client device 106) in response to receiving user input at client device 1061 (or another client device within client device 106) after receiving one or more tokens. For example, the conversation session may be paused at client device 1061 in response to receiving user input that requires user interaction with non-assistant platform 140, and remain paused until user interaction is completed at client device 1061 (or another client device within client device 106) and via non-assistant platform 140. When user interaction is completed, one or more client devices in an ecosystem with multiple client devices (e.g., client device 106) may provide indication that one or more tokens have been received (e.g., audible indications and / or graphical elements), and the conversation session may resume based on one or more tokens. However, the conversation session may not be resumed until user input to resume the conversation is received.
[0050] although Figure 1The description is intended to illustrate a specific configuration with components implemented by a client device 106 and / or a server communicating via a specific network; however, it should be understood that this is for illustrative purposes and not restrictive. For example, the client device 106, the cloud-based automation assistant component 119, the token engine, and / or the non-assistant platform 140 may communicate via a single network or any combination of networks. Furthermore, conversation states and conversation state databases 115 and / or user information stored in user information databases 135 may be accessible via these networks. In embodiments where data including conversation states and / or user information is transmitted via any of these networks, the data may be encrypted, filtered, or otherwise protected in any way to ensure the privacy of any user.
[0051] Now for reference Figure 2 State diagram 200 and Figure 3 State diagram 300 provides information on Figure 1 Additional descriptions of the various components. For illustrative purposes, it is assumed that the user associated with client device 106 (alone or as part of a group) is currently... Figure 2 252 places and Figure 3 352 of them participate in a conversational session with the automation assistant 120 via an automation assistant platform (e.g., an automation assistant application) operating on client device 1061. Although in Figure 3 The token engine 130 is depicted as being implemented separately from the client device 106; however, it should be understood that this is for illustrative purposes, and one or more client devices 106 may include the token engine 130. Furthermore, although some operations are described herein as being performed by certain devices, engines, and / or systems, it should be understood that the operations described herein may be performed by the automation assistant 120.
[0052] exist Figure 2 In state diagram 200, client device 1061 receives user input from a user of client device 1061 at 252 during a conversation session. User input may be, for example, spoken words, touch input, and / or typed input received at client device 1061 from a user who may or may not be associated with client device 1061 before the user input is provided.
[0053] At point 254, client device 1061 determines whether the user input received at point 252 requires user interaction with a non-assistant platform different from the automation assistant platform. The non-assistant platform may be associated with a first party (e.g., a first-party platform) that shares a public publisher with the automation assistant 120, or with a third party (e.g., a third-party platform) that does not share a public publisher with the automation assistant 120. The non-assistant platform may include, for example, a web-based platform (e.g., a web browser), an application-based platform (e.g., an IoT device application, a navigation application, an email application), and / or a platform different from the one currently used by the automation assistant 120. Figure 2 The automated assistant 120 may utilize any other platform of the automated assistant platform during the dialogue session. For example, assume the user input is the user's spoken words. In this example, the automated assistant 120 may cause the spoken words to be processed (e.g., using the corresponding speech capture / text-to-speech (TTS) / speech-to-text (STT) module 114 of the client device 1061 and / or the STT module 117 of the cloud-based automated assistant component) to generate discriminated text corresponding to the spoken words. Furthermore, the automated assistant may cause the natural language speech processor 122 to process the discriminated text to determine the user's intent (e.g., actions and corresponding slot values of parameters associated with the actions) based on the spoken words. If the client device 1061 determines that the user input (or the intent included therein) does not require interaction with the user on a non-assistant platform, the automated assistant 120 may continue the dialogue session with the user based on the user input received at 252. However, if the client device 1061 determines that the user input (or the intent included therein) requires interaction with the user on a non-assistant platform, it may store the state of the dialogue session in the dialogue state database 115 and pause the dialogue session at 256. It is worth noting that when the conversation is paused, the client device 1061 and / or the automation assistant 120 can still be utilized.
[0054] At point 258, client device 1061 generates a request. Automation assistant 120 may generate the request based on user input. The request may include, for example, instructions from a non-assistant platform and / or instructions for user interactions to be completed via a non-assistant platform (e.g., determined based on user intent).
[0055] At 260, client device 1061 will forward the request generated at 258 to token engine 130. In some implementations, token engine 130 may be implemented locally at client device 1061. In some additional or alternative implementations, token engine 130 may be implemented remotely at a remote computing device (e.g., one or more servers).
[0056] At 262, the non-assistant platform 140 receives the request generated at 258 on the client device 1061. At 264, the non-assistant platform 140 generates one or more prompts based on the request received at 264. The one or more prompts may be generated based on, for example, user input received at 252 (or intent determined based on user input) and / or a response rendered by the automation assistant 120 during the conversation session. Furthermore, the one or more prompts may be transmitted as, for example, electronic communications associated with the non-assistant platform (e.g., text messages, instant messages, emails, etc.), software application notifications from software applications associated with the non-assistant platform, as part of data sent to an application programming interface associated with the non-assistant platform, and / or other representations. It is noteworthy that in these embodiments, the request transmitted from the client device 1061 indirectly causes the rendering of one or more prompts at the client device 1062.
[0057] At 266, the non-assistant platform 140 will transmit one or more prompts generated at 264 to the user's client device 1062. Before transmitting one or more prompts, the automation assistant 120 may determine where the request should be transmitted. In some implementations, the automation assistant 120 determines where to transmit the request based on the device capabilities of client device 1061 and / or other client devices communicating with client device 1061 via one or more networks 1101. For example, if user interaction requires touch or typed input, but client device 1061 lacks a display, one or more prompts may be transmitted to different client devices that include a display (e.g., such as...). Figure 2 and 3 (See client device 1062 described herein). However, if client device 1061 includes a display required to complete user interaction, one or more prompts can be transmitted back to client device 1061. In some additional or alternative embodiments, automation assistant 120 determines where to transmit the request based on the user interaction required by the non-assistant platform.
[0058] At 268, client device 1062 receives one or more prompts, and at 270, renders one or more prompts on client device 1062. Client device 1062 may render the one or more prompts audibly and / or visually. How client device 1062 renders the one or more prompts may be based on the type of the one or more prompts generated. For example, if one or more prompts are included in electronic communication, the electronic communication may be delivered to client device 1062 as email, text message, instant message, etc. As another example, if one or more prompts are included in software application notification, the notification may be delivered to client device 1062 as a pop-up notification, banner notification, and / or any type of notification based on the settings of client device 1062.
[0059] At 272, client device 1062 receives additional user input from the user of client device 106. The additional user input may be received in response to rendering one or more prompts at 270. For example, additional user input may be requested to complete a user interaction. User interactions may include, for example, confirming a user account linked to an account on client device 1061, entering new credit card information, verifying current credit card information, confirming an address, checking into accommodations, prompting a service provider, purchasing game credits for gaming, and / or other user interactions that may require the user to specify one or more values for user information or transactions. In various implementations, one or more prompts may additionally or alternatively require the user to provide biometric information in addition to or in lieu of the additional user input. Biometric information may include, for example, fingerprint recognition via a fingerprint scanner associated with the user, voice recognition via one or more microphones associated with the user, facial recognition via one or more visual components associated with the user, and / or other biometric information.
[0060] At 274, client device 1062 transmits the additional user input received at 272 to token engine 130. At 276, token engine 130 receives the additional user input. At 278, token engine 130 generates one or more tokens based on the additional user input. At 280, token engine 130 transmits one or more of the tokens to client device 1061. One or more tokens can be generated based on additional user input. The one or more tokens can be, for example, data objects that include the result of a user interaction completed via a non-assistant platform, one or more values associated with the user interaction, and / or other information provided by the non-assistant platform. In other embodiments, an instance of token engine 130 can be implemented at client device 1062, and tokens can be generated locally at client device 1062. In those embodiments, tokens can be transmitted directly from client device 1062 to client device 1061.
[0061] At 282, client device 1061 receives one or more tokens transmitted at 280. At 284, client device 1061 stores one or more tokens (e.g., in dialogue state database 215) in association with the dialogue state stored at 256. Then, at 286A, automation assistant 120 can load the stored state of the dialogue session based on one or more tokens at client device 1061 (e.g., from dialogue state database 215) to resume the dialogue session. In some embodiments, the dialogue session can be resumed automatically in response to client device 1061 receiving one or more tokens, while in other embodiments, the dialogue can be resumed only in response to receiving additional user input at client device 1061 (e.g., as per [reference to...]). Figures 4A-4C (As described in more detail in 5A-5C). For example, if a conversation session is initially paused at client device 1061 because the user needs to purchase more game credits to continue playing the game, the conversation session can be resumed only in response to receiving a token instructing the user to purchase more game credits via a non-assistant platform, and the automation assistant 120 can automatically resume the game. In some additional or alternative implementations, the conversation state can be on additional client device 1061. N The session is loaded at the specified location. In some versions of these implementations, the client device loading the session can be the client device closest to the user, as indicated by a presence sensor. For example, it is assumed that the user is closer to client device 106 than client device 1061 when the token is received. N In this example, client device 106 N It can load the state of a conversation session and resume the conversation session based on one or more tokens (e.g., as per [the relevant token]). Figures 5A-5C The above).
[0062] exist Figure 2 In some implementations, the non-assistant platform 140 may not include a corresponding instance of the token engine 130. However, in some implementations, the non-assistant platform 140 may implement a corresponding instance of the token engine 130. For example, in Figure 3 In state diagram 300, 352-372 can be respectively related to the above regarding... Figure 2 The state diagrams 252-272 described in state diagram 200 are the same as or similar to those described above. Furthermore, 382, 384, 386A, and 386B can be respectively related to the states described above. Figure 2 The state diagrams 200 describe 282, 284, 286A, and 286B as identical or similar. Figure 2 Conversely, steps 376-380 can be executed by the non-assistant platform 140 using the corresponding instance of the token engine 130, and these steps can be executed in the same or similar manner as steps 276-280.
[0063] By using the information about Figure 2 and 3 The described method of pausing and asynchronously resuming a conversation allows the automated assistant platform to resume the conversation using the stored conversation context. This eliminates the need for the user to initiate a new conversation before pausing to achieve the same conversation state, thereby reducing user input received on the client device and saving computational resources. Furthermore, the automated assistant platform can also resume the conversation based on one or more tokens associated with user interactions completed outside of the automated assistant platform, allowing the stored conversation to move beyond its stored state.
[0064] Now go to Figure 4A , 4B Figures 4C and 4C depict an example dialogue session between user 401 and automation assistant 120. Automation assistant 120 can be accessed from one or more client devices 106. 1-2 Implement the token engine locally and / or remotely on one or more servers (e.g., Figure 1 The token engine 130), one or more servers communicate over a network (e.g., as per the token engine 130), and one or more servers communicate over a network (e.g., as per the token engine 130). Figure 1 Network 1102 (as described) and client device 106 1-2 One or more communications are used to pause and resume the conversation in response to determining that user input received during the conversation requires interaction with a non-assistant platform that is different from the automated assistant platform where the conversation took place.
[0065] Figure 4A and 4CThe client device 1061 depicted may include various user interface components, including, for example, a microphone that generates audio data based on spoken words and / or other audible input, a speaker that audibly renders synthesized speech and / or other audible output, and a display 1801 that receives touch input and / or visually renders transcription and / or other visual output. Figure 4A , 4B The client device 1062 depicted in 4C may include various user interface components, including, for example, a microphone that generates audio data based on spoken words and / or other audible input, a speaker that audibly renders synthesized speech and / or other audible output, and a display 1802 that receives touch input and / or visually renders transcription and / or other visual output.
[0066] Furthermore, the display 1802 of the client device 1062 includes various system interface elements 191, 192, and 193 (e.g., hardware and / or software interface elements) that can be interacted with by the user 401 to cause the client device 1062 to perform one or more actions. The display 1802 of the client device 1062 enables the user 401 to interact with the content rendered on the display 1802 via touch input (e.g., by directing user input to the display 1802 or a portion thereof) and / or via verbal input (e.g., by selecting the microphone interface element 194 at the client device 1062—or simply by speaking without having to select the microphone interface element 194 (i.e., the automation assistant 120 can monitor one or more terms or phrases, gestures, gazes, mouth movements, lip movements, and / or other conditions to activate verbal input)
[0067] For details, please refer to the following: Figure 4A Assuming user 401 has arrived at their accommodation, such as a hotel, vacation rental hotel, timeshare, and / or any other form of temporary accommodation. In some implementations, the automated assistant 120 can (e.g., via...) Figure 1 The presence sensor 105 initiates a conversation with the user when it detects the user's presence, while in other embodiments, the conversation may be initiated in response to receiving user input at the client device 1061. For example... Figure 4AAs depicted, further assuming that upon arrival, client device 1061 receives a spoken utterance 452A from user 401, “Assistant, I'm here to check into my lodging,” to initiate a conversational session. Automation assistant 120 may enable client device 1061 to audibly render a synthesized speech 454A, “Okay, would you like to link your account?” and user 401 may respond with an additional spoken utterance 456A, “Yes.” Automation assistant 120 may determine, based on the additional spoken utterance 456A instructing user 401 to want to link user 401's user account to client device 1061, that the additional spoken utterance 456A requires the user to interact with someone other than the currently active user. Figure 4A The user interaction during the conversation utilizes a non-assistant platform other than the automated assistant platform, thus associating user 401 with client device 1061. The non-assistant platform may include, for example, a web-based platform (e.g., a web browser), an application-based platform (e.g., an IoT device application, a navigation application, an email application), and / or [other platforms]. Figure 4A The non-assistant platform is any other platform that is currently being used during the dialogue session, different from the automated assistant platform. In some implementations, the non-assistant platform may be a first-party platform that shares a public publisher with the automated assistant 120, while in other implementations, the non-assistant platform may be a third-party platform that does not share a public publisher with the automated assistant 120.
[0068] In response to determining that spoken utterance 456A requires user interaction with a non-assistant platform, the automation assistant 120 may, as instructed by 458A1, store the state of the conversation session, as instructed by 458A2, pause the conversation session, and / or generate a request as instructed by 458A3, and transmit the request to the non-assistant platform and / or token engine 130. For example, the automation assistant 120 may do so in one or more databases (e.g., Figure 1 The dialogue state database 115 stores Figure 4AThe state of the conversation session. The state of the conversation session may include, for example, spoken words 452A and 456A, synthesized speech 454A, contextual information associated with the conversation session (e.g., the user's location (and optionally at home or in a residence), the time when the conversation session was initiated and / or paused, the duration of the conversation session), and / or other data associated with the conversation session. Furthermore, the automation assistant 120 may pause the conversation session until it is resumed at the client device 1061 by disabling one or more components of the client device 1061 (e.g., disabling speech detection, natural language processing, and / or other components), displaying one or more visual instructions related to user interactions to be performed by a non-assistant platform (e.g., "Please link your account using your mobile device") and / or providing other instructions related to the conversation session. Notably, when the conversation session is paused, the client device 1061 can still be used to perform other assistant-related actions, including, for example, initiating other conversation sessions with the user 401 and / or performing other actions on behalf of the user 401 (e.g., playing music, providing weather or search results, etc.). Furthermore, the automation assistant 120 can generate requests to be sent to non-assistant platforms. For example, the automation assistant 120 can identify the non-assistant platform required to perform user interactions (e.g., with...). Figure 4A The platform associated with the account link in the system generates a request that includes user information stored in one or more databases (e.g., user ID associated with user 401, client ID associated with client device 1062 detected through one or more networks 110, and / or other user information stored in user information database 135), and transmits the request to the non-assistant platform and / or token engine.
[0069] In some implementations, the token engine 130 and / or the non-assistant platform can process the received request to generate one or more prompts to complete the user interaction. In some implementations where the non-assistant platform required to complete the user interaction is a third-party platform, the request can be sent directly to the third-party platform, allowing it to generate one or more prompts. In other implementations where the non-assistant platform required to complete the user interaction is a third-party platform, the request can be sent to the third-party platform, allowing it to generate one or more prompts. In implementations where the non-assistant platform required to complete the user interaction is a first-party platform, the request can be sent to the first-party platform.
[0070] One or more prompts may be generated based on, for example, synthesized speech 454A and / or additional spoken words 456A instructing user 401 to link their user account to client device 1061. Furthermore, one or more prompts may be transmitted as, for example, electronic communications associated with a non-assistant platform (e.g., text messages, instant messages, emails, etc.), software application notifications from software applications associated with a non-assistant platform, as part of data sent to an application programming interface associated with a non-assistant platform, and / or other representations. It is noteworthy that in these embodiments, a request transmitted from client device 1061 indirectly causes one or more prompts to be rendered at client device 1062.
[0071] The non-assistant platform allows client device 1061 and / or client device 1062 to present one or more prompts to complete user interaction. For example, now refer to Figure 4B Unlike Figure 4A The account linking platform of the automated assistant platform used in the dialogue session can generate and / or receive one or more prompts, as indicated by 452B1. Figure 4B As shown, the account linking platform can render a given prompt 452B2, “Are you sure you want you link your account with the lodging clientdevice?”, at the display 1802 of the client device 1062 (visually and optionally audibly). Assume a user response 454B of “Yes” is received in response to the non-assistant platform rendering prompt 452B2. Response 454B can be based on, for example, touch input to a first selectable element 471 pointing to “Yes” (among multiple selectable elements including at least a second selectable element 472 for preventing account linking), a verbal “Yes” received in response to selection of microphone interface element 194 (or another statement confirming account linking), and / or typed input of “Yes” received via display 1802 (or other typed input confirming account linking). In various embodiments, one or more prompts may additionally or alternatively request biometric information from user 401. Biometric information may include, for example, fingerprint recognition of a fingerprint associated with user 401 detected by a fingerprint scanner of client device 1061, voice recognition of a voice associated with user 401 detected by one or more microphones of client device 1061, facial recognition of a face associated with user 401 detected by one or more visual components of client device 1061, and / or other biometric information.
[0072] In some implementations, a non-assistant platform, distinct from the automation assistant 120, can generate a token associated with the user interaction and transmit that token to the token engine 130, client device 1061, and / or an additional client device not depicted. In some additional or alternative implementations, the non-assistant platform is used to transmit the result of the user interaction to the token engine 130 to generate a token associated with the user interaction, and the token engine can transmit the token to client device 1061 and / or an additional client device not depicted. In these implementations, the token can be, for example, a data object that includes the result of a user interaction performed via the non-assistant platform, one or more values associated with the user interaction, and / or other information provided by the non-assistant platform. For example, in Figure 4B In the example, the token can instruct the user to authorize permission to link user 401's user account with client device 1061. It is worth noting that in various implementations, the conversation session may not be resumed at client device 1061 (or any other client device 106) until the token is received. Therefore, a token may be required to advance the conversation session through its stored state.
[0073] For details, please refer to the following: Figure 4C Client device 1061 can receive tokens from token engine 130 and / or non-assistant platform, as indicated by 452A1 (and optionally store the tokens in association with the state of the conversation session, e.g., in...). Figure 1 (in the dialogue state database 115). Furthermore, the client device 1061 can load stored dialogue sessions (e.g., from...). Figure 4A As indicated by 452C2, the conversation session is resumed based on the loaded conversation session and the received token, as indicated by 452C3. In some implementations, the conversation session may be automatically resumed at the client device 1061 in response to the receipt of a token. For example, Figure 4AThe dialogue session can be paused at client device 1061 in response to receiving a spoken utterance 456A requiring user interaction with a non-assistant platform, and remain paused until user interaction is completed at client device 1062. When user interaction is completed at client device 1062, the paused dialogue session can be resumed at client device 1061. For example, client device 1062 can render an additional synthesized voice 454C saying "Your account is successfully linked, you can now control the smart appliances for your lodging," and the user can provide the spoken utterance 456C saying "(Thank you, please open the smart blinds)," and the automation assistant can render an additional synthesized voice 458C saying "Okay, opening the smart blinds," and actuate the smart blinds to the open position. Notably, the synthesized voice can also include a list of smart appliances that user 401 can control. For example, user 401 may be able to control a smart alarm, a smart coffee machine, and a smart TV associated with the accommodation, but not a smart thermostat associated with the accommodation. In some versions of those implementations, the conversation session can only be resumed when user 401 is present in the vicinity of client device 1062 (e.g., determined based on the presence sensor 105 of client device 1061).
[0074] In some additional or alternative implementations, the conversation session can be resumed at the client device 1061 in response to receiving user input to resume the conversation session after receiving a token. For example, Figure 4A The conversation session can be paused at client device 1061 in response to receiving a verbal utterance 456A indicating that interaction with a user on a non-assistant platform is required, and remain paused until user interaction is completed at client device 1062. When user interaction is completed at client device 1062, client device 1061 can provide an indication that a token has been received (e.g., an audible indication and / or a graphical element). However, the conversation session may not be resumed until user input to resume the conversation is received. For example, in an embodiment where client device 1061 includes a display 1801, optional graphical elements may be displayed, and when selected by user 401, the conversation session may be resumed at client device 1061. As another example, client device 1061 may receive a verbal utterance to resume the conversation session (e.g., "resume," "continue," etc.). In these examples, (e.g., using...) Figure 1 The speech processing of the natural language processor 122 can be biased towards resuming the spoken utterance of the conversation in the event that a token is received at the client device 1061.
[0075] It should be noted that Figure 4C The resumption of the conversation can vary based on a token received at client device 1061. For example, user 401 may be prompted to enter membership information associated with accommodation, and the user may be able to control different smart devices based on the membership status of a user with accommodation. For example, if user 401 is a user with... Figure 4A , 4B For Platinum members of the accommodation service described in 4C, user 401 can control every smart device included in the accommodation (e.g., smart TV, smart curtains, smart coffee machine, smart thermostat, etc.). However, if user 401 only has information about... Figure 4A , 4B For a Bronze member of the accommodation service described in 4C, user 401 may only be able to control certain smart devices included in the accommodation (e.g., smart TVs, smart curtains, and smart coffee machines, but not smart thermostats). In other words, resuming a conversation based on a token may not be just about authentication and / or verification. Instead, upon resumption, the automation assistant 120 may manipulate the conversation in different directions based on the received token.
[0076] although Figure 4A , 4B The examples of 4C described in this document relate to user 401 linking their account to client device 1061; however, it should be understood that this is for illustrative purposes and not restrictive. For example, the techniques described herein can be used to enter or verify payment information associated with accommodation at check-in, to verify purchases charged to the accommodation at check-in, to add specific values associated with the operator requesting accommodation for the rendered service. In these examples, the non-assistant platform could be a software application associated with user 401's bank, a software application associated with the accommodation provider, the user's email account that receives a prompt from the user to verify payment information, etc.
[0077] Furthermore, despite Figure 4A and 4C The client device 1061 described herein is a standalone client device with a display and Figure 4A , 4B The client device 1062 depicted in 4C is the mobile device of user 401, but it should be understood that this is for illustrative purposes and not restrictive. For example, client device 106... 1-2Other client devices may include, for example, stand-alone speakers without displays, home automation or Internet of Things (IoT) devices, vehicle client systems, laptops, computers, and / or any other devices capable of participating in human-machine dialogue sessions with the user. Furthermore, although... Figure 4A , 4B The diagrams in 4C illustrate specific interactions as either voice-based or touch-based, but it should be understood that this is for illustrative purposes and not restrictive. For example, Figure 4A and 4C The voice-based interaction of the depicted conversation could be between user 401 and automation assistant 120, and via typing on client device 1061. As another example, user 401 and... Figure 4B The touch-based interaction between non-assistant platforms shown can be voice-based. As yet another example, Figure 4A , 4B Any interaction described in 4C can be a typing-based interaction via a virtual keyboard and / or a hardware keyboard.
[0078] Furthermore, although client device 1062 is used to perform the interaction with... Figure 4B The user interaction shown is for a non-assistant platform, but it should be understood that this is for illustrative purposes and not limiting. In various embodiments, both the conversation and the user interaction with the non-assistant platform can be performed via client device 1061, or both can be performed via client device 1062. In various embodiments, the automation assistant 120 can determine which client devices 106 to utilize when completing user interactions requiring a non-assistant platform, based on the capabilities of the client devices 106. For example, assume that client device 1061 is a standalone speaker device capable of implementing an automation assistant platform, but lacks a display, and further assume that client device 1062 is the mobile device of user 401. In this example, the automation assistant 120 can enable the user interaction to be completed at client device 1062 to utilize the capabilities of touchscreen display 1802. Conversely, assume that client device 1061 includes touchscreen display 1801. In some of these examples, the automation assistant 120 can enable the user interaction to be completed at client device 1061 because client device 1061 includes touchscreen display 1802.
[0079] Now go to Figure 5A , 5B And 5C, depicting another example of a conversational session between a user and an automated assistant. Figure 5A and 5C The document depicts a family floor plan. The floor plan includes multiple rooms (550-562). It also includes multiple client devices (106). 1-4Deployed in at least some rooms (e.g., from) Figure 1 Client device 106 1-4 Each of the client devices 506 can implement an automation assistant client configured with selected aspects of this disclosure (e.g., ...). Figure 1 The corresponding instance of the automation assistant client 118, and may include one or more user interface components (e.g., Figure 1 User interface components 107, such as a microphone capable of capturing speech spoken by nearby people, a speaker capable of audibly rendering synthesized speech, a display capable of visually rendering visual content, and / or other user interface components. For example, a first client device 1061, in the form of a standalone interactive speaker lacking a display, is deployed in room 550, which in this example is a kitchen. A second client device 1062, in the form of a mobile device, is held by user 501 (e.g., the same or similar client device 1062 is in...). Figure 4A , 4B And described in 4C and about Figure 5B (Description), User 501 is located in Figure 5A kitchen and Figure 5C In the living room, and about Figure 5B Description. A third client device 1063, in the form of an interactive independent speaker, is deployed in room 554, which in this example is a bedroom. A fourth client device 1064, in the form of another interactive independent speaker, is deployed in room 556, which in this example is a living room.
[0080] Although Figure 5A , 5B Not depicted in 5C, but multiple client devices 106 can communicatively couple with each other and / or via one or more wired or wireless LANs (e.g., as per [reference]). Figure 1 As described in 1101, it is coupled to communicate with other resources (e.g., the Internet). Additionally, other client devices may also exist, particularly other smartphones, tablets, laptops, wearable devices, etc., carried by one or more people in the household, and may or may not be connected to the same LAN. It should be understood that... Figure 5A , 5B The configuration of client device 106 described in 5C is merely an example; more or fewer and / or different client devices 106 can be deployed in any number of other rooms and / or areas outside the home.
[0081] exist Figure 5A and 5C The home floor plan also depicts multiple IoT devices. 1-4For example, a first IoT device 5101, in the form of a smart doorbell, is deployed outside the home near the front door. A second IoT device 5102, in the form of a smart third-party alarm, is coupled to the front door. A third IoT device 5103, in the form of a smart washing machine, is deployed in room 562, which in this example is the laundry room. A fourth IoT device 5104, in the form of an attached smart third-party alarm, is coupled to the back door. A fifth IoT device 5105, in the form of a smart thermostat, is deployed in room 552, which in this example is the den.
[0082] Each IoT device 510 can (e.g., via the Internet) communicate with a corresponding IoT-based platform to provide data to the IoT-based platform and optionally control it based on commands provided by the IoT-based platform in response to user input detected at one or more client devices 106. It should be understood that... Figure 5A , 5B The configuration of IoT device 510 described in 5C is merely an example; more or fewer and / or different IoT devices 510 can be deployed in any number of other rooms and / or areas outside the home.
[0083] For details, please refer to the following: Figure 5A Assume user 501 participates in a conversation with client device 1061 via an automated assistant platform to establish a morning routine, as instructed by 582A. This routine causes multiple assistant-based actions to be performed when invoked (typically in the morning) using one or more trigger terms or phrases (e.g., “Assistant, good morning,” “Tell me about my morning,” etc.). Further assume user 501 provides a verbal utterance 584A in the morning routine that “includes disabling third-party alarms” (e.g., third-party alarm 5102 and additional third-party alarm 5104). In response to determining that verbal utterance 584A requires user interaction with a non-assistant platform (e.g., a third-party platform associated with third-party alarm 5102 and additional third-party alarm 5104), the automated assistant may store the state of the conversation, pause the conversation, and / or generate a request and transmit the request to the non-assistant platform and / or token engine 130 (e.g., as described above regarding…). Figure 4A The non-assistant platform and / or token engine 130 can generate one or more prompts to complete the user interaction, and cause one or more prompts to be rendered at the client device 1062 (e.g., as described above regarding...). Figure 4A(as described above). For example, in response to determining that one or more prompts require a display but client device 1061 lacks a display, one or more prompts may be rendered at client device 1062.
[0084] For example, see specific references Figure 5B With third-party alarm 5102 and / or different Figure 5A In the dialogue session, the associated third-party alarm 5104 of the automated assistant platform can generate and / or receive one or more prompts, as indicated by 552B1. Figure 5B As shown, a third-party platform associated with third-party alarm 5102 and / or additional third-party alarm 5104 can cause a given prompt 552B2 to be rendered (visually and optionally audibly) at the display 1802 of client device 1062, stating "Confirm that you would like to allow the assistant to disable the third-party alarm." It is assumed that a user response 554B of "Confirm" is received in response to a non-assistant platform causing prompt 552B2 to be rendered. Response 554B may be based on touch input, for example, a first selectable element 571 pointing to “Confirm” (among a plurality of selectable elements including at least a second selectable element 572 for preventing account linking), a verbal “Confirm” received in response to selection of microphone interface element 194 (or another statement from the assistant confirming that the assistant can control a third-party alarm), and / or typed input of “Confirm” received via display 1802 (or other typed input from the assistant confirming that the assistant can control a third-party alarm). In various embodiments, one or more prompts may additionally or alternatively request biometric information from user 401. Biometric information may include, for example, fingerprint recognition of a fingerprint associated with user 401 detected by a fingerprint scanner of client device 1061, voice recognition of a voice associated with user 401 detected by one or more microphones of client device 1061, facial recognition of a face associated with user 401 detected by one or more visual components of client device 1061, and / or other biometric information.
[0085] In some implementations, a non-assistant platform, distinct from the automation assistant 120, can generate a token associated with the user interaction and transmit that token to the token engine 130, client device 1061, and / or an undescribed additional client device. In some additional or alternative implementations, the non-assistant platform is used to transmit the result of the user interaction to the token engine 130 to generate a token associated with the user interaction, and the token engine can transmit the token to client device 1061 and / or an undescribed additional client device. In these implementations, the token can be, for example, a data object comprising, the result of a user interaction performed via the non-assistant platform. For example, in Figure 5B In the example, the token could instruct the user to confirm that the automation assistant can control a third-party alarm. It is worth noting that in various implementations, the conversation session may not be resumed at client device 1061 (or any other client device 106) until the token is received. Therefore, a token may be needed to advance the conversation session through its stored state.
[0086] For details, please refer to the following: Figure 5C The client device 1064 can receive tokens from the token engine 130 and / or the non-assistant platform (and optionally store the tokens in association with the state of the conversation session). Figure 1 (in the dialogue state database 115). Furthermore, the client device 1064 can load stored dialogue sessions (e.g., from...). Figure 5A And through Figure 1 One or more networks 1101), and restore the conversation session based on the loaded conversation session and the received token. As a result, user 501 and the automation assistant can resume setting the morning routine, as instructed by 582C. It is worth noting that restoring the conversation session... Figure 4C The client device 1064 described herein is associated with the suspension of the conversation session. Figure 4A The depicted client device 1061 differs. In various embodiments, the conversation session can be resumed at the given client device (e.g., determined based on presence sensor data generated by presence sensor 105) where user 501 is closest to the user in client device 106 upon completion of user interaction. For example, suppose user 501 starts from... Figure 5A Room 550 shown – which in this example is the kitchen – leads to... Figure 5C The room shown is 556—in this example, it is the living room, while Figure 5B User interaction is completed via client device 1062. This allows users to set up their morning routine while... Figure 5A and 5C The houses depicted roam freely.
[0087] In some implementations, the conversation session can be automatically resumed at the client device 1064 in response to receiving a token (e.g., as mentioned above regarding...). Figure 4A (as described). For example, Figure 5A The conversation session can be paused at client device 1061 in response to receiving a verbal utterance 584A indicating that user interaction with a non-assistant platform is required, and remains paused until user interaction is completed at client device 1062. When user interaction is completed at client device 1062, the paused conversation session can be resumed at client device 1064 based on the presence of a user at client device 1064. In some additional or alternative implementations, the conversation session can be resumed at client device 1064 in response to receiving user input (at client device 1064) to resume the conversation session after receiving a token (e.g., as described above regarding...). Figure 4A (as described). For example, Figure 5A The conversation session may be paused at client device 1061 in response to receiving a verbal utterance 584A indicating that interaction with a user on a non-assistant platform is required, and remain paused until the user interaction is completed at client device 1062. When the user interaction is completed at client device 1062, one or more client devices 106 in the room may provide an indication that a token has been received (e.g., an audible indication and / or a graphical element). However, the conversation session may not be resumed until user input to resume the conversation is received.
[0088] although Figure 5A , 5B The specific client device 106 is described in 5C, but it should be understood that this is for illustrative purposes and not limiting. For example, client device 106 can be other client devices, including, for example, a standalone speaker without a display, a home automation or Internet of Things (IoT) device, a vehicle client system, a laptop computer, a computer, and / or any other device capable of participating in a human-machine dialogue session with user 501. Furthermore, although... Figure 5A , 5B The illustrations in 5C depict specific interactions as either voice-based or touch-based, but it should be understood that these are for illustrative purposes and not restrictive.
[0089] Furthermore, although client device 1061 is used to perform and Figure 5BThe user interaction with the non-assistant platform described herein is for illustrative purposes only and not for limitation. In various embodiments, both the conversation and the user interaction with the non-assistant platform can be performed via client device 1061, or both via client device 1062. In various embodiments, the automation assistant 120 can determine which client devices 106 to utilize when performing user interactions requiring a non-assistant platform based on the capabilities of the client devices 106. For example, suppose client device 1061 is a standalone speaker device capable of implementing an automation assistant platform, but lacks a display, such as... Figure 5A As shown, and further assumed that client device 1062 is a mobile device of user 501. In this example, automation assistant 120 may enable user interaction to be completed at client device 1062 to utilize the capabilities of touchscreen display 1802. Conversely, it is assumed that client device 1061 includes touchscreen display 1801. In some of these examples, automation assistant 120 may enable user interaction to be completed at client device 1061 because client device 1061 includes touchscreen display 1802.
[0090] Turn now Figure 6 This document depicts a block diagram of an example computing device 610 that may be optionally utilized to perform one or more aspects of the techniques described herein. In some embodiments, one or more of the client device 106, token engine 130, third-party platform 140, and / or other components may include one or more components of the example computing device 610.
[0091] Computing device 610 typically includes at least one processor 614, which communicates with a number of peripheral devices via a bus subsystem 612. These peripheral devices may include a storage subsystem 624, which includes, for example, a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow users to interact with computing device 610. The network interface subsystem 616 provides an interface to external networks and is coupled to corresponding interface devices in other computing devices.
[0092] User interface input device 622 may include: a keyboard, a pointing device such as a mouse, trackball, touchpad, or graphics tablet, a scanner, a touchscreen integrated into a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. Generally, the term "input device" is used to encompass all possible types of devices and methods for inputting information into computing device 610 or a communication network.
[0093] User interface output device 620 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and methods for outputting information from computing device 610 to a user or other machine or computing device.
[0094] Storage subsystem 624 stores the functional programming and data construction of some or all of the modules described herein. For example, storage subsystem 624 may include logic for selecting aspects of the methods described herein, as well as implementations... Figure 1 The logic of the various components described.
[0095] These software modules are typically executed by processor 614 alone or in combination with other processors. Memory 625 used in storage subsystem 624 may include a plurality of memories, including main random access memory (RAM) 630 for storing instructions and data during program execution and read-only memory (ROM) 632 for storing fixed instructions. File storage subsystem 626 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives with associated removable media, CD-ROM drives, optical disk drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored in file storage subsystem 626 within storage subsystem 624, or may be stored in other machines accessible by processor 614.
[0096] Bus subsystem 612 provides a mechanism for enabling various components and subsystems of computing device 610 to communicate with each other as intended. Although bus subsystem 612 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0097] The computing device 610 can be of various types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the constantly changing nature of computers and networks, Figure 6 The description of the computing device 610 depicted herein is intended only as a specific example for illustrating some implementations. Many other configurations of the computing device 610 may have... Figure 6 The computing device depicted in the text has more or fewer components.
[0098] In some implementations discussed herein, where personal information about users can be collected or used (e.g., user data extracted from other electronic communications, information about user social networks, user location, user time, user biometric information, user activity and demographic information, relationships between users, etc.), users are provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about users is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use user personal information only upon receiving explicit authorization from the relevant user.
[0099] For example, users can be given control over whether a program or feature collects user information about that specific user or other users associated with that program or feature. One or more options will be presented to each user whose personal information is collected, allowing control over the collection of information related to that user, providing permission or authorization regarding whether information is collected and which parts of the information are collected. For example, one or more such control options can be provided to users on a communication network. Furthermore, data can be processed in one or more ways before it is stored or used to remove personally identifiable information. As an example, a user's identity can be processed so that no personally identifiable information can be determined. As another example, a user's geographic location can be generalized to a larger area, making it impossible to determine the user's specific location.
[0100] In some implementations, a method implemented by one or more processors is provided and includes: receiving user input from a user on a client device during a conversation between the user and an automation assistant and via an automation assistant platform; and, in response to determining that the user input requires user interaction with a user on a non-assistant platform other than the automation assistant platform: storing the state of the conversation between the user and the automation assistant in one or more databases accessible to at least the client device; transmitting a request to the non-assistant platform to initiate user interaction; and receiving a token associated with the user interaction from the non-assistant platform and in response to the user completing the user interaction via an attached client device. Transmitting the request to the non-assistant platform causes the user's attached client device to render a prompt for completing the user interaction with the non-assistant platform. In response to receiving the token associated with the user interaction: resuming the conversation between the user and the automation assistant at the client device or another attached client device based on the state of the conversation and based on the token associated with the user interaction.
[0101] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0102] In some embodiments, the method may further include, in response to receiving a token associated with a user interaction and before resuming the conversation session between the user and the automation assistant: storing the token in association with the state of the conversation session between the user and the automation assistant in one or more databases accessible to the client device. In some versions of those embodiments, resuming the conversation session between the user and the automation assistant at the client device may include an indication to load the state of the conversation session between the user and the automation assistant, along with the token, at the client device or an additional client device and via the automation assistant platform. In some other versions of those embodiments, the indication to load the state of the conversation session, along with the token, at the client device or an additional client device is in response to receiving a token associated with a user interaction performed via an additional client device. In some additional or alternative versions of those further embodiments, the indication to load the state of the conversation session, along with the token, at the client device or an additional client device is in response to receiving additional user input from the user of the client device to resume the conversation session.
[0103] In some implementations, the method may further include, after storing the state of the conversation between the user and the automation assistant, pausing the conversation between the user and the automation assistant.
[0104] In some implementations, in response to determining that user input requires interaction with a non-assistant platform different from the automated assistant platform, the method may further include determining whether the client device is capable of facilitating user interaction with the non-assistant platform. Sending a request to an additional client device to initiate user interaction may be in response to determining that the client device is not capable of facilitating user interaction with the non-assistant platform.
[0105] In some implementations, the automation assistant can be accessed at each of the client device, the additional client device, and another additional client device. The additional client device can be in addition to the client device and the other additional client device. In some versions of those implementations, the client device and the other additional client device can be corresponding independent automation assistant devices associated with a user, and the additional client device can be a mobile device associated with the user.
[0106] In some implementations, the non-assistant platform can be an authentication or authorization platform that has a public publisher with the automation assistant platform. In some versions of those implementations, the non-assistant platform can be a third-party platform that does not have a public publisher with the automation assistant platform.
[0107] In some implementations, the non-assistant platform may enable the user's attached client device to respond to a received request and render prompts via the attached client device to complete user interactions with the non-assistant platform.
[0108] In some implementations, a method implemented by one or more processors is provided and includes: receiving from a user's client device a request to initiate user interaction with a non-assistant platform, the non-assistant platform being different from the automation assistant platform being utilized during a conversation between the user and an automation assistant accessible at the client device, and the request being received in response to determining that user input received during the conversation requires user interaction with a non-assistant platform different from the automation assistant platform. The method further includes, in response to receiving the request from the client device: causing a prompt to be rendered at the user's additional client device to complete the user interaction with the non-assistant platform, and generating a token associated with the user interaction completed via the additional client device. The token is generated based on user input received from the user at the additional client device in response to rendering the prompt. The method may further include, in response to generating the token associated with the user interaction: transmitting the token associated with the user interaction to the user's client device. Transmitting the token associated with the user interaction causes the user's client device to store the token in association with a stored state of the conversation between the user and the automation assistant.
[0109] These and other implementations of the technology disclosed herein may include one or more of the following features.
[0110] In some implementations, transmitting the token associated with the user interaction can also cause the user's client device to load an indication of the state of the conversation session between the user and the automation assistant at the client device, along with the token, via the automation assistant platform. In some versions of those implementations, loading the conversation session state along with the token at the client device can be in response to causing the client device to store the token associated with the completed user interaction. In some additional or alternative versions of those implementations, loading the conversation session state along with the token at the client device can be in response to receiving additional user input from the user at the client device to resume the conversation session.
[0111] In some implementations, the conversation session may be paused at the client device in response to determining that user input received during the conversation session requires interaction with a user on a non-assistant platform different from the automated assistant platform.
[0112] In some implementations, the automation assistant can be accessed at each of the client device, the additional client device, and another additional client device. The additional client device may be in addition to the client device and the other additional client device. In some versions of those implementations, one or more processors may belong to the additional client device.
[0113] Furthermore, some implementations include one or more processors (e.g., a central processing unit (CPU), a graphics processing unit (GPU), and / or a tensor processing unit (TPU)) of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein the instructions are configured to cause performance of any of the foregoing methods. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by the one or more processors to perform any of the foregoing methods. Some implementations also include a computer program product comprising instructions executable by the one or more processors to perform any of the foregoing methods.
[0114] It should be understood that all combinations of the foregoing and additional concepts described in more detail herein are considered part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are considered part of the subject matter disclosed herein.
Claims
1. A method implemented by one or more processors, the method comprising: The user on the client device receives user input during the dialogue session between the user and the automation assistant, and via the automation assistant platform; In response to determining that the user input requires interaction with a user on a non-assistant platform different from the automated assistant platform: The state of the conversation between the user and the automated assistant is stored in one or more databases accessible to at least the client device. After storing the state of the conversation between the user and the automated assistant, the conversation between the user and the automated assistant is paused. A request to initiate the user interaction is sent to the non-assistant platform, wherein sending the request to the non-assistant platform causes the user's additional client device to render a prompt for completing the user interaction with the non-assistant platform, and Receive a token associated with the user interaction from the non-assistant platform and in response to the user completing the user interaction via the additional client device; In response to receiving the token associated with the user interaction: Based on the state of the conversation and the token associated with the user interaction, the conversation between the user and the automation assistant is resumed on the client device or an additional client device.
2. The method according to claim 1, further comprising: In response to receiving a token associated with the user interaction and before the conversation session between the user and the automated assistant is resumed: The token is stored in association with the state of the conversation between the user and the automation assistant in one or more of the databases accessible to at least the client device.
3. The method according to claim 1, wherein, Enabling the resumption of the conversation between the user and the automated assistant at the client device includes: The state of the conversation between the user and the automation assistant, along with the indication of the token, is loaded at the client device or the additional client device and via the automation assistant platform.
4. The method according to claim 3, wherein, The indication that the state of the conversation session, along with the token, is loaded at the client device or the additional client device is in response to receiving the token associated with the user interaction completed via the additional client device.
5. The method according to claim 3, wherein, The state of the conversation session loaded at the client device or the additional client device, along with the indication of the token, is in response to receiving additional user input from the user of the client device to resume the conversation session.
6. The method of claim 1, further comprising, in response to determining that the user input requires interaction with a user on a non-assistant platform different from the automated assistant platform: Determine whether the client device can facilitate user interaction with the non-assistant platform; as well as Specifically, sending a request to the additional client device to initiate the user interaction is in response to determining that the client device is unable to facilitate the user interaction with the non-assistant platform.
7. The method according to claim 1, wherein, The automation assistant can be accessed at each of the client device, the additional client device, and the additional client device, wherein the additional client device is an addition to the client device and the additional client device, and wherein the additional client device is an addition to the client device and the additional client device.
8. The method according to claim 7, wherein, The client device and the additional client device are corresponding independent automation assistant devices associated with the user, and the additional client device is a mobile device associated with the user.
9. The method according to claim 1, wherein, The non-assistant platform is an authentication or authorization platform that shares a common publisher with the automated assistant platform.
10. The method according to claim 9, wherein, The non-assistant platform is a third-party platform that does not have the public publisher mentioned in the automated assistant platform.
11. The method according to any one of claims 1 to 10, wherein, The non-assistant platform causes the user's additional client device to respond to the received request and render the prompt for completing the user interaction with the non-assistant platform via the additional client device.
12. A method implemented by one or more processors, the method comprising: The system receives a request from the user's client device to initiate user interaction with a non-assistant platform, which is different from the automated assistant platform being used during a conversation between the user and an automated assistant accessible on the client device, and the request is received in response to determining that user input received during the conversation requires user interaction with the non-assistant platform, which is different from the automated assistant platform. In response to receiving the request from the client device: This causes prompts to be rendered on the user's additional client device for completing the user's interaction with the non-assistant platform, and Generate a token associated with the user interaction completed via the attached client device, wherein the token is generated based on user input received from the user on the attached client device in response to rendering the prompt. In response to generating the token associated with the user interaction: The token associated with the user interaction is transmitted to the user's client device, wherein transmitting the token associated with the user interaction causes the user's client device to store the token in association with the stored state of the conversation between the user and the automation assistant.
13. The method according to claim 12, wherein, Transmitting the token associated with the user interaction also causes the user's client device to load the state of the conversation between the user and the automation assistant on the client device via the automation assistant platform, along with the indication of the token.
14. The method according to claim 13, wherein, The indication that the state of the conversation session is loaded at the client device, along with the token, is in response to causing the client device to store the token associated with the completed user interaction.
15. The method according to claim 13, wherein, The state of the conversation session loaded on the client device, along with the indication of the token, is in response to receiving additional user input from the user on the client device to resume the conversation session.
16. The method according to claim 12, wherein, The conversation is paused on the client device in response to determining that the user input received during the conversation requires interaction with the user on the non-assistant platform, which is different from the automated assistant platform.
17. The method according to any one of claims 12 to 16, wherein, The automation assistant can be accessed at each of the client device, the additional client device, and the additional client device, wherein the additional client device is an addition to the client device and the additional client device, and wherein the additional client device is an addition to the client device and the additional client device.
18. The method according to claim 17, wherein, The one or more processors belong to the additional client device.
19. A system for pausing a conversational session between a user and an automated assistant, comprising: At least one processor; as well as At least one memory storing instructions, which, when executed, cause the at least one processor to perform the method according to any one of claims 1 to 18.
20. A non-transitory computer-readable storage medium storing instructions, which, when executed, cause at least one processor to perform the method of any one of claims 1 to 18.
Citation Information
Patent Citations
Balance modifications of audio-based computer program output
US20180358010A1
Systems and methods for facilitating network voice authentication
US20200153821A1