Proactive incorporation of unsolicited content into human-machine conversations
By configuring an automated assistant to proactively incorporate unrequested content under specific conditions, the problem of users needing initial input is solved, achieving resource savings and convenient information delivery.
Patent Information
- Application Number
- CN202511501345.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2017-03-22
- Filing Date
- 2017-11-17
- Publication Date
- 2026-03-03
AI Technical Summary
Existing automated assistants cannot proactively provide information that users may be interested in during human-computer dialogues. Users are required to provide initial natural language input to start the conversation, which leads to wasted resources and inconvenience to users.
Configure automated assistants to proactively incorporate unsolicited content into human-computer dialogue when specific conditions or events of the user are detected. Utilize natural language processing and various signals to identify information or actions that the user may be interested in and provide them within the existing dialogue.
It saves users' computing resources, reduces the need for natural language input, provides useful information that users may be interested in but have not requested, and improves the user experience.
Smart Images

Figure CN121597882A_ABST
Abstract
Description
Case Analysis
[0001] This application is a divisional application of Chinese invention patent application 201711146949.2, filed on November 17, 2017. Technical Field
[0002] This application involves proactively incorporating non-requested content into human-computer dialogue. Background Technology
[0003] Humans can participate in human-computer dialogues through what is referred to herein as an “automated assistant” (also known as a “chat room,” “interactive personal assistant,” “intelligent personal assistant,” “conversational agent,” etc.). For example, a human (who may be referred to as a “user” when interacting with an automated assistant) can provide commands and / or requests using spoken natural language input (i.e., utterances), which in some cases is converted to text and then processed, and / or by providing text-based (e.g., typed) natural language input. Automated assistants are typically passive rather than active. For example, at the start of a human-computer dialogue session between a user and an automated assistant (e.g., when there is no currently applicable conversational context), the automated assistant can at most provide general greetings such as “hey,” “good morning,” etc. Automated assistants cannot proactively acquire and provide specific information that the user may be interested in. Therefore, the user must provide initial natural language input (e.g., spoken or typed) before the automated assistant responds with substantial information and / or initiates one or more tasks on behalf of the user. Summary of the Invention
[0004] This document describes techniques for configuring automated assistants to proactively incorporate non-requested content that a user may be interested in into existing or newly initiated human-computer dialogue sessions. In some implementations, an automated assistant configured with selected aspects of this disclosure—and / or one or more other components that work in collaboration with the automated assistant—determines that, in an existing human-computer dialogue session, the automated assistant has effectively fulfilled its obligations to the user (e.g., the automated assistant is awaiting further instructions). This can be as simple as a user saying "Good morning" and the automated assistant providing the common response "Good morning." In this case, the user may still be (at least briefly) engaged in the human-computer dialogue session (e.g., the chatbot screen displaying the ongoing transcription of the human-computer dialogue may still be on, the user may still be within audible range of the audio input / output device enabling the human-computer dialogue, etc.). Therefore, any non-requested content incorporated into the human-computer dialogue session is likely to be consumed by the user (e.g., heard, seen, perceived, understood, etc.).
[0005] Incorporating non-requested content that users might find interesting into human-computer dialogue sessions can have several technical advantages. Users can confidently request this content, which can save computational resources that would otherwise be used to process the user's natural language input, and / or can help users who lack the ability to provide input (e.g., while driving, physically limited, etc.). Additionally, users can receive potentially useful content that they wouldn't normally expect to request. As another example, incorporating non-requested content can provide users with information they might be looking for by submitting additional requests. Avoiding these additional requests can save computational resources (e.g., network bandwidth, processing cycles, battery power) required to parse and / or interpret them.
[0006] In some implementations, in response to various events, the automated assistant can initiate a human-computer dialogue (for incorporating unrequested content) and / or incorporate unrequested content into an existing human-computer dialogue. In some implementations, the event may include the determination that the user is within the audible range of the automated assistant. For example, the independent interactive speaker operating the automated assistant may use various types of sensors (e.g., IP network cameras, or motion sensors / cameras incorporated into instruments such as smart thermostats, smoke detectors, carbon monoxide detectors, etc.), or detect the user's proximity by detecting the coexistence of another computing device carried by the user. In response, the automated assistant may provide the user with unrequested content such as "Don't forget your umbrella today, it's expected to rain," "Don't forget your sister's birthday today," "Did you hear that the power forward of <the sports team> was injured last night?" or "<the stock> has risen by 8% in the past few hours."
[0007] In some implementations, an automated assistant operating on a first computing device within a collaborative ecosystem of computing devices associated with a user can receive one or more signals from another computing device in the ecosystem. These signals may include the user's computational interactions (e.g., the user is performing a search, researching a topic, or reading a specific article), the state of an application operating on the other computing device (e.g., consuming media, playing a game, etc.), etc. For example, suppose a user is listening to a specific music artist on a separate interactive speaker (an instance that may or may not operate the automated assistant). The automated assistant on the user's smartphone can audibly detect the music and / or receive one or more signals from the separate interactive speaker, and in response, incorporate unrequested content into newly initiated or pre-existing human-computer dialogue, such as additional information about the artist (or song), suggestions about other similar artists / songs, notifications of upcoming travel dates or other artist-related events, etc.
[0008] Non-requested content incorporated into the human-computer dialogue may include information that the user may be interested in (e.g., weather, scores, traffic information, answers to questions, reminders, etc.) and / or actions that the user may be interested in (e.g., playing music, creating a reminder, adding items to a shopping list, creating, etc.). The selection of potentially interesting information and / or actions is based on various signals. In some implementations, these signals may include past human-computer dialogue between the user and the automated assistant. Assume that during the first human-computer session, the user researched flights to a specific destination but did not purchase any tickets. Further assume that a subsequent human-computer dialogue is triggered between the automated assistant and the user, and the automated assistant determines that it has responded to all natural language input from the user. In this case, the user has not yet provided any additional natural language input to the automated assistant. Therefore, the automated assistant may proactively incorporate non-requested content, including information related to the user's previous flight searches, such as "Have you purchased a ticket to your destination?" or "I don't know if you are still looking for flights, but I found some good deals on <website>."
[0009] Other signals that can be used to select information and / or actions to be incorporated into human-computer dialogue as unsolicited content include, but are not limited to, user location (e.g., this could prompt the automated assistant to proactively suggest specific menu items, special offers, etc.), calendar entries (e.g., "Don't forget your anniversary is next Monday"), appointments (e.g., an upcoming flight could prompt the automated assistant to proactively remind the user to perform online check-in and / or start packing), reminders, search history, browsing history, and topics of interest (e.g., interest in a particular sports team could prompt the automated assistant to proactively ask the user, "Did you see last night's score?"). Documents (e.g., emails including invitations to upcoming events prompt automated assistants to proactively remind users of upcoming events), application status (e.g., "I've performed updates for these three applications," "I see you still have several apps open, which is draining your device's resources," "I see you're currently streaming <Movie> to your TV. Do you know <Movie Trivia>?" etc.), newly available widgets (e.g., "Welcome back. While you were away, I learned how to call a taxi. Let me know if you need it," weather (e.g., "I see it's nice outside. Would you like me to search for restaurants with outdoor seating?"), etc.
[0010] In some embodiments, a method executed by one or more processors is provided, comprising: determining, by the one or more processors, in an existing human-computer dialogue session between a user and an automated assistant, that the automated assistant has responded to all natural language input received from the user during the human-computer dialogue session; identifying, by the one or more processors, information or one or more actions that the user may be interested in based on one or more characteristics of the user; generating, by the one or more processors, unrequested content representing the information or one or more actions that the user may be interested in; and incorporating the unrequested content into the existing human-computer dialogue session by the automated assistant. In various embodiments, at least the incorporation is performed in response to determining that the automated assistant has responded to all natural language input received from the user during the human-computer dialogue session.
[0011] These and other embodiments of the technology disclosed herein may optionally include one or more of the following features.
[0012] In various embodiments, the unrequested content includes unrequested natural language content. Throughout the embodiments, the identification is based at least in part on one or more signals obtained from one or more computing devices operated by the user. In various embodiments, the one or more computing devices operated by the user may include a designated computing device currently operated by the user.
[0013] In various embodiments, the one or more signals are received from another computing device, which is different from the designated computing device currently operated by the user, and is one of one or more computing devices operated by the user. In various embodiments, the one or more signals may include an indication of the status of an application executing on the other computing device. In various embodiments, the indication of the application's status may include an indication that the application is providing media playback. In various embodiments, the indication of the application's status may include an indication that the application has received a search query from the user or provided search results to the user.
[0014] In various embodiments, the method may further include one or more of the processors determining a desire metric representing that the user wants to receive unrequested content, wherein the desire metric is determined based on one or more signals, and wherein at least the incorporation is performed in response to determining that the desire metric satisfies one or more thresholds. In various embodiments, the unrequested content may include one or more user interface elements, each selectable by the user, to cause the automation assistant to provide information that the user may be interested in or to trigger one or more actions that the user may be interested in.
[0015] In another aspect, a method may include determining, based on one or more signals, that a user is within audible range of the one or more audio output devices; identifying, at least in part, information or actions that the user may be interested in, based on one or more characteristics of the user; generating unrequested content representing the potentially interesting information or actions; and incorporating the unrequested content into an audible human-computer dialogue session between the automated assistant and the user. In various embodiments, at least the incorporation may be performed by the automated assistant in response to determining that the user is within audible range of the one or more audio output devices.
[0016] Furthermore, some embodiments include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein the instructions are configured to cause any of the methods described above to be performed. Some embodiments also include one or more non-transitory computer-readable storage media storing computer instructions executable by one or more processors to implement any of the methods described above.
[0017] It should be recognized that all combinations of the foregoing concepts and the additional concepts described in more detail herein are contemplated as part of the subject matter disclosed herein. For example, all combinations of the claimed concepts appearing at the end of this disclosure are contemplated as part of the subject matter disclosed herein. Attached Figure Description
[0018] Figure 1 This is a block diagram of an exemplary environment in which the embodiments disclosed herein can be implemented.
[0019] Figures 2, 3, 4, 5, 6 and 7 illustrate exemplary dialogues between various users and automation assistants according to various implementations.
[0020] Figure 8 and Figure 9 This is a flowchart illustrating an exemplary method according to an embodiment disclosed herein.
[0021] Figure 10 An exemplary architecture of a computing device is shown. Detailed Implementation
[0022] Now go to Figure 1 This illustrates an exemplary environment in which the techniques described herein can be implemented. The exemplary environment includes multiple client computing devices 106. 1-N And Automation Assistant 120. Despite in Figure 1 In the diagram, the automation assistant 120 is shown as being connected to the client computing device 106. 1-NSeparately, but in some implementations, it can be handled by the client computing device 106. 1-N One or more of these can implement all or all aspects of the automation assistant 120. For example, client device 1061 can implement one or more aspects of the automation assistant 120, as well as client device 106... N Individual examples of one or more aspects of the automation assistant 120 can also be implemented. (This is achieved by a computing device 106 located remotely from the client.) 1-N In one or more embodiments of the automation assistant 120 implemented by one or more computing devices, the client computing device 106 1-N Those aspects of the Automation Assistant 120 can communicate via one or more networks, such as a Local Area Network (LAN) and / or a Wide Area Network (WAN) (e.g., the Internet).
[0023] Client device 106 1-N This may include one or more of the following: desktop computing devices, portable computing devices, tablet computing devices, mobile phone computing devices, computing devices in the user's vehicle (e.g., in-vehicle communication systems, in-vehicle entertainment systems, in-vehicle navigation systems), stand-alone interactive speakers, and / or wearable devices of the user that include computing devices (e.g., a user's watch with a computing device, glasses with a computing device, virtual or augmented reality computing devices). Additional and / or alternative client computing devices may be provided. In some embodiments, a designated user may communicate with the automation assistant 120 using multiple client computing devices of users that collectively form a coordinated "ecosystem" of computing devices. In some such embodiments, the automation assistant 120 may be considered to be "serving" that particular user, for example, granting the automation assistant 120 enhanced access to resources (e.g., content, documents, etc.) controlled by the "serving" user. However, for simplicity, some examples described herein will focus on the user operating a single client computing device 106.
[0024] Client computing device 106 1-N Each of them can operate multiple different applications, such as message exchange client 107. 1-N The corresponding one in the message exchange client 107. 1-N It can appear in various forms and the form can be on the client computing device 106 1-N The above changes and / or multiple forms can be operated on the client computing device 106 1-N On a single one. In some implementations, message exchange client 107 1-NOne or more of these can take the form of a Short Message Service (“SMS”) / Multimedia Messaging Service (“MMS”) client, an online chat client (e.g., an instant messaging service, internet relay chat, or “IRC”), a messaging application associated with a social network, a personal assistant messaging service dedicated to conversations with the automation assistant 120, etc. In some implementations, the message exchange client 107 can be implemented via a webpage or other resources rendered by a web browser (not shown) or other applications on the client computing device 106. 1-N One or more of them.
[0025] In addition to message exchange client 107, client computing device 106 1-N Each of them can also operate various other applications ( Figure 1 MISC.APP in 109 1-N These other applications may include, but are not limited to, game applications, media playback applications (such as music players, video players, etc.), productivity applications (such as word processing programs, spreadsheet applications, etc.), web browsers, map applications, reminder applications, cloud storage applications, camera applications, etc. As will be described in more detail below, in some embodiments, these other applications 109 1-N The various states are used as signals to prompt the automation assistant 120 to incorporate non-requested content into the human-computer dialogue.
[0026] As described in more detail herein, the automation assistant 120 is transmitted via one or more client devices 106 1-N The user interface input and output devices are used to join a human-computer dialogue session with one or more users. In some implementations, the user responds via client device 106. 1-N The automation assistant 120 can join a human-computer dialogue session with the user through one or more user interface input devices. In some of these embodiments, the user interface input is explicitly directed to the automation assistant 120. For example, message exchange client 107 1-N One of these could be a personal assistant messaging service dedicated to communicating with the automation assistant 120, and user interface input provided via this personal assistant messaging service could be automatically provided to the automation assistant 120. Furthermore, for example, based on specific user interface input indicating the automation assistant 120 to be invoked, the user interface input could be explicitly directed to one or more message exchange clients 107. 1-NThe automation assistant 120 is described in the text. For example, specific user interface input can be one or more type characters (e.g., @AutomatedAssistant), user interaction with hardware buttons and / or virtual buttons (e.g., tap, long tap), a password (e.g., "Hey, Automation Assistant"), and / or other specific user interface input. In some implementations, the automation assistant 120 may join a conversation even when the user interface input is not explicitly directed to it. For example, the automation assistant 120 may check the content of the user interface input and join a conversation in response to the presence of certain words in the input and / or based on other prompts. In many implementations, the automation assistant 120 may use interactive voice response ("IVR"), allowing the user to issue commands, search, etc., and the automation assistant may utilize natural language processing and / or one or more grammars to convert the speech into text and respond accordingly with the text.
[0027] Client computing device 106 1-N Each of the automation assistants 120 may include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components to facilitate communication over a network. It can be powered by one or more client computing devices 106. 1-N And / or the operations performed by the automation assistant 120 are distributed across multiple computer systems. The automation assistant 120 can be implemented as, for example, a computer program running on one or more computers in one or more locations coupled to each other via a network.
[0028] The automation assistant 120 may include a natural language processor 122 and a response content engine 130. In some implementations, one or more of the engine and / or modules of the automation assistant 120 may be omitted, combined, and / or implemented in components separate from the automation assistant 120. The automation assistant 120 may be implemented via a client device 106. 1-N Join a human-computer dialogue session with one or more users to provide response content generated and / or maintained by the response content engine 130.
[0029] In some implementations, the response content engine 130 responds during a human-computer dialogue session with the automation assistant 120, via the client device 106. 1-N The response content engine 130 generates response content from various user-generated inputs. The response content engine 130 provides response content (e.g., when disconnected from the user's client device, on one or more networks) to be presented to the user as part of a conversational session. For example, the response content engine 130 may respond to responses via client device 106. 1-NOne of the free-form natural language inputs provided in this paper is used to generate response content. As used in this paper, free-form input is input that is user-defined and not limited to a set of options selected by the user to be presented.
[0030] As used herein, a “dialogue session” can include a logically self-contained exchange of one or more messages between the user and the automation assistant 120. The automation assistant 120 can distinguish different dialogue sessions with the user based on various signals, such as the passage of time between sessions, changes in the user context between sessions (e.g., location, before / during / after a scheduled meeting), detecting one or more intervention interactions between the user and the client device (e.g., the user switching applications for a period of time, the user leaving and then returning to a standalone voice-activated product), locking / hibernating the client device between sessions, changing the client device used to interface with one or more instances of the automation assistant 120, etc.
[0031] In some implementations, when the automation assistant 120 provides a prompt requesting user feedback, the automation assistant 120 may preemptively activate one or more components of a client device (via which the prompt is provided) configured to process user interface input received in response to the prompt. For example, in the case where user interface input is provided via the microphone of client device 1061, the automation assistant 120 may provide one or more commands such that: the microphone is preemptively "turned on" (thus avoiding the need to hit an interface element or "hot word" to turn on the microphone), the local speech-to-text processor of client device 1061 is preemptively activated, a communication session between client device 1061 and a remote speech-to-text processor is preemptively established, and / or a graphical user interface is rendered on client device 1061 (e.g., an interface including one or more optional elements that can be selected to provide feedback). This allows for faster provisioning and / or processing of user interface input compared to components that are not preemptively activated.
[0032] The natural language processor 122 of the automation assistant 120 processes data sent by the user via client device 106. 1-N The generated natural language input can produce annotated output for use by one or more other components of the automation assistant 120, such as the response content engine 130. For example, the natural language processor 122 can process free-form natural language input generated by a user via one or more user interface input devices of the client device 1061. The generated annotated output includes one or more annotations to the natural language input, and optionally, one or more (e.g., all) terms of the natural language input.
[0033] In some implementations, the natural language processor 122 is configured to identify and annotate various grammatical information in the natural language input. For example, the natural language processor 122 may include a part-of-speech tagger configured to annotate terms based on their grammatical roles. For example, the part-of-speech tagger can label a term by its part of speech, such as "noun," "verb," "adjective," "pronoun," etc. Additionally, in some implementations, the natural language processor 122 may additionally and / or alternatively include a relevance parser configured to determine syntactic relationships between terms in the natural language input. For example, the relevance parser can determine which terms, topics, and verbs (e.g., parse trees) modify in a sentence—and can annotate these relevances.
[0034] In some implementations, the natural language processor 122 may additionally and / or alternatively include an entity tagger configured to annotate entity references, such as references to people (including, for example, literary characters), organizations, locations (real and virtual), etc., in one or more segments. The entity tagger may annotate entity references at a high-granularity level (e.g., to identify all references to an entity class, such as people) and / or a low-granularity level (e.g., to identify references to a specific entity, such as a specific person). The entity tagger may rely on the content of the natural language input to interpret specific entities and / or may optionally communicate with a knowledge graph or other entity database to interpret specific entities.
[0035] In some implementations, the natural language processor 122 may additionally and / or alternatively include a referential interpreter configured to cluster or “aggregate” references to the same entity based on one or more contextual cues. For example, the referential interpreter may be used to interpret the term “there” as “Hypothetical Café” in the natural language input “I liked the Hypothetical Café I went there last time”.
[0036] In some implementations, one or more components of the natural language processor 122 may rely on annotations from one or more other components of the natural language processor 122. For example, in some implementations, a designated entity tagger relies on annotations from a referential interpreter and / or a relevance parser when annotating all references to a particular entity. Meanwhile, for example, in some implementations, a referential interpreter relies on annotations from a relevance parser when aggregating references to the same entity. In some implementations, while processing a particular natural language input, one or more components of the natural language processor 122 may use relevant previous inputs and / or other relevant data besides the particular natural language input to determine one or more annotations.
[0037] As described above, the response content engine 130 uses one or more resources to generate suggestions and / or other content to communicate with the client device 106. 1-N Provided during a user's conversation session. In various implementations, the responsive content engine 130 may include an action module 132, an entity module 134, and a content module 136.
[0038] The action module 132 of the responsive content engine 130 utilizes the client computing device 106 1-N The received natural language input and / or annotations of the natural language input provided by the natural language processor 122 determine at least one action in response to the natural language input. In some embodiments, the action module 132 may determine the action based on one or more terms included in the natural language input. For example, the action module 132 may determine the action based on actions mapped to one or more terms included in the natural language input in one or more computer-readable media. For example, the action "add <item> to my shopping list" may be mapped to one or more terms such as "I need <item> in the market...", "I need to get <item>", "We've run out of <item>", etc.
[0039] Entity module 134 determines candidate entities based on input provided by one or more users via a user interface input device during a dialogue session between the user and the automation assistant 120. Entity module 134 uses one or more resources to determine and / or refine these candidate entities. For example, entity module 134 may utilize the natural language input itself and / or annotations provided by natural language processor 122.
[0040] The proactive content module 136 can be configured to proactively incorporate non-requested content that the user may be interested in into an existing or newly initiated human-computer dialogue session. For example, in some implementations, the proactive content module 136 can determine—for example, based on data received from other modules such as the natural language processor 122, the action module 132, and / or the entity module 134—that in an existing human-computer dialogue session between the user and the automation assistant 120, the automation assistant 120 has already responded to all natural language input received from the user during the human-computer dialogue session. Suppose a user operates the client device 106 to request a search for specific information, and the automation assistant 120 performs the search (or causes the search to be performed) and returns response information as part of the human-computer dialogue. At this point, unless the user requests additional information, the automation assistant 120 has fully responded to the user's request. In some implementations, the proactive content module 136 can wait for the automation assistant 120 for some predetermined time interval (e.g., two seconds, five seconds, etc.) to receive additional user input. If no input is received during the time interval, the active content module 136 can determine that it has responded to all natural language input received from the user during the human-computer dialogue session.
[0041] The proactive content module 136 can be further configured to identify information or actions that the user may be interested in (collectively referred to herein as "content that the user may be interested in") based on one or more characteristics of the user. In some embodiments, this identification of content that the user may be interested in can be performed by the proactive content module 136 at various time intervals (e.g., periodic, continuous, recurring, etc.). Thus, in some such embodiments, the proactive content module 136 can be "stimulated" continuously (or at least periodically) to provide unsolicited content that the user may be interested in. Additionally or alternatively, in some embodiments, the identification of content that may be interested in can be performed by the proactive content module 136 in response to various events. One such event could be determining that the automated assistant 120 has responded to all natural language input received from the user during the human-computer dialogue and has not received any additional user input before the aforementioned time interval expires. Other events that may trigger the active content module 136 to identify content that the user may be interested in may include, for example, the user performing a search using the client device 106, the user interacting with a specific application on the client device 106, the user moving to a new location (e.g., detected by the location coordinate sensor of the client device or by the location the user “checks in” on social media), the user being detected within the audible range of a speaker under the control of an automation assistant, etc.
[0042] User characteristics that can be used by the proactive content module 136 to determine content that a user may be interested in can take various forms and can be determined from various sources. For example, topics of interest to a user can be determined from sources such as the user's search history, browsing history, user settings preferences, location, media playback history, travel history, past human-computer dialogue sessions between the user and the automation assistant 120, etc. Therefore, in some embodiments, the proactive content module 136 can access various signals or other data from one or more client devices 106 operated by the user, for example, directly from client devices 106 and / or indirectly via one or more computing systems operating as a so-called "cloud". Topics of interest to the user can include, for example, specific interests (e.g., golf, skiing, games, painting, etc.), literature, movies, music genres, specific entities (e.g., artists, athletes, sports teams, companies), etc. Other user characteristics may include, for example, age, location (e.g., determined from a location coordinate sensor of client device 106, such as a Global Positioning System (“GPS”) sensor or other triangulation-based location coordinate sensor), user settings preferences, whether the user is currently in a moving vehicle (e.g., as determined by an accelerometer of client device 106), the user’s scheduled events (e.g., as determined by one or more calendar entries), etc.
[0043] In various implementations, the proactive content module 136 can be configured to generate unrequested content representing information and / or one or more actions that a user may be interested in, and to incorporate the unrequested content into the human-computer dialogue. This unrequested content can appear in various forms that can be incorporated into an existing human-computer dialogue session. For example, in some implementations where the user is interacting with the automation assistant 120 using a text-based messaging client 107, the unrequested content generated by the proactive content module 136 can take the form of text, images, video, or any combination thereof, and can be incorporated into a transcript of the human-computer dialogue rendered by the messaging client 107. In some implementations, the unrequested content can include or take the form of so-called “deep links,” which can be selected by the user to expose different application programming interfaces (APIs) to the user. For example, a deep link, when selected by the user, can cause the client device 106 to launch (or activate) a specific application 109 in a specific state. In other implementations where the user is interacting with the automation assistant 120 using a voice interface (e.g., when the automation assistant 120 operates on a stand-alone interactive speaker or on an in-vehicle system), the unrequested content can take the form of audible natural language output provided to the user.
[0044] In some implementations, the incorporation of unrequested content may be performed in response to, for example, a determination by the active content module 136 that the automation assistant 120 has responded to all natural language received from the user during the human-computer dialogue session. In some implementations, one or more of the other operations described above with respect to the active content module 136 may also be performed in response to such an event. Alternatively, as described above, these operations may be performed periodically or continuously by the active content module 136, such that the active content module 136 (and therefore the automation assistant 120) remains "stimulated" to rapidly incorporate unrequested content that the user may be interested in into the existing human-computer dialogue session.
[0045] In some implementations, the automation assistant 120 may provide unsolicited output even before the user initiates a human-computer dialogue session. For example, in some implementations, the active content module 136 is configured to determine, based on one or more signals, that the user is within audible range of one or more audio output devices (e.g., independent interactive or passive speakers operatively coupled to all or part of the client device 106 operating the automation assistant 120). These signals may include, for example, the coexistence of one or more client devices 106 carried by the user with the audio output devices, detection of the physical presence of the user (e.g., using passive infrared, sound detection (e.g., detecting the user's voice), etc.).
[0046] Once the active content module 136 has determined that the user is within audible range of one or more audio output devices, the active content module 136 may: at least in part, identify information or one or more actions that the user may be interested in (as described above) based on one or more characteristics of the user; generate unrequested content representing the information or one or more actions that the user may be interested in; and / or incorporate the unrequested content into an audible human-computer dialogue session between the automation assistant 120 and the user. As described above, one or more of these additional operations may be performed in response to determining that the user is within audible range of the audio output devices. Additionally or alternatively, one or more of these operations may be performed periodically or continuously, such that the active content module 136 is always (or at least usually) "stimulated" to incorporate unrequested content into the human-computer dialogue.
[0047] Figure 2 This shows an example of user 101 interacting with an automation assistant. Figure 1 120 in the middle, Figure 2 An example of a human-computer dialogue session between (not shown in the image). Figure 2This illustration shows an example of a conversational session occurring via a microphone and speaker between a user 101 (described as a standalone interactive speaker, but not limited thereto) and an automation assistant 120, according to the embodiments described herein. One or more aspects of the automation assistant 120 may be implemented on the computing device 210 and / or on one or more computing devices that communicate with the computing device 210 via a network.
[0048] exist Figure 2 In the example, user 101 provides natural language input 280, “Good morning. What’s your schedule for today?” to initiate a human-computer dialogue session between user 101 and automation assistant 120. Responding to natural language input 280, automation assistant 120 provides natural language output 282, “You have an appointment with the dentist at 9:30 AM and a meeting at the Hypothetical Café at 11:00 AM.” Assuming these are the only two events in the user’s schedule for the day, automation assistant 120 (e.g., via action module 132) has already fully responded to the user’s natural language input. However, instead of waiting for additional user input, automation assistant 120 (e.g., via proactive content module 136) can proactively incorporate additional content that might be of interest. Figure 2 In a human-computer dialogue, for example, the automation assistant 120 can search (or request another component to search) one or more routes between the dentist and the meeting location, for instance, to determine the most direct route in case of major construction. Since the two appointments are relatively close, the automation assistant 120 proactively incorporates the following unrequested content (shown in italics) into the human-computer dialogue box: “There is major construction on the direct route between the dentist and the Hypothetical Café. Is an <alternative route> recommended?”
[0049] Figure 3 The illustration shows another exemplary dialogue between user 101 and automation assistant 120 operating on computing device 210 during different sessions. At 380, user 101 says the phrase “What’s the outdoor temperature?”. After the outdoor temperature is determined by one or more sources (e.g., a weather-related web service), at 382, automation assistant 120 can answer “75 degrees Fahrenheit.” Again, automation assistant 120 (e.g., via proactive content module 136) can determine that it has fully responded to the user’s natural language input. Therefore, based on user 101’s interest in a particular team and determining that the team won the game the previous night, automation assistant 120 can proactively incorporate the following unrequested content into the human-machine dialog box: “Did you see that <team> won by twenty points last night?”
[0050] Figure 4The illustration depicts another exemplary dialogue between user 101 and automation assistant 120 operating on computing device 210 during different sessions. In this example, user 101 does not provide natural language input. Instead, automation assistant 120, or another component operating on computing device 210, determines the presence of user 101 with computing device 210 based on one or more signals provided by client device 406 (a smartphone in this example), thereby placing them within audible range of audible output provided by computing device 210. Therefore, at 482, automation assistant 120 will output non-requested content (with...) Figure 3 The same unrequested content is proactively incorporated into a new human-computer dialogue initiated by the automation assistant based on the coexistence of user 101 and computing device 210. One or more signals provided by client device 406 to computing device 210 may include, for example, wireless signals (e.g., Wi-Fi, Bluetooth), network sharing (e.g., client device 406 joining the same Wi-Fi network as computing device 210, etc.).
[0051] In some implementations, the automation assistant 120 may proactively incorporate other content that the user 101 might be interested in into the human-computer dialogue when it determines that the user 101 is coexisting with the computing device 210. In some implementations, this other content may be determined, for example, based on the state of an application running on the client device 406. Suppose the user 101 is playing a game on the client device 406. The automation assistant 120 on the computing device 210 can determine that the client device 406 is in a specific game state and can provide various unsolicited content that the user might be interested in, such as tips, tricks, recommendations for similar games, etc., as part of the human-computer dialogue. In some implementations where the computing device 210 is a standalone interactive speaker, the computing device 210 may even output background music (e.g., copied or added background music) and / or sound effects associated with the game played on the client device 406, at least as long as the user 101 remains coexisting with the computing device 210.
[0052] Figure 5 The illustration depicts an exemplary human-computer dialogue between user 101 and an instance of an automated assistant 120 operating on client device 406. In this example, user 101 again does not provide natural language input. Instead, computing device 210 (again in the form of an independent interactive speaker) is playing music. This music is detected at one or more audio sensors (e.g., microphones) on client device 406. One or more components of client device 406, such as software applications configured to analyze the audibly detected music, can identify one or more attributes of the detected music, such as artist / song, etc. Another component, such as Figure 1The entity module 134 can use these attributes to search for information about the entity in one or more online sources. Then, the automation assistant 120 operating on client device 406 can provide (at 582) – for example via… Figure 5 One or more speakers of the client device 406 loudly output non-requested content, which informs the user of various information about the entity. For example, in Figure 5 In 582, the automation assistant 120 says, “I see you are listening to <The Artist>. Do you know if <The Artist> has a tour itinerary in <Your Town>?”. A similar technique can be applied when an instance of the automation assistant 120 operating on a client device (e.g., a smartphone, tablet, laptop, or standalone interactive speaker) detects (via sound and / or vision) audiovisual content (e.g., movies, TV shows, sporting events, etc.) presented on the user’s television.
[0053] exist Figure 5 In this context, computing device 210 audibly outputs music that is "heard" by client device 406. However, it is assumed that user 101 is listening to music using client device 406 instead of computing device 210. Further, it is assumed that user 101 is using earphones to listen to music, making the music audible only to user 101 and not necessarily to other computing devices such as computing device 210. In various implementations, particularly where client device 406 and computing device 210 are part of the same ecosystem of computing devices associated with user 101, computing device 210 can determine that the music playback application of client device 406 is currently in a state where it is playing music. For example, client device 406 can, for instance, use wireless communication technologies such as Wi-Fi, Bluetooth, etc., to provide an indication of the status of the music playback application (and / or other application user status) to nearby devices (such as computing device 210). Additionally or alternatively, for the ecosystem of computing devices operated by user 101, a global index of currently executing applications and their respective states can be maintained (e.g., by an automated assistant serving user 101) and used among the computing devices in the ecosystem. In either case, as long as the automation assistant 120 associated with the computing device 210 knows the status of the music playback application on the client device 406, the automation assistant 120 can, via the computing device 210 (which can be triggered by the automation assistant 120), transfer the music playback application to the client device 406. Figure 5 Similar content, as illustrated in 582 instances, is proactively incorporated into, for example, the human-computer dialogue between user 101 and automation assistant 120.
[0054] Figure 2-5The illustration shows user 101 joining a human-computer dialogue with automation assistant 120 using audio input / output. However, this does not imply limitation. As mentioned above, in various implementations, users can join the automation assistant using other means such as message exchange client 107. Figure 6 The illustration shows an example of a client device 606 in the form of a smartphone or tablet (but not limited thereto), including a touchscreen 640. Visually rendered on the touchscreen 640 is the user of the client device 606. Figure 6 The transcription 642 of the human-computer dialogue between the user ("you") and an instance of the automated assistant 120 executed on the client device 606. An input field 644 is also provided, whereby the user can provide natural language content as well as other types of input such as images and sounds.
[0055] exist Figure 6 In the process, users initiate a human-computer dialogue session by asking the question, "What are the opening hours of the store?" (Automation Assistant 120) Figure 6 For example, via action module 132 or another component, the "AA" in the example performs one or more searches for information related to the store's opening hours and responds with "<store> opens at 10:00 AM". At this point, the automation assistant 120 has responded to the only natural language input provided by the user in the current human-computer dialogue session. However, in this example, suppose the user recently operated client device 606 or another client device in the ecosystem that includes client device 606 to search for flight tickets to New York. The user performs this search by joining one or more human-computer dialogue sessions with the automation assistant 120 via a web browser or a combination thereof.
[0056] Based on this past search activity, in some implementations, the automation assistant 120 (e.g., via the proactive content module 136) can—periodically / continuously or responsively—search for information relevant to the search, and thus potentially of interest to the user, from one or more online sources after determining that the automation assistant 120 has responded to all received natural language input in the current human-computer dialogue session. The automation assistant 120 can then proactively incorporate the following unrequested content into... Figure 6In the illustrated human-computer dialogue session: "Have you bought a ticket to New York? I found a deal on direct flights and hotels." The automated assistant 120 (e.g., via the proactive content module 136) can then incorporate additional non-requested content into the dialogue in the form of user interface elements (e.g., deep links) 646 that can be selected by the user to open a travel application installed on client device 606. If user interface element 646 is selected, the travel application can open to a booking state, for example, pre-populated input about booking a flight to New York, so the user can edit the pre-populated information and / or issue relatively few additional input commands to purchase the ticket. Other non-requested content can also be presented as optional options, such as one or more hyperlinks 648 on a webpage, for example, for booking flights or general information about the destination.
[0057] although Figure 6 User interface element 646 is a visual element that can be selected by clicking or otherwise touching, but this does not imply limitation. Figure 6 Similar human-computer dialogues can occur audibly between the user and an audio output device (e.g., a separate interactive speaker shown in a previous figure). In some such embodiments, the user interface elements can take the form of audible prompts, such as questions or options that can be “selected” if answered affirmatively by the user. For example, instead of presenting a visual user interface element 646, the automation assistant 120 could audibly output something like, “Tell me if you want me to open the <Travel App> to book a flight to New York.” In some embodiments, the travel app itself may include its own automation assistant, specifically tailored to join the human-computer dialogue with the user to book travel arrangements. In some such embodiments, the user can be “transferred” to the travel app-specific automation assistant. In other embodiments, the automation assistant 120 can utilize various information and states associated with the travel app to formulate natural language output for using the travel app, from the user’s request for information to book a flight. The automation assistant 120 can then interact with the travel app on behalf of the user (e.g., responding to spoken natural language input provided by the user).
[0058] Figure 7The diagram again illustrates a client device 606 with a touchscreen 640, user input fields 644, and transcription 742 of the human-computer dialogue session. In this example, the user (“you”) initiates the human-computer dialogue by typing and / or speaking (which can be recognized and converted into text) the natural language input “Good morning”. The automation assistant 120 (“AA”) responds: “Good morning”. At this point, there are no further pending natural language requests from the user that need to be completed. Therefore, using the techniques described herein, the automation assistant 120 determines, for example, from the user’s associated calendar, correspondences exchanged between the user and others (e.g., emails, text messages), a restaurant app installed on the client device 6060, etc., that the user has a reservation at <restaurant> tonight. Therefore, the automation assistant 120 proactively incorporates the following statement into the human-computer dialogue: “Don’t forget you have a reservation at <restaurant> tonight.” The automation assistant 120 then proactively incorporates a user interface element 746 (“My Reservations”) that the user can choose to open the restaurant app to a state where the user can view and / or edit / cancel the reservation. In addition, the automation assistant 120 proactively incorporates other content that the user may be interested in, such as hyperlinks 748 to websites associated with restaurant reservations and / or nearby attractions. In some implementations, as soon as the user arrives at the restaurant, the automation assistant 120 can incorporate different unsolicited content into the same human-computer dialogue session or a new human-computer dialogue session, such as photos, reviews, suggestions, special offers, etc., previously taken at the restaurant (by the user and / or others).
[0059] The examples of proactively incorporated unrequested content described above are not intended to be limiting. Other unrequested content that a user may be interested in can be proactively incorporated into human-computer dialogue using the techniques described herein. For example, in some implementations where the user has an upcoming scheduled flight (or train departure or other travel arrangements), the automation assistant 120 may proactively incorporate unrequested content into the human-computer dialogue session with the user. This unrequested content may include, for example, a reminder that the user's flight is about to arrive, one or more user interface elements of an application that can be optionally opened (via touch, voice, gesture, etc.) to allow the user to view or edit the booked flight, information about (or connected to optional user interface elements) the time to travel to the airport, etc. Alternatively, if the automation assistant 120 determines (e.g., based on the user's schedule, location coordinate sensors, etc.) that the user's flight has arrived at its destination, the automation assistant 120 may proactively incorporate various information and / or user interface elements that the user may be interested in into a new or existing human-computer dialogue session, such as information / user interface related to calling a service (or running a ride-sharing application), directions to hotels or other attractions, nearby restaurants, etc.
[0060] As another example, the automation assistant 120 can determine that changes have been made to one or more computing devices operated by the user (in some cases, this could be part of a coordinated ecosystem of computing devices associated with the user). For instance, the automation assistant 120 can determine that one or more applications (including the automation assistant 120 itself) installed on one or more client devices associated with the user have been updated since the last human-computer interaction session with the user. Because the user may be interested in being informed of these updates, the automation assistant can incorporate non-requested content, such as “Welcome back. I learned how to call a taxi when you left. Let me know whenever you need it,” into the human-computer interaction.
[0061] As another example, in some implementations, the automation assistant 120 can determine various information that a user might be interested in at a specific time (e.g., based on one or more topics of general interest to the user, the user's browsing history, etc.) and can proactively incorporate unrequested content related to this information into the human-computer dialogue session with the user. For example, suppose a particular user is interested in history and electronics. In various implementations, when the automation assistant 120 determines that it has responded to all natural language input received from the user during an existing human-computer dialogue session, the automation assistant 120 can proactively incorporate, for example, information that is currently relevant and that the user might be interested in. For example, on Nicola Tesla's birthday, a user interested in history and electronics could be presented with user interface elements that the user can choose to open Tesla-related applications or web pages. As another example, suppose today is a user's wedding anniversary. The automation assistant can proactively incorporate graphic elements or other information that the user might be interested in on the anniversary, such as links to flower websites or restaurants, into the existing human-computer dialogue session.
[0062] As yet another example, in some implementations, the user's location (e.g., determined by location coordinate sensors on a computing device carried by the user) can prompt the automation assistant 120 to proactively incorporate unrequested content into the human-computer dialogue session with the user. For example, suppose the user is at or near a grocery store. The automation assistant 120 can determine, for example, based on one or more shopping lists associated with the user (e.g., locally stored on the client device or based in the cloud) the items the user should obtain at the grocery store. The automation assistant 120 can then proactively incorporate unrequested content into the human-computer dialogue with the user, where unrequested content includes the desired items, information about the items, transactions available for the items, etc.
[0063] As another example, in some implementations, frequently requested information or actions can be proactively incorporated into the human-computer dialogue session as unrequested content. For instance, suppose a user is discussing various topics with an automation assistant 120, and the time is close to the user's usual dinner time. In some implementations, the automation assistant 120 can incorporate unrequested content related to dining, such as user interface elements that the user can choose to order pizza, open a recipe (from local storage or from a frequently visited recipe webpage), into the existing human-computer dialogue session. In other implementations, unrequested content that can be incorporated into the existing human-computer dialogue session may include, but is not limited to, trending news stories, popular searches, and updated search results for previously submitted search queries by the user.
[0064] Of course, users may not always expect unsolicited content. For example, a user may be driving in heavy traffic, in an emergency, or operating a computing device in a manner that implies the user does not want to receive unsolicited content (e.g., during a video call). Therefore, in some implementations, the automation assistant 120 may be configured to determine (e.g., based on signals such as location signals, the context of the session, the state of one or more applications, accelerometer signals, etc.) a metric of the user's desire to receive unsolicited content, and only provide the unsolicited content when that metric meets one or more thresholds.
[0065] Similarly, in some implementations, the automation assistant 120 may provide unsolicited content (as part of a new or existing human-computer dialogue session) during a specific time period. For example, if a user is detected within audible range of the client device operating the automation assistant 120 between 7:00 AM and 8:00 AM, the automation assistant 120 may automatically output unsolicited greetings such as "Good morning," "Don't forget your umbrella, it's going to rain," "Traffic is heavy on Route 405," "These are today's headlines," "This is your schedule for today," etc.
[0066] As another example, in some implementations, the automation assistant 120 may consider the activities of multiple users at a specific time and / or location to determine which particular user necessarily wants to receive unsolicited content. In various implementations, the automation assistant 120 may analyze search queries from multiple users to identify spikes, popular and / or other patterns in searches associated with a specific location, specific time, etc. For example, suppose many users visiting a landmark perform similar web searches on their mobile devices, such as "how many floors does it have," "when was it built," "how many years ago," etc. Upon detecting patterns or popularity in these searches, the automation assistant 120 may proactively offer unsolicited content to new users when they arrive at the landmark.
[0067] Figure 8 This is a flowchart illustrating an exemplary method 800 according to an embodiment disclosed herein. For convenience, the operations of the flowchart are described with reference to a system performing the operations. This system may include various components of various computer systems, such as one or more components of the automation assistant 120. Furthermore, although the operations of method 800 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.
[0068] At box 802, the system can determine that, in an existing human-computer dialogue session between the user and the automated assistant, the automated assistant has responded to all natural language input received from the user during the dialogue session. In some implementations, this may include waiting for a predetermined time interval after responding to all natural language input, although this is not required.
[0069] In some implementations, at block 804, in response to further determining that the user is likely to expect unrequested content (i.e., as can be represented by the “desire metric” described above), the system may proceed only with one or more of operations 806-810. This further determination may be based on various sources, such as the conversational context of the human-computer dialogue session, the user context determined by signals unrelated to the human-computer dialogue session (e.g., position signals, accelerometer signals, etc.), or a combination thereof. For example, if it is determined based on the user’s accelerometer and / or position coordinate sensor signals that the user is currently driving (e.g., after the user inquires about traffic updates or directions), the system may determine that the user is unlikely to be disturbed by unrequested content. As another example, the context of the human-computer dialogue session may imply that the user does not want to be disturbed by unrequested content. For example, if the user inquires with the automated assistant about the location of the nearest emergency room or about treatment for an injury, the determined desire metric will be relatively low (e.g., not meeting a threshold), and the automated assistant may prohibit the provision of unrequested content after the requested information. As another example, if a user asks an automated assistant to trigger an action that may take some time to complete and requires the user's attention (e.g., starting a video call, initiating a phone call, playing a movie, etc.), the user is less likely to be disturbed by additional unsolicited content.
[0070] In box 806, the system can identify information or actions that the user may be interested in based on one or more user characteristics. As described above, the operation of box 806 can be performed in response to the determinations in boxes 802-804, or it can be performed on a continuous basis, causing the automated assistant to be "stimulated" to provide unrequested content at any specified point in time. In various implementations, the automated assistant can identify information or actions that the user may be interested in based on a variety of sources, including but not limited to the user's search history, browsing history, human-computer interaction history (including the same session and / or previous sessions on the same or different client devices), the user's location (e.g., determined from the user's schedule, social network status (e.g., check-in), location coordinate sensors, etc.), timetable / calendar, common topics of interest to the user (which can be manually set by the user and / or learned based on the user's activities), etc.
[0071] In box 808, the system can generate non-requested content representing information or one or more actions that the user may be interested in. This non-requested content may include, for example, natural language output providing information that the user may be interested in, such as in a natural language format (e.g., audible output or visual form), or interface elements (graphical or audible) that the user can choose to obtain additional information and / or trigger one or more tasks (e.g., setting reminders, creating calendar entries, creating appointments, opening the application in a scheduled state, etc.).
[0072] In box 810, the system can incorporate the unsolicited content generated in box 808 into the existing human-computer dialogue session. For example, the unsolicited content can be presented as automated speech output from the automation assistant, user interface elements such as cards, hyperlinks, audible prompts, etc. Incorporating unsolicited content into the existing human-computer dialogue differs from simply presenting information to the user (e.g., as a card or drop-down menu on the lock screen). The user is already engaged in a human-computer dialogue session with the automation assistant, so the unsolicited content is more likely to be seen / heard and interacted with by the user, rather than simply being presented to the user's lock screen (where users often ignore and / or are overwhelmed by too many notifications).
[0073] Figure 9 This is a flowchart illustrating an exemplary method 900 according to an embodiment disclosed herein. For convenience, the operations of the flowchart are described with reference to a system performing the operations. This system may include various components of various computer systems, such as one or more components of the automation assistant 120. Furthermore, although the operations of method 900 are shown in a specific order, this is not intended to be limiting. One or more operations may be reordered, omitted, or added.
[0074] At box 902, the system can determine that a user is within audible range of one or more audio output devices (e.g., one or more speakers of a computing device operatively coupled to an instance of an automated assistant, a separate interactive speaker of an instance of an automated assistant, etc.) based on one or more signals. These signals can take various forms. In some embodiments, one or more signals may be triggered by a computing device operated by the user, separate from the system, and received at one or more communication interfaces operatively coupled to one or more processors. For example, a computing device may push notifications to other computing devices about the user's participation in a specific activity, such as driving, operating a specific application (e.g., playing music or a movie), etc. In some embodiments, one or more signals may include detecting the coexistence of the system and the computing device. In some embodiments, one or more signals may include indications of the state of an application running on a computing device separate from the system, such as the user preparing a document, performing various searches, playing media, viewing photos, joining a phone / video call, etc. In some embodiments, a human-computer interaction may be initiated in response to determining that the user is within audible range of one or more audio output devices, although this is not necessary. Figure 9 Boxes 904-908 can be similar to Figure 8 Boxes 804-808. Although in Figure 9 Not shown in the figure, but in various implementations, the automation assistant may determine whether the user might expect unrequested content before it is provided, as described in reference box 804 above.
[0075] Figure 10 This is a block diagram of an exemplary computing device 1010, which may be used to perform one or more aspects of the techniques described herein. In some embodiments, one or more of a client computing device, an automation assistant 120, and / or other components may include one or more components of the exemplary computing device 1010.
[0076] Computing device 1010 typically includes at least one processor 1014 that communicates with multiple peripheral devices via a bus subsystem 1012. These peripheral devices may include a storage subsystem 1024, such as a memory subsystem 1025 and a file storage subsystem 1026, a user interface output device 1020, a user interface input device 1022, and a network interface subsystem 1016. The input and output devices allow users to interact with computing device 1010. The network interface subsystem 1016 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.
[0077] User interface input device 1022 may include a keyboard, pointing devices such as a mouse, trackball, or graphics tablet, a scanner, a touchscreen integrated into a display, audio input devices such as a voice recognition system, a microphone, and / or other types of input devices. Generally, the term "input device" is used to encompass all possible types of devices and methods for inputting information into computing device 1010 or a communication network.
[0078] User interface output device 1020 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating visible images. The display subsystem may also provide non-visual displays via an audio output device. Generally, the term "output device" is used to encompass all possible types of devices and methods for outputting information from computing device 1010 to a user or another machine or computing device.
[0079] Storage subsystem 1024 stores programs and data structures that provide some or all of the functionality of the modules described herein. For example, storage subsystem 1024 may include programs that execute... Figure 8 and Figure 9 The selection of methods and their implementation Figure 1 The logic of each component is shown.
[0080] These software modules are typically executed by processor 1014 alone or in conjunction with other processors. The memory 1025 used in the storage subsystem may include various types of memory, including main random access memory (RAM) 1030 for storing instructions and data during program execution and read-only memory (ROM) 1032 for storing fixed instructions therein. The file storage subsystem 1026 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives with associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Modules implementing the functionality of certain embodiments may be stored in storage subsystem 1024 by file storage subsystem 1026 or in other machines accessible to processor 1014.
[0081] Bus subsystem 1012 provides a mechanism that enables the various components and subsystems of computing device 1010 to communicate with each other as desired. Although bus subsystem 1012 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0082] The computing device 1010 can be of various types, including workstations, servers, computing clusters, blade servers, server groups, or any other data processing system or computing device. Due to the constantly changing properties of computers and networks, Figure 10 The description of the computing device 1010 depicted herein is intended only as a specific example to illustrate some implementation methods. Many other configurations of the computing device 1010 are possible, which have more... Figure 10 The computing device depicted has more or fewer components.
[0083] In certain implementations discussed herein that allow the collection or use of personal information about users (e.g., user data extracted from other electronic communications, information about users' social networks, user location, user time, user biometric information, user activity and demographic information, relationships between users, etc.), users are provided with one or more opportunities to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about users is collected, stored, and used. In other words, the systems and methods discussed herein may only collect, store, and / or use user personal information upon receiving explicit authorization from the relevant user.
[0084] For example, users may be provided with control over whether a program or feature collects user information about that particular user or other users associated with that program or feature. One or more options may be presented to each user whose personal information will be collected, allowing control over the collection of information related to that user, providing permission or authorization regarding whether information is collected, and regarding which parts of the information will be collected. For example, one or more such control options may be provided to users on a communication network. Furthermore, certain data may be processed in one or more ways before being stored or used to erase personally identifiable information. As an example, a user's identity may be processed to the point that personally identifiable information cannot be determined. As another example, a user's geographic location may be generalized to a larger area, making it impossible to determine the user's specific location. In the context of this disclosure, any relationships captured by the system, such as parent-child relationships, may be maintained in a secure manner, for example, such that these relationships cannot be used outside of automated assistants to parse / or interpret natural language input.
[0085] In addition to the benefits mentioned above, it should be recognized that the techniques described in this article can make automated assistants appear more “realistic” or “human” to users, which may incentivize increased interaction with automated assistants.
[0086] Although several implementations have been described and illustrated herein, various other means and / or structures may be used to perform the functions and / or obtain the results and / or one or more advantages described herein, each such variation and / or modification being considered within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and constructions described herein are intended to be exemplary, and actual parameters, dimensions, materials, and / or constructions will depend on the specific application or use of the teachings. Those skilled in the art will recognize or be able to determine many equivalents of the specific implementations described herein using only conventional experimentation. Therefore, it should be understood that the foregoing implementations are given by way of example only, and that implementations may be practiced in ways different from those specifically described and claimed within the scope of the appended claims and their equivalents. Implementations of this disclosure relate to each individual feature, system, article, material, kit, and / or method described herein. Furthermore, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of this disclosure if these features, systems, articles, materials, kits, and / or methods do not contradict each other.
Claims
1. A method implemented using one or more processors, comprising: During a human-computer dialogue session between a user and an automated assistant, natural language input is received from the user. The natural language input is processed to determine if the user has requested information about a trip involving two or more events; Generate natural language output and have the automated assistant incorporate the natural language output into the human-computer dialogue session, the natural language output conveying requested information about the schedule of two or more events; Based on one or more of the events, identify information or one or more actions that the user may be interested in, wherein the identification includes searching one or more routes between the locations of the two or more events; Generate non-requested content indicating information or one or more actions that the user may be interested in, wherein the information that the user may be interested in includes response content obtained by searching the one or more travel routes; and Based on the fact that the time interval between the two or more events meets a threshold, the non-requested content is conditionally merged into the existing human-computer dialogue session.
2. The method according to claim 1, wherein, The response content includes information about traffic between the locations of the two or more events.
3. The method according to claim 1, wherein, The response content includes information about construction occurring between the locations of the two or more events.
4. The method according to claim 1, wherein, The information that the user may be interested in includes facts about entities that will exist at one or more of the events in the event.
5. The method according to claim 1, wherein, The information that the user may be interested in includes facts about the entities that host one or more of the events in the event.
6. The method according to claim 1, wherein, One or more actions that may be of interest include launching a ride-sharing application.
7. The method according to claim 1, wherein, At least the merging is performed in response to determining that the automated assistant has responded to all natural language input received from the user during the human-computer dialogue session.
8. A system comprising one or more processors and a memory storing instructions, the instructions causing the one or more processors to perform the following operations in response to execution of the instructions: During a human-computer dialogue session between a user and an automated assistant, natural language input is received from the user. The natural language input is processed to determine if the user has requested information about a trip involving two or more events; Generate natural language output and have the automated assistant incorporate the natural language output into the human-computer dialogue session, the natural language output conveying requested information about the schedule of two or more events; Based on one or more of the events, identify information or one or more actions that the user may be interested in, wherein the instructions for identification include instructions for searching one or more routes of travel between the locations of the two or more events; Generate non-requested content indicating information or one or more actions that the user may be interested in, wherein the information that the user may be interested in includes response content obtained by searching the one or more travel routes; and Based on the fact that the time interval between the two or more events meets a threshold, the non-requested content is conditionally merged into the existing human-computer dialogue session.
9. The system according to claim 8, wherein, The response content includes information about traffic between the locations of the two or more events.
10. The system according to claim 8, wherein, The response content includes information about construction occurring between the locations of the two or more events.
11. The system according to claim 8, wherein, The information that the user may be interested in includes facts about entities that will exist at one or more of the events in the event.
12. The system according to claim 8, wherein, The information that the user may be interested in includes facts about the entities that host one or more of the events in the event.
13. A non-transitory computer-readable medium comprising instructions, the instructions being responsive to execution by a processor to cause the processor to perform the following operations: During a human-computer dialogue session between a user and an automated assistant, natural language input is received from the user. The natural language input is processed to determine if the user has requested information about a trip involving two or more events; Generate natural language output and have the automated assistant incorporate the natural language output into the human-computer dialogue session, the natural language output conveying requested information about the schedule of two or more events; Based on one or more of the events, identify information or one or more actions that the user may be interested in, wherein the instructions for identification include instructions for searching one or more routes of travel between the locations of the two or more events; Generate non-requested content indicating information or one or more actions that the user may be interested in, wherein the information that the user may be interested in includes response content obtained by searching the one or more travel routes; and Based on the fact that the time interval between the two or more events meets a threshold, the non-requested content is conditionally merged into the existing human-computer dialogue session.
14. A method implemented using one or more processors, comprising: Determine all natural language input received from the user during an existing human-computer dialogue session between the user and an automated assistant that occurred at one or more computing devices operated by the user. Determine the user's current context; Analyze search queries submitted by other users to identify spikes, popular or other patterns in the search queries submitted by other users in contexts similar to the current context of the user; Based on the peak, popular, or other patterns, select one or more search queries from the search queries submitted by other people; Search one or more online sources in response to one or more selected search queries; One or more of the processors generate unrequested content, which indicates information in response to one or more selected search queries; as well as The automated assistant merges the unrequested content into the existing human-computer dialogue session; Wherein, at least the merging is performed in response to determining that the automation assistant has responded to all natural language input received from the user during the human-computer dialogue session.
15. The method according to claim 14, wherein, The non-requested content includes non-requested natural language content.
16. The method of claim 14, further comprising determining a desire metric indicative of the user's desire to receive unrequested content, wherein, The desire metric is determined at least in part based on the user's current context, and wherein at least the merging is performed in response to determining that the desire metric satisfies one or more thresholds.
17. The method of claim 14, wherein, The user's current environment is determined based on one or more signals generated by one or more sensors integrated with one or more computing devices in the computing device.
18. The method according to claim 17, wherein, The user's current environment includes position coordinates generated by the position coordinate sensors of one or more computing devices operated by the user.
19. The method of claim 17, wherein, The user's current environment includes sensor data generated by the accelerometers of one or more computing devices operated by the user.
20. The method of claim 14, wherein, The user's current context includes the user's location.
21. The method according to claim 20, wherein, The user's location is determined based on the user's schedule or calendar.
22. The method according to claim 20, wherein, The selected search query or one search query retrieves information about the user's location.
23. A system comprising one or more processors and a memory storing instructions, the instructions causing the one or more processors to perform the following operations in response to execution by the one or more processors: Determine all natural language input received from the user during an existing human-computer dialogue session between the user and an automated assistant that occurred at one or more computing devices operated by the user. Determine the user's current context; Analyze search queries submitted by other users to identify spikes, popular or other patterns in the search queries submitted by other users in contexts similar to the current context of the user; Based on the peak, popular, or other patterns, select one or more search queries from the search queries submitted by other people; Search one or more online sources in response to one or more selected search queries; One or more of the processors generate unrequested content, which indicates information in response to one or more selected search queries; as well as The automated assistant merges the unrequested content into the existing human-computer dialogue session; Wherein, at least the merging is performed in response to determining that the automation assistant has responded to all natural language input received from the user during the human-computer dialogue session.
24. The system according to claim 23, wherein, The non-requested content includes non-requested natural language content.
25. The system of claim 23, further comprising instructions for: determining a desire metric indicative of the user's desire to receive unrequested content, wherein, The desire metric is determined at least in part based on the user's current context, and wherein at least the merging is performed in response to determining that the desire metric satisfies one or more thresholds.
26. The system according to claim 23, wherein, The user's current environment is determined based on one or more signals generated by one or more sensors integrated with one or more computing devices in the computing device.
27. The system according to claim 26, wherein, The user's current environment includes position coordinates generated by the position coordinate sensors of one or more computing devices operated by the user.
28. The system according to claim 26, wherein, The user's current environment includes sensor data generated by the accelerometers of one or more computing devices operated by the user.
29. The system according to claim 23, wherein, The user's current context includes the user's location.
30. The system according to claim 29, wherein, The user's location is determined based on the user's schedule or calendar.
31. The system according to claim 29, wherein, The selected search query or one search query retrieves information about the user's location.
32. At least one non-transitory computer-readable medium comprising instructions that, in response to execution by one or more processors, cause the one or more processors to perform the following operations: Determine all natural language input received from the user during an existing human-computer dialogue session between the user and an automated assistant that occurred at one or more computing devices operated by the user. Determine the user's current context; Analyze search queries submitted by other users to identify spikes, popular or other patterns in the search queries submitted by other users in contexts similar to the current context of the user; Based on the peak, popular, or other patterns, select one or more search queries from the search queries submitted by other people; Search one or more online sources in response to one or more selected search queries; One or more of the processors generate unrequested content, which indicates information in response to one or more selected search queries; as well as The automated assistant merges the unrequested content into the existing human-computer dialogue session; Wherein, at least the merging is performed in response to determining that the automation assistant has responded to all natural language input received from the user during the human-computer dialogue session.
33. The at least one non-transitory computer-readable medium according to claim 32, wherein, The non-requested content includes non-requested natural language content.
34. A system comprising one or more processors and a memory operably coupled to said one or more processors, and one or more audio output devices operably coupled to said one or more processors, wherein, The memory stores instructions, which, in response to execution by the one or more processors, cause the one or more processors to operate the automation assistant to perform the following operations: Based on one or more signals, determine that the user is within the audible distance of the one or more audio output devices; A desire metric is determined that indicates the user wants to receive unrequested content, wherein the desire metric is determined based on one or more signals indicating the user’s current context, including position signals and / or accelerometer signals; Based at least in part on one or more characteristics of the user, identify information that the user may be interested in or one or more actions that the user may be interested in; Generate non-requested content indicating the information or one or more actions that may be of interest; and The non-requested content is incorporated into the audible human-computer dialogue session between the automated assistant and the user; Wherein, at least the merging is performed by the automation assistant in response to determining that the user is within a possible distance of the one or more audio output devices; Wherein, at least the merging is performed in response to determining that the desire metric satisfies one or more thresholds; The human-computer dialogue is initiated in response to determining that the user is within audible distance of the one or more audio output devices.
35. The system according to claim 34, wherein, The one or more signals are triggered by a user-operated computing device different from the system and received at one or more communication interfaces operatively coupled to the one or more processors.
36. The system according to claim 35, wherein, The one or more signals include the detection of the coexistence of the system and the computing device.
37. The system according to claim 34, wherein, The signal indicating the user's current context further includes the status of one or more applications.
38. The system according to claim 34, wherein, The signal indicating the user's current situation also includes an indication that the application is providing a video call.
39. The system according to claim 34, wherein, The information that the user may be interested in, or the one or more actions that the user may be interested in, are further identified based on one or more of the signals.
40. The system according to claim 34, wherein, One or more of the signals indicate the identity of the user, and the information that the user may be interested in or the one or more actions that the user may be interested in are identified at least in part based on the user's identity.
41. The system according to claim 34, wherein, The non-requested content includes non-requested natural language content.
42. A computer-implemented method, comprising: Based on one or more signals, determine that the user is within the audible distance of the one or more audio output devices; A desire metric is determined that indicates the user wants to receive unrequested content, wherein the desire metric is determined based on one or more signals indicating the user’s current context, including position signals and / or accelerometer signals; Based at least in part on one or more characteristics of the user, identify information that the user may be interested in or one or more actions that the user may be interested in; Generate non-requested content indicating the information or one or more actions that may be of interest; and The non-requested content is incorporated into the audible human-computer dialogue session between the automated assistant and the user; Wherein, at least the merging is performed by the automation assistant in response to determining that the user is within a possible distance of the one or more audio output devices; Wherein, at least the merging is performed in response to determining that the desire metric satisfies one or more thresholds; The human-computer dialogue is initiated in response to determining that the user is within audible distance of the one or more audio output devices.
43. The method according to claim 42, wherein, The signal indicating the user's current context further includes the status of one or more applications.
44. The method according to claim 42, wherein, The signal indicating the user's current situation also includes an indication that the application is providing a video call.
45. The method according to claim 42, wherein, The information that the user may be interested in, or the one or more actions that the user may be interested in, are further identified based on one or more of the signals.
46. The method according to claim 42, wherein, One or more of the signals indicate the identity of the user, and the information that the user may be interested in or the one or more actions that the user may be interested in are identified at least in part based on the user's identity.
47. A non-transitory computer-readable medium comprising instructions that, when executed by a computing device, cause the computing device to perform the following operations: Based on one or more signals, determine that the user is within the audible distance of the one or more audio output devices; Determine a desire metric that indicates the user's desire to receive unrequested content, wherein, The desire metric is determined based on one or more signals that indicate the user’s current context, including position signals and / or accelerometer signals. Based at least in part on one or more characteristics of the user, identify information that the user may be interested in or one or more actions that the user may be interested in; Generate non-requested content indicating the information or one or more actions that may be of interest to the individual; as well as The non-requested content is incorporated into the audible human-computer dialogue session between the automated assistant and the user; Wherein, at least the merging is performed by the automation assistant in response to determining that the user is within a possible distance of the one or more audio output devices; Wherein, at least the merging is performed in response to determining that the desire metric satisfies one or more thresholds; The human-computer dialogue is initiated in response to determining that the user is within audible distance of the one or more audio output devices.
48. The non-transitory computer-readable medium according to claim 47, wherein, The one or more signals are triggered by a user-operated computing device different from the system and received at one or more communication interfaces operatively coupled to the one or more processors.
49. The non-transitory computer-readable medium according to claim 48, wherein, The one or more signals include the detection of the coexistence of the system and the computing device.
50. The non-transitory computer-readable medium according to claim 47, wherein, The signal indicating the user's current context further includes the status of one or more applications.
51. The non-transitory computer-readable medium according to claim 47, wherein, The signal indicating the user's current situation also includes an indication that the application is providing a video call.
52. The non-transitory computer-readable medium according to claim 47, wherein, The information that the user may be interested in, or the one or more actions that the user may be interested in, are further identified based on one or more of the signals.
53. The non-transitory computer-readable medium according to claim 47, wherein, One or more of the signals indicate the identity of the user, and the information that the user may be interested in or the one or more actions that the user may be interested in are identified at least in part based on the user's identity.