Proactive Incorporation of Non-Requestable Content into Human-Computer Dialogs
By enabling automated assistants to proactively incorporate relevant non-requested content into human-computer dialog sessions, the inefficiencies of reactive systems are addressed, improving interaction efficiency and user engagement.
Patent Information
- Application Number
- JP2024001367
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-11-29
- Filing Date
- 2024-01-09
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2038-04-30
AI Technical Summary
Existing automatic assistants are reactive, requiring explicit user input to provide substantial information or initiate tasks, leading to inefficiencies and missed opportunities to offer relevant content.
Configuring an automated assistant to proactively incorporate non-requested content of potential interest into human-computer dialog sessions, using entity data and user profiles to determine relevance and interest criteria.
Enhances dialog efficiency by reducing the need for explicit user requests, providing potentially useful or interesting information that users might not otherwise seek, and improving the interaction experience by making the assistant appear more natural and engaging.
Smart Images

Figure 0007690620000001 
Figure 0007690620000002 
Figure 0007690620000003
Abstract
Description
Technical Field
[0001] The present invention relates to the proactive incorporation of unsolicited content into human-computer dialogs.
Background Art
[0002] Humans can engage in human-computer dialogs using a dialog software application called an "automatic assistant" (which may also be referred to herein as a "chatbot", "conversational personal assistant", "intelligent personal assistant", "conversational agent", etc.). For example, a human (who may be referred to as a "user" when interacting with the automatic assistant) can provide commands and / or requests by using spoken natural language input (i.e., utterances) that can be converted to text and then processed, and / or by providing text (e.g., typed) natural language input. Automatic assistants are generally reactive, as opposed to proactive, which can lead to inefficiencies when providing available services. For example, an automatic assistant does not proactively obtain and provide specific information that may be potentially interesting to the user. Thus, the user must provide an initial natural language input (e.g., spoken or typed) before the automatic assistant will respond with substantial information and / or before it will initiate one or more tasks on behalf of the user.
Summary of the Invention
Means for Solving the Problems
[0003] This specification describes techniques for configuring an automated assistant to proactively incorporate non-requested content that may be potentially interesting to a user within a human-computer dialog session. In some implementations, an automated assistant configured in accordance with selected aspects of the present disclosure, and / or one or more other components that work in conjunction with the automated assistant, can perform such an incorporation in an existing human-computer dialog session when it is determined that the automated assistant has effectively fulfilled its obligation to the user (e.g., the automated assistant is waiting for further instructions). This can be as simple as the user saying "good morning" and the automated assistant providing a common response "good morning". In such a scenario, the user may still be (at least briefly) involved in the human-computer dialog session (e.g., the chatbot screen showing the transcript of the ongoing human-computer dialog may still be open, and the user may still be within earshot of the audio output device through which the human-computer dialog is implemented, etc.). Thus, any non-requested content incorporated within the human-computer dialog session is likely to be consumed by the user.
[0004] In some implementations, the non-requested content can be presented as observations (e.g., natural language descriptions) that do not directly respond to the user's query (e.g., an aside, incidentally, etc.). In some such implementations, the non-requested content may be prefaced with an appropriate description that identifies the content as something that the user did not specifically request but may, in some cases, be relevant to the current conversation context, such as "by the way...", "did you know...", "incidentally...". Although not directly responding to the user's query, the non-requested content may in some cases be tangential to the conversation context in various ways.
[0005] Unsolicited (e.g., off-topic) content that is potentially interesting to the user and is incorporated within a human-computer dialog can be selected in various ways. In some implementations, the content may take the form of one or more facts selected based on one or more entities mentioned in the conversation by the user and / or an automated assistant. For example, if the user or an automated assistant mentions a particular celebrity, the automated assistant can incorporate one or more facts about that celebrity, about another similar celebrity, etc. Facts about or in some cases related to an entity can appear in various forms, such as recent news items about that entity (e.g., "Did you know that <entity> turned 57 last week?", "By the way, <entity> is coming to your town next month", "By the way, there has been a recall of a new product from <entity>", etc.), general facts (e.g., birthday, political party, net worth, topics, spouse, family, awards, price, availability, etc.).
[0006] In various implementations, multiple facts about an entity mentioned in a human-computer dialog may be determined and ranked, for example, based on a so-called "criterion potentially interesting to the user", and the fact ranked highest can be presented to the user. The criterion for facts of potential interest can be determined in various ways based on various signals.
[0007] In some implementations, the criteria for potential interest in specific facts about an entity can be determined based on data related to the user's unique user profile. The user profile can be associated with a user account used by the user when operating one or more client devices (e.g., forming an adjusted "ecosystem" of client devices related to the user). Various data can be associated with the user's profile, such as search history (including patterns detectable from the search history), messaging history (e.g., including past human-computer dialogs between the user and an automated assistant), personal preferences, browsing history, sensor signals (e.g., global positioning system, i.e., "GPS", location), etc.
[0008] When a user searches for a celebrity (or other public figure), it is assumed that the user also tends to search for the political party to which those celebrities belong. Based on the detection of this search pattern, the automated assistant can assign a relatively large criterion of potential interest to the political party to which a celebrity (or other public figure) mentioned in the current human-computer dialog between the automated assistant and the user belongs. In other words, the automated assistant can generalize the user's specific interest in the political party of a particular celebrity searched by the user to the political parties of all celebrities.
[0009] As another example, when researching flights (e.g., by interacting with an automated assistant and / or by operating a web browser), assume that the user tends to search for flights from the nearest airport, as well as flights departing from different airports that are somewhat farther away, for example, to compare prices. This search pattern can be detected by the automated assistant. Assume that the user later asks the automated assistant to provide the price for a flight from the nearest airport to a specific destination. By default, the automated assistant can directly respond to the user's query by providing the price for a flight from the nearest airport. However, based on the detected search pattern, the automated assistant can provide unsolicited content that includes one or more prices for flights from different airports that are somewhat farther away to the destination.
[0010] In some implementations, the criteria for potential interest in specific facts about the entities mentioned (or related entities) can be determined based on aggregated data generated by multiple users, who may or may not be members of the user, or based on the behavior of those multiple users. For example, assume that generally, after searching for an entity itself, a user tends to search for specific facts about that entity. The automated assistant can detect and use that aggregation pattern to determine the criteria for potential interest in facts about the entity currently being discussed in a human-computer dialog with an individual user.
[0011] In some implementations, one or more corpora of "online conversations" among people (e.g., message exchange threads, comment threads, etc.) can be analyzed to detect, for example, references to entities and references to facts about the entities being referred to. For example, assume that an entity (e.g., a person) is being discussed in a comment thread, and participants mention a particular fact about that entity (e.g., the entity's political party) within a particular proximity of the reference to the entity (e.g., within the same thread, within x days of it, within the same forum, etc.). In some implementations, when that entity is discussed in a subsequent human-computer dialogue, the fact about that entity (which may, in some cases, be confirmed against a knowledge graph including the entity and the verified fact) can potentially be flagged as being of interest, or in some cases, that fact can be indicated (e.g., within the same knowledge graph). If multiple participants tend to mention the same fact in multiple different online conversations when the same entity is being referred to, the criteria for potentially interesting facts can be further enhanced.
[0012] Similar to the above example, correspondingly, the relationship between a particular fact and a particular entity can be generalized so that criteria for potential interest can be assigned to similar facts about different entities. For example, assume that online conversation participants frequently mention, for example, as an aside, the net worth of a celebrity being discussed. The concept that the personal assets of a celebrity are frequently mentioned as an aside or incidentally when discussing the celebrity can be used to determine the criteria for potential interest in the net worth of the celebrity (or other public figure) currently being discussed in a human-computer dialogue between a user and an automated assistant.
[0013] The unsolicited content discussed in this specification is not limited to facts about specific entities mentioned in a human-computer dialog between a user and an automated assistant. In some implementations, other entities related to the mentioned entity may also be considered. These other entities can include, for example, entities that share one or more attributes with the mentioned entity (such as a musician, artist, actor, athlete, restaurant, point of interest, etc.), or entities that share one or more attributes (such as being located nearby, having temporally approximate events, etc.). For example, assume that a user asks an automated assistant a question about a specific actor. After answering the user's question, the automated assistant can incorporate unsolicited facts about another actor (or a director with whom the actor has worked, etc.) into the conversation. As another example, assume that a user engages an automated assistant 120 to make a reservation at a restaurant. The automated assistant can proactively recommend another nearby restaurant as an alternative, for example, because it is less expensive, better rated, less likely to be crowded, etc. As yet another example, assume that a user likes two musicians (which can be determined, for example, from a playlist related to the user's profile or from an aggregated user playback history / playlist). Further assume that the user asks when one of the two musicians will be touring nearby soon. After answering the user's question, the automated assistant can determine that the other musician will be touring nearby soon and can incorporate that unsolicited fact into the human-computer dialog.
[0014] Incorporating unsolicited content that may potentially be of interest to a user within a human-computer dialog session can have several technical advantages. By presenting unsolicited content without the need for direct commands from the user, it is possible to achieve a reduction in additional user requests to access such content, thereby enhancing the efficiency of a given dialog session and reducing the load on the resources required to interpret multiple user requests. For example, the user can be liberated from having to affirmatively request information, which can save computing resources that would otherwise be used to process the user's natural language input. Moreover, an automated assistant can thus provide functionality that would otherwise be difficult to discover or access, improving the effectiveness of the interaction with the user. The automated assistant can be made to appear more "natural" or "human" to the user, which can encourage an increase in interaction with the automated assistant. Additionally, the incorporated content can be selected based on one or more entities (e.g., people, places, knowledge graphs, etc., things documented in one or more databases) mentioned by the user or the automated assistant during the human-computer dialog, so the incorporated content may be of interest to the user. In addition, the user can receive potentially useful or interesting information that the user might not otherwise have thought to request.
[0015] In some implementations, based on the content of an existing human-computer dialog session between a user and an automated assistant, identifying entities mentioned by the user or the automated assistant; based on entity data contained within one or more databases, identifying one or more facts about the entity or another entity related to the entity; for each of the one or more facts, determining a corresponding criterion potentially of interest to the user; generating unsolicited natural language content, the unsolicited natural language content including one or more of the facts selected based on the corresponding one or more criteria potentially of interest; and the automated assistant incorporating the unsolicited natural language content into the existing human-computer dialog session or a subsequent human-computer dialog session. A method is provided that is executed by one or more processors and includes these steps.
[0016] These and other implementations of the technology disclosed herein may optionally include one or more of the following features.
[0017] In various implementations, the step of determining corresponding criteria of potential interest may be based on data related to a user profile associated with a user. In various implementations, data related to a user profile may include a search history associated with the user. In various implementations, the step of determining corresponding criteria of potential interest may be based on an analysis of a corpus of one or more online conversations among one or more people. In various implementations, this analysis may include detecting one or more references to one or more entities, as well as detecting one or more references to facts regarding one or more entities within a particular proximity of one or more references to one or more entities. In various implementations, one or more entities may include entities mentioned by a user or an automated assistant. In various implementations, one or more entities may share an entity class with an entity mentioned by a user or an automated assistant. In various implementations, one or more entities may share one or more attributes with an entity mentioned by a user or an automated assistant.
[0018] In various implementations, the step of determining corresponding criteria of potential interest may include detecting that a given fact among one or more facts has been previously referenced in an existing human-computer dialog between a user and an automated assistant, or in a previous human-computer dialog between the user and the automated assistant. In various implementations, the potentially interesting criteria determined for a given fact may reflect the detection.
[0019] In various implementations, the method may further include excluding a given fact among one or more facts from consideration based on detecting that the given fact has been previously referenced in an existing human-computer dialog between a user and an automated assistant, or in a previous human-computer dialog between the user and the automated assistant.
[0020] In various implementation forms, the entity may be a location, and a given fact among one or more facts may be an event that occurs at or near that location. In various implementation forms, the entity may be a person, and a given fact among one or more facts may be the current event involving that person.
[0021] In some implementations, based on the content of an existing human-computer dialog session between a user and an automated assistant, identifying entities mentioned by the user or the automated assistant; determining that the automated assistant has fulfilled one or more outstanding obligations of the automated assistant to the user; based on entity data included in one or more databases, identifying one or more facts about the entity, wherein the entity is associated with an entity class in one or more of the databases; for each of the one or more facts, determining a corresponding criterion potentially of interest to the user, wherein the corresponding criterion potentially of interest to each fact is determined based on the relationship between the entity and the fact and the interest of other users in the same relationship between the entity and other entities sharing the same entity class and their respective facts; generating unsolicited natural language content, wherein the unsolicited natural language content includes one or more of the facts selected based on the corresponding one or more criteria potentially of interest; and after determining that the automated assistant has fulfilled one or more outstanding obligations of the automated assistant to the user, the automated assistant incorporating the unsolicited natural language content into the existing human-computer dialog session, wherein the incorporating step automatically outputs the unsolicited natural language content to the user as part of the existing human-computer dialog session. A method is provided that is executed by one or more processors and includes the foregoing steps.
[0022] In some implementations, a method performed by one or more processors includes: identifying an entity based on the state of a media application or game application operating on a first client device that a user is operating, where the entity is identified without using explicit input from the user; determining that an automated assistant operating on a second client device associated with the user has no outstanding obligations to the user; identifying one or more facts about the entity based on entity data included in one or more databases; determining, for each of the one or more facts, a corresponding criterion that may be of potential interest to the user; generating unsolicited natural language content that includes one or more of the facts selected based on the corresponding criterion that may be of potential interest; and after determining that the automated assistant has no outstanding obligations to the user, the automated assistant incorporating the unsolicited natural language content into a new or existing human-computer dialog session, where incorporating includes automatically outputting the unsolicited natural language content to the user as part of the new or existing human-computer dialog session.
[0023] These and other implementations of the technology disclosed herein may optionally include one or more of the following features.
[0024] In various implementations, the step of determining corresponding criteria of potential interest can be based on an analysis of a corpus of one or more online conversations among one or more people. In various implementations, the analysis can include detecting one or more references to one or more entities, as well as one or more references to facts regarding one or more entities within a particular proximity of the one or more references to the one or more entities. In various implementations, the one or more entities can include entities determined based on the state of a media application or a game application. In various implementations, the one or more entities can share an entity class with an entity determined based on the state of a media application or a game application. In various implementations, the one or more entities can share one or more attributes with an entity determined based on the state of a media application or a game application.
[0025] In various implementations, an entity can be related to a video game and a given one of the one or more facts is a fact regarding the video game.
[0026] In addition, some implementations include one or more processors of one or more computing devices, the one or more processors being operable to execute instructions stored in an associated memory, the instructions being configured to cause execution of any of the methods described above. Some implementations also include one or more non-transitory computer-readable storage media storing computer instructions executable by the one or more processors to execute any of the methods described above.
[0027] It should be understood that all combinations of the foregoing concepts and additional concepts described in more detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter that appear at the end of this disclosure are intended to be part of the subject matter disclosed herein.
Brief Description of the Drawings
[0028]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Best Mode for Carrying Out the Invention
[0029] Next, referring to FIG. 1, an exemplary environment in which the techniques disclosed herein may be implemented is shown. The exemplary environment includes a plurality of client computing devices 106 1~N and an automatic assistant 120. The automatic assistant 120 is shown separately from the client computing device 106 in FIG. 1 1~N but in some implementations, all or aspects of the automatic assistant 120 may be implemented by one or more of the client computing devices 106 1~N . For example, a client device 106 1 may implement one instance of one or more aspects of the automatic assistant 120, and a client device 106 N may also implement a separate instance of those one or more aspects of the automatic assistant 120. In implementations where one or more aspects of the automatic assistant 120 are implemented by one or more computing devices remote from the client computing device 106 1-N , the client computing device 106 1-N and those aspects of the automatic assistant 120 may communicate via one or more networks such as a local area network (LAN) and / or a wide area network (WAN) (e.g., the Internet).
[0030] Client device 106 1~Nmay include one or more of, for example, a desktop computing device, a laptop computing device, a tablet computing device, a mobile phone computing device, a computing device of the user's vehicle (e.g., an in-vehicle communication system, an in-vehicle entertainment system, a vehicle navigation system), a stand-alone interactive speaker, and / or a wearable device of the user that includes a computing device (e.g., the user's watch having a computing device, the user's glasses having a computing device, a virtual reality computing device or an augmented reality computing device). Additional and / or alternative client computing devices may be provided. In some implementations, a given user may communicate with the automatic assistant 120 using a plurality of client computing devices that collectively form an adjusted "ecosystem" of computing devices. In some such implementations, the automatic assistant 120 may be considered to "serve" that given user, e.g., by giving the automatic assistant 120 extended access to resources (e.g., content, documents, etc.) that are controlled by the user to whom the access is "served". However, for the sake of brevity, some of the examples described herein will focus on a user operating a single client computing device 106.
[0031] Client computing device 106 1~N Each of which can operate various different applications, such as a corresponding one of a plurality of message exchange clients 107 1~N The message exchange client 107 1~N can appear in various forms, these forms may vary depending on the client computing device 106 1~N and / or multiple forms may vary depending on the client computing device 106 1~Nmay be operated on a single client computing device among them. In some implementations, the message exchange client 107 1~N one or more of which may appear in the form of a Short Message Service (「SMS」) and / or Multimedia Messaging Service (「MMS」) client, an online chat client (e.g., an instant messenger, Internet Relay Chat, or 「IRC」, etc.), a messaging application related to a social network, a personal assistant messaging service dedicated to conversations with the automatic assistant 120, etc. In some implementations, the message exchange client 107 1~N one or more of which may be implemented via a web browser (not shown) or a web page or other resource rendered by another application of the client computing device 106.
[0032] As will be described in more detail herein, the automatic assistant 120 participates in a human-computer dialog session with one or more users via the user interface input device and the user interface output device of one or more client devices 106. In some implementations, the automatic assistant 120 can participate in a human-computer dialog session with a user in response to a user interface input provided by the user via one or more user interface input devices of one of the client devices 106. In some of these implementations, the user interface input explicitly targets the automatic assistant 120. For example, the message exchange client 107 1~N one or more of which may be implemented via a web browser (not shown) or a web page or other resource rendered by another application of the client computing device 106. 1~N one or more of which may be implemented via a web browser (not shown) or a web page or other resource rendered by another application of the client computing device 106. 1~NOne of them may be a personal assistant messaging service dedicated to conversations with the automated assistant 120, and user interface inputs provided via the personal assistant messaging service may be automatically provided to the automated assistant 120. Also, for example, the user interface input may be, explicitly, based on a specific user interface input indicating that the automated assistant 120 should be launched, the messaging exchange client 107 1~N may explicitly target one or more of the automated assistants 120 among them. For example, the specific user interface input may be one or more typed characters (e.g., @AutomatedAssistant), user interaction using a hardware button and / or a virtual button (e.g., tap, long tap), a verbal command (e.g., "Hey, automated assistant"), and / or other specific user interface inputs. In some implementations, the automated assistant 120 can participate in a dialog session in response to a user interface input even when the user interface input does not explicitly target the automated assistant 120. For example, the automated assistant 120 can inspect the content of the user interface input and participate in a dialog session in response to the presence of some terms within the user interface input and / or based on other cues. In many implementations, the automated assistant 120 can participate in an interactive voice response ("IVR") so that the user can issue commands, searches, etc., and the automated assistant can utilize natural language processing and / or one or more grammars to convert the utterance into text and respond to the text accordingly.
[0033] Client computing device 106 1~NEach of and the automatic assistant 120 may include one or more memories for storing data and software applications, one or more processors for accessing the data and executing the applications, and other components for facilitating communication over a network. The client computing device 106 1~N The operations performed by one or more of and / or by the automatic assistant 120 may be distributed across multiple computer systems. The automatic assistant 120 may be implemented, for example, as a computer program running on one or more computers within one or more locations coupled to each other through a network.
[0034] The automatic assistant 120 may include a natural language processor 122 and a responsive content engine 130. In some implementations, one or more of the engines and / or modules of the automatic assistant 120 may be omitted, combined, and / or implemented within components separate from the automatic assistant 120. The automatic assistant 120 may participate in a human-computer dialog session with one or more users via the associated client device 106 1~N and provide responsive content generated and / or maintained by the responsive content engine 130.
[0035] In some implementations, the responsive content engine 130 generates responsive content in response to various inputs generated by a user of one of the client devices 106 during a human-computer dialog session with the automatic assistant 120. The responsive content engine 130 provides the responsive content (e.g., via one or more networks when separate from the user's client device) for presentation to the user as part of the dialog session. For example, the responsive content engine 130 may the client device 106 1~N The operations performed by one or more of and / or by the automatic assistant 120 may be distributed across multiple computer systems. The automatic assistant 120 may be implemented, for example, as a computer program running on one or more computers within one or more locations coupled to each other through a network. 1~NResponsive content can be generated in response to free-form natural language input provided via one of them. The free-form input used herein is input that is constructed by the user and not restricted to a group of options presented for user selection.
[0036] As used herein, a "dialog session" can include one or more logically independent exchanges of messages between a user and the automated assistant 120 (and, in some cases, other human participants within the thread). The automated assistant 120 can distinguish multiple dialog sessions with a user based on various signals, such as the passage of time between sessions, changes in the user context between sessions (e.g., location, before / during / after a scheduled meeting, etc.), detection of one or more intervening conversations between the user and the client device other than the dialog between the user and the automated assistant (e.g., the user switches applications for a period of time, the user leaves a stand-alone voice-activated product and then returns later), locking / sleeping of the client device between sessions, changes in the client device used to interact with one or more instances of the automated assistant 120, etc.
[0037] In some implementations, when the automated assistant 120 provides a prompt requesting user feedback, the automated assistant 120 can preemptively activate one or more components of the client device (through which the prompt is provided) that are configured to process the user interface input to be received in response to the prompt. For example, if the user interface input is to be provided via the microphone of the client device 106 1 the automated assistant 120 causes the microphone to "open" preemptively (thereby eliminating the need to hit an interface element or speak a "hot word" to open the microphone), and the client device 106 1The client device 106 for the text processor that pre-emptively activates local voice for the text processor 1 Pre-emptively establish a communication session between the remote voice and / or provide one or more commands for rendering on the client device 106 a graphical user interface (e.g., an interface including one or more selectable elements that can be selected to provide feedback). This allows user interface inputs to be provided and / or processed more quickly than if these components were not pre-emptively activated. 1 This may be provided. This allows user interface inputs to be provided and / or processed more quickly than if these components were not pre-emptively activated.
[0038] The natural language processor 122 of the automatic assistant 120 can process natural language input generated by the user via the client device 106 and generate an annotated output for use by one or more other components of the automatic assistant 120, such as the responsive content engine 130. For example, the natural language processor 122 can process natural language free-form input generated by the user via one or more user interface input devices of the client device 106. The generated, annotated output includes one or more annotations of the natural language input and, in some cases, one or more (e.g., all) of the terms of the natural language input. 1~N For example, the natural language processor 122 can process natural language free-form input generated by the user via one or more user interface input devices of the client device 106. The generated, annotated output includes one or more annotations of the natural language input and, in some cases, one or more (e.g., all) of the terms of the natural language input. 1 The generated, annotated output includes one or more annotations of the natural language input and, in some cases, one or more (e.g., all) of the terms of the natural language input.
[0039] In some implementations, the natural language processor 122 is configured to identify and annotate various types of grammatical information within the natural language input. For example, the natural language processor 122 may include a portion of a speech tagger configured to annotate terms using their grammatical roles. For example, a portion of the speech tagger can tag each term in that portion of the speech, such as "noun", "verb", "adjective", "pronoun", etc. Also, for example, in some implementations, the natural language processor 122 may include, additionally and / or alternatively, a dependency parser (not shown) configured to determine syntactic relationships between terms within the natural language input. For example, the dependency parser can determine which terms modify other terms, the subject and verb of the sentence, etc. (e.g., a parse tree), and can annotate such dependencies.
[0040] In some implementations, the natural language processor 122 may additionally and / or alternatively include an entity tagger (not shown) configured to annotate entity references within one or more segments, such as references to people (including, for example, literary characters, celebrities, public figures, etc.), organizations, locations (real and imaginary), etc. In some implementations, data about entities may be stored in one or more databases, such as the knowledge graph 124. In some implementations, the knowledge graph 124 may include nodes representing known entities (and, in some cases, entity attributes), as well as edges connecting the nodes and representing relationships between the entities. For example, a "banana" node may be connected (e.g., as a child) to a "fruit" node, which may in turn be connected (e.g., as a child) to "agricultural product" and / or "food" nodes. As another example, a restaurant called "Hypothetical Cafe" may be represented by nodes that also include attributes such as its address, the types of food served, business hours, contact information, etc. The "Hypothetical Cafe" node may, in some implementations, be connected by edges (e.g., representing a child-to-parent relationship) to one or more other nodes, such as a "restaurant" node, a "business" node, and nodes representing the city and / or state in which the restaurant is located.
[0041] The entity tagger of the natural language processor 122 can annotate references to entities at a high level of granularity (e.g., to enable identification of all references to an entity class, such as people) and / or at a low level of granularity (e.g., to enable identification of all references to a specific entity, such as a particular person). The entity tagger may depend on the content of the natural language input to resolve a particular entity and / or may, in some cases, communicate with the knowledge graph or other entity databases to resolve a particular entity.
[0042] In some implementations, the natural language processor 122 may additionally and / or alternatively include a coreference resolver (not shown) configured to group or "cluster" references to the same entity based on one or more context cues. For example, the coreference resolver may be utilized to resolve the terms from "there" to "Hypothetical Cafe" in the natural language input "I liked Hypothetical Cafe where I had a meal last time."
[0043] In some implementations, one or more components of the natural language processor 122 may depend on annotations from one or more other components of the natural language processor 122. For example, in some implementations, the named entity tagger may depend on annotations from the coreference resolver and / or the dependency parser when annotating all mentions of a particular entity. Also, for example, in some implementations, the coreference resolver may depend on annotations from the dependency parser when clustering references to the same entity. In some implementations, when processing a particular natural language input, one or more components of the natural language processor 122 can use related previous inputs and / or other related data external to the particular natural language input to determine one or more annotations.
[0044] As described above, the automatic assistant 120 can utilize one or more resources, for example, by the responsive content engine 130, to generate suggestions and / or other non-requested content for providing during a human-computer dialog session with one of the N users of the client device 106 1~ In various implementations, the responsive content engine 130 may include an action module 132, an entity module 134, and a proactive content module 136.
[0045] The action module 132 of the responsive content engine 130 acts on the client computing device 106 1~N and / or annotations of the natural language input provided by the natural language processor 122 to determine at least one action responsive to the natural language input. In some implementations, the action module 132 can determine the action based on one or more terms included in the natural language input. For example, the action module 132 can determine the action based on a mapping, in one or more computer-readable media, of the action to one or more terms included in the natural language input. For example, the action of "add <item> to my shopping list" can be mapped to one or more terms, such as "I need <item> from the market...", "I need to get <item>", "We're out of <item>", etc.
[0046] The entity module 134 may be configured to identify entities mentioned by a user or the automated assistant 120 based on the content (and / or annotations thereof) of an existing human-computer dialogue session between the user and the automated assistant 120. This content may include input provided by one or more users via user interface input devices during the human-computer dialogue session between the user and the automated assistant 120, as well as content incorporated within the dialogue session by the automated assistant 120. The entity module 134 may utilize one or more resources in identifying the referenced entity and / or in refining the candidate entities. For example, the entity module 134 may utilize the natural language input itself, annotations provided by the natural language processor 122, and / or information from the knowledge graph 124. In some cases, the entity module 134 may be integral with, e.g., the same as, the aforementioned entity tagger that forms part of the natural language processor 122.
[0047] The proactive content module 136 can be configured to proactively incorporate unsolicited content of potential interest to the user into the human-computer dialogue session. In some implementations, the proactive content module 136 can determine, based on data received from other modules, such as, for example, the natural language processor 122, the behavior module 132, and / or the entity module 134, that in an existing human-computer dialogue session between the user and the automated assistant 120, the automated assistant 120 has responded to all natural language inputs received from the user during the human-computer dialogue session. Assume that the user operates the client device 106 to request a search for specific information, and the automated assistant 120 performs (or causes a search to be performed) the search and returns response information as part of the human-computer dialogue. At this point, unless the user has also requested other information, the automated assistant 120 has sufficiently responded to the user's request. In some implementations, the proactive content module 136 can wait for some pre-determined time interval (e.g., 2 seconds, 5 seconds, etc.) for the automated assistant 120 to receive additional user input. If none is received during that time interval, the proactive content module 136 may determine that it has responded to all natural language input received from the user during the human-computer dialogue session and is now free to incorporate unsolicited content.
[0048] Based on one or more entities identified by the entity module 134 as being mentioned (or relating to) in the human-to-computer dialog session, the proactive content module 136 may be configured to identify one or more facts about the entity or about another entity related to the entity based on entity data contained in one or more databases (e.g., the knowledge graph 124). The proactive content module 136 may then determine corresponding criteria of potential interest to the user for each of the one or more facts. Based on the one or more criteria of potential interest corresponding to the one or more facts, the proactive content module 136 may select one or more of the facts to be included in the unsolicited natural language content that it generates. The proactive content module 136 may then incorporate the unsolicited natural language content into an existing human-to-computer dialog session or a subsequent human-to-computer dialog session.
[0049] In some implementations, the criteria of potential interest for facts about an entity may be determined, for example, by the proactive content module 136, based on data obtained from one or more user profile databases 126. The data contained in the user profile database 126 may relate to user profiles associated with human participants in a human-to-computer dialogue. In some implementations, the user profiles may each be associated with a user account used by the user when operating one or more client devices. Various data may be associated with the user's profile (and thus stored in the user profile database 126), such as search history (including patterns detectable from the search history), messaging history (including, for example, past human-to-computer dialogues between the user and an automated assistant), personal preferences, browsing history, etc. Other information related to an individual user profile may also be stored in the user profile database 126 or may be determined based on data stored in the user profile database 126. This other user-related information may include, for example, topics of interest to the user (which may be stored directly within database 126 or determined from other data stored therein), search history, browsing history, user setting preferences, current / past location, media play history, travel history, content of past human-computer dialogue sessions, etc.
[0050] Thus, in some implementations, the proactive content module 136 may have access to various signals or other data from one or more client devices 106 operated by the user, e.g., directly from the client device 106, directly from the user profile database 126, and / or indirectly via one or more computing systems operating as a so-called "cloud." Topics of interest to a user may include, for example, a particular hobby (e.g., golf, skiing, games, painting, etc.), literature, movies, music genres, particular entities (e.g., artists, athletes, sports teams, companies), etc. Other information that may be relevant to a user's profile may include, for example, age, the user's scheduled events (e.g., determined from one or more calendar entries), etc.
[0051] In various implementations, the proactive content module 136 may be configured to generate unsolicited content that indicates (e.g., includes) facts of potential interest to the user and incorporate the unsolicited content into the human-to-computer dialogue. This unsolicited content may appear in various forms that may be incorporated into an existing human-to-computer dialogue session. For example, in some implementations where a user is interacting with the automated assistant 120 using a text-based messaging client 107, the unsolicited content generated by the proactive content module 136 may take the form of text, images, videos, or any combination thereof that may be incorporated into a transcript of the human-to-computer dialogue rendered by the messaging client 107. In some implementations, the unsolicited content may include or take the form of so-called "deep links" that are selectable by the user to expose a different application interface to the user. For example, a deep link, when selected by the user, may cause the client device 106 to launch (or run) a particular application in a particular state. In other implementations where the user is interacting with the automated assistant 120 using a voice interface (e.g., when the automated assistant 120 operates on a standalone interactive speaker or on an in-vehicle system), the unsolicited content may take the form of natural language output that is audibly provided to the user. As noted above, in many cases the unsolicited content may be prefaced with language, such as "by the way," "did you know," "as an aside," etc.
[0052] In some implementations, the incorporation of unsolicited content can be performed in response to a determination by, for example, the proactive content module 136 that the auto assistant 120 has responded to all natural language inputs received from the user during a human-computer dialog session. In some implementations, one or more of the other operations described above with respect to the proactive content module 136 can also be performed in response to such an event. Or, in some implementations, those operations can be performed periodically or continuously by the proactive content module 136 such that the proactive content module 136 (and thus the auto assistant 120) remains in a "ready" state to quickly incorporate unsolicited content potentially of interest to the user into an existing human-computer dialog session.
[0053] The proactive content module 136 may also have access to (e.g., obtain facts from) components other than the user's profile, such as a fresh content module 138 and one or more miscellaneous domain modules 140. The fresh content module 138 may provide the proactive content module 136 with access to data regarding current events, news, current schedules (e.g., tour dates for performers), current prices (e.g., for goods or services), trending news / searches (e.g., indicated by so-called "hashtags"), and the like. In some implementations, the fresh content module 138 may be part of a larger search engine system and may be configured to return temporally relevant search results. For example, some search engine interfaces include a "news" filter where a user may select to limit search results to information published by various news sources. The fresh content module 138 may have access to such information. The miscellaneous domain module 140 may provide data from a variety of other domains and may therefore operate similarly to other search engine filters. For example, a "weather" domain module can return facts about weather, a "history" domain module can return data about historical facts, a "topical" module can return random facts about entities, etc. In some implementations, a "manual" fact module can be configured to receive manually entered facts, for example from a paid advertising company, along with indications of the entities for which those facts apply.
[0054] In some implementations, when the entity module 134 identifies one or more entities mentioned during the human-computer dialogue, the proactive content module 136 can derive various facts about those one or more entities from one or more sources, such as the knowledge graph 124, or one or more modules 138-140. The proactive content module 136 can then rank the returned facts, for example, by determining the aforementioned criteria of user interest related to the facts.
[0055] The measure of potential user interest in a fact may be determined by the proactive content module 136 in a variety of ways based on a variety of information. As mentioned above, in some implementations, the proactive content module 136 may determine the measure of potential user interest in a fact based on individual user information contained, for example, in the user profile database 126. For example, if a particular user tends to search for upcoming tour dates for a musician whenever they also search for information about that musician, any facts related to upcoming tour dates may be assigned a relatively high measure of potential user interest. In some implementations, facts related to tour dates that are close to the user's location (e.g., as determined using GPS) may be assigned a higher measure of potential interest than facts related to tour dates that are farther away.
[0056] In some implementations, the proactive content module 136 can generate a measure of potential user interest for facts based on aggregate data related to searches and / or behavior of a population of users. If users in general tend to search for upcoming tour dates when searching for musicians, then for similar reasons as above, a relatively large measure of potential user interest may be assigned to facts about upcoming tour dates, particularly tour dates close to the user, when musicians are mentioned in the human-to-computer dialogue. If users in general tend to search for rental cars at the same time as searching for flight reservations, then a relatively large measure of potential user interest may be assigned to facts about rental cars (e.g., price, availability) when one or more flights are mentioned in the human-to-computer dialogue. If participants in an online conversation frequently mention the reliability of a particular product when discussing that product, then a relatively large measure of potential user interest may be assigned to facts about the reliability of that product when that product is mentioned in the human-to-computer dialogue, and so on.
[0057] Figure 2 illustrates an example of a human-computer dialog session between a user 101 and an instance of an automated assistant (120 in Figure 1, not shown in Figure 2). Figures 2-5 illustrate an example of a dialog session that may occur between a user 101 of a computing device 210 (shown as a standalone interactive speaker, but this is not meant to be limiting) and the automated assistant 120 via a microphone and speaker according to implementations described herein. One or more aspects of the automated assistant 120 may be implemented on the computing device 210 and / or on one or more computing devices in network communication with the computing device 210.
[0058] In FIG. 2, user 101 provides natural language input 280 of "How many concertos did Mozart compose?" in a human-to-computer dialogue session between user 101 and automated assistant 120. In response to natural language input 280, automated assistant 120 provides responsive natural language output 282 of "Mozart's concertos for piano and orchestra are numbers 1 through 27." Rather than waiting for additional user input, automated assistant 120 can proactively incorporate (e.g., via proactive content module 136) additional content of potential interest to the user into the human-to-computer dialogue of FIG. 2. For example, automated assistant 120 can obtain / receive one or more facts about the entity "Mozart" from various sources (e.g., 124, 138-140). The automated assistant 120 proactively incorporates the following unsolicited content (shown in italics in FIG. 2 and other figures) into the human-to-computer dialogue: "By the way, did you know that the Louisville Orchestra will be performing a concert next month featuring works composed by Mozart?" Other facts that may have been presented include, but are not limited to, Mozart's date of birth, birthplace, date of death, etc. In this example, the upcoming performance may have been selected, for example, by the proactive content module 136, because the user 101 tends to search for opportunities to see musical compositions performed by musicians that he searches for.
[0059] 3 illustrates another exemplary dialog between user 101 and automated assistant 120 running on computing device 210 during different sessions. At 380, user 101 speaks the phrase, "What is the name of my hotel in Louisville?" After searching various data sources (e.g., calendar entries, travel itinerary emails, etc.) related to user profile of user 101, automated assistant 120 responds at 382, "You're staying at the Brown Hotel." Automated assistant 120 can then determine (e.g., via proactive content module 136) that it has fully responded to the user's natural language input. Accordingly, based on user 101's interest in the flying troupe known as the "Blue Angels," and a determination that the Blue Angels will soon be performing in Louisville (which is a location entity), automated assistant 120 can proactively incorporate the following unsolicited content into the human-to-computer dialogue: "By the way, while you're in Louisville, the Blue Angels will be performing. Want me to find tickets?" If user 101 responds affirmatively, automated assistant 120 can engage user 101 in additional dialogue to, for example, obtain tickets to the event.
[0060] 4 illustrates another exemplary dialogue between a user 101 and an automated assistant 120 running on a computing device 210 during different sessions. In this example, at 480, the user 101 issues a command to provide natural language input:<performing_artist> What is the name of ?'s first album? After performing any necessary searches, the automated assistant 120 may, at 482,<name_of_studio_album> " (The terms in <brackets> are simply meant to generally refer to the entity). After the automated assistant 120 has also determined various facts about the entity, assigned a measure of potential user interest in the facts, and ranked them, the automated assistant 120 may respond with the unsolicited content "By the way,<performance_artist> is playing near you next month. Want more information about tickets?'
[0061] In some implementations, the entity about which the proactive content module 136 determines a fact of potential interest may not be explicitly mentioned in the human-to-computer dialogue. In some implementations, this proactively incorporated content may be determined, for example, based on the state of an application running on the client device. Assume that the user 101 is playing a game on the client device. The automated assistant 120 on the computing device 210 may determine that other client devices are in a particular gameplay state and may provide various unsolicited content of potential interest to the user, such as tips, tricks, recommendations of similar games, etc., as part of the human-to-computer dialogue. In some implementations where the computing device 210 is a standalone interactive speaker, the computing device 210 may even output background music (e.g., duplicating or adding background music) and / or sound effects related to the game being played on the other client device.
[0062] FIG. 5 illustrates an exemplary human-computer dialogue between a user 101 and an instance of an automated assistant 120 running on a client device 406 executed by the user 101. In this example, the user 101 does not provide natural language input. Instead, the computing device 210 (again, taking the form of a standalone interactive speaker) is playing music. This music is detected at one or more audio sensors (e.g., microphones) of the client device 406. One or more components of the client device 406, such as a software application configured to analyze the audibly detected music, can identify one or more attributes of the detected music, such as artist / song. Another component, such as the entity module 134 of FIG. 1, can use these attributes to search one or more online sources (e.g., the knowledge graph 124) to identify entities and associated facts. The automated assistant 120 running on the client device 406 can then provide (at 582) unsolicited content, e.g., audibly or visually, on the client device 406 of FIG. 5, informing the user 101 of various information about the entity. For example, at 582 in Figure 5, the automated assistant 120 states, "You're listening to <artist>. Did you know <artist> has a tour date in <your city> on <date>?" Similar techniques can be applied by an instance of the automated assistant 120 running on a client device (e.g., a smartphone, tablet, laptop, standalone interactive speaker) when the instance detects (by sound or visual detection) that audiovisual content (e.g., a movie, a television program, a sporting event, etc.) is being presented on the user's television.
[0063] 2-5 illustrate a human-computer dialogue in which the user 101 is engaged by the automated assistant 120 using audio input and output. However, this is not meant to be limiting. As mentioned above, in various implementations, the user can engage the automated assistant using other means, such as a messaging client 107. FIG. 6 illustrates an example in which a client device 606 in the form of a smartphone or tablet (not meant to be limiting) includes a touch screen 640. Visually rendered on the touch screen 640 is a transcript 642 of the human-computer dialogue between the user of the client device 606 ("you" in FIG. 6) and an instance of the automated assistant 120 running (at least in part) on the client device 606. An input field 644 is also provided in which the user can provide natural language content, as well as other types of input, such as images, sounds, etc.
[0064] In FIG. 6, the user starts a human-computer dialog session with the question "<Store> what time does it open?". The automatic assistant 120 ( "AA" in FIG. 6) performs one or more searches for information regarding the business hours of the store and answers "The <Store> opens at 10:00 am". The automatic assistant 120 then determines one or more facts regarding the entity <Store>, ranks those facts according to criteria that may be of potential interest to the user, and proactively incorporates the following content into the human-computer dialog: "By the way, the <Store> is going to hold a spring clearance sale next month". The automatic assistant 120 then incorporates user interface elements in the form of a "card" 646 regarding the sale. The card 646 can include various content, such as a link to the store's website for a shopping application installed on the client device 606 that is operable to purchase items from the store, a so-called "deeplink". In some implementations, the automatic assistant 120 can also incorporate other non-requested content as selectable options, such as one or more hyperlinks 648 to web pages, for example, to the store-related web page and / or the competitor's web page.
[0065] The card 646 in FIG. 6 is a visual element that can be selected by tapping it or, in some cases, touching it, although this is not meant to be limiting. A human-computer dialog similar to that shown in FIG. 6 may occur audibly between the user and an audio output device (e.g., the stand-alone interactive speaker shown in the previous figure). In some such implementations, user interface elements may take the form of audible prompts, such as questions or options that can be “selected” if positively answered by the user. For example, instead of presenting the card 646, the automatic assistant 120 may audibly output something like “Would you like a list of items that will soon be on sale?” In some implementations, the shopping application itself may include its own automatic assistant specifically tuned to engage in a human-computer dialog with the user to enable the user to learn about sale items. In some such implementations, the user may be “transferred” to the shopping-application-specific automatic assistant. In other implementations, the automatic assistant 120 can construct a natural language output that requests from the user the information necessary to obtain items using the shopping application, leveraging various information and states related to the shopping application. The automatic assistant 120 can then interact with the shopping application on behalf of the user (e.g., in response to spoken natural language input provided by the user).
[0066] FIG. 7 also shows a client device 606 having a touch screen 640 and user input field 644, as well as a transcript 742 of a human-computer dialog session. In this example, the user (“you”) initiates a human-computer dialog by typing and / or speaking a natural language input “Song by <artist>” (which can be recognized and converted to text). The automatic assistant 120 (“AA”) responds by providing a series of selectable user interface elements 746, and the user can select (e.g., tap) any one of those elements to play back the corresponding song by the artist, for example, on the client device 606 or another nearby client device (not shown, e.g., a stand-alone interactive speaker that forms part of an ecosystem of tuned client devices like 606). Using the techniques described herein, the automatic assistant 120 determines, for example, from the knowledge graph 124 that the artist turned 40 last week. In response, the automatic assistant 120 proactively incorporates the following statement into the human-computer dialog: “Did you know that <artist> turned 40 last week?” In some implementations, if the user was going to answer affirmatively or not answer at all, the automatic assistant 120 can refrain from providing additional proactive content, for example, based on the assumption that the user is not interested in receiving additional content. However, assume that the user provided a natural language response suggesting interest, such as “Oh really? I didn't know! Getting old!” In such a case, the automatic assistant 120 can detect such sentiment (e.g., by performing sentiment analysis on the user's input) and provide additional facts about that artist or, perhaps, another closely related artist (e.g., an artist of a similar age).
[0067] The examples of unsolicited content proactively incorporated as described above are not meant to be limiting. Using the techniques described herein, other unsolicited content potentially of interest to the user can be proactively incorporated within a human-computer dialog. For example, in some implementations where the user mentions their next scheduled flight (or train departure, or other travel arrangements), the automated assistant 120 can proactively incorporate unsolicited content within a human-computer dialog session with the user. This unsolicited content can include, for example, traffic patterns on the route to the airport, one or more user interface elements selectable (by touch, voice, gesture, etc.) to open an application that enables the user to view or edit their scheduled flight, information about alternative flights that may be less expensive (or selectable user interface elements linking to that flight), and the like.
[0068] Of course, the user does not always desire unsolicited content. For example, the user may be driving in heavy traffic, may be in an emergency situation, and may be operating the computing device in a manner that suggests they do not want to receive unsolicited content (e.g., during a video call). Thus, in some implementations, the automated assistant 120 can be configured to determine criteria for user desire (e.g., based on location signals, context of the conversation, state of one or more applications, accelerometer signals, sentiment analysis of the user's natural language input, etc.) for receiving unsolicited content, and can provide the unsolicited content only if this criteria meets one or more thresholds.
[0069] FIG. 8 also shows a client device 606 having a touch screen 640 and a user input field 644, as well as a transcript 842 of a human-computer dialog session. In this example, the user ("you") starts a human-computer dialog by typing and / or speaking a natural language input "flight to Louisville on Monday next week" (which can be recognized and converted to text). The automatic assistant 120 ("AA") responds by providing a series of selectable user interface elements 846, and the user can select (e.g., tap) any one of those elements to participate in the dialog, for example, using the automatic assistant or a travel application installed on the client device 606 to obtain the corresponding ticket. Using the techniques described herein, the automatic assistant 120 also determines that a cheaper flight is available, for example, by the proactive content module 136, if the user selects to depart on the next Tuesday instead of the next Monday. In response, the automatic assistant 120 proactively incorporates the following statement into the human-computer dialog: "By the way, if you fly one day later, you can save $145."
[0070] FIG. 9 is a flow diagram showing an exemplary method 900 according to the implementations disclosed herein. For convenience, the operations of the flow diagram are described with reference to a system that performs these operations. This system can include various components of various computer systems, such as one or more components of the automatic assistant 120. Moreover, although the operations of method 900 are shown in a particular order, this does not mean that it is limiting. One or more operations may be reordered, omitted, or added.
[0071] In block 902, the system can identify, by the entity module 134, entities mentioned by the user or the automated assistant based on, for example, the content of an existing human-computer dialog session between the user and the automated assistant. As suggested above, entities can appear in various forms such as people (e.g., celebrities, public figures, writers, artists, etc.), places (cities, states, countries, points of interest, intersections, restaurants, businesses, hotels, etc.), things (e.g., flights, train trips, products, services, songs, albums, films, books, poems, etc.). The entity module 134 (or, more generally, the automated assistant 120) can use various data sources such as the knowledge graph 124, annotations from an entity tagger associated with the natural language processor 122, a fresh content module 138, various other miscellaneous domain modules 140, etc. to identify entities.
[0072] In block 904, the system can identify, by the entity module 134, one or more facts about that entity or another entity related to that entity based on entity data contained within one or more databases. For example, the entity module 134 can examine the knowledge graph 124 for nodes, attributes, edges (which can represent relationships to other entities), etc. that enable another component of the entity module 134 or the automated assistant 120 to identify facts about either the mentioned entity or another entity that is related in some way to the mentioned entity. For example, if the user or the automated assistant mentions Mozart, the automated assistant 120 can identify facts related to another similar composer of the same or a similar era in addition to, or instead of, facts related to Mozart.
[0073] In some implementations, the automatic assistant 120 can identify entities and / or facts related to another entity depending on data related to the user profile. For example, assume that the user mentions a first musician (e.g., asks a question about the first musician, requests playback of a song composed by the first musician) in a human-computer dialog with the automatic assistant 120. Assume that the first musician is a musician that the user originally listened to frequently (e.g., determined by a playlist or playback history related to the user's user profile), and further assume that a second musician that the user also frequently listens to follows immediately thereafter. In some implementations, at block 904, the automatic assistant 120 can identify facts related to the first musician and / or facts related to the second musician.
[0074] At block 906, the system can determine (i.e., score) corresponding criteria that may be of potential interest to the user for each of one or more facts determined at block 904. The criteria for potential interest may appear in various forms (e.g., a percentage, a value within a range, a numerical value, etc.) and can be determined in various ways. In some implementations, the criteria for potential interest can be determined based on the user's unique user profile. For example, if the user frequently searches for flights from two different airports to compare costs and then mentions the first airport in a query about flights directed to the automatic assistant 120, the automatic assistant 120 can assign (e.g., promote) a relatively large criterion of potential interest to facts about flights from the second airport even if the user does not explicitly mention the second airport.
[0075] Additionally or alternatively, in various implementations, the fact can be scored using aggregated user data and / or behavior. For example, when discussing an entity, data can be collected from an online conversation to determine, for example, as an aside, which entity attributes are often brought up. As another example, when a user searches for information about an entity or, in some cases, consumes that information, the aggregated user search query logs can be analyzed to determine which entity attributes are often searched for, clicked on, or, in some cases, interacted with. As yet another example, in some implementations, the system can analyze trend searches and / or news from, for example, a fresh content module 138 to determine which facts about an entity may be currently trendy (and thus, potentially, a higher criterion of interest can be assigned than if that were not the case). As yet another example, the user's own context information (e.g., location data generated by a location coordinate sensor integrated with the user's smartphone) can be used to assign a criterion of potential interest. For example, if the entity being discussed is having an upcoming event nearby (as determined from the user's current location), that fact can be assigned a greater criterion of potential interest than if the upcoming event related to the entity were taking place at a distant location.
[0076] In block 908, the system can generate unsolicited natural language content that includes one or more of the facts selected based on one or more corresponding criteria that may be of interest. In some implementations, the system can select only the top-ranked facts that would be included within the unsolicited content. In other implementations, the system can select the facts ranked in the top n (a positive integer). In block 910, the system can incorporate the unsolicited natural language content into an existing human-computer dialog session or a subsequent human-computer dialog session. For example, the automated assistant 120 can generate natural language descriptions that are prefaced by phrases such as "by the way," "did you know," "incidentally," etc. As described above, in some implementations, this unsolicited content can include selectable graphical elements, such as deep links, that can be selected by the user to initiate various tasks, such as participating in additional dialog, and / or it may be attached thereto.
[0077] Although not explicitly shown in FIG. 9, in various implementations, the automatic assistant 120 can avoid incorporating unsolicited content that may already be known to the user into the human-computer dialog session with the user. For example, in some implementations, the system can detect that a given fact among one or more facts identified at block 904 has been previously referenced in an existing human-computer dialog between the user and the automatic assistant or in a previous human-computer dialog between the user and the automatic assistant (optionally, going back a predetermined time interval such as 30 days). In some implementations, the criteria for potential interest determined by the system for a given fact may reflect this detection, for example, by having a relatively low score (or even a zero score in some cases) assigned. In other implementations, the system can exclude a given fact from consideration based on the detection that the given fact has been previously referenced in an existing human-computer dialog between the user and the automatic assistant or in a previous human-computer dialog between the user and the automatic assistant.
[0078] FIG. 10 is a block diagram of an exemplary computing device 1010 that can optionally be utilized to perform one or more of the techniques described herein. In some implementations, one or more of the client computing device, the automatic assistant 120, and / or other components can include one or more components of the exemplary computing device 1010.
[0079] Computing device 1010 generally includes at least one processor 1014 that communicates with several peripheral devices via a bus subsystem 1012. These peripheral devices can include, for example, a storage subsystem 1024 that includes a memory subsystem 1025 and a file storage subsystem 1026, a user interface output device 1020, a user interface input device 1022, and a network interface subsystem 1016. The input / output devices enable a user to interact with the computing device 1010. The network interface subsystem 1016 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.
[0080] The user interface input device 1022 can include a keyboard, a mouse, a trackball, a pointing device such as a touchpad or a graphics tablet, a scanner, a touch screen incorporated within a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. Generally, the use of the term "input device" is intended to include all conceivable types of devices and ways of inputting information into the computing device 1010 or onto a communication network.
[0081] The user interface output device 1020 may include a non-visual display such as a display subsystem, a printer, a facsimile machine, or an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visual image. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally, the use of the term "output device" is intended to include all conceivable types of devices and ways for outputting information from the computing device 1010 to the user or to another machine or computing device.
[0082] The memory subsystem 1024 stores programming constructs and data constructs that provide some or all of the functionality of the modules described herein. For example, the memory subsystem 1024 may include logic for executing selected aspects of the method of FIG. 9 and for implementing the various components shown in FIG. 1.
[0083] These software modules are generally executed by the processor 1014 alone or in combination with other processors. The memory 1025 used within the memory subsystem 1024 may include several memories including a main random access memory (RAM) 1030 for storing instructions and data during program execution and a read-only memory (ROM) 1032 in which fixed instructions are stored. The file storage subsystem 1026 can provide a persistent storage device for program files and data files and may include a hard disk drive, a floppy disk drive, a CD-ROM drive, an optical drive, or a removable media cartridge together with associated removable media. Modules implementing the functionality of some implementations may be stored by the file storage subsystem 1026 within the memory subsystem 1024 or in another machine accessible by the processor 1014.
[0084] The bus subsystem 1012 provides a mechanism for enabling the various components and subsystems of the computing device 1010 to communicate with each other as intended. Although the bus subsystem 1012 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.
[0085] The computing device 1010 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the ever-changing nature of computers and networks, the description of the computing device 1010 shown in FIG. 10 is intended only as a specific example for illustrating some implementations. Many other configurations of the computing device 1010 with more or fewer components than the computing device shown in FIG. 10 are possible.
[0086] In situations where some of the implementations discussed herein may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, as well as the user's activities and demographic information, relationships between users, etc.), one or more opportunities are provided to the user to control whether information is collected, whether personal information is stored, whether personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed herein collect, store, and / or use user personal information only when an explicit permission to do so has been received from the relevant user.
[0087] For example, a user is provided with control over whether a program or feature collects user information about a particular user or other users associated with that program or feature. Each user from whom personal information is collected is presented with one or more options for providing permission or authorization regarding whether information is to be collected and which portions of the information are to be collected in order to enable control over the collection of information related to that user. For example, one or more such control options may be provided to the user via a communication network. Additionally, some data may be handled in one or more ways before the data is stored or used such that personally identifiable information is removed. In one example, a user's identifying information may be handled such that personally identifiable information cannot be determined. As another example, a user's geographical location may be generalized to a broader area such that the user's specific location cannot be determined.
[0088] Although several implementations have been described and illustrated herein, various other means and / or structures for performing functions and / or obtaining results, and / or one or more of the advantages described herein may be utilized, and such variations and / or modifications are each considered to be within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are illustrative, and the actual parameters, dimensions, materials, and / or configurations will depend on the particular one or more applications in which the present teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the foregoing implementations are presented by way of example only, and it should be understood that implementations other than those specifically described and claimed may be practiced within the scope of the appended claims and their equivalents. Implementations of the present disclosure relate to each individual feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods do not mutually conflict.
Explanation of Signs
[0089] 101 User 106 1~N Client Computing Device, Client Device 107 1~N Message Exchange Client 109 1~N Application, MISC.APP 120 Automatic Assistant 122 Natural Language Processor 124 Knowledge Graph 126 User Profile Database, Database, User Profile 130 Responsive Content Engine 132 Behavior Module 134 Entity Module 136 Proactive Content Module 138 Fresh Content Module 140 Miscellaneous Domain Module 210 Computing Device 280 Natural Language Input 282 Responsive Natural Language Output 406 Client Device 606 Client Device 640 Touch Screen 642 Transcript 644 Input Field, User Input Field 646 Card 648 Hyperlink 842 Transcript 900 Method 1010 Computing Device 1012 Bus Subsystem 1014 Processor 1016 Network Interface Subsystem 1020 User Interface Output Device 1022 User Interface Input Device 1024 Memory Subsystem 1025 Memory Subsystem 1026 File Memory Subsystem 1030 Main Random Access Memory (RAM) 1032 Read-Only Memory (ROM)
Claims
1. 1. A method implemented using one or more processors, comprising: processing a speech input provided by a user as part of a dialogue session involving the user and an automated assistant executed by one or more of the processors; generating solicited natural language content, the solicited natural language content being responsive to requests identified in the speech input based on the processing; incorporating, by the automated assistant, the solicited natural language content into the dialogue session involving the user and the automated assistant, the solicited natural language content including a fact; determining criteria for a desire by the user to receive unsolicited content based on the request identified in the speech input and contextual information associated with the user; responsive to determining that the criteria of desire meets a threshold, causing unsolicited natural language content to be automatically output to the user without the user specifically requesting the unsolicited natural language content, the unsolicited natural language content including a query that derails from the desire identified in the speech input based on the processing; the query incorporates the fact, the automated assistant engaging the user in additional dialogue related to the query if the user responds affirmatively to the query; in response to determining that the criteria of desire does not meet a threshold, refraining from automatically outputting unsolicited natural language content to the user. A method comprising:
2. The method of claim 1 , wherein the contextual information includes traffic detected near the user's current location.
3. The method of claim 1 , wherein the context information includes past human-to-computer dialogues between the user and the automated assistant.
4. The method of claim 1 , wherein the context information includes one or more applications with which the user is currently interacting.
5. The method of claim 1 , wherein the context information includes a state of an application running on a computing device controlled by the user.
6. The method of claim 1 , wherein the contextual information comprises an accelerometer signal generated by a computing device carried by the user.
7. The method of claim 1 , wherein the contextual information includes a sentiment analysis of a speech recognition output of the speech input.
8. 1. A system comprising one or more processors and a memory, the memory being configured to, in response to execution of a computer program by the one or more processors, provide to the one or more processors: processing a speech input provided by a user as part of a dialogue session involving the user and an automated assistant executed by one or more of the processors to identify a request; generating solicited natural language content, the solicited natural language content being responsive to the request; incorporating, by the automated assistant, the solicited natural language content into the dialogue session involving the user and the automated assistant, the solicited natural language content including a fact; determining criteria for a desire by the user to receive unsolicited content based on the request identified in the speech input and contextual information associated with the user; responsive to determining that the criteria of desire meets a threshold, causing unsolicited natural language content to be automatically output to the user without the user specifically requesting the unsolicited natural language content, the unsolicited natural language content including a query that derails the desire identified in the speech input; the query incorporates the fact, the automated assistant engaging the user in additional dialogue related to the query if the user responds affirmatively to the query; in response to determining that the criteria of desire does not meet a threshold, refraining from automatically outputting unsolicited natural language content to the user. A system that stores a computer program that causes the
9. The system of claim 8 , wherein the contextual information includes traffic detected near the user's current location.
10. 10. The system of claim 8, wherein the context information includes past human-computer dialogues between the user and the automated assistant.
11. The system of claim 8 , wherein the context information includes one or more applications with which the user is currently interacting.
12. The system of claim 8 , wherein the context information includes a state of an application running on a computing device controlled by the user.
13. The system of claim 8 , wherein the contextual information comprises an accelerometer signal generated by a computing device carried by the user.
14. The system of claim 8 , wherein the contextual information includes a sentiment analysis of a speech recognition output of the speech input.
15. A non-transitory computer-readable storage medium comprising a computer program, the computer program causing a processor to, in response to execution of the computer program by the processor, processing a speech input provided by a user as part of a dialogue session involving the user and an automated assistant executed by one or more of the processors to identify a request; generating solicited natural language content, the solicited natural language content being responsive to the request; incorporating, by the automated assistant, the solicited natural language content into the dialogue session involving the user and the automated assistant, the solicited natural language content including a fact; determining criteria for a desire by the user to receive unsolicited content based on the request identified in the speech input and contextual information associated with the user; responsive to determining that the criteria of desire meets a threshold, causing unsolicited natural language content to be automatically output to the user without the user specifically requesting the unsolicited natural language content, the unsolicited natural language content including a query that derails the desire identified in the speech input; the query incorporates the fact, the automated assistant engaging the user in additional dialogue related to the query if the user responds affirmatively to the query; in response to determining that the criteria of desire does not meet a threshold, refraining from automatically outputting unsolicited natural language content to the user. A non-transitory computer-readable storage medium that causes
16. 20. The non-transitory computer-readable storage medium of claim 15, wherein the contextual information comprises traffic detected near the user's current location.
17. 16. The non-transitory computer-readable storage medium of claim 15, wherein the context information includes past human-to-computer dialogues between the user and the automated assistant.
18. 16. The non-transitory computer-readable storage medium of claim 15, wherein the contextual information includes one or more applications with which the user is currently interacting.
19. 16. The non-transitory computer-readable storage medium of claim 15, wherein the contextual information includes a state of an application running on a computing device controlled by the user.
20. 16. The non-transitory computer-readable storage medium of claim 15, wherein the contextual information comprises an accelerometer signal generated by a computing device carried by the user.
Citation Information
Patent Citations
Information dissemination method, server, information terminal device, system, and voice interaction system
WO2016129276A1