Generating natural language output based on natural language user interface input

By generating personal database entries based on free-form natural language input and using a search parameter matching mechanism, the problem of precise matching dependence and user intervention requirements in existing note-taking applications is solved, achieving more efficient natural language input management and response.

CN117235335BActive Publication Date: 2026-02-03GOOGLE LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311109978.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2016-05-17
Filing Date
2016-12-27
Publication Date
2026-02-03
Estimated Expiration
2036-12-27

AI Technical Summary

Technical Problem

Existing note-taking computer applications rely on precise keyword matching when searching and managing note entries, lack automatic filtering and ranking, require users to explicitly specify content, and cannot handle free-form natural language input and contextual features.

Method used

By generating personal database entries based on users' free-form natural language input, including words and descriptive metadata, and utilizing search parameter matching and ranking mechanisms, the system processes users' natural language input to generate response outputs.

Benefits of technology

It enables effective management and searching of free-form natural language input, improves the matching accuracy and response efficiency of note entries, reduces user intervention, and enhances the processing capabilities of natural language input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117235335B_ABST
    Figure CN117235335B_ABST
Patent Text Reader

Abstract

This application discloses generating natural language output based on natural language user interface input. Some implementations involve generating personal database entries for a user based on free-form natural language input made by the user via a user interface input device of a computing device of the user. The generated personal database entries can include terms of the natural language input and descriptive metadata determined based on the terms of the natural language input and / or based on contextual features associated with receiving the natural language input. Some implementations involve generating output responsive to another free-form natural language input of a user based on one or more personal database entries of the user. For example, one or more entries responsive to another natural language input can be identified based on matching content of the entries to search parameters determined based on the other input. Some implementations involve improved automated personal assistants.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Case Analysis

[0002] This application is a divisional application of Chinese invention patent application 201680085778.4, filed on December 27, 2016. Technical Field

[0003] This specification describes the generation of natural language output based on natural language user interface input. Some implementations relate to improved automated personal assistants. Background Technology

[0004] Some note-taking computer applications and / or other computer applications enable users to create note entries that include content explicitly set by the user. For example, a user can create a note for "Calling Bob". Furthermore, some computer applications allow for limited searching of user-created notes.

[0005] However, these and / or other technologies may suffer from one or more drawbacks. For example, note-taking computer applications may search for note entries solely based on an exact keyword match between the search terms and the words in the notes. For instance, if the search includes “call” and / or “Bob,” only “call Bob” will be identified. Furthermore, some note-taking computer applications may not automatically filter and / or rank note entries in response to a search and / or may only do so based on a limited amount of user-provided content. Additionally, some note-taking computer applications may require the user to explicitly specify that user-provided content will be used to create note entries, thus not creating any entries for many types of user input (such as input provided during conversations with a personal assistant and / or during conversations with one or more additional users). Additional and / or alternative drawbacks may arise. Summary of the Invention

[0006] Some embodiments of this specification relate to improved automated personal assistants. Some embodiments relate to generating a user's personal database entries based on free-form natural language input made by the user via one or more user interface input devices on the user's computing device (such as natural language input provided to the automated personal assistant and / or to one or more computing devices of these additional users during communication with one or more additional users). In some of these embodiments, the entries include one or more words of natural language input and optionally include descriptive metadata determined based on one or more words of natural language input and / or based on contextual features associated with receiving the natural language input. As used herein, free-form input is input made by the user and is not limited to input presented to the user as a set of options to choose from. For example, it could be typing input provided by the user via a physical or virtual keyboard on the computing device, voice input provided by the user to a microphone on the computing device, and / or images captured by the user via a camera (e.g., the camera of the user's computing device).

[0007] Some embodiments of this specification additionally and / or alternatively relate to generating output in response to another free-form natural language input from a user, based on one or more personal database entries from that user. In some of these embodiments, one or more entries in response to another natural language input from the user may be identified based on matching the (soft and / or precise) content of these entries (e.g., descriptive metadata and / or words of these entries) with one or more search parameters determined based on the other natural language input. Further, in some of these embodiments, the content of the identified entries may be used when generating output and / or when ranking multiple identified entries relative to each other (in the case of identifying multiple entries). Ranking may be used to select a subset of entries (e.g., one entry) to use and / or determine the presentation order of multiple pieces of content in the generated response (based on the ranking of the entries from which multiple pieces of content are derived).

[0008] Some embodiments of this specification additionally and / or alternatively involve analyzing a given natural language input to determine a measure of the likelihood that the input is intended as input that the user expects the automated personal assistant to subsequently recall and / or to determine a measure of the likelihood that the input is intended as a request for personal information from the user's personal database.

[0009] Natural language input from users can be received and processed in various scenarios. For example, natural language input can be input provided by the user during communication with one or more other users (e.g., via chat, SMS, and / or other message exchanges). As another example, natural language input can be provided to an automated personal assistant that participates in a conversation with the user via one or more user interface input and output devices. For example, the automated personal assistant can be wholly or partially integrated into the user's computing device (e.g., a mobile phone, tablet computer, device dedicated to automated assistance functions) and can include one or more user interface input devices (e.g., a microphone, a touchscreen) and one or more user interface output devices (e.g., a speaker, a display screen). Moreover, for example, the automated personal assistant can be wholly or partially implemented in one or more computing devices that are separate from but communicate with the user's client computing device.

[0010] In some embodiments, a method executed by one or more processors is provided, the method comprising: receiving first natural language input, the first natural language input being free-form input made by a user via a user interface input device of a user's computing device. The method further comprises: generating an entry for the first natural language input in a user's personal database stored in one or more computer-readable media, the generation comprising: storing one or more given words or identifiers of given words from the first natural language input in the entry; generating descriptive metadata based on at least one of the words from the first natural language input, and one or more contextual features associated with receiving the first natural language input; and storing the descriptive metadata in the entry. The method further comprises: receiving second natural language input after receiving the first natural language input. The second natural language input is free-form input made by a user via a user interface input device or an additional user interface input device of a user's additional computing device. The method further comprises: determining at least one search parameter based on the second natural language input; searching the personal database based on the search parameter; and determining an entry in response to the second natural language input based on the search and at least in part on matching the search parameter with at least some of the descriptive metadata. The method further includes: generating natural language output, the natural language output including one or more natural language output words based on entries; and, in response to a second natural language input, providing the natural language output to be presented to a user via a user interface output device of a computing device or a user interface output device of an additional computing device.

[0011] This method and other implementations of the technology disclosed herein may optionally include one or more of the following features.

[0012] In some implementations, the generation of descriptive metadata is based on both a given word of the first natural language input and one or more contextual features associated with receiving the first natural language input.

[0013] In some embodiments, generating descriptive metadata is based on one or more contextual features associated with receiving a first natural language input, and generating descriptive metadata based on one or more contextual features includes generating temporal metadata indicating the date or time the first natural language input was received. In some of these embodiments, in generating temporal metadata, at least one search parameter is a time search parameter, and determining an entry in response to a second natural language input includes matching the time search parameter with the temporal metadata. In some of these embodiments, in generating temporal metadata, the method further includes determining an additional entry in response to the second natural language input; and selecting an entry instead of the additional entry based on the consistency of the current date or time with the temporal metadata of the entry. In some versions of these embodiments, natural language output including at least some of the given words of the entry is provided in response to selecting the entry, and the additional entry is not used when generating the natural language output, and no output based on the additional entry is provided in response to the second natural language input. In some versions of these embodiments, generating natural language output includes generating one or more time words of the natural language output based on the temporal metadata of the entry.

[0014] In some implementations, generating descriptive metadata is based on one or more contextual features associated with receiving a first natural language input, and generating descriptive metadata includes generating location metadata indicating the user's location at the time the first natural language input is received. In some implementations of these embodiments, in generating location metadata, at least one search parameter is a location search parameter, and determining an entry in response to a second natural language input includes matching the location search parameter with the location metadata. In some implementations of these embodiments, in generating location metadata, the method further includes determining an additional entry in response to the second natural language input; and selecting an entry instead of the additional entry based on the consistency between the user's current location and the location metadata of the entry. In some versions of these implementations, natural language output including at least some of the given words of the entry is provided in response to selecting the entry, and the additional entry is not used when generating the natural language output, and no output based on the additional entry is provided in response to the second natural language input. In some versions of these implementations, generating natural language output includes generating one or more location words of the natural language output based on the location metadata of the entry.

[0015] In some implementations, during communication between the user and at least one additional user, first natural language input is provided, descriptive metadata is generated based on one or more contextual features associated with receiving the first natural language input, and the generation of descriptive metadata includes additional user metadata that generates the descriptive metadata. The additional user metadata identifies the additional user communicating with the user upon receiving the first natural language input. In some of these implementations, in the case of generating additional metadata, communication between the user and the additional user is via a first message exchange client of the user's computing device and a second message exchange client of the additional user's computing device.

[0016] In some implementations, generating descriptive metadata is based on one or more words from a first natural language input, and generating descriptive metadata based on one or more words from the first natural language input includes generating semantic tags based on one or more words from the first natural language input. Semantic tags indicate the categories to which one or more words belong, as well as the categories to which additional words not included in the first natural language output also belong. In some implementations of these embodiments, in the case of generating semantic tags, at least one search parameter is a semantic search parameter, and determining an entry in response to a second natural language input includes matching the semantic search parameter with the semantic tags. In some versions of these implementations, determining the search parameter is based on a prefix of the second natural language input. In some implementations of these embodiments, in the case of generating semantic tags, the method further includes: determining additional entries in response to the second natural language input; and selecting an entry instead of an additional entry based on the consistency between the user's second natural language input and the semantic tags of the entry. In some versions of these implementations, natural language output including at least some of the given words of the entry is provided in response to selecting the entry, and additional entries are not used when generating natural language output, and no output based on additional entries is provided in response to the second natural language input.

[0017] In some embodiments, the method further includes: determining a measure based on at least one of the words in the second natural language input that indicates the second natural language input is a request for information from a personal database. In some of these embodiments, at least one of the search and provision depends on the magnitude of the measure.

[0018] In some embodiments, a method executed by one or more processors is provided, the method comprising: receiving first natural language input, the first natural language input being free-form input made by a user via a user interface input device of a user's computing device; and generating an entry for the first natural language input in a user's personal database stored in one or more computer-readable media. The generation includes: storing one or more given words or identifiers of given words from the first natural language input in the entry; generating descriptive metadata based on at least one of the words from the first natural language input, and one or more contextual features associated with receiving the first natural language input; and storing the descriptive metadata in the entry. The method further includes: receiving a second natural language input after receiving a first natural language input; determining at least one search parameter based on the second natural language input; searching a personal database based on the search parameter; determining entries that respond to the second natural language input based on the search; determining additional entries that also respond to the second natural language input based on the search; ranking the entries relative to the additional entries based on at least some of the descriptive metadata; generating natural language output including one or more natural language output words based on the entries; and providing natural language output to a user via a user interface output device in response to the second natural language input, wherein providing natural language output is based on ranking the entries relative to the additional entries.

[0019] This method and other implementations of the technology disclosed herein may optionally include one or more of the following features.

[0020] In some implementations, providing natural language output based on ranking entries relative to additional entries includes providing natural language output in response to a second natural language input without providing any output based on additional entries.

[0021] In some implementations, providing natural language output based on ranking entries relative to additional entries includes providing natural language output and providing additional output based on additional entries, and providing natural language output and additional output in a ranking-based order.

[0022] Additionally, some embodiments include one or more processors of one or more computing devices, wherein the one or more processors are operable to execute instructions stored in associated memory, and wherein the instructions are configured to perform any of the methods described above. Some embodiments also include a non-transitory computer-readable storage medium storing computer instructions executable by one or more processors to perform any of the methods described above.

[0023] It should be understood that all combinations of the foregoing and additional concepts described in detail herein are considered part of the subject matter disclosed herein. For example, all combinations of the claimed topics appearing at the end of this disclosure are considered part of the subject matter disclosed herein. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of an example environment in which the implementation methods disclosed herein can be carried out.

[0025] Figure 2A The illustration shows an example of generating a user's personal database entries based on free-form natural language input made by the user via one or more user interface input devices on the user's computing device.

[0026] Figure 2B The illustration shows an example of generating natural language output in response to another form of natural language input from a user, based on one or more personal database entries of the user.

[0027] Figure 3A An example client computing device with a display screen according to an embodiment described herein is illustrated, the display screen showing an example of a conversation that may occur between a user of the client computing device and an automated personal assistant.

[0028] Figure 3B , Figure 3C , Figure 3D , Figure 3E , Figure 3F and Figure 3G The figures illustrate a display screen according to the embodiments described herein. Figure 3A As part of an example client computing device, the display shows what can happen. Figure 3A The dialogue that follows the first dialogue.

[0029] Figure 4A The illustration shows a display screen according to an embodiment described herein. Figure 3A The client computing device displays another example of a conversation that can occur between the user of the client computing device and an automated personal assistant.

[0030] Figure 4B and Figure 4C The figures illustrate a display screen according to the embodiments described herein. Figure 4A As part of an example client computing device, the display shows what can happen. Figure 4A The dialogue that follows the first dialogue.

[0031] Figure 5A and Figure 5BThe flowchart presents an example method for generating a user's personal database entries based on free-form natural language input made by the user, and generating natural language output in response to another free-form natural language input from the user based on one or more of the personal database entries.

[0032] Figure 6 An example architecture of a computing device is illustrated. Detailed Implementation

[0033] This specification relates to the generation of a user's personal database entries based on free-form natural language input made by the user via one or more user interface input devices on the user's computing device. For example, a personal database entry can be generated by an automated personal assistant in response to natural language input provided to the automated personal assistant. As used herein, a personal database entry is an entry in a personal database accessible to the user but inaccessible to multiple additional users. For example, only the user and one or more applications and / or systems specified by the user can access the personal database, while any other user, application, and / or system cannot. Furthermore, for example, the user and certain additional users specified by the user can access the personal database. As described herein, any personal database can be stored on one or more computer-readable media and can be associated with access information to allow access only to users, applications, and / or systems authorized to access this content and to prevent access by all other users, applications, and / or systems.

[0034] In some implementations, the generated personal database entries include one or more words from natural language input, and optionally include one or more words based on natural language input and / or descriptive metadata determined based on contextual features associated with receiving the natural language input.

[0035] Descriptive metadata based on word determination of natural language input can include, for example, semantic labels indicating the category to which one or more words belong, as well as additional words not included in the natural language output that also belong to that category. For example, the semantic label for "person" can be determined based on the word "Bob," and the semantic label for "place" can be determined based on the word "home," etc. Another example of descriptive metadata based on word determination of natural language input can include "memory-based" metadata, which indicates the likelihood (e.g., true / false, or range of values) that the natural language input is intended to represent as input that the user expects an automated personal assistant to be able to recall subsequently. For example, descriptive memory-based metadata indicating a high probability of the user's expected future recall ability can be generated based on the presence of certain keywords in the input (e.g., "remember," "don't forget") and / or based on the output received in response to providing the input to a classifier trained to predict whether and / or to what extent the natural language input indicates the expectation of "remembering" the natural language input.

[0036] Descriptive metadata determined based on contextual features associated with the received natural language input may include, for example, time metadata, location metadata, and / or additional user metadata. Time metadata may indicate the date (e.g., 5 / 1 / 16; May; 2016; last week; Sunday) and / or time (e.g., 8:00 AM, morning, between 7:00 and 9:00). Location metadata may indicate the user's location at the time the natural language input was received (e.g., postal code, "home," "work," city, neighborhood). Additional user metadata may indicate one or more additional users to whom the natural language input is directed and / or one or more additional users who are communicating with the user when the natural language input is received. For example, when natural language input is received from an ongoing message exchange thread (e.g., chat, SMS) between the user and an additional user, the additional user metadata may indicate that additional user.

[0037] Implementations of this specification additionally and / or alternatively relate to generating output in response to another free-form natural language input from a user, based on one or more personal database entries from that user. In some of these implementations, one or more entries in response to another natural language input from the user may be identified by matching the (soft and / or precise) content of these entries (e.g., descriptive metadata and / or natural language input terms of these entries) with one or more search parameters determined based on the other natural language input. Further, in some of these implementations, the content of the identified entries may be used when generating output and / or when ranking multiple identified entries relative to each other (in the case of identifying multiple entries). Ranking may be used to select a subset of entries for use in generating a response and / or to determine the presentation order of multiple pieces of content in the generated response (based on the ranking of the entries from which multiple pieces of content are derived).

[0038] Now go to Figure 1 The illustration shows an example environment in which the techniques disclosed herein can be implemented. The example environment includes multiple client computing devices 106. 1-N And automated personal assistant 120. Despite in Figure 1 The illustration of the Lieutenant General Automated Personal Assistant 120 is shown with a client computing device 106. 1-N Separate, but in some implementations, all or all aspects of the automated personal assistant 120 may be handled by the client computing device 106. 1-N One or more implementations of the automated personal assistant 120. For example, client computing device 1061 may implement one or more instances of the automated personal assistant 120, and client computing device 106 N Individual instances of one or more of these aspects of the automated personal assistant 120 may also be implemented. One or more aspects of the automated personal assistant 120 are controlled remotely by the client computing device 106. 1-N In one or more computing devices implemented in the embodiment, the client computing device 106 1-N These aspects of the automated personal assistant 120 can communicate via one or more networks (such as local area networks (LANs) and / or wide area networks (WANs) (e.g., the Internet)).

[0039] Client computing device 106 1-NThis may include, for example, one or more of the following: desktop computing devices, laptop computing devices, tablet computing devices, dedicated computing devices for enabling dialogue with a user (e.g., standalone devices including a microphone and speaker but without a display), mobile phone computing devices, computing devices in the user's vehicle (e.g., in-vehicle communication systems, in-vehicle entertainment systems, in-vehicle navigation systems), or wearable devices of the user that include computing devices (e.g., a user's watch with computing devices, glasses with computing devices, virtual or augmented reality computing devices). Additional and / or alternative client computing devices may be provided. In some embodiments, a given user may communicate with the automated personal assistant 120 using multiple client computing devices that collectively form a coordinated "ecosystem" of computing devices. For example, a personal database entry may be generated via a first client computing device based on the user's natural language input, and output may be generated via a second client computing device in response to another natural language input from the user based on that entry. However, for the sake of brevity, many of the examples described in this disclosure will focus on operating client computing device 106. 1-N A given user on a single client computing device.

[0040] Client computing device 106 1-N Each of these can operate various different applications, such as message exchange clients 107. 1-N A corresponding message exchange client in the context. Message exchange client 107 1-N It can be presented in various forms, and these forms can be displayed on the client computing device 106. 1-N Changes between and / or may occur on the client computing device 106 1-N Multiple forms of operation are performed on a single client computing device. In some implementations, message exchange client 107 1-N One or more of these may be presented as a Short Message Service (“SMS”) and / or Multimedia Messaging Service (“MMS”) client, an online chat client (e.g., instant messaging, Internet Relay Chat, or “IRC”, etc.), a messaging application associated with a social network, a personal assistant messaging service dedicated to dialogue with the automated personal assistant 120 (e.g., implemented on a dedicated computing device for implementing dialogue with the user), etc. In some embodiments, message exchange client 107 1-N One or more of these can be implemented via web pages or other resources rendered by a web browser (not shown) or other applications of the client computing device 106.

[0041] As described in more detail herein, the automated personal assistant 120 is transmitted via one or more client devices 106 1-NThe user interface input and output device receives input from one or more users and / or provides output to one or more users. The embodiments described herein relate to an automated personal assistant 120 that generates a user's personal database entry based on natural language input provided by the user and generates output to present to the user based on one or more of the generated user's personal database entries. However, it is to be understood that in many embodiments, the automated personal assistant 120 may also perform one or more additional functions. For example, the automated personal assistant 120 may respond to some of the user's natural language interface input using output generated wholly or partially based on information not based on personal database entries generated based on the user's past natural language input. For example, the automated personal assistant may use one or more documents from a public database 154 when generating a response to some of the user's natural language interface input. For example, in response to the user's natural language input "What does dexterity mean?", the automated personal assistant may generate output "Dexterity is an adjective describing the ability to use both the left and right hands with ease" based on a definition of "dexterity" that is not personal to the user.

[0042] In some implementations, the user interface input described herein explicitly points to the automated personal assistant 120. For example, message exchange client 107. 1-N One of them could be a personal assistant messaging service dedicated to communicating with the automated personal assistant 120, and user interface input provided via this personal assistant messaging service could be automatically provided to the automated personal assistant 120. For example, message exchange client 107 1-N One or more instances of the personal assistant 120 can be implemented at least partially on a standalone client computing device that at least selectively monitors voice user interface input and responds with audible user interface output. Furthermore, for example, user interface input can be based on specific user interface input via message exchange client 107. 1-N One or more of the user interface input explicitly refers to the automated personal assistant 120, indicating that the specific user interface input is to invoke the automated personal assistant 120. For example, the specific user interface input can be one or more typed characters (e.g., @PersonalAssistant), user interaction with a virtual button (e.g., tap, long press), verbal commands (e.g., "Hey, personal assistant"), etc. In some implementations, the automated personal assistant 120 can perform one or more actions in response to user interface input, even when the user interface input does not explicitly refer to the automated personal assistant 120. For example, the automated personal assistant 120 can examine the content of the user interface input and perform one or more actions in response to certain words present in the user interface input and / or based on other prompts.

[0043] Client computing device 106 1-N Each of the automated personal assistants 120 may include one or more memories for storing data and software applications, one or more processors for accessing data and executing applications, and other components for facilitating communication via a network. (By client computing device 106) 1-N One or more of the operations performed by the automated personal assistant 120 can be distributed across multiple computer systems. The automated personal assistant 120 can be implemented as, for example, a computer program running on one or more computers in one or more locations coupled to each other via a network.

[0044] The automated personal assistant 120 may include an input processing engine 122, a context engine 124, an entry generation engine 126, a search engine 128, a ranking engine 130, and an output generation engine 132. In some embodiments, one or more of engines 122, 124, 126, 128, 130, and / or 132 may be omitted. In some embodiments, all or aspects of one or more engines 122, 124, 126, 128, 130, and / or 132 may be combined. In some embodiments, one or more engines 122, 124, 126, 128, 130, and / or 132 may be implemented in a component separate from the automated personal assistant 120. For example, one or more engines 122, 124, 126, 128, 130, and / or 132, or any operable portion thereof, may be implemented in a component controlled by a client computing device 106. 1-N In the executing component.

[0045] Input processing engine 122 processes user input via client computing device 106 1-N The generated natural language input is processed, and annotated output is generated for use by one or more other components of the automated personal assistant 120. For example, the input processing engine 122 can process free-form natural language input generated by the user via one or more user interface input devices of the client device 1061. The generated annotated output includes one or more annotations to the natural language input and optionally includes one or more (e.g., all) words from the natural language input.

[0046] In some implementations, the input processing engine 122 is configured to identify and annotate various types of syntactic information in the natural language input. For example, the input processing engine 122 may include a part-of-speech tagger configured to annotate words using their syntactic roles. For example, the part-of-speech tagger may use its part of speech to tag each word, such as "noun," "verb," ​​"adjective," "pronoun," etc. Furthermore, for example, in some implementations, the input processing engine 122 may additionally and / or alternatively include a dependency parser configured to determine syntactic relations between words in the natural language input. For example, the dependency parser may determine which words modify other words, subjects, and verbs in the sentence (e.g., parse trees), and may annotate such dependency syntax.

[0047] In some implementations, the input processing engine 122 may additionally and / or alternatively include an entity annotator configured to annotate entity references, such as references to people, organizations, locations, etc., in one or more sections. The entity annotator may annotate entity references at a high-granularity level (e.g., to achieve identification of all references to an entity class (such as people)) and / or at a low-granularity level (e.g., to achieve identification of all references to a specific entity (such as a specific person)). The entity annotator may rely on the content of the natural language input to parse specific entities and / or may optionally communicate with a knowledge graph or other entity database to parse specific entities.

[0048] In some implementations, the input processing engine 122 may additionally and / or alternatively include a reference parser configured to group or “cluster” references to the same entity based on one or more contextual cues. For example, a reference parser could be used to resolve the word “it” in the natural language input “I’m looking for my bicycle helmet. Do you know where it is?” to “bicycle helmet”.

[0049] In some implementations, one or more components of the input processing engine 122 may rely on annotations from one or more other components of the input processing engine 122. For example, in some implementations, the named entity annotator may rely on annotations from the referential parser and / or dependency parser when annotating all references to a particular entity. Furthermore, for example, in some implementations, the referential parser may rely on annotations from the dependency parser when clustering references to the same entity. In some implementations, when processing a particular natural language input, one or more components of the input processing engine 122 may use relevant data beyond related previous inputs and / or the particular natural language input to determine one or more annotations. For example, a first user in a message exchange thread may provide the input "I will leave the key under the doormat for you," and a second user may provide the input "Remember that" to an automated messaging system via the message exchange thread. When processing "Remember that," the referential parser may resolve "that" to "I will leave the key under the doormat for you," and may further resolve "I will" to the first user and "you" to the second user.

[0050] Scene engine 124 determines the connection from client computing device 106 1-N One or more context features associated with natural language input received by a client computing device. In some embodiments, the context engine 124 determines one or more context features independently of the words in the natural language input. For example, a context feature may indicate the date and / or time when the natural language input was received, and may be determined independently of the words in the natural language input. Furthermore, for example, a context feature may indicate the user's location at the time the natural language input was received (e.g., based on GPS or other location data), and may be determined independently of the words in the natural language input. Moreover, for example, a context feature may indicate one or more additional users to whom the natural language input is directed and / or one or more additional users who are communicating with the user at the time the natural language input was received, and may be determined independently of the words in the natural language input. In some embodiments, the context engine 124 determines one or more context features independently of any explicit input from the user providing the natural language input (such as explicit user interface input defining the context feature and / or explicit input confirming that the automatically determined context feature is correct).

[0051] Item generation engine 126 responds to user input via client computing device 106 1-NThe entry generation engine 126 generates personal database entries for a user by using at least some examples of natural language input produced by a corresponding client computing device. When generating entries for the received natural language input, the entry generation engine 126 may use one or more words from the natural language input, annotations of the natural language input provided by the input processing engine 122, and / or context features provided by the context engine 124.

[0052] As another example, suppose the natural language input of "Remember my locker combination 10, 20, 30" is provided by the user via client computing device 106 1-N A client computing device is manufactured and provided to the automated personal assistant 120. Further, assume that the context engine 124 determines that natural language input is provided at 6:00 AM on 5 / 1 / 16, while the user's location is "at the gym". The user's location "at the gym" can be defined using any one or more of various granularity levels, such as specific latitude / longitude, street address, entity identifier identifying a specific gym, entity class identifier for "gym", etc. In this example, the entry generation engine 126 can generate an entry including: the words "locker combination is 10, 20, 30"; time metadata: time metadata for 6:00 AM, "early morning" and / or some other indication of 6:00 AM and some other identifier of "May 1st", "Sunday in May" and / or "5 / 1 / 16"; and location metadata as an indication of the user's location "at the gym".

[0053] In some implementations, the entry generation engine 126 may generate personal database entries based on received natural language input in response to natural language input having one or more characteristics. For example, in some implementations of these embodiments, the entry generation engine 126 may generate personal database entries in response to natural language input that includes one or more keywords (e.g., “remember,” “remember this,” and / or “don’t forget”) and / or includes these keywords at certain locations in the natural language input (e.g., as a prefix). Moreover, for example, in some implementations of these embodiments, the entry generation engine 126 may additionally and / or alternatively generate personal database entries in response to providing at least some of the natural language input to a classifier and receiving natural language input indicating an expectation to remember the natural language input and / or indicating an expectation at least to a threshold degree as output from the classifier, which is trained to predict whether and / or to what extent the natural language input indicates an expectation to “remember” the natural language input.

[0054] In some implementations, the entry generation engine 126 may also generate “memory-based” metadata for the entry, which indicates the likelihood that the automated personal assistant 120 expects to subsequently recall at least some of the natural language input used to generate the entry. For example, the entry generation engine 126 may generate memory-based metadata that is a first value (e.g., “1”) when the natural language input includes one or more keywords (e.g., “remember”, “remember this”, and / or “don’t forget”) and a second value (e.g., “0”) when the natural language input does not include any of these keywords. Furthermore, for example, the entry generation engine 126 may generate memory-based metadata selected from one of three or more values ​​based on one or more characteristics of the natural language input and / or associated contextual features. For example, entry generation engine 126 may provide a classifier with at least some and / or one or more contextual features from the received natural language input, the classifier being trained to predict whether and / or to what extent the natural language input indicates an expectation to "remember" the natural language input, and receive a measure of the likelihood that the natural language input indicates an expectation to remember the natural language input as output from the classifier. Engine 126 may use this measure or another indication based on the measure as memory metadata.

[0055] Personal database entries generated by entry generation engine 126 can be stored in personal database 152. Personal database 152 can be located on one or more non-transitory computer-readable media, such as client computing device 106. 1-N One of the client computing devices is local, the automated personal assistant 120 is local, and / or remote from the client computing device 106. 1-N And / or one or more media of the automated personal assistant 120. Each personal database entry in the personal database 152 may include the underlying content itself and / or one or more index entries linked to the underlying content. In some embodiments, the personal database 152 may be located on the client computing device 106. 1-N The client computing device is local to one of the client computing devices, and the client computing device may implement one or more (e.g., all) aspects of the automated personal assistant 120. For example, the client computing device may include one or more user interface input devices (e.g., microphones), one or more user interface output devices (e.g., speakers), and a personal database 152.

[0056] Personal database 152 enables the searching of personal database entries to determine whether all or all aspects of an entry are relevant to another natural language input by the user described herein. In some embodiments, personal database 152 may be provided in multiple iterations, each iteration being specific to a particular user and may include personal database entries that are accessible to the user and / or systems specified by the user, but inaccessible to multiple additional users and / or systems other than the user. For example, personal database 152 may be accessible only to that user, and inaccessible to any other user. Moreover, for example, the user and certain additional users specified by the user may have access to personal database 152. In some embodiments, personal database 152 may be associated with access information to allow access to the content by users authorized to do so and to prevent access by all other users. In some embodiments, personal database 152 includes personal database entries for multiple users, each entry and / or group of entries being associated with access information to allow access only to users and / or systems authorized to access these index entries and to prevent access by all other users and / or systems. Thus, personal database 152 may include entries for multiple users, but each entry and / or group of entries may include access information to prevent access by any user not authorized to access these entries. Additional and / or alternative techniques may be used to restrict access to entries in the personal database 152.

[0057] Search engine 128 responds to user input via client computing device 106 1-N The search engine 128 uses at least some instances of natural language input made by a corresponding client computing device to search for entries in a personal database 152 associated with the user. The search engine 128 generates one or more search parameters based on the generated natural language input and optionally on annotations provided by the input processing engine 122, and uses these search parameters to determine one or more entries in the personal database 152 that include content (e.g., descriptive metadata and / or natural language input words) that matches (e.g., all) one or more of the search parameters (e.g., soft matches and / or exact matches).

[0058] The search parameters generated by the search engine 128 may include, for example, one or more words input in natural language, synonyms of one or more words input in natural language, semantic tags of one or more words input in natural language, positional restrictions based on one or more words input in natural language, and time restrictions based on one or more words input in natural language. For example, for the natural language input "Who did I give my bicycle helmet to last month?", the search engine 128 can determine the search parameters as "bicycle helmet" (a word based on the natural language input), "bicycle helmet" (a synonym based on the natural language input), semantic tags indicating "person" (based on the presence of "who" in the natural language input), and time restrictions corresponding to "last month".

[0059] In some implementations, search engine 128 searches personal database 152 based on natural language input having one or more characteristics. For example, in some implementations of these implementations, search engine 128 may search personal database 152 in response to natural language input that includes one or more keywords (e.g., "what", "who", "where", "when", "my", and / or "I") and / or includes these keywords at certain locations in the natural language input. For example, search engine may search personal database 152 in response to natural language input that includes query terms (e.g., "what", "who", or "where") that have syntactic relations with one or more words (e.g., as indicated by annotations of input processing engine 122). Moreover, for example, in some implementations, the search engine 128 may additionally and / or alternatively respond to providing at least some of the natural language input to a classifier and receiving natural language input indicating expectation and / or indication of expectation at least to a threshold degree as output from the classifier to search the personal database 152, the classifier being trained to predict whether and / or to what extent the natural language input is intended as a request for personal information from a user's personal database.

[0060] In some implementations, the search engine 128 may also respond to requests from users via client computing device 106. 1-NThe search engine 128 searches public database 154 using at least some examples of natural language input made by a corresponding client computing device. The search engine 128 may search both public database 154 and personal database 152 and / or only one of public database 154 and personal database 152 in response to some natural language input. Typically, public database 154 may include various types of content, not limited to users and / or other users and / or systems specified by the user, and may optionally be used by search engine 128 in response to various types of natural language input. For example, public database 154 may include content items from publicly accessible documents on the World Wide Web, scripted responses to certain natural language inputs, etc. As an example, public database 154 may include a definition of “smart” so that search engine 128 can recognize the definition in response to natural language input such as “What does smart mean?”. As another example, public database 154 may include a scripted response such as “Hello, how are you?” recognized by search engine 128 in response to natural language input such as “Hi” or “Hello”.

[0061] Ranking engine 130 may optionally rank one or more entries (personal and / or public) identified by search engine 128 based on one or more criteria. In some implementations, ranking engine 130 may optionally use one or more criteria to calculate a score for each of the one or more entries in response to a query. An entry's score may be used on its own as a ranking, or it may be compared with the scores of other entries (where other entries are identified) to determine the entry's ranking. Each criterion used to rank the entries provides information about the entry itself and / or the relationship between the entry and the natural language input.

[0062] In some implementations, the ranking of a given entry can be based on the consistency between the semantic tags of the entry's descriptive metadata and the semantic tags generated for the natural language input. For example, if the semantic tags of one or more words in the natural language input match the semantic tags in the descriptive metadata of a given entry, it can positively influence the ranking of the given entry.

[0063] In some implementations, the ranking of a given entry may additionally and / or alternatively be based on the consistency between the location metadata of the given entry and the user's current location (e.g., determined by the context engine 124) when natural language input is received. For example, if the user's current location is the same as or within a threshold distance of the location metadata of the given entry, the ranking of the given entry can be positively influenced.

[0064] In some implementations, the ranking of a given entry may additionally and / or alternatively be based on the consistency between the given entry's time metadata and the current date and / or time at the time the natural language input is received (e.g., determined by the context engine 124). For example, a given entry may receive a greater boost if the current date is within a threshold of the date indicated by the given entry's time data than if the current date is not within the threshold.

[0065] In some implementations, the ranking of a given entry may additionally and / or alternatively be based on the memory metadata of the given entry. For example, if a metric included in or defined by the memory metadata meets a threshold, the ranking of a given entry may receive a greater boost than if the metric does not meet the threshold. In some implementations where search engine 128 identifies a first entry from public database 154 and a second entry from personal database 152, the ranking of entries may be based on whether and / or to what extent the natural language input is intended as a request for personal information from a user's personal database (e.g., as described above with respect to search engine 128). For example, if the characteristics of the natural language input indicate a high probability of being a request for personal information from a user's personal database, any entry identified from public database 154 may be downgraded.

[0066] Output generation engine 132 utilizes the content of one or more entries identified by search engine 128, and optionally utilizes the ranking of these entries determined by ranking engine 130, to generate output to be presented to the user as a response to the provided natural language input. For example, output generation engine 132 may select a subset (e.g., one) of the identified entries based on the ranking and may generate output based on the selected subset. Moreover, for example, output generation engine 132 may additionally and / or alternatively utilize ranking to determine the presentation order of multiple pieces of content in the generated response (based on the ranking of the entries from which multiple pieces of content are derived).

[0067] When generating output based on a given personal database entry, the output generation engine 132 may include one or more words from the given entry derived from the natural language input for which the given entry was generated and / or may include output based on descriptive metadata of the generated entry.

[0068] Now go to Figure 2A and Figure 2B Additional descriptions of the various components of the Automated Personal Assistant 120 are provided.

[0069] exist Figure 2AIn this embodiment, a user uses one or more user interface input devices of computing device 1061 to provide natural language input 201A to message exchange client 1071, which transmits the natural language input 201A to input processing engine 122. Natural language input 201A can be free-form input as described herein, and can be, for example, typing input provided by the user via a physical or virtual keyboard of the computing device, or voice input provided by the user to a microphone of the computing device. In embodiments where natural language input 201B is voice input, the computing device and / or input processing engine 122 may optionally transcribe it into text input.

[0070] The input processing engine 122 processes the natural language input 201A and generates various annotations for the natural language input. The input processing engine 122 provides annotation input (e.g., words of the natural language input 201A and the generated annotations) 202A to the entry generation engine 126.

[0071] The context engine 124 determines contextual features associated with the received natural language input 201A. For example, the context engine 124 may determine contextual features based on data provided by the messaging client 1071 (e.g., timestamps, location data) and / or the operating system of other applications and / or the client computing device 1061. In some embodiments, contextual features 204 may be used by the entry generation engine 126 described below and / or may be used by the input processing engine 122 (e.g., to improve entity parsing and / or other determinations based on user location and / or other contextual features).

[0072] The entry generation engine 126 generates personal database entries 203 for the user and stores them in the personal database 152. When generating entries 203, the entry generation engine 126 may use one or more words from natural language input 201A, annotation input 202A from natural language input provided by input processing engine 122, and / or context features 204 provided by context engine 124.

[0073] Figure 2B The illustration shows an example of generating natural language output in response to another form of natural language input from a user, based on one or more personal database entries of the user. Figure 2B Examples can occur Figure 2A Following the example, and can be used in Figure 2B Entry 203 was generated in the middle.

[0074] exist Figure 2BIn this context, a user uses one or more user interface input devices of computing device 1061 to provide another natural language input 201B to message exchange client 1071, which transmits the natural language input 201B to input processing engine 122. The natural language input 201B can be any form of input described herein, and can be, for example, typing input provided by the user via a physical or virtual keyboard of the computing device, or voice input provided by the user to a microphone of the computing device.

[0075] The input processing engine 122 processes the natural language input 201A and generates various annotations for the natural language input. The input processing engine 122 provides the annotated output (e.g., words of the natural language input 201B and the generated annotations) 202B to the search engine 128.

[0076] Search engine 128 generates one or more search parameters based on annotation input 202B and uses these search parameters to determine one or more entries in personal database 152 and / or public database 154, including content that matches one or more of the search parameters. As described herein, in some embodiments, search engine 128 may search only personal database 152 in response to natural language input 201B based on one or more characteristics of natural language input 201B. Search engine 128 determines entries 203 in personal database 152 in response to a search by matching one or more search parameters with the content of entry 203.

[0077] Search engine 128 provides entry 203 to ranking engine 130, which optionally generates a ranking for entry 203. For example, ranking engine 130 may determine whether the ranking score of entry 203 meets a threshold that ensures output generation engine 132 uses entry 203 to generate natural language output 205. Context engine 124 determines contextual features associated with the received natural language input 201B and may provide these contextual features to ranking engine 130 for generating the rankings described herein. In some embodiments where search engine 128 determines multiple response entries, ranking engine 130 may rank these entries relative to each other.

[0078] Output generation engine 132 uses the content of entry 203 to generate natural language output 205 to be presented to the user as a response to natural language input 201B. For example, output generation engine 132 may generate output 205 that includes one or more words from entry 203 of natural language input 201A and / or includes one or more words based on descriptive metadata of the generated entry 203. Natural language output 205 is provided to message exchange client 1071 for audible and / or visual presentation to the user via the user interface output device of client computing device 1061. In some embodiments where search engine 128 identifies multiple response entries, output generation engine 132 may select a subset (e.g., one) of the identified entries based on ranking and may generate output based on the selected subset. Moreover, for example, output generation engine 132 may additionally and / or alternatively utilize ranking to determine the presentation order of multiple pieces of content in the generated response (based on the ranking of the entries from which multiple pieces of content are derived).

[0079] Now go to Figures 3A to 3G Additional descriptions of the various components and technologies described herein are provided. Figure 3A The illustration shows a display screen 140 according to an embodiment described herein. Figure 1 The client computing device 1061 displays an example of a dialogue that may occur between a user and an automated personal assistant 120. Specifically, the dialogue includes natural language input 380A, 380B, and 380C provided by the user to the automated personal assistant, and responses 390A, 390B, and 390C that may optionally be provided by the automated personal assistant 120 to confirm that the assistant has responded to the input generated entry.

[0080] In some other implementations, alternative responses may be provided by the automated personal assistant 120. For example, in some implementations where instructions from the natural language input may be necessary and / or beneficial in determining a more specific entry, the automated personal assistant may respond with prompts to seek such further instructions. For example, in response to natural language input 380A, the automated personal assistant 120 may identify both “Bob Smith” and “Bob White” in the user’s contact entries and may generate prompts such as “Do you mean Bob Smith or Bob White?” to request further input from the user to determine a specific contact and further refine the entry based on input 380A. Moreover, for example, in response to natural language input 380D, the automated personal assistant 120 may determine that input 380D contradicts the previous input 380B, and the assistant 120 may prompt “Should this replace the bicycle lock combination you provided on 5 / 5 / 16?” to determine whether the entry generated based on input 380D should replace the entry generated based on input 380D.

[0081] Figure 3A The date and time for each of the various inputs (380A to 380D) provided by the user are also shown in parentheses. As described herein, dates and times can be used when generating descriptive metadata. It is optional that the date and / or time not be displayed on the user's display, but for clarity, they are shown in [the parentheses / instructions]. Figure 3A The diagram shows the date and / or time.

[0082] The display screen 140 further includes a text response interface element 384 that the user can choose to generate user input via a virtual keyboard, and a voice response interface element 385 that the user can choose to generate user input via a microphone. The display screen 140 also includes system interface elements 381, 382, ​​and 383 that can interact with the user to cause the computing device 1061 to perform one or more actions.

[0083] Figure 3A , Figure 3B , Figure 3C , Figure 3D , Figure 3E , Figure 3F and Figure 3G A portion of an example client computing device 1061 having a display screen 140 according to an embodiment described herein is illustrated, the display screen 140 showing what can happen. Figure 3A The dialogue that follows the first dialogue.

[0084] exist Figure 3BIn the input, the user then provides another natural language input, "Who took my bicycle helmet?", 380E. The automated personal assistant 120 provides a response, 390E, "On May 1st, you told me Bob took your bicycle helmet." The response 390E can be generated by: issuing a search based on "Where is my bicycle helmet?"; and, as a response to that search, identifying information based on... Figure 3A The entry created from user input 380A; and the response 390E generated using the entry. For example, the entry generated for user input 380A may include the phrase “Bob took my bicycle helmet” and descriptive metadata: a semantic tag for “person” associated with “Bob”, a semantic tag for “object” associated with “bicycle helmet”, and time metadata indicating May 1st and / or 8:00 AM. Moreover, for example, a search may be issued based on user input 380E, such as a search that includes a first search parameter of “bicycle helmet” and a second search parameter of the semantic tag of “person” (based on the presence of “who” in user input 380E). The entry for input 380A may be identified based on words that match the first search parameter and semantic tags that match the second search parameter. The response 390E may be generated to include time metadata (“on May 1st”) from the entry for input 380A and text (“Bob took your bicycle helmet”) from the entry for input 380A.

[0085] exist Figure 3C In the process, the user then provides another natural language input, "When will I install the tires on my bicycle?", for input 380F. The automated personal assistant provides a response, "May 7, 2016," for input 390F. The response 390F can be generated by: issuing a search based on "When will I install the tires on my bicycle?"; and, as a response to that search, identifying... Figure 3AThe entry is created from user input 380B; and the entry is used to generate a response 390F. For example, an entry generated for user input 380B may include the phrase “New tires were installed on my bicycle on May 7, 2016”. Based on the entry generation engine 126 replacing “today” with the date the user input 380B was received (e.g., based on a note provided by the input processing engine 122), the entry may include the phrase “May 7, 2016” instead of “today” (as provided in user input 380B). The entry generated for user input 380B may also include descriptive metadata: time metadata indicating May 7, 2016, and a semantic tag of “when” associated with the phrase “May 7, 2016” in the entry. Moreover, for example, a search may be issued based on user input 380F, such as a search that includes a first search parameter of “tires”, a second search parameter of “bicycle”, and a third search parameter of the semantic tag of “when” (based on the presence of “when” in user input 380F). Entries generated for input 380B can be identified based on words that match the first and second search parameters and semantic tag metadata that match the third search parameter. Response 390F can be generated based on these words associated with the semantic tag metadata that matches the third search parameter, to include only the word "May 7, 2016" from the entries in input 380B.

[0086] Figure 3D and Figure 3E The diagram illustrates another natural language input, 380G1, and 380G2, for the question "What is my bicycle lock combination?". For user input 380B (… Figure 3A The first entry generated and the 380D input for user input ( Figure 3A The generated second entry can be identified based on a search of the natural language inputs 380G1 and 380G2. Figure 3D In this context, without providing any output generated based on the first entry (380B), a response 390G generated based on the second entry (380D) is provided. This can be based on ranking the first and second entries according to one or more criteria and determining that the second entry ranks higher than the first entry. For example, based on the fact that the time metadata of the second entry is closer to the current time than the time metadata of the first entry, the second entry can be ranked higher.

[0087] exist Figure 3DIn the example, a response 390G is generated based on a second entry of user input 380D without providing any output generated based on the first entry of user input 380B. Providing content from only one of the two responses can be based on various factors (such as system settings, the size of display 140) and / or on the processing of input 380G1 to understand that only a single response is expected. However, in some implementations, the response may include output generated based on both the first and second entries.

[0088] For example, in Figure 3E In the process, a response 390H is generated based on a second entry based on user input 380D and a first entry based on user input 380B. In response 390H, the portion of response 390H corresponding to the second entry is presented before the portion based on the first entry. This can be based on, for example, that the second entry ranks higher than the first entry (e.g., the time metadata based on the second entry is closer to the current time than the time metadata based on the first entry).

[0089] exist Figure 3F The user then provides another natural language input 380H: "How do I install the tire on my bicycle?". The automated personal assistant 120 provides a response 390I, which includes a portion of the steps for installing the new tire on the bicycle. Response 390I is an example of a response that can be generated by the automated personal assistant based on public entries in the public database 154. While the other natural language input 380H includes, it can also include, based on user input 380C (… Figure 3A The generated personal entry contains the words "tire" and "bicycle," but the automated personal assistant 120 does not provide response content based on the personal entry. This can be based on one or more of various factors. For example, the semantic tags of input 380H (e.g., tags associated with "how") may not match the semantic tags of the personal entry generated based on user input 380C. This may result in the inability to determine the entry and / or may negatively impact the entry's ranking. As another example, the public entry used to generate response 390I may be ranked higher than the personal entry, and response 390I generated solely based on public entries according to rules for a specified response is generated solely based on the highest-ranked entry and / or solely based on the highest-ranked entry whose ranking score meets a threshold relative to the next highest-ranked entry.

[0090] exist Figure 3G In the next step, the user provides another natural language input 380I: "Put the new tires on my bicycle?". The automated personal assistant 120 provides a response 390J, which includes content from a private database entry and content from a public database. Specifically, response 390J includes content based on input 380C (…). Figure 3A The content of the generated personal entry (“You installed a new tire on your bicycle on May 7, 2016”). Response 390J also includes content that can be based on public entries from public database 154 (“Do you want to learn how to install a new bicycle tire?”). The automated personal assistant 120 can include entry-based content generated for input 380C based on matching the search parameters “new,” “tire,” and “bicycle” with the content of the entry. In some implementations, content from a private database can be presented before content from a public database based on the ranking of entries at the underlying content level.

[0091] Figure 4A A client computing device 1061 with a display screen 140 is illustrated according to an embodiment described herein. The display screen 140 shows another example of a dialogue that may occur between a user of the client computing device 1061 and an automated personal assistant 120. Specifically, the dialogue includes natural language input 480A and 480B provided by the user to the automated personal assistant 120, and responses 490A and 490B that may optionally be provided by the automated personal assistant 120 to confirm that the assistant 120 has responded to the input generation entry. Figure 4A The date and time for each of the user-provided inputs 480A and 480B are also shown in parentheses. As described herein, dates and times can be used when generating descriptive metadata. It is optional that the date and / or time not be displayed on the user's monitor, but for clarity, they are shown in [the parentheses]. Figure 4A The diagram shows the date and / or time.

[0092] The display screen 140 further includes a text response interface element 484 that the user can choose to generate user input via a virtual keyboard, and a voice response interface element 485 that the user can choose to generate user input via a microphone. The display screen 140 also includes system interface elements 481, 482, and 483 that can interact with the user to cause the computing device 1061 to perform one or more actions.

[0093] exist Figure 4A In the example, the user provides initial user input 480A including "My locker combination is 7, 22, 40", and provides it at "6:00 AM" on "5 / 1 / 16" at the "Gym". A first entry can be created for this initial user input, which includes one or more words from the input, location metadata indicating the "Gym", and time metadata indicating 5 / 1 / 16 and / or 6:00 AM. Figure 4BIn this context, the user provides a second user input 480B including "My locker password is 40, 20, 10", and provides it at "8:00 AM" on "5 / 10 / 16" when "working". A second entry can be created for the second user input, which includes one or more words from the input, location metadata indicating "working", and time metadata indicating 5 / 10 / 16 and / or 8:00 AM.

[0094] Figure 4B and 4C A portion of an example client computing device 1061 having a display screen 140 according to an embodiment described herein is illustrated, the display screen 140 showing what can happen. Figure 4A The dialogue that follows the first dialogue.

[0095] exist Figure 4B In the middle, the user provides another input, 480C, for "What is my locker password?", and provides it at 5:45 AM when the user is at the gym. The automated personal assistant 120 provides based on... Figure 4A The response 490C is generated from the first entry of user input 480A, and this response 490C does not include any content generated based on the second entry of user input 480B. This can be based on, for example, the consistency of the location metadata of the first entry with the user's current location and / or the consistency of the time metadata with the current time. This is compared to the case where the user is "working" at "8:00 AM". In this case, the response can be generated based on the second entry and independently of the first entry.

[0096] exist Figure 4C In the process, the user provides another input 480D for "What is my locker password?", which is different from the input 480D. Figure 4B The input is the same as 480C. However, in Figure 4C In the context of the system, when the user is "at home," the user provides input at 6:00 PM (480D). The automated personal assistant 120 provides input based on... Figure 4A The response 490D is generated from both the first entry of user input 480A and the second entry of user input 480B. Generating a response based on two entries instead of just one can be based on, for example, the user's current location not being perfectly consistent with the location metadata of either entry and / or the current time not being perfectly consistent with the time metadata of either entry. In response 490D, if the date based on the time metadata of the second entry is closer to the current date than the date based on the time metadata of the first entry, the content generated based on the second entry can be presented before the content generated based on the first entry.

[0097] although Figures 3A to 3G and Figures 4A to 4CThe illustration depicts an example graphical user interface that can be provided to a user; however, it is to be understood that in various implementations, responses provided by the automated personal assistant 120 for audible presentation (e.g., via a speaker) may be provided instead of graphical indications of these responses. For example, a user may provide natural language input via a microphone, and the automated personal assistant 120 may respond to such input using output configured for audible presentation.

[0098] Figure 5A and Figure 5B A flowchart illustrating an example method 500 is presented, illustrating how to generate a user's personal database entries based on free-form natural language input created by the user, and how to generate natural language output in response to another free-form natural language input from the user based on one or more of these personal database entries. For convenience, the operations in the flowchart are described with reference to a system performing the operations. This system may include various components of various computer systems, such as an automated personal assistant 120. Furthermore, although the operations of method 500 are shown in a specific order, this is not intended as a limitation. One or more operations may be rearranged, omitted, or added.

[0099] In box 550, the system waits for natural language input.

[0100] In box 552, the system receives natural language input from the user via one or more user interface input devices.

[0101] In box 554, the system determines whether the natural language input in box 552 indicates a request for personal information, such as a request for personal information from a user's personal database. For example, in some implementations, the system may determine that the input indicates a request for personal information if the input includes one or more keywords (e.g., "what," "who," "where," "my," "I," and / or "don't forget") and / or includes these keywords at certain locations in the input. Moreover, for example, in some implementations, the system may additionally and / or alternatively determine that the input indicates a request for personal information based on the output received in response to providing at least some of the input to a classifier trained to predict whether and / or to what extent the natural language input is intended as a request for personal information. In some implementations, box 554 may be omitted (e.g., one or more of boxes 562 to 570 may be performed for each input) or it may be instantiated (first or second) after one or more of boxes 562 to 570 (e.g., the decision for box 554 may be based on whether any entry was identified in box 570 and / or may be based on the content and / or ranking of any returned entries).

[0102] If the system determines that the natural language input in box 552 does not indicate a request for information, the system proceeds to box 556, where the system determines whether the input indicates an expectation of recall. For example, in some implementations, the system may determine that the input indicates an expectation of recall based on the presence of certain words in the input (e.g., "remember," "don't forget") and / or based on the output received in response to providing the input to a classifier trained to predict whether and / or to what extent the natural language input indicates an expectation of "remembering" the natural language input. In some implementations, box 556 may be omitted. For example, one or more of boxes 558 and 560 may be performed for each input or for each input that, in box 554, is determined not to indicate a request for personal information and / or meets one or more other criteria.

[0103] If the system determines in box 556 that the input does not indicate an expectation of recall, it may optionally return to box 550 and wait for another natural language output after providing an incorrect output (e.g., “I do not understand your request”) and / or proceeding to another unillustrated box (e.g., a box for searching a public database and / or providing other outputs).

[0104] If the system determines in box 556 that the natural language input instruction of box 552 is expected to recall, the system proceeds to box 558, where the system generates an entry for the input of box 552 in the user's personal database.

[0105] In some implementations, when executing box 558, the system generates descriptive metadata for the entry in box 558A (e.g., based on the input words and / or contextual features associated with the received input) and stores one or more words of the input from box 552 and the generated descriptive metadata in the entry in box 558B. The system may then optionally perform one or more other actions in box 560, such as providing an output confirming that an entry has been generated for the input, before returning to box 550.

[0106] Returning to box 554, if the system determines in box 554 that the input instruction request of box 552 was received, the system proceeds to box 562. Figure 5B ).

[0107] In box 562, the system determines one or more search parameters based on the input in box 552. In box 564, the system searches the user's personal database based on the search parameters determined in box 562.

[0108] In box 566, the system identifies at least one personal entry in a personal database in response to a search. In box 568, the system generates natural language output based on the content of the at least one personal entry identified in box 566. In some implementations, the system may generate output based only on a personal entry whose ranking meets a threshold. This ranking may be for the personal entry itself, or for the personal entry relative to one or more other personal entries that may be identified in box 566 in response to a search.

[0109] In box 570, the system provides the output of box 568 to be presented to the user via one or more user interface output devices. Then, the system returns to box 550. Figure 5A And wait for another natural language input.

[0110] Although not in Figure 5A and Figure 5B As illustrated in the diagram, in some implementations, the system may optionally search one or more public databases in method 500, and may optionally output content from the public databases. This public database content may be combined with and / or included in the output in place of content from individual entries.

[0111] Figure 6 This is a block diagram of an example computing device 610 that can be optionally used to perform one or more aspects of the techniques described herein. In some embodiments, client computing device 160 1-N One or more of the automated personal assistant 120 and / or other components may include one or more components of the example computing device 610.

[0112] Computing device 610 typically includes at least one processor 614 that communicates with a plurality of peripheral devices via a bus subsystem 612. These peripheral devices may include a storage subsystem 624, including, for example, a memory subsystem 625 and a file storage subsystem 626, a user interface output device 620, a user interface input device 622, and a network interface subsystem 616. The input and output devices allow users to interact with computing device 610. The network interface subsystem 616 provides an interface to an external network and is coupled to corresponding interface devices in other computing devices.

[0113] User interface input device 622 may include a keyboard, pointing device (such as a mouse, trackball, touchpad, or drawing tablet), scanner, touchscreen integrated into the display, audio input device (such as a speech recognition system, microphone), and / or other types of input device. Generally, the term "input device" is used to include all possible types of means and methods for inputting information into computing device 610 or into a communication network.

[0114] User interface output device 620 may include a display subsystem, a printer, a fax machine, or a non-visual display (such as an audio output device). The display subsystem may include a cathode ray tube (CRT), a flat panel device (such as a liquid crystal display (LCD)), a projection device, or other mechanisms for creating visible images. The display subsystem may also provide a non-visual display, such as via an audio output device. Generally, the term "output device" is used to encompass all possible types of means and methods for outputting information from computing device 610 to a user or to another machine or computing device.

[0115] Storage subsystem 624 stores the programming and data construction functions of some or all of the modules described herein. For example, storage subsystem 624 may include functions for performing... Figure 5A and Figure 5B The logic of the selected aspect of the method.

[0116] These software modules are typically executed by processor 614 alone or by processor 614 in conjunction with other processors. Memory 625 used in storage subsystem 624 may include multiple memories, including main random access memory (RAM) 630 for storing instructions and data during program execution and read-only memory (ROM) 632 for storing fixed instructions. File storage subsystem 626 can provide persistent storage for program and data files and may include hard disk drives, floppy disk drives with associated removable media, CD-ROM drives, optical drives, or removable media cartridges. Modules implementing specific implementation functions may be stored in storage subsystem 624 or in other machines accessible by processor 614 by file storage subsystem 626.

[0117] The bus subsystem 612 provides a mechanism for allowing various components and subsystems of the computing device 610 to communicate with each other as intended. Although the bus subsystem 612 is schematically shown as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0118] The computing device 610 can be of different types, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Due to the constantly evolving nature of computers and networks, for the purposes of illustrating some embodiments, Figure 6 The description of the computing device 610 depicted is intended only as a specific example. Figure 6 Compared to the computing device depicted in the diagram, computing device 610 has many other configurations with more or fewer components.

[0119] In the case of a system described herein collecting or utilizing personal information about a user, the system can provide the user with the following opportunities: control whether a program or feature collects user information (e.g., information about the user's social networks, social actions or activities, occupation, user preferences, or the user's current geographic location) or control whether and / or how content that may be more relevant to the user is received from a content server. Furthermore, before storing or using specific data, it can be processed in one or more ways to remove personally identifiable information. For example, the user's identity can be processed to the point that the user's personally identifiable information cannot be determined, or the user's geographic location information (such as city, postal code, or state / county level) can be generalized to make it impossible to determine the user's specific geographic location. Therefore, the user can control how information about themselves is collected and / or how that information is used.

[0120] While several embodiments have been described and illustrated herein, various other means and / or structures may be utilized for performing functions and / or obtaining results and / or one or more of the advantages described herein, and each of such variations and / or modifications is considered to be within the scope of the embodiments described herein. More generally, all parameters, dimensions, materials, and configurations described herein are intended as examples, and actual parameters, dimensions, materials, and / or configurations will depend on the specific application or use of the teachings. Those skilled in the art will recognize, or can determine, several embodiments equivalent to the specific embodiments described herein simply by using conventional experimentation. Therefore, it is to be understood that the foregoing embodiments are merely exemplary, and embodiments may be practiced in ways different from those specifically described and claimed within the scope of the appended claims and their equivalents. Embodiments of this disclosure relate to each individual feature, system, clause, material, apparatus, and / or method described herein. Furthermore, any combination of two or more such features, systems, clauses, materials, apparatus, and / or methods is included within the scope of this disclosure if such features, systems, clauses, materials, apparatus, and / or methods do not contradict each other.

Claims

1. A method executed by one or more processors, comprising: Receive first natural language input, wherein the first natural language input is a free-form input generated by the user via a user interface input device of the user's computing device; Generate an entry for the first natural language input in the user's personal database stored in one or more computer-readable media, wherein generating the entry includes: Generate descriptive metadata for the entry, wherein generating the descriptive metadata for the entry includes: Generate semantic labels associated with entities based on one or more words from the first natural language input, and Generate additional semantic tags associated with the object based on one or more words from the words in the first natural language input; and The descriptive metadata is stored in the entry; After receiving the first natural language input, a second natural language input is received, wherein the second natural language input is a free-form input made by the user via the user interface input device or an additional user interface input device of the user's additional computing device, and wherein one or more additional words of the second natural language input include at least the semantic tags associated with the entity and the object; At least one search parameter is determined based on the second natural language input, wherein the at least one search parameter includes the object; Search the personal database based on at least one search parameter; The entry is determined to respond to the second natural language input based on the search, wherein determining the entry to respond to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata; and In response to determining the entry in response to the second natural language input: Generate natural language output, the natural language output including one or more natural language output words based on the entry; and The natural language output is provided to the computing device for audible presentation to the user via the user interface output device of the computing device.

2. The method according to claim 1, wherein, Generating semantic tags associated with the entity based on one or more words from the first natural language input includes: Process the first natural language input to classify one or more words from the words in the first natural language input; and The semantic labels associated with entities are generated based on the classification of one or more words in the first natural language input.

3. The method according to claim 2, wherein, The semantic tags associated with the entity include one or more of the following: person semantic tags or location semantic tags.

4. The method according to claim 2, wherein, Generating additional semantic tags associated with the object based on one or more words from the first natural language input includes: The additional semantic tags associated with the object are generated based on the classification of one or more words in the first natural language input.

5. The method according to claim 1, wherein, The first natural language input is received as part of a conversational session with the automated personal assistant, wherein the second natural language input is received as part of a subsequent conversational session with the automated personal assistant, and wherein the subsequent conversational session follows the initial conversational session.

6. The method of claim 1, further comprising: In response to determining the entry in response to the second natural language input: The natural language output is provided to be visually presented to the user via an additional user interface output device of the computing device.

7. The method according to claim 6, wherein, The representation of the natural language output includes a transcription corresponding to the natural language output.

8. The method of claim 1, further comprising: Additional entries are determined in response to the second natural language input based on the search, wherein the determination of the additional entries in response to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata.

9. The method according to claim 8, wherein, Generating the natural language output further includes one or more additional natural language output words based on the additional entries.

10. The method according to claim 8, wherein, The descriptive metadata used to generate the entry further includes generating time metadata indicating one or more of the following: the date on which the first natural language input was received, or the time on which the first natural language input was received.

11. The method of claim 10, further comprising: Both the determined entry and the additional entry are in response to the second natural language input: The entry is selected in response to the second natural language input based on the time metadata.

12. The method according to claim 11, wherein, Selecting the entry in response to the second natural language input based on the time metadata includes: The current date or time at which the second natural language input is received is determined to be consistent with one or more of the following: the date on which the first natural language input is received, or the time at which the first natural language input is received.

13. The method according to claim 8, wherein, The descriptive metadata for generating the entry further includes generating location metadata indicating the location of the user's computing device when the first natural language input is received.

14. The method of claim 13, further comprising: Both the determination of the entry and the additional entry are in response to the second natural language input: The entry is selected in response to the second natural language input based on the location metadata.

15. The method according to claim 14, wherein, Selecting the entry in response to the second natural language input based on the location metadata includes: It is determined that the current position of the user's computing device when the second natural language input is received is the same as the position of the user's computing device when the first natural language input is received.

16. A system comprising: At least one processor; as well as At least one memory storing instructions, which, when executed, cause the at least one processor to: Receive first natural language input, wherein the first natural language input is a free-form input generated by the user via a user interface input device of the user's computing device; An entry for the first natural language input is generated in the user's personal database stored in one or more computer-readable media, wherein the instructions for generating the entry cause the at least one processor to: Generate descriptive metadata for the entry, wherein the instructions for generating the descriptive metadata for the entry cause the at least one processor to: Generate semantic labels associated with entities based on one or more words from the first natural language input, and Generate additional semantic tags associated with the object based on one or more words from the words in the first natural language input; and The descriptive metadata is stored in the entry; After receiving the first natural language input, a second natural language input is received, wherein the second natural language input is a free-form input made by the user via the user interface input device or an additional user interface input device of the user's additional computing device, and wherein one or more additional words of the second natural language input include at least the semantic tags associated with the entity and the object; At least one search parameter is determined based on the second natural language input, wherein the at least one search parameter includes the object; Search the personal database based on at least one search parameter; The entry is determined to respond to the second natural language input based on the search, wherein determining the entry to respond to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata; and In response to determining the entry in response to the second natural language input: Generate natural language output, the natural language output including one or more natural language output words based on the entry; and The natural language output is provided to the computing device for audible presentation to the user via the user interface output device of the computing device.

17. A non-transitory computer-readable storage medium storing instructions, said instructions, when executed, causing at least one processor to perform operations, said operations including: Receive first natural language input, wherein the first natural language input is a free-form input generated by the user via a user interface input device of the user's computing device; Generate an entry for the first natural language input in the user's personal database stored in one or more computer-readable media, wherein generating the entry includes: Generate descriptive metadata for the entry, wherein generating the descriptive metadata for the entry includes: Generate semantic labels associated with entities based on one or more words from the first natural language input, and Generate additional semantic tags associated with the object based on one or more words from the words in the first natural language input; and The descriptive metadata is stored in the entry; After receiving the first natural language input, a second natural language input is received, wherein the second natural language input is a free-form input made by the user via the user interface input device or an additional user interface input device of the user's additional computing device, and wherein one or more additional words of the second natural language input include at least the semantic tags associated with the entity and the object; At least one search parameter is determined based on the second natural language input, wherein the at least one search parameter includes the object; Search the personal database based on at least one search parameter; The entry is determined to respond to the second natural language input based on the search, wherein determining the entry to respond to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata; and In response to determining the entry in response to the second natural language input: Generate natural language output, the natural language output including one or more natural language output words based on the entry; and The natural language output is provided to the computing device for audible presentation to the user via the user interface output device of the computing device.

18. A method executed by one or more processors, comprising: Receive first natural language input, wherein the first natural language input is a free-form input generated by the user via a user interface input device of the user's computing device; Generate an entry for the first natural language input in the user's personal database stored in one or more computer-readable media, wherein generating the entry includes: Generate descriptive metadata for the entry, wherein generating the descriptive metadata for the entry includes: Generate semantic labels associated with objects based on one or more words from the first natural language input, and Generate additional semantic tags associated with location based on one or more words from the words in the first natural language input; and The descriptive metadata is stored in the entry; After receiving the first natural language input, a second natural language input is received, wherein the second natural language input is a free-form input made by the user via the user interface input device or an additional user interface input device of the user's additional computing device, and wherein one or more additional words of the second natural language input include at least the semantic tag associated with the object; At least one search parameter is determined based on the second natural language input, wherein the at least one search parameter includes the object; Search the personal database based on at least one search parameter; The entry is determined to respond to the second natural language input based on the search, wherein determining the entry to respond to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata; and In response to determining the entry in response to the second natural language input: Generate natural language output, the natural language output including one or more natural language output words based on the entry; and The natural language output is provided to the computing device for audible presentation to the user via the user interface output device of the computing device.

19. The method according to claim 18, wherein, Generating semantic tags associated with the object based on one or more words from the first natural language input includes: Process the first natural language input to classify one or more words from the words in the first natural language input; and The semantic labels associated with the object are generated based on the classification of one or more words in the first natural language input.

20. The method according to claim 19, wherein, The semantic tag associated with the object includes a reference to the object.

21. The method according to claim 19, wherein, Generating additional semantic tags associated with the location based on one or more words from the first natural language input includes: The additional semantic label associated with the location is generated based on the classification of one or more words in the first natural language input.

22. The method according to claim 18, wherein, The first natural language input is received as part of a conversational session with the automated personal assistant, wherein the second natural language input is received as part of a subsequent conversational session with the automated personal assistant, and wherein the subsequent conversational session follows the initial conversational session.

23. The method of claim 18, further comprising: In response to determining the entry in response to the second natural language input: The natural language output is provided to be visually presented to the user via an additional user interface output device of the computing device.

24. The method according to claim 23, wherein, The representation of the natural language output includes a transcription corresponding to the natural language output.

25. The method of claim 18, further comprising: Additional entries are determined in response to the second natural language input based on the search, wherein the determination of the additional entries in response to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata.

26. The method of claim 25, wherein, Generating the natural language output further includes one or more additional natural language output words based on the additional entries.

27. The method according to claim 25, wherein, The descriptive metadata used to generate the entry further includes generating time metadata indicating one or more of the following: the date on which the first natural language input was received, or the time on which the first natural language input was received.

28. The method of claim 27, further comprising: Both the determination of the entry and the additional entry are in response to the second natural language input: The entry is selected in response to the second natural language input based on the time metadata.

29. The method according to claim 28, wherein, Selecting the entry in response to the second natural language input based on the time metadata includes: The current date or time at which the second natural language input is received is determined to be consistent with one or more of the following: the date on which the first natural language input is received, or the time at which the first natural language input is received.

30. A system comprising: At least one processor; as well as At least one memory storing instructions, which, when executed, cause the at least one processor to: Receive first natural language input, wherein the first natural language input is a free-form input generated by the user via a user interface input device of the user's computing device; An entry for the first natural language input is generated in the user's personal database stored in one or more computer-readable media, wherein the instructions for generating the entry cause the at least one processor to: Generate descriptive metadata for the entry, wherein the instructions for generating the descriptive metadata for the entry cause the at least one processor to: Generate semantic labels associated with objects based on one or more words from the first natural language input, and Generate additional semantic tags associated with location based on one or more words from the words in the first natural language input; and The descriptive metadata is stored in the entry; After receiving the first natural language input, a second natural language input is received, wherein the second natural language input is a free-form input made by the user via the user interface input device or an additional user interface input device of the user's additional computing device, and wherein one or more additional words of the second natural language input include at least the semantic tag associated with the object; At least one search parameter is determined based on the second natural language input, wherein the at least one search parameter includes the object; Search the personal database based on at least one search parameter; The entry is determined to respond to the second natural language input based on the search, wherein determining the entry to respond to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata; and In response to determining the entry in response to the second natural language input: Generate natural language output, the natural language output including one or more natural language output words based on the entry; and The natural language output is provided to the computing device for audible presentation to the user via the user interface output device of the computing device.

31. A non-transitory computer-readable storage medium storing instructions, said instructions, when executed, causing at least one processor to perform operations, said operations including: Receive first natural language input, wherein the first natural language input is a free-form input generated by the user via a user interface input device of the user's computing device; Generate an entry for the first natural language input in the user's personal database stored in one or more computer-readable media, wherein generating the entry includes: Generate descriptive metadata for the entry, wherein generating the descriptive metadata for the entry includes: Generate semantic labels associated with objects based on one or more words from the first natural language input, and Generate additional semantic tags associated with location based on one or more words from the words in the first natural language input; and The descriptive metadata is stored in the entry; After receiving the first natural language input, a second natural language input is received, wherein the second natural language input is a free-form input made by the user via the user interface input device or an additional user interface input device of the user's additional computing device, and wherein one or more additional words of the second natural language input include at least the semantic tag associated with the object; At least one search parameter is determined based on the second natural language input, wherein the at least one search parameter includes the object; Search the personal database based on at least one search parameter; The entry is determined to respond to the second natural language input based on the search, wherein determining the entry to respond to the second natural language input is based at least in part on matching the at least one search parameter with at least some of the descriptive metadata; and In response to determining the entry in response to the second natural language input: Generate natural language output, the natural language output including one or more natural language output words based on the entry; and The natural language output is provided to the computing device for audible presentation to the user via the user interface output device of the computing device.

Citation Information

Patent Citations

  • System and method of reduction of irrelevant information during search

    CN103838815A

  • Method for adaptive conversation state management with filtering operators applied dynamically as part of a conversational interface

    CN104969173A