Ambiguity resolution using conversation search history
By using previously trained code generators to generate data flow programs containing search history functions, process user verbs and resolve ambiguity, the problem of limited processing capabilities of session computing interfaces in the prior art is solved, and more efficient and flexible user interaction is achieved.
Patent Information
- Application Number
- CN202080052259.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-23
- Filing Date
- 2020-06-01
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2040-06-01
AI Technical Summary
Existing session computing interfaces are limited to processing a predefined set of hard-coded templates, limiting actions that can be performed by the session computing interface, and making it difficult to handle complex or novel user discourses.
By generating a data flow program using a previously trained code generator, processing user utterances, the data flow program contains search history functions for selecting the highest confidence-based disambiguation concept from context-specific dialogue history to resolve ambiguity.
Improves user experience, enhances the responsiveness and efficiency of automation assistants, can handle different user verbals more effectively, and remain robust through error recovery mechanisms when errors are encountered.
Smart Images

Figure CN114127710B_ABST
Abstract
Description
Background Art
[0001] Conversational computing interfaces process user utterances and respond by automatically performing actions, such as answering questions, calling application programming interfaces (APIs), or otherwise assisting the user based on the user utterances. Many conversational computing interfaces are limited to processing a predefined set of hard-coded templates, which limits the actions that can be performed by the conversational computing interface. Summary of the invention
[0002] This summary is provided to introduce a set of concepts in a simplified form, which are further described in the detailed description below. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to implementations that solve any or all disadvantages noted in any part of this disclosure.
[0003] A method includes identifying a user utterance that includes an ambiguity. The method also includes generating a dataflow program including a search history function based on the user utterance using a previously trained code generator. The search history function is configured to select a highest confidence disambiguation concept from one or more candidate concepts stored in a context-specific conversation history. BRIEF DESCRIPTION OF THE DRAWINGS
[0004] Figures 1A-1D An exemplary data flow architecture for an automated assistant is shown.
[0005] Figure 1E An exemplary conversation between a user and an automated assistant is shown.
[0006] Figure 2 A method for processing user utterances containing ambiguity is shown.
[0007] Figures 3A-3C An exemplary user utterance containing ambiguity and a corresponding data flow program for processing the exemplary user utterance are shown.
[0008] Figure 4 A method of handling errors during processing of a user utterance is shown.
[0009] Figures 5A-5F An exemplary user utterance, a corresponding data flow program for processing an exemplary user utterance that may result in an error, and an alternative data flow program for resolving the error are shown.
[0010] Figure 6 An exemplary computing system is shown. DETAILED DESCRIPTION
[0011] A conversational computing interface may be used to interact with a user via natural language (e.g., via voice and / or submitted text). As an example, an automated assistant may be used to assist a user via natural language interaction. Although the present disclosure uses an automated assistant as an exemplary conversational computing interface, this example is non-limiting, and a conversational computing interface may be implemented in accordance with the present disclosure for any suitable purpose, e.g., allowing a user to interact with any suitable computer hardware and / or computer software using natural language. Thus, each reference to an automated assistant in the present disclosure applies equally to any other conversational computer interface or other computing framework configured to respond to voice or text input.
[0012] The automated assistant can use natural language processing (NLP) techniques (e.g., machine learning classifiers) to process input user utterances (e.g., user voice and / or submitted text) to perform predefined, hard-coded actions associated with the input user utterances. For example, the automated assistant can support a predefined plurality of hard-coded templates, wherein each template has a plurality of time slots that can be filled in to parameterize the hard-coded actions. As an example, the automated assistant can support predefined interactions to call an application programming interface (API), such as to reserve a seat at a restaurant, call a ride-hailing service, or check the weather. However, although the automated assistant can support multiple different predefined actions via predefined templates, an automated assistant that only supports predefined actions via templates may not be configured to perform more complex or novel behaviors.
[0013] The present disclosure relates to an automated assistant that uses a dataflow program in a dataflow programming language to process a user utterance (e.g., in addition to or in lieu of using a template). The automated assistant uses a previously trained code generator to generate and / or output a dataflow program for a user utterance, wherein the dataflow program uses a plurality of predefined functions to define individual steps for processing the user utterance.
[0014] Processing user utterances using a data flow program generated by a previously trained code generator can result in an improved user experience, improved efficiency (e.g., improved storage usage and / or improved processing time) of an automated assistant or other interactive computer service, and / or an improved ability to respond to different user utterances. As an example, the data flow program can encode various different processing strategies for different user utterances, including: performing calculations based on the user utterances, accessing APIs for responding to the user utterances, etc. The code generator can generate a data flow program specific to the user utterance being processed, which can enable the user utterance to be processed more efficiently (e.g., not executing irrelevant steps) and with improved user satisfaction (e.g., by generating a program that better solves the request expressed in the user utterance). In addition, a data flow program that encounters one or more errors can be paused so that the error can be handled using an error recovery program, such as a modified version of the data flow program or a substitute data flow program. In this way, the method of using a previously trained code generator and a data flow program can be more robustly applicable to different situations where errors may occur, because different types of errors can be handled by running an error recovery data flow program.
[0015] therefore, Figure 1A A data flow architecture for an automated assistant system 100 is shown. The automated assistant system 100 is configured to process a user utterance 102 by operating a previously trained code generator 104, which is configured to output a data flow program 106 for the user utterance 102. Although the present disclosure focuses on interactions via natural language (e.g., speech and submitted text), a conversational computing interface such as the automated assistant system 100 may also allow interactions via any other suitable input method, for example, interactions via touch screen input and / or button presses. Similarly, although the present disclosure focuses on processing user utterances 102, other inputs such as button press events may be processed in a similar manner. For example, a previously trained code generator 104 may be configured to output a data flow program 106 for one or more button press events. The previously trained code generator 104 may be trained to recognize different kinds of non-verbal input events, for example, based on the specific button pressed, the timing of the button press relative to other input events, and the like.
[0016] The data flow program 106 is shown as a graph including a plurality of function nodes, wherein the function nodes are depicted with inputs and outputs shown by arrows. The data flow program is configured to produce a return value indicated by the bottom-most arrow. The previously trained code generator 104 is configured to add any of a plurality of predefined functions 110 to the data flow program based on the user utterance. Each predefined function defines one or more individual steps for processing the user utterance 102. The data flow program 106 is executable to cause the automated assistant to respond to the user utterance, for example, by performing any suitable response action. The predefined functions of the data flow program 106 may be executable to cause the automated assistant to perform any suitable response action, for example, outputting a response as speech and / or text (e.g., outputting as in Figure 1A ), calling an API to perform an action using the API (e.g., ordering food from a restaurant, scheduling a ride using a ride-hailing service, scheduling a meeting in a calendar service, making a phone call). Although the present disclosure focuses on examples in which the automated assistant responds to an utterance by outputting an assistant response (e.g., as speech and / or text), these examples are non-limiting, and the predefined functions and data flow programs described herein can be configured to cause the automated assistant to respond to an utterance in any suitable manner, such as by performing one or more actions using an API instead of or in addition to outputting an assistant response.
[0017] The previously described code generator 104 described in this article can be used to respond to user speech in any suitable form. For example, the previously described code generator 104 can be configured to identify user speech, and generate a data flow program that defines an executable plan for responding to the user speech according to the user speech. The resulting data flow program can be run, for example, by an automated assistant to process the user speech. In some examples, the data flow program for responding to the user speech can be run to respond to the user speech without using the code generator 104 to generate any additional code. In other words, the code generator 104 is configured to output a complete plan for processing user speech. Alternatively or additionally, the code generator 104 can be configured to output a data flow program for responding to user speech, wherein some or all of the data flow program is run before the code generator 104 is used to generate additional codes that define additional plans for completing the processing of the user speech. The code generator 104 can be used to plan and run a data flow program in any suitable manner, including completing planning before running and / or interweaving planning and running in any suitable form.
[0018] The previously trained code generator 104 can be based on any suitable technology, such as existing or future machine learning (ML), artificial intelligence (AI), and / or natural language processing (NLP) technology. In some examples, the previously trained code generator 104 includes: an encoder machine configured to encode the user utterance 102 into semantic features (e.g., vectors in a semantic vector space learned by the previously trained code generator 104); and a decoder machine configured to decode the semantic features by outputting one or more functions from a plurality of predefined functions 110. In some examples, the decoder machine is configured to output one or more functions according to a typed syntax for combining functions from a plurality of predefined functions 110, thereby limiting the output of the decoder machine to a well-typed, fully executable data flow program. The plurality of predefined, composable functions 110 support a range of different automated assistant behaviors, such as calling an API, answering questions, looking up user data, and / or utilizing historical context from a context-specific conversation history 130 maintained for the automated assistant. As shown in Figure 1A As shown by the dashed arrows in , the context-specific conversation history 130 can include any suitable aspects of the conversation, such as user utterances 102 , data flow programs 106 , and / or resulting assistant responses 120 .
[0019] Thus, the user utterance 102, the data flow program 106 generated for the user utterance 102, and / or any relevant assistant responses 120 of the automated assistant may be stored in a context-specific dialog history 130. Thus, the context-specific dialog history 130 defines a plurality of concepts (e.g., concept 130A, and any suitable number of concepts including concept 130N). "Concept" is used herein to refer to any relevant or potentially relevant aspect of an interaction between a user and an automated assistant. For example, a concept may include an entity (e.g., a person, a place, a thing, a number, a date), the intent of a user query (e.g., an intent to order food, an intent to check the weather, an intent to schedule a meeting), an action performed by an automated assistant (e.g., ordering food, checking the weather, calling an API, finding information related to an entity, recognizing a particular user utterance, performing a composite action consisting of more than one action), or any other suitable feature. Concepts may be defined in any suitable manner, such as based on the textual content of the user utterance 102. In some examples, concept 130A is defined according to a data flow program fragment 132A. For example, the data flow program fragment 132A may include one or more functions configured to look up information related to a particular entity, and / or one or more functions configured to cause the automated assistant to perform a particular action.
[0020] In some examples, the plurality of predefined functions 110 include one or more history access functions that are configured to access a context-specific conversation history 130. Thus, the data flow program 106 may include such a history access function. For example, both the plurality of predefined functions 110 and the data flow program 106 include a history access function 112 that is configured to access a context-specific conversation history 130 as indicated by an arrow. In some examples, the plurality of predefined functions 110 include a search history function that is configured to search for concepts from the context-specific conversation history 130 (e.g., a getSalient() function that is configured to search for previously discussed entities or search for previously performed actions), such as searching for a previously discussed entity or a previously performed action. Figure 1B , 2 and 3A-3F further discussed. In some examples, a plurality of predefined functions 110 include a program rewrite function, which is parameterized by a specified concept stored in the context-specific dialog history and is configured to generate a new data flow program fragment related to a specified concept based on a concept from a context-specific dialog history 130 (e.g., a Clobber() function, which is configured to generate a new rewrite program based on a specified concept, for example, to write a new program for performing a previous action on a new entity or for performing a new action on a previously discussed entity). In some examples, the specified concept stored in the context-specific dialog history includes a target subconcept, and the program rewrite function is also parameterized by a replacement subconcept for replacing the target subconcept. Therefore, the new data flow program fragment corresponds to a specified concept that has replaced a target subconcept by replacing a subconcept. Specifically, in some examples, the new data flow program fragment includes a rewritten data flow program fragment based on a historical data flow program fragment corresponding to the specified concept, wherein the sub-program fragment corresponding to the target sub-concept is replaced by a different sub-program fragment corresponding to the replacement sub-concept. Figure 1C , 2 3F-3G further discuss the program rewriting function.
[0021] By looking up and / or rewriting concepts from the context-specific dialog history 130, the automated assistant is able to repeat an action, perform an action with modifications, look up related entities or other details related to a previous action, and the like. For example, the automated assistant can repeat an action or perform an action with modifications by re-running code based on a program fragment 132A corresponding to a concept 130A from the context-specific dialog history 130. In addition, if any error conditions are reached during the execution of the data flow program 106 before outputting an assistant response 120, the data flow program 106 and / or other program fragments stored in the context-specific dialog history 130 can be re-run (e.g., with or without modifications) to recover from the error, such as by re-running the program fragment 132A corresponding to the concept 130A from the context-specific dialog history 130. Figure 1D , 4 and 5A-5H as further described.
[0022] therefore, Figure 1B A different view of automated assistant system 100 focusing on search history function 112 among multiple predefined functions 110 is shown. Search history function 112 is configured to handle ambiguous user utterance 102 by resolving any ambiguities using concepts from context-specific dialog history 130. As indicated by the dashed arrow leading from context-specific dialog history 130 to search history function 112, search history function 112 is configured to access one or more concepts from context-specific dialog history 130 to determine disambiguation concepts 134.
[0023] In some examples, the disambiguation concept 134 is defined by a program snippet 136. As an example, if the user utterance 102 refers to an ambiguous entity (e.g., by a pronoun or a partial name, such as "Tom"), the search history function 112 can be configured to search the ambiguous entity in the context-specific conversation history 130 (e.g., based on the partial name "Tom") to find a clarifying entity that matches the ambiguous entity (e.g., an entity with a name "Tom" or a related name such as "Thomas"). In some examples, the clarifying entity can be defined by a full name (e.g., "Thomas Jones"). In other examples, the clarifying entity can be defined by code configured to find the clarifying entity (e.g., code that finds a person named "Tom" in a user's address book).
[0024] In some examples, the disambiguation concept 134 indicates a bridging backreference between a user utterance and a previous user utterance from the context-specific conversation history. As an example, if a user asks "When am I having lunch with Charles?" and then follows up with "How long does it take to get there?" the word "there" in the subsequent user utterance refers to the location of having lunch with Charles. Thus, the search history function 112 may be configured to search for concepts corresponding to the location of having lunch with Charles. For example, the concept may include a data flow program fragment that recursively uses the search history function 112 to find a salient event (i.e., meeting with Charles), and further instructions for obtaining the location of the salient event. More generally, the concept found by the search history function 112 may include a data flow program that recursively calls the search history function 112 to find any suitable sub-concepts, for example, to define a salient concept based on searching for other salient sub-concepts.
[0025] About Figure 2 and Figures 3A-3C The search history function 112 is further described.
[0026] Figure 1C Another different view of the automated assistant system 100 focusing on a program rewrite function 112' is shown. The program rewrite function 112' is configured to generate a rewrite concept 140 (e.g., defined by a rewrite program snippet 142) by starting with a specified concept 152 from a context-specific dialog history 130, wherein the specified concept 152 includes at least one replacement target sub-concept 154 to be modified / replaced with a different replacement sub-concept 156. As an example, the specified concept 152 can refer to an action of the automated assistant, including "make a reservation at a local sushi restaurant." Thus, the target sub-concept 154 can be the description "sushi restaurant" and the replacement sub-concept 156 can be the alternative description "burrito restaurant." Specifically, the specified concept 152 can be defined by a program snippet for calling an API to make a restaurant reservation, wherein the API is parameterized by a program snippet corresponding to the target sub-concept 154 indicating a restaurant for making a reservation. Thus, program rewrite function 112' is configured to output a rewrite concept 140 corresponding to "make a reservation at the local burrito restaurant," such that rewrite concept 140 can be used to perform a new action consisting of making a reservation at the local burrito restaurant (e.g., by running program fragment 142). Figure 2 and Figure 3C The program rewriting function 112' is further described.
[0027] Figure 1DAnother view of the automated assistant system 100 is shown that focuses on handling errors that may occur during the execution of the data flow program 106. Figure 1D Not shown, but as in Figures 1A-1C As shown in , user utterances 102 , data flow programs 106 , and / or assistant responses 120 may be stored in a context-specific conversation history 130 as the user utterance is processed.
[0028] If an error condition is reached during execution of the data flow program 106 before outputting the assistant response 120, the data flow program may be paused and saved as a paused execution 160. The error handling execution engine 170 is configured to recover from the error condition in order to generate the assistant response 120'. For example, the error handling execution engine 170 may implement the method 400, described below with respect to Figure 4 Method 400 is described. In order to recover from the error condition, the error handling execution engine 170 can modify and / or re-execute the suspended operation 160. Alternatively or additionally, the error handling execution engine 170 can run a substitute program fragment 180. For example, the error handling execution engine 170 can replace the program fragment of the data flow program 106 with the substitute program fragment 180 by using a program rewrite function, and / or by running the substitute program fragment instead of the data flow program 106 to modify the suspended operation 160. The substitute program fragment 180 can be derived according to the context-specific dialogue history 130 (for example, the substitute program fragment 180 can be a previously run program fragment), constructed according to a plurality of predefined functions 110 and / or output by a previously trained code generator 104. In some examples, the previously trained code generator 104 is configured to recognize an error condition and output a new program fragment configured to recover from the error. For example, the previously trained code generator 104 can be trained with respect to one or more training examples, each training example including an exemplary error condition and an exemplary data flow program fragment for responding to the error. For example, the previously trained code generation engine 104 may be trained on a large number of annotated dialog histories (as will be described below), where some or all of the annotated dialog histories include occurrences of error conditions.
[0029] In some examples, the error condition may arise due to ambiguity in the user utterance 102, where insufficient information provided by the user utterance 102 is sufficient to fully serve the user based on the user utterance. As an example, if the user utterance is "schedule a meeting with Tom," but there is more than one "Tom" in the user's address book, it may be unclear with whom to schedule a meeting. Therefore, in some examples, the error handling executor 170 is configured to run code to generate an initial assistant response 120' that includes a clarification question 122. The clarification question 122 is output for the user to respond to the new, clarifying user utterance 102. Thus, the clarifying user utterance can be processed using a previously trained code generator 104 to generate a new data flow program 106 for responding to the clarifying user utterance. Error recovery involving clarification questions will be discussed below. Figure 4 and Figure 5A -5H further discussed. The previously trained code generator 104 can be trained via supervised training on a plurality of annotated dialog histories. In many examples, the previously trained code generator 104 is trained on a large number of annotated dialog histories (e.g., hundreds, thousands, tens of thousands, or more). The annotated dialog histories describe state information associated with a user interacting with the automated assistant, annotated with an exemplary data flow program for responding to the user interaction. For example, the annotated dialog histories may include state information associated with the present disclosure (as described below with respect to Figures 3A-3C and Figures 5A-5F A context-specific dialog history annotated with an exemplary data flow program suitable for being run by an automated assistant in response to a context established in the context-specific dialog history is provided. In the example, the context-specific dialog history includes a plurality of events (e.g., timestamped events) arranged in chronological order, including user utterances, data flow programs run by the automated assistant, responses output by the automated assistant, and / or error conditions reached when running the data flow program.
[0030] As a non-limiting example, the annotated conversation history may include a context-specific conversation history, wherein the most recent event is an exemplary user utterance, which is annotated with an exemplary data flow program for responding to the exemplary user utterance with respect to the context established by the context-specific conversation history. Therefore, the previously trained code generator 104 can be trained to reproduce the exemplary data flow program given the exemplary user utterance and the context-specific conversation history. The exemplary data flow program may contain any suitable function in any suitable sequence / arrangement, so that through training, the code generator is configured to output a suitable function in a suitable sequence. For example, the exemplary data flow program may include a search history function, and therefore, the code generator may be trained to output the search history function and other functions in a suitable sequence for responding to the user utterance (e.g., so that the automated assistant performs any suitable response action, such as outputting the response as text and / or voice, or calling an API). In some examples, the annotated conversation history includes a context-specific conversation history, wherein the most recent event is the occurrence of an error condition (e.g., not the user utterance as the most recent event), which is annotated with a data flow program for recovering from the error. Thus, the previously trained code generation engine 104 may be trained on such annotated dialog histories to generate suitable data flow programs for recovering from error conditions.
[0031] The annotated conversation histories may be obtained, for example, from a human presenter in any suitable manner. For example, one or more human presenters may be shown context-specific conversation histories (e.g., context-specific conversation histories derived from usage data obtained from interactions with humans, and / or machine-generated context-specific conversation histories), and for each context-specific conversation history, be asked to provide a suitable data flow program to respond to the context-specific conversation history. The data flow program provided by the human presenter may perform a wide range of tasks in any suitable manner using predefined functions in response to user utterances and / or error conditions.
[0032] For example, based on the illustrated exemplary user utterances or exemplary error conditions, the human presenter may provide an exemplary data flow program that performs any suitable computations, outputs responses (e.g., asks a user a clarifying question, or answers a query issued by a user), listens to utterances from a user (e.g., obtains clarification from a user), calls an API, etc. In addition, the exemplary data flow program may include a search history function (e.g., "getSalient()") and / or a program rewrite function (e.g., "Clobber()") that is called with any suitable parameters, e.g., to execute a data flow program fragment according to the context-specific dialog history. Thus, by training on multiple annotated dialog histories, the code generator 104 may be trained to generate data flow programs similar to those provided by a human presenter in order to respond to user utterances and / or recover from errors.
[0033] The data flow program (e.g., a data flow program generated by a previously trained code generator and / or an exemplary data flow program) is constructed based on a plurality of predefined combinable functions. The predefined combinable functions can be combined into a program that calls the predefined combinable functions in any suitable order and parameterizes the predefined combinable functions in any suitable manner. Therefore, based on the exemplary data flow program provided during training, the previously trained code generator can be trained to output a suitable data flow program for user utterances. The previously trained code generator is not limited to hard-coded behavior. For example, instead of or in addition to responding to exemplary user utterances seen during training, the previously trained code generator is configured to handle novel user utterances (which may not have been provided during training) by generating corresponding, novel data flow programs (which may not have been provided during training).
[0034] In order to generalize from the specific training examples seen during training and to respond to novel user utterances using novel data flow procedures, the data flow procedures may be used in any suitable manner (e.g., as described below with respect to Figure 6As described, the previously trained code generator is trained on any suitable training data, for example, on a large amount of annotated conversation history, using any suitable ML, AI and / or NLP model. In some examples, the previously trained code generator can be trained on a loss function to evaluate whether a data flow program is an appropriate response to a user utterance, wherein the loss function is configured to indicate zero loss or a relatively small loss (e.g., no adjustment is required, or a relatively small adjustment to training parameters) when the code generator successfully reproduces the training example (e.g., given a context-specific conversation history, by producing the same data flow program as provided by a human annotator). However, although the loss function can be configured for a relatively small loss when the code generator successfully reproduces the training example, the loss function can also be configured for a relatively small loss when the code generator reproduces a different data flow program, for example, a data flow program that has a similar effect when run, and / or a data flow program indicated as satisfactory by a human user (e.g., a human presenter and / or an end user of an automated assistant device). These exemplary schemes for generalization during training are non-limiting, and any suitable AI, ML, and / or NLP techniques may be utilized to appropriately train a code generation engine to generate appropriate data flow programs in response to various user utterances.
[0035] In some examples, error conditions can be identified by operating a previously trained error detection model. For example, the previously trained error detection model can be trained via supervised training on a plurality of annotated conversation histories, wherein the annotated conversation histories are annotated to indicate when errors occur. For example, the annotated conversation histories can be obtained by showing context-specific conversation histories to one or more human presenters and asking the human presenters to indicate when the context-specific conversation histories indicate an error state.
[0036] In addition to history access functions such as search history functions and program rewrite functions, the plurality of predefined combinable functions may include any suitable functions, for example, a listening function configured to listen for specific user utterances before continuing execution of a data flow program, a response function configured to output a description of a value determined during execution of a data flow program, and / or a primitive computation function for processing values obtained from user utterances and / or values calculated during execution of a data flow program (e.g., data structure operations, such as forming tuples from data, or arithmetic operations).
[0037] In some examples, the plurality of predefined combinable functions include external functions configured to call external (i.e., third-party) APIs. For example, an external API may be called to interact with a real-world service (e.g., to arrange a car, order a meal, or make a reservation at a restaurant in a ride-hailing service). In some examples, the plurality of predefined combinable functions include reasoning functions configured to perform calculations on the results of external functions. The reasoning function may encapsulate high-level behaviors about the API, which would otherwise require the use of multiple different low-level functions of the API. For example, an external ride-hailing API may support functions for arranging a car, adding stops on a route, and ultimately deciding the route to be arranged. Thus, the reasoning function may be configured to: receive a destination, and arrange a car based on the destination, add stops corresponding to the pickup location for the user, add stops corresponding to the destination, and ultimately decide the route to be arranged including stops corresponding to the pickup location and the destination. By encapsulating high-level behaviors using reasoning functions, the code generator can easily output well-typed code for performing high-level behaviors using an external API without having to use low-level functions of the external API to output individual steps. In some examples, an inference function can be defined with respect to one or more constraints, and running the inference function can include running a constraint satisfaction program to satisfy the one or more constraints, and then calling an external API with parameters defined by a solution to the constraints. In some examples, the one or more constraints can include "fuzzy" or "soft" constraints, and thus, solving the constraints can include running an inference program suitable for "fuzzy" logic reasoning, such as a Markov logic reasoning program.
[0038] In some examples, the plurality of predefined combinable functions include a user-customized function configured to access user-customized settings and perform calculations based on the user-customized settings. For example, the user-customized function can be configured to determine whether a user is free based on a user-customized schedule (e.g., calendar data). The user-customized function can be implemented using an external function configured to call an external API (e.g., an API for finding calendar data).
[0039] In some examples, the plurality of predefined combinable functions include an intelligent decision function, wherein the intelligent decision function is configured to perform calculations using a previously trained machine learning model. As an example, the search history function may be an intelligent decision function configured to search the context-specific conversation history using a previously trained correlation detector. As another example, the plurality of predefined combinable functions may include an intelligent decision function configured to evaluate whether it is "morning" in a user-specific and / or crowd-specific manner, for example, the function may be configured to recognize that the time a user considers to be "morning" may vary depending on the day of the week or the time of the year. For example, the intelligent decision function may be configured to evaluate "morning" between 6AM and 11AM on weekdays, but may be configured to evaluate morning between 9AM and 12PM on weekends. The intelligent decision function may be trained in any suitable manner, such as based on labeled time examples and whether the user considers the time to be morning. As with user-customized functions, the intelligent decision function can take into account auxiliary information such as the user's work schedule, calendar, and / or phone usage, for example to determine whether it is "morning" based on whether the user is likely to have woken up on a given day. In some examples, the intelligent decision function can be configured to evaluate ambiguity (e.g., ambiguous user utterances or ambiguous constraints) and select a disambiguation data flow program to respond to the ambiguity.
[0040] In some examples, the plurality of predefined combinable functions include a macro function, wherein the macro function includes a plurality of other predefined combinable functions and is configured to run the plurality of other predefined combinable functions. For example, macro functions can be used to sort and organize related low-level steps of a high-level task using low-level functions. By encapsulating high-level behaviors using macro functions, the code generator can easily output well-typed code for performing the high-level behaviors without using low-level functions to output individual steps.
[0041] Briefly go to Figure 1E , Figure 1E A first example of a conversation between a user and an automated assistant is shown, wherein the automated assistant processes a user utterance 102' in which there is no ambiguity. A context-specific conversation history 130 is shown as including concepts 130A, 130B, and further concepts up to 130N. However, in this example, the user utterance 102' is unambiguous and can be processed without reference to the context-specific conversation history 130. The user asks "When is my next meeting with Tom Jones?" and gets a previously trained code generator output data flow program 106'.
[0042] The data flow program 106' is shown with a non-limiting example syntax, where square brackets indicate the return value of an expression, for example, [events] indicates the return value of an expression for finding all events that match a set of criteria, and [time] indicates the start time of the first such event in the set. The example syntax includes various functions, including a search history function (e.g., "getSalient()") and a program rewrite function (e.g., "Clobber()"), as well as other functions (such as original functions, API functions, etc., as described herein). The example functions are shown with a function call syntax indicating the name of the function (e.g., "Find", "getSalient", "Clobber", and other named functions) and brackets containing parameters for calling the function. The example function call syntax is non-limiting, and functions can be called and parameterized in any suitable manner (e.g., using any suitable formal language syntax). The example function names represent pre-defined functions with implementations not shown herein. For example, each predefined function can be implemented by any suitable sequence of one or more instructions executable by the automated assistant to perform any suitable steps (e.g., produce the behavior indicated by the function name, the behavior set forth in this disclosure, and / or any other suitable behavior, such as performing a calculation, calling an API, outputting audio, presenting information visually via a display, etc.). For example, a "Find" function can be implemented in any suitable manner, such as by calling an API to find information stored in a user's calendar.
[0043] As shown, the data flow program looks for events with an attendee named "Tom Jones". The data flow program calculates values such as [events] and [time], and then outputs a response 120' using a "describe" function that is configured to output a description of the value [time], for example as speech via a speaker of an automated assistant device. Thus, the response 120' indicates the meeting time for the next meeting with Tom Jones, i.e., "at 12:30".
[0044] Figure 2 An exemplary method 200 for handling ambiguous user utterances is shown. The method uses a search history function to utilize information in a context-specific conversation history to resolve the ambiguity. Figure 3A As shown in , the context-specific dialog history may include any suitable information tracked for a dialog, such as concepts 334 indicating previous user utterances 304 and concepts 324 indicating previously executed dataflow program fragments 326 .
[0045] At 202, method 200 includes identifying a user utterance that contains ambiguity. For example, Figure 3A An exemplary ambiguous user utterance 302 is shown, in which the user asks "what's after that?" The utterance itself does not give enough information for the user to respond, for example, because "that" does not refer to any specific event without further context.
[0046] In some examples, at 204, method 200 includes identifying constraints associated with the ambiguity. Figure 3A , the utterance includes the word "after that", and therefore, method 200 includes a constraint to identify that the user is referring to a related time. For example, the previously trained code generator may be trained on one or more training examples in which the user utterance includes a word related to time, such as "after that" when the user is referring to a time. For example, the time may be a start time defined by a schedule for a scheduled event, such as an event scheduled in a user's calendar.
[0047] Therefore, briefly returning to Figure 2 At 206, method 200 includes using the previously trained code generator to generate a data flow program for responding to the user utterance, wherein the data flow program includes the search history function, which is Figure 3A The search history function may be run to find relevant information from the context-specific conversation history to resolve the ambiguity. As described at 208, the search history function is configured to select a highest confidence disambiguation concept from one or more concepts in the context-specific conversation history to resolve the ambiguity using the highest confidence disambiguation concept.
[0048] For example, in Figure 3A , after the user talks about meeting Tom Jones, the context-specific conversation history 130 includes concepts 334 that include a previous user utterance 304 in which the user asks “When is my next meeting with Tom Jones?” and a data flow program fragment 326 that was previously executed in response to the previous user utterance 304 (e.g., as in Figure 1E324 defined in the data flow program 106' shown in . The data flow program 322 is configured to: determine a specific event (e.g., the next meeting with Tom), determine a start time for the event, and describe the time. Therefore, the search history function "getSalient" is configured to search for a relevant time from the context-specific conversation history, so as to use the relevant time to find an event that occurs after the time and output a response 330 indicating what event occurs after the time. The relevant time can be the [time] value defined in the data flow program 326 related to the meeting with Tom. Therefore, the response 330 indicates that after the meeting with Tom, there is another meeting with Richard Brown.
[0049] Return to Figure 2 In some examples, as described at 210, the search history function is configured to search the context-specific dialog history for a subset of one or more candidate concepts that satisfy the constraints associated with the ambiguity, wherein the highest confidence disambiguation concept is found within the subset. Figure 3A In the example, based on the utterance 302 including the word "after that", the identified constraint may include that the user is referring to the constraint of the relevant time defined by the schedule for the scheduled event (as described above), and therefore the relevant concept from the context-specific dialog history 130 should be a value of the "time" type. Therefore, the search history function "getSalient(Time())" is configured to search for a subset of relevant values of the "time" type from the context-specific dialog history 130, for example, including the [time] value.
[0050] In some examples, as described at 212, the search history function is configured to use a previously trained correlation detector to identify the ambiguous constraints, and select a disambiguation data flow program fragment corresponding to a disambiguation concept based on such identified constraints. Figure 3A As shown in , based on the "time" type constraint, a data flow program fragment 326 can be selected for disambiguation. The previously trained correlation detection engine can include any suitable combination of existing and / or future ML, AI and / or NLP technologies.
[0051] In some examples, the previously trained correlation detector can be trained via supervised training on multiple annotated conversation histories, wherein the annotated conversation histories include unresolved search history functions that are labeled with disambiguation concepts that will resolve the unresolved search history functions. Thus, given an unresolved search history function, the previously trained correlation detector can be trained to predict an appropriate disambiguation concept. For example, the annotated conversation history can include a data flow program fragment that uses the search history function (e.g., a data flow program 322 that includes "getSalient") and an exemplary disambiguation concept that will resolve the ambiguity (e.g., a concept 324 that includes a data flow program fragment 326 that defines [time] = [vents][0].start).
[0052] In some examples, the annotated conversation history can be provided by a human presenter, who can be shown an ambiguous user utterance and a corresponding data flow program including the search history function, and is asked to provide an exemplary data flow program fragment that appropriately matches the search history function in the context of the ambiguous user utterance. In some examples, a disambiguation concept that will resolve the unresolved search history function is selected by a human presenter from the context-specific conversation history. In some examples, the disambiguation concept is an exemplary data flow program fragment received from a human presenter. For example, a human presenter can be asked to provide the data flow program fragment by combining one or more of the predefined combinable functions and / or one or more data flow program fragments from the context-specific conversation history. For example, a human presenter can be asked to select one or more disambiguation concepts from the context-specific conversation history using a graphical user interface (GUI), and / or to combine a new program using such concepts and predefined combinable functions that can be selected from a menu. In some examples, the disambiguation concept is independent of the context-specific conversation history, for example, a human presenter may indicate that the significant date for resolving a search history function constrained to find a date is "today," regardless of the concepts in the context-specific conversation history.
[0053] After generating the data flow program including the search history function, at 214, method 200 optionally further includes running the data flow program. Therefore, at 216, method 200 optionally further includes outputting a response obtained by running the data flow program. For example, Figure 3AAs shown in , the data flow program 322 can be run, including running the search history function to search for significant "time" type values from the history to determine the value [time2], and find the event that starts after [time2] (for example, by calling an API to access the user's calendar). Therefore, a response 330 can be output to tell the user "After the meeting with Tom, you have a meeting with Richard Brown, starting at 1:30".
[0054] Figure 3B-3C Another example of resolving ambiguous user utterances is shown. Figure 3B As shown in , in some examples, the constraints for the ambiguity indicate ambiguous entities and entity attributes of the ambiguous entities, and a subset of one or more candidate concepts includes only those candidate entities from the context-specific conversation history that have the entity attributes. For example, in user utterance 302', the user asks to "arrange a meeting with him at 2:30." The descriptor "he" may be ambiguous without further context because it can refer to anyone using the masculine pronoun "he." However, the context-specific conversation history 130 includes potentially related entities, including Tom Jones indicated by concept 350A and Jane Smith indicated by concept 350B. As described above (and not shown in detail here), concepts 350A and 350B can be represented in any suitable form, for example, concept 350A can correspond to a data flow program fragment for accessing a user's address book using an API to find an individual named "Tom Jones." Based on the user utterance, data flow program 322' is configured to find salient individuals based on attributes with masculine pronouns ("he"). Thus, if the history includes the related entities "Tom Jones" and "Jane Smith", and "Tom Jones" uses masculine pronouns and "Jane Smith" uses feminine pronouns, then the search history function ("getSalient") may find the more relevant entity "Tom Jones". Thus, response 360 indicates that the meeting is scheduled for 2:30 and Tom Jones is being invited.
[0055] In some examples, such as in Figure 3CAs shown in , the constraint indicates an ambiguous reference to an action performed by the automated assistant and a constraint attribute associated with the action. Thus, the subset of the one or more candidate concepts includes a plurality of candidate actions with constraint attributes, wherein each candidate action is defined by a candidate data flow program fragment. For example, user utterance 302" indicates that the user wishes to invite "Jane" instead, but it is ambiguous as to what event to invite Jane to based on user utterance 302". Thus, the constraint attribute associated with the action is that the action is related to inviting individuals to events. Therefore, there may be related actions performed by the automated assistant, which are stored in the context-specific dialog history 130 as concepts 306. As shown, concept 306 includes a data flow program 308 corresponding to an action performed by the assistant, for example, in Figure 3B The code corresponding to the meeting set in the interaction shown in . The data flow program 308 for resolving the ambiguity is configured to use the search history function ("getSalienf") to search for related events.
[0056] In some examples, concepts indicated by a user in a user utterance can be indicated based on extrinsic and / or situational attributes and / or based on the provenance of the concept within a conversation between the user and the automated assistant. As an example, a user utterance can refer to a meeting event by mentioning a "second meeting," which can refer to a meeting via extrinsic and / or situational attributes, such as a second meeting on the user's calendar. Alternatively or additionally, the "second meeting" can refer to the second meeting discussed in the conversation between the user and the automated assistant. As a result of the supervised training herein, the code generator can be configured to appropriately identify from context what a user is referring to using a phrase such as "second meeting."
[0057] In some examples, such as Figure 3CAs shown in FIG. 1 , the data flow program 322 ″ includes a program rewriting function (referred to as “Clobber” in the exemplary syntax) that is parameterized by a designated concept (“[designated]”) stored in the context-specific conversation history and is configured to generate a new data flow program fragment related to the designated concept (e.g., write a new program for performing a previous action with respect to a new entity) based on the concept from the context-specific conversation history 130. In some examples, the designated concept (“[designated]”, i.e., the arrangement of a new meeting with Tom Jones) includes a target sub-concept (“replacementTarget”, i.e., the invitee to the event, Tom Jones). Therefore, the program rewriting function is further parameterized by a replacement sub-concept that replaces the target sub-concept, i.e., “[replacing]” indicates a new invitee, Jane, for the meeting. Therefore, the new data flow program fragment corresponds to the designated concept in which the target sub-concept is replaced with the replacement sub-concept. In other words, the program rewriting function produces a new data flow program that is configured to invite Jane Smith to the meeting that was previously scheduled with Tom Jones. Jones. Specifically, in some examples, the new data flow program fragment includes a rewritten data flow program fragment based on a historical data flow program fragment corresponding to the specified concept, wherein the sub-program fragment corresponding to the target sub-concept is replaced by a different sub-program fragment corresponding to the replacement sub-concept.
[0058] The program rewrite function can be parameterized in any suitable manner. For example, the specified concept, replacement target subconcept, and replacement subconcept can be derived from a context-specific dialog history 130 (e.g., the replacement program fragment 180 can be a previously executed program fragment), constructed according to a plurality of predefined functions 110, and / or output by a previously trained code generator 104. For example, a previously trained code generator 104 can be trained on one or more training examples, the one or more training examples including an exemplary program rewrite function, and an exemplary data flow program fragment for each parameter of the program rewrite function. In some examples, a previously trained code generator is trained on a large number of different training examples including a program rewrite function.
[0059] In the above example, the user utterance is processed using the context-specific dialog history, resulting in outputting an automated assistant response. However, in some examples, errors may occur during the processing of the user utterance. However, a data flow program according to the present disclosure may be configured for conditionless evaluation of one or more data values including a return value. Therefore, processing a data flow program may include running the data flow program to obtain the return value. However, although the data flow program may be configured for conditionless evaluation of the data values including the return value, errors may occur during the processing of the data flow program, for example due to ambiguity in the user utterance, which prevents the data flow program from being fully resolved based on the user utterance. Therefore, in response to detecting any error condition when running the data flow program, the running of the data flow program may be suspended. After suspending the running of the data flow program, the previously trained code generator may be used to generate an alternative error handling data flow program based on the suspended running of the data flow program. Therefore, an alternative error handling data flow program may be run to recover from an error condition.
[0060] Figure 4 An exemplary method 400 for processing a user utterance by an automated assistant when errors may occur during such processing is shown. At 402, method 400 includes identifying a user utterance for processing. For example, Figure 5A An exemplary dialog is shown including a user utterance 502, in which the user asks "Who is Dan's manager?" As in previous examples, while processing the user utterance, the automated assistant maintains a context-specific dialog history 130, which is used to track concepts related to previous interactions between the user and the automated assistant, such as concepts 540A, 540B, and 540C, which are not shown in detail in this example.
[0061] At 404, method 400 includes generating a data flow program (e.g., Figure 5A ). The data flow program is configured to produce a return value when successfully executed, for example, a description of Dan's manager. Thus, the data flow program includes calculating a first value [rl], which is configured as a list of all relevant persons named, for example, "Dan", where the list of expected results is a single-item list including exactly one entry, for example, a unique result for a person named, for example, "Dan". Assuming that a single individual named "Dan" is found and saved as [rl], the data flow program also includes calculating another value [r2] by finding the manager of [rl] and describing the result.
[0062] At 405, method 400 includes running the data flow program. Thus, at 406, method 400 includes starting to run the data flow program. Figure 5A , evaluating [rl] includes searching for an individual named "Dan", for example, by accessing the context-specific conversation history 130 and / or accessing the user's address book. However, when evaluating [rl], an error condition 532 is reached because there is more than one individual named, for example, "Dan", namely, "Dan A" and "Dan B", so the list of people named, for example, "Dan" is not a single item. Therefore, the difference is detected as an error condition 532, which includes a data flow program fragment describing an event in which the error condition is detected. The error event indicates that the error is due to an ambiguity.
[0063] At 408, in response to reaching an error condition caused by the execution of the dataflow program, method 400 includes handling the error condition. Handling the error condition includes: at 410, before the dataflow program generates the return value, before calculating and describing the return value (e.g., return value [r2]), suspending the execution of the dataflow program (e.g., as in Figure 5A 522). To resolve the error condition (e.g., error condition 532), the automated assistant is configured to describe the error, e.g., to obtain disambiguation information from the user. Thus, optionally at 418, handling the error condition may include outputting a clarifying question and identifying a clarifying user utterance, e.g., by looping back to 402 to identify further user utterances. As in Figure 5A As shown in , the automated assistant output includes a response 534 to the clarifying question “Do you mean Dan A or Dan B?”
[0064] Now go to Figure 5B After pausing the execution of the data flow program and outputting the response including the clarification question, the automated assistant can receive a subsequent user utterance to obtain disambiguation information for processing the error. Thus, the erroneous data flow program 522 and the error condition 532 are saved as concepts 520 and 530 (respectively) in the context-specific dialog history while processing a new user utterance 502' indicating "Dan A".
[0065] Thus, at 412, method 400 also includes generating an error handling data flow program using the previously trained code generator, wherein the error handling data flow program 552 is configured to produce the return value (e.g., the error handling data flow program 552 is an alternative means of calculating the expected value corresponding to [r2]).
[0066] In some examples, reaching the error condition includes detecting an error in running a problematic program segment of the data flow program, e.g., Figure 5A The error condition shown in occurs when evaluating the problematic program fragment that defines the [rl] value, namely, "Find(Person(name:like("Dan"))).results.singleton()". Therefore, generating the error handling data flow program includes outputting a new data flow program based on the problematic program fragment [rl]. As shown in the example, the error handling data flow program 552 uses the search history function ("getSalient") to find a related person named "Dan A" and saves the result as [r3]. In some examples, a data flow program may include an "intensionOf" function configured to return an "intention", which refers to a reference to a data flow program fragment in this document. The error handling data flow program 552 uses the data flow program fragment corresponding to the [error] event, namely, the "intensionOf[rl]" defined in the data flow program fragment of the error condition 532.
[0067] As in Figure 5B As shown in , the new data flow program based on the problematic program fragment includes a program rewriting function (e.g., "Clobber") configured to replace the problematic program fragment with an alternative program fragment in the data flow program. The error handling data flow program 552 defines a new calculation [newCompl] by using the program rewriting function ("Clobber") to rewrite the original return value [r2] by replacing the target sub-concept defined by [r4] (i.e., the problematic program fragment that caused the error) with a different sub-concept [r3].
[0068] In some examples, such as in Figure 5B As shown in , the alternative program fragment includes a search history function configured to search for the alternative program fragment in the context-specific conversation history based on its relationship with the problematic program fragment. Thus, the sub-concept [r3] is defined by code for finding a related individual named "Dan A" using the search history function.
[0069] At 414, method 400 includes starting execution of the error handling data flow program to generate a return value. For example, error handling data flow program 522 is configured to return a value [r2'] by executing a new computation [newCompl]. At 420, method 400 includes outputting the return value. After outputting the return value, at 422, method 400 also includes outputting a response based on the return value, for example, Figure 5BResponse 514 is shown, indicating "Dan's manager is Tom Jones."
[0070] In some examples, running of the dataflow program may result in reaching more than one different error condition. For example, after reaching and fully resolving a first error condition, another error condition may be reached. Alternatively or additionally, the automated assistant may be configured to detect more than one simultaneous error condition, and handle any such error condition simultaneously by generating an error handling dataflow program for all such error conditions.
[0071] In some examples, method 400 includes identifying additional errors during execution of the error handling data flow programs, and executing an error recovery loop, including generating and executing additional error handling data flow programs until one of the additional error handling data flow programs generates the return value. Thus, handling the error condition in response to reaching the error condition at 408 includes returning to 408 in response to detecting an additional error condition at 416 to handle such additional error condition. By detecting each error in turn and generating an error handling data flow program to recover from the error(s), any number of errors may be addressed sequentially and / or simultaneously.
[0072] Figure 5C Examples of handling two different error conditions are shown. Figure 5C 504, the user requests to schedule a meeting with Dan and Tom. However, there may be more than one relevant Dan, and similarly, there may be more than one relevant Tom. Thus, in a manner similar to that in Figure 5A-5B , the automated assistant is configured to output a response 534' which asks the user which Dan they would like to schedule a meeting with. The user responds in user utterance 506 that they mean "Dan A." The automated assistant is configured to output a subsequent response 536 which asks the user which Tom they would like to schedule a meeting with. The user responds in user utterance 508 that they mean "Tom J." Therefore, after "Dan" and "Tom" are disambiguated, the automated assistant is configured to schedule the meeting and output a response 538 indicating that the meeting has been scheduled.
[0073] In some examples, such as in Figure 5DAs shown in , in response to user utterance 510, the data flow program generated by the previously trained code generator uses a search history function to specify a concept for resolution, namely, finding a person [r1] by searching for significant matches of "he" mentioned in the user utterance 510. However, the search history function may find two or more related concepts, as indicated by error condition 534". For example, the context-specific conversation history 130 may include a concept 530A corresponding to "Dan A" and another concept 530B corresponding to "Tom J". For example, the concept corresponding to "Dan A" may be specified by a function such as in Figure 5B , or defined by any other suitable dataflow program fragment such as “[r3]=getSalient(Person(name:like(“Dan A”)))”, and a concept corresponding to “Tom J” may be similarly defined. Thus, the automated assistant is configured to output a response 536” containing a clarification question, wherein the clarification question requires the user to provide a clarification response indicating one of two or more related concepts (e.g., indicating “Dan A” or “Tom J”). The error handling dataflow program is also configured to receive a clarification user utterance 512, and resolve the ambiguity based on the clarification utterance. Thus, after resolving that the user referred to “Dan A” in the original user utterance 510, the automated assistant is configured to perform an action [r2] to dial the phone number of “Dan A” and output a response 538” describing the action.
[0074] In some examples, such as in Figure 5E , the error is an ambiguity, wherein the dataflow program specifies a concept for resolution using a search history function, the search history function finds zero relevant concepts, and the error handling dataflow program is operable to output a clarification question requiring the user to: provide a clarification response indicating a relevant concept, receive a clarification user utterance, and resolve the ambiguity based on the clarification utterance. For example, in Figure 5E, the user requests in user utterance 524 to "call him to confirm the meeting." The resulting data flow program 526 can be executed to find the relevant individual and call the individual. However, the context-specific conversation history 130 may not contain any concepts corresponding to the relevant individual (e.g., because male persons have not been discussed recently). Therefore, error condition 534'" indicates that no match was found to find a significant individual to calculate [rl]. Therefore, the automated assistant is configured to run an error handling data flow program (not shown), which includes obtaining a clarification response 582 to find out who to call and calling that person (e.g., by using the program to rewrite the function to create a new calculation based on [r1] and / or [r2], similar to in Figure 5A-5B ). After the user clarifies that they meant "Dan A" in user utterance 516, error handling data flow program 552 is configured to dial the phone number of "Dan A" with respect to the clarification provided by the user to complete the same steps as original data flow program 522. Thus, the data flow program is configured to output response 584 indicating that the call is being dialed.
[0075] In some examples, such as in Fig. 5F , the error is a lack of confirmation, wherein the data flow program specifies an action to be performed only after receiving a user confirmation response and the user utterance does not include such a user confirmation response, and the error handling data flow program is operable to output a confirmation question that requires the user to provide the user confirmation response before performing the action. For example, the user asks in user utterance 518 to arrange a ride to a meeting. Data flow program 562 is able to find a significant meeting and determine that a ride is now available to reach the meeting, storing the resulting value in [r2]. For example, [r2] may include any suitable information related to arranging a ride, such as a resource descriptor published by a ride-hailing service API. Data flow program 562 is configured to arrange a ride based on the resulting value, but only after obtaining a confirmation from the user. The lack of confirmation can be detected as an error condition 554. Therefore, the error handling data flow program (not shown) can be configured to output a response 556 that indicates the cost of the ride and requires the user to confirm. In user utterance 520, the user indicates that they do want to arrange a ride. Thus, the error handling dataflow program can continue to schedule a ride, such as by re-running computation [r3], and output a response 558 indicating that the action was performed.
[0076] In some examples, by asking the user one or more follow-up questions, the error handling mechanism described in this article can be used to determine the information required for processing the user's speech. For example, if the user speech includes a request for performing a specific action parameterized by one or more constraints, and the user speech specifies some but not all of the constraint values for specifying the constraints, the action may not be performed before specifying the remaining constraints. Therefore, the missing constraint parameters can be detected as errors, and accordingly, the previously trained code generator is configured to output an error handling data flow program. For example, the error handling data flow program can be run to ask questions to guide the user to provide further information for specifying the remaining constraints, and use such further information received from the user to perform the requested action. Alternatively or additionally, the error handling data flow program can be run to specify one or more of the remaining constraints using default values. As an example, if the user requires "arrange a meeting with Charles at 10 am next Thursday", the resulting data flow program can be configured to execute a meeting arrangement parameterized by date, time, invitee, and duration. Since the user utterance mentions "Charles", "Thursday" and "10 am", constraints for date, time and invitees can be specified. However, constraints for duration may not be specified, resulting in an error condition. Therefore, the previously trained code generator can be configured to output an error handling data flow program to specify the remaining constraints for the duration. The error handling data flow program can be configured to ask the user a follow-up question "How long should the meeting be?" in order to obtain another user utterance that can determine the duration of the meeting. Alternatively, the error handling data flow program can be configured to determine a default value for specifying the constraints, for example, a default meeting duration of 30 minutes. The default value can be determined by the code generator based on supervised training of the exemplary default value, using the search history function, using any other intelligent decision function, using a user-customized function (for example, to access the user's default meeting duration preference), or in any other suitable manner. In some examples, the actions defined by the data flow program are implemented using an API inference function, which can be parameterized by one or more constraints. Thus, if any such parameters are not defined by user utterances, some or all of the parameters may be inferred by running the constraint satisfaction procedure. For example, for an API inference function for scheduling a meeting that requires a meeting start time and a meeting end time, the meeting end time may be inferred based on the meeting start time and the duration. Similarly, the meeting start time may be inferred based on the meeting end time and the duration.Therefore, before asking questions, assuming default values, or otherwise handling missing constraints, the error handling data flow program can be configured to infer as many constraints as possible for the API inference function. In some examples, instead of generating a single default value for specifying a constraint, the error handling data flow program can be configured to generate multiple different candidate values for specifying the constraint and require the user to select a specific candidate value. In some examples, the error handling data flow program can be parameterized with multiple different candidate values and run multiple different API calls (e.g., using an API inference function), and require the user to select a candidate value based on the results of the API call. For example, the error handling data flow program can try to schedule multiple different meetings with different durations, and require the user to select one of the resulting arrangements. In some examples, the error handling data flow program can filter candidate options based on the results of the API call, for example, trying to schedule multiple different meetings with different durations and requiring the user to choose between meetings that do not cause scheduling conflicts, while omitting candidates with durations that will cause scheduling conflicts as indicated by the meeting scheduling API.
[0077] In some examples, the alternative program fragment is output by the previously trained code generator. Therefore, the previously trained code generator can be based on multiple training examples to output the alternative program fragment using supervised training, wherein the training examples include exemplary problematic program fragments that will lead to an error condition, and alternative program fragments that will not lead to a solution to the error condition. In other words, the previously trained code generator can be trained in a similar manner to that for generating code for user utterances and for generating code for responding to errors. In either case, the previously trained code generator is configured to generate a program using multiple predefined, combinable functions, which are arranged in any suitable sequence for processing user utterances and / or errors. In some examples, the previously trained code generator is trained on a large number of training examples, in which an error condition is reached and in which the alternative program fragment can be run to recover from the error, for example, by successfully generating a return value.
[0078] The methods and processes described herein may be bound to a computing system of one or more computing devices. Specifically, such methods and processes may be implemented as an executable computer application, a network accessible computing service, an application programming interface (API), a library, or a combination of the above and / or other computing resources.
[0079] Figure 6Schematically illustrated is a simplified representation of a computing system 600 configured to provide any to all of the computing functions described herein. The computing system 600 may take the form of one or more personal computers, network accessible server computers, tablet computers, home entertainment computers, gaming devices, mobile computing devices, mobile communication devices (e.g., smart phones), virtual / augmented / mixed reality computing devices, wearable computing devices, Internet of Things (IoT) devices, embedded computing devices, and / or other computing devices.
[0080] The computing system 600 includes a logic subsystem 602 and a storage subsystem 604. The computing system 600 may optionally include an input / output subsystem 606, a communication subsystem 608, and / or Figure 6 Other subsystems not shown.
[0081] The logic subsystem 602 includes one or more physical devices configured to run instructions. For example, the logic subsystem may be configured to run instructions that are part of one or more applications, services, or other logical constructs. The logic subsystem may include one or more hardware processors configured to run software instructions. Additionally or alternatively, the logic subsystem may include one or more hardware or firmware devices configured to run hardware or firmware instructions. The processor of the logic subsystem may be single-core or multi-core, and the instructions run thereon may be configured for sequential, parallel, and / or distributed processing. The individual components of the logic subsystem may optionally be distributed in two or more separate devices, which may be remotely located and / or configured for coordinated processing. Various aspects of the logic subsystem may be virtualized and run by a remotely accessible networked computing device configured in a cloud computing configuration.
[0082] The storage subsystem 604 includes one or more physical devices configured to temporarily and / or permanently store computer information, such as data and instructions that can be run by the logic subsystem. When the storage subsystem includes two or more devices, the devices can be co-located and / or remotely located. The storage subsystem 604 can include volatile, non-volatile, dynamic, static, read / write, read-only, random access, sequential access, location addressable, file addressable, and / or content addressable devices. The storage subsystem 604 can include removable and / or built-in devices. When the logic subsystem runs instructions, the state of the storage subsystem 604 can be converted—for example, to save different data.
[0083] Aspects of logic subsystem 602 and storage subsystem 604 may be integrated together into one or more hardware logic components. For example, such hardware logic components may include program-specific and application-specific integrated circuits (PASIC / ASIC), program-specific and application-specific standard products (PSSP / ASSP), systems on chips (SOC), and complex programmable logic devices (CPLD).
[0084] The logic subsystem and the storage subsystem can collaborate to instantiate one or more logic machines. As used herein, the term "machine" is used to collectively refer to a combination of hardware, firmware, software, instructions, and / or any other components that collaborate to provide a computer function. In other words, "machine" is never an abstract concept and always has a tangible form. A machine can be instantiated by a single computing device, or a machine can include two or more subcomponents instantiated by two or more different computing devices. In some implementations, a machine includes a local component (e.g., a software application run by a computer processor) that collaborates with a remote component (e.g., a cloud computing service provided by a network of server computers). The software and / or other instructions that give a specific machine its function can be optionally stored on one or more suitable storage devices as one or more unexecuted modules. For example, the previously trained code generator, the previously trained correlation detection machine, and / or the error handling execution machine are examples of machines according to the present disclosure.
[0085] The machine may be implemented using any suitable combination of existing and / or future machine learning (ML), artificial intelligence (AI), and / or natural language processing (NLP) techniques. For example, the previously trained code generator and / or previously trained correlation detection machine may be combined with any suitable ML, AI, and / or NLP techniques, including any suitable language model.
[0086] Non-limiting examples of techniques that may be incorporated into implementations of one or more machines include support vector machines, multi-layer neural networks, convolutional neural networks (e.g., including spatial convolutional networks for processing images and / or videos, temporal convolutional neural networks for processing audio signals and / or natural language sentences, and / or any other suitable convolutional neural network configured to convolve and aggregate features across one or more temporal and / or spatial dimensions), recurrent neural networks (e.g., long short-term memory networks), associative memories (e.g., lookup tables, hash tables, Bloom filters, neural Turing machines, and / or neural random access memories), word embedding models (e.g., GloVe or Word2Vec), unsupervised spatial and / or clustering methods (e.g., nearest neighbor algorithms, topological data analysis, and / or k-means clustering), graphical models (e.g., (hidden) Markov models, Markov random fields, (hidden) conditional random fields, and / or AI knowledge bases), and / or natural language processing techniques (e.g., tokenization, stemming, constituency and / or dependency parsing and / or intent recognition, segmentation models, and / or super-segmentation models (e.g., hidden dynamic models)).
[0087] In some examples, the methods and processes described herein can be implemented using one or more differentiable functions, where the gradient of the differentiable function can be calculated and / or estimated with respect to the input and / or output of the differentiable function (e.g., with respect to training data, and / or with respect to an objective function). Such methods and processes can be determined, at least in part, by a set of trainable parameters. Thus, the trainable parameters for a particular method or process can be adjusted by any suitable training procedure to continuously improve the functionality of the method or process.
[0088] Non-limiting examples of training procedures for adjusting trainable parameters include: supervised training (e.g., using gradient descent or any other suitable optimization method), zero-shot, few-shot, unsupervised learning methods (e.g., classification based on categories derived from unsupervised clustering methods), reinforcement learning (e.g., feedback-based deep Q-learning) and / or generative adversarial neural network training methods, belief propagation, RANSAC (random sample consistency), contextual bandit methods, maximum likelihood methods and / or expectation maximization. In some examples, components of multiple methods, processes and / or systems described herein can be trained simultaneously with respect to an objective function that measures the performance of the collective function of multiple components (e.g., with respect to reinforced feedback and / or with respect to labeled training data). Training multiple methods, processes and / or components simultaneously can improve such collective functions. In some examples, one or more methods, processes and / or components can be trained independently of other components (e.g., offline training on historical data).
[0089] The previously trained code generator and / or previously trained related detection machine can be incorporated into any suitable language model. The language model can use vocabulary features to guide sampling / searching words to recognize speech. For example, the language model can be defined at least in part by the statistical distribution of words or other vocabulary features. For example, the language model can be defined by the statistical distribution of n-grams, and the transition probability between candidate words is defined according to vocabulary statistics. The language model can also be based on any other appropriate statistical features, and / or utilize one or more machine learning and / or statistical algorithms to process the results of statistical features (for example, the confidence values generated by such processing). In some examples, for example, based on the assumption that the words in the audio signal come from a specific vocabulary, the statistical model can constrain which words can be identified for the audio signal.
[0090] Alternatively or additionally, the language model can be based on one or more neural networks previously trained to represent audio input and words in a shared latent space, for example, a vector space learned by one or more audio and / or word models (e.g., wav2letter and / or word2vec). Thus, finding candidate words can include searching the shared latent space based on vectors encoded by the audio model for the audio input to find candidate word vectors for decoding using the word model. The shared latent space can be used to evaluate, for one or more candidate words, the confidence that the candidate word is represented in the speech.
[0091] The language model can be used in conjunction with an acoustic model, which is configured to evaluate the confidence of the candidate words in the speech contained in the audio signal based on the acoustic features of the words (e.g., mel-frequency cepstral coefficients, formants, etc.) for the candidate words and the audio signal. Optionally, in some examples, the language model can be incorporated into the acoustic model (e.g., the evaluation and / or training of the language model can be based on the acoustic model). The acoustic model defines, for example, a mapping between an acoustic signal and a basic sound unit such as a phoneme based on labeled speech. The acoustic model can be based on any suitable combination of existing or future machine learning (ML) and / or artificial intelligence (AI) models, such as: deep neural networks (e.g., long short-term memory, time convolutional neural networks, restricted Boltzmann machines, deep belief networks), hidden Markov models (HMMs), conditional random fields (CRFs) and / or Markov random fields, Gaussian mixture models, and / or other graphical models (e.g., deep Bayesian networks). The audio signal to be processed using the acoustic model can be preprocessed in any suitable manner, for example, encoded at any suitable sampling rate, Fourier transformed, bandpass filtered, etc. The acoustic model can be trained to identify the mapping between the acoustic signal and the sound unit based on training using labeled audio data. For example, the acoustic model can be trained based on labeled audio data including speech and correction text to learn the mapping between the speech signal and the sound unit represented by the correction text. Therefore, the acoustic model can be continuously improved to improve its utility for correctly recognizing speech.
[0092] In some examples, the language model can incorporate any suitable graphical model, such as a hidden Markov model (HMM) or a conditional random field (CRF), in addition to a statistical model, a neural network, and / or an acoustic model. The graphical model can utilize statistical features (e.g., transition probabilities) and / or confidence values to determine the probability of recognizing a word given speech and / or other words recognized to date. Thus, the graphical model can utilize statistical features, previously trained machine learning models, and / or acoustic models to define transition probabilities between states represented in the graphical model.
[0093] When included, the input / output subsystem 606 may include one or more displays that can be used to present a visual representation of the data stored by the storage subsystem 604. The visual representation can take the form of a graphical user interface (GUI). The input / output subsystem 606 may include one or more display devices that utilize virtually any type of technology. In some implementations, the display subsystem may include one or more virtual reality, augmented reality, or mixed reality displays. When included, the input / output subsystem 606 may also include one or more speakers that are configured to output speech, for example, to present an audible representation of the data stored by the storage subsystem 604, such as an automated assistant response.
[0094] When included, the input / output subsystem 606 may include one or more input devices or interface with one or more input devices. An input device may include a sensor device or a user input device. Examples of user input devices include: a keyboard, a mouse, a touch screen, or a game controller. In some embodiments, the input subsystem may include a selected natural user input (NUI) component or interface with a selected natural user input (NUI) component. Such components may be integrated or peripheral, and the conversion and / or processing of input actions may be processed on-board or off-board. Exemplary NUI components may include microphones for voice and / or speech recognition; infrared, color, stereo and / or depth cameras for machine vision and / or gesture recognition; head trackers, eye trackers, accelerometers and / or gyroscopes for motion detection and / or intent recognition.
[0095] When included, the communication subsystem 608 can be configured to communicatively couple the computing system 600 with one or more other computing devices. The communication subsystem 608 can include wired and / or wireless communication devices compatible with one or more different communication protocols. The communication subsystem can be configured to communicate via a personal area network, a local area network, and / or a wide area network.
[0096] The present disclosure is presented by way of example and with reference to the associated drawings. Components, process steps and other elements that may be substantially the same in one or more of the drawings are identified in a coordinated manner and described with minimal repetition. However, it will be noted that the elements identified in a coordinated manner may also be different to some extent. It will also be noted that some of the drawings may be schematic and not drawn to scale. The various drawing scales, aspect ratios and number of components shown in the drawings may be deliberately distorted to make specific features or relationships easier to see.
[0097] In an example, a method includes: recognizing a user utterance; generating, using a previously trained code generator, from the user utterance a dataflow program defining an executable plan for responding to the user utterance; and executing the dataflow program.
[0098] In an example, a method includes: identifying a user utterance containing an ambiguity; using a previously trained code generator to generate a data flow program containing a search history function based on the user utterance; wherein the previously trained code generator is configured to add any of a plurality of predefined combinable functions to the data flow program based on the user utterance containing the ambiguity; and wherein the search history function is configured to select a disambiguated concept with the highest confidence from one or more candidate concepts stored in a context-specific conversation history.
[0099] In an example, a method includes: identifying a user utterance containing ambiguity; using a previously trained code generator to generate a data flow program containing a search history function based on the user utterance; wherein the search history function is configured to select a disambiguation concept with the highest confidence from one or more candidate concepts stored in a context-specific dialogue history; and the method further includes: identifying constraints related to the ambiguity, wherein the search history function is configured to select a disambiguation data flow program fragment corresponding to the disambiguation concept from the context-specific dialogue history based on such identified constraints using a previously trained correlation detector.
[0100] In an example, a method includes: identifying a user utterance containing ambiguity; using a previously trained code generator to generate a data flow program containing a search history function based on the user utterance; wherein the search history function is configured to select a disambiguation concept with the highest confidence from one or more candidate concepts stored in a context-specific conversation history. In this example or any other example, the method also includes: identifying constraints related to the ambiguity, wherein the search history function is configured to search the context-specific conversation history for a subset of one or more candidate concepts that satisfy the constraints related to the ambiguity, and wherein the disambiguation concept with the highest confidence is one of the subset of candidate concepts. In this example or any other example, the constraints indicate ambiguous entities and entity attributes of the ambiguous entities, and wherein the subset of the one or more candidate concepts includes candidate entities from the context-specific conversation history that have the entity attributes. In this example or any other example, the constraint indicates an ambiguous reference to an action performed by an automated assistant and a constraint attribute associated with the action, and wherein a subset of the one or more candidate concepts includes a plurality of candidate actions having the constraint attribute, wherein each candidate action is defined by a candidate data flow program fragment. In this example or any other example, the method further includes identifying a constraint associated with the ambiguity, wherein the search history function is configured to use a previously trained relevance detector to select a disambiguation data flow program fragment corresponding to the disambiguation concept from the context-specific dialog history based on such identified constraints. In this example or any other example, the previously trained relevance detector is trained via supervised training on a plurality of annotated dialog histories, wherein the annotated dialog histories include unresolved search history functions, the unresolved search history functions being labeled with a disambiguation concept that will resolve the unresolved search history function. In this example or any other example, the disambiguation concept that will resolve the unresolved search history function is selected by a human presenter from the context-specific dialog history. In this example or any other example, the exemplary disambiguation concept for the exemplary search history function includes an exemplary program fragment received from a human presenter. In this example or any other example, the data flow program is configured for conditionless evaluation of one or more data values including a return value, and the method also includes running the data flow program, and in response to reaching any error condition while running the data flow program: pausing the running of the data flow program; using the previously trained code generator, based on the suspended running of the data flow program, generating an error handling data flow program; and resuming running using the error handling data flow program.In this example or any other example, the previously trained code generator is configured to add any combinable function of a plurality of predefined combinable functions to the data flow program based on the user utterance containing the ambiguity. In this example or any other example, the plurality of predefined combinable functions include an intelligent decision function, wherein the intelligent decision function is configured to perform calculations using a previously trained machine learning model. In this example or any other example, the plurality of predefined combinable functions include user-customized functions, which are configured to access user-customized settings and perform calculations based on user-customized settings. In this example or any other example, the plurality of predefined combinable functions include external functions, which are configured to call external application programming interfaces (APIs). In this example or any other example, the plurality of predefined combinable functions include inference functions, which are configured to perform calculations on the results of the external functions. In this example or any other example, the plurality of predefined combinable functions include macro functions, wherein the macro functions include a plurality of other predefined combinable functions, and wherein the macro functions are configured to run a plurality of other predefined combinable functions. In this example or any other example, the plurality of predefined combinable functions include a program rewrite function parameterized by a specified concept stored in the specific context conversation history and configured to generate a new data flow program fragment associated with the specified concept. In this example or any other example, the specified concept stored in the context-specific conversation history includes a target subconcept, wherein the program rewrite function is further parameterized by a replacement subconcept, and wherein the new data flow program fragment corresponds to the specified concept, wherein the target subconcept is replaced by the replacement subconcept. In this example or any other example, the previously trained code generator is trained via supervised training on a plurality of annotated conversation histories, wherein the annotated conversation histories include exemplary user utterances and an exemplary data flow program including the search history function.
[0101] In an example, a computer system includes: a microphone; a logic device; and a storage device, the storage device storing instructions that can be executed by the logic device to: receive speech from the microphone; identify ambiguous user utterances in the speech; and use a previously trained code generator to generate a data flow program including a search history function based on the user utterances, wherein the search history function is configured to select a disambiguated concept with the highest confidence from one or more candidate concepts stored in a context-specific conversation history.
[0102] In an example, a method includes: identifying a user utterance containing ambiguity; identifying constraints related to the ambiguity; using a previously trained code generator to generate a data flow program containing a search history function based on the user utterance, wherein the search history function is configured to: search for one or more candidate concepts stored in a context for a subset of the one or more candidate concepts that satisfy the constraints related to the ambiguity; using a previously trained correlation detection engine to select a disambiguation concept with the highest confidence from a subset of the one or more candidate concepts; and selecting a disambiguation data flow program fragment corresponding to the disambiguation concept.
[0103] It will be understood that the configuration and / or method described in this article are exemplary in nature, and these specific embodiments or examples should not be considered restrictive, because many variations are possible. The specific routine or method described in this article can represent one or more of any number of processing strategies. Therefore, the various actions shown and / or described can be performed in parallel or omitted in the sequence shown and / or described, in other sequences. Likewise, the order of the above process can be changed.
Claims
1. A method, include: Identify ambiguous user utterances; generating a dataflow program including a search history function based on the user utterance using a previously trained code generator, wherein the dataflow program defines an executable plan for responding to the user utterance and is configured to cause an automated assistant to respond to the user utterance; and Identify the constraints associated with the ambiguity, wherein the search history function is configured to: select a disambiguation concept with the highest confidence from one or more candidate concepts stored in a context-specific dialog history and select a disambiguation data flow program fragment corresponding to the disambiguation concept from the context-specific dialog history based on such identified constraints using a previously trained correlation detection engine, and wherein the previously trained code generator is configured to add any of a plurality of predefined composable functions to the data flow program based on the user utterance containing the ambiguity.
2. The method according to claim 1, in, The search history function is configured to search the context-specific dialog history for a subset of the one or more candidate concepts that satisfies the constraint associated with the ambiguity, and wherein the highest confidence disambiguated concept is a candidate concept in the subset of candidate concepts.
3. The method according to claim 2, in, The constraints indicate ambiguous entities and entity attributes of the ambiguous entities, and wherein the subset of the one or more candidate concepts includes candidate entities from the context-specific conversation history having the entity attributes.
4. The method according to claim 2, in, The constraint indicates an ambiguous reference to an action performed by the automated assistant and a constraint attribute associated with the action, and wherein the subset of the one or more candidate concepts includes a plurality of candidate actions having the constraint attribute, wherein each candidate action is defined by a candidate dataflow program fragment.
5. The method according to claim 1, in, The previously trained relevance detector is trained via supervised training on a plurality of annotated conversation histories, wherein the annotated conversation histories include unresolved search history functions labeled with disambiguation concepts that will resolve the unresolved search history functions.
6. The method according to claim 5, in, The disambiguated concept that will resolve the unresolved search history function is selected by a human presenter from the context-specific dialog history.
7. The method according to claim 5, in, An exemplary disambiguation concept for an exemplary search history function includes an exemplary program snippet received from a human presenter.
8. The method according to claim 1, in, The data flow program is configured for conditionless evaluation of one or more data values including a return value, and the method further comprises running the data flow program, and in response to reaching any error condition while running the data flow program: Pause the running of the data stream program; generating an error handling dataflow program based on a suspended run of the dataflow program using the previously trained code generator; as well as The error handling data flow program is utilized to resume operation.
9. The method according to claim 1, in, The plurality of predefined composable functions include a program rewrite function that is parameterized by a specified concept stored in the context-specific dialog history and is configured to generate a new data flow program fragment associated with the specified concept.
10. The method according to claim 9, in, The specified concept stored in the context-specific dialog history includes a target sub-concept, wherein the program rewrite function is also parameterized by a replacement sub-concept, and wherein the new data flow program fragment corresponds to the specified concept in which the target sub-concept is replaced by the replacement sub-concept.
11. The method according to claim 1, in, The previously trained code generator is trained via supervised training on a plurality of annotated dialog histories, wherein the annotated dialog histories include exemplary user utterances and exemplary data flow programs including the search history function.
Citation Information
Patent Citations
Intelligent WeChat bank system based on natural language automatic scheduler and intelligent scheduling method for computer system through natural language
CN103902248A
Anaphora Resolution Using Linguisitic Cues, Dialogue Context, and General Knowledge
US20140257792A1