Automated Assistant Call for Appropriate Agents
By employing a machine learning-based agent selection model within automated assistants, the system effectively selects and invokes the appropriate agent for user inputs, improving interaction efficiency and resource utilization.
Patent Information
- Application Number
- JP2023172099
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-04-18
- Filing Date
- 2023-10-03
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2037-04-18
AI Technical Summary
Existing automated assistants struggle to efficiently select and invoke the appropriate agent in response to user input that cannot specify a particular agent, often leading to suboptimal interactions and resource wastage.
Implementing a system that utilizes a machine learning-based agent selection model to identify the most suitable agent from a plurality of available agents based on natural language input and contextual factors, ensuring that only the appropriate agent is invoked to handle the user's intent.
This approach significantly enhances the probability that the selected agent can appropriately handle user requests, reducing the risk of failed interactions and conserving network and processor resources by avoiding unnecessary agent calls.
Smart Images

Figure 0007698013000001 
Figure 0007698013000002 
Figure 0007698013000003
Abstract
Description
Background Art
[0001] Automated assistants (also known as "personal assistants", "mobile assistants", etc.) can be interactively operated by a user via various client devices such as smartphones, tablet computers, wearable devices, automotive systems, stand-alone personal assistant devices, and the like. An automated assistant receives input from a user (e.g., typed and / or spoken natural language input), and responds with response content (e.g., visual and / or auditory natural language output). An automated assistant that is interactively operated via a client device can be implemented via the client device itself and / or via one or more remote computing devices that communicate with the client device (e.g., computing devices in the "cloud") over a network.
Summary of the Invention
Problems to be Solved by the Invention
[0002] This specification is generally directed to methods, systems, and computer-readable media for invoking an agent in an interaction between a user and an automated assistant. The step of invoking an agent can include transmitting a call request that includes values of call parameters (e.g., values for intent parameters, values for intent slot parameters, and / or values for other parameters), and that causes the agent to generate content to be presented to the user via one or more user interface output devices (e.g., via one or more of the user interface output devices utilized in the interaction with the automated assistant) (e.g., by using an application programming interface (API)). The response content generated by the agent can be tailored to the call parameters of the call request.
Means for Solving the Problems
[0003] Some implementations are directed to receiving a natural language input from a user that indicates a desire to involve an agent in a human / automation assistant conversation, but that cannot specify a particular agent to be involved. For example, "book me a hotel in Chicago" indicates a desire to involve an agent with "hotel booking" intent parameters and "Chicago" location parameters, but does not specify a particular agent to call. Those implementations are further directed to selecting a particular agent from a plurality of available agents, and transmitting a call request to the selected particular agent. For example, the call request may be transmitted to the selected particular agent without transmitting the call request to the other of the available agents. In some of those implementations, a particular agent and a particular intent for the particular agent are selected (e.g., when the particular agent is operable to generate response content as a response to one of a plurality of different intents).
[0004] In some of the implementations for selecting a specific agent from a plurality of available agents, an agent selection model is utilized when selecting a specific agent. In some versions of those implementations, the agent selection model includes at least one machine learning model, such as a deep neural network model. The machine learning model can be trained to enable the generation of an output that indicates, for each of the plurality of available agents (and optionally the intents for those agents), the probability that the available agent (and optionally the intent) will generate appropriate response content. The generated output is based on the input applied to the machine learning model, and the input is based on the current interaction with the automated assistant and optionally additional context values. For example, the input based on the current interaction can include various values based on the latest natural language input provided to the automated assistant in the current interaction and / or past natural language inputs provided in the current interaction. Also, the optional additional context values can include client device context values, such as values based on, for example, the user's interaction history on the client device, the currently rendered and / or most recently rendered content on the client device, the location of the client device, the current date and time, etc.
[0005] In some implementations, when the agent selection model includes at least one machine learning model, at least one of the machine learning models can be trained based on training instances based on past interactions with the available agents.
[0006] As an example, a plurality of training instances can be generated based on corresponding agent requests that are each generated based on natural language inputs (e.g., natural language inputs that could not identify a particular agent) provided in corresponding human / automation assistant dialogues. The agent requests may be transmitted to each of a plurality of available agents (e.g., all available agents), and the responses can be received from one or more of the available agents to which the agent requests were transmitted. Each of the training instances can include a training instance input based on the agent request (e.g., the corresponding natural language input and optionally a context value) and a training instance output based on the response. The responses can each indicate the ability of the corresponding one of these agents to resolve the agent request. For example, a response from a given agent can be a binary indication (e.g., "resolvable", "not resolvable", or "responsive content", "no responsive content / error"), a non-binary confidence measure (e.g., "70% likely resolvable"), the actual response content (or no content / error), etc. Also, for example, receiving a response from a given agent can indicate that it can respond, while not receiving a response from a given agent can indicate that it cannot respond. The agent requests can be transmitted to the agents without an active call to the agents available in the dialogue. For example, the agent requests may be similar to call requests, but can include a "non-call" flag and / or other indication that the agent should not be called immediately. Also, for example, the responses can be processed by the automation assistant, in addition to or alternatively, without providing the corresponding content in the dialogue.
[0007] Once trained, such a machine learning model may be utilized to predict the probability for each of a plurality of available agents (and optionally intents) based on the current conversation (e.g., the natural language input of the current conversation and optionally context values), each of these probabilities indicating the probability (e.g., binary or non-binary) that an agent can appropriately handle a call request based on the conversation. The selection of a particular agent may be at least partly based on such probabilities, and the call request may be transmitted only to a particular agent.
[0008] In some implementations, at least one of the machine learning models can be generated based on natural language input provided to the agent after the agent is called. For example, the natural language input provided to the agent immediately after the agent is called can be stored under the context of the agent. For example, the natural language input can be the input provided to the agent immediately after a "naked" call. A naked call of an agent is a call of the agent based on a call request that is directed to the agent but does not include a value for an intent parameter and / or does not include a value for an intent slot parameter. For example, in response to a natural language input of "open Agent X", a naked call of "Agent X" can occur in response to a call request transmitted to "Agent X" that does not include a value for an intent parameter and does not include a value for an intent slot parameter. Also, for example, in response to a natural language input of "set a reminder with Agent X", a naked call of "Agent X" can occur in response to a call request transmitted to "Agent X" that includes a "reminder" value for an intent parameter but does not include a value for an intent slot parameter. A selection model can be generated that includes a mapping (or other association) between the natural language input and the corresponding agent (and optionally the intent) to which the natural language input was provided. In this way, the mapping is based on the initial interaction provided by the user after a naked call of the agent, which enables the generation of an agent selection model that provides the most likely insights for the agent to perform skillfully in response to various natural language inputs. Additional and / or alternative selection models can be utilized when selecting a particular agent. As an example, a selection model generated based on past explicit selections of agents by various users can be utilized in addition to, or alternatively to, selecting a particular agent.
[0009] In some implementations, various additional and / or alternative criteria are utilized when selecting a particular agent (and optionally an intent). As an example, agent requests are transmitted "live" to multiple candidate agents (as described above), and the responses from those agents can be analyzed when determining which particular agent to call. As another example, additional and / or alternative criteria can include the history of user interactions with the client device (e.g., how frequently a particular agent has been used by the user, how recently a particular agent has been used by the user), the currently rendered and / or most recently rendered content on the client device (e.g., the content corresponds to an agent feature), the location of the client device, the current date and / or time, the ranking of a particular agent (e.g., ranking by the user population), the popularity of a particular agent (e.g., popularity among the user population), and the like. In implementations where a machine learning model is utilized to select a particular agent, such criteria can be applied as inputs to the machine learning model and / or considered in combination with the outputs generated on the machine learning model.
[0010] The various techniques described above and / or elsewhere in this specification enable the selection of a particular agent and can increase the probability that the selected particular agent can appropriately handle a call request. This can reduce the risk that the particular agent selected for a call cannot execute the intent of the call request (optionally, along with values for additional parameters of the call request) in a way that can save various computing resources. For example, this can save network and / or processor resources that might otherwise be consumed by an initial failed attempt to execute an intent using an agent, which is followed by subsequent calls to alternative agents in another attempt to execute the intent. Further, in implementations where a particular agent is selected without prompting the user to choose from among a plurality of available agents, this can reduce the number of "turns" of human / automation assistant interaction required before a call. This can also save various network and / or processor resources that would otherwise be consumed by such turns. Further, in implementations that utilize a trained machine learning model, the trained machine learning model can be used to determine the probability that an agent can handle a particular call without requiring network resources to be consumed through "live" interactions with one or more of the agents to make such a determination. This can also save various network and / or processor resources that would otherwise be consumed by such live interactions.
[0011] In some situations, in response to the invocation of a particular agent by the techniques disclosed herein, the human / automation assistant dialogue may be transferred (either actually or effectively) to the particular agent, at least temporarily. For example, the output based on the response content of the particular agent may be provided to the user in furthering the dialogue, and further user input may be received in response to the output. Further user input (or its transformation) may be provided to the particular agent. The particular agent may utilize its semantic engine and / or other components in generating further response content that can be used to generate further output to be provided in furthering the dialogue. This general process may continue, for example, until the additional user interface input of the user ends the particular agent dialogue (e.g., provides an answer or solution instead of a prompt), the additional user interface input of the user ends the particular agent dialogue (e.g., instead calls for a response from an automation assistant or another agent), etc.
[0012] In some situations, the automation assistant can still act as a mediator when the conversation is effectively transferred to a particular agent. For example, when acting as a mediator when the user's natural language input is voice input, the automation assistant converts the voice input to text, provides the text (and optionally annotations of the text) to a particular agent, receives response content from the particular agent, and may provide an output based on the particular response content for presentation to the user. Also, for example, when acting as a mediation means, the automation assistant can analyze the user input and / or response content of a particular agent to determine whether the conversation with the particular agent should end, whether the user should be transferred to an alternative agent, whether global parameter values should be updated based on the particular agent conversation, and so on. In some situations, the conversation is actually transferred to a particular agent (without the automation assistant acting as a mediator after the transfer), and optionally, can be transferred back to the automation assistant after the occurrence of one or more conditions such as termination by the particular agent (e.g., in response to the completion of an intent via the particular agent).
[0013] In the implementations described herein, an automation assistant can select an appropriate agent based on an interaction with a user and call the agent to achieve the user's intent as directed by the user in the interaction. In these implementations, the user may be able to engage an agent via an interaction with the automation assistant without needing to know a "call phrase" to explicitly trigger the agent and / or without even needing to know initially that the agent exists. Further, the implementations may enable the user to call any one of a plurality of heterogeneous agents that can perform actions across a plurality of different intents using a common automation assistant interface (e.g., an audible / voice-based interface and / or a graphical interface). For example, the common automation assistant interface may be used to engage any one of a plurality of agents handling a "restaurant reservation" intent, any one of a plurality of agents handling a "purchasing professional services" intent, any one of a plurality of agents handling a "telling jokes" intent, any one of a plurality of agents handling a "reminder" intent, any one of a plurality of agents handling a "purchasing travel services" intent, and / or any one of a plurality of agents handling an "interactive game" intent.
[0014] As used herein, "agent" refers to one or more computing devices and / or software that are separate from the automated assistant. In some situations, the agent can be a third-party (3P) agent in that it is managed by a party that is separate from the party that manages the automated assistant. The agent is configured to receive call requests (e.g., over a network and / or via an API) from the automated assistant. In response to receiving a call request, the agent generates response content based on the call request and transmits the response content for providing an output based on the response content. For example, the agent can transmit the response content to the automated assistant for the automated assistant to provide an output based on the response content. As another example, the agent can itself provide an output. For example, a user can interactively operate the automated assistant via a client device (e.g., the automated assistant can be implemented on and / or communicate with the client device over a network), and the agent can be an application installed on the client device or an application executable remotely from the client device, but may be "streamable" on the client device. When the application is called, this can be executed by the client device and / or brought to the forefront by the client device (e.g., its content can take over the display of the client device).
[0015] This specification describes various types of inputs that can be provided to an automated assistant and / or agent by a user via a user interface input device. In some cases, the input may be free-form natural language input, such as text input based on user interface input generated by the user via one or more user interface input devices (e.g., based on typed input provided via a physical or virtual keyboard, or based on spoken input provided via a microphone). As used herein, free-form input is input that is formulated by the user and not restricted to a group of options presented for user selection (e.g., not restricted to a group of options presented in a drop-down menu).
[0016] In some implementations, a method is provided that is executed by one or more processors and includes receiving a natural language input instance generated based on a user interface input in a human / automation assistant dialogue. The method includes generating an agent request based on the natural language input instance before invoking an agent in response to the natural language input instance, selecting a set of multiple agents from a corpus of available agents for the agent request, transmitting the agent request to each of the multiple agents in the set, receiving a corresponding response to the request from at least a subset of the multiple agents in response to the transmission, determining, from each of the responses, the relative capabilities of the agents that provide the responses to generate response content in response to the agent request, selecting a particular agent from among the multiple agents based on at least one of the responses, and further including invoking the particular agent based on the step of selecting the particular agent in response to the natural language input. The step of invoking the particular agent causes the response content generated by the particular agent to be provided for presentation via one or more user interface output devices. In some implementations, only the particular agent selected is invoked in response to receiving the natural language input.
[0017] These methods and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0018] In some implementations, the method further includes storing, on one or more computer-readable media, an association of an agent request with at least one of the agents determined to be able to respond to the agent request, and generating an agent selection model based on the stored association between the agent request and at least one of the agents determined to be able to respond to the agent request. In some of those implementations, the method includes, after generating the agent selection model, receiving additional natural language input in an additional human / automation assistant dialogue, selecting an additional one of the plurality of agents based on the additional natural language input and the agent selection model, and transmitting an additional call request to the additional agent based on selecting the additional agent in response to the additional natural language input. The additional call request is a request to call the additional agent. In response to receiving the additional natural language input, the additional call request is optionally transmitted only to the selected additional agent.
[0019] In some implementations, the step of selecting a particular agent is further based on the amount of interaction of the user involved in the dialogue with the particular agent, the recency of the interaction of the user with the particular agent, and / or the ranking or popularity of the particular agent within the user population.
[0020] In some implementations, a method is provided that is executed by one or more processors, which, for each of a plurality of natural language input instances generated based on user interface inputs in a human / automation assistant interaction, includes generating an agent request based on the natural language input instance, selecting a set of a plurality of agents from a corpus of available agents for the agent request, transmitting the agent request via one or more application programming interfaces to each of the plurality of agents in the set, receiving, in response to the transmission, a corresponding response to the request from each of the plurality of agents, and storing, on one or more computer-readable media, one or more associations between the agent request and the responses to the agent request. Each response can indicate the ability of the corresponding one of the plurality of agents to generate response content in response to the agent request. The method further includes generating an agent selection model based on the stored associations between the agent requests and their responses. After generating the agent selection model, the method includes receiving a subsequent natural language input from the user directed to the automation assistant as part of an interaction between the user and the automation assistant, selecting a particular agent based on the subsequent natural language input and the agent selection model, where the particular agent is one of the available agents, and transmitting, in response to receiving the subsequent natural language input and in response to selecting the particular agent, a call request via one or more of the application programming interfaces to the selected particular agent. The call request causes the selected particular agent to be called and to generate particular response content to be presented to the user via one or more user interface output devices. In some implementations, in response to receiving the subsequent natural language input, the call request is transmitted only to the selected particular agent.
[0021] These methods and other implementations of the technology disclosed in this specification may optionally include one or more of the following features.
[0022] In some implementations, for a given natural language input instance among a plurality of natural language input instances, a first subset of responses each indicates an ability to generate response content, and a second subset of responses each indicates an inability to generate response content. In some of those implementations, the responses of the second subset indicate that inability based on a step that indicates an error or a confidence measure that fails to meet a threshold.
[0023] In some implementations, the agent selection model is a machine learning model. In some of those implementations, the step of generating the machine learning model includes generating a plurality of training instances based on agent requests and their responses, and training the machine learning model based on the training instances. The step of generating each training instance can include generating a training instance input of the training instance based on the corresponding agent request of the agent request, and generating a training instance output of the training instance based on the response stored in relation to the corresponding agent request. In some of the implementations, the step of selecting a particular agent based on a subsequent natural language input and the agent selection model includes applying an input feature based on the subsequent natural language input as an input to the machine learning model, generating an output including a value for the particular agent on the machine learning model based on the input, and selecting the particular agent based on the value for the particular agent. In some versions of those implementations, the step of selecting a particular agent is further based on one or more context values. For example, the step of selecting a particular agent based on one or more context values can include applying the one or more context values as additional input to the machine learning model.
[0024] In some implementations, the method further includes selecting a plurality of natural language input instances based on a step of determining that the plurality of natural language input instances cannot specify an agent.
[0025] In some implementations, the method further includes, for a given natural language input instance of a plurality of natural language input instances, selecting a given agent of the plurality of agents using a response to an agent request, and transmitting a selected call request to the selected given agent, the selected call request being based on the given natural language input instance.
[0026] In some implementations, the set of a plurality of agents is selected from a corpus of available agents based on a set of a plurality of agents each associated with a value for an intent parameter represented by a natural language input instance.
[0027] In some implementations, a method is provided that is executed by one or more processors and includes, for each of a plurality of natural language input instances generated based on user interface inputs in a human / automation assistant dialogue, generating an agent request based on the natural language input instance; selecting a set of a plurality of agents from a corpus of available agents for the agent request; transmitting the agent request to each of the plurality of agents in the set; receiving, in response to the transmission, corresponding responses from at least a subset of the plurality of agents; determining, from each of the responses, the relative capabilities of the agents providing the responses to generate response content in response to the agent request; storing an association of the agent request with at least one of the agents determined to be able to respond to the agent request in one or more computer-readable media; generating an agent selection model based on the stored association between the agent request and the agents determined to be able to respond to the agent request; after generating the agent selection model, receiving a subsequent natural language input from the user directed to the automation assistant as part of a dialogue between the user and the automation assistant; selecting a particular agent based on the subsequent natural language input and the agent selection model, the particular agent being one of the available agents; and transmitting a call request to the selected particular agent in response to selecting the particular agent, the call request causing the selected particular agent to generate particular response content to be presented to the user via one or more user interface output devices.
[0028] These and other implementations of the techniques disclosed herein may optionally include one or more of the following features.
[0029] In some implementations, the step of selecting a particular agent occurs without providing the user with any output that explicitly asks the user to select between the particular agent and one or more other agents among the available agents.
[0030] In some implementations, the agent selection model is a machine learning model. In some of those implementations, generating the machine learning model includes generating a plurality of training instances based on the agent request and the agents determined to be able to respond to the agent request, and training the machine learning model based on the training instances. Generating each of the training instances can include generating a training instance input of the training instance based on the corresponding agent request of the agent request, and generating a training instance output of the training instance based on at least one of the agents determined to be able to respond to the request. In some of the implementations, selecting a particular agent based on the subsequent natural language input and the agent selection model includes applying an input feature based on the subsequent natural language input as an input to the machine learning model, generating an output including a value for the particular agent on the machine learning model based on the input, and selecting a particular agent based on the value for the particular agent.
[0031] In addition, some implementations include one or more processors of one or more computing devices, the one or more processors operable to execute instructions stored in associated memory, the instructions configured to cause any of the aforementioned methods to be performed. Some implementations also include one or more non-transitory computer-readable storage media that store computer instructions executable by the one or more processors to perform any of the aforementioned methods.
[0032] It should be understood that all combinations of the foregoing concepts and additional concepts that are more particularly described herein are intended to be part of the subject matter disclosed herein. For example, all combinations of the claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter disclosed herein.
Brief Description of the Drawings
[0033]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0034] In some situations, in order to call a particular agent for a particular intent via an automation assistant, the user must provide an input that explicitly calls that particular agent. For example, to call an agent named "Hypothetical Agent" for the "restaurant reservation" intent, the user must know to utter a "call phrase" such as "book a restaurant with Hypothetical Agent" to the agent. In such explicit calls, first, the user needs to know which agent is most appropriate for the intent, and directly send the user to that agent for the attempted resolution of the intent via the interaction with the agent.
[0035] However, in many cases, users are not aware that various agents are available, and it is often not practical and / or desirable to explicitly provide users with a list of available agents and associated functions in an automated assistant interface that is often constrained. For example, some automated assistant interfaces are "voice only", and it may not be practical and / or desirable to "read a list" of agents and associated functions to the user. Further, it is possible that the automated assistant may not be aware of the capabilities of various available agents.
[0036] The various implementations disclosed herein enable the selection and invocation of an appropriate agent in response to a user's "vague" natural language input, such as natural language input that desires to involve an agent but cannot specify a particular agent to be involved.
[0037] Next, referring to FIG. 1, an exemplary environment in which the technology disclosed herein may be implemented is shown. The exemplary environment includes a client device 106, an automated assistant 110, and a plurality of agents 140A-N. The client device 106 may be, for example, a stand-alone voice-activated speaker device, a desktop computing device, a laptop computing device, a tablet computing device, a cellular phone computing device, a computing device of the user's vehicle, and / or a wearable device of the user that includes a computing device (e.g., a user's wristwatch having a computing device, a user's glasses having a computing device, a virtual or augmented reality computing device). Additional and / or alternative client devices may be implemented.
[0038] Although the automation assistant 110 is illustrated in FIG. 1 as being separate from the client device 106, in some implementations, all or some aspects of the automation assistant 110 may be implemented by the client device 106. For example, in some implementations, the input processing engine 112 may be implemented by the client device 106. In implementations where one or more (e.g., all) aspects of the automation assistant 110 are implemented by one or more computing devices remote from the client device 106, those aspects of the client device 106 and the automation assistant 110 communicate via one or more networks, such as a wide area network (WAN) (e.g., the Internet).
[0039] Although only one client device 106 is illustrated in combination with the automation assistant 110, in many implementations, the automation assistant 110 may be remote and may interface with each of a plurality of client devices of a plurality of users. For example, the automation assistant 110 may manage communication with each of a plurality of devices via different sessions and may manage multiple sessions in parallel. For example, the automation assistant 110 in some implementations may be implemented as a cloud-based service that employs cloud infrastructure, such as using a server farm or cluster of high-performance computers that execute software suitable for handling a large number of requests from a plurality of users. However, for simplicity, many examples in this specification are described with respect to a single client device 106.
[0040] The automation assistant 110 is separated from agents 140A - N and communicates with agents 140A - N via an API and / or via one or more communication channels (e.g., internal communication channels of the client device 106 and / or a network such as the WAN). In some implementations, one or more of agents 140A - N are each managed by a respective party that is separate from the party that manages the automation assistant 110.
[0041] One or more of agents 140A - N may each optionally provide data directly or indirectly for storage in agent database 152. However, some agents 140A - N often do not provide certain data, provide incomplete data, and / or provide inaccurate data. Some implementations disclosed herein may mitigate these situations by utilizing various additional techniques when selecting an appropriate agent for ambiguous user input. The data provided for a given agent may, for example, define intents that can be solved by the given agent. Further, the data provided for a given agent may, for each intent, define available values that can be handled by the agent for a plurality of intent slot parameters defined for the intent. In some implementations, automation assistant 110 and / or other components may define acceptable values that can be defined for each of an intent and an intent slot parameter. For example, such criteria may be defined via an API maintained by automation assistant 110. Then, one or more of agents 140A - N may provide the intent and its available values for the intent slot parameter to automation assistant 110 and / or other components (e.g., transmitted via a WAN) that may verify the validity of the data and store the data in agent database 152. Agent database 152 may, in addition or alternatively, store various other characteristics for the agents, such as agent rankings, agent popularity metrics, etc.
[0042] The automated assistant 110 includes an input processing engine 112, a local content engine 130, an agent engine 120, and an output engine 135. In some implementations, one or more of the engines of the automated assistant 110 may be omitted, combined, and / or implemented as components separate from the automated assistant 110. Additionally, the automated assistant 110 may include additional engines not illustrated herein for simplicity.
[0043] The automated assistant 110 receives an instance of user input from the client device 106. For example, the automated assistant 110 may receive free-form natural language voice input in the form of a streaming audio recording. The streaming audio recording may be generated by the client device 106 in response to a signal received from a microphone of the client device 106 that captures the spoken input of the user of the client device 106. As another example, the automated assistant 110 may receive free-form natural language typed input and / or even structured (non-free-form) input in some implementations. In some implementations, the user input may be generated by the client device 106 and / or provided to the automated assistant 110 in response to an explicit invocation of the automated assistant 110 by the user of the client device 106. For example, the invocation may be the detection by the client device 106 of a specific voice input of the user (e.g., a hot word / phrase of the automated assistant 110 such as "Hey Assistant"), a user interaction with a hardware button and / or virtual button (e.g., a tap of a hardware button, a selection of a graphical interface element displayed by the client device 106), and / or other specific user interface input.
[0044] Automation assistant 110 provides an instance of output in response to receiving an instance of user input from client device 106. The instance of output may be, for example, audio provided auditorily by device 106 (e.g., output via the speakers of client device 106), text and / or graphic content presented graphically by device 106 (e.g., rendered via the display of client device 106), and the like. As described herein, some instances of output may be based on local response content generated by automation assistant 110, while other instances of output may be based on response content generated by a selected agent among agents 140A - N.
[0045] The input processing engine 112 of the automation assistant 110 processes natural language input and / or other user input received via the client device 106 and generates an annotated output for use by one or more other components of the automation assistant 110, such as the local content engine 130 and / or the agent engine 120. For example, the input processing engine 112 may process natural language free-form input generated by a user via one or more user interface input devices of the client device 106. The generated annotated output includes one or more annotations of the natural language input and optionally one or more (e.g., all) of the words of the natural language input. As another example, the input processing engine 112 may, in addition to or alternatively, receive an instance of voice input (e.g., in the form of digital audio data) and include a voice / text module that converts the voice input into text including one or more text words or phrases. In some implementations, the voice / text module is a streaming voice / text engine. The voice / text module may depend on one or more stored voice / text models (also referred to as language models) that can each model the relationship between an audio signal and the speech units in a word, along with the word sequence in the word.
[0046] In some implementations, the input processing engine 112 is configured to identify and annotate various types of grammatical information in natural language input. For example, the input processing engine 112 may include a part-of-speech tagger configured to attach its grammatical role to a word as an annotation. For example, the part-of-speech tagger may tag each word with a part of speech such as "noun", "verb", "adjective", "pronoun", etc. Also, for example, in some implementations, the input processing engine 112 may additionally and / or alternatively include a dependency parser configured to determine syntactic relationships between words in natural language input. For example, the dependency parser may determine which words modify other words, the subject, and the verb in a sentence (e.g., parse tree), and may annotate such dependency relationships.
[0047] In some implementations, the input processing engine 112 may additionally and / or alternatively include an entity tagger configured to annotate entity references in one or more clauses such as references to people, organizations, locations, etc. The entity tagger may annotate references to entities at a high level of granularity (e.g., enabling identification of all references to entity classes such as people) and / or at a low level of granularity (e.g., enabling identification of all references to specific entities such as a particular person). The entity tagger may depend on the content of the natural language input to resolve a particular entity and / or optionally interact with a knowledge graph or other entities to resolve a particular entity.
[0048] In some implementations, the input processing engine 112 may additionally and / or alternatively include an identity resolver configured to group or "cluster" references to the same entity based on cues obtained from one or more contexts. For example, the identity resolver may be used to resolve the word "it" in an instance of user input to the preceding mention of "Restaurant A" in the instance immediately preceding the user input.
[0049] In some implementations, one or more components of the input processing engine 112 may depend on annotations from one or more other components of the input processing engine 112. For example, in some implementations, a named entity tagger may depend on annotations from an identity resolver and / or a dependency parser when annotating all references to a particular entity. Also, for example, in some implementations, an identity resolver may depend on annotations from a dependency parser when clustering references to the same entity. In some implementations, when processing a particular natural language input, one or more components of the input processing engine 112 may determine one or more annotations using related previous inputs outside of the particular natural language input and / or other related data.
[0050] The input processing engine 112 may attempt to recognize semantic or meaning differences in user input and provide the semantic directives of the user input for use by the local content engine 130 and / or the agent engine 120. The input processing engine 112 may map text (or other input) to a particular action and may depend on one or more stored grammar models to identify attributes that constrain input variables for the execution of such an action, for example.
[0051] The local content engine 130 may be configured to generate a response to user input received when the user input is associated with a "local action" (as opposed to an agent action). In some implementations, the input processing engine 112 determines whether the user input is associated with a local action or an agent action. The local content engine 130 may cooperate with the input processing engine 112 and execute one or more actions as indicated by the parsed text (e.g., an action and action parameters) provided by the input processing engine 112. For local intents, the local content engine 130 may generate local response content, provide the local response content to the output engine 135 to provide a corresponding output, and present it to the user via the device 106. The local content engine 130 may utilize one or more stored local content models 154 to generate local content and / or execute other actions. The local content model 154 may incorporate various rules, for example, to create local response content. In some implementations, the local content engine 130 may communicate with one or more other "local" components, such as a local dialogue module (which may be the same as or similar to the dialogue module 126), when generating local response content.
[0052] Output engine 135 supplies instances of output to client device 106. The instances of output may be based on local response content (from local content engine 130) and / or response content from one of agents 140A - N (when automation assistant 110 operates as a mediator). In some implementations, output engine 135 may comprise a text - to - speech engine that converts the text component of the response content to an audio format, and the output supplied by output engine 135 is in audio format (e.g., as streaming audio). In some implementations, the response content may already be in audio format. In some implementations, output engine 135 additionally or alternatively provides text response content as output (optionally for conversion to audio by device 106), and / or provides other graphic content as output for graphic display by client device 106.
[0053] Agent engine 120 comprises a parameter module 122, an agent selection module 124, a dialogue module 126, and a call module 128. In some implementations, the modules of agent engine 120 may be omitted, combined, and / or implemented by components separated from agent engine 120. Further, agent engine 120 may comprise additional modules not illustrated herein for simplicity.
[0054] The parameter module 122 determines values for parameters such as intent parameters, intent slot parameters, context parameters, etc. The parameter module 122 determines values based on the input provided by the user in the interaction with the automation assistant 110 and optionally based on the client device context. The value for an intent parameter indicates the intent indicated by the user-provided input in the interaction and / or indicated by other data. For example, the value for the intent of the interaction may be one of a plurality of available intents such as "booking", "booking a restaurant reservation", "booking a hotel", "purchasing professional services", "telling jokes", "reminder", "purchasing travel services", and / or one of other intents. The parameter module 122 can determine the intent based on the most recent natural language input provided to the automation assistant in the interaction and / or past natural language input provided in the interaction.
[0055] The values for the intent slot parameters indicate values for parameters with a high level of granularity of the intent. For example, the "booking restaurant reservation" intent may have intent slot parameters for "number of people", "date", "time", "cuisine type", "particular restaurant", "restaurant area", etc. The parameter module 122 can determine the values for the intent slot parameters based on user-provided input in the conversation and / or based on other considerations (such as user-set preferences, past user interactions). For example, the values for one or more intent slot parameters for the "booking restaurant reservation" intent may be based on the most recent natural language input provided to the automated assistant in the conversation and / or past natural language input provided in the conversation. For example, the natural language input "book me a restaurant for tonight at 6:30" can be used by the parameter module 122 to determine the value of "today's date" for the "date" intent slot parameter and the value of "18:30" for the "time" intent slot parameter. It is understood that for many conversations, the parameter module 122 may not be able to resolve the values for all of the intent slot parameters of an intent (or for any parameter for which a value is desired). Such values may be resolved by an agent called in response to the intent (if so).
[0056] The values for the context parameters may include client device context values such as, for example, the user's interaction history on the client device, currently rendered and / or most recently rendered content on the client device, the location of the client device, the current date and / or time, etc.
[0057] When the dialogue module 126 interacts with the user interactively via the client device 106 to select a specific agent, it can utilize one or more grammar models, rules, and / or annotations from the input processing engine 112. The parameter module 122 and / or the agent selection module 124 can optionally generate a prompt for further user input related to interacting with the dialogue module 126 to select a specific agent. The prompt generated by the dialogue module 126 is provided by the output engine 135 for presentation to the user, and further response user input can be received. The further user inputs are each analyzed by the parameter module 122 (optionally as annotated by the input processing engine 112) and / or the agent selection module 124, thereby assisting in the selection of a specific agent. As an example, three candidate agents may first be selected by the agent selection module 124 as potential agents according to the techniques described herein, and the dialogue module 126 can present one or more prompts to the user asking the user to make a selection of a specific one of the three candidate agents. Then, the agent selection module 124 can select a specific agent based on the user selection.
[0058] The agent selection module 124 selects a specific agent to be called from agents 140A - N using the values determined by the parameter module 122. The agent selection module 124 can in addition or alternatively utilize other criteria when selecting a specific agent. For example, the agent selection module 124 may be configured to utilize one or more selection modules of the selection module database 156 and / or may use the agent database 152 when selecting a specific agent.
[0059] Referring to FIG. 2, examples of various components 124A - E that may be included in the agent selection module 124 are shown.
[0060] The request component 124A can transmit a "live" agent request to one or more of agents 140A - N and utilize the responses (and / or lack of responses) from those agents to determine the ability of each agent to generate response content in response to the agent request. For example, the request component 124A can transmit a live agent request to agents 140A - N that are associated with an intent (e.g., within the agent database 152) that matches the value for the intent parameter determined by the parameter module 122. As described herein, the agent request may be based on the values for the parameters determined by the parameter module 122. The agent request may be similar to (or the same as) a call request, but as a result, does not cause a direct call to any agent. The response (or lack of response) from the agent responding to the agent request can directly or indirectly indicate the ability of that agent to generate response content in response to the agent request.
[0061] The agent context component 124B determines characteristics for one or more of agents 140A - N, such as agents 140A - N that are associated with an intent (e.g., within the agent database 152) that matches the value for the intent parameter determined by the parameter module 122. The agent context component 124B can determine the characteristics from the agent database 152. Characteristics for an agent can include, for example, the stored ranking of a particular agent (e.g., ranking by the user population), generally the popularity of a particular agent, the popularity of a particular agent for the intent determined by the parameter module 122, etc.
[0062] The client context component 124C determines features associated with the user of the client device 106 (and / or the client device 106 itself), such as features based on the history of interactions between the user of the client device and the automation assistant 110. For example, these features can include how frequently each of various agents is used by the user, when each of the various agents was last used by the user, the currently rendered and / or most recently rendered content on the client device (e.g., entities present within the rendered content, most recently used applications), the location of the client device, the current date and / or time, and the like.
[0063] The selection model component 124D utilizes one or more selection models of the selection model database 156 to determine one or more of the agents 140A - N that seem appropriate for invocation. For example, the selection model component 124D can utilize the selection model to determine, for each of the multiple agents 140A - N, one or more probabilities or other metrics that indicate whether it is appropriate to invoke the agent. The selection model component 124D can apply to each of the selection models one or more of the parameter values determined by the parameter module 122 and / or one or more of the values determined by components 124A, 124B, and / or 124C.
[0064] The selection component 124E utilizes the output provided by components 124A, 124B, 124C, and / or 124D when selecting one or more agents 140A - N. In some implementations and / or situations, the selection component 124E selects only a single agent out of agents 140A - N without prompting the user to select from among the plurality of agents. In some other implementations and / or situations, the selection component 124E may provide a prompt to the user (e.g., via a prompt generated by the dialogue module 126 in a dialogue) asking the user to select a subset of agents 140A - N and provide user interface input to select one of the agents in that subset. The selection component 124E can provide an indication of the selected agent to the invocation module 128.
[0065] The calling module 128 transmits a call request including the parameters determined by the parameter module 122 to the agent (among agents 140A to N) selected by the agent selection module 124. The transmitted call request calls a specific agent. As described herein, in some situations, the automated assistant 110 can still act as a mediator when a specific agent is called. For example, when acting as a mediator when the user's natural language input is voice input, the input processing engine 112 of the automated assistant 110 converts the voice input into text, and the automated assistant 110 transmits the text (and optionally, the annotation of the text from the input processing engine 112) to a specific agent, receives response content from the specific agent, and the output engine 135 may provide an output based on the response content for presentation to the user via the client device 106. Also, for example, when acting as a mediator, the automated assistant 110 can, in addition or alternatively, analyze the user input and / or the response content to determine whether the interaction with the agent should end, be transferred to an alternative agent, etc. As described herein, in some situations, the interaction is actually transferred to the agent (without the automated assistant 110 acting as a mediator after the transfer) and can be transferred back to the automated assistant 110 after the occurrence of one or more conditions. Further, as also described herein, in some situations, the called agent may be executed by the client device 106 and / or brought to the forefront by the client device 106 (for example, its content can take over the display of the client device 106).
[0066] Each of agents 140A - N may include a context parameter engine, a content engine, and / or other engines. Further, in many implementations, an agent may access various stored models and / or other resources (e.g., its own grammar model and / or content model) when generating response content.
[0067] Also illustrated in FIG. 1 are model engine 150 and record database 158. As described in more detail herein, record database 158 may include stored information based on various interactions between automated assistant 110 and agents 140A - N. Model engine 150 can utilize such information when generating one or more selection models of selection model database 156. Additional explanation is provided here.
[0068] Referring now to FIGS. 3 - 10, additional explanation of the various components of the environment of FIG. 1 is described.
[0069] FIG. 3 is an example diagram showing how agent requests and responses can be used in selecting a particular agent and / or stored in record database 158 for use in generating an agent selection model.
[0070] In FIG. 3, natural language input 171 is received by input processing engine 112 of automated assistant 110. As an example, natural language input 171 may be "table for 4, outdoor seating, Restaurant A". Input processing engine 112 generates annotated input 172 and supplies annotated input 172 to parameter module 122.
[0071] Parameter module 122 generates a value for parameter 173 based on the annotated input 172 and / or based on the client device context. Continuing with this example, parameter module 122 can generate a "restaurant booking" value for the intent parameter, an "outdoor" value for the seating arrangement preference intent slot parameter, and a "Restaurant A" value for the restaurant location intent slot parameter.
[0072] Request component 124A generates an agent request based on the value for parameter 173. As shown by the "AR" directed arrow in FIG. 3, request component 124A transmits the agent request to each of the plurality of agents 140A - D. In some implementations, request component 124A can select agents 140A - D based on determining that they are associated with the "restaurant booking" intent. In some other implementations, request component 124A can transmit the agent request to agents 140A - D and / or additional agents (e.g., it may be sent to all of agents 140A - N) regardless of the intent. For example, the intents that can be handled by various agents may be unknown, and / or the intent may not be derivable from the conversation.
[0073] As indicated by the "R" directed arrow in FIG. 3, upon receiving responses from each of agents 140A - D in response to transmitting an agent request to agents 140A - D, request component 124A receives the responses. Each response indicates the ability of the corresponding one of these agents 140A - D to resolve the agent request. For example, a response from a given agent may be an indication taking two values, a non - binary confidence scale, actual response content (or no content / error), etc. Although responses from each of the agents are shown in FIG. 3, in some implementations or situations, one or more of agents 140A - D may not be able to respond, which may indicate that the corresponding agent is unable to respond (e.g., the agent is unable to process the agent request and / or is offline). The agent request is transmitted to agents 140A - D without an active call to agents 140A - D.
[0074] Request component 124A stores, within record database 158, the responses (and / or decisions made based on the responses) and the agent requests. For example, request component 124A can store the agent requests and indications of responses for each of agents 140A - D. For example, request component 124A can store an agent request and an indication showing that agent 140A was unable to respond, agent 140B was able to respond, agent 140C was unable to respond, etc.
[0075] Request component 124A provides response 174 (and / or decisions made based on the response) to selection component 124E.
[0076] The selection component 124E selects a particular agent 176 using the response 174 and, optionally, in addition thereto, may utilize features provided by the agent context component 124B and / or the client context component 124C. As an example, the selection component 124E may select a particular agent 176 based only on the response 174 (e.g., select the agent having the response that best demonstrates the ability to respond). As another example, the selection component 124E may utilize the response 174 and the ranking of one or more of the agents 140A - D provided by the agent context component 124B. For example, the selection component 124E may first select two of the agents 140A - D whose response best demonstrates the ability to respond, and then select only one of them based on the selected agent having a higher ranking than the non - selected agents. In yet another example, the selection component 124E may utilize the response 174 and the usage feature history provided by the client context component 124C. For example, the selection component 124E may first select two of the agents 140A - D whose response best demonstrates the ability to respond, and then select only one of them based on the selected agent being more frequently and / or more recently utilized by the user interacting with the automated assistant 110.
[0077] The selection component 124E provides a specific agent 176 to the invocation module 128. The invocation module 128 can invoke a specific agent. For example, the invocation module 128 can invoke a specific agent by transmitting an invocation request based on the value for the parameter 173 to the specific agent. Also, for example, the invocation module 128 can cause the output engine 135 to provide an output based on the response content already received from the specific agent, and optionally, can invoke the specific agent to generate further response content in response to further received user interface input (if any) received in response to providing the output.
[0078] FIG. 4 is a flowchart showing an exemplary method 400 according to the implementations disclosed herein. For convenience, the operations of the flowchart of FIG. 4 are described with reference to a system that executes the operations. This system can comprise various components of various computer systems, such as one or more components of the automated assistant 110. Further, although the operations of method 400 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0079] In block 450, the system receives user input. In some implementations, the user input received in block 450 is free-form natural language input.
[0080] In block 452, the system determines whether the received user input indicates an agent action. For example, the system may parse the text of the received natural language input (or the text converted from the received voice natural language input) and determine whether the parsed text maps to an agent action. For example, the system may determine whether the parsed text maps to an agent action based on whether the words / phrases included in the text match the words / phrases stored in relation to the agent action. Also, for example, the system may determine whether one or more entities derived from the text match one or more entities stored in relation to the agent action. In yet another example, the system may assume that the input maps to an agent action if the system determines that it cannot generate local response content in response to the input. Note that in some implementations, the system may determine that the input maps to an agent action even when the system can generate local response content in response to the input. For example, the system may determine that it can generate local response content and that one or more agents may also potentially generate agent response content. In some of those examples, the system may include local agents among the agents to be considered in blocks 462, 464, 466, and 468 (described below).
[0081] If the system determines at block 452 that the action intended by the agent is not instructed, the system proceeds to blocks 454, 456, and 458. At block 454, the system generates local response content without invoking the agent. For example, the system may utilize the system's local grammar model and / or local content model to generate local response content. At block 456, the system provides an output based on the local response content. For example, the output may be the local response content or a transformation of the local response content (e.g., text-to-speech transformation). The output is provided for presentation via a client device (e.g., by voice or image). At block 458, the system waits for additional voice input and returns to block 450 after receiving the additional voice input.
[0082] If the system determines during the iteration of block 452 that an agent action is instructed, the system proceeds to block 460. At block 460, the system determines whether a single agent is specified in the user input of block 450 and / or whether a single agent can be ambiguously resolved when it is not.
[0083] If the system determines at block 460 that a single agent is instructed, the system proceeds to block 474.
[0084] If the system determines at block 460 that no single agent has been instructed, the system proceeds to blocks 462, 464, 466, and 468. At block 462, the system generates an agent request based on the user input in the most recent iteration of block 450 and / or based on previous user input and / or other criteria. At block 464, the system selects a plurality of agents from the corpus of available agents. For example, the system may select all of the agents in the corpus or a subset of the agents (e.g., agents having the intent indicated by the agent request). The agents may include only non-local agents and / or both local and non-local agents. At block 466, the system transmits the agent request from block 462 to each of the agents selected at block 464.
[0085] At block 468, the system receives one or more responses from the agents in response to the transmission of block 466. At block 470, the system stores an association between the response received at block 468 (or a decision made based on the response) and the agent request transmitted at block 466. Such stored associations may be used when generating an agent selection model (e.g., in method 500 of FIG. 5). At block 472, the system selects a single agent based on the response received at block 468 and / or based on other criteria.
[0086] In some implementations, at block 472, the system selects a subset of agents using the responses received at block 468 and provides, as output for presenting instructions of the agents in the subset to the user, a user selection in response to the output is utilized to select a single agent from the subset of agents. In some versions of those implementations, the system may also, at block 470, store the instructions of the single agent selected by the user. The instructions of the agent may be, for example, the name or other identifier of the agent and / or the instructions of the agent response content included in the response received at block 468.
[0087] At block 474, the system transmits the call request to a single agent. The single agent can be the agent selected at block 472 (if the determination at block 460 is "no"), or the single agent indicated by user input (if the determination at block 460 is "yes"). For example, the system can transmit the call request over one or more communication channels and, optionally, utilize an API. In some implementations, the call request includes values for various call parameters as described herein. In some implementations, the system may first prompt the user to confirm that the user desires to utilize a single agent before proceeding from block 472 to block 474. In those implementations, the system may require affirmative user input in response to that prompt before proceeding from block 472 to block 474. In other implementations, the system may automatically proceed from block 472 to block 474 without first prompting the user for confirmation.
[0088] In an optional block 476, the system may update global parameter values based on the interaction between the user and a single agent. For example, when the user provides further natural language input to a single agent, as a result, the single agent may define a value for a global parameter that is still undefined (e.g., a specific intent slot parameter), or the single agent may modify a previously defined value for the parameter. The single agent may well provide such updated values to the system, and the system may update the global parameter values to reflect the updated values.
[0089] In block 478, the system receives an agent switch input. The agent switch input is an input (e.g., natural language input) from the user indicating a desire to switch to a different agent. For example, inputs such as "talk to another agent", "different agent", "try it with Agent X", etc. may be agent switch inputs.
[0090] In response to receiving an agent switch input, at block 480, the system transmits a call request to an alternative agent. The call request to the alternative agent calls the alternative agent in place of the single agent called at block 474. The call request may optionally include the updated global parameter values updated at block 476. In this way, values derived through interaction with the first agent are transferred to the subsequently called agent, thereby enhancing the efficiency of interaction with the subsequently called agent. In some situations, which alternative agent is selected to transmit an additional call request may be based on the agent switch input itself (e.g., when referring to one of the alternative agents by name or by characteristics) and / or based on other factors (e.g., the agent request may be selected based on the updated global parameter values again and the response used to select the alternative agent). In some implementations, before transmitting the call request to the alternative agent at block 480, the system may check to confirm that the alternative agent is likely to be able to generate response content. For example, the system can send the agent request to the alternative agent, make such a determination based on the response, and / or make such a determination based on information about the agent in the agent database 152.
[0091] Figure 5 is a flowchart showing another exemplary method 500 according to the implementations disclosed herein. Figure 5 shows an example of generating a selection model of the selection model database 156 based on agent requests and associated responses. In the example of Figure 5, the selection model is a machine learning model such as a deep neural network model.
[0092] For the sake of simplicity, the operations of the flowchart of FIG. 5 are described with reference to the system that performs the operations. This system may comprise various components of various computer systems, such as model engine 150. Further, although the operations of method 500 are shown in a particular order, this is not meant to be limiting. One or more operations may be changed in order, omitted, or added.
[0093] In block 552, the system selects an agent request and the associated response. As an example, the agent request and the associated response may be selected from record database 158. In some implementations, the selected agent request and the associated response may have been generated by request component 124A as illustrated in FIG. 3 and / or stored in an iteration of block 470 of method 400 of FIG. 4. In some other implementations, agent requests and / or associated responses generated through other techniques may be utilized. For example, the agent request may be transmitted to the agent and the response may be received in a “non-live” manner. In other words, the agent request and the response need not necessarily have been generated when actively selecting a particular agent in past human / automation assistant interactions. As a non-limiting example, the agent request may include those based on natural language input provided to the agent immediately following a call that is a “naked” call. For example, the natural language input provided to “Agent A” immediately following a naked call to “Agent A” may be utilized to generate an agent request that is then transmitted to a plurality of additional agents.
[0094] As an example, assume that the agent requests selected in the iteration of block 552 include a "restaurant booking" value for the intent parameter, an "outdoor" value for the seating arrangement preference intent slot parameter, and a "Restaurant A" value for the restaurant location intent slot parameter. Further assume that the selected and associated responses include a binary response of "yes" (able to generate response content) or "no" (unable to generate response content) to the agent request. In particular, assume that the associated responses indicate that Agents 1 - 5 generated a "yes" response and Agents 6 - 200 generated a "no" response.
[0095] In block 554, the system generates a training instance based on the selected agent request and the associated response. Block 554 includes sub - blocks 5541 and 5542.
[0096] In sub - block 5541, the system generates a training instance input for the training instance based on the agent request and optionally based on additional values. Continuing with the example, the system can generate a training instance input that includes the values of the agent request for the intent parameter, the seating arrangement preference intent slot parameter, and the restaurant location intent slot parameter. The system may include a "null" value (or other value) in the training instance for dimensions of the training instance input that are undefined by the agent request. For example, if the input dimensions of the machine learning model to be trained include inputs for other intent slot parameters (of the same intent and / or other intents), the null value can be utilized in the training instance for such inputs.
[0097] In sub-block 5542, the system generates a training instance output of the training instance based on the response. Continuing with the example, the system can generate a training instance output that includes a "1" (or other "positive value") for each output dimension corresponding to Agents 1 - 5 (which generated a "yes" response) and a "0" (or other "negative value") for each output dimension corresponding to Agents 6 - 200 (which generated a "no" response).
[0098] In block 556, the system determines whether there are additional agent requests and associated responses. If so, the system returns to block 552, selects another agent request and associated response, and then generates another training instance based on the selected agent request and associated response.
[0099] Blocks 558 - 566 can be executed following multiple iterations of blocks 552, 554, and 556, or in parallel.
[0100] In block 558, the system selects the training instances generated in the iteration of block 554.
[0101] In block 560, the system applies the training instance as an input to the machine learning model. For example, the machine learning model can have input dimensions corresponding to the dimensions of the training instance input generated in block 5541.
[0102] In block 562, the system generates an output on the machine learning model based on the applied training instance input. For example, the machine learning model can have output dimensions corresponding to the dimensions of the training instance output generated in block 5541 (e.g., each dimension of the output can correspond to an agent and / or an agent and intent).
[0103] In block 564, the system updates the machine learning model based on the generated output and the training instance output. For example, the system can determine an error based on the output generated in block 562 and the training instance output, and backpropagate the error on the machine learning model.
[0104] In block 566, the system determines whether there are one or more additional unprocessed training instances. If so, the system returns to block 558, selects an additional training instance, and then executes blocks 560, 562, and 564 based on the additional unprocessed training instance. In some implementations, at block 566, the system can determine not to process additional unprocessed training instances if one or more training criteria are met (e.g., a threshold number of epochs have occurred and / or training of a threshold duration has occurred). Although method 500 is described with respect to non-batch learning techniques, batch learning can be utilized in addition to and / or alternatively.
[0105] The machine learning model trained by method 500 may hereinafter be used to predict the probability for each of a plurality of available agents (and optionally intents) based on the current conversation, each of these probabilities indicating the probability that an agent can appropriately handle a call request based on the conversation. For example, values based on the current conversation are applied as input to the trained machine learning model, and outputs can be generated on the model, the outputs each including a plurality of values corresponding to agents, the values each indicating the probability (e.g., a value between 0 and 1) that the corresponding agent can generate appropriate response content if called. For example, if 200 available agents are represented by the model, 200 values may be included in the output, each value corresponding to one of the agents and indicating the probability that the agent can generate appropriate response content. In this way, the trained machine learning model effectively provides insights into the capabilities of various agents through training based on their responses to various real-world agent requests. The trained machine learning model can be used to determine the capabilities of various agents to generate responses to an input even when the agent database 152 and / or other resources do not explicitly indicate their capabilities for the input based on the input.
[0106] FIG. 5 shows an example of an agent selection model that can be generated and used. However, as described herein, additional and / or alternative agent selection models may be used when selecting a particular agent. Such additional and / or alternative agent selection models may optionally be machine learning models trained based on training instances different from those described with respect to FIG. 5.
[0107] As an example, a selection model may be generated based on the past explicit selections of agents by various users, and such a selection model may, in addition or alternatively, be used when selecting a particular agent. For example, as described with respect to block 472 in FIG. 4, in some implementations, instructions for a plurality of agents may be presented to a user, and a user selection of a single agent among the plurality of agents may be used to select a single agent from the plurality of agents. Such explicit selections by multiple users can be used for the selection model to generate. For example, a training instance similar to that described above with respect to method 500 may be generated, but the training instance output of each training instance may be generated based on the agent selected by the user. For example, for a training instance, a "1" (or other "positive value") may be used for the output dimension corresponding to the selected agent, and a "0" (or other "negative" value) may be used for each of the output dimensions corresponding to all other agents. Also, for example, for a training instance, a "1" (or other "positive value") may be used for the output dimension corresponding to the selected agent, a "0.5" (or other "intermediate value") may be used for the output dimension corresponding to the other agents presented to the user but not selected, and a "0" (or other "negative" value) may be used for each of the output dimensions corresponding to all other agents. In this and other ways, exemplary selections of agents by users can be used when generating one or more agent selection models.
[0108] FIG. 6 is a flowchart showing another exemplary method 600 according to the implementations disclosed herein. FIG. 6 shows an example of using an agent selection model, such as an agent selection model generated based on method 500 of FIG. 5, when selecting a single agent to call.
[0109] For the sake of simplicity, the operations of the flowchart of FIG. 6 are described with reference to the system that executes the operations. This system may include various components of various computer systems, such as one or more components of the automation assistant 110. Further, although the operations of method 600 are shown in a particular order, this does not mean a limitation. One or more operations may be changed in order, omitted, or added.
[0110] In block 650, the system receives user input. Block 650 may share one or more aspects common to block 450 of FIG. 4.
[0111] In block 652, the system determines whether the received user input indicates an agent action. Block 652 may share one or more aspects common to block 452 of FIG. 4.
[0112] If the system determines in block 652 that the action intended by the agent is not indicated, the system proceeds to blocks 654, 656, and 658. In block 654, the system generates local response content without invoking the agent. In block 656, the system provides an output based on the local response content. In block 658, the system waits for additional voice input, and after receiving the additional voice input, returns to block 650. Blocks 654, 656, and 658 may share one or more aspects common to blocks 454, 456, and 458 of FIG. 4.
[0113] If the system determines in the iteration of block 652 that an agent action is indicated, the system proceeds to block 660. In block 660, the system determines whether a single agent is specified in the user input 650 and / or whether a single agent can be ambiguously resolved if not. Block 660 may share one or more aspects common to block 460 of FIG. 4.
[0114] If the system determines at block 660 that a single agent is being directed, the system proceeds to block 680.
[0115] If the system determines at block 660 that a single agent is not being directed, the system proceeds to blocks 672, 674, 676, and 678. At block 672, the system generates input features based on the user input in the most recent iteration of 650 and / or based on previous user input and / or other criteria. For example, the system can generate input features that include values for parameters determined based on user input, such as values for intent parameters, intent slot parameters, etc. Also, for example, the system can generate values based on the current client device context.
[0116] At block 674, the system applies the input features to an agent selection model.
[0117] At block 676, the system generates probabilities for each of a plurality based on the application of the input to the agent selection model. Each of these probabilities indicates the ability of the corresponding agent to generate appropriate response content.
[0118] In block 678, the system selects a single agent based on probability and / or other criteria. In some implementations, the system selects a single agent based on it having the highest probability of generating appropriate response content. In some other implementation, the system selects a single agent based on additional criteria. For example, the system can select an initial subset of agents based on probability, transmit “live” agent requests to the agents in the subset, and utilize the “live” responses to the agent requests in selecting a single agent. As another example, the system can, in addition to or alternatively, select a single agent based on the history of the user's interactions with the client device (e.g., how frequently a single agent has been utilized by the user, how recently a single agent has been utilized by the user), the currently rendered and / or most recently rendered content on the client device, the location of the client device, the current date and / or time, the ranking of a single agent (e.g., ranking by the user population), the popularity of a single agent (e.g., popularity among the user population), etc.
[0119] In block 680, the system transmits a call request to a single agent. The single agent can be the agent selected in block 678 (if the determination in block 660 was “no”), or the single agent indicated by user input (if the determination in block 660 was “yes”). Block 680 can share one or more aspects common with block 474 of FIG. 4.
[0120] FIG. 7 is a flowchart showing another exemplary method 700 according to the implementations disclosed herein. FIG. 7 shows an example of a method that can be performed by one or more of the agent selection models.
[0121] For the sake of simplicity, the operation of the flowchart of FIG. 7 is described with reference to the system that executes the operation. This system may include various components of various computer systems, such as one or more components of one of the agents 140A - N. Further, although the operations of method 700 are shown in a particular order, this is not meant to be limiting. One or more operations may be reordered, omitted, or added.
[0122] In block 752, the system receives an agent request from the automation assistant. In some implementations, the agent request is flagged as an agent request or otherwise indicated.
[0123] In block 754, the system generates a response based on the agent request. For example, the system analyzes the content of the agent request, determines whether it can respond to the request, and may generate a response based on whether it can respond to the request. For example, the request may include a value for an intent parameter, and the system may determine whether it can respond to the request based on whether it can respond to the intent indicated by the value. For example, if the value for the intent parameter is "booking" but the system can only handle the "gaming" intent, it may determine that it cannot respond and generate a response (e.g., "0" or other "negative" response value) indicating that it cannot fully respond. Also, for example, the request may include a value for an intent slot parameter, and the system may determine whether it can respond to the request based on whether it can respond to the intent slot parameter and / or the value. For example, if the intent slot parameter is not supported by the system but other parameters and values of the agent request are supported, the system may generate a response (e.g., "0.5" or other "partial" response value) indicating that it can handle some but not all of the parameters of the agent request. As yet another example, the intent slot parameter may be supported by the system, but a particular value may not be supported by the system. For example, the intent slot parameter may be a "geographic region" parameter, and the value may be a geographic region that does not receive services from the system. In such a scenario, the system may generate a response indicating that it cannot respond, or more specifically, indicating that it can handle some but not all of the values of the agent request.
[0124] In block 756, the system provides (e.g., transmits) the agent response to the automated assistant without being called by the automated assistant.
[0125] In block 758, the system can subsequently receive a call request from the automation assistant. In some implementations, the call request may indicate that the system should actually or virtually take over the conversation. If the call request indicates that the system should actually take over the conversation, the system may establish a direct network communication session with the corresponding client device. If the call request indicates that the system should virtually take over the conversation, the system can take over the conversation while communicating directly with the component that provided the call request and / or related components.
[0126] In block 760, the system generates response content based on the values of the parameters included in the call request.
[0127] In block 762, the system provides the response content. For example, if the call request indicates that the system should virtually take over the conversation and / or only perform the intended action without participating in the conversation, the system can transmit the response content to the component that sent the call request (or related components). Also, for example, if the call request indicates that the system should actually take over the conversation, the system can transmit the response content to the corresponding client device.
[0128] Blocks 758, 760, and 762 are shown in dashed lines in FIG. 7 to indicate that they may not be executed in some situations. For example, as described herein, in some implementations, the system may receive an agent request without even receiving the corresponding call request.
[0129] Figures 8 and 9 each show an example of a conversation that can take place between user 101, voice-responsive client device 806, and user 101, the automated assistant associated with client device 806, and an agent. Client device 806 includes one or more microphones and one or more speakers. One or more aspects of the automated assistant 110 of FIG. 1 can be implemented on client device 806 and / or on one or more computing devices in network communication with client device 806. Thus, for simplicity of explanation, automated assistant 110 is referenced in the description of FIGS. 8 and 9.
[0130] In FIG. 8, the user provides an utterance input 880A of "Assistant, deliver flowers to my house today". A voice input corresponding to the utterance input is generated by device 806 and provided to automated assistant 110 (e.g., as a streaming voice input). Even if utterance input 880A does not specify a particular agent, automated assistant 110 can use utterance input 880A to select a single agent from among a plurality of available agents based on one or more of the techniques described herein (e.g., based on an agent selection model).
[0131] In response to utterance input 880A and to having selected a single agent, automated assistant 110 can generate and provide an output 882A of "Sure, Agent 1 can handle that". Further, automated assistant 110 invokes "Agent 1", which then provides an agent output 882B of "Hi, this is Agent 1. What kind of flowers?"
[0132] In response to agent output 882B, the user provides a further utterance input 880B of "12 red roses". The voice input corresponding to the utterance input is generated by device 806 and provided to automated assistant 110, which transfers the utterance input (or its conversion and / or annotation) to "Agent 1". The further utterance input 880B specifies a value for the "flower type" intent slot parameter of the yet unspecified "order flowers" intent. Automated assistant 110 may update the global value for the "flower type" intent slot parameter based on the further utterance input 880B (either directly or based on the indication of the value provided by "Agent 1").
[0133] In response to the further utterance input 880B, "Agent 1" provides a further agent output 882C of "I can have them delivered at 5:00 for a total of $60. Want to order?".
[0134] In response to the further agent output 882C, the user provides a further utterance input 880C of "Assistant, switch me to another flower agent". Automated assistant 110 can recognize such a further utterance as a switch input and select an appropriate alternative agent. For example, based on an agent selection model and / or "live" agent requests, automated assistant 110 can determine that "Agent 2" can handle the intent with various values for the intent slot parameters (including the value for "flower type").
[0135] In response to further utterance input 880C and also to the selection of an alternative agent, the automated assistant 110 may generate and provide the output 882D "Sure, Agent 2 can also handle". Further, the automated assistant 110 may call "Agent 2" with a call request that includes updated global values for the "flower type" intent slot parameter. Then "Agent 2" provides the agent output 882E "Hi, this is Agent 2. I can have the 12 red roses delivered at 5:00 for $50. Order?". In particular, this output is generated based on updated global values for the "flower type" intent slot parameter that were updated in response to the interaction with the previously called "Agent 1".
[0136] The user then provides further utterance input 880F of "Yes" to have "Agent 2" satisfy the intent with the specified values for the intent slot parameters.
[0137] In FIG. 9, the user provides the utterance input 980A "Assistant, table for 2, outdoor seating, 6:00 tonight at Hypothetical Cafe". The voice input corresponding to the utterance input is generated by the device 806 and provided to the automated assistant 110 (e.g., as streaming voice input). Even if the utterance input 980A does not specify a particular agent, the automated assistant 110 may utilize the utterance input 980A to select a single agent from among a plurality of available agents based on one or more of the techniques described herein (e.g., based on an agent selection model).
[0138] In response to the speech input 980A and also to the selection of a single agent, the automated assistant 110 may invoke "Agent 3". "Agent 3" can satisfy an intent with a specified value for the intent slot parameter without participating in further dialogue with the user. "Agent 3" generates response content indicating that the reservation is complete and provides it to the automated assistant 110, and a text confirmation will be sent to the user 101. In response thereto, the automated assistant 110 provides the output 982A of "Reserved with Agent 3. Text confirmation on its way."
[0139] FIG. 10 shows another client device 1006 and a display screen 1040 of the client device 1006. The client device 1006 may include and / or communicate with the automated assistant 110. The display screen 1040 includes a response interface element 1088 that a user may select to generate user input via a virtual keyboard, and an audio response interface element 1089 that a user may select to generate user input via a microphone. In some implementations, the user may generate user input via the microphone without selecting the audio response interface element 1089. For example, during a dialogue, active monitoring of audible user interface input via the microphone may be performed to eliminate the need for the user to select the audio response interface element 1089. In some and / or other implementations of those implementations, the audio response interface element 1089 may be omitted. Further, in some implementations, the response interface element 1088 may be omitted in addition to and / or alternatively (e.g., the user may provide only audible user interface input). The display screen 1040 also includes system interface elements 1081, 1082, 1083 that a user may interactively operate to cause the client device 1006 to perform one or more actions.
[0140] In FIG. 10, input 1080A is the same as input 880A of FIG. 8. Similarly, outputs 1082A and 1082B are the same as outputs 882A and 882B of FIG. 8. FIG. 10 is shown as an example of a graphical interaction with the automation assistant 110 and / or an agent (instead of the spoken / audible interaction of FIG. 8). Although FIG. 10 shows an example of a graphical interaction, different examples are contemplated. For example, instead of “Agent 1” providing output 1082B via a common interface with an automation assistant as illustrated in FIG. 10, “Agent 1” could instead be a separate application that is launched and / or brought to the foreground in response to a call. For example, after output 1082A in FIG. 10, the “Agent 1” application interface could replace the automation assistant interface of FIG. 10 and then provide an output similar to output 1082B.
[0141] FIG. 11 is a block diagram of an exemplary computing device 1110 that may optionally be utilized to implement one or more aspects of the techniques described herein. In some implementations, one or more of devices 106, automation assistant 110, 3P agent, and / or other components may include one or more components of exemplary computing device 1110.
[0142] Computing device 1110 typically includes at least one processor 1114 that communicates with a number of peripheral devices via a bus subsystem 1112. These peripheral devices can include, for example, a storage device subsystem 1124, which includes a memory subsystem 1125 and a file storage device subsystem 1126, a user interface output device 1120, a user interface input device 1122, and a network interface subsystem 1116. The input and output devices enable interaction between the user and the computing device 1110. The network interface subsystem 1116 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.
[0143] The user interface input device 1122 can include a keyboard, a mouse, a trackball, a pointing device such as a touchpad or a graphics tablet, a scanner, a touch screen incorporated into a display, a voice input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into the computing device 1110 or onto a communication network.
[0144] The user interface output device 1120 may include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for generating a visible image. The display subsystem may also include a non-visual display via an audio output device or the like. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 1110 to the user or another machine or computing device.
[0145] The storage device subsystem 1124 stores programming and data structures that implement some or all of the functions of the modules described herein. For example, the storage device subsystem 1124 may include logic circuitry for executing selected aspects of the methods of FIGS. 4, 5, 6, and / or 7.
[0146] These software modules are generally executed by the processor 1114, either alone or in combination with other processors. The memory 1125 used in the storage device subsystem 1124 can include a number of memories, including a main random access memory (RAM) 1130 for storing instructions and data during program execution, and a read only memory (ROM) 1132 in which fixed instructions are stored. The file storage device subsystem 1126 can include a persistent storage area for program and data files, and can include a hard disk drive, a floppy disk drive with an associated removable medium, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functions of some implementations can be stored in the storage device subsystem 1124 by the file storage device subsystem 1126, or in another machine accessible by the processor 1114.
[0147] The bus subsystem 1112 comprises a mechanism for enabling the various components and subsystems of the computing device 1110 to communicate with each other as intended. Although the bus subsystem 1112 is schematically illustrated as a single bus, multiple buses may be used in alternative implementations of the bus subsystem.
[0148] The computing device 1110 may be of various types including a workstation, server, computing cluster, blade server, server farm, or other data processing system or computing device. Since computers and networks are constantly changing in nature, the description of the computing device 1110 shown in FIG. 11 is intended only as a specific example for purposes of illustrating some implementations. Many other configurations of the computing device 1110 are possible having more or fewer components than the computing device shown in FIG. 11.
[0149] In some implementations described in this specification, where personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, and the user's activities and demographic information) may be collected or used, the user is provided with one or more opportunities to control whether the information is collected, whether the personal information is stored, whether the personal information is used, and how that information is collected, stored, and used with respect to the user. That is, implementations of the systems and methods described in this specification collect, store, and / or use user personal information only after receiving explicit permission from the relevant user to do so. For example, the user is provided with a control function regarding whether a program or feature collects user information about that particular user or other users related to that program or feature. Each user from whom personal information is collected is presented with one or more options that enable the user to control the information collection related to the user and give consent or permission regarding whether the information is collected and which portions of the information should be collected. For example, the user may be provided with one or more such control options on a communication network. In addition, specific data may be well processed in one or more ways before it is stored or used, and thus personally identifiable information is removed. As an example, a user's identity may be handled such that no personally identifiable information can be determined. As another example, a user's geographical location may be generalized to a broader area such that the user's specific location cannot be determined.
Explanation of Signs
[0150] 101 User 106 Client Device 110 Automated Assistant 112 Input Processing Engine 171 Natural Language Input 120 Agent Engine 122 Parameter Module 124 Agent Selection Module 124A~E Components 126 Dialogue Module 128 Call Module 130 Local Content Engine 135 Output Engine 140A~N Agents 150 Model Engine 152 Agent Database 154 Local Content Model 156 Selection Module Database, Selection Model Database 158 Record Database 171 Natural Language Input 172 Annotated Input 173 Parameters 174 Responses 176 Agents 400 Method 500 Method 600 Method 806 Client Device 880A Utterance Input 880B Utterance Input 880C Utterance Input 880F Utterance Input 882A Output 882B Agent Output 882C Agent Output 882D Output 882E Agent Output 980A Utterance Input 982A Output 1006 Client Device 1040 Display Screen 1080A Input 1081, 1082, 1083 System Interface Elements 1082A and 1082B Output 1088 Response Interface Element 1089 Voice Response Interface Element 1110 Computing device 1112 Bus subsystem 1114 Processor 1116 Network interface subsystem 1120 User interface output device 1122 User interface input device 1124 Storage device subsystem 1125 Memory subsystem 1126 File storage device subsystem 1130 Main random access memory 1132 Read-only memory
Claims
Claim 1 A method executed by one or more processors, comprising: receiving, in a human / automation assistant dialogue, a natural language input instance generated based on a user interface input provided by a user, wherein the natural language input instance does not explicitly indicate an agent invoked based on the natural language input instance; before invoking an agent in response to the natural language input instance, selecting a specific agent from a plurality of candidate agents, wherein selecting the specific agent is based on the natural language input instance and an agent selection model, and one or more context values added to the natural language input instance and added to other natural language input instances provided in the human / automation assistant dialogue, the one or more context values being based on at least one entity present in currently rendered content on a client device on which the user interface input is received; and wherein selecting the specific agent is performed without providing a user interface output that explicitly asks the user to select the specific agent from one or more other agents of the plurality of candidate agents; in response to receiving the natural language input instance and in response to selecting the specific agent, transmitting, via an application programming interface, a call request to the selected specific agent, wherein the call request causes the specific agent to be invoked and to generate specific response content to be presented to the user via one or more user interface output devices; and only the selected specific agent is invoked in response to receiving the natural language input instance, receiving an additional natural language input instance generated based on additional natural language input provided by the user while the specific agent is being invoked. Determining that the additional natural language input instance desires to switch to an alternative agent among the plurality of candidate agents; In response to determining that the additional natural language input instance desires to switch to the alternative agent, Transmitting an additional call request to the alternative agent, wherein the additional call request includes at least one value based on the natural language input instance or an additional natural language input instance received during the invocation of the particular agent; A method further comprising. **Claim 2** The method according to claim 1, wherein the at least one value included in the additional call request transmitted to the alternative agent is based on an additional natural language input instance received during the invocation of the particular agent. **Claim 3** The method according to claim 1, wherein the at least one value included in the additional call request transmitted to the alternative agent is based on the natural language input instance. **Claim 4** The method according to claim 1, wherein the one or more context values further include a given context value based on the amount of interaction of the user involved in the human / automation assistant dialogue with the particular agent. **Claim 5** The method according to claim 1, wherein the one or more context values further include a given context value based on the licensee of the interaction of the user involved in the human / automation assistant dialogue with the particular agent. **Claim 6** The method according to claim 1, wherein the one or more context values further include a given context value based on the amount of interaction of the user with the particular agent and an additional given context value based on the licensee of the interaction of the user with the particular agent. **Claim 7** The method according to claim 1, wherein the one or more context values further include the location of the client device where the user interface input is received. **Claim 8** A memory storing instructions; One or more processors for executing the instructions to perform the method according to any one of claims 1 to 7; A system comprising.
Citation Information
Patent Citations
System and method for interactive control
JP2005242243A
Voice interaction device and method
JP2008090545A
Voice interaction method, and device
WO2014203495A1