Providing related query to secondary automated assistant based on past interaction

By employing a primary automatic assistant to manage and direct queries to secondary assistants based on past interactions, the inefficiencies and inconsistencies in existing automatic assistant systems are mitigated, resulting in improved query processing efficiency and consistent responses.

JP2025094132APending Publication Date: 2025-06-24GOOGLE LLC

Patent Information

Application Number
JP2025045558
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-11-23
Filing Date
2025-03-19
Publication Date
2025-06-24

Smart Images

  • Figure 2025094132000001_ABST
    Figure 2025094132000001_ABST
Patent Text Reader

Abstract

To provide a system and a method for providing audio data from an initially invoked automated assistant to a subsequently invoked automated assistant.SOLUTION: An initially invoked automated assistant may be invoked by a user utterance which is followed by audio data that includes a query. The query is provided to a secondary automated assistant for processing. Subsequently, a user can submit a query related to the first query. In response, the initially invoked automated assistant provides the query to the secondary automated assistant instead of providing the query to another secondary automated assistant on the basis of similarity between the first query and the subsequent query.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to providing relevant queries to a secondary automatic assistant based on past conversations.

Background Art

[0002] Humans can participate in human-computer dialogues with a conversational software application, referred to herein as an "automatic assistant" (also called a "digital agent", "chatbot", "conversational personal assistant", "intelligent personal assistant", "assistant application", "conversational agent", etc.). For example, a human (sometimes called a "user" when interacting with an automatic assistant) can provide commands and / or requests to the automatic assistant using oral natural language input (i.e., utterances) that may be converted to text and then processed and / or by providing text (e.g., typed) natural language input. The automatic assistant responds to the request by providing a response user interface output that can include audible and / or visual user interface outputs.

[0003] As described above, many automatic assistants are configured to be interacted with via verbal utterances such as call instructions and subsequent verbal queries. To protect the user's privacy and / or to conserve resources, often the user must explicitly invoke the automatic assistant before the automatic assistant fully processes the verbal utterance. The explicit invocation of the automatic assistant typically occurs in response to a specific user interface input received at the client device. The client device includes an assistant interface that provides an interface for the user of the client device to interface with the automatic assistant (e.g., receive verbal and / or typed input from the user and provide audible and / or graphical responses), and interfaces with one or more additional components that implement the automatic assistant (e.g., a remote server device that processes user input and generates an appropriate response).

[0004] Some user interface inputs that can be used to invoke an automatic assistant via a client device include hardware and / or virtual buttons on the client device for invoking the automatic assistant (e.g., tapping a hardware button, selecting a graphical interface element displayed by the client device). Many automatic assistants can additionally or alternatively be invoked in response to one or more verbal call phrases, also known as "hot words / phrases" or "trigger words / phrases". For example, verbal call phrases such as "Hey, Assistant", "OK Assistant", and / or "Assistant" can be spoken to invoke the automatic assistant.

[0005] Often, a client device that includes an assistant interface includes one or more locally stored models that the client device utilizes to monitor for the occurrence of a spoken invocation phrase. Such a client device can utilize the locally stored models to locally process received audio data and discard any audio data that does not include a spoken invocation phrase. However, if the local processing of the received audio data indicates the occurrence of a spoken invocation phrase, then the client device causes the audio data and / or subsequent audio data to be further processed by the auto assistant. For example, if the spoken invocation phrase is "Hey, Assistant" and the user says "Hey, Assistant, what time is it", the audio data corresponding to "what time is it" is processed by the auto assistant based on the detection of "Hey, Assistant" and can be utilized to provide an auto assistant response for the current time. On the other hand, if the user simply says "what time is it" (without first speaking the invocation phrase or providing an alternative invocation input), no response from the auto assistant is provided as a result of there being no invocation phrase (or other invocation input) prior to "what time is it". SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0006] Techniques for providing dialog data from a primary automatic assistant to one or more secondary assistants based on a user's previous interactions with the one or more secondary automatic assistants are described herein. For example, various techniques are directed to having a primary automatic assistant receive a dialog between the user and the primary assistant, provide a first query that is part of the dialog to one or more secondary assistants, and then provide a query related to the first query to the same secondary assistant that the first query was provided to for processing, rather than providing the query to another secondary assistant. The user can first invoke the primary automatic assistant and issue one or more queries that the invoked primary automatic assistant can provide to the secondary assistants. The secondary assistants can process the queries and provide responses to the user via their own interfaces or via the primary automatic assistant. At some point in the dialog between the user and the primary automatic assistant, a query related to the first processed query can be received. The primary automatic assistant can identify that the second received query is related to the first query and provide the second query to the same secondary automatic assistant rather than another secondary assistant, thereby facilitating continuity in the communication between the user and the secondary automatic assistant and preventing related queries from being processed by different secondary automatic assistants.

[0007] In some implementations, the user may utter a call phrase, such as "OK Assistant", to invoke an auto-assistant (also referred to herein as the "first-invoked auto-assistant") without explicitly invoking any other auto-assistants that can at least selectively interact when processing a query received in relation to the invocation (e.g., immediately after, immediately before). Thus, rather than individually invoking one of the other auto-assistants based on providing a call input specific to the first-invoked auto-assistant, the user may specify to utilize the first-invoked auto-assistant. For example, a first call phrase (e.g., "OK Assistant A") can, when detected, exclusively invoke the first auto-assistant without invoking any other auto-assistants. Similarly, a second call phrase (e.g., "OK Assistant B") can, when detected, exclusively invoke the second auto-assistant without invoking the first-invoked auto-assistant and / or any other auto-assistants. The first-invoked assistant, when invoked, can at least selectively interact with other auto-assistants (i.e., "secondary assistants") when processing the input provided in relation to the invocation.

[0008] The user issues a call phrase such as an "OK Assistant" that invokes a primary automatic assistant and can perform one or more actions in another way without explicitly invoking other automatic assistants with which the primary automatic assistant can at least selectively interact when processing queries received in relation to (e.g., immediately after, immediately before) the call phrase. Thus, rather than individually invoking only one of the other automatic assistants based on providing a call input specific to the primary automatic assistant, the user can specify to utilize the primary automatic assistant. For example, when a first call phrase (e.g., "OK Assistant A") is detected, it can exclusively invoke the first automatic assistant without invoking the primary automatic assistant and / or any other automatic assistants. Similarly, when a second call phrase (e.g., "OK Assistant B") is detected, it can exclusively invoke the second automatic assistant without invoking the primary automatic assistant and / or any other automatic assistants. Other call phrases (e.g., "OK Assistant") can, when detected, invoke the primary assistant. The primary assistant, when invoked, can at least selectively interact with the first automatic assistant and / or the second automatic assistant (i.e., the "secondary assistant") when processing the input provided in relation to the call. In some implementations, the primary assistant can always interact with one or both of the first and second automatic assistants and can be a "meta-assistant" or a "general automatic assistant" that lacks one or more automatic assistant capabilities such as speech recognition, natural language understanding, and / or fulfillment capabilities itself. In other examples, the primary automatic assistant can also interact with the secondary assistant while performing its own query processing to determine a response to the query.

[0009] A primary automated assistant can provide a query to only one of the secondary assistants based on the user's historical interactions with the primary and / or secondary assistants. For example, the primary automated assistant can receive a query and provide the query to a first automated assistant. The response from the first automated assistant can be provided directly from the first automated assistant or via the primary automated assistant to the user. Subsequently, the primary automated assistant can receive a query related to the first query. The primary automated assistant can determine that the new query is related to the first query and can provide the second query to only the first automated assistant and not to the second automated assistant. Thus, if the intervening query is processed by the first automated assistant, the user can be provided with a consistent response to related queries instead of having related queries processed by different secondary automated assistants.

[0010] Accordingly, the use of the techniques described herein reduces the need for a user to explicitly invoke each of a plurality of automated assistants based on information the user is interested in being provided by a particular automated assistant. For example, the implementations described herein reduce the need for a user to explicitly invoke a first automated assistant to process a particular type of query and a second automated assistant to process other types of queries. Further, by identifying portions of a dialog previously performed by the first automated assistant and providing additional relevant queries for that automated assistant to process, consistent results can be provided without the need to invoke multiple automated assistants to process related queries. By utilizing a primary assistant to perform an initial analysis of the received query, preprocessing of the query (e.g., natural language understanding, text-to-speech processing, automatic speech recognition) need only be performed once for a given query rather than being processed individually by each secondary assistant. Further, the techniques described herein reduce the need for a user to extend the overall duration of a dialog via follow-up queries to obtain a response from an automated assistant that the user is interested in processing the query. Accordingly, these implementations attempt to determine which automated assistant will process related queries and provide an objective improvement even if the related queries are not consecutive within the dialog, whereas otherwise, inconsistent or incomplete results may be provided and additional processing may be required.

[0011] A primary automatic assistant can be invoked by an invocation phrase that invokes the primary automatic assistant and does not invoke other automatic assistants in the same environment. For example, the primary automatic assistant is invoked with the phrase "OK Assistant", and one or more secondary assistants are invoked with one or more other invocation phrases (e.g., "OK Assistant A", "Hey Assistant B"). When invoked, the primary assistant processes the queries associated with the invocation (e.g., queries preceding and / or following the invocation phrase). In some implementations, the primary assistant can identify multiple queries associated with the invocation. For example, the primary assistant can continue to receive audio data for a certain period of time after the invocation phrase and process any additional utterances of the user during that period. Thus, the user may not need to invoke the primary automatic assistant for each query, and instead, can utter multiple queries, and for each of the multiple queries, a response from the primary automatic assistant and / or one or more other secondary assistants may follow.

[0012] In some implementations, the received spoken utterance can be processed to determine an intent, terms within the query, and / or a classification of the query. For example, a primary automated assistant can perform speech recognition, natural language processing, and / or one or more other techniques to convert the audio data into data that can be further processed to determine the user's intent. In some implementations, one or more secondary assistants can further process the incoming spoken utterance even if the processing secondary assistant is not invoked by the user. For example, the user can invoke the primary automated assistant by uttering a phrase such as "OK Assistant" followed by a query. The primary automated assistant can activate the microphone and capture the utterance, along with the user and / or one or more other secondary assistants within the same environment as the primary automated assistant. In some implementations, only the primary automated assistant is invoked and can provide an instruction of the query (e.g., audio data, text representation of the audio, user intent) to one or more of the secondary automated assistants.

[0013] In some implementations, the primary automated assistant can continue to execute on a client device along with one or more secondary assistants. For example, the device may be executing a primary automated assistant, as well as one or more other automated assistants that are invoked using a call phrase different from the primary automated assistant. In some implementations, one or more secondary automated assistants can continue to execute on a device different from the device executing the primary automated assistant, and the primary automated assistant can communicate with the secondary automated device through one or more communication channels. For example, the primary automated assistant can provide a text representation of the query via Wi-Fi, an ultrasonic signal perceptible by the secondary automated assistant via a speaker, and / or one or more other communication channels.

[0014] In some implementations, the primary automatic assistant can classify queries based on the intent of the query. Classification can indicate to the primary automatic assistant which type of response can satisfy the query. For example, for a query of "Set an alarm", the primary automatic assistant can perform natural language understanding on the query to determine the intent of the query (i.e., the user wants to set an alarm). Query classification can include, for example, a request to answer a question, a request to execute an application, and / or an intent to perform some other task that one or more of the automatic assistants can perform.

[0015] When a query is received by the primary automatic assistant, the primary automatic assistant can determine which secondary automatic assistant can best satisfy the requested intent. For example, the primary automatic assistant can process the query, determine that the intent of the query is to set an alarm, and provide the intent to one or more secondary automatic assistants that are communicating with the primary automatic assistant. Also, for example, the primary automatic assistant can be further configured to process some or all of the incoming queries. For example, the primary automatic assistant can receive a query, process the query, determine whether it has the ability to fulfill the query, and execute one or more tasks in response to the query.

[0016] In some implementations, one or more of the secondary automatic assistants can notify the primary automatic assistant of the tasks they can perform. For example, the secondary assistant can provide a signal to the primary automatic assistant indicating that it can access a music application, set an alarm, and respond to answer questions. Thus, the primary automatic assistant can check which of the set of secondary automatic assistants can perform the task before providing the query and / or an indication of the intent of the query to the secondary automatic assistant.

[0017] When a primary automated assistant receives a query and processes the query to determine the intent of the query, the primary automated assistant may access one or more databases to determine whether the user prefers a particular secondary automated assistant with respect to that classification of the query. For example, in previous interactions with the automated assistant, the user may utilize Assistant A to set an alarm, but may prefer Assistant B if the user is interested in receiving an answer to a question. The user's preferences may be determined based on explicit instructions by the user (e.g., setting a preference for Assistant A to set an alarm) and / or based on previous usage situations (e.g., the user calls Assistant B more frequently than Assistant A when asking questions).

[0018] The primary automated assistant can then provide the query to a secondary automated assistant for further processing. For example, the user may utter a query, and the primary automated assistant can determine the intent from the query and then provide the intent of the query to a secondary automated assistant that can process it (and that has been indicated to the primary automated assistant as being able to do so). The secondary automated assistant can then perform additional processing of the query and execute one or more tasks as needed. Tasks can include, for example, determining an answer to a question, setting an alarm, playing music via a music application, and utilizing information in the query to execute one or more other applications. In some implementations, the secondary automated assistant can provide a response (or other fulfillment) to the primary automated assistant regarding completion of the task. For example, the user may utter a query that includes a question, and the query can be provided to the secondary automated assistant. The secondary automated assistant can process the query, generate a response, provide the response to the primary automated assistant, and the primary automated assistant can provide an audible response to the user for the question.

[0019] In some cases, the user may utter a second query related to a previous query. For example, the user may utter a query such as "How tall is the president" that Assistant A can fulfill. Subsequently, immediately after an answer is provided to the query, or as a subsequent query with intervening unrelated queries, the user may utter "How old is he" related to the president mentioned in the previous query. One or more of the primary automatic assistant and / or secondary assistants may identify the intent of the first query (e.g., the interest in having the president's height provided), and determine an answer to the follow-up query by determining that the second query is an additional question about the president mentioned in the first query. The dialog between the primary automatic assistant and the user may continue with additional queries and responses that are either related or unrelated to each other.

[0020] If the user utters a query related to a previous query, the user may be interested in having the same secondary assistant that fulfilled the first request fulfill the subsequent request. However, in some cases, the subsequent query and the previous query are part of a dialog that includes intervening queries that are not related to the previous query and the subsequent query. For example, the user may call the primary automatic assistant and utter "How tall is the president", and then utter "Set an alarm for 9am" (unrelated to the first query), and subsequently utter "How old is he" related to "How tall is the president". The user may be interested in having both the first query and the third query fulfilled by the same secondary automatic assistant, and the intervening query is processed by a different automatic assistant (depending on the intent of the query) or otherwise processed separately from the related queries.

[0021] A primary automated assistant can receive an utterance from a user that includes a query and identify a secondary automated assistant that can process the query. For example, for a query such as "How tall is the president", the primary automated assistant can determine that the first automated assistant is configured to process "question / answer" type queries and provide the query to the first secondary automated assistant. The query can be provided to the first automated assistant via one or more network protocols (e.g., Wi-Fi, Bluetooth, LAN) and / or one or more other channels. For example, the primary automated assistant can generate an ultrasonic signal that can be captured by a microphone of the device on which the first automated assistant is running and is inaudible to humans. This signal can be, for example, a processed version of the audio data that includes the query (e.g., raw audio data processed to a higher frequency and / or a text representation of the audio data that is transmitted via a speaker and is inaudible to humans). Also, for example, the primary automated assistant and the secondary automated assistant that is receiving the query can continue to run on the same device and the query can be provided directly to the secondary assistant. In some implementations, the query can be provided to the secondary automated assistant as audio data, as a text representation of the query, and / or as a text representation of a portion of the information included in the query, as the intent of the query, and / or as other automatic speech recognition output.

[0022] The selected secondary automated assistant can provide a response to the user, such as an audible response, and the primary automated assistant can store the conversation between the user and the secondary automated assistant. For example, the primary automated assistant can store the query and response from the secondary automated assistant, the type of response provided by the secondary automated assistant (e.g., question and answer), and an indication of the secondary automated assistant that provided the response.

[0023] Following the initial query, the primary automated assistant can receive additional queries from users not related to the first query. These queries can be provided to the same secondary assistant that provided the first query, and / or one or more other secondary automated assistants (or, if so configured, the primary automated assistant). Following intervening queries, the user may utter a query related to the first query (e.g., a query of the same type, similar intent, one or more similar keywords within the query that reference the previous query), and based on the stored user-secondary automated assistant dialog, determine that the user has previously interacted with a secondary automated assistant that responded to a similar query. Based on the determination that the user has previously interacted with a particular secondary automated assistant, the primary automated assistant can provide subsequent queries to the same secondary automated assistant.

[0024] In some implementations, providing relevant queries to a secondary automatic assistant may be based on the time since the user interacted with the secondary automatic assistant. For example, a primary automatic assistant can remember the time of the interaction along with the instructions of the dialog between the user and the secondary automatic assistant. Then, when the user utters a query related to a previous query, the primary automatic assistant can determine the amount of time elapsed since the last interaction between the user and the secondary automatic assistant. If the primary automatic assistant determines that the last interaction was within a threshold time before the current query, the primary automatic assistant can provide the new query to the same secondary automatic assistant for further processing. Conversely, if the user utters a query related to a previous query but it is determined that the time elapsed from the historical dialog between the user and the secondary automatic assistant exceeds the threshold time, the primary automatic assistant may determine that the same automatic assistant should not be called to process the query even if it is related to the previous query. However, regardless of whether the threshold time has elapsed, if the primary automatic assistant determines that the secondary automatic assistant can process the new query, it can still provide the query to the secondary automatic assistant.

[0025] In some implementations, regardless of whether previous related queries were processed by different automatic assistants, the user may decide that a particular type of query should be processed by different automatic assistants based on the user's preference for one of the secondary assistants. For example, in some cases, the primary assistant may identify the secondary automatic assistant that the user previously used for a particular type of query and / or a particular context (e.g., the user's location, time, subject of the query). In those cases, even if the primary automatic assistant could provide the query to a different automatic assistant based on other factors, the primary automatic assistant may provide the query to the preferred secondary automatic assistant.

[0026] As an example, the primary automatic assistant may identify that both secondary automatic assistants "Assistant A" and "Assistant B" can fulfill the query "Play music". In past conversations, the user may be able to instruct that "Assistant A" should play music using a music application it can access. The previous query "What is the current number 1 song" may have been provided to "Assistant B". Thus, the query may be provided to either "Assistant B" that processed the relevant query or "Assistant A" that the user has designated as the preferred automatic assistant for playing music. The primary automatic assistant may provide the query to "Assistant A" using the context of the previous response from "Assistant B" based on the user's preference, even though an ongoing dialog with "Assistant B" may avoid switching between automatic assistants.

[0027] In some cases, the secondary automatic assistant may not be able to provide a response to the query. For example, the query may be provided to the secondary automatic assistant, and the secondary automatic assistant may not respond and / or provide an indication that the response is unavailable. In those cases, a different secondary automatic assistant may be provided with the query. For example, candidate secondary automatic assistants may be ranked, and the highest-ranked secondary assistant may be provided with the query first. If the response is not successful, the next highest-ranked secondary assistant may be provided with the query.

[0028] In some implementations, the ranking of candidate auto assistants can be based on the output from a machine learning model. For example, queries provided to one or more secondary assistants can be embedded within an embedding space, and candidate secondary assistants can be ranked based on the similarity between the embedding of previous queries and the current query. Other metrics that can be utilized to rank auto assistants can include, for example, the text and / or environmental context associated with the provided query, the time, the location, the application running on the client device when the query was provided, and / or other actions of the user and / or other actions of the computing device running the candidate secondary assistant at both the time the query was first provided and the current state.

[0029] When a secondary auto assistant provides a response, the query and response directives can be stored by the primary auto assistant for further processing of queries related to those queries. For example, the secondary assistant can be provided with a "reply" query that is a request for additional information. Then, the same secondary assistant can be provided with a next "reply" query even if it has provided intervening queries to different secondary auto assistants as described above. The directives can be stored along with context information and / or other information such as the time the query was provided, the location of the secondary auto assistant, the state of one or more applications running on the user's computing device, and / or other signals that can be utilized to determine which secondary auto assistant to provide subsequent queries to.

[0030] The above description is provided as an overview of some implementations of the present disclosure. Further description of those implementations and other implementations are described in more detail below.

Brief Description of the Drawings

[0031]

Figure 1

Figure 2

Figure 3A

Figure 3B

Figure 4A

Figure 4B

Figure 5A

Figure 5B

Figure 6

Figure 7

DETAILED DESCRIPTION OF THE INVENTION

[0032] Referring to FIG. 1, an exemplary environment is provided that includes a plurality of auto assistants that can be invoked by user 101. The environment includes a first stand-alone interactive speaker 105 having a microphone (not shown), and a second stand-alone interactive speaker 110 having a microphone (likewise not shown). The first speaker can at least partially execute a first auto assistant that can be invoked with an invocation phrase. The second speaker 110 may execute a second auto assistant that can be invoked with an invocation phrase that is either the same invocation phrase as the first auto assistant or one of a different set of phrases to enable the user to select which auto assistant to invoke based on the spoken phrase. In the exemplary environment, user 101 is speaking a verbal utterance 115 of "OK Assistant, set an alarm" near the first speaker 105 and the second speaker 110. If one of the first and / or second auto assistants is configured to be invoked with the phrase "OK Assistant", the invoked assistant may process the query following the invocation phrase (i.e., "set an alarm").

[0033] In some implementations, a device such as the first speaker 105 may run multiple virtual assistants. Referring to FIG. 2, an exemplary environment is shown that includes multiple client devices running multiple virtual assistants. The system includes a first client device 205 that runs a first virtual assistant 215 and a second virtual assistant 220. Each of the first and second virtual assistants can be invoked by uttering an invocation phrase (unique to each assistant or the same phrase that invokes both assistants) near the client device 205 such that audio is captured by a microphone 225 of the client device 205. For example, user 101 can invoke the first virtual assistant 215 by uttering "OK Assistant 1" near the client device 205, and can further invoke the second virtual assistant 220 by uttering the phrase "OK Assistant 2" near the client device 205. Based on which invocation phrase was uttered, the user can indicate which of the multiple assistants running on the first client device 205 the user is interested in processing a verbal query. The exemplary environment further includes a second client device 210 that runs a third virtual assistant 245. The third virtual assistant 245 can be configured to be invoked using a third invocation phrase such as "OK Assistant 3" such that it can be captured by a microphone 230. In some implementations, one or more of the virtual assistants in FIG. 2 may be absent. Further, the exemplary environment may include additional virtual assistants not present in FIG. 2. For example, the system may include a third device that runs an additional virtual assistant, and / or the client device 210 and / or the client device 205 may run additional virtual assistants and / or fewer virtual assistants than shown.Each of the automatic assistants 215, 220, and 225 can include one or more components of the automatic assistants described herein. For example, the automatic assistant 215 can include its own audio capture component for processing incoming queries, a visual capture component for processing incoming visual data, a hotword detection engine, and / or other components. In some implementations, automatic assistants that are running on the same device, such as the automatic assistants 215 and 220, can share one or more components that can be utilized by both automatic assistants. For example, the automatic assistant 215 and the automatic assistant 220 can share one or more of an on-device speech recognizer, an on-device NLU engine, and / or other components.

[0034] In some implementations, one or more of the virtual assistants can be invoked by a common invocation phrase, such as “OK Assistant,” that does not separately and individually invoke any of the other virtual assistants. When the user utters the common invocation phrase, one or more of the virtual assistants function as a primary virtual assistant and can coordinate responses among the other virtual assistants. Referring to FIG. 3A, a primary virtual assistant 305 is shown along with secondary virtual assistants 310 and 315. The primary virtual assistant can be invoked with the phrase “OK Assistant” or another common invocation phrase, which may indicate that the user is interested in providing queries to multiple virtual assistants. Similarly, secondary virtual assistants 310 and 315 can each have one or more alternative invocation phrases that, when spoken by the user, invoke the corresponding virtual assistant. For example, secondary virtual assistant 310 can be invoked with the invocation phrase “OK Assistant A,” and secondary virtual assistant 315 can be invoked with the invocation phrase “OK Assistant B.” In some implementations, the primary virtual assistant 305, and one or more of the secondary virtual assistants 310 and 315, may be running on the same client device, as shown in FIG. 2 with respect to first virtual assistant 215 and second virtual assistant 220. In some implementations, one or more of the secondary virtual assistants may be running on a separate device, as shown in FIG. 2 with respect to third virtual assistant 245.

[0035] The environment shown in FIG. 3A further includes a database 320. The database 320 can be utilized to store the conversations between a user and one or more automated assistants. In some implementations, the primary automated assistant 305 can receive queries from the user, such as a query spoken after a call phrase that invokes the primary automated assistant 305. For example, the user can say, "OK Assistant, how tall is the president?" The phrase "OK Assistant" can be a call phrase specific to the primary automated assistant 305, and once invoked, the primary automated assistant 305 can determine the intent of the query that follows the call phrase (i.e., "how tall is the president"). In some implementations, the primary automated assistant 305 can provide the query to one or more of the secondary automated assistants 310 and 315 for further processing. Additionally, the primary automated assistant 305 can store the query instructions in the database 320 along with the assistant to which the query was provided for future access by the primary automated assistant 305.

[0036] In some implementations, the primary automated assistant 305 can always interact with one or both of the secondary automated assistants 310 and 315 and can be a "meta-assistant" that itself can lack one or more automated assistant capabilities such as speech recognition, natural language understanding, and / or fulfillment capabilities. In other examples, the primary automated assistant can also interact with the secondary assistant while performing its own query processing to determine a response to the query. For example, as further described herein, the primary automated assistant 305 can include a query processing engine, or the primary automated assistant 305 may not be configured to process the query and provide a response to the user.

[0037] Other components of the automatic assistants 305, 310, and 315 are optional and can include, for example, a local speech-to-text (“STT”) engine (which converts captured audio to text), a local text-to-speech (“TTS”) engine (which converts text to speech), a local natural language processor (which determines the semantic meaning of the audio and / or text converted from the audio), and / or other local components. Since client devices that execute the automatic assistants may be relatively constrained in terms of computing resources (e.g., processor cycles, memory, battery, etc.), the local components may have limited functionality compared to any counterparts included within any remotely-executing cloud-based automatic assistant components that cooperate with the automatic assistants.

[0038] Referring again to FIG. 2, in some implementations, one or more of the automatic assistants 210, 215, and 220 can be invoked by one or more gestures indicating that the user is interested in interacting with the primary automatic assistant. For example, the user can indicate an intention to invoke the automatic assistant by interacting with the device, such as pressing a button or a touch screen, can perform movements that can be captured by a visible and image capture device such as a camera, and / or can be seen by the device such that the image capture device can recognize the user's movements and / or positioning. When the user performs a gesture or an action, the automatic assistant is invoked and can begin capturing audio data following the gesture or action, as described above.

[0039] In some implementations, one automatic assistant can be selected as the primary assistant, and one or more other automatic assistants can be designated as secondary assistants. For example, the user can utter a call phrase common to a plurality of automatic assistants near the user. One or more components can determine which of the plurality of devices running the automatic assistants is closest to the user, and that closest automatic assistant can be designated as the primary automatic assistant, and the other automatic assistants can be designated as secondary assistants. Also, for example, when the user calls an automatic assistant, one or more components can determine which automatic assistant has been most frequently used by the user and can designate that automatic assistant as the primary automatic assistant.

[0040] In some implementations, the user can call a particular automatic assistant using a call phrase specific to that automatic assistant, and that automatic assistant can be designated as the primary automatic assistant. For example, the user can utter the call phrase "OK Assistant 1" to call the first assistant 215, and then the first assistant 215 is designated as the primary automatic assistant. Then, other automatic assistants such as the second automatic assistant 220 and the third automatic assistant 245 can be called by the primary automatic assistant, can be provided with queries by the primary automatic assistant, and / or can receive responses from other automatic assistants as described herein.

[0041] In some implementations, one or more automatic assistants, such as first automatic assistant 215 and second automatic assistant 220, may share one or more modules, such as a natural language processor and / or the results of a natural language TTS and / or STT processor. For example, referring again to FIG. 2, when client device 205 receives audio data, the audio data is once processed into text, and then the text can be provided to both automatic assistants 215 and 220, so that both the first automatic assistant 215 and the second automatic assistant 220 may share natural language processing. Also, for example, one or more components of client device 205 may process the audio data into text and provide a text representation of the audio data to a third automatic assistant 245, as further described below. In some implementations, the audio data is not processed into text and instead may be provided as raw audio data to one or more of the automatic assistants.

[0042] Referring to FIG. 3B, components of a primary automatic assistant 305 and a secondary automatic assistant 310 are shown. As described above, the primary automatic assistant 305 and / or one or more secondary automatic assistants 310 may include additional components other than those shown in FIG. 3B. Additionally or alternatively, one or more components may be absent from the primary automatic assistant 305 and / or one or more secondary automatic assistants 310. For example, the primary automatic assistant 305 may include a query processing engine 340 and may execute as both a primary and a secondary automatic assistant.

[0043] The calling engine 335 is operable to detect one or more verbal calling phrases and, in response to the detection of one of the verbal calling phrases, call the primary automatic assistant 305. For example, the calling engine 335 can call the primary automatic assistant 305 in response to the detection of verbal calling phrases such as "Hey, Assistant", "OK Assistant", and / or "Assistant". The calling engine 335 can continuously process a stream of audio data frames (e.g., when not in the "inactive" mode) based on the output from one or more microphones of the client device that executes the primary automatic assistant 305 to monitor for the occurrence of verbal calling phrases. While monitoring for the occurrence of verbal calling phrases, the calling engine 335 discards any audio data frames that do not contain a verbal calling phrase (e.g., after temporarily storing them in a buffer). However, when the calling engine 335 detects the occurrence of a verbal calling phrase in the processed audio data frame, the calling engine 335 can call the primary automatic assistant 305.

[0044] Referring to FIGS. 4A and 4B, a flowchart illustrating a method of providing a query to a secondary automatic assistant instead of providing a query to another automatic assistant is shown. For convenience, the operation of the method will be described with reference to a system that performs operations such as the primary automatic assistant 305 and the secondary automatic assistants 310 and 315 shown in FIGS. 3A and 3B. The system of the illustrated method includes one or more processors and / or other components of a client device. Further, the operations of the illustrated method are shown in a particular order, but this is not meant to be limiting. One or more operations can be rearranged, omitted, or added.

[0045] At 405, the primary automatic assistant is invoked. The primary automatic assistant can share one or more features with the primary automatic assistant 305. For example, invoking the primary automatic assistant can include the user uttering a call phrase specific to the primary automatic assistant (e.g., "OK Assistant"), performing a gesture detected by one or more cameras of the client device running the primary automatic assistant 305, pressing a button on the client device running the primary automatic assistant, and / or one or more other actions by the user indicating that the user is interested in one or more subsequent (or previous) utterances being processed by the primary automatic assistant. In some implementations, the primary automatic assistant is a "meta-assistant" that performs only limited operations, such as determining which of a plurality of secondary automatic assistants to provide a query to. For example, the primary automatic assistant can include one or more of the components of the primary automatic assistant 305 shown in FIG. 3B (e.g., the query classifier 330, the assistant determination module 325) without including a query processing engine for processing queries and providing responses. In some implementations, the primary automatic assistant 305 also performs the operations of the secondary automatic assistant and can include a query processing engine such as the query processing engine 340 of the secondary automatic assistant 310. Verbal queries can be provided by the user using a call phrase. For example, referring to FIG. 5A, a dialog between the user 101 and a plurality of automatic assistants is shown. To initiate the dialog, the user uttered a verbal utterance 505 including the call phrase "OK Assistant" and the query "how tall is the president".

[0046] At 410, the primary automatic assistant 305 determines the intent of a query provided verbally by a user. In some implementations, the query classifier 330 of the primary automatic assistant 305 can determine the classification of one or more queries provided to the primary automatic assistant 305 by a user. For example, after invocation, the primary automatic assistant 305 can receive an utterance containing a query and then perform NLU on the audio data to determine the intent of the query. Subsequently, the query classifier 330 can determine the classification of the query. As an example, the primary automatic assistant 305 can receive audio data including a user uttering a query such as "how tall is the president". After NLU processing, the query classifier 330 can determine that the query is a "response" query (i.e., a request from the user for which an answer to the query is provided). Also, for example, the primary automatic assistant 305 can receive audio data including a user uttering a query such as "set an alarm". After NLU processing, the query classifier 330 can determine that the query is a "productivity" query (i.e., a request from the user to perform a task). Other examples of query classification can include requesting a device such as a smart light and / or home appliance to perform a task, requesting a third-party application to perform a task (e.g., sending an email and / or text message), and / or other classifications of queries that can be performed by the automatic assistant. Further, the categories of query classification can include additional information related to the query such as the topic of the query, additional information related to the query (e.g., the current state of the device while the query is provided), and / or other fine-grained classifications that can be utilized to identify past queries related to the current query.

[0047] Once the query is classified, the assistant determination module 325 can determine which automated assistant to provide the query to for further processing. As described above, the database 320 can include indications of past interactions of the user with the secondary automated assistants 310 and 315. Each indication can include an identifier of the query classification and the automated assistant that responded to the query. As an example, the query classifier 330 can determine that a query of "how tall is the president" is an "answer" query and provide it to the secondary automated assistant 310. Further, the primary automated assistant 305 can store the query and / or an indication of the query classification in the database 320 along with an indication that the query was provided to the secondary automated assistant 310.

[0048] In some implementations, the user can indicate a preference for a particular automated assistant to process queries of a particular classification. Referring again to FIG. 4A, at 415, the assistant determination module 325 can determine whether the user has indicated a preference for a particular secondary automated assistant to process queries of the classification determined at 410. If the user 101 has indicated a preference, at 420, the preferred secondary automated assistant can be provided the query. For example, the user can explicitly or based on past interactions with the secondary automated assistant indicate that all text messaging (e.g., queries classified as "messaging") is to be processed by the automated assistant 310. The indication of the user's preference can be stored in the database 320 and for subsequent queries classified as "messaging", the query can be provided to the automated assistant 310 instead of other secondary automated assistants.

[0049] If the user has not indicated a preference for a particular secondary automatic assistant, in decision block 425, the assistant decision module 325 determines, for the classified query, whether the secondary automatic assistant provided a query of the same classification as the current query within a threshold time period. In some implementations, the assistant decision module 325 can identify from the database 320 which automatic assistant previously processed a query of the same classification as the current query. For example, a query of "how many kids does he have" can be received by the primary automatic assistant 305 and classified as a "response" query by the query classifier 330. The assistant decision module 325 can determine, based on one or more indications within the database 320, that a previous "response" query was processed by the secondary automatic assistant 310 and / or that the user indicated a preference for the "response" query to be processed by the secondary automatic assistant 310. In response, the assistant decision module 325 can determine in 430 to provide the query "how many kids does he have" to the secondary automatic assistant 310 instead of providing the query to another automatic assistant.

[0050] In some implementations, one or more communication protocols may be utilized to facilitate communication between the automatic assistants. For example, referring to FIG. 3B, the secondary automatic assistant 310 includes an API 345 that may be utilized by the primary automatic assistant 305 to provide a query to the secondary automatic assistant 310. In some implementations, one or more other communication protocols may be utilized. For example, in an example where the primary automatic assistant 305 and the secondary automatic assistant 310 are running on separate client devices, the primary automatic assistant 305 may utilize the speaker of the client device on which the assistant is running to transmit a signal (e.g., an ultrasonic audio signal including audio data of an oral utterance and / or an encoding of natural language processing of a query) that can be detected by the microphone of the client device.

[0051] In some implementations, the secondary automatic assistant 310 can include a capabilities communicator 350 that can provide the primary automatic assistant 305 with an indication of the classification of queries that can be processed by the secondary automatic assistant 310. For example, the secondary automatic assistant 310 may be configured to process "answer" queries (i.e., may include a query processing engine 340 that can respond to queries including questions), but may not be configured to process "message" requests. Thus, the capabilities communicator 350 can provide an indication of the classification of queries that can be processed, and / or an indication of the classification of queries for which to opt out of responding. In some implementations, the secondary automatic assistant 310 may utilize the API 390 of the primary automatic assistant 305 and / or one or more other communication protocols (e.g., ultrasonic signals) to indicate its capabilities to the primary automatic assistant 305.

[0052] In block 445, the interaction between the user and the secondary automated assistant provided with the query is stored in database 320. In some implementations, the time can be stored along with the query (and / or the classification of the query) and the automated assistant that processed the query. Thus, in some implementations, when determining which secondary automated assistant to provide the query to, the assistant decision module 325 can determine whether one of the secondary automated assistants has processed a query of the same classification within a threshold time period. As an example, the user may submit a query of the "Answer" classification that is provided to secondary automated assistant 310 for processing. The indication of the query (or its classification) can be stored by assistant decision module 325 along with a timestamp indicating the time the query was provided to secondary automated assistant 315. Subsequently, when another query (e.g., the next query, or a query after one or more intervening queries) is provided by the user and query classifier 330 determines that it is an "Answer" query, assistant decision module 325 can determine how much time has elapsed since secondary automated assistant 315 last processed an "Answer" query. If secondary automated assistant 315 processed an "Answer" query within the threshold time, subsequent "Answer" queries can be provided to secondary automated assistant 315.

[0053] In some implementations, the query processing engine 340 may provide at least a portion of the query and / or information included within the query to a third party application to determine a response. For example, a user may submit a query of "Make me a dinner reservation tonight at 6:30" that is provided to a secondary auto assistant for further processing. The secondary auto assistant is provided with audio data and / or a textual representation of the audio data, performs natural language understanding on the query, and may provide some portion of the information to a third party application such as a meal reservation application. In response, the third party application may determine a response (e.g., check whether a meal reservation is available) and provide the response to the secondary auto assistant, and the secondary auto assistant may generate a response based on the response from the third party application.

[0054] In some implementations, remembering the conversation can include storing the context of the query in database 320 along with the secondary automatic assistant that processed the query. Additionally or alternatively, a timestamp, and / or other information related to the query can be stored in database 320 along with the instructions for the conversation. For example, a query such as "how tall is the president" can be received by primary automatic assistant 305. The query can be provided to secondary automatic assistant 310, and the classification of the query, the timestamp when the query was provided to secondary automatic assistant 310, and / or the query (e.g., audio data of the user's verbal utterance, NLU data of the query, text representation of the query) can be stored in database 320. Further, the context of the query, such as an indication that the query is related to "the president", can be stored. Subsequently, the user can submit a query such as "how many kids does he have", and assistant decision module 325 can determine, based on the indication stored in database 320, that the user has submitted a query of the same classification (e.g., an "answer" query) within a threshold time period. Primary automatic assistant 305 can provide the query (i.e., "how many kids does he have") to secondary automatic assistant 310 along with the context (e.g., the previous query was related to "the president") for further processing. In some implementations, upon receiving the query "how many kids does he have", secondary automatic assistant 310 can store the indication of the context of the previous query separately so that secondary automatic assistant 310 can determine, based on the previous context, that "he" refers to "the president".

[0055] Referring again to FIG. 4, in decision block 425, in some implementations, the assistant decision module 325 may determine that the secondary automatic assistant has not processed a query of the same classification as the current query within a threshold time period. For example, during the current session and / or a query that is the first query of that classification received by the primary automatic assistant 305 within a certain time period may be submitted. If the assistant decision module 325 determines, based on instructions in the database 320, that the user has not indicated a preference for the automatic assistant to process queries of the current classification and / or that the secondary automatic assistant has not processed queries of the current classification, the assistant decision module 325 may select a secondary automatic assistant to process the query based on the capabilities of the available secondary automatic assistants. For example, the primary automatic assistant 305 can determine one or more candidate secondary automatic assistants that provide queries based on the current classification based on the capabilities of each secondary automatic assistant provided by the capability communicator 350 as described herein. For example, an "answer" query may not have been handled by either the secondary automatic assistant 310 or the secondary automatic assistant 315 within the threshold time period. The assistant decision module 325 may determine, based on the capability communicator of each of the available secondary automatic assistants, that the secondary automatic assistant 310 is configured to process "answer" queries and that the secondary automatic assistant 315 is not configured to process "answer" queries. Thus, instead of providing the query to the secondary automatic assistant 315 that is not configured to process such queries, the query may be provided to the secondary automatic assistant 310.

[0056] In some implementations, the secondary automatic assistant can be selected based on the user who uttered the query. For example, if the primary automatic assistant cannot determine that a relevant query was provided to the secondary automatic assistant within a threshold amount of time, the primary automatic assistant can decide to provide the query to the same secondary automatic assistant that previously received a query from the same user. In one example, a first user may provide a query of "Do I have anything on my calendar" to the primary automatic assistant, and this query is provided to the secondary automatic assistant for processing. Subsequently, another user may provide additional queries that are processed by a different secondary automatic assistant. Subsequently, the first user may submit a query of "How tall is the president", but after the first query was submitted and after exceeding the threshold amount of time. In that case, to improve the continuity of the conversation, the same secondary automatic assistant can be provided with the subsequent query based on identifying the speaker as the same speaker as the first query.

[0057] In decision block 430, the assistant decision module 325 can determine that the automatic assistant has not processed a query of the current classification within a threshold time period. In response, in block 440, the assistant decision module 325 can provide the query to a secondary automatic assistant that can process the query of the current classification. For example, the assistant decision module 325 can select an automatic assistant that has indicated that it can process a query of the current classification via the automatic assistant's capabilities communicator 350. When the query is provided to a capable secondary automatic assistant, in block 445, the conversation can be stored in the database 320 for later use by the assistant decision module 325.

[0058] The assistant decision module 325 may be configured such that a plurality of secondary automatic assistants are to process queries of a particular classification, but may determine that none of them have processed such a query within a threshold time period (or that a plurality of secondary automatic assistants are currently processing the current query). Additionally or alternatively, a plurality of secondary automatic assistants may have the potential to process queries of the same classification within a threshold time period. For example, the secondary automatic assistant 310 may have processed a query of "set an alarm" classified as a "productivity" query by the query classifier 330, and further, the secondary automatic assistant 315 may have processed a query of "turn off the bedroom lights" also classified as a "productivity" query by the query classifier 330. When a plurality of secondary automatic assistants are identified by the assistant decision module 325 to process a query, candidate secondary automatic assistants may be ranked and / or scored so that one secondary automatic assistant may be selected in place of the other secondary automatic assistants.

[0059] Referring to FIG. 5A, an interaction between a user 101 and two automatic assistants 305 and 310 is shown. In the illustrated environment, the primary automatic assistant 305 as well as the secondary automatic assistants 310 and 315 are each shown to be executed on separate devices. However, in some implementations, one or more of the automatic assistants may be executed on the same device, as shown in FIG. 2. Further, the illustrated dialog is received by the primary automatic assistant 305 that does not fulfill the query. Thus, the verbal utterance of the user 101 may be detected by the speaker of the primary automatic assistant 305, which provides the query to the secondary automatic assistants 310 and / or 315 for further processing. Thus, the user 110 can interact with the primary automatic assistant 305, but as shown, the response to the query may be provided by the secondary automatic assistants 310 and 315.

[0060] In dialog turn 505, user 101 makes an utterance: "OK Assistant, how tall is the president". As described herein, the invocation phrase "OK Assistant" can call the primary automatic assistant 305, and further, the primary automatic assistant 305 receives the verbal query "how tall is the president" via one or more microphones. The query classifier 330 of the primary automatic assistant 305 can determine the classification of the query, such as the classification of "answer" indicating that user 101 is requesting an answer to the query. Further, based on the instructions in database 320, the primary automatic assistant 305 can determine that user 101 has not previously submitted a query classified as an "answer" within the threshold time period in the current dialog session. In response, the assistant decision module 325 can provide the query to one of the candidate secondary automatic assistants, such as secondary automatic assistant 310. The decision to provide the query to one automatic assistant instead of providing it to other automatic assistants can be based on, for example, the user's preference for secondary automatic assistant 310 when providing "answer" queries identified based on past interactions between the user and the secondary automatic assistant, the time, the user's location, and / or other factors indicating that the user is interested in having secondary automatic assistant 310 process the query instead of secondary automatic assistant 315 processing the query. Further, the capabilities communicator 350 of the secondary automatic assistant 310 can provide an indication to the primary automatic assistant 305 that it can process the "answer" query, and the primary automatic assistant 305 can provide the query to the secondary automatic assistant 310 based on the indicated capabilities of the automatic assistant.

[0061] Referring to FIG. 5B, a diagram is provided that includes the dialog turn of FIG. 5A and the components for processing the query. As shown, the spoken utterance 505 is first provided to the query classifier 330, and the query classifier 330 determines the classification of "answer" at 505A. The classification of "answer" and / or other information related to the query (e.g., context) is provided to the assistant decision module 325, and the assistant decision module 325 checks the database 320 (at 505B) to determine whether the secondary automatic assistant has previously processed the "answer" query. Since this is the first query classified and processed as an "answer", the database 320 provides a "n / a" response 505C without including instructions for the secondary automatic assistant. In response, the assistant decision module 325 determines to provide the query to the secondary automatic assistant 310 (at 505D) instead of providing the query to the secondary automatic assistant 315 based on one or more factors (e.g., previous interactions between the user and the secondary automatic assistant, machine learning model output).

[0062] The primary automated assistant provides a query to the secondary automated assistant 310 for further processing. In response, the query processing engine 340 of the secondary automated assistant determines a response 510 to the query. For example, referring again to FIG. 5A, in response to a query of "how tall is the president", the secondary automated assistant 310 responds with "the president is 6 feet, 3 inches". In some implementations, the response may be provided by a speaker specific to the secondary automated assistant 310 (e.g., a speaker integrated with the device running the automated assistant). In some implementations, the response may be provided by a speaker of one or more other devices such as the primary automated assistant 305. In those cases, even if the speaker is utilized by more than one automated assistant in providing the response, a voice profile may be utilized to provide a response specific to the secondary automated assistant 310 so that the user can recognize, based on the audio characteristics of the response, that the response was generated by the secondary automated assistant 310 and not by another automated assistant. Further, referring again to FIG. 5B, when a query is provided to the secondary automated assistant 310 (and / or when the secondary automated assistant 310 provides a response 510), the assistant decision module 325 can identify the previous conversation and utilize the conversation to determine which automated assistant to provide a subsequent "answer" query to. To that end, the assistant decision module 325 can store in the database 320 a conversation 505E that indicates the query, the query classification, the timestamp when the query was provided, the context of the query, and / or other information related to the query.

[0063] In some implementations, the assistant determination module 325 may determine that a plurality of secondary automatic assistants have processed the query of the current classification. In those cases, at block 450, the assistant determination module 325 can rank the candidate secondary automatic assistants to determine which secondary automatic assistant to provide the query to. As described above, in some implementations, a query having the current classification may not have been processed previously (or not at all) within a threshold time period. Thus, any secondary automatic assistant configured to process the query can be a candidate automatic assistant. In some implementations, a plurality of secondary automatic assistants may be processing queries of the same classification as the current query, and the assistant determination module 325 can determine which of the plurality of assistants to provide the query to. Further, in some implementations, related queries of a different classification than the current query may be processed, and based on determining that a previous query is relevant, one of the secondary automatic assistants may be selected to process the query preferentially over other secondary automatic assistants.

[0064] In some implementations, to determine which secondary automatic assistant to provide the current query to, it is determined whether one or more of the previous queries are relevant to the current query, and thus, the current query and the previous queries can be embedded within an embedding space to rank the candidate secondary automatic assistants. For example, if a plurality of secondary automatic assistants have previously processed queries of the same classification as the current query, the previous queries can be embedded within the embedding space and the relevance between the current query and the previously processed queries. Thus, the secondary automatic assistant that previously processed the query most relevant to the current query can be selected as the secondary automatic assistant to process the current query (and / or the relevance of the previously processed queries can be a factor in determining which automatic assistant to provide the current query to).

[0065] Referring back to FIG. 5A, an oral utterance 515 is provided by user 101. The utterance "OK Assistant, what's on my calendar" is received by primary automatic assistant 305, classified, and provided to secondary automatic assistant 310 for further processing. In response, secondary automatic assistant 310 provides a response of "you have a meeting at 5". Referring back to FIG. 5B, oral utterance 515 is first provided to query classifier 305, and query classifier 305 determines that the query is a "reply" query 515A. The classification is provided to assistant decision module 325, and assistant decision module 325 determines whether user 101 has previously interacted with the automatic assistant that processed the "reply" query based on the instructions stored in database 320. After checking database 320 at 515B, a previous interaction between user 101 and secondary automatic assistant 305 is identified (515C). However, assistant decision module 325 can further determine that this query is not related to the previous "reply" query based on the context information identified from the previous "reply" query 505. In response, based on the determination that the context of oral utterance 515 is not related to the context of the previous "reply" query (i.e., utterance 515 is not a continuation of the dialog started by utterance 505), assistant decision module 325 may not select secondary automatic assistant 305 as the automatic assistant to process the query.The assistant decision module 325 can determine, for example, based on the capabilities indicated by each of the automatic assistants, that, for example, the secondary automatic assistant 310 is configured to access the user's calendar, identify that the user has previously used it to access the user's calendar, and / or provide a query to the secondary automatic assistant 310 based on the indicated user preferences of the secondary automatic assistant 310 that prioritize the secondary automatic assistant 310 over the secondary automatic assistant 305 when processing calendar-related queries. When the query is provided to the secondary automatic assistant 310, the conversation is stored in the database 320 (515E), and the secondary automatic assistant 310 can provide a response. Similar to the previous conversations stored in the database 320, the conversation with the secondary automatic assistant 310 can be stored along with context information, a timestamp, vector embeddings of the query and / or response provided by the secondary automatic assistant, and / or other information related to the query and / or response for use in determining whether to provide subsequent queries to the secondary automatic assistant. In some implementations, for subsequent queries, the previous conversations can be stored for a period of time so that only the conversations that occurred within a threshold time period are reexamined to determine whether the query is related to a previous query.

[0066] Referring again to FIG. 5A, the user 101 provides the verbal utterance 525 of "how many kids does he have". This verbal utterance is a continuation of the dialog started at utterance 505, and at least a portion of the context of the previous dialog is required to generate a response. For example, the verbal utterance 525 includes the term "he" which refers to "the president" as indicated by the user 101 in the utterance 525. In response to this, the secondary automatic assistant 310 responds with "the president has two children".

[0067] Referring again to FIG. 5B, the spoken utterance 525 is provided to the query classifier 330, which classifies the query as an "answer" query at 525A. Similar to previous queries, the assistant decision module 325 checks within the database 320 (at 525B) to identify (at 525C) that both the secondary automatic assistant 310 and the secondary automatic assistant 315 have processed the "answer" query within a threshold time period. However, based on the content of the current query (e.g., having the term "he" without indicating what the term refers to) and based on context information that may be stored with the dialogue 505E, the assistant decision module 325 can decide to provide the query to the secondary automatic assistant 310 (at 525D) instead of providing the query to the secondary automatic assistant 315 that has processed an "answer" query not relevant to the current query. Further, the dialogue with the secondary automatic assistant 310 is stored within the database 320 (at 525E) along with query information, context, and / or other information as described herein.

[0068] Referring back to FIG. 5A, the user then provides the verbal utterance 535 of "OK, Assistant, send Bob a message of 'Meeting is at 5'". In response, the secondary automatic assistant 310 executes the request and responds to the user 101 with "I sent Bob the message via Message Application". In some implementations, the primary automatic assistant may determine, as described above, that the verbal utterance is not related to other previous queries, and thus, based on the capabilities provided by each automatic assistant, user preferences, and / or other factors, etc., can determine the automatic assistant that provides the query. In some implementations, the primary automatic assistant 305 may determine that the utterance 535 is related to a meeting (based on the terms in the query) and that the user 101 has previously interacted with the secondary automatic assistant 315 regarding meetings within the calendar application. Thus, for continuity, the primary automatic assistant 305 may provide the query to the secondary automatic assistant 315 preferentially over the secondary automatic assistant 310 based on determining that the context of the query is more related to previous queries processed by the secondary automatic assistant 315 than those previously processed by the secondary automatic assistant 310.

[0069] Referring again to FIG. 5B, the spoken utterance is first provided to the query classifier 330, which determines (535A) that the query is a "messaging" query (i.e., a request to send a message). In response, the assistant decision module 325 checks the database 320 (535B) and does not identify a previous interaction between the user 101 and the secondary automatic assistant that processed the "messaging" query, but identifies (at 535C) a previous interaction between the user 101 and the secondary automatic assistant 315 regarding the calendar (i.e., the stored interaction 515E). Thus, the assistant decision module 325 determines that the previous interaction with the secondary automatic assistant 315 is relevant to the current query and can provide the query to the secondary automatic assistant 315 instead of providing the query to the secondary automatic assistant 310 (535D). For example, the assistant decision module 325 can utilize one or more machine learning models to determine that the utterance 535 is most similar to the response (e.g., the utterance 535 and the response 520 include a reference to "meeting"), and embed the utterances 505, 515, 525, and 535, and / or the responses 510, 520, and 530 into the embedding space to provide the utterance 535 to the same secondary automatic assistant 315 that processed the utterance 515 to provide the response 520. In response, the query processing engine of the secondary automatic assistant 315 can determine a response, execute the requested action (i.e., send a message to Bob), and provide the response 540 to the user 101. The subsequent interaction 535E with the secondary automatic assistant can then be stored in the database 320 for subsequent use in determining when to provide the "messaging" query to the secondary automatic assistant.

[0070] In some implementations, the machine learning model may receive one or more other signals as input and, if no previous query of the same classification has been provided to the secondary automatic assistant within a threshold time period, provide as output a ranking of candidate secondary automatic assistants that may provide the query. For example, the current context, and the corresponding historical context when previous queries were provided to the secondary automatic assistant, may be utilized as input to the machine learning model to determine which automatic assistant to provide the query to. The context of the query can include, for example, an image and / or interface currently presented to the user via one or more computing devices, the current application accessed by the user when providing the query, the time at which the user provided the query, and / or one or more other user or device characteristics that can indicate the user's intention to provide the query to a certain secondary automatic assistant instead of providing the query to other secondary automatic assistants.

[0071] Referring again to FIG. 4B, the query may be provided to the highest-ranked candidate secondary automatic assistant at block 455. At decision block 460, if the selected secondary automatic assistant processes the query successfully, the interaction with the secondary automatic assistant is stored in database 320 (at block 465). If the selected automatic assistant does not process the query successfully (e.g., cannot generate a response, cannot perform the intended action), at block 470, the query may be provided to the next highest-ranked candidate automatic assistant. This can continue until providing the query to all candidate automatic assistants fails and / or until a candidate secondary automatic assistant processes the query successfully.

[0072] In some implementations, when a query determined to be related to previous queries provided to different secondary assistants is provided to a secondary assistant for the first time, the context of the previous query can be provided along with the query. For example, for a first query of "How tall is the president", a text context including an indication that the query was related to "the president" is stored in database 320. Subsequently, when a query of "how many kids does he have" is received and provided to a secondary assistant different from the first query, the text context of "the president" can be provided to assist the secondary assistant in resolving "he" in the subsequent query. In some implementations, other context information such as time, the state of one or more computing devices, the applications currently being executed by the user, and / or previous queries and / or other contexts of the current query that can assist in resolving one or more terms in the provided query and that can assist the selected secondary assistant can be provided along with the subsequent query.

[0073] FIG. 6 shows a flowchart illustrating another exemplary method 600 for providing a query to a secondary assistant instead of providing the query to other assistants. For convenience, the operations of method 600 will be described with reference to the system that performs the operations. This system of method 600 includes one or more processors and / or other components of a client device. Further, although the operations of method 600 are shown in a particular order, this is not meant to be limiting. One or more operations can be reordered, omitted, or added.

[0074] In step 605, the call is received by the primary assistant. In some implementations, the call can be an utterance that transitions the primary automatic assistant from an inactive state (e.g., does not process incoming spoken utterances) to an active state. In some implementations, the call can include the user selecting a button via a user gesture, a client device and / or an interface of the client device, and / or one or more other actions indicating the user's interest in activating the primary automatic assistant.

[0075] In some implementations, the primary automatic assistant can share one or more features with the primary automatic assistant 305. For example, in some implementations, the primary automatic assistant can receive spoken utterances via one or more microphones and determine a secondary automatic assistant that processes the spoken utterances. In some implementations, the primary automatic assistant can include a query processing engine that enables the primary automatic assistant to generate responses to spoken queries and / or other utterances. In some implementations, the primary automatic assistant can perform processing of spoken utterances, such as ASR, TTS, NLU, and / or other processing to determine the content of the spoken utterance.

[0076] In some implementations, the invocation can indicate a particular secondary assistant to which audio data is provided. For example, the primary assistant can be invoked by a general invocation such as "OK Assistant", and further can be invoked by a secondary invocation indicating a particular assistant such as "OK Assistant A". In this case, the general assistant can provide audio (or provide an indication of audio data) specifically to "Assistant A" instead of providing the audio data to other secondary assistants. Additionally or alternatively, when the invocation indicates a particular assistant, the invocation can be used with one or more other signals to rank the assistant when determining which assistant to provide the audio data to.

[0077] In step 610, an oral query is received by the primary assistant. In some implementations, the oral query can include a request for the automated assistant to perform one or more actions. For example, the user can submit an oral query such as "How tall is the president" expecting a response that includes an answer to the presented question. Also, for example, the user can submit an oral query such as "turn off the lights" intending to transition one or more smart lights from on to off.

[0078] In step 615, the user's historical data is identified. The historical data can include information related to one or more secondary assistants and queries processed by each of the secondary assistants. For example, the historical data can include the text representation of the query, the content when the query was processed, the assistant that generated a response to the query, the generated response, and / or other indications of the historical interaction between the user and one or more assistants.

[0079] In step 620, the relationship between at least a portion of the history dialogue and the verbal query is determined. In some implementations, the relationship can be determined based on identifying previous dialogues between the user and one or more of the secondary automatic assistants and determining, based on the current query, which secondary automatic assistant processed the previous query that is most similar to the current query. For example, the user may have submitted a previous query of "How tall is the president". Subsequently, the user may submit a query of "how many kids does he have", where "he" refers to "the president". The relationship between the queries can be determined based on the context of the second query following the first query (either directly or having one or more intervening queries). Further, based on instructions stored in a database regarding the previous query, the same secondary automatic assistant that generated a response to the first query can be identified so that the second query can be provided to the same automatic assistant.

[0080] In step 625, audio data including the verbal query is provided to the secondary automatic assistant. In some implementations, the audio data can be provided as a transcription of the audio data. For example, a general automatic assistant can perform automatic speech recognition on the audio data and provide the transcription to the secondary automatic assistant. In some implementations, the audio data can be provided as audio (e.g., raw audio data and / or pre-processed audio data), and then the audio can be processed by a particular secondary automatic assistant.

[0081] FIG. 7 is a block diagram of an exemplary computing device 710 that can optionally be utilized to execute one or more aspects of the techniques described herein. Computing device 710 typically includes at least one processor 714 that communicates with several peripheral devices via a bus subsystem 712. These peripheral devices can include, for example, a storage subsystem 724 that includes a memory subsystem 725 and a file storage subsystem 726, a user interface output device 720, a user interface input device 722, and a network interface subsystem 716. The input and output devices enable a user to interact with computing device 710. Network interface subsystem 716 provides an interface to an external network and is coupled to a corresponding interface device within other computing devices.

[0082] User interface input device 722 can include a keyboard, a mouse, a trackball, a pointing device such as a touchpad or a graphics tablet, a scanner, a touch screen incorporated in a display, an audio input device such as a voice recognition system, a microphone, and / or other types of input devices. In general, the use of the term "input device" is intended to include all possible types of devices and methods for inputting information into computing device 710 or a communication network.

[0083] The user interface output device 720 may include a non-visual display such as a display subsystem, a printer, a fax machine, or an audio output device. The display subsystem may include a cathode ray tube (CRT), a flat panel device such as a liquid crystal display (LCD), a projection device, or any other mechanism for creating a visible image. The display subsystem may also provide a non-visual display via, for example, an audio output device. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 710 to a user or another machine or computing device.

[0084] The memory subsystem 724 stores programming structures and data structures that provide some or all of the functionality of some of the modules described herein. For example, the memory subsystem 724 may include logic for executing selected aspects of the methods of FIGS. 5 and 6 and / or for implementing the various components shown in FIGS. 2 and 3.

[0085] These software modules are generally executed by the processor 714 alone or in combination with other processors. The memory 725 used within the memory subsystem 724 can include several memories including a main random access memory (RAM) 730 for storing instructions and data during program execution and a read-only memory (ROM) 732 in which fixed instructions are stored. The file storage subsystem 726 can provide permanent storage for programs and data files and can include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. The modules implementing the functionality of a particular implementation may be stored by the file storage subsystem 726 within the memory subsystem 724 or within another machine accessible by the processor 714.

[0086] The bus subsystem 712 provides a mechanism for enabling various components and subsystems of the computing device 710 to communicate with each other as intended. Although the bus subsystem 712 is shown schematically as a single bus, alternative implementations of the bus subsystem may use multiple buses.

[0087] The computing device 710 can be of various types, including a workstation, server, computing cluster, blade server, server farm, or any other data processing system or computing device. Due to the constantly changing nature of computers and networks, the description of the computing device 710 shown in FIG. 7 is intended only as a specific example for the purpose of explaining some implementations. Many other configurations of the computing device 710 can have more or fewer components than the computing device shown in FIG. 7.

[0088] In some implementations, a method implemented by one or more processors is provided, including the steps of receiving a call by a general purpose assistant, where receiving the call causes the general purpose assistant to be invoked; receiving, via the invoked general purpose assistant, a verbal query captured within audio data generated by one or more microphones of a client device; identifying user historical dialogue data, where the historical dialogue data is generated based on one or more past queries of the user and one or more responses generated by a plurality of secondary assistants within a time period in response to the one or more past queries; determining that a relationship exists between the verbal query and a portion of the historical dialogue data related to a particular one of the plurality of secondary assistants based on comparing a transcription of the verbal query with the historical dialogue data; and in response to determining that a relationship exists between the verbal query and the particular assistant, providing an instruction of the audio data to the particular assistant instead of providing it to any other of the plurality of secondary assistants, where providing the instruction of the audio data causes the particular assistant to generate a response to the verbal query.

[0089] These and other implementations of the techniques disclosed herein can include one or more of the following features.

[0090] In some implementations, the method further includes determining, based on the historical conversation data, that there is an additional relationship between the verbal query and an additional portion of the historical conversation data associated with a second particular automatic assistant among the plurality of secondary automatic assistants; determining, based on the historical conversation data, a second assistant time for a second assistant historical conversation between the user and the second particular automatic assistant; determining, based on the historical conversation data, a first assistant time for a first assistant historical conversation between the user and a particular automatic assistant; ranking the particular automatic assistant and the second particular automatic assistant based on the first assistant time and the second assistant time; selecting a particular automatic assistant based on the ranking; and providing an instruction of the audio data to the particular automatic assistant instead of providing it to any other automatic assistant among the secondary automatic assistants, where the step of providing the instruction of the audio data to the particular automatic assistant is further based on the step of selecting a particular automatic assistant based on the ranking. In some of those implementations, the method further includes determining a first similarity score for the relationship and a second similarity score for the additional relationship, where the step of ranking the particular automatic assistant and the second particular automatic assistant is further based on the first similarity score and the second similarity score.

[0091] In some implementations, the method further includes determining a classification of the verbal query, and providing the transcription of the verbal query to a particular assistant is further based on the classification. In some of those implementations, the method further includes identifying a user's indicated preference to utilize a particular assistant for one or more of the past queries that also have the classification, based on historical conversation data, and providing the audio data to the particular assistant is further based on the user's preference. In other ones of those implementations, the method further includes generating an instruction for a general assistant to provide the audio data to a particular assistant, and storing the instruction along with the classification instruction. In some of those examples, the method further includes determining a current context of the audio data, and storing the current context along with an instruction for a general assistant to provide the audio data to a particular assistant.

[0092] In some implementations, the particular assistant performs automatic speech recognition on the instruction of the audio data to generate a text representation of the audio data, generates a natural language process of the text representation, and generates a response based on the natural language process. In some of those implementations, generating the response includes providing the result of the natural language process to a third party application, receiving a response from the third party application, and generating a response at least partially based on the response from the third party.

[0093] In some implementations, the call indicates a particular assistant, and the instruction of the audio data is an audio representation of the audio data.

[0094] In situations where the particular implementations discussed in this specification may collect or use personal information about a user (e.g., user data extracted from other electronic communications, information about the user's social network, the user's location, the user's time, the user's biometric information, as well as the user's activities and demographic information, relationships between users, etc.), the user is provided with one or more opportunities to control whether the information is collected, whether the personal information is stored, whether the personal information is used, and how information about the user is collected, stored, and used. That is, the systems and methods discussed in this specification collect, store, and / or use a user's personal information only when they have received explicit permission to do so from the relevant user.

[0095] For example, the user is provided with control over whether a program or function collects user information about that particular user or other users related to the program or function. Each user from whom personal information is collected is presented with one or more options that enable control over the collection of information related to that user in order to provide permission or approval regarding whether the information is collected and which portions of the information should be collected. For example, the user may be provided with one or more such control options via a communication network. Additionally, certain data may be processed in one or more ways before being stored or used such that information that can identify an individual is removed. As an example, a user's identifying information may be processed such that it cannot be used to determine information that can identify an individual. As another example, a user's geographical location may be generalized to a larger area such that the user's specific location cannot be identified.

[0096] Although several implementations are described and illustrated herein, various other means and / or structures may be utilized for performing the functions and / or for obtaining one or more of the results and / or advantages described herein, and each such variation and / or modification is to be regarded as being within the scope of the implementations described herein. More generally, all parameters, dimensions, materials, and configurations described herein are meant to be illustrative, and the actual parameters, dimensions, materials, and / or configurations will depend upon the particular application for which the teachings are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific implementations described herein. Accordingly, the foregoing implementations are presented by way of example only, and it is to be understood that within the scope of the appended claims and their equivalents, implementations may be practiced otherwise than as specifically described and claimed. Implementations of the present disclosure are directed to each of the individual features, systems, articles, materials, kits, and / or methods described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the present disclosure if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.

Explanation of Signs

[0097] 101 User 105 First stand-alone interactive speaker, first speaker 110 Second stand-alone interactive speaker, second speaker 115 Spoken utterance 205 First client device, client device 210 Second client device 215 First automatic assistant, first assistant 220 Second automatic assistant 225 Microphone 230 Microphone 245 Third automatic assistant 305 Primary Automatic Assistant 310 Secondary Automatic Assistant 315 Secondary Automatic Assistant 320 Database 325 Assistant Decision Module 330 Query Classifier 335 Call Engine 340 Query Processing Engine 345 API 350 Capability Communicator 390 API 505 Spoken Utterance, Utterance 510 Response 515 Spoken Utterance, Utterance 515E Stored Dialogue 520 Response 525 Spoken Utterance, Utterance 530 Response 535 Spoken Utterance, Utterance 535E Dialogue 540 Response 710 Computing Device 712 Bus Subsystem 714 Processor 716 Network Interface Subsystem 720 User Interface Output Device 722 User Interface Input Device 724 Memory Subsystem 725 Memory Subsystem 726 File Memory Subsystem 730 Main Random Access Memory (RAM) 732 Read-Only Memory (ROM)

Claims

1. 1. A method implemented by one or more processors of a client device, the method comprising: receiving a call by a generic automated assistant, where receiving the call causes the generic automated assistant to be invoked; receiving, via the invoked generic automated assistant, a verbal query captured in audio data generated by one or more microphones of the client device; identifying historical interaction data for a user, the historical interaction data being generated based on one or more past queries of the user and one or more responses generated by a plurality of secondary automated assistants within a period of time in response to the one or more past queries; determining, based on comparing a transcription of the spoken query to the historical interaction data, that a relationship exists between the spoken query and a portion of the historical interaction data associated with a particular automated assistant among the plurality of secondary automated assistants; In response to determining that the relationship exists between the spoken query and the particular automated assistant, providing an indication of the audio data to the particular automated assistant in lieu of providing the indication to any other automated assistants of the secondary automated assistants, where providing the indication of the audio data causes the particular automated assistant to generate a response to the verbal query; Including, method.

2. Determining, based on the historical interaction data, that an additional relationship exists between the verbal query and an additional portion of the historical interaction data associated with a second particular automated assistant of the plurality of secondary automated assistants; Determining a second assistant time for a second assistant historical interaction between the user and the second particular automated assistant based on the historical interaction data; Determining a first assistant time for a first assistant historical interaction between the user and the particular automated assistant based on the historical interaction data; ranking the particular automated assistant and the second particular automated assistant based on the first assistant time and the second assistant time; selecting the particular automated assistant based on the ranking; Further comprising: providing the indication of the audio data to the particular automated assistant instead of providing the indication to any other automated assistant of the secondary automated assistants is further based on selecting the particular automated assistant based on the ranking; The method of claim 1.

3. The method further includes determining a first similarity score for the relationship and a second similarity score for the additional relationship, and the step of ranking the specific automated assistant and the second specific automated assistant is further based on the first similarity score and the second similarity score. The method of claim 2.

4. determining a classification of the spoken query, and providing the transcription of the spoken query to the specific automated assistant is further based on the classification; 4. The method according to any one of claims 1 to 3.

5. Further comprising: identifying an indicated preference of the user to utilize the particular automated assistant for one or more of the past queries that also have the classification based on the historical interaction data; and providing the audio data to the particular automated assistant further based on the user's preference. The method of claim 4.

6. generating instructions for the general automated assistant to provide the audio data to the specific automated assistant; storing said indication together with an indication of said classification; 5. The method of claim 4, further comprising:

7. determining a current context of the audio data; storing the current context together with the instructions of the general automated assistant to provide the audio data to the specific automated assistant; 7. The method of claim 6, further comprising:

8. 8. The method of claim 1, wherein the particular automated assistant performs automatic speech recognition on the indication of the audio data to generate a textual representation of the audio data, generates natural language processing of the textual representation, and generates a response based on the natural language processing.

9. It is possible to generate a response, providing results of the natural language processing to a third party application; and receiving a response from the third party application; and generating the response based at least in part on the response from the third party.

10. 10. The method of claim 1, wherein the invocation indicates the particular automated assistant and the indication of the audio data is an audio representation of the audio data.

11. 1. A system comprising one or more processors and a memory storing instructions, the instructions causing the one or more processors to, in response to execution of the instructions by the one or more processors: an act of receiving a call by a generic automated assistant, the act of receiving the call causing the generic automated assistant to be invoked; receiving, via the invoked generic automated assistant, a verbal query captured within audio data generated by one or more microphones of a client device; an operation of identifying historical interaction data for a user, the historical interaction data being generated based on one or more past queries of the user and one or more responses generated by a plurality of secondary automated assistants within a period of time in response to the one or more past queries; determining, based on comparing a transcription of the spoken query to the historical interaction data, that a relationship exists between the spoken query and a portion of the historical interaction data associated with a particular automated assistant among the plurality of secondary automated assistants; In response to determining that the relationship exists between the spoken query and the particular automated assistant, providing an indication of the audio data to the particular automated assistant in lieu of providing the indication to any other automated assistant of the secondary automated assistants, the act of providing the indication of the audio data causing the particular automated assistant to generate a response to the verbal query; A system that executes the above.

12. The instructions cause the one or more processors to: determining, based on the historical interaction data, that an additional relationship exists between the verbal query and an additional portion of the historical interaction data associated with a second particular automated assistant of the plurality of secondary automated assistants; determining a second assistant time for a second assistant historical interaction between the user and the second particular automated assistant based on the historical interaction data; determining a first assistant time for a first assistant historical interaction between the user and the particular automated assistant based on the historical interaction data; ranking the particular automated assistant and the second particular automated assistant based on the first assistant time and the second assistant time; Selecting the particular automated assistant based on the ranking; Then, providing the indication of the audio data to the particular automated assistant instead of providing the indication to any other automated assistant of the secondary automated assistants is further based on selecting the particular automated assistant based on the ranking; 12. The system of claim 11.

13. 13. The system of claim 11 or 12, wherein the particular automated assistant performs automatic speech recognition on the indication of the audio data to generate a text representation of the audio data, generates natural language processing of the text representation, and generates a response based on the natural language processing.

14. It is possible to generate a response, providing results of the natural language processing to a third party application; and receiving a response from the third party application; and generating the response based at least in part on the response from the third party.

15. 15. The system of claim 11, wherein the invocation indicates the particular automated assistant and the indication of the audio data is an audio representation of the audio data.

16. in response to execution of instructions by one or more processors, an act of receiving a call by a generic automated assistant, the act of receiving the call causing the generic automated assistant to be invoked; receiving, via the invoked generic automated assistant, a verbal query captured within audio data generated by one or more microphones of a client device; an operation of identifying historical interaction data for a user, the historical interaction data being generated based on one or more past queries of the user and one or more responses generated by a plurality of secondary automated assistants within a period of time in response to the one or more past queries; determining, based on comparing a transcription of the spoken query to the historical interaction data, that a relationship exists between the spoken query and a portion of the historical interaction data associated with a particular automated assistant among the plurality of secondary automated assistants; In response to determining that the relationship exists between the spoken query and the particular automated assistant, providing an indication of the audio data to the particular automated assistant in lieu of providing the indication to any other automated assistant of the secondary automated assistants, the act of providing the indication of the audio data causing the particular automated assistant to generate a response to the verbal query; At least one non-transitory computer-readable medium comprising instructions for executing the method.

17. The instruction: Determining, based on the historical interaction data, that an additional relationship exists between the verbal query and an additional portion of the historical interaction data associated with a second particular automated assistant of the plurality of secondary automated assistants; Determining a second assistant time for a second assistant historical interaction between the user and the second particular automated assistant based on the historical interaction data; Determining a first assistant time for a first assistant historical interaction between the user and the particular automated assistant based on the historical interaction data; ranking the particular automated assistant and the second particular automated assistant based on the first assistant time and the second assistant time; selecting the particular automated assistant based on the ranking; Further comprising: Providing the indication of the audio data to the particular automated assistant instead of providing the indication to any other automated assistant of the secondary automated assistants is further based on selecting the particular automated assistant based on the ranking.

17. At least one non-transitory computer-readable medium as recited in claim 16.

18. 18. At least one non-transitory computer-readable medium as described in claim 16 or 17, wherein the particular automated assistant performs automatic speech recognition on the indication of the audio data to generate a text representation of the audio data, generates natural language processing of the text representation, and generates a response based on the natural language processing.

19. It is possible to generate a response, providing results of the natural language processing to a third party application; and receiving a response from the third party application; and generating the response based at least in part on the response from the third party.

20. 20. At least one non-transitory computer-readable medium according to any one of claims 16 to 19, wherein the invocation indicates the particular automated assistant and the indication of the audio data is an audio representation of the audio data.

Citation Information

Patent Citations

  • Information processing device, information processing system, information processing method, and program

    JP2021110768A

  • Agent system, agent server, and agent program

    JP2021117302A

  • Session processing interaction between two or more virtual assistants

    US20170269975A1

  • Speaker command and key phrase management for muli -virtual assistant systems

    US20190013019A1

  • Information processing device and information processing method

    WO2020105466A1

Cited By

  • Stapler with composite cardan and screw drive

    US12359696B2