Dynamically adapt the on-device model of grouped assistant devices for collaborative processing of assistant requests.

By adapting on-device models and processing roles among multiple assistant devices based on individual capabilities, the solution enhances robustness and accuracy, reducing latency and network usage while improving data security.

JP7861192B2Active Publication Date: 2026-05-18GOOGLE LLC
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
GOOGLE LLC
Filing Date
2025-04-10
Publication Date
2026-05-18

AI Technical Summary

Technical Problem

Existing assistant devices, particularly older or less expensive ones, lack the processing power and memory capacity to run robust and accurate local models, leading to reduced performance and increased reliance on cloud-based components, which can compromise data security and increase network usage.

Method used

Dynamically adapt on-device models and processing roles among multiple heterogeneous assistant devices in a group, considering individual capabilities, to distribute and collaborate on processing tasks, enhancing robustness and accuracy.

Benefits of technology

This approach increases the robustness and accuracy of processing assistant requests, reduces latency, improves data security, and decreases network usage by allowing local processing to handle more requests independently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861192000001
    Figure 0007861192000001
  • Figure 0007861192000002
    Figure 0007861192000002
  • Figure 0007861192000003
    Figure 0007861192000003
Patent Text Reader

Abstract

To provide implementations directed to dynamically adapting assistant on-device models locally stored at assistant devices of an assistant device group and / or dynamically adapting assistant processing roles of the assistant devices of the assistant device group.SOLUTION: In some of the implementations, the corresponding on-device models and / or corresponding processing roles, for each of the assistant devices of the group, are determined based on collectively considering individual processing capabilities of the assistant devices of the group. The implementations are additionally or alternatively directed to cooperatively utilizing assistant devices of a group, and their associated post-adaptation on-device models and / or post-adaptation processing roles, in cooperatively processing assistant requests that are directed to any one of the assistant devices of the group.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Humans can engage in human-computer dialogue with interactive software applications referred to herein as “Automated Assistants” (also known as “Chatbots,” “Interactive Personal Assistants,” “Intelligent Personal Assistants,” “Personal Voice Assistants,” “Conversational Agents,” etc.). For example, a human (who may be referred to as a “User” when interactively operating an Automated Assistant) can provide commands and / or requests to the Automated Assistant using natural language speech input (i.e., spoken utterance), which in some cases may be converted to text and then processed. Commands and / or requests may, in addition to or alternatively, be provided through one or more other input modalities, such as text (e.g., typed) natural language input, touchscreen input, and / or touch-free gesture input (e.g., detected by the camera of the corresponding Assistant device). Automated Assistants typically respond to commands or requests by providing responsive user interface output (e.g., auditory and / or visual user interface output), controlling smart devices, and / or performing other actions.

[0002] Automated assistants typically rely on a pipeline of components when processing user requests. For example, a wake word detection engine can process audio data when monitoring for the occurrence of an audio wake word (e.g., "OK Assistant") and cause other components to process in response to the detection of the occurrence. As another example, an automatic speech recognition (ASR) engine can be used to process audio data including spoken utterances to generate a transcript of the user's utterance (i.e., a sequence of words and / or other tokens). The ASR engine can process the audio data based on the next occurrence of an audio wake word when detected by the wake word detection engine and / or in response to other invocations of the automated assistant. As another example, a natural language understanding (NLU) engine can be used to process the text of a request (e.g., text converted from a spoken utterance using ASR) to generate a symbolic representation, or belief state, which is a semantic representation of the text. For example, the belief state can include an intent corresponding to the text and optionally parameters (e.g., slot values) for the intent. The belief state represents the actions to be performed in response to the spoken utterance after being fully formed over one or more dialog turns (e.g., after all required parameters have been resolved). Then, another fulfillment component can utilize the fully formed belief state to execute the actions corresponding to the belief state.

[0003] When a user interactively operates an automated assistant, the user utilizes one or more assistant devices (client devices having an automated assistant interface). The pipeline of components utilized when processing requests provided at the assistant device can include components executed locally at the assistant device and / or components executed at one or more remote servers in network communication with the assistant device.

[0004] Efforts have been made to increase the number of locally running components in assistant devices, and / or to improve the robustness and / or accuracy of such components. These efforts are motivated by considerations such as reducing latency, improving data security, reducing network usage, and / or achieving other technological advantages. As an example, some assistant devices may feature a local wake word engine and / or a local ASR engine.

[0005] However, due to the limited processing power of various assistant devices, components implemented locally on assistant devices may be less robust and / or accurate than their cloud-based counterparts. This is especially true for older and / or less expensive assistant devices, which may lack (a) the processing power and / or memory capacity to run various components and / or utilize their associated models, and (b) the disk space to store and / or the various associated models. [Overview of the Initiative] [Problems that the invention aims to solve]

[0006] The implementations disclosed herein are directed toward dynamically adapting locally stored assistant on-device models of assistant devices in an assistant device group, and / or adapting assistant processing roles of assistant devices in an assistant device group. In some of these implementations, the corresponding on-device models and / or corresponding processing roles are determined for each assistant device in the group, based on a collective consideration of the individual processing capabilities of the assistant devices in the group. For example, the on-device models and / or processing roles for a given assistant device may be determined based on the individual processing capabilities of a given assistant device (e.g., the given assistant device can store those on-device models and execute those processing roles, taking into account processor, memory, and / or storage constraints), taking into account the corresponding processing capabilities of other assistant devices in the group (e.g., other devices can store other required on-device models and / or execute other required processing roles). In some implementations, utilization data may also be used when determining the corresponding on-device models and / or corresponding processing roles for each of the assistant devices in the group.

[0007] The implementations disclosed herein, in addition to or alternatively, are directed to collaboratively utilize the group's assistant devices, as well as their associated post-adaptive on-device models and / or post-adaptive processing roles, when coordinating the processing of an assistant request directed to any one of the group's assistant devices. [Means for solving the problem]

[0008] In these and other schemes, on-device models and on-device processing roles may be distributed among multiple heterogeneous assistant devices in a group, taking into account the processing capabilities of those assistant devices. Furthermore, the collective robustness and / or capabilities of on-device models and on-device processing roles, when distributed among multiple assistant devices in a group, exceed those that would be possible individually by any one of the assistant devices. In other words, the implementations disclosed herein can effectively implement on-device pipelines of assistant components distributed among a group of assistant devices. The robustness and / or accuracy of such distributed pipelines far exceeds the robustness and / or accuracy capabilities of any pipeline implemented instead on only a single device among the group of assistant devices. By increasing robustness and / or accuracy as disclosed herein, latency for a larger volume of assistant requests can be reduced. Furthermore, as a result of the increased robustness and / or accuracy, less (or no) data is transmitted to remote automation assistant components when resolving assistant requests. As a direct result of this, user data security is improved, network usage is reduced, the amount of data transmitted over the network is reduced, and / or latency is reduced when resolving assistant requests (for example, those assistant requests resolved locally may be resolved with less latency than when remote assistant components are involved).

[0009] As a result of some adaptations, one or more assistant devices in a group may lack the engine and / or model necessary for that assistant device to process many assistant requests independently. For example, as a result of some adaptations, an assistant device may lack any wake word detection capability and / or ASR capability. However, when within a group and adapted according to the implementation forms disclosed herein, an assistant device can cooperate with other assistant devices in the group to perform its own processing role, and in doing so, utilize its own on-device model to collaboratively process assistant requests. Thus, an assistant request directed to an assistant device can be processed in cooperation with other assistant devices in the group.

[0010] In various implementations, adaptation of a group to its assistant devices is performed in response to the creation or modification of a group (e.g., the inclusion or removal of an assistant device from a group). As described herein, groups may be created based on explicit user input indicating a desire for a group, and / or automatically based on, for example, the determination that the group's assistant devices meet proximity conditions with respect to each other. Implementations that perform adaptation only when a group is created in such a manner can mitigate the occurrence of two assistant requests received simultaneously by two separate devices in a group (which may not be able to be processed concurrently and cooperatively). For example, when proximity conditions are considered when creating a group, it is unlikely that two different assistant devices in a group will receive two heterogeneous simultaneous requests. For example, this is less likely to happen when all the assistant devices in a group are in the same room, as opposed to when the group's assistant devices are scattered across multiple floors of a house. As another example, when user input explicitly indicates that a group should be created, it is likely that non-duplicate assistant requests will be provided to the group's assistant devices.

[0011] As one concrete example of various implementation forms, let us assume that an assistant device group is created consisting of a first assistant device and a second assistant device. Furthermore, let us assume that at the time the assistant device group is created, the first assistant device includes a wake word engine and corresponding wake word model, a warm cue engine and corresponding warm cue model, an authentication engine and corresponding authentication model, and a corresponding on-device ASR engine and corresponding ASR model. Furthermore, let us assume that at the time the assistant device group is created, the second assistant device includes the same engines and models as the first assistant device (or its modified form), and in addition, an on-device NLU engine and corresponding NLU model, on-device fulfillment and corresponding fulfillment model, and an on-device TTS engine and corresponding TTS model.

[0012] In response to the grouping of the first and second assistant devices, the assistant on-device models stored locally in the first and second assistant devices, as well as the corresponding processing roles for the first and second assistant devices, may be adapted. The on-device models and processing roles may be determined for each of the first and second assistant devices based on considering the first processing capacity of the first assistant device and the first processing capacity of the second assistant device. For example, the corresponding processing role / engine may determine a set of on-device models, including a first subset that can be stored and utilized on the first assistant device. Furthermore, the set may include a second subset that can be stored and utilized on the second assistant device, depending on the corresponding processing role. For example, the first subset may include only ASR models, but the ASR models of the first subset may be more robust and / or accurate than the pre-adapted ASR models of the first assistant device. Furthermore, they may require larger computing resources when performing ASR. However, the first assistant device may only have the first subset of ASR models and ASR engines, and the pre-adapted models and engines can be purged, thereby freeing up the computational resources used when running ASR using the first subset of ASR models. Continuing in this example, the second subset may contain the same models as the previously included second assistant device, except that the ASR models may be omitted and more robust and / or accurate NLU models may replace the pre-adapted NLU models. More robust and / or accurate NLU models may require more resources than the pre-adapted NLU models, but these can be freed up through purging the pre-adapted ASR models (and omitting any ASR models from the second subset).

[0013] Next, the first and second assistant devices can collaboratively utilize their associated post-adaptive on-device models and / or post-adaptive processing roles when collaboratively processing an assistant request directed to any of the group's assistant devices. For example, suppose there is a spoken utterance, "OK Assistant, increase the temperature two degrees." The wake cue engine of the second assistant device can detect the appearance of the wake cue "OK Assistant." In response, the wake cue engine can transmit a command to the first assistant device to the ASR engine of the first assistant device to perform speech recognition on the captured speech data following the wake cue. The transcript generated by the ASR engine of the first assistant device may be transmitted to the second assistant device, where the NLU engine of the second assistant device can perform NLU on the transcript. The results of the NLU may be transmitted to the fulfillment engine of the second assistant device, which can use those NLU results to determine a command to increase the temperature by two degrees to transmit to the smart thermostat.

[0014] The above description is presented as an overview of only a few implementations. These and other implementations are disclosed in more detail herein.

[0015] Furthermore, some implementations comprise a system comprising one or more user devices, each device comprising one or more processors and memory operably coupled to one or more processors, the memory of one or more user devices storing instructions that cause one or more processors to perform one or more of the methods described herein in response to the execution of instructions by one or more processors of one or more user devices. Some implementations also comprise at least one non-temporary computer-readable medium containing instructions that cause one or more processors to perform one or more of the methods described herein in response to the execution of instructions by one or more processors.

[0016] It should be understood that all combinations of the aforementioned concepts and any additional concepts described in more detail herein are intended to be part of the subject matter disclosed herein. For example, all combinations of claimed subject matter appearing at the end of this disclosure are intended to be part of the subject matter disclosed herein. [Brief explanation of the drawing]

[0017] [Figure 1A] This is a block diagram of an exemplary assistant ecosystem without grouped and adapted assistant devices in the implementation forms disclosed herein. [Figure 1B1] Figure 1A illustrates an exemplary assistant ecosystem in which the first and second assistant devices are grouped together, and different examples of adaptations can be implemented. [Figure 1B2] Figure 1A illustrates an exemplary assistant ecosystem in which the first and second assistant devices are grouped together, and different examples of adaptations can be implemented. [Figure 1B3]Figure 1A illustrates an exemplary assistant ecosystem in which the first and second assistant devices are grouped together, and different examples of adaptations can be implemented. [Figure 1C] Figure 1A illustrates an exemplary assistant ecosystem in which a first assistant device, a second assistant device, and a third assistant device are grouped together, and an example of adaptation can be implemented. [Figure 1D] Figure 1A illustrates an exemplary assistant ecosystem in which a third and fourth assistant device are grouped together, and an example of adaptation can be implemented. [Figure 2] This flowchart illustrates an exemplary method for adapting the on-device model and / or processing role of assistant devices within a group. [Figure 3] This flowchart illustrates one exemplary method that can be performed by each of the multiple assistant devices in a group when adapting the on-device model and / or processing role of the assistant devices within the group. [Figure 4] This is a diagram illustrating an exemplary architecture of a computing device. [Modes for carrying out the invention]

[0018] Many users may engage with an automated assistant using any one of several assistant devices. For example, some users may own a collaborative “ecosystem” of assistant devices that can receive user input directed to the automated assistant and / or be controlled by the automated assistant, among other things, including one or more smartphones, one or more tablet computers, one or more vehicle computing systems, one or more wearable computing devices, one or more smart TVs, one or more interactive standalone speakers, one or more interactive standalone speakers with displays, one or more IoT devices, and so on.

[0019] Users can use any of these assistant devices to engage in human-computer dialogue with an automated assistant (assuming the automated assistant client is installed and the assistant device is able to receive input). In some cases, these assistant devices may be scattered around the user's primary residence, second home, workplace, and / or other structures. For example, mobile assistant devices such as smartphones, tablets, and smartwatches may be where the user is wearing and / or where the user last put them. Other assistant devices such as traditional desktop computers, smart TVs, interactive standalone speakers, and IoT devices may often be relatively stationary, but nevertheless, they may be placed in various locations (e.g., rooms) within the user's home or workplace.

[0020] Referring first to Figure 1A, an exemplary assistant ecosystem is illustrated. The exemplary assistant ecosystem includes a first assistant device 110A, a second assistant device 110B, a third assistant device 110C, and a fourth assistant device 110D. Assistant devices 110A-D may all be deployed in a home, business, or other environment. Furthermore, all assistant devices 110A-D may be linked together in one or more data structures, or otherwise associated with one another. For example, all four assistant devices 110A-D may all be registered under the same user account, registered under the same set of user accounts, registered under a specific structure, and / or all assigned to a specific structure in a device topology representation. The device topology representation may include a corresponding unique identifier for each of the assistant devices 110A-D, and optionally may include a corresponding unique identifier for other devices that are not assistant devices (but can be interactively operated through an assistant device), such as IoT devices that do not have an assistant interface. Furthermore, the device topology representation can specify device attributes associated with each assistant device 110A-D. Device attributes for a given assistant device may include, for example, one or more input and / or output modalities supported by each assistant device, the processing capacity of each assistant device, the manufacturing type, model, and / or unique identifier (e.g., serial number) of each assistant device (on which processing capacity may be determined), and / or other attributes. As another example, the four assistant devices may all be linked together or otherwise associated with each other as a function of connecting to the same wireless network, such as a secure access wireless network, and / or as a function of collectively (e.g., via Bluetooth, after pairing) communicating peer-to-peer with one another.Stated another way, in some implementations, the plurality of assistant devices are linked together, as a function of making a secure network connection with each other, and not necessarily associated with each other in any data structure, and can potentially be adapted by the implementations disclosed herein.

[0021] In non-limiting embodiments, the first assistant device 110A may be a first type of assistant device, such as a specific model of an interactive standalone speaker equipped with a display and a camera. The second assistant device 110B may be a second type of assistant device, such as a first model of an interactive standalone speaker without a display or camera. Assistant devices 110C and 110D may each be a third type of assistant device, such as a third model of an interactive standalone speaker without a display. The third type (assistant devices 110C and 110D) may have lower processing power than the second type (assistant device 110D). For example, the third type may have a processor with lower processing power than the second type processor. For example, the third type processor may not include a GPU, while the first type processor does. Also, for example, the third type processor may have a smaller cache and / or a lower operating frequency than the second type processor. As another example, the size of the third type of on-device memory can be smaller than that of the second type of on-device memory (for example, 1GB compared to 2GB). Yet another example: the available disk space of the third type can be smaller than that of the first type. The available disk space may differ from the currently available disk space. For example, the available disk space may be determined by adding the currently available disk space to the disk space currently occupied by one or more on-device models. As yet another example, the available disk space may be the total disk space minus the space occupied by the operating system and / or other specific software. Continuing in this embodiment, the first and second types may have the same processing power.

[0022] In addition to being linked together in a data structure, two or more (e.g., all) of the assistant devices 110A - D communicate with each other at least selectively via a local area network (LAN) 108. The LAN 108 can include a wireless network such as one that utilizes Wi-Fi, a direct peer-to-peer network such as one that utilizes Bluetooth, and / or other communication topologies that utilize other communication protocols.

[0023] Assistant device 110A includes an assistant client 120A, which can be a stand-alone application on top of an operating system or can form all or part of the operating system of assistant device 110A. Assistant client 120A includes, in FIG. 1A, a wake / call engine 121A1 and one or more associated on-device wake / call models 131A1. Wake / call engine 121A1 monitors for the occurrence of one or more wake or call queues and, in response to detecting one or more queues, can call one or more previously inactive functions of assistant client 120A. For example, calling assistant client 120A can include starting an ASR engine 122A1, an NLU engine 123A1, and / or other engines. For example, this can cause the ASR engine 122A1 to process additional audio data frames following a wake or call queue (whereas prior to the call, no further processing of audio data frames occurred), and / or cause assistant client 120A to transmit additional audio data frames and / or other data to be transmitted for processing to a cloud-based assistant component 140 (e.g., processing of audio data frames by a remote ASR engine of the cloud-based assistant component 140).

[0024] In some implementations, the wake queue engine 121A can continuously process a stream of audio data frames based on the output from one or more microphones of the client device 110A (for example, when not in "inactive" mode) and monitor for the occurrence of an audio wake word or call phrase (e.g., "OK Assistant", "Hey Assistant"). Processing may be performed by the wake queue engine 121A utilizing one or more of the wake models 131A1. For example, one of the wake models 131A1 may be a neural network model trained to process frames of audio data and generate an output indicating whether one or more wake words are present in the audio data. While monitoring for wake word occurrences, the wake queue engine 121A1 discards any audio data frames that do not contain wake words (e.g., after temporary storage in a buffer). In addition to monitoring for wake word occurrences, or instead, the wake queue engine 121A1 can monitor for occurrences of other call queues. For example, the wake queue engine 121A1 can also monitor presses of call hardware buttons and / or call software buttons. As another example, in a subsequent embodiment, when the assistant device 110A is equipped with a camera, the wake queue engine 121A1 may also optionally process image frames from the camera when monitoring for the appearance of calling gestures such as hand gestures while the user's gaze is directed towards the camera and / or the appearance of other calling cues such as the user's gaze being directed towards the camera along with an instruction that the user is speaking.

[0025] The assistant client 120A also includes an automatic speech recognition (ASR) engine 122A1 and one or more associated on-device ASR models 132A1, as shown in Figure 1A. The ASR engine 122A1 may be used to process speech data, including vocalized utterances, and to generate a transcript of the user's speech (i.e., a sequence of words and / or other tokens). The ASR engine 122A1 can process speech data using the on-device ASR model 132A1. The on-device ASR model 132A1 may include, for example, a two-pass ASR model, which is a neural network model, and is used by the ASR engine 122A1 to generate a sequence of probabilities on tokens (and probabilities used to generate the transcript). As another example, the on-device ASR model 132A1 may include an acoustic model, which is a neural network model, and a language model, which includes mappings of phoneme sequences to words. The ASR engine 122A1 can process speech data using an acoustic model, generate sequences of phonemes, and map those phoneme sequences to specific words using a language model. Additional or alternative ASR models may be available.

[0026] The assistant client 120A, as shown in Figure 1A, also includes a natural language understanding (NLU) engine 123A1 and one or more associated on-device NLU models 133A1. The NLU engine 123A1 can generate symbolic representations, which are semantic representations of natural language text, such as transcript text generated by the ASR engine 122A1 or typed text (for example, typed using the virtual keyboard of the assistant device 110A), or belief states. For example, a belief state may include an intent corresponding to the text and optionally parameters for the intent (for example, slot values). The belief state represents an action to be performed in response to a spoken utterance after it has been fully formed through one or more dialogue turns (for example, after all required parameters have been resolved). When generating symbolic representations, the NLU engine 123A1 can utilize one or more on-device NLU models 133A1. The NLU model 133A1 may include one or more neural network models trained to process text and generate an output indicating the intent represented by the text, and / or instructions on which parts of the text correspond to which parameters of the intent. In addition to or alternatively, the NLU model may include one or more models that include mappings of text and / or templates to their corresponding symbolic representations. For example, these mappings may include a mapping of the text "what time is it" to the intent "current time" with the parameter "current location". As another example, the mapping may include a mapping of the template "add [item(s)] to my shopping list" to the intent "insert in shopping list" with the parameter of the items in the actual natural language corresponding to [item(s)] in the template.

[0027] The assistant client 120A also includes a fulfillment engine 124A1 and one or more associated on-device fulfillment models 134A1 in Figure 1A. The fulfillment engine 124A1 can use a fully formed symbolic representation to perform or cause the NLU engine 123A1 to perform an action corresponding to the symbolic representation. The action may include providing responsive user interface output (e.g., auditory and / or visual user interface output), controlling a smart device, and / or performing other actions. When performing or causing an action to be performed, the fulfillment engine 124A1 can utilize the fulfillment model 134A1. For example, in response to a "turn on" intent with parameters specifying a particular smart light, the fulfillment engine 124A1 can utilize the fulfillment model 134A1 to identify the network address of the particular smart light and / or the command to transmit to bring the particular smart light into the "on" state. As another example, for a "current" intent with the parameter "current location", the fulfillment engine 124A1 may use the fulfillment model 134A1 to identify that the current time on the client device 110A should be retrieved and (using the TTS engine 125A1) rendered audibly.

[0028] The assistant client 120A also includes a text-to-speech (TTS) engine 125A1 and one or more associated on-device TTS models 135A1 in Figure 1A. The TTS engine 125A1 can process text (or its speech representation) using the on-device TTS models 135A1 and generate synthesized speech. The synthesized speech may be audibly rendered through the speaker of the assistant device 110A's local text-to-speech ("TTS") engine (which converts text to speech). The synthesized speech may be generated and rendered as all or part of a response from the automated assistant, and / or when prompting the user to define and / or clarify parameters and / or intents (for example, so as to be orchestrated by the NLU engine 123A1 and / or another dialogue state engine).

[0029] The assistant client 120A also includes an authentication engine 126A1 and one or more associated on-device authentication models 136A1 in Figure 1A. The authentication engine 126A1 can utilize one or more authentication techniques to verify which of multiple registered users is interactively operating the assistant device 110, or, if only a single user is registered with the assistant device 110, whether the user interactively operating the assistant device 110 is a registered user (or a guest / unregistered user instead). As an example, a text-dependent speaker verification (TD-SV) may be generated and stored for each registered user (for example, in relation to a corresponding user profile) with permission from the user concerned. The authentication engine 126A1 can utilize the TD-SV model of the on-device authentication model 136A1 when generating a corresponding TD-SV and / or when processing a corresponding portion of the voice data TD-SV and generating a corresponding current TD-SV that can then be compared with a stored TD-SV to determine if there is a match. As another example, the authentication engine 126A1 may, in addition to or alternatively, utilize text-independent speaker verification (TI-SV) ​​technology, speaker verification technology, facial verification technology, and / or other verification technologies (e.g., PIN entry), and the corresponding on-device authentication model 136A1 when authenticating a particular user.

[0030] The assistant client 120A also includes a warm queue engine 127A1 and one or more associated on-device warm queue models 137A1 in Figure 1A. The warm queue engine 127A1 selectively monitors for the occurrence of one or more warm words or other warm queues and can cause the assistant client 120A to perform a specific action in response to the detection of one or more warm queues. A warm queue may be added to any wake word or other wake queue, and each warm queue may be at least selectively active. In particular, detecting the occurrence of a warm queue can trigger the performance of a specific action even if there was no wake queue prior to the detected occurrence. Therefore, when a warm queue is one or more specific words, the user may simply speak those words without needing to add a wake queue, and this can trigger the performance of the corresponding specific action.

[0031] For example, a “stop” warm queue may be active at least when a timer or alarm is being audibly rendered by the assistant device 110A via the automation assistant 120A. For instance, in such a situation, the warm queue engine 127A may continuously (or at least when the VAD engine 128A1 detects vocal activity) process a stream of audio data frames based on the output from one or more microphones of the client device 110A to monitor for occurrences of “stop,” “halt,” or other limited sets of specific warm words. Processing may be performed by the warm queue engine 127A utilizing one of the warm queue models 137A1, such as a neural network model trained to process frames of audio data and generate an output indicating whether a “stop” utterance occurrence is present in the audio data. In response to the detection of a “stop” occurrence, the warm queue engine 127A may trigger the execution of a command to stop the timer or alarm that is sounding. In such a situation, the warm cue engine 127A can continuously process a stream of images from the camera of the assistant device 110A (or at least when the motion sensor detects presence) to monitor for the appearance of a hand in a "stop" position. Processing may be performed by the warm cue engine 127A utilizing one of the warm cue models 137A1, such as a neural network model trained to process frames of vision data and generate an output indicating whether a hand is present in a "stop" position. In response to the detection of the appearance of a "stop" position, the warm cue engine 127A can be instructed to execute a command to stop any timers or alarms that are sounding.

[0032] As another example, the “volume up,” “volume down,” and “next” warm cues may be active at least when music is being audibly rendered by the assistant device 110A via the automation assistant 120A. For example, in such a case, the warm cue engine 127A can sequentially process a stream of audio data frames based on the output from one or more microphones of the client device 110A. Processing may include monitoring for the occurrence of “volume up” using a first of the warm cue models 137A1, monitoring for the occurrence of “volume down” using a second of the warm cue models 137A1, and monitoring for the occurrence of “next” using a third of the warm cue models 137A1. In response to the detection of the occurrence of "volume up", the warm queue engine 127A can execute a command to increase the volume of the music being rendered; in response to the detection of the occurrence of "volume down", the warm queue engine can execute a command to decrease the volume of the music; and in response to the detection of the occurrence of "volume down", the warm queue engine can execute a command that causes the next track to be rendered instead of the current music track.

[0033] The assistant client 120A also includes a voice activity detector (VAD) engine 128A1 and one or more associated on-device VAD models 138A1 in Figure 1A. The VAD engine 128A1 can selectively monitor for the occurrence of voice activity in voice data and, in response to detecting such occurrences, can trigger one or more functions to be executed by the assistant client 120A. For example, the VAD engine 128A1 can activate the wake queue engine 121A1 in response to detecting voice activity. Alternatively, the VAD engine 128A1 may be used in continuous listening mode, which can monitor for the occurrence of voice activity in voice data and, in response to detecting such occurrences, activate the ASR engine 122A1. The VAD engine 128A1 can process the voice data using the VAD model 138A1 when determining whether voice activity is present in the voice data.

[0034] Specific engines and corresponding models have been described for the Assistant Client 120A. However, it should be noted that some engines may be omitted, and / or additional engines may be included. Also, it should be noted that the Assistant Client 120A can fully process many assistant requests, including many assistant requests provided as spoken utterances, through its various on-device engines and corresponding models. However, since the client device 110A is relatively limited in terms of processing capacity, there are still many assistant requests that cannot be fully processed locally in the Assistant Device 110A. For example, the NLU engine 123A1 and / or the corresponding NLU model 133A1 may only cover a subset of all available intents and / or parameters available through the automated assistant. As another example, the fulfillment engine 124A1 and / or the corresponding fulfillment model may only cover a subset of available fulfillments. As yet another example, the ASR engine 122A1 and its corresponding ASR model 132A1 may not have sufficient robustness and / or accuracy to correctly transcribe various speech utterances.

[0035] Considering these and other considerations, the cloud-based assistant component 140 may still be used, at least selectively, when performing processing of at least some of the assistant requests received by the assistant device 110A. The cloud-based automated assistant component 140 may include engines and / or models (and / or additional or alternative forms) that are paired with the engine of the assistant device 110A. However, because the cloud-based automated assistant component 140 can leverage the virtually unlimited resources of the cloud, one or more cloud-based counterparts may be more robust and / or accurate than those of the assistant client 120A. For example, in response to a spoken utterance requesting the execution of an assistant action not supported by the local NLU engine 123A1 and / or the local fulfillment engine 124A1, the assistant client 120A may transmit the voice data for the spoken utterance and / or a transcript of it generated by the ASR engine 122A1 to the cloud-based automated assistant component 140. A cloud-based automation assistant component 140 (for example, its NLU engine and / or fulfillment engine) may perform more robust processing of such data and enable the resolution and / or execution of assistant actions. The transmission of data to the cloud-based automation assistant component 140 is via one or more wide area networks (WANs) 109, such as the Internet or a private WAN.

[0036] The second assistant device 110B includes an assistant client 120B, which may be a standalone application on top of an operating system, or may form all or part of the operating system of the assistant device 110B. Similar to assistant client 120A, assistant client 120B includes a wake / call engine 121B1 and one or more associated on-device wake / call models 131B1, an ASR engine 122B1 and one or more associated on-device ASR models 132B1, an NLU engine 123B1 and one or more associated on-device NLU models 133B1, a fulfillment engine 124B1 and one or more associated on-device fulfillment models 134B1, a TTS engine 125B1 and one or more associated on-device TTS models 135B1, an authentication engine 126B1 and one or more associated on-device authentication models 136B1, a warm queue engine 127B1 and one or more associated on-device warm queue models 137B1, and a VAD engine 128B1 and one or more associated on-device VAD models 138B1.

[0037] The engine and / or model of Assistant Client 120B may be the same as that of Assistant Client 120A, and the engine and / or model may be different in some or all respects. For example, Wake Queue Engine 121B1 may not have the capability to detect wake queues in an image, and / or Wake Model 131B1 may not have a model for processing an image to detect wake queues, while Wake Queue Engine 121A1 has such capability, and Wake Model 131B1 includes such a model. This may be due, for example, to Assistant Device 110A including a camera, while Assistant Device 110B not including a camera. As another example, ASR Model 131B1 used by ASR Engine 122B1 may be different from ASR Model 131A1 used by ASR Engine 122A1. This may be due, for example, to different models being optimized for different processors and / or memory capabilities between Assistant Device 110A and Assistant Device 110B.

[0038] Specific engines and corresponding models have been described for the Assistant Client 120B. However, it should be noted that some engines may be omitted, and / or additional engines may be included. Also, it should be noted that the Assistant Client 120B can fully process many assistant requests, including many assistant requests provided as spoken utterances, through its various on-device engines and corresponding models. However, since the client device 110B is relatively limited in terms of processing capacity, there are still many assistant requests that cannot be fully processed locally on the Assistant Device 110B. In light of these and other considerations, the cloud-based assistant component 140 may still be used, at least selectively, to perform processing of at least some of the assistant requests received on the Assistant Device 110B.

[0039] The third assistant device 110C includes an assistant client 120C, which may be a standalone application on top of an operating system or may form all or part of the operating system of the assistant device 110C. Similar to assistant clients 120A and 120B, assistant client 120C includes a wake / call engine 121C1 and one or more associated on-device wake / call models 131C1, an authentication engine 126C1 and one or more associated on-device authentication models 136C1, a warm queue engine 127C1 and one or more associated on-device warm queue models 137C1, and a VAD engine 128C1 and one or more associated on-device VAD models 138C1. Some or all of the engines and / or models of assistant client 120C may be the same as those of assistant clients 120A and / or assistant client 120B, and some or all of the engines and / or models may be different.

[0040] However, unlike assistant clients 120A and 120B, it should be noted that assistant client 120C does not include any ASR engine or associated model, any NLU engine or associated model, any fulfillment engine or associated model, and any TTS engine or associated model. Furthermore, it should be noted that assistant client 120B, through its various on-device engines and corresponding models, can only fully process some assistant requests (i.e., those that fit into the warm queue detected by the warm queue engine 127C1) and cannot process many assistant requests, such as those provided as spoken utterances that do not fit into the warm queue. In light of these and other considerations, the cloud-based assistant component 140 can still be used, at least selectively, when performing processing of at least some of the assistant requests received in the assistant device 110C.

[0041] The fourth assistant device 110D includes an assistant client 120D, which may be a standalone application on top of an operating system or may form all or part of the operating system of the assistant device 110D. Similar to assistant clients 120A, 120B, and 120C, assistant client 120D includes a wake / call engine 121D1 and one or more associated on-device wake / call models 131D1, an authentication engine 126D1 and one or more associated on-device authentication models 136D1, a warm queue engine 127D1 and one or more associated on-device warm queue models 137D1, and a VAD engine 128D1 and one or more associated on-device VAD models 138D1. Some or all of the engines and / or models of assistant client 120C may be the same as those of assistant clients 120A, 120B, and / or 120C, and some or all of the engines and / or models may be different.

[0042] However, it should be noted that, unlike Assistant Clients 120A and 120B, and similar to Assistant Client 120C, Assistant Client 120D does not include any ASR engine or associated model, any NLU engine or associated model, any fulfillment engine or associated model, and any TTS engine or associated model. Furthermore, it should be noted that Assistant Client 120D can only fully process a limited number of assistant requests (i.e., those that fit into the warm queue detected by the warm queue engine 127D1) through its various on-device engines and corresponding models, and cannot process many assistant requests, such as those provided as spoken utterances that do not fit into the warm queue. In light of these and other considerations, the cloud-based assistant component 140 can still be used, at least selectively, when performing processing of at least some of the assistant requests received in Assistant Device 110D.

[0043] Referring next to Figures 1B1, 1B2, 1B3, 1C, and 1D, different non-limiting examples of assistant device groups are illustrated along with different non-limiting examples of adaptations that may be implemented in response to the generation of the assistant device groups. Through each adaptation, the grouped assistant devices may be used collectively when processing various assistant requests, and through collectively using them, the processing of those various assistant requests can be performed with greater robustness and / or accuracy than if any one of the assistant devices in the group could be performed individually before the adaptation. As a result, various technical advantages, such as those described herein, can be obtained.

[0044] In Figures 1B1, 1B2, 1B3, 1C, and 1D, the engines and models of assistant clients having the same reference numbers as in Figure 1A are not applicable with respect to Figure 1A. For example, in Figures 1B1, 1B2, and 1B3, the engines and models of assistant client devices 110C and 110D are not applicable because they are not included in group 101B in Figures 1B1, 1B2, and 1B3. However, in Figures 1B1, 1B2, 1B3, 1C, and 1D, the engines and models of assistant clients having different reference numbers than those in Figure 1A (i.e., ending in "2", "3", or "4" instead of "1") indicate that they are applicable with respect to their counterparts in Figure 1. Furthermore, an engine or model having a reference number ending in "2" in one figure and "3" in the other means that different applications of the engine or model have been made between the figures. Similarly, an engine or model with a reference number ending in "4" in the figure means that the adaptation for that engine or model is different from that in a figure where the reference number ends in "2" or "3".

[0045] Referring first to Figure 1B1, a device group 101B has been created, and assistant devices 110A and 110B are included in device group 101B. In some implementations, device group 101B may be generated in response to a user interface input that explicitly indicates a desire to group assistant devices 110A and 110B. For example, a user may provide the voice utterance “group [label for assistant device 110A] and [label for assistant device 110B]” for any one of assistant devices 110A-D. Such a voice utterance is processed by the respective assistant device and / or cloud-based assistant component 140 and interpreted as a request to group assistant devices 110A and 110B, and group 101B may be generated in response to such interpretation. As another example, a registered user of assistant devices 110A-D may provide touch input in an application that allows configuring the settings of assistant devices 110A-D. These touch inputs can explicitly specify that assistant devices 110A and 110B should be grouped, and group 101B may be generated in response. As yet another example, when automatically generating device group 101B, one of the exemplary techniques described below may be used instead to determine that device group 101B should be generated, but user input explicitly authorizing the generation of device group 101B may be requested before device group 101B is generated. For example, a prompt indicating that device group 101B should be generated may be rendered on one or more of the assistant devices 110A-D, and device group 101B is actually generated only if a positive user interface input is received in response to the prompt (and optionally verified to be coming from a registered user).

[0046] In some implementations, device group 101B may be automatically generated instead. In some of these implementations, a user interface output indicating the generation of device group 101B may be rendered on one or more of the assistant devices 110A-D to notify the corresponding user of the group, and / or a registered user may override the automatic generation of device group 101B via a user interface input. However, when device group 101B is automatically generated, the corresponding adaptation is made without initially requiring a user interface input that explicitly indicates that device group 101B is to be generated and that a specific device group 101B should be created (although an input earlier in time may indicate approval to create the group in general). In some implementations, device group 101B may be automatically generated in response to assistant devices 110A and 110B determining that one or more proximity conditions are met with respect to each other. For example, proximity conditions may include the fact that assistant devices 110A and 110B are assigned to the same structure (e.g., a specific house, a specific villa, a specific office) and / or the same room (e.g., a kitchen, living room, dining room) or other areas within the same structure in the device topology. Another example of proximity conditions may include the fact that sensor signals from each of the assistant devices 110A and 110B indicate that they are geographically close to each other. For example, if both assistant devices 110A and 110B consistently detect the appearance of a wake word at the same time or close to it (e.g., within 1 second) (e.g., beyond 70% of the time or other thresholds), this may indicate that they are geographically close to each other. Alternatively, for example, one of the assistant devices 110A and 110B may emit a signal (e.g., ultrasound), and the other assistant device 110A and 110B may attempt to detect the emitted signal.If the other of the assistant devices 110A and 110B detects the emitted signal by a threshold strength, it may indicate that they are spatially close to each other. Additional and / or alternative techniques may be used to determine temporal proximity and / or to automatically generate device groups.

[0047] Regardless of how group 101B was generated, Figure 1B1 shows an example of possible adaptations that may be made to assistant devices 110A and 110B in response to their inclusion in group 101B. In various implementations, one or both of assistant clients 120A and 120B may determine and trigger adaptations to be made. In other implementations, one or more engines of the cloud-based assistant component 140 may, in addition to or alternatively, determine and trigger adaptations to be made. As described herein, adaptations to be made may be determined based on considering the processing power of both assistant clients 120A and 120B. For example, adaptations may attempt to utilize the collective processing power as much as possible while ensuring that the individual processing power of each assistant device is sufficient for the engine and / or model to be stored and utilized locally in the assistant. Furthermore, adaptations to be made may also be determined based on utilization data that reflects metrics relating to the actual usage of the group's assistant devices and / or other ungrouped assistant devices in the ecosystem. For example, if processing power allows for either a larger, more accurate wake queue model or a larger, more robust warm queue model, but not both, utilization data can be used to select one of the two options. For instance, if utilization data reflects infrequent (or non-existent) use of warm words, and / or wake word detection often barely exceeds the threshold, and / or false negatives for wake words occur frequently, a larger, more accurate wake queue model may be chosen. On the other hand, if utilization data reflects frequent use of warm words, and / or wake word detection consistently exceeds the threshold, and / or false negatives for wake words are rare, a larger, more accurate warm queue model may be chosen.Consideration of such usage data may determine whether the adaptation shown in Figure 1B1, Figure 1B2, or Figure 1B3 is selected (Figures 1B1, 1B2, or 1B3 each show different adaptations for the same group 101B).

[0048] In Figure 1B1, the fulfillment engine 124A1 and fulfillment model 134A1, and the TTS engine 125A1 and TTS model 135A1 have been purged from the assistant device 110A. Furthermore, the assistant device 110A has a different ASR engine 122A2 and a different on-device ASR model 132A2, as well as a different NLU engine 123A2 and a different on-device NLU model 133A2. Different engines can be downloaded from a local model repository 150, which is accessible in the assistant device 110A via interaction with the cloud-based assistant component 140. In Figure 1B1, the wake queue engine 121B1, the ASR engine 122B1, and the authentication engine 126B1, as well as their corresponding models 131B1, 133B1, and 136B1, have already been purged from the assistant device 110B. Furthermore, the assistant device 110B has different NLU engines 123B2 and different on-device NLU models 133B2, different fulfillment engines 124B2 and fulfillment models 133B2, and different warm word engines 127B2 and warm word models 137B2. Different engines can be downloaded from a local model repository 150, which is accessible in the assistant client 110B via interaction with the cloud-based assistant component 140.

[0049] The ASR engine 122A2 and ASR model 132A2 of the assistant device 110A are more robust and / or accurate than the ASR engine 122A1 and ASR model 132A1, but may occupy more disk space, utilize more memory, and / or require more processor resources. For example, the ASR model 132A1 may only include a single-pass model, while the ASR model 132A2 may include a two-pass model.

[0050] Similarly, while the NLU engine 123A2 and NLU model 133A2 are more robust and / or accurate than the NLU engine 123A1 and NLU model 133A1, they may occupy more disk space, utilize more memory, and / or require more processor resources. For example, NLU model 133A1 may already contain intents and parameters for a first classification such as "lighting control," while NLU model 133A2 may include intents not only for "lighting control" but also for "thermostat control," "smart lock control," and "reminders."

[0051] Therefore, the ASR engine 122A2, ASR model 132A2, NLU engine 123A2, and NLU model 133A2 are improvements with respect to their replaced counterparts. However, it should be noted that the processing capacity of the assistant device 110A may prevent it from storing and / or using the ASR engine 122A2, ASR model 132A2, NLU engine 123A2, and NLU model 133A2 without first purging the Fulfillment Engine 124A1 and Fulfillment Model 134A1 and the TTS Engine 125A1 and TTS Model 135A1. Simply purging such models from the assistant device 110A without complementary adaptations for the assistant device 110B and cooperative processing with the assistant device 110B would result in the assistant client 120A lacking the ability to process various assistant requests entirely locally (i.e., without requiring the use of one or more cloud-based assistant components 140).

[0052] Therefore, complementary adaptations are made to the assistant device 110B, and after the adaptations, cooperative processing takes place between the assistant devices 110A and 110B. The NLU engine 123B2 and NLU model 133B2 for the assistant device 110B are more robust and / or accurate than the NLU engine 123B1 and NLU model 133B1, but may occupy more disk space, utilize more memory, and / or require more processor resources. For example, the NLU model 133B1 may already include intents and parameters for a first classification such as "lighting control". However, the NLU model 133B2 may cover more intents and parameters. Note that the intents and parameters covered by the NLU model 133B2 may be limited to intents not already covered by the NLU model 133A2 of the assistant client 120A. This prevents functional overlap between assistant clients 120A and 120B and extends the collective capability of assistant clients 120A and 120B when they collaboratively process assistant requests.

[0053] Similarly, while the fulfillment engine 124B2 and fulfillment model 134B2 are more robust and / or accurate than the fulfillment engine 124B1 and fulfillment model 124B1, they may occupy more disk space, utilize more memory, and / or require more processor resources. For example, the fulfillment model 124B1 may already have the fulfillment capability for a single classification of the NLU model 133B1, but the fulfillment model 124B2 may include the fulfillment capability for all classifications of the NLU model 133B2, and even for the NLU model 133A2.

[0054] Therefore, the fulfillment engine 124B2, fulfillment model 134B2, NLU engine 123B2, and NLU model 133B2 are improvements with respect to their replaced counterparts. However, the processing capacity of the assistant device 110B may prevent it from storing and / or using the fulfillment engine 124B2, fulfillment model 134B2, NLU engine 123B2, and NLU model 133B2 without first purging the purged models and purged engines from the assistant device 110B. Simply purging such models from the assistant device 110B without complementary adaptations for the assistant device 110A and cooperative processing with the assistant device 110A would result in the assistant client 120B lacking the ability to process various assistant requests entirely locally.

[0055] The warm queue engine 127B2 and warm queue model 137B2 of client device 110B occupy less additional disk space, use less memory, and require less processor resources compared to the warm queue engine 127B1 and warm queue model 127B1. For example, the processing power they require may be the same or even less. However, the warm queue engine 127B2 and warm queue model 137B2 cover warm queues in addition to those covered by the warm queue engine 127B1 and warm queue model 127B1—and in addition to those covered by the warm queue engine 127A1 and warm queue model 127A1 of assistant client 120A.

[0056] In the configuration shown in Figure 1B1, Assistant Client 120A may be assigned processing roles for monitoring the wake queue, performing ASR, performing NLU for a first set of classifications, performing authentication, monitoring a first set of warm queues, and performing VAD. Assistant Client 120B may be assigned processing roles for performing NLU for a second set of classifications, performing fulfillment, performing TTS, and monitoring a second set of warm queues. Processing roles may be communicated and stored within each Assistant Client 120A, and coordination of processing for various Assistant requests may be performed by either or both Assistant Clients 120A and 120B.

[0057] As an example of coordinated processing of an assistant request using the adaptation in Figure 1B1, suppose the voice utterance "OK Assistant, turn on the kitchen lights" is provided, and the assistant device 120A is the lead device coordinating the processing. The wake queue engine 121A1 of the assistant client 120A can detect the appearance of the wake queue "OK Assistant". In response, the wake queue engine 121A1 can cause the ASR engine 122A2 to process the captured voice data following the wake queue. The wake queue engine 121A1 can also optionally transmit a command to the assistant device 110B locally to transition from a low-power state to a high-power state so that the assistant client 120B can immediately perform specific processing of the assistant request. The voice data processed by the ASR engine 122A2 may be voice data captured by the microphone of the assistant device 110A and / or voice data captured by the microphone of the assistant device 110B. For example, a command transmitted to the assistant device 110B to transition to a higher power state may also cause it to capture audio data locally and optionally transmit such audio data to the assistant client 120A. In some implementations, the assistant client 120A can decide whether to use the received audio data, or instead the locally captured audio data, based on an analysis of the characteristics of each instance of the audio data. For example, an instance of audio data may be used over another instance based on the instance that has a lower signal-to-noise ratio and / or captures vocal utterances at a higher volume.

[0058] The transcript generated by the ASR engine 122A2 is transmitted to the NLU engine 123A2 to perform an NLU on the transcript for the first set of classifications, and can also be transmitted to the assistant client 120B to perform an NLU on the transcript for the second set of classifications. The results of the NLU performed by the NLU engine 123B2 are transmitted to the assistant client 120A, which can then decide which results to use, if any, based on those results and the results from the NLU engine 123A2. For example, the assistant client 120A can use the result with the highest probability intent, as long as that probability meets a certain threshold. For example, a result containing the intent "turn on" and a parameter specifying an identifier for "kitchen lights" may be used. Note that if the probability does not meet the threshold, the NLU engine of the cloud-based assistant component 140 may, at its discretion, be used to perform an NLU. The assistant client 120A can transmit the NLU result with the highest probability to the assistant client 120B. The fulfillment engine 124B2 of the assistant client 120B can use those NLU results to decide to transmit a command to the kitchen lights to transition them to the "on" state, and can transmit such a command over LAN 108. Optionally, the fulfillment engine 124B2 can use the TTS engine 125B1 to generate synthesized speech confirming the execution of "turning on the kitchen lights". In such a situation, the synthesized speech can be rendered by the assistant client 120B on the assistant device 110B and / or transmitted to the assistant device 110A for rendering by the assistant client 120A.

[0059] As another example of cooperative processing of assistant requests, suppose assistant client 120A is rendering an alarm for its local timer that has just finished. Furthermore, suppose a warm queue monitored by warm queue engine 127B2 contains "stop", and a warm queue monitored by warm queue engine 127A1 does not contain "stop". Finally, suppose a spoken utterance of "stop" is provided when the alarm is being rendered and is captured in audio data detected via the microphone of assistant client 120B. Warm word engine 127B2 can process the audio data and determine the occurrence of the word "stop". Furthermore, warm word engine 127B2 can determine that the occurrence of the word "stop" maps directly to a command that cancels the timer or alarm that is sounding. That command is transmitted by assistant client 120B to assistant client 120A, thereby causing assistant client 120A to execute the command and cancel the timer or alarm that is sounding. In some implementations, the warm word engine 127B2 may only need to monitor for the appearance of "stop" in certain situations. In those implementations, the assistant client 120A can transmit a command to the warm word engine 127B2 to monitor for the appearance of "stop" in response to or in anticipation of the rendering of an alarm. This command can cause monitoring to continue for a specific period of time, or alternatively, until a monitoring stop command is sent.

[0060] Next, referring to Figure 1B2, the same group 101B is illustrated. In Figure 1B2, the same adaptation as in Figure 1B1 is made, except that the ASR engine 121A1 and ASR model 132A1 are not replaced by the ASR engine 122A2 and ASR model 132A2. Rather, the ASR engine 121A1 and ASR model 132A1 remain, and an additional ASR engine 122A3 and an additional ASR model 132A3 are provided.

[0061] The ASR engine 121A1 and ASR model 132A1 may be for speech recognition of utterances in a first language (e.g., English), while the additional ASR engine 122A3 and additional ASR model 132A3 may be for speech recognition of utterances in a second language (e.g., Spanish). The ASR engine 121A2 and ASR model 132A2 in Figure 1B1 are also for the first language and may be more robust and / or accurate than the ASR engine 121A1 and ASR model 132A1. However, the processing power of the assistant client 120A may prevent the ASR engine 121A2 and ASR model 132A2 from being stored locally along with the ASR engine 122A3 and additional ASR model 132A3. Nevertheless, this processing power allows for the storage and utilization of both the ASR engine 121A1 and ASR model 132A1, as well as additional ASR engines 122A3 and additional ASR model 132A3.

[0062] The decision to locally store both ASR engine 121A1 and ASR model 132A1, as well as additional ASR engine 122A3 and additional ASR model 132A3, instead of ASR engine 121A2 and ASR model 132A2 in Figure 1B1, may be based on usage statistics showing that the spoken utterances provided by assistant devices 110A and 110B (and / or assistant devices 110C and 110D) include both first-language and second-language spoken utterances, as in the example in Figure 1B2. In the example in Figure 1B1, the usage statistics show only first-language spoken utterances, and as a result, the more robust ASR engine 121A2 and ASR model 132A2 may be selected in Figure 1B1.

[0063] Next, referring to Figure 1B3, the same group 101B is illustrated again. In Figure 1B3, the same adaptations as in Figure 1B1 are made, except that (1) the ASR engine 121A1 and ASR model 132A1 are replaced by the ASR engine 122A4 and ASR model 132A4 instead of the ASR engine 122A2 and ASR model 132A2, (2) there is no warm queue engine or warm queue model on the assistant device 110B, and (3) the ASR engine 122B4 and ASR model 132B4 are stored and used locally on the assistant device 110B.

[0064] The ASR engine 122A4 and ASR model 132A4 are used to perform the first part of speech recognition, and the ASR engine 122B4 and ASR model 132B4 may be used to perform the second part of speech recognition. For example, the ASR engine 122A4 may use the ASR model 132A4 to generate an output, which is transmitted to the assistant client 120B, and the ASR engine 122B4 may process the output when generating speech recognition. In one specific example, the output may be a graph representing candidate recognition, and the ASR engine 122B4 may perform beam search on the graph when generating speech recognition. In another specific example, the ASR model 132A4 may be the initial / downstream part (i.e., the first neural network layer) of an end-to-end speech recognition model, and the ASR model 132B4 may be the later / upstream part (i.e., the second neural network layer) of an end-to-end speech recognition model. In such an example, the end-to-end model is split between two assistant devices 110A and 110B, and the output may be the state of the final layer of the initial part after processing (e.g., the embedding). In yet another example, ASR model 132A4 may be an acoustic model and ASR model 132B4 may be a language model. In such an example, the output may represent a sequence of phonemes or a sequence of probability distributions for phonemes, and the ASR engine 122B4 can utilize the language model to select a transcript / recognition corresponding to that sequence.

[0065] The robustness and / or accuracy of the ASR engine 122A4, ASR model 132A4, ASR engine 122B4, and ASR model 132B4 working in coordination may surpass that of the ASR engine 122A2 and ASR model 132A2 shown in Figure 1B1. Furthermore, the processing power of assistant clients 120A and 120B may prevent the ASR models 132A4 and 132B4 from being stored and utilized on either of those devices independently. However, the processing power may allow for model splitting and splitting of processing roles between the ASR engines 122A4 and 122B4, as described herein. Note that on assistant device 110B, purging of the warm queue engine and warm queue model may enable the storage and utilization of the ASR engine 122B4 and ASR model 132B4. In other words, the processing power of this assistant device 110B would not be sufficient to store and / or utilize the worm queue engine and worm queue model, along with the other engines and models illustrated in Figure 1B3.

[0066] The decision to locally store ASR engine 122A4, ASR model 132A4, ASR engine 122B4, and ASR model 132B4 instead of ASR engine 121A2 and ASR model 132A2 in Figure 1B1 may be based on usage statistics showing that speech recognition in assistant devices 110A and 110B (and / or assistant devices 110C and 110D) is often unreliable and / or often inaccurate, as in the example in Figure 1B3. For example, usage statistics may show that the confidence metric for recognition is below average (e.g., average based on a group of users) and / or that recognition is often corrected by the user (e.g., through editing the display of the transcript).

[0067] Referring to Figure 1C, we see that device group 101C has been created, and assistant devices 110A, 110B, and 110C are included in device group 101C. In some implementations, device group 101C may be generated in response to a user interface input that explicitly indicates the desire to group assistant devices 110A, 110B, and 110C. For example, the user interface input may indicate the desire to create device group 101C from scratch, or alternatively, to add assistant device 110C to device group 101B (Figures 1B1, 1B2, and 1B3) and thereby create a modified group 101C. In some implementations, device group 101C may be generated automatically instead. For example, device group 101B (Figures 1B1, 1B2, and 1B3) may have been previously generated based on the determination that assistant devices 110A and 110B are in close proximity, and after the creation of device group 101B, assistant device 110C may be moved by the user to be in proximity to devices 110A and 110B. As a result, assistant device 110C may be automatically added to device group 101B, thereby creating a modified group 101C.

[0068] Regardless of how group 101C was generated, Figure 1C shows an example of possible adaptations that may be made to assistant devices 110A, 110B, and 110C in response to them being included in group 101C.

[0069] In Figure 1C, the assistant device 110B has the same adaptation as in Figure 1B3. Furthermore, the assistant device 110A has the same adaptation as in Figure 1B3, except that (1) the authentication engine 126A1 and VAD engine 128A1, and their corresponding models 13A1 and 138A1 are purged, (2) the wake queue engine 121A1 and wake queue model 131A1 are replaced with wake queue engine 121A2 and wake queue model 131A2, and (3) the warm queue engine or warm queue model is not present on the assistant device 110B. The models and engines stored on the assistant device 110C are not adapted. However, the assistant client 120C may be adapted to enable cooperative processing of assistant requests with assistant clients 120A and 120B.

[0070] In Figure 1C, the authentication engine 126A1 and the VAD engine 128A1 are purged from the assistant device 110A because their corresponding components already exist on the assistant device 110C. In some implementations, the authentication engine 126A1 and / or the VAD engine 128A1 may only be purged after some or all of the data from their components has been merged with their corresponding components already present on the assistant device 110C. For example, the authentication engine 126A1 can store voice embeddings for both a first and a second user, while the authentication engine 126C1 can only store voice embeddings for the first user. Before purging the authentication engine 126A1, the voice embeddings for the second user may be transmitted locally to the authentication engine 126C1, thereby ensuring that such voice embeddings are utilized by the authentication engine 126C1 and that pre-adaptation capabilities are maintained after adaptation. As another example, authentication engine 126A1 may store instances of speech data used to capture the utterances of each second user and generate speech embeddings for the second user, while authentication engine 126C1 may lack speech embeddings for the second user. Before purging authentication engine 126A1, instances of speech data are transmitted locally from authentication engine 126A1 to authentication engine 126C1, thereby ensuring that instances of speech data are utilized by authentication engine 126C1 when generating speech embeddings for the second user using the on-device authentication model 136C1, and that pre-adaptation capability is maintained after adaptation. Furthermore, wake queue engine 121A1 and wake queue model 131A1 have been replaced by wake queue engine 121A2 and wake queue model 131A2 with smaller storage sizes. For example, the wake cue engine 121A1 and wake cue model 131A1 enable the detection of both vocal wake cues and image-based wake cues, while the wake cue engine 121A2 and wake cue model 131A2 enable the detection of image-based wake cues only.Optionally, personalization, training instances, and / or other settings from the image-based wake cue portion of Wake Cue Engine 121A1 and Wake Cue Model 131A1 may be merged with Wake Cue Engine 121A2 and Wake Cue Model 131A2 before purging Wake Cue Engine 121A1 and Wake Cue Model 131A1, or otherwise shared in some way. Wake Cue Engine 121C1 and Wake Cue Model 131C1 enable the detection of only vocalized wake cues. Thus, Wake Cue Engine 121A2 and Wake Cue Model 131A2, as well as Wake Cue Engine 121C1 and Wake Cue Model 131C1, collectively enable the detection of both vocalized and image-based wake cues. Optionally, personalization and / or other settings from the voice cue portion of the wake cue engine 121A1 and wake cue model 131A1 may be transmitted to the client device 110C to be merged with the wake cue engine 121C1 and wake cue model 131C1, or otherwise shared in some way.

[0071] Furthermore, extra storage space can be obtained by replacing the wake queue engine 121A1 and wake queue model 131A1 with the smaller storage size wake queue engine 121A2 and wake queue model 131A2. This extra storage space, as well as the extra storage space obtained by purging the authentication engine 126A1 and VAD engine 128A1, and their corresponding models 13A1 and 138A1, provides space for the warm queue engine 127A2 and warm queue model 137A2 (which are collectively larger than the replaced warm queue engine 127A1 and warm queue model). The warm queue engine 127A2 and warm queue model 137A2 can be used to monitor warm queues different from those monitored using the warm queue engine 127C1 and warm queue model 137C1.

[0072] As an example of coordinated processing of an assistant request using the adaptation in Figure 1C, suppose the voice utterance "OK Assistant, turn on the kitchen lights" is provided, and the assistant device 120A is the lead device that coordinates the processing. The wake cue engine 121C1 of the assistant client 110C can detect the appearance of the wake cue "OK Assistant". In response, the wake cue engine 121C1 can transmit a command to the assistant devices 110A and 110B to cause the ASR engines 122A4 and 122B4 to coordinately process the captured voice data following the wake cue. The voice data to be processed may be voice data captured by the microphone of assistant device 110C, and / or voice data captured by the microphone of assistant device 110B and / or by assistant device 110C.

[0073] The transcript generated by the ASR engine 122B4 is transmitted to the NLU engine 123B2 to perform an NLU on the transcript for the second set of classifications, and can also be transmitted to the assistant client 120A to perform an NLU on the transcript for the first set of classifications. The results of the NLU performed by the NLU engine 123B2 are transmitted to the assistant client 120A, which can decide which results, if any, to use based on those results and the results from the NLU engine 123A2. The assistant client 120A can transmit the NLU result with the highest probability to the assistant client 120B. The fulfillment engine 124B2 of the assistant client 120B can use those NLU results to decide to transmit a command to "kitchen lights" to transition "kitchen lights" to the "on" state, and can transmit such a command over LAN 108. Optionally, the fulfillment engine 124B2 can use the TTS engine 125B1 to generate synthesized speech confirming the execution of "turning on the kitchen lights". In such a scenario, the synthesized speech may be rendered by the assistant client 120B at the assistant device 110B, transmitted to the assistant device 110A for rendering by the assistant client 120A, and / or transmitted to the assistant device 110C for rendering by the assistant client 120C.

[0074] Next, referring to Figure 1D, we see that device group 101D has been created, and assistant devices 110C and 110D are included in device group 101D. In some implementations, device group 101D may be generated in response to a user interface input that explicitly indicates the desire to group assistant devices 110C and 110D. In some implementations, device group 101D may be generated automatically instead.

[0075] Regardless of how group 101D was generated, Figure 1D shows an example of possible adaptations that may be made to assistant devices 110C and 110D in response to their inclusion in group 101D.

[0076] In Figure 1D, the wake queue engine 121C1 and wake queue model 131C1 of the assistant device 110C are replaced by the wake queue engine 121C2 and wake queue model 131C2. Furthermore, the authentication engine 126D1 and authentication model 136D1, as well as the VAD engine 128D1 and VAD model 138D2, ​​are purged from the assistant device 110D. Moreover, the wake queue engine 121D1 and wake queue model 131D1 of the assistant device 110D are replaced by the wake queue engine 121D2 and wake queue model 131D2, and the worm queue engine 127D1 and worm queue model 137D1 are replaced by the worm queue engine 127D2 and worm queue model 137D2.

[0077] The previous wake cue engine 121C1 and wake cue model 131C1 could only be used to detect a first set of one or more wake words, such as "Hey Assistant" and "OK Assistant". On the other hand, the wake cue engine 121C2 and wake cue model 131C2 can only detect an alternative second set of one or more wake words, such as "Hey Computer" and "OK Computer". Similarly, the previous wake cue engine 121D1 and wake cue model 131D1 could only be used to detect a first set of one or more wake words, and the wake cue engine 121D2 and wake cue model 131D2 can only be used to detect a first set of one or more wake words. However, the wake cue engine 121D2 and wake cue model 131D2 are larger and more robust (e.g., more robust to background noise) and / or more accurate than their replaced counterparts. By purging the engine and model from the assistant device 110D, it may be possible to utilize larger wake cue engines 121D2 and wake cue models 131D2. Furthermore, collectively, the wake cue engine 121C2 and wake cue model 131C2 and the wake cue engine 121D2 and wake cue model 131D2 enable the detection of two sets of wake words, although each of the assistant clients 120C and 120D was only able to detect the first set before adaptation.

[0078] The warm queue engine 127D2 and warm queue model 137D2 of the assistant device 110D may require more computing power than the replaced wake queue engine 127D1 and wake queue model 137D1. However, this power is available through purging the engines and models from the assistant device 110D. Furthermore, the warm queues monitored by the warm queue engine 127D2 and warm queue model 137D1 can be added to those monitored by the warm queue engine 127C1 and warm queue model 137D1. Before adaptation, the wake queues monitored by the wake queue engine 127D1 and wake queue model 137D1 were the same as those monitored by the warm queue engine 127C1 and warm queue model 137D1. Therefore, through cooperative processing, the assistant clients 120C and 120D can monitor more wake queues.

[0079] Note that in the example in Figure 1D, there are many assistant requests that cannot be fully processed collaboratively on-device by assistant clients 120C and 120D. For example, assistant clients 120C and 120D lack an ASR engine, an NLU engine, and a fulfillment engine. This may be due to the fact that the processing power of assistant devices 110C and 110D is insufficient to support any such engine or model. Therefore, for non-warm queued voice utterances supported by assistant clients 120C and 120D, the cloud-based assistant component 140 still needs to be utilized to fully process many assistant requests. However, the adaptation and adaptive-based collaborative processing in Figure 1D can still be more robust and / or accurate than any processing performed individually on the devices before adaptation. For example, adaptation enables the detection of additional wake queues and additional warm queues.

[0080] As an example of possible collaborative processing, let's assume the spoken utterance "OK Computer, play some music." In such an example, the wake queue engine 121D2 can detect the wake queue "OK Computer." In response, the wake queue engine 121D2 can transmit the voice data corresponding to the wake queue to the assistant client 120C. The authentication engine 126C1 of the assistant client 120C can use the voice data to determine whether the wake queue utterance can be authenticated for the registered user. The wake queue engine 121D2 can further stream the voice data following the spoken utterance to the cloud-based assistant component 140 for further processing. The voice data can be captured by the assistant device 110D or assistant device 110C (for example, the assistant client 120 can transmit a command to the assistant client 120C to capture the voice data in response to the wake queue engine 121D2 detecting the wake queue). Furthermore, authentication data based on the output of the authentication engine 126C1 can also be transmitted along with the voice data. For example, if the authentication engine 126C1 authenticates a wake queue utterance against a registered user, the authentication data may include the registered user's identifier. Alternatively, if the authentication engine 126C1 does not authenticate a wake queue utterance against any registered user, the authentication data may include an identifier that reflects the utterance provided by the guest user.

[0081] Various specific examples have been described so far with reference to Figures 1B1, 1B2, 1B3, 1C, and 1D. However, it should be noted that various additional or alternative groups may be generated, and / or various additional or alternative adaptations may be performed in response to the generation of groups.

[0082] Figure 2 is a flowchart illustrating an exemplary method 200 for adapting the on-device model and / or processing role of assistant devices within a group. For convenience, the operations in the flowchart are described with reference to the system in which the operations are performed. This system may include various components of various computer systems, such as one or more of the assistant clients 120A-D in Figure 1, and / or components of the cloud-based assistant component 140 in Figure 1. Furthermore, although the operations of method 200 are shown in a specific order, this is not intended to be restrictive. One or more operations may be reordered, omitted, or added.

[0083] In block 252, the system generates a group of assistant devices. For example, the system may generate a group of assistant devices in response to a user interface input that explicitly indicates that it wants to generate a group. As another example, the system may automatically generate a group in response to determining that one or more conditions are met. As yet another example, the system may automatically determine that a group should be generated in response to determining that conditions are met, provide a user interface output suggesting the generation of a group, and then generate the group in response to a positive user interface being received in response to the user interface output.

[0084] In block 254, the system obtains the processing power for each of the group's assistant devices. For example, the system may be one of the group's assistant devices. In such an example, an assistant device can obtain its own processing power, and other assistant devices in the group can communicate their processing power to it. In another example, the processing power of assistant devices is stored in a device topology, and the system can retrieve them from the device topology. In yet another example, the system may be a cloud-based component, and each of the group's assistant devices can communicate its processing power to the system.

[0085] The processing power of an assistant device may include a corresponding processor value based on the capabilities of one or more on-device processors, a corresponding memory value based on the size of on-device memory, and / or a corresponding disk space value based on available disk space. For example, the processor value may include details about the operating frequency of one or more processors, details about the size of the processor cache, whether each processor is a GPU, CPU, or DSP, and / or other details. As another example, the processor value may, in addition to or alternatively, include a higher level classification of processor capabilities such as high, medium, or low, or GPU+CPU+DSP, high-power CPU+DSP, medium-power CPU+DSP, or low-power CPU+DSP. As yet another example, the memory value may include memory details such as a specific size of memory, or a higher level classification of memory such as high, medium, or low. As yet another example, the disk space value may include details about available disk space such as a specific size of disk space, or a higher level classification of available disk space such as high, medium, or low.

[0086] In block 256, the system utilizes the processing power of block 254 when determining a collective set of on-device models for a group. For example, the system may determine a set of on-device models that maximizes the use of collective processing power, ensuring that each on-device model in that set is locally stored and available on a device that can store and utilize on-device models. The system may also, if possible, try to ensure that the selected set contains a complete (or more complete than other candidate sets) pipeline of on-device models. For example, the system may choose a set containing ASR models but less robust NLU models over a set containing highly robust NLU models but not ASR models.

[0087] In some implementations, block 256 includes subblock 256A, where the system uses utilization data when selecting a collective set of on-device models for a group. Past utilization data may be data relating to past assistant interactions in one or more assistant devices in the group and / or in one or more additional assistant devices in the ecosystem. In some implementations, in subblock 256A, the system considers utilization data, along with the above considerations, when selecting on-device models to include in the set. For example, if processing power allows the set to include a high-precision ASR model (instead of a low-precision ASR model) or a highly robust NLU model (instead of a less robust NLU model), but not both, utilization data may be used to determine which to select. For example, if utilization data reflects that past assistant interactions were directed exclusively to intents covered by a less robust NLU model that could be included in the set along with a high-precision ASR model, then the high-precision ASR model might be selected for inclusion in the set. On the other hand, if the usage data reflects that past assistant interactions are covered by a robust NLU model but include many directed towards intents not covered by a less robust NLU model, then the robust NLU model may be selected to be included in the set. In some implementations, the candidate set is initially determined based on processing power, without considering the usage data, and then, if multiple valid candidate sets exist, the usage data may be used to select one over the other.

[0088] In block 258, the system causes each assistant device to locally store a corresponding subset of a collective set of on-device models. For example, the system can communicate corresponding instructions to each of the group's assistant devices about which on-device models should be downloaded. Based on the received instructions, each assistant device can download the corresponding models from a remote database. As another example, the system can retrieve an on-device model and push the corresponding on-device model to each of the group's assistant devices. As yet another example, for any on-device model that is stored in one of the group's assistant devices before adaptation and to be stored in the other of the group's assistant devices during adaptation, such models can be communicated directly between the devices. For example, suppose the first assistant device stores an ASR model before adaptation, and that same ASR model is stored on the second assistant device and purged from the first assistant device during adaptation. In such a case, the system can instruct the first assistant device to transmit the ASR model to the second assistant device's local storage (and / or to download it from the first assistant device), and the first assistant device can then purge the ASR model. In addition to disrupting WAN traffic, transmitting the pre-adapted models locally allows for the preservation of any personalization of those on-device models previously performed during transmission. Personalized models may have higher accuracy for ecosystem users compared to their unpersonalized counterparts in remote storage.As another example, for any assistant device that has a stored training instance for personalizing an on-device model on it before adaptation, such training instance can be propagated to an assistant device that will have a corresponding model downloaded from a remote database after adaptation. The assistant device with the on-device model after adaptation can then use the training instance to personalize the corresponding model downloaded from the remote database. The corresponding model downloaded from the remote database may be different (for example, smaller or larger) from the corresponding model used by the training instance before adaptation, but the training instance can be used as is when personalizing a different downloaded on-device model.

[0089] In block 260, the system assigns a corresponding role to each of the assistant devices. In some implementations, assigning a corresponding role involves having each assistant device download and / or implement an engine corresponding to an on-device model stored locally in the assistant device. Each engine can utilize the corresponding on-device model when performing its corresponding processing role, such as performing all or some ASR, performing wake word recognition for at least some wake words, performing warm word recognition for some warm words, and / or performing authentication. In some implementations, one or more of the processing roles are executed only when there is a command from the lead device of the group of assistant devices. For example, an NLU processing role performed by a given device using an on-device NLU model may be executed only in response to the lead device transmitting the corresponding text for NLU processing and / or a specific command to cause the given device to perform NLU processing. As another example, a warm word monitoring processing role, executed by a given device using an on-device warm queue engine and on-device warm queue model, may only be executed in response to a command transmitted by a lead device to the given device to perform warm word processing. For example, a lead device may cause a given device to monitor for the occurrence of the “stop” warm word in response to an alarm sounding on the lead device or other devices in the group. In some implementations, one or more processing roles may be executed independently of any commands from the lead assistant device, at least selectively. For example, a wake queue monitoring role, executed by a given device using an on-device wake queue engine and on-device wake queue model, may run continuously unless explicitly disabled by the user.As another example, a warm queue monitoring role performed by a given device using an on-device warm queue engine and on-device warm queue model may be performed continuously or based on conditions under which monitoring is detected locally on the given device.

[0090] In block 262, the system causes subsequent vocal utterances detected at one or more of the group's devices to be processed locally and cooperatively by the group's assistant devices according to their roles. Various non-exclusive examples of such cooperative processing are described herein. For example, examples are described with reference to Figures 1B1, 1B2, 1B3, 1C, and 1D.

[0091] In block 264, the system determines whether any changes have been made to the group, such as adding a device to the group, removing a device from the group, or dissolving the group. If not, the system continues executing block 262. If it has been changed, the system proceeds to block 266.

[0092] In block 266, the system determines whether the change to a group means that one or more assistant devices that were in the group are now solitary (i.e., no longer assigned to a group). If so, the system proceeds to block 268, where each solitary device is instructed to locally store its pre-grouping on-device model and assume its pre-grouping on-device processing role. In other words, if a device is no longer in a group, it may be reverted to the state it was in before the adaptation was performed in response to being included in a group. In these and other ways, after reverting to that state, the solitary device can functionally process a variety of assistant requests operating in solitary capacity. Before reverting to that state, the solitary device may not have been able to functionally process any assistant requests, or at least fewer assistant requests than before reverting to that state.

[0093] In block 270, the system determines whether there are two or more devices remaining in the modified group. If so, the system returns to block 254 and performs another iteration of blocks 254, 256, 258, 260, and 262 based on the modified group. For example, if the modified group includes an additional assistant device without losing any of the previous assistant devices in the group, an adaptation may be made to account for the additional processing power of the additional assistant device. If the decision in block 270 is no, the group is disbanded and the system proceeds to block 272, where method 200 terminates (until another group is generated).

[0094] Figure 3 is a flowchart illustrating exemplary method 300 that may be performed by each of several assistant devices in a group when adapting the on-device model and / or processing role of the assistant devices in the group. The actions of method 300 are shown in a specific order, but this is not intended to be restrictive. One or more actions may be reordered, omitted, or added.

[0095] The operation of Method 300 is a specific example of Method 200 that can be performed by each of the assistant devices in the group. Therefore, the operation will be described with reference to an assistant device that performs the operation, such as one or more of the assistant clients 120A-D in Figure 1. Each of the assistant devices in the group can perform Method 300 in response to receiving input indicating that it is included in the group.

[0096] In block 352, the assistant device receives a grouping instruction indicating that it was included in a group. In block 352, the assistant device also receives identifiers of other assistant devices in the group. Each identifier may be, for example, a MAC address, an IP address, a label assigned to the device (for example, the use assigned in the device topology), a serial number, or other identifiers.

[0097] In an optional block 354, the assistant device transmits data to other assistant devices in the group. The data is transmitted to the other device using an identifier received in block 352. In other words, the identifier may be a network address or may be used to find the network address to which the data is to be transmitted. The transmitted data may include one or more processing values, another device identifier, and / or other data as described herein.

[0098] In an optional block 356, the assistant device receives data transmitted by another device in block 354.

[0099] In block 358, the assistant device determines whether it is a read device based on data optionally received in block 356 or an identifier received in block 352. For example, a device may choose to be a reader if its own identifier is the lowest (or alternatively highest) compared to other identifiers received in block 352. As another example, a device may choose to be a reader if its processing value exceeds the processing value received in the data in block 356 of all other processing values. Other data may be transmitted in block 354 and received in block 356, and such other data may also enable the assistant device to make an objective decision on whether it should be a reader. More generally, in block 358, the assistant device may use one or more objective criteria in determining whether it should be a read device.

[0100] In block 360, the assistant device determines whether it was determined to be a read device in block 358. If the assistant device was not determined to be a read device, it then moves down from block 360 to the "no" branch. If the assistant device was determined to be a read device, it moves down from block 360 to the "yes" branch.

[0101] In the "yes" branch, the assistant device utilizes processing power received from other assistant devices in the group, as well as its own processing power, when determining a collective set of on-device models for the group in block 360. Processing power may be transmitted to the lead device by other assistant devices in optional block 354 or block 370 (described later) if optional block 354 is not executed or the data in block 354 does not contain processing power. In some implementations, block 362 may share one or more aspects common to block 256 of method 200 in Figure 2. For example, in some implementations, block 362 may also include considering past usage data when determining a collective set of on-device models.

[0102] In block 364, an assistant device transmits to each of the other assistant devices in the group instructions for each of the collective set of on-device models that the other assistant devices should download. Optionally, in block 364, an assistant device may also transmit to each of the other assistant devices in the group instructions for each of the processing roles that should be performed by the assistant device using the on-device model.

[0103] In block 366, the assistant device downloads and stores the set of on-device models assigned to the assistant device. In some implementations, blocks 364 and 366 may share one or more embodiments common to block 258 of method 200 in Figure 2.

[0104] In block 368, the assistant device coordinates the cooperative processing of the assistant request, including utilizing its own on-device model when performing a portion of the cooperative processing. In some implementations, block 368 may share one or more embodiments common to block 262 of method 200 in Figure 2.

[0105] Next, turning to the "no" branch, in the optional block 370, the assistant device communicates its processing capacity to the read device. Block 370 may be omitted, for example, when block 354 is executed and the processing capacity is included in the data transmitted in block 354.

[0106] In block 372, the assistant device receives instructions from the read device regarding the on-device model to be downloaded and, optionally, instructions regarding the processing role.

[0107] In block 374, the assistant device downloads and stores the on-device model reflected in the on-device model instruction received in block 372. In some implementations, blocks 372 and 374 may share one or more embodiments common to block 258 of method 200 in Figure 2.

[0108] In block 376, the assistant device utilizes its on-device model when performing part of the cooperative processing of the assistant request. In some implementations, block 376 can share one or more embodiments common to block 262 of method 200 in Figure 2.

[0109] Figure 4 is a block diagram of an exemplary computing device 410 that may be optionally used to perform one or more embodiments of the techniques described herein. In some implementations, one or more of the assistant devices and / or other components may include one or more components of the exemplary computing device 410.

[0110] The computing device 410 typically comprises at least one processor 414 that communicates with numerous peripheral devices via a bus subsystem 412. These peripheral devices may include, for example, a storage subsystem 425, a user interface output device 420, a user interface input device 422, and a network interface subsystem 416, including a memory subsystem 425 and a file storage subsystem 426. The input and output devices enable interaction between the user and the computing device 410. The network interface subsystem 416 provides an interface to an external network and is coupled to a corresponding interface device in another computing device.

[0111] The user interface input device 422 may include pointing devices such as keyboards, mice, trackballs, touchpads, or graphics tablets, scanners, touchscreens integrated into displays, voice input devices such as voice recognition systems, microphones, and / or other types of input devices. In general, the use of the term “input device” is intended to include all possible types of devices and methods for inputting information into the computing device 410 or onto a communication network.

[0112] The user interface output device 420 may include a display subsystem, a printer, a fax machine, or a non-visual display such as an audio output device. The display subsystem may include a flat panel device such as a cathode ray tube ("CRT") or a liquid crystal display ("LCD"), a projection device, or any other mechanism for generating visible images. The display subsystem may also include a non-visual display via an audio output device, etc. In general, the use of the term "output device" is intended to include all possible types of devices and methods for outputting information from the computing device 410 to a user or another machine or computing device.

[0113] The storage subsystem 425 stores programming and data structures that implement some or all of the functions of the modules described herein. For example, the storage subsystem 425 may include logic circuits for performing one or more selected embodiments of the methods described herein and / or for implementing the various components depicted herein.

[0114] These software modules are generally executed by processor 414, either alone or in combination with other processors. The memory 425 used in the storage subsystem 425 may comprise a number of memories, including primary random access memory ("RAM") 430 for storing instructions and data during program execution, and read-only memory ("ROM") 432 for storing fixed instructions. The file storage subsystem 426 may comprise persistent storage for program and data files and may include a hard disk drive, a floppy disk drive with associated removable media, a CD-ROM drive, an optical drive, or a removable media cartridge. Modules implementing functions in several implementation forms may be stored within the storage subsystem 425 by the file storage subsystem 426, or in other machines accessible by processor 414.

[0115] The bus system 412 provides a mechanism for various components and subsystems of the computing device 410 to communicate with each other as intended. Although the bus system 412 is schematically illustrated as a single bus, alternative implementations of the bus system may use multiple buses.

[0116] The computing device 410 may be any type of device, including workstations, servers, computing clusters, blade servers, server farms, or any other data processing system or computing device. Because computers and networks are constantly changing by their very nature, the description of the computing device 410 shown in Figure 4 is intended only as a specific example to illustrate several implementations. Many other configurations of the computing device 410 may have more or fewer components than the computing device shown in Figure 4.

[0117] Where the systems described herein may collect or use personal information relating to a user (or, more often, a “participant” herein), the user may be given the opportunity to control whether the programs or functions collect user information (e.g., information about the user’s social networks, social behavior or activities, professional occupation, user preferences, or the user’s current geographical location), or whether and / or how content is received from the content server that is deemed to be of higher relevance to the user. Furthermore, certain data may be processed in one or more ways before it is stored or used, so that personally identifiable information is removed. For example, a user’s identity may be processed in such a way that personally identifiable information cannot be determined about the user, or the user’s geographical location may be generalized when geographical location information (such as city name, postal code, or national level) is available, so that the user’s specific geographical location does not need to be determined. Thus, the user may control how information is collected and / or used relating to them.

[0118] In several implementations, a method is provided for generating an assistant device group of heterogeneous assistant devices. The heterogeneous assistant devices include at least a first assistant device and a second assistant device. At the time of generating the group, the first assistant device includes a first set of locally stored on-device models used when locally processing assistant requests directed to the first assistant device. Furthermore, at the time of generating the group, the second assistant device includes a second set of locally stored on-device models used when locally processing assistant requests directed to the second assistant device. The method further includes determining a collective set of locally stored on-device models for use when collaboratively locally processing assistant requests directed to any of the heterogeneous assistant devices in the assistant device group, based on the corresponding processing capabilities of each heterogeneous assistant device in the assistant device group. The method further includes causing each heterogeneous assistant device to locally store a corresponding subset of the collective set of locally stored on-device models in response to the generation of the assistant device group, and assigning one or more corresponding processing roles to each heterogeneous assistant device in the assistant device group. Each processing role utilizes one or more corresponding on-device models stored locally. Furthermore, having each heterogeneous assistant device locally store a corresponding subset includes purging one or more first on-device models from the first set from the first assistant device to provide storage space for the corresponding subset stored locally on the first assistant device, and purging one or more second on-device models from the second set from the second assistant device to provide storage space for the corresponding subset stored locally on the second assistant device.This method further includes assigning a corresponding processing role to each of the heterogeneous assistant devices in the assistant device group, detecting a spoken utterance via the microphone of at least one of the heterogeneous assistant devices in the assistant device group, and, in response to the detection of the spoken utterance via the microphone of the assistant device group, causing the spoken utterance to be processed collaboratively and locally by the heterogeneous assistant devices in the assistant device group using their corresponding processing roles.

[0119] These and other implementations of the technology disclosed herein may optionally include one or more of the following features:

[0120] In some implementations, purging one or more of a first set of first on-device models to a first assistant device includes purging a first set of first device wake word detection models used when detecting a first wake word. In those implementations, a corresponding subset stored locally on a second assistant device includes a second device wake word detection model used when detecting a first wake word, and assigning a corresponding processing role includes assigning a first wake word detection role to the second assistant device that utilizes the second device wake word detection model when monitoring for the occurrence of a first wake word. In some of those implementations, a voice utterance includes a first wake word followed by an assistant command, and in the first wake word detection role, the second assistant device detects the occurrence of a first wake word and triggers the execution of an additional role of the corresponding processing role in response to the detection of the occurrence of a first wake word. In some versions of these implementations, the additional roles of the corresponding processing roles are performed by the first assistant device, and the second assistant device triggers the execution of the additional roles of the corresponding processing roles by transmitting an instruction to the first assistant device to detect the first wake word.

[0121] In some implementations, the corresponding subset stored locally on the first assistant device includes the first wake word detection model of the first device used to detect one or more first wake words, but excludes the wake word detection model used to detect one or more second wake words. In some of those implementations, the corresponding subset stored locally on the second assistant device includes the second wake word detection model of the second device used to detect one or more second wake words, but excludes the wake word detection model used to detect one or more first wake words. In some versions of those implementations, assigning the corresponding processing roles includes assigning the first assistant device a first wake word detection role that utilizes the wake word detection model of the first device when monitoring the occurrence of one or more first wake words, and assigning the second assistant device a second wake word detection role that utilizes the wake word detection model of the second device when monitoring the occurrence of one or more second wake words.

[0122] In some implementations, the corresponding subset stored locally on the first assistant device includes a first language speech recognition model used for speech recognition in the first language, and excludes any speech recognition models used for speech recognition in the second language. In some of those implementations, the corresponding subset stored locally on the second assistant device includes a second language speech recognition model used for speech recognition in the second language, and excludes any speech recognition models used for speech recognition in the second language. In some versions of those implementations, assigning the corresponding processing roles includes assigning the first assistant device a first language speech recognition role that uses the first language speech recognition model for speech recognition in the first language, and assigning the second assistant device a second language speech recognition role that uses the second language speech recognition model for speech recognition in the second language.

[0123] In some implementations, the corresponding subset stored locally on the first assistant device includes the first part of the speech recognition model used to perform the first part of speech recognition, but excludes the second part of the speech recognition model. In some of those implementations, the corresponding subset stored locally on the second assistant device includes the second part of the speech recognition model used to perform the second part of speech recognition, but excludes the first part of the speech recognition model. In some versions of those implementations, assigning the corresponding processing role includes assigning the first part of a language speech recognition role to the first assistant device, which uses the first part of the speech recognition model when generating the corresponding embedding of the corresponding speech; transmitting the corresponding embedding to the second assistant device; and assigning the second language speech recognition role to the second assistant device, which uses the corresponding embedding from the first assistant device and the second language speech recognition model when generating the corresponding recognition of the corresponding speech.

[0124] In some implementations, the corresponding subset stored locally on the first assistant device includes a speech recognition model used when performing the first part of speech recognition. In some of these implementations, assigning the corresponding processing role includes assigning the first assistant device a first part of a language speech recognition role that uses the speech recognition model when generating an output, transmitting the corresponding output to the second assistant device, and assigning the second assistant device a second language speech recognition role that performs beam search on the corresponding output from the first assistant device when generating the corresponding recognition of the corresponding speech.

[0125] In some implementations, the corresponding subset stored locally on the first assistant device includes one or more pre-adapted natural language understanding models used when performing semantic analysis of natural language input, with the one or more pre-adapted natural language understanding models occupying a first amount of local disk space on the first assistant device. In some of these implementations, the corresponding subset stored locally on the first assistant device includes at least one additional natural language understanding model that complements the one or more pre-adapted natural language understanding models, occupying a second amount of local disk space on the first assistant device, with the second amount being greater than the first amount, and includes one or more post-adapted natural language understanding models.

[0126] In some implementations, the corresponding subset stored locally on the first assistant device includes the first device's natural language understanding model used in semantic analysis for one or more first classifications, but excludes the natural language understanding model used in semantic analysis for a second classification. In some of these implementations, the corresponding subset stored locally on the second assistant device includes at least the second device's natural language understanding model used in semantic analysis for a second classification.

[0127] In some implementations, the corresponding processing capability for each heterogeneous assistant device in an assistant device group includes a corresponding processor value based on the capability of one or more on-device processors, a corresponding memory value based on the size of the on-device memory, and / or a corresponding disk space value based on the available disk space.

[0128] In some implementations, generating a group of heterogeneous assistant devices responds to user interface input that explicitly indicates a desire to group heterogeneous assistant devices.

[0129] In some implementations, the generation of a group of heterogeneous assistant devices is performed automatically in response to a determination that the heterogeneous assistant devices satisfy one or more proximity conditions with respect to each other.

[0130] In some implementations, generating an assistant device group of heterogeneous assistant devices is performed in response to receiving a positive user interface input in response to a recommendation to create an assistant device group, which is automatically generated in response to a determination that the heterogeneous assistant devices satisfy one or more proximity conditions with respect to each other.

[0131] In some implementations, this method further includes assigning a corresponding processing role to each of the heterogeneous assistant devices in the assistant device group, determining that the first assistant device is no longer in the group, and, in response to the determination that the first assistant device is no longer in the group, causing the first assistant device to replace the corresponding subset stored locally on the first assistant device with the first on-device model of the first set.

[0132] In some implementations, determining this collective set is further based on usage data that reflects past usage of one or more of the group's assistant devices.

[0133] In some of these implementations, determining a collective set involves determining multiple candidate sets, each of which can be collectively stored locally and collectively used by the group's assistant devices, based on the corresponding processing capabilities of each heterogeneous assistant device in the assistant device group, and selecting a collective set from the candidate sets based on the usage data.

[0134] In several implementations, a method is provided that is implemented by one or more processors of an assistant device. This method includes determining that the assistant device is a lead device to a group of assistant devices, in response to a determination that the assistant device is included in a group of assistant devices, which includes the assistant device and one or more additional assistant devices. In response to the determination that the assistant device is a lead device to a group, the method further includes determining, based on the processing capacity of the assistant device and the processing capacity received for each of the one or more additional assistant devices, a collective set of on-device models for use in cooperatively processing assistant requests directed to any of the heterogeneous assistant devices in the assistant device group locally, and for each on-device model, a corresponding designation indicating which of the assistant devices in the group locally stores the on-device model. The method further includes, in response to a determination that an assistant device is the lead device for a group, communicating with one or more additional assistant devices to cause each of the one or more additional assistant devices to locally store one of on-device models having a corresponding designation for the additional assistant devices; an assistant device locally storing an on-device model having a corresponding designation for the assistant device; and assigning one or more corresponding processing roles to each of the group's assistant devices for the cooperative local processing of assistant requests directed to the group.

[0135] These and other implementations of the technology disclosed herein may optionally include one or more of the following features:

[0136] In some implementations, determining that an assistant device is a lead device for a group involves comparing the processing capacity of the assistant device with the received processing capacity of one or more additional assistant devices, and determining, based on the comparison, that the assistant device is a lead device for the group.

[0137] In some implementations, groups of assistant devices are created in response to user interface input that explicitly indicates the desire to group assistant devices.

[0138] In some implementations, the method further includes determining that an assistant device is a lead device for a group, receiving an assistant request on one or more of the group's assistant devices, and, in response, coordinating the cooperative local processing of the assistant request using the corresponding processing role assigned to the assistant device.

[0139] In several implementations, a method is provided that is implemented by one or more processors of the assistant device, which includes determining that the assistant device has been removed from a group of heterogeneous assistant devices, the group which included the assistant device and at least one additional assistant device. At the time the assistant device was removed from the group, the assistant device locally stores a set of on-device models, the set of on-device models was insufficient locally for the assistant device to fully process the spoken utterances directed to the automated assistant. The method further includes, in response to the determination that the assistant device has been removed from the group of assistant devices, causing the assistant device to purge one or more of the on-device models from the set and retrieve one or more additional on-device models to be stored locally. Following the retrieval and local storage of one or more additional on-device models, the assistant device, the one or more additional on-device models, and any remaining on-device models from the set may be used locally by the assistant device to fully process the spoken utterances directed to the automated assistant. [Explanation of Symbols]

[0140] 101B Device Group 101C Device Group 101D Device Group 108 Local Area Network (LAN) 109 Wide Area Network (WAN) 110A First Assistant Device 110B Second Assistant Device 110C Third Assistant Device 110D: The fourth assistant device 120A Assistant Client 120B Assistant Client 120C Assistant Client 120D Assistant Client 121 Wake Queue Engine 121A1 Wake / Call Engine 121A2 Wake Queue Engine 121B1 Wake / Call Engine 121C1 Wake / Call Engine 121D1 Wake / Call Engine 121D2 Wake Queue Engine 122A1 ASR engine 122A2 ASR engine 122A3 ASR engine 122A4 ASR engine 122B1 ASR engine 122B4 ASR engine 123A1 NLU engine 123A2 NLU engine 123B1 NLU engine 123B2 NLU engine 124A1 Fulfillment Engine 124B1 Fulfillment Engine 124B2 Fulfillment Engine 125A1 Text-to-Speech (TTS) Engine 125B1 TTS engine 126A1 Authentication Engine 126B1 Authentication Engine 126C1 Authentication Engine 126D1 Authentication Engine 127A1 Worm Cube Engine 127A2 Worm Cube Engine 127B1 Worm Cube Engine 127B2 Warm Worm Engine 127C1 Worm Cube Engine 127D1 Worm Cube Engine 127D2 Worm Cube Engine 128A1 VAD engine 128B1 VAD engine 128C1 VAD engine 128D1 VAD engine 131A1 On-device wake / call model 131A2 Wake Cue Model 131B1 On-device wake / call model 131C1 On-device wake / call model 131D1 On-device wake / call model 131D2 Wake Cue Model 132A1 On-device ASR model 132A2 On-Device ASR Model 132A3 ASR model 132A4 ASR model 132B1 On-device ASR model 132B4 ASR Model 133A1 On-device NLU model 133A2 On-device NLU model 133B1 On-device NLU model 133B2 On-device NLU model, Fulfillment model 134A1 On-Device Fulfillment Model 134B1 On-Device Fulfillment Model 135A1 On-device TTS model 135B1 On-device TTS model 136A1 On-device authentication model 136B1 On-device authentication model 136C1 On-Device Authentication Model 136D1 On-Device Authentication Model 137A1 On-Device Warm Queue Model 137A2 Warm Cue Model 137B1 On-Device Warm Queue Model 137B2 Warmward Model 137C1 On-Device Warm Queue Model 137D1 On-Device Warm Queue Model 137D2 Warm Cue Model 138A1 On-device VAD model 138B1 On-Device VAD Model 138C1 On-Device VAD Model 138D1 On-Device VAD Model 138D2 VAD Model 140 Cloud-Based Automation Assistant Components 150 Local Model Repositories 200 ways 300 ways 410 Computing Devices 412 Bus subsystem 414 processors 416 Network Interface Subsystem 420 User Interface Output Devices 422 User Interface Input Devices 425 Memory subsystem. Storage subsystem. 426 File Storage Subsystem 430. Primary Random Access Memory ("RAM") 432 Read-only memory ("ROM")

Claims

[Claim 1] A method implemented by one or more processors in an assistant device, A step in determining that an assistant device has been removed from a group of heterogeneous assistant devices, The group already includes the assistant device and at least one additional assistant device. At the point when the assistant device is removed from the group, The aforementioned assistant device locally stores a set of on-device models. The aforementioned set of on-device models is insufficient to fully process the spoken utterances directed to the automated assistant locally on the assistant device, step, In response to the decision that the assistant device has been removed from the group of assistant devices, The steps include causing the assistant device to purge one or more of the on-device models in the set, and to retrieve and locally store one or more additional on-device models, After retrieving the one or more additional on-device models and storing them locally, the assistant device, the one or more additional on-device models, and any remaining on-device models in the set are used locally by the assistant device to fully process the spoken utterances directed to the automated assistant, and Methods that include...